#32492: Add django API for Postgres pg_trgm word_similarity, fuzzy full-text
search.
-------------------------------+--------------------------------------
     Reporter:  Taneli         |                    Owner:  nobody
         Type:  Uncategorized  |                   Status:  new
    Component:  Uncategorized  |                  Version:  3.1
     Severity:  Normal         |               Resolution:
     Keywords:                 |             Triage Stage:  Unreviewed
    Has patch:  0              |      Needs documentation:  0
  Needs tests:  0              |  Patch needs improvement:  0
Easy pickings:  0              |                    UI/UX:  0
-------------------------------+--------------------------------------
Description changed by Taneli:

Old description:

> Use case: filter queryset by an inexact substring. For example, retrieve
> an object with the following headline ' Dogecoin is following bitcoin in
> its dramatic rise' when searching for 'dogge'.
>
> Currently it is possible to filter a queryset on the basis of trigram
> similarity between the search string and full text stored in a column.
>  {{{#!python
> Author.objects.annotate(similarity=TrigramSimilarity('name',
> test),).filter(similarity__gt=0.3)
> }}}
>
> This is a wrapper around the '''similarity''' function of the
> '''pg_trgm''' extension. While it allows comparing full strings, i.e.
> searching for 'doge'  would find 'dogs' or 'dogge' it is useless for
> fuzzy searching of substrings.
>
> {{{
> SELECT similarity('dogge', 'doge');
> ---------
> 0.57
>
> SELECT similarity('dogge', 'dogecoin is following bitcoin');
> ------------
> 0.1
> }}}
>
> '''word_similarity''' does take into account the word boundaries
>
> {{{
> SELECT word_similarity('doge', 'dogecoin is following bitcoin');
> --------------
> 0.5
> }}}
>
> Adding a django API to '''word_similarity''' would allow for better fuzzy
> fulltext search without a need to use either raw SQL or external tools
> like elasticsearch.

New description:

 Use case: filter queryset by an inexact substring. For example, retrieve
 an object with the following headline ''Dogecoin is following bitcoin in
 its dramatic rise'' when searching for ''dogge''.

 Currently it is possible to filter a queryset on the basis of trigram
 similarity between the search string and full text stored in a column.
  {{{#!python
 Author.objects.annotate(similarity=TrigramSimilarity('name',
 test),).filter(similarity__gt=0.3)
 }}}

 This is a wrapper around the '''similarity''' function of the
 '''pg_trgm''' extension. While it allows comparing full strings, i.e.
 searching for 'doge'  would find 'dogs' or 'dogge' it is useless for fuzzy
 searching of substrings.

 {{{
 SELECT similarity('dogge', 'doge');
 ---------
 0.57

 SELECT similarity('dogge', 'dogecoin is following bitcoin');
 ------------
 0.1
 }}}

 '''word_similarity''' does take into account the word boundaries

 {{{
 SELECT word_similarity('doge', 'dogecoin is following bitcoin');
 --------------
 0.5
 }}}

 Adding a django API to '''word_similarity''' would allow for better fuzzy
 fulltext search without a need to use either raw SQL or external tools
 like elasticsearch.

--

-- 
Ticket URL: <https://code.djangoproject.com/ticket/32492#comment:1>
Django <https://code.djangoproject.com/>
The Web framework for perfectionists with deadlines.

-- 
You received this message because you are subscribed to the Google Groups 
"Django updates" group.
To unsubscribe from this group and stop receiving emails from it, send an email 
to [email protected].
To view this discussion on the web visit 
https://groups.google.com/d/msgid/django-updates/065.2327a67a755bcf90353b2f4a65cf9724%40djangoproject.com.

Reply via email to