[ 
https://issues.apache.org/jira/browse/LUCENE-8250?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16436960#comment-16436960
 ] 

Jim Ferenczi commented on LUCENE-8250:
--------------------------------------

I attached a small test that I hope illustrate the issue. The synonym rule is 
"twd, the walking dead, the zombie show" and removing "the" from the stream 
after the synonym graph makes "zombie show" a following path of "walking" so 
the output of the graph is "twd, walking dead, walking zombie show". It's 
unclear to me if the FilteringTokenFilter is doing the right thing here. I 
added the dot output of TokenStreamToAutomaton in the test, this class is able 
to fill the hole when a stop filter removes a token but in this case I don't 
see how we can infer that "zombie show" is not after "walking". 

> Should FilteringTokenFilter handle positionLength
> -------------------------------------------------
>
>                 Key: LUCENE-8250
>                 URL: https://issues.apache.org/jira/browse/LUCENE-8250
>             Project: Lucene - Core
>          Issue Type: Improvement
>            Reporter: Jim Ferenczi
>            Priority: Major
>         Attachments: LUCENE-8250.patch
>
>
> FilteringTokenFilter does not handle the position length graph attribute when 
> removing a token from the stream. This doesn't work well with graph token 
> stream that sets position length since removing a token from the stream can 
> invalidate the position length set on the previous tokens. 
> This issue was first discussed in 
> https://issues.apache.org/jira/browse/LUCENE-4065 but it has a different 
> purpose which is why I am opening a new issue here.



--
This message was sent by Atlassian JIRA
(v7.6.3#76005)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to