[ https://issues.apache.org/jira/browse/LUCENE-3907?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=13652770#comment-13652770 ]
Jack Krupansky commented on LUCENE-3907: ---------------------------------------- Why not make the big changes for trunk/5.0, but leave the existing filters/tokenziers in 4.x as deprecated. Add the "leading" replacements as well in 4x, but be sure to preserve the existing stuff with support for "back" in 4x - as deprecated. > Improve the Edge/NGramTokenizer/Filters > --------------------------------------- > > Key: LUCENE-3907 > URL: https://issues.apache.org/jira/browse/LUCENE-3907 > Project: Lucene - Core > Issue Type: Improvement > Reporter: Michael McCandless > Assignee: Adrien Grand > Labels: gsoc2013 > Fix For: 4.3 > > Attachments: LUCENE-3907.patch > > > Our ngram tokenizers/filters could use some love. EG, they output ngrams in > multiple passes, instead of "stacked", which messes up offsets/positions and > requires too much buffering (can hit OOME for long tokens). They clip at > 1024 chars (tokenizers) but don't (token filters). The split up surrogate > pairs incorrectly. -- This message is automatically generated by JIRA. If you think it was sent incorrectly, please contact your JIRA administrators For more information on JIRA, see: http://www.atlassian.com/software/jira --------------------------------------------------------------------- To unsubscribe, e-mail: dev-unsubscr...@lucene.apache.org For additional commands, e-mail: dev-h...@lucene.apache.org