[ https://issues.apache.org/jira/browse/LUCENE-1824?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=13099157#comment-13099157 ]
Robert Muir commented on LUCENE-1824: ------------------------------------- Thanks for adding breakiterator implementations! the implementation seems independent of what type of breakiterator it uses, so maybe its simpler for it to just be BreakIteratorBoundaryScanner(BreakIterator bi), and then the user can create the breakiterator however they like (they could even pass in a custom subclass, for expert control) ? > FastVectorHighlighter truncates words at beginning and end of fragments > ----------------------------------------------------------------------- > > Key: LUCENE-1824 > URL: https://issues.apache.org/jira/browse/LUCENE-1824 > Project: Lucene - Java > Issue Type: Improvement > Components: modules/highlighter > Environment: any > Reporter: Alex Vigdor > Assignee: Koji Sekiguchi > Priority: Minor > Fix For: 4.0 > > Attachments: LUCENE-1824.patch, LUCENE-1824.patch, LUCENE-1824.patch, > LUCENE-1824.patch > > > FastVectorHighlighter does not take word boundaries into consideration when > building fragments, so that in most cases the first and last word of a > fragment are truncated. This makes the highlights less legible than they > should be. I will attach a patch to BaseFragmentBuilder that resolves this > by expanding the start and end boundaries of the fragment to the first > whitespace character on either side of the fragment, or the beginning or end > of the source text, whichever comes first. This significantly improves > legibility, at the cost of returning a slightly larger number of characters > than specified for the fragment size. -- This message is automatically generated by JIRA. For more information on JIRA, see: http://www.atlassian.com/software/jira --------------------------------------------------------------------- To unsubscribe, e-mail: dev-unsubscr...@lucene.apache.org For additional commands, e-mail: dev-h...@lucene.apache.org