[jira] Commented: (LUCENE-1522) another highlighter

Michael McCandless (JIRA) Mon, 16 Mar 2009 15:40:13 -0700

    [ 
https://issues.apache.org/jira/browse/LUCENE-1522?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=12682494#action_12682494
 ]


Michael McCandless commented on LUCENE-1522:
--------------------------------------------


{quote}
But, if you have for example a document 'a b c a b c' and the query
'a AND b', then this approach would only highlight the first two terms,
no?
{quote}

Ahh right -- in fact, nothing would be highlighted because the scorer
for AND queries doesn't visit positions at all (it doesn't need to).

I guess we'd have to ask such scorers to forcefully visit positions &
enumerate all matches within one doc, when running in "highlight"
mode.  Hmm, feeling like a big change...

But maybe it could work.  It'd be sort of like a positional-aware
"explain", ie "show me the term occurrences that allowed the full
query to accept this document".

Imagine query "(a AND b) OR (c AND d)".  When looking at the fragments
for each doc, I would want to see both a AND b, or both c AND d, but
never (for example) just a and d.

But, flattening could produce just a and d (I think?); and I think H1
could do the same even with SpanScorer (Mark is that true?  I don't
fully understand the Query -> SpanQuery conversion).

Whereas if we could ask for positions of the "real" matches I think it
would work correctly?


> another highlighter
> -------------------
>
>                 Key: LUCENE-1522
>                 URL: https://issues.apache.org/jira/browse/LUCENE-1522
>             Project: Lucene - Java
>          Issue Type: Improvement
>          Components: contrib/highlighter
>            Reporter: Koji Sekiguchi
>            Assignee: Michael McCandless
>            Priority: Minor
>             Fix For: 2.9
>
>         Attachments: colored-tag-sample.png, LUCENE-1522.patch, 
> LUCENE-1522.patch
>
>
> I've written this highlighter for my project to support bi-gram token stream 
> (general token stream (e.g. WhitespaceTokenizer) also supported. see test 
> code in patch). The idea was inherited from my previous project with my 
> colleague and LUCENE-644. This approach needs highlight fields to be 
> TermVector.WITH_POSITIONS_OFFSETS, but is fast and can support N-grams. This 
> depends on LUCENE-1448 to get refined term offsets.
> usage:
> {code:java}
> TopDocs docs = searcher.search( query, 10 );
> Highlighter h = new Highlighter();
> FieldQuery fq = h.getFieldQuery( query );
> for( ScoreDoc scoreDoc : docs.scoreDocs ){
>   // fieldName="content", fragCharSize=100, numFragments=3
>   String[] fragments = h.getBestFragments( fq, reader, scoreDoc.doc, 
> "content", 100, 3 );
>   if( fragments != null ){
>     for( String fragment : fragments )
>       System.out.println( fragment );
>   }
> }
> {code}
> features:
> - fast for large docs
> - supports not only whitespace-based token stream, but also "fixed size" 
> N-gram (e.g. (2,2), not (1,3)) (can solve LUCENE-1489)
> - supports PhraseQuery, phrase-unit highlighting with slops
> {noformat}
> q="w1 w2"
> <b>w1 w2</b>
> ---------------
> q="w1 w2"~1
> <b>w1</b> w3 <b>w2</b> w3 <b>w1 w2</b>
> {noformat}
> - highlight fields need to be TermVector.WITH_POSITIONS_OFFSETS
> - easy to apply patch due to independent package (contrib/highlighter2)
> - uses Java 1.5
> - looks query boost to score fragments (currently doesn't see idf, but it 
> should be possible)
> - pluggable FragListBuilder
> - pluggable FragmentsBuilder
> to do:
> - term positions can be unnecessary when phraseHighlight==false
> - collects performance numbers

-- 
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

[jira] Commented: (LUCENE-1522) another highlighter

Reply via email to