Hello Doug , In the attachment, you may find my implementation. I use the second way (I have no idea which way is better without testing against performance, precision, recall etc). i.e.,
"blah C1C2C3 blah C4 blah" ->
"blah", "C1", "C1C2"+0, "C2", "C2C3"+0, "C3", "blah",
"C4", "blah"
(+0 to indicate a positionIncrement of zero, rather
than 1.)
By the way, is there any way in lucene API which can
tell you the occurrences of a Term in the whole
collection, not just the "idf" value. I try to
implement a language model (probabilistic) based
similarity measure.
Best regards,
Jun Cai
======= At 2004-06-09, 11:36:43 you wrote: =======
>Jun Cai wrote:
>> I feel interest on the task you gave to me and I
>> started to look into the code in analysis package.
>
>Great! Thanks for looking into this.
>
>> I found something strange when I input some CJK
>> characters to NutchDocumentTokenizer:
>>
>> Nothing is output. And program does not even go
into
>> next() routine.
>>
>> Is it a bug? Because I am absolute beginer of
Javacc,
>> I can not figure out where the problem is
immediately.
>
>Yes, this is bug. Thanks for finding it.
>
>It looks like 'UNICODE_INPUT = true' option must be
specified. The
>JavaCC documentation says that this is ignored when
'USER_CHAR_STREAM =
>true', but that doesn't seem to be the case.
>
>When I set 'DEBUG_TOKEN_MANAGER = true', then I see:
>
>$ ant -Djavacc.home=c:\\javacc-3.2 generate-src
compile-core
>$ bin/nutch net.nutch.analysis.NutchDocumentTokenizer
>Tokens: Current character : \u8521 (34081)
> No string literal matches possible.
> Starting NFA to match one of : { <WORD>,
<ACRONYM>, <SIGRAM> }
>Current character : \u8521 (34081)
> Current character matched as a <WHITE> token.
>****** FOUND A <WHITE> MATCH (\u8521) ******
>
>Returning the <EOF> token.
>
>But when I then define 'UNICODE_INPUT = true', I see:
>
>$ ant -Djavacc.home=c:\\javacc-3.2 generate-src
compile-core
>$ bin/nutch net.nutch.analysis.NutchDocumentTokenizer
>Tokens: Current character : \u8521 (34081)
> No string literal matches possible.
> Starting NFA to match one of : { <WORD>,
<ACRONYM>, <SIGRAM> }
>Current character : \u8521 (34081)
> Currently matched the first 1 characters as a
<SIGRAM> token.
>****** FOUND A <SIGRAM> MATCH (\u8521) ******
>
>? Returning the <EOF> token.
>
>That '?' is the \u8521, as displayed by my shell.
>
>JavaCC's DEBUG_TOKEN_MANAGER option is very handy
here.
>
>I'll commit the fix.
>
>> To create Bi-gram for CJK, is it ok that just
change
>> the next() method? When it see CJK character, it
looks
>> one character ahead and concatenate two characters
>> together for create a new lucene Token.
>
>I think it would be better to implement this as a
separate filter, like
>CommonGrams.java. This should be simpler than
CommonGrams.java, which
>implements n-grams, not just bigrams. Since we're
only interested in
>bigrams, only one token needs to be queued, so a
buffered token field
>will suffice, instead of a queue. Does that make
sense?
>
>Also, ff C1, C2, etc. represent CJK characters, given
the text:
> "blah C1C2C3 blah C4 blah"
>should we index:
> "blah", "C1C2", "C2C3", "blah" "C4", "blah"
>or
> "blah", "C1", "C1C2"+0, "C2", "C2C3"+0, "C3",
"blah", "C4", "blah"
>?
>
>(I use +0 to indicate a positionIncrement of zero,
rather than 1.)
>
>The latter would permit folks to still search for C1,
C2 or C3 as
>one-word queries. But if they search for C1C2C3,
then we'd convert the
>query into a search for the bigram phrase "C1C2
C2C3". I think that's
>the better approach. What do others think?
>
>Once NutchDocumentAnalysis has been modified to
include a
>CJKBigramFilter or somesuch, then the query parser
(NutchAnalysis.jj)
>and/or translator (QueryTranslator.java) must also be
modified to
>construct bigram phrases when appropriate.
>
>Cheers,
>
>Doug
>
>
>
>
>
>-------------------------------------------------------
>This SF.Net email is sponsored by: GNOME Foundation
>Hackers Unite! GUADEC: The world's #1 Open Source
Desktop Event.
>GNOME Users and Developers European Conference,
28-30th June in Norway
>http://2004/guadec.org
>_______________________________________________
>Nutch-developers mailing list
>[EMAIL PROTECTED]
>https://lists.sourceforge.net/lists/listinfo/nutch-developers
>.
= = = = = = = = = = = = = = = = = = = =
Jun Cai
[EMAIL PROTECTED]
2004-06-10
__________________________________
Do you Yahoo!?
Friends. Fun. Try the all-new Yahoo! Messenger.
http://messenger.yahoo.com/
CJKBigramFilterTest.java
Description: CJKBigramFilterTest.java
CJKBigramFilter.java
Description: CJKBigramFilter.java
