[
https://issues.apache.org/jira/browse/LUCENE-2167?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=12887741#action_12887741
]
Steven Rowe commented on LUCENE-2167:
-------------------------------------
I ran it three more times, and it appears that the difference between
ClassicTokenizer, UAX29Tokenizer, and the new StandardTokenizer is in the noise:
||Operation||recsPerRun||rec/s||elapsedSec||
|ClassicTokenizer|1262799|665,682.12|1.90|
|ICUTokenizer|1268451|553,666.94|2.29|
|RBBITokenizer|1268451|575,261.25|2.20|
|StandardTokenizer|1268450|658,935.06|1.92|
|UAX29Tokenizer|1268451|642,579.00|1.97|
||Operation||recsPerRun||rec/s||elapsedSec||
|ClassicTokenizer|1262799|668,501.31|1.89|
|ICUTokenizer|1268451|546,275.19|2.32|
|RBBITokenizer|1268451|563,255.31|2.25|
|StandardTokenizer|1268450|651,824.25|1.95|
|UAX29Tokenizer|1268451|664,806.62|1.91|
||Operation||recsPerRun||rec/s||elapsedSec||
|ClassicTokenizer|1262799|674,932.69|1.87|
|ICUTokenizer|1268451|541,841.50|2.34|
|RBBITokenizer|1268451|586,431.38|2.16|
|StandardTokenizer|1268450|635,814.56|2.00|
|UAX29Tokenizer|1268451|650,487.69|1.95|
> Implement StandardTokenizer with the UAX#29 Standard
> ----------------------------------------------------
>
> Key: LUCENE-2167
> URL: https://issues.apache.org/jira/browse/LUCENE-2167
> Project: Lucene - Java
> Issue Type: New Feature
> Components: contrib/analyzers
> Affects Versions: 3.1
> Reporter: Shyamal Prasad
> Assignee: Robert Muir
> Priority: Minor
> Attachments: LUCENE-2167-jflex-tld-macro-gen.patch,
> LUCENE-2167-jflex-tld-macro-gen.patch, LUCENE-2167-jflex-tld-macro-gen.patch,
> LUCENE-2167-lucene-buildhelper-maven-plugin.patch,
> LUCENE-2167.benchmark.patch, LUCENE-2167.benchmark.patch,
> LUCENE-2167.benchmark.patch, LUCENE-2167.patch, LUCENE-2167.patch,
> LUCENE-2167.patch, LUCENE-2167.patch, LUCENE-2167.patch, LUCENE-2167.patch,
> LUCENE-2167.patch, LUCENE-2167.patch, LUCENE-2167.patch, LUCENE-2167.patch,
> LUCENE-2167.patch, LUCENE-2167.patch, LUCENE-2167.patch, LUCENE-2167.patch,
> standard.zip
>
> Original Estimate: 0.5h
> Remaining Estimate: 0.5h
>
> It would be really nice for StandardTokenizer to adhere straight to the
> standard as much as we can with jflex. Then its name would actually make
> sense.
> Such a transition would involve renaming the old StandardTokenizer to
> EuropeanTokenizer, as its javadoc claims:
> bq. This should be a good tokenizer for most European-language documents
> The new StandardTokenizer could then say
> bq. This should be a good tokenizer for most languages.
> All the english/euro-centric stuff like the acronym/company/apostrophe stuff
> can stay with that EuropeanTokenizer, and it could be used by the european
> analyzers.
--
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]