[
https://issues.apache.org/jira/browse/LUCENE-1799?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=12893413#action_12893413
]
Robert Muir commented on LUCENE-1799:
-------------------------------------
yeah it did (it didnt seem 'stable' but the first run was much different than
yours, e.g. 3300 vs 3500 or so).
I just ran with -server also [using my same 64-bit 1.6.0_19 as before]:
there is more of a difference, however not as much as yours
ret=704032704 UTF-8 encode: 32134
ret=704032704BOCU-1 encode: 36391
but go figure, if i run with my 32-bit [same jdk: 1.6.0_19], i get horrible
numbers!
here is -client
ret=684832704 UTF-8 encode: 26237
ret=684832704BOCU-1 encode: 54662
here is -server
ret=697132704 UTF-8 encode: 30062
ret=697132704BOCU-1 encode: 46293
so there is definitely an issue with 32-bit jvm, sure yours is 64-bit?
> Unicode compression
> -------------------
>
> Key: LUCENE-1799
> URL: https://issues.apache.org/jira/browse/LUCENE-1799
> Project: Lucene - Java
> Issue Type: New Feature
> Components: Store
> Affects Versions: 2.4.1
> Reporter: DM Smith
> Priority: Minor
> Attachments: Benchmark.java, Benchmark.java, Benchmark.java,
> LUCENE-1779.patch, LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799.patch,
> LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799.patch,
> LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799.patch,
> LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799_big.patch
>
>
> In lucene-1793, there is the off-topic suggestion to provide compression of
> Unicode data. The motivation was a custom encoding in a Russian analyzer. The
> original supposition was that it provided a more compact index.
> This led to the comment that a different or compressed encoding would be a
> generally useful feature.
> BOCU-1 was suggested as a possibility. This is a patented algorithm by IBM
> with an implementation in ICU. If Lucene provide it's own implementation a
> freely avIlable, royalty-free license would need to be obtained.
> SCSU is another Unicode compression algorithm that could be used.
> An advantage of these methods is that they work on the whole of Unicode. If
> that is not needed an encoding such as iso8859-1 (or whatever covers the
> input) could be used.
--
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]