[
https://issues.apache.org/jira/browse/TIKA-2790?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16869004#comment-16869004
]
Ken Krugler commented on TIKA-2790:
-----------------------------------
Hi [[email protected]] - I finally got around to looking at your branch, and
test code.
# You're right, I'd forgotten that Yalder master branch has code that checks
for early termination every 11th normalization, so after 110 known ngrams. In
the branch I've been working in for a while, I've been trying a different
approach to decide when to terminate (if N% of text "segments" are in
agreement).
# I was a bit confused by the timing results above, as yalder is slower than
Optimaize & OpenNLP when early termination is disabled, and even slower on
short text with early termination. But I assume you're using your modified
version of Yalder, that loads all 199 language models. If so, then this will
significantly slow down the processing (and reduce accuracy), versus using the
same (smaller) set of languages that Optimaize supports. For an
apples-to-apples comparison with OpenNLP, I guess you'd have to load the same
103 language models that they support (or some intersection of the same?)
> Consider switching lang-detection in tika-eval to open-nlp
> ----------------------------------------------------------
>
> Key: TIKA-2790
> URL: https://issues.apache.org/jira/browse/TIKA-2790
> Project: Tika
> Issue Type: Improvement
> Reporter: Tim Allison
> Priority: Major
> Attachments: fra_mixed_100000_0.0_0.txt, hasEnough.png,
> langid_20190509.zip, langid_20190510.zip, langid_20190514.zip,
> langid_20190514_plus_minus_1.zip, timeVsLength.png
>
>
--
This message was sent by Atlassian JIRA
(v7.6.3#76005)