[ 
https://issues.apache.org/jira/browse/TIKA-2790?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16869004#comment-16869004
 ] 

Ken Krugler commented on TIKA-2790:
-----------------------------------

Hi [[email protected]] - I finally got around to looking at your branch, and 
test code.

 
 # You're right, I'd forgotten that Yalder master branch has code that checks 
for early termination every 11th normalization, so after 110 known ngrams. In 
the branch I've been working in for a while, I've been trying a different 
approach to decide when to terminate (if N% of text "segments" are in 
agreement).
 # I was a bit confused by the timing results above, as yalder is slower than 
Optimaize & OpenNLP when early termination is disabled, and even slower on 
short text with early termination. But I assume you're using your modified 
version of Yalder, that loads all 199 language models. If so, then this will 
significantly slow down the processing (and reduce accuracy), versus using the 
same (smaller) set of languages that Optimaize supports. For an 
apples-to-apples comparison with OpenNLP, I guess you'd have to load the same 
103 language models that they support (or some intersection of the same?)

 

> Consider switching lang-detection in tika-eval to open-nlp
> ----------------------------------------------------------
>
>                 Key: TIKA-2790
>                 URL: https://issues.apache.org/jira/browse/TIKA-2790
>             Project: Tika
>          Issue Type: Improvement
>            Reporter: Tim Allison
>            Priority: Major
>         Attachments: fra_mixed_100000_0.0_0.txt, hasEnough.png, 
> langid_20190509.zip, langid_20190510.zip, langid_20190514.zip, 
> langid_20190514_plus_minus_1.zip, timeVsLength.png
>
>




--
This message was sent by Atlassian JIRA
(v7.6.3#76005)

Reply via email to