[jira] [Commented] (TIKA-2293) Tess4jOCRParser - A simpler Java version of TesseractOCRParser

Thejan Wijesinghe (JIRA) Fri, 24 Mar 2017 14:27:07 -0700

    [ 
https://issues.apache.org/jira/browse/TIKA-2293?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15941164#comment-15941164
 ]


Thejan Wijesinghe commented on TIKA-2293:
-----------------------------------------

Thank you, [[email protected]] and [~thammegowda] for your responses.
Tim, 
I understand your concern of trained data files taking so much space. Doubling 
the size of tika-app and tika-server for a single component of TIKA is not at 
all a practical thing to do. I am happy that you are willing to promote this 
parser :)

Thamme,
Yes :), I will make this an independent parser that can be pluggable to TIKA, 
then I will write a wiki page linking it to my repo.   

>  Tess4jOCRParser - A simpler Java version of TesseractOCRParser
> ---------------------------------------------------------------
>
>                 Key: TIKA-2293
>                 URL: https://issues.apache.org/jira/browse/TIKA-2293
>             Project: Tika
>          Issue Type: Improvement
>          Components: ocr
>            Reporter: Thejan Wijesinghe
>             Fix For: 1.15
>
>
> Right now, TesseractOCRParser calls tesseract and imagemagick from command 
> line. Intention of this new parser "Tess4jOCRParser" is to use the Tess4J API 
> instead of the runtime.exec way to executing tesseract out of process.  



--
This message was sent by Atlassian JIRA
(v6.3.15#6346)

[jira] [Commented] (TIKA-2293) Tess4jOCRParser - A simpler Java version of TesseractOCRParser

Reply via email to