[
https://issues.apache.org/jira/browse/TIKA-2293?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15941164#comment-15941164
]
Thejan Wijesinghe commented on TIKA-2293:
-----------------------------------------
Thank you, [[email protected]] and [~thammegowda] for your responses.
Tim,
I understand your concern of trained data files taking so much space. Doubling
the size of tika-app and tika-server for a single component of TIKA is not at
all a practical thing to do. I am happy that you are willing to promote this
parser :)
Thamme,
Yes :), I will make this an independent parser that can be pluggable to TIKA,
then I will write a wiki page linking it to my repo.
> Tess4jOCRParser - A simpler Java version of TesseractOCRParser
> ---------------------------------------------------------------
>
> Key: TIKA-2293
> URL: https://issues.apache.org/jira/browse/TIKA-2293
> Project: Tika
> Issue Type: Improvement
> Components: ocr
> Reporter: Thejan Wijesinghe
> Fix For: 1.15
>
>
> Right now, TesseractOCRParser calls tesseract and imagemagick from command
> line. Intention of this new parser "Tess4jOCRParser" is to use the Tess4J API
> instead of the runtime.exec way to executing tesseract out of process.
--
This message was sent by Atlassian JIRA
(v6.3.15#6346)