3.0 is available by getting the latest sources using svn.  I do not
believe there is an official release package yet.  I do not see
Chinese or Japanese yet.  3.0 should handle the Spanish sentence.
Before 3.0, tesseract performance was lower on small amounts of text.
There is a wiki with documentation in the same area with the source
distributions.  It includes a doc about training.  It is not updated
yet with all the new info for 3.0, though.  For Spanish, I suspect it
would do fine without training.

On Sep 23, 4:43 am, Barney <[email protected]> wrote:
> I tried a very simple Spanish sentence today and the OCR spit out one
> very odd spelling. A bit concerning.
>
> Input:
> [IMG]http://i36.tinypic.com/35i5chj.jpg[/IMG]
>
> Output:
> [i]El ctoiic se estrena con inundaciones[/i]
>
> Chinese is coming for version 3.0 right? Do you have an ETA? Since
> Japanese is derived from Chinese how would I train the engine for
> Japanese? Or any other languages for that matter.
>
> Lastly, I just want to confirm that this is wholly open source. How
> does the licensing work?
>
> Thank you
--~--~---------~--~----~------------~-------~--~----~
You received this message because you are subscribed to the Google Groups 
"tesseract-ocr" group.
To post to this group, send email to [email protected]
To unsubscribe from this group, send email to 
[email protected]
For more options, visit this group at 
http://groups.google.com/group/tesseract-ocr?hl=en
-~----------~----~----~----~------~----~------~--~---

Reply via email to