3.0 is available by getting the latest sources using svn. I do not believe there is an official release package yet. I do not see Chinese or Japanese yet. 3.0 should handle the Spanish sentence. Before 3.0, tesseract performance was lower on small amounts of text. There is a wiki with documentation in the same area with the source distributions. It includes a doc about training. It is not updated yet with all the new info for 3.0, though. For Spanish, I suspect it would do fine without training.
On Sep 23, 4:43 am, Barney <[email protected]> wrote: > I tried a very simple Spanish sentence today and the OCR spit out one > very odd spelling. A bit concerning. > > Input: > [IMG]http://i36.tinypic.com/35i5chj.jpg[/IMG] > > Output: > [i]El ctoiic se estrena con inundaciones[/i] > > Chinese is coming for version 3.0 right? Do you have an ETA? Since > Japanese is derived from Chinese how would I train the engine for > Japanese? Or any other languages for that matter. > > Lastly, I just want to confirm that this is wholly open source. How > does the licensing work? > > Thank you --~--~---------~--~----~------------~-------~--~----~ You received this message because you are subscribed to the Google Groups "tesseract-ocr" group. To post to this group, send email to [email protected] To unsubscribe from this group, send email to [email protected] For more options, visit this group at http://groups.google.com/group/tesseract-ocr?hl=en -~----------~----~----~----~------~----~------~--~---

