My system is Ubuntu in version Jaunty and the tesseract installed is tesseract-ocr version 2.03-2.
By tesseract-gui.py with your image I got exactly: El otoño Se estrena con inundaciones Then I converted your image to spanish.tif and I tried again by tesseract-gui.py with the tif image and the result was the same. But in a terminal, using the tif image with '$ tesseract spanish.tif spanish -l spa' I got: El ¢:'l:<:•řî •:<:•n1 i run. Is the command wrong? On Sep 23, 1:43 pm, Barney <[email protected]> wrote: > I tried a very simple Spanish sentence today and the OCR spit out one > very odd spelling. A bit concerning. > > Input: > [IMG]http://i36.tinypic.com/35i5chj.jpg[/IMG] > > Output: > [i]El ctoiic se estrena con inundaciones[/i] > > Chinese is coming for version 3.0 right? Do you have an ETA? Since > Japanese is derived from Chinese how would I train the engine for > Japanese? Or any other languages for that matter. > > Lastly, I just want to confirm that this is wholly open source. How > does the licensing work? > > Thank you --~--~---------~--~----~------------~-------~--~----~ You received this message because you are subscribed to the Google Groups "tesseract-ocr" group. To post to this group, send email to [email protected] To unsubscribe from this group, send email to [email protected] For more options, visit this group at http://groups.google.com/group/tesseract-ocr?hl=en -~----------~----~----~----~------~----~------~--~---

