My system is Ubuntu in version Jaunty and the tesseract installed is
tesseract-ocr version 2.03-2.

By tesseract-gui.py with your image I got exactly:

El otoño Se estrena
con inundaciones

Then I converted your image to spanish.tif and I tried again by
tesseract-gui.py with the tif image and the result was the same.

But in a terminal, using the tif image with '$ tesseract spanish.tif
spanish -l spa' I got:

El ¢:'l:<:•řî
•:<:•n1 i run.

Is the command wrong?

On Sep 23, 1:43 pm, Barney <[email protected]> wrote:
> I tried a very simple Spanish sentence today and the OCR spit out one
> very odd spelling. A bit concerning.
>
> Input:
> [IMG]http://i36.tinypic.com/35i5chj.jpg[/IMG]
>
> Output:
> [i]El ctoiic se estrena con inundaciones[/i]
>
> Chinese is coming for version 3.0 right? Do you have an ETA? Since
> Japanese is derived from Chinese how would I train the engine for
> Japanese? Or any other languages for that matter.
>
> Lastly, I just want to confirm that this is wholly open source. How
> does the licensing work?
>
> Thank you
--~--~---------~--~----~------------~-------~--~----~
You received this message because you are subscribed to the Google Groups 
"tesseract-ocr" group.
To post to this group, send email to [email protected]
To unsubscribe from this group, send email to 
[email protected]
For more options, visit this group at 
http://groups.google.com/group/tesseract-ocr?hl=en
-~----------~----~----~----~------~----~------~--~---

Reply via email to