That ought to work.  And you don't need your fictive letter, since the 
tesseract training allows for one blob to become two characters.  See 
http://code.google.com/p/tesseract-ocr/wiki/TrainingTesseract


Svetlin Nakov wrote, On 2009-09-23 11:12:

> I have the following idea: add few fictive letters in the training 
> alphabet. For example I could add “TY” as a single letter and use some 
> not occupied Unicode character for this fictive letter, e.g. $. Later 
> when tessearct finds $ as result of OCR operation I could replace back 
> $ with the letter sequence “TY”. This should work but I still believe 
> there should be more simple way to overcome such recognition errors.
>


--~--~---------~--~----~------------~-------~--~----~
You received this message because you are subscribed to the Google Groups 
"tesseract-ocr" group.
To post to this group, send email to [email protected]
To unsubscribe from this group, send email to 
[email protected]
For more options, visit this group at 
http://groups.google.com/group/tesseract-ocr?hl=en
-~----------~----~----~----~------~----~------~--~---

Reply via email to