Hello Svetlin,
Look for the part that says "or if the pair is common, put both 
characters at the start of the line..."


Svetlin Nakov wrote, On 2009-09-23 16:37:

>Hello Eugene,
>
>Please can you explain better how to tell tessaract that a single blob has
>two characters? The link you sent was nothing to do with multiple letter
>blobs.
>
>Thanks,
>
>Svetlin Nakov
>Development Manager
>Intelligent Software Consulting (ISC)
>-----Original Message-----
>From: [email protected] [mailto:[email protected]]
>On Behalf Of Eugene Reimer
>Sent: Wednesday, September 23, 2009 9:28 PM
>To: [email protected]
>Subject: Re: Agglutinated letters
>
>
>That ought to work.  And you don't need your fictive letter, since the 
>tesseract training allows for one blob to become two characters.  See 
>http://code.google.com/p/tesseract-ocr/wiki/TrainingTesseract
>
>
>Svetlin Nakov wrote, On 2009-09-23 11:12:
>
>  
>
>>I have the following idea: add few fictive letters in the training 
>>alphabet. For example I could add "TY" as a single letter and use some 
>>not occupied Unicode character for this fictive letter, e.g. $. Later 
>>when tessearct finds $ as result of OCR operation I could replace back 
>>$ with the letter sequence "TY". This should work but I still believe 
>>there should be more simple way to overcome such recognition errors.
>>
>-~----------~----~----~----~------~----~------~--~---
>
>
>  
>


--~--~---------~--~----~------------~-------~--~----~
You received this message because you are subscribed to the Google Groups 
"tesseract-ocr" group.
To post to this group, send email to [email protected]
To unsubscribe from this group, send email to 
[email protected]
For more options, visit this group at 
http://groups.google.com/group/tesseract-ocr?hl=en
-~----------~----~----~----~------~----~------~--~---

Reply via email to