Hello Eugene, Please can you explain better how to tell tessaract that a single blob has two characters? The link you sent was nothing to do with multiple letter blobs.
Thanks, Svetlin Nakov Development Manager Intelligent Software Consulting (ISC) -----Original Message----- From: [email protected] [mailto:[email protected]] On Behalf Of Eugene Reimer Sent: Wednesday, September 23, 2009 9:28 PM To: [email protected] Subject: Re: Agglutinated letters That ought to work. And you don't need your fictive letter, since the tesseract training allows for one blob to become two characters. See http://code.google.com/p/tesseract-ocr/wiki/TrainingTesseract Svetlin Nakov wrote, On 2009-09-23 11:12: > I have the following idea: add few fictive letters in the training > alphabet. For example I could add "TY" as a single letter and use some > not occupied Unicode character for this fictive letter, e.g. $. Later > when tessearct finds $ as result of OCR operation I could replace back > $ with the letter sequence "TY". This should work but I still believe > there should be more simple way to overcome such recognition errors. > --~--~---------~--~----~------------~-------~--~----~ You received this message because you are subscribed to the Google Groups "tesseract-ocr" group. To post to this group, send email to [email protected] To unsubscribe from this group, send email to [email protected] For more options, visit this group at http://groups.google.com/group/tesseract-ocr?hl=en -~----------~----~----~----~------~----~------~--~---

