Hello Svetlin, Look for the part that says "or if the pair is common, put both characters at the start of the line..."
Svetlin Nakov wrote, On 2009-09-23 16:37: >Hello Eugene, > >Please can you explain better how to tell tessaract that a single blob has >two characters? The link you sent was nothing to do with multiple letter >blobs. > >Thanks, > >Svetlin Nakov >Development Manager >Intelligent Software Consulting (ISC) >-----Original Message----- >From: [email protected] [mailto:[email protected]] >On Behalf Of Eugene Reimer >Sent: Wednesday, September 23, 2009 9:28 PM >To: [email protected] >Subject: Re: Agglutinated letters > > >That ought to work. And you don't need your fictive letter, since the >tesseract training allows for one blob to become two characters. See >http://code.google.com/p/tesseract-ocr/wiki/TrainingTesseract > > >Svetlin Nakov wrote, On 2009-09-23 11:12: > > > >>I have the following idea: add few fictive letters in the training >>alphabet. For example I could add "TY" as a single letter and use some >>not occupied Unicode character for this fictive letter, e.g. $. Later >>when tessearct finds $ as result of OCR operation I could replace back >>$ with the letter sequence "TY". This should work but I still believe >>there should be more simple way to overcome such recognition errors. >> >-~----------~----~----~----~------~----~------~--~--- > > > > --~--~---------~--~----~------------~-------~--~----~ You received this message because you are subscribed to the Google Groups "tesseract-ocr" group. To post to this group, send email to [email protected] To unsubscribe from this group, send email to [email protected] For more options, visit this group at http://groups.google.com/group/tesseract-ocr?hl=en -~----------~----~----~----~------~----~------~--~---

