On May 17, 9:29 am, Falke <[email protected]> wrote: > On May 17, 5:50 am, Galt <[email protected]> wrote: > > I am assuming you only used one type of font? I mean -- no font > variation at all, right? I ran a test training recently, with very > little data, but mixed italic with normal and bold, and very good > recognition for the normal font, but very poor for the italics. > Yes, one kind of font of one size, for a book. The TOC (table of contents) used a smaller font, but Tess recognized it fine.
I am sorry, but I have no experience with OCR for italics. > What dpi scans did you use? 300dpi? 600dpi? (higher? i once tried > 1200dpi, and it didn't seem to improve recognition results from > 600dpi) > I am using 600 dpi and getting great results. Amongst the benefits is that it can sometimes help keep letters from touching that would be so at a lower dpi. > However, as Zdenko already said, 3.02 is already functional, and, as I > understand it, has significant improvements in accuracy. I tried Tess2.04 but it had the same issue, so this aspect of training has been around awhile. It's very likely to apply to 3.02 too. > All of your diacritics were recognized correctly? Every darn one of them. Pretty impressive. > > [...] Make sure > > your training letters look good though, solid, > > connected, and clear. When I ran this > > What exactly do you mean by "connected" ? You're not talking about > overlaps, right? (Some italic glyphs actually "lean into/across" each > other's x-space) Sorry, I mean don't use letters from scans that are of poor quality for your training. Usually, diacriticals aside, letters are connected. I suppose the dot over the i in english is one odd case. But for other letters, they are connected, so one could write with a pen or brush in a continuous fashion. Old or poor quality scans when you look at them closely can have broken letters where the ink did not run right on the page or some other issue, leaves letters that should be continuous fragmented. Tess does a great job of finding these and putting out the right letter at runtime. But your training example does NOT need to be and should not be damaged like that. If you are stuck with a rare damaged letter, fix it up with gimp or some graphic program before training. My font was not italics, and my font did not have any letters that were invading each other's space and requiring special help to separate them adequately. Although that is not super hard to achieve with various methods if it is an issue. -- You received this message because you are subscribed to the Google Groups "tesseract-ocr" group. To post to this group, send email to [email protected] To unsubscribe from this group, send email to [email protected] For more options, visit this group at http://groups.google.com/group/tesseract-ocr?hl=en

