On May 17, 5:50 am, Galt <[email protected]> wrote: > SUCCESS AT LAST! > > I have used this simple training text and the output is highly > accurate. > I am very happy to have succeeded at last. I only wish the > documentation > had warned more explicitly what is needed in training. > > Here is what worked for me:
A thousand blessings on your head, and then some! It sounds amazingly good that you achieved such accuracy with such a small amount of training data. I am assuming you only used one type of font? I mean -- no font variation at all, right? I ran a test training recently, with very little data, but mixed italic with normal and bold, and very good recognition for the normal font, but very poor for the italics. What dpi scans did you use? 300dpi? 600dpi? (higher? i once tried 1200dpi, and it didn't seem to improve recognition results from 600dpi) What you say about starting each line with a distinct uppercase letter is very interesting; i will be trying this, too. And all the other details, about not needing that much repetition, etc., etc. Sounds like the existing documentation DOES need to be updated/ clarified. However, as Zdenko already said, 3.02 is already functional, and, as I understand it, has significant improvements in accuracy. I do hope your discoveries also apply to 3.02. I will try to experiment with your methods on 3.02 (i would be interested to know how 3.02 works for you, too, in light of your new findings) All of your diacritics were recognized correctly? > Start each line with a capital letter that you need. > Try hard to avoid any other capitals in the line. > Use punctuation, but very naturally. > Don't make the lines too short. > Lots of repetition did not seem to be required. > There seems to be no real need to train Tess > on screwed up scans with damaged letters -- > if anything it just confuses tess. Make sure > your training letters look good though, solid, > connected, and clear. When I ran this What exactly do you mean by "connected" ? You're not talking about overlaps, right? (Some italic glyphs actually "lean into/across" each other's x-space) > on scans, it performed extremely accurately, > even when the real scans had breaks in > the letters or parts of quotes were missing. > Finally I see why Tess has something to offer! > If the training documentation were better > about guiding people from getting a messed up > model, that would save them a lot of time. > > Arán ar maidin! > Áḃar ar biṫ ba ṁaiṫ leat. > Ba é an fear cliste é. > Ḃí muid ag iarraiḋ dul ann. > Cé hé an duine úd ṫall? > Ċonaic siḃ gaċ rud: > Druid an doras, le do ṫoil! > Ḋein sí rud air. > Earrach - an séasúr is fearr. > Éire: is grá liom ṫu. > Fuair siad an dea-ṗost. > Máthair Ḟinn mac Cuṁaill is ea í. > Go raiḃ maiṫ agaiḃ, a ḋaoine uaisle? > Ġeall sé dúinn go raiḃ sé fíor. > Haló! An ḃfuil duine sa teaċ? > Is é atá pósta lena ḃean. > Íde béil a ṫug sé don ḟear. > Leipreacán a dúirt liom é. > Ṁeas an ḃean ṡaiḃir nach raiḃ siad go breá. > Ná taḃair aird ar bith dó. > Oraiste a ṫug sé don ġasúr. > Ón droch-rud a ṫagann olc! > Páid is ainm dó. > Ṗós siad go luath ina ḋiaiḋ sin. > Rith sí léi go gasta. > “Seal ṫuas, seal ṫíos.” > Ṡíl mé go dtiocfainn, ach níor ṫangas. > Tá an cáilín fós ann. > Ṫall is aḃus, sin an áit a bí siad. > Uinnsean Morlei is ainm do. > Úll dón ṁúinteoir, a ṁic léinn! > ’Sé an bealach ceart. > Ċuaiġ sé im’ intinn ḟéin. > “Is maith an rud é.” > Duirt sé, “Ciúnas!” > > I ran this on a 75 page book > and virtually every page was perfect, > even though the scans themselves were not. > I admit that I took the time to manually > go through each page and remove specks. > > Actually there were about 3 pages on which > Tess still got confused by the high quotes, > but after this success, I can fix those manually. -- You received this message because you are subscribed to the Google Groups "tesseract-ocr" group. To post to this group, send email to [email protected] To unsubscribe from this group, send email to [email protected] For more options, visit this group at http://groups.google.com/group/tesseract-ocr?hl=en

