Hi, I do not know how to create image training files ..even with the seemingly helpful notes in trainging Tesseract 3.
I want to know what the minimum number of files is..??? Is it two files or one file ??? eg 1)Is it an image file with a text file containing a copy of what is in the image file with a closest match font. or 2) just an image file alone. Then .. how does one name the file/s ? Is this important ? If there are two files minimum eg image file and complementary text file can I please see some examples ? Cheers Richard On Thursday, May 17, 2012 7:50:36 PM UTC+10, Galt wrote: > > SUCCESS AT LAST! > > I have used this simple training text and the output is highly > accurate. > I am very happy to have succeeded at last. I only wish the > documentation > had warned more explicitly what is needed in training. > > Here is what worked for me: > > Start each line with a capital letter that you need. > Try hard to avoid any other capitals in the line. > Use punctuation, but very naturally. > Don't make the lines too short. > Lots of repetition did not seem to be required. > There seems to be no real need to train Tess > on screwed up scans with damaged letters -- > if anything it just confuses tess. Make sure > your training letters look good though, solid, > connected, and clear. When I ran this > on scans, it performed extremely accurately, > even when the real scans had breaks in > the letters or parts of quotes were missing. > Finally I see why Tess has something to offer! > If the training documentation were better > about guiding people from getting a messed up > model, that would save them a lot of time. > > Arán ar maidin! > Áḃar ar biṫ ba ṁaiṫ leat. > Ba é an fear cliste é. > Ḃí muid ag iarraiḋ dul ann. > Cé hé an duine úd ṫall? > Ċonaic siḃ gaċ rud: > Druid an doras, le do ṫoil! > Ḋein sí rud air. > Earrach - an séasúr is fearr. > Éire: is grá liom ṫu. > Fuair siad an dea-ṗost. > Máthair Ḟinn mac Cuṁaill is ea í. > Go raiḃ maiṫ agaiḃ, a ḋaoine uaisle? > Ġeall sé dúinn go raiḃ sé fíor. > Haló! An ḃfuil duine sa teaċ? > Is é atá pósta lena ḃean. > Íde béil a ṫug sé don ḟear. > Leipreacán a dúirt liom é. > Ṁeas an ḃean ṡaiḃir nach raiḃ siad go breá. > Ná taḃair aird ar bith dó. > Oraiste a ṫug sé don ġasúr. > Ón droch-rud a ṫagann olc! > Páid is ainm dó. > Ṗós siad go luath ina ḋiaiḋ sin. > Rith sí léi go gasta. > “Seal ṫuas, seal ṫíos.” > Ṡíl mé go dtiocfainn, ach níor ṫangas. > Tá an cáilín fós ann. > Ṫall is aḃus, sin an áit a bí siad. > Uinnsean Morlei is ainm do. > Úll dón ṁúinteoir, a ṁic léinn! > ’Sé an bealach ceart. > Ċuaiġ sé im’ intinn ḟéin. > “Is maith an rud é.” > Duirt sé, “Ciúnas!” > > > I ran this on a 75 page book > and virtually every page was perfect, > even though the scans themselves were not. > I admit that I took the time to manually > go through each page and remove specks. > > Actually there were about 3 pages on which > Tess still got confused by the high quotes, > but after this success, I can fix those manually. > -- You received this message because you are subscribed to the Google Groups "tesseract-ocr" group. To post to this group, send email to [email protected] To unsubscribe from this group, send email to [email protected] For more options, visit this group at http://groups.google.com/group/tesseract-ocr?hl=en

