Heelo,could you tell me where (using cygwin and your schema of directories from manual) should i put tif files with box files to train and tif with box files to eval? I just cant get trough that part.
W dniu wtorek, 5 lutego 2019 16:35:07 UTC+1 użytkownik Kristóf Horváth napisał: > > The way you know how to train with boxes seems very much like version > 3.0.5 learning. Im currently only interested in LSTM learning (aka > Tesseract 4.*). Since i last wrote i managed to dig up something that > actually helped me a lot in both understanding the process and actually > training data. > So I was pointed at the following github repo > <https://github.com/OCR-D/ocrd-train>. I run my training on Win10 64bit > with cygwin. I just had to put the files from repo in folder and inside > execute (order counts) make leptonica and then make tesseract, with these > steps you compiled tesseract and leptonica. > (other links: 1 > <https://groups.google.com/forum/#!topic/tesseract-ocr/CLFM6dw5HIM>, 2 > <https://groups.google.com/forum/#!searchin/tesseract-ocr/l.bolzani%7Csort:date/tesseract-ocr/be4-rjvY2tQ/32evtMHlAQAJ>, > > 3 > <https://groups.google.com/forum/#!msg/tesseract-ocr/COJ4IjcrL6s/C1OeE9bWBgAJ>) > > Before you start training there are some changes to be made on makefile and > generate_line_box.py. I opened them with notepad++. In makefile you have to > change variables according to your stuff and in .py you have to change a > few lines (here <https://github.com/OCR-D/ocrd-train/issues/18> you can > see why). If you make these changes you can run make training and it will > get your tiff/text file pairs from ground-truth and auto create every > training file. Essentially if you understand what the makefile does you > will understand the process of training and the key to make it for your own > specifications is to edit that makefile to your needs. What i mean for > example its hardcoded for max 10000 iterations so if you want something > else you have to change it in the script. > The script itself executes codes that you can find on wiki but for me the > script gave more details because i could see what steps are there and what > for do i need them. > > 2019. február 5., kedd 15:32:05 UTC+1 időpontban WidmoPL a következőt írta: >> >> Hello, >> Thanks for quick reply. >> Yeah, from what i found, or rather didn't find, there is very little >> documentation or it is poorly written,especially for version 4. >> I only know how to train using boxes and then editing them in some >> ''editor'' when you must check letters with rectangulares. This is I know >> well by now,so if you need some info to this part I will gladly help. But >> this method is very time expensive, despite results looks promising. Also >> noticed that when you recognize text having trained data done this way its >> most time better not to use standard trained data + your trained data, just >> only use your trained data. >> Example: >> tesseract qwe.arial.exp0.png qwe.arial.exp0 -l xyz -> where >> xyz is name of your trained data file >> >> AND NOT -> tesseract qwe.arial.exp0.png qwe.arial.exp0 -l eng+xyz >> >> >> Thats all interesting i can say for now. >> Thanks for reply! >> >> >> W dniu poniedziałek, 4 lutego 2019 14:33:09 UTC+1 użytkownik Kristóf >> Horváth napisał: >>> >>> Oh boy, where to start! So first of all you are not alone not finding >>> any information. Currently i am a week ahead of you, so im gonna share what >>> i found out. >>> Lets start with training_files.txt. Whats inside? >>> /home/kh/tesstutorial/engtrain/eng.Arial.exp0.lstmf/home/kh/tesstutorial/engtrain/eng.Impact_Condensed.exp0.lstmf >>> What is it for? it holds the location of .lstmf files, which are >>> required for training (tesseract 4.*) >>> How do i get it? if you manage to create starter training files it will >>> create it with it so you dont have to write it on your own. >>> Whats the difference between eval and train? as far as i know the only >>> difference is their place in the training proccess. So i think you compare >>> results from train training_files to eval training_files. (Sadly im still >>> stuck on this question) >>> >>> Thats all i know about these training files >>> >>> And for your second question: If thats possible i wanna know how. >>> Currently for me it is a controversal question because there is no >>> documentation and what you can find says that its not supported but others >>> say oh you can do it ( >>> https://groups.google.com/forum/#!topic/tesseract-ocr/0r8QvV3j8ew) >>> >>> Hopefully i could help, currently im working on a newbie friendly >>> documentation for tesseract 4.* training so if you get some info on topic, >>> i would appreciate if you would send it. >>> Good luck for going further into topic. >>> >>> 2019. február 4., hétfő 12:15:13 UTC+1 időpontban Krzysztof Kanafa a >>> következőt írta: >>>> >>>> Helllo >>>> >>>> I'm completely new in tesseract, first version I'm using is 4.0.0. >>>> Sorry for noob question, but I really didn't find answer despite quite >>>> long >>>> searching. >>>> Its about these two options -> eval_listfile and train_listfile. What >>>> exactly should be in these files? >>>> Is there in train_listfile.txt should be list of images that we want to >>>> be learned? for example images with scanned letters: >>>> 1. first image "image1.tif" -> ABCDEF >>>> 2. second "image2.tif" -> GHIJ >>>> And in eval_listfile.txt a list of same named as images text files that >>>> each contain good solution? >>>> Example: >>>> 1.first txt file "image1.txt" with text inside -> ABCDEF >>>> 2. second txt file "image2.txt" ->GHIJ >>>> >>>> Or is it about completely something different? >>>> >>>> >>>> >>>> And second question ..: >>>> when i ve already described examples... let's say i have such two TIF >>>> files as mentioned before (photos with some serial numbers from a kind of >>>> labels - each photo has one serial no). >>>> And in two txt files I wrote by hand these serial numbers (just read it >>>> from photo - one serial in one txt file). >>>> >>>> *Is there a way to just point these files *saying tesseract -> >>>> >>>> "here are some photo and here are some txt files telling what is on >>>> each photo" >>>> >>>> and make him learn such way? for example if i have 60 of such tif+txt >>>> paris, in this case selecting one by one letter in each tif (using BOX >>>> files of course) will take forever, and what if i would have 50 pairs but >>>> each tiff have 50 pages, that would take forever.. >>>> Again sorry foor noob question, I read about tesseract for few days >>>> and it all mixed up in head. >>>> >>>> Greetings >>>> Krzysiek >>>> >>> -- You received this message because you are subscribed to the Google Groups "tesseract-ocr" group. To unsubscribe from this group and stop receiving emails from it, send an email to [email protected]. To post to this group, send email to [email protected]. Visit this group at https://groups.google.com/group/tesseract-ocr. To view this discussion on the web visit https://groups.google.com/d/msgid/tesseract-ocr/3c7085ac-66ad-4252-98ef-c1b87e905abe%40googlegroups.com. For more options, visit https://groups.google.com/d/optout.

