Hello/Weytk, I'd like to submit a new set of training data to the project. It is data for Secwepemctsín. [1] I used two sets of tifs one was a serif and one was a sans font. I've tested the font on some books I have it and it works pretty well. I think I need to included more English langauge text because a lot of material is bilingual. The training data is very new and could and will be expanded, but I wanted to follow the philosophy of release early and release often. Should I open an issue to make this available? If you wanted to check over the work I've uploaded it here for the tim being [2].
[1] - http://languagegeek.com/salishan/secwtext.html [2] - http://secpewt.sd73.bc.ca/tesseract/tesseract-2.01.shs.tar.gz --~--~---------~--~----~------------~-------~--~----~ You received this message because you are subscribed to the Google Groups "tesseract-ocr" group. To post to this group, send email to [email protected] To unsubscribe from this group, send email to [email protected] For more options, visit this group at http://groups.google.com/group/tesseract-ocr?hl=en -~----------~----~----~----~------~----~------~--~---

