It may be possible due to compression artifacts in Image. You can achieve a lot of improvement using lossless compression like PNG than JPEG.
On Sunday, 26 November 2017 02:36:17 UTC+5:30, [email protected] wrote: > > Hi, > > I'm using OCR for my Raspberry Pi book reader project and I seem to have > this issue where the OCR produces extra characters along with the text > every time. I tried using some of the config parameters to see if it makes > any difference which hasn't been the case so far. > > For example, here is my output for the following text from a Dr.Seuss > book: (The text that was on the page was just Everyone wants > > a big green kangaroo.) How can I get rid of the extra characters? Thanks > > 4441‘ 'muw-u» > > > A; Wmme > > \ K r'f'. > > ,\ > > 3 I , > > E v \y r > > 3 \g\\ A ' > > '3 I» ’1,” ‘ > > ,x > > > i, 1/ / > > > I > > ¢ 1 > > > “. > > / ‘x > > m- WWW-““a”..- > > > Everyone wants > > a big green kangaroo. > > > > : is C-Bpfright Cnnwnucns. > > > r-wwrwmlaai.vn~is&‘-M'ikiww. _ . A > -- You received this message because you are subscribed to the Google Groups "tesseract-ocr" group. To unsubscribe from this group and stop receiving emails from it, send an email to [email protected]. To post to this group, send email to [email protected]. Visit this group at https://groups.google.com/group/tesseract-ocr. To view this discussion on the web visit https://groups.google.com/d/msgid/tesseract-ocr/bc14fb2c-31ad-47e5-be19-42fcd6d59c3a%40googlegroups.com. For more options, visit https://groups.google.com/d/optout.

