I am experiencing the same issue. Did you ever find a resolution for this?
On Wednesday, 25 April 2018 10:59:34 UTC-4, Youcef wrote:
>
> Hi,
>
>
> Tesseract seems to post process its prediction.
>
> Here after, what I get after OCRizing images (same font, same size images
> generated with text2image):
>
> - an image containing "12345678I" => `123456781`
> - an image containing "GLOTHUVFI" => `GLOTHUVFI`
> - an image containing "12345678H" => `12345678H`
> - an image containing "GLOTHUVFH" => `GLOTHUVFH`
> - an image containing "12345678A" => `123456784`
> - an image containing "GLOTHUVFA" => `GLOTHUVFA`
>
> It looks like Tesseract doesn't like a word with a some numbers and one
> letter at the end. In fact, if the letter looks like a number ("I" and "A"
> looks like "1" and "4" respectively), it replaces it by the closest number.
> I have tried to tune following parameters without any changement in the
> result:
>
> - segment_penalty_dict_frequent_word
> - language_model_penalty_chartype
>
> Thanks for any help.
>
> Regards
>
>
--
You received this message because you are subscribed to the Google Groups
"tesseract-ocr" group.
To unsubscribe from this group and stop receiving emails from it, send an email
to [email protected].
To post to this group, send email to [email protected].
Visit this group at https://groups.google.com/group/tesseract-ocr.
To view this discussion on the web visit
https://groups.google.com/d/msgid/tesseract-ocr/e0bb43eb-8663-4ce2-afb8-95f38c0744ce%40googlegroups.com.
For more options, visit https://groups.google.com/d/optout.