In addition to Macrons, I will also request addition of other accented
letters for Indic text transliterations.

Ray,

Will using -l eng+<new training> be the best way to handle these?

 I tried to do an add layer training, but the recognition is worse, since I
did not use many fonts for the test training. I am attaching the training
sample I used. Thanks.

Please see the following links for the transliteration schemes showing the
letters to be included.

https://en.wikipedia.org/wiki/National_Library_at_Kolkata_romanisation

https://en.wikipedia.org/wiki/International_Alphabet_of_Sanskrit_Transliteration

https://en.wikipedia.org/wiki/ISO_15919

Here are various sites that have Sanskrit corpus in transliteration.

https://github.com/cltk/sanskrit_text_dcs

http://sarit.indology.info/exist/apps/sarit/works/ (IAST)

You can see
http://gretil.sub.uni-goettingen.de/gretil_elib/Suk9441__Sukthankar_MemorialEd_1_CritStud_Mbh_1944.pdf

eg. pages 87-91 - showing use of both devanagari script as well as
trasliterated sanskrit as part of mainly English text



ShreeDevi
____________________________________________________________
भजन - कीर्तन - आरती @ http://bhajans.ramparivar.com

On Fri, Jan 20, 2017 at 10:52 AM, <[email protected]> wrote:

> Dear all,
> I frequently use Tesseract (3.04) and it’s great.
> Still, I can’t find a way to get Tesseract recognize macrons (āĀēĒīĪōŌūŪ).
> There was a discussion
> <https://groups.google.com/forum/#!searchin/tesseract-ocr/macron$20tesseract/tesseract-ocr/ZlQ3xlz_4As/XpnhTSlV3n8J>
> here about it 5 years ago but at the time, there wasn’t much of a solution.
> Things may have changed since then and I’m wondering if somebody would
> have some hints.
> Macrons are used among other things when doing recognition from japanese
> transcribed in latin alphabet (rōmaji).
> Thanks in advance for all possible ideas.
> For now, using fra or deu as one of the language, I get ô or ö…
> Best,
> Nicolas
>
> --
> You received this message because you are subscribed to the Google Groups
> "tesseract-ocr" group.
> To unsubscribe from this group and stop receiving emails from it, send an
> email to [email protected].
> To post to this group, send email to [email protected].
> Visit this group at https://groups.google.com/group/tesseract-ocr.
> To view this discussion on the web visit https://groups.google.com/d/
> msgid/tesseract-ocr/74805b35-b70b-47f8-b287-ddcd34d216e2%
> 40googlegroups.com
> <https://groups.google.com/d/msgid/tesseract-ocr/74805b35-b70b-47f8-b287-ddcd34d216e2%40googlegroups.com?utm_medium=email&utm_source=footer>
> .
> For more options, visit https://groups.google.com/d/optout.
>

-- 
You received this message because you are subscribed to the Google Groups 
"tesseract-ocr" group.
To unsubscribe from this group and stop receiving emails from it, send an email 
to [email protected].
To post to this group, send email to [email protected].
Visit this group at https://groups.google.com/group/tesseract-ocr.
To view this discussion on the web visit 
https://groups.google.com/d/msgid/tesseract-ocr/CAG2NduWBmmbBLEUBST3j%3DbgwZZnKa_GeceOjPfOhC5%2B70htdbg%40mail.gmail.com.
For more options, visit https://groups.google.com/d/optout.

Attachment: san_latn.training_text
Description: Binary data

Reply via email to