El ds 02 de 04 de 2011 a les 17:53 -0400, en/na Ahsan va escriure: > Hi all, > > > This is regarding the adopting of the existing ur-hi language pair. > Currently, the ur-hi pair has much work done with both urdu and hindi > languages. > > > == Motivation == > > > The field of machine translation has gained much traction in the past > couple of decades. Its use of translating between different language > pairs is of political, military and international importance. > Apertium was quite interesting to me when I first came across it given > that it has an excellent documentation. As a newbie to MT, it walked > me through some basic concepts on its wiki which has provided me a > good base to build upon. >
Thanks for the compliment ! > > > == Why must my project be selected? == > > > Looking at the state of this language pair, it can be made stable > considering that all the initial work is done. Knowing both the > languages personally I would be able to comprehend the transfer rules > between them more easily. I think "must" here sounds a bit presumptuous. > == Goals == > > > *Convert the words from the apertium_hi_WX.dix from WX notation to > unicode encoding. Unicode encoding makes it consistent to use with the > apertium engine. Also, make the tags in the IIIT morphologic analyzer > to be along the same lines as the tags in apertium language pair. This should be done before the application is in. E.g. the deadline is the 8th April. We've had a lot of people try this, but even after weeks they have not succeeded. And I can't quite work out why. If you are familiar with Hindi and the tagset, it shouldn't take more than a day. > > *Make the bidix more complete by adding more words to the bilingual > dictionary. And, finally test it. > The bidix file is apertium-ur-hi.ur-hi.dix How are you going to do this ? > *There is presently no parts of speech tagger for this. This tagger > will be added to which will be the new file. > apertium-ur-hi.ur-hi.tsx This doesn't sound right. Or are you planning to make a single tagger for both Hindi and Urdu ? This is feasible, but then you will need to make sure that the tagsets are identical. > > *Adding tranfer rules before target language tagger training. The > rules need to be added to the file. > apertium-ur-hi.ur-hi.t1x. > > > *Train parts of speech tagger for the target language by using three > stage transfer command, > apertium-transfer > apertium-tagger-tl-trainer > > > *Finally, quality control tests would be run. The three test to be > carried out are: > -corpus test > -regression test > -testvoc Which corpus are you planning to use ? [...] > This is not complete by any means. Any feedback to improve it will be > much appreciated. Thanks for your suggestions. I suggest you read also my comments to n0nick (Sagie), and his project proposal too. Fran ------------------------------------------------------------------------------ Create and publish websites with WebMatrix Use the most popular FREE web apps or write code yourself; WebMatrix provides all the features you need to develop and publish your website. http://p.sf.net/sfu/ms-webmatrix-sf _______________________________________________ Apertium-stuff mailing list [email protected] https://lists.sourceforge.net/lists/listinfo/apertium-stuff
