I have been giving some thought to Erik's proposal and while already fascinating, I would like to put it in different terms. Instead of asking "Could open source MT be such a strategic investment?", I would ask "is there a way to have Wikimedia's technology and people involved collaborate with MT systems?" The first can be seen as entering areas quite out of reach, the second would be more about paving the way for other actors that are already in the field. Our strength has always been based around human collaboration empowered by technology, and if MT is wished, then we should consider approaching it from our areas of expertise.
One of the biggest problems in MT is word disambiguation. Wikidata's item properties could be a way of setting the general context for article translation, and if that results not to be reliable enough, users should have the opportunity to specify on the source text the intended meaning of a certain word. While that could be less than ideal for literary works, where double meanings and other subtleties must be taken into account, it might be quite useful for Wikipedia, providing MT software a fertile soil where to grow. The standards for specifying word meanings for MT software are unknown to me, but it might be worth exploring. Another interesting hurdle for MT is dictionary building. OmegaWiki seems like a system that could be used for bridging the gap between pairs of languages, in such a way that if we know the exact use of the word in the source language, a user could seamlessly fill in the missing word and definition in the target language. That could be a unique way of collaboration between source-language speakers providing precision about the meaning being used, and target-language speakers filling the gaps. Dictionaries alone are not enough. Grammar rules would need to be wikified. All in all, OmegaWiki/Wiktionary could become the front-end and repository for external MT systems, either to be used in Wikipedia or with other pages. It wouldn't be needed to create a new MT system, because the rule-based MT programs that could make use of such infraestructure already exist. Some of them are open-source too. If you are interested, I could ask for opinions about the feasability in the Apertium lists. In my opinion, they also fit into the "smartest, well-intentioned group of people" category that Erik was asking about. Cheers, David --Micru On Tue, Apr 30, 2013 at 4:15 AM, Chris Tophe <[email protected]> wrote: > 2013/4/29 Mathieu Stumpf <[email protected]> > > > Le 2013-04-26 20:27, Milos Rancic a écrit : > > > > OmegaWiki is a masterpiece from the perspective of one [computational] > >> linguist. Erik made the structure so well, that it's the best starting > >> point to create a contemporary multilingual dictionary. I didn't see > >> anything better in concept. (And, yes, when I was thinking about > >> creating such software by my own, I was always at the dead end of > >> "but, OmegaWiki is already that".) > >> > > > > Where can I find documentation about this structure, please ? > > > > Here (planned structure): > http://meta.wikimedia.org/wiki/OmegaWiki_data_design > > and also there (current structure): > http://www.omegawiki.org/Help:OmegaWiki_database_layout > > And a gentle reminder that comments are requested ;-) > http://meta.wikimedia.org/wiki/Requests_for_comment/Adopt_OmegaWiki > _______________________________________________ > Wikimedia-l mailing list > [email protected] > Unsubscribe: https://lists.wikimedia.org/mailman/listinfo/wikimedia-l > -- Etiamsi omnes, ego non _______________________________________________ Wikimedia-l mailing list [email protected] Unsubscribe: https://lists.wikimedia.org/mailman/listinfo/wikimedia-l
