2011/4/4 Antonio Toral <[email protected]>: > Hi Roman, > > thanks for your application and for your interest in the topic in > particular and in Apertium in general. > > My main concern is the lack of detail regarding what I think is the main > goal of this project: what linguistic information are you going to > induce? (e.g. word correspondences, parts of speech, conjugations, > singular/plural affixes, etc) > > These sentences, for example, seem too vague: > > "Improving DBPedia extraction framework to retrieve more information > about words and add more languages form Wiktionaries" > > what kind of info? which languages? > > "get data required to build dictionaries ( words frequency, > alphabets, ...)" > > the type of data should be more precise > > > Hope you find these comments useful. > > cheers, > Antonio
Thank You for response. I would like to concentrate my effort on improving system retrieving information from Wiktionary, capable of using different language templates. It should retrieve most of information contained inside Wiktionary pages, but I think most important parts will be (as You wrote before): translations, parts of speech,singular/plural affixes. Also, I would like to create templates for two Wiktionaries: en.wiktionary and pl.dictionary . I know Polish Wiktionary it is my native language, and I think I can make it in the fastest and most precise way. 2011/4/4 Jimmy O'Regan <[email protected]>: > One thing that's missing is that, because the DBPedia framework > extracts RDF, you will need to give some thought to the ontology that > will be used. For most of the parts of Wiktionary that are directly > useful to Apertium, you should use the Gold ontology > (http://linguistics-ontology.org/), though there may be other data in > Wiktionary for which new ontological properties (and thus a new > ontology) may need to be defined. You should consider that during the > community bonding period. > I included the Gold ontology, into community bonding period. > 'Work on creating interface for DBPedia allowing user to retrieve data > in language neutral way (as far > as it is possible)' > > I think the existing DBPedia method of using a wiki to define > extraction templates is the most natural interface, given its nature, > so you should perhaps stick to that. Instead, you could use this time > to develop extraction 'templates' for the most frequently used > templates on en.wiktionary (and perhaps one other, of your choice), > which can then be used for testing. > > The same goes for the contents of weeks 5-9. Developing the extraction > will, I think, take longer than you have planned for, but the extra > interface is not necessary, while testing and extraction of > dictionaries should be concurrent to development. > I modified work plan after Your suggestions. I would spend more time on developing DBPedia extraction framework. I plan to spend maximum a week on creating mapping for English Wiktionary. I think it rather take less time with use of tools created for DBPedia (like MappingTool and Extraction Tester). I would like to implement template for Polish language (should take me less time than English). This version includes improved work plan: http://dl.dropbox.com/u/11351380/Roman_Zegarski_proposal_ver2.pdf In addition - improved work plan only: Community Bonding Period: - get more familiar with Apertium community - retrieve more information about DBPedia - DBPedia mappings - DBPedia ontology - get to know GOLD ontology - get to know Scala language - read documentation related to the project - Week 1 - 3: - Improving DBPedia extraction framework - creating code in Scala, which could handle more languages - create new ontology class - creating basic templates to English Wiktionary - Week 4: - expansion of templates for en.wiktionary - create templates for pl.wiktionary Deliverable #1 ‹- improved DBPedia extraction framework - Week 5 - 7 : - create module for dixtools retrieving data from DBPedia - Week 8 - create dictionaries in Apertium format Deliverable #2 ‹- Aperitum-dixtools module creating dictionaries from data extracted from DBPedia - Week 9: - improving existing OmegaWiki data retriever implemented in apertium-dixtools - retrieve dictionaries data from OmegaWiki using dixtools - Week 10 - 11: - find if some data from OmegaWiki and DBPedia are complementary - merge complementary data retrieved from OmegaWiki and DBPedia - Week 12: - final amendments - creation of documentation for the project Project completed ‹- dictionaries created, new features in dixtools, improved DBPedia extraction framework Best regards, Roman Zegarski ------------------------------------------------------------------------------ Xperia(TM) PLAY It's a major breakthrough. An authentic gaming smartphone on the nation's most reliable network. And it wants your games. http://p.sf.net/sfu/verizon-sfdev _______________________________________________ Apertium-stuff mailing list [email protected] https://lists.sourceforge.net/lists/listinfo/apertium-stuff
