2011/4/4 Antonio Toral <[email protected]>:
> Hi Roman,
>
> thanks for your application and for your interest in the topic in
> particular and in Apertium in general.
>
> My main concern is the lack of detail regarding what I think is the main
> goal of this project: what linguistic information are you going to
> induce? (e.g. word correspondences, parts of speech, conjugations,
> singular/plural affixes, etc)
>
> These sentences, for example, seem too vague:
>
> "Improving DBPedia extraction framework to retrieve more information
> about words and add more languages form Wiktionaries"
>
> what kind of info? which languages?
>
> "get data required to build dictionaries ( words frequency,
> alphabets, ...)"
>
> the type of data should be more precise
>
>
> Hope you find these comments useful.
>
> cheers,
> Antonio

Thank You for response.
I would like to concentrate my effort on improving system retrieving
information from Wiktionary, capable of using different language
templates. It should retrieve most of information contained inside
Wiktionary pages, but I think most important parts will be (as You
wrote before): translations, parts of speech,singular/plural affixes.
Also, I would like to create templates for two Wiktionaries:
en.wiktionary  and  pl.dictionary . I know Polish Wiktionary it is my
native language, and I think I can make it in the fastest and most
precise way.

2011/4/4 Jimmy O'Regan <[email protected]>:
> One thing that's missing is that, because the DBPedia framework
> extracts RDF, you will need to give some thought to the ontology that
> will be used. For most of the parts of Wiktionary that are directly
> useful to Apertium, you should use the Gold ontology
> (http://linguistics-ontology.org/), though there may be other data in
> Wiktionary for which new ontological properties (and thus a new
> ontology) may need to be defined. You should consider that during the
> community bonding period.
>
I included the Gold ontology, into community bonding period.

> 'Work on creating interface for DBPedia allowing user to retrieve data
> in language neutral way (as far
> as it is possible)'
>
> I think the existing DBPedia method of using a wiki to define
> extraction templates is the most natural interface, given its nature,
> so you should perhaps stick to that. Instead, you could use this time
> to develop extraction 'templates' for the most frequently used
> templates on en.wiktionary (and perhaps one other, of your choice),
> which can then be used for testing.
>
> The same goes for the contents of weeks 5-9. Developing the extraction
> will, I think, take longer than you have planned for, but the extra
> interface is not necessary, while testing and extraction of
> dictionaries should be concurrent to development.
>

I modified work plan after Your suggestions. I would spend more time
on developing DBPedia extraction framework. I plan to spend maximum a
week on creating mapping for English Wiktionary. I think it rather
take less time with use of tools created for DBPedia (like MappingTool
and Extraction Tester). I would like to implement template for Polish
language (should take me less time than English).



This version includes improved work plan:

http://dl.dropbox.com/u/11351380/Roman_Zegarski_proposal_ver2.pdf

In addition - improved work plan only:

Community Bonding Period:
        - get more familiar with Apertium community
        - retrieve more information about DBPedia
                - DBPedia mappings
                - DBPedia ontology
        - get to know GOLD ontology
        - get to know Scala language
        - read documentation related to the project

- Week 1 - 3:
        - Improving DBPedia extraction framework
                - creating code in Scala, which could handle more languages
                - create new ontology class
                - creating basic templates to English Wiktionary
- Week 4:
        - expansion of templates for en.wiktionary
        - create templates for pl.wiktionary
                
Deliverable #1  ‹-    improved DBPedia extraction framework

- Week 5 - 7 :
        - create module for dixtools retrieving data from DBPedia
- Week 8
        - create dictionaries in Apertium format
        
Deliverable #2 ‹- Aperitum-dixtools module creating dictionaries from
data extracted from DBPedia

- Week 9:
        - improving existing OmegaWiki data retriever implemented in 
apertium-dixtools
        - retrieve dictionaries data from OmegaWiki using dixtools
- Week 10 - 11:
        - find if some data from OmegaWiki and DBPedia are complementary
        - merge complementary data retrieved from OmegaWiki and DBPedia
- Week 12:
        - final amendments
        - creation of documentation for the project
        
Project completed ‹- dictionaries created, new features in dixtools,
improved DBPedia extraction framework

Best regards,
Roman Zegarski

------------------------------------------------------------------------------
Xperia(TM) PLAY
It's a major breakthrough. An authentic gaming
smartphone on the nation's most reliable network.
And it wants your games.
http://p.sf.net/sfu/verizon-sfdev
_______________________________________________
Apertium-stuff mailing list
[email protected]
https://lists.sourceforge.net/lists/listinfo/apertium-stuff

Reply via email to