On 5 April 2011 02:01, Roman Zegarski <[email protected]> wrote:
> 2011/4/4 Antonio Toral <[email protected]>:
>> Hi Roman,
> Thank You for response.
> I would like to concentrate my effort on improving system retrieving
> information from Wiktionary, capable of using different language
> templates. It should retrieve most of information contained inside
> Wiktionary pages, but I think most important parts will be (as You
> wrote before): translations, parts of speech,singular/plural affixes.
> Also, I would like to create templates for two Wiktionaries:
> en.wiktionary and pl.dictionary . I know Polish Wiktionary it is my
> native language, and I think I can make it in the fastest and most
> precise way.
>
Actually, I somewhat anticipated this :) Targeting pl.wiktionary will
require some specific code enhancements that would probably not be
instantly visible in terms of the other large Wiktionaries.
Inflection information is one of the more important things for us to
be able to extract, and pl.wiktionary is the best source of inflection
information for several languages, but the way inflection is
represented will require specific support in code:
== pies ({{język polski}}) ==
[snip]
{{odmiana}}
: (1.1-3) {{lp}} pies, psa, psu, psa, psem, psie, psie; {{lm}} ps|y,
~ów, ~om, ~y, ~ami, ~ach, ~y
The obvious extra work is the tilde replacement, but as inflection is
laid out positionally, which varies from language to language, a
language specific template should be invoked. I'm relatively sure that
support for that will need to be implemented.
> 2011/4/4 Jimmy O'Regan <[email protected]>:
>> One thing that's missing is that, because the DBPedia framework
>> extracts RDF, you will need to give some thought to the ontology that
>> will be used. For most of the parts of Wiktionary that are directly
>> useful to Apertium, you should use the Gold ontology
>> (http://linguistics-ontology.org/), though there may be other data in
>> Wiktionary for which new ontological properties (and thus a new
>> ontology) may need to be defined. You should consider that during the
>> community bonding period.
>>
> I included the Gold ontology, into community bonding period.
>
Try to make time to get an idea of SUMO too
(http://www.ontologyportal.org/). Mostly, this can be inferred from
wiktionary articles which include a link to a wikipedia article
(because DBPedia has been (mostly) mapped to SUMO).
>> 'Work on creating interface for DBPedia allowing user to retrieve data
>> in language neutral way (as far
>> as it is possible)'
>>
>> I think the existing DBPedia method of using a wiki to define
>> extraction templates is the most natural interface, given its nature,
>> so you should perhaps stick to that. Instead, you could use this time
>> to develop extraction 'templates' for the most frequently used
>> templates on en.wiktionary (and perhaps one other, of your choice),
>> which can then be used for testing.
>>
>> The same goes for the contents of weeks 5-9. Developing the extraction
>> will, I think, take longer than you have planned for, but the extra
>> interface is not necessary, while testing and extraction of
>> dictionaries should be concurrent to development.
>>
>
> I modified work plan after Your suggestions.
Po angielsku 'you', 'your', itd. piszą się małymi literami (tylko 'I'
pisze się wielkim) - wy polacy są greczniejsi niż my :)
> I would spend more time
> on developing DBPedia extraction framework. I plan to spend maximum a
> week on creating mapping for English Wiktionary. I think it rather
> take less time with use of tools created for DBPedia (like MappingTool
> and Extraction Tester). I would like to implement template for Polish
> language (should take me less time than English).
Ah. Well, the tools will need to be adapted, I think.
> This version includes improved work plan:
>
> http://dl.dropbox.com/u/11351380/Roman_Zegarski_proposal_ver2.pdf
>
> In addition - improved work plan only:
>
> Community Bonding Period:
> - get more familiar with Apertium community
> - retrieve more information about DBPedia
> - DBPedia mappings
> - DBPedia ontology
> - get to know GOLD ontology
Aside from the basics (which are already handled), there isn't a huge
need for this - you'll mostly need to have it at hand as a reference.
> - get to know Scala language
Scala looks a little odd, but for most of the task it should be relatively easy.
> - read documentation related to the project
>
> - Week 1 - 3:
> - Improving DBPedia extraction framework
> - creating code in Scala, which could handle more languages
> - create new ontology class
> - creating basic templates to English Wiktionary
> - Week 4:
> - expansion of templates for en.wiktionary
> - create templates for pl.wiktionary
>
> Deliverable #1 ‹- improved DBPedia extraction framework
>
> - Week 5 - 7 :
> - create module for dixtools retrieving data from DBPedia
> - Week 8
> - create dictionaries in Apertium format
>
> Deliverable #2 ‹- Aperitum-dixtools module creating dictionaries from
> data extracted from DBPedia
>
> - Week 9:
> - improving existing OmegaWiki data retriever implemented in
> apertium-dixtools
> - retrieve dictionaries data from OmegaWiki using dixtools
> - Week 10 - 11:
> - find if some data from OmegaWiki and DBPedia are complementary
> - merge complementary data retrieved from OmegaWiki and DBPedia
> - Week 12:
> - final amendments
> - creation of documentation for the project
>
> Project completed ‹- dictionaries created, new features in dixtools,
> improved DBPedia extraction framework
>
--
<Leftmost> jimregan, that's because deep inside you, you are evil.
<Leftmost> Also not-so-deep inside you.
------------------------------------------------------------------------------
Xperia(TM) PLAY
It's a major breakthrough. An authentic gaming
smartphone on the nation's most reliable network.
And it wants your games.
http://p.sf.net/sfu/verizon-sfdev
_______________________________________________
Apertium-stuff mailing list
[email protected]
https://lists.sourceforge.net/lists/listinfo/apertium-stuff