On 4 April 2011 16:07, Antonio Toral <[email protected]> wrote:
> Hi Roman,
>
> thanks for your application and for your interest in the topic in
> particular and in Apertium in general.
>
> My main concern is the lack of detail regarding what I think is the main
> goal of this project: what linguistic information are you going to
> induce? (e.g. word correspondences, parts of speech, conjugations,
> singular/plural affixes, etc)
>
That's probably the essence of the difference between your take on the
project, and mine: I think you're thinking in terms of extracting
specific items, whereas I'm thinking of an extensible mechanism for
extraction. (That said, there should be more detail on the features to
be extracted).
I've explained the DBPedia project to Roman, but I think I should give
some more information, so here's a braindump:
Rather than manually parsing individual items, the dbpedia framework
(where possible) uses user definable templates, written in Wikimedia's
template language. (There is code for some things, where writing a
template would be overkill. Part of speech, etc., are represented by a
map, effectively making them localisable).
For example, collocations on de.wiktionary are represented like this:
{{Charakteristische Wortkombinationen}}
:[[räudig]]er Hund, [[streunen]]der Hund, [[tot]]er Hund, [[bissig]]er
Hund, [[Pawlowscher]] Hund, [[schwarz]]er Hund, [[andalusisch]]er
Hund, [[tollwütig]]er Hund, [[scharf]]er Hund, [[faul]]er Hund,
[[treu]]er Hund, [[toll]]er Hund, [[nass]]er Hund, [[dick]]er Hund,
[[herrenlos]]er Hund, Hund [[dressieren]], [[bunt]]er Hund,
For which the extraction template is:
{{Charakteristische Wortkombinationen}}
{{extractiontpl|list-start}}:{{extractiontpl|var|usage}}
{{extractiontpl|list-end}}
The set of example templates for de.wiktionary is far from exhaustive,
covering only the most basic cases. The project should include a good
set of initial templates for at least en.wiktionary and one other, to
ensure that the code is general (and to get some useful results!)
> These sentences, for example, seem too vague:
>
> "Improving DBPedia extraction framework to retrieve more information
> about words and add more languages form Wiktionaries"
>
> what kind of info? which languages?
>
> "get data required to build dictionaries ( words frequency,
> alphabets, ...)"
>
> the type of data should be more precise
>
Word frequency, on en.wiktionary at least, is just a template away... :)
This is vague, though, and it's not the sort of thing wiktionary is
typically good for. It should probably be removed.
>
> Hope you find these comments useful.
>
> cheers,
> Antonio
>
>> Hello,
>>
>> In attachment there is first version of my application. I hope I
>> included all things mentioned earlier about this project. What do You
>> think about it?
>>
>> Best Regards,
>> Roman Zegarski
>
>
> ------------------------------------------------------------------------------
> Create and publish websites with WebMatrix
> Use the most popular FREE web apps or write code yourself;
> WebMatrix provides all the features you need to develop and
> publish your website. http://p.sf.net/sfu/ms-webmatrix-sf
> _______________________________________________
> Apertium-stuff mailing list
> [email protected]
> https://lists.sourceforge.net/lists/listinfo/apertium-stuff
>
--
<Leftmost> jimregan, that's because deep inside you, you are evil.
<Leftmost> Also not-so-deep inside you.
------------------------------------------------------------------------------
Create and publish websites with WebMatrix
Use the most popular FREE web apps or write code yourself;
WebMatrix provides all the features you need to develop and
publish your website. http://p.sf.net/sfu/ms-webmatrix-sf
_______________________________________________
Apertium-stuff mailing list
[email protected]
https://lists.sourceforge.net/lists/listinfo/apertium-stuff