2011/4/5 Jimmy O'Regan <[email protected]>: > On 5 April 2011 02:01, Roman Zegarski <[email protected]> wrote: >> 2011/4/4 Antonio Toral <[email protected]>: >>> Hi Roman, >> Thank You for response. >> I would like to concentrate my effort on improving system retrieving >> information from Wiktionary, capable of using different language >> templates. It should retrieve most of information contained inside >> Wiktionary pages, but I think most important parts will be (as You >> wrote before): translations, parts of speech,singular/plural affixes. >> Also, I would like to create templates for two Wiktionaries: >> en.wiktionary and pl.dictionary . I know Polish Wiktionary it is my >> native language, and I think I can make it in the fastest and most >> precise way. >> > > Actually, I somewhat anticipated this :) Targeting pl.wiktionary will > require some specific code enhancements that would probably not be > instantly visible in terms of the other large Wiktionaries. > > Inflection information is one of the more important things for us to > be able to extract, and pl.wiktionary is the best source of inflection > information for several languages, but the way inflection is > represented will require specific support in code: > > == pies ({{język polski}}) == > [snip] > {{odmiana}} > : (1.1-3) {{lp}} pies, psa, psu, psa, psem, psie, psie; {{lm}} ps|y, > ~ów, ~om, ~y, ~ami, ~ach, ~y > > The obvious extra work is the tilde replacement, but as inflection is > laid out positionally, which varies from language to language, a > language specific template should be invoked. I'm relatively sure that > support for that will need to be implemented. >
So maybe I should try Spanish? My knowledge about it is very general, but making templates not necessary require a good knowledge of language. There aren't so many heading types, and I can always use Apertium to translate from Spanish to English :) . Spanish looks like good language for project purposes. It is,as far as I know, used by many Apertium developers, also it's forth language on Wiktionary (counting number of entries) and contains twice as many "form-of definitions" as English language (from: http://en.wiktionary.org/wiki/Wiktionary:Statistics). I think that I could spend some time of community bounding period on getting know more about templates used in different languages. That would help me to improve DBPedia framework in more deliberate way later.I made quick look onto German inflection, and I found that even the entities uses basically the same pattern, but there are still inconsistencies, for example: {{Charakteristische Wortkombinationen}} :[[räudig]]er Hund, [[streunen]]der Hund, [[tot]]er Hund, [[bissig]]er Hund, [[Pawlowscher]] Hund, [[schwarz]]er Hund, [[andalusisch]]er Hund, [[tollwütig]]er Hund, [[scharf]]er Hund, [[faul]]er Hund, [[treu]]er Hund, [[toll]]er Hund, [[nass]]er Hund, [[dick]]er Hund, [[herrenlos]]er Hund, Hund [[dressieren]], [[bunt]]er Hund, {{Charakteristische Wortkombinationen}} :[1] ein Fenster öffnen; durch ein Fenster schauen; ein offenes Fenster :[3] das Fenster schließt sich {{Charakteristische Wortkombinationen}} : [[sitzen]], [[Stuhl]], [[gedeckt]], [[wischen]], [[Bett]], [[rund|runder]], [[auf]], [[Thema]], [[unter|unterm]], [[endgültig]], [[stehen|stand]], [[Schrank]], [[Teller]], [[klein]], [[Karo]], [[lang]], [[herum]], [[Sessel]], [[Restaurant]], [[Zimmer]], [[Sofa]], [[Plan|Pläne]], [[Kellner]], [[groß]] As I looked briefly on couple of words in Spanish Wiktionary. I didn't found such inconsistencies, but I think they can exist. So I think it would be a good idea, to get know what can I expect from Wiktionary entities. :) >> 2011/4/4 Jimmy O'Regan <[email protected]>: >>> One thing that's missing is that, because the DBPedia framework >>> extracts RDF, you will need to give some thought to the ontology that >>> will be used. For most of the parts of Wiktionary that are directly >>> useful to Apertium, you should use the Gold ontology >>> (http://linguistics-ontology.org/), though there may be other data in >>> Wiktionary for which new ontological properties (and thus a new >>> ontology) may need to be defined. You should consider that during the >>> community bonding period. >>> >> I included the Gold ontology, into community bonding period. >> > > Try to make time to get an idea of SUMO too > (http://www.ontologyportal.org/). Mostly, this can be inferred from > wiktionary articles which include a link to a wikipedia article > (because DBPedia has been (mostly) mapped to SUMO). > I think I have an overall view of how should new ontology classes looks like. I would like to spend some time in community bounding period to make sure that I would create new ontology classes appropriate to the situation. 2011/4/5 Jimmy O'Regan <[email protected]>: > On 5 April 2011 02:59, Jimmy O'Regan <[email protected]> wrote: >> Po angielsku 'you', 'your', itd. piszą się małymi literami (tylko 'I' >> pisze się wielkim) - wy polacy są greczniejsi niż my :) > > Raczej, Wy Polacy są grzeczniejsi, a chyba wiesz o co mi chodziło :D > Tak, wiem o co chodziło. :) Jesteśmy bardziej przywiązani do form grzecznościowych i mi osobiście ciężko jest się przełamać. Cały czas pisząc wiadomości po angielsku, mam wrażenie, że popełniam nietakt pisząc 'you' albo 'your' małymi literami. Pozostaje mi oswoić się z tym, że jest to poprawna forma. :) Best Reagrds, Roman Zegarski ------------------------------------------------------------------------------ Xperia(TM) PLAY It's a major breakthrough. An authentic gaming smartphone on the nation's most reliable network. And it wants your games. http://p.sf.net/sfu/verizon-sfdev _______________________________________________ Apertium-stuff mailing list [email protected] https://lists.sourceforge.net/lists/listinfo/apertium-stuff
