El dg 03 de 04 de 2011 a les 01:15 +0300, en/na Sagie Maoz va escriure:
> Dear All,
> 
> 
> I'm happy to attach my application draft for a GSoC project, a
> Maltese-Hebrew language pair.
> I will be happy to hear your thoughts and comments, particularly on
> when should I submit it at the Google Melange site.
> 
> 
> Please find my application available
> at http://sagie.maoz.info/apertium-application.pdf.

Hi, nice application... I think you are underselling yourself though.
I've chatted with you a bit on IRC, and I think you have already
gathered quite a lot of references, possible sources of information for
both languages. So it might be a good idea to have a section in your
proposal "Resources" (and include here grammars, dictionaries etc.)

If you are able to get access to the morphological analyser for Hebrew
before the 8th it should be included in your application. If you are not
able to, then it should be assumed that it is _not_ available, and the
time plan should be adjusted accordingly. Note that "promised access to"
is not quite the same as "will be released under the GPL" -- which would
be our requirement for using it.

I'm not sure if a morphological analyser for Maltese exists, but if it
does, then the same goes for that as for the Hebrew one.

I'm not sure about the "Renaissance of Hebrew and Maltese" reference.
Perhaps there is a more scholarly reference available ? 

* bellow -> below

Where are you expecting to find a parallel corpus for Hebrew and
Maltese ? One person you might try contacting on this question is Adam
Ussishkin, who I met at SEPLN and is a nice guy.

I'm missing a bit of a description of the languages, and some potential
transfer problems. I'm not expecting you to have done all the work, but
some examples would be nice :) Perhaps you could demonstrate by
translating the examples from
http://wiki.apertium.eu/index.php/Appendix_A:_Frequency as a computer
would ? 

==Week plan==

* Are you going to be using separate frequency lists for Hebrew and
  Maltese ? If so, then might you not end up with having to take the
  intersection and losing words -- thus not completely effectively using
  your time ? 

* Another idea might be to just use a frequency list from the Maltese
  side ?

* I think you are optimistic about getting useful handtagged corpora for
  both Maltese and Hebrew -- in order to include them in your proposal,
  you should have them _before_ the 8th April. Otherwise assume that 
  they don't exist... remember Zigglebottom.

* Even if you do get hand tagged corpora, you would need to convert them
  to the tagset of your analyser(s), and also you might need to fix
  tokenisation issues.

* I think that a week for working on the bilingual dictionary is too
  little. When working from a language that you don't know into one 
  that you do know, you need to be particularly careful. For 
  Macedonian--English, it took me approx 8 days. And I have quite a lot
  of experience. http://permalink.gmane.org/gmane.comp.nlp.apertium/285

The thanks in Maltese/Hebrew is a really nice touch :)

Fran


------------------------------------------------------------------------------
Create and publish websites with WebMatrix
Use the most popular FREE web apps or write code yourself; 
WebMatrix provides all the features you need to develop and 
publish your website. http://p.sf.net/sfu/ms-webmatrix-sf
_______________________________________________
Apertium-stuff mailing list
[email protected]
https://lists.sourceforge.net/lists/listinfo/apertium-stuff

Reply via email to