Hi Per and the rest,
Back in EAMT 2011 we presented a paper about a pretty aggressive pruning
strategy for the phrase table, independently from the test set and without any
loss in translation quality. It basically boils down to keeping only those
phrases that would be used when translating the training data. Of course, that
means you will have to translate the training data, but it can be done offline
or on a computer which is not the one used for decoding. Or else you can also
split it in much smaller chunks, which, by the way, is also a possibility for
reducing the memory requirements when decoding. Not really sure if this will
work for you, but I thought it might be worth the email just in case ;)
I think nobody answered the question about what tuning actually does
(conceptually), so: the tuning stage is in charge of adjusting the relative
weight of each one of the translation models involved, and it may well imply
quite a difference in translation quality. Short: I would not recommend
skipping it.
Cheers,
Germán
Barry Haddow <[email protected]> escribió:
>Hi Per
>
>The tuning script will filter the phrase table (leaving only the
>entries
>required for the tuning set) and then binarise it (if you give it a
>binariser) before running the actual tuning. So no, the whole table
>doesn't need to be loaded into memory during tuning.
>
>You could prune before tuning, and I don't know how this will compare
>to
>pruning after tuning. I'm not sure if anyone has tried it. The
>signifcance filtering (in advanced features) works (afaik), although
>there are a few steps involved in building the code. The relent
>filtering has bit-rotted a bit, but you can run it with Moses v0.91.
>
>cheers - Barry
>
>On 11/03/13 14:37, Per Tunedal wrote:
>> Hi Barry,
>> Binarise, yes. But before that I plan to prune the translation table
>> (Advanced features). Any hints?
>>
>> Back to my original question, can pruning be done before tuning? Is
>it
>> possible to binarise too, before tuning? (The base-line page suggests
>> binarisiation after tuning.) I fear that the tuning might be an
>> overwhelming task for my poor computer.
>>
>> Yours,
>> Per Tunedal
>>
>> BTW What is actually done when tuning? Has all the tables to be
>loaded
>> into memory?
>>
>> On Mon, Mar 11, 2013, at 10:42, Barry Haddow wrote:
>>> Hi Per
>>>
>>> You need to binarise the models (phrase table, reordering table and
>>> language model) before running Moses
>>> http://www.statmt.org/moses/?n=Moses.AdvancedFeatures#ntoc3
>>> If you don't binarise then Moses will load all the tables into
>memory,
>>> so the memory requirement will be at least as large as the on-disk
>size,
>>> in fact a lot more since it doesn't store them efficiently.
>>>
>>> The complexity of training is not easy to calculate since there are
>a
>>> number of steps, but since one step involves sorting the list of
>>> extracted phrases the complexity must be at least as bad as that.
>>>
>>> cheers - Barry
>>>
>>> On 11/03/13 08:24, Per Tunedal wrote:
>>>> Hi Barry,
>>>> it turns out that me too have succeeded to build my model in one
>day. I
>>>> forced a restart and checked the log and the working directory. All
>is
>>>> fine! The moses.ini file was created only 7 hours after submitting
>the
>>>> command to build the model. I don't understand why the computer
>didn't
>>>> respond, though.
>>>>
>>>> How does the time to build a model vary with the size of the
>corpus?
>>>> Linearly? Or quadratic? Or what?
>>>>
>>>> I've now tried to do a test translation, without any tuning. I soon
>ran
>>>> out of memory: even the virtual memory was exhausted after a while.
>Any
>>>> way to predict the memory needed?
>>>>
>>>> Yours,
>>>> Per Tunedal
>>>>
>>>>
>>>> On Sun, Mar 10, 2013, at 11:39, Barry Haddow wrote:
>>>>> Hi Per
>>>>>
>>>>> I would suggest starting from a smallish corpus, then building up
>to a
>>>>> larger one, to get experience with the process. Using the news
>commentary
>>>>> corpus described in the Moses baseline page, I was able to train
>and tune
>>>>> in an evening on my laptop.
>>>>>
>>>>> There have been papers on predicting quality given corpus size,
>but
>>>>> there's not an easy answer. Look for Marco Turchi at last year's
>EAMT, or
>>>>> (I think) one by Xerox Grenoble from last year.
>>>>>
>>>>> As regards europarl, yes there's noise, but the models are quite
>robust
>>>>> to it
>>>>>
>>>>> Cheers - Barry
>>>>>
>>>>> Per Tunedal <[email protected]> wrote:
>>>>>
>>>>>> Hi,
>>>>>> Is there any way to predict the time for training and/or tuning,
>given
>>>>>> the corpus size and the computer specifications? It would be nice
>to
>>>>>> know what would be a reasonable time for accomplishing the
>tasks. Now
>>>>>> my computer has been running for 3 days and nights and doesn't
>respond
>>>>>> any more: I cannot "wake it" to see what's going on. I don't know
>if
>>>>>> it's normal or if something has gone havoc.
>>>>>>
>>>>>> I agree with Ken Fasano, that it would be very useful to know how
>big a
>>>>>> corpus is needed to get meaningful results. I would like to be
>able to
>>>>>> judge the quality of the translation, to see if it would be
>useful to
>>>>>> continue with Moses in some more serious manner.
>>>>>>
>>>>>> I'm a bit puzzled by the parameter limiting sentence length to,
>say 80
>>>>>> (characters?), giving that e.g. the Europarl corpus contains
>mainly VERY
>>>>>> long sentences. Skipping long sentences probably implicates that
>many
>>>>>> typical expressions are lost in the model. Wouldn't it be more
>sensible
>>>>>> to skip short sentences? Or to make a representative sample of
>the
>>>>>> corpus, by doing a random sample of a sufficient size or
>something?
>>>>>> Yours,
>>>>>> Per Tunedal
>>>>>>
>>>>>> PS I've noticed that the Europarl corpus contains some very bad,
>>>>>> completely incomprehensible, translations. That makes me question
>the
>>>>>> quality of that corpus. How are the translations actually done?
>By
>>>>>> humans relying heavily on machine translation? Sometimes letting
>some
>>>>>> strange MT-translation pass?
>>>>>>
>>>>>> On Sat, Mar 9, 2013, at 18:43, Ken Fasano wrote:
>>>>>>> I'd like to respond to this thread. I, too, have limited
>resources (at
>>>>>>> work, at least) - a 3 GB RAM 32-bit Windows i5 machine with
>Linux running
>>>>>>> on VMWare with 2.5 GB RAM. Training and tuning take many hours;
>the
>>>>>>> machine is running BitParl over the weekend and may be done with
>>>>>>> NewsCommentary de-en (DE) on Monday when I get back to work. I'm
>afraid
>>>>>>> after all that I won't be able to subsequently run Collins on
>the
>>>>>>> English, train, tune, and decode all that and, even if it takes
>forever,
>>>>>>> expect it to run with the limited memory resources available.
>>>>>>> What I think I need to do is trim the corpora according to some
>criteria
>>>>>>> that isn't too complicated. Is it enough for learning purposes
>(we are a
>>>>>>> long way from any sort of real comparison of results, let alone
>>>>>>> production - this will receive proper hardware) to take the
>first n
>>>>>>> sentences, or every nth sentence? The result is simply to get a
>feel for
>>>>>>> the various modes of tree-based SMT, run hierarchical phrase,
>>>>>>> string-to-tree, tree-to-string and tree-to-tree without worrying
>which
>>>>>>> one is the best - the idea is just to get some experience with
>it.What
>>>>>>> would be a good number of sentences to take so that it runs
>relatively
>>>>>>> quickly, without killing RAM, but produces results that aren't
>useless?
>>>>>>> Thanks - and I'd like to thank everyone on the group for their
>eager
>>>>>>> helpfulness, and for discussing things that I as a newbie find
>very
>>>>>>> useful!
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>> Date: Sat, 9 Mar 2013 11:25:45 -0500
>>>>>>>> From: [email protected]
>>>>>>>> To: [email protected]
>>>>>>>> Subject: Re: [Moses-support] Accelerate the tuning
>>>>>>>>
>>>>>>>> Hi,
>>>>>>>>
>>>>>>>> It won't fix everything, but there is a long-term TODO to
>rewrite
>>>>>>>> phrase table scoring to use binary files with vocabulary ids
>instead of
>>>>>>>> text files.
>>>>>>>>
>>>>>>>> Kenneth
>>>>>>>>
>>>>>>>> On 03/09/13 08:30, Per Tunedal wrote:
>>>>>>>>> Hi,
>>>>>>>>> the training seems to be an overwhelming task for my computer.
>If it
>>>>>>>>> ever succeeds, I will have to undertake the even more
>demanding task of
>>>>>>>>> tuning. Can anything be done to accelerate it?
>>>>>>>>>
>>>>>>>>> Specifically, I wonder if it's feasible to prune the
>translation table
>>>>>>>>> before doing the tuning.
>>>>>>>>>
>>>>>>>>> Yours,
>>>>>>>>> Per Tunedal
>>>>>>>>>
>>>>>>>>> PS I've abandoned the idea of building a Hierarchical phrase
>model, I'm
>>>>>>>>> now trying to make a phrase-based system. I suppose that would
>use less
>>>>>>>>> resources.
>>>>>>>>>
>>>>>>>>> _______________________________________________
>>>>>>>>> Moses-support mailing list
>>>>>>>>> [email protected]
>>>>>>>>> http://mailman.mit.edu/mailman/listinfo/moses-support
>>>>>>>>>
>>>>>>>> _______________________________________________
>>>>>>>> Moses-support mailing list
>>>>>>>> [email protected]
>>>>>>>> http://mailman.mit.edu/mailman/listinfo/moses-support
>>>>>>>
>>>>>>> _______________________________________________
>>>>>>> Moses-support mailing list
>>>>>>> [email protected]
>>>>>>> http://mailman.mit.edu/mailman/listinfo/moses-support
>>>>>> _______________________________________________
>>>>>> Moses-support mailing list
>>>>>> [email protected]
>>>>>> http://mailman.mit.edu/mailman/listinfo/moses-support
>>>>>>
>>>>> --
>>>>> The University of Edinburgh is a charitable body, registered in
>>>>> Scotland, with registration number SC005336.
>>>>>
>>>> _______________________________________________
>>>> Moses-support mailing list
>>>> [email protected]
>>>> http://mailman.mit.edu/mailman/listinfo/moses-support
>>>>
>>>
>>> --
>>> The University of Edinburgh is a charitable body, registered in
>>> Scotland, with registration number SC005336.
>>>
>> _______________________________________________
>> Moses-support mailing list
>> [email protected]
>> http://mailman.mit.edu/mailman/listinfo/moses-support
>>
>
>
>--
>The University of Edinburgh is a charitable body, registered in
>Scotland, with registration number SC005336.
>
>_______________________________________________
>Moses-support mailing list
>[email protected]
>http://mailman.mit.edu/mailman/listinfo/moses-support
--
Enviado desde mi teléfono Android con K-9 Mail. Disculpa mi brevedad
_______________________________________________
Moses-support mailing list
[email protected]
http://mailman.mit.edu/mailman/listinfo/moses-support