Hi Per If you add the following to mert-moses.pl then it should work:
-filtercmd '<path-to-filter-script> -Binarizer "<path-to-processPhraseTable>"' Obviously you replace the angle-bracket expressions with the paths, and you might have to fiddle with the quotes, cheers - Barry On 11/03/13 16:23, Per Tunedal wrote: > Hi Barry, > unfortunately, I don't fully understand your advice, as I'm not familiar > with how the scripts work etc. > At this stage I just copy some commands, and make some minor changes I > find appropriate. I'm at the script-kiddie stage! > > 1. What command shall I submit if I will let the mert-moses.pl script do > the job, as I first intended? > 2. What commands shall I submit if I choose to do some of the job before > running the mert-moses.pl script? > > Yours, > Per Tunedal > > > On Mon, Mar 11, 2013, at 16:40, Barry Haddow wrote: >> Hi Per >> >> It's probably as easy to filter and binarise the input yourself before >> tuning using the filter-model-given-input.pl script in Moses, and using >> the -Binarizer argument to pass it the path to the processPhraseTable >> binary. >> >> By default, mert-moses.pl will filter (so the command below will filter) >> and this may reduce the size of the phrase table sufficiently for you. >> You can add binarisation to mert-moses.pl by using the --filtercmd to >> give it the filter script with its binariser argument, >> >> cheers - Barry >> >> On 11/03/13 15:27, Per Tunedal wrote: >>> Hi, >>> thanks for your kind advice. I plan to use the tuning command displayed >>> at the Baseline page: >>> nohup nice ~/mosesdecoder/scripts/training/mert-moses.pl >>> ~/corpus/news-test2008.true.fr ~/corpus/news-test2008.true.en \ >>> ~/mosesdecoder/bin/moses train/model/moses.ini --mertdir >>> ~/mosesdecoder/bin/ &> mert.out & >>> with my own "small" tuning-corpus (10% of the training-corpus size). >>> >>> Will this "give [the tuning script] a binariser" ? If not, what shall I >>> add to the command? >>> >>> Yours, >>> Per Tunedal >>> >>> On Mon, Mar 11, 2013, at 16:00, Barry Haddow wrote: >>>> Hi Per >>>> >>>> The tuning script will filter the phrase table (leaving only the entries >>>> required for the tuning set) and then binarise it (if you give it a >>>> binariser) before running the actual tuning. So no, the whole table >>>> doesn't need to be loaded into memory during tuning. >>>> >>>> You could prune before tuning, and I don't know how this will compare to >>>> pruning after tuning. I'm not sure if anyone has tried it. The >>>> signifcance filtering (in advanced features) works (afaik), although >>>> there are a few steps involved in building the code. The relent >>>> filtering has bit-rotted a bit, but you can run it with Moses v0.91. >>>> >>>> cheers - Barry >>>> >>>> On 11/03/13 14:37, Per Tunedal wrote: >>>>> Hi Barry, >>>>> Binarise, yes. But before that I plan to prune the translation table >>>>> (Advanced features). Any hints? >>>>> >>>>> Back to my original question, can pruning be done before tuning? Is it >>>>> possible to binarise too, before tuning? (The base-line page suggests >>>>> binarisiation after tuning.) I fear that the tuning might be an >>>>> overwhelming task for my poor computer. >>>>> >>>>> Yours, >>>>> Per Tunedal >>>>> >>>>> BTW What is actually done when tuning? Has all the tables to be loaded >>>>> into memory? >>>>> >>>>> On Mon, Mar 11, 2013, at 10:42, Barry Haddow wrote: >>>>>> Hi Per >>>>>> >>>>>> You need to binarise the models (phrase table, reordering table and >>>>>> language model) before running Moses >>>>>> http://www.statmt.org/moses/?n=Moses.AdvancedFeatures#ntoc3 >>>>>> If you don't binarise then Moses will load all the tables into memory, >>>>>> so the memory requirement will be at least as large as the on-disk size, >>>>>> in fact a lot more since it doesn't store them efficiently. >>>>>> >>>>>> The complexity of training is not easy to calculate since there are a >>>>>> number of steps, but since one step involves sorting the list of >>>>>> extracted phrases the complexity must be at least as bad as that. >>>>>> >>>>>> cheers - Barry >>>>>> >>>>>> On 11/03/13 08:24, Per Tunedal wrote: >>>>>>> Hi Barry, >>>>>>> it turns out that me too have succeeded to build my model in one day. I >>>>>>> forced a restart and checked the log and the working directory. All is >>>>>>> fine! The moses.ini file was created only 7 hours after submitting the >>>>>>> command to build the model. I don't understand why the computer didn't >>>>>>> respond, though. >>>>>>> >>>>>>> How does the time to build a model vary with the size of the corpus? >>>>>>> Linearly? Or quadratic? Or what? >>>>>>> >>>>>>> I've now tried to do a test translation, without any tuning. I soon ran >>>>>>> out of memory: even the virtual memory was exhausted after a while. Any >>>>>>> way to predict the memory needed? >>>>>>> >>>>>>> Yours, >>>>>>> Per Tunedal >>>>>>> >>>>>>> >>>>>>> On Sun, Mar 10, 2013, at 11:39, Barry Haddow wrote: >>>>>>>> Hi Per >>>>>>>> >>>>>>>> I would suggest starting from a smallish corpus, then building up to a >>>>>>>> larger one, to get experience with the process. Using the news >>>>>>>> commentary >>>>>>>> corpus described in the Moses baseline page, I was able to train and >>>>>>>> tune >>>>>>>> in an evening on my laptop. >>>>>>>> >>>>>>>> There have been papers on predicting quality given corpus size, but >>>>>>>> there's not an easy answer. Look for Marco Turchi at last year's EAMT, >>>>>>>> or >>>>>>>> (I think) one by Xerox Grenoble from last year. >>>>>>>> >>>>>>>> As regards europarl, yes there's noise, but the models are quite robust >>>>>>>> to it >>>>>>>> >>>>>>>> Cheers - Barry >>>>>>>> >>>>>>>> Per Tunedal <[email protected]> wrote: >>>>>>>> >>>>>>>>> Hi, >>>>>>>>> Is there any way to predict the time for training and/or tuning, given >>>>>>>>> the corpus size and the computer specifications? It would be nice to >>>>>>>>> know what would be a reasonable time for accomplishing the tasks. Now >>>>>>>>> my computer has been running for 3 days and nights and doesn't respond >>>>>>>>> any more: I cannot "wake it" to see what's going on. I don't know if >>>>>>>>> it's normal or if something has gone havoc. >>>>>>>>> >>>>>>>>> I agree with Ken Fasano, that it would be very useful to know how big >>>>>>>>> a >>>>>>>>> corpus is needed to get meaningful results. I would like to be able to >>>>>>>>> judge the quality of the translation, to see if it would be useful to >>>>>>>>> continue with Moses in some more serious manner. >>>>>>>>> >>>>>>>>> I'm a bit puzzled by the parameter limiting sentence length to, say 80 >>>>>>>>> (characters?), giving that e.g. the Europarl corpus contains mainly >>>>>>>>> VERY >>>>>>>>> long sentences. Skipping long sentences probably implicates that many >>>>>>>>> typical expressions are lost in the model. Wouldn't it be more >>>>>>>>> sensible >>>>>>>>> to skip short sentences? Or to make a representative sample of the >>>>>>>>> corpus, by doing a random sample of a sufficient size or something? >>>>>>>>> Yours, >>>>>>>>> Per Tunedal >>>>>>>>> >>>>>>>>> PS I've noticed that the Europarl corpus contains some very bad, >>>>>>>>> completely incomprehensible, translations. That makes me question the >>>>>>>>> quality of that corpus. How are the translations actually done? By >>>>>>>>> humans relying heavily on machine translation? Sometimes letting some >>>>>>>>> strange MT-translation pass? >>>>>>>>> >>>>>>>>> On Sat, Mar 9, 2013, at 18:43, Ken Fasano wrote: >>>>>>>>>> I'd like to respond to this thread. I, too, have limited resources >>>>>>>>>> (at >>>>>>>>>> work, at least) - a 3 GB RAM 32-bit Windows i5 machine with Linux >>>>>>>>>> running >>>>>>>>>> on VMWare with 2.5 GB RAM. Training and tuning take many hours; the >>>>>>>>>> machine is running BitParl over the weekend and may be done with >>>>>>>>>> NewsCommentary de-en (DE) on Monday when I get back to work. I'm >>>>>>>>>> afraid >>>>>>>>>> after all that I won't be able to subsequently run Collins on the >>>>>>>>>> English, train, tune, and decode all that and, even if it takes >>>>>>>>>> forever, >>>>>>>>>> expect it to run with the limited memory resources available. >>>>>>>>>> What I think I need to do is trim the corpora according to some >>>>>>>>>> criteria >>>>>>>>>> that isn't too complicated. Is it enough for learning purposes (we >>>>>>>>>> are a >>>>>>>>>> long way from any sort of real comparison of results, let alone >>>>>>>>>> production - this will receive proper hardware) to take the first n >>>>>>>>>> sentences, or every nth sentence? The result is simply to get a feel >>>>>>>>>> for >>>>>>>>>> the various modes of tree-based SMT, run hierarchical phrase, >>>>>>>>>> string-to-tree, tree-to-string and tree-to-tree without worrying >>>>>>>>>> which >>>>>>>>>> one is the best - the idea is just to get some experience with >>>>>>>>>> it.What >>>>>>>>>> would be a good number of sentences to take so that it runs >>>>>>>>>> relatively >>>>>>>>>> quickly, without killing RAM, but produces results that aren't >>>>>>>>>> useless? >>>>>>>>>> Thanks - and I'd like to thank everyone on the group for their eager >>>>>>>>>> helpfulness, and for discussing things that I as a newbie find very >>>>>>>>>> useful! >>>>>>>>>> >>>>>>>>>> >>>>>>>>>> >>>>>>>>>> >>>>>>>>>>> Date: Sat, 9 Mar 2013 11:25:45 -0500 >>>>>>>>>>> From: [email protected] >>>>>>>>>>> To: [email protected] >>>>>>>>>>> Subject: Re: [Moses-support] Accelerate the tuning >>>>>>>>>>> >>>>>>>>>>> Hi, >>>>>>>>>>> >>>>>>>>>>> It won't fix everything, but there is a long-term TODO to >>>>>>>>>>> rewrite >>>>>>>>>>> phrase table scoring to use binary files with vocabulary ids >>>>>>>>>>> instead of >>>>>>>>>>> text files. >>>>>>>>>>> >>>>>>>>>>> Kenneth >>>>>>>>>>> >>>>>>>>>>> On 03/09/13 08:30, Per Tunedal wrote: >>>>>>>>>>>> Hi, >>>>>>>>>>>> the training seems to be an overwhelming task for my computer. If >>>>>>>>>>>> it >>>>>>>>>>>> ever succeeds, I will have to undertake the even more demanding >>>>>>>>>>>> task of >>>>>>>>>>>> tuning. Can anything be done to accelerate it? >>>>>>>>>>>> >>>>>>>>>>>> Specifically, I wonder if it's feasible to prune the translation >>>>>>>>>>>> table >>>>>>>>>>>> before doing the tuning. >>>>>>>>>>>> >>>>>>>>>>>> Yours, >>>>>>>>>>>> Per Tunedal >>>>>>>>>>>> >>>>>>>>>>>> PS I've abandoned the idea of building a Hierarchical phrase >>>>>>>>>>>> model, I'm >>>>>>>>>>>> now trying to make a phrase-based system. I suppose that would use >>>>>>>>>>>> less >>>>>>>>>>>> resources. >>>>>>>>>>>> >>>>>>>>>>>> _______________________________________________ >>>>>>>>>>>> Moses-support mailing list >>>>>>>>>>>> [email protected] >>>>>>>>>>>> http://mailman.mit.edu/mailman/listinfo/moses-support >>>>>>>>>>>> >>>>>>>>>>> _______________________________________________ >>>>>>>>>>> Moses-support mailing list >>>>>>>>>>> [email protected] >>>>>>>>>>> http://mailman.mit.edu/mailman/listinfo/moses-support >>>>>>>>>> >>>>>>>>>> _______________________________________________ >>>>>>>>>> Moses-support mailing list >>>>>>>>>> [email protected] >>>>>>>>>> http://mailman.mit.edu/mailman/listinfo/moses-support >>>>>>>>> _______________________________________________ >>>>>>>>> Moses-support mailing list >>>>>>>>> [email protected] >>>>>>>>> http://mailman.mit.edu/mailman/listinfo/moses-support >>>>>>>>> >>>>>>>> -- >>>>>>>> The University of Edinburgh is a charitable body, registered in >>>>>>>> Scotland, with registration number SC005336. >>>>>>>> >>>>>>> _______________________________________________ >>>>>>> Moses-support mailing list >>>>>>> [email protected] >>>>>>> http://mailman.mit.edu/mailman/listinfo/moses-support >>>>>>> >>>>>> -- >>>>>> The University of Edinburgh is a charitable body, registered in >>>>>> Scotland, with registration number SC005336. >>>>>> >>>>> _______________________________________________ >>>>> Moses-support mailing list >>>>> [email protected] >>>>> http://mailman.mit.edu/mailman/listinfo/moses-support >>>>> >>>> -- >>>> The University of Edinburgh is a charitable body, registered in >>>> Scotland, with registration number SC005336. >>>> >>> _______________________________________________ >>> Moses-support mailing list >>> [email protected] >>> http://mailman.mit.edu/mailman/listinfo/moses-support >>> >> >> -- >> The University of Edinburgh is a charitable body, registered in >> Scotland, with registration number SC005336. >> -- The University of Edinburgh is a charitable body, registered in Scotland, with registration number SC005336. _______________________________________________ Moses-support mailing list [email protected] http://mailman.mit.edu/mailman/listinfo/moses-support
