Hi Barry, Binarise, yes. But before that I plan to prune the translation table (Advanced features). Any hints?
Back to my original question, can pruning be done before tuning? Is it possible to binarise too, before tuning? (The base-line page suggests binarisiation after tuning.) I fear that the tuning might be an overwhelming task for my poor computer. Yours, Per Tunedal BTW What is actually done when tuning? Has all the tables to be loaded into memory? On Mon, Mar 11, 2013, at 10:42, Barry Haddow wrote: > Hi Per > > You need to binarise the models (phrase table, reordering table and > language model) before running Moses > http://www.statmt.org/moses/?n=Moses.AdvancedFeatures#ntoc3 > If you don't binarise then Moses will load all the tables into memory, > so the memory requirement will be at least as large as the on-disk size, > in fact a lot more since it doesn't store them efficiently. > > The complexity of training is not easy to calculate since there are a > number of steps, but since one step involves sorting the list of > extracted phrases the complexity must be at least as bad as that. > > cheers - Barry > > On 11/03/13 08:24, Per Tunedal wrote: > > Hi Barry, > > it turns out that me too have succeeded to build my model in one day. I > > forced a restart and checked the log and the working directory. All is > > fine! The moses.ini file was created only 7 hours after submitting the > > command to build the model. I don't understand why the computer didn't > > respond, though. > > > > How does the time to build a model vary with the size of the corpus? > > Linearly? Or quadratic? Or what? > > > > I've now tried to do a test translation, without any tuning. I soon ran > > out of memory: even the virtual memory was exhausted after a while. Any > > way to predict the memory needed? > > > > Yours, > > Per Tunedal > > > > > > On Sun, Mar 10, 2013, at 11:39, Barry Haddow wrote: > >> Hi Per > >> > >> I would suggest starting from a smallish corpus, then building up to a > >> larger one, to get experience with the process. Using the news commentary > >> corpus described in the Moses baseline page, I was able to train and tune > >> in an evening on my laptop. > >> > >> There have been papers on predicting quality given corpus size, but > >> there's not an easy answer. Look for Marco Turchi at last year's EAMT, or > >> (I think) one by Xerox Grenoble from last year. > >> > >> As regards europarl, yes there's noise, but the models are quite robust > >> to it > >> > >> Cheers - Barry > >> > >> Per Tunedal <[email protected]> wrote: > >> > >>> Hi, > >>> Is there any way to predict the time for training and/or tuning, given > >>> the corpus size and the computer specifications? It would be nice to > >>> know what would be a reasonable time for accomplishing the tasks. Now > >>> my computer has been running for 3 days and nights and doesn't respond > >>> any more: I cannot "wake it" to see what's going on. I don't know if > >>> it's normal or if something has gone havoc. > >>> > >>> I agree with Ken Fasano, that it would be very useful to know how big a > >>> corpus is needed to get meaningful results. I would like to be able to > >>> judge the quality of the translation, to see if it would be useful to > >>> continue with Moses in some more serious manner. > >>> > >>> I'm a bit puzzled by the parameter limiting sentence length to, say 80 > >>> (characters?), giving that e.g. the Europarl corpus contains mainly VERY > >>> long sentences. Skipping long sentences probably implicates that many > >>> typical expressions are lost in the model. Wouldn't it be more sensible > >>> to skip short sentences? Or to make a representative sample of the > >>> corpus, by doing a random sample of a sufficient size or something? > >>> Yours, > >>> Per Tunedal > >>> > >>> PS I've noticed that the Europarl corpus contains some very bad, > >>> completely incomprehensible, translations. That makes me question the > >>> quality of that corpus. How are the translations actually done? By > >>> humans relying heavily on machine translation? Sometimes letting some > >>> strange MT-translation pass? > >>> > >>> On Sat, Mar 9, 2013, at 18:43, Ken Fasano wrote: > >>>> I'd like to respond to this thread. I, too, have limited resources (at > >>>> work, at least) - a 3 GB RAM 32-bit Windows i5 machine with Linux running > >>>> on VMWare with 2.5 GB RAM. Training and tuning take many hours; the > >>>> machine is running BitParl over the weekend and may be done with > >>>> NewsCommentary de-en (DE) on Monday when I get back to work. I'm afraid > >>>> after all that I won't be able to subsequently run Collins on the > >>>> English, train, tune, and decode all that and, even if it takes forever, > >>>> expect it to run with the limited memory resources available. > >>>> What I think I need to do is trim the corpora according to some criteria > >>>> that isn't too complicated. Is it enough for learning purposes (we are a > >>>> long way from any sort of real comparison of results, let alone > >>>> production - this will receive proper hardware) to take the first n > >>>> sentences, or every nth sentence? The result is simply to get a feel for > >>>> the various modes of tree-based SMT, run hierarchical phrase, > >>>> string-to-tree, tree-to-string and tree-to-tree without worrying which > >>>> one is the best - the idea is just to get some experience with it.What > >>>> would be a good number of sentences to take so that it runs relatively > >>>> quickly, without killing RAM, but produces results that aren't useless? > >>>> Thanks - and I'd like to thank everyone on the group for their eager > >>>> helpfulness, and for discussing things that I as a newbie find very > >>>> useful! > >>>> > >>>> > >>>> > >>>> > >>>>> Date: Sat, 9 Mar 2013 11:25:45 -0500 > >>>>> From: [email protected] > >>>>> To: [email protected] > >>>>> Subject: Re: [Moses-support] Accelerate the tuning > >>>>> > >>>>> Hi, > >>>>> > >>>>> It won't fix everything, but there is a long-term TODO to > >>>>> rewrite > >>>>> phrase table scoring to use binary files with vocabulary ids instead of > >>>>> text files. > >>>>> > >>>>> Kenneth > >>>>> > >>>>> On 03/09/13 08:30, Per Tunedal wrote: > >>>>>> Hi, > >>>>>> the training seems to be an overwhelming task for my computer. If it > >>>>>> ever succeeds, I will have to undertake the even more demanding task of > >>>>>> tuning. Can anything be done to accelerate it? > >>>>>> > >>>>>> Specifically, I wonder if it's feasible to prune the translation table > >>>>>> before doing the tuning. > >>>>>> > >>>>>> Yours, > >>>>>> Per Tunedal > >>>>>> > >>>>>> PS I've abandoned the idea of building a Hierarchical phrase model, I'm > >>>>>> now trying to make a phrase-based system. I suppose that would use less > >>>>>> resources. > >>>>>> > >>>>>> _______________________________________________ > >>>>>> Moses-support mailing list > >>>>>> [email protected] > >>>>>> http://mailman.mit.edu/mailman/listinfo/moses-support > >>>>>> > >>>>> _______________________________________________ > >>>>> Moses-support mailing list > >>>>> [email protected] > >>>>> http://mailman.mit.edu/mailman/listinfo/moses-support > >>>> > >>>> _______________________________________________ > >>>> Moses-support mailing list > >>>> [email protected] > >>>> http://mailman.mit.edu/mailman/listinfo/moses-support > >>> _______________________________________________ > >>> Moses-support mailing list > >>> [email protected] > >>> http://mailman.mit.edu/mailman/listinfo/moses-support > >>> > >> -- > >> The University of Edinburgh is a charitable body, registered in > >> Scotland, with registration number SC005336. > >> > > _______________________________________________ > > Moses-support mailing list > > [email protected] > > http://mailman.mit.edu/mailman/listinfo/moses-support > > > > > -- > The University of Edinburgh is a charitable body, registered in > Scotland, with registration number SC005336. > _______________________________________________ Moses-support mailing list [email protected] http://mailman.mit.edu/mailman/listinfo/moses-support
