Hi Germán, this indeed sounds very interesting. Unfortunately, as a beginner I'm not yet savvy enough to do such equilibristik stunts. But, I hope to be able to test later on. I'm interested in a short description of the algo for tuning, i.e. the steps. Yours, Per Tunedal
On Mon, Mar 11, 2013, at 16:21, Germán Sanchis Trilles wrote: > Hi Per and the rest, > > Back in EAMT 2011 we presented a paper about a pretty aggressive pruning > strategy for the phrase table, independently from the test set and > without any loss in translation quality. It basically boils down to > keeping only those phrases that would be used when translating the > training data. Of course, that means you will have to translate the > training data, but it can be done offline or on a computer which is not > the one used for decoding. Or else you can also split it in much smaller > chunks, which, by the way, is also a possibility for reducing the memory > requirements when decoding. Not really sure if this will work for you, > but I thought it might be worth the email just in case ;) > > I think nobody answered the question about what tuning actually does > (conceptually), so: the tuning stage is in charge of adjusting the > relative weight of each one of the translation models involved, and it > may well imply quite a difference in translation quality. Short: I would > not recommend skipping it. > > Cheers, > > Germán > > Barry Haddow <[email protected]> escribió: > > >Hi Per > > > >The tuning script will filter the phrase table (leaving only the > >entries > >required for the tuning set) and then binarise it (if you give it a > >binariser) before running the actual tuning. So no, the whole table > >doesn't need to be loaded into memory during tuning. > > > >You could prune before tuning, and I don't know how this will compare > >to > >pruning after tuning. I'm not sure if anyone has tried it. The > >signifcance filtering (in advanced features) works (afaik), although > >there are a few steps involved in building the code. The relent > >filtering has bit-rotted a bit, but you can run it with Moses v0.91. > > > >cheers - Barry > > > >On 11/03/13 14:37, Per Tunedal wrote: > >> Hi Barry, > >> Binarise, yes. But before that I plan to prune the translation table > >> (Advanced features). Any hints? > >> > >> Back to my original question, can pruning be done before tuning? Is > >it > >> possible to binarise too, before tuning? (The base-line page suggests > >> binarisiation after tuning.) I fear that the tuning might be an > >> overwhelming task for my poor computer. > >> > >> Yours, > >> Per Tunedal > >> > >> BTW What is actually done when tuning? Has all the tables to be > >loaded > >> into memory? > >> > >> On Mon, Mar 11, 2013, at 10:42, Barry Haddow wrote: > >>> Hi Per > >>> > >>> You need to binarise the models (phrase table, reordering table and > >>> language model) before running Moses > >>> http://www.statmt.org/moses/?n=Moses.AdvancedFeatures#ntoc3 > >>> If you don't binarise then Moses will load all the tables into > >memory, > >>> so the memory requirement will be at least as large as the on-disk > >size, > >>> in fact a lot more since it doesn't store them efficiently. > >>> > >>> The complexity of training is not easy to calculate since there are > >a > >>> number of steps, but since one step involves sorting the list of > >>> extracted phrases the complexity must be at least as bad as that. > >>> > >>> cheers - Barry > >>> > >>> On 11/03/13 08:24, Per Tunedal wrote: > >>>> Hi Barry, > >>>> it turns out that me too have succeeded to build my model in one > >day. I > >>>> forced a restart and checked the log and the working directory. All > >is > >>>> fine! The moses.ini file was created only 7 hours after submitting > >the > >>>> command to build the model. I don't understand why the computer > >didn't > >>>> respond, though. > >>>> > >>>> How does the time to build a model vary with the size of the > >corpus? > >>>> Linearly? Or quadratic? Or what? > >>>> > >>>> I've now tried to do a test translation, without any tuning. I soon > >ran > >>>> out of memory: even the virtual memory was exhausted after a while. > >Any > >>>> way to predict the memory needed? > >>>> > >>>> Yours, > >>>> Per Tunedal > >>>> > >>>> > >>>> On Sun, Mar 10, 2013, at 11:39, Barry Haddow wrote: > >>>>> Hi Per > >>>>> > >>>>> I would suggest starting from a smallish corpus, then building up > >to a > >>>>> larger one, to get experience with the process. Using the news > >commentary > >>>>> corpus described in the Moses baseline page, I was able to train > >and tune > >>>>> in an evening on my laptop. > >>>>> > >>>>> There have been papers on predicting quality given corpus size, > >but > >>>>> there's not an easy answer. Look for Marco Turchi at last year's > >EAMT, or > >>>>> (I think) one by Xerox Grenoble from last year. > >>>>> > >>>>> As regards europarl, yes there's noise, but the models are quite > >robust > >>>>> to it > >>>>> > >>>>> Cheers - Barry > >>>>> > >>>>> Per Tunedal <[email protected]> wrote: > >>>>> > >>>>>> Hi, > >>>>>> Is there any way to predict the time for training and/or tuning, > >given > >>>>>> the corpus size and the computer specifications? It would be nice > >to > >>>>>> know what would be a reasonable time for accomplishing the > >tasks. Now > >>>>>> my computer has been running for 3 days and nights and doesn't > >respond > >>>>>> any more: I cannot "wake it" to see what's going on. I don't know > >if > >>>>>> it's normal or if something has gone havoc. > >>>>>> > >>>>>> I agree with Ken Fasano, that it would be very useful to know how > >big a > >>>>>> corpus is needed to get meaningful results. I would like to be > >able to > >>>>>> judge the quality of the translation, to see if it would be > >useful to > >>>>>> continue with Moses in some more serious manner. > >>>>>> > >>>>>> I'm a bit puzzled by the parameter limiting sentence length to, > >say 80 > >>>>>> (characters?), giving that e.g. the Europarl corpus contains > >mainly VERY > >>>>>> long sentences. Skipping long sentences probably implicates that > >many > >>>>>> typical expressions are lost in the model. Wouldn't it be more > >sensible > >>>>>> to skip short sentences? Or to make a representative sample of > >the > >>>>>> corpus, by doing a random sample of a sufficient size or > >something? > >>>>>> Yours, > >>>>>> Per Tunedal > >>>>>> > >>>>>> PS I've noticed that the Europarl corpus contains some very bad, > >>>>>> completely incomprehensible, translations. That makes me question > >the > >>>>>> quality of that corpus. How are the translations actually done? > >By > >>>>>> humans relying heavily on machine translation? Sometimes letting > >some > >>>>>> strange MT-translation pass? > >>>>>> > >>>>>> On Sat, Mar 9, 2013, at 18:43, Ken Fasano wrote: > >>>>>>> I'd like to respond to this thread. I, too, have limited > >resources (at > >>>>>>> work, at least) - a 3 GB RAM 32-bit Windows i5 machine with > >Linux running > >>>>>>> on VMWare with 2.5 GB RAM. Training and tuning take many hours; > >the > >>>>>>> machine is running BitParl over the weekend and may be done with > >>>>>>> NewsCommentary de-en (DE) on Monday when I get back to work. I'm > >afraid > >>>>>>> after all that I won't be able to subsequently run Collins on > >the > >>>>>>> English, train, tune, and decode all that and, even if it takes > >forever, > >>>>>>> expect it to run with the limited memory resources available. > >>>>>>> What I think I need to do is trim the corpora according to some > >criteria > >>>>>>> that isn't too complicated. Is it enough for learning purposes > >(we are a > >>>>>>> long way from any sort of real comparison of results, let alone > >>>>>>> production - this will receive proper hardware) to take the > >first n > >>>>>>> sentences, or every nth sentence? The result is simply to get a > >feel for > >>>>>>> the various modes of tree-based SMT, run hierarchical phrase, > >>>>>>> string-to-tree, tree-to-string and tree-to-tree without worrying > >which > >>>>>>> one is the best - the idea is just to get some experience with > >it.What > >>>>>>> would be a good number of sentences to take so that it runs > >relatively > >>>>>>> quickly, without killing RAM, but produces results that aren't > >useless? > >>>>>>> Thanks - and I'd like to thank everyone on the group for their > >eager > >>>>>>> helpfulness, and for discussing things that I as a newbie find > >very > >>>>>>> useful! > >>>>>>> > >>>>>>> > >>>>>>> > >>>>>>> > >>>>>>>> Date: Sat, 9 Mar 2013 11:25:45 -0500 > >>>>>>>> From: [email protected] > >>>>>>>> To: [email protected] > >>>>>>>> Subject: Re: [Moses-support] Accelerate the tuning > >>>>>>>> > >>>>>>>> Hi, > >>>>>>>> > >>>>>>>> It won't fix everything, but there is a long-term TODO to > >rewrite > >>>>>>>> phrase table scoring to use binary files with vocabulary ids > >instead of > >>>>>>>> text files. > >>>>>>>> > >>>>>>>> Kenneth > >>>>>>>> > >>>>>>>> On 03/09/13 08:30, Per Tunedal wrote: > >>>>>>>>> Hi, > >>>>>>>>> the training seems to be an overwhelming task for my computer. > >If it > >>>>>>>>> ever succeeds, I will have to undertake the even more > >demanding task of > >>>>>>>>> tuning. Can anything be done to accelerate it? > >>>>>>>>> > >>>>>>>>> Specifically, I wonder if it's feasible to prune the > >translation table > >>>>>>>>> before doing the tuning. > >>>>>>>>> > >>>>>>>>> Yours, > >>>>>>>>> Per Tunedal > >>>>>>>>> > >>>>>>>>> PS I've abandoned the idea of building a Hierarchical phrase > >model, I'm > >>>>>>>>> now trying to make a phrase-based system. I suppose that would > >use less > >>>>>>>>> resources. > >>>>>>>>> > >>>>>>>>> _______________________________________________ > >>>>>>>>> Moses-support mailing list > >>>>>>>>> [email protected] > >>>>>>>>> http://mailman.mit.edu/mailman/listinfo/moses-support > >>>>>>>>> > >>>>>>>> _______________________________________________ > >>>>>>>> Moses-support mailing list > >>>>>>>> [email protected] > >>>>>>>> http://mailman.mit.edu/mailman/listinfo/moses-support > >>>>>>> > >>>>>>> _______________________________________________ > >>>>>>> Moses-support mailing list > >>>>>>> [email protected] > >>>>>>> http://mailman.mit.edu/mailman/listinfo/moses-support > >>>>>> _______________________________________________ > >>>>>> Moses-support mailing list > >>>>>> [email protected] > >>>>>> http://mailman.mit.edu/mailman/listinfo/moses-support > >>>>>> > >>>>> -- > >>>>> The University of Edinburgh is a charitable body, registered in > >>>>> Scotland, with registration number SC005336. > >>>>> > >>>> _______________________________________________ > >>>> Moses-support mailing list > >>>> [email protected] > >>>> http://mailman.mit.edu/mailman/listinfo/moses-support > >>>> > >>> > >>> -- > >>> The University of Edinburgh is a charitable body, registered in > >>> Scotland, with registration number SC005336. > >>> > >> _______________________________________________ > >> Moses-support mailing list > >> [email protected] > >> http://mailman.mit.edu/mailman/listinfo/moses-support > >> > > > > > >-- > >The University of Edinburgh is a charitable body, registered in > >Scotland, with registration number SC005336. > > > >_______________________________________________ > >Moses-support mailing list > >[email protected] > >http://mailman.mit.edu/mailman/listinfo/moses-support > > -- > Enviado desde mi teléfono Android con K-9 Mail. Disculpa mi brevedad > _______________________________________________ > Moses-support mailing list > [email protected] > http://mailman.mit.edu/mailman/listinfo/moses-support _______________________________________________ Moses-support mailing list [email protected] http://mailman.mit.edu/mailman/listinfo/moses-support
