What about specifying percentage to keep or split instead?

On Wed, Jul 22, 2009 at 7:27 PM, Grant Ingersoll <[email protected]>wrote:

> What's the best way, in a M/R program, to keep track of the number of items
> one has produced/consumed?  For instance, for the WikipediaDatasetCreator,
> how would you go about putting a maximum number of examples to include in an
> input file (or to split some number between train and test)?  In the old
> message passing days, one would have to communicate how many elements one
> has to the other nodes or something like that, right?  In theory, it seems
> like each Mapper could just do take max/(num mappers) but I'm not 100%
> certain that works as I am not sure that matches Hadoop's semantics, as I
> thought the value was just a suggestion.
>
> -Grant
>



-- 
Ted Dunning, CTO
DeepDyve

Reply via email to