Don't order, that's expensive.
Just group by rand(), specify parallelism on the group by, and store the
result of "foreach grouped generate FLATTEN(name_of_original_relation);"

On Fri, Apr 1, 2011 at 11:22 AM, Xiaomeng Wan <[email protected]> wrote:

> Hi Jameson,
>
> Do you mind to add something like this:
>
> c = order b by $0 parallel n;
> store c into '20110331-ab';
>
> you can order on anything. it will add a reduce and give you less files.
>
> Regards,
> Shawn
> On Fri, Apr 1, 2011 at 1:57 AM, Jameson Li <[email protected]> wrote:
> > Hi,
> >
> > When I run the below pig codes:
> > a = load '/logs/2011-03-31';
> > b = filter a by $1=='a' and $2=='b';
> > store b into '20110331-ab';
> >
> > It runs a M/R that have thousands maps, and then create a output store
> > directory that have the same number so many files.
> >
> > I have a doubt that how I could store less files when I use pig to store
> > files in the HDFS.
> >
> >
> > Thanks,
> > Jameson Li.
> >
>

Reply via email to