it seems you can imitate RDD.top()'s implementation. for each partition, you get the number of records, and the total sum of key, and in the final result handler, you add all the sum together, and add the number of records together, then you can get the mean, I mean, arithmetic mean.
On Tue, Apr 1, 2014 at 10:55 AM, Jaonary Rabarisoa <jaon...@gmail.com>wrote: > Hi all; > > Can someone give me some tips to compute mean of RDD by key , maybe with > combineByKey and StatCount. > > Cheers, > > Jaonary > -- Dachuan Huang Cellphone: 614-390-7234 2015 Neil Avenue Ohio State University Columbus, Ohio U.S.A. 43210