Sorry, didn't mean to sound like that. I am happy to build the data sets!
I can also demo R for you this evening. On Wed, May 21, 2008 at 10:36 AM, Jeff Eastman < [EMAIL PROTECTED]> wrote: > Ok, ok, UNCLE! > > Things to do: > - Install R > - Learn R > > Now I have at least two reasons to do that<grin> > Jeff > > > > Ted Dunning wrote: > >> It is also the work of a moment to build some synthetic data sets using R. >> Real data is cooler, though. >> >> On Wed, May 21, 2008 at 10:25 AM, Jeff Eastman < >> [EMAIL PROTECTED]> wrote: >> >> >> >>> Thanks, Ted, and most are small enough to run on a single node. I'm >>> investigating further... >>> >>> Jeff >>> >>> >>> Ted Dunning wrote: >>> >>> >>> >>>> Do these 5 suffice: >>>> >>>> >>>> >>>> http://archive.ics.uci.edu/ml/datasets.html?format=&task=clu&att=&area=&numAtt=&numIns=&type=&sort=attUp&view=table >>>> >>>> The classification data sets are also reasonable to try with clustering. >>>> The Irises dataset and the Japanese vowels are both plausible for >>>> clustering >>>> (inter alia, of course). >>>> >>>> On Wed, May 21, 2008 at 8:10 AM, Jeff Eastman < >>>> [EMAIL PROTECTED]> wrote: >>>> >>>> >>>> >>>> >>>> >>>>> Does anybody have some links to datasets we can use for clustering >>>>> examples? I'm thinking we could publish an EC2 AMI that includes Hadoop >>>>> and >>>>> Mahout, along with a script to deploy it on a cluster, upload the >>>>> examples >>>>> and run clustering on it. Is that too ambitious? I'm kinda hoping that >>>>> we >>>>> can use 0.17 which advertises simpler EC2 deployment than 0.16. If that >>>>> won't meet our schedule then maybe I should work through the 0.16 >>>>> deployment. >>>>> >>>>> Jeff >>>>> >>>>> >>>>> Grant Ingersoll wrote: >>>>> >>>>> >>>>> >>>>> >>>>> >>>>>> I was thinking we should get the Taste stuff in (seems to be pretty >>>>>> close >>>>>> to done) and I would like to get Mahout-9 (Naive Bayes) in. This >>>>>> would >>>>>> give >>>>>> us a pretty nice release, I think. Namely, a couple of clustering >>>>>> implementations, a classifier, and, of course, Taste. I think I can >>>>>> finish >>>>>> up my part in the next week or so. Then, we will need to start to >>>>>> figure >>>>>> out all the fun of releases (signatures, notices.txt, etc.) I'd also >>>>>> like >>>>>> to see us have an easy to use demo of the clustering stuff, but it is >>>>>> all >>>>>> right if we don't. >>>>>> >>>>>> -Grant >>>>>> >>>>>> On May 21, 2008, at 1:23 AM, Sean Owen wrote: >>>>>> >>>>>> Just curious, what are people thinking about the timeline for a >>>>>> first, >>>>>> >>>>>> >>>>>> >>>>>> >>>>>>> very early release, like an 0.1 release? any open tasks that I could >>>>>>> pick up to help? >>>>>>> >>>>>>> Without rushing anything, I'm keen to retire my current project site >>>>>>> and forward everybody that's interested to Mahout. As long as there's >>>>>>> a .jar distro someone can pick up and use, that's cool. >>>>>>> >>>>>>> Sean >>>>>>> >>>>>>> >>>>>>> >>>>>>> >>>>>>> >>>>>> >>>>>> >>>>>> >>>>>> >>>>> >>>> >>>> >>> >>> >> >> >> >> > > -- ted
