http://statwww.epfl.ch/davison/teaching/Microarrays/lab/clustering.html
http://research.nhgri.nih.gov/microarray/Gene_Expression_Supplement/ This provides some sample micro array data with a lab project for clustering the data. On Wed, May 21, 2008 at 10:44 AM, Ted Dunning <[EMAIL PROTECTED]> wrote: > > Sorry, didn't mean to sound like that. > > I am happy to build the data sets! > > I can also demo R for you this evening. > > > On Wed, May 21, 2008 at 10:36 AM, Jeff Eastman < > [EMAIL PROTECTED]> wrote: > >> Ok, ok, UNCLE! >> >> Things to do: >> - Install R >> - Learn R >> >> Now I have at least two reasons to do that<grin> >> Jeff >> >> >> >> Ted Dunning wrote: >> >>> It is also the work of a moment to build some synthetic data sets using >>> R. >>> Real data is cooler, though. >>> >>> On Wed, May 21, 2008 at 10:25 AM, Jeff Eastman < >>> [EMAIL PROTECTED]> wrote: >>> >>> >>> >>>> Thanks, Ted, and most are small enough to run on a single node. I'm >>>> investigating further... >>>> >>>> Jeff >>>> >>>> >>>> Ted Dunning wrote: >>>> >>>> >>>> >>>>> Do these 5 suffice: >>>>> >>>>> >>>>> >>>>> http://archive.ics.uci.edu/ml/datasets.html?format=&task=clu&att=&area=&numAtt=&numIns=&type=&sort=attUp&view=table >>>>> >>>>> The classification data sets are also reasonable to try with >>>>> clustering. >>>>> The Irises dataset and the Japanese vowels are both plausible for >>>>> clustering >>>>> (inter alia, of course). >>>>> >>>>> On Wed, May 21, 2008 at 8:10 AM, Jeff Eastman < >>>>> [EMAIL PROTECTED]> wrote: >>>>> >>>>> >>>>> >>>>> >>>>> >>>>>> Does anybody have some links to datasets we can use for clustering >>>>>> examples? I'm thinking we could publish an EC2 AMI that includes >>>>>> Hadoop >>>>>> and >>>>>> Mahout, along with a script to deploy it on a cluster, upload the >>>>>> examples >>>>>> and run clustering on it. Is that too ambitious? I'm kinda hoping that >>>>>> we >>>>>> can use 0.17 which advertises simpler EC2 deployment than 0.16. If >>>>>> that >>>>>> won't meet our schedule then maybe I should work through the 0.16 >>>>>> deployment. >>>>>> >>>>>> Jeff >>>>>> >>>>>> >>>>>> Grant Ingersoll wrote: >>>>>> >>>>>> >>>>>> >>>>>> >>>>>> >>>>>>> I was thinking we should get the Taste stuff in (seems to be pretty >>>>>>> close >>>>>>> to done) and I would like to get Mahout-9 (Naive Bayes) in. This >>>>>>> would >>>>>>> give >>>>>>> us a pretty nice release, I think. Namely, a couple of clustering >>>>>>> implementations, a classifier, and, of course, Taste. I think I can >>>>>>> finish >>>>>>> up my part in the next week or so. Then, we will need to start to >>>>>>> figure >>>>>>> out all the fun of releases (signatures, notices.txt, etc.) I'd also >>>>>>> like >>>>>>> to see us have an easy to use demo of the clustering stuff, but it is >>>>>>> all >>>>>>> right if we don't. >>>>>>> >>>>>>> -Grant >>>>>>> >>>>>>> On May 21, 2008, at 1:23 AM, Sean Owen wrote: >>>>>>> >>>>>>> Just curious, what are people thinking about the timeline for a >>>>>>> first, >>>>>>> >>>>>>> >>>>>>> >>>>>>> >>>>>>>> very early release, like an 0.1 release? any open tasks that I could >>>>>>>> pick up to help? >>>>>>>> >>>>>>>> Without rushing anything, I'm keen to retire my current project site >>>>>>>> and forward everybody that's interested to Mahout. As long as >>>>>>>> there's >>>>>>>> a .jar distro someone can pick up and use, that's cool. >>>>>>>> >>>>>>>> Sean >>>>>>>> >>>>>>>> >>>>>>>> >>>>>>>> >>>>>>>> >>>>>>> >>>>>>> >>>>>>> >>>>>>> >>>>>> >>>>> >>>>> >>>> >>>> >>> >>> >>> >>> >> >> > > > -- > ted > > -- ted
