Do these 5 suffice:

http://archive.ics.uci.edu/ml/datasets.html?format=&task=clu&att=&area=&numAtt=&numIns=&type=&sort=attUp&view=table

The classification data sets are also reasonable to try with clustering.
The Irises dataset and the Japanese vowels are both plausible for clustering
(inter alia, of course).

On Wed, May 21, 2008 at 8:10 AM, Jeff Eastman <
[EMAIL PROTECTED]> wrote:

> Does anybody have some links to datasets we can use for clustering
> examples? I'm thinking we could publish an EC2 AMI that includes Hadoop and
> Mahout, along with a script to deploy it on a cluster, upload the examples
> and run clustering on it. Is that too ambitious? I'm kinda hoping that we
> can use 0.17 which advertises simpler EC2 deployment than 0.16. If that
> won't meet our schedule then maybe I should work through the 0.16
> deployment.
>
> Jeff
>
>
> Grant Ingersoll wrote:
>
>> I was thinking we should get the Taste stuff in (seems to be pretty close
>> to done) and I would like to get Mahout-9 (Naive Bayes) in.  This would give
>> us a pretty nice release, I think.  Namely, a couple of clustering
>> implementations, a classifier, and, of course, Taste.  I think I can finish
>> up my part in the next week or so.  Then, we will need to start to figure
>> out all the fun of releases (signatures, notices.txt, etc.)  I'd also like
>> to see us have an easy to use demo of the clustering stuff, but it is all
>> right if we don't.
>>
>> -Grant
>>
>> On May 21, 2008, at 1:23 AM, Sean Owen wrote:
>>
>>  Just curious, what are people thinking about the timeline for a first,
>>> very early release, like an 0.1 release? any open tasks that I could
>>> pick up to help?
>>>
>>> Without rushing anything, I'm keen to retire my current project site
>>> and forward everybody that's interested to Mahout. As long as there's
>>> a .jar distro someone can pick up and use, that's cool.
>>>
>>> Sean
>>>
>>
>>
>>
>>
>


-- 
ted

Reply via email to