Task 1 is completed and R is running :). Maybe this afternoon you can give me the cooks tour?
Jeff

Ted Dunning wrote:
Sorry, didn't mean to sound like that.

I am happy to build the data sets!

I can also demo R for you this evening.

On Wed, May 21, 2008 at 10:36 AM, Jeff Eastman <
[EMAIL PROTECTED]> wrote:

Ok, ok, UNCLE!

Things to do:
- Install R
- Learn R

Now I have at least two reasons to do that<grin>
Jeff



Ted Dunning wrote:

It is also the work of a moment to build some synthetic data sets using R.
Real data is cooler, though.

On Wed, May 21, 2008 at 10:25 AM, Jeff Eastman <
[EMAIL PROTECTED]> wrote:



Thanks, Ted, and most are small enough to run on a single node. I'm
investigating further...

Jeff


Ted Dunning wrote:



Do these 5 suffice:



http://archive.ics.uci.edu/ml/datasets.html?format=&task=clu&att=&area=&numAtt=&numIns=&type=&sort=attUp&view=table

The classification data sets are also reasonable to try with clustering.
The Irises dataset and the Japanese vowels are both plausible for
clustering
(inter alia, of course).

On Wed, May 21, 2008 at 8:10 AM, Jeff Eastman <
[EMAIL PROTECTED]> wrote:





Does anybody have some links to datasets we can use for clustering
examples? I'm thinking we could publish an EC2 AMI that includes Hadoop
and
Mahout, along with a script to deploy it on a cluster, upload the
examples
and run clustering on it. Is that too ambitious? I'm kinda hoping that
we
can use 0.17 which advertises simpler EC2 deployment than 0.16. If that
won't meet our schedule then maybe I should work through the 0.16
deployment.

Jeff


Grant Ingersoll wrote:





I was thinking we should get the Taste stuff in (seems to be pretty
close
to done) and I would like to get Mahout-9 (Naive Bayes) in.  This
would
give
us a pretty nice release, I think.  Namely, a couple of clustering
implementations, a classifier, and, of course, Taste.  I think I can
finish
up my part in the next week or so.  Then, we will need to start to
figure
out all the fun of releases (signatures, notices.txt, etc.)  I'd also
like
to see us have an easy to use demo of the clustering stuff, but it is
all
right if we don't.

-Grant

On May 21, 2008, at 1:23 AM, Sean Owen wrote:

 Just curious, what are people thinking about the timeline for a
first,




very early release, like an 0.1 release? any open tasks that I could
pick up to help?

Without rushing anything, I'm keen to retire my current project site
and forward everybody that's interested to Mahout. As long as there's
a .jar distro someone can pick up and use, that's cool.

Sean












begin:vcard
fn:Jeff Eastman
n:Eastman;Jeff
org:Windward Solutions Inc.
adr:;;;Los Altos;CA;;USA
email;internet:[EMAIL PROTECTED]
title:President
tel;pager:http://[EMAIL PROTECTED]
tel;cell:+1.415.298.0023
x-mozilla-html:TRUE
url:http://www.windwardsolutions.com
version:2.1
end:vcard

Attachment: PGP.sig
Description: PGP signature

Reply via email to