As soon as I can get this puppy into production, I'll add our name to that
list.

That idea has a lot of potential.  We get anywhere between 1,000 and 2,000
articles per hour.  Since we're looking at news, articles published within,
say, 24 hours of each other have a lot of affinity to form a cluster.
Articles that are days apart have almost no affinity to each other.  Every 6
hours or so, I could run the clustering on just new articles since the last
clustering.  In between, as articles come in, use the headline to query
against the cluster set that I have indexed in solr, and add the article to
the most similar cluster.

I think I just repeated, in different words, exactly what you said below.
Thanks in any case for helping me reason this out.

On Fri, Jul 16, 2010 at 12:06 PM, Grant Ingersoll <[email protected]>wrote:

>
> On Jul 16, 2010, at 11:27 AM, Asif Rahman wrote:
>
> > Can anyone provide some advice on how to update an existing clustering
> with
> > new data points.  Our data set is approximately 1mm newspaper headlines
> over
> > the course of a month.  I'm able to get a high quality clustering using
> the
> > existing mahout tasks (I'm just using canopy in this instance)
>
> [OT] Care to share more (since you've already said you are using it)?
> https://cwiki.apache.org/confluence/display/MAHOUT/Powered+By+Mahout
>
> > but I'd like
> > to update the clusters on an hourly basis.  Given the hardware that is
> > available to me, I won't be able to run the clustering to completion over
> > the entire data set every hour.  Are there any methods for completing
> such a
> > task?
>
> How many new docs are you talking in that hour?  I'm sure others can add
> here, but AIUI, people in this situation often calculate the clusters and
> then for new docs in some time period, they just see which cluster that new
> document is closest to and add it there, then, offline or "later" they
> recluster the whole set.  So, for instance, perhaps nightly or every 6 hours
> or whatever you can afford, you do the whole job, but then in between you
> just do the lighter weight calculation.  I imagine there are probably ways
> of calculating when a new cluster is needed or when quality has dropped too
> much, so perhaps that could be used to trigger a new full run, too.
>
> >
> > Since I'm not a mahout or linear algebra expert at this point, ideally
> the
> > solution would involve a combination of the existing mahout tasks.  That
> > being said, I'd be appreciative of any and all advice.
> >
> > Thanks,
> >
> > Asif
> >
> >
> > --
> > Asif Rahman
> > Lead Engineer - NewsCred
> > [email protected]
> > http://platform.newscred.com
>
>


-- 
Asif Rahman
Lead Engineer - NewsCred
[email protected]
http://platform.newscred.com

Reply via email to