Can anyone provide some advice on how to update an existing clustering with new data points. Our data set is approximately 1mm newspaper headlines over the course of a month. I'm able to get a high quality clustering using the existing mahout tasks (I'm just using canopy in this instance) but I'd like to update the clusters on an hourly basis. Given the hardware that is available to me, I won't be able to run the clustering to completion over the entire data set every hour. Are there any methods for completing such a task?
Since I'm not a mahout or linear algebra expert at this point, ideally the solution would involve a combination of the existing mahout tasks. That being said, I'd be appreciative of any and all advice. Thanks, Asif -- Asif Rahman Lead Engineer - NewsCred [email protected] http://platform.newscred.com
