Sean Dean wrote:
As it stands now with whats in trunk under 0.9-dev, one of the biggest problems is the 
version of Hadoop we have included. It fails on anything above 200k URLs, and should be 
considered a "blocker" issue.
Its my understanding that Andrzej has a newer Hadoop JAR with some custom patches applied, but hasn't had the time yet to commit them back to Nutch. When he does get the chance, some testing will need to be initiated and I can be a small help there. Not to make it sound like the end of the world, but since almost everything in Nutch revolves around Hadoop we should get this issue corrected before we make other "big" plans for fixes and changes.

To be precise, the version is 0.11.2 release and it's been committed just now (rev. 515791). Your help in testing would be most welcome ...

--
Best regards,
Andrzej Bialecki     <><
___. ___ ___ ___ _ _   __________________________________
[__ || __|__/|__||\/|  Information Retrieval, Semantic Web
___|||__||  \|  ||  |  Embedded Unix, System Integration
http://www.sigram.com  Contact: info at sigram dot com


Reply via email to