We've got a website that is causing our crawler to slow down (from 20mbits down to 3-5) - 400K pages that are basically not available, we're just getting 404's. I'd like to remove them from the DB to get our crawl speed back up again.

Here's what our developer told me - I'm stumped, that seems really odd. Is there a better way to remove a URL so that it doesn't get crawled?

Running nutch 0.71 on a dual xeon with 8 gigs of ram.
-------------------------
There are more than 400,000 urls in the webdb. It takes ~4 hours to remove a url from the webdb. That means that it'll take ~1,600,000 hours (~66,666 days, or ~2222 months, ~185 years) to remove 400,000 CAA urls from the webdb. Do you really want to remove them in this way?




-------------------------------------------------------
This SF.Net email is sponsored by xPML, a groundbreaking scripting language
that extends applications into web and mobile media. Attend the live webcast
and join the prime developer group breaking into this new coding territory!
http://sel.as-us.falkag.net/sel?cmd=lnk&kid=110944&bid=241720&dat=121642
_______________________________________________
Nutch-general mailing list
[email protected]
https://lists.sourceforge.net/lists/listinfo/nutch-general

Reply via email to