We've got a website that is causing our crawler to slow down (from
20mbits down to 3-5) - 400K pages that are basically not available,
we're just getting 404's. I'd like to remove them from the DB to get
our crawl speed back up again.
Here's what our developer told me - I'm stumped, that seems really odd.
Is there a better way to remove a URL so that it doesn't get crawled?
Running nutch 0.71 on a dual xeon with 8 gigs of ram.
-------------------------
There are more than 400,000 urls in the webdb. It takes ~4 hours
to remove a url from the webdb. That means that it'll take ~1,600,000
hours (~66,666 days, or ~2222 months, ~185 years) to remove 400,000 CAA
urls from the webdb. Do you really want to remove them in this way?
-------------------------------------------------------
This SF.Net email is sponsored by xPML, a groundbreaking scripting language
that extends applications into web and mobile media. Attend the live webcast
and join the prime developer group breaking into this new coding territory!
http://sel.as-us.falkag.net/sel?cmd=lnk&kid=110944&bid=241720&dat=121642
_______________________________________________
Nutch-general mailing list
[email protected]
https://lists.sourceforge.net/lists/listinfo/nutch-general