Hello all, just have a question on the fetcher in 0.8 dev.
A big problem I've been noticing with fetching is the number of max delays exceeded errors. I recently did a fetch of about 600,000 pages and 1 in 6 of those was never downloaded because of max delays exceeded. This seems like an excessive amount, so I decided to look into the code to see if I could find out how the url's are sorted for fetching, so that maybe I could improve this to spread out the attempts to fetch from each host. I saw that the urllist is partitioned by host at one point, but I gather that is just so that all the urls from one host are sent to the same server for fetching, to ensure politeness? I couldn't track down any other sorting. Am I missing something here? Is there any other way that the fetcher could be improved so that I don't get so many max delays exceeded errors?
Thanks in advance. -Matt Zytaruk
