jnioche commented on issue #2185: URL: https://github.com/apache/stormcrawler/issues/2185#issuecomment-5812625892
thanks @GGraziadei > With URLFrontier the order of the crawl is decided entirely by the frontier: it rotates between queues (one per host) and serves each queue FIFO this needs checking but if I remember correctly, this is might not be the case. The priority of a queue in URLFrontier is determined by the earliest `nextFetchDate` of its URLs. The depth of a URL is usually reflected by its `nextFetchDate`. On a more general note, the advantage of using URLFrontier is that we delegate a lot of the logic away from SC - what you are suggesting it would reintroduce some of it. If we want to change the way URLs are served by URLFrontier then it should happen there. Having said that, if we just want to observe things, this is totally OK but I would argue that it does not have to be done within the spout if we can avoid it. Can this be done in a separate bolt? a bit like https://github.com/apache/stormcrawler/blob/main/external/opensearch/src/main/java/org/apache/stormcrawler/opensearch/metrics/StatusMetricsBolt.java Better still, you can do similar measurements within URLFrontier itself and let it report the metrics. Other crawlers would benefit from it and not just SC. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
