jnioche commented on issue #2185:
URL: https://github.com/apache/stormcrawler/issues/2185#issuecomment-5812625892

   thanks @GGraziadei 
   
   > With URLFrontier the order of the crawl is decided entirely by the 
frontier: it rotates between queues (one per host) and serves each queue FIFO
   
   this needs checking but if I remember correctly, this is might not be the 
case. The priority of a queue in URLFrontier is determined by the earliest 
`nextFetchDate` of its URLs.  The depth of a URL is usually reflected by its 
`nextFetchDate`.
   
   On a more general note, the advantage of using URLFrontier is that we 
delegate a lot of the logic away from SC - what you are suggesting it would 
reintroduce some of it. If we want to change the way URLs are served by 
URLFrontier then it should happen there.
   
   Having said that, if we just want to observe things, this is totally OK but 
I would argue that it does not have to be done within the spout if we can avoid 
it. Can this be done in a separate bolt? a bit like 
https://github.com/apache/stormcrawler/blob/main/external/opensearch/src/main/java/org/apache/stormcrawler/opensearch/metrics/StatusMetricsBolt.java
   
   Better still, you can do similar measurements within URLFrontier itself and 
let it report the metrics. Other crawlers would benefit from it and not just SC.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to