rzo1 commented on issue #2185:
URL: https://github.com/apache/stormcrawler/issues/2185#issuecomment-5876570479

   Late to the party, but I am also wondering what is missing in practice.
   
   Depth is tracked by default (`metadata.track.depth`) and `MaxDepthFilter` 
already caps it, globally or per seed via `max.depth`. That covers the case of 
a crawl sinking into pagination or facets on a few large hosts, although as a 
hard limit and not as an ordering.
   
   If the goal is ordering, like in a focused crawl, that is possible today 
with the OpenSearch backend: the sort fields of the spout are configurable 
(`opensearch.status.bucket.sort.field`, `opensearch.status.global.sort.field`), 
so one can sort on depth or on a score instead of `nextFetchDate`.
   
   The gap is with URLFrontier only. There is no priority in its API, and for 
DISCOVERED URLs we do not even send a date, so nothing on the SC side can 
influence the order in which new URLs are served. That would have to be added 
in URLFrontier, so +1 to discuss it there.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to