rzo1 commented on issue #2185: URL: https://github.com/apache/stormcrawler/issues/2185#issuecomment-5876570479
Late to the party, but I am also wondering what is missing in practice. Depth is tracked by default (`metadata.track.depth`) and `MaxDepthFilter` already caps it, globally or per seed via `max.depth`. That covers the case of a crawl sinking into pagination or facets on a few large hosts, although as a hard limit and not as an ordering. If the goal is ordering, like in a focused crawl, that is possible today with the OpenSearch backend: the sort fields of the spout are configurable (`opensearch.status.bucket.sort.field`, `opensearch.status.global.sort.field`), so one can sort on depth or on a score instead of `nextFetchDate`. The gap is with URLFrontier only. There is no priority in its API, and for DISCOVERED URLs we do not even send a date, so nothing on the SC side can influence the order in which new URLs are served. That would have to be added in URLFrontier, so +1 to discuss it there. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
