GGraziadei opened a new issue, #2185: URL: https://github.com/apache/stormcrawler/issues/2185
With URLFrontier the order of the crawl is decided entirely by the frontier: it rotates between queues (one per host) and serves each queue FIFO. The depth of a URL plays no role, so a crawl that starts broad can sink into the deep pages of a few large hosts (pagination, facets) without anything in the topology showing it. `GetURLs` gives no way to ask for shallow URLs either, so before adding any steering it must be possible to measure the shape of the crawl. Proposal: have the URLFrontier `Spout` count the URLs handed out by the frontier per depth level, using the `depth` metadata written when `metadata.track.depth` is enabled: - a `depth` counter with one scope per depth (`depth.0`, `depth.1`, ...), an `unknown` scope for URLs without a numeric depth, and a single `N+` scope from `urlfrontier.depth.metric.max` (default 10) upwards so that the number of scopes stays bounded; - a `depth_le` counter with cumulative counts in the style of Prometheus histograms (`depth_le.X` = URLs with depth <= X, `depth_le.inf` = total), so that P(depth <= X) is a plain ratio in a dashboard; - `probabilityDepthAtMost(int)` and `windowProbabilityDepthAtMost(int)` on the spout, returning that ratio in constant time since the spout was opened and over a sliding window of `urlfrontier.depth.window.secs` (default 300), for a future control loop in `populateBuffer` (e.g. more queues with fewer URLs each when the recent share of shallow URLs drops). No change to the behaviour of the crawl, only observability. Steering (adaptive `max.buckets` / `urls.per.bucket`, or explicit queue selection with the `key` parameter) can follow in a separate issue once the metric shows the drift on real crawls. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
