dpol1 opened a new pull request, #2184: URL: https://github.com/apache/stormcrawler/pull/2184
Closes #2058. A URL which waits in the queue longer than `fetcher.timeout.queue` is acked without being fetched and without a status, and until now left only an INFO line in the log. `FetcherBolt` now counts it under `queue.timeout` in `fetcher_counter`. The URL is handled as before: no `FETCH_ERROR`, since the fetch was never attempted, and it keeps its `nextFetchDate`, so the spout emits it again. The code comment claimed that Storm had already failed such a tuple, which is not always the case. The comment and the `fetcher.timeout.queue` row in the configuration docs now state what the fetcher does, including that the check runs after robots.txt processing. The test delays robots.txt past the timeout and checks that the tuple is acked, the page is never requested, nothing is emitted, and the counter is at 1. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
