dpol1 opened a new pull request, #2184:
URL: https://github.com/apache/stormcrawler/pull/2184

   Closes #2058.
   
   A URL which waits in the queue longer than `fetcher.timeout.queue` is acked 
without being fetched and without a status, and until now left only an INFO 
line in the log. `FetcherBolt` now counts it under `queue.timeout` in 
`fetcher_counter`. The URL is handled as before: no `FETCH_ERROR`, since the 
fetch was never attempted, and it keeps its `nextFetchDate`, so the spout emits 
it again.
   
   The code comment claimed that Storm had already failed such a tuple, which 
is not always the case. The comment and the `fetcher.timeout.queue` row in the 
configuration docs now state what the fetcher does, including that the check 
runs after robots.txt processing.
   
   The test delays robots.txt past the timeout and checks that the tuple is 
acked, the page is never requested, nothing is emitted, and the counter is at 1.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to