bryancall opened a new issue, #13779:
URL: https://github.com/apache/trafficserver/issues/13779

   ## Summary
   
   When an HTTP/2 origin sends GOAWAY, ATS closes the session immediately. That 
aborts every stream at or below `last_stream_id`, even though RFC 9113 ยง6.8 
says the peer may still finish them. Graceful drains (GOAWAY `NO_ERROR` with 
`last_stream_id` 2^31-1, or a real last id) therefore fail every request in 
flight on the connection. Requests with non-idempotent methods aren't retried, 
so each one goes to the client as `502` with `ERR_CONNECT_FAIL`.
   
   ## Cause
   
   `Http2ConnectionState::rcv_goaway_frame` on master:
   
   ```cpp
   this->rx_error_code = {ProxyErrorClass::SSN, 
static_cast<uint32_t>(goaway.error_code)};
   this->session->get_proxy_session()->do_io_close();
   ```
   
   1. Every stream gets EOS from `Http2Stream::initiating_close`.
   2. For a stream still waiting for its response header, 
`HttpSM::state_read_server_response_header` handles VC_EVENT_EOS by going to 
`handle_server_setup_error`, then `handle_post_failure` when a request body is 
in flight.
   3. The state ends up `CONNECTION_CLOSED`. 
`HttpTransact::is_request_retryable` refuses to retry a non-idempotent method 
once `server_request_hdr_bytes > 0`, so the client gets a 502.
   
   HTTP/1.1 has no equivalent drain signal, so an HTTP/1.1 POST hit by an 
abrupt close fails the same way. What's specific to HTTP/2 is that the origin 
announces it will finish these streams, and ATS throws them away.
   
   ## Observed
   
   On a 10.0.x build with HTTP/2 to origin enabled. `rcv_goaway_frame` closes 
the same way on master.
   
   - `502 ERR_CONNECT_FAIL` with origin status `000` and zero retries, almost 
only on POST, PROPFIND, PUT, REPORT and DELETE.
   - It comes in bursts: 183 requests across several origin sessions ended 
within the same 100 ms, while other sessions to the same origin kept returning 
200.
   - `proxy.process.http2.session_die_default` tracks 
`proxy.process.http2.goaway_frames_in` almost one-for-one (1,469 / 1,468 on one 
host). Both counters include client sessions as well as origin sessions.
   - Over HTTP/2 the sessions are long-lived and heavily multiplexed (thousands 
of transactions per session), so one drain fails dozens of requests at once.
   - A build without HTTP/2 session reuse in the same site and window, where 
each connection carried only a few requests, had 0 failures where about 5 were 
expected.
   
   We haven't yet confirmed from a packet capture that every burst was a GOAWAY 
rather than a TCP close, but the handler behaves as described whichever it was.
   
   ## Suggested fix
   
   On GOAWAY from an origin:
   1. Mark the session as draining and take it out of the session pool, so no 
new streams go on it.
   2. Reset streams with id > `last_stream_id` and make them retryable (they 
were never processed).
   3. Let streams at or below `last_stream_id` run to completion, and close the 
session once the active stream count reaches 0 or a drain timeout passes.
   
   The inbound (client) GOAWAY path could follow the same approach.
   
   Related, and may be worth its own issue: ATS keeps opening new streams on an 
origin session that has answered nothing for many seconds. In the bursts above, 
sessions silent for about 16 s were still taking new streams until about 300 ms 
before they closed. A PING-based liveness check would limit how many requests 
one stalled session can take down.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to