bryancall opened a new issue, #13765:
URL: https://github.com/apache/trafficserver/issues/13765

   **Problem**
   
   With the default `proxy.config.http2.flow_control.policy_out: 0`, the 
receive window ATS advertises on an HTTP/2 connection to an origin is the same 
size for the whole connection as for a single stream: 
`initial_window_size_out`, 65,535 bytes. RFC 9113 lets a receiver enlarge the 
connection-level window beyond that with a `WINDOW_UPDATE` on stream 0, and 
policy 0 never does.
   
   That was a reasonable default when a connection carried one request at a 
time. With HTTP/2 to origin, one connection now multiplexes many concurrent 
streams, and all of them share that single 64 KiB connection window:
   
   - Aggregate response throughput on the connection is bounded by roughly 64 
KiB per window-update round trip. At 150 ms RTT that is about 440 KB/s for 
**every stream combined**.
   - One large or slow-draining response can consume the whole connection 
window and stall every other stream on that connection until ATS sends a 
connection-level `WINDOW_UPDATE`.
   
   **Proposal**
   
   Change the default of `proxy.config.http2.flow_control.policy_out` from `0` 
to `1`. Under policy 1 the connection window is `initial_window_size_out × 
max_concurrent_streams_out` (65,535 × 100 ≈ 6.25 MiB with the other defaults), 
while each stream's window stays at `initial_window_size_out`. So per-stream 
behaviour is unchanged, and streams stop contending for one small 
connection-level budget.
   
   **Trade-off to discuss**
   
   A larger connection window lets an origin send more before ATS pushes back. 
If the downstream client drains slowly, ATS may hold more response data per 
origin connection. The per-stream window still caps each stream, and 
`max_concurrent_streams_out` bounds the product, but this is the reason to 
change the default deliberately rather than quietly. Policy 2, which also 
rebalances stream windows dynamically, is out of scope here. It changes 
per-stream behaviour and deserves its own discussion.
   
   **Production experience**
   
   Yahoo's cloud proxy tier has run `policy_out: 1` (with 
`max_concurrent_streams_out: 1000`, and `policy_in: 1`) on a set of production 
hosts since 2026-09-30 00:30 UTC. Those hosts proxy gRPC over HTTP/2 to mTLS 
origins, and no issues have been observed so far. The setting was in 
records.yaml before that restart, so it is in effect despite #13694. Mean 
origin latency on those hosts fell substantially over the same period. However, 
the same rollout moved to a build carrying other HTTP/2-to-origin fixes, and at 
least one of those hosts had previously been configured with `policy_out: 2`. 
So **the improvement cannot be attributed to this setting**, and this issue 
does not rest on it. The argument above holds on its own.
   
   **Questions**
   
   - Should `policy_in` get the same default change? Inbound connections have 
the same shape, but different memory exposure, since the peers are arbitrary 
clients.
   - Is there a reason policy 0 was kept as the default when HTTP/2 to origin 
landed in #9366?
   
   Related: #8199 (separate stream and session windows in HTTP/2), #9085 (added 
HTTP/2 flow control configuration, including `policy_in`), #9366 (Http2 to 
origin, which added `policy_out`), #13694 (the flow-control policies do not 
take effect on reload, so a changed default only applies after a restart).
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to