On 2026/7/15 23:30, Eric Dumazet wrote: > On Wed, Jul 15, 2026 at 5:26 PM Leon Hwang <[email protected]> wrote: >> >> On 2026/7/15 23:15, Eric Dumazet wrote: >>> On Wed, Jul 15, 2026 at 4:54 PM Leon Hwang <[email protected]> wrote: >> >> [...] >> >>>> diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c >>>> index 61045a8886e4..4f1027173e95 100644 >>>> --- a/net/ipv4/tcp_input.c >>>> +++ b/net/ipv4/tcp_input.c >>>> @@ -4853,6 +4853,7 @@ void tcp_done_with_error(struct sock *sk, int err) >>>> /* When we get a reset we do this. */ >>>> void tcp_reset(struct sock *sk, struct sk_buff *skb) >>>> { >>>> + const struct net *net = sock_net(sk); >>>> int err; >>>> >>>> trace_tcp_receive_reset(sk); >>>> @@ -4869,6 +4870,27 @@ void tcp_reset(struct sock *sk, struct sk_buff *skb) >>>> err = ECONNREFUSED; >>>> break; >>>> case TCP_CLOSE_WAIT: >>>> + /* RFC9293 3.10.7.4. Other States >>>> + * Second, check the RST bit: >>>> + * CLOSE-WAIT STATE >>>> + * >>>> + * If the RST bit is set, then any outstanding RECEIVEs and >>>> + * SEND should receive "reset" responses. All segment >>>> queues >>>> + * should be flushed. Users should also receive an >>>> unsolicited >>>> + * general "connection reset" signal. Enter the CLOSED >>>> state, >>>> + * delete the TCB, and return. >>>> + * >>>> + * If net.ipv4.tcp_purge_receive_queue is enabled, >>>> + * sk_receive_queue will be flushed too. >>>> + */ >>>> + if >>>> (unlikely(READ_ONCE(net->ipv4.sysctl_tcp_purge_receive_queue))) { >>>> + struct tcp_sock *tp = tcp_sk(sk); >>>> + >>>> + skb_queue_purge(&sk->sk_receive_queue); >>>> + WRITE_ONCE(tp->copied_seq, tp->rcv_nxt); >>>> + WRITE_ONCE(tp->urg_data, 0); >>>> + sk_set_peek_off(sk, -1); >>>> + } >>>> err = EPIPE; >>>> break; >>>> case TCP_CLOSE: >>>> -- >>>> 2.55.0 >>>> >>> >>> My thoughts are: >>> >>> out_of_order_queue has been forgotten. skbs could be there and still >>> 'block devmem' >>> >>> WRITE_ONCE(tp->copied_seq, tp->rcv_nxt) is certainly wrong, because >>> read() will return 0, instead of -1 (errno = EPIPE or ECONNRESET) >>> So the application will not know a RST was received :/ >>> >>> I think that BSD and linux implementations have historically retained >>> acknowledged, >>> buffered receive data upon RST to allow applications to drain data >>> already ACKed prior to the reset. >>> >>> Adding a narrow sysctl specifically for CLOSE_WAIT creates >>> inconsistent behavior across TCP states. >> >> >> Got it. I won't pursue this sysctl approach in the future. Thanks for >> the review. > > My intention was not to kill your proposal, only to start a conversation...
Thanks for clarifying. I agree this approach needs more thought. Thanks, Leon

