On 2026/7/15 23:15, Eric Dumazet wrote:
> On Wed, Jul 15, 2026 at 4:54 PM Leon Hwang <[email protected]> wrote:
[...]
>> diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
>> index 61045a8886e4..4f1027173e95 100644
>> --- a/net/ipv4/tcp_input.c
>> +++ b/net/ipv4/tcp_input.c
>> @@ -4853,6 +4853,7 @@ void tcp_done_with_error(struct sock *sk, int err)
>> /* When we get a reset we do this. */
>> void tcp_reset(struct sock *sk, struct sk_buff *skb)
>> {
>> + const struct net *net = sock_net(sk);
>> int err;
>>
>> trace_tcp_receive_reset(sk);
>> @@ -4869,6 +4870,27 @@ void tcp_reset(struct sock *sk, struct sk_buff *skb)
>> err = ECONNREFUSED;
>> break;
>> case TCP_CLOSE_WAIT:
>> + /* RFC9293 3.10.7.4. Other States
>> + * Second, check the RST bit:
>> + * CLOSE-WAIT STATE
>> + *
>> + * If the RST bit is set, then any outstanding RECEIVEs and
>> + * SEND should receive "reset" responses. All segment queues
>> + * should be flushed. Users should also receive an
>> unsolicited
>> + * general "connection reset" signal. Enter the CLOSED
>> state,
>> + * delete the TCB, and return.
>> + *
>> + * If net.ipv4.tcp_purge_receive_queue is enabled,
>> + * sk_receive_queue will be flushed too.
>> + */
>> + if
>> (unlikely(READ_ONCE(net->ipv4.sysctl_tcp_purge_receive_queue))) {
>> + struct tcp_sock *tp = tcp_sk(sk);
>> +
>> + skb_queue_purge(&sk->sk_receive_queue);
>> + WRITE_ONCE(tp->copied_seq, tp->rcv_nxt);
>> + WRITE_ONCE(tp->urg_data, 0);
>> + sk_set_peek_off(sk, -1);
>> + }
>> err = EPIPE;
>> break;
>> case TCP_CLOSE:
>> --
>> 2.55.0
>>
>
> My thoughts are:
>
> out_of_order_queue has been forgotten. skbs could be there and still
> 'block devmem'
>
> WRITE_ONCE(tp->copied_seq, tp->rcv_nxt) is certainly wrong, because
> read() will return 0, instead of -1 (errno = EPIPE or ECONNRESET)
> So the application will not know a RST was received :/
>
> I think that BSD and linux implementations have historically retained
> acknowledged,
> buffered receive data upon RST to allow applications to drain data
> already ACKed prior to the reset.
>
> Adding a narrow sysctl specifically for CLOSE_WAIT creates
> inconsistent behavior across TCP states.
Got it. I won't pursue this sysctl approach in the future. Thanks for
the review.
Leon