On Wed, Jul 15, 2026 at 5:26 PM Leon Hwang <[email protected]> wrote:
>
> On 2026/7/15 23:15, Eric Dumazet wrote:
> > On Wed, Jul 15, 2026 at 4:54 PM Leon Hwang <[email protected]> wrote:
>
> [...]
>
> >> diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
> >> index 61045a8886e4..4f1027173e95 100644
> >> --- a/net/ipv4/tcp_input.c
> >> +++ b/net/ipv4/tcp_input.c
> >> @@ -4853,6 +4853,7 @@ void tcp_done_with_error(struct sock *sk, int err)
> >>  /* When we get a reset we do this. */
> >>  void tcp_reset(struct sock *sk, struct sk_buff *skb)
> >>  {
> >> +       const struct net *net = sock_net(sk);
> >>         int err;
> >>
> >>         trace_tcp_receive_reset(sk);
> >> @@ -4869,6 +4870,27 @@ void tcp_reset(struct sock *sk, struct sk_buff *skb)
> >>                 err = ECONNREFUSED;
> >>                 break;
> >>         case TCP_CLOSE_WAIT:
> >> +               /* RFC9293 3.10.7.4. Other States
> >> +                *   Second, check the RST bit:
> >> +                *     CLOSE-WAIT STATE
> >> +                *
> >> +                * If the RST bit is set, then any outstanding RECEIVEs and
> >> +                * SEND should receive "reset" responses.  All segment 
> >> queues
> >> +                * should be flushed.  Users should also receive an 
> >> unsolicited
> >> +                * general "connection reset" signal.  Enter the CLOSED 
> >> state,
> >> +                * delete the TCB, and return.
> >> +                *
> >> +                * If net.ipv4.tcp_purge_receive_queue is enabled,
> >> +                * sk_receive_queue will be flushed too.
> >> +                */
> >> +               if 
> >> (unlikely(READ_ONCE(net->ipv4.sysctl_tcp_purge_receive_queue))) {
> >> +                       struct tcp_sock *tp = tcp_sk(sk);
> >> +
> >> +                       skb_queue_purge(&sk->sk_receive_queue);
> >> +                       WRITE_ONCE(tp->copied_seq, tp->rcv_nxt);
> >> +                       WRITE_ONCE(tp->urg_data, 0);
> >> +                       sk_set_peek_off(sk, -1);
> >> +               }
> >>                 err = EPIPE;
> >>                 break;
> >>         case TCP_CLOSE:
> >> --
> >> 2.55.0
> >>
> >
> > My thoughts are:
> >
> > out_of_order_queue has been forgotten. skbs could be there and still
> > 'block devmem'
> >
> > WRITE_ONCE(tp->copied_seq, tp->rcv_nxt) is certainly wrong, because
> > read() will return 0, instead of -1 (errno = EPIPE or ECONNRESET)
> > So the application will not know a RST was received :/
> >
> > I think that BSD and linux implementations have historically retained
> > acknowledged,
> > buffered receive data upon RST to allow applications to drain data
> > already ACKed prior to the reset.
> >
> > Adding a narrow sysctl specifically for CLOSE_WAIT creates
> > inconsistent behavior across TCP states.
>
>
> Got it. I won't pursue this sysctl approach in the future. Thanks for
> the review.

My intention was not to kill your proposal, only to start a conversation...

Reply via email to