On Thu, Aug 6, 2026 at 6:12 AM Hyunjung Ko <[email protected]> wrote:
>
> tcf_ct_handle_fragments() runs its header sanity checks before handing
> anything to the defragmentation engine:
>
>         if (family == NFPROTO_IPV4)
>                 err = tcf_ct_ipv4_is_fragment(skb, &frag);
>         else
>                 err = tcf_ct_ipv6_is_fragment(skb, &frag);
>         if (err || !frag)
>                 return err;
>
> tcf_ct_ipv4_is_fragment() returns -EINVAL or -ENOMEM;
> tcf_ct_ipv6_is_fragment() adds -EPROTO when ipv6_find_hdr() fails. None of
> them frees or queues the skb, so on that path the caller still owns it.
>
> tcf_ct_act() however funnels every non-zero return into the
> ownership-transfer exit:
>
>         err = tcf_ct_handle_fragments(net, skb, family, p->zone, &defrag);
>         if (err)
>                 goto out_frag;
>         ...
> out_frag:
>         if (err != -EINPROGRESS)
>                 tcf_action_inc_drop_qstats(&c->common);
>         return TC_ACT_CONSUMED;
>
> TC_ACT_CONSUMED means the action took ownership of the skb, so no caller
> frees it - sch_handle_ingress(), sch_handle_egress() and
> tcf_qevent_handle() all deliberately skip the free for that verdict. The
> skb is therefore orphaned: one sk_buff plus its data buffer is leaked per
> malformed packet, unbounded. Note the drop counter is already incremented
> for these errors, so the statistics claim a drop that never happens.
>
> Three different ownership states reach out_frag: today - the skb may be
> queued by the defrag engine (-EINPROGRESS), already freed by
> nf_ct_handle_fragments(), or still owned by us. Tell the caller which of
> those it is, and free the packet ourselves in the last case, which
> restores the TC_ACT_SHOT behaviour that predated the Fixes: commit.
>
> Reproduced on v7.2-rc6 with a 54-byte frame carrying a 40-byte IPv6
> header with nexthdr = 0 (hop-by-hop) and nothing after it, on a
> clsact ingress chain with "action ct". kmemleak reports one leaked
> 232-byte skbuff_head_cache object plus its 704-byte data buffer per
> packet; with this patch it reports none.
>
> Fixes: 3f14b377d01d ("net/sched: act_ct: fix skb leak and crash on ooo frags")
> Cc: [email protected] # v6.8+
> Assisted-by: Anthropic-Claude-Code:Claude-Opus-5
> Signed-off-by: Hyunjung Ko <[email protected]>

I actually have tested and reviewed this. So:

Acked-by: Jamal Hadi Salim <[email protected]>

cheers,
jamal
> ---
>  net/sched/act_ct.c | 29 +++++++++++++++++++++++++----
>  1 file changed, 25 insertions(+), 4 deletions(-)
>
> v2:
>  - add a tdc selftest (patch 2/2), as requested by Jamal
>  - add Assisted-by: tag
>  - no functional change to the fix itself
> v1: https://lore.kernel.org/netdev/[email protected]/
>
> Reproducer needs CONFIG_NET_ACT_CT, plus CONFIG_DEBUG_KMEMLEAK and
> kmemleak=on to observe it:
>
>   ip link add veth0 type veth peer name veth1
>   ip link set veth0 up; ip link set veth1 up
>   tc qdisc add dev veth0 clsact
>   tc filter add dev veth0 ingress matchall action ct
>
> then inject at veth1 a 54-byte frame: ethertype 0x86DD, a 40-byte IPv6
> header with nexthdr = 0 (hop-by-hop) and nothing after it, so
> ipv6_find_hdr() fails with -EBADMSG and tcf_ct_ipv6_is_fragment()
> returns -EPROTO.
>
> Before, one sk_buff plus its data buffer per packet:
>
>   kmemleak: 50 new suspected memory leaks
>   unreferenced object 0xffff888103ed13c0 (size 232):
>     kmem_cache_alloc_node_noprof+0x2f1/0x3e0
>     __alloc_skb+0xe5/0x860
>     alloc_skb_with_frags+0x82/0x750
>     sock_alloc_send_pskb+0x658/0x7e0
>     packet_sendmsg+0x1833/0x4860
>     __x64_sys_sendto+0xe0/0x1c0
>     do_syscall_64+0x102/0x5a0
>
> After: kmemleak reports no unreferenced objects.
>
> Note /proc/slabinfo is not a usable check here on a KASAN build -
> skbuff_head_cache active_objs still grows because the quarantine holds
> the freed objects. kmemleak is the reliable signal.
>
> Patch 2/2 turns the same case into a tdc test, using the clsact drop
> counter as the discriminator: before the fix act_ct returns
> TC_ACT_CONSUMED, so tc_run() never reaches its TC_ACT_SHOT arm and the
> counter stays at zero while the skbs leak.
>
> diff --git a/net/sched/act_ct.c b/net/sched/act_ct.c
> index be535a261fa0..e250969c84ac 100644
> --- a/net/sched/act_ct.c
> +++ b/net/sched/act_ct.c
> @@ -840,8 +840,15 @@ static int tcf_ct_ipv6_is_fragment(struct sk_buff *skb, 
> bool *frag)
>         return 0;
>  }
>
> +/* On error, tells the caller whether it still owns @skb and must free it
> + * itself.  @skb is ours only when the header checks below reject the packet
> + * before it is handed to the defragmentation engine; once nf_ct_handle_
> + * fragments() has been called the skb is either queued (-EINPROGRESS) or has
> + * already been freed by it.
> + */
>  static int tcf_ct_handle_fragments(struct net *net, struct sk_buff *skb,
> -                                  u8 family, u16 zone, bool *defrag)
> +                                  u8 family, u16 zone, bool *defrag,
> +                                  bool *skb_is_ours)
>  {
>         enum ip_conntrack_info ctinfo;
>         struct tc_skb_cb cb;
> @@ -859,8 +866,12 @@ static int tcf_ct_handle_fragments(struct net *net, 
> struct sk_buff *skb,
>                 err = tcf_ct_ipv4_is_fragment(skb, &frag);
>         else
>                 err = tcf_ct_ipv6_is_fragment(skb, &frag);
> -       if (err || !frag)
> +       if (err) {
> +               *skb_is_ours = true;
>                 return err;
> +       }
> +       if (!frag)
> +               return 0;
>
>         cb = *tc_skb_cb(skb);
>         err = nf_ct_handle_fragments(net, skb, zone, family, &proto, &cb.mru);
> @@ -977,6 +988,7 @@ TC_INDIRECT_SCOPE int tcf_ct_act(struct sk_buff *skb, 
> const struct tc_action *a,
>         int nh_ofs, err, retval;
>         struct tcf_ct_params *p;
>         bool add_helper = false;
> +       bool skb_is_ours = false;
>         bool skip_add = false;
>         bool defrag = false;
>         struct nf_conn *ct;
> @@ -1012,9 +1024,18 @@ TC_INDIRECT_SCOPE int tcf_ct_act(struct sk_buff *skb, 
> const struct tc_action *a,
>          */
>         nh_ofs = skb_network_offset(skb);
>         skb_pull_rcsum(skb, nh_ofs);
> -       err = tcf_ct_handle_fragments(net, skb, family, p->zone, &defrag);
> -       if (err)
> +       err = tcf_ct_handle_fragments(net, skb, family, p->zone, &defrag,
> +                                     &skb_is_ours);
> +       if (err) {
> +               /* The skb is still ours only when the header checks rejected
> +                * it; returning TC_ACT_CONSUMED for such a packet would leak
> +                * it, since no caller frees an skb it was told it no longer
> +                * owns.
> +                */
> +               if (skb_is_ours)
> +                       goto drop;
>                 goto out_frag;
> +       }
>
>         err = nf_ct_skb_network_trim(skb, family);
>         if (err)
> --
> 2.43.0

Reply via email to