https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298935

            Bug ID: 298935
           Summary: IPv4 multicast routing: pending MFC expiry leaks
                    nexpire and can stall bucket expiry
           Product: Base System
           Version: 15.1-RELEASE
          Hardware: Any
                OS: Any
            Status: New
          Severity: Affects Only Me
          Priority: ---
         Component: kern
          Assignee: [email protected]
          Reporter: [email protected]

Created attachment 275199
  --> https://bugs.freebsd.org/bugzilla/attachment.cgi?id=275199&action=edit
ip_mroute-nexpire-accounting

In sys/netinet/ip_mroute.c, IPv4 multicast routing stores a byte-sized
pending-expiry count in mfct->nexpire[hash].

Creation of an unresolved MFC increments this count, while successful
resolution through MRT_ADD_MFC decrements it. Two removal paths omit the
corresponding decrement:

1. expire_upcalls() when a pending MFC expires.
2. del_mfc() when explicitly removing a still-pending MFC.

Repeated unresolved creation/timeout cycles therefore leak counts.

After 255 timeouts in an otherwise empty bucket, the next pending insertion
changes the byte-sized count from 255 to 0. expire_upcalls() skips buckets
whose count is zero, so the newly created pending MFC is no longer aged.

Subsequent packets for the same flow find that existing pending entry and fill
its pending queue instead of generating new NOCACHE notifications.

A different flow colliding with the same hash bucket could increment the
counter again and allow scanning to resume, so the stall is not unconditional
for every possible traffic pattern.


Affected source
===============

The problem was reproduced on an OPNsense FreeBSD 15.1-derived kernel built
from:

    083dc7025377cba0776aa2686f4720d1df27e693

The same omissions were also inspected in:

    FreeBSD main:
    c2f66b6616425e0e8edab3e893bc4cb58869140d

    FreeBSD stable/15:
    d1c4838a746489dc2f53bef5964f8ad40ced7173

File:

    sys/netinet/ip_mroute.c

The issue was *NOT* independently reproduced on a stock, unmodified FreeBSD
main or stable/15 binary; those branches were inspected at source level.


Relevant lifecycle
==================

Pending creation:

    rt->mfc_expire = UPCALL_EXPIRE;
    mfct->nexpire[hash]++;
    rt->mfc_parent = -1;

Successful MRT_ADD_MFC resolution correctly does:

    rt->mfc_expire = 0;
    mfct->nexpire[hash]--;

However, del_mfc() removes either pending or installed entries through:

    expire_mfc(rt);

without adjusting nexpire when the removed entry is pending.

Similarly, expire_upcalls() first skips buckets whose count is zero:

    if (mfct->nexpire[i] == 0)
        continue;

and when a pending entry expires it calls:

    expire_mfc(rt);

without decrementing mfct->nexpire[i].

The analogous IPv6 timeout path in sys/netinet6/ip6_mroute.c already
decrements its pending-expiry count before removing the entry:

    mfct->nexpire[i]--;


Reproduction and evidence
=========================

The problem was reproduced in an isolated VM environment with a multicast
router, a directly connected UDP multicast source, and a receiver.

Forwarding was first verified with an active receiver. The receiver then left
the group while the source continued transmitting at approximately
10 packets/second with TTL 2.

With no downstream outgoing interface, source traffic resulted in successive
NOCACHE-created pending MFCs which were allowed to expire.

After more than 256 such cycles, a late receiver join was attempted.

Expected behavior:

Each pending timeout should restore the bucket's pending-entry count. New
source packets should continue generating NOCACHE notifications, and a later
receiver join should be able to establish the required (S,G) forwarding state.

Observed behavior on the unmodified kernel:

A directly sampled bucket contained no MFC entries but had:

    nexpire = 255

The next source packet created one pending MFC and changed the counter to:

    nexpire = 0

The new pending entry had:

    mfc_expire = 6
    mfc_parent = -1

Its pending queue filled to its three usable slots, while mfc_expire remained
at 6 for more than 242 seconds even though the multicast expiry callout itself
continued to run.

Over an approximately 10-second interval:

    multicast cache misses        +100
    pending queue overflows       +100
    routing-daemon upcalls          +0
    cache cleanups                  +0
    upcall socket-buffer drops      +0

Before leaving, the receiver had received 180 forwarded packets at TTL 1.
After the stuck state developed, a late receiver rejoin received zero packets
despite continued source traffic.

Daemon history for the affected flow showed:

    257 NOCACHE-triggering pending creations
      1 successful resolution

Therefore:

    (257 - 1) mod 256 = 0

which matches the directly observed zero counter.

Read-only kernel inspection also found the converse inconsistency: empty hash
buckets with nonzero leaked nexpire values.

An unresolved entry has mfc_parent = -1. Since vifi_t is an unsigned short,
netstat displays this as:

    In-Vif 65535

65535 is therefore normal temporarily while an MFC is unresolved. The failure
condition is its persistence together with a frozen expiry, a full pending
queue, and a zero nexpire count.


Proposed correction
===================

The pending-entry count should be decremented on both missing removal paths:

1. When expire_upcalls() removes a timed-out pending MFC.
2. When del_mfc() explicitly removes a pending MFC.

For del_mfc(), the pending state can be identified by a nonempty stall ring.

Using mfc_expire != 0 as the sole predicate in a common cleanup function would
not be sufficient because expire_upcalls() has already decremented
mfc_expire to zero before removing an expired entry.

A minimal two-hunk patch implementing only these two accounting corrections is
attached as:

    ip_mroute-nexpire-accounting.patch


Independent validation of the minimal patch
===========================================

The exact two-hunk patch was independently built and booted from the affected
source revision.

The original:

    u_char *nexpire

and its allocation size were intentionally left unchanged so that the
accounting fix could be tested independently of counter widening.

Native MRT API tests confirmed:

    pending timeout:          nexpire 1 -> 0
    pending MRT_DEL_MFC:      nexpire 1 -> 0
    successful MRT_ADD_MFC:   nexpire 1 -> 0

Deleting an already resolved entry did not change nexpire, and deleting a
nonexistent entry also left the count unchanged.

In a separate single-source test with no receiver, the daemon recorded:

    408 NOCACHE events
    407 pending-entry cleanups

during 650 seconds.

The bucket continued aging normally beyond the 256th sequential expiry.

Across the complete run, all 7,384 stable read-only samples showed nexpire
consistent with the actual number of pending entries.

After 650 seconds without a receiver, a receiver rejoin successfully forwarded
193 packets at TTL 1.

Three subsequent short leave/rejoin cycles also restored forwarding normally.

These tests were performed using the FreeBSD 15.1-derived OPNsense kernel
source described above. The minimal patch was not separately runtime-tested on
a stock FreeBSD main or stable/15 binary.


Separate counter-width consideration
====================================

Correcting lifecycle accounting does not address a separate theoretical limit:
because nexpire is eight bits wide, 256 simultaneously pending entries hashing
to the same bucket could still wrap the counter.

Widening the counter and increasing its allocation size would address that
independent limitation.

Counter widening alone, however, does not fix the accounting leaks described
in this report and is intentionally not part of the proposed minimal patch.

-- 
You are receiving this mail because:
You are the assignee for the bug.

Reply via email to