https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298646
Bug ID: 298646
Summary: pf: pfsync_in_upd_c/pfsync_in_upd apply a stale peer's
dst TCP state, regressing the sequence window and
causing permanent PFRES_BADSTATE drops
Product: Base System
Version: CURRENT
Hardware: Any
OS: Any
Status: New
Severity: Affects Some People
Priority: ---
Component: kern
Assignee: [email protected]
Reporter: [email protected]
## Summary
On a CARP pair with pfsync state synchronisation enabled, `pfsync_in_upd_c()`
(and `pfsync_in_upd()`) overwrite the local state's **destination-side** TCP
sequence window with the peer's copy *even when that copy has already been
judged stale*. The local window regresses, `pf_tcp_track_full()` then rejects
all subsequent data and ACKs with `PFRES_BADSTATE`, and because a dropped
packet never updates a state, the connection can never recover. It hangs until
the application times out.
The drop is silent: no `pflog` record, no ICMP, no counter other than
`state-mismatch` in `pfctl -si`, and no kernel message unless `pf` debug is
raised to `misc`.
## The defect
`sys/netpfil/pf/if_pfsync.c`. `pfsync_upd_tcp()` is careful — for each side it
either accepts the update or counts it as stale and leaves the local copy
alone:
```c
if ((st->dst.state > dst->state) ||
(st->dst.state >= TCPS_SYN_SENT &&
SEQ_GT(st->dst.seqlo, ntohl(dst->seqlo))))
sync++;
else
pf_state_peer_ntoh(dst, &st->dst);
```
The caller then undoes that decision:
```c
if (st->key[PF_SK_WIRE]->proto == IPPROTO_TCP)
sync = pfsync_upd_tcp(st, &up->src, &up->dst);
else {
sync = 0;
if (st->src.state > up->src.state)
sync++;
else
pf_state_peer_ntoh(&up->src, &st->src);
if (st->dst.state > up->dst.state)
sync++;
else
pf_state_peer_ntoh(&up->dst, &st->dst);
}
if (sync < 2) {
pfsync_alloc_scrub_memory(&up->dst, &st->dst);
pf_state_peer_ntoh(&up->dst, &st->dst); /* <-- unconditional
*/
st->expire = pf_get_uptime();
st->timeout = timeout;
}
```
`sync` is a **count of how many sides were stale**, not a per-side flag. When
exactly one side is stale — `sync == 1` — the guard still passes and `dst` is
copied unconditionally.
That is the common case for any asymmetric transfer. During a bulk download the
client side (`src`) barely moves, so it compares equal and contributes nothing
to `sync`; the server side (`dst`) is the one advancing, and therefore the one
whose stale copy gets applied. The local `dst.seqlo`/`seqhi` jump *backwards*.
Note also that the `pf_state_peer_ntoh(&up->dst, &st->dst)` inside the `sync <
2` block is **redundant in both branches above** — each already copies `dst`
when the update is acceptable. So it is never needed and is actively harmful
when `dst` was the stale side.
Both `pfsync_in_upd_c()` and `pfsync_in_upd()` contain this block.
## Consequence in pf
With `dst.seqlo`/`seqhi` behind reality, `pf_tcp_track_full()` fails on the
ordinary checks:
- server data fails `data_end > src->seqhi`
- client ACKs fail `ackskew < -MAXACKWINDOW`
Both return `PF_DROP` with `PFRES_BADSTATE`. Since pf does not update a state
from a packet it drops, the window never catches up: the state is permanently
wedged. The peer's retransmissions — which by definition sit at sequence
numbers the regressed window rejects — are discarded on arrival, forever.
## Evidence from a production pair
Measured with DTrace on the live MASTER, ~380 Mbit/s of pfsync traffic between
the nodes:
- **292,647** TCP `dst.seqlo` regressions in 160 s, across ~20 GB of sequence
space; single regressions up to 2.6 MB.
- One stalled flow took **77,316 regressions (11.2 GB)** across its states
within 5 s, with writes such as `dst.seqlo 1584797355 -> 1584753195 (-44160)`.
- `pfctl -si` `state-mismatch`: **20.6/s idle**, **8,000–12,000/s during a
download**. Over one 170 s window the counter grew by 3,513 while DTrace
attributed 3,309 drops to `pf_tcp_track_full`.
- `netstat -sp pfsync` "stale states": **54.8 M** on the master, **109.8 M** on
the backup. Each node, on receiving a copy older than its own, applies the
older `dst` and pushes its state back — the loop is self-sustaining while the
flow advances.
- Packet capture confirms the shape: the server retransmits one segment at 0.9
/ 1.7 / 3.4 / 7 / 14 / 27 / 58 s; every retransmission arrives on the WAN
interface and none is forwarded. Nothing is sent back, because the client has
nothing to acknowledge.
The BACKUP passes no traffic for these states (its copies show `0:0 pkts`) yet
emitted ~5,400 pfsync frames/s toward the master.
## Reproduction
1. Two hosts in a CARP pair with `pfsync` enabled and a dedicated sync
interface.
2. Pass sustained, asymmetric TCP through the MASTER — a repeated large HTTP
download (200 MB+) through NAT is sufficient.
3. Watch `pfctl -si | grep state-mismatch` climb, and observe a fraction of
transfers hang mid-stream and never resume.
Likelihood scales with the number of states per connection (NAT64, which
creates four rather than two, failed ~83% of transfers versus ~10–40% for plain
NAT) and with anything that perturbs ACK/window positions, such as queue drops.
## Confirmation
Disabling pfsync state synchronisation on both nodes, changing nothing else:
| Measure | pfsync enabled | pfsync disabled |
| --- | --- | --- |
| `state-mismatch`, idle | 20.6/s | 0.42/s |
| 200 MB download, NAT64 path | ~83% fail | 12/12 clean |
| 200 MB download, two CI hosts | ~50–90% ok | 30/30 clean |
| 2 MB download, two CI hosts | intermittent | 20/20 clean |
## Proposed fix
Delete the redundant `pf_state_peer_ntoh(&up->dst, &st->dst)` from the `sync <
2` block in both `pfsync_in_upd_c()` and `pfsync_in_upd()`, keeping
`pfsync_alloc_scrub_memory()` and the expire/timeout refresh. Both branches
that compute `sync` already copy `dst` when the update is acceptable, so no
legitimate update is lost, and a stale `dst` is no longer applied.
If the copy is wanted for some case I have not identified, the alternative is
to make `pfsync_upd_tcp()` return per-side staleness (a bitmask rather than a
count) and skip the `dst` copy when the `dst` side was the stale one.
---
*Investigated and drafted with assistance from Claude (Anthropic). All
measurements were taken on live hardware; the code analysis was verified
against the FreeBSD sources rather than inferred. Reported by a human who
reviewed the findings.*
--
You are receiving this mail because:
You are the assignee for the bug.