On Thu, 30 Jul 2026 17:47:17 -0700 Mingming Cao wrote: > After aggressive ethtool -L cycling, PHYP can leave a VALID RX descriptor > with a correlator that no longer matches the per-queue buffer pools. Poll > treated this as fatal: ibmveth_rxq_get_buffer() WARNed and returned NULL > without advancing the ring, then restart_poll retried the same slot > forever. > > Advance past bad correlators instead of spinning: validate correlators > without WARN_ON, skip invalid slots in poll (count as invalid_buffers), > and advance the RX ring when remove_buffer_from_pool cannot map the > correlator. Rate-limit the bad correlator message. > > Complete NAPI when the interface is down or napi_disable is pending so > close/quiesce can finish. Do not restart_poll in that window. Close > keeps hypervisor IRQ disable before napi_disable (via > cleanup_rx_interrupts() / related cleanup helpers). > > Also validate descriptor length against skb tailroom before skb_put(), > and after napi_complete_done() on the budget-exhausted shutdown path > return a value less than budget so NAPI does not immediately reschedule.
This patch looks very odd, like there's a bug somewhere and instead of fixing it we're making the rx path more resilient. BTW minor coding note - you're adding a bunch of heavily indented code to an already long function. Please consider factoring things out. Similar comment can be applied to other patches.
