On Tue, Jul 21, 2026 at 3:35 PM Frederic Weisbecker <[email protected]> wrote: > > Le Wed, Jun 24, 2026 at 06:23:52AM -0700, Puranjay Mohan a écrit : > > Even when rcu_pending() triggers rcu_core(), the normal callback > > advancement path through note_gp_changes() -> __note_gp_changes() bails > > out when rdp->gp_seq == rnp->gp_seq (no normal GP change). Since > > expedited GPs do not update rnp->gp_seq, rcu_advance_cbs() is never > > called and callbacks remain stuck in RCU_WAIT_TAIL. > > > > Add a direct callback advancement block in rcu_core() that checks for GP > > completion via rcu_segcblist_nextgp() combined with > > poll_state_synchronize_rcu_full(). When detected, trylock rnp and call > > rcu_advance_cbs() to move completed callbacks to RCU_DONE_TAIL. Wake the > > GP kthread if rcu_advance_cbs() requests a new grace period. > > > > Uses trylock to avoid adding contention on rnp->lock. If the lock is > > contended, callbacks will be advanced on the next tick. > > > > Reviewed-by: Paul E. McKenney <[email protected]> > > Signed-off-by: Puranjay Mohan <[email protected]> > > --- > > kernel/rcu/tree.c | 17 +++++++++++++++++ > > 1 file changed, 17 insertions(+) > > > > diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c > > index b01d7bf6b57b1..f42e01ef479c4 100644 > > --- a/kernel/rcu/tree.c > > +++ b/kernel/rcu/tree.c > > @@ -2891,6 +2891,23 @@ static __latent_entropy void rcu_core(void) > > /* Update RCU state based on any recent quiescent states. */ > > rcu_check_quiescent_state(rdp); > > > > + /* Advance callbacks if an expedited GP has completed. */ > > + if (!rcu_rdp_is_offloaded(rdp) && > > rcu_segcblist_is_enabled(&rdp->cblist)) { > > + struct rcu_gp_seq gp_state; > > + > > + if (rcu_segcblist_nextgp(&rdp->cblist, &gp_state) && > > + poll_state_synchronize_rcu_full(&gp_state)) { > > + guard(irqsave)(); > > + if (raw_spin_trylock_rcu_node(rnp)) { > > + bool needwake = rcu_advance_cbs(rnp, rdp); > > + > > + raw_spin_unlock_rcu_node(rnp); > > + if (needwake) > > + rcu_gp_kthread_wake(); > > + } > > + } > > + } > > Should that go as an improvement to note_gp_changes() instead?
note_gp_changes() only reconciles rdp->gp_seq against rnp->gp_seq, and the expedited path never advances rnp->gp_seq. So the gap this closes is exactly rdp->gp_seq == rnp->gp_seq, where note_gp_changes() and __note_gp_changes() both short-circuit, the expedited completion isn't visible there at all. It's detected from the cblist's stored gp_seq (rcu_segcblist_nextgp()) confirmed with poll_state_synchronize_rcu_full(), so hosting it in note_gp_changes() would mean running that in the lockless preamble for every caller, including the off-tick call_rcu_core() path. In rcu_core() it's already gated by rcu_pending(), which does the barrier-free detection.
