> On Oct 3, 2026, at 9:14 PM, Cunlong Li <[email protected]> wrote: > > The kernel.panic_on_rcu_stall and kernel.max_rcu_stall_to_panic sysctls > count RCU CPU stalls since system boot, and invoke panic() once > max_rcu_stall_to_panic stalls have elapsed. This count is never reset, > so stalls caused by transient incidents keep consuming the budget of a > long-running system, and a later unrelated stall can
Please provide examples of specific transient events you ran into, and recovered from? > immediately > trigger panic() instead of allowing the intended fresh window of > stalls. > > This commit therefore introduces the kernel.rcu_stall_panic_count > sysctl. Reading this file reports the number of stalls counted so far, > and writing 0 to it resets the count. This allows the cumulative stall > history to be cleared after an incident has been resolved, and also > allows userspace to clear the count periodically, so that panic() is > only triggered by a burst of stalls occurring within a short period. Could you provide details of such an incident that got resolved without requiring a reboot? > > Signed-off-by: Cunlong Li <[email protected]> If the commit is AI assisted, please add an assisted tag. Thanks, Joel > --- > Documentation/admin-guide/sysctl/kernel.rst | 15 ++++++++++++++- > kernel/rcu/tree_stall.h | 28 ++++++++++++++++++++++++++-- > 2 files changed, 40 insertions(+), 3 deletions(-) > > diff --git a/Documentation/admin-guide/sysctl/kernel.rst > b/Documentation/admin-guide/sysctl/kernel.rst > index ffea61d448eb..64fe2985e358 100644 > --- a/Documentation/admin-guide/sysctl/kernel.rst > +++ b/Documentation/admin-guide/sysctl/kernel.rst > @@ -959,7 +959,20 @@ max_rcu_stall_to_panic > When ``panic_on_rcu_stall`` is set to 1, this value determines the > number of times that RCU can stall before panic() is called. > > -When ``panic_on_rcu_stall`` is set to 0, this value is has no effect. > +When ``panic_on_rcu_stall`` is set to 0, this value has no effect. > + > +rcu_stall_panic_count > +===================== > + > +Indicates the number of RCU CPU stalls that have been counted since > +system boot or since the counter was reset. When ``panic_on_rcu_stall`` > +is set to 1, this count is compared against ``max_rcu_stall_to_panic`` > +to decide whether panic() should be called. > + > +Writing 0 to this file resets the counter to zero, which restarts the > +``max_rcu_stall_to_panic`` window of stalls. This allows system > +administrators to clear the cumulative stall count after an incident > +has been resolved, without requiring a system restart. > > perf_cpu_time_max_percent > ========================= > diff --git a/kernel/rcu/tree_stall.h b/kernel/rcu/tree_stall.h > index 091e7850ab6e..a80f03e1c7ea 100644 > --- a/kernel/rcu/tree_stall.h > +++ b/kernel/rcu/tree_stall.h > @@ -19,6 +19,20 @@ > /* panic() on RCU Stall sysctl. */ > static int sysctl_panic_on_rcu_stall __read_mostly; > static int sysctl_max_rcu_stall_to_panic __read_mostly; > +static unsigned long sysctl_rcu_stall_panic_count; > + > +/* Reset the RCU stall panic count when written to. */ > +static int proc_do_rcu_stall_panic_count(const struct ctl_table *table, int > write, > + void *buffer, size_t *lenp, loff_t *ppos) > +{ > + if (!write) > + return proc_doulongvec_minmax(table, write, buffer, lenp, ppos); > + > + WRITE_ONCE(sysctl_rcu_stall_panic_count, 0); > + *ppos += *lenp; > + > + return 0; > +} > > static const struct ctl_table rcu_stall_sysctl_table[] = { > { > @@ -39,6 +53,13 @@ static const struct ctl_table rcu_stall_sysctl_table[] = { > .extra1 = SYSCTL_ONE, > .extra2 = SYSCTL_INT_MAX, > }, > + { > + .procname = "rcu_stall_panic_count", > + .data = &sysctl_rcu_stall_panic_count, > + .maxlen = sizeof(sysctl_rcu_stall_panic_count), > + .mode = 0644, > + .proc_handler = proc_do_rcu_stall_panic_count, > + }, > }; > > static int __init init_rcu_stall_sysctl(void) > @@ -161,7 +182,7 @@ early_initcall(check_cpu_stall_init); > /* If so specified via sysctl, panic, yielding cleaner stall-warning output. > */ > static void panic_on_rcu_stall(const struct cpumask *stalled_mask) > { > - static int cpu_stall; > + unsigned long count; > > /* > * Attempt to kick out the BPF scheduler if it's installed and defer > @@ -170,7 +191,10 @@ static void panic_on_rcu_stall(const struct cpumask > *stalled_mask) > if (scx_rcu_cpu_stall(stalled_mask)) > return; > > - if (++cpu_stall < sysctl_max_rcu_stall_to_panic) > + /* A lost RMW update only delays the panic by one stall. */ > + count = READ_ONCE(sysctl_rcu_stall_panic_count) + 1; > + WRITE_ONCE(sysctl_rcu_stall_panic_count, count); > + if (count < (unsigned long)READ_ONCE(sysctl_max_rcu_stall_to_panic)) > return; > > if (sysctl_panic_on_rcu_stall) > > --- > base-commit: ce1e0223d8ad4211275c82a17ed6d43ab81e13d9 > change-id: 20261003-rcu-375496d7704e > > Best regards, > -- > Cunlong Li <[email protected]> >

