On Mon, Sep 14, 2026 at 2:18 AM Paul E. McKenney <[email protected]> wrote:
>
> On Sun, Sep 13, 2026 at 06:36:56PM +0800, KunWu Chan wrote:
> > On Sun, Sep 13, 2026 at 2:21 PM Zqiang <[email protected]> wrote:
> > >
> > > >
> > > > On Sat, Sep 12, 2026 at 08:37:33PM +0800, KunWu Chan wrote:
> > > >
> > > > >
> > > > > On Sat, Sep 12, 2026 at 12:47 AM Paul E. McKenney 
> > > > > <[email protected]> wrote:
> > > > >
> > > > >  On Fri, Sep 11, 2026 at 05:35:15PM +0800, Kunwu Chan wrote:
> > > > >  > Add the lockdep annotation, same-type SRCU nesting warning, and
> > > > >  > early-boot check used by __synchronize_srcu().
> > > > >  >
> > > > >  > Suggested-by: Zqiang <[email protected]>
> > > > >  > Signed-off-by: Kunwu Chan <[email protected]>
> > > > >
> > > > >  Queued for review and testing, thank you both!
> > > > >
> > > > >  Interestingly enough, it is now the case that there is a grace-period
> > > > >  wait that can be placed in a normal RCU read-side critical section.
> > > > >  Does this mean that we should also adjust the --do-srcu-lockdep 
> > > > > testing
> > > > >  in tools/testing/selftests/rcutorture/bin/torture.sh?
> > > > >
> > > > >  Thanks, Paul. Good point.
> > > > >
> > > > >  I’ll check the current `--do-srcu-lockdep` coverage, including the
> > > > >  case where `synchronize_srcu_atomic()` is called from a normal RCU
> > > > >  read-side critical section, and follow up with the necessary torture
> > > > >  testing changes.
> > > > >
> > > > Sounds good!
> > > >
> > > > Perhaps you and Zqiang can work together on this. Co-developed-by,
> > > > for example.
> > >
> > > Hi, Paul and KunWu
> >
> > Hi Zqiang,
> >
> > Thanks for pointing out these cases.
> >
> > >
> > > Should we also consider the following situations ?
> > >
> > >
> > > idx = srcu_read_lock_atomic(srcu)
> > >
> > > by interrupt run hardirq context:
> > >    synchronize_rcu_atomic(srcu)
> > >
> > > srcu_read_unlock_atomic(srcu, idx)
> >
> > For the same-CPU interrupt case, the existing check in
> > synchronize_srcu_atomic() already catches it:
> >
> >     synchronize_srcu_atomic(ssp)
> >       RCU_LOCKDEP_WARN(lockdep_is_held(ssp),
> >         "Illegal synchronize_srcu_atomic() in same-type SRCU ...");
> >
> > lockdep_is_held() resolves to __lock_is_held() (lockdep.c:5612),
> > which checks current->held_locks[]. The interrupt handler runs with
> > the same current, so it sees the SRCU dep_map acquired by
> > srcu_read_lock_atomic().
>
> I am not sure whether or not it is worth checking in lockdep, but the
> general rule is (almost) that if synchronize_rcu_atomic() runs in a given
> context, then srcu_read_lock_atomic() must be invoked from a context
> that is at least as strict.  So if synchronize_srcu_atomic() is invoked
> with interrupts disabled, then all of the SRCU-atomic readers for that
> same srcu_struct structure must also be invoked with interrupts disabled.
>
> For the "almost" part, note that preemption-enabled context is
> treated the same as is preemption-disabled context because both
> synchronize_srcu_atomic() and srcu_read_lock_atomic() disable preemption.
>

Thanks, Paul, for the detailed explanation. The context rule makes
the two cases much clearer.

> > > or:
> > >
> > >
> > >
> > >            CPU0:                                                          
> > >         CPU1:
> > >
> > > idx = srcu_read_lock_atomic(srcu)
> > >
> > > smp_call_function_single(CPU1, som_func, NULL, 1)
> > > to send IPI to CPU1, and sync wait complete.
> > >                                                                    
> > > hardirq context or ide task context:
> > >
> > >                                                                           
> > > some_func()
> > >                                                                           
> > > ->synchronize_rcu_atomic(srcu)
> > >
> > > srcu_read_unlock_atomic(srcu, idx)
> >
> > The cross-CPU case is different. current->held_locks[] is part of
> > struct task_struct (sched.h:1302), so the existing
> > __lock_is_held() check can only see the current task's held locks.
> > It cannot see the SRCU read-side lock held by the task running on
> > another CPU. The same limitation applies to the
> > lock_is_held(&rcu_lock_map) check in synchronize_srcu() at
> > srcutree.c:1665.
>
> The rule stated above also prevents this situation.  The problem in
> the above scenario is that srcu_read_lock_atomic() was invoked with
> interrupts enabled, which means that invoking synchronize_srcu_atomic()
> for that same srcu_struct structure from the interrupts-disabled IPI
> handler is a usage bug.

The cross-CPU case initially looked like a lockdep limitation to me.
Your context rule clarifies that the usage itself is invalid
regardless:
the reader is in a less strict context than synchronize_srcu_atomic().
The rule is a more direct way of catching the problem than trying to
make lockdep see across CPUs.

The preemption detail is also helpful here, since both
srcu_read_lock_atomic() and synchronize_srcu_atomic() disable
preemption.

>
> > > Add WARN_ON(irqs_disabled()) to synchronize_rcu_atomic() ?
> > >
> > > Any thoughts?
> >
> > WARN_ON(irqs_disabled()) wouldn't help with the cross-CPU case:
> > CPU1 could be running in process context with interrupts enabled, so
> > the WARN would not trigger. It would also add a false positive for
> > legitimate hardirq calls. synchronize_srcu_atomic() omits
> > might_sleep() (compare __synchronize_srcu() at srcutree.c:1676)
> > because it is designed to work in contexts where sleeping is not
> > allowed, including hardirq context.
> >
> > Whether a general cross-CPU read-side-hold check is feasible is an
> > open question. It would need to account for the read-side state across
> > CPUs without adding too much overhead to the SRCU read-side fast path.
> >
> > I'm happy to discuss and explore whether there is a reasonable way to
> > handle this cross-CPU case.
>
> If lockdep could check for the rule stated above, that would be quite
> nice.  ;-)

I'll investigate how this context rule could be checked by lockdep.
I'll also check the IRQ-state handling in synchronize_srcu_atomic()
based on Zqiang's suggestion.

Thanks for the clarification and for pointing out the direction here.

Thanks,
Kunwu

>
>                                                         Thanx, Paul
>
> > Thanks,
> > Kunwu
> >
> > >
> > > Thanks
> > > Zqiang
> > >
> > >
> > > >
> > > >  Thanx, Paul
> > > >
> > > > >
> > > > > Thanks,
> > > > >  Kunwu
> > > > >
> > > > >
> > > > >  Thanx, Paul
> > > > >
> > > > >  > ---
> > > > >  > kernel/rcu/srcutree.c | 8 ++++++++
> > > > >  > 1 file changed, 8 insertions(+)
> > > > >  >
> > > > >  > diff --git a/kernel/rcu/srcutree.c b/kernel/rcu/srcutree.c
> > > > >  > index 6a9c432a3bd0..6c729e805fb3 100644
> > > > >  > --- a/kernel/rcu/srcutree.c
> > > > >  > +++ b/kernel/rcu/srcutree.c
> > > > >  > @@ -2123,6 +2123,14 @@ void synchronize_srcu_atomic(struct 
> > > > > srcu_struct *ssp)
> > > > >  > unsigned long rdm0, rdm1;
> > > > >  > unsigned long unlocks0, unlocks1;
> > > > >  >
> > > > >  > + srcu_lock_sync(&ssp->dep_map);
> > > > >  > +
> > > > >  > + RCU_LOCKDEP_WARN(lockdep_is_held(ssp),
> > > > >  > + "Illegal synchronize_srcu_atomic() in same-type SRCU read-side 
> > > > > critical section");
> > > > >  > +
> > > > >  > + if (rcu_scheduler_active == RCU_SCHEDULER_INACTIVE)
> > > > >  > + return;
> > > > >  > +
> > > > >  > // Initialize. Either init_srcu_struct() was invoked or
> > > > >  > // DEFINE_SRCU() or similar was used. Therefore, no allocation
> > > > >  > // will be done here.
> > > > >  > --
> > > > >  > 2.43.0
> > > > >  >
> > > > >
> > > >

Reply via email to