Le Wed, Sep 16, 2026 at 04:35:50PM +0200, Frederic Weisbecker a écrit : > Le Wed, Sep 16, 2026 at 07:26:35AM -0700, Paul E. McKenney a écrit : > > On Wed, Sep 16, 2026 at 02:40:20PM +0200, Frederic Weisbecker wrote: > > > Le Tue, Sep 15, 2026 at 04:56:41PM -0700, Paul E. McKenney a écrit : > > > > > Alternatively the approach could be generalized to vanilla RCU, it > > > > > could be > > > > > possible to define a .text.rcu_no_qs section within which code > > > > > running is > > > > > considered as an RCU reader (with a pause while on the explicit RCU > > > > > tasks > > > > > section). It would be forbidden to voluntary sleep inside > > > > > and to put explicit preemption points (CONFIG_PROVE_RCU could report > > > > > misuses). > > > > > > > > > > Based on IP, RCU could consider those interrupted section as readers. > > > > > This would > > > > > require PREEMPT_RCU though. > > > > > > > > > > And then synchronize_rcu() would do the 1, 2, 4 jobs. > > > > > > > > If I am following correctly (ha!), sleepable BPF programs rule out use > > > > of RCU in this manner. > > > > > > > > But your point is nevertheless valid, in that SRCU could be used. > > > > And because rcu_read_lock_trace() is a thin wrapper around SRCU-fast, we > > > > *might* be able to instead use rcu_read_lock_tasks_trace(), which would > > > > skip the task-struct increment and decrement, saving a few instructions. > > > > Then, instead of waiting for each task's counter to go to zero, instead > > > > just invoke synchronize_rcu_tasks_trace(). > > > > > > > > Which is pretty close to what Josef is proposing, just with the new RCU > > > > Tasks Trace read-side primitives. I think. ;-) > > > > > > > > This assumes that we do not need to flatten partially overlapping RCU > > > > Tasks Trace readers into one big reader. > > > > > > > > Or am I missing something here? > > > > > > Yes I think that's what Josef does in this patchset. The problem is about > > > handling the few instructions: > > > > > > 1) between the begining of the trampoline and the call to > > > rcu_read_lock_trace() > > > > > > 2) between the call to rcu_read_unlock_trace() and the end of the > > > trampoline > > > > > > So what I'm proposing is to make those two parts implicit RCU read lock > > > sections. > > > > > > So the whole trampoline would be .text.rcu_no_qs: > > > > > > .text.rcu_no_qs trampoline: > > > > > > __________________________________________________________________________________________ > > > |Few instructions 1 | rcu_read_lock_trace() .... rcu_read_unlock_trace | > > > Few instructions 2| > > > > > > ___________________________________________________________________________________________ > > > > > > Then when a tick fires, rcu_flavor_sched_clock_irq() discards the > > > interrupted > > > code as QS if the IP was within .text.rcu_no_qs _unless_ it is in the > > > rcu_read_lock_trace. Both are easy and quick to verify. > > > > > > Also preempt_schedule_irq() would make sure to verify the same condition > > > and > > > enqueue the task as a GP blocker if preempting inside "Few instructions 1" > > > or "Few instructions 2". > > > > Ah, OK, I might be following now. ;-) > > > > We also need both versions of rcu_exp_handler() to check the IP as well, > > given that sooner or later someone is going to want trampoline removal > > to go faster. Or am I still missing a turn in here somewhere? > > Yes indeed, missed the exp part!
What remains to handle also is non-preemptible RCU because if the task is preempted by an IRQ while in the .text.rcu_no_qs, we may still need to keep track of that somewhere. -- Frederic Weisbecker SUSE Labs
