On 30 September 2026 21:29:51 BST, "Paul E. McKenney" <[email protected]> wrote: >Adding Sunho Park on CC for his fix: > >b7e310c3e43e ("srcu: Fix WARN_ON() for rcu_segcblist_n_cbs() in >cleanup_srcu_struct()")
Already done? Aw no > >On Wed, Sep 30, 2026 at 06:45:09PM +0000, syzbot wrote: >> From: Bradley Morgan <[email protected]> >> >> In cleanup_srcu_struct(), commit 78a38cbf6f20 ("srcu: Queue sdp->work >when >> the delay timer is successfully deleted") added logic to queue sdp->work >if >> timer_delete_sync(&sdp->delay_work) successfully canceled a pending >timer >> while callbacks remained pending on sdp->srcu_cblist. However, it >wrapped >> this check in a WARN_ON() under the assumption that callers invoking >> srcu_barrier() prior to cleanup_srcu_struct() would prevent the warning >> from triggering. >> >> This warning can be spuriously triggered during valid teardown paths >where >> srcu_barrier() is properly invoked before cleanup_srcu_struct(), such as >> when releasing blk-mq tag sets: >> >> WARNING: kernel/rcu/srcutree.c:707 at cleanup_srcu_struct+0x3d6/0x8b0 >> kernel/rcu/srcutree.c:706 >> Call Trace: >> <TASK> >> blk_mq_free_tag_set+0x617/0x790 block/blk-mq.c:4976 >> scsi_mq_free_tags+0x16/0x30 drivers/scsi/scsi_lib.c:2167 >> scsi_remove_host+0x243/0x730 drivers/scsi/hosts.c:193 >> uas_disconnect+0x135/0x3e0 drivers/usb/storage/uas.c:1236 >> usb_unbind_interface+0x295/0x9f0 drivers/usb/core/driver.c:461 >> device_release_driver_internal+0x4f5/0x880 drivers/base/dd.c:1372 >> bus_remove_device+0x444/0x560 drivers/base/bus.c:664 >> device_del+0x524/0x8f0 drivers/base/core.c:3965 >> usb_disconnect+0x346/0x9a0 drivers/usb/core/hub.c:2350 >> hub_event+0x1bbb/0x4d30 drivers/usb/core/hub.c:5966 >> process_scheduled_works+0xc3d/0x1630 kernel/workqueue.c:3479 >> worker_thread+0xa47/0xfb0 kernel/workqueue.c:3560 >> kthread+0x38b/0x480 kernel/kthread.c:436 >> ret_from_fork+0x514/0xb70 arch/x86/kernel/process.c:158 >> ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245 >> </TASK> >> >> This false positive occurs due to a race between callback invocation and >> the teardown thread. First, sdp->delay_work can legitimately remain >armed: >> when an SRCU grace period completes, srcu_gp_end() arms sdp->delay_work >> with a delay, and if callbacks are later scheduled with zero delay via >> srcu_schedule_cbs_sdp(sdp, 0), queue_work_on() queues sdp->work directly >> without deleting sdp->delay_work. Second, in srcu_invoke_callbacks(), >> callbacks (including the one queued by srcu_barrier()) are extracted >into a >> local list and executed, but sdp->srcu_cblist length is only decremented >> via rcu_segcblist_add_len(&sdp->srcu_cblist, -len) after all callbacks >> finish executing. When srcu_barrier_cb() executes, it wakes the waiting >> srcu_barrier() thread, which proceeds immediately to >cleanup_srcu_struct(). >> At that point, the worker thread is still executing callbacks and has >not >> yet decremented the callback count, so >timer_delete_sync(&sdp->delay_work) >> returns 1 and rcu_segcblist_n_cbs(&sdp->srcu_cblist) is non-zero, firing >> the WARN_ON(). Immediately thereafter, flush_work(&sdp->work) waits for >the >> worker to finish, and the subsequent check on rcu_segcblist_n_cbs() >> correctly sees no remaining callbacks. >> >> Because WARN_ON() must not be used for conditions that can legitimately >> happen, and pr_err() should be used instead if an actual error needs to >be >> reported (which is not applicable here as this is normal recovery >> behavior), remove the WARN_ON() check and its accompanying comment. Keep >> the queue_work_on() recovery logic so that flush_work() properly waits >for >> remaining callbacks to complete. Genuine callback leaks remain caught by >> the subsequent authoritative >> WARN_ON(rcu_segcblist_n_cbs(&sdp->srcu_cblist)) check performed after >work >> has been flushed. >> >> Fixes: 78a38cbf6f20 ("srcu: Queue sdp->work when the delay timer is >successfully deleted") >> Assisted-by: Gemini:gemini-3.8-flash Gemini:gemini-3.1-pro-preview >syzbot >> Reported-by: [email protected] >> Closes: https://syzkaller.appspot.com/bug?extid=02b37e31e64ea5cb6d29 >> Link: >https://syzkaller.appspot.com/ai_job?id=5a402d0b-b407-49a6-b0f3-b06197c1398f >> Signed-off-by: Bradley Morgan <[email protected]> >> >> --- >> diff --git a/kernel/rcu/srcutree.c b/kernel/rcu/srcutree.c >> index ed204b3f4..7f30a5587 100644 >> --- a/kernel/rcu/srcutree.c >> +++ b/kernel/rcu/srcutree.c >> @@ -701,10 +701,8 @@ void cleanup_srcu_struct(struct srcu_struct *ssp) >> for_each_possible_cpu(cpu) { >> struct srcu_data *sdp = per_cpu_ptr(ssp->sda, cpu); >> >> - // Call srcu_barrier() before this cleanup_srcu_struct() >> - // to avoid triggering this WARN_ON(). >> - if (WARN_ON(timer_delete_sync(&sdp->delay_work) && >> - rcu_segcblist_n_cbs(&sdp->srcu_cblist)) && >> + if (timer_delete_sync(&sdp->delay_work) && >> + rcu_segcblist_n_cbs(&sdp->srcu_cblist) && >> rcu_cpu_beenfullyonline(sdp->cpu)) >> queue_work_on(sdp->cpu, rcu_gp_wq, &sdp->work); >> flush_work(&sdp->work); >> >> >> base-commit: 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e >> -- >> See https://goo.gle/syzbot-ai-patches for information about AI-generated >patches. >> The person who has signed off on the patch is responsible for >> addressing comments. >> syzbot engineers can be reached at [email protected]. --- Thanks! "I'm not a very positive person" - Linus torvalds

