> diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c
> index 1dae1c9b53fe..64583fb1f380 100644
> --- a/kernel/bpf/btf.c
> +++ b/kernel/bpf/btf.c
[ ... ]
> @@ -9165,8 +9175,14 @@ static int btf_module_sysfs_add(struct btf_module
> *btf_mod, const char *name,
> return 0;
> }
>
> +/*
> + * Called with btf_module_mutex NOT held: removing the sysfs file waits for
> + * readers to leave, and a deferred reader may need the mutex to get there.
> + */
> static void btf_module_free(struct btf_module *btf_mod)
> {
> + WRITE_ONCE(btf_mod->gone, true);
> + wake_up_all(&btf_module_wq);
> if (btf_mod->sysfs_attr)
> sysfs_remove_bin_file(btf_kobj, btf_mod->sysfs_attr);
[ ... ]
> @@ -9221,15 +9237,57 @@ static int btf_vmlinux_module_coming(struct module
> *mod)
> return 0;
> }
>
> +/*
> + * sysfs reader for a module kept aside with a .BTF.base section: its .BTF is
> + * split against the distilled base and only becomes valid split BTF against
> + * the vmlinux BTF once relocated, which rewrites the buffer in place. So
> + * first make sure the vmlinux BTF is loaded (which parses and relocates the
> + * kept modules), then wait until this module's BTF is published. The size
> + * does not change: relocation only rewrites ids and string offsets.
> + */
> +static bool btf_module_published(struct btf_module *btf_mod)
> +{
> + /* Pairs with the smp_store_release() of @ready after btf_mod->btf is
> set */
> + return smp_load_acquire(&btf_mod->ready);
> +}
> +
> +static ssize_t btf_module_sysfs_read_deferred(struct file *filp, struct
> kobject *kobj,
> + const struct bin_attribute *attr,
> + char *buf, loff_t off, size_t
> count)
> +{
> + struct btf_module *btf_mod = attr->private;
> + int err;
> +
> + if (IS_ERR_OR_NULL(bpf_get_btf_vmlinux()))
> + return -ENODEV;
[ ... ]
> @@ -9393,8 +9454,10 @@ static int btf_module_notify(struct notifier_block
> *nb, unsigned long op,
> if (btf_mod->btf)
> btf_free_id(btf_mod->btf);
> list_del(&btf_mod->list);
> + mutex_unlock(&btf_module_mutex);
> + /* off the list, nobody else can find it now */
> btf_module_free(btf_mod);
> - break;
> + goto out;
> }
> mutex_unlock(&btf_module_mutex);
> break;
Can btf_module_notify() deadlock here during MODULE_STATE_GOING?
The MODULE_STATE_GOING notifier runs from
blocking_notifier_call_chain(&module_notify_list, ...), which holds
module_notify_list.rwsem for read. This patch moves btf_module_free()
outside btf_module_mutex, and btf_module_free() calls
sysfs_remove_bin_file(), which waits uninterruptibly in kernfs_drain()
until every active reader has returned.
This patch made that reader btf_module_sysfs_read_deferred(), which
starts with:
if (IS_ERR_OR_NULL(bpf_get_btf_vmlinux()))
return -ENODEV;
If the vmlinux BTF has not been loaded yet, bpf_get_btf_vmlinux() calls
btf_vmlinux_data(&size, true), which calls request_module("btf_vmlinux").
That waits (UMH_WAIT_PROC | UMH_KILLABLE) for modprobe.
modprobe's load_module() calls prepare_coming_module(), which calls
blocking_notifier_call_chain_robust(&module_notify_list,
MODULE_STATE_COMING, ...) and needs down_read() on the same rwsem.
If a writer is queued on that rwsem, new readers block once the handoff
bit is set. Writers come from register_module_notifier() or
unregister_module_notifier(), called for example from kunit_init() when
kunit is modular, or from a module exit that unregisters its notifier.
The resulting cycle is:
rmmod X (GOING, holds rwsem for read)
-> kernfs_drain waits for the reader of /sys/kernel/btf/X
reader of X
-> request_module("btf_vmlinux") waits for modprobe
modprobe btf_vmlinux
-> down_read(module_notify_list.rwsem) waits for the queued writer
writer
-> down_write waits for rmmod X to release its read side
Everything stays in D state until the reader gets SIGKILL. Meanwhile
every other module load or unload also stalls on the notifier chain.
Before this patch, a deferred .BTF.base module had no sysfs file until
btf_parse_deferred_modules() had relocated it. No reader could block, so
the drain in GOING could not wait on a module load.
The new code conflicts with the rationale stated in btf_module_defer():
"The vmlinux BTF ... must not be loaded from the module notifier (that
would nest a module load into a module load)". The GOING notifier now
waits on exactly that nested load, through the reader.
Possible fixes: do not trigger the vmlinux BTF load from this reader
(serve -ENODEV/-EAGAIN until the BTF is published, or only wait when
btf_vmlinux is already set); or remove the sysfs file in a way that does
not wait under the notifier chain (for example,
kernfs_break_active_protection() in the reader around the
request_module()/wait, with the btf_mod lifetime guarded separately).
Also, this patch changes the lock ordering that affects a comment
elsewhere in the file. The comment for btf_vmlinux_regs_mutex justifies
the separate mutex by saying that purge_cand_cache() takes
cand_cache_mutex under btf_module_mutex. That was only true through
btf_module_free() on the GOING path, which this patch moves outside
btf_module_mutex (as shown above).
The remaining purge_cand_cache() callers, in btf_module_notify() COMING
and at the end of btf_parse_deferred_modules(), already run without
btf_module_mutex. No btf_module_mutex -> cand_cache_mutex chain is left,
so the stated reason for the ordering rule is gone.
The ordering that matters now is kernfs active reference ->
btf_vmlinux_lock/btf_module_mutex, which comes from the new deferred
sysfs reader. Should that comment be updated to reflect the current lock
ordering rationale?
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36198628965