> diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c
> index 1dae1c9b53fe..64583fb1f380 100644
> --- a/kernel/bpf/btf.c
> +++ b/kernel/bpf/btf.c

[ ... ]

> @@ -9165,8 +9175,14 @@ static int btf_module_sysfs_add(struct btf_module 
> *btf_mod, const char *name,
>       return 0;
>  }
>
> +/*
> + * Called with btf_module_mutex NOT held: removing the sysfs file waits for
> + * readers to leave, and a deferred reader may need the mutex to get there.
> + */
>  static void btf_module_free(struct btf_module *btf_mod)
>  {
> +     WRITE_ONCE(btf_mod->gone, true);
> +     wake_up_all(&btf_module_wq);
>       if (btf_mod->sysfs_attr)
>               sysfs_remove_bin_file(btf_kobj, btf_mod->sysfs_attr);

[ ... ]

> @@ -9221,15 +9237,57 @@ static int btf_vmlinux_module_coming(struct module 
> *mod)
>       return 0;
>  }
>
> +/*
> + * sysfs reader for a module kept aside with a .BTF.base section: its .BTF is
> + * split against the distilled base and only becomes valid split BTF against
> + * the vmlinux BTF once relocated, which rewrites the buffer in place.  So
> + * first make sure the vmlinux BTF is loaded (which parses and relocates the
> + * kept modules), then wait until this module's BTF is published.  The size
> + * does not change: relocation only rewrites ids and string offsets.
> + */
> +static bool btf_module_published(struct btf_module *btf_mod)
> +{
> +     /* Pairs with the smp_store_release() of @ready after btf_mod->btf is 
> set */
> +     return smp_load_acquire(&btf_mod->ready);
> +}
> +
> +static ssize_t btf_module_sysfs_read_deferred(struct file *filp, struct 
> kobject *kobj,
> +                                           const struct bin_attribute *attr,
> +                                           char *buf, loff_t off, size_t 
> count)
> +{
> +     struct btf_module *btf_mod = attr->private;
> +     int err;
> +
> +     if (IS_ERR_OR_NULL(bpf_get_btf_vmlinux()))
> +             return -ENODEV;

[ ... ]

> @@ -9393,8 +9454,10 @@ static int btf_module_notify(struct notifier_block 
> *nb, unsigned long op,
>                       if (btf_mod->btf)
>                               btf_free_id(btf_mod->btf);
>                       list_del(&btf_mod->list);
> +                     mutex_unlock(&btf_module_mutex);
> +                     /* off the list, nobody else can find it now */
>                       btf_module_free(btf_mod);
> -                     break;
> +                     goto out;
>               }
>               mutex_unlock(&btf_module_mutex);
>               break;

Can btf_module_notify() deadlock here during MODULE_STATE_GOING?

The MODULE_STATE_GOING notifier runs from
blocking_notifier_call_chain(&module_notify_list, ...), which holds
module_notify_list.rwsem for read. This patch moves btf_module_free()
outside btf_module_mutex, and btf_module_free() calls
sysfs_remove_bin_file(), which waits uninterruptibly in kernfs_drain()
until every active reader has returned.

This patch made that reader btf_module_sysfs_read_deferred(), which
starts with:

    if (IS_ERR_OR_NULL(bpf_get_btf_vmlinux()))
            return -ENODEV;

If the vmlinux BTF has not been loaded yet, bpf_get_btf_vmlinux() calls
btf_vmlinux_data(&size, true), which calls request_module("btf_vmlinux").
That waits (UMH_WAIT_PROC | UMH_KILLABLE) for modprobe.

modprobe's load_module() calls prepare_coming_module(), which calls
blocking_notifier_call_chain_robust(&module_notify_list,
MODULE_STATE_COMING, ...) and needs down_read() on the same rwsem.

If a writer is queued on that rwsem, new readers block once the handoff
bit is set. Writers come from register_module_notifier() or
unregister_module_notifier(), called for example from kunit_init() when
kunit is modular, or from a module exit that unregisters its notifier.

The resulting cycle is:

  rmmod X (GOING, holds rwsem for read)
    -> kernfs_drain waits for the reader of /sys/kernel/btf/X

  reader of X
    -> request_module("btf_vmlinux") waits for modprobe

  modprobe btf_vmlinux
    -> down_read(module_notify_list.rwsem) waits for the queued writer

  writer
    -> down_write waits for rmmod X to release its read side

Everything stays in D state until the reader gets SIGKILL. Meanwhile
every other module load or unload also stalls on the notifier chain.

Before this patch, a deferred .BTF.base module had no sysfs file until
btf_parse_deferred_modules() had relocated it. No reader could block, so
the drain in GOING could not wait on a module load.

The new code conflicts with the rationale stated in btf_module_defer():
"The vmlinux BTF ... must not be loaded from the module notifier (that
would nest a module load into a module load)". The GOING notifier now
waits on exactly that nested load, through the reader.

Possible fixes: do not trigger the vmlinux BTF load from this reader
(serve -ENODEV/-EAGAIN until the BTF is published, or only wait when
btf_vmlinux is already set); or remove the sysfs file in a way that does
not wait under the notifier chain (for example,
kernfs_break_active_protection() in the reader around the
request_module()/wait, with the btf_mod lifetime guarded separately).

Also, this patch changes the lock ordering that affects a comment
elsewhere in the file. The comment for btf_vmlinux_regs_mutex justifies
the separate mutex by saying that purge_cand_cache() takes
cand_cache_mutex under btf_module_mutex. That was only true through
btf_module_free() on the GOING path, which this patch moves outside
btf_module_mutex (as shown above).

The remaining purge_cand_cache() callers, in btf_module_notify() COMING
and at the end of btf_parse_deferred_modules(), already run without
btf_module_mutex. No btf_module_mutex -> cand_cache_mutex chain is left,
so the stated reason for the ordering rule is gone.

The ordering that matters now is kernfs active reference ->
btf_vmlinux_lock/btf_module_mutex, which comes from the new deferred
sysfs reader. Should that comment be updated to reflect the current lock
ordering rationale?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36198628965

Reply via email to