On Mon, Sep 28, 2026 at 04:11 PM [email protected] <[email protected]> 
wrote:
> irq_work_is_busy() cannot see the per-CPU send_signal_work while
> it is being filled: the check only matches after irq_work_queue()
> has claimed the work. An NMI interrupting the fill therefore passes
> it, both callers race for the same irq_work, and the loser's signal
> is silently lost along with its task reference while the queued
> work runs with a mix of both callers' fields.

kprobe, tracepoint and perf_event progs exclude each other on a cpu
via bpf_prog_active, so one of the two progs has to be raw_tp or fentry.
And since commit 87c544108b61 ("bpf: Send signals asynchronously if
!preemptible") this path runs with irqs enabled too, so hard irq
can do the same. Not only NMI.
Pls describe it in the commit log.
Did you reproduce it or was it found by code inspection?

>  struct send_signal_irq_work {
>       struct irq_work irq_work;
> +     /* Covers the fill-to-run span which irq_work_is_busy() cannot see. */
> +     atomic_t claimed;
>       struct task_struct *task;

can work->task be the claim ?
cmpxchg(&work->task, NULL, task) instead of irq_work_is_busy() and
set it back to NULL at the end of do_bpf_send_signal().
Then no need for extra field.

> -             irq_work_queue(&work->irq_work);
> +             if (unlikely(!irq_work_queue(&work->irq_work))) {
> +                     /* Unreachable while the claim is held. */
> +                     put_task_struct(task);
> +                     atomic_set_release(&work->claimed, 0);
> +                     return -EBUSY;
> +             }

Drop this hunk. It's dead code.
irq_work_queue() fails only when IRQ_WORK_PENDING is set.
irq_work_single() clears it before calling do_bpf_send_signal()
and the claim is released at the end of it.
bpf_mmap_unlock_mm() doesn't check it either after
commit fa9dcacdcdf4 ("bpf: Fix mmap_lock leak in irq_work path").

Pls tag the respin as [PATCH v2 bpf-next].

pw-bot: cr

Reply via email to