Thank you for your contribution! Sashiko AI review found 1 potential issue(s) 
to consider:
- [High] sched/core: FAIR kthreads cause stopper thread livelock on 
non-preferred CPUs
--

commit a6718dc77694942ae8b63a9eb25ace9c01b3c5ef
Author: Shrikanth Hegde <[email protected]>

sched/core: Push current task from non preferred CPU

This patch actively pushes the currently running task off a non-preferred
CPU using a stopper thread, specifically targeting tasks in the FAIR
scheduling class. It adds the npc_push_work_pending flag to avoid queueing
multiple stopper threads.

> diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> --- a/kernel/sched/core.c
> +++ b/kernel/sched/core.c
[ ... ]
> +static int sched_non_preferred_cpu_push_stop(void *arg)
> +{
> +     struct task_struct *p = arg;
> +     struct rq *rq = this_rq();
> +     struct rq_flags rf;
> +     int cpu;
> +
> +     if (cpu_preferred(rq->cpu)) {
> +             scoped_guard(rq_lock_irqsave, rq)
> +                     rq->npc_push_work_pending = false;
> +             put_task_struct(p);
> +             return 0;
> +     }
> +
> +     scoped_guard (raw_spinlock_irq, &p->pi_lock) {
> +             /*
> +              * select_fallback_rq() may acquire the rq lock in case of
> +              * fallback. So call it before grabbing rq lock. If the task
> +              * migrates to another CPU before the rq lock is acquired,
> +              * subsequent validation of task's current rq will help to
> +              * safely bail out.
> +              */
> +             cpu = select_fallback_rq(rq->cpu, p);
[ ... ]
> +void sched_push_current_non_preferred_cpu(struct rq *rq)
> +{
> +     struct task_struct *push_task = rq->curr;
> +
> +     scoped_guard(rq_lock, rq) {
> +             /* Push the task if its explicit affinity allows */
> +             if (!task_can_migrate_to_preferred(push_task, rq->cpu))
> +                     return;

[Severity: High]
Can this result in a livelock with endless stopper thread preemptions when
the task being pushed is a FAIR kthread (such as an unbound kworker)?

When task_can_migrate_to_preferred() returns true for a FAIR kthread,
this function successfully queues a stopper thread. However, when the
stopper thread runs sched_non_preferred_cpu_push_stop() and calls
select_fallback_rq(), it ultimately evaluates is_cpu_allowed(), which
skips the preferred-CPU check for kthreads:

kernel/sched/core.c:is_cpu_allowed() {
    ...
    /* Non kernel threads are not allowed during either online or offline. */
    if (!(p->flags & PF_KTHREAD)) {
        /* Try to use preferred CPU if task's affinity allows */
        if (task_can_migrate_to_preferred(p, cpu))
            return false;
        return cpu_active(cpu);
    }
    ...
}

Because the kthread bypasses the preferred-CPU enforcement,
select_fallback_rq() can pick the first online CPU in the local node,
which might be the exact same non-preferred CPU we are trying to push it
away from.

If that happens, the migration becomes a no-op, npc_push_work_pending gets
cleared, and sched_tick() will just restart the same sequence on the next
tick, wasting CPU cycles in an endless loop.

> +
> +             /* There is already a stopper thread. Don't race with it. */
> +             if (rq->npc_push_work_pending)
> +                     return;
[ ... ]

-- 
Sashiko AI review ยท 
https://sashiko.dev/#/patchset/[email protected]?part=8

Reply via email to