Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider: - [High] sched/core: FAIR kthreads cause stopper thread livelock on non-preferred CPUs --
commit a6718dc77694942ae8b63a9eb25ace9c01b3c5ef Author: Shrikanth Hegde <[email protected]> sched/core: Push current task from non preferred CPU This patch actively pushes the currently running task off a non-preferred CPU using a stopper thread, specifically targeting tasks in the FAIR scheduling class. It adds the npc_push_work_pending flag to avoid queueing multiple stopper threads. > diff --git a/kernel/sched/core.c b/kernel/sched/core.c > --- a/kernel/sched/core.c > +++ b/kernel/sched/core.c [ ... ] > +static int sched_non_preferred_cpu_push_stop(void *arg) > +{ > + struct task_struct *p = arg; > + struct rq *rq = this_rq(); > + struct rq_flags rf; > + int cpu; > + > + if (cpu_preferred(rq->cpu)) { > + scoped_guard(rq_lock_irqsave, rq) > + rq->npc_push_work_pending = false; > + put_task_struct(p); > + return 0; > + } > + > + scoped_guard (raw_spinlock_irq, &p->pi_lock) { > + /* > + * select_fallback_rq() may acquire the rq lock in case of > + * fallback. So call it before grabbing rq lock. If the task > + * migrates to another CPU before the rq lock is acquired, > + * subsequent validation of task's current rq will help to > + * safely bail out. > + */ > + cpu = select_fallback_rq(rq->cpu, p); [ ... ] > +void sched_push_current_non_preferred_cpu(struct rq *rq) > +{ > + struct task_struct *push_task = rq->curr; > + > + scoped_guard(rq_lock, rq) { > + /* Push the task if its explicit affinity allows */ > + if (!task_can_migrate_to_preferred(push_task, rq->cpu)) > + return; [Severity: High] Can this result in a livelock with endless stopper thread preemptions when the task being pushed is a FAIR kthread (such as an unbound kworker)? When task_can_migrate_to_preferred() returns true for a FAIR kthread, this function successfully queues a stopper thread. However, when the stopper thread runs sched_non_preferred_cpu_push_stop() and calls select_fallback_rq(), it ultimately evaluates is_cpu_allowed(), which skips the preferred-CPU check for kthreads: kernel/sched/core.c:is_cpu_allowed() { ... /* Non kernel threads are not allowed during either online or offline. */ if (!(p->flags & PF_KTHREAD)) { /* Try to use preferred CPU if task's affinity allows */ if (task_can_migrate_to_preferred(p, cpu)) return false; return cpu_active(cpu); } ... } Because the kthread bypasses the preferred-CPU enforcement, select_fallback_rq() can pick the first online CPU in the local node, which might be the exact same non-preferred CPU we are trying to push it away from. If that happens, the migration becomes a no-op, npc_push_work_pending gets cleared, and sched_tick() will just restart the same sequence on the next tick, wasting CPU cycles in an endless loop. > + > + /* There is already a stopper thread. Don't race with it. */ > + if (rq->npc_push_work_pending) > + return; [ ... ] -- Sashiko AI review ยท https://sashiko.dev/#/patchset/[email protected]?part=8
