Thank you for your contribution! Sashiko AI review found 1 potential issue(s) 
to consider:
- [High] sched: capacity mismatch in load balance due to restricted env->cpus
--

commit 614fa1366931a4c3a90955d4b7c85dd2e0e3bb49
Author: Shrikanth Hegde <[email protected]>
Subject: sched/fair: Load balance only among preferred CPUs

This commit restricts load balancing to preferred CPUs to avoid conflicting
with the tick's push task mechanism. It replaces cpu_active_mask with
cpu_preferred_mask during load balance CPU mask initialization and bails
out early in newidle balancing if the CPU is not preferred.

> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 03206e15e6fe4..1c687c3c70f11 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -13467,7 +13467,7 @@ static int sched_balance_rq(int this_cpu, struct rq 
> *this_rq,
>       };
>       bool need_unlock = false;
>  
> -     cpumask_and(cpus, sched_domain_span(sd), cpu_active_mask);
> +     cpumask_and(cpus, sched_domain_span(sd), cpu_preferred_mask);

[Severity: High]
Will this change create a capacity mismatch in the load balancer?

When env->cpus is restricted to cpu_preferred_mask here, update_sg_lb_stats()
will only aggregate group load for those preferred CPUs:

kernel/sched/fair.c:update_sg_lb_stats() {
    ...
    for_each_cpu_and(i, sched_group_span(group), env->cpus) {
        struct rq *rq = cpu_rq(i);
        unsigned long load = cpu_load(rq);

        sgs->group_load += load;
    ...
}

However, the group's capacity is set using the full capacity of all active
CPUs:

kernel/sched/fair.c:update_sg_lb_stats() {
    ...
    sgs->group_capacity = group->sgc->capacity;
    ...
}

Could the mismatch between the deflated group_load and the full group_capacity
artificially lower sgs->avg_load?

This phantom spare capacity might cause the load balancer to aggressively pull
tasks from other groups into the preferred CPUs of the destination group,
overloading them.

Can this also prevent tasks on non-preferred CPUs from ever being load
balanced?

kernel/sched/fair.c:sched_balance_find_src_rq() {
    ...
    for_each_cpu_and(i, sched_group_span(group), env->cpus) {
    ...
}

Because this loop uses env->cpus, the load balancer ignores non-preferred
CPUs entirely, even if they are overloaded.

[ ... ]

-- 
Sashiko AI review ยท 
https://sashiko.dev/#/patchset/[email protected]?part=7

Reply via email to