On 7/25/26 1:30 AM, Yury Norov wrote:
On Fri, Jul 24, 2026 at 07:37:23PM +0530, Shrikanth Hegde wrote:
Provide cpu_preferred_mask infrastructure. Define get/set macros
which could be used to get/set CPU state as preferred.

PREFERRED_CPU config will be selected by the driver which handles
steal time values. It is going to set/clear preferred CPU state.
This driver will be called steal_governor and it is introduced in
subsequent patches. It periodically samples the steal time and
decides on preferred CPU state.

A CPU is set to preferred when it becomes active. Later it may be
marked as non-preferred depending on steal time values with
steal_governor being enabled.

Always maintain design construct of preferred is subset of active.
i.e. preferred ⊆ active ⊆ online ⊆ present ⊆ possible

With PREFERRED_CPU=n, ensure set_cpu_preferred is a nop and get
method returns the active state in that case.

Signed-off-by: Shrikanth Hegde <[email protected]>
---
  include/linux/cpumask.h | 24 ++++++++++++++++++++++++
  kernel/Kconfig.preempt  |  4 ++++
  kernel/cpu.c            |  6 ++++++
  kernel/sched/core.c     |  5 +++++
  4 files changed, 39 insertions(+)

diff --git a/include/linux/cpumask.h b/include/linux/cpumask.h
index d3cda0544954..34d08a3d80e1 100644
--- a/include/linux/cpumask.h
+++ b/include/linux/cpumask.h
@@ -122,12 +122,20 @@ extern struct cpumask __cpu_enabled_mask;
  extern struct cpumask __cpu_present_mask;
  extern struct cpumask __cpu_active_mask;
  extern struct cpumask __cpu_dying_mask;
+
+#ifdef CONFIG_PREFERRED_CPU
+extern struct cpumask __cpu_preferred_mask;
+#else
+#define __cpu_preferred_mask __cpu_active_mask
+#endif
+
  #define cpu_possible_mask ((const struct cpumask *)&__cpu_possible_mask)
  #define cpu_online_mask   ((const struct cpumask *)&__cpu_online_mask)
  #define cpu_enabled_mask   ((const struct cpumask *)&__cpu_enabled_mask)
  #define cpu_present_mask  ((const struct cpumask *)&__cpu_present_mask)
  #define cpu_active_mask   ((const struct cpumask *)&__cpu_active_mask)
  #define cpu_dying_mask    ((const struct cpumask *)&__cpu_dying_mask)
+#define cpu_preferred_mask ((const struct cpumask *)&__cpu_preferred_mask)
extern atomic_t __num_online_cpus;
  extern unsigned int __num_possible_cpus;
@@ -1164,6 +1172,12 @@ void init_cpu_possible(const struct cpumask *src);
  #define set_cpu_active(cpu, active)   assign_cpu((cpu), &__cpu_active_mask, 
(active))
  #define set_cpu_dying(cpu, dying)     assign_cpu((cpu), &__cpu_dying_mask, 
(dying))
+#ifdef CONFIG_PREFERRED_CPU
+#define set_cpu_preferred(cpu, preferred) assign_cpu((cpu), 
&__cpu_preferred_mask, (preferred))
+#else
+#define set_cpu_preferred(cpu, preferred) do { } while (0)
+#endif
+
  void set_cpu_online(unsigned int cpu, bool online);
  void set_cpu_possible(unsigned int cpu, bool possible);
@@ -1258,6 +1272,11 @@ static __always_inline bool cpu_dying(unsigned int cpu)
        return cpumask_test_cpu(cpu, cpu_dying_mask);
  }
+static __always_inline bool cpu_preferred(unsigned int cpu)
+{
+       return cpumask_test_cpu(cpu, cpu_preferred_mask);
+}
+
  #else
#define num_online_cpus() 1U
@@ -1296,6 +1315,11 @@ static __always_inline bool cpu_dying(unsigned int cpu)
        return false;
  }
+static __always_inline bool cpu_preferred(unsigned int cpu)
+{
+       return cpu == 0;
+}
+
  #endif /* NR_CPUS > 1 */
#define cpu_is_offline(cpu) unlikely(!cpu_online(cpu))
diff --git a/kernel/Kconfig.preempt b/kernel/Kconfig.preempt
index 88c594c6d7fc..de789b274ba3 100644
--- a/kernel/Kconfig.preempt
+++ b/kernel/Kconfig.preempt
@@ -192,3 +192,7 @@ config SCHED_CLASS_EXT
          For more information:
            Documentation/scheduler/sched-ext.rst
            https://github.com/sched-ext/scx
+
+config PREFERRED_CPU
+       bool
+       depends on SMP && PARAVIRT
diff --git a/kernel/cpu.c b/kernel/cpu.c
index b3c8553d7bd6..376d297a6292 100644
--- a/kernel/cpu.c
+++ b/kernel/cpu.c
@@ -3103,6 +3103,11 @@ EXPORT_SYMBOL(__cpu_dying_mask);
  atomic_t __num_online_cpus __read_mostly;
  EXPORT_SYMBOL(__num_online_cpus);
+#ifdef CONFIG_PREFERRED_CPU
+struct cpumask __cpu_preferred_mask __read_mostly;
+EXPORT_SYMBOL_GPL(__cpu_preferred_mask);
+#endif
+
  void init_cpu_present(const struct cpumask *src)
  {
        cpumask_copy(&__cpu_present_mask, src);
@@ -3160,6 +3165,7 @@ void __init boot_cpu_init(void)
        /* Mark the boot cpu "present", "online" etc for SMP and UP case */
        set_cpu_online(cpu, true);
        set_cpu_active(cpu, true);
+       set_cpu_preferred(cpu, true);
        set_cpu_present(cpu, true);
        set_cpu_possible(cpu, true);
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index 2e7cde033a31..a45f7c308329 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -8690,6 +8690,9 @@ int sched_cpu_activate(unsigned int cpu)
         */
        sched_set_rq_online(rq, cpu);
+ /* preferred is subset of active and follows its state */
+       set_cpu_preferred(cpu, true);
+
        return 0;
  }
@@ -8703,6 +8706,8 @@ int sched_cpu_deactivate(unsigned int cpu)
        if (ret)
                return ret;
+ set_cpu_preferred(cpu, false);
+

Is it possible that this CPU would be the last preferred CPU in the
system? If so, you'll make the preferred mask empty.


Possible case is, say there are 80 CPUs and all CPUs are part of housekeeping.
driver marked 40-80 as non-preferred and before driver gets a chance to run 
again,
user disabled 0-39. Now preferred mask is empty. if steal time is low, it might 
recover
without check broken in the next sampling, but it stays in between or high, 
then that
check is broken.

I don't think there is any side effect in core mechanism since is_cpu_allowed 
will pass
due to empty preferred mask. In driver, further reduction will not happen. But 
yes, it
will break the design checks.


In v9 you disabled integrity check while the steal time is withing the
threshold, so this condition may stay undetected quite a long.


I think simplest solution is do the design checks always and restore the 
preferred state
if such case happens. I.e drop the optimization that was done in v9 compared to 
v8.

Can you add another integrity check here? If you're going to remove
the last preferred CPU, you need to force-enable some alternative.
Something like:

         if (cpumask_nth(1, cpu_preferred_mask) >= nr_cpu_ids) {
                 new_cpu = cpumask_any_andnot_but(cpu_active_mask, 
cpu_preferred_mask, cpu);
                 if (!WARN_ON(new_cpu >= nr_cpu_ids))
                         set_cpu_preferred(new_cpu);
         }

         set_cpu_preferred(cpu, false);


I think we shouldn't do such change. The design constraints are of driver.
Hotplug mechanism just ensure to set preferred after setting active and clear 
preferred
before clearing the active. That's all.

Driver runs only once in 100ms at the very least and enforcing design checks of 
driver
into core hotplug/scheduler mechanism is not right IMHO. It should be the role 
of driver
to either actively recover or gracefully shut.

That is user triggered edge case, i think simplest solution is gracefully shut 
the driver and
let user to load the driver again. Always run the design checks. I can add this 
corner case details
to the driver change log.

What do you think?

Reply via email to