On Wed, 2026-07-22 at 18:13 +0800, Chen, Yu C wrote: > Hi Yangyu, > > On 7/22/2026 5:10 PM, Yangyu Chen wrote: > > Cache aware scheduling is currently controlled only through global > > debugfs knobs, but the right aggressiveness is workload and platform > > specific. A multi-threaded Verilator run is one example: its RSS is > > large while only a small part of it is hot, so an RSS-based footprint > > estimate should not decide whether it is aggregated; and packing its > > threads onto the SMT siblings of one LLC beats spreading them across > > LLCs on some platforms (e.g. AMD EPYC Turin) but not on others (e.g. > > EPYC Milan). Such choices cannot be made globally for the whole > > machine. Add a prctl interface to override the knobs per process > > (per mm_struct): > > > > prctl(PR_SCHED_CACHE, PR_SCHED_CACHE_SET, attr, value, 0); > > prctl(PR_SCHED_CACHE, PR_SCHED_CACHE_GET, attr, &value, 0); > > > > A single prctl command implements both directions, like > > PR_RSEQ_SLICE_EXTENSION. The attributes are a per-process enable > > (effective only while the feature is globally active), the two > > aggregation tolerances, the overaggr percentage (applied where a > > task's own migration is admitted; group level statistics span many > > processes and keep using the global value), and an inherit mask > > selecting which attributes an mm created by execve() keeps. fork() > > always inherits everything, and the mask itself lives on the > > task_struct so it survives both, which lets a numactl-like launcher > > configure a workload and exec it. > > > > The overrides live in mm->sc_stat with -1 meaning "follow the global > > default"; GET stores the raw value through an int pointer so this > > sentinel round-trips without being mistaken for an errno. > > mm_init_sched() gains the creating task to tell fork (p != current) > > from exec (p == current) apart. A disabled mm has its preferred LLC > > invalidated at the existing invalidation points, so all group-level > > statistics self-neutralize. > > > > Also sync the tools/perf/trace/beauty copy of prctl.h. > > > > Assisted-by: Claude:claude-fable-5 > > Signed-off-by: Yangyu Chen <[email protected]> > > [ ... ] > > > +static int sched_cache_set_attr(unsigned long attr, unsigned long val) > > +{ > > + struct mm_struct *mm = current->mm; > > As preparation work, should we first decouple sc_stat from > mm_struct and tie this stat to per-task task_struct? In this > way, we could have per-task cache preference control and extend > it to tasks/threads/process/cgroup if needed, which looks more > flexible IMO. We have a proposal here: > https://github.com/chen-yu-surf/linux/commit/bd43a0b6dd189d5091fb88630208cb7bf67b3165.patch > > which introduces a pointer in task_struct: > struct sched_cache_group __rcu *sched_cache_grp; > > > + bool def = sched_cache_val_default(val); > > + int ival = def ? -1 : (int)val; > > + > > + switch (attr) { > > + case PR_SCHED_CACHE_ENABLE: > > + if (!def && val > 1) > > + return -EINVAL; > > + WRITE_ONCE(mm->sc_stat.user_enabled, ival); > > + /* > > + * Drop the preferred LLC hint on any change: a process > > + * that became disabled must stop being honored right > > + * away, and one that became enabled re-establishes the > > + * hint within an epoch anyway. This is best effort: an > > + * in-flight task_cache_work() scan re-checks the enable > > + * before publishing a new preference, and a lost race > > + * is corrected at the next tick. > > + */ > > + WRITE_ONCE(mm->sc_stat.cpu, -1); > > + break; > > After we switching from per mm_struct to per task control, we could provide > fine-gain control at task/process/process group granularity(similar to > core-scheduling)
We are planning to introduce the concept of a sched_group. And tasks in a sched group can be grouped by mm, or using prctl to explicitly group them together. We could enhance prctl to introduce per sched_group parameters like aggr_tolerance* if it makes sense. Tim > > int prctl(PR_SCHED_CACHE, unsigned long subop, pid_t pid, > unsigned long cookie, unsigned long type); > pid argument: the PID of the target task. 0 means "the calling task." > pid_type : PIDTYPE_PID targets the single thread, > PIDTYPE_TGID the whole thread group and PIDTYPE_PGID the process > group of the target task. > > And the proposal is here: > https://github.com/chen-yu-surf/linux/commit/17718b7cef1d03948e9fd3bcd0b5a49aba7aae2d.patch > > thanks, > Chenyu

