Re: [PATCH 2/4] sched/cache: Add PR_SCHED_CACHE prctl for per-mm control

From: Tim Chen

Date: Wed Jul 22 2026 - 16:57:07 EST


On Wed, 2026-07-22 at 18:13 +0800, Chen, Yu C wrote:
> Hi Yangyu,
>
> On 7/22/2026 5:10 PM, Yangyu Chen wrote:
> > Cache aware scheduling is currently controlled only through global
> > debugfs knobs, but the right aggressiveness is workload and platform
> > specific. A multi-threaded Verilator run is one example: its RSS is
> > large while only a small part of it is hot, so an RSS-based footprint
> > estimate should not decide whether it is aggregated; and packing its
> > threads onto the SMT siblings of one LLC beats spreading them across
> > LLCs on some platforms (e.g. AMD EPYC Turin) but not on others (e.g.
> > EPYC Milan). Such choices cannot be made globally for the whole
> > machine. Add a prctl interface to override the knobs per process
> > (per mm_struct):
> >
> > prctl(PR_SCHED_CACHE, PR_SCHED_CACHE_SET, attr, value, 0);
> > prctl(PR_SCHED_CACHE, PR_SCHED_CACHE_GET, attr, &value, 0);
> >
> > A single prctl command implements both directions, like
> > PR_RSEQ_SLICE_EXTENSION. The attributes are a per-process enable
> > (effective only while the feature is globally active), the two
> > aggregation tolerances, the overaggr percentage (applied where a
> > task's own migration is admitted; group level statistics span many
> > processes and keep using the global value), and an inherit mask
> > selecting which attributes an mm created by execve() keeps. fork()
> > always inherits everything, and the mask itself lives on the
> > task_struct so it survives both, which lets a numactl-like launcher
> > configure a workload and exec it.
> >
> > The overrides live in mm->sc_stat with -1 meaning "follow the global
> > default"; GET stores the raw value through an int pointer so this
> > sentinel round-trips without being mistaken for an errno.
> > mm_init_sched() gains the creating task to tell fork (p != current)
> > from exec (p == current) apart. A disabled mm has its preferred LLC
> > invalidated at the existing invalidation points, so all group-level
> > statistics self-neutralize.
> >
> > Also sync the tools/perf/trace/beauty copy of prctl.h.
> >
> > Assisted-by: Claude:claude-fable-5
> > Signed-off-by: Yangyu Chen <cyy@xxxxxxxxxxxx>
>
> [ ... ]
>
> > +static int sched_cache_set_attr(unsigned long attr, unsigned long val)
> > +{
> > + struct mm_struct *mm = current->mm;
>
> As preparation work, should we first decouple sc_stat from
> mm_struct and tie this stat to per-task task_struct? In this
> way, we could have per-task cache preference control and extend
> it to tasks/threads/process/cgroup if needed, which looks more
> flexible IMO. We have a proposal here:
> https://github.com/chen-yu-surf/linux/commit/bd43a0b6dd189d5091fb88630208cb7bf67b3165.patch
>
> which introduces a pointer in task_struct:
> struct sched_cache_group __rcu  *sched_cache_grp;
>
> > + bool def = sched_cache_val_default(val);
> > + int ival = def ? -1 : (int)val;
> > +
> > + switch (attr) {
> > + case PR_SCHED_CACHE_ENABLE:
> > + if (!def && val > 1)
> > + return -EINVAL;
> > + WRITE_ONCE(mm->sc_stat.user_enabled, ival);
> > + /*
> > + * Drop the preferred LLC hint on any change: a process
> > + * that became disabled must stop being honored right
> > + * away, and one that became enabled re-establishes the
> > + * hint within an epoch anyway. This is best effort: an
> > + * in-flight task_cache_work() scan re-checks the enable
> > + * before publishing a new preference, and a lost race
> > + * is corrected at the next tick.
> > + */
> > + WRITE_ONCE(mm->sc_stat.cpu, -1);
> > + break;
>
> After we switching from per mm_struct to per task control, we could provide
> fine-gain control at task/process/process group granularity(similar to
> core-scheduling)

We are planning to introduce the concept of a sched_group. And tasks in a sched
group can be grouped by mm, or using prctl to explicitly group them together.

We could enhance prctl to introduce per sched_group parameters like aggr_tolerance*
if it makes sense.

Tim

>
> int prctl(PR_SCHED_CACHE, unsigned long subop, pid_t pid,
> unsigned long cookie, unsigned long type);
> pid argument: the PID of the target task. 0 means "the calling task."
> pid_type : PIDTYPE_PID targets the single thread,
> PIDTYPE_TGID the whole thread group and PIDTYPE_PGID the process
> group of the target task.
>
> And the proposal is here:
> https://github.com/chen-yu-surf/linux/commit/17718b7cef1d03948e9fd3bcd0b5a49aba7aae2d.patch
>
> thanks,
> Chenyu