Re: [PATCH v12 12/25] arm,x86,fs/resctrl: Allocate maximum needed rmid_ptrs[]

From: Reinette Chatre

Date: Thu Sep 24 2026 - 12:06:40 EST


Hi Tony,

On 9/16/26 4:13 PM, Tony Luck wrote:
> File system code allocates the rmid_ptrs[] array once during

(no need to say "array" when using "[]")

> initialization. The number of entries needed in this array is currently

initialization -> "first mount"?

> constant. When changes are made to allow Application Energy Telemetry

constant implies a constant value but I think this actually intends to
say that the value does not change from mount to mount?

> (AET) to run with the pmt_telemetry driver as a module, then number of
> entries needed may change from one mount to the next.

tip preference is to separate the context from the problem description

How about something like this _draft_:

resctrl keeps per-RMID state in rmid_ptrs[]. It is allocated on
first mount, sized for the number of RMIDs available at that
time, and reused by every subsequent mount.

Application Energy Telemetry (AET) requires the pmt_telemetry
driver to be built in. Allowing it to be built as a module means the
number of RMIDs can change from one mount to the next, so a later mount
may need more entries than the first mount allocated. Reallocating per
mount is not possible because the limbo handler continues to access
rmid_ptrs[] after resctrl is unmounted.

Size rmid_ptrs[] for the maximum number of RMIDs the system can ever need,
so that it is large enough for any future mount.

>
> Allocate rmid_ptrs[] with enough entries for any future mount.
>
> Signed-off-by: Tony Luck <tony.luck@xxxxxxxxx>
> ---
> v12:
> New patch. Split out from old patch 12.
> New global pqr_assoc_num_rmid to avoid repeat CPUID calls.
> ---
> include/linux/resctrl.h | 1 +
> arch/x86/kernel/cpu/resctrl/core.c | 27 +++++++++++++++++++++++++++
> drivers/resctrl/mpam_resctrl.c | 9 +++++++++
> fs/resctrl/monitor.c | 2 +-
> 4 files changed, 38 insertions(+), 1 deletion(-)
>
> diff --git a/include/linux/resctrl.h b/include/linux/resctrl.h
> index fbf737e884db..5535bde7b925 100644
> --- a/include/linux/resctrl.h
> +++ b/include/linux/resctrl.h
> @@ -447,6 +447,7 @@ static inline u32 resctrl_get_default_ctrl(struct rdt_resource *r)
> /* The number of closid supported by this resource regardless of CDP */
> u32 resctrl_arch_get_num_closid(struct rdt_resource *r);
> u32 resctrl_arch_system_num_rmid_idx(void);
> +u32 resctrl_arch_system_max_rmid_idx(void);
> int resctrl_arch_update_domains(struct rdt_resource *r, u32 closid);
>
> /**
> diff --git a/arch/x86/kernel/cpu/resctrl/core.c b/arch/x86/kernel/cpu/resctrl/core.c
> index 2e3b9c16cbda..a9109f2bc43e 100644
> --- a/arch/x86/kernel/cpu/resctrl/core.c
> +++ b/arch/x86/kernel/cpu/resctrl/core.c
> @@ -45,6 +45,9 @@ static DEFINE_MUTEX(domain_list_lock);
> */
> DEFINE_PER_CPU(struct resctrl_pqr_state, pqr_state);
>
> +/* Number of RMIDS values that can be written to IA32_PQR_ASSOC.RMID */

"Number of RMIDS values" does not sound right. Also, please use grep friendly
names for registers. For example, "Number of RMIDs supported by MSR_IA32_PQR_ASSOC.RMID"?


> +static u32 pqr_assoc_num_rmid;
> +
> static void mba_wrmsr_intel(struct msr_param *m);
> static void cat_wrmsr(struct msr_param *m);
> static void mba_wrmsr_amd(struct msr_param *m);
> @@ -124,6 +127,28 @@ u32 resctrl_arch_system_num_rmid_idx(void)
> return num_rmids == U32_MAX ? 0 : num_rmids;
> }
>
> +/**
> + * resctrl_arch_system_max_rmid_idx - Largest possible number of RMIDs

To match function name and later quest for "maximum RMID value" perhaps
"Largest possible number of RMIDs" -> "Largest possible RMID index"?

> + *
> + * Return: Maximum possible number of RMIDs used for boot time allocations.

This function goes from "maximum index" to "largest number of RMIDs" and then
comment below goes back to "maximum RMID value" with the caller finally using it
to guide allocation, not used as an index. Pick one usage and stick with it please.

> + */
> +u32 resctrl_arch_system_max_rmid_idx(void)
> +{
> + struct rdt_resource *r = &rdt_resources_all[RDT_RESOURCE_L3].r_resctrl;
> + u32 num_rmid = pqr_assoc_num_rmid;
> +
> + /*
> + * If the system is capable of L3 monitoring the maximum RMID value may
> + * be lower than the system maximum. Either because the L3 monitoring
> + * feature supports fewer RMIDs, or because SNC (Sub-NUMA Cluster)
> + * is enabled and divides RMIDs per cluster.
> + */
> + if (r->mon_capable)
> + num_rmid = r->mon.num_rmid;
> +
> + return num_rmid;
> +}
> +
> struct rdt_resource *resctrl_arch_get_resource(enum resctrl_res_level l)
> {
> if (l >= RDT_NUM_RESOURCES)
> @@ -967,6 +992,8 @@ static __init bool get_rdt_mon_resources(void)
> if (!cpu_feature_enabled(X86_FEATURE_CQM))
> return false;
>
> + pqr_assoc_num_rmid = cpuid_ebx(0xf) + 1;
> +
> /* Any of the L3 monitoring features? */
> if (!cpu_feature_enabled(X86_FEATURE_CQM_LLC))
> return false;
> diff --git a/drivers/resctrl/mpam_resctrl.c b/drivers/resctrl/mpam_resctrl.c
> index 0db62dd2a71c..0ffa25199f74 100644
> --- a/drivers/resctrl/mpam_resctrl.c
> +++ b/drivers/resctrl/mpam_resctrl.c
> @@ -252,6 +252,15 @@ u32 resctrl_arch_system_num_rmid_idx(void)
> return (mpam_pmg_max + 1) * (mpam_partid_max + 1);
> }
>
> +/*
> + * File system calls this for one-time allocation of structures
> + * during initialization. Return the largest possible value.

Perhaps just drop "during initialization" since it is not accurate
(allocation is on first mount) and architecture need not be concerned
about this detail.

> + */
> +u32 resctrl_arch_system_max_rmid_idx(void)
> +{
> + return resctrl_arch_system_num_rmid_idx();
> +}
> +
> u32 resctrl_arch_rmid_idx_encode(u32 closid, u32 rmid)
> {
> return closid * (mpam_pmg_max + 1) + rmid;
> diff --git a/fs/resctrl/monitor.c b/fs/resctrl/monitor.c
> index 2a28fe04284b..e8775e08aa18 100644
> --- a/fs/resctrl/monitor.c
> +++ b/fs/resctrl/monitor.c
> @@ -978,7 +978,7 @@ int setup_rmid_lru_list(void)
> if (rmid_ptrs)
> return 0;
>
> - idx_limit = resctrl_arch_system_num_rmid_idx();
> + idx_limit = resctrl_arch_system_max_rmid_idx();
> rmid_ptrs = kzalloc_objs(struct rmid_entry, idx_limit);
> if (!rmid_ptrs)
> return -ENOMEM;

This split does not look right. It intentionally allocates the maximum as
the changelog describes but then it also uses this maximum to guide how many
RMIDs are added to the free list that should still be guided by the number of
RMIDs available during this mount as obtained from resctrl_arch_system_num_rmid_idx(),
no? This is what is done before this change ... and then changed back later.
Even more, the comment that accompanies this change ("Allocate the largest number
of RMIDs that this system will ever need.") only appears in later patch?

Reinette