Re: [PATCH v6 9/9] x86/resctrl: Add MMIO-based LLC occupancy monitoring support

From: Chen, Yu C

Date: Tue Aug 25 2026 - 12:12:04 EST


Hi Reinette,

On 8/20/2026 7:10 AM, Reinette Chatre wrote:
Hi Chenyu,

On 7/25/26 2:23 AM, Chen Yu wrote:
Add erdt_mon_read() to read LLC occupancy via MMIO and use it when
the platform supports ERDT. Register the L3 occupancy event with
ERDT enabled when available, falling back to the MSR-based path
otherwise.

Use the CMRC (Cache Monitoring Registers for CPU Agents Description)
ACPI sub-table to read LLC occupancy counters for each RMID via MMIO
when ERDT is enabled. This CMRC information is stored in the
rdt_hw_l3_mon_domain, which could be accessed directly.

Please write in imperative tone.


OK, let me try to rewrite it(I suppose you were refereeing to:
This CMRC information is stored -> Store the CMRC information)


Currently, the per-domain limbo handler is still in use. There is no need
to switch to a global limbo handler, because even after such a switch, the
worker thread would still have to iterate through all domains one by one.
The per-domain handler already accomplishes this using a worker thread rather
than costly IPIs, so there is no clear benefit to switching to a global handler.

Could you please elaborate how a global limbo handler would require IPIs?


The global limbo handler does not require IPIs. Previously, I wondered what the
benefit would be of switching from a per-domain limbo handler to a global one.

per-domain handler:
N workers, each worker calculates the current occupancy of that domain
for a rmid. If the occupancy of all the domains drops below the threshold,
recycle that rmid. No IPI involved.

global handler:
One worker iterates over every domain. If the occupancy of all domains drops
below the threshold, it recycles the RMID - with no IPI involved.

For both the per-domain and the global handler, no IPI is involved, and
we still have to iterate over every domain. So it seems that there is not much
benefit in switching to the global handler, IIUC.

diff --git a/arch/x86/include/asm/resctrl.h b/arch/x86/include/asm/resctrl.h
index 5491853113dd..0948f64856ef 100644
--- a/arch/x86/include/asm/resctrl.h
+++ b/arch/x86/include/asm/resctrl.h
@@ -132,7 +132,13 @@ static inline void __resctrl_sched_in(struct task_struct *tsk)
static inline unsigned int resctrl_arch_round_mon_val(unsigned int val)
{
- unsigned int scale = boot_cpu_data.x86_cache_occ_scale;
+ unsigned int scale = boot_cpu_data.x86_cache_occ_scale, escale;

related to earlier topic, "scale" being unsigned int is ok since
x86_cache_occ_scale is initialized from 32bits. As I understand it the
ERDT scale value is initialized from 64 bits instead so the existing
types do not seem to accommodate?


As you mentioned in another thread, there seems to be an inconsistency
in the spec, I'll check with the team.

@@ -39,6 +43,9 @@ static int erdt_scale;
bool erdt_support(int flag)
{
+ if (flag == X86_FEATURE_CQM_OCCUP_LLC)
+ return valid_subtbl_mask & BIT(ACPI_ERDT_TYPE_CMRC);
+

Is the plan to keep adding more if() statements as new flags need to be tested?



Yes. For example, to also support MBM:

if (flag == X86_FEATURE_CQM_OCCUP_LLC)
return valid_subtbl_mask & BIT(ACPI_ERDT_TYPE_CMRC);

if (flag == X86_FEATURE_CQM_MBM_TOTAL)
return valid_subtbl_mask & BIT(ACPI_ERDT_TYPE_MMRC);

return false;
}

@@ -430,12 +434,15 @@ int __init rdt_get_l3_mon_config(struct rdt_resource *r)
struct rdt_hw_resource *hw_res = resctrl_to_arch_res(r);
unsigned int threshold;
u32 eax, ebx, ecx, edx;
+ int max_rmid;
snc_nodes_per_l3_cache = snc_get_config();
+ max_rmid = erdt_cpu_has(X86_FEATURE_CQM_OCCUP_LLC) ?
+ erdt_get_max_rmid() : boot_cpu_data.x86_cache_max_rmid;

This does not look right. Wouldn't this use the ERDT supported RMID for the MBM events also
even though they are read via MSR?


Got it, this is a bug that might impact the MBM. Let me use min() to get
the minimal rmid between the erdt and the legacy one.

resctrl_rmid_realloc_limit = boot_cpu_data.x86_cache_size * 1024;
hw_res->mon_scale = boot_cpu_data.x86_cache_occ_scale / snc_nodes_per_l3_cache;

Should the scale used by ERDT also be adjusted when SNC enabled?


My understanding is that the reason hw_res->mon_scale is divided by
snc_nodes_per_l3_cache is that one LLC is composed of several SNC nodes.
Therefore, when we sum up the monitor data from all SNC domains, we need
to scale down mon_scale per domain to avoid "over-counting". For the platform
on which we are enabling MMIO-based CMT, I am not sure whether SNC will be
supported. But we can still adjust the scale for each SNC configuration to
ensure future compatibility. Let me change the code.

thanks,
Chenyu