Re: [RFC] mpam,x86,fs/resctrl: Generic schema description Proof of Concept
From: Ben Horgan
Date: Tue Aug 04 2026 - 05:31:58 EST
Hi Babu,
On 7/22/26 18:02, Babu Moger wrote:
> Hi Ben,
>
> On 7/22/26 05:47, Ben Horgan wrote:
>> Hi Babu,
>>
>> On 7/21/26 21:02, Babu Moger wrote:
>>> Hi Ben/Reinette,
>>>
>>> On 7/21/26 12:30, Reinette Chatre wrote:
>>>> Hi Ben,
>>>>
>>>> On 7/21/26 6:23 AM, Ben Horgan wrote:
>>>>> On 7/20/26 23:54, Reinette Chatre wrote:
>>>>>> On 7/20/26 6:30 AM, Ben Horgan wrote:
>>>> ...>
>>>>> The former, info/ contains a directory for each allocation scope of each resource.
>
>>>>>>
>>>>>> The other x86 feature to consider is AMD's upcoming "Global" MBA/SMBA that exposes memory
>>>>>> bandwidth allocation
>>>>>> in "groups of L3" that I understand could usually be mapped to NODE scope (but it remains
>>>>>> controlled at L3 scope),
>>>>>> except for one configuration where it is "SYSTEM"(?) scope.
>>>>>> Ref.: https://lore.kernel.org/
>>>>>> lkml/8f77f498b1c77fa8fd8f5d5687f03ae598068544.1776980182.git.babu.moger@xxxxxxx/
>>>>>
>>>>> Hmmm, I'm not sure that the scope can be considered to be NODE scope for GMBA. To me it seems
>>>>> to be
>>>>> accidental that it maps to the NUMA node but really the scope is just a grouping of L3 instances.
>>>>> For a control to NUMA scope I would expect the resctrl domains to go offline and online in sync
>>>>> with
>>>>> the NUMA nodes. For GMBA it looks like it would just going offline/online based on whether any of
>>>>> the CPUs and so L3 instances in the group are online. Am I correct here?
>>>>>
>>>>> Assuming the domains are on L3 groups rather than NUMA also changes which end of the link the
>>>>> traffic is regulated and so how cross-NUMA traffic behaves differently. If the domain is an L3
>>>>> group
>>>>> then a task running on a CPU affine to that L3 group won't be throttled unless that particular
>>>>> domain is throttled but with NUMA node domains it may be throttled if it has traffic going to that
>>>>> domain.
>>>>
>>>> I'll defer to Babu for accurate answers about this hardware capability.
>>>>
>>>
>>> To me, Global MBA should be considered a NODE-scoped resource. In some configurations it may appear
>>> as SYSTEM-scoped, but that is effectively equivalent to a single-node encompassing the entire
>>> system. In such cases, there is only one schemata entry controlling the whole system.
>>>
>>> Yes, multiple L3 instances are grouped together to form a NODE. Internally, programming is still
>>> performed at the L3 level, but that implementation detail can be hidden from users and does not need
>>> to be exposed through the interface.
>>
>> We seem to have two things that can both, somewhat reasonably, be called NODE scope in the resctrl
>> user interface but the behaviour required for an MPAM system and an AMD system appears different
>> from the point of view of lifecycle of the resctrl domain.
>>
>> For MPAM NUMA scope the MSC instance (MPAM hardware interface) is at the memory controller and so
>> goes on and offline based on whether the NUMA node is offline or online. For AMD NUMA scope it looks
>> to me that the lifecycle of the resctrl domains would be tied to the CPUs associated with the NUMA
>> node. To me it does seem odd that a control with a domain associated with an offline NUMA node can
>> continue to throttle (cross-NUMA) traffic.
>>
>> Is there any GLBE Control Domain ID or similar that is exposed to the user, e.g. is sysfs, or is
>> this just implicitly the NUMA id?
>>
> Yes, the GLBE Control Domain ID is exposed to the user. It is essentially equivalent to the NUMA ID.
What's the on/off lifecycle of these nodes? Does it follow the NUMA lifecycle as managed by the NUMA
node notifiers documented in Documentation/core-api/memory-hotplug.rst or is it just linked to the
cpu hotplug as is done currently for the resctrl cache based domains.
Thanks,
Ben
>
> Thanks
> Babu