Re: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy

From: Zi Yan

Date: Thu Oct 01 2026 - 07:18:24 EST


On 1 Oct 2026, at 6:54, David Hildenbrand (Arm) wrote:

> On 9/30/26 17:10, Zi Yan wrote:
>> On 30 Sep 2026, at 10:50, Gregory Price wrote:
>>
>>> On Wed, Sep 30, 2026 at 10:02:23PM +0800, Li Zhe wrote:
>>>> The use case is closer to the second one, but the main motivation is not
>>>> only startup-time performance.  The more important point is that each
>>>> workload has its own DDR/CXL budget assigned by the workload manager.
>>>>
>>>
>>> mempolicy is the wrong interface to do budget policy, you'd be better
>>> off looking at Joshua's memcg tiered node solutions for that.
>>>
>>>> MPOL_WEIGHTED_INTERLEAVE is useful because it lets new allocations
>>>> follow that per-workload budget from the beginning, instead of placing
>>>> everything on one tier first and correcting the placement later. This is
>>>> important for workload orchestration because the workload starts from a
>>>> placement close to its assigned DDR/CXL ratio.
>>>>
>>>
>>> This however is reasonable to me - you'd prefer to spread out the cost
>>> of initial faulting placement explicitly, rather than simply take
>>> fallbacks when the top-tier budget becomes pressured.
>>>
>>> i.e. w/o interleave:
>>>
>>> [node 0 ]
>>> ^^^^^^^^^^^^^^ alloc until full
>>> vvvvvvvvv fallback
>>> [node 1 ]
>>>
>>> in this scenario you end up with considerable hot memory
>>> on the remote node consolidated in time-space (everything
>>> allocated after node0 becomes full skews heavily toward
>>> node1)
>>>
>>> Tiering then likely takes many faults to rebalance after
>>> you've already reached node0 limits.
>>>
>>> w/ interleave
>>>
>>> [node 0 ]
>>> ^^^vv^^^vv^^^vv^^^vv^^^vv^^^vv....
>>> [node 1 ]
>>>
>>> In this scenario you do an initial fill distributed by weight
>>> and then let tiering figure it out without consolidating all
>>> of the pressure to the point where node0 has no space left.
>>>
>>> That said - this seems like mostly an initial-fill problem, after
>>> you initially fill your memory, a new allocation largely implies
>>> the memory is hot - and you probably prefer that to be local.
>>
>> If this is a initial-fill problem, can userspace set weighted interleave
>> initially? The program or harness can observe memory usage of the program
>> or related NUMA nodes and switch the policy to numa balancing via
>> set_mempolicy() or mbind() without MPOL_MF_MOVE after certain threshold
>> is met?
>
> You mean: use the weighted policy initially and then switch to a NUMA-balancing
> one which doesn't involve the weights anymore?
>
> That makes more sense to me. Although I struggle to see why an effectively
> "let's put random memory on slow and others at hot" is a good starting point to
> later let if be fixed up by actual balancing/tiering.

Statistically speaking, unless the hot/cold page distribution follows a power-law
distribution, random placement usually provides an average placement result.
So it is not a bad start point.

>
> It all sounds a bit hackish. :)


Best Regards,
Yan, Zi