Re: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy

From: Li Zhe

Date: Fri Oct 02 2026 - 04:20:08 EST


On 10/2/26 3:56 PM, David Hildenbrand (Arm) wrote:
> On 10/1/26 15:26, Gregory Price wrote:
>> On Thu, Oct 01, 2026 at 12:54:22PM +0200, David Hildenbrand (Arm) wrote:
>>>> If this is a initial-fill problem, can userspace set weighted interleave
>>>> initially? The program or harness can observe memory usage of the program
>>>> or related NUMA nodes and switch the policy to numa balancing via
>>>> set_mempolicy() or mbind() without MPOL_MF_MOVE after certain threshold
>>>> is met?
>>> You mean: use the weighted policy initially and then switch to a NUMA-balancing
>>> one which doesn't involve the weights anymore?
>>>
>>> That makes more sense to me. Although I struggle to see why an effectively
>>> "let's put random memory on slow and others at hot" is a good starting point to
>>> later let if be fixed up by actual balancing/tiering.
>>>
>>> It all sounds a bit hackish. :)
>>>
>> It is a bit of a non-combo (Nonbo). You're using weighted interleave
>> with the intent of spreading out the bandwidth utilization (and maybe to
>> offset some reclaim behavior? *shrug*) but then undo all the placement
>> with tiering.
> I can understand the "random initial placement will help if you cross your
> fingers" argument from Zi.
>
> But then the question really is whether the app should then change the policy
> after the initial placement was done and the weighted stuff no longer makes a
> lot of sense.
Yes, that model makes sense if the application or runtime can
cooperate with the policy switch.

One limitation is that this is not fully transparent for existing
workloads. set_mempolicy() updates the calling task's policy, and
mbind() updates VMAs in the calling mm. move_pages() and
migrate_pages() can move pages of another process, but they do not
change that process's future allocation policy.

So I agree this staged approach is worth considering, but it also has
some deployment cost for workloads that cannot participate in the
policy switch.

Thanks,
Zhe

>
>> But, in defense of the hackery - I have seen strategies like this work
>> to optimize startup times and then let tiering optimize runtime. Phased
>> execution gets funky like that.
> No doubt about that.