Re: [PATCH v2 0/4] sched/numa: stop VMA scan filters from gating promotion

From: Zi Yan

Date: Fri Sep 18 2026 - 16:56:49 EST


On Thu Sep 10, 2026 at 8:18 PM EDT, Gregory Price wrote:
> NUMA balancing uses hinting faults for both task placement and memory-tier
> promotion. Several filters designed to avoid unproductive socket-placement
> faults can also prevent promotion.
>
> Read-only file mappings and VMAs without recent PID activity may never be
> scanned, while lower MM layers reject some shared folios. Hot memory in
> these mappings can therefore remain on a slow tier indefinitely.
>
> Separate promotion-only scans from socket-placement scans.
>
> MM_CP_PROT_NUMA_PROMO_ONLY carries that choice for one protection walk, so
> the PTE and PMD paths can restrict hinting faults to promotion candidates.
> Keeping this transient state in the protection flags avoids per-mm state and
> its associated lifetime, concurrency, and VMA identity problems.
>
> - Allow eligible shared folios to be promoted to a fast tier.
> - Scan read-only file mappings and PID-inactive VMAs for promotion
> without re-enabling placement sampling.
> - Track the last placement scan separately so promotion-only scans
> cannot postpone the existing placement-starvation fallback.
>
> The series is ordered as follows:
>
> 1. Add promotion-only NUMA protection walks without changing behavior.
> 2. Permit eligible shared folios to be promoted to a fast tier.
> 3. Scan read-only file mappings using promotion-only scans.
> 4. Scan PID-inactive VMAs for promotion and account for placement scans
> separately.
>
> Tested on a host with 768GB/256GB DRAM/CXL.
> Ran 2 ~430GB database workloads with large (>300GB) shmem VMAs.

Do you just launch them without any additional NUMA control or NUMA
policy configuration, like using numactl? Just want to understand the
scope of the issue and how well the fixes are tested. It might be good
to have a list of expected behaviors for people/AI to check against.

>
> Before change:
> - 150-200GB/s DRAM bandwidth usage
> - 40-45GB/s CXL bandwidth usage (maxed out)
> - request latencies over 5ms (longer tails)
>
> After chage:
> - 250GB/s+ sustained DRAM bandwidth usage
> - ~10GB/s sustained CXL bandwidth usage
> - request latencies 800us-2ms.
>
> Functional observation:
> A 20 GB hash table VMA that previously remained entirely on CXL was
> split evenly between DRAM and CXL after the changes - and tier
> residency tracked hotness. This was previously affected by the
> stavation issue caused by the "unaccessed VMA" filter.
>
> Gregory Price (Meta) (4):
> mm: support promotion-only NUMA hinting scans
> mm: allow shared folios to be promoted to a fast tier
> sched/numa: scan read-only file mappings in tiering mode
> sched/numa: do not let VMA PID activity gate promotion
>
> include/linux/mm.h | 4 ++-
> include/linux/mm_types.h | 7 ++++
> kernel/sched/fair.c | 72 +++++++++++++++++++++++++++++-----------
> mm/huge_memory.c | 3 +-
> mm/internal.h | 5 +--
> mm/mempolicy.c | 30 +++++++++++------
> mm/migrate.c | 6 ++--
> mm/mprotect.c | 4 ++-
> 8 files changed, 93 insertions(+), 38 deletions(-)




--
Best Regards,
Yan, Zi