Re: [PATCH v2 0/4] sched/numa: stop VMA scan filters from gating promotion

From: Andrew Morton

Date: Thu Sep 17 2026 - 01:35:34 EST


On Thu, 10 Sep 2026 20:18:22 -0400 Gregory Price <gourry@xxxxxxxxxx> wrote:

> NUMA balancing uses hinting faults for both task placement and memory-tier
> promotion. Several filters designed to avoid unproductive socket-placement
> faults can also prevent promotion.
>
> Read-only file mappings and VMAs without recent PID activity may never be
> scanned, while lower MM layers reject some shared folios. Hot memory in
> these mappings can therefore remain on a slow tier indefinitely.
>
> Separate promotion-only scans from socket-placement scans.
>
> MM_CP_PROT_NUMA_PROMO_ONLY carries that choice for one protection walk, so
> the PTE and PMD paths can restrict hinting faults to promotion candidates.
> Keeping this transient state in the protection flags avoids per-mm state and
> its associated lifetime, concurrency, and VMA identity problems.
>
> - Allow eligible shared folios to be promoted to a fast tier.
> - Scan read-only file mappings and PID-inactive VMAs for promotion
> without re-enabling placement sampling.
> - Track the last placement scan separately so promotion-only scans
> cannot postpone the existing placement-starvation fallback.
>
> The series is ordered as follows:
>
> 1. Add promotion-only NUMA protection walks without changing behavior.
> 2. Permit eligible shared folios to be promoted to a fast tier.
> 3. Scan read-only file mappings using promotion-only scans.
> 4. Scan PID-inactive VMAs for promotion and account for placement scans
> separately.
>
> Tested on a host with 768GB/256GB DRAM/CXL.
> Ran 2 ~430GB database workloads with large (>300GB) shmem VMAs.
>
> Before change:
> - 150-200GB/s DRAM bandwidth usage
> - 40-45GB/s CXL bandwidth usage (maxed out)
> - request latencies over 5ms (longer tails)
>
> After chage:
> - 250GB/s+ sustained DRAM bandwidth usage
> - ~10GB/s sustained CXL bandwidth usage
> - request latencies 800us-2ms.
>
> Functional observation:
> A 20 GB hash table VMA that previously remained entirely on CXL was
> split evenly between DRAM and CXL after the changes - and tier
> residency tracked hotness. This was previously affected by the
> stavation issue caused by the "unaccessed VMA" filter.

This seems very significant?

Why cc:stable and Fixes:? Is this something which ran at these sorts
of speeds before the offending commits?