[PATCH v3 3/7] sched/numa: scan read-only file mappings in tiering mode
From: Gregory Price
Date: Tue Sep 22 2026 - 15:02:31 EST
From: "Gregory Price (Meta)" <gourry@xxxxxxxxxx>
Commit 4591ce4f2d22 ("sched/numa: Do not trap hinting faults for
shared libraries") excludes file-backed read-only VMAs from NUMA
hint faulting to prevent east-west placement bouncing.
This filter hides hot file folios on slow memory from promotion.
Scan those VMAs when tiering is enabled, but make their scans promotion
only to retains the existing restriction. Keep the historical VMA
predicate unchanged for backport-ability.
Read the balancing mode once per task_numa_work() invocation and build
the protection flags for each VMA from that snapshot. This keeps the PTE
and PMD paths on the same policy for an entire protection walk (which
may be split across multiple scanning periods).
On a host with 768 GB of DRAM and 256 GB of CXL memory running two
database services using ~430GB each, 169 MB of their shared 185 MB
main binary accumulated on CXL before permanently stuck there.
With the change, the binary tier residency tracks its runtime hotness.
Fixes: c574bbe91703 ("NUMA balancing: optimize page placement for memory tiering system")
Cc: stable@xxxxxxxxxxxxxxx
Assisted-by: LLM
Signed-off-by: Gregory Price (Meta) <gourry@xxxxxxxxxx>
---
kernel/sched/fair.c | 31 +++++++++++++++++++++----------
1 file changed, 21 insertions(+), 10 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index dc78d24ed8bc0..8a4687f67d82f 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -4119,21 +4119,21 @@ static bool vma_is_accessed(struct mm_struct *mm, struct vm_area_struct *vma)
*/
static void task_numa_work(struct callback_head *work)
{
+ const unsigned int numab_mode = READ_ONCE(sysctl_numa_balancing_mode);
+ const bool tiering = numab_mode & NUMA_BALANCING_MEMORY_TIERING;
unsigned long migrate, next_scan, now = jiffies;
struct task_struct *p = current;
struct mm_struct *mm = p->mm;
u64 runtime = p->se.sum_exec_runtime;
struct vm_area_struct *vma;
- unsigned long cp_flags = MM_CP_PROT_NUMA;
+ unsigned long cp_flags;
unsigned long start, end;
unsigned long nr_pte_updates = 0;
long pages, virtpages;
struct vma_iterator vmi;
bool vma_pids_skipped;
bool vma_pids_forced = false;
-
- if (!(READ_ONCE(sysctl_numa_balancing_mode) & NUMA_BALANCING_NORMAL))
- cp_flags |= MM_CP_PROT_NUMA_PROMO_ONLY;
+ bool placement_scan;
WARN_ON_ONCE(p != container_of(work, struct task_struct, numa_work));
@@ -4221,13 +4221,19 @@ static void task_numa_work(struct callback_head *work)
}
/*
- * Shared library pages mapped by multiple processes are not
- * migrated as it is expected they are cache replicated. Avoid
- * hinting faults in read-only file-backed mappings or the vDSO
- * as migrating the pages will be of marginal benefit.
+ * Shared library pages mapped by multiple processes are limited
+ * to south->north migrations as it is expected they are cache
+ * replicated. The benefit of east-west migration in this case
+ * is at best marginal and may be harmful due to TLB/cache
+ * invalidation.
+ *
+ * Allow promotion as a cold page incurring many cache-misses
+ * under cache pressure can drive considerable bandwidth.
*/
- if (!vma->vm_mm ||
- (vma->vm_file && (vma->vm_flags & (VM_READ|VM_WRITE)) == (VM_READ))) {
+ placement_scan = !(vma->vm_file &&
+ (vma->vm_flags & (VM_READ | VM_WRITE)) == VM_READ);
+
+ if (!vma->vm_mm || (!placement_scan && !tiering)) {
trace_sched_skip_vma_numa(mm, vma, NUMAB_SKIP_SHARED_RO);
continue;
}
@@ -4307,6 +4313,11 @@ static void task_numa_work(struct callback_head *work)
continue;
}
+ placement_scan &= numab_mode & NUMA_BALANCING_NORMAL;
+ cp_flags = MM_CP_PROT_NUMA;
+ if (!placement_scan)
+ cp_flags |= MM_CP_PROT_NUMA_PROMO_ONLY;
+
do {
start = max(start, vma->vm_start);
end = ALIGN(start + (pages << PAGE_SHIFT), HPAGE_SIZE);
--
2.53.0-Meta