Re: [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure
From: Davidlohr Bueso
Date: Thu Sep 24 2026 - 22:20:24 EST
On Tue, 28 Jul 2026, Bharata B Rao wrote:
Performance summary
===================
All results are on 3-node tiered systems (DRAM top tier + CPU-less CXL
lower tier). Speedups below are normalized to the base kernel with no
tiering (NUMAB=0). Columns compare mainline hint-fault tiering
(base NUMAB=2) against pghot in two modes: hint-fault driven (pghot-hf,
NUMAB=2) and HW-hint driven via the AMD IBS memory profiler (pghot-hw,
NUMAB=0, no NUMA scanning). Single run per config (abench is avg of 3).
Benchmark Metric (higher=better) base-NUMAB2 pghot-hf pghot-hw
-------------------------------------------------------------------------
NAS BT (MPI) Mop/s total 2.26x 2.26x 2.13x
Graph500 BFS harmonic-mean TEPS 2.64x 2.49x 2.68x
llama.cpp decode tok/s (tg128) 1.09x 1.41x 1.21x
Redis+memtier ops/sec 1.05x 1.05x 1.00x
Microbench completion time (1/t) 2.41x 2.40x 2.35x
-------------------------------------------------------------------------
(baseline = base kernel, no tiering = 1.00x)
Summary: pghot hint-faults reproduces mainline NUMAB=2 with no
regression on every workload; pghot HW-hints/IBS matches it mostly.
The IBS doesn't seem to bring anything to the table, other than just
serve as an example - we've discussed this in the past this. And for
NUMAB2 I would expect this to be the very best case in that you are
not suffering from loosing the locality info.
This is a lot of added complexity and code just to break even. And
hardware sources should really outperform NUMAB2 for this to be
worthwhile, methinks.
I know Joshua mentioned some improved PSI by doing the promotion async;
which makes me wonder that perhaps just a simpler promoting kthread is
enough, so NUMAB2 just kicks the migration to it and not done in the
accessing tasks' context etc. And Gregory also recently improved this
path vs NUMAB1.
I mention this because I have integrated the chmu as a source of
hotness for pghot and this was tested on a real device vs NUMAB3.
This was based on Jonathan's rfc for the perf driver, only that I carve
out the first chmu instance for this, and leave the rest of them (if any
for perf).
For example, 'XSBench' improved by a factor of ~2.7x (I would like to
report a wider range of benchmarks at some point, but this is what I have
now). But a lot of this bypasses most of this framework and just piggy
backs on kmigrated moving memory. chmu does not do well with
pghot_record_access(), which is designed for sampling... so I just set
MIGRATE_READY right away. I also added a pghot call/api to make use of
the regions chmu deals with instead of per access/page. The cxl driver
just takes the dpa range given by the chmu hotlist, coalesces into adjacent
hpa ranges (just 1 for the case of no interleaving) and feeds that to the
pghot subsystem as a list of these contiguous physical chunks ranked/sorted
by access density (accesses per page per second, so a large lukewarm region
and a smaller hotter one can be differentiated and promote the latter
first).
So perhaps any mm interface for this stuff should just not be designed
around sampling, and instead proper hardware sources?
A lot of the complexity in pghot can be removed without this imo (per-pfn
metadata, accumulate+decay phase). Of course we still have the problem
to evaluate the cost of replacing a chunk from the top tier(s) with the
what the low tier is reporting has "hot". And this is particularly true
with chmu with limited full system memory visibility (as opposed to IBS,
for example). I don't really have a good answer for this, other than the
user could use proactive reclaim along with the per-device tunables
(ie unit sizes and thresholds for chmu) to configure things realistically
for their use cases - prob. easier said than done.
Anyway these are some thoughts prior to LPC.
Thanks,
Davidlohr