Re: [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure
From: Davidlohr Bueso
Date: Tue Sep 29 2026 - 16:14:02 EST
On Sun, 27 Sep 2026, Bharata B Rao wrote:
Hotness promotion engine
------------------------
This is not something that was written for pghot from scratch, but instead it is
the same hot page promotion engine that is part of NUMAB2 which is now
generalized and moved to pghot. So this is not the complexity that pghot
introduces afresh.
So considering all these, I see pghot as a light-weight and low-overhead
mechanism to track per-PFN hotness and do async batch migration. Initial
versions had fancy double data structures; a large hash a small binary tree of
promotion-ready records and associated synchronization mechanism, but that is
all past now.
I think we are all in agreement that the async batch migration is wanted.
pghot interface for sampling
============================
pghot_record_access() interface was designed keeping the existing NUMAB2 source
in mind. It fits that and it fits other sources like IBS Memory Profiler. So
sampling sources report an access and the shared promotion engine acts upon it.
But for sources like CHMU, from what you describe, I gather that a bulk
reporting interface plus an indication to bypass the engine to treat the PFNs as
migrate-ready, is what is required. Should those migrate-ready PFNs go through
the regular pghot tracking (getting into section hotmaps to be picked up by
kmigrated) or even that should be bypassed?
However, in the context of PTE A bit based source, I have often thought about
extending the interface for
- bulk reporting where more than one PFN gets reported.
- indicating the bypass options (frequency check bypass, recency check bypass etc)
NUMAB2 has to perform better in pghot
=====================================
pghot is about a sub-system that makes it possible to have multiple sources to
coexist with reuse of common hot page promotion engine.
NUMAB2 source resides within the scheduler and the promotion engine is also part
of the scheduler. Through pghot, I am separating the source (NUMA hint faults)
from the engine and moving that existing engine into pghot, to a common place
where it gets reused for other sources as well.
It is the same NUMA hint faults and more or less the same engine and hence my
main objective is to ensure that there is no regression during this move.
Additional performance optimizations can be done to the engine itself separately
but that shouldn't be the baseline expectation from pghot.
Is moving hot page promotion out of scheduler into a dedicated system, a good
thing in general? I believe so as scheduler isn't the right place for it to
reside. However I would like to hear from scheduler folks on this.
So two of the autonuma balancing og authors are scheduler experts - and iirc
*the* reason back then was locality. And Peter has already nacked the IBS
stuff in the past. But indeed the batch async part would be good to get nack/ack;
albeit the cgroup charging situation.
Do we even need a centralized hot page promotion engine?
========================================================
NUMAB2 is good and serves as a good baseline for any new source that comes up.
But with different kinds of sources becoming available, do we want all of them
to duplicate the hot page heuristics and promote hot pages on their own? I
thought that may not be preferable and hence started this effort.
What are these sources that will become available?
CXL HMU may not need the promotion engine, but IBS Memory Profiler needs. It
needs a promotion engine with full recency and frequency considerations before
promoting. I don't think an arch driver like IBS Memory Profiler should be doing
hotness heuristics within itself but instead be using the existing engine. In
fact in my early posts, the driver based on primary IBS instance was feeding the
"access sample" as "NUMA hint fault" to NUMAB1/B2 so that rest of NUMA
Balancing/Hot page promotion engine just worked. But I think pghot is a better
approach than that.
I agree that the IBS driver should not be doing hotness heuristics, it's the hw
that should.
Why full pghot? Isn't async batch migration enough?
===================================================
Some of the above reasons apply but during the course of iterations, I have had
implementations of just the migrator (kmigrated [1]).
If every sub-system/source has intelligence of its own and just wants to
handover a list of pages to async migrator thread, that's not much of an effort
as this implementation showed.
But then if some source needs rate-limiting and another source needs only
hottest pages to be promoted, then again we overlap with the existing NUMAB2 engine.
Then if we want to be slightly generic and want two sources to complement each
other or the promoter to differentiate between lukewarm vs hot pages/regions
then we may have to maintain hotness records and may soon end up with something
similar to pghot's hotness tracking and reporting mechanism.
Hardware sources have to out-perform NUMAB2
===========================================
Different sources will have different characteristics and capabilities and will
help different workloads differently. So it is the choice that one could
provide. Sometimes sources can complement each other as well.
What IBS Memory Profiler has shown is that it can match and/or exceed (for
Graph500, ptr-chase, llama) NUMAB2 with no hint faults overhead [2]. Also please
check the initial XSBench numbers in my inline reply.
It's all about the numbers, and the ones you have just don't really sell - which is
one of the reasons this has been going on for years. If numa balancing didn't
exist, then maybe adding all this would make sense. For the XSBench I don't think
making decisions based on the non-overcommitted case is worthwhile.
Cost of an unused source
========================
Not all the sources are required for every situation. Sources can be disabled at
compile time or not enabled at run time with no cost or effect on other sources.
But hotmap allocations would remain as a static cost even when no source is
enabled (built out at compile time, the map is gone; sources off at runtime, the
map remains)
Tracking granularity: per-PFN vs region
=======================================
Often times this question comes up when pghot is compared with DAMON.
Firstly, the baseline that we have (NUMAB2), tracks hotness at page granularity.
To support this, pghot started with per-PFN granularity. Naturally two concerns
come up:
1. Memory overhead: I have shown the numbers above. It is lower-tier only and
not much IMHO.
2. Scan/CPU overhead: pghot started with scanning all PFNs, that was expensive.
Then it introduced hotness bit per memory section so that only those sections
which are marked hot are scanned. Now I have added (yet to be posted) a
sub-section level hotness tracking where a hotness bit is maintained for each
fixed 2M region within a section. This has considerably reduced the CPU overhead
for kmigrated thread. While the ptr-chase numbers that I shared with the
separated out IBS RFC v0 post of IBS Memory Profiler [2] does show the kmigrated
utilization numbers with sub-section tracking, I plan to have some more numbers
ready for LPC.
I look forward to seeing any new numbers you have. Region granularity is certainly
more aligned with willy as well as chmu.
I am beginning to feel that this may be a good middle ground between real
region-level tracking (where all pages of region are promoted irrespective of
their real hotness) vs scanning at sub-section granularity but performing
per-PFN precise promotion.
[...]
I would be interested to understand more about how you ran XSBench, the
parameters used, the promotion stats, if demotion stats etc. Do share when you
get time.
The actual command is 'XSBench -g 90424' so only set the gridpoints, everything
else is default, so total ~44Gb footprint. I don't have the vmstats currently
(I do not run these benchmarks) but will share them once I get them - I can
affirm that demotion is in fact enabled, so full TPP up and down.
For the chmu: 32GB device (1:1 dram and cxl), this is with a 4k unit size, 1s
epoch, reporting mode is always on, threshold value is 1024.
The numa balancing mode was set to 3, the rest used the default values:
pghot_freq_threshold=2, pghot_promote_freq_window_ms=3000,
pghot_promote_rage_limit_MBps=65536, kmigrated_sleep_ms=100, kmigrated_batch_nr=512.
I also have numbers for two more benchmarks with a real chmu, with basically
the same numab parameters:
(i) TaoBench almost 2x throughput, going from ~250 qps to ~480 qps, vs NUMAB3.
(dram:cxl is 16:32Gb with a 32Gb memsize, num_clients=2, clients_per_thread=75)
(ii) Graph500 only shows a smaller ~20% improvement vs NUMAB3 (bfs mean time drops
from 5.0 to 4.15 secs). This was for a memory ratio dram:cxl as 64:32Gb. The
algorithm is BFS, edge factor 512, 16 mpi processes.
(mpiexec.openmpi -n 16 ./graph500_reference_bfs 22 512)
I did a round of testing to compare base-NUMAB2, pghot-hintfaults(NUMAB2) and
pghot-hwhints (IBS Memory Profiler). I am still experimenting with options,
placement etc but here are my initial numbers:
Non-Overcommitted case: XSBench working set fits fully within toptier but starts
on lower tier before the measurement phase
Those are nice numbers, but I don't think this is the methodology to use...
It is more representative for the workload's working set to be > total dram and
therefore spill into slower tier(s), instead of artificially starting in the slow
memory and moving up. Do you have data for the over committed case?
Thanks,
Davidlohr
(XSBench -s XL -g 238847 -t 64 -p 60000000 -l 34, footprint ~116.56 GiB)
Runtime (s) lookups/s Promotions (pages)
base-NUMAB0 498.6 4.09M 0
base-NUMAB2 422.4 4.83M 30.5M
pghot-hintfaults 329.6 6.20M 30.5M
pghot-hwhints 109.2 18.69M 2.7M
Same amount of pages promoted by both base-NUMAB2 and pghot-hintfaults but
runtime is better with pghot. This could be the benefit of async batched
migration showing and no adverse effect of losing cache locality.
pghot-hwhints shows good results. As I said this is just a first peek to the
experimental numbers, I should have more concrete numbers and conclusion in LPC.
[1] Kmigrated -
https://lore.kernel.org/linux-mm/20250616133931.206626-1-bharata@xxxxxxx/#t
[2] IBS Memory Profiler RFC v0 -
https://lore.kernel.org/linux-mm/20260924062206.319314-1-bharata@xxxxxxx/
Regards,
Bharata.