Re: [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure

From: Bharata B Rao

Date: Sun Sep 27 2026 - 11:52:58 EST


Thanks Davidlohr for using pghot and providing your comments. The timing is also
perfect as we can continue the discussions in LPC too. Let me try to address
your concerns by with top-posting first where I cover other general concerns as
well and would reply inline for your XSBench observation.

Complexity of pghot
===================
Let me start with lines-of-code and functionality breakup.

1135 mm/pghot.c
86 mm/pghot-default.c
85 mm/pghot-precise.c
208 include/linux/pghot.h
45 mm/migrate.c
1559 total plus some small bits in mm/vmstat.c, mm/init.c, Kconfig, Makefile etc
That includes around 160 lines of NUMAB2 engine code moved from
kernel/sched/fair.c

The above includes the following distinct functionalities and parts:

Hotness recording mechanism (pghot_record_access())
Hotness promotion engine (which is moved from kernel/sched/fair.c)
PFN walk and async migration (kmigrated)
CPU, memory and node hotplug handling
Sysctl and debugfs tunables, vmstat counters, enabling/disabling parts etc

Memory overhead
---------------
Default mode: 1 byte per lower-tier PFN => 256MB overhead per 1TB of lower-tier
memory.
Precision mode: 4 bytes per lower-tier PFN => 1GB for 1TB of lower-tier memory.

In not-yet-posted version, I am tracking hotness at sub-section granularity (in
addition to the existing per-section bit) which makes PFN walk more targeted
and reduces the CPU overhead of kmigrated thread. This adds 64KB per 1TB of
lower tier memory overhead for 2M sub-section tracking.

Runtime overhead
----------------
Recording hotness is mostly a constant time operation function. Hotness record
lookup involves PFN to section, PFN to section->hot_map. Hotness record update
involves a few atomic bit operations and an atomic cmpxchg operation.

Kmigrated walks only those sections that are marked hot. In the next version, it
would even skip those sub-section ranges that aren't marked hot. Within a
sub-section of 2M it walks all the PFNs (512 for x86). Hotness record
consumption again involves an atomic cmpxchg operation.

migrate_pages() already supports batched migrations, kmigrated just feeds it
with batched hot folio list extracted during the PFN walk.

No fancy data structures to track hotness, no complicated synchronization
mechanisms between producer (pghot_record_access()) and consumer (kmigrated)

Hotness promotion engine
------------------------
This is not something that was written for pghot from scratch, but instead it is
the same hot page promotion engine that is part of NUMAB2 which is now
generalized and moved to pghot. So this is not the complexity that pghot
introduces afresh.

So considering all these, I see pghot as a light-weight and low-overhead
mechanism to track per-PFN hotness and do async batch migration. Initial
versions had fancy double data structures; a large hash a small binary tree of
promotion-ready records and associated synchronization mechanism, but that is
all past now.

pghot interface for sampling
============================
pghot_record_access() interface was designed keeping the existing NUMAB2 source
in mind. It fits that and it fits other sources like IBS Memory Profiler. So
sampling sources report an access and the shared promotion engine acts upon it.

But for sources like CHMU, from what you describe, I gather that a bulk
reporting interface plus an indication to bypass the engine to treat the PFNs as
migrate-ready, is what is required. Should those migrate-ready PFNs go through
the regular pghot tracking (getting into section hotmaps to be picked up by
kmigrated) or even that should be bypassed?

However, in the context of PTE A bit based source, I have often thought about
extending the interface for

- bulk reporting where more than one PFN gets reported.
- indicating the bypass options (frequency check bypass, recency check bypass etc)

NUMAB2 has to perform better in pghot
=====================================
pghot is about a sub-system that makes it possible to have multiple sources to
coexist with reuse of common hot page promotion engine.

NUMAB2 source resides within the scheduler and the promotion engine is also part
of the scheduler. Through pghot, I am separating the source (NUMA hint faults)
from the engine and moving that existing engine into pghot, to a common place
where it gets reused for other sources as well.

It is the same NUMA hint faults and more or less the same engine and hence my
main objective is to ensure that there is no regression during this move.
Additional performance optimizations can be done to the engine itself separately
but that shouldn't be the baseline expectation from pghot.

Is moving hot page promotion out of scheduler into a dedicated system, a good
thing in general? I believe so as scheduler isn't the right place for it to
reside. However I would like to hear from scheduler folks on this.

Do we even need a centralized hot page promotion engine?
========================================================
NUMAB2 is good and serves as a good baseline for any new source that comes up.
But with different kinds of sources becoming available, do we want all of them
to duplicate the hot page heuristics and promote hot pages on their own? I
thought that may not be preferable and hence started this effort.

CXL HMU may not need the promotion engine, but IBS Memory Profiler needs. It
needs a promotion engine with full recency and frequency considerations before
promoting. I don't think an arch driver like IBS Memory Profiler should be doing
hotness heuristics within itself but instead be using the existing engine. In
fact in my early posts, the driver based on primary IBS instance was feeding the
"access sample" as "NUMA hint fault" to NUMAB1/B2 so that rest of NUMA
Balancing/Hot page promotion engine just worked. But I think pghot is a better
approach than that.

Why full pghot? Isn't async batch migration enough?
===================================================
Some of the above reasons apply but during the course of iterations, I have had
implementations of just the migrator (kmigrated [1]).

If every sub-system/source has intelligence of its own and just wants to
handover a list of pages to async migrator thread, that's not much of an effort
as this implementation showed.

But then if some source needs rate-limiting and another source needs only
hottest pages to be promoted, then again we overlap with the existing NUMAB2 engine.

Then if we want to be slightly generic and want two sources to complement each
other or the promoter to differentiate between lukewarm vs hot pages/regions
then we may have to maintain hotness records and may soon end up with something
similar to pghot's hotness tracking and reporting mechanism.

Hardware sources have to out-perform NUMAB2
===========================================
Different sources will have different characteristics and capabilities and will
help different workloads differently. So it is the choice that one could
provide. Sometimes sources can complement each other as well.

What IBS Memory Profiler has shown is that it can match and/or exceed (for
Graph500, ptr-chase, llama) NUMAB2 with no hint faults overhead [2]. Also please
check the initial XSBench numbers in my inline reply.

Cost of an unused source
========================
Not all the sources are required for every situation. Sources can be disabled at
compile time or not enabled at run time with no cost or effect on other sources.
But hotmap allocations would remain as a static cost even when no source is
enabled (built out at compile time, the map is gone; sources off at runtime, the
map remains)

Tracking granularity: per-PFN vs region
=======================================
Often times this question comes up when pghot is compared with DAMON.

Firstly, the baseline that we have (NUMAB2), tracks hotness at page granularity.
To support this, pghot started with per-PFN granularity. Naturally two concerns
come up:

1. Memory overhead: I have shown the numbers above. It is lower-tier only and
not much IMHO.

2. Scan/CPU overhead: pghot started with scanning all PFNs, that was expensive.
Then it introduced hotness bit per memory section so that only those sections
which are marked hot are scanned. Now I have added (yet to be posted) a
sub-section level hotness tracking where a hotness bit is maintained for each
fixed 2M region within a section. This has considerably reduced the CPU overhead
for kmigrated thread. While the ptr-chase numbers that I shared with the
separated out IBS RFC v0 post of IBS Memory Profiler [2] does show the kmigrated
utilization numbers with sub-section tracking, I plan to have some more numbers
ready for LPC.

I am beginning to feel that this may be a good middle ground between real
region-level tracking (where all pages of region are promoted irrespective of
their real hotness) vs scanning at sub-section granularity but performing
per-PFN precise promotion.

Losing cache locality due to async migration
============================================
I see that async batched migration is important to free up the heavy migration
related activities from the process context. Losing cache locality is inherent
to that mechanism I believe. Given that pghot-NUMAB2 didn't show any
significant regression with the benchmark numbers I presented, it would be good
to find other benchmarks that show up this locality cost in a tiered memory setup.


On 25-Sep-26 7:42 AM, Davidlohr Bueso wrote:
> On Tue, 28 Jul 2026, Bharata B Rao wrote:
>
>> Performance summary
>> ===================
>> All results are on 3-node tiered systems (DRAM top tier + CPU-less CXL
>> lower tier). Speedups below are normalized to the base kernel with no
>> tiering (NUMAB=0). Columns compare mainline hint-fault tiering
>> (base NUMAB=2) against pghot in two modes: hint-fault driven (pghot-hf,
>> NUMAB=2) and HW-hint driven via the AMD IBS memory profiler (pghot-hw,
>> NUMAB=0, no NUMA scanning). Single run per config (abench is avg of 3).
>>
>> Benchmark      Metric (higher=better)   base-NUMAB2  pghot-hf  pghot-hw
>> -------------------------------------------------------------------------
>> NAS BT (MPI)   Mop/s total                  2.26x     2.26x     2.13x
>> Graph500 BFS   harmonic-mean TEPS           2.64x     2.49x     2.68x
>> llama.cpp      decode tok/s (tg128)         1.09x     1.41x     1.21x
>> Redis+memtier  ops/sec                      1.05x     1.05x     1.00x
>> Microbench     completion time (1/t)        2.41x     2.40x     2.35x
>> -------------------------------------------------------------------------
>> (baseline = base kernel, no tiering = 1.00x)
>>
>> Summary: pghot hint-faults reproduces mainline NUMAB=2 with no
>> regression on every workload; pghot HW-hints/IBS matches it mostly.
>
> The IBS doesn't seem to bring anything to the table, other than just
> serve as an example - we've discussed this in the past this. And for
> NUMAB2 I would expect this to be the very best case in that you are
> not suffering from loosing the locality info.
>
> This is a lot of added complexity and code just to break even. And
> hardware sources should really outperform NUMAB2 for this to be
> worthwhile, methinks.
>
> I know Joshua mentioned some improved PSI by doing the promotion async;
> which makes me wonder that perhaps just a simpler promoting kthread is
> enough, so NUMAB2 just kicks the migration to it and not done in the
> accessing tasks' context etc. And Gregory also recently improved this
> path vs NUMAB1.
>
> I mention this because I have integrated the chmu as a source of
> hotness for pghot and this was tested on a real device vs NUMAB3.
> This was based on Jonathan's rfc for the perf driver, only that I carve
> out the first chmu instance for this, and leave the rest of them (if any
> for perf).
> For example, 'XSBench' improved by a factor of ~2.7x (I would like to
> report a wider range of benchmarks at some point, but this is what I have
> now). But a lot of this bypasses most of this framework and just piggy
> backs on kmigrated moving memory. chmu does not do well with
> pghot_record_access(), which is designed for sampling... so I just set
> MIGRATE_READY right away. I also added a pghot call/api to make use of
> the regions chmu deals with instead of per access/page. The cxl driver
> just takes the dpa range given by the chmu hotlist, coalesces into adjacent
> hpa ranges (just 1 for the case of no interleaving) and feeds that to the
> pghot subsystem as a list of these contiguous physical chunks ranked/sorted
> by access density (accesses per page per second, so a large lukewarm region
> and a smaller hotter one can be differentiated and promote the latter
> first).

I would be interested to understand more about how you ran XSBench, the
parameters used, the promotion stats, if demotion stats etc. Do share when you
get time.

I did a round of testing to compare base-NUMAB2, pghot-hintfaults(NUMAB2) and
pghot-hwhints (IBS Memory Profiler). I am still experimenting with options,
placement etc but here are my initial numbers:

Non-Overcommitted case: XSBench working set fits fully within toptier but starts
on lower tier before the measurement phase
(XSBench -s XL -g 238847 -t 64 -p 60000000 -l 34, footprint ~116.56 GiB)

Runtime (s) lookups/s Promotions (pages)
base-NUMAB0 498.6 4.09M 0
base-NUMAB2 422.4 4.83M 30.5M
pghot-hintfaults 329.6 6.20M 30.5M
pghot-hwhints 109.2 18.69M 2.7M

Same amount of pages promoted by both base-NUMAB2 and pghot-hintfaults but
runtime is better with pghot. This could be the benefit of async batched
migration showing and no adverse effect of losing cache locality.

pghot-hwhints shows good results. As I said this is just a first peek to the
experimental numbers, I should have more concrete numbers and conclusion in LPC.

[1] Kmigrated -
https://lore.kernel.org/linux-mm/20250616133931.206626-1-bharata@xxxxxxx/#t
[2] IBS Memory Profiler RFC v0 -
https://lore.kernel.org/linux-mm/20260924062206.319314-1-bharata@xxxxxxx/

Regards,
Bharata.