Re: [RFC PATCH v0 0/3] pghot: x86: IBS Memory Profiler for hot page promotion

From: Bharata B Rao

Date: Thu Sep 24 2026 - 02:34:46 EST


On 24-Sep-26 11:52 AM, Bharata B Rao wrote:
>
> Detailed per-benchmark tables (throughput/latency + vmstat and pghot
> promotion counters) are posted as replies to this thread.
======================================================================
Tiered-memory promotion: baseline vs NUMA-balancing vs DAMON vs pghot
hwhints (IBS) - multithreaded runs
======================================================================

Runs taken on one machine, same workload and sizes. This is a self-contained
comparison of four ways to promote a hot working set from a slow (CXL) tier
to a fast (DRAM) tier:

- baseline : no promotion.
- NUMAB2 : the kernel's NUMA-balancing tier promotion
(numa_balancing=2), driven by NUMA hint faults.
- DAMON fine+quota: in-kernel access monitor with a migrate_hot
scheme (paddr), driven by the kdamond thread.
- pghot hwhints : per-PFN hotness from the IBS Memory Profiler
(hardware sampling), promoted by the kmigrated
thread; shown at two IBS sample periods.

Machine / topology
------------------
- 256 CPUs, 792 GB RAM.
- NUMA: node 0, node 1 = DRAM (fast tier); node 2 = CXL (slow tier,
256 GB). Promotion target = node 0; cold data starts on node 2.
- Kernels: base = 7.3.0-rc2-base+ (baseline, NUMAB2, DAMON);
pghot = 7.3.0-rc2-pghot+ (pghot hwhints).

Workload
--------
- 64-thread pointer chase, 64 GB total buffer, 16 GB hot set, 240 s
timed phase. Buffer relocated to node 2 before the timed phase;
worker threads bound to node 0. "hot set" = 16 GB;
over = (node0 - 16 GB) / 16 GB.

Configurations
--------------
- baseline : numa_balancing=0, no promotion.
- NUMAB2 : numa_balancing=2 (tier promotion via hint faults).
- DAMON : sysfs defaults (sample 5ms, aggr 100ms, update 60s) +
migrate_hot (paddr, nr_accesses>=1) + max_nr_regions=
100000 (min default 10) + quota 128 MiB/1000 ms with
nr_accesses weighting (hottest first).
- hwhints : IBS Memory Profiler as the sole hotness source
(numa_balancing=0), pghot_freq_threshold=1, IBS
l3miss-only=1; IBS sample period 10000 and 5008.

Metric notes
------------
- lat : steady access latency, ns/access, mean of last 25% of samples
(lower is better).
- node0: end-of-run placement on node 0 (numastat -p).
- promoted: bytes migrated to the fast tier (DAMON scheme sz_applied;
NUMAB2/pghot pgpromote_success).
- CPU : cumulative cputime over the 240 s run of the single promoter/
monitor thread - kd = kdamond (DAMON), km = kmigrated (pghot).
NUMAB2 has no single promoter thread; its cost is distributed
across the application threads (scanning + hint faults).
- over: The "over" column is measured against the 16 GB HOT SET, not
against B0: over = (node0 - 16 GB) / 16 GB. B0 promotes
nothing, so its node0 is 0.0 and its over is n/a.


Table. Multithreaded results (64 GB total / 16 GB hot)
------------------------------------------------------
id config kern lat node0 over promoted promoter
(ns) (GB) (GB) CPU
--- ------------------ ----- --- ----- ----- -------- ----------
B0 baseline base 320 0.0 - - -
N2 NUMAB2 base 143 16.0 0% ~14.6* distributed
D3 DAMON fine+quota base 119 23.9 +49% 24.7 kd 13s(5%)
H10 hwhints IBS p10000 pghot 215 4.9 -69% 4.9 km 20s(8%)
H5 hwhints IBS p5008 pghot 92 11.6 -27% 12.0 km 38s(16%)

* N2 pgpromote_success over the timed window is ~14.6 GB; some
promotion happens during the load phase before the vmstat baseline,
so the reliable figure is the placement: node 0 = 16.0 GB (exact).

Supporting counters
-------------------
- N2 : numa_hint_faults 3.73M, numa_pte_updates 12.9M.
- D3 : sz_applied 24.7 GB, qt_exceeds 229, regions settled ~21-71.
- H10: hwhint_total_events 1.40M, pgpromote_success 1.29M.
- H5 : hwhint_total_events 5.91M, pgpromote_success 3.15M.

Findings
--------
1. Accuracy (placement):
- NUMAB2 places exactly the 16 GB hot set (0% over).
- DAMON over-promotes (+49%, 23.9 GB): a coarse-then-quota scheme
bounds volume but still moves ~1.5x the hot set.
- hwhints under-promotes within 240 s (-69% at p10000, -27% at
p5008) because hardware sampling observes only a fraction of
accesses; it had not fully converged at 240 s.

2. IBS sample period is the coverage knob. 10000 -> 5008 roughly
tripled events (1.40M -> 5.91M), nearly tripled placement (4.9 ->
11.6 GB) and cut latency 215 -> 92 ns. period=10000 under-samples
this workload.

3. Latency does not track placement linearly:
- hwhints p5008 reaches the lowest latency (92 ns) with only 11.6 GB
on node 0 - IBS targets the hottest pages, so its smaller
placement is well-chosen.
- DAMON needs 23.9 GB to reach 119 ns.
- NUMAB2 places the exact 16 GB but its steady latency (143 ns) is
higher than its own first interval (119 ns): numa_balancing keeps
scanning and faulting the workload's pages after promotion, an
ongoing per-access tax (12.9M pte updates, 3.73M hint faults).

4. Promoter CPU:
- DAMON kdamond and pghot kmigrated are both out-of-band and cheap
here (kd 13 s; km 20-38 s over 240 s, i.e. ~8-16% of one core).
- NUMAB2's cost is not a single thread; it is paid by the workload
threads via scanning and hint-fault handling.

Caveats
-------
- Two kernels: base (B0/N2/D3) and pghot (H10/H5). No separate
no-promotion baseline was taken on the pghot kernel in this set, so
cross-kernel absolute latency should be read with that in mind;
placement / over-promotion is the primary accuracy metric.
- The CPU column is the concentrated promoter/monitor thread only. It
excludes distributed costs: NUMAB2's scanning + hint-fault handling,
and IBS interrupt/sample handling for hwhints.
- Latency is not isolated to a single cause (bandwidth, working set,
selection quality all plausible); placement is the primary accuracy
metric and latency is supporting.
- hwhints coverage is still ~27% under the 16 GB hot set at 240 s
(p5008); it was still climbing at the end of the run.