Re: [RFC PATCH v0 0/3] pghot: x86: IBS Memory Profiler for hot page promotion

From: Bharata B Rao

Date: Thu Sep 24 2026 - 02:39:34 EST


On 24-Sep-26 11:52 AM, Bharata B Rao wrote:
>
> Detailed per-benchmark tables (throughput/latency + vmstat and pghot
> promotion counters) are posted as replies to this thread.
==========================================================================
Redis + memtier: hot-page promotion on a CXL-tiered system
==========================================================================

Benchmark: an in-memory Redis server is loaded with a ~64 GB dataset
(62.2M keys x 1 KB) whose pages are then explicitly migrated to the CXL
node (lower tier); the Redis server and the memtier client run on a top-
tier DRAM node (node 1). The measurement phase drives GET traffic with
memtier (16 threads x 100 conns, PASSES=12) over 50% of the keyspace so
that repeatedly-accessed lower-tier pages become promotion candidates.
Promotions target node 1 (local to the accessing threads): NUMAB uses
the accessing node, pghot uses pghot_target_nid=1.

System : AMD Zen6. Nodes 0,1 = DRAM (top tier);
node 2 = CXL (lower tier). Redis/client pinned to node 1.
Kernels: base=7.3.0-rc2-base+ pghot=7.3.0-rc2-pghot+

Cases:
C1 base/NUMAB0 base kernel, promotion OFF (baseline)
C2 base/NUMAB2 base kernel, NUMAB tiering promotion (hint
faults); promotes to local accessing node
C3 pghot/hwhints-10k pghot, NUMAB=0, source=IBS mprof,
period=10000, freq_thr=1, target_nid=1
C4 pghot/hwhints-5k as C3 but IBS period=5008 (kernel min,
~2x sampling; 5008)

==========================================================================
Table 1: Benchmark metrics (memtier)
==========================================================================
Case Ops/sec vs C1 Avg p50 p99 p99.9
latency (ms) ->
----------------------------------------------------------------------
C1 base/NUMAB0 285,288 ref 179.46 177.15 344.06 358.40
C2 base/NUMAB2 295,475 +3.57% 173.18 168.96 325.63 364.54
C3 pghot/hwhints-10k 285,923 +0.22% 178.91 177.15 344.06 358.40
C4 pghot/hwhints-5k 285,290 +0.00% 179.24 177.15 346.11 358.40
----------------------------------------------------------------------

==========================================================================
Table 2: Page-migration / hotness metrics (vmstat delta)
==========================================================================
Legend: C1=base/NUMAB0 C2=base/NUMAB2 C3=hwhints p=10000
C4=hwhints p=5008 (both hwhints: target_nid=1)
('-' = counter not present on base kernel)

metric C1 C2 C3 C4
-----------------------------------------------------------------
pgpromote_success 0 10,433,506 575,010 1,107,371
numa_pte_updates 0 20,333,730 0 0
numa_hint_faults 0 10,433,506 0 0
numa_pages_migrated 0 10,433,506 575,006 1,107,369
pgmigrate_success 0 10,433,506 575,006 1,107,369
pghot_recorded_accesses - - 575,677 1,110,760
pghot_reported_hwhints - - 964,304 1,967,044
hwhint_total_events - - 964,324 1,967,078
hwhint_dram_accesses - - 388,027 854,679
hwhint_extmem_accesses - - 575,674 1,110,758
hwhint_useful_events - - 964,304 1,967,044
pgdemote_kswapd 0 0 0 0
-----------------------------------------------------------------

==========================================================================
Key observations
==========================================================================
1. NUMAB tiering promotion (C2) helps only marginally: 295,475 vs 285,288
ops/sec (+3.57%), avg latency 179.5 -> 173.2 ms, promoting the full hot set
(10.43M pages / ~39.8 GiB) to the local node.

2. pghot with the IBS hwhints source is flat vs baseline at both periods
(C3 +0.22%, C4 +0.00%). numa_pte_updates / numa_hint_faults are
0 (no NUMA balancing); promotion is purely hardware-sample driven.

3. Sampling density scales promotion linearly (period 10000 -> 5008
doubles reported hwhints 964,304 -> 1,967,044 and promotions 575,010 ->
1,107,371 pages, ~2.19 -> ~4.22 GiB); ext-mem/CXL samples map ~1:1 to
promotions (freq_threshold=1), dram/already-toptier samples are not
promotable.

4. But even ~4.22 GiB is only ~10% of the ~40 GiB hot set, so throughput
does not move. Unlike C2 (which promotes the whole hot set), sampling-
based hwhints at these periods covers too little of the working set within
the run. Reaching C2's gain needs far denser sampling and/or a longer run
so hwhints promotes a large fraction of the hot set.