Re: [RFC PATCH v0 0/3] pghot: x86: IBS Memory Profiler for hot page promotion

From: Bharata B Rao

Date: Thu Sep 24 2026 - 02:36:24 EST


On 24-Sep-26 11:52 AM, Bharata B Rao wrote:
>
> Detailed per-benchmark tables (throughput/latency + vmstat and pghot
> promotion counters) are posted as replies to this thread.
llama.cpp tiering benchmark: base vs pghot-hwhints (kernel 7.3.0-rc2)
====================================================================
3-run average per configuration.

Workload: llama-bench (llama.cpp), Mixtral-8x22B-Instruct Q4_K_M (140.6B
params, 79.7 GB), -t 64 -p 512 -n 128 -r 5 --mmap 0. pp512 = prefill
(compute-bound); tg128 = decode (memory-bandwidth-bound).

Setup: 3-node box, N0/N1 DRAM (256 GB each), N2 CXL (distances 10/12/50).
Bench pinned to N1 CPUs with MPOL_PREFERRED_MANY({1}); a 200 GB hot hog on
N1 forces kswapd to naturally demote ~1/3 of the model to CXL; each mode is
then measured while the hog keeps N1 under pressure. pghot target_nid=1,
freq_threshold=1, promote_window=3000ms, rate_limit=65536 MBps, kmigrated
100ms/512. hwhints arms AMD IBS (l3miss-only=1, period=10000).

Legend (columns)
----------------
r1 = base / notier 7.3.0-rc2-base nb=0 src=-
r2 = base / tier (NUMAB2) 7.3.0-rc2-base nb=2 src=-
r3 = pghot / hwhints(IBS) 7.3.0-rc2-pghot nb=0 src=0x2
(nb = kernel.numa_balancing; src = pghot_enabled_sources; 3 runs each)

Table 1 - Throughput, 3-run mean +/- stdev (llama-bench tokens/s)
----------------------------------------------------------------
metric r1 r2 r3
-------------------------------------------------
pp512 mean 69.43 61.34 57.90
stdev 0.85 1.78 2.21
tg128 mean 4.094 4.883 4.975
stdev 0.045 0.053 0.023
tg128 vs r1 1.00x 1.19x 1.22x
pp512 vs r1 1.00x 0.88x 0.83x

Table 1b - per-run values (3 runs), shows consistency
-----------------------------------------------------
r1 r2 r3
-------------------------------------------------
tg128 run1 4.099 4.856 4.949
tg128 run2 4.037 4.836 5.005
tg128 run3 4.146 4.957 4.970
pp512 run1 70.37 63.37 61.01
pp512 run2 68.31 59.03 56.61
pp512 run3 69.60 61.64 56.07

Table 2 - Key vmstat counters, 3-run mean of run-phase deltas (millions)
-----------------------------------------------------------------------
metric r1 r2 r3
------------------------------------------------------
pgpromote_success 0.00 3.07 1.74
pgpromote_candidate 0.00 12.80 12.86
pgdemote_kswapd 2.58 4.78 4.40
pgmigrate_success 2.58 7.85 6.13
numa_pte_updates 0.00 26.06 0.00
numa_hint_faults 0.00 25.12 0.00
pghot_recorded_accesses 0.00 0.00 12.86
pghot_reported_hwhints 0.00 0.00 30.79
hwhint_total_events 0.00 0.00 30.79
hwhint_dram_accesses 0.00 0.00 17.74
hwhint_extmem_accesses 0.00 0.00 12.86
start N2 % (at SIGCONT) 28.3 28.9 27.6

Key observations (3-run averages)
---------------------------------
1. base/tier (r2) tg128 4.883 is +19.3% over base/notier (r1)
4.094; pghot/hwhints (r3) 4.975 is +21.5% over r1 and
+1.9% over r2.
2. Consistency is tight: r3 tg128 4.949-5.005 (stdev 0.023);
r2 4.836-4.957 (stdev 0.053).
3. hwhints armed IBS: ~35M events, ~13M on CXL; pgpromote_candidate
12.9M matches pghot_recorded_accesses 12.9M (1.00x).
4. Prefill (pp512) pays a tiering tax: r2 0.88x, r3 0.83x of r1.

Caveats: base and pghot are different kernels (7.3.0-rc2-base vs -pghot);
the natural-overflow setup settled at ~30% CXL at SIGCONT in all runs;
3 runs per configuration (per-run values in Table 1b).