Re: [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure - llama-bench

From: Bharata B Rao

Date: Tue Jul 28 2026 - 02:27:18 EST


llama-bench is the micro-benchmark shipped with llama.cpp. It loads a GGUF
model through the production inference path and reports two throughput numbers
per run: pp512 (one batched 512-token prefill, compute-bound) and tg128 (128
sequential autoregressive decode steps, bandwidth-bound).

Workload: llama-bench (llama.cpp), Mixtral-8x22B-Instruct Q4_K_M (140.6B
params, 79.7 GiB), -t 64 -p 512 -n 128 -r 5 --mmap 0. pp512 = prefill
(compute-bound); tg128 = decode (memory-bandwidth-bound).

Setup: 3-node box, N0/N1 DRAM (256 GiB each), N2 CXL (distances 10/12/50).
Bench pinned to N1 CPUs with MPOL_PREFERRED_MANY({1}); a 200 GiB hot hog on
N1 forces kswapd to naturally demote ~1/3 of the model to CXL; each mode is
then measured while the hog keeps N1 under pressure. pghot target_nid=1,
freq_threshold=1, freq_window=3000ms, rate_limit=65536 MBps, kmigrated
100ms/512. hwhints runs arm AMD IBS (l3miss-only=1, period=10000).

Legend (columns)
----------------
r1 = base / notier 7.2.0-rc2-base nb=0 src=-
r2 = base / tier (NUMAB2) 7.2.0-rc2-base nb=2 src=-
r3 = pghot / hintfaults 7.2.0-rc2-pghot nb=2 src=0x1
r4 = pghot / hwhints(IBS) 7.2.0-rc2-pghot nb=0 src=0x2
r5 = pghot / both 7.2.0-rc2-pghot nb=2 src=0x3
(nb = kernel.numa_balancing; src = pghot_enabled_sources)

Table 1 - Throughput (llama-bench, tokens/s over 5 reps)
-------------------------------------------------------
metric r1 r2 r3 r4 r5
-------------------------------------------------------------------
pp512 t/s 69.31 62.46 54.53 53.95 51.49
stddev 2.70 6.65 10.41 8.12 4.79
tg128 t/s 4.262 4.625 5.993 5.161 6.851
stddev 0.004 0.313 1.255 0.535 1.807
tg128 vs r1 1.00x 1.09x 1.41x 1.21x 1.61x
pp512 vs r1 1.00x 0.90x 0.79x 0.78x 0.74x

Table 1b - tg128 per-iteration (t/s), shows convergence
-------------------------------------------------------
iter r1 r2 r3 r4 r5
-------------------------------------------------------------------
iter1 4.268 4.183 4.393 4.462 4.415
iter2 4.261 4.457 5.336 4.820 5.602
iter3 4.260 4.678 5.995 5.218 7.322
iter4 4.261 4.841 6.501 5.494 8.152
iter5 4.259 4.969 7.740 5.813 8.763

Table 2 - Key vmstat counters, run-phase deltas (millions of pages/events)
--------------------------------------------------------------------------
metric r1 r2 r3 r4 r5
------------------------------------------------------------------
pgpromote_success 0.00 3.05 6.95 2.17 6.53
pgpromote_candidate 0.00 14.34 24.12 13.83 21.09
pgdemote_kswapd 0.69 5.02 10.84 3.89 8.05
pgmigrate_success 0.69 8.07 17.80 6.06 14.59
numa_pte_updates 0.00 28.77 24.56 0.00 12.04
numa_hint_faults 0.00 27.77 24.12 0.00 11.39
pghot_recorded_accesses 0.00 0.00 24.12 13.83 21.28
pghot_recorded_hintfaults 0.00 0.00 24.12 0.00 11.39
pghot_recorded_hwhints 0.00 0.00 0.00 34.03 32.42
hwhint_total_events 0.00 0.00 0.00 34.03 32.42
hwhint_dram_accesses 0.00 0.00 0.00 20.03 22.34
hwhint_extmem_accesses 0.00 0.00 0.00 13.83 9.89

Key observations
----------------
1. pghot-both (both hintfaults and hwhints enabled) improves decode over both
base baselines.
2. Prefill (pp512) suffers from tiering promotion (0.79x hintfaults, 0.78x hwhints,
0.74x both vs notier); pghot/hf and /hw pp512 are noisy (stddev ~8-10).