Re: [RFC PATCH v0 0/3] pghot: x86: IBS Memory Profiler for hot page promotion
From: Bharata B Rao
Date: Thu Sep 24 2026 - 02:28:25 EST
On 24-Sep-26 11:52 AM, Bharata B Rao wrote:
>
> Detailed per-benchmark tables (throughput/latency + vmstat and pghot
> promotion counters) are posted as replies to this thread.
Graph500 tiering comparison: base vs pghot (hwhints)
====================================================
Kernel(s) : base = 7.3.0-rc2-base+
pghot = 7.3.0-rc2-pghot+
Benchmark : Graph500 reference BFS, SCALE=28, edgefactor=16, 128 ranks
Topology : top-tier NUMA node=1 (CPUs+DRAM), lower-tier node=2 (mem-only/CXL)
Note : SKIP_VALIDATION=1 (timing-only); all TEPS carry Graph500 (!) flag.
Figure of merit is harmonic_mean_TEPS (hmean).
Config legend
-------------
C1 : base kernel, no tiering (numa_balancing=0)
C2 : base kernel, NUMAB2 tiering (numa_balancing=2)
C3 : pghot kernel, hwhints source (pghot_enabled_sources=2,
pghot_freq_threshold=1, pghot_target_nid=0 [default],
IBS mem-profiler on, numa_balancing=0, promotion=off)
Column legend
-------------
hmean : harmonic_mean_TEPS (primary Graph500 metric)
hstddev : harmonic_stddev_TEPS
median : median_TEPS
bfs_t : mean BFS time (seconds)
spdup : speedup of hmean vs C1 baseline
Table 1: Performance
--------------------
+--------+------------+----------+------------+---------+--------+
| Config | hmean | hstddev | median | bfs_t | spdup |
| | (TEPS) | (TEPS) | (TEPS) | (sec) | |
+--------+------------+----------+------------+---------+--------+
| C1 | 5.543e+08 | 5.51e+05 | 5.548e+08 | 7.748 | 1.00x |
| C2 | 1.298e+09 | 6.80e+07 | 1.394e+09 | 3.308 | 2.34x |
| C3 | 1.755e+09 | 2.32e+07 | 1.805e+09 | 2.447 | 3.17x |
+--------+------------+----------+------------+---------+--------+
Table 2: Relevant kernel counters (/proc/vmstat deltas over the run)
--------------------------------------------------------------------
Values are accumulated deltas (before -> after) for the whole run.
+----------------------------+-----------+-----------+-----------+
| Counter | C1 | C2 | C3 |
+----------------------------+-----------+-----------+-----------+
| numa_pte_updates | 0 | 25867526 | 0 |
| numa_hint_faults | 0 | 13318442 | 0 |
| numa_pages_migrated | 0 | 13318248 | 3702779 |
| pgpromote_success | 0 | 13317996 | 3702779 |
| pghot_recorded_accesses | 0 | 0 | 3709015 |
| pghot_reported_hintfaults | 0 | 0 | 0 |
| pghot_reported_hwhints | 0 | 0 | 21385228 |
| hwhint_total_events | 0 | 0 | 21385281 |
| hwhint_dram_accesses | 0 | 0 | 16747814 |
| hwhint_extmem_accesses | 0 | 0 | 3697760 |
| hwhint_cache_accesses | 0 | 0 | 0 |
| hwhint_useful_events | 0 | 0 | 21385228 |
| hwhint_dropped_events | 0 | 0 | 0 |
| pgmigrate_success | 26841377 | 40161043 | 30548434 |
+----------------------------+-----------+-----------+-----------+
Key findings
------------
1. Tiering is the dominant win: C2 (base NUMAB2) reaches 2.34x and C3
(pghot hwhints) 3.17x over the untiered baseline (C1), where the
working set is stranded on the lower-tier/CXL node 2.
2. pghot hwhints now clearly leads base NUMAB2 on the official metric:
C3 1.755e9 (3.17x) vs C2 1.298e9 (2.34x) -> ~+35% hmean_TEPS.
3. hwhints/IBS is far more efficient and stable. C3 reaches its higher
hmean while:
- promoting only ~3.70M pages, ~1/4 of C2 (13.32M);
- issuing zero NUMA hint faults / PTE scans (numa_pte_updates=0);
- being ~3x more consistent (hstddev 2.32e7 vs 6.80e7).
IBS reported 21.39M hwhint events (16.75M DRAM + 3.70M ext-mem),
with hwhint_dropped_events=0.
4. median vs harmonic-mean: C2 median (1.394e9) sits well above its
hmean (1.298e9), reflecting high per-BFS variance in the fault-driven
path. C3 median (1.805e9) and hmean (1.755e9) sit close together ->
low variance.
Caveats
-------
* Single run per configuration; the median-vs-hmean spread for C2
indicates non-trivial run-to-run variance. Repeat runs are advisable
before drawing firm quantitative conclusions.
* SKIP_VALIDATION=1 was used (timing-only), so TEPS values are flagged
invalid (!) by Graph500 and are intended for relative comparison only.