Re: [PATCH v2 2/4] mm: allow shared folios to be promoted to a fast tier

From: Bharata B Rao

Date: Fri Sep 18 2026 - 00:15:40 EST


> On Thu, Sep 17, 2026 at 12:18:24PM -0400, Gregory Price wrote:
> > On Thu, Sep 17, 2026 at 06:08:27PM +0200, Peter Zijlstra wrote:
> >
> > Honestly I'm starting to think hint faults are a big hammer making up
> > for the lack of hardware support for getting this data.
> >
> > Would be nice to just have the hardware report what's hot (and how hot)
> > rather than depending on a software-heuristic like deriving hotness
> > from a page fault.
>
> Yeah, there is/was this patch-set from AMD that uses their IBS counters
> for this, but 'ab'-using the performance counters for this also has ick.
> PMU data isn't ideal either. Mostly they generate a ton of data that
> needs to be analyzed as well. Its not clear cut and easy.

Here are my experiments with using IBS for memory access profiling:

1. My very early attempt was to replace NUMA hint faults with IBS
provided memory access information to drive NUMA Balancing. This
worked for all the modes of NUMA Balancing (kernel.numa_balancing=1
or 2) [1]

2. The next attempt was to use IBS data to drive only hot page promotion.
In this approach, IBS was used as one of the sources of page hotness
information to pghot (existing NUMA hintfaults being the other) [2]

Both the above approaches used the primary IBS instance that was being
used by perf sub-system also and I had made the use mutually exclusive.

>
> I'm not sure there's been proposals for better hardware support.

Then AMD Zen6 processors introduced a 2nd light-weight IBS instance
called IBS Memory Profiler which is separate, works independently
of the primary IBS instance (which continues to be used by perf)
and which is primarily targeted for memory access profiling.

I have used this as source of page hotness with pghot (pghot-hwhints)
and the benchmark results are encouraging. [3]

Also it is worth reiterating here that neither primary IBS nor this
new IBS Memory Profiler have got anything to do with PMU sub-system.

In this context, I would also like to point out that pghot patchset [4]
is trying to move hot page promotion from scheduler to its own dedicated
sub-system. I have done the following till now:

- Extracted out hot page promotion engine and moved it to pghot
so that the same gets used for other sources of page hotness.
- Moved fault-time migration to async and batched kernel-thread driven
migration (kmigrated)
- Used NUMA hint faults as page hotness source to pghot (pghot-hintfaults)

I have often wondered if it makes sense to move out complete hint faulting
mechanism out of scheduler but then I see that task-follows-memory part,
the scanning logic, fault stats heuristics are tightly tied to the scheduler.

Also apt is to remember the PTE-A bit scanning approach [5] that was started
as a potential replacement to NUMA hint faults based scanning. We are
planning to revive that effort and make it as another source for pghot
if we get good results with different benchmarks.

Regards,
Bharata.

[1] Primary IBS instance driving NUMA Balancing
https://lore.kernel.org/lkml/20230208073533.715-1-bharata@xxxxxxx/
[2] The last pghot version (v5) which used primary IBS instance as page hotness source
https://lore.kernel.org/linux-mm/20260129144043.231636-1-bharata@xxxxxxx/
[3] IBS Memory Profiler as page hotness source for pghot
https://lore.kernel.org/linux-mm/92c26cce-0608-4c0d-bb13-fe87afc225ba@xxxxxxx/T/#m6d17d0d58a5026c16d63c34e1abfd67a57a66256
[4] The last posted pghot (v8) patchset
https://lore.kernel.org/linux-mm/20260728054356.291998-1-bharata@xxxxxxx/
[5] Kscand - PTE A bit based scanning
https://lore.kernel.org/linux-mm/20250814153307.1553061-1-raghavendra.kt@xxxxxxx/