Re: [PATCH 2/2] perf: arm_spe: Prefer large AUX mappings

From: Leo Yan

Date: Mon Oct 05 2026 - 13:54:47 EST


On Thu, Oct 01, 2026 at 08:16:09AM +0100, Will Deacon wrote:
> On Wed, Sep 30, 2026 at 05:43:11PM +0100, Leo Yan wrote:

> > How about adding a field to struct pmu to specify a preferred maximum
> > page order for the AUX buffer? The perf core could try that order first
> > and fall back to smaller orders if the allocation fails.
>
> I'm not sure that's thr right place for it, really. The driver has no
> clue about whether it makes sense to use large contiguous mappings or
> not, so I'd have thought that decision should be driven from userspace
> (e.g. like MADV_HUGEPAGE) because it really depends on the user's
> preference and isn't a fixed property of the hardware.

Here MADV_HUGEPAGE cannot directly apply on this case: perf allocates
the AUX pages during mmap, while TRBE accesses them through a separate
kernel vmap() mapping.

MADV_HUGEPAGE is applied after mmap, but a preference (or flag) would
need to be specified before the AUX mmap.

> > For example, the Neoverse V2 TRM documents:
> >
> > L1 Trace Buffer Extension (TRBE) TLB: 1 entry
>
> Wow, they really pulled out the stops for that implementation. I bet
> we're supposed to be grateful for that entry!

Yeah, Neoverse V3 was improved to have two entries. Even so, I was told
it still suffers from TLB misses, so still needs a large mapping
granule.

> > Given the single L1 TRBE TLB entry, the TRBE driver could prefer
> > PMD_ORDER (2 MiB with 4 KiB pages) to reduce TLB pressure. This reflects
> > the hardware characteristic.
> >
> > This could be a trade-off instead of using PERF_PMU_CAP_AUX_PREFER_LARGE,
> > avoiding large contiguous allocations that could reintroduce the Android
> > OOM issue. I did a quick test with this approach and the results look
> > positive.
>
> I really don't want the driver to second-guess userspace based on whatever
> information it happens to have hard-coded about the specific CPU it's
> running on.

The kernel already takes the PERF_PMU_CAP_AUX_PREFER_LARGE flag from a
PMU; a preferred maximum order would let the driver give it a bounded
value.

I do not intend to hard-code or guess a preference for a particular CPU
variant. We can map TRBE or SPE buffer at PMD granularity. On a 4
KiB-page system, PMD_ORDER is order 9, or 2 MiB. Requesting a larger
contiguous chunk cannot increase the mapping granule, so I would cap the
preference there. It remains a preference: perf can fall back to smaller
orders when allocation fails.

Exposing the preference to userspace also seems problematic. Users
generally lack the hardware details needed to choose an appropriate
value. Even if tools provide a default, the same policy would need to
be duplicated across perf, simpleperf, and proprietary tools. I don't
think userspace tools are the right place for this policy.

Thanks,
Leo