Re: [PATCH 2/2] perf: arm_spe: Prefer large AUX mappings
From: Will Deacon
Date: Thu Oct 01 2026 - 03:16:35 EST
On Wed, Sep 30, 2026 at 05:43:11PM +0100, Leo Yan wrote:
> On Mon, Aug 10, 2026 at 04:10:48PM +0100, Will Deacon wrote:
> > On Mon, Aug 10, 2026 at 03:44:42PM +0100, Leo Yan wrote:
> > > Commit 18049c8cff9c ("perf/aux: Allocate non-contiguous AUX pages by
> > > default") made the AUX allocator use order-0 pages by default unless a
> > > PMU explicitly asks for contiguous allocations.
> >
> > But that commit specifically calls out SPE as benefitting from
> > non-contiguous pages:
> >
> > "For instance, ARM SPE and TRBE operate with virtual pages, and
> > Coresight ETR allocates a separate buffer. For these PMUs,
> > allocating contiguous AUX pages unnecessarily exacerbates memory
> > fragmentation. This fragmentation can prevent their use on
> > long-running devices."
> >
> > so why doesn't passing PERF_PMU_CAP_AUX_PREFER_LARGE reintroduce the
> > problems that 18049c8cff9c was trying to solve?
>
> How about adding a field to struct pmu to specify a preferred maximum
> page order for the AUX buffer? The perf core could try that order first
> and fall back to smaller orders if the allocation fails.
I'm not sure that's thr right place for it, really. The driver has no
clue about whether it makes sense to use large contiguous mappings or
not, so I'd have thought that decision should be driven from userspace
(e.g. like MADV_HUGEPAGE) because it really depends on the user's
preference and isn't a fixed property of the hardware.
> For example, the Neoverse V2 TRM documents:
>
> L1 Trace Buffer Extension (TRBE) TLB: 1 entry
Wow, they really pulled out the stops for that implementation. I bet
we're supposed to be grateful for that entry!
> Given the single L1 TRBE TLB entry, the TRBE driver could prefer
> PMD_ORDER (2 MiB with 4 KiB pages) to reduce TLB pressure. This reflects
> the hardware characteristic.
>
> This could be a trade-off instead of using PERF_PMU_CAP_AUX_PREFER_LARGE,
> avoiding large contiguous allocations that could reintroduce the Android
> OOM issue. I did a quick test with this approach and the results look
> positive.
I really don't want the driver to second-guess userspace based on whatever
information it happens to have hard-coded about the specific CPU it's
running on.
Will