Re: [PATCH 2/2] perf: arm_spe: Prefer large AUX mappings
From: Dev Jain
Date: Thu Oct 08 2026 - 06:45:38 EST
On 10/08/26 11:11 pm, Leo Yan wrote:
> On Mon, Aug 10, 2026 at 04:10:48PM +0100, Will Deacon wrote:
>> On Mon, Aug 10, 2026 at 03:44:42PM +0100, Leo Yan wrote:
>>> Commit 18049c8cff9c ("perf/aux: Allocate non-contiguous AUX pages by
>>> default") made the AUX allocator use order-0 pages by default unless a
>>> PMU explicitly asks for contiguous allocations.
>>
>> But that commit specifically calls out SPE as benefitting from
>> non-contiguous pages:
>>
>> "For instance, ARM SPE and TRBE operate with virtual pages, and
>> Coresight ETR allocates a separate buffer. For these PMUs,
>> allocating contiguous AUX pages unnecessarily exacerbates memory
>> fragmentation. This fragmentation can prevent their use on
>> long-running devices."
>>
>> so why doesn't passing PERF_PMU_CAP_AUX_PREFER_LARGE reintroduce the
>> problems that 18049c8cff9c was trying to solve?
>
> The question is how "allocating contiguous AUX pages unnecessarily
> exacerbates memory fragmentation." The relevant information I could find
> is [1]:
>
> "On Android, we collect ETM data periodically on internal user devices
> for AutoFDO optimization (for both userspace libraries and the
> kernel). Allocating a large chunk of contiguous AUX pages (4M for each
> CPU) periodically is almost unbearable. The kernel may need to kill
> many processes to fulfill the request. It affects user experience even
> after using PMU."
>
> We might have missed chance to clarify how the fragmentation issue
> occurs in the first place. Let's say, a phone with 8 CPUs, allocating
> 4MB per CPU requires 32MB in total, which is a relatively small
> portion of 4GiB or 8GiB of RAM commonly found in phones. Moreover, once
> contiguous pages are freed, the buddy allocator can coalesce them
> again into buddy list. It is not obvious to me that PREFER_LARGE
> directly causes fragmentation.
>
> One case where AUX allocation could exacerbate fragmentation is when the
> system is already fragmented. If a high-order allocation fails and the
> allocator falls back to smaller-order blocks, those allocations may
> consume free blocks scattered across different buddy regions and make
> subsequent high-order allocations more difficult.
>
> If this is the main concern, I'd suggest using a smaller AUX buffer
> (e.g. 1MB or even 512KB) for TRBE/SPE to reduce memory pressure.
> Snapshot mode '-S' could also be considered, as it allows the buffer to
> be allocated once and reused for subsequent recordings by signals.
>
> OTOH, using only order-0 pages can significantly increase TTW overhead
> on the trace path and lead to overflows, we observe this causes huge
> trace discontinuity. In the end, we need to trace-off the fragmentation
> concern against the trace discontinuity.
>
> Thanks,
> Leo
I don't have the full context here (can look in detail later) so drive-by comment:
In general, doing large order allocations should *reduce* fragmentation. If you allocate
16 order-0 pages as opposed to an order-4 compound page, you may end up spreading your
allocations across different blocks, thus preventing them from merging later.
A recent example can be found in vmalloc [1].
[1] https://lore.kernel.org/all/aPjrRkjiIt6HmXmT@xxxxxxxxxxxxxxxxxxxx/
>
> [1] https://lore.kernel.org/lkml/CALJ9ZPNLgEBxOmDim-vztUknEETwdL-Z2gJ8K9s44TiPgKZgHg@xxxxxxxxxxxxxx/