Re: [PATCH] perf/arm-cmn: Allow userspace to select the PMU's CPU
From: Robin Murphy
Date: Fri Oct 02 2026 - 12:45:10 EST
On 30/09/2026 8:29 pm, Okanovic, Haris wrote:
Hi Robin,
All of the PMU's recurring work therefore lands on one CPU. This is
problematic on systems which reserve particular CPUs for latency
sensitive work or confine background activity to a chosen set of
housekeeping CPUs.
There is no way to set it explicitly. perf_event_open() and 'perf stat
-C' have no effect because arm_cmn_event_init() overwrites event->cpu;
/proc/irq/*/smp_affinity is refused for the DTC interrupts, which are
requested with IRQF_NOBALANCING because their affinity has to follow the
owning CPU.
All system PMU drivers have the same concern in this regard - why
should arm-cmn be special?
As I mentioned earlier, we can improve performance on certain system
topologies by assigning PMU work to housekeeping CPUs.
arm-cmn has no constraints around CPU assignment, so this is possible.
Every register access is MMIO and the DTC interrupts can be affinitised
anywhere, so any online CPU can own it. That's what makes "write any
online CPU" a sound interface here. I've moved it to CPUs in both NUMA
nodes on the two platforms I tested.
You're right that the concern is general, but a single interface isn't
easily shared, because the set of CPUs that a PMU may be driven from
is platform/device-specific:
arm_dsu_pmu, for instance, can only be driven from the CPUs attached to
the DSU, which it already exposes as "associated_cpus" and enforces in
event_init(); hisi_uncore_pmu publishes the same attribute.
Intel uncore is per-die: MSR-accessed boxes must be read from a CPU on
the target die, since rdmsr reads the executing CPU. Accepting an
arbitrary CPU there would silently read a different die's counters.
You seem to have missed my point. Pretty much all system/uncore PMU drivers - other than the trivial ones with non-programmable free-running counters and no interrupt - have to pick a CPU to associate with, irrespective of whether it's from the whole system or some specific subset, and they all do so effectively arbitrarily, whether that's by cpumask_local_spread(), cpumask_any(), cpumask_any_and() or whatever. If one driver picking an arbitrary CPU is a problem that needs fixing, why do the dozens of other drivers which also pick an arbitrary CPU not also need fixing?
And the answer is that they do! Irrespective of whether systems might want particular CPUs to do particular things, we've already seen that in systems with lots of PMU instances, when they all end up on the same default "any CPU", it ends up severely over-serialising event scheduling to a degree that can start to significantly impact event runtimes.
Do you have an alternate API in mind?
I don't see this being practical without first generalising the notion of CPU affinity so that it can at least be dealt with from a single place in the perf core. Inconsistent, ad-hoc hacks in individual drivers will not scale or be maintainable. Perf core also then has a reasonable chance of being able to reliably avoid racing against itself, whereas attempting to mitigate racy calls from within those calls when it's already too late is never really going to be a proper solution.
Thanks,
Robin.