Re: [PATCH] cpufreq: Don't track pressure if policy->max equals top of freq_table

From: Christian Loehle

Date: Fri Oct 09 2026 - 05:40:46 EST


On 10/9/26 09:29, Giovanni Gherdovich wrote:
> The acpi-cpufreq driver stores non-boost frequencies in
> policy->freq_table, but policy->cpuinfo.max_freq contains the max
> boost frequency.
>
> policy->max is constrained to be, at most, the largest value in
> freq_table. This means that under nominal conditions, policy->max is
> substantially lower than policy->cpuinfo.max_freq, resulting in large
> values of cpufreq_pressure, on all CPUs, for no reason.
> The result is large performance regressions on compute/memory intensive
> benchmarks that use a limited number of CPUs; tasks migrate more and
> lose locality. My tests are with the "NAS Parallel Benchmarks" suite,
> limiting to 1/4 of available CPUs.
>
> This change sets cpufreq_pressure to zero if policy->max is equal
> (or above) the maximum frequency in freq_table.
>
> Here a concrete example from a 1st generation AMD EPYC (Naples) using
> acpi-cpufreq. Under nominal conditions (no capping, no throttling):
>
> freq_table : 1200 MHz, 1700 MHz, 2200 MHz
> policy->max : 2200 MHz (maximum from freq_table)
> policy->cpuinfo.max_freq : 3200 MHz (max boost)
>
> Before this change, in the example above we'd get a cpufreq_pressure
> value of 320 (32% pressure); with the change, the pressure is zero as
> it's expected to be.
>
> Signed-off-by: Giovanni Gherdovich <ggherdovich@xxxxxxx>
> ---
> Clarifying the recipients list, beyond the cpufreq maintainers:
>
> Mario Limonciello, Huang Rui: the patch isn't for amd-pstate, but
> AMD EPYC up to Genoa (Zen4) defaults to acpi-cpufreq, so they may come
> across this if they haven't already.
> Vincent Guittot: for the cpufreq_pressure and scheduler load balancer
> implications.
> Ricardo Neri: he's been looking at cpufreq_pressure on x86, albeit
> this patch isn't for intel_pstate.
> Pierre Gondois: he looked at policy->max in freq_table.c earlier this year.


I'm assuming this isn't on top of
f341dc4d934a ("cpufreq: intel_pstate: Fix max_freq fallback in cpufreq_update_pressure()")
Care to retry?

>
> drivers/cpufreq/cpufreq.c | 3 +++
> drivers/cpufreq/freq_table.c | 1 +
> include/linux/cpufreq.h | 1 +
> 3 files changed, 5 insertions(+)
>
> diff --git a/drivers/cpufreq/cpufreq.c b/drivers/cpufreq/cpufreq.c
> index 54dde8419bdc..aa652750f711 100644
> --- a/drivers/cpufreq/cpufreq.c
> +++ b/drivers/cpufreq/cpufreq.c
> @@ -2601,6 +2601,9 @@ static void cpufreq_update_pressure(struct cpufreq_policy *policy)
> */
> if (max_freq <= capped_freq) {
> pressure = 0;
> + } else if (policy->freq_table && policy->max_table_freq &&
> + policy->max_table_freq <= capped_freq) {
> + pressure = 0;
> } else {
> max_capacity = arch_scale_cpu_capacity(cpu);
> pressure = max_capacity -
> diff --git a/drivers/cpufreq/freq_table.c b/drivers/cpufreq/freq_table.c
> index ea994647abc8..55c3293685c2 100644
> --- a/drivers/cpufreq/freq_table.c
> +++ b/drivers/cpufreq/freq_table.c
> @@ -50,6 +50,7 @@ int cpufreq_frequency_table_cpuinfo(struct cpufreq_policy *policy)
> }
>
> policy->cpuinfo.min_freq = min_freq;
> + policy->max_table_freq = max_freq;
> /*
> * If the driver has set its own cpuinfo.max_freq above max_freq, leave
> * it as is.
> diff --git a/include/linux/cpufreq.h b/include/linux/cpufreq.h
> index d3d0d9d02aa4..4a3316e47f4e 100644
> --- a/include/linux/cpufreq.h
> +++ b/include/linux/cpufreq.h
> @@ -67,6 +67,7 @@ struct cpufreq_policy {
> unsigned int max; /* in kHz */
> unsigned int cur; /* in kHz, only needed if cpufreq
> * governors are used */
> + unsigned int max_table_freq; /* max freq in the table */
> unsigned int suspend_freq; /* freq to set during suspend */
>
> unsigned int policy; /* see above */