Re: [PATCH] sched/fair: Only apply cpufreq pressure where frequency is invariant
From: jong wu
Date: Mon Sep 07 2026 - 10:59:23 EST
在 2026/9/7 10:32, Hongyan Xia 写道:
On 9/3/2026 10:04 AM, jong wu wrote:
在 2026/9/2 17:37, Hongyan Xia 写道:
On 9/2/2026 4:49 PM, jong wu wrote:
在 2026/8/25 21:05, Vincent Guittot 写道:
On Mon, 24 Aug 2026 at 15:06, Jianyong Wu <jianyong.wu@xxxxxxxxxxx>
wrote:
Hi Vincent, Hongyan,
Thanks for your comments.
My original commit message did not clearly describe the concrete issue
being fixed, and its explanation based on frequency invariance was not
correct. After looking into this further, I found that the issue I
observed has a different cause: the cpuinfo.max_freq fallback added by
d2d5c129d07e ("cpufreq: Make cpufreq_update_pressure() fall back to
cpuinfo.max_freq").
The commit message says:
However, in the absence of arch_scale_freq_ref(), it is reasonable
to assume that cpuinfo.max_freq is the maximum sustainable
frequency
for the given cpufreq policy.
That assumption does not always hold.
On an x86 server using acpi-cpufreq, cpuinfo.max_freq includes the
autonomous boost frequency, while policy->max is resolved to the
highest selectable _PSS state. With boost enabled and policy->max
unchanged at that state, the measured CPU frequency can still exceed
policy->max. Thus, policy->max does not represent an effective
hardware
maximum-frequency cap in this case.
Nevertheless, the cpuinfo.max_freq fallback makes
cpufreq_update_pressure() calculate positive pressure for every
policy,
although no effective maximum-frequency restriction has been applied.
The underlying issue is that cpuinfo.max_freq is the maximum possible
operating frequency and may include an autonomous boost frequency,
whereas policy->max may represent the highest selectable _PSS state.
Consequently, policy->max < cpuinfo.max_freq does not necessarily mean
that the available CPU capacity has been capped.
IIUC, cpuinfo.max_freq == boost freq and policy->max reflects the
correct highest frequency reachable by the CPU when boost is disabled
so the cpufreq_pressure is correct. But your policy->max is not
updated when boot is enable and doesn't reflect the highest freq
reachable by the CPU.
Exactly. I think the root cause is that policy->max has different
semantics across cpufreq drivers. On Intel and most AMD machines it is
the maximum attainable frequency, i.e. the boost frequency when boost
is enabled. For acpi-cpufreq, however, policy->max is resolved from the
ACPI _PSS table, which does not contain the boost frequency.
Proper solutions aside, I vaguely remember investigating scheduler
issues on a Ryzen 7840U. That has acpi-cpufreq with only 3 OPPs. Boost
frequencies are not included in those 3 and are much higher than the
ACPI OPPs. I certainly managed to disable pstate and switched to
acpi-cpufreq on it. I wonder if you can reproduce such issues on such a
machine. Maybe this is a broader issue than we realize.
That may well be the case, but so far I have only observed it on my own
machine, so I would rather not claim more than that yet.
My current understanding is that the trigger would be the driver rather
than the vendor: if a system runs acpi-cpufreq and its boost frequency
is not present in the _PSS table, the same reasoning should apply. The
7840U you describe -- 3 OPPs, boost well above the highest one -- looks
like it could fit that shape, but that remains a guess until it is
actually measured.
I do not have a 7840U at hand. I will try to reproduce it on an Intel
or AMD box by forcing acpi-cpufreq (intel_pstate=disable /
amd_pstate=disable), then comparing the measured frequency against
policy->max with boost enabled and checking whether a non-zero cpufreq
pressure shows up while the system is unconstrained. I will report back
with the numbers once I have them.
I managed to reproduce the problem on an AMD 5900X. I added trace_printk
outputs on cpufreq pressure updates. Under AMD pstate with boost
frequencies I get:
[006] ..... 34.096492: cpufreq_set_policy: CPU 1 has max_freq
4683471, max 4683471
You can see policy->cpuinfo.max_freq and policy->max are the same.
If I force disable pstate and use ACPI OPPs with schedutil but still
with boost frequencies, I get:
[003] ..... 4.697753: cpufreq_set_policy: CPU 1 has max_freq
4680714, max 3300000
So you can see policy->max includes no boost frequencies (3300000 is the
highest OPP) and will trigger policy->max < policy->cpuinfo.max_freq,
hence applying pressure when there is actually no pressure.
The conclusion is that yes, this is a wider problem than we realize. I
suspect this might also be present in Intel CPUs with ACPI cpufreq.
Thanks for testing. I reproduced it on an AMD box too, and I agree it is
not AMD-specific: any driver that reports a non-boost policy->max while
cpuinfo.max_freq includes boost will hit the same path, so ACPI cpufreq
on Intel should be affected as well.
I will address that in a separate patch and we can continue the
discussion there.
Thanks,
Jianyong