[PATCHSET wq/for-7.4] workqueue: Make flush_workqueue() cost scale with active pwqs
From: Tejun Heo
Date: Tue Sep 01 2026 - 17:10:35 EST
Hello,
Since 636b927eba5b ("workqueue: Make unbound workqueues to use per-cpu
pool_workqueues"), flush_workqueue() walks one pwq per possible CPU, cycling
each pool lock, even when the workqueue is idle. Yao Kai reported the XFS
CIL workqueue, flushed on every log force, spending up to 64us per flush in
that walk on a 128-CPU machine.
This patchset makes flushes visit only the pwqs which have been active since
the last flush by tracking them on per-node lists. On a 192-CPU 2-node
machine, flushing an idle per-cpu workqueue goes from 40k to 4.2M per second
and 16 threads doing queue+flush in a loop go from 30k to 110k flushes per
second. Dense flushes with every pwq active get 15-20% slower on unbound
workqueues.
Yao, can you please test whether this resolves the fsync latencies you
reported?
This patchset contains the following three patches:
0001-workqueue-Add-for_each_node_with_fallback.patch
0002-workqueue-Maintain-pwq-total_in_flight.patch
0003-workqueue-Make-flush_workqueue-visit-only-pwqs-activ.patch
0001-0002 are prep patches. 0003 implements the active pwq lists.
The patchset is on top of wq/for-7.4 (ab85b68e630e) and also available in
the following git branch:
git://git.kernel.org/pub/scm/linux/kernel/git/tj/wq.git wq-flush-active-pwq-lists
diffstat follows. Thanks.
kernel/workqueue.c | 301 ++++++++++++++++++++++++++++++++++++++++++-----------
1 file changed, 238 insertions(+), 63 deletions(-)
--
tejun