Re: [PATCH] writeback: bound cleanup_offline_cgwb() rescans by rotating b_attached
From: Roman Gushchin
Date: Wed Sep 09 2026 - 17:38:59 EST
"Patrick Lu (Anthropic)" <perf.patrick.lu@xxxxxxxxx> writes:
> cleanup_offline_cgwb() prepares at most WB_MAX_INODES_PER_ISW inodes
> per call and is called again until the dying wb is drained, but every
> call walks wb->b_attached from the head. Inodes already prepared (they
> stay on the list with I_WB_SWITCH set until the switch worker runs) and
> inodes that cannot be switched (I_FREEING, I_WILL_FREE, !SB_ACTIVE,
> DAX, already on the target wb) stay at the head, so each pass rescans a
> growing prefix under wb->list_lock and a full drain is quadratic in the
> number of attached inodes. With ~17M inodes attached to one dying cgwb
> we have seen this end in soft lockups, with CPUs reported stuck for
> 21-48s.
>
> Move every scanned inode to the tail of b_attached, so the next pass
> starts where the previous one stopped and the drain becomes linear.
> b_attached is unordered and isw_prepare_wbs_switch() is its only
> walker, so nobody else sees the reorder. b_dirty_time is ordered by
> expiry for move_expired_inodes() and keeps its current scan.
>
> Fixes: c22d70a162d3 ("writeback, cgroup: release dying cgwbs by switching attached inodes")
> Cc: stable@xxxxxxxxxxxxxxx
> Assisted-by: LLM
> Signed-off-by: Patrick Lu (Anthropic) <perf.patrick.lu@xxxxxxxxx>
> ---
> Seen in production on a 6.18-based kernel: with ~17M inodes attached
> to one dying cgwb, a node spent 36 minutes in back-to-back
> wb->list_lock holds by the cleanup scanner (~6ms each, ~46% of wall
> time, starving writeback on that wb); with this patch the same workload
> drains in ~30 seconds. Also seen on stock Amazon Linux 2023 6.12.68 as
> isw workers spinning on the list_lock in inode_switch_wbs_work_fn()
> while cleanup_offline_cgwbs_workfn() runs.
>
> Tested with a QEMU A/B setup at 100k attached inodes and patched vs
> unpatched on production-class hardware at ~17M attached inodes.
>
> Josef Bacik's patch making the drain loop report a Tasks-RCU quiescent
> state [1] fixes BPF/ftrace detach stalls behind the same drain; this
> patch bounds the walk itself. The two are independent.
>
> [1] https://lore.kernel.org/linux-mm/20260909-cgwb-tasks-rcu-qs-v1-1-967a7754771f@xxxxxxxxxxxxxx/
> ---
> fs/fs-writeback.c | 39 ++++++++++++++++++++++++++++-----------
> 1 file changed, 28 insertions(+), 11 deletions(-)
Acked-by: Roman Gushchin <roman.gushchin@xxxxxxxxx>
Thanks