Re: [PATCH] writeback: bound cleanup_offline_cgwb() rescans by rotating b_attached

From: Jan Kara

Date: Thu Sep 10 2026 - 07:35:27 EST


On Wed 09-09-26 18:50:27, Patrick Lu (Anthropic) wrote:
> cleanup_offline_cgwb() prepares at most WB_MAX_INODES_PER_ISW inodes
> per call and is called again until the dying wb is drained, but every
> call walks wb->b_attached from the head. Inodes already prepared (they
> stay on the list with I_WB_SWITCH set until the switch worker runs) and
> inodes that cannot be switched (I_FREEING, I_WILL_FREE, !SB_ACTIVE,
> DAX, already on the target wb) stay at the head, so each pass rescans a
> growing prefix under wb->list_lock and a full drain is quadratic in the
> number of attached inodes. With ~17M inodes attached to one dying cgwb
> we have seen this end in soft lockups, with CPUs reported stuck for
> 21-48s.

Hum, somehow I've missed this case when fixing slow inode switching couple
months ago. Likely because I was more focused on the b_dirty_time case back
then :) since that was what the user was hitting.

> Move every scanned inode to the tail of b_attached, so the next pass
> starts where the previous one stopped and the drain becomes linear.
> b_attached is unordered and isw_prepare_wbs_switch() is its only
> walker, so nobody else sees the reorder. b_dirty_time is ordered by
> expiry for move_expired_inodes() and keeps its current scan.

OK, but isn't there the very same quadratic behavior problem with
b_dirty_time scan which you don't touch (and where your trick cannot work)?

Honza

>
> Fixes: c22d70a162d3 ("writeback, cgroup: release dying cgwbs by switching attached inodes")
> Cc: stable@xxxxxxxxxxxxxxx
> Assisted-by: LLM
> Signed-off-by: Patrick Lu (Anthropic) <perf.patrick.lu@xxxxxxxxx>
> ---
> Seen in production on a 6.18-based kernel: with ~17M inodes attached
> to one dying cgwb, a node spent 36 minutes in back-to-back
> wb->list_lock holds by the cleanup scanner (~6ms each, ~46% of wall
> time, starving writeback on that wb); with this patch the same workload
> drains in ~30 seconds. Also seen on stock Amazon Linux 2023 6.12.68 as
> isw workers spinning on the list_lock in inode_switch_wbs_work_fn()
> while cleanup_offline_cgwbs_workfn() runs.
>
> Tested with a QEMU A/B setup at 100k attached inodes and patched vs
> unpatched on production-class hardware at ~17M attached inodes.
>
> Josef Bacik's patch making the drain loop report a Tasks-RCU quiescent
> state [1] fixes BPF/ftrace detach stalls behind the same drain; this
> patch bounds the walk itself. The two are independent.
>
> [1] https://lore.kernel.org/linux-mm/20260909-cgwb-tasks-rcu-qs-v1-1-967a7754771f@xxxxxxxxxxxxxx/
> ---
> fs/fs-writeback.c | 39 ++++++++++++++++++++++++++++-----------
> 1 file changed, 28 insertions(+), 11 deletions(-)
>
> diff --git a/fs/fs-writeback.c b/fs/fs-writeback.c
> index e744f9f9d43f..69a452b12b12 100644
> --- a/fs/fs-writeback.c
> +++ b/fs/fs-writeback.c
> @@ -725,21 +725,37 @@ static void inode_switch_wbs(struct inode *inode, int new_wb_id)
>
> static bool isw_prepare_wbs_switch(struct bdi_writeback *new_wb,
> struct inode_switch_wbs_context *isw,
> - struct list_head *list, int *nr)
> + struct list_head *list, bool rotate, int *nr)
> {
> - struct inode *inode;
> + struct inode *inode, *tmp;
> + LIST_HEAD(scanned);
> + bool full = false;
> +
> + list_for_each_entry_safe(inode, tmp, list, i_io_list) {
> + /*
> + * Rotate scanned inodes to the tail so the next scan resumes
> + * at unscanned ones instead of re-walking an ever-growing
> + * prefix of prepared and skipped inodes. b_dirty_time is
> + * expiry-ordered and so must not be rotated.
> + */
> + if (rotate)
> + list_move_tail(&inode->i_io_list, &scanned);
>
> - list_for_each_entry(inode, list, i_io_list) {
> if (!inode_prepare_wbs_switch(inode, new_wb))
> continue;
>
> isw->inodes[*nr] = inode;
> (*nr)++;
>
> - if (*nr >= WB_MAX_INODES_PER_ISW - 1)
> - return true;
> + if (*nr >= WB_MAX_INODES_PER_ISW - 1) {
> + full = true;
> + break;
> + }
> }
> - return false;
> + if (rotate)
> + list_splice_tail(&scanned, list);
> +
> + return full;
> }
>
> /**
> @@ -747,8 +763,9 @@ static bool isw_prepare_wbs_switch(struct bdi_writeback *new_wb,
> * @wb: target wb
> *
> * Switch all inodes attached to @wb to a nearest living ancestor's wb in order
> - * to eventually release the dying @wb. Returns %true if not all inodes were
> - * switched and the function has to be restarted.
> + * to eventually release the dying @wb. Returns %true if the scan stopped
> + * early after making progress; the caller should call again to continue
> + * draining.
> */
> bool cleanup_offline_cgwb(struct bdi_writeback *wb)
> {
> @@ -783,13 +800,13 @@ bool cleanup_offline_cgwb(struct bdi_writeback *wb)
> * bandwidth restrictions, as writeback of inode metadata is not
> * accounted for.
> */
> - restart = isw_prepare_wbs_switch(new_wb, isw, &wb->b_attached, &nr);
> + restart = isw_prepare_wbs_switch(new_wb, isw, &wb->b_attached, true, &nr);
> if (!restart)
> restart = isw_prepare_wbs_switch(new_wb, isw, &wb->b_dirty_time,
> - &nr);
> + false, &nr);
> spin_unlock(&wb->list_lock);
>
> - /* no attached inodes? bail out */
> + /* nothing to switch? bail out */
> if (nr == 0) {
> atomic_dec(&isw_nr_in_flight);
> wb_put(new_wb);
>
> ---
> base-commit: e14d4302cbd0de773960bec33c2281508c8d8855
> change-id: 20260909-wb-cgwb-rotate-f17a75facfdc
>
> Best regards,
> --
> Patrick Lu (Anthropic) <perf.patrick.lu@xxxxxxxxx>
>
--
Jan Kara <jack@xxxxxxxx>
SUSE Labs, CR