[PATCH] writeback: bound cleanup_offline_cgwb() rescans by rotating b_attached
From: Patrick Lu (Anthropic)
Date: Wed Sep 09 2026 - 14:53:52 EST
cleanup_offline_cgwb() prepares at most WB_MAX_INODES_PER_ISW inodes
per call and is called again until the dying wb is drained, but every
call walks wb->b_attached from the head. Inodes already prepared (they
stay on the list with I_WB_SWITCH set until the switch worker runs) and
inodes that cannot be switched (I_FREEING, I_WILL_FREE, !SB_ACTIVE,
DAX, already on the target wb) stay at the head, so each pass rescans a
growing prefix under wb->list_lock and a full drain is quadratic in the
number of attached inodes. With ~17M inodes attached to one dying cgwb
we have seen this end in soft lockups, with CPUs reported stuck for
21-48s.
Move every scanned inode to the tail of b_attached, so the next pass
starts where the previous one stopped and the drain becomes linear.
b_attached is unordered and isw_prepare_wbs_switch() is its only
walker, so nobody else sees the reorder. b_dirty_time is ordered by
expiry for move_expired_inodes() and keeps its current scan.
Fixes: c22d70a162d3 ("writeback, cgroup: release dying cgwbs by switching attached inodes")
Cc: stable@xxxxxxxxxxxxxxx
Assisted-by: LLM
Signed-off-by: Patrick Lu (Anthropic) <perf.patrick.lu@xxxxxxxxx>
---
Seen in production on a 6.18-based kernel: with ~17M inodes attached
to one dying cgwb, a node spent 36 minutes in back-to-back
wb->list_lock holds by the cleanup scanner (~6ms each, ~46% of wall
time, starving writeback on that wb); with this patch the same workload
drains in ~30 seconds. Also seen on stock Amazon Linux 2023 6.12.68 as
isw workers spinning on the list_lock in inode_switch_wbs_work_fn()
while cleanup_offline_cgwbs_workfn() runs.
Tested with a QEMU A/B setup at 100k attached inodes and patched vs
unpatched on production-class hardware at ~17M attached inodes.
Josef Bacik's patch making the drain loop report a Tasks-RCU quiescent
state [1] fixes BPF/ftrace detach stalls behind the same drain; this
patch bounds the walk itself. The two are independent.
[1] https://lore.kernel.org/linux-mm/20260909-cgwb-tasks-rcu-qs-v1-1-967a7754771f@xxxxxxxxxxxxxx/
---
fs/fs-writeback.c | 39 ++++++++++++++++++++++++++++-----------
1 file changed, 28 insertions(+), 11 deletions(-)
diff --git a/fs/fs-writeback.c b/fs/fs-writeback.c
index e744f9f9d43f..69a452b12b12 100644
--- a/fs/fs-writeback.c
+++ b/fs/fs-writeback.c
@@ -725,21 +725,37 @@ static void inode_switch_wbs(struct inode *inode, int new_wb_id)
static bool isw_prepare_wbs_switch(struct bdi_writeback *new_wb,
struct inode_switch_wbs_context *isw,
- struct list_head *list, int *nr)
+ struct list_head *list, bool rotate, int *nr)
{
- struct inode *inode;
+ struct inode *inode, *tmp;
+ LIST_HEAD(scanned);
+ bool full = false;
+
+ list_for_each_entry_safe(inode, tmp, list, i_io_list) {
+ /*
+ * Rotate scanned inodes to the tail so the next scan resumes
+ * at unscanned ones instead of re-walking an ever-growing
+ * prefix of prepared and skipped inodes. b_dirty_time is
+ * expiry-ordered and so must not be rotated.
+ */
+ if (rotate)
+ list_move_tail(&inode->i_io_list, &scanned);
- list_for_each_entry(inode, list, i_io_list) {
if (!inode_prepare_wbs_switch(inode, new_wb))
continue;
isw->inodes[*nr] = inode;
(*nr)++;
- if (*nr >= WB_MAX_INODES_PER_ISW - 1)
- return true;
+ if (*nr >= WB_MAX_INODES_PER_ISW - 1) {
+ full = true;
+ break;
+ }
}
- return false;
+ if (rotate)
+ list_splice_tail(&scanned, list);
+
+ return full;
}
/**
@@ -747,8 +763,9 @@ static bool isw_prepare_wbs_switch(struct bdi_writeback *new_wb,
* @wb: target wb
*
* Switch all inodes attached to @wb to a nearest living ancestor's wb in order
- * to eventually release the dying @wb. Returns %true if not all inodes were
- * switched and the function has to be restarted.
+ * to eventually release the dying @wb. Returns %true if the scan stopped
+ * early after making progress; the caller should call again to continue
+ * draining.
*/
bool cleanup_offline_cgwb(struct bdi_writeback *wb)
{
@@ -783,13 +800,13 @@ bool cleanup_offline_cgwb(struct bdi_writeback *wb)
* bandwidth restrictions, as writeback of inode metadata is not
* accounted for.
*/
- restart = isw_prepare_wbs_switch(new_wb, isw, &wb->b_attached, &nr);
+ restart = isw_prepare_wbs_switch(new_wb, isw, &wb->b_attached, true, &nr);
if (!restart)
restart = isw_prepare_wbs_switch(new_wb, isw, &wb->b_dirty_time,
- &nr);
+ false, &nr);
spin_unlock(&wb->list_lock);
- /* no attached inodes? bail out */
+ /* nothing to switch? bail out */
if (nr == 0) {
atomic_dec(&isw_nr_in_flight);
wb_put(new_wb);
---
base-commit: e14d4302cbd0de773960bec33c2281508c8d8855
change-id: 20260909-wb-cgwb-rotate-f17a75facfdc
Best regards,
--
Patrick Lu (Anthropic) <perf.patrick.lu@xxxxxxxxx>