Re: [PATCH v4 2/2] writeback: switch a replaced cgwb's inodes to its successor
From: Tejun Heo
Date: Thu Oct 01 2026 - 14:43:59 EST
Hello, Liz.
On Wed, Sep 30, 2026 at 07:35:37PM +0000, Liz Fong-Jones wrote:
> + new_wb = wb_get_lookup(wb->bdi, wb->memcg_css);
> + if (!new_wb)
> + return false;
If the successor was already released, the slot is empty and the inodes
stay on a wb that foreign flushes can't find. Maybe fall back to
&wb->bdi->wb like cleanup_offline_cgwb() does?
> + spin_lock(&wb->list_lock);
> + restart = isw_prepare_wbs_switch(new_wb, isw, &wb->b_attached, &nr) ||
> + isw_prepare_wbs_switch(new_wb, isw, &wb->b_dirty, &nr) ||
> + isw_prepare_wbs_switch(new_wb, isw, &wb->b_io, &nr) ||
> + isw_prepare_wbs_switch(new_wb, isw, &wb->b_more_io, &nr) ||
> + isw_prepare_wbs_switch(new_wb, isw, &wb->b_dirty_time, &nr);
> + spin_unlock(&wb->list_lock);
Switching a dirty inode away leaves WB_has_dirty_io set on the old wb.
inode_do_switch_wbs() moves it onto new_wb->b_dirty through
inode_io_list_move_locked(), which only updates new_wb, and nothing calls
wb_io_lists_depopulated() on old_wb. A live wb clears it on its next dirty
to clean transition, but a replaced wb never gets another inode, so it's
freed with its avg_write_bandwidth still in bdi->tot_write_bandwidth,
which wb_split_bdi_pages() and wb_min_max_ratio() divide by. Can you add a
prep patch which calls wb_io_lists_depopulated(old_wb) after the switch
loop in process_inode_switch_wbs(), while old_wb->list_lock is still held?
> + while (switch_replaced_cgwb(wb))
> + cond_resched();
cond_resched() is a no-op on PREEMPTION kernels, so a wb with a lot of
inodes keeps this worker from reporting a Tasks-RCU quiescent state. See
407a5d205179 ("writeback: report a Tasks-RCU quiescent state per cgwb
drain pass"). Can you use the same do { } while () shape with
cond_resched_tasks_rcu_qs() as cleanup_offline_cgwbs_workfn()?
Thanks.
--
tejun