Re: [PATCH 1/2] jbd2: check need_resched() when skipping busy checkpoint buffers

From: Jan Kara

Date: Mon Jul 13 2026 - 08:28:09 EST


On Mon 13-07-26 12:22:28, Max Kellermann wrote:
> journal_shrink_one_cp_list() skips busy checkpoint buffers when called
> with JBD2_SHRINK_BUSY_SKIP. The continue statement on this path also
> skips the need_resched() check at the end of the loop body.
>
> Consequently, when a checkpoint list contains mostly busy buffers, the
> shrinker can walk the entire list while holding journal->j_list_lock,
> even when a reschedule has been requested. Large checkpoint lists under
> memory pressure can therefore cause long lock hold times and leave other
> CPUs spinning on j_list_lock, resulting in soft lockups or RCU stalls.
>
> Route the busy-buffer path through the need_resched() check so that the
> shrinker can release j_list_lock and reschedule promptly, restoring
> parity with the clean-buffer path, which already checks need_resched().
> This does not change which checkpoint buffers are eligible for removal.
>
> Fixes: b98dba273a0e ("jbd2: remove journal_clean_one_cp_list()")
> Cc: stable@xxxxxxxxxxxxxxx
> Signed-off-by: Max Kellermann <max.kellermann@xxxxxxxxx>

Good catch! Thanks for fixing this. Feel free to add:

Reviewed-by: Jan Kara <jack@xxxxxxx>

Honza

> ---
> fs/jbd2/checkpoint.c | 3 ++-
> 1 file changed, 2 insertions(+), 1 deletion(-)
>
> diff --git a/fs/jbd2/checkpoint.c b/fs/jbd2/checkpoint.c
> index 1508e2f54462..5266017565ac 100644
> --- a/fs/jbd2/checkpoint.c
> +++ b/fs/jbd2/checkpoint.c
> @@ -389,7 +389,7 @@ static unsigned long journal_shrink_one_cp_list(struct journal_head *jh,
> ret = jbd2_journal_try_remove_checkpoint(jh);
> if (ret < 0) {
> if (type == JBD2_SHRINK_BUSY_SKIP)
> - continue;
> + goto next;
> break;
> }
> }
> @@ -400,6 +400,7 @@ static unsigned long journal_shrink_one_cp_list(struct journal_head *jh,
> break;
> }
>
> +next:
> if (need_resched())
> break;
> } while (jh != last_jh);
> --
> 2.47.3
>
--
Jan Kara <jack@xxxxxxxx>
SUSE Labs, CR