[PATCH] md: fix soft lockup during resync when sync is repeatedly skipped
From: Yunye Zhao
Date: Fri Jul 17 2026 - 02:44:03 EST
md_do_sync()'s main loop advances io_sectors only when I/O is actually
issued (skipped == 0). When sync_request() keeps returning skipped == 1,
io_sectors never increases, the "last_check + window > io_sectors" test
stays true, and every iteration takes the continue branch:
sectors = mddev->pers->sync_request(mddev, j, max_sectors, &skipped);
...
if (!skipped)
io_sectors += sectors;
j += sectors;
...
if (last_check + window > io_sectors || j == max_sectors)
continue;
During recovery or resync of a large array with a sparse bitmap, many
regions that need no syncing are skipped:
raid10_sync_request()
md_bitmap_start_sync() -> must_sync = false (no bitmap page)
/* every mirror skipped */
biolist == NULL -> *skipped = 1; return max_sync;
j then traverses the whole skipped range while io_sectors stays
unchanged. On a non-preemptive kernel the resync thread (mdX_resync)
hogs the CPU for a long time and eventually triggers a soft lockup:
watchdog: BUG: soft lockup - CPU#149 stuck for 313s! [mdX_resync]
md_bitmap_start_sync+0x6f/0xe0
raid10_sync_request+0x2c9/0x1530 [raid10]
md_do_sync+0x810/0x1030
md_thread+0xa7/0x150
Add cond_resched() to this continue path.
Signed-off-by: Yunye Zhao <yunye.zhao@xxxxxxxxxxxxxxxxx>
---
drivers/md/md.c | 5 +++--
1 file changed, 3 insertions(+), 2 deletions(-)
diff --git a/drivers/md/md.c b/drivers/md/md.c
index d1465bcd86c8..e7411b033490 100644
--- a/drivers/md/md.c
+++ b/drivers/md/md.c
@@ -9881,9 +9881,10 @@ void md_do_sync(struct md_thread *thread)
*/
md_new_event();
- if (last_check + window > io_sectors || j == max_sectors)
+ if (last_check + window > io_sectors || j == max_sectors) {
+ cond_resched();
continue;
-
+ }
last_check = io_sectors;
repeat:
if (time_after_eq(jiffies, mark[last_mark] + SYNC_MARK_STEP )) {
--
2.19.1.6.gb485710b