Re: [syzbot] [block?] INFO: task hung in read_cache_folio (6) (extid 9db0864859224b833108)

From: clkernel

Date: Mon Sep 28 2026 - 03:45:58 EST


On Fri, 03 Jul 2026, syzbot wrote:
> syzbot found the following issue on:
> ...
> folio_wait_bit_common+0x6b8/0xa30 mm/filemap.c:1324
> folio_put_wait_locked mm/filemap.c:1493 [inline]
> do_read_cache_folio+0x23c/0x5a8 mm/filemap.c:4089
> read_cache_folio+0x68/0x88 mm/filemap.c:4139
> read_part_sector+0xcc/0x708 block/partitions/core.c:724

A second observation of the same wait site: on x86-64 servers we see
reader threads (a database reader in a busy server workload) parked
D-state in folio_wait_bit_common's io_schedule, entered through the same
folio_put_wait_locked -> do_read_cache_folio -> read_cache_folio path
this report shows. Observed on 7.0.0-30, 7.0.0-31 and 7.2.5, reproduced N=3.

Three measurements, per repro:

1. the pinned folio's lock is FREE while the waiter sleeps: a fresh read
of the exact pinned pages answers in 4-108 ms while the thread stays D
in the same wait;
2. each later lock/unlock of that folio wakes the waiter for one ~4KB
read: the blocked read's offsets crawl one page per unlock cycle (~one
page per 10 s under churn), so the syscall can sit in the same wait for
minutes to days;
3. the discriminator: read /proc/<tid>/syscall while the waiter sits
inside the wait, then dd the same page range fresh. The dd returns
immediately; the waiter does not move.

Root cause (established from source, then observed directly on a local
build; measured): folio_wake_bit() wakes the hashed folio waitqueue
through __wake_up_locked_key(), i.e. __wake_up_common(...,
nr_exclusive=1). The walk in kernel/sched/wait.c stops at the first
WQ_FLAG_EXCLUSIVE entry it wakes:

if (ret && (flags & WQ_FLAG_EXCLUSIVE) && !--nr_exclusive)
break;

folio_wait_bit_common() queued every waiter with
__add_wait_queue_entry_tail(), so a non-exclusive waiter could sit
behind an exclusive folio_lock() waiter on the same folio. When the fill
completed, folio_end_read()'s wake woke the exclusive waiter and
stopped. The shared/DROP waiter behind it was never visited:
WQ_FLAG_WOKEN never set, although the wake itself ran. PG_locked was
already clear and the folio uptodate (so a fresh read of that page
answers through the uptodate fast path, measurement 1), and PG_waiters
stayed set, so the sleeper only ran when some later lock/unlock of the
same folio address reached the queue (measurement 2: one page per wake,
crawling until traffic arrived).

Same mechanism explains this report's hang: the waiters are the DROP
ones (folio_put_wait_locked), which drop their folio reference before
sleeping and re-lookup only after a wake that arrives late or never.

Observation on a debug build, same queue after folio_wake_bit() returned
(a read-only walk counting matching non-exclusive entries left
unvisited): 363 events per 10-minute run of readers racing
punch/fault/truncate churn under reclaim pressure on an unmodified
queue; 1 per run after queueing non-exclusive waiters at the head side
only; 0 after the same change at softleaf_entry_wait_on_locked (the
other non-exclusive enqueue on this queue).

The fix (mm/filemap.c, attached): queue non-exclusive waiters with
__add_wait_queue() (wait.h's own default head-side insert) and keep
exclusive folio_lock waiters on __add_wait_queue_entry_tail(). With that
order every non-exclusive full-match waiter is visited before any walk
break can fire, which restores the "wake all shared waiters, then one
exclusive" contract of nr_exclusive=1; exclusive waiters keep their FIFO
tail and their one-at-a-time wake. Same change at both non-exclusive
enqueue sites (folio_wait_bit_common and softleaf_entry_wait_on_locked).

One known sibling waits on a follow-up: __folio_lock_async's entry
(io_uring buffered reads with ki_waitq) enqueues non-exclusive at the
tail. Same class, unreachable from the workloads above.

Reproducer shape: threads pread a file while others block in folio_lock
on the same pages (hole punch, truncate, or mmap faults) under reclaim
churn. On an unmodified queue the waitqueue walk above counts the
skipped waiters within seconds. The harness, traces or configs are yours
on request.

The same change cherry-picks clean onto v7.2.6 (mm/filemap.o builds
there). Happy to send a [PATCH] for stable@ if the fix is wanted on the
7.2.x queue.

---
mm/filemap.c | 16 +++++++++++++---
1 file changed, 13 insertions(+), 3 deletions(-)

diff --git a/mm/filemap.c b/mm/filemap.c
index 00fd89cf6f55..383701407843 100644
--- a/mm/filemap.c
+++ b/mm/filemap.c
@@ -1294,8 +1294,18 @@ static inline int folio_wait_bit_common(struct
folio *folio, int bit_nr,
*/
spin_lock_irq(&q->lock);
folio_set_waiters(folio);
- if (!folio_trylock_flag(folio, bit_nr, wait))
- __add_wait_queue_entry_tail(q, wait);
+ if (!folio_trylock_flag(folio, bit_nr, wait)) {
+ /*
+ * A non-exclusive (shared/DROP) waiter must queue ahead of
+ * exclusive folio_lock waiters: the nr_exclusive=1 wake walk
+ * stops at the first exclusive match, so a non-exclusive
+ * waiter behind one is never visited and never woken.
+ */
+ if (behavior == EXCLUSIVE)
+ __add_wait_queue_entry_tail(q, wait);
+ else
+ __add_wait_queue(q, wait);
+ }
spin_unlock_irq(&q->lock);
/*
@@ -1429,7 +1439,7 @@ void softleaf_entry_wait_on_locked(softleaf_t
entry, spinlock_t *ptl)
spin_lock_irq(&q->lock);
folio_set_waiters(folio);
if (!folio_trylock_flag(folio, PG_locked, wait))
- __add_wait_queue_entry_tail(q, wait);
+ __add_wait_queue(q, wait);
spin_unlock_irq(&q->lock);
/*