[PATCH] afs: Don't flush dirty data when zapping an invalidated vnode

From: Yuanfu Xie

Date: Sat Sep 26 2026 - 13:56:22 EST


afs_validate() holds vnode->validate_lock for write while it
revalidates the vnode. If it determines that the vnode's data must be
zapped, it calls afs_zap_data(), which for a regular file calls
filemap_invalidate_inode() with flush=true. That submits WB_SYNC_ALL
writeback, which re-enters afs_writepages() - and afs_writepages()
takes validate_lock for read before it will write anything back. The
task thus deadlocks against itself in an uninterruptible sleep and
hangs forever; with the hung-task watchdog set to panic, the kernel
panics:

Kernel panic - not syncing: hung_task: blocked tasks
task:repro state:D pid:1
Call Trace:
<TASK>
rwsem_down_read_slowpath+0x4d0/0xde0
down_read+0xb6/0x220
afs_writepages+0xb8/0xd0
do_writepages+0x22c/0x550
filemap_writeback+0x1dc/0x260
afs_validate+0xd1d/0xfc0
afs_fsync+0xcb/0x1d0
do_fsync+0xac/0x200
__x64_sys_fsync+0x32/0x50
entry_SYSCALL_64_after_hwframe+0x77/0x7f
</TASK>

Trigger: fsync() on a dirty file whose server-side data version
changed (callback break, v_break change or callback expiry) - a normal
concurrency scenario, no malicious server required. Everything stuck
behind WB_SYNC_ALL (sync(), syncfs, umount, shutdown flush) hangs with
it, and the task cannot be killed.

The read acquisition in afs_writepages() cannot be dropped: it orders
writeback against afs_setattr() truncating the pagecache under the
write lock. So the flush has to go: invalidate the pages without
flushing, which is what the directory and symlink branches of
afs_zap_data() have always done. Any queued writes are based on a
stale data version at this point (that is why the zap was judged) and
cannot be stored coherently anyway.

The workload that hangs the unpatched kernel was run on a build of
7.3-rc3-135-g4982d3552a3bf with only this change applied: the fsync()
that used to hang now completes cleanly, and several hundred ordinary
file operations on the same build also completed without a crash.

Fixes: d73065e60dcc8 ("afs: Use alternative invalidation to using launder_folio")
Cc: stable@xxxxxxxxxxxxxxx
Signed-off-by: Yuanfu Xie <yuanfuxie@xxxxxxxxxxxxxx>
---
fs/afs/validation.c | 14 +++++++++-----
1 file changed, 9 insertions(+), 5 deletions(-)

diff --git a/fs/afs/validation.c b/fs/afs/validation.c
index e997563af658b..20d654bbf7e8b 100644
--- a/fs/afs/validation.c
+++ b/fs/afs/validation.c
@@ -372,11 +372,15 @@ static void afs_zap_data(struct afs_vnode *vnode)

/* nuke all the non-dirty pages that aren't locked, mapped or being
* written back in a regular file and completely discard the pages in a
- * directory or symlink */
- if (S_ISREG(vnode->netfs.inode.i_mode))
- filemap_invalidate_inode(&vnode->netfs.inode, true, 0, LLONG_MAX);
- else
- filemap_invalidate_inode(&vnode->netfs.inode, false, 0, LLONG_MAX);
+ * directory or symlink.
+ *
+ * Don't flush here: we hold validate_lock for write, and writeback
+ * would take validate_lock for read again in afs_writepages(),
+ * deadlocking against ourselves. The queued writes are also based
+ * on a stale data version at this point and cannot be stored
+ * coherently anyway.
+ */
+ filemap_invalidate_inode(&vnode->netfs.inode, false, 0, LLONG_MAX);
}

/*
--
2.43.0