Re: [PATCH v2 0/4] kernfs: don't hold kernfs_rwsem across dir_emit()

From: Christian Brauner

Date: Thu Sep 17 2026 - 06:26:17 EST


> kernfs_fop_readdir() takes kernfs_rwsem for reading and holds it for the
> whole listing, dir_emit() included. dir_emit() copies into a userspace
> buffer, so it can fault, and under memory pressure that fault goes to
> reclaim.
>
> We are hitting this case on the Meta fleet very regularly. Many times it
> is below[1], a monitoring daemon. It walks the cgroup tree and faults on
> its own getdents(2) buffer with the lock held for reading.
>
> below: page allocation stall for 120 secs: order:0,
> mode:0x140dca(GFP_HIGHUSER_MOVABLE|__GFP_ZERO|__GFP_COMP)
> nodemask=(null),cpuset=hostcritical.slice,mems_allowed=0
> Call Trace:
> <TASK>
> dump_stack_lvl+0x5d/0x80
> __alloc_frozen_pages_noprof+0x5f4d/0x6300
> ? memcg_list_lru_alloc+0x73/0x320
> ? ima_file_check+0xd0/0x7d0
> vma_alloc_folio_noprof+0x145/0x560
> handle_mm_fault+0x17c9/0x2720
> ? find_vma+0x27/0x30
> do_user_addr_fault+0x39f/0x6e0
> exc_page_fault+0x8f/0x110
> asm_exc_page_fault+0x22/0x30
> RIP: 0010:filldir64+0xd7/0x1a0
> [Code:/RSP:/RAX:..R15: register block elided]
> kernfs_fop_readdir+0x2de/0x420
> iterate_dir+0x8c/0x1f0
> __se_sys_getdents64+0x61/0xe0
> ? copy_page_from_iter+0x860/0x860
> do_syscall_64+0x6a/0x250
> entry_SYSCALL_64_after_hwframe+0x4b/0x53
> </TASK>
>
> kernfs_rwsem is the lock every create, remove and rename in the hierarchy
> needs, and sysfs and cgroupfs have one per machine. More importantly, the
> userspace OOM killers such as systemd-oomd also traverse the cgroupfs
> hierarchy and can get stuck behind below, which is stuck in reclaim. They
> are supposed to relieve memory pressure on the system, but can get stuck
> themselves.
>
> The fix is to not hold the lock across dir_emit(). Copy the name while the
> lock is held, drop the lock, emit, and take it again.
>
> That opens a window. With the lock dropped, the entry the listing stopped
> on can be removed or renamed before the listing picks up again, and readdir
> only remembers its place as a name hash. Two entries in one directory can
> share a hash, so the hash alone does not say where the listing stopped.
>
> Patch 1 makes the fallback search land after the missing entry rather than
> on either side of it. Patch 2 keeps the name of the entry the listing is
> on and resumes on the full key. Patch 3 is the fix. Patch 4 adds tests.
>
> Patch 1 is worth having on its own. Without the rest, the same search can
> land on the entry before the missing one and report it a second time,
> between two getdents(2) calls.
>
> The series applies to the vfs-7.4.kernfs branch of the vfs tree.

Looks fine to me although it spaghettifies the code quite a bit.
I'll pull it but wait for Tejun to give it a nod.

--