Re: [PATCH] fsnotify: avoid unrelated mark reaper waits during group teardown
From: Jia Zhu
Date: Thu Oct 08 2026 - 08:02:40 EST
On Thu 10-09-26 10:34:09, Jan Kara wrote:
> Yes, but is it really practically relevant? The optimization you add really
> only works in the case when you have no notification marks in a group when
> entering inotify_release(). Usually though you actually want to get
> notified about something and thus do have some notification marks :) and in
> that case fsnotify_clear_marks_by_group() -> fsnotify_put_mark() will add
> marks to the list for the reaper to process. So please explain why this
> corner case practically matters to you.
Fair question. We've been chasing a class of fsnotify hangs in production,
recurring across several kernels, all the same shape: one thread blocked
holding a global fsnotify resource, cascading into unrelated tasks. This patch
tackles one end of that -- skipping the global reaper flush -- but as you say
the refcount==1 skip only helps a genuinely empty group, so it doesn't cover
the common case. Tracing the production reports pointed at the root instead:
inotify delivery holds fsnotify_mark_srcu across the event
kmalloc(GFP_KERNEL_ACCOUNT | __GFP_RETRY_MAYFAIL); under memcg pressure that
allocation sits in direct reclaim (NMI catches it running in
shrink_node/mem_cgroup_iter), so synchronize_srcu() in
fsnotify_mark_destroy_workfn can't complete and the global reaper stalls (D,
327s). SRCU and the reaper being global, the stall isn't confined to that
group: unrelated processes closing an inotify instance on exit all serialize
behind the one reaper in fsnotify_destroy_group -> __flush_work -- in one
report both inotifywait and a systemd process, blocked 327s. So an ordinary
workload writing to watched files under memcg pressure can wedge unrelated
process exits host-wide.
The fix that actually addresses this is to not hold fsnotify_mark_srcu across
that allocation: pin the iter marks, drop SRCU around the delivery, re-acquire
after -- so a delivery sleeping in reclaim no longer holds the reaper. Teardown
still waits on the group's own marks, so ordering holds. Reproducible on stock
upstream by holding an fsnotify_mark_srcu reader and closing an unrelated group.
So I'd send that SRCU-drop as the primary patch; this destroy_group change is
at most a small optimization on top of it, and I'm happy to drop it.
Jia