Re: [PATCH 0/1] fsnotify: Check parent watches without child dentry flags

From: Jan Kara

Date: Thu Oct 08 2026 - 05:42:21 EST


On Wed 07-10-26 15:20:50, Partha Sarathi Satapathy wrote:
> This patch removes the cached-child dentry walk performed when a directory
> first starts watching child events. Instead, the event path takes a reference
> to the parent dentry and checks its child-watch mask. It also removes the
> DCACHE_FSNOTIFY_PARENT_WATCHED flag and its update sites.

If we can get rid of it, it would be a nice simplification but I have
doubts this is feasible.

> On the test host, 12 million files were populated and their positive dentries
> were warmed before each run. With eight concurrent unlink/recreate workers
> and 100 sole-watch add/close cycles, inotify_add_watch() measured:
>
> Before After
> run 1 average 107.420 ms 0.004 ms
> run 2 average 125.433 ms 0.004 ms
> run 1 maximum 122.811 ms 0.113 ms
> run 2 maximum 141.198 ms 0.095 ms
>
> The displayed 0.004 ms averages are rounded to three decimal places.
> Maximum unlink/recreate latency was 777.852 and 770.569 ms before, versus
> 8.525 and 15.624 ms after. Those mutation maxima are supporting observations:
> workers ran only during the watch-add loops, so the pre-change and post-change
> mutation measurement windows had very different lengths.

This looks good and is kind of expected. We don't have to traverse the huge
children list.

> One 100,000-iteration, single-worker open/close event-path run per mode gave:
>
> Before After
> none: p50/p99 ns 9564 / 15703 9684 / 15023
> throughput 102221 ops/s 101432 ops/s
> other: p50/p99 ns 9554 / 15573 9684 / 14962
> throughput 99442 ops/s 101497 ops/s
> parent: p50/p99 ns 12529 / 19179 12499 / 17716
> throughput 78196 ops/s 78655 ops/s
>
> The 'other' mode watches a different directory on the same filesystem, so it
> exercises the superblock watcher path without target-file event delivery.
> These are single runs; they do not establish a small event-path cost or its
> absence. No open/close regression is apparent at this measurement resolution.
> The test used worker CPU 2 and listener CPU 4 on kernels
> 6.19.0-rc8.V_fsn0.el9.omm0.x86_64 and 6.19.0-rc8.fsn3.el9.omm3.x86_64,
> respectively. The test host reported XFS for its working directory.
>
> The inotify correctness test and broader fsnotify functional test passed on
> both kernels. The fanotify permission subtest skipped with EPERM in both runs,
> so FAN_OPEN_PERM remains unvalidated by these results.

This isn't a load for which the DCACHE_FSNOTIFY_PARENT_WATCHED optimization
is that interesting. Try the following:

Create a directory on tmpfs with say 1024 files, each 4k large.
Place IN_MODIFY inotify watch on the directory.
Start X clients (where X can be 1,2,4,8,...,1024 to see the scaling)
The client will open the file based on its number (so different clients
work on different file)
Do 1000000 writes of 1 byte at offset 0 to the file => measure time for
this
Close the file

I would bet you would see the difference already at 1 client and as the
number of clients grows and the cacheline contention on the parent's
refcount increases, it will get progressively worse. In particular if you
try on a NUMA machine where the cacheline will ping-pong across NUMA nodes.

We've got performance regression reports for similar loads already when we
added one cacheline load to fsnotify_parent(). Doing the "grab parent
refcount" dance is much more expensive than that.

Honza

>
> Partha Sarathi Satapathy (1):
> fsnotify: Check parent watches without child dentry flags
>
> fs/dcache.c | 4 --
> fs/notify/fsnotify.c | 99 ++++++++------------------------
> fs/notify/fsnotify.h | 6 --
> fs/notify/mark.c | 35 -----------
> include/linux/dcache.h | 1 -
> include/linux/fsnotify.h | 7 +--
> include/linux/fsnotify_backend.h | 29 ++--------
> 7 files changed, 29 insertions(+), 152 deletions(-)
--
Jan Kara <jack@xxxxxxxx>
SUSE Labs, CR