Re: [PATCH 1/3] kernfs: take kernfs_rename_lock for same-parent renames too

From: Shakeel Butt

Date: Thu Sep 03 2026 - 13:26:54 EST


On Wed, Sep 02, 2026 at 11:15:16PM -0700, Shakeel Butt wrote:
> On Thu, Sep 03, 2026 at 07:41:53AM +0200, Greg Kroah-Hartman wrote:
> > On Wed, Sep 02, 2026 at 10:37:21PM -0700, Shakeel Butt wrote:
> > > On Thu, Sep 03, 2026 at 06:31:58AM +0200, Greg Kroah-Hartman wrote:
> > > > On Wed, Sep 02, 2026 at 09:02:51PM -0700, Shakeel Butt wrote:
> > > > > kernfs_rename_ns() only takes kernfs_rename_lock when the rename moves
> > > > > the node to a new parent. A rename that keeps the same parent, like
> > > > > renaming a network interface, changes kernfs_node::name with only
> > > > > kernfs_rwsem held. So the lock protects ->__parent but not ->name, and
> > > > > a reader that wants a stable name has to take kernfs_rwsem, the same
> > > > > lock every path lookup needs.
> > > > >
> > > > > That also makes for a small but real bug. kernfs_path_from_node()
> > > > > takes kernfs_rename_lock for reading, and kernfs_path_from_node_locked()
> > > > > then reads the name of each ancestor. It reads each one once, so a
> > > > > single same-parent rename only moves the answer from the old path to the
> > > > > new one, but two of them landing inside one walk build a path that never
> > > > > existed:
> > > > >
> > > > > CPU0 CPU1
> > > > > kernfs_path_from_node() on /a/b/c
> > > > > reads the name of a, gets "a"
> > > > > renames a to a2
> > > > > renames b to b2
> > > > > reads the name of b, gets "b2"
> > > > > returns "/a/b2/c"
> > > > >
> > > > > This hits roots without KERNFS_ROOT_INVARIANT_PARENT: sysfs, where the
> > > > > bad path can reach sysfs_warn_dup() and pr_cont_kernfs_path(), and
> > > > > resctrl, which renames a mon group inside its mon_groups directory.
> > > > > cgroup sets the flag, so it skips the lock and reads names under RCU
> > > > > alone; that case needs something else and is not addressed here.
> > > > >
> > > > > So take the lock in both cases, and let kernfs_rcu_name() accept it the
> > > > > way kernfs_parent() already does for ->__parent. Same-parent renames
> > > > > are rare, the lock is per filesystem, and the locked section is at most
> > > > > three stores. It also gives a future rename sequence counter one place
> > > > > to sit that covers every rename.
> > > > >
> > > > > Fixes: 741c10b096bc ("kernfs: Use RCU to access kernfs_node::name.")
> > > > > Signed-off-by: Shakeel Butt <shakeel.butt@xxxxxxxxx>
> > > > > ---
> > > > > fs/kernfs/dir.c | 28 +++++++++++++++-------------
> > > > > fs/kernfs/kernfs-internal.h | 9 ++++++++-
> > > > > 2 files changed, 23 insertions(+), 14 deletions(-)
> > > >
> > > > How was this found and tested? Did you forget an Assisted-by: tag?
> > >
> > > I am working on a series to improve kernfs_rwsem and going through
> > > review-prompt with AI to review my series and these were existing
> > > issues AI found. I have created reproducers with AI for these and
> > > tested that these patches those.
> >
> > Then please read our documentation for how to properly document this
> > usage of a LLM tool.
>
> Sure
>
> >
> > If you have reproducers, please add them to the kernfs tests as well as
> > patches part of this series when you resend them.
> >
>
> The reproducers are like stress tests and are targeting race conditions.
> In one case delay was added to fully expose the race. I am not sure
> selftests is the right place for this kind of tests. I can just publish
> the reproducer on the list to have them on record if that is what you
> are looking for.

Greg, let me know what would you prefer. I can add selftests which execise the
paths these bugs are on but to trigger the bug, more stress would be needed and
still will not trigger the bug always.

Also I have inflight kernfs selftest patch [1] as well. I can combine that to
this series.

[1] https://lore.kernel.org/all/20260902014050.499002-1-shakeel.butt@xxxxxxxxx/