On Wed, Sep 02, 2026 at 11:15:16PM -0700, Shakeel Butt wrote:
> On Thu, Sep 03, 2026 at 07:41:53AM +0200, Greg Kroah-Hartman wrote:
> > On Wed, Sep 02, 2026 at 10:37:21PM -0700, Shakeel Butt wrote:
> > > On Thu, Sep 03, 2026 at 06:31:58AM +0200, Greg Kroah-Hartman wrote:
> > > > On Wed, Sep 02, 2026 at 09:02:51PM -0700, Shakeel Butt wrote:
> > > > > kernfs_rename_ns() only takes kernfs_rename_lock when the rename moves
> > > > > the node to a new parent.  A rename that keeps the same parent, like
> > > > > renaming a network interface, changes kernfs_node::name with only
> > > > > kernfs_rwsem held.  So the lock protects ->__parent but not ->name, 
> > > > > and
> > > > > a reader that wants a stable name has to take kernfs_rwsem, the same
> > > > > lock every path lookup needs.
> > > > > 
> > > > > That also makes for a small but real bug.  kernfs_path_from_node()
> > > > > takes kernfs_rename_lock for reading, and 
> > > > > kernfs_path_from_node_locked()
> > > > > then reads the name of each ancestor.  It reads each one once, so a
> > > > > single same-parent rename only moves the answer from the old path to 
> > > > > the
> > > > > new one, but two of them landing inside one walk build a path that 
> > > > > never
> > > > > existed:
> > > > > 
> > > > >   CPU0                                   CPU1
> > > > >   kernfs_path_from_node() on /a/b/c
> > > > >     reads the name of a, gets "a"
> > > > >                                          renames a to a2
> > > > >                                          renames b to b2
> > > > >     reads the name of b, gets "b2"
> > > > >     returns "/a/b2/c"
> > > > > 
> > > > > This hits roots without KERNFS_ROOT_INVARIANT_PARENT: sysfs, where the
> > > > > bad path can reach sysfs_warn_dup() and pr_cont_kernfs_path(), and
> > > > > resctrl, which renames a mon group inside its mon_groups directory.
> > > > > cgroup sets the flag, so it skips the lock and reads names under RCU
> > > > > alone; that case needs something else and is not addressed here.
> > > > > 
> > > > > So take the lock in both cases, and let kernfs_rcu_name() accept it 
> > > > > the
> > > > > way kernfs_parent() already does for ->__parent.  Same-parent renames
> > > > > are rare, the lock is per filesystem, and the locked section is at 
> > > > > most
> > > > > three stores.  It also gives a future rename sequence counter one 
> > > > > place
> > > > > to sit that covers every rename.
> > > > > 
> > > > > Fixes: 741c10b096bc ("kernfs: Use RCU to access kernfs_node::name.")
> > > > > Signed-off-by: Shakeel Butt <[email protected]>
> > > > > ---
> > > > >  fs/kernfs/dir.c             | 28 +++++++++++++++-------------
> > > > >  fs/kernfs/kernfs-internal.h |  9 ++++++++-
> > > > >  2 files changed, 23 insertions(+), 14 deletions(-)
> > > > 
> > > > How was this found and tested?  Did you forget an Assisted-by: tag?
> > > 
> > > I am working on a series to improve kernfs_rwsem and going through
> > > review-prompt with AI to review my series and these were existing 
> > > issues AI found. I have created reproducers with AI for these and
> > > tested that these patches those.
> > 
> > Then please read our documentation for how to properly document this
> > usage of a LLM tool.
> 
> Sure
> 
> > 
> > If you have reproducers, please add them to the kernfs tests as well as
> > patches part of this series when you resend them.
> >
> 
> The reproducers are like stress tests and are targeting race conditions.
> In one case delay was added to fully expose the race. I am not sure
> selftests is the right place for this kind of tests. I can just publish
> the reproducer on the list to have them on record if that is what you
> are looking for.

Greg, let me know what would you prefer. I can add selftests which execise the
paths these bugs are on but to trigger the bug, more stress would be needed and
still will not trigger the bug always.

Also I have inflight kernfs selftest patch [1] as well. I can combine that to
this series. 

[1] https://lore.kernel.org/all/[email protected]/

Reply via email to