On Wed, Sep 30, 2026 at 10:22 AM Andrei Vagin <[email protected]> wrote:
>
> On Sat, Sep 26, 2026 at 12:01 AM Kir Kolyshkin <[email protected]> wrote:
> >
> > A recursive open_tree(OPEN_TREE_CLONE) copies the nsfs mounts of any
> > mount namespaces pinned below the source. When the clone is attached
> > inside a mount namespace younger than a pinned one, move_mount(2) fails
> > with ELOOP from check_for_nsfs_mounts(), per the rule from commit
> > 8823c079ba71 ("vfs: Add setns support for the mount namespace") that
> > prevents mount namespace reference loops.
> >
> > This breaks container runtimes that use OPEN_TREE_NAMESPACE, such as
> > crun [1]. The new namespace is younger than everything else, so bind
> > mounting e.g. the host root fails on any host that pins a mount namespace
> > (snapd does, under /run/snapd/ns). By the time it fails, setns() has
> > already run and userspace cannot recover: the host tree is out of reach,
> > and the offending mounts cannot be unmounted from the detached copy.
> >
> > copy_mnt_ns() and create_new_namespace() already leave these mounts
> > behind; only get_detached_copy() copies them. Add an open_tree() flag to
> > drop them from the clone, as suggested by Aleksa [2]. Locked nsfs mounts
> > are dropped the same way copy_mnt_ns() does it, so nothing new is exposed.
> >
> > Keep it opt-in: open_tree(OPEN_TREE_CLONE) plus move_mount(2) is how
> > mount --rbind works through a file descriptor, and in the caller's own
> > or an older namespace such mounts remain usable.
> >
> > The flag requires OPEN_TREE_CLONE or OPEN_TREE_NAMESPACE. With the
> > latter it is a no-op, so a runtime can pass it unconditionally.
> >
> > Link: https://github.com/containers/crun/issues/2262 [1]
> > Link:
> > https://lore.kernel.org/all/[email protected]/
> > [2]
> > Assisted-by: Claude:claude-opus-5
> > Signed-off-by: Kir Kolyshkin <[email protected]>
> > ---
> > fs/namespace.c | 14 ++++++++++++--
> > include/uapi/linux/mount.h | 1 +
> > 2 files changed, 13 insertions(+), 2 deletions(-)
> >
> > diff --git a/fs/namespace.c b/fs/namespace.c
> > index ae5dc64f8b45..a95173996a8f 100644
> > --- a/fs/namespace.c
> > +++ b/fs/namespace.c
> > @@ -3062,7 +3062,8 @@ static struct mnt_namespace *get_detached_copy(const
> > struct path *path, unsigned
> > ns->seq_origin = src_mnt_ns->ns.ns_id;
> > }
> >
> > - mnt = __do_loopback(path, (flags & AT_RECURSIVE),
> > CL_COPY_MNT_NS_FILE);
> > + mnt = __do_loopback(path, (flags & AT_RECURSIVE),
> > + (flags & OPEN_TREE_DROP_MNTNS_MOUNTS) ? 0 :
> > CL_COPY_MNT_NS_FILE);
>
> Overall, the patch looks good. One small thing is that open_tree() with
> OPEN_TREE_DROP_MNTNS_MOUNTS doesn't fail on a mntns file if AT_RECURSIVE
> isn't set:
>
> open_tree(AT_FDCWD, "/proc/self/ns/mnt",
> OPEN_TREE_CLONE|OPEN_TREE_CLOEXEC|0x4) = 3
>
> With AT_RECURSIVE, __do_loopback() calls copy_tree(), which checks
> !(flag & CL_COPY_MNT_NS_FILE) && is_mnt_ns_file(dentry) and returns
> -EINVAL, but without AT_RECURSIVE it calls clone_mnt(), which doesn't
> have such check.
Thanks! A call with arguments like this does not make any sense*, as you're
asking to clone a pinned mount namespace while skipping it.
* except maybe to check if the path is indeed a pinned mount namespace,
but there are cheaper ways to do it.
That said, we can either keep things as is, or
1. require AT_RECURSIVE with OPEN_TREE_DROP_MNTNS_MOUNTS
when checking valid flags; probably the easiest thing to do;
2. add a check to a non-recursive path in __do_loopback, something like:
if (!(copy_flags & CL_COPY_MNT_NS_FILE) && is_mnt_ns_file(old_path->dentry))
return ERR_PTR(-EINVAL);
I'm going to implement either one in the next iteration (v3); do you have
a preference?
Kir.
>
> Thanks,
> Andrei