On Sat, Sep 26, 2026 at 12:01 AM Kir Kolyshkin <[email protected]> wrote:
>
> A recursive open_tree(OPEN_TREE_CLONE) copies the nsfs mounts of any
> mount namespaces pinned below the source.  When the clone is attached
> inside a mount namespace younger than a pinned one, move_mount(2) fails
> with ELOOP from check_for_nsfs_mounts(), per the rule from commit
> 8823c079ba71 ("vfs: Add setns support for the mount namespace") that
> prevents mount namespace reference loops.
>
> This breaks container runtimes that use OPEN_TREE_NAMESPACE, such as
> crun [1].  The new namespace is younger than everything else, so bind
> mounting e.g. the host root fails on any host that pins a mount namespace
> (snapd does, under /run/snapd/ns).  By the time it fails, setns() has
> already run and userspace cannot recover: the host tree is out of reach,
> and the offending mounts cannot be unmounted from the detached copy.
>
> copy_mnt_ns() and create_new_namespace() already leave these mounts
> behind; only get_detached_copy() copies them.  Add an open_tree() flag to
> drop them from the clone, as suggested by Aleksa [2].  Locked nsfs mounts
> are dropped the same way copy_mnt_ns() does it, so nothing new is exposed.
>
> Keep it opt-in: open_tree(OPEN_TREE_CLONE) plus move_mount(2) is how
> mount --rbind works through a file descriptor, and in the caller's own
> or an older namespace such mounts remain usable.
>
> The flag requires OPEN_TREE_CLONE or OPEN_TREE_NAMESPACE.  With the
> latter it is a no-op, so a runtime can pass it unconditionally.
>
> Link: https://github.com/containers/crun/issues/2262 [1]
> Link: 
> https://lore.kernel.org/all/[email protected]/
>  [2]
> Assisted-by: Claude:claude-opus-5
> Signed-off-by: Kir Kolyshkin <[email protected]>
> ---
>  fs/namespace.c             | 14 ++++++++++++--
>  include/uapi/linux/mount.h |  1 +
>  2 files changed, 13 insertions(+), 2 deletions(-)
>
> diff --git a/fs/namespace.c b/fs/namespace.c
> index ae5dc64f8b45..a95173996a8f 100644
> --- a/fs/namespace.c
> +++ b/fs/namespace.c
> @@ -3062,7 +3062,8 @@ static struct mnt_namespace *get_detached_copy(const 
> struct path *path, unsigned
>                         ns->seq_origin = src_mnt_ns->ns.ns_id;
>         }
>
> -       mnt = __do_loopback(path, (flags & AT_RECURSIVE), 
> CL_COPY_MNT_NS_FILE);
> +       mnt = __do_loopback(path, (flags & AT_RECURSIVE),
> +                           (flags & OPEN_TREE_DROP_MNTNS_MOUNTS) ? 0 : 
> CL_COPY_MNT_NS_FILE);

Overall, the patch looks good. One small thing is that open_tree() with
OPEN_TREE_DROP_MNTNS_MOUNTS doesn't fail on a mntns file if AT_RECURSIVE
isn't set:

  open_tree(AT_FDCWD, "/proc/self/ns/mnt",
            OPEN_TREE_CLONE|OPEN_TREE_CLOEXEC|0x4) = 3

With AT_RECURSIVE, __do_loopback() calls copy_tree(), which checks
!(flag & CL_COPY_MNT_NS_FILE) && is_mnt_ns_file(dentry) and returns
-EINVAL, but without AT_RECURSIVE it calls clone_mnt(), which doesn't
have such check.

Thanks,
Andrei

Reply via email to