A recursive open_tree(OPEN_TREE_CLONE) copies the nsfs mounts of any
mount namespaces pinned below the source.  When the clone is attached
inside a mount namespace younger than a pinned one, move_mount(2) fails
with ELOOP from check_for_nsfs_mounts(), per the rule from commit
8823c079ba71 ("vfs: Add setns support for the mount namespace") that
prevents mount namespace reference loops.

This breaks container runtimes that use OPEN_TREE_NAMESPACE, such as
crun [1].  The new namespace is younger than everything else, so bind
mounting e.g. the host root fails on any host that pins a mount namespace
(snapd does, under /run/snapd/ns).  By the time it fails, setns() has
already run and userspace cannot recover: the host tree is out of reach,
and the offending mounts cannot be unmounted from the detached copy.

copy_mnt_ns() and create_new_namespace() already leave these mounts
behind; only get_detached_copy() copies them.  Add an open_tree() flag to
drop them from the clone, as suggested by Aleksa [2].  Locked nsfs mounts
are dropped the same way copy_mnt_ns() does it, so nothing new is exposed.

Keep it opt-in: open_tree(OPEN_TREE_CLONE) plus move_mount(2) is how
mount --rbind works through a file descriptor, and in the caller's own
or an older namespace such mounts remain usable.

The flag requires OPEN_TREE_CLONE or OPEN_TREE_NAMESPACE.  With the
latter it is a no-op, so a runtime can pass it unconditionally.

Link: https://github.com/containers/crun/issues/2262 [1]
Link: 
https://lore.kernel.org/all/[email protected]/
 [2]
Assisted-by: Claude:claude-opus-5
Signed-off-by: Kir Kolyshkin <[email protected]>
---
 fs/namespace.c             | 14 ++++++++++++--
 include/uapi/linux/mount.h |  1 +
 2 files changed, 13 insertions(+), 2 deletions(-)

diff --git a/fs/namespace.c b/fs/namespace.c
index ae5dc64f8b45..a95173996a8f 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -3062,7 +3062,8 @@ static struct mnt_namespace *get_detached_copy(const 
struct path *path, unsigned
                        ns->seq_origin = src_mnt_ns->ns.ns_id;
        }
 
-       mnt = __do_loopback(path, (flags & AT_RECURSIVE), CL_COPY_MNT_NS_FILE);
+       mnt = __do_loopback(path, (flags & AT_RECURSIVE),
+                           (flags & OPEN_TREE_DROP_MNTNS_MOUNTS) ? 0 : 
CL_COPY_MNT_NS_FILE);
        if (IS_ERR(mnt)) {
                emptied_ns = ns;
                return ERR_CAST(mnt);
@@ -3204,7 +3205,16 @@ static struct file *vfs_open_tree(int dfd, const char 
__user *filename, unsigned
 
        if (flags & ~(AT_EMPTY_PATH | AT_NO_AUTOMOUNT | AT_RECURSIVE |
                      AT_SYMLINK_NOFOLLOW | OPEN_TREE_CLONE |
-                     OPEN_TREE_CLOEXEC | OPEN_TREE_NAMESPACE))
+                     OPEN_TREE_CLOEXEC | OPEN_TREE_NAMESPACE |
+                     OPEN_TREE_DROP_MNTNS_MOUNTS))
+               return ERR_PTR(-EINVAL);
+
+       /*
+        * Only meaningful when a tree is copied.  OPEN_TREE_NAMESPACE never
+        * copies pinned mount namespaces, so there the flag is a no-op.
+        */
+       if ((flags & OPEN_TREE_DROP_MNTNS_MOUNTS) &&
+           !(flags & (OPEN_TREE_CLONE | OPEN_TREE_NAMESPACE)))
                return ERR_PTR(-EINVAL);
 
        if ((flags & (AT_RECURSIVE | OPEN_TREE_CLONE | OPEN_TREE_NAMESPACE)) ==
diff --git a/include/uapi/linux/mount.h b/include/uapi/linux/mount.h
index 2204708dbf7a..ac865edac517 100644
--- a/include/uapi/linux/mount.h
+++ b/include/uapi/linux/mount.h
@@ -63,6 +63,7 @@
  */
 #define OPEN_TREE_CLONE                (1 << 0)        /* Clone the target 
tree and attach the clone */
 #define OPEN_TREE_NAMESPACE    (1 << 1)        /* Clone the target tree into a 
new mount namespace */
+#define OPEN_TREE_DROP_MNTNS_MOUNTS    (1 << 2)        /* Drop mntns mounts 
from the clone */
 #define OPEN_TREE_CLOEXEC      O_CLOEXEC       /* Close the file on execve() */
 
 /*

-- 
2.55.0


Reply via email to