A recursive open_tree(OPEN_TREE_CLONE) copies the nsfs mounts of any
mount namespaces pinned below the source. When the clone is attached
inside a mount namespace younger than a pinned one, move_mount(2) fails
with ELOOP from check_for_nsfs_mounts(), per the rule from commit
8823c079ba71 ("vfs: Add setns support for the mount namespace") that
prevents mount namespace reference loops.
This breaks container runtimes that use OPEN_TREE_NAMESPACE, such as
crun [1]. The new namespace is younger than everything else, so bind
mounting e.g. the host root fails on any host that pins a mount namespace
(snapd does, under /run/snapd/ns). By the time it fails, setns() has
already run and userspace cannot recover: the host tree is out of reach,
and the offending mounts cannot be unmounted from the detached copy.
copy_mnt_ns() and create_new_namespace() already leave these mounts
behind; only get_detached_copy() copies them. Add an open_tree() flag to
drop them from the clone, as suggested by Aleksa [2]. Locked nsfs mounts
are dropped the same way copy_mnt_ns() does it, so nothing new is exposed.
Keep it opt-in: open_tree(OPEN_TREE_CLONE) plus move_mount(2) is how
mount --rbind works through a file descriptor, and in the caller's own
or an older namespace such mounts remain usable.
The flag requires OPEN_TREE_CLONE or OPEN_TREE_NAMESPACE. With the
latter it is a no-op, so a runtime can pass it unconditionally.
Link: https://github.com/containers/crun/issues/2262 [1]
Link:
https://lore.kernel.org/all/[email protected]/
[2]
Assisted-by: Claude:claude-opus-5
Signed-off-by: Kir Kolyshkin <[email protected]>
---
fs/namespace.c | 14 ++++++++++++--
include/uapi/linux/mount.h | 1 +
2 files changed, 13 insertions(+), 2 deletions(-)
diff --git a/fs/namespace.c b/fs/namespace.c
index ae5dc64f8b45..a95173996a8f 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -3062,7 +3062,8 @@ static struct mnt_namespace *get_detached_copy(const
struct path *path, unsigned
ns->seq_origin = src_mnt_ns->ns.ns_id;
}
- mnt = __do_loopback(path, (flags & AT_RECURSIVE), CL_COPY_MNT_NS_FILE);
+ mnt = __do_loopback(path, (flags & AT_RECURSIVE),
+ (flags & OPEN_TREE_DROP_MNTNS_MOUNTS) ? 0 :
CL_COPY_MNT_NS_FILE);
if (IS_ERR(mnt)) {
emptied_ns = ns;
return ERR_CAST(mnt);
@@ -3204,7 +3205,16 @@ static struct file *vfs_open_tree(int dfd, const char
__user *filename, unsigned
if (flags & ~(AT_EMPTY_PATH | AT_NO_AUTOMOUNT | AT_RECURSIVE |
AT_SYMLINK_NOFOLLOW | OPEN_TREE_CLONE |
- OPEN_TREE_CLOEXEC | OPEN_TREE_NAMESPACE))
+ OPEN_TREE_CLOEXEC | OPEN_TREE_NAMESPACE |
+ OPEN_TREE_DROP_MNTNS_MOUNTS))
+ return ERR_PTR(-EINVAL);
+
+ /*
+ * Only meaningful when a tree is copied. OPEN_TREE_NAMESPACE never
+ * copies pinned mount namespaces, so there the flag is a no-op.
+ */
+ if ((flags & OPEN_TREE_DROP_MNTNS_MOUNTS) &&
+ !(flags & (OPEN_TREE_CLONE | OPEN_TREE_NAMESPACE)))
return ERR_PTR(-EINVAL);
if ((flags & (AT_RECURSIVE | OPEN_TREE_CLONE | OPEN_TREE_NAMESPACE)) ==
diff --git a/include/uapi/linux/mount.h b/include/uapi/linux/mount.h
index 2204708dbf7a..ac865edac517 100644
--- a/include/uapi/linux/mount.h
+++ b/include/uapi/linux/mount.h
@@ -63,6 +63,7 @@
*/
#define OPEN_TREE_CLONE (1 << 0) /* Clone the target
tree and attach the clone */
#define OPEN_TREE_NAMESPACE (1 << 1) /* Clone the target tree into a
new mount namespace */
+#define OPEN_TREE_DROP_MNTNS_MOUNTS (1 << 2) /* Drop mntns mounts
from the clone */
#define OPEN_TREE_CLOEXEC O_CLOEXEC /* Close the file on execve() */
/*
--
2.55.0