On Wed, Sep 30, 2026 at 12:49 PM Kirill Kolyshkin <[email protected]> wrote: > > On Wed, Sep 30, 2026 at 10:22 AM Andrei Vagin <[email protected]> wrote: > > > > On Sat, Sep 26, 2026 at 12:01 AM Kir Kolyshkin <[email protected]> wrote: > > > > > > A recursive open_tree(OPEN_TREE_CLONE) copies the nsfs mounts of any > > > mount namespaces pinned below the source. When the clone is attached > > > inside a mount namespace younger than a pinned one, move_mount(2) fails > > > with ELOOP from check_for_nsfs_mounts(), per the rule from commit > > > 8823c079ba71 ("vfs: Add setns support for the mount namespace") that > > > prevents mount namespace reference loops. > > > > > > This breaks container runtimes that use OPEN_TREE_NAMESPACE, such as > > > crun [1]. The new namespace is younger than everything else, so bind > > > mounting e.g. the host root fails on any host that pins a mount namespace > > > (snapd does, under /run/snapd/ns). By the time it fails, setns() has > > > already run and userspace cannot recover: the host tree is out of reach, > > > and the offending mounts cannot be unmounted from the detached copy. > > > > > > copy_mnt_ns() and create_new_namespace() already leave these mounts > > > behind; only get_detached_copy() copies them. Add an open_tree() flag to > > > drop them from the clone, as suggested by Aleksa [2]. Locked nsfs mounts > > > are dropped the same way copy_mnt_ns() does it, so nothing new is exposed. > > > > > > Keep it opt-in: open_tree(OPEN_TREE_CLONE) plus move_mount(2) is how > > > mount --rbind works through a file descriptor, and in the caller's own > > > or an older namespace such mounts remain usable. > > > > > > The flag requires OPEN_TREE_CLONE or OPEN_TREE_NAMESPACE. With the > > > latter it is a no-op, so a runtime can pass it unconditionally. > > > > > > Link: https://github.com/containers/crun/issues/2262 [1] > > > Link: > > > https://lore.kernel.org/all/[email protected]/ > > > [2] > > > Assisted-by: Claude:claude-opus-5 > > > Signed-off-by: Kir Kolyshkin <[email protected]> > > > --- > > > fs/namespace.c | 14 ++++++++++++-- > > > include/uapi/linux/mount.h | 1 + > > > 2 files changed, 13 insertions(+), 2 deletions(-) > > > > > > diff --git a/fs/namespace.c b/fs/namespace.c > > > index ae5dc64f8b45..a95173996a8f 100644 > > > --- a/fs/namespace.c > > > +++ b/fs/namespace.c > > > @@ -3062,7 +3062,8 @@ static struct mnt_namespace > > > *get_detached_copy(const struct path *path, unsigned > > > ns->seq_origin = src_mnt_ns->ns.ns_id; > > > } > > > > > > - mnt = __do_loopback(path, (flags & AT_RECURSIVE), > > > CL_COPY_MNT_NS_FILE); > > > + mnt = __do_loopback(path, (flags & AT_RECURSIVE), > > > + (flags & OPEN_TREE_DROP_MNTNS_MOUNTS) ? 0 : > > > CL_COPY_MNT_NS_FILE); > > > > Overall, the patch looks good. One small thing is that open_tree() with > > OPEN_TREE_DROP_MNTNS_MOUNTS doesn't fail on a mntns file if AT_RECURSIVE > > isn't set: > > > > open_tree(AT_FDCWD, "/proc/self/ns/mnt", > > OPEN_TREE_CLONE|OPEN_TREE_CLOEXEC|0x4) = 3 > > > > With AT_RECURSIVE, __do_loopback() calls copy_tree(), which checks > > !(flag & CL_COPY_MNT_NS_FILE) && is_mnt_ns_file(dentry) and returns > > -EINVAL, but without AT_RECURSIVE it calls clone_mnt(), which doesn't > > have such check. > > Thanks! A call with arguments like this does not make any sense*, as you're > asking to clone a pinned mount namespace while skipping it. > > * except maybe to check if the path is indeed a pinned mount namespace, > but there are cheaper ways to do it. > > That said, we can either keep things as is, or > 1. require AT_RECURSIVE with OPEN_TREE_DROP_MNTNS_MOUNTS > when checking valid flags; probably the easiest thing to do; > 2. add a check to a non-recursive path in __do_loopback, something like: > if (!(copy_flags & CL_COPY_MNT_NS_FILE) && > is_mnt_ns_file(old_path->dentry)) > return ERR_PTR(-EINVAL); > > I'm going to implement either one in the next iteration (v3); do you have > a preference?
I prefer the first option. Thanks, Andrei

