[PATCH 06/21] namespace: handle mount locking for automounts correctly

From: Christian Brauner

Date: Fri Oct 02 2026 - 09:57:43 EST


If mounts are propagated across user namespaces, attach_recursive_mnt()
locks every copy of the source mount to protect overmounts from
vanishing and revealing the underlying files or directories.

The user namespace is taken from the caller's mount namespaces since
this is where the mounts end up. Except, that's not always true.
Automounts may legitimately get popped in by tasks located in a
different mount and user namespace during path lookup.

Then check doesn't make sense at that point. The copy in the namespace
of the parent - which may be the host's - is now locked and the host
cannot change the flags of its own mounts anymore.

Here's the reproducer:

- host has debugfs mounted nosuid,nodev,noexec and shared
- tracefs automount below it is not yet active
- hand a child process a directory descriptor on that mount
- child process enters auser and mount namespace
- child process' namespace now holds a locked copy of the host's debugfs
mount that receives propagation from it
- child process stats "tracing/." through the descriptor
- lookup runs on the host's mount so the automount lands below the host's mount
- propagation puts a copy below the child's copy
- both try to clear the flags on the mount they got, with a bind remount:

host, on its own automount: MS_REMOUNT|MS_BIND = EPERM
child, on the copy in its namespace: MS_REMOUNT|MS_BIND = 0

So the lock landed on the host's mount instead of the child's copy.
Congrats. So we need to compare with the owner of the namespace the
mount actually gets mounted on. For all regular cases that is the
caller's mount namespace and so nothing changes.

Detached trees in anonymous mount namespaces by be handed over via
SCM_RIGHTS or inherited in other ways on purpose so the attaching task's
mount namespace is authoritative, not the creator of the detached tree.

Fixes: 132c94e31b8b ("vfs: Carefully propogate mounts across user namespaces")
Cc: stable@xxxxxxxxxxxxxxx
Signed-off-by: Christian Brauner (Amutable) <brauner@xxxxxxxxxx>
---
fs/namespace.c | 18 +++++++++++++++++-
1 file changed, 17 insertions(+), 1 deletion(-)

diff --git a/fs/namespace.c b/fs/namespace.c
index 60b57572fc64..27bf8665ed58 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -2604,11 +2604,11 @@ enum mnt_tree_flags_t {
static int attach_recursive_mnt(struct mount *source_mnt,
const struct pinned_mountpoint *dest)
{
- struct user_namespace *user_ns = current->nsproxy->mnt_ns->user_ns;
struct mount *dest_mnt = dest->parent;
struct mountpoint *dest_mp = dest->mp;
HLIST_HEAD(tree_list);
struct mnt_namespace *ns = dest_mnt->mnt_ns;
+ struct user_namespace *user_ns = ns->user_ns;
struct pinned_mountpoint root = {};
struct mountpoint *shorter = NULL;
struct mount *child, *p;
@@ -2617,6 +2617,22 @@ static int attach_recursive_mnt(struct mount *source_mnt,
int err = 0;
bool moving = mnt_has_parent(source_mnt);

+ /*
+ * A caller in an unprivileged mount namespaces may trigger an
+ * automount and propagate locked mounts into privileged mount
+ * namespaces. Take ownership from the target mount namespace.
+ * It's equivalent for everything but the automount case.
+ *
+ * Detached trees in anonymous mount namespaces by be handed
+ * over via SCM_RIGHTS or inherited in other ways on purpose
+ * the attaching task's mount namespace is authoritative, not
+ * the creator of the detached tree.
+ */
+ if (is_anon_ns(ns))
+ user_ns = current->nsproxy->mnt_ns->user_ns;
+ else
+ user_ns = ns->user_ns;
+
/*
* Preallocate a mountpoint in case the new mounts need to be
* mounted beneath mounts on the same mountpoint.

--
2.53.0