[PATCH 09/21] namespace: keep the lock on a mount that a propagated copy is moved beneath

From: Christian Brauner

Date: Fri Oct 02 2026 - 09:54:24 EST


attach_recursive_mnt() transfers MNT_LOCKED from the top mount to the
mount that is moved beneath it with MOVE_MOUNT_BENEATH. This allows the
owner of a user namespace to replace its locked /proc or its root. The
mount beneath takes on the job of covering the underlying mountpoint
allowing the top mount to be unmounted.

Consider two mount namespaces:

(H) The host H has a shared mount P with a secret in P/d
(Z) Z is a user namespace made from M. Its copy of P receives
propagation from the host's P and its copy of M's cover on P/d is
locked Z's owner must not get to see P/d.

Now (H) mounts X on P/d. This propagates into (Z). The copy of X lands
beneath (Z)'s locked cover. The locked property is transfered from (Z)'s
cover to the copy of X propagated beneath it. The cover is now unlocked.

Now (H) unmounts X again. The copy of X in (Z) gets unmounted and the
covering mount is left unlocked on top of P/d. (Z) can now unmount it:

Z: umount2(P/d) = EINVAL /* the cover is locked */
H: mount X on P/d, umount X /* both propagate into Z /*
Z: umount2(P/d) = 0 /* the cover is now unlocked */
Z: read P/d/secret = "covered-by-root" /* secret revealed */

So only transfer the locked property to the mount beneath for mounts the
caller has placed. A propagated copy that lands beneath a locked mount
is locked as well so that the mount at the bottom of the stack carries a
lock the way every check expects. The mount on top of it remains locked
to ensure that it keeps covering even if the propagated mount is
unmounted again.

Fixes: c62a4766937e ("move_mount: transfer MNT_LOCKED")
Cc: stable@xxxxxxxxxxxxxxx # v7.1+
Signed-off-by: Christian Brauner (Amutable) <brauner@xxxxxxxxxx>
---
fs/namespace.c | 14 ++++++++------
1 file changed, 8 insertions(+), 6 deletions(-)

diff --git a/fs/namespace.c b/fs/namespace.c
index e74e63466c24..bb0183ec2aaf 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -2712,15 +2712,17 @@ static int attach_recursive_mnt(struct mount *source_mnt,
/*
* If @q was locked it was meant to hide
* whatever was under it. Let @child take over
- * that job and lock it, then we can unlock @q.
- * That'll allow another namespace to shed @q
- * and reveal @child. Clearly, that mounter
- * consented to this by not severing the mount
- * relationship. Otherwise, what's the point.
+ * that job and lock it. If @child is the mount
+ * the caller placed we can then unlock @q:
+ * nothing another namespace does removes it
+ * again. A propagated copy goes away when the
+ * mounter of the original unmounts it, so @q
+ * keeps its lock.
*/
if (IS_MNT_LOCKED(q)) {
child->mnt.mnt_flags |= MNT_LOCKED;
- q->mnt.mnt_flags &= ~MNT_LOCKED;
+ if (child == source_mnt)
+ q->mnt.mnt_flags &= ~MNT_LOCKED;
}
mnt_change_mountpoint(r, mp, q);
}

--
2.53.0