[PATCH 3/5] userns: check the writer too before mapping uid 0
From: Josef Bacik
Date: Tue Oct 06 2026 - 11:50:31 EST
verify_root_map() only looks at file->f_cred. For every other privileged
id mapping new_idmap_permitted() wants the capability from the opener of
the map file and from the task that calls write(), but a map that consists
of the single line "0 0 1" takes the unprivileged branch, where only the
opener's euid counts. So an open uid_map file is a token for mapping
uid 0:
task A: uid 0, full caps task B: uid 0, no CAP_SETFCAP
fd = open("/proc/<pid>/uid_map")
B gets fd by inheritance
or SCM_RIGHTS
write(fd, "0 0 1")
opener had CAP_SETFCAP -> allowed
B set up a mapping that it is not allowed to set up.
Apply to the writer what is applied to the opener: in the namespace that
is being mapped, its credentials must have come in with CAP_SETFCAP over
the parent; anywhere else it needs CAP_SETFCAP over the parent now.
This only affects maps that contain uid 0 of the parent namespace, and
only when the file was opened by a task with CAP_SETFCAP and is written to
by one without. Those writes now fail with -EPERM.
Fixes: db2e718a4798 ("capabilities: require CAP_SETFCAP to map uid 0")
Assisted-by: LLM
Signed-off-by: Josef Bacik <josef@xxxxxxxxxxxxxx>
---
kernel/user_namespace.c | 32 +++++++++++++++++++++-----------
1 file changed, 21 insertions(+), 11 deletions(-)
diff --git a/kernel/user_namespace.c b/kernel/user_namespace.c
index bd5f9cea7430..421769e2d24f 100644
--- a/kernel/user_namespace.c
+++ b/kernel/user_namespace.c
@@ -891,8 +891,9 @@ EXPORT_SYMBOL_IF_KUNIT(uid_gid_map_sort);
* @new_map: requested idmap
*
* If a process requests mapping parent uid 0 into the new ns, verify that the
- * process writing the map had the CAP_SETFCAP capability as the target process
- * will be able to write fscaps that are valid in ancestor user namespaces.
+ * process that opened the map file and the process writing the map had the
+ * CAP_SETFCAP capability as the target process will be able to write fscaps
+ * that are valid in ancestor user namespaces.
*
* Return: true if the mapping is allowed, false if not.
*/
@@ -928,19 +929,28 @@ static bool verify_root_map(const struct file *file,
* when it unshared, and that the opener, which may have come
* in later with setns(), had it as well when it entered.
*/
- if (!file_ns->parent_could_setfcap)
+ if (!file_ns->parent_could_setfcap ||
+ file->f_cred->setfcap_level > level)
+ return false;
+ } else {
+ /* Process p1 is writing to uid_map of p2, who is in a child
+ * user namespace to p1's. Verify that the opener of the map
+ * file has CAP_SETFCAP against the parent of the new map
+ * namespace, and not just because it entered that.
+ */
+ if (!file_ns_capable(file, map_ns->parent, CAP_SETFCAP) ||
+ cap_setfcap_level(file->f_cred, map_ns->parent) > level)
return false;
- return file->f_cred->setfcap_level <= level;
}
- /* Process p1 is writing to uid_map of p2, who is in a child
- * user namespace to p1's. Verify that the opener of the map
- * file has CAP_SETFCAP against the parent of the new map
- * namespace, and not just because it entered that.
+ /* The file may have been handed to someone else since it was opened,
+ * so the same goes for the process that is doing the write.
*/
- if (!file_ns_capable(file, map_ns->parent, CAP_SETFCAP))
- return false;
- return cap_setfcap_level(file->f_cred, map_ns->parent) <= level;
+ if (map_ns == current_user_ns())
+ return current_cred()->setfcap_level <= level;
+
+ return ns_capable(map_ns->parent, CAP_SETFCAP) &&
+ cap_setfcap_level(current_cred(), map_ns->parent) <= level;
}
static ssize_t map_write(struct file *file, const char __user *buf,
--
2.55.0