[PATCH 4/5] capabilities: limit fscaps to where CAP_SETFCAP reaches
From: Josef Bacik
Date: Tue Oct 06 2026 - 11:52:35 EST
Commit db2e718a4798 ("capabilities: require CAP_SETFCAP to map uid 0")
keeps a root process without CAP_SETFCAP from creating a user namespace
that maps uid 0, in which it could attach capabilities to a file that are
then honoured outside. It does not have to create one, though. Any
namespace will do that maps uid 0 and that it may enter, and it may enter
all those that were created by uid 0:
task A: uid 0, full caps task B: uid 0, no CAP_SETFCAP
unshare(CLONE_NEWUSER)
write "0 0 1" to uid_map, gid_map
setns(A's user ns)
full capability set
setxattr(file, "security.capability")
rootid = kuid 0
(back in the initial namespace)
execve(file)
file capabilities apply
The map does not even have to show uid 0 of the parent as uid 0, since a
v3 xattr can name any mapped uid as the root user.
Check at the place where it matters, in cap_convert_nscap(). On an
idmapped mount the root user named in the xattr has two identities: the
kuid seen through this mount, and the kuid that is stored and seen
through every other mount of the filesystem. If either of them is the
root user of an ancestor of the caller's namespace too, CAP_SETFCAP has
to reach up to there, which cred->setfcap_level tells.
Tasks in the initial namespace, tasks that held CAP_SETFCAP when they
entered their namespace, and file capabilities for a root user that only
exists in namespaces the caller came into with CAP_SETFCAP or below are
not affected; that includes everything an unprivileged user does in his
own namespaces. The rest gets -EPERM.
Fixes: 8db6c34f1dbc ("Introduce v3 namespaced file capabilities")
Assisted-by: LLM
Signed-off-by: Josef Bacik <josef@xxxxxxxxxxxxxx>
---
security/commoncap.c | 14 ++++++++++++++
1 file changed, 14 insertions(+)
diff --git a/security/commoncap.c b/security/commoncap.c
index b406ede2fadc..26798d62e6b1 100644
--- a/security/commoncap.c
+++ b/security/commoncap.c
@@ -644,10 +644,24 @@ int cap_convert_nscap(const struct mnt_idmap *idmap, struct dentry *dentry,
if (!vfsuid_valid(vfsrootid))
return -EINVAL;
+ /*
+ * The root user may be the root user of ancestors of our namespace as
+ * well. CAP_SETFCAP that we got for entering it doesn't cover those.
+ * On an idmapped mount the root user is vfsrootid as seen through
+ * this mount and rootid as seen through every other mount of the
+ * filesystem, so both have to stay within reach.
+ */
+ if (cap_root_level(vfsuid_into_kuid(vfsrootid), task_ns) <
+ current_cred()->setfcap_level)
+ return -EPERM;
+
rootid = from_vfsuid(idmap, fs_ns, vfsrootid);
if (!uid_valid(rootid))
return -EINVAL;
+ if (cap_root_level(rootid, task_ns) < current_cred()->setfcap_level)
+ return -EPERM;
+
nsrootid = from_kuid(fs_ns, rootid);
if (nsrootid == -1)
return -EINVAL;
--
2.55.0