[PATCH 5/5] capabilities: don't let ptrace borrow CAP_SETFCAP

From: Josef Bacik

Date: Tue Oct 06 2026 - 11:52:09 EST


A task that entered a user namespace while it held CAP_SETFCAP outside may
map uid 0 of the parent and write file capabilities that are honoured
there. That is the one privilege it keeps outside of its namespace, and
the ptrace checks don't know about it: for a target in another user
namespace cap_ptrace_access_check() is satisfied with CAP_SYS_PTRACE over
the target's namespace, which the owner rule hands to any task with the
right euid, and inside the namespace everybody has a full set.

task A: uid 0, full caps task B: uid 0, no CAP_SETFCAP
unshare(CLONE_NEWUSER)
ptrace(PTRACE_ATTACH, A)
uids match, euid == ns->owner
make A write "0 0 1" to its uid_map,
or set file capabilities
A is entitled -> allowed

The same goes for /proc/<pid>/mem, process_vm_writev() and pidfd_getfd().

Require for PTRACE_MODE_ATTACH and PTRACE_TRACEME that CAP_SETFCAP of the
tracer reaches as far up as the target can make use of: to the topmost
ancestor, within the target's setfcap_level, whose root user is mapped
into the target's namespace, or can still be mapped because the namespace
has no uid map yet. CAP_SYS_PTRACE over that ancestor is accepted too,
since it gives control of tasks there that have CAP_SETFCAP anyway.

Targets in the initial namespace, in namespaces that were entered without
CAP_SETFCAP (what unprivileged users create) and in namespaces that don't
map the root user of an ancestor (the usual container) are not affected.
What remains is a tracer that gave up CAP_SETFCAP and CAP_SYS_PTRACE and a
target that didn't, in a namespace that shares its root user with the
tracer's. PTRACE_MODE_READ is unchanged.

Fixes: db2e718a4798 ("capabilities: require CAP_SETFCAP to map uid 0")
Assisted-by: LLM
Signed-off-by: Josef Bacik <josef@xxxxxxxxxxxxxx>
---
security/commoncap.c | 56 ++++++++++++++++++++++++++++++++++++++++++++++++----
1 file changed, 52 insertions(+), 4 deletions(-)

diff --git a/security/commoncap.c b/security/commoncap.c
index 26798d62e6b1..7263c78b6397 100644
--- a/security/commoncap.c
+++ b/security/commoncap.c
@@ -195,14 +195,52 @@ int cap_settime(const struct timespec64 *ts, const struct timezone *tz)
return 0;
}

+/*
+ * CAP_SETFCAP of a task can count in ancestors of its user namespace, see
+ * cap_setfcap_level(). That is a privilege outside of the namespace, so
+ * capabilities over the namespace are not enough to take control of the task.
+ */
+static bool cap_covers_setfcap(const struct cred *cred,
+ const struct cred *child_cred)
+{
+ struct user_namespace *ns = child_cred->user_ns, *seen, *p, *top = NULL;
+
+ if (child_cred->setfcap_level >= ns->level)
+ return true;
+
+ /*
+ * It is of use only for root users that are mapped into the namespace,
+ * or that can still be because there is no map yet.
+ */
+ seen = READ_ONCE(ns->uid_map.nr_extents) ? ns : ns->parent;
+ for (p = ns->parent; p; p = p->parent) {
+ if (p->level < child_cred->setfcap_level)
+ break;
+ if (kuid_has_mapping(seen, make_kuid(p, 0)))
+ top = p;
+ }
+ if (!top)
+ return true;
+
+ if (cred->user_ns == ns)
+ return cred->setfcap_level <= top->level;
+
+ /* CAP_SYS_PTRACE up there gives control of tasks with CAP_SETFCAP. */
+ return cap_setfcap_level(cred, ns->parent) <= top->level ||
+ !cap_capable(cred, top, CAP_SYS_PTRACE, CAP_OPT_NOAUDIT);
+}
+
/**
* cap_ptrace_access_check - Determine whether the current process may access
* another
* @child: The process to be accessed
* @mode: The mode of attachment.
*
- * If we are in the same or an ancestor user_ns and have all the target
- * task's capabilities, then ptrace access is allowed.
+ * For PTRACE_MODE_ATTACH, if the target task may make use of CAP_SETFCAP in
+ * an ancestor of its user_ns that our CAP_SETFCAP or CAP_SYS_PTRACE doesn't
+ * reach, then ptrace access is denied.
+ * Otherwise, if we are in the same or an ancestor user_ns and have all the
+ * target task's capabilities, then ptrace access is allowed.
* If we have the ptrace capability to the target user_ns, then ptrace
* access is allowed.
* Else denied.
@@ -223,11 +261,15 @@ int cap_ptrace_access_check(struct task_struct *child, unsigned int mode)
caller_caps = &cred->cap_effective;
else
caller_caps = &cred->cap_permitted;
+ if ((mode & PTRACE_MODE_ATTACH) &&
+ !cap_covers_setfcap(cred, child_cred))
+ goto deny;
if (cred->user_ns == child_cred->user_ns &&
cap_issubset(child_cred->cap_permitted, *caller_caps))
goto out;
if (ns_capable(child_cred->user_ns, CAP_SYS_PTRACE))
goto out;
+deny:
ret = -EPERM;
out:
rcu_read_unlock();
@@ -238,8 +280,11 @@ int cap_ptrace_access_check(struct task_struct *child, unsigned int mode)
* cap_ptrace_traceme - Determine whether another process may trace the current
* @parent: The task proposed to be the tracer
*
- * If parent is in the same or an ancestor user_ns and has all current's
- * capabilities, then ptrace access is allowed.
+ * If current may make use of CAP_SETFCAP in an ancestor of its user_ns that
+ * parent's CAP_SETFCAP or CAP_SYS_PTRACE doesn't reach, then ptrace access
+ * is denied.
+ * Otherwise, if parent is in the same or an ancestor user_ns and has all
+ * current's capabilities, then ptrace access is allowed.
* If parent has the ptrace capability to current's user_ns, then ptrace
* access is allowed.
* Else denied.
@@ -255,11 +300,14 @@ int cap_ptrace_traceme(struct task_struct *parent)
rcu_read_lock();
cred = __task_cred(parent);
child_cred = current_cred();
+ if (!cap_covers_setfcap(cred, child_cred))
+ goto deny;
if (cred->user_ns == child_cred->user_ns &&
cap_issubset(child_cred->cap_permitted, cred->cap_permitted))
goto out;
if (has_ns_capability(parent, child_cred->user_ns, CAP_SYS_PTRACE))
goto out;
+deny:
ret = -EPERM;
out:
rcu_read_unlock();

--
2.55.0