Re: [PATCH 0/5] capabilities: close the ways around the CAP_SETFCAP rule for uid 0
From: Josef Bacik
Date: Thu Oct 08 2026 - 11:46:13 EST
On Thu, Oct 08, 2026 at 08:57:56AM -0500, Serge E. Hallyn wrote:
> Would you mind describing what other solutions you considered? I've been
> looking over this set since Tuesday, and finding it hard to reason about.
> (Part of that is certainly the nature of the problem, and it's possible
> that this is the best/simplest solution.)
Everything we looked at kept the state in the cred. We model checked
the variants before writing the code, and these fell over:
- a bool per cred for "had CAP_SETFCAP over the parent when it entered".
It breaks on two hops: setns() into a namespace that maps 0, unshare
again, and the bool says yes for the second namespace. Hence the
level.
- checking only at uid_map write time. That misses setxattr of
security.capability in a namespace that already maps 0, hence patch 4.
We didn't look at keeping the state on the namespace.
> If we replaced the userns->parent_could_setfcap bool with a ref to the
> creator's cred, then at both setns and write we could check the actor's
> credentials, right? There are probably issues with that specific idea,
> but that's why it would be good to see what else you've considered.
Checking at write time alone doesn't work: once a task is inside the
namespace its cred says nothing about what it could do outside, so a
joiner and the creator look the same.
It does work if setns() refuses to join a namespace that maps, or can
still map, the parent's uid 0 unless the joiner has CAP_SETFCAP over the
parent. Then everybody in a namespace has the same reach, and it can be
a level stored on the namespace at create time instead of a cred ref,
which would pin keyrings and the rest for the life of the namespace.
The checks would be setns(), the map write (opener and writer are in the
parent, so a plain capable check), setxattr of security.capability, and
ptrace.
The difference in behaviour is that the -EPERM moves to setns(): a root
task without CAP_SETFCAP couldn't enter a root-owned container that maps
host uid 0 at all, where with this series it can enter and is refused
only for the map, fscaps and ptrace. I can prototype it if you prefer
that.
> Of course UID 0 will always continue to carry privileges even with an
> empty cap_eff. Here we're stopping it from writing filecaps to uid 0
> owned files, but if it can open a 0 owned file on the host, like
> /bin/sh or a systemd init file, or ptrace a process (in a child ns that
> maps parent uid 0) doing so, it can still cause damage. My point being,
> we do need to keep in mind the tradeoff of keeping the code simple
> versus the realistic threat of the problem being addressed.
On the same kernels, the restricted root task can copy a binary and
chmod 4755 it (it owns it, no capability needed), and a uid 1000 user
runs it with a full CapEff. With SECBIT_NOROOT that setuid copy gives
uid 1000 nothing, while the fscap file still gives it what's in the
xattr on an unpatched kernel. So SECBIT_NOROOT is the case the series
adds anything for, the same case db2e718a4798 covers.
A smaller version is patches 1, 4 and 5, with cap_root_level() moved
from 2 into 4. I built that and ran the same flows: every route that
ends in a file capability still gets -EPERM and uid 1000 gets nothing.
The uid 0 map writes refused by 2 and 3 go through again, but the fscap
write after them is refused.
Let me know which way you'd like to go and I'll rework it.
Thanks,
Josef