Re: [PATCH 0/5] capabilities: close the ways around the CAP_SETFCAP rule for uid 0
From: Serge E. Hallyn
Date: Thu Oct 08 2026 - 12:22:05 EST
On Thu, Oct 08, 2026 at 03:39:30PM +0000, Josef Bacik wrote:
> On Thu, Oct 08, 2026 at 08:57:56AM -0500, Serge E. Hallyn wrote:
> > Would you mind describing what other solutions you considered? I've been
> > looking over this set since Tuesday, and finding it hard to reason about.
> > (Part of that is certainly the nature of the problem, and it's possible
> > that this is the best/simplest solution.)
>
> Everything we looked at kept the state in the cred. We model checked
> the variants before writing the code, and these fell over:
>
> - a bool per cred for "had CAP_SETFCAP over the parent when it entered".
> It breaks on two hops: setns() into a namespace that maps 0, unshare
> again, and the bool says yes for the second namespace. Hence the
> level.
> - checking only at uid_map write time. That misses setxattr of
> security.capability in a namespace that already maps 0, hence patch 4.
>
> We didn't look at keeping the state on the namespace.
>
> > If we replaced the userns->parent_could_setfcap bool with a ref to the
> > creator's cred, then at both setns and write we could check the actor's
> > credentials, right? There are probably issues with that specific idea,
> > but that's why it would be good to see what else you've considered.
>
> Checking at write time alone doesn't work: once a task is inside the
> namespace its cred says nothing about what it could do outside, so a
> joiner and the creator look the same.
>
> It does work if setns() refuses to join a namespace that maps, or can
> still map, the parent's uid 0 unless the joiner has CAP_SETFCAP over the
> parent. Then everybody in a namespace has the same reach, and it can be
Yeah, that's what I was thinking. Or even stricter: ensure that to join
any user namespace, a process must have a superset of the namespace
creator's capabilities.
> a level stored on the namespace at create time instead of a cred ref,
> which would pin keyrings and the rest for the life of the namespace.
> The checks would be setns(), the map write (opener and writer are in the
> parent, so a plain capable check), setxattr of security.capability, and
> ptrace.
>
> The difference in behaviour is that the -EPERM moves to setns(): a root
> task without CAP_SETFCAP couldn't enter a root-owned container that maps
> host uid 0 at all, where with this series it can enter and is refused
> only for the map, fscaps and ptrace. I can prototype it if you prefer
> that.
Sorry let me think about it (or let us talk about it) a bit more.
> > Of course UID 0 will always continue to carry privileges even with an
> > empty cap_eff. Here we're stopping it from writing filecaps to uid 0
> > owned files, but if it can open a 0 owned file on the host, like
> > /bin/sh or a systemd init file, or ptrace a process (in a child ns that
> > maps parent uid 0) doing so, it can still cause damage. My point being,
> > we do need to keep in mind the tradeoff of keeping the code simple
> > versus the realistic threat of the problem being addressed.
>
> On the same kernels, the restricted root task can copy a binary and
> chmod 4755 it (it owns it, no capability needed), and a uid 1000 user
> runs it with a full CapEff. With SECBIT_NOROOT that setuid copy gives
> uid 1000 nothing, while the fscap file still gives it what's in the
> xattr on an unpatched kernel. So SECBIT_NOROOT is the case the series
> adds anything for, the same case db2e718a4798 covers.
>
> A smaller version is patches 1, 4 and 5, with cap_root_level() moved
> from 2 into 4. I built that and ran the same flows: every route that
> ends in a file capability still gets -EPERM and uid 1000 gets nothing.
> The uid 0 map writes refused by 2 and 3 go through again, but the fscap
> write after them is refused.
>
> Let me know which way you'd like to go and I'll rework it.
>
> Thanks,
> Josef