Re: [RFC PATCH 0/7] cgroup: charge kernel work to the cgroup it is done for
From: Tejun Heo
Date: Thu Sep 24 2026 - 16:28:31 EST
Hello, Shakeel.
On Thu, Sep 24, 2026 at 11:47:04AM -0700, Shakeel Butt wrote:
> This series lets a kernel thread say which cgroup it is working for.
> That cgroup then sees the CPU time in its cpu.stat and the stalls in
> its memory.pressure, and the CPU time comes out of its cpu.max quota.
> The first user is the memcg reclaim that runs from high_work.
This doesn't translate to net rx, which is another major source of
displaced CPU usage. Switching membership on each packet isn't going to
work there. Attribution can't happen that way. We'd much rather count
per-cgroup received packets and prorate the CPU consumption. If at all
possible, I think it'd be better to adopt an approach which can cover
both use cases.
> - The debt is capped at one period's quota, so one long piece of
> work cannot starve the cgroup for long. Time over the cap still
> shows up in cpu.stat. Writing cpu.max or cpu.max.burst clears the
> debt.
I don't like the debt capping. Having debt doesn't have to mean that
the cgroup doesn't get any bandwidth at all. The cgroup just needs to
be slowed down enough that the generation of new work is throttled and
the whole thing doesn't go out of control. IO control already does
this: when IO debt is accumulated, userspace is heavily throttled, but
not completely stalled, until the whole cgroup's consumption comes
under control. I don't see why the debts would need to be forgiven
unconditionally. The cgroup can keep paying them while running at a
minimal rate to avoid triggering stall failures, and if the situation
doesn't resolve quickly, that will most likely trigger pressure based
kills in any reasonable setup anyway.
Thanks.
--
tejun