[PATCH bpf-next 0/4] bpf: BPF-driven proactive memcg reclaim
From: Hui Zhu
Date: Fri Aug 07 2026 - 03:13:16 EST
From: Hui Zhu <zhuhui@xxxxxxxxxx>
This series lets a BPF program decide when to trigger memcg reclaim
and how aggressively to do it, based on whatever runtime signal it
chooses to observe -- rather than reclaim only being triggered once a
cgroup's usage crosses a fixed threshold. The core idea is a new kfunc,
bpf_try_to_free_mem_cgroup_pages(), which gives BPF direct access to
the reclaim path so this decision can be made in BPF policy rather than
hard-coded threshold logic.
This was originally part of a larger series posted here [1].
That series also adds a memcg BPF struct_ops (memcg_charged,
memcg_uncharged, below_low, below_min) for synchronous, in-line memory
protection decisions. That mechanism and this one solve different
problems -- struct_ops hooks run inline on the charge/reclaim path,
while the kfunc here is for asynchronous, out-of-band reclaim decided
independently by a BPF program -- so I think they're better reviewed
as separate series rather than bundled together. This series carries
only the async reclaim piece: the bpf_try_to_free_mem_cgroup_pages
kfunc (patch 1), plus a new bpf_thread_wq mechanism (patch 2) that
grew out of discussion here [2].
Patch 1 adds bpf_try_to_free_mem_cgroup_pages(), a sleepable kfunc
wrapping try_to_free_mem_cgroup_pages(). With it, a BPF program can
reclaim from a given cgroup on its own terms -- any condition it can
observe at runtime -- instead of being limited to "usage hit
threshold X".
Patch 2 adds bpf_thread_wq, a bpf_wq-like map field backed by a
dedicated kthread_worker rather than the system workqueue, which can
be attached to a specific cgroup at init time. The motivation is
accounting: reclaim work triggered from BPF still costs CPU (and can
itself touch memory), and running it via a plain irq_work/system
workqueue callback would charge that cost to whatever context happens
to run it, not to the cgroup the policy cares about. bpf_thread_wq
lets that cost be attributed to a chosen cgroup instead -- e.g. the
low-priority cgroup being reclaimed from.
Patch 3 is a selftest that exercises bpf_thread_wq's cgroup attachment
in isolation: verifying the callback observes the target cgroup when
one is given, and does not when it isn't.
Patch 4 (selftests/bpf: add memcg async reclaim test for
bpf_wq/bpf_thread_wq) ties patches 1 and 2 together as a worked
example: it watches the WORKINGSET_REFAULT_FILE counter of a
high-priority cgroup as a proxy for memory-pressure impact, and once
it starts climbing, proactively reclaims pages from a low-priority
cgroup via bpf_try_to_free_mem_cgroup_pages, running that reclaim
inside a bpf_thread_wq attached to the low-priority cgroup so the
reclaim cost lands on it rather than leaking into an unrelated
context. This demonstrates the end-to-end use case: BPF observes
pressure on the cgroup it wants to protect, and reclaims from the
cgroup it wants to charge, in one self-contained mechanism.
[1] https://sashiko.dev/#/message/cover.1779760876.git.zhuhui%40kylinos.cn
[2] https://sashiko.dev/#/message/1b58d56976202f26818d31dbd0da2ecb2e2460f5%40linux.dev
Hui Zhu (4):
mm/bpf: Add bpf_try_to_free_mem_cgroup_pages kfunc
bpf: add bpf_thread_wq kthread-backed workqueue with cgroup placement
selftests/bpf: add thread_wq cgroup test
selftests/bpf: add memcg async reclaim test for bpf_wq/bpf_thread_wq
include/linux/bpf.h | 15 +-
include/linux/cgroup.h | 2 +
include/uapi/linux/bpf.h | 4 +
kernel/bpf/btf.c | 7 +
kernel/bpf/helpers.c | 418 +++++++++++++++
kernel/bpf/syscall.c | 15 +-
kernel/bpf/verifier.c | 44 +-
kernel/cgroup/cgroup.c | 13 +
mm/bpf_memcontrol.c | 58 +++
.../testing/selftests/bpf/bpf_experimental.h | 7 +
.../bpf/prog_tests/memcg_async_reclaim.c | 479 ++++++++++++++++++
.../bpf/prog_tests/thread_wq_cgroup.c | 87 ++++
.../selftests/bpf/progs/memcg_async_reclaim.c | 255 ++++++++++
.../selftests/bpf/progs/thread_wq_cgroup.c | 56 ++
14 files changed, 1455 insertions(+), 5 deletions(-)
create mode 100644 tools/testing/selftests/bpf/prog_tests/memcg_async_reclaim.c
create mode 100644 tools/testing/selftests/bpf/prog_tests/thread_wq_cgroup.c
create mode 100644 tools/testing/selftests/bpf/progs/memcg_async_reclaim.c
create mode 100644 tools/testing/selftests/bpf/progs/thread_wq_cgroup.c
--
2.53.0