Re: [PATCH v2] selftests/cgroup: account for zswap shrinker writeback
From: Joshua Hahn
Date: Thu Sep 03 2026 - 11:44:45 EST
On Wed, 2 Sep 2026 18:04:08 -0700 Andrew Morton <akpm@xxxxxxxxxxxxxxxxxxxx> wrote:
> On Wed, 2 Sep 2026 12:45:20 -0700 Joshua Hahn <joshua.hahnjy@xxxxxxxxx> wrote:
>
> > The test_no_invasive_cgroup_shrink selftest checks that when a cgroup
> > has zswapped out more memory than memory.zswap.max, it does not
> > trigger writeback for other cgroups. To do this, it compares the
> > writeback count in a control cgroup and makes sure that it is 0,
> > and then checks the writeback count in an aggressor cgroup who does
> > expect to see writeback.
> >
> > However, when the zswap shrinker is enabled, the victim cgroup can see
> > legitimate writebacks not triggered by the aggressor. In some Meta CI
> > tests, we have seen this failure mode happen.
> >
> > Instead of checking that the victim cgroup has 0 writeback, compare the
> > writeback values before and after the aggressor runs and check that
> > the victim cgroup did not perform any additional writeback. Note that
> > this can still lead to probabilistic failures if writebacks take longer
> > than 5 seconds, but this should fix the systematic failure case.
> >
> > Fixes: b5ba474f3f51 ("zswap: shrink zswap pool based on memory pressure")
Hello Andrew, I hope you're doing well!
> Thanks.
>
> I do support backporting selftests fixes. We want selftests to work
> well in the kernel with which they are shipped. But, as ever, we
> should include a clear explanation of the userspace-visible runtime
> effects of the bug.
>
> Seems the short answer is "false-positive test failures" but we can
> perhaps do a little better than that. Quoting the failure messages
> would be a good way - very recognizable.
That sounds good to me. Unfortunately it seems like the selftests
aren't too expressive in terms of why it failed, so this is all that
I was able to pick up from the selftest logs:
ok 1 test_zswap_usage
ok 2 test_swapin_nozswap
ok 3 test_zswapin
ok 4 test_zswap_writeback_enabled
ok 5 test_zswap_writeback_disabled
ok 6 # SKIP test_no_kmem_bypass
not ok 7 test_no_invasive_cgroup_shrink
ok 8 test_zswap_incompressible
# 1 skipped test(s) detected. Consider enabling relevant config options to improve coverage.
# Totals: pass:6 fail:1 xfail:0 xpass:0 skip:1 error:0
I'm not too sure how helpful this will be, but maybe it would be better
to amend the last line of the commit message to say:
...
than 5 seconds, but this should fix the systematic failure case
and make "not ok test_no_invasive_cgroup_shrink" less likely.
> > Reported-by: Krush Chavan <krushchavan@xxxxxxxxxxx>
>
> Relatedly, is there a Link:/Closes:?
Unfortunately this was detected in our private CI workflow and raised
by Krush, so I'm not sure if there is anything we can share.
If there is anything else that I can add to make this clearer, please
let me know. I hope you have a great day Andrew!
Joshua