Re: [RFC PATCH 0/2] mm: zsmalloc: make shrinker compaction budget-aware
From: Xueyuan Chen
Date: Tue Aug 11 2026 - 05:06:56 EST
On 8/8/2026 12:25 AM, Nhat Pham wrote:
On Fri, Aug 7, 2026 at 5:12 AM Sergey Senozhatsky
<senozhatsky@xxxxxxxxxxxx> wrote:
On (26/08/07 18:57), Xueyuan Chen wrote:Hmm we'd need to collect more data in our workloads to determine, but
Hi Sergey,Let's bring in heavy artillery to this discussion, in addition to Andrew
Here is some additional data:
I used the following definitions:
compactable ratio = freeable_pages / total_pages
memory reclaimed = pages_freed * PAGE_SIZE
freeable_pages is the estimate before compaction, based on the same
calculation as zs_shrinker_count(), while pages_freed is the actual
number of backing pages released.
There were 264 callbacks in the trace:
callback elapsed time:
median: 5.77 ms
p95: 55.82 ms
maximum: 271.36 ms
compactable ratio before compaction:
median: 0.32%
p95: 2.86%
maximum: 8.33%
memory reclaimed per callback:
median: 3.80 MiB
p95: 30.45 MiB
maximum: 92.73 MiB
The longest callback took 271.36 ms. Its compactable ratio was 3.25%,
and it released 7,650 pages, or about 29.88 MiB.
There was also a 241.91 ms callback (with 30 schedule-outs) with a
compactable ratio of 0.44%. It released 1,019 pages, or about
3.98 MiB.
Based on this data, it seems better to remove the shrinker.
Would you prefer that I change v2 to remove the zsmalloc shrinker
callbacks directly?
and Minchan, adding Nhat, Yosry, Barry, Johannes, Brian (random order).
Folks, I'm bullish on removal of zsmalloc shrinker callbacks.
I don't think those buy us much apart from memcpy-s and lock
contention. Systems that want to compact zsmalloc have a sysfs
knob (and API) to do so (based on zram mm_stat numbers).
if we can compute compactable ratio (freeable_pages / total_pages) on
a per size class basis, can we just skip the size class whose ratio is
too bad? Would that at least cut down on the vast majority of
fruitless memcpys and lock acquisitions etc?
I'm always a bit hesitant to over-rely on userspace, especially when
kernel has information to do something smart about it. It might not
react in time, and many proactive reclaiming schemes back off under
heavy memory pressure.
Hi Nhat,
I collected per size-class data. 53.7% of size classes have freeable==0,
so skipping them saves some time. However, 94.7% of zs_compact() time is
spent on classes where freeable > 0 — meaning the cost is dominated by the actual compaction work, not by scanning empty classes.
That said, skipping freeable == 0 classes is still a worthwhile
optimization on its own — it avoids unnecessary write_lock,spin_lock.
But it would not meaningfully reduce the worst-case latency.
Thanks,
Xueyuan
Any thoughts?