Re: [PATCH 0/2] mm: zswap: reduce request contention on loads
From: Nhat Pham
Date: Tue Oct 06 2026 - 05:36:12 EST
On Tue, Oct 6, 2026 at 2:23 AM Usama Arif <usama.arif@xxxxxxxxx> wrote:
>
> Stores and loads share a per-CPU acomp request and mutex. A low-priority
> store can be preempted right after the compressor drops its stream
> lock, while it still holds the zswap mutex, and a higher-priority load
> on that CPU then waits for the store to run again. This follows the work
> from Sergey Senozhatsky's zram series which splits it for the same
> reason [1].
>
> Patch 1 gives compression and decompression separate requests, waits
> and mutexes, so loads no longer wait for stores, though they can still
> wait for each other. Patch 2 decompresses with an on-stack request when
> the algorithm is synchronous and needs no request context, which covers
> all in-tree software compressors, so those loads take no zswap lock.
> Asynchronous algorithms keep the per-CPU request and mutex. For software
> compressors the series allocates the same number of requests as before;
> each per-CPU context grows by 72 bytes, and the load path is about 270
> bytes deeper on x86-64.
>
> The series does not fix two related cases:
> - Stores still serialize on the compression mutex, so a high-priority
> task that reclaims (direct reclaim, MADV_PAGEOUT) can still wait for
> a preempted store.
Any reasons why we cannot tackle this too? Or just one at a time?
> - On PREEMPT_RT the codec stream locks are preemptible, so a load can
> still wait for a preempted store inside the codec.
Acked.
>
> The numbers below are the slowest read per run, as a median (min-max)
> of 5 runs. Each run is 12 seconds in a zstd VM with lazy preemption,
> vm.page-cluster=0 and swap on /dev/ram0. With 1 vCPU, four nice +10
> workers page memory out and read it back while a nice 0 task spins. A
> nice -19 reader pages out its own buffer and measures how long each
> read of it takes. With 8 vCPUs there are 16 workers, 8 spinning tasks
> and 8 readers.
>
> Before series (ms) With series (ms)
> 1 vCPU 22.3 (21.6-22.6) 0.97 (0.72-1.4)
> 8 vCPUs 314 (97-2542) 7.0 (5.0-98)
>
> Reads over 10 ms fell from 26-35 per run to none with 1 vCPU, and from
> 3-18 per run to at most one with 8 vCPUs. The benchmark and test programs
> were written with the help of an LLM.
Great find, Usama!