Re: [PATCH 0/2] mm: zswap: reduce request contention on loads
From: Usama Arif
Date: Tue Oct 06 2026 - 05:19:38 EST
On 06/10/2026 01:22, Usama Arif wrote:
> Stores and loads share a per-CPU acomp request and mutex. A low-priority
> store can be preempted right after the compressor drops its stream
> lock, while it still holds the zswap mutex, and a higher-priority load
> on that CPU then waits for the store to run again. This follows the work
> from Sergey Senozhatsky's zram series which splits it for the same
> reason [1].
>
> Patch 1 gives compression and decompression separate requests, waits
> and mutexes, so loads no longer wait for stores, though they can still
> wait for each other. Patch 2 decompresses with an on-stack request when
> the algorithm is synchronous and needs no request context, which covers
> all in-tree software compressors, so those loads take no zswap lock.
> Asynchronous algorithms keep the per-CPU request and mutex. For software
> compressors the series allocates the same number of requests as before;
> each per-CPU context grows by 72 bytes, and the load path is about 270
> bytes deeper on x86-64.
>
> The series does not fix two related cases:
> - Stores still serialize on the compression mutex, so a high-priority
> task that reclaims (direct reclaim, MADV_PAGEOUT) can still wait for
> a preempted store.
> - On PREEMPT_RT the codec stream locks are preemptible, so a load can
> still wait for a preempted store inside the codec.
>
> The numbers below are the slowest read per run, as a median (min-max)
> of 5 runs. Each run is 12 seconds in a zstd VM with lazy preemption,
> vm.page-cluster=0 and swap on /dev/ram0. With 1 vCPU, four nice +10
> workers page memory out and read it back while a nice 0 task spins. A
> nice -19 reader pages out its own buffer and measures how long each
> read of it takes. With 8 vCPUs there are 16 workers, 8 spinning tasks
> and 8 readers.
>
> Before series (ms) With series (ms)
> 1 vCPU 22.3 (21.6-22.6) 0.97 (0.72-1.4)
> 8 vCPUs 314 (97-2542) 7.0 (5.0-98)
>
> Reads over 10 ms fell from 26-35 per run to none with 1 vCPU, and from
> 3-18 per run to at most one with 8 vCPUs. The benchmark and test programs
> were written with the help of an LLM.
>
In Meta fleet, looking at lock profiler in the last day, the longest observed
mutex hold was 137.6 ms, including 137.5 ms during which the holder was runnable
but off-CPU.