Re: [PATCH 0/2] mm: zswap: reduce request contention on loads

From: Usama Arif

Date: Wed Oct 07 2026 - 07:35:56 EST




On 06/10/2026 11:35, Nhat Pham wrote:
> On Tue, Oct 6, 2026 at 2:23 AM Usama Arif <usama.arif@xxxxxxxxx> wrote:
>>
>> Stores and loads share a per-CPU acomp request and mutex. A low-priority
>> store can be preempted right after the compressor drops its stream
>> lock, while it still holds the zswap mutex, and a higher-priority load
>> on that CPU then waits for the store to run again. This follows the work
>> from Sergey Senozhatsky's zram series which splits it for the same
>> reason [1].
>>
>> Patch 1 gives compression and decompression separate requests, waits
>> and mutexes, so loads no longer wait for stores, though they can still
>> wait for each other. Patch 2 decompresses with an on-stack request when
>> the algorithm is synchronous and needs no request context, which covers
>> all in-tree software compressors, so those loads take no zswap lock.
>> Asynchronous algorithms keep the per-CPU request and mutex. For software
>> compressors the series allocates the same number of requests as before;
>> each per-CPU context grows by 72 bytes, and the load path is about 270
>> bytes deeper on x86-64.
>>
>> The series does not fix two related cases:
>> - Stores still serialize on the compression mutex, so a high-priority
>> task that reclaims (direct reclaim, MADV_PAGEOUT) can still wait for
>> a preempted store.
>
> Any reasons why we cannot tackle this too? Or just one at a time?

Stores need more than a request. The compression mutex also protects
the per-CPU PAGE_SIZE output buffer, which has to stay ours until
zs_obj_write() copies it out, since zs_malloc() needs the compressed
length first. Loads stopped using that buffer in e2c3b6b21c77f, so an
on-stack request was enough for them, but the buffer is too big for
the stack.

>
>> - On PREEMPT_RT the codec stream locks are preemptible, so a load can
>> still wait for a preempted store inside the codec.
>
> Acked.
>
>>
>> The numbers below are the slowest read per run, as a median (min-max)
>> of 5 runs. Each run is 12 seconds in a zstd VM with lazy preemption,
>> vm.page-cluster=0 and swap on /dev/ram0. With 1 vCPU, four nice +10
>> workers page memory out and read it back while a nice 0 task spins. A
>> nice -19 reader pages out its own buffer and measures how long each
>> read of it takes. With 8 vCPUs there are 16 workers, 8 spinning tasks
>> and 8 readers.
>>
>> Before series (ms) With series (ms)
>> 1 vCPU 22.3 (21.6-22.6) 0.97 (0.72-1.4)
>> 8 vCPUs 314 (97-2542) 7.0 (5.0-98)
>>
>> Reads over 10 ms fell from 26-35 per run to none with 1 vCPU, and from
>> 3-18 per run to at most one with 8 vCPUs. The benchmark and test programs
>> were written with the help of an LLM.
>
> Great find, Usama!

Thanks for the reviews!