[PATCH v2 00/11] zram: redesign zcomp and rework backends

From: Sergey Senozhatsky

Date: Fri Oct 09 2026 - 03:13:28 EST


The series has a number of memory usage optimisations and fixes
that result in savings from double digits KB to triple digits
KB per-context/per-CPU, depending on backend and its configuration.
More details are in the corresponding commit messages.

The zcomp redesign and backend rework happen in the final three
patches. zcomp part, basically, attempts to address the priority
inversion problem, that is seen on some setups. In zcomp, we have
always kept Read and Write contexts together in per-CPU streams.
With preemptible zram this becomes a bit of a problem, as a higher
priority Reader can preempt a Writer, which may hold a stream
lock, but since the compression stream is locked by the preempted
Writer, the Reader gets blocked on the very same stream lock.
However, we don't have any reasons to have Read and Write contexts
share and compete for the stream, because Read (decompression) and
Write (compression) paths are completely independent, they never
touch each other's buffers. Thus, we can split compression streams
into per-CPU Read streams and per-CPU Write streams, so that Readers
and Writers don't block each other anymore. Kudos to Barry Song
for the idea. Decoupling Reads and Writes results in noticeable
performance improvements on synthetic tests.

"use a singleton compression context for recompression" patch
builds atop of decoupled Read and Write streams and reworks the
way secondary streams are handled. Secondary compression streams
are only used from recompression, which is serialized by device
lock, IOW it's single-threaded. However, we still allocated
secondary streams per-CPU (we needed to permit concurrent Reads
of recompressed objects). Because we now have dedicated R/W
streams, we can allocate a singleton compression context (for
recompression) yet still have per-CPU decompression contexts.
This saves a notable amount of memory with some backends and
configurations (e.g. zstd, deflate, lz4hc).

The final patch in the series pushes performance gains even
further by switching to rw-semaphores for Read streams. This
allows multiple Readers to concurrently decompress objects (with
some caveats).

v1 -> v2:
-- Switched to rw-sem for Read streams (Brian)

Sergey Senozhatsky (11):
zram: prefix backends printk-s
zram: remove debugfs entry on init failure
zram: zstd: do not allocate empty C/D-dictionaries
zram: lz4hc: pre-initialize dictionary stream in setup_params
zram: lz4: use LZ4_decompress_safe_usingDict()
zram: lz4hc: use LZ4_decompress_safe_usingDict()
zram: zstd: do not use kvzalloc() for zstd allocations
zram: reset writeback state in zram_reset_device()
zram: split zcomp into separate R/W streams
zram: use a singleton compression context for recompression
zram: use rw_semaphore for decompression streams

drivers/block/zram/backend_842.c | 6 +-
drivers/block/zram/backend_deflate.c | 98 ++++++++-----
drivers/block/zram/backend_lz4.c | 85 +++--------
drivers/block/zram/backend_lz4hc.c | 105 +++++--------
drivers/block/zram/backend_lzo.c | 6 +-
drivers/block/zram/backend_lzorle.c | 6 +-
drivers/block/zram/backend_zstd.c | 133 +++++++++++------
drivers/block/zram/zcomp.c | 212 ++++++++++++++++++++++-----
drivers/block/zram/zcomp.h | 43 ++++--
drivers/block/zram/zram_drv.c | 55 ++++---
10 files changed, 456 insertions(+), 293 deletions(-)

--
2.56.0.385.gd3acb90ef8-goog