[PATCH 00/11] Convert barrier pairs to acquire/release for better performance

From: Jinjie Ruan

Date: Tue Aug 25 2026 - 05:54:54 EST


Hi,

This series converts some existing smp_wmb()/smp_rmb() barrier pairs to
smp_store_release()/smp_load_acquire() across various subsystems.

Background
==========

Many architectures support load acquire and store release instructions
which can replace explicit memory barriers and save cycles. As noted
in the ARM architecture reference [1]:

"Weaker ordering requirements that are imposed by Load-Acquire and
Store-Release instructions allow for micro-architectural
optimizations, which could reduce some of the performance impacts
that are otherwise imposed by an explicit memory barrier.

If the ordering requirement is satisfied using either a Load-Acquire
or Store-Release, then it would be preferable to use these
instructions instead of a DMB."

On arm64, a typical seqcount [2] read loop requires 13 cycles with DMB
barriers. Replacing the read barrier with smp_load_acquire() reduces
this to 8 cycles on an Ampere Altra.

We also observed significant barrier overhead while profiling Unxibench
syscall test on arm64: a single getuid() call is ~8ns slower than on
a comparable x86 system, with the dominant cost in map_id_up()'s smp_rmb(),
which is a DMB ISHLD on arm64. Converting it to smp_load_acquire() allows
the use of LDAR, eliminating the measurable overhead.

This motivated a broader search for existing barrier pairs that can
be converted to the lighter acquire/release semantics.

Changes
=======

Each patch in this series targets a specific barrier pair where the
publish/subscribe pattern is already present:

- Writers populate data, then publish a flag/count/pointer via
smp_store_release()

- Readers load the flag/count/pointer via smp_load_acquire(), then
consume the data

This preserves the existing memory ordering guarantees while allowing
architectures with native acquire/release instructions (e.g. arm64's
STLR/LDAR) to avoid the cost of full one-way barriers (DMB ISHST/ISHLD).
On architectures without native support, the generated code is
generally no worse than the explicit barrier pair.

The conversions are mechanical and no functional change is intended.

[1]: https://support.arm.com/documentation/102336/0100/Load-Acquire-and-Store-Release-instructions
[2]: https://github.com/torvalds/linux/commit/d0dd066a0fa26d55c19ace9e89dedd9504c5bcba

Jinjie Ruan (11):
user_namespace: Use acquire/release for nr_extents synchronization
lib/vsprintf: Use acquire/release for ptr_key publication
fs: aio: Use acquire/release for ring->tail publication
fs: Use acquire/release for fdtable resize synchronization
pidfs: Use test_bit_acquire() for attr flag tests
super: Use acquire for SB_BORN check in super_cache_count()
ext4: Convert group-count barrier protocol to acquire/release
soreuseport: publish num_socks with acquire/release
net: sched: act_gact: use acquire/release for tcfg_ptype
8021q: publish vlan_devices_arrays entries with acquire/release
can: isotp: publish tx.state with smp_store_release()

fs/aio.c | 12 +++++-------
fs/ext4/ext4.h | 10 +++-------
fs/ext4/mballoc.c | 6 ++----
fs/ext4/resize.c | 19 +++++++++++--------
fs/file.c | 10 ++++------
fs/pidfs.c | 6 ++----
fs/super.c | 7 +++----
kernel/user_namespace.c | 24 +++++++++++++-----------
lib/vsprintf.c | 11 ++++-------
net/8021q/vlan.c | 6 ++----
net/8021q/vlan.h | 8 +++-----
net/can/isotp.c | 4 ++--
net/core/sock_reuseport.c | 20 ++++++++------------
net/sched/act_gact.c | 12 ++++--------
14 files changed, 66 insertions(+), 89 deletions(-)

--
2.34.1