Re: [PATCH v2 00/12] Convert barrier pairs to acquire/release for better performance

From: Jinjie Ruan

Date: Mon Aug 31 2026 - 23:16:12 EST




在 2026/9/1 11:06, Kuniyuki Iwashima 写道:
> On Mon, Aug 31, 2026 at 7:42 PM Jinjie Ruan <ruanjinjie@xxxxxxxxxx> wrote:
>>
>> Hi,
>>
>> This series converts some existing smp_wmb()/smp_rmb() barrier pairs to
>> smp_store_release()/smp_load_acquire() across various subsystems.
>>
>> Background
>> ==========
>>
>> Many architectures support load acquire and store release instructions
>> which can replace explicit memory barriers and save cycles. As noted
>> in the ARM architecture reference [1]:
>>
>> "Weaker ordering requirements that are imposed by Load-Acquire and
>> Store-Release instructions allow for micro-architectural
>> optimizations, which could reduce some of the performance impacts
>> that are otherwise imposed by an explicit memory barrier.
>>
>> If the ordering requirement is satisfied using either a Load-Acquire
>> or Store-Release, then it would be preferable to use these
>> instructions instead of a DMB."
>>
>> On arm64, a typical seqcount [2] read loop requires 13 cycles with DMB
>> barriers. Replacing the read barrier with smp_load_acquire() reduces
>> this to 8 cycles on an Ampere Altra.
>>
>> We also observed significant barrier overhead while profiling Unxibench
>> syscall test on arm64: a single getuid() call is ~8ns slower than on
>> a comparable x86 system, with the dominant cost in map_id_up()'s smp_rmb(),
>> which is a DMB ISHLD on arm64. Converting it to smp_load_acquire() allows
>> the use of LDAR, eliminating the measurable overhead.
>>
>> This motivated a broader search for existing barrier pairs that can
>> be converted to the lighter acquire/release semantics.
>>
>> Changes
>> =======
>>
>> Each patch in this series targets a specific barrier pair where the
>> publish/subscribe pattern is already present:
>>
>> - Writers populate data, then publish a flag/count/pointer via
>> smp_store_release()
>>
>> - Readers load the flag/count/pointer via smp_load_acquire(), then
>> consume the data
>>
>> This preserves the existing memory ordering guarantees while allowing
>> architectures with native acquire/release instructions (e.g. arm64's
>> STLR/LDAR) to avoid the cost of full one-way barriers (DMB ISHST/ISHLD).
>> On architectures without native support, the generated code is
>> generally no worse than the explicit barrier pair.
>>
>> The conversions are mechanical and no functional change is intended.
>>
>> Testing (arm64 Kunpeng HIP09 server)
>> ================
>>
>> 1. UNIXBENCH syscall
>> Baseline: 715.27
>> Patched: 718.83
>> Improvement: +0.50%
>>
>> 2. fs/aio (fio + null_blk, 4 jobs):
>> Baseline: 1441k IOPS, 86.46us
>> Patched: 1452k IOPS, 85.80us
>> Improvement: ~0.8%
>>
>> 3. soreuseport (wrk, 8 servers):
>> Baseline: 162.6k req/s, 452.5us
>> Patched: 164.2k req/s, 449.4us
>> Improvement: ~1.0%
>>
>> Both improvements are consistent across runs and align with the
>> expected savings from replacing DMB with LDAR/STLR on arm64.
>>
>> [1]: https://support.arm.com/documentation/102336/0100/Load-Acquire-and-Store-Release-instructions
>> [2]: https://github.com/torvalds/linux/commit/d0dd066a0fa26d55c19ace9e89dedd9504c5bcba
>>
>> Changes in v2:
>> - Fix pre-existing issue for ext4 and 8021q [3].
>> - Fix missing copy_mnt_idmap() udapte [3].
>> - Drop nacked isotp patch.
>> - Add test data.
>> - Add Reviewed-by and update fs patch as Jan suggested.
>>
>> [3]: https://sashiko.dev/#/patchset/20260825095422.3166067-1-ruanjinjie%40huawei.com
>>
>> Jinjie Ruan (12):
>> user_namespace: Use acquire/release for nr_extents synchronization
>> lib/vsprintf: Use acquire/release for ptr_key publication
>> fs: aio: Use acquire/release for ring->tail publication
>> fs: Use acquire/release for fdtable resize synchronization
>> pidfs: Use test_bit_acquire() for attr flag tests
>> super: Use acquire for SB_BORN check in super_cache_count()
>> ext4: Fix out-of-bounds read in ext4_get_group_info()
>> ext4: Convert group-count barrier protocol to acquire/release
>> soreuseport: publish num_socks with acquire/release
>> net: sched: act_gact: use acquire/release for tcfg_ptype
>> 8021q: Fix data race when publishing vlan net_device pointers
>> 8021q: publish vlan_devices_arrays entries with acquire/release
>
> Please post networking patches separately with the target tree specified:
>
> Subject: [PATCH vX net-next] soreuseport: ...
>
> 8021q changes can be posted a series.

Thanks for the review. I will split the series as suggested — the
networking patches will be posted separately.