Re: [PATCH] perf bench: Add atomic CAS benchmark

From: David Laight

Date: Thu Oct 08 2026 - 03:58:58 EST


On Wed, 30 Sep 2026 17:16:17 +0800
Changbin Du <changbin.du@xxxxxxxxx> wrote:

> Add a new 'atomic' collection to perf bench for benchmarking
> compare-and-swap (CAS) atomic operations with multi-threaded
> contention testing.
>
> The benchmark tests __atomic_compare_exchange_n operations
> with configurable thread count and iteration count to measure
> atomic contention effects.
>
> Why this benchmark is needed:
> - CAS operations are fundamental to lock-free algorithms and data
> structures. Understanding their performance characteristics under
> contention is critical for designing high-performance concurrent
> applications.
> - The benchmark helps identify atomic operation latency and
> scalability issues across different thread counts, revealing
> contention patterns that are not visible in single-threaded tests.
> - Useful for evaluating atomic implementation quality on different
> architectures and for regression testing after changes to atomic
> primitives or memory ordering.
>
> Measurement methodology:
> - Each thread starts a private timer (clock_gettime CLOCK_MONOTONIC)
> after synchronizing on a pthread_barrier, ensuring all threads
> begin simultaneously.
> - Each thread performs a hot loop of atomic compare-and-swap on a
> shared u64 counter, incrementing from 0 to iterations.
> - The shared counter is cache-line aligned (64 bytes) to isolate
> contention to the target cache line and avoid false sharing.
> - The wall-clock time is measured as the max of all per-thread
> runtimes (the time for the slowest thread to finish).
> - The first repeat is excluded from statistics as a warmup phase
> to avoid cache-cold effects.
> Example usage:
> $ perf bench atomic cas --threads 2
> # Running 'atomic/cas' benchmark:
>
> Threads: 2, iterations/thread: 100000000, repeats: 10 (warmup: 1)
> Avg wall-clock time: 7365.480 msec (stddev 66.014 msec)
> Total ops: 200,000,000
> Throughput total: 27,153,697 ops/sec
> Per-thread times and throughput (last repeat):
> fastest: 7510.031 msec (13315525 ops/sec)
> slowest: 7581.678 msec (13189692 ops/sec)
> avg: 7545.854 msec (13252310 ops/sec)
>
> Output fields explained:
> - Threads: number of contending threads
> - iterations/thread: CAS operations each thread performs
> - repeats: number of test runs (first is warmup)
> - Avg wall-clock time: mean time for all threads to complete
> - stddev: standard deviation across repeats
> - Total ops: threads x iterations/thread
> - Throughput total: aggregate ops/sec across all threads
> - Per-thread times: fastest/slowest/avg thread completion time
> - Per-thread throughput: per-thread ops/sec (shows scheduling imbalance)
>

I think you need to default to one thread per cpu.
Also try to run the test for a fixed time period rather than a very
large count.
You should be able to see that some systems completely fail to make
progress under very heavy contention.
(This isn't one thread getting starved, none of them make progress.)

David