Re: [PATCH 11/11] can: isotp: publish tx.state with smp_store_release()
From: Oliver Hartkopp
Date: Wed Aug 26 2026 - 04:10:21 EST
On 26.08.26 05:33, Jinjie Ruan wrote:
在 2026/8/25 19:46, Oliver Hartkopp 写道:
On 25.08.26 11:54, Jinjie Ruan wrote:
The writer already pairs with the smp_load_acquire() readers
in isotp_tx_timeout()/isotp_tx_gen_done(); convert
the smp_wmb() + WRITE_ONCE() into a release store.
Assisted-by: DeepSeek:DeepSeek-V3
Signed-off-by: Jinjie Ruan <ruanjinjie@xxxxxxxxxx>
---
net/can/isotp.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
diff --git a/net/can/isotp.c b/net/can/isotp.c
index 155530aedce2..11b653ba7c10 100644
--- a/net/can/isotp.c
+++ b/net/can/isotp.c
@@ -1156,8 +1156,8 @@ static int isotp_sendmsg(struct socket *sock,
struct msghdr *msg, size_t size)
my_gen = isotp_inc_tx_gen(READ_ONCE(so->tx_gen));
isotp_set_tx_result(so, my_gen, ECOMM); /* prevent stale slot
matching */
WRITE_ONCE(so->tx_gen, my_gen);
- smp_wmb(); /* see smp_load_acquire() in isotp_tx_[timeout|
gen_done] */
- WRITE_ONCE(so->tx.state, ISOTP_SENDING);
+ /* Pairs with smp_load_acquire() in isotp_tx_[timeout|gen_done] */
+ smp_store_release(&so->tx.state, ISOTP_SENDING);
WRITE_ONCE(so->cfecho, 0);
spin_unlock_bh(&so->rx_lock);
Hi Jinjie,
Hi Oliver,
thank you for the patch, but I think this breaks the barrier logic.
The original smp_wmb() ensures that so->tx_gen is visible before both
subsequent writes (so->tx.state and so->cfecho).
Right!
By converting only the first write into smp_store_release(), the
WRITE_ONCE(so->cfecho, 0) is no longer protected. The compiler or CPU
could reorder and execute the cfecho write before the release store of
Indeed, that's true.
so->tx.state, introducing a race condition with the concurrent readers.
My rough understanding is as follows:
All lock-free readers fall into two disjoint sets:
- `isotp_tx_timeout()` and `isotp_tx_gen_done()`: read only `tx.state`
(acquire) and `tx_gen`
- `isotp_txfr_timer_handler()` the timer path of `isotp_send_cframe()`,
and the post-claim path of `isotp_sendmsg()` touch `cfecho` but never
`tx_gen`.
So no lock-free reader observes both `tx_gen` and `cfecho`.
`isotp_rcv_echo()` is the only function reading both, and it runs under
`so->rx_lock`, which serializes it with the claim.
So the ordering the `smp_wmb()` provided on top of the new release store
`tx_gen` before `cfecho` — is unobservable to every reader.
Moreover, `tx.state` and `cfecho` were never ordered against each other
by the original barrier: both followed the `smp_wmb()`, so the `(state,
cfecho)` visibility seen by the lock-free timer readers is bit-for-bit
identical before and after this change.
So the release store preserves the one ordering that matters: a reader
observing `ISOTP_SENDING` sees the new `tx_gen`.
Best regards,
Jinjie
Thanks for your explanation, but this assumption is too narrow and misses how so->cfecho interacts with the rest of the ISO-TP state machine, especially under concurrent TX and RX traffic.
Even if isotp_tx_timeout() and isotp_tx_gen_done() don't read cfecho, other functions do. For example, cfecho is heavily involved in the RX path—such as in isotp_rcv_cf()—where incoming Consecutive Frames are validated against the current transmission state.
By converting the code to smp_store_release(), you only guarantee that the write to so->tx_gen happens before so->tx.state. However, you completely lose the ordering guarantee for WRITE_ONCE(so->cfecho, 0). The compiler or CPU is now free to reorder and execute the cfecho clear before the smp_store_release().
If a concurrent CAN frame arrives and triggers the RX path exactly at this microsecond, it could see an updated, cleared cfecho value while tx.state is still in its old state, or vice versa. This breaks the atomicity of the state transition in isotp_sendmsg().
The original smp_wmb() acts as a clear fence: it ensures that both subsequent writes (tx.state and cfecho) become visible to all concurrent readers strictly after the new tx_gen generation is visible.
We cannot weaken this guarantee. It makes the lockless design extremely fragile and prone to hard-to-debug race conditions.
Therefore:
Nacked-by: Oliver Hartkopp socketcan@xxxxxxxxxxxx
Best regards,
Oliver