Re: [PATCH] timers/nohz: Annotate lockless accesses to got_idle_tick
From: Thomas Gleixner
Date: Wed Sep 30 2026 - 18:16:54 EST
On Wed, Sep 30 2026 at 13:00, Paul E. McKenney wrote:
> On Tue, Sep 29, 2026 at 02:21:32PM +0200, Thomas Gleixner wrote:
>> On Sun, Sep 20 2026 at 16:54, Kunwu Chan wrote:
>> > tick_sched_do_timer(), called from the tick interrupt handler, sets
>> > ts->got_idle_tick when a tick fires while the CPU is idle. The idle
>> > path reads and clears it via tick_nohz_idle_got_tick() to detect
>> > whether the tick handler has run.
>> >
>> > The flag is accessed locklessly from hardirq and task context. A
>> > concurrent set can be overwritten by the clear.
>> >
>> > Use READ_ONCE() and WRITE_ONCE() to annotate the lockless accesses.
>>
>> The following scenarios still happen:
>>
>> A)
>>
>> if (READ_ONCE(got_tick)) // reads true
>> -> interrupt
>> WRITE_ONCE(got_tick, true);
>>
>> WRITE_ONCE(got_tick, false);
>>
>> B)
>>
>> if (READ_ONCE(got_tick)) // reads false
>> -> interrupt
>> WRITE_ONCE(got_tick, true);
>>
>> #A is harmless because got_tick is already true and #B is harmless as
>> well because the got_tick read result is only a hint for the next idle
>> invocation.
>>
>> As this is stricly per CPU the READ/WRITE_ONCE() does not change
>> anything in terms of ordering. So what exactly is solved by the
>> READ/WRITE_ONCE()?
>
> This is a step towards allowing more-strict KCSAN checks to be enabled,
> namely data races between base code and interrupt handlers. KCSAN has
> found a few bugs of this type in RCU over the past year or so, which
> suggests that similar bugs might be lurking elsewhere.
>
> Yes, in this particular case, current compilers don't have a huge amount
> of freedom to mess things up, but the standard really does permit the
> compiler to use a to-be-stored-to location as a temporary just prior to
> that store.
>
> This is like the situation with lockdep, where we tell it about
> deadlock-free situations that it does not see as being deadlock-free.
>
> Seem reasonable?
Yes, but then the change log should clearly state exactly that along
with a corresponding comment in the code. These boilerplate copy&pasta
changes are annoying as hell because they fail to provide any reasonable
rationale for the proposed change. See Documentation/process/...
I'm refusing to accept this slop whether it's generated by humans or any
form of tools.
That said, I further have to say that just reusing READ/WRITE_ONCE() for
this is completely wrong.
A strict and harmless per CPU modification/race is very much different
from cross CPU races. So avoiding any potential confusion and thereby
allowing tools like KCSAN to treat them differently is useful.
Using the same annotation for something which is strictly a per CPU
situation prevents to detect the accidential cross CPU access which
might not be safe at all.
Even if the momentary implementation falls back to the same mechanism
for the tools the value of implied code documentation is valuable. It
avoids extensive comments and it also allows tools to differentiate the
meaning in the future.
READ/WRITE_ONCE() are just non-distinguishable "paper over all race
problems which tools might complain about" hammers.
But we all should know by now that "Birmingham screwdrivers" are more
than suboptimal.
Thanks,
tglx