Re: [PATCH 2/5] PM: runtime: Avoid racy clock checks for autosuspend-retry

From: Brian Norris

Date: Mon Oct 05 2026 - 15:01:02 EST


On Fri, Oct 02, 2026 at 05:27:23PM -0700, Doug Anderson wrote:
> On Tue, Sep 29, 2026 at 11:47 AM Brian Norris <briannorris@xxxxxxxxxxxx> wrote:
> >
> > @@ -242,7 +242,12 @@ static inline bool pm_runtime_has_no_callbacks(struct device *dev)
> > */
> > static inline void pm_runtime_mark_last_busy(struct device *dev)
> > {
> > - WRITE_ONCE(dev->power.last_busy, ktime_get_mono_fast_ns());
> > + u64 now = ktime_get_mono_fast_ns();
> > +
> > + if (now == READ_ONCE(dev->power.last_busy))
> > + now++;
> > +
> > + WRITE_ONCE(dev->power.last_busy, now);
>
> While we definitely want to fix the problem identified in this patch,
> the proposed logic doesn't sit right with me. Let's say that
> 'power.last_busy" starts out as a given value, let's say "1020". Now,
> while the clock hasn't ticked you call pm_runtime_mark_last_busy(). It
> detects that 1020 == 1020 so it sets the time to 1021 so it's
> different. Now you call pm_runtime_mark_last_busy() again when the
> clock hasn't ticked. Now 1020 != 1021, so it goes back to 1020. This
> could cause the whole heuristic to fail, can't it?
>
> Maybe I'm just worrying about something that can't happen, but the
> logic still seems odd.

Thanks for pointing this all out. The above solution was trying to be
too clever for its own good, and left a different set of holes. It's
definitely odd, and I think it points toward:

1) either we somehow have to be even more clever or

2) we really need an additional state variable.

I don't think #1 is a good idea.

And #2 was what I was pointing toward below the "---" fold, mentioning a
|last_busy_counter|. I thought we could avoid it, but I believe I'm
wrong.

> It felt to me like we could just have a "bool". We set it to false
> before we call runtime_suspend() and we check it after
> runtime_suspend() returns. If the bool is set then we know they called
> pm_runtime_mark_last_busy().

Yes, I think it doesn't even need to be a counter -- just a bool is
probably fine.

> There shouldn't even be any weird
> problems with weakly ordered memory since this boolean should always
> be cleared, set, and checked in the same thread, right?

Well, it *can* be set in other threads (there's intentionally no locking
on pm_runtime_mark_last_busy()), but as long as the setters we care
about are in-thread, I think the only problem we'd invite is reading a
false "last_busy was set", and performing a spurious retry.

But potentially-spurious retry is already baked into this protocol.

I'll let this sit for a bit in case somebody else has other
thoughts/suggestions, but otherwise, I'll probably end up adding a
"last_busy_updated" field for v2.

Brian