[PATCH 2/5] PM: runtime: Avoid racy clock checks for autosuspend-retry
From: Brian Norris
Date: Tue Sep 29 2026 - 14:48:22 EST
The runtime PM docs clearly tell how a .runtime_suspend() implementation
can update the last-busy time (pm_runtime_mark_last_busy()) and return a
transient error code (-EBUSY or -EAGAIN) in order to rearm the
autosuspend timer. However, this check is racy, such that it will not
always rearm the timer properly.
Consider:
1. runtime_suspend() updates last_busy
2. long delay (scheduling, etc.)
3. runtime_suspend() returns -EBUSY
4. pm_runtime_autosuspend_expiration() looks up the current time
(ktime_get_mono_fast_ns())
5. the time from #4 suggests the "expiry" time (based on #1) is already
in the past
6. rpm_suspend() quits without either retry or rearm
This violates the documented expectations.
The primary problem is the racy lookup of the current time in #4. We can
avoid that by:
* Checking whether the expiration time changed before/after the suspend
attempt, *without* consulting the current clock for past-expiry
* Improving pm_runtime_mark_last_busy() so that it guards against
producing the same expiry (e.g., consider a system with a
low-resolution clock)
Noticed by inspection, and proved out with KUnit tests which follow this
patch.
I see this pattern is utilized in at least a few drivers:
* blk_pre_runtime_suspend() (block/blk-pm.c)
* serial_port_runtime_suspend() (drivers/tty/serial/serial_port.c)
* omap4_keypad_runtime_suspend() (drivers/input/keyboard/omap4-keypad.c)
* omap_rproc_runtime_suspend() (drivers/remoteproc/omap_remoteproc.c)
* more...
Signed-off-by: Brian Norris <briannorris@xxxxxxxxxxxx>
---
There's a vanishingly small corner case in here still, if:
1) the monotonic clock gets an update while we're suspending
2) there are *two* mark_last_busy() calls
3) the result of those updates (backward in time) and mark_last_busy()
(forward in time) puts us back at the same expiration.
We could potentially solve this by updating a |last_busy_counter|, and
tracking that instead. I'm not sure if it's worth the complexity though.
There's a similar corner case if somebody calls
pm_runtime_set_autosuspend_delay(), such that a shortened autosuspend
delay lands on the same expiry. This also seems vanishingly unlikely;
but this also could be solved by using a |last_busy_counter|, and
ignoring the delay entirely.
drivers/base/power/runtime.c | 43 +++++++++++++++++++++++++-----------
include/linux/pm_runtime.h | 7 +++++-
2 files changed, 36 insertions(+), 14 deletions(-)
diff --git a/drivers/base/power/runtime.c b/drivers/base/power/runtime.c
index bc5905ffc45f..2a99b7509744 100644
--- a/drivers/base/power/runtime.c
+++ b/drivers/base/power/runtime.c
@@ -162,6 +162,26 @@ static void pm_runtime_cancel_pending(struct device *dev)
dev->power.request = RPM_REQ_NONE;
}
+/*
+ * Expiration time without considering whether it has already passed.
+ */
+static u64 __pm_runtime_autosuspend_expiration(struct device *dev)
+{
+ int autosuspend_delay;
+ u64 expires;
+
+ if (!dev->power.use_autosuspend)
+ return 0;
+
+ autosuspend_delay = READ_ONCE(dev->power.autosuspend_delay);
+ if (autosuspend_delay < 0)
+ return 0;
+
+ expires = READ_ONCE(dev->power.last_busy);
+ expires += (u64)autosuspend_delay * NSEC_PER_MSEC;
+ return expires;
+}
+
/**
* pm_runtime_autosuspend_expiration - Get a device's autosuspend-delay expiration time.
* @dev: Device to handle.
@@ -176,18 +196,10 @@ static void pm_runtime_cancel_pending(struct device *dev)
*/
u64 pm_runtime_autosuspend_expiration(struct device *dev)
{
- int autosuspend_delay;
- u64 expires;
+ u64 expires = __pm_runtime_autosuspend_expiration(dev);
- if (!dev->power.use_autosuspend)
- return 0;
-
- autosuspend_delay = READ_ONCE(dev->power.autosuspend_delay);
- if (autosuspend_delay < 0)
+ if (!expires)
return 0;
-
- expires = READ_ONCE(dev->power.last_busy);
- expires += (u64)autosuspend_delay * NSEC_PER_MSEC;
if (expires > ktime_get_mono_fast_ns())
return expires; /* Expires in the future */
@@ -586,6 +598,7 @@ static int rpm_suspend(struct device *dev, int rpmflags)
{
int (*callback)(struct device *);
struct device *parent = NULL;
+ u64 old_expires;
int retval;
trace_rpm_suspend(dev, rpmflags);
@@ -689,6 +702,7 @@ static int rpm_suspend(struct device *dev, int rpmflags)
}
__update_runtime_status(dev, RPM_SUSPENDING);
+ old_expires = __pm_runtime_autosuspend_expiration(dev);
callback = RPM_GET_CALLBACK(dev, runtime_suspend);
@@ -760,9 +774,12 @@ static int rpm_suspend(struct device *dev, int rpmflags)
* autosuspend expiration time, automatically reschedule another
* autosuspend.
*/
- if (!dev->power.runtime_error && (rpmflags & RPM_AUTO) &&
- pm_runtime_autosuspend_expiration(dev) != 0)
- goto repeat;
+ if (!dev->power.runtime_error && (rpmflags & RPM_AUTO)) {
+ u64 new_expires = __pm_runtime_autosuspend_expiration(dev);
+
+ if (new_expires && new_expires != old_expires)
+ goto repeat;
+ }
goto out;
}
diff --git a/include/linux/pm_runtime.h b/include/linux/pm_runtime.h
index 322e3b17f987..424ac9da94ef 100644
--- a/include/linux/pm_runtime.h
+++ b/include/linux/pm_runtime.h
@@ -242,7 +242,12 @@ static inline bool pm_runtime_has_no_callbacks(struct device *dev)
*/
static inline void pm_runtime_mark_last_busy(struct device *dev)
{
- WRITE_ONCE(dev->power.last_busy, ktime_get_mono_fast_ns());
+ u64 now = ktime_get_mono_fast_ns();
+
+ if (now == READ_ONCE(dev->power.last_busy))
+ now++;
+
+ WRITE_ONCE(dev->power.last_busy, now);
}
/**
--
2.56.0.rc1.315.gc6ed9934b7-goog