[PATCH] ceph: bound the mon command wait

From: Yuanfu Xie

Date: Mon Sep 28 2026 - 07:44:55 EST


ceph_update_snap_trace() reacts to a corrupted snap trace by fencing
the mount and blocklisting the client through ceph_monc_blocklist_add().
That issues the "osd blocklist add" mon command and waits for the reply
synchronously, from the messenger dispatch path, with session->s_mutex
held. Dispatch runs in a kworker, where no signal will ever arrive,
and the wait has no timeout: a monitor that doesn't reply leaves that
worker asleep holding the mutex for good. ceph_cap_release_work() and
anything else queued behind the session mutex pile up on it, and the
WARN(1, "...do remount to continue") that is supposed to tell the admin
to remount sits after the wait and never runs.

Bound the mon command wait with CEPH_MONC_PING_TIMEOUT (30s). That is
already how long an unanswered monitor keepalive may last before the
session is reopened. On timeout the existing error handling takes
over: the blocklist failure is reported, the warning is issued and the
mount stays fenced, so the admin can remount as intended.

The blocked worker on the unpatched kernel, trimmed:

INFO: task kworker/0:1:11 blocked for more than 10 seconds.
Workqueue: ceph-msgr ceph_con_workfn
Call Trace:
<TASK>
__schedule+0x1850/0x4a80
schedule+0x7a/0x1f0
schedule_timeout+0x217/0x250
__wait_for_common+0x2af/0x460
wait_for_completion_interruptible+0x1f/0x40
wait_generic_request+0x1f/0x270
do_mon_command+0x4b6/0x600
ceph_monc_blocklist_add+0x88/0x1a0
ceph_update_snap_trace+0x222/0x3420
ceph_con_process_message+0x1e5/0x270
[...]
</TASK>

INFO: task kworker/0:0:9 blocked for more than 10 seconds.
Workqueue: ceph-cap ceph_cap_release_work
mutex_lock+0xce/0xe0
[...]

With the watchdog set to panic this ends in "Kernel panic - not
syncing: hung_task: blocked tasks"; with the production defaults it
just stays there. With this patch the same run prints:

ceph: [client 1]: error -5
ceph: [client 1]: failed to blocklist ...: -110
[client.1] ceph_update_snap_trace do remount to continue after corrupted snaptrace is fixed
WARNING: fs/ceph/snap.c:926 at ceph_update_snap_trace+0x2d1/0x3420

(thirty seconds after the snap trace error) and the pending requests
on the fenced mount fail instead of hanging.

Verified on 7.3-rc4-00537-ga3ff15db6820 with only this change applied:
the blocklist failure and the warning arrive thirty seconds after the
snap trace error, and the blocked worker returns. An unresponsive
monitor plus a corrupted snap trace used to wedge the mount for good;
when the wait expires the blocklist attempt returns -ETIMEDOUT, the
warning fires and the mount stays fenced until remount.

Fixes: a68e564adcaa6 ("ceph: blocklist the kclient when receiving corrupted snap trace")
Signed-off-by: Yuanfu Xie <yuanfuxie@xxxxxxxxxxxxxx>
---
net/ceph/mon_client.c | 27 ++++++++++++++++++++++++++-
1 file changed, 26 insertions(+), 1 deletion(-)

diff --git a/net/ceph/mon_client.c b/net/ceph/mon_client.c
index c56457378d00e..435c7f851ab26 100644
--- a/net/ceph/mon_client.c
+++ b/net/ceph/mon_client.c
@@ -705,6 +705,24 @@ static int wait_generic_request(struct ceph_mon_generic_request *req)
return ret;
}

+static int wait_generic_request_timeout(struct ceph_mon_generic_request *req,
+ unsigned long timeout)
+{
+ long ret;
+
+ dout("%s greq %p tid %llu\n", __func__, req, req->tid);
+ ret = wait_for_completion_interruptible_timeout(&req->completion,
+ timeout);
+ if (ret == 0)
+ ret = -ETIMEDOUT;
+ if (ret < 0) {
+ cancel_generic_request(req);
+ return ret;
+ }
+
+ return req->result; /* completed */
+}
+
static struct ceph_msg *get_generic_reply(struct ceph_connection *con,
struct ceph_msg_header *hdr,
int *skip)
@@ -1006,7 +1024,14 @@ int do_mon_command_vargs(struct ceph_mon_client *monc, const char *fmt,
send_generic_request(monc, req);
mutex_unlock(&monc->mutex);

- ret = wait_generic_request(req);
+ /*
+ * ceph_monc_blocklist_add() runs from messenger dispatch, in a
+ * kworker, with session->s_mutex held. No signal arrives there.
+ * CEPH_MONC_PING_TIMEOUT is how long an unanswered monitor
+ * keepalive may last before the session is reopened; use that
+ * so this worker cannot sleep for good.
+ */
+ ret = wait_generic_request_timeout(req, CEPH_MONC_PING_TIMEOUT);
out:
put_generic_request(req);
return ret;
--
2.43.0