Re: [PATCH v3] RDMA/srpt: Clamp the CQ size request to max_cqe

From: Leon Romanovsky

Date: Mon Sep 07 2026 - 02:33:14 EST


On Sun, Sep 06, 2026 at 07:12:35PM +0300, Serhat Kumral wrote:
> srpt_create_ch_ib() asks ib_cq_pool_get() for ch->rq_size + sq_size
> completion queue entries. rq_size is bounded by max_qp_wr, but sq_size
> comes from the per-port srp_sq_size configfs attribute, which is only
> validated against MAX_SRPT_SRQ_SIZE (65535) and never compared with
> dev->attrs.max_cqe. The largest configuration srpt accepts therefore
> asks for 128 + 65535 = 65663 entries, while rxe caps max_cqe at 32767
> and cxgb4 and ionic are in the same range.
>
> Such a request cannot be met, and ib_cq_pool_get() does not reject it.
> It clamps every CQ it creates to max_cqe, so no CQ it adds to the pool
> can ever fit the request, and it keeps allocating batches until the
> allocation fails. A single SRP login against such a target exhausts
> memory:
>
> Out of memory and no killable processes...
> Kernel panic - not syncing: System is deadlocked on memory
> Workqueue: ib_cm cm_work_handler
> Call Trace:
> __vmalloc_node_range_noprof
> vmalloc_user_noprof
> rxe_queue_init
> rxe_cq_from_init
> rxe_create_cq
> __ib_alloc_cq
> ib_cq_pool_get
> srpt_cm_req_recv.cold
>
> Before commit c804af2c1d31 ("IB/srpt: use new shared CQ mechanism") the
> same request went to ib_alloc_cq_any(), which rejected it with -EINVAL.
>
> The existing backoff, which halves sq_size when queue pair creation
> fails, only runs after ib_cq_pool_get() has returned, so shrink sq_size
> before asking for the CQ. With the clamp the same login proceeds exactly
> like a correctly sized target.
>
> Fixes: c804af2c1d31 ("IB/srpt: use new shared CQ mechanism")
> Signed-off-by: Serhat Kumral <serhatkumral1@xxxxxxxxx>
> ---
> Changes since v2:
> - Add a WARN_ON_ONCE().
>
> Changes since v1:
> - Move the fix from the RDMA core to ib_srpt, as requested by Leon
> Romanovsky.
> - Clamp sq_size before the CQ request instead of rejecting oversized
> requests in ib_cq_pool_get().
>
> v1: https://lore.kernel.org/linux-rdma/20260831171354.72140-1-serhatkumral1@xxxxxxxxx/
> v2: https://lore.kernel.org/linux-rdma/20260904200240.48976-1-serhatkumral1@xxxxxxxxx/
>
> drivers/infiniband/ulp/srpt/ib_srpt.c | 9 +++++++++
> 1 file changed, 9 insertions(+)
>
> diff --git a/drivers/infiniband/ulp/srpt/ib_srpt.c b/drivers/infiniband/ulp/srpt/ib_srpt.c
> index 7197d95f2216..7ac526e3d2c2 100644
> --- a/drivers/infiniband/ulp/srpt/ib_srpt.c
> +++ b/drivers/infiniband/ulp/srpt/ib_srpt.c
> @@ -1867,6 +1867,15 @@ static int srpt_create_ch_ib(struct srpt_rdma_ch *ch)
> if (!qp_init)
> goto out;
>
> + /* The send and receive queues share a single CQ. */
> + if (ch->rq_size + sq_size > attrs->max_cqe) {
> + /* Catch drivers that incorrectly set ch->rq_size */
> + WARN_ON_ONCE(ch->rq_size > attrs->max_cqe);
> + sq_size = attrs->max_cqe - ch->rq_size;
> + pr_debug("reduced sq_size to %u because max_cqe is %u\n",
> + sq_size, attrs->max_cqe);
> + }

All these Sashiko reports suggest that sq_size is being changed in the wrong
place. Can you check for the correct value in srpt_tpg_attrib_srp_sq_size_store()?

Bart, can we simply decrease MAX_SRPT_SRQ_SIZE by, say, 100?

Thanks

> +
> retry:
> ch->cq = ib_cq_pool_get(sdev->device, ch->rq_size + sq_size, -1,
> IB_POLL_WORKQUEUE);
>
> base-commit: cee9395acd8043be0644b25c34bfa86623f2b935
> --
> 2.53.0
>