mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v3] RDMA/srpt: Clamp the CQ size request to max_cqe
@ 2026-09-06 16:12 Serhat Kumral
  2026-09-07  6:32 ` Leon Romanovsky
  0 siblings, 1 reply; 4+ messages in thread
From: Serhat Kumral @ 2026-09-06 16:12 UTC (permalink / raw)
  To: Bart Van Assche, Jason Gunthorpe, Leon Romanovsky
  Cc: Sagi Grimberg, linux-rdma, target-devel, linux-kernel, Serhat Kumral

srpt_create_ch_ib() asks ib_cq_pool_get() for ch->rq_size + sq_size
completion queue entries. rq_size is bounded by max_qp_wr, but sq_size
comes from the per-port srp_sq_size configfs attribute, which is only
validated against MAX_SRPT_SRQ_SIZE (65535) and never compared with
dev->attrs.max_cqe. The largest configuration srpt accepts therefore
asks for 128 + 65535 = 65663 entries, while rxe caps max_cqe at 32767
and cxgb4 and ionic are in the same range.

Such a request cannot be met, and ib_cq_pool_get() does not reject it.
It clamps every CQ it creates to max_cqe, so no CQ it adds to the pool
can ever fit the request, and it keeps allocating batches until the
allocation fails. A single SRP login against such a target exhausts
memory:

  Out of memory and no killable processes...
  Kernel panic - not syncing: System is deadlocked on memory
  Workqueue: ib_cm cm_work_handler
  Call Trace:
   __vmalloc_node_range_noprof
   vmalloc_user_noprof
   rxe_queue_init
   rxe_cq_from_init
   rxe_create_cq
   __ib_alloc_cq
   ib_cq_pool_get
   srpt_cm_req_recv.cold

Before commit c804af2c1d31 ("IB/srpt: use new shared CQ mechanism") the
same request went to ib_alloc_cq_any(), which rejected it with -EINVAL.

The existing backoff, which halves sq_size when queue pair creation
fails, only runs after ib_cq_pool_get() has returned, so shrink sq_size
before asking for the CQ. With the clamp the same login proceeds exactly
like a correctly sized target.

Fixes: c804af2c1d31 ("IB/srpt: use new shared CQ mechanism")
Signed-off-by: Serhat Kumral <serhatkumral1@gmail.com>
---
Changes since v2:
- Add a WARN_ON_ONCE().

Changes since v1:
- Move the fix from the RDMA core to ib_srpt, as requested by Leon
  Romanovsky.
- Clamp sq_size before the CQ request instead of rejecting oversized
  requests in ib_cq_pool_get().

v1: https://lore.kernel.org/linux-rdma/20260831171354.72140-1-serhatkumral1@gmail.com/
v2: https://lore.kernel.org/linux-rdma/20260904200240.48976-1-serhatkumral1@gmail.com/

 drivers/infiniband/ulp/srpt/ib_srpt.c | 9 +++++++++
 1 file changed, 9 insertions(+)

diff --git a/drivers/infiniband/ulp/srpt/ib_srpt.c b/drivers/infiniband/ulp/srpt/ib_srpt.c
index 7197d95f2216..7ac526e3d2c2 100644
--- a/drivers/infiniband/ulp/srpt/ib_srpt.c
+++ b/drivers/infiniband/ulp/srpt/ib_srpt.c
@@ -1867,6 +1867,15 @@ static int srpt_create_ch_ib(struct srpt_rdma_ch *ch)
 	if (!qp_init)
 		goto out;
 
+	/* The send and receive queues share a single CQ. */
+	if (ch->rq_size + sq_size > attrs->max_cqe) {
+		/* Catch drivers that incorrectly set ch->rq_size */
+		WARN_ON_ONCE(ch->rq_size > attrs->max_cqe);
+		sq_size = attrs->max_cqe - ch->rq_size;
+		pr_debug("reduced sq_size to %u because max_cqe is %u\n",
+			 sq_size, attrs->max_cqe);
+	}
+
 retry:
 	ch->cq = ib_cq_pool_get(sdev->device, ch->rq_size + sq_size, -1,
 				 IB_POLL_WORKQUEUE);

base-commit: cee9395acd8043be0644b25c34bfa86623f2b935
-- 
2.53.0


^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH v3] RDMA/srpt: Clamp the CQ size request to max_cqe
  2026-09-06 16:12 [PATCH v3] RDMA/srpt: Clamp the CQ size request to max_cqe Serhat Kumral
@ 2026-09-07  6:32 ` Leon Romanovsky
  2026-09-07 11:47   ` Serhat Kumral
  2026-09-09 16:16   ` Bart Van Assche
  0 siblings, 2 replies; 4+ messages in thread
From: Leon Romanovsky @ 2026-09-07  6:32 UTC (permalink / raw)
  To: Serhat Kumral, Bart Van Assche
  Cc: Jason Gunthorpe, Sagi Grimberg, linux-rdma, target-devel, linux-kernel

On Sun, Sep 06, 2026 at 07:12:35PM +0300, Serhat Kumral wrote:
> srpt_create_ch_ib() asks ib_cq_pool_get() for ch->rq_size + sq_size
> completion queue entries. rq_size is bounded by max_qp_wr, but sq_size
> comes from the per-port srp_sq_size configfs attribute, which is only
> validated against MAX_SRPT_SRQ_SIZE (65535) and never compared with
> dev->attrs.max_cqe. The largest configuration srpt accepts therefore
> asks for 128 + 65535 = 65663 entries, while rxe caps max_cqe at 32767
> and cxgb4 and ionic are in the same range.
> 
> Such a request cannot be met, and ib_cq_pool_get() does not reject it.
> It clamps every CQ it creates to max_cqe, so no CQ it adds to the pool
> can ever fit the request, and it keeps allocating batches until the
> allocation fails. A single SRP login against such a target exhausts
> memory:
> 
>   Out of memory and no killable processes...
>   Kernel panic - not syncing: System is deadlocked on memory
>   Workqueue: ib_cm cm_work_handler
>   Call Trace:
>    __vmalloc_node_range_noprof
>    vmalloc_user_noprof
>    rxe_queue_init
>    rxe_cq_from_init
>    rxe_create_cq
>    __ib_alloc_cq
>    ib_cq_pool_get
>    srpt_cm_req_recv.cold
> 
> Before commit c804af2c1d31 ("IB/srpt: use new shared CQ mechanism") the
> same request went to ib_alloc_cq_any(), which rejected it with -EINVAL.
> 
> The existing backoff, which halves sq_size when queue pair creation
> fails, only runs after ib_cq_pool_get() has returned, so shrink sq_size
> before asking for the CQ. With the clamp the same login proceeds exactly
> like a correctly sized target.
> 
> Fixes: c804af2c1d31 ("IB/srpt: use new shared CQ mechanism")
> Signed-off-by: Serhat Kumral <serhatkumral1@gmail.com>
> ---
> Changes since v2:
> - Add a WARN_ON_ONCE().
> 
> Changes since v1:
> - Move the fix from the RDMA core to ib_srpt, as requested by Leon
>   Romanovsky.
> - Clamp sq_size before the CQ request instead of rejecting oversized
>   requests in ib_cq_pool_get().
> 
> v1: https://lore.kernel.org/linux-rdma/20260831171354.72140-1-serhatkumral1@gmail.com/
> v2: https://lore.kernel.org/linux-rdma/20260904200240.48976-1-serhatkumral1@gmail.com/
> 
>  drivers/infiniband/ulp/srpt/ib_srpt.c | 9 +++++++++
>  1 file changed, 9 insertions(+)
> 
> diff --git a/drivers/infiniband/ulp/srpt/ib_srpt.c b/drivers/infiniband/ulp/srpt/ib_srpt.c
> index 7197d95f2216..7ac526e3d2c2 100644
> --- a/drivers/infiniband/ulp/srpt/ib_srpt.c
> +++ b/drivers/infiniband/ulp/srpt/ib_srpt.c
> @@ -1867,6 +1867,15 @@ static int srpt_create_ch_ib(struct srpt_rdma_ch *ch)
>  	if (!qp_init)
>  		goto out;
>  
> +	/* The send and receive queues share a single CQ. */
> +	if (ch->rq_size + sq_size > attrs->max_cqe) {
> +		/* Catch drivers that incorrectly set ch->rq_size */
> +		WARN_ON_ONCE(ch->rq_size > attrs->max_cqe);
> +		sq_size = attrs->max_cqe - ch->rq_size;
> +		pr_debug("reduced sq_size to %u because max_cqe is %u\n",
> +			 sq_size, attrs->max_cqe);
> +	}

All these Sashiko reports suggest that sq_size is being changed in the wrong
place. Can you check for the correct value in srpt_tpg_attrib_srp_sq_size_store()?

Bart, can we simply decrease MAX_SRPT_SRQ_SIZE by, say, 100?

Thanks

> +
>  retry:
>  	ch->cq = ib_cq_pool_get(sdev->device, ch->rq_size + sq_size, -1,
>  				 IB_POLL_WORKQUEUE);
> 
> base-commit: cee9395acd8043be0644b25c34bfa86623f2b935
> -- 
> 2.53.0
> 

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH v3] RDMA/srpt: Clamp the CQ size request to max_cqe
  2026-09-07  6:32 ` Leon Romanovsky
@ 2026-09-07 11:47   ` Serhat Kumral
  2026-09-09 16:16   ` Bart Van Assche
  1 sibling, 0 replies; 4+ messages in thread
From: Serhat Kumral @ 2026-09-07 11:47 UTC (permalink / raw)
  To: leon
  Cc: bvanassche, jgg, sagi, linux-rdma, target-devel, linux-kernel,
	Serhat Kumral

On Mon, Sep 07, 2026 at 09:32:51AM +0300, Leon Romanovsky wrote:
> All these Sashiko reports suggest that sq_size is being changed in the wrong
> place. Can you check for the correct value in srpt_tpg_attrib_srp_sq_size_store()?

I think sashiko could be triggered again here due to theoretical max < sq+rq
probability because in the default (4096) state of sq, store() is not called.
max_cqe <= 127 and max_cqe < 4224 points to the same issue and it is not
resolved at this layer either.

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH v3] RDMA/srpt: Clamp the CQ size request to max_cqe
  2026-09-07  6:32 ` Leon Romanovsky
  2026-09-07 11:47   ` Serhat Kumral
@ 2026-09-09 16:16   ` Bart Van Assche
  1 sibling, 0 replies; 4+ messages in thread
From: Bart Van Assche @ 2026-09-09 16:16 UTC (permalink / raw)
  To: Leon Romanovsky, Serhat Kumral
  Cc: Jason Gunthorpe, Sagi Grimberg, linux-rdma, target-devel, linux-kernel

On 9/6/26 11:32 PM, Leon Romanovsky wrote:
> All these Sashiko reports suggest that sq_size is being changed in the wrong
> place. Can you check for the correct value in srpt_tpg_attrib_srp_sq_size_store()?
> 
> Bart, can we simply decrease MAX_SRPT_SRQ_SIZE by, say, 100?

Sure, that's fine with me.

Note: the default in the ib_srpt driver is not to use SRQ. SQR is only 
used if enabled explicitly via configfs. I'm not sure anyone is still
using the SRQ support in the ib_srpt driver.

Bart.

^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-09-09 16:16 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-06 16:12 [PATCH v3] RDMA/srpt: Clamp the CQ size request to max_cqe Serhat Kumral
2026-09-07  6:32 ` Leon Romanovsky
2026-09-07 11:47   ` Serhat Kumral
2026-09-09 16:16   ` Bart Van Assche

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®