mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Leon Romanovsky <leon@kernel.org>
To: Praveen Kumar Kannoju <praveen.kannoju@oracle.com>
Cc: jgg@ziepe.ca, saeedm@nvidia.com, tariqt@nvidia.com,
	mbloch@nvidia.com, davem@davemloft.net, kuba@kernel.org,
	pabeni@redhat.com, jiri@resnulli.us,
	kalesh-anakkur.purayil@broadcom.com, ynachum@amazon.com,
	kees@kernel.org, mrgolin@amazon.com, parav@nvidia.com,
	linux-rdma@vger.kernel.org, linux-kernel@vger.kernel.org,
	netdev@vger.kernel.org, anand.a.khoje@oracle.com
Subject: Re: [PATCH v2] RDMA/mlx5: Add poll-EQ callback for ULP recovery
Date: Mon, 28 Sep 2026 22:08:27 +0300	[thread overview]
Message-ID: <20260928190827.GB563127@unreal> (raw)
In-Reply-To: <20260921085006.1443240-1-praveen.kannoju@oracle.com>

On Mon, Sep 21, 2026 at 08:50:06AM +0000, Praveen Kumar Kannoju wrote:
> Some upper layer protocols, such as RDS, use mlx5 RDMA CQs whose
> completion EQs are not shared with mlx5e queues. If an EQ notification is
> missed for one of those CQs, the ULP can remain idle until another event
> arrives or a driver health recovery path polls the EQ.
> 
> mlx5e already has a recovery path that polls a completion EQ from the tx
> timeout handler through mlx5_eq_poll_irq_disabled(). That mechanism is
> internal to the mlx5 core and is not reachable from RDMA ULPs such as RDS
> when they need to recover a CQ-associated EQ.
> 
> Add an optional RDMA device operation, reap_eq, and an ib_reap_eq() helper
> so ULPs can ask the provider to poll the event queue associated with a CQ.
> Providers that do not implement the callback return -EOPNOTSUPP.
> 
> The recovery poll path is sleepable: mlx5 disables the IRQ synchronously
> and serializes recovery polling with a mutex. Document that ib_reap_eq()
> must be called from process context, assert the contract with
> might_sleep(), and reject interrupt-context or IRQ-disabled callers with
> -EWOULDBLOCK before invoking the provider callback.
> 
> Validate the device and CQ passed to ib_reap_eq() before dereferencing the
> callback table or provider-private CQ storage. The mlx5 callback also
> checks that the CQ belongs to the supplied device and that the mlx5 device
> and completion EQ are present before polling.
> 
> Implement the callback for mlx5 by mapping the ib_cq to the mlx5 CQ and
> polling the CQ's completion EQ through a new exported mlx5_eq_reap()
> helper. mlx5_eq_reap() logs the EQ state, invokes
> mlx5_eq_poll_irq_disabled(), and reports any recovered EQEs.
> 
> Serialize mlx5_eq_poll_irq_disabled() with a per-EQ mutex. The recovery
> poll path disables the IRQ, runs the EQ handler, advances the EQ consumer
> index, and updates the CI doorbell. Multiple recovery callers polling the
> same EQ concurrently could race on that state and reap the same EQ in
> parallel.
> 
> Use this only as a recovery path for missed EQ notifications. It is not a
> normal completion polling path.
> 
> Signed-off-by: Praveen Kumar Kannoju <praveen.kannoju@oracle.com>
> ---
> v1: https://lore.kernel.org/linux-rdma/20260919100609.732391F000FF@smtp.kernel.org/T/#t
> 
> Changes in v2:
> - Document ib_reap_eq() as process-context-only because the mlx5 recovery
>   path may sleep while synchronously disabling the IRQ.
> - Reject interrupt-context or IRQ-disabled callers with -EWOULDBLOCK before
>   invoking the provider callback.
> - Add might_sleep() assertions in ib_reap_eq() and
>   mlx5_eq_poll_irq_disabled().
> - Validate device/CQ input and CQ ownership before provider-private
>   dereferences.
> - Serialize mlx5_eq_poll_irq_disabled() with a per-EQ mutex so concurrent
>   recovery callers cannot reap the same EQ in parallel.
> 
>  drivers/infiniband/core/device.c              |  1 +
>  drivers/infiniband/hw/mlx5/main.c             | 18 ++++++++++
>  drivers/net/ethernet/mellanox/mlx5/core/eq.c  | 25 +++++++++++++-
>  .../net/ethernet/mellanox/mlx5/core/lib/eq.h  |  2 ++
>  include/linux/mlx5/eq.h                       |  2 ++
>  include/rdma/ib_verbs.h                       | 34 +++++++++++++++++++
>  6 files changed, 81 insertions(+), 1 deletion(-)

As noted for v1, we should not add an API for ULPs to work around
driver bugs.

Thanks

      reply	other threads:[~2026-09-28 19:08 UTC|newest]

Thread overview: 2+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-21  8:50 Praveen Kumar Kannoju
2026-09-28 19:08 ` Leon Romanovsky [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260928190827.GB563127@unreal \
    --to=leon@kernel.org \
    --cc=anand.a.khoje@oracle.com \
    --cc=davem@davemloft.net \
    --cc=jgg@ziepe.ca \
    --cc=jiri@resnulli.us \
    --cc=kalesh-anakkur.purayil@broadcom.com \
    --cc=kees@kernel.org \
    --cc=kuba@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-rdma@vger.kernel.org \
    --cc=mbloch@nvidia.com \
    --cc=mrgolin@amazon.com \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=parav@nvidia.com \
    --cc=praveen.kannoju@oracle.com \
    --cc=saeedm@nvidia.com \
    --cc=tariqt@nvidia.com \
    --cc=ynachum@amazon.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®