From: Leon Romanovsky <leon@kernel.org>
To: Praveen Kumar Kannoju <praveen.kannoju@oracle.com>
Cc: jgg@ziepe.ca, saeedm@nvidia.com, tariqt@nvidia.com,
mbloch@nvidia.com, davem@davemloft.net, kuba@kernel.org,
pabeni@redhat.com, jiri@resnulli.us,
kalesh-anakkur.purayil@broadcom.com, ynachum@amazon.com,
kees@kernel.org, mrgolin@amazon.com, parav@nvidia.com,
linux-rdma@vger.kernel.org, linux-kernel@vger.kernel.org,
netdev@vger.kernel.org, anand.a.khoje@oracle.com
Subject: Re: [PATCH v2] RDMA/mlx5: Add poll-EQ callback for ULP recovery
Date: Mon, 28 Sep 2026 22:08:27 +0300 [thread overview]
Message-ID: <20260928190827.GB563127@unreal> (raw)
In-Reply-To: <20260921085006.1443240-1-praveen.kannoju@oracle.com>
On Mon, Sep 21, 2026 at 08:50:06AM +0000, Praveen Kumar Kannoju wrote:
> Some upper layer protocols, such as RDS, use mlx5 RDMA CQs whose
> completion EQs are not shared with mlx5e queues. If an EQ notification is
> missed for one of those CQs, the ULP can remain idle until another event
> arrives or a driver health recovery path polls the EQ.
>
> mlx5e already has a recovery path that polls a completion EQ from the tx
> timeout handler through mlx5_eq_poll_irq_disabled(). That mechanism is
> internal to the mlx5 core and is not reachable from RDMA ULPs such as RDS
> when they need to recover a CQ-associated EQ.
>
> Add an optional RDMA device operation, reap_eq, and an ib_reap_eq() helper
> so ULPs can ask the provider to poll the event queue associated with a CQ.
> Providers that do not implement the callback return -EOPNOTSUPP.
>
> The recovery poll path is sleepable: mlx5 disables the IRQ synchronously
> and serializes recovery polling with a mutex. Document that ib_reap_eq()
> must be called from process context, assert the contract with
> might_sleep(), and reject interrupt-context or IRQ-disabled callers with
> -EWOULDBLOCK before invoking the provider callback.
>
> Validate the device and CQ passed to ib_reap_eq() before dereferencing the
> callback table or provider-private CQ storage. The mlx5 callback also
> checks that the CQ belongs to the supplied device and that the mlx5 device
> and completion EQ are present before polling.
>
> Implement the callback for mlx5 by mapping the ib_cq to the mlx5 CQ and
> polling the CQ's completion EQ through a new exported mlx5_eq_reap()
> helper. mlx5_eq_reap() logs the EQ state, invokes
> mlx5_eq_poll_irq_disabled(), and reports any recovered EQEs.
>
> Serialize mlx5_eq_poll_irq_disabled() with a per-EQ mutex. The recovery
> poll path disables the IRQ, runs the EQ handler, advances the EQ consumer
> index, and updates the CI doorbell. Multiple recovery callers polling the
> same EQ concurrently could race on that state and reap the same EQ in
> parallel.
>
> Use this only as a recovery path for missed EQ notifications. It is not a
> normal completion polling path.
>
> Signed-off-by: Praveen Kumar Kannoju <praveen.kannoju@oracle.com>
> ---
> v1: https://lore.kernel.org/linux-rdma/20260919100609.732391F000FF@smtp.kernel.org/T/#t
>
> Changes in v2:
> - Document ib_reap_eq() as process-context-only because the mlx5 recovery
> path may sleep while synchronously disabling the IRQ.
> - Reject interrupt-context or IRQ-disabled callers with -EWOULDBLOCK before
> invoking the provider callback.
> - Add might_sleep() assertions in ib_reap_eq() and
> mlx5_eq_poll_irq_disabled().
> - Validate device/CQ input and CQ ownership before provider-private
> dereferences.
> - Serialize mlx5_eq_poll_irq_disabled() with a per-EQ mutex so concurrent
> recovery callers cannot reap the same EQ in parallel.
>
> drivers/infiniband/core/device.c | 1 +
> drivers/infiniband/hw/mlx5/main.c | 18 ++++++++++
> drivers/net/ethernet/mellanox/mlx5/core/eq.c | 25 +++++++++++++-
> .../net/ethernet/mellanox/mlx5/core/lib/eq.h | 2 ++
> include/linux/mlx5/eq.h | 2 ++
> include/rdma/ib_verbs.h | 34 +++++++++++++++++++
> 6 files changed, 81 insertions(+), 1 deletion(-)
As noted for v1, we should not add an API for ULPs to work around
driver bugs.
Thanks
prev parent reply other threads:[~2026-09-28 19:08 UTC|newest]
Thread overview: 2+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-21 8:50 Praveen Kumar Kannoju
2026-09-28 19:08 ` Leon Romanovsky [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260928190827.GB563127@unreal \
--to=leon@kernel.org \
--cc=anand.a.khoje@oracle.com \
--cc=davem@davemloft.net \
--cc=jgg@ziepe.ca \
--cc=jiri@resnulli.us \
--cc=kalesh-anakkur.purayil@broadcom.com \
--cc=kees@kernel.org \
--cc=kuba@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-rdma@vger.kernel.org \
--cc=mbloch@nvidia.com \
--cc=mrgolin@amazon.com \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=parav@nvidia.com \
--cc=praveen.kannoju@oracle.com \
--cc=saeedm@nvidia.com \
--cc=tariqt@nvidia.com \
--cc=ynachum@amazon.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®