From: Patrisious Haddad <phaddad@nvidia.com>
To: Jason Gunthorpe <jgg@nvidia.com>
Cc: Edward Srouji <edwards@nvidia.com>,
Leon Romanovsky <leon@kernel.org>,
Chiara Meiohas <cmeiohas@nvidia.com>,
Maor Gottlieb <maorg@mellanox.com>,
Dennis Dalessandro <dennis.dalessandro@cornelisnetworks.com>,
Gal Pressman <galpress@amazon.com>,
Steve Wise <larrystevenwise@gmail.com>,
Mark Bloch <markb@mellanox.com>,
Mark Zhang <markzhang@nvidia.com>,
Neta Ostrovsky <netao@nvidia.com>,
linux-rdma@vger.kernel.org, linux-kernel@vger.kernel.org,
Michael Guralnik <michaelgur@nvidia.com>
Subject: Re: [PATCH rdma-next 0/6] RDMA: Fix restrack UAF in QP/CQ/SRQ destroy
Date: Sun, 14 Jun 2026 16:04:25 +0300 [thread overview]
Message-ID: <451aef3a-dafd-4754-8b43-066a1f204d2b@nvidia.com> (raw)
In-Reply-To: <20260612115205.GE1962447@nvidia.com>
On 6/12/2026 2:52 PM, Jason Gunthorpe wrote:
> On Fri, Jun 12, 2026 at 11:53:23AM +0300, Patrisious Haddad wrote:
>> On 6/11/2026 10:11 PM, Jason Gunthorpe wrote:
>>> On Sun, Jun 07, 2026 at 09:18:07PM +0300, Edward Srouji wrote:
>>>> The resource-tracking (restrack) database is the back-end for the netlink
>>>> "rdma resource show" interface which pins objects with
>>>> rdma_restrack_get().
>>>> The QP/CQ/SRQ destroy flows call rdma_restrack_del() at the end of
>>>> ib_destroy_*_user(), after device->ops.destroy_*() had already freed the
>>>> vendor object. Therefore, a concurrent netlink dump could look the
>>>> object up and touch freed memory, causing a use-after-free via
>>>> ib_query_qp() for instance.
>>>>
>>>> Fix this by splitting the delete into a begin/commit/abort sequence:
>>>> begin_del() parks the entry as XA_ZERO_ENTRY (so lookups return NULL),
>>>> drops the birth reference and waits for in-flight readers to drain,
>>>> while keeping the index reserved. The destroy paths run begin_del()
>>>> first, then commit_del() on success or abort_del() on error.
>>>> abort_del() re-inserts into the reserved slot, so it needs no allocation
>>>> and cannot fail.
>>>>
>>>> The first two patches remove DCT and raw RSS QP restrack tracking as
>>>> they have never worked (their ID is unset/reserved at create time).
>>>>
>>>> Signed-off-by: Edward Srouji <edwards@nvidia.com>
>>>> ---
>>>> Patrisious Haddad (6):
>>>> RDMA/mlx5: Remove DCT restrack tracking
>>>> RDMA/mlx5: Remove raw RSS QP restrack tracking
>>>> RDMA/core: Add rdma_restrack_begin/abort/commit_del() operations
>>>> RDMA/core: Fix use after free in ib_query_qp()
>>>> RDMA/core: Fix potential use after free in ib_destroy_cq_user()
>>>> RDMA/core: Fix potential use after free in ib_destroy_srq_user()
>>> The pre-existing sashiko issues look real too, can you fix them also:
>> Sure but one of them is a false-positive though:
>> Before destroy_qp() is called, the counter is unconditionally unbound:
>> rdma_counter_unbind_qp(qp, qp->port, true);
>> ret = qp->device->ops.destroy_qp(qp, udata);
>> If destroy_qp() fails and we abort destruction here, the kref on the
>> counter was dropped in rdma_counter_unbind_qp(), but qp->counter is never
>> set to NULL.
>>
>> This is actually wrong the qp->counter is actually set to NULL(inside the
>> driver though not the core) so subsequent calls will hit the NULL check and
>> return safely.
> That doesn't sound very good why is it like that, why is a driver
> making any permanent change on destroy failure?
No idea but thats totally unrelated issue so not fixing it with this series.
will follow up on it separately.
>
>> I wonder what about places where switching to commit/abort doesnt fix an
>> actual bug.
> I would change them all.
Will do and resend in next kernel cycle.
>
>> and ib_dereg_mr_user() actually calls the delete at start so it doesnt have
>> this issue (but it also doesnt readd it to restrack when driver OP fails) -
>> but here I think thats by design since the MR would be in weird state and we
>> dont want to track it ?
> That doesn't sound right either.
>
>>> I don't think this should NULL the task on abort either, it doesn't
>>> seem necessary.
>> I dont NULL the task on abort(I do NULL it on commit_del() though , or were
>> you talking about the restrack_sync() ?
> Since begin_del calls rdma_restrack_put() which calls through
> rdma_restrack_del it NULL's it.
Oh okay you mean rdma_restrack_put calls release() not del , but yeah I
resolved that by checking if the resource is still valid to not
release/null the task.
I'll refactor my code on top of restrack_sync to use same shared
functions and add few other places I missed and resend.
>
> Jason
next prev parent reply other threads:[~2026-06-14 13:04 UTC|newest]
Thread overview: 12+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-06-07 18:18 Edward Srouji
2026-06-07 18:18 ` [PATCH rdma-next 1/6] RDMA/mlx5: Remove DCT restrack tracking Edward Srouji
2026-06-07 18:18 ` [PATCH rdma-next 2/6] RDMA/mlx5: Remove raw RSS QP " Edward Srouji
2026-06-07 18:18 ` [PATCH rdma-next 3/6] RDMA/core: Add rdma_restrack_begin/abort/commit_del() operations Edward Srouji
2026-06-07 18:18 ` [PATCH rdma-next 4/6] RDMA/core: Fix use after free in ib_query_qp() Edward Srouji
2026-06-07 18:18 ` [PATCH rdma-next 5/6] RDMA/core: Fix potential use after free in ib_destroy_cq_user() Edward Srouji
2026-06-07 18:18 ` [PATCH rdma-next 6/6] RDMA/core: Fix potential use after free in ib_destroy_srq_user() Edward Srouji
2026-06-11 19:11 ` [PATCH rdma-next 0/6] RDMA: Fix restrack UAF in QP/CQ/SRQ destroy Jason Gunthorpe
2026-06-12 8:53 ` Patrisious Haddad
2026-06-12 11:52 ` Jason Gunthorpe
2026-06-14 13:04 ` Patrisious Haddad [this message]
2026-06-11 19:14 ` Jason Gunthorpe
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=451aef3a-dafd-4754-8b43-066a1f204d2b@nvidia.com \
--to=phaddad@nvidia.com \
--cc=cmeiohas@nvidia.com \
--cc=dennis.dalessandro@cornelisnetworks.com \
--cc=edwards@nvidia.com \
--cc=galpress@amazon.com \
--cc=jgg@nvidia.com \
--cc=larrystevenwise@gmail.com \
--cc=leon@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-rdma@vger.kernel.org \
--cc=maorg@mellanox.com \
--cc=markb@mellanox.com \
--cc=markzhang@nvidia.com \
--cc=michaelgur@nvidia.com \
--cc=netao@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®