From: Lizhi Hou <lizhi.hou@amd.com>
To: Mario Limonciello <mario.limonciello@amd.com>,
<ogabbay@kernel.org>, <quic_jhugo@quicinc.com>,
<dri-devel@lists.freedesktop.org>,
<maciej.falkowski@linux.intel.com>
Cc: <linux-kernel@vger.kernel.org>, <max.zhen@amd.com>,
<sonal.santan@amd.com>
Subject: Re: [PATCH V1] accel/amdxdna: Fix race where send ring appears full due to delayed head update
Date: Fri, 12 Dec 2025 09:56:35 -0800 [thread overview]
Message-ID: <b17b259f-626a-c62a-3d57-777edbf303d9@amd.com> (raw)
In-Reply-To: <15ee6051-0639-4c4a-bc1c-8762a5ae5128@amd.com>
Applied to drm-misc-next.
On 12/12/25 08:43, Mario Limonciello wrote:
> On 12/11/25 10:41 PM, Lizhi Hou wrote:
>>
>> On 12/11/25 13:48, Mario Limonciello wrote:
>>> On 12/10/25 10:51 PM, Lizhi Hou wrote:
>>>> The firmware sends a response and interrupts the driver before
>>>> advancing
>>>> the mailbox send ring head pointer.
>>>
>>> What's the point of the interrupt then? Is this possible to improve
>>> in future firmware or is this really a hardware issue? If it can be
>>> fixed in firmware it would be ideal to conditionalize such behavior
>>> on firmware version.
>>
>> This has been found recently with a muti-thread stress test with
>> released firmware. My understanding is that this is a firmware issue.
>> I am not sure if it could be improved in future firmware.
>
> OK, thanks for explaining.
>
>>
>> I believe this driver change is more robust. It does not change the
>> logic but adds an polling instead of returning error. If the firmware
>> updates header quick, there will not be any polling. It works for
>> both current and future firmware.
>>
>>
>
> Yeah I don't think it's harmful, it's just unfortunate to have to land
> code like this when it sounds like a firmware bug with interrupt
> handling. If the firmware can be improved at least it will be right
> on the very first read in the polling.
>
> Reviewed-by: Mario Limonciello (AMD)
>
>> Thanks,
>>
>> Lizhi
>>
>>>> As a result, the driver may observe
>>>> the response and attempt to send a new request before the firmware has
>>>> updated the head pointer. In this window, the send ring still appears
>>>> full, causing the driver to incorrectly fail the send operation.
>>>>
>>>> This race can be triggered more easily in a multithreaded environment,
>>>> leading to unexpected and spurious "send ring full" failures.
>>>>
>>>> To address this, poll the send ring head pointer for up to 100us
>>>> before
>>>> returning a full-ring condition. This allows the firmware time to
>>>> update
>>>> the head pointer.
>>>>
>>>> Fixes: b87f920b9344 ("accel/amdxdna: Support hardware mailbox")
>>>> Signed-off-by: Lizhi Hou <lizhi.hou@amd.com>
>>>> ---
>>>> drivers/accel/amdxdna/amdxdna_mailbox.c | 27
>>>> +++++++++++++++----------
>>>> 1 file changed, 16 insertions(+), 11 deletions(-)
>>>>
>>>> diff --git a/drivers/accel/amdxdna/amdxdna_mailbox.c
>>>> b/drivers/accel/ amdxdna/amdxdna_mailbox.c
>>>> index a60a85ce564c..469242ed8224 100644
>>>> --- a/drivers/accel/amdxdna/amdxdna_mailbox.c
>>>> +++ b/drivers/accel/amdxdna/amdxdna_mailbox.c
>>>> @@ -191,26 +191,34 @@ mailbox_send_msg(struct mailbox_channel
>>>> *mb_chann, struct mailbox_msg *mb_msg)
>>>> u32 head, tail;
>>>> u32 start_addr;
>>>> u32 tmp_tail;
>>>> + int ret;
>>>> head = mailbox_get_headptr(mb_chann, CHAN_RES_X2I);
>>>> tail = mb_chann->x2i_tail;
>>>> - ringbuf_size = mailbox_get_ringbuf_size(mb_chann, CHAN_RES_X2I);
>>>> + ringbuf_size = mailbox_get_ringbuf_size(mb_chann,
>>>> CHAN_RES_X2I) - sizeof(u32);
>>>> start_addr = mb_chann->res[CHAN_RES_X2I].rb_start_addr;
>>>> tmp_tail = tail + mb_msg->pkg_size;
>>>> - if (tail < head && tmp_tail >= head)
>>>> - goto no_space;
>>>> -
>>>> - if (tail >= head && (tmp_tail > ringbuf_size - sizeof(u32) &&
>>>> - mb_msg->pkg_size >= head))
>>>> - goto no_space;
>>>> - if (tail >= head && tmp_tail > ringbuf_size - sizeof(u32)) {
>>>> +check_again:
>>>> + if (tail >= head && tmp_tail > ringbuf_size) {
>>>> write_addr = mb_chann->mb->res.ringbuf_base + start_addr
>>>> + tail;
>>>> writel(TOMBSTONE, write_addr);
>>>> /* tombstone is set. Write from the start of the
>>>> ringbuf */
>>>> tail = 0;
>>>> + tmp_tail = tail + mb_msg->pkg_size;
>>>> + }
>>>> +
>>>> + if (tail < head && tmp_tail >= head) {
>>>> + ret = read_poll_timeout(mailbox_get_headptr, head,
>>>> + tmp_tail < head || tail >= head,
>>>> + 1, 100, false, mb_chann, CHAN_RES_X2I);
>>>> + if (ret)
>>>> + return ret;
>>>> +
>>>> + if (tail >= head)
>>>> + goto check_again;
>>>> }
>>>> write_addr = mb_chann->mb->res.ringbuf_base + start_addr +
>>>> tail;
>>>> @@ -222,9 +230,6 @@ mailbox_send_msg(struct mailbox_channel
>>>> *mb_chann, struct mailbox_msg *mb_msg)
>>>> mb_msg->pkg.header.id);
>>>> return 0;
>>>> -
>>>> -no_space:
>>>> - return -ENOSPC;
>>>> }
>>>> static int
>>>
>
prev parent reply other threads:[~2025-12-12 17:56 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-12-11 4:51 Lizhi Hou
2025-12-11 21:48 ` Mario Limonciello
2025-12-12 4:41 ` Lizhi Hou
2025-12-12 16:43 ` Mario Limonciello
2025-12-12 17:56 ` Lizhi Hou [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=b17b259f-626a-c62a-3d57-777edbf303d9@amd.com \
--to=lizhi.hou@amd.com \
--cc=dri-devel@lists.freedesktop.org \
--cc=linux-kernel@vger.kernel.org \
--cc=maciej.falkowski@linux.intel.com \
--cc=mario.limonciello@amd.com \
--cc=max.zhen@amd.com \
--cc=ogabbay@kernel.org \
--cc=quic_jhugo@quicinc.com \
--cc=sonal.santan@amd.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®