From: Lizhi Hou <lizhi.hou@amd.com>
To: Max Zhen <max.zhen@amd.com>, <ogabbay@kernel.org>,
<quic_jhugo@quicinc.com>, <dri-devel@lists.freedesktop.org>,
<mario.limonciello@amd.com>, <karol.wachowski@linux.intel.com>
Cc: Wendy Liang <wendy.liang@amd.com>, <linux-kernel@vger.kernel.org>,
<sonal.santan@amd.com>
Subject: Re: [PATCH V1] accel/amdxdna: Fix command timeout race
Date: Sun, 19 Jul 2026 18:35:05 -0700 [thread overview]
Message-ID: <215a5b1b-7acc-08f3-8db4-1b46db1d66ec@amd.com> (raw)
In-Reply-To: <5ed0acbc-ec50-4055-9e03-cc4aa730c293@amd.com>
Applied to drm-misc-fixes
On 7/19/26 15:29, Max Zhen wrote:
>
>
> On 7/18/2026 Sat 01:34, Lizhi Hou wrote:
>> From: Wendy Liang <wendy.liang@amd.com>
>>
>> When two commands enter aie2_sched_job_timedout() concurrently, both
>> check the timeout detection state. The first scheduler thread observes
>> tdr_status as SIGNALED and updates it to WAIT. The second thread then
>> observes the updated state instead of the original SIGNALED state, which
>> may cause the command timeout to be handled incorrectly.
>>
>> Replace tdr_status with last_signal_ts, which records the timestamp of
>> the last driver signal. Timeout detection now only reads
>> last_signal_ts and never modifies it, allowing multiple serialized
>> detect() calls under dev_lock to evaluate the same signal timestamp
>> independently. If there is not any new job scheduled or completed
>> within tdr_timeout_ms, the command will timeout.
>>
>> Fixes: 9022f010977f ("accel/amdxdna: Check for device hang on job
>> timeout")
>> Signed-off-by: Wendy Liang <wendy.liang@amd.com>
>> Signed-off-by: Lizhi Hou <lizhi.hou@amd.com>
> Reviewed-by: Max Zhen <max.zhen@amd.com>
>> ---
>> drivers/accel/amdxdna/aie2_ctx.c | 22 +++++++++++++++-------
>> drivers/accel/amdxdna/aie2_pci.c | 1 +
>> drivers/accel/amdxdna/aie2_pci.h | 7 +------
>> 3 files changed, 17 insertions(+), 13 deletions(-)
>>
>> diff --git a/drivers/accel/amdxdna/aie2_ctx.c
>> b/drivers/accel/amdxdna/aie2_ctx.c
>> index 101f324ee178..94dfee7263bd 100644
>> --- a/drivers/accel/amdxdna/aie2_ctx.c
>> +++ b/drivers/accel/amdxdna/aie2_ctx.c
>> @@ -43,20 +43,22 @@ struct aie2_ctx_health {
>> static inline void aie2_tdr_signal(struct amdxdna_dev *xdna)
>> {
>> - WRITE_ONCE(xdna->dev_handle->tdr_status, AIE2_TDR_SIGNALED);
>> + WRITE_ONCE(xdna->dev_handle->last_signal_ts, jiffies);
>> }
>> static bool aie2_tdr_detect(struct amdxdna_dev *xdna)
>> {
>> struct amdxdna_dev_hdl *ndev = xdna->dev_handle;
>> + unsigned long last = READ_ONCE(ndev->last_signal_ts);
>> - if (READ_ONCE(ndev->tdr_status) == AIE2_TDR_WAIT) {
>> - XDNA_ERR(xdna, "TDR timeout detected");
>> - return true;
>> - }
>> + if (!tdr_timeout_ms)
>> + return false;
>> +
>> + if (!time_after(jiffies, last + msecs_to_jiffies(tdr_timeout_ms)))
>> + return false;
>> - WRITE_ONCE(ndev->tdr_status, AIE2_TDR_WAIT);
>> - return false;
>> + XDNA_ERR(xdna, "TDR timeout detected");
>> + return true;
>> }
>> static void aie2_cmd_release(struct kref *ref)
>> @@ -434,6 +436,12 @@ aie2_sched_job_run(struct drm_sched_job *sched_job)
>> mmput(job->mm);
>> fence = ERR_PTR(ret);
>> } else {
>> + /*
>> + * Command is successfully posted to hardware, update the
>> + * tdr timestamp. The total pending commands are limited.
>> + * So there will not be a case that driver keeps posting
>> + * commands without getting any hardware respond.
>> + */
>> aie2_tdr_signal(hwctx->client->xdna);
>> }
>> trace_xdna_job(sched_job, hwctx->name, "sent to device",
>> diff --git a/drivers/accel/amdxdna/aie2_pci.c
>> b/drivers/accel/amdxdna/aie2_pci.c
>> index 22f66c7f534d..daec1f6b4907 100644
>> --- a/drivers/accel/amdxdna/aie2_pci.c
>> +++ b/drivers/accel/amdxdna/aie2_pci.c
>> @@ -420,6 +420,7 @@ static int aie2_hw_start(struct amdxdna_dev *xdna)
>> goto stop_fw;
>> }
>> + WRITE_ONCE(ndev->last_signal_ts, jiffies);
>> ndev->dev_status = AIE2_DEV_START;
>> return 0;
>> diff --git a/drivers/accel/amdxdna/aie2_pci.h
>> b/drivers/accel/amdxdna/aie2_pci.h
>> index 77648cc548b6..ea1dac106400 100644
>> --- a/drivers/accel/amdxdna/aie2_pci.h
>> +++ b/drivers/accel/amdxdna/aie2_pci.h
>> @@ -143,11 +143,6 @@ struct aie2_exec_msg_ops {
>> u32 (*get_chain_msg_op)(u32 cmd_op);
>> };
>> -enum aie2_tdr_status {
>> - AIE2_TDR_WAIT,
>> - AIE2_TDR_SIGNALED,
>> -};
>> -
>> struct amdxdna_dev_hdl {
>> struct aie_device aie;
>> const struct amdxdna_dev_priv *priv;
>> @@ -179,7 +174,7 @@ struct amdxdna_dev_hdl {
>> u32 hwctx_num;
>> struct amdxdna_async_error last_async_err;
>> - enum aie2_tdr_status tdr_status;
>> + unsigned long last_signal_ts;
>> };
>> struct aie2_hw_ops {
>
prev parent reply other threads:[~2026-07-20 1:35 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-18 8:34 Lizhi Hou
2026-07-19 22:29 ` Max Zhen
2026-07-20 1:35 ` Lizhi Hou [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=215a5b1b-7acc-08f3-8db4-1b46db1d66ec@amd.com \
--to=lizhi.hou@amd.com \
--cc=dri-devel@lists.freedesktop.org \
--cc=karol.wachowski@linux.intel.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mario.limonciello@amd.com \
--cc=max.zhen@amd.com \
--cc=ogabbay@kernel.org \
--cc=quic_jhugo@quicinc.com \
--cc=sonal.santan@amd.com \
--cc=wendy.liang@amd.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®