From: Igor Paunovic <royalnet026@gmail.com>
To: Tomeu Vizoso <tomeu@tomeuvizoso.net>, Oded Gabbay <ogabbay@kernel.org>
Cc: Sidong Yang <sidong.yang@furiosa.ai>,
Heiko Stuebner <heiko@sntech.de>,
Jiaxing Hu <gahing@gahingwoo.com>,
dri-devel@lists.freedesktop.org,
linux-rockchip@lists.infradead.org,
linux-arm-kernel@lists.infradead.org,
linux-kernel@vger.kernel.org,
Igor Paunovic <royalnet026@gmail.com>,
stable@vger.kernel.org
Subject: [PATCH] accel/rocket: search every core slot when looking up a scheduler
Date: Sat, 5 Sep 2026 17:04:32 +0200 [thread overview]
Message-ID: <20260905150432.7477-1-royalnet026@gmail.com> (raw)
sched_to_core() walks rdev->cores[] up to rdev->num_cores, and
rocket_remove() decrements num_cores for every core it removes. Unbind a
core that is not the last one and the cores behind it fall outside the
search, so sched_to_core() returns NULL for a core that is still bound and
still running jobs. Neither caller checks the result:
rocket_job_run(): rocket_fence_create(core), core->dev
rocket_job_timedout(): dev_err(core->dev, "NPU job timed out")
Unbinding the middle core of the three on an RK3588 while three clients are
submitting to all of them faults twice, once from the surviving core's
job queue and once from its reset work:
KASAN: null-ptr-deref in range [0x0000000000000220-0x0000000000000227]
Workqueue: fdad0000.npu drm_sched_run_job_work [gpu_sched]
pc : rocket_job_run+0x234/0x838 [rocket]
Call trace:
rocket_job_run+0x234/0x838 [rocket]
drm_sched_run_job_work+0x2cc/0xad8 [gpu_sched]
process_one_work+0x640/0x14f0
KASAN: null-ptr-deref in range [0x0000000000000000-0x0000000000000007]
Workqueue: rocket-reset-2 drm_sched_job_timedout [gpu_sched]
pc : rocket_job_timedout+0xf0/0x1e0 [rocket]
Call trace:
rocket_job_timedout+0xf0/0x1e0 [rocket]
drm_sched_job_timedout+0x188/0x6a0 [gpu_sched]
Both are the third core: the workqueue names are its device and its
core->index, and it was left at slot 2 while num_cores had dropped to 2.
Search all the slots that were allocated, the way find_core_for_dev() now
does. A core that is still bound is then found, and the two callers get
the pointer they already assume they have.
This does not make unbinding one core out of several safe. An open client
keeps an entity pointing at the scheduler of the core that went away:
drm_sched reports it as not ready for every job that lands on it, and the
client waits in dma_fence_default_wait for a fence that will never signal.
Stopping the NULL dereference is what belongs in a fix; the rest wants
more thought.
Reported-by: Sidong Yang <sidong.yang@furiosa.ai>
Closes: https://lore.kernel.org/dri-devel/apwUewaRnoTNXHCt@rock-5b-plus/
Fixes: 0810d5ad88a1 ("accel/rocket: Add job submission IOCTL")
Cc: stable@vger.kernel.org
Signed-off-by: Igor Paunovic <royalnet026@gmail.com>
Assisted-by: LLM sparse checkpatch
---
drivers/accel/rocket/rocket_job.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/drivers/accel/rocket/rocket_job.c b/drivers/accel/rocket/rocket_job.c
index 3141f210fcd1b..a6c24dfe0563a 100644
--- a/drivers/accel/rocket/rocket_job.c
+++ b/drivers/accel/rocket/rocket_job.c
@@ -283,7 +283,7 @@ static struct rocket_core *sched_to_core(struct rocket_device *rdev,
{
unsigned int core;
- for (core = 0; core < rdev->num_cores; core++) {
+ for (core = 0; core < rdev->max_cores; core++) {
if (&rdev->cores[core].sched == sched)
return &rdev->cores[core];
}
base-commit: a9f09b5ea0c3db1e2d4c0f8d3ebdd612d8aa0366
prerequisite-patch-id: 519bcdfdde80d902309c8346f749ebc4bb6b29c0
--
2.43.0
next reply other threads:[~2026-09-05 15:04 UTC|newest]
Thread overview: 2+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-05 15:04 Igor Paunovic [this message]
[not found] ` <20260905151815.3BE7F1F00A3A@smtp.kernel.org>
2026-09-05 15:27 ` Igor Paunovic
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260905150432.7477-1-royalnet026@gmail.com \
--to=royalnet026@gmail.com \
--cc=dri-devel@lists.freedesktop.org \
--cc=gahing@gahingwoo.com \
--cc=heiko@sntech.de \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-rockchip@lists.infradead.org \
--cc=ogabbay@kernel.org \
--cc=sidong.yang@furiosa.ai \
--cc=stable@vger.kernel.org \
--cc=tomeu@tomeuvizoso.net \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®