mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Igor Paunovic <royalnet026@gmail.com>
To: Tomeu Vizoso <tomeu@tomeuvizoso.net>,
	Oded Gabbay <ogabbay@kernel.org>,
	Heiko Stuebner <heiko@sntech.de>
Cc: "Rob Herring" <robh@kernel.org>,
	"Krzysztof Kozlowski" <krzk+dt@kernel.org>,
	"Conor Dooley" <conor+dt@kernel.org>,
	"Jeff Hugo" <jeff.hugo@oss.qualcomm.com>,
	"Robert Foss" <rfoss@kernel.org>,
	"Sidong Yang" <sidong.yang@furiosa.ai>,
	"Diederik de Haas" <diederik@cknow-tech.com>,
	"Sebastian Reichel" <sebastian.reichel@collabora.com>,
	"Jiaxing Hu" <gahing@gahingwoo.com>,
	"Nicolas Dufresne" <nicolas@ndufresne.ca>,
	"Jonas Karlman" <jonas@kwiboo.se>,
	"Guangshuo Li" <lgs201920130244@gmail.com>,
	"Hüseyin BIYIK" <boogiepop@gmx.com>,
	dri-devel@lists.freedesktop.org,
	linux-rockchip@lists.infradead.org,
	linux-arm-kernel@lists.infradead.org, devicetree@vger.kernel.org,
	linux-kernel@vger.kernel.org,
	"Igor Paunovic" <royalnet026@gmail.com>
Subject: [PATCH v2 03/11] accel/rocket: search every core slot when looking up a scheduler
Date: Tue, 22 Sep 2026 10:01:06 +0200	[thread overview]
Message-ID: <20260922080114.44662-4-royalnet026@gmail.com> (raw)
In-Reply-To: <20260922080114.44662-1-royalnet026@gmail.com>

sched_to_core() walks rdev->cores[] up to rdev->num_cores, and
rocket_remove() decrements num_cores for every core it removes. Unbind a
core that is not the last one and the cores behind it fall outside the
search, so sched_to_core() returns NULL for a core that is still bound and
still running jobs. Neither caller checks the result:

  rocket_job_run():      rocket_fence_create(core), core->dev
  rocket_job_timedout(): dev_err(core->dev, "NPU job timed out")

Unbinding the middle core of the three on an RK3588 while three clients are
submitting to all of them faults twice, once from the surviving core's
job queue and once from its reset work:

  KASAN: null-ptr-deref in range [0x0000000000000220-0x0000000000000227]
  Workqueue: fdad0000.npu drm_sched_run_job_work [gpu_sched]
  pc : rocket_job_run+0x234/0x838 [rocket]
  Call trace:
   rocket_job_run+0x234/0x838 [rocket]
   drm_sched_run_job_work+0x2cc/0xad8 [gpu_sched]
   process_one_work+0x640/0x14f0

  KASAN: null-ptr-deref in range [0x0000000000000000-0x0000000000000007]
  Workqueue: rocket-reset-2 drm_sched_job_timedout [gpu_sched]
  pc : rocket_job_timedout+0xf0/0x1e0 [rocket]
  Call trace:
   rocket_job_timedout+0xf0/0x1e0 [rocket]
   drm_sched_job_timedout+0x188/0x6a0 [gpu_sched]

Both are the third core: the workqueue names are its device and its
core->index, and it was left at slot 2 while num_cores had dropped to 2.

Search all the slots that were allocated, the way find_core_for_dev() now
does. A core that is still bound is then found, and the two callers get
the pointer they already assume they have.

This does not make unbinding one core out of several safe. An open client
keeps an entity pointing at the scheduler of the core that went away:
drm_sched reports it as not ready for every job that lands on it, and the
client waits in dma_fence_default_wait for a fence that will never signal.
Stopping the NULL dereference is what belongs in a fix; the rest wants
more thought.

Reported-by: Sidong Yang <sidong.yang@furiosa.ai>
Closes: https://lore.kernel.org/dri-devel/apwUewaRnoTNXHCt@rock-5b-plus/
Fixes: 0810d5ad88a1 ("accel/rocket: Add job submission IOCTL")
Cc: stable@vger.kernel.org
Assisted-by: LLM sparse checkpatch
Signed-off-by: Igor Paunovic <royalnet026@gmail.com>
---
Unchanged from the standalone posting, which this series supersedes:
https://lore.kernel.org/r/20260905150432.7477-1-royalnet026@gmail.com

 drivers/accel/rocket/rocket_job.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/drivers/accel/rocket/rocket_job.c b/drivers/accel/rocket/rocket_job.c
index f404355058185..4bc4f9c8ee403 100644
--- a/drivers/accel/rocket/rocket_job.c
+++ b/drivers/accel/rocket/rocket_job.c
@@ -283,7 +283,7 @@ static struct rocket_core *sched_to_core(struct rocket_device *rdev,
 {
 	unsigned int core;
 
-	for (core = 0; core < rdev->num_cores; core++) {
+	for (core = 0; core < rdev->max_cores; core++) {
 		if (&rdev->cores[core].sched == sched)
 			return &rdev->cores[core];
 	}
-- 
2.43.0


  parent reply	other threads:[~2026-09-22  8:01 UTC|newest]

Thread overview: 22+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-22  8:01 [PATCH v2 00/11] accel/rocket: DVFS for the RK3588 NPU Igor Paunovic
2026-09-22  8:01 ` [PATCH v2 01/11] accel/rocket: search every core slot when a core is removed Igor Paunovic
2026-09-22  8:01 ` [PATCH v2 02/11] accel/rocket: number the cores by devicetree position, not bind order Igor Paunovic
2026-09-22  8:01 ` Igor Paunovic [this message]
2026-09-22  8:01 ` [PATCH v2 04/11] accel/rocket: keep core slots stable across unbind and rebind Igor Paunovic
2026-09-22  8:01 ` [PATCH v2 05/11] accel/rocket: request the core clocks by name Igor Paunovic
2026-09-22  8:01 ` [PATCH v2 06/11] dt-bindings: npu: rockchip: allow DVFS and thermal properties Igor Paunovic
2026-09-22 16:06   ` Rob Herring
2026-09-23  8:57     ` Igor Paunovic
2026-09-23  9:15       ` Diederik de Haas
2026-09-23  9:43         ` Igor Paunovic
2026-09-22  8:01 ` [PATCH v2 07/11] arm64: dts: rockchip: rk3588: add an OPP table for the NPU Igor Paunovic
2026-09-22  8:01 ` [PATCH v2 08/11] accel/rocket: restore the NPU clock boot rate before powering the cores down Igor Paunovic
     [not found]   ` <20260922081326.B46651F000FF@smtp.kernel.org>
2026-09-22  8:55     ` Igor Paunovic
2026-09-22  8:01 ` [PATCH v2 09/11] accel/rocket: add devfreq support Igor Paunovic
     [not found]   ` <20260922081855.160451F00893@smtp.kernel.org>
2026-09-22  8:56     ` Igor Paunovic
2026-09-23 13:14   ` Sidong Yang
2026-09-23 14:26     ` Igor Paunovic
2026-09-22  8:01 ` [PATCH v2 10/11] accel/rocket: register a devfreq cooling device Igor Paunovic
2026-09-22  8:01 ` [PATCH v2 11/11] arm64: dts: rockchip: rk3588: add passive cooling to the NPU thermal zone Igor Paunovic
2026-09-23 19:29 ` [PATCH v2 00/11] accel/rocket: DVFS for the RK3588 NPU Nicolas Dufresne
2026-09-23 19:54   ` Igor Paunovic

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260922080114.44662-4-royalnet026@gmail.com \
    --to=royalnet026@gmail.com \
    --cc=boogiepop@gmx.com \
    --cc=conor+dt@kernel.org \
    --cc=devicetree@vger.kernel.org \
    --cc=diederik@cknow-tech.com \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=gahing@gahingwoo.com \
    --cc=heiko@sntech.de \
    --cc=jeff.hugo@oss.qualcomm.com \
    --cc=jonas@kwiboo.se \
    --cc=krzk+dt@kernel.org \
    --cc=lgs201920130244@gmail.com \
    --cc=linux-arm-kernel@lists.infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-rockchip@lists.infradead.org \
    --cc=nicolas@ndufresne.ca \
    --cc=ogabbay@kernel.org \
    --cc=rfoss@kernel.org \
    --cc=robh@kernel.org \
    --cc=sebastian.reichel@collabora.com \
    --cc=sidong.yang@furiosa.ai \
    --cc=tomeu@tomeuvizoso.net \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®