* [PATCH v2 0/3] nouveau: fix 3 null-ptr derefs
@ 2026-10-09 17:44 Jim Cromie via B4 Relay
2026-10-09 17:44 ` [PATCH v2 1/3] nouveau: Fix NULL pointer dereferences in GETPARAM ioctl Jim Cromie via B4 Relay
` (2 more replies)
0 siblings, 3 replies; 11+ messages in thread
From: Jim Cromie via B4 Relay @ 2026-10-09 17:44 UTC (permalink / raw)
To: Lyude Paul, Danilo Krummrich, Maarten Lankhorst, Maxime Ripard,
Thomas Zimmermann, David Airlie, Simona Vetter, Dave Airlie
Cc: dri-devel, nouveau, linux-kernel, Daniel Campos Ramos, Jim Cromie
These 3 patches fix NULL ptr derefs I hit while booting kernels with
unrelated work (dynamic-debug) on a 2019 amdgpu+noueau laptop. One of
them is possibly hard to reproduce, the others are "dependable".
Signed-off-by: Jim Cromie <jim.cromie@gmail.com>
---
Changes in v2:
- added check of gsp->rm into nvkm_gsp_rm(gsp)
- change if-then-else to if-early-fail, fallthru
- Link to v1: https://lore.kernel.org/r/20261002-my-fixups-v1-0-a83d20f9d3fe@gmail.com
---
Jim Cromie (3):
nouveau: Fix NULL pointer dereferences in GETPARAM ioctl
nouveau: Fix NULL pointer dereference in GET_ZCULL_INFO ioctl
nouveau: Fix NULL dereference in nvkm_gsp_gcx_ready()
drivers/gpu/drm/nouveau/include/nvkm/subdev/gsp.h | 2 +-
drivers/gpu/drm/nouveau/nouveau_abi16.c | 24 ++++++++++++++++++++---
drivers/gpu/drm/nouveau/nvkm/engine/device/user.c | 2 +-
drivers/gpu/drm/nouveau/nvkm/subdev/gsp/base.c | 2 +-
4 files changed, 24 insertions(+), 6 deletions(-)
---
base-commit: 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e
change-id: 20261002-my-fixups-324d25f3bd05
Best regards,
--
Jim Cromie <jim.cromie@gmail.com>
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH v2 1/3] nouveau: Fix NULL pointer dereferences in GETPARAM ioctl
2026-10-09 17:44 [PATCH v2 0/3] nouveau: fix 3 null-ptr derefs Jim Cromie via B4 Relay
@ 2026-10-09 17:44 ` Jim Cromie via B4 Relay
2026-10-09 21:24 ` lyude
2026-10-09 21:47 ` Danilo Krummrich
2026-10-09 17:44 ` [PATCH v2 2/3] nouveau: Fix NULL pointer dereference in GET_ZCULL_INFO ioctl Jim Cromie via B4 Relay
2026-10-09 17:44 ` [PATCH v2 3/3] nouveau: Fix NULL dereference in nvkm_gsp_gcx_ready() Jim Cromie via B4 Relay
2 siblings, 2 replies; 11+ messages in thread
From: Jim Cromie via B4 Relay @ 2026-10-09 17:44 UTC (permalink / raw)
To: Lyude Paul, Danilo Krummrich, Maarten Lankhorst, Maxime Ripard,
Thomas Zimmermann, David Airlie, Simona Vetter, Dave Airlie
Cc: dri-devel, nouveau, linux-kernel, Daniel Campos Ramos, Jim Cromie
From: Jim Cromie <jim.cromie@gmail.com>
When hardware or firmware initialization fails, the graphics engine
(gr) or device functions may remain NULL. Attempting to access these
during the GETPARAM ioctl (e.g., NOUVEAU_GETPARAM_GRAPH_UNITS) results
in a kernel NULL pointer dereference, causing a crash when userspace
(GNOME/Mesa) attempts to probe the device.
Add safety checks for 'gr', 'gr->func', and 'nvkm_device->func' in the
ioctl handler. Return -ENODEV to signal the missing hardware state to
userspace, and use NV_ERROR_ONCE to provide diagnostic proof in the
kernel log without risking a console flood.
NOTE:
I encountered these crashes while booting unrelated code (in
dynamic-debug) on real HW. There was also some fwupd updates around
this time, so it could all be transient/unreproducible. But it did
happen, and could happen again in unforseen circumstances.
@Capitain_Jack tested the error path in v1 by deleting the FW
Signed-off-by: Jim Cromie <jim.cromie@gmail.com>
Tested-by: Daniel Campos Ramos <Capitain_Jack@yahoo.com>
---
drivers/gpu/drm/nouveau/nouveau_abi16.c | 22 ++++++++++++++++++++--
1 file changed, 20 insertions(+), 2 deletions(-)
diff --git a/drivers/gpu/drm/nouveau/nouveau_abi16.c b/drivers/gpu/drm/nouveau/nouveau_abi16.c
index 4542d5f4ded8..3f130cd4fbcd 100644
--- a/drivers/gpu/drm/nouveau/nouveau_abi16.c
+++ b/drivers/gpu/drm/nouveau/nouveau_abi16.c
@@ -306,6 +306,11 @@ nouveau_abi16_ioctl_getparam(ABI16_IOCTL_ARGS)
getparam->value = 1;
break;
case NOUVEAU_GETPARAM_GRAPH_UNITS:
+ if (!gr || !gr->func) {
+ NV_ERROR_ONCE(drm, "GETPARAM_GRAPH_UNITS: no gr engine or func\n");
+ return -ENODEV;
+ }
+
getparam->value = nvkm_gr_units(gr);
break;
case NOUVEAU_GETPARAM_EXEC_PUSH_MAX: {
@@ -315,10 +320,23 @@ nouveau_abi16_ioctl_getparam(ABI16_IOCTL_ARGS)
break;
}
case NOUVEAU_GETPARAM_VRAM_BAR_SIZE:
- getparam->value = nvkm_device->func->resource_size(nvkm_device, NVKM_BAR1_FB);
+ if (!nvkm_device || !nvkm_device->func || !nvkm_device->func->resource_size) {
+ NV_ERROR_ONCE(drm, "GETPARAM_VRAM_BAR_SIZE: no device func\n");
+ return -ENODEV;
+ }
+
+ getparam->value =
+ nvkm_device->func->resource_size(nvkm_device, NVKM_BAR1_FB);
break;
case NOUVEAU_GETPARAM_VRAM_USED: {
- struct ttm_resource_manager *vram_mgr = ttm_manager_type(&drm->ttm.bdev, TTM_PL_VRAM);
+ struct ttm_resource_manager *vram_mgr =
+ ttm_manager_type(&drm->ttm.bdev, TTM_PL_VRAM);
+
+ if (!vram_mgr) {
+ NV_ERROR_ONCE(drm, "GETPARAM_VRAM_USED: no vram mgr\n");
+ return -ENODEV;
+ }
+
getparam->value = (u64)ttm_resource_manager_usage(vram_mgr);
break;
}
--
2.56.0
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH v2 2/3] nouveau: Fix NULL pointer dereference in GET_ZCULL_INFO ioctl
2026-10-09 17:44 [PATCH v2 0/3] nouveau: fix 3 null-ptr derefs Jim Cromie via B4 Relay
2026-10-09 17:44 ` [PATCH v2 1/3] nouveau: Fix NULL pointer dereferences in GETPARAM ioctl Jim Cromie via B4 Relay
@ 2026-10-09 17:44 ` Jim Cromie via B4 Relay
2026-10-09 21:25 ` lyude
2026-10-09 21:48 ` Danilo Krummrich
2026-10-09 17:44 ` [PATCH v2 3/3] nouveau: Fix NULL dereference in nvkm_gsp_gcx_ready() Jim Cromie via B4 Relay
2 siblings, 2 replies; 11+ messages in thread
From: Jim Cromie via B4 Relay @ 2026-10-09 17:44 UTC (permalink / raw)
To: Lyude Paul, Danilo Krummrich, Maarten Lankhorst, Maxime Ripard,
Thomas Zimmermann, David Airlie, Simona Vetter, Dave Airlie
Cc: dri-devel, nouveau, linux-kernel, Daniel Campos Ramos, Jim Cromie
From: Jim Cromie <jim.cromie@gmail.com>
When graphics engine firmware fails to load or initialization aborts
early, nvxx_gr(drm) returns NULL. Calling DRM_IOCTL_NOUVEAU_GET_ZCULL_INFO
causes nouveau_abi16_ioctl_get_zcull_info() to dereference gr at offset
0xf0 without checking for NULL, triggering a kernel page fault.
Validate that gr is non-NULL before inspecting gr->has_zcull_info.
Signed-off-by: Jim Cromie <jim.cromie@gmail.com>
---
Aug 14 09:21:52 frodo kernel: BUG: kernel NULL pointer dereference, address: 00000000000000f0
Aug 14 09:21:52 frodo kernel: #PF: error_code(0x0000) - not-present page
Aug 14 09:21:52 frodo kernel: Oops: Oops: 0000 [#1] SMP NOPTI
Aug 14 09:21:52 frodo kernel: RIP: 0010:nouveau_abi16_ioctl_get_zcull_info+0x17/0xa0 [nouveau]
Aug 14 09:21:52 frodo kernel: Code: 00 00 00 90 90 90 90 90 90 90 90 90 90 90 90 90 90 90 90 f3 0f 1e fa 0f 1f 44 00 00 48 8b 47 40 48 8b 00 48 8b 80 70 02 00 00 <80> b8 f0 00 00 00 00 74 72 8b 90 c0 00 00 00 89 16 8b 90 c4 00 00
Aug 14 09:21:52 frodo kernel: Call Trace:
Aug 14 09:21:52 frodo kernel: <TASK>
Aug 14 09:21:52 frodo kernel: drm_ioctl_kernel+0xae/0x100
Aug 14 09:21:52 frodo kernel: drm_ioctl+0x2e0/0x560
Aug 14 09:21:52 frodo kernel: ? __pfx_nouveau_abi16_ioctl_get_zcull_info+0x10/0x10 [nouveau]
Aug 14 09:21:52 frodo kernel: nouveau_drm_ioctl+0x58/0xc0 [nouveau]
Aug 14 09:21:52 frodo kernel: __x64_sys_ioctl+0xb9/0x100
Aug 14 09:21:52 frodo kernel: do_syscall_64+0xe2/0x560
Aug 14 09:21:52 frodo kernel: entry_SYSCALL_64_after_hwframe+0x76/0x7e
Aug 14 09:21:52 frodo kernel: </TASK>
Reviewed-by: Lyude Paul <lyude@redhat.com>
Tested-by: Daniel Campos Ramos <Capitain_Jack@yahoo.com>
---
drivers/gpu/drm/nouveau/nouveau_abi16.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/drivers/gpu/drm/nouveau/nouveau_abi16.c b/drivers/gpu/drm/nouveau/nouveau_abi16.c
index 3f130cd4fbcd..c7ddc54a3c4d 100644
--- a/drivers/gpu/drm/nouveau/nouveau_abi16.c
+++ b/drivers/gpu/drm/nouveau/nouveau_abi16.c
@@ -358,7 +358,7 @@ nouveau_abi16_ioctl_get_zcull_info(ABI16_IOCTL_ARGS)
struct nvkm_gr *gr = nvxx_gr(drm);
struct drm_nouveau_get_zcull_info *out = data;
- if (gr->has_zcull_info) {
+ if (gr && gr->has_zcull_info) {
const struct nvkm_gr_zcull_info *i = &gr->zcull_info;
out->width_align_pixels = i->width_align_pixels;
--
2.56.0
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH v2 3/3] nouveau: Fix NULL dereference in nvkm_gsp_gcx_ready()
2026-10-09 17:44 [PATCH v2 0/3] nouveau: fix 3 null-ptr derefs Jim Cromie via B4 Relay
2026-10-09 17:44 ` [PATCH v2 1/3] nouveau: Fix NULL pointer dereferences in GETPARAM ioctl Jim Cromie via B4 Relay
2026-10-09 17:44 ` [PATCH v2 2/3] nouveau: Fix NULL pointer dereference in GET_ZCULL_INFO ioctl Jim Cromie via B4 Relay
@ 2026-10-09 17:44 ` Jim Cromie via B4 Relay
2026-10-09 21:26 ` lyude
2026-10-09 21:53 ` Danilo Krummrich
2 siblings, 2 replies; 11+ messages in thread
From: Jim Cromie via B4 Relay @ 2026-10-09 17:44 UTC (permalink / raw)
To: Lyude Paul, Danilo Krummrich, Maarten Lankhorst, Maxime Ripard,
Thomas Zimmermann, David Airlie, Simona Vetter, Dave Airlie
Cc: dri-devel, nouveau, linux-kernel, Daniel Campos Ramos, Jim Cromie
From: Jim Cromie <jim.cromie@gmail.com>
Commit cb4c7603678c ("drm/nouveau/gsp/r570: Add support for
INTERNAL_GCX_ENTRY_PREREQUISITE") hooked nvif_device_gcx_ready() into
nouveau_pmops_runtime_suspend() to consult GSP before runtime suspend.
On GPUs running without GSP-RM (e.g. Volta GV100, NvGspRm=0, or missing
GSP firmware blobs), device->gsp is instantiated but gsp->rm remains
NULL. nvkm_udevice_gcx_ready() checked "!gsp" instead of verifying
nvkm_gsp_rm(gsp), allowing non-RM instances to invoke
nvkm_gsp_gcx_ready(). nvkm_gsp_gcx_ready() then unconditionally
dereferenced gsp->rm->api, causing a kernel oops:
RIP: 0010:nvkm_gsp_gcx_ready+0x10/0x30 [nouveau]
Code: ... 48 8b 87 70 0e 00 00 <48> 8b 40 18 ...
RAX: 0000000000000000 RDI: ffff8c9555fca000
Call Trace:
nvkm_udevice_mthd+0xa5/0x150 [nouveau]
nvkm_ioctl+0xcc/0x1d0 [nouveau]
nvif_object_mthd+0x110/0x1e0 [nouveau]
nvif_device_gcx_ready+0x36/0x60 [nouveau]
nouveau_pmops_runtime_suspend+0x39/0x140 [nouveau]
Instruction decode shows `mov rax, [rdi+0xe70]` loads gsp->rm (0x0),
and `<48> 8b 40 18` (`mov rax, [rax+0x18]`) faults accessing rm->api.
Fix by:
0. Checking !nvkm_gsp_rm(gsp) in nvkm_udevice_gcx_ready() so non-RM
devices report GC6/GcOff ready and bypass the GSP call entirely.
1. Guarding gsp and gsp->rm in nvkm_gsp_gcx_ready() before inspecting
gsp->rm->api->gsp->gcx_ready.
Fixes: cb4c7603678c ("drm/nouveau/gsp/r570: Add support for INTERNAL_GCX_ENTRY_PREREQUISITE")
Signed-off-by: Jim Cromie <jim.cromie@gmail.com>
---
drivers/gpu/drm/nouveau/include/nvkm/subdev/gsp.h | 2 +-
drivers/gpu/drm/nouveau/nvkm/engine/device/user.c | 2 +-
drivers/gpu/drm/nouveau/nvkm/subdev/gsp/base.c | 2 +-
3 files changed, 3 insertions(+), 3 deletions(-)
diff --git a/drivers/gpu/drm/nouveau/include/nvkm/subdev/gsp.h b/drivers/gpu/drm/nouveau/include/nvkm/subdev/gsp.h
index ed5c6e0e68d3..933f13ed2a50 100644
--- a/drivers/gpu/drm/nouveau/include/nvkm/subdev/gsp.h
+++ b/drivers/gpu/drm/nouveau/include/nvkm/subdev/gsp.h
@@ -274,7 +274,7 @@ struct nvkm_gsp {
static inline bool
nvkm_gsp_rm(struct nvkm_gsp *gsp)
{
- return gsp && (gsp->fws.rm || gsp->fw.img);
+ return gsp && gsp->rm && (gsp->fws.rm || gsp->fw.img);
}
#include <rm/rm.h>
diff --git a/drivers/gpu/drm/nouveau/nvkm/engine/device/user.c b/drivers/gpu/drm/nouveau/nvkm/engine/device/user.c
index f78e6b9b4292..d469437eb991 100644
--- a/drivers/gpu/drm/nouveau/nvkm/engine/device/user.c
+++ b/drivers/gpu/drm/nouveau/nvkm/engine/device/user.c
@@ -201,7 +201,7 @@ nvkm_udevice_gcx_ready(struct nvkm_udevice *udev, void *data, u32 size)
} *args = data;
int ret = -ENOSYS;
- if (!gsp) {
+ if (!nvkm_gsp_rm(gsp)) {
args->v0.ready = NV_DEVICE_GC6_READY | NV_DEVICE_GCOFF_READY;
return 0;
}
diff --git a/drivers/gpu/drm/nouveau/nvkm/subdev/gsp/base.c b/drivers/gpu/drm/nouveau/nvkm/subdev/gsp/base.c
index e475d0e8fa7b..ea895f04293a 100644
--- a/drivers/gpu/drm/nouveau/nvkm/subdev/gsp/base.c
+++ b/drivers/gpu/drm/nouveau/nvkm/subdev/gsp/base.c
@@ -51,7 +51,7 @@ nvkm_gsp_intr_stall(struct nvkm_gsp *gsp, enum nvkm_subdev_type type, int inst)
int
nvkm_gsp_gcx_ready(struct nvkm_gsp *gsp)
{
- if (!gsp->rm->api->gsp->gcx_ready)
+ if (!nvkm_gsp_rm(gsp) || !gsp->rm->api->gsp->gcx_ready)
return NV_DEVICE_GC6_READY | NV_DEVICE_GCOFF_READY;
return gsp->rm->api->gsp->gcx_ready(gsp);
--
2.56.0
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH v2 1/3] nouveau: Fix NULL pointer dereferences in GETPARAM ioctl
2026-10-09 17:44 ` [PATCH v2 1/3] nouveau: Fix NULL pointer dereferences in GETPARAM ioctl Jim Cromie via B4 Relay
@ 2026-10-09 21:24 ` lyude
2026-10-09 21:47 ` Danilo Krummrich
1 sibling, 0 replies; 11+ messages in thread
From: lyude @ 2026-10-09 21:24 UTC (permalink / raw)
To: jim.cromie, Danilo Krummrich, Maarten Lankhorst, Maxime Ripard,
Thomas Zimmermann, David Airlie, Simona Vetter, Dave Airlie
Cc: dri-devel, nouveau, linux-kernel, Daniel Campos Ramos
On Fri, 2026-10-09 at 11:44 -0600, Jim Cromie via B4 Relay wrote:
> From: Jim Cromie <jim.cromie@gmail.com>
>
> When hardware or firmware initialization fails, the graphics engine
> (gr) or device functions may remain NULL. Attempting to access these
> during the GETPARAM ioctl (e.g., NOUVEAU_GETPARAM_GRAPH_UNITS)
> results
> in a kernel NULL pointer dereference, causing a crash when userspace
> (GNOME/Mesa) attempts to probe the device.
>
> Add safety checks for 'gr', 'gr->func', and 'nvkm_device->func' in
> the
> ioctl handler. Return -ENODEV to signal the missing hardware state to
> userspace, and use NV_ERROR_ONCE to provide diagnostic proof in the
> kernel log without risking a console flood.
>
> NOTE:
>
> I encountered these crashes while booting unrelated code (in
> dynamic-debug) on real HW. There was also some fwupd updates around
> this time, so it could all be transient/unreproducible. But it did
> happen, and could happen again in unforseen circumstances.
>
> @Capitain_Jack tested the error path in v1 by deleting the FW
I can just fix this up before pushing but JFYI - if you put comments
below the --- that's down below, they will show up in the email but not
the commit message.
Either way -
Reviewed-by: Lyude Paul <lyude@redhat.com>
>
> Signed-off-by: Jim Cromie <jim.cromie@gmail.com>
> Tested-by: Daniel Campos Ramos <Capitain_Jack@yahoo.com>
> ---
> drivers/gpu/drm/nouveau/nouveau_abi16.c | 22 ++++++++++++++++++++--
> 1 file changed, 20 insertions(+), 2 deletions(-)
>
> diff --git a/drivers/gpu/drm/nouveau/nouveau_abi16.c
> b/drivers/gpu/drm/nouveau/nouveau_abi16.c
> index 4542d5f4ded8..3f130cd4fbcd 100644
> --- a/drivers/gpu/drm/nouveau/nouveau_abi16.c
> +++ b/drivers/gpu/drm/nouveau/nouveau_abi16.c
> @@ -306,6 +306,11 @@ nouveau_abi16_ioctl_getparam(ABI16_IOCTL_ARGS)
> getparam->value = 1;
> break;
> case NOUVEAU_GETPARAM_GRAPH_UNITS:
> + if (!gr || !gr->func) {
> + NV_ERROR_ONCE(drm, "GETPARAM_GRAPH_UNITS: no
> gr engine or func\n");
> + return -ENODEV;
> + }
> +
> getparam->value = nvkm_gr_units(gr);
> break;
> case NOUVEAU_GETPARAM_EXEC_PUSH_MAX: {
> @@ -315,10 +320,23 @@ nouveau_abi16_ioctl_getparam(ABI16_IOCTL_ARGS)
> break;
> }
> case NOUVEAU_GETPARAM_VRAM_BAR_SIZE:
> - getparam->value = nvkm_device->func-
> >resource_size(nvkm_device, NVKM_BAR1_FB);
> + if (!nvkm_device || !nvkm_device->func ||
> !nvkm_device->func->resource_size) {
> + NV_ERROR_ONCE(drm, "GETPARAM_VRAM_BAR_SIZE:
> no device func\n");
> + return -ENODEV;
> + }
> +
> + getparam->value =
> + nvkm_device->func-
> >resource_size(nvkm_device, NVKM_BAR1_FB);
> break;
> case NOUVEAU_GETPARAM_VRAM_USED: {
> - struct ttm_resource_manager *vram_mgr =
> ttm_manager_type(&drm->ttm.bdev, TTM_PL_VRAM);
> + struct ttm_resource_manager *vram_mgr =
> + ttm_manager_type(&drm->ttm.bdev,
> TTM_PL_VRAM);
> +
> + if (!vram_mgr) {
> + NV_ERROR_ONCE(drm, "GETPARAM_VRAM_USED: no
> vram mgr\n");
> + return -ENODEV;
> + }
> +
> getparam->value =
> (u64)ttm_resource_manager_usage(vram_mgr);
> break;
> }
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH v2 2/3] nouveau: Fix NULL pointer dereference in GET_ZCULL_INFO ioctl
2026-10-09 17:44 ` [PATCH v2 2/3] nouveau: Fix NULL pointer dereference in GET_ZCULL_INFO ioctl Jim Cromie via B4 Relay
@ 2026-10-09 21:25 ` lyude
2026-10-09 21:48 ` Danilo Krummrich
1 sibling, 0 replies; 11+ messages in thread
From: lyude @ 2026-10-09 21:25 UTC (permalink / raw)
To: jim.cromie, Danilo Krummrich, Maarten Lankhorst, Maxime Ripard,
Thomas Zimmermann, David Airlie, Simona Vetter, Dave Airlie
Cc: dri-devel, nouveau, linux-kernel, Daniel Campos Ramos
Just for patchwork's sake, since the tags got stuck below the ---
Reviewed-by: Lyude Paul <lyude@redhat.com>
On Fri, 2026-10-09 at 11:44 -0600, Jim Cromie via B4 Relay wrote:
> From: Jim Cromie <jim.cromie@gmail.com>
>
> When graphics engine firmware fails to load or initialization aborts
> early, nvxx_gr(drm) returns NULL. Calling
> DRM_IOCTL_NOUVEAU_GET_ZCULL_INFO
> causes nouveau_abi16_ioctl_get_zcull_info() to dereference gr at
> offset
> 0xf0 without checking for NULL, triggering a kernel page fault.
>
> Validate that gr is non-NULL before inspecting gr->has_zcull_info.
>
> Signed-off-by: Jim Cromie <jim.cromie@gmail.com>
> ---
> Aug 14 09:21:52 frodo kernel: BUG: kernel NULL pointer dereference,
> address: 00000000000000f0
> Aug 14 09:21:52 frodo kernel: #PF: error_code(0x0000) - not-present
> page
> Aug 14 09:21:52 frodo kernel: Oops: Oops: 0000 [#1] SMP NOPTI
> Aug 14 09:21:52 frodo kernel: RIP:
> 0010:nouveau_abi16_ioctl_get_zcull_info+0x17/0xa0 [nouveau]
> Aug 14 09:21:52 frodo kernel: Code: 00 00 00 90 90 90 90 90 90 90 90
> 90 90 90 90 90 90 90 90 f3 0f 1e fa 0f 1f 44 00 00 48 8b 47 40 48 8b
> 00 48 8b 80 70 02 00 00 <80> b8 f0 00 00 00 00 74 72 8b 90 c0 00 00
> 00 89 16 8b 90 c4 00 00
> Aug 14 09:21:52 frodo kernel: Call Trace:
> Aug 14 09:21:52 frodo kernel: <TASK>
> Aug 14 09:21:52 frodo kernel: drm_ioctl_kernel+0xae/0x100
> Aug 14 09:21:52 frodo kernel: drm_ioctl+0x2e0/0x560
> Aug 14 09:21:52 frodo kernel: ?
> __pfx_nouveau_abi16_ioctl_get_zcull_info+0x10/0x10 [nouveau]
> Aug 14 09:21:52 frodo kernel: nouveau_drm_ioctl+0x58/0xc0 [nouveau]
> Aug 14 09:21:52 frodo kernel: __x64_sys_ioctl+0xb9/0x100
> Aug 14 09:21:52 frodo kernel: do_syscall_64+0xe2/0x560
> Aug 14 09:21:52 frodo kernel:
> entry_SYSCALL_64_after_hwframe+0x76/0x7e
> Aug 14 09:21:52 frodo kernel: </TASK>
>
> Reviewed-by: Lyude Paul <lyude@redhat.com>
> Tested-by: Daniel Campos Ramos <Capitain_Jack@yahoo.com>
> ---
> drivers/gpu/drm/nouveau/nouveau_abi16.c | 2 +-
> 1 file changed, 1 insertion(+), 1 deletion(-)
>
> diff --git a/drivers/gpu/drm/nouveau/nouveau_abi16.c
> b/drivers/gpu/drm/nouveau/nouveau_abi16.c
> index 3f130cd4fbcd..c7ddc54a3c4d 100644
> --- a/drivers/gpu/drm/nouveau/nouveau_abi16.c
> +++ b/drivers/gpu/drm/nouveau/nouveau_abi16.c
> @@ -358,7 +358,7 @@
> nouveau_abi16_ioctl_get_zcull_info(ABI16_IOCTL_ARGS)
> struct nvkm_gr *gr = nvxx_gr(drm);
> struct drm_nouveau_get_zcull_info *out = data;
>
> - if (gr->has_zcull_info) {
> + if (gr && gr->has_zcull_info) {
> const struct nvkm_gr_zcull_info *i = &gr-
> >zcull_info;
>
> out->width_align_pixels = i->width_align_pixels;
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH v2 3/3] nouveau: Fix NULL dereference in nvkm_gsp_gcx_ready()
2026-10-09 17:44 ` [PATCH v2 3/3] nouveau: Fix NULL dereference in nvkm_gsp_gcx_ready() Jim Cromie via B4 Relay
@ 2026-10-09 21:26 ` lyude
2026-10-09 21:53 ` Danilo Krummrich
1 sibling, 0 replies; 11+ messages in thread
From: lyude @ 2026-10-09 21:26 UTC (permalink / raw)
To: jim.cromie, Danilo Krummrich, Maarten Lankhorst, Maxime Ripard,
Thomas Zimmermann, David Airlie, Simona Vetter, Dave Airlie
Cc: dri-devel, nouveau, linux-kernel, Daniel Campos Ramos
Reviewed-by: Lyude Paul <lyude@redhat.com>
On Fri, 2026-10-09 at 11:44 -0600, Jim Cromie via B4 Relay wrote:
> From: Jim Cromie <jim.cromie@gmail.com>
>
> Commit cb4c7603678c ("drm/nouveau/gsp/r570: Add support for
> INTERNAL_GCX_ENTRY_PREREQUISITE") hooked nvif_device_gcx_ready() into
> nouveau_pmops_runtime_suspend() to consult GSP before runtime
> suspend.
>
> On GPUs running without GSP-RM (e.g. Volta GV100, NvGspRm=0, or
> missing
> GSP firmware blobs), device->gsp is instantiated but gsp->rm remains
> NULL. nvkm_udevice_gcx_ready() checked "!gsp" instead of verifying
> nvkm_gsp_rm(gsp), allowing non-RM instances to invoke
> nvkm_gsp_gcx_ready(). nvkm_gsp_gcx_ready() then unconditionally
> dereferenced gsp->rm->api, causing a kernel oops:
>
> RIP: 0010:nvkm_gsp_gcx_ready+0x10/0x30 [nouveau]
> Code: ... 48 8b 87 70 0e 00 00 <48> 8b 40 18 ...
> RAX: 0000000000000000 RDI: ffff8c9555fca000
> Call Trace:
> nvkm_udevice_mthd+0xa5/0x150 [nouveau]
> nvkm_ioctl+0xcc/0x1d0 [nouveau]
> nvif_object_mthd+0x110/0x1e0 [nouveau]
> nvif_device_gcx_ready+0x36/0x60 [nouveau]
> nouveau_pmops_runtime_suspend+0x39/0x140 [nouveau]
>
> Instruction decode shows `mov rax, [rdi+0xe70]` loads gsp->rm (0x0),
> and `<48> 8b 40 18` (`mov rax, [rax+0x18]`) faults accessing rm->api.
>
> Fix by:
> 0. Checking !nvkm_gsp_rm(gsp) in nvkm_udevice_gcx_ready() so non-RM
> devices report GC6/GcOff ready and bypass the GSP call entirely.
> 1. Guarding gsp and gsp->rm in nvkm_gsp_gcx_ready() before inspecting
> gsp->rm->api->gsp->gcx_ready.
>
> Fixes: cb4c7603678c ("drm/nouveau/gsp/r570: Add support for
> INTERNAL_GCX_ENTRY_PREREQUISITE")
> Signed-off-by: Jim Cromie <jim.cromie@gmail.com>
> ---
> drivers/gpu/drm/nouveau/include/nvkm/subdev/gsp.h | 2 +-
> drivers/gpu/drm/nouveau/nvkm/engine/device/user.c | 2 +-
> drivers/gpu/drm/nouveau/nvkm/subdev/gsp/base.c | 2 +-
> 3 files changed, 3 insertions(+), 3 deletions(-)
>
> diff --git a/drivers/gpu/drm/nouveau/include/nvkm/subdev/gsp.h
> b/drivers/gpu/drm/nouveau/include/nvkm/subdev/gsp.h
> index ed5c6e0e68d3..933f13ed2a50 100644
> --- a/drivers/gpu/drm/nouveau/include/nvkm/subdev/gsp.h
> +++ b/drivers/gpu/drm/nouveau/include/nvkm/subdev/gsp.h
> @@ -274,7 +274,7 @@ struct nvkm_gsp {
> static inline bool
> nvkm_gsp_rm(struct nvkm_gsp *gsp)
> {
> - return gsp && (gsp->fws.rm || gsp->fw.img);
> + return gsp && gsp->rm && (gsp->fws.rm || gsp->fw.img);
> }
>
> #include <rm/rm.h>
> diff --git a/drivers/gpu/drm/nouveau/nvkm/engine/device/user.c
> b/drivers/gpu/drm/nouveau/nvkm/engine/device/user.c
> index f78e6b9b4292..d469437eb991 100644
> --- a/drivers/gpu/drm/nouveau/nvkm/engine/device/user.c
> +++ b/drivers/gpu/drm/nouveau/nvkm/engine/device/user.c
> @@ -201,7 +201,7 @@ nvkm_udevice_gcx_ready(struct nvkm_udevice *udev,
> void *data, u32 size)
> } *args = data;
> int ret = -ENOSYS;
>
> - if (!gsp) {
> + if (!nvkm_gsp_rm(gsp)) {
> args->v0.ready = NV_DEVICE_GC6_READY |
> NV_DEVICE_GCOFF_READY;
> return 0;
> }
> diff --git a/drivers/gpu/drm/nouveau/nvkm/subdev/gsp/base.c
> b/drivers/gpu/drm/nouveau/nvkm/subdev/gsp/base.c
> index e475d0e8fa7b..ea895f04293a 100644
> --- a/drivers/gpu/drm/nouveau/nvkm/subdev/gsp/base.c
> +++ b/drivers/gpu/drm/nouveau/nvkm/subdev/gsp/base.c
> @@ -51,7 +51,7 @@ nvkm_gsp_intr_stall(struct nvkm_gsp *gsp, enum
> nvkm_subdev_type type, int inst)
> int
> nvkm_gsp_gcx_ready(struct nvkm_gsp *gsp)
> {
> - if (!gsp->rm->api->gsp->gcx_ready)
> + if (!nvkm_gsp_rm(gsp) || !gsp->rm->api->gsp->gcx_ready)
> return NV_DEVICE_GC6_READY | NV_DEVICE_GCOFF_READY;
>
> return gsp->rm->api->gsp->gcx_ready(gsp);
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH v2 1/3] nouveau: Fix NULL pointer dereferences in GETPARAM ioctl
2026-10-09 17:44 ` [PATCH v2 1/3] nouveau: Fix NULL pointer dereferences in GETPARAM ioctl Jim Cromie via B4 Relay
2026-10-09 21:24 ` lyude
@ 2026-10-09 21:47 ` Danilo Krummrich
1 sibling, 0 replies; 11+ messages in thread
From: Danilo Krummrich @ 2026-10-09 21:47 UTC (permalink / raw)
To: Jim Cromie via B4 Relay
Cc: jim.cromie, Lyude Paul, Maarten Lankhorst, Maxime Ripard,
Thomas Zimmermann, David Airlie, Simona Vetter, Dave Airlie,
dri-devel, nouveau, linux-kernel, Daniel Campos Ramos
On Fri Oct 9, 2026 at 7:44 PM CEST, Jim Cromie via B4 Relay wrote:
> From: Jim Cromie <jim.cromie@gmail.com>
>
> When hardware or firmware initialization fails, the graphics engine
> (gr) or device functions may remain NULL. Attempting to access these
> during the GETPARAM ioctl (e.g., NOUVEAU_GETPARAM_GRAPH_UNITS) results
> in a kernel NULL pointer dereference, causing a crash when userspace
> (GNOME/Mesa) attempts to probe the device.
>
> Add safety checks for 'gr', 'gr->func', and 'nvkm_device->func' in the
> ioctl handler. Return -ENODEV to signal the missing hardware state to
> userspace, and use NV_ERROR_ONCE to provide diagnostic proof in the
> kernel log without risking a console flood.
>
> NOTE:
>
> I encountered these crashes while booting unrelated code (in
> dynamic-debug) on real HW. There was also some fwupd updates around
> this time, so it could all be transient/unreproducible. But it did
> happen, and could happen again in unforseen circumstances.
>
> @Capitain_Jack tested the error path in v1 by deleting the FW
>
> Signed-off-by: Jim Cromie <jim.cromie@gmail.com>
> Tested-by: Daniel Campos Ramos <Capitain_Jack@yahoo.com>
Do they need to be separate patches? Do we need a Fixes: tag? Why not CC stable?
> ---
> drivers/gpu/drm/nouveau/nouveau_abi16.c | 22 ++++++++++++++++++++--
> 1 file changed, 20 insertions(+), 2 deletions(-)
>
> diff --git a/drivers/gpu/drm/nouveau/nouveau_abi16.c b/drivers/gpu/drm/nouveau/nouveau_abi16.c
> index 4542d5f4ded8..3f130cd4fbcd 100644
> --- a/drivers/gpu/drm/nouveau/nouveau_abi16.c
> +++ b/drivers/gpu/drm/nouveau/nouveau_abi16.c
> @@ -306,6 +306,11 @@ nouveau_abi16_ioctl_getparam(ABI16_IOCTL_ARGS)
> getparam->value = 1;
> break;
> case NOUVEAU_GETPARAM_GRAPH_UNITS:
> + if (!gr || !gr->func) {
> + NV_ERROR_ONCE(drm, "GETPARAM_GRAPH_UNITS: no gr engine or func\n");
> + return -ENODEV;
> + }
> +
> getparam->value = nvkm_gr_units(gr);
> break;
> case NOUVEAU_GETPARAM_EXEC_PUSH_MAX: {
> @@ -315,10 +320,23 @@ nouveau_abi16_ioctl_getparam(ABI16_IOCTL_ARGS)
> break;
> }
> case NOUVEAU_GETPARAM_VRAM_BAR_SIZE:
> - getparam->value = nvkm_device->func->resource_size(nvkm_device, NVKM_BAR1_FB);
> + if (!nvkm_device || !nvkm_device->func || !nvkm_device->func->resource_size) {
How can any of those ever be NULL?
> + NV_ERROR_ONCE(drm, "GETPARAM_VRAM_BAR_SIZE: no device func\n");
> + return -ENODEV;
> + }
> +
> + getparam->value =
> + nvkm_device->func->resource_size(nvkm_device, NVKM_BAR1_FB);
> break;
> case NOUVEAU_GETPARAM_VRAM_USED: {
> - struct ttm_resource_manager *vram_mgr = ttm_manager_type(&drm->ttm.bdev, TTM_PL_VRAM);
> + struct ttm_resource_manager *vram_mgr =
> + ttm_manager_type(&drm->ttm.bdev, TTM_PL_VRAM);
> +
> + if (!vram_mgr) {
Same here, how can this ever be NULL?
> + NV_ERROR_ONCE(drm, "GETPARAM_VRAM_USED: no vram mgr\n");
> + return -ENODEV;
> + }
> +
> getparam->value = (u64)ttm_resource_manager_usage(vram_mgr);
> break;
> }
>
> --
> 2.56.0
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH v2 2/3] nouveau: Fix NULL pointer dereference in GET_ZCULL_INFO ioctl
2026-10-09 17:44 ` [PATCH v2 2/3] nouveau: Fix NULL pointer dereference in GET_ZCULL_INFO ioctl Jim Cromie via B4 Relay
2026-10-09 21:25 ` lyude
@ 2026-10-09 21:48 ` Danilo Krummrich
1 sibling, 0 replies; 11+ messages in thread
From: Danilo Krummrich @ 2026-10-09 21:48 UTC (permalink / raw)
To: Jim Cromie via B4 Relay
Cc: jim.cromie, Lyude Paul, Maarten Lankhorst, Maxime Ripard,
Thomas Zimmermann, David Airlie, Simona Vetter, Dave Airlie,
dri-devel, nouveau, linux-kernel, Daniel Campos Ramos
On Fri Oct 9, 2026 at 7:44 PM CEST, Jim Cromie via B4 Relay wrote:
> From: Jim Cromie <jim.cromie@gmail.com>
>
> When graphics engine firmware fails to load or initialization aborts
> early, nvxx_gr(drm) returns NULL. Calling DRM_IOCTL_NOUVEAU_GET_ZCULL_INFO
> causes nouveau_abi16_ioctl_get_zcull_info() to dereference gr at offset
> 0xf0 without checking for NULL, triggering a kernel page fault.
>
> Validate that gr is non-NULL before inspecting gr->has_zcull_info.
>
> Signed-off-by: Jim Cromie <jim.cromie@gmail.com>
Please add a Fixes: tag and CC stable.
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH v2 3/3] nouveau: Fix NULL dereference in nvkm_gsp_gcx_ready()
2026-10-09 17:44 ` [PATCH v2 3/3] nouveau: Fix NULL dereference in nvkm_gsp_gcx_ready() Jim Cromie via B4 Relay
2026-10-09 21:26 ` lyude
@ 2026-10-09 21:53 ` Danilo Krummrich
2026-10-09 22:01 ` lyude
1 sibling, 1 reply; 11+ messages in thread
From: Danilo Krummrich @ 2026-10-09 21:53 UTC (permalink / raw)
To: Jim Cromie via B4 Relay
Cc: jim.cromie, Lyude Paul, Maarten Lankhorst, Maxime Ripard,
Thomas Zimmermann, David Airlie, Simona Vetter, Dave Airlie,
dri-devel, nouveau, linux-kernel, Daniel Campos Ramos
On Fri Oct 9, 2026 at 7:44 PM CEST, Jim Cromie via B4 Relay wrote:
> From: Jim Cromie <jim.cromie@gmail.com>
>
> Commit cb4c7603678c ("drm/nouveau/gsp/r570: Add support for
> INTERNAL_GCX_ENTRY_PREREQUISITE") hooked nvif_device_gcx_ready() into
> nouveau_pmops_runtime_suspend() to consult GSP before runtime suspend.
>
> On GPUs running without GSP-RM (e.g. Volta GV100, NvGspRm=0, or missing
> GSP firmware blobs), device->gsp is instantiated but gsp->rm remains
> NULL. nvkm_udevice_gcx_ready() checked "!gsp" instead of verifying
> nvkm_gsp_rm(gsp), allowing non-RM instances to invoke
> nvkm_gsp_gcx_ready(). nvkm_gsp_gcx_ready() then unconditionally
> dereferenced gsp->rm->api, causing a kernel oops:
>
> RIP: 0010:nvkm_gsp_gcx_ready+0x10/0x30 [nouveau]
> Code: ... 48 8b 87 70 0e 00 00 <48> 8b 40 18 ...
> RAX: 0000000000000000 RDI: ffff8c9555fca000
> Call Trace:
> nvkm_udevice_mthd+0xa5/0x150 [nouveau]
> nvkm_ioctl+0xcc/0x1d0 [nouveau]
> nvif_object_mthd+0x110/0x1e0 [nouveau]
> nvif_device_gcx_ready+0x36/0x60 [nouveau]
> nouveau_pmops_runtime_suspend+0x39/0x140 [nouveau]
>
> Instruction decode shows `mov rax, [rdi+0xe70]` loads gsp->rm (0x0),
> and `<48> 8b 40 18` (`mov rax, [rax+0x18]`) faults accessing rm->api.
>
> Fix by:
> 0. Checking !nvkm_gsp_rm(gsp) in nvkm_udevice_gcx_ready() so non-RM
> devices report GC6/GcOff ready and bypass the GSP call entirely.
> 1. Guarding gsp and gsp->rm in nvkm_gsp_gcx_ready() before inspecting
> gsp->rm->api->gsp->gcx_ready.
>
> Fixes: cb4c7603678c ("drm/nouveau/gsp/r570: Add support for INTERNAL_GCX_ENTRY_PREREQUISITE")
> Signed-off-by: Jim Cromie <jim.cromie@gmail.com>
Isn't this fixed in commit 4d0e27041185 ("nouveau: check gsp->rm as well as gsp
pointer before gcx ready") already?
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH v2 3/3] nouveau: Fix NULL dereference in nvkm_gsp_gcx_ready()
2026-10-09 21:53 ` Danilo Krummrich
@ 2026-10-09 22:01 ` lyude
0 siblings, 0 replies; 11+ messages in thread
From: lyude @ 2026-10-09 22:01 UTC (permalink / raw)
To: Danilo Krummrich, Jim Cromie via B4 Relay
Cc: jim.cromie, Maarten Lankhorst, Maxime Ripard, Thomas Zimmermann,
David Airlie, Simona Vetter, Dave Airlie, dri-devel, nouveau,
linux-kernel, Daniel Campos Ramos
On Fri, 2026-10-09 at 23:53 +0200, Danilo Krummrich wrote:
> On Fri Oct 9, 2026 at 7:44 PM CEST, Jim Cromie via B4 Relay wrote:
> > From: Jim Cromie <jim.cromie@gmail.com>
> >
> > Commit cb4c7603678c ("drm/nouveau/gsp/r570: Add support for
> > INTERNAL_GCX_ENTRY_PREREQUISITE") hooked nvif_device_gcx_ready()
> > into
> > nouveau_pmops_runtime_suspend() to consult GSP before runtime
> > suspend.
> >
> > On GPUs running without GSP-RM (e.g. Volta GV100, NvGspRm=0, or
> > missing
> > GSP firmware blobs), device->gsp is instantiated but gsp->rm
> > remains
> > NULL. nvkm_udevice_gcx_ready() checked "!gsp" instead of verifying
> > nvkm_gsp_rm(gsp), allowing non-RM instances to invoke
> > nvkm_gsp_gcx_ready(). nvkm_gsp_gcx_ready() then unconditionally
> > dereferenced gsp->rm->api, causing a kernel oops:
> >
> > RIP: 0010:nvkm_gsp_gcx_ready+0x10/0x30 [nouveau]
> > Code: ... 48 8b 87 70 0e 00 00 <48> 8b 40 18 ...
> > RAX: 0000000000000000 RDI: ffff8c9555fca000
> > Call Trace:
> > nvkm_udevice_mthd+0xa5/0x150 [nouveau]
> > nvkm_ioctl+0xcc/0x1d0 [nouveau]
> > nvif_object_mthd+0x110/0x1e0 [nouveau]
> > nvif_device_gcx_ready+0x36/0x60 [nouveau]
> > nouveau_pmops_runtime_suspend+0x39/0x140 [nouveau]
> >
> > Instruction decode shows `mov rax, [rdi+0xe70]` loads gsp->rm
> > (0x0),
> > and `<48> 8b 40 18` (`mov rax, [rax+0x18]`) faults accessing rm-
> > >api.
> >
> > Fix by:
> > 0. Checking !nvkm_gsp_rm(gsp) in nvkm_udevice_gcx_ready() so non-RM
> > devices report GC6/GcOff ready and bypass the GSP call entirely.
> > 1. Guarding gsp and gsp->rm in nvkm_gsp_gcx_ready() before
> > inspecting
> > gsp->rm->api->gsp->gcx_ready.
> >
> > Fixes: cb4c7603678c ("drm/nouveau/gsp/r570: Add support for
> > INTERNAL_GCX_ENTRY_PREREQUISITE")
> > Signed-off-by: Jim Cromie <jim.cromie@gmail.com>
>
> Isn't this fixed in commit 4d0e27041185 ("nouveau: check gsp->rm as
> well as gsp
> pointer before gcx ready") already?
(gonna hold off on pushing all 3 of these patches, you're right - I
completely forgot that we had already pushed a commit for this)
^ permalink raw reply [flat|nested] 11+ messages in thread
end of thread, other threads:[~2026-10-09 22:01 UTC | newest]
Thread overview: 11+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-09 17:44 [PATCH v2 0/3] nouveau: fix 3 null-ptr derefs Jim Cromie via B4 Relay
2026-10-09 17:44 ` [PATCH v2 1/3] nouveau: Fix NULL pointer dereferences in GETPARAM ioctl Jim Cromie via B4 Relay
2026-10-09 21:24 ` lyude
2026-10-09 21:47 ` Danilo Krummrich
2026-10-09 17:44 ` [PATCH v2 2/3] nouveau: Fix NULL pointer dereference in GET_ZCULL_INFO ioctl Jim Cromie via B4 Relay
2026-10-09 21:25 ` lyude
2026-10-09 21:48 ` Danilo Krummrich
2026-10-09 17:44 ` [PATCH v2 3/3] nouveau: Fix NULL dereference in nvkm_gsp_gcx_ready() Jim Cromie via B4 Relay
2026-10-09 21:26 ` lyude
2026-10-09 21:53 ` Danilo Krummrich
2026-10-09 22:01 ` lyude
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®