* [PATCH 0/4] remoteproc: core: Fix the error unwind paths
@ 2026-09-29 7:54 Yonghao Zhang
2026-09-29 7:54 ` [PATCH 1/4] remoteproc: core: Reset freed vring entries to FW_RSC_ADDR_ANY Yonghao Zhang
` (3 more replies)
0 siblings, 4 replies; 5+ messages in thread
From: Yonghao Zhang @ 2026-09-29 7:54 UTC (permalink / raw)
To: andersson, mathieu.poirier; +Cc: linux-remoteproc, linux-kernel, Yonghao Zhang
The remoteproc core unwinds failed boots, failed attaches and failed
recoveries in a handful of places, and those paths have drifted over
the years: an unconditional stop() here, a missing cleanup there, a
reset value a running remote processor could misread. This series
fixes four of them. The success paths are untouched, and no in-tree
driver changes behaviour when nothing fails.
The single-sided ops configuration that two of the patches guard
against is not hypothetical. Version 9 of the TI k3 refactor series
briefly introduced it [1]: ti_k3_dsp registered in IPC-only mode
with start() kept and stop() cleared, on the reasoning that
rproc_attach() never invokes start(). The shape sat on the list for
a month before the rework that finally landed removed the ops
overrides altogether (commit 41d746b3423a ("remoteproc: k3-dsp: Don't
override rproc ops in IPC-only mode")). The point stands: nothing in
the core would have caught the intermediate version, and a driver
sent today with the same single-sided ops would pass rproc_validate()
just the same.
Patch 3 mostly benefits the attach-on-recovery users (imx_rproc when
the Mcore resource is owned by another partition, xlnx_r5 when the
RPU is already running firmware): since attach recovery was introduced
(commit ba194232edc0 ("remoteproc: Support attach recovery after rproc
crash")), a failed re-attach was unwound with an unconditional
ops->stop() -- powering off a processor which, per the feature's own
contract, "does not need help from Linux to recover... Linux just needs
to attach", and which zynqmp_r5_rproc_stop() would additionally flip out
of attach recovery mode by clearing the feature bit. With the unwind
detaching instead, the processor keeps running under its original owner
and the next attempt attaches again; attach()/detach() stay paired
one-to-one with the attach() calls that succeeded.
- 1/4: rproc_free_vring() writes FW_RSC_ADDR_ANY rather than 0 into
a freed vring entry, for the callers that reach it while the
remote processor is running and the reset lands in the installed
resource table.
- 2/4: rproc_start() no longer calls stop() unguarded when it
unrolls a failed subdevice registration.
- 3/4: __rproc_attach() rolls a failed attach back with detach()
when available, keeping the resource table accounting and the
state consistent, instead of unconditionally stopping a processor
that was started by another entity.
- 4/4: a failed recovery releases the resources of the boot it was
trying to restore and voids the power count, which nothing else
can drain once the processor is offline or detached.
Testing: on an Allwinner T153 (sun8iw22) with an E907 remote core,
driven by an out-of-tree platform driver for the Allwinner remoteproc
hardware.
The error paths of patches 1 and 3 were hit on this hardware rather
than constructed. During development, a subdevice registration
failure during rproc_attach(), unwound by an intermediate version
that called detach() without the resource table reset, left the
freed vring entries at da 0 in the installed table: the next
rproc_attach() failed in rproc_check_carveout_da(), called from
rproc_alloc_vring(), with "Registered carveout doesn't fit da request"
That trace motivated the reset bookkeeping of patch 3 and the
FW_RSC_ADDR_ANY reset value of patch 1; with the series applied, a
failed attach can simply be retried.
The failure path of patch 4 was exercised by renaming the firmware
file once the processor was up and then crashing the remote core:
recovery fails in request_firmware() as expected, and after the
file is restored the next rproc_boot() performs a real new boot
instead of silently returning success on the stale power count,
with no "already associated to resource table" leftovers from the
failed recovery.
[1] [PATCH v9 17/26] remoteproc: k3: Refactor .start rproc ops into
common driver, 2025-03-17:
https://lore.kernel.org/all/20250317120622.1746415-18-b-padhi@ti.com/
Yonghao Zhang (4):
remoteproc: core: Reset freed vring entries to FW_RSC_ADDR_ANY
remoteproc: core: Guard against a missing stop() in rproc_start()
remoteproc: core: Roll a failed attach back with detach() when
available
remoteproc: core: Clean up after a failed recovery
drivers/remoteproc/remoteproc_core.c | 144 +++++++++++++++++++++++----
1 file changed, 127 insertions(+), 17 deletions(-)
base-commit: 586a2fadadb160295a93e5d5d32db1a2558a4d3f
--
2.34.1
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH 1/4] remoteproc: core: Reset freed vring entries to FW_RSC_ADDR_ANY
2026-09-29 7:54 [PATCH 0/4] remoteproc: core: Fix the error unwind paths Yonghao Zhang
@ 2026-09-29 7:54 ` Yonghao Zhang
2026-09-29 7:54 ` [PATCH 2/4] remoteproc: core: Guard against a missing stop() in rproc_start() Yonghao Zhang
` (2 subsequent siblings)
3 siblings, 0 replies; 5+ messages in thread
From: Yonghao Zhang @ 2026-09-29 7:54 UTC (permalink / raw)
To: andersson, mathieu.poirier; +Cc: linux-remoteproc, linux-kernel, Yonghao Zhang
rproc_free_vring() resets a vring entry in the resource table as it
releases it. Resetting da to 0 leaves the entry looking like a vring
at device address 0; write FW_RSC_ADDR_ANY instead, the marker for an
unallocated address, which is what the entry means from then on.
That matters on the error paths of rproc_start() and rproc_attach(),
where rproc_free_vring() runs while the remote processor is already
running: the failure originates in rp_find_vq(), called through
virtio_find_vqs() from e.g. rpmsg_probe(), which the core only
reaches through rproc_start_subdevices(), after ops->start() or
ops->attach() have completed. At that point table_ptr is the
resource table installed in remote processor memory, and whatever
the reset writes is visible to the remote side. Whether and when a
running processor reads the entry cannot be known, so the value must
be safe for one: 0 is a plausible device address, FW_RSC_ADDR_ANY is
not. On the teardown paths, stop and detach, the write lands in the
cached table or in a copy of the installed one and reaches no one.
The comment replaced along with it was written for a call graph that
no longer exists. It described the teardown callers only:
rproc_stop() has run, table_ptr points at the cached table, and that
table is NULL for a processor started by another entity, so there is
nothing to clear. Two facts have overtaken it:
- rp_find_vq() has been calling rproc_free_vring() when
vring_new_virtqueue() fails since commit 6db20ea8d850 ("remoteproc:
allocate vrings on demand, free when not needed"), and that call
was present, unchanged, when commit 9dc9507f1880 ("remoteproc:
Properly deal with the resource table when detaching") landed: it
runs from rproc_start_subdevices(), after ops->start() or
ops->attach() have completed, with table_ptr at the table installed
in remote processor memory and the processor running. The comment
never described this caller, not on the day it was written.
- commit 9dc9507f1880 ("remoteproc: Properly deal with the resource
table when detaching") and commit 8088dd4d9316 ("remoteproc: Properly
deal with the resource table when stopping") later taught the
detach/stop paths to take a kmemdup() copy of the installed table for
processors started by another entity, so the NULL case the guard was
deciding between is gone: table_ptr is always valid where the teardown
callers run. (Those callers have since moved to the rproc-virtio
platform driver, 9c31255ce5fe.)
Fixes: c0d631570ad5 ("remoteproc: set vring addresses in resource table")
Fixes: 4d3ebb3b9990 ("remoteproc: Refactor function rproc_free_vring()")
Signed-off-by: Yonghao Zhang <hyz3367@gmail.com>
---
drivers/remoteproc/remoteproc_core.c | 40 +++++++++++++++++++---------
1 file changed, 28 insertions(+), 12 deletions(-)
diff --git a/drivers/remoteproc/remoteproc_core.c b/drivers/remoteproc/remoteproc_core.c
index 263e12f022ea..123aadb467a0 100644
--- a/drivers/remoteproc/remoteproc_core.c
+++ b/drivers/remoteproc/remoteproc_core.c
@@ -403,6 +403,33 @@ rproc_parse_vring(struct rproc_vdev *rvdev, struct fw_rsc_vdev *rsc, int i)
return 0;
}
+/*
+ * rproc_free_vring() runs while a remote processor is being torn down
+ * (rproc_stop()/rproc_detach() paths) or while a boot or attach attempt
+ * is failing. Whether the reset below can reach the remote processor
+ * depends on what rproc->table_ptr refers to at that point:
+ *
+ * - teardown paths: the call comes from the rvdev cleanup in
+ * rproc_resource_cleanup(), by which time rproc_stop()/__rproc_detach()
+ * have switched table_ptr to the table the core was booted with, or
+ * to a copy of the installed table when the remote processor was
+ * started by another entity (rproc_reset_rsc_table_on_{stop,detach}()),
+ * so the write cannot reach the remote processor.
+ *
+ * - error paths of rproc_start()/rproc_attach(): the failure originates
+ * in rp_find_vq() (virtio_find_vqs() failing in e.g. rpmsg_probe()),
+ * which rproc_start_subdevices() calls only once ops->start() or
+ * ops->attach() have completed. table_ptr still points at the
+ * resource table installed in remote processor memory and the remote
+ * processor is already running.
+ *
+ * Especially in the last case the reset has to tell the remote
+ * processor that the vring entry is not usable: with da set to the
+ * invalid FW_RSC_ADDR_ANY a running remote processor sees the address
+ * as unallocated, and notifyid is invalidated along with it.
+ *
+ * Reset the virtio device section only if there is a table to work with.
+ */
void rproc_free_vring(struct rproc_vring *rvring)
{
struct rproc *rproc = rvring->rvdev->rproc;
@@ -411,20 +438,9 @@ void rproc_free_vring(struct rproc_vring *rvring)
idr_remove(&rproc->notifyids, rvring->notifyid);
- /*
- * At this point rproc_stop() has been called and the installed resource
- * table in the remote processor memory may no longer be accessible. As
- * such and as per rproc_stop(), rproc->table_ptr points to the cached
- * resource table (rproc->cached_table). The cached resource table is
- * only available when a remote processor has been booted by the
- * remoteproc core, otherwise it is NULL.
- *
- * Based on the above, reset the virtio device section in the cached
- * resource table only if there is one to work with.
- */
if (rproc->table_ptr) {
rsc = (void *)rproc->table_ptr + rvring->rvdev->rsc_offset;
- rsc->vring[idx].da = 0;
+ rsc->vring[idx].da = FW_RSC_ADDR_ANY;
rsc->vring[idx].notifyid = -1;
}
}
--
2.34.1
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH 2/4] remoteproc: core: Guard against a missing stop() in rproc_start()
2026-09-29 7:54 [PATCH 0/4] remoteproc: core: Fix the error unwind paths Yonghao Zhang
2026-09-29 7:54 ` [PATCH 1/4] remoteproc: core: Reset freed vring entries to FW_RSC_ADDR_ANY Yonghao Zhang
@ 2026-09-29 7:54 ` Yonghao Zhang
2026-09-29 7:54 ` [PATCH 3/4] remoteproc: core: Roll a failed attach back with detach() when available Yonghao Zhang
2026-09-29 7:54 ` [PATCH 4/4] remoteproc: core: Clean up after a failed recovery Yonghao Zhang
3 siblings, 0 replies; 5+ messages in thread
From: Yonghao Zhang @ 2026-09-29 7:54 UTC (permalink / raw)
To: andersson, mathieu.poirier; +Cc: linux-remoteproc, linux-kernel, Yonghao Zhang
When the subdevice registration fails after a successful ops->start(),
rproc_start() unrolls with an unconditional ops->stop() call: an
implementation without stop() turns that error path into a NULL
dereference, and there is no other way to undo a start once the
processor is up.
Nothing rules that implementation out. rproc_validate() checks the
callbacks against the state a processor registers in: start() for
an offline one, attach() for a detached one, which never look at
stop(); the only written rule, Documentation/staging/remoteproc.rst
("Every remoteproc implementation should at least provide the ->start
and ->stop handlers"), is a should the core does not enforce. The
in-tree implementations all provide both, or neither when they only
attach (commit 1168af40b1ad ("remoteproc: k3-r5: Add support for
IPC-only mode for all R5Fs")), so none of them can reach the call
today.
Skip the rollback call when there is no stop(): with nothing to roll
the start back with, the processor stays running and the failure is
reported by the boot attempt itself.
Fixes: 7bdc9650f036 ("remoteproc: Introduce subdevices")
Signed-off-by: Yonghao Zhang <hyz3367@gmail.com>
---
drivers/remoteproc/remoteproc_core.c | 5 ++++-
1 file changed, 4 insertions(+), 1 deletion(-)
diff --git a/drivers/remoteproc/remoteproc_core.c b/drivers/remoteproc/remoteproc_core.c
index 123aadb467a0..19e0ea3e7240 100644
--- a/drivers/remoteproc/remoteproc_core.c
+++ b/drivers/remoteproc/remoteproc_core.c
@@ -1336,7 +1336,10 @@ static int rproc_start(struct rproc *rproc, const struct firmware *fw)
return 0;
stop_rproc:
- rproc->ops->stop(rproc);
+ if (rproc->ops->stop)
+ rproc->ops->stop(rproc);
+ else
+ dev_err(dev, "can't roll %s back: no stop()\n", rproc->name);
unprepare_subdevices:
rproc_unprepare_subdevices(rproc);
reset_table_ptr:
--
2.34.1
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH 3/4] remoteproc: core: Roll a failed attach back with detach() when available
2026-09-29 7:54 [PATCH 0/4] remoteproc: core: Fix the error unwind paths Yonghao Zhang
2026-09-29 7:54 ` [PATCH 1/4] remoteproc: core: Reset freed vring entries to FW_RSC_ADDR_ANY Yonghao Zhang
2026-09-29 7:54 ` [PATCH 2/4] remoteproc: core: Guard against a missing stop() in rproc_start() Yonghao Zhang
@ 2026-09-29 7:54 ` Yonghao Zhang
2026-09-29 7:54 ` [PATCH 4/4] remoteproc: core: Clean up after a failed recovery Yonghao Zhang
3 siblings, 0 replies; 5+ messages in thread
From: Yonghao Zhang @ 2026-09-29 7:54 UTC (permalink / raw)
To: andersson, mathieu.poirier; +Cc: linux-remoteproc, linux-kernel, Yonghao Zhang
When the subdevice registration that follows ops->attach() fails,
__rproc_attach() rolls back with an unconditional ops->stop() call.
That is wrong on three counts.
An attach-only implementation, one without stop() such as
commit 1168af40b1ad ("remoteproc: k3-r5: Add support for IPC-only
mode for all R5Fs"), dereferences NULL right there. A processor that
is being attached to was started by another entity and is not ours to
power off: detach() is the matching undo of attach(), and stop()
should only be used as a last resort, when there is no detach() or
it fails. This is reachable today: when the re-attach of an
RPROC_FEAT_ATTACH_ON_RECOVERY processor fails (imx_rproc and
xlnx_r5 use the feature), the unwind stops a processor which,
per the feature's own contract, "does not need help from Linux to
recover... Linux just needs to attach". And the unwind leaves
the accounting inconsistent -- no resource table bookkeeping
is done, unlike on the rproc_stop() and __rproc_detach() paths,
and a processor powered off through the fallback keeps its
RPROC_DETACHED state, so the next rproc_boot() tries to attach to a
core that is no longer running.
Roll the attach back with detach() first, along with the same
rproc_reset_rsc_table_on_detach() bookkeeping __rproc_detach() does,
and fall back to stop() only when detach() is unavailable or failed.
A failed resource table reset does not abort the unwind: this is an
error path, and detaching from the remote processor, or powering it
off as the last resort, takes precedence over the bookkeeping. A
successful fallback moves the processor to RPROC_OFFLINE so the next
boot reloads firmware instead of attaching to a dead core;
implementations with neither handler keep the processor running and
untouched, which is all an attach-only core needs. The rollback is
factored into rproc_unwind_attach().
The unwind runs the same resource table resets as the detach and
stop paths, which free clean_table and leave a cached copy of the
installed table in rproc->cached_table. Make rproc_attach()'s error
cleanup, which runs right after, null clean_table after freeing it
and release that copy along with table_ptr, or a failed attach
double-frees clean_table and leaks the copy.
Fixes: d848a4819d85 ("remoteproc: Introducing function rproc_attach()")
Signed-off-by: Yonghao Zhang <hyz3367@gmail.com>
---
drivers/remoteproc/remoteproc_core.c | 65 ++++++++++++++++++++++++++--
1 file changed, 62 insertions(+), 3 deletions(-)
diff --git a/drivers/remoteproc/remoteproc_core.c b/drivers/remoteproc/remoteproc_core.c
index 19e0ea3e7240..c52212a1d180 100644
--- a/drivers/remoteproc/remoteproc_core.c
+++ b/drivers/remoteproc/remoteproc_core.c
@@ -1348,6 +1348,59 @@ static int rproc_start(struct rproc *rproc, const struct firmware *fw)
return ret;
}
+static int rproc_reset_rsc_table_on_detach(struct rproc *rproc);
+static int rproc_reset_rsc_table_on_stop(struct rproc *rproc);
+
+/*
+ * Undo an attach whose subdevice registration failed. The remote
+ * processor was started by another entity and is not ours to power
+ * off, so roll back with detach() when available and fall back to
+ * stop() when there is no detach() or when it failed. stop() is
+ * the only rollback that does not rely on the remote side. A
+ * processor powered off through the fallback is marked RPROC_OFFLINE,
+ * so the next boot reloads firmware instead of attaching to a dead
+ * core; with neither handler there is nothing to roll back with and
+ * the processor is left running.
+ *
+ * The resource table resets are best-effort: when one fails, the
+ * unwind carries on with detach()/stop() anyway. Unlike
+ * __rproc_detach() and rproc_stop(), which bail out before touching
+ * the processor when their reset fails, this is already an error
+ * path, and detaching the remote processor -- or, failing that,
+ * powering it off -- is the minimum it must still deliver.
+ */
+static void rproc_unwind_attach(struct rproc *rproc)
+{
+ struct device *dev = &rproc->dev;
+ int ret;
+
+ if (rproc->ops->detach) {
+ ret = rproc_reset_rsc_table_on_detach(rproc);
+ if (ret)
+ dev_err(dev, "can't reset rsc table on detach: %d\n",
+ ret);
+
+ ret = rproc->ops->detach(rproc);
+ if (!ret)
+ return;
+
+ dev_err(dev, "can't detach from rproc %s: %d\n",
+ rproc->name, ret);
+ }
+
+ if (rproc->ops->stop) {
+ ret = rproc_reset_rsc_table_on_stop(rproc);
+ if (ret)
+ dev_err(dev, "can't reset rsc table on stop: %d\n",
+ ret);
+
+ if (rproc->ops->stop(rproc))
+ dev_err(dev, "can't stop rproc %s\n", rproc->name);
+ else
+ rproc->state = RPROC_OFFLINE;
+ }
+}
+
static int __rproc_attach(struct rproc *rproc)
{
struct device *dev = &rproc->dev;
@@ -1373,7 +1426,7 @@ static int __rproc_attach(struct rproc *rproc)
if (ret) {
dev_err(dev, "failed to probe subdevices for %s: %d\n",
rproc->name, ret);
- goto stop_rproc;
+ goto unwind_attach;
}
rproc->state = RPROC_ATTACHED;
@@ -1382,8 +1435,8 @@ static int __rproc_attach(struct rproc *rproc)
return 0;
-stop_rproc:
- rproc->ops->stop(rproc);
+unwind_attach:
+ rproc_unwind_attach(rproc);
unprepare_subdevices:
rproc_unprepare_subdevices(rproc);
out:
@@ -1562,6 +1615,7 @@ static int rproc_reset_rsc_table_on_detach(struct rproc *rproc)
* rproc_set_rsc_table().
*/
kfree(rproc->clean_table);
+ rproc->clean_table = NULL;
return 0;
}
@@ -1597,6 +1651,7 @@ static int rproc_reset_rsc_table_on_stop(struct rproc *rproc)
* won't be needed. Allocated in rproc_set_rsc_table().
*/
kfree(rproc->clean_table);
+ rproc->clean_table = NULL;
out:
/*
@@ -1675,6 +1730,10 @@ static int rproc_attach(struct rproc *rproc)
/* release HW resources if needed */
rproc_unprepare_device(rproc);
kfree(rproc->clean_table);
+ rproc->clean_table = NULL;
+ kfree(rproc->cached_table);
+ rproc->cached_table = NULL;
+ rproc->table_ptr = NULL;
disable_iommu:
rproc_disable_iommu(rproc);
return ret;
--
2.34.1
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH 4/4] remoteproc: core: Clean up after a failed recovery
2026-09-29 7:54 [PATCH 0/4] remoteproc: core: Fix the error unwind paths Yonghao Zhang
` (2 preceding siblings ...)
2026-09-29 7:54 ` [PATCH 3/4] remoteproc: core: Roll a failed attach back with detach() when available Yonghao Zhang
@ 2026-09-29 7:54 ` Yonghao Zhang
3 siblings, 0 replies; 5+ messages in thread
From: Yonghao Zhang @ 2026-09-29 7:54 UTC (permalink / raw)
To: andersson, mathieu.poirier; +Cc: linux-remoteproc, linux-kernel, Yonghao Zhang
When rproc_boot_recovery() fails to bring a crashed processor back
(the firmware request fails, or rproc_start() fails), it returns
with the processor stopped but two kinds of state still held.
The resources of the boot the recovery was trying to restore are
never released: rproc_stop() does not clean them up, and unlike
rproc_attach_recovery(), which releases everything when its
re-attach fails, boot_recovery just returned. Stale carveout
entries then fail every later firmware boot at "already associated
to resource table", until one of those failing boots happens to run
the cleanup of rproc_fw_boot().
The power references fare worse: nothing can release them anymore.
Take a processor with two outstanding rproc_boot() references whose
recovery stops it and then fails to restart it -- RPROC_OFFLINE,
with the count still at two. rproc_shutdown() drops exactly one
reference per call, and only once past its state gate; the gate
admits RPROC_RUNNING, RPROC_ATTACHED and RPROC_CRASHED, so an
offline processor never passes and no amount of shutdown() calls
releases anything. With the count still above zero, rproc_boot()
then short-circuits on atomic_inc_return(&rproc->power) > 1 and
returns success without doing anything: the users of a dead
processor are told it is running. The reference count and the
state machine are misaligned for good.
Release the resources on the failure paths of rproc_boot_recovery(),
the same way rproc_shutdown() does, and void the power count in
rproc_trigger_recovery() when the recovery failed without leaving
the processor crashed, offline or detached. The service the count
was tracking is gone, so every outstanding reference is dead, which
decrementing instead would leave the survivors stranded exactly as
above. A processor that is still crashed keeps its references, as
rproc_shutdown() can still drain them in that state.
Fixes: ba194232edc0 ("remoteproc: Support attach recovery after rproc crash")
Signed-off-by: Yonghao Zhang <hyz3367@gmail.com>
---
drivers/remoteproc/remoteproc_core.c | 34 +++++++++++++++++++++++++++-
1 file changed, 33 insertions(+), 1 deletion(-)
diff --git a/drivers/remoteproc/remoteproc_core.c b/drivers/remoteproc/remoteproc_core.c
index c52212a1d180..b138680b1905 100644
--- a/drivers/remoteproc/remoteproc_core.c
+++ b/drivers/remoteproc/remoteproc_core.c
@@ -1900,7 +1900,7 @@ static int rproc_boot_recovery(struct rproc *rproc)
ret = request_firmware(&firmware_p, rproc->firmware, dev);
if (ret < 0) {
dev_err(dev, "request_firmware failed: %d\n", ret);
- return ret;
+ goto clean_up_resources;
}
/* boot the remote processor up again */
@@ -1908,6 +1908,24 @@ static int rproc_boot_recovery(struct rproc *rproc)
release_firmware(firmware_p);
+ if (ret < 0)
+ goto clean_up_resources;
+
+ return 0;
+
+clean_up_resources:
+ /*
+ * rproc_stop() has already switched the remote processor off, but
+ * unlike rproc_shutdown() nothing releases the resources of the
+ * boot this recovery was trying to restore.
+ */
+ rproc_resource_cleanup(rproc);
+ kfree(rproc->cached_table);
+ rproc->cached_table = NULL;
+ rproc->table_ptr = NULL;
+ /* release HW resources if needed */
+ rproc_unprepare_device(rproc);
+ rproc_disable_iommu(rproc);
return ret;
}
@@ -1948,6 +1966,20 @@ int rproc_trigger_recovery(struct rproc *rproc)
else
ret = rproc_boot_recovery(rproc);
+ /*
+ * A failed recovery leaves the remote processor in a state from which
+ * rproc_shutdown() refuses to release the outstanding power references
+ * (RPROC_OFFLINE or RPROC_DETACHED), so every rproc_boot() would
+ * free-ride on them and silently do nothing. The service those
+ * references were tracking is gone: void them all. Failures that
+ * leave the processor crashed keep the references, as rproc_shutdown()
+ * can still drain them in that state.
+ */
+ if (ret && rproc->state != RPROC_CRASHED) {
+ dev_err(dev, "failed to recover %s: %d\n", rproc->name, ret);
+ atomic_set(&rproc->power, 0);
+ }
+
unlock_mutex:
mutex_unlock(&rproc->lock);
return ret;
--
2.34.1
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-09-29 7:55 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-29 7:54 [PATCH 0/4] remoteproc: core: Fix the error unwind paths Yonghao Zhang
2026-09-29 7:54 ` [PATCH 1/4] remoteproc: core: Reset freed vring entries to FW_RSC_ADDR_ANY Yonghao Zhang
2026-09-29 7:54 ` [PATCH 2/4] remoteproc: core: Guard against a missing stop() in rproc_start() Yonghao Zhang
2026-09-29 7:54 ` [PATCH 3/4] remoteproc: core: Roll a failed attach back with detach() when available Yonghao Zhang
2026-09-29 7:54 ` [PATCH 4/4] remoteproc: core: Clean up after a failed recovery Yonghao Zhang
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®