From: Steven Price <steven.price@arm.com>
To: "Adrián Larumbe" <adrian.larumbe@collabora.com>
Cc: Boris Brezillon <boris.brezillon@collabora.com>,
Rob Herring <robh@kernel.org>,
Maarten Lankhorst <maarten.lankhorst@linux.intel.com>,
Maxime Ripard <mripard@kernel.org>,
Thomas Zimmermann <tzimmermann@suse.de>,
David Airlie <airlied@gmail.com>, Simona Vetter <simona@ffwll.ch>,
Faith Ekstrand <faith.ekstrand@collabora.com>,
"Marty E. Plummer" <hanetzer@startmail.com>,
Tomeu Vizoso <tomeu@tomeuvizoso.net>,
Eric Anholt <eric@anholt.net>,
Robin Murphy <robin.murphy@arm.com>,
Philipp Zabel <p.zabel@pengutronix.de>,
dri-devel@lists.freedesktop.org, linux-kernel@vger.kernel.org,
Collabora Kernel Team <kernel@collabora.com>,
Neil Armstrong <neil.armstrong@linaro.org>
Subject: Re: [PATCH v12 09/15] drm/panfrost: Add warning messages to fatal error conditions
Date: Wed, 7 Oct 2026 13:51:17 +0100 [thread overview]
Message-ID: <5bd07106-e9dd-487a-be85-8e6a73c67967@arm.com> (raw)
In-Reply-To: <179129938072.1018806.14660127756555690690.b4-reply@b4>
On 06/10/2026 16:09, Adrián Larumbe wrote:
> On 2026-10-02 15:59:20+01:00, Steven Price wrote:
>> On 29/09/2026 04:44, Adrián Larumbe wrote:
>>
>>> Rather than just failing silently, let's warn the user of device remove not
>>> being able to take an PM reference or the PM suspend path still reporting
>>> inflight jobs. Neither situation should ever happen.
>>>
>>> Reviewed-by: Boris Brezillon <boris.brezillon@collabora.com>
>>> Signed-off-by: Adrián Larumbe <adrian.larumbe@collabora.com>
>>> ---
>>> drivers/gpu/drm/panfrost/panfrost_device.c | 5 +++--
>>> 1 file changed, 3 insertions(+), 2 deletions(-)
>>>
>>> diff --git a/drivers/gpu/drm/panfrost/panfrost_device.c b/drivers/gpu/drm/panfrost/panfrost_device.c
>>> index c6bf3d0663df..09a5752a3f40 100644
>>> --- a/drivers/gpu/drm/panfrost/panfrost_device.c
>>> +++ b/drivers/gpu/drm/panfrost/panfrost_device.c
>>> @@ -9,6 +9,7 @@
>>> #include <linux/pm_runtime.h>
>>> #include <linux/regulator/consumer.h>
>>> #include <drm/drm_drv.h>
>>> +#include <drm/drm_print.h>
>>>
>>> #include "panfrost_device.h"
>>> #include "panfrost_devfreq.h"
>>> @@ -357,7 +358,7 @@ int panfrost_device_init(struct panfrost_device *pfdev)
>>>
>>> void panfrost_device_fini(struct panfrost_device *pfdev)
>>> {
>>> - pm_runtime_get_sync(pfdev->base.dev);
>>> + drm_WARN_ON(&pfdev->base, pm_runtime_get_sync(pfdev->base.dev) < 0);
>>
>> This seems fine.
>>
>>> pm_runtime_dont_use_autosuspend(pfdev->base.dev);
>>> pm_runtime_disable(pfdev->base.dev);
>>> @@ -516,7 +517,7 @@ static int panfrost_device_runtime_suspend(struct device *dev)
>>> {
>>> struct panfrost_device *pfdev = dev_get_drvdata(dev);
>>>
>>> - if (!panfrost_jm_is_idle(pfdev))
>>> + if (drm_WARN_ON(&pfdev->base, !panfrost_jm_is_idle(pfdev)))
>>
>> I'm a bit wary that this might be something that user space can trigger.
>> My AI says:
>>
>> The runtime-suspend WARN can be reached by ordinary userspace job
>> submissions. The DRM scheduler increments credit_count before calling
>> Panfrost’s job runner (drivers/gpu/drm/scheduler/sched_main.c:1044).
>> Panfrost takes the job’s PM reference later in hardware submission
>> (drivers/gpu/drm/panfrost/panfrost_job.c:213). If autosuspend runs in
>> that interval, the new WARN
>> (drivers/gpu/drm/panfrost/panfrost_device.c:525) sees the credit and
>> fires, even though this is a timing race rather than a broken job.
>> Repeated submissions near the autosuspend boundary could therefore
>> produce repeated stack traces. The PM core treats the resulting -EBUSY
>> as a transient failure.
>>
>> Now I have to admit I don't trust it that much - but I'd want a
>> convincing argument on why panfrost_jm_is_idle() will never be false here.
>
> You're right. I was in the belief that autosuspend kicking in was proof of no
> inflight or pending jobs present in the scheduler queues, so I came to treat
> this check as things having gone awry.
>
> I guess its value lies in the ability of the PM runtime suspend handler
> to cancel itself at an autosuspend event, like you said.
>
> However, it just made me wonder: what would happen in the event that autosuspend
> kicks in and runs panfrost_device_runtime_suspend() right at the same time that
> a scheduler job is picked up by drm_sched_run_job_work(), but hasn't yet reached
> the statement where it does an atomic increment on the credit_count? I guess
> nothing, because panfrost_job_hw_submit() is getting a PM reference before
> accessing any HW registers, and that should take care of dealing with any
> ongoing autosuspend events.
>
> In that case I'll just delete that warning. However, I'd say it's bad practice
> to have DRM drivers access the internal state of the DRM scheduler. At present,
> only Panfrost and Etnaviv poke it in their RPM suspend handlers, and I've been
> wondering whether we should get rid of this check altogether, or else maybe ask
> the scheduler maintainers whether it makes sense to have a non-racy way to
> query the presence of pending jobs in their queues?
Yes, I'm not sure whether we actually need that check. As you say
there's still a race where panfrost_jm_is_idle() returns true, but
afterwards credit_count is incremented. I think this is safe -
panfrost_job_hw_submit() takes a PM reference which will cause the GPU
to be woken up again. So we could just drop the panfrost_jm_is_idle()
function completely.
Whether that has any performance impact - i.e. do we often race the
autosuspend operation? - I've no idea. Presumably you didn't hit it when
you had the WARN in place. So if you'd prefer to just remove the code
then that's fine by me.
Thanks,
Steve
next prev parent reply other threads:[~2026-10-07 12:51 UTC|newest]
Thread overview: 48+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-29 3:44 [PATCH v12 00/15] Collection of fixes for Panfrost: Perfcnt, RPM, refactorings Adrián Larumbe
2026-09-29 3:44 ` [PATCH v12 01/15] drm/panfrost: Move shrinker initialization and unplug one level down Adrián Larumbe
2026-10-02 13:58 ` Steven Price
2026-09-29 3:44 ` [PATCH v12 02/15] drm/panfrost: Move lock and modparam initialisations into their subsystems Adrián Larumbe
2026-10-02 14:06 ` Steven Price
2026-09-29 3:44 ` [PATCH v12 03/15] drm/panfrost: Move debugfs initialisation to relevant subsystems Adrián Larumbe
2026-10-02 14:10 ` Steven Price
2026-10-06 0:50 ` Adrián Larumbe
2026-09-29 3:44 ` [PATCH v12 04/15] drm/panfrost: Skip NULL checks for clock enable/disabling Adrián Larumbe
2026-10-02 14:14 ` Steven Price
2026-09-29 3:44 ` [PATCH v12 05/15] drm/panfrost: Consolidate device clock management and reset Adrián Larumbe
2026-10-02 14:20 ` Steven Price
2026-09-29 3:44 ` [PATCH v12 06/15] drm/panfrost: Fix PM refcnt and autosuspend issues at device probe/remove Adrián Larumbe
2026-10-02 14:21 ` Steven Price
2026-09-29 3:44 ` [PATCH v12 07/15] drm/panfrost: Explicitly enable MMU interrupts at device init Adrián Larumbe
2026-10-02 14:26 ` Steven Price
2026-09-29 3:44 ` [PATCH v12 08/15] drm/panfrost: Move all DRM device initialisation into device_init() Adrián Larumbe
2026-10-02 14:37 ` Steven Price
2026-09-29 3:44 ` [PATCH v12 09/15] drm/panfrost: Add warning messages to fatal error conditions Adrián Larumbe
2026-10-02 14:59 ` Steven Price
2026-10-06 15:09 ` Adrián Larumbe
2026-10-07 12:51 ` Steven Price [this message]
2026-09-29 3:44 ` [PATCH v12 10/15] drm/panfrost: Add debugfs knob for manually triggering a GPU reset Adrián Larumbe
2026-10-02 15:10 ` Steven Price
2026-10-06 13:38 ` Adrián Larumbe
2026-10-07 12:54 ` Steven Price
2026-09-29 3:44 ` [PATCH v12 11/15] drm/panfrost: Move perfcnt GPU disable sequence into a helper Adrián Larumbe
2026-09-29 3:44 ` [PATCH v12 12/15] drm/panfrost: Skip cache flush/invalidate when enabling perfcnt Adrián Larumbe
2026-10-02 15:14 ` Steven Price
2026-10-05 8:38 ` Boris Brezillon
2026-10-05 15:06 ` Steven Price
2026-10-05 15:37 ` Boris Brezillon
2026-10-05 16:05 ` Steven Price
2026-10-05 16:57 ` Boris Brezillon
2026-10-06 17:42 ` Adrián Larumbe
2026-10-06 17:35 ` Adrián Larumbe
2026-10-07 13:28 ` Steven Price
2026-10-06 17:18 ` Adrián Larumbe
2026-10-07 13:21 ` Steven Price
2026-10-06 16:45 ` Adrián Larumbe
2026-10-07 13:09 ` Steven Price
2026-09-29 3:44 ` [PATCH v12 13/15] drm/panfrost: Avoid cache flush after perfcnt sample in fully coherent systems Adrián Larumbe
2026-10-02 15:28 ` Steven Price
2026-09-29 3:44 ` [PATCH v12 14/15] drm/panfrost: Introduce a reset lock Adrián Larumbe
2026-10-02 15:34 ` Steven Price
2026-09-29 3:44 ` [PATCH v12 15/15] drm/panfrost: Fix races between perfcnt and reset sequence Adrián Larumbe
2026-10-02 15:44 ` Steven Price
2026-10-06 14:14 ` Adrián Larumbe
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=5bd07106-e9dd-487a-be85-8e6a73c67967@arm.com \
--to=steven.price@arm.com \
--cc=adrian.larumbe@collabora.com \
--cc=airlied@gmail.com \
--cc=boris.brezillon@collabora.com \
--cc=dri-devel@lists.freedesktop.org \
--cc=eric@anholt.net \
--cc=faith.ekstrand@collabora.com \
--cc=hanetzer@startmail.com \
--cc=kernel@collabora.com \
--cc=linux-kernel@vger.kernel.org \
--cc=maarten.lankhorst@linux.intel.com \
--cc=mripard@kernel.org \
--cc=neil.armstrong@linaro.org \
--cc=p.zabel@pengutronix.de \
--cc=robh@kernel.org \
--cc=robin.murphy@arm.com \
--cc=simona@ffwll.ch \
--cc=tomeu@tomeuvizoso.net \
--cc=tzimmermann@suse.de \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®