From: "Adrián Larumbe" <adrian.larumbe@collabora.com>
To: Boris Brezillon <boris.brezillon@collabora.com>
Cc: Rob Herring <robh@kernel.org>,
Steven Price <steven.price@arm.com>,
Maarten Lankhorst <maarten.lankhorst@linux.intel.com>,
Maxime Ripard <mripard@kernel.org>,
Thomas Zimmermann <tzimmermann@suse.de>,
David Airlie <airlied@gmail.com>,
Simona Vetter <simona@ffwll.ch>,
Faith Ekstrand <faith.ekstrand@collabora.com>,
"Marty E. Plummer" <hanetzer@startmail.com>,
Tomeu Vizoso <tomeu@tomeuvizoso.net>,
Eric Anholt <eric@anholt.net>,
Alyssa Rosenzweig <alyssa.rosenzweig@collabora.com>,
Robin Murphy <robin.murphy@arm.com>,
Philipp Zabel <p.zabel@pengutronix.de>,
dri-devel@lists.freedesktop.org, linux-kernel@vger.kernel.org,
Collabora Kernel Team <kernel@collabora.com>,
Neil Armstrong <neil.armstrong@linaro.org>
Subject: Re: [PATCH v9 06/16] drm/panfrost: Fix PM refcnt and autosuspend issues at device probe/remove
Date: Tue, 22 Sep 2026 20:51:48 +0100 [thread overview]
Message-ID: <arGhZayS8zH4y-61@sobremesa> (raw)
In-Reply-To: <20260914112237.70dedb9c@fedora-21.home>
On 14.09.2026 11:22, Boris Brezillon wrote:
> On Sat, 12 Sep 2026 00:28:07 +0100
> Adrián Larumbe <adrian.larumbe@collabora.com> wrote:
>
> > During device probe(), failure to do a PM get() will leave the usage_count
> > set to 0, which is the value assigned at device creation time. That means
> > when the autosuspend delay expires, runtime suspend callback won't be
> > invoked, so the device will remain powered on forever.
> >
> > On top of that, failure to call PM put() during device unplug means
> > Panfrost device's PM usage_count increases monotonically for every new
> > module reload.
> >
> > The outcome of both of the above meant that:
> >
> > - Devfreq OPP transition notifications would be printed all the time,
> > even when no jobs are being submitted. This quickly fills the kernel
> > ring buffer with junk.
> > - Because MMU interrupts are only enabled when the device is reset,
> > the very first job targeting the tiler heap BO after device probe()
> > would always time out, since the driver's PM runtime resume callback
> > would not be invoked.
> >
> > To fix the above:
> > - Manually adjust the PM refcnt at device probe and removal time.
> > - Ensure pm_runtime_dont_use_autosuspend is called in the wind-down path.
> > - Call pm_runtime_put_autosuspend() when device is ready to accept jobs
> > - Move pm_runtime_set_suspended() before panfrost_device_fini() so that
> > resource unwinding happens in the opposite order as initialisation.
> >
> > Signed-off-by: Adrián Larumbe <adrian.larumbe@collabora.com>
> > Fixes: 635430797d3f ("drm/panfrost: Rework runtime PM initialization")
> > Fixes: 876b15d2c88d ("drm/panfrost: Fix module unload")
> > ---
> > drivers/gpu/drm/panfrost/panfrost_drv.c | 10 ++++++++--
> > 1 file changed, 8 insertions(+), 2 deletions(-)
> >
> > diff --git a/drivers/gpu/drm/panfrost/panfrost_drv.c b/drivers/gpu/drm/panfrost/panfrost_drv.c
> > index 55fc22e8d4d4..a3eff77add55 100644
> > --- a/drivers/gpu/drm/panfrost/panfrost_drv.c
> > +++ b/drivers/gpu/drm/panfrost/panfrost_drv.c
> > @@ -854,6 +854,7 @@ static int panfrost_probe(struct platform_device *pdev)
> >
> > pm_runtime_set_active(pfdev->base.dev);
> > pm_runtime_mark_last_busy(pfdev->base.dev);
> > + pm_runtime_get_noresume(pfdev->base.dev);
> > pm_runtime_enable(pfdev->base.dev);
> > pm_runtime_set_autosuspend_delay(pfdev->base.dev, 50); /* ~3 frames */
> > pm_runtime_use_autosuspend(pfdev->base.dev);
> > @@ -866,13 +867,16 @@ static int panfrost_probe(struct platform_device *pdev)
> > if (err < 0)
> > goto err_out1;
> >
> > + pm_runtime_put_autosuspend(pfdev->base.dev);
> >
> > return 0;
> >
> > err_out1:
> > + pm_runtime_dont_use_autosuspend(pfdev->base.dev);
> > pm_runtime_disable(pfdev->base.dev);
> > - panfrost_device_fini(pfdev);
> > + pm_runtime_put_noidle(pfdev->base.dev);
> > pm_runtime_set_suspended(pfdev->base.dev);
> > + panfrost_device_fini(pfdev);
>
> Not an issue per-se, because put_noidle() is a NOP, but I think it'd be
> easier to reason about with this order:
>
> pm_runtime_dont_use_autosuspend(pfdev->base.dev);
> pm_runtime_disable(pfdev->base.dev);
> panfrost_device_fini(pfdev);
> pm_runtime_set_suspended(pfdev->base.dev);
> pm_runtime_put_noidle(pfdev->base.dev);
>
> This makes it clear that panfrost_device_fini() assumes the device is
> resumed when it's called and suspended when it returns.
Your suggestion feels more natural, which makes me wonder why I decided to move
panfrost_device_fini() to the end of the block.
IIRC we chatted about where to place pm_runtime_put_noidle() and pm_runtime_set_suspended()
in panfrost_device_init() when one of the subsystem initialisations failed. My intuition
had been to leave them until the very bottom of it, in imitation of what you hinted here,
but then it'd go against the notion that winding down must be done in the opposite order
as initialisation. And so because panfrost_device_init() happens before all the pm_runtime_*
incantations, I thought panfrost_device_fini() should be placed at the end in the error
and driver remove paths.
That said, I think a more intuitive init order would be one that imitates what happens in
Panthor when runtime PM is enabled:
/* First thing done in pm_runtime_resume_and_get -> pm_runtime_get_active -> __pm_runtime_resume */
pm_runtime_get_noresume(pfdev->base.dev);
/* Now come subsystem initialisations */
panfrost_gpu_init(pfdev);
panfrost_mmu_init(pfdev);
panfrost_jm_init(pfdev);
panfrost_perfcnt_init(pfdev);
panfrost_gem_init(pfdev);
/* __pm_runtime_resume -> rpm_resume first invokes the resume callback and then
* updates runtime status to RPM_ACTIVE. */
pm_runtime_set_active(pfdev->base.dev);
pm_runtime_enable(ptdev->base.dev);
pm_runtime_mark_last_busy(pfdev->base.dev);
pm_runtime_set_autosuspend_delay(pfdev->base.dev, 50); /* ~3 frames */
pm_runtime_use_autosuspend(pfdev->base.dev);
And then in the error and device remove path would be almost just like you suggest.
pm_runtime_dont_use_autosuspend(pfdev->base.dev);
pm_runtime_disable(pfdev->base.dev);
pm_runtime_put_noidle(pfdev->base.dev);
panfrost_device_fini(pfdev);
pm_runtime_set_suspended(pfdev->base.dev);
The only difference is that pm_runtime_put_noidle() happens before device_fini
because that's the same order as in __pm_runtime_suspend().
However, that means in the commit where I move all DRM initialisation into
panfrost_device_init(), in order to keep things consistent, I should move
pm_runtime_set_suspended() to the very end of the error path.
> > err_out0:
> > return err;
> > }
> > @@ -884,9 +888,11 @@ static void panfrost_remove(struct platform_device *pdev)
> > drm_dev_unregister(&pfdev->base);
> >
> > pm_runtime_get_sync(pfdev->base.dev);
> > + pm_runtime_dont_use_autosuspend(pfdev->base.dev);
> > + pm_runtime_put_noidle(pfdev->base.dev);
> > pm_runtime_disable(pfdev->base.dev);
>
> Let's keep the order consistent with the probe path:
>
> pm_runtime_dont_use_autosuspend(pfdev->base.dev);
> pm_runtime_disable(pfdev->base.dev);
> panfrost_device_fini(pfdev);
> pm_runtime_set_suspended(pfdev->base.dev);
> pm_runtime_put_noidle(pfdev->base.dev);
>
> I also think this deserves comments to explain the noresume/noidle
> dance (device_init/fini take care of clks internally, and when they
> return the device is resumed/suspended, so all we have to is update the
> state, and acquire a ref).
I don't mind doing this, but I think the fact we're calling noresume/noidle
functions already implies device power-up is being handled manually rather
by having RPM invoke resume and suspend callbacks behind the scenes, so
we've no choice other than adjusting PM user counter and status ourselves.
> > - panfrost_device_fini(pfdev);
> > pm_runtime_set_suspended(pfdev->base.dev);
> > + panfrost_device_fini(pfdev);
> > }
> >
> > static ssize_t profiling_show(struct device *dev,
> >
Adrian Larumbe
next prev parent reply other threads:[~2026-09-22 19:52 UTC|newest]
Thread overview: 35+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-11 23:28 [PATCH v9 00/16] Collection of fixes for Panfrost: Perfcnt, RPM, refactorings Adrián Larumbe
2026-09-11 23:28 ` [PATCH v9 01/16] drm/panfrost: Move shrinker initialization and unplug one level down Adrián Larumbe
2026-09-14 8:36 ` Boris Brezillon
2026-09-22 19:48 ` Adrián Larumbe
2026-09-23 6:58 ` Boris Brezillon
2026-09-11 23:28 ` [PATCH v9 02/16] drm/panfrost: Move lock and modparam initialisations into their subsystems Adrián Larumbe
2026-09-11 23:28 ` [PATCH v9 03/16] drm/panfrost: Move debugfs initialisation to relevant subsystems Adrián Larumbe
2026-09-11 23:28 ` [PATCH v9 04/16] drm/panfrost: Skip NULL checks for clock enable/disabling Adrián Larumbe
2026-09-11 23:28 ` [PATCH v9 05/16] drm/panfrost: Consolidate device clock management and reset Adrián Larumbe
2026-09-14 8:45 ` Boris Brezillon
2026-09-22 19:50 ` Adrián Larumbe
2026-09-11 23:28 ` [PATCH v9 06/16] drm/panfrost: Fix PM refcnt and autosuspend issues at device probe/remove Adrián Larumbe
2026-09-14 9:22 ` Boris Brezillon
2026-09-22 19:51 ` Adrián Larumbe [this message]
2026-09-23 8:00 ` Boris Brezillon
2026-09-11 23:28 ` [PATCH v9 07/16] drm/panfrost: Explicitly enable MMU interrupts at device init Adrián Larumbe
2026-09-11 23:28 ` [PATCH v9 08/16] drm/panfrost: Move all DRM device initialisation into device_init() Adrián Larumbe
2026-09-14 9:31 ` Boris Brezillon
2026-09-11 23:28 ` [PATCH v9 09/16] drm/panfrost: Add warning messages to fatal error conditions Adrián Larumbe
2026-09-11 23:28 ` [PATCH v9 10/16] drm/panfrost: Add debugfs knob for manually triggering a GPU reset Adrián Larumbe
2026-09-14 9:39 ` Boris Brezillon
2026-09-22 19:54 ` Adrián Larumbe
2026-09-23 7:10 ` Boris Brezillon
2026-09-11 23:28 ` [PATCH v9 11/16] drm/panfrost: Move perfcnt GPU disable sequence into a helper Adrián Larumbe
2026-09-11 23:28 ` [PATCH v9 12/16] drm/panfrost: Skip cache flush/invalidate when enabling perfcnt Adrián Larumbe
2026-09-14 9:41 ` Boris Brezillon
2026-09-11 23:28 ` [PATCH v9 13/16] drm/panfrost: Avoid cache flush after perfcnt sample in fully coherent systems Adrián Larumbe
2026-09-11 23:28 ` [PATCH v9 14/16] drm/panfrost: Introduce a reset lock Adrián Larumbe
2026-09-11 23:28 ` [PATCH v9 15/16] drm/panfrost: Fix races between perfcnt and reset sequence Adrián Larumbe
2026-09-14 10:16 ` Boris Brezillon
2026-09-23 1:39 ` Adrián Larumbe
2026-09-23 8:08 ` Boris Brezillon
2026-09-11 23:28 ` [PATCH v9 16/16] drm/panfrost: Bump driver minor to reflect new DUMP IOCTL req field Adrián Larumbe
2026-09-14 10:19 ` Boris Brezillon
2026-09-22 19:54 ` Adrián Larumbe
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=arGhZayS8zH4y-61@sobremesa \
--to=adrian.larumbe@collabora.com \
--cc=airlied@gmail.com \
--cc=alyssa.rosenzweig@collabora.com \
--cc=boris.brezillon@collabora.com \
--cc=dri-devel@lists.freedesktop.org \
--cc=eric@anholt.net \
--cc=faith.ekstrand@collabora.com \
--cc=hanetzer@startmail.com \
--cc=kernel@collabora.com \
--cc=linux-kernel@vger.kernel.org \
--cc=maarten.lankhorst@linux.intel.com \
--cc=mripard@kernel.org \
--cc=neil.armstrong@linaro.org \
--cc=p.zabel@pengutronix.de \
--cc=robh@kernel.org \
--cc=robin.murphy@arm.com \
--cc=simona@ffwll.ch \
--cc=steven.price@arm.com \
--cc=tomeu@tomeuvizoso.net \
--cc=tzimmermann@suse.de \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®