* [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT
@ 2026-08-08 23:33 Luke.Wildhardt
2026-08-10 8:28 ` Philipp Stanner
` (2 more replies)
0 siblings, 3 replies; 21+ messages in thread
From: Luke.Wildhardt @ 2026-08-08 23:33 UTC (permalink / raw)
To: matthew.brost, dakr, phasta, ckoenig.leichtzumerken, dri-devel
Cc: regressions, amd-gfx, linux-kernel
Since the DRM scheduler default policy was switched to FAIR in 7.2, sustained 100% GPU load in-game on my RX 9070 XT causes severe performance degradation. In the example title running through Proton, Project Silverfish, the foreground application degrades to roughly 10 fps or freezes outright while audio continues, and the KDE Plasma Wayland session can lock up entirely, requiring a reboot or killing the compositor to recover. The problem is not limited to games. Running the game in the background + alt tabbing, then playing a YouTube video can also cause the entire desktop to freeze.
Then, if the game starts to stutter, either alt-tabbing, or opening up a menu in game (which reduced the load on the GPU) immediately stops the stuttering.
The issue is reliably reproducible under sustained GPU saturation, although the time-to-failure varies. Increasing shadow load (which decreases in-game frame rate) substantially shortens the time-to-failure.
In game, MangoHud does not reveal any frametime discrepancy.
Bisection / test matrix
Reproducer: Project Silverfish (Proton) at settings that saturate the GPU, with a second GPU client active (browser video). Other titles show the same symptom under sustained load; I have not yet re-tested those specific titles against the fix.
Kernel Result
torvalds/master stock FAIR (default) FAIL
same as above + reverts below gpu_sched.sched_policy=2 (FAIR) FAIL
same as above + reverts below gpu_sched.sched_policy=1 (FIFO) PASS
7.1.5 release PASS
drm-next FAIR (default) FAIL
Reverted Commits:
d09339388b77 drm/sched: Remove drm_sched_init_args->num_rqs
2833a0512b4c drm/sched: Remove drm_sched_init_args->num_rqs usage
16e7698bc04d drm/sched: Embed run queue singleton into the scheduler
77a6809f1dc3 drm/sched: Remove FIFO and RR and simplify to a single run queue
45c211ddf92a drm/sched: Switch default policy to fair
2462a0ce23b0 drm/amdgpu: Remove drm_sched_init_args->num_rqs usage
Kernel (failing): 7.2.0-rc6 (stock), also current drm-next
Kernel (working): 7.2.0-rc6-default-revert-00006-g7e5de9b07e3d with sched_policy=1
7.1 release
Distro: Arch Linux
Session: KDE Plasma 6.7.4 / Wayland (KWin), Qt 6.11.1, KF 6.28.0
CPU: AMD Ryzen 9 9950X3D
RAM: 64 GiB
GPU: Sapphire Pulse Radeon RX 9070 XT
03:00.0 [1002:7550] rev c0, subsys Sapphire 1478
Motherboard: ASUS (AM5 / 800-series chipset)
Mesa: 26.1.6-arch1.1 (RADV, driverVersion 26.1.6)
Vulkan: instance 1.4.357 / device 1.4.354
On the same reverted kernel build, changing only gpu_sched.sched_policy from 2 (FAIR) to 1 (FIFO) changes the result from reliably failing to passing extended stress testing.
Possible mitigation options
Given that 7.2 is late in the release cycle and 77a6809f1dc3 removes FIFO/RR as selectable policies, there is currently no runtime workaround for affected systems.
Revert 45c211ddf92a so FIFO remains the default while FAIR remains available for further testing.
If necessary, also revert 77a6809f1dc3 and dependent changes to restore runtime policy selection.
Alternatively, fix the underlying FAIR regression if a suitable fix can be identified in time for 7.2.
I mention the revert options because the justification for removing FIFO/RR included the absence of known regressions relative to those policies. This report appears to provide at least one counterexample.
I am willing to test patches responsively and collect data if necessary.
^ permalink raw reply [flat|nested] 21+ messages in thread* Re: [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT 2026-08-08 23:33 [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT Luke.Wildhardt @ 2026-08-10 8:28 ` Philipp Stanner 2026-08-10 12:45 ` Danilo Krummrich 2026-08-10 11:36 ` Tvrtko Ursulin 2026-08-11 7:46 ` Tvrtko Ursulin 2 siblings, 1 reply; 21+ messages in thread From: Philipp Stanner @ 2026-08-10 8:28 UTC (permalink / raw) To: Luke.Wildhardt, matthew.brost, dakr, phasta, ckoenig.leichtzumerken, dri-devel Cc: regressions, amd-gfx, linux-kernel, Tvrtko Ursulin +Cc Tvrtko On Sat, 2026-08-08 at 23:33 +0000, Luke.Wildhardt@proton.me wrote: > Since the DRM scheduler default policy was switched to FAIR in 7.2, sustained 100% GPU load in-game on my RX 9070 XT causes severe performance degradation. In the example title running through Proton, Project Silverfish, the foreground application degrades to roughly 10 fps or freezes outright while audio continues, and the KDE Plasma Wayland session can lock up entirely, requiring a reboot or killing the compositor to recover. The problem is not limited to games. Running the game in the background + alt tabbing, then playing a YouTube video can also cause the entire desktop to freeze. > > Then, if the game starts to stutter, either alt-tabbing, or opening up a menu in game (which reduced the load on the GPU) immediately stops the stuttering. > > The issue is reliably reproducible under sustained GPU saturation, although the time-to-failure varies. Increasing shadow load (which decreases in-game frame rate) substantially shortens the time-to-failure. > > In game, MangoHud does not reveal any frametime discrepancy. > > Bisection / test matrix > > Reproducer: Project Silverfish (Proton) at settings that saturate the GPU, with a second GPU client active (browser video). Other titles show the same symptom under sustained load; I have not yet re-tested those specific titles against the fix. > > Kernel Result > torvalds/master stock FAIR (default) FAIL > same as above + reverts below gpu_sched.sched_policy=2 (FAIR) FAIL > same as above + reverts below gpu_sched.sched_policy=1 (FIFO) PASS > 7.1.5 release PASS > drm-next FAIR (default) FAIL > > Reverted Commits: > > d09339388b77 drm/sched: Remove drm_sched_init_args->num_rqs > 2833a0512b4c drm/sched: Remove drm_sched_init_args->num_rqs usage > 16e7698bc04d drm/sched: Embed run queue singleton into the scheduler > 77a6809f1dc3 drm/sched: Remove FIFO and RR and simplify to a single run queue > 45c211ddf92a drm/sched: Switch default policy to fair > 2462a0ce23b0 drm/amdgpu: Remove drm_sched_init_args->num_rqs usage > > > > Kernel (failing): 7.2.0-rc6 (stock), also current drm-next > Kernel (working): 7.2.0-rc6-default-revert-00006-g7e5de9b07e3d with sched_policy=1 > 7.1 release > Distro: Arch Linux > Session: KDE Plasma 6.7.4 / Wayland (KWin), Qt 6.11.1, KF 6.28.0 > CPU: AMD Ryzen 9 9950X3D > RAM: 64 GiB > GPU: Sapphire Pulse Radeon RX 9070 XT > 03:00.0 [1002:7550] rev c0, subsys Sapphire 1478 > Motherboard: ASUS (AM5 / 800-series chipset) > Mesa: 26.1.6-arch1.1 (RADV, driverVersion 26.1.6) > Vulkan: instance 1.4.357 / device 1.4.354 > > > > > On the same reverted kernel build, changing only gpu_sched.sched_policy from 2 (FAIR) to 1 (FIFO) changes the result from reliably failing to passing extended stress testing. > > > > Possible mitigation options > > Given that 7.2 is late in the release cycle and 77a6809f1dc3 removes FIFO/RR as selectable policies, there is currently no runtime workaround for affected systems. > > > Revert 45c211ddf92a so FIFO remains the default while FAIR remains available for further testing. That might be the wisest thing to do; but let's hear if Tvrtko has an idea for a hotfix first. > > If necessary, also revert 77a6809f1dc3 and dependent changes to restore runtime policy selection. > > Alternatively, fix the underlying FAIR regression if a suitable fix can be identified in time for 7.2. > > > I mention the revert options because the justification for removing FIFO/RR included the absence of known regressions relative to those policies. This report appears to provide at least one counterexample. > > > I am willing to test patches responsively and collect data if necessary. Thanks for your effort and help. I had dared to hope that we could get this one through without a regression, but reality always catches up.. P. ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT 2026-08-10 8:28 ` Philipp Stanner @ 2026-08-10 12:45 ` Danilo Krummrich 2026-08-10 13:48 ` Tvrtko Ursulin 0 siblings, 1 reply; 21+ messages in thread From: Danilo Krummrich @ 2026-08-10 12:45 UTC (permalink / raw) To: Philipp Stanner Cc: phasta, Luke.Wildhardt, matthew.brost, ckoenig.leichtzumerken, dri-devel, regressions, amd-gfx, linux-kernel, Tvrtko Ursulin On Mon Aug 10, 2026 at 10:28 AM CEST, Philipp Stanner wrote: >> Reverted Commits: >> >> d09339388b77 drm/sched: Remove drm_sched_init_args->num_rqs >> 2833a0512b4c drm/sched: Remove drm_sched_init_args->num_rqs usage >> 16e7698bc04d drm/sched: Embed run queue singleton into the scheduler >> 77a6809f1dc3 drm/sched: Remove FIFO and RR and simplify to a single run queue >> 45c211ddf92a drm/sched: Switch default policy to fair >> 2462a0ce23b0 drm/amdgpu: Remove drm_sched_init_args->num_rqs usage [...] > That might be the wisest thing to do; but let's hear if Tvrtko has an idea for > a hotfix first. -rc7 was released yesterday; even with a working hotfix today it'd be tricky to ensure the hotfix does not regress other drivers or workloads with 7.2 being just a few days ahead. I suggest to not wait and get the reverts ready. Thanks, Danilo ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT 2026-08-10 12:45 ` Danilo Krummrich @ 2026-08-10 13:48 ` Tvrtko Ursulin 2026-08-10 16:40 ` Luke.Wildhardt 0 siblings, 1 reply; 21+ messages in thread From: Tvrtko Ursulin @ 2026-08-10 13:48 UTC (permalink / raw) To: Danilo Krummrich, Philipp Stanner Cc: phasta, Luke.Wildhardt, matthew.brost, ckoenig.leichtzumerken, dri-devel, regressions, amd-gfx, linux-kernel On 10/08/2026 13:45, Danilo Krummrich wrote: > On Mon Aug 10, 2026 at 10:28 AM CEST, Philipp Stanner wrote: >>> Reverted Commits: >>> >>> d09339388b77 drm/sched: Remove drm_sched_init_args->num_rqs >>> 2833a0512b4c drm/sched: Remove drm_sched_init_args->num_rqs usage >>> 16e7698bc04d drm/sched: Embed run queue singleton into the scheduler >>> 77a6809f1dc3 drm/sched: Remove FIFO and RR and simplify to a single run queue >>> 45c211ddf92a drm/sched: Switch default policy to fair >>> 2462a0ce23b0 drm/amdgpu: Remove drm_sched_init_args->num_rqs usage > > [...] > >> That might be the wisest thing to do; but let's hear if Tvrtko has an idea for >> a hotfix first. > > -rc7 was released yesterday; even with a working hotfix today it'd be tricky to > ensure the hotfix does not regress other drivers or workloads with 7.2 being > just a few days ahead. > > I suggest to not wait and get the reverts ready. Yes reverts would be safer. I have them in a branch at people.freedesktop.org/~tursulin/drm-intel drm-sched-fair-reverts, mostly straightforward apart from one easy conflict in amdgpu_xcp_release_sched. Smoke tested on Steam Deck looks fine. Having said that, the fix for min_vruntime handling I provided earlier in the thread also looks fine in my testing and is simple. No regressions found with synthetic unit test workloads or messing around on the Steam Deck. But we need to wait to hear from Luke since I haven't been able to repro his report locally yet. Regards, Tvrtko ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT 2026-08-10 13:48 ` Tvrtko Ursulin @ 2026-08-10 16:40 ` Luke.Wildhardt 2026-08-11 7:29 ` Philipp Stanner ` (2 more replies) 0 siblings, 3 replies; 21+ messages in thread From: Luke.Wildhardt @ 2026-08-10 16:40 UTC (permalink / raw) To: Tvrtko Ursulin Cc: Danilo Krummrich, Philipp Stanner, phasta, matthew.brost, ckoenig.leichtzumerken, dri-devel, regressions, amd-gfx, linux-kernel Oops, forgot to reply all in my last message. I did a quick test of the patch before I left this morning, the freezing seems to gone, but when the GPU is under the same condition, the stuttering is more uniform if that makes sense. Granted this was only a 5 minute test but the freeze was rather reliable before, I can do more testing later. I know these things can be hard to describe, if that helps. If there's any profiling I could do or detailed logging that would help don't hesitate to ask. -------- Original Message -------- On Monday, 08/10/26 at 06:48 Tvrtko Ursulin <tvrtko.ursulin@igalia.com> wrote: On 10/08/2026 13:45, Danilo Krummrich wrote: > On Mon Aug 10, 2026 at 10:28 AM CEST, Philipp Stanner wrote: >>> Reverted Commits: >>> >>> d09339388b77 drm/sched: Remove drm_sched_init_args->num_rqs >>> 2833a0512b4c drm/sched: Remove drm_sched_init_args->num_rqs usage >>> 16e7698bc04d drm/sched: Embed run queue singleton into the scheduler >>> 77a6809f1dc3 drm/sched: Remove FIFO and RR and simplify to a single run queue >>> 45c211ddf92a drm/sched: Switch default policy to fair >>> 2462a0ce23b0 drm/amdgpu: Remove drm_sched_init_args->num_rqs usage > > [...] > >> That might be the wisest thing to do; but let's hear if Tvrtko has an idea for >> a hotfix first. > > -rc7 was released yesterday; even with a working hotfix today it'd be tricky to > ensure the hotfix does not regress other drivers or workloads with 7.2 being > just a few days ahead. > > I suggest to not wait and get the reverts ready. Yes reverts would be safer. I have them in a branch at people.freedesktop.org/~tursulin/drm-intel drm-sched-fair-reverts, mostly straightforward apart from one easy conflict in amdgpu_xcp_release_sched. Smoke tested on Steam Deck looks fine. Having said that, the fix for min_vruntime handling I provided earlier in the thread also looks fine in my testing and is simple. No regressions found with synthetic unit test workloads or messing around on the Steam Deck. But we need to wait to hear from Luke since I haven't been able to repro his report locally yet. Regards, Tvrtko ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT 2026-08-10 16:40 ` Luke.Wildhardt @ 2026-08-11 7:29 ` Philipp Stanner 2026-08-11 7:49 ` Tvrtko Ursulin 2026-08-11 17:59 ` Luke.Wildhardt 2 siblings, 0 replies; 21+ messages in thread From: Philipp Stanner @ 2026-08-11 7:29 UTC (permalink / raw) To: Luke.Wildhardt, Tvrtko Ursulin Cc: Danilo Krummrich, phasta, matthew.brost, ckoenig.leichtzumerken, dri-devel, regressions, amd-gfx, linux-kernel On Mon, 2026-08-10 at 16:40 +0000, Luke.Wildhardt@proton.me wrote: > Oops, forgot to reply all in my last message. > > I did a quick test of the patch before I left this morning, the > freezing seems to gone, but when the GPU is under the same condition, > the stuttering is more uniform if that makes sense. Granted this was > only a 5 minute test but the freeze was rather reliable before, I can > do more testing later. That, unfortunately, indeed sounds as if we should go for the revert ASAP. > > I know these things can be hard to describe, if that helps. If > there's any profiling I could do or detailed logging that would help > don't hesitate to ask. > Would a screencast / filmed screen be helpful? IIRC Steam also has some fps counter etc in a submenu somewhere? In any case, if we want to go for the CFS rework in another respin, I suppose that it would be cool if @Tvrtko and you stay in contact Thank you very much for the report! Philipp ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT 2026-08-10 16:40 ` Luke.Wildhardt 2026-08-11 7:29 ` Philipp Stanner @ 2026-08-11 7:49 ` Tvrtko Ursulin 2026-08-11 17:59 ` Luke.Wildhardt 2 siblings, 0 replies; 21+ messages in thread From: Tvrtko Ursulin @ 2026-08-11 7:49 UTC (permalink / raw) To: Luke.Wildhardt Cc: Danilo Krummrich, Philipp Stanner, phasta, matthew.brost, ckoenig.leichtzumerken, dri-devel, regressions, amd-gfx, linux-kernel On 10/08/2026 17:40, Luke.Wildhardt@proton.me wrote: > Oops, forgot to reply all in my last message. > > I did a quick test of the patch before I left this morning, the freezing seems to gone, but when the GPU is under the same condition, the stuttering is more uniform if that makes sense. Granted this was only a 5 minute test but the freeze was rather reliable before, I can do more testing later. > > I know these things can be hard to describe, if that helps. If there's any profiling I could do or detailed logging that would help don't hesitate to ask. I am assuming this more uniform stuttering is something not present with FIFO? Would it be possible to make two short videos, one before the regression kicks in, and one after. And for both cases a snapshot of gputop (from igt-gpu-tools) output before and after. Regards, Tvrtko > -------- Original Message -------- > On Monday, 08/10/26 at 06:48 Tvrtko Ursulin <tvrtko.ursulin@igalia.com> wrote: > > On 10/08/2026 13:45, Danilo Krummrich wrote: >> On Mon Aug 10, 2026 at 10:28 AM CEST, Philipp Stanner wrote: >>>> Reverted Commits: >>>> >>>> d09339388b77 drm/sched: Remove drm_sched_init_args->num_rqs >>>> 2833a0512b4c drm/sched: Remove drm_sched_init_args->num_rqs usage >>>> 16e7698bc04d drm/sched: Embed run queue singleton into the scheduler >>>> 77a6809f1dc3 drm/sched: Remove FIFO and RR and simplify to a single run queue >>>> 45c211ddf92a drm/sched: Switch default policy to fair >>>> 2462a0ce23b0 drm/amdgpu: Remove drm_sched_init_args->num_rqs usage >> >> [...] >> >>> That might be the wisest thing to do; but let's hear if Tvrtko has an idea for >>> a hotfix first. >> >> -rc7 was released yesterday; even with a working hotfix today it'd be tricky to >> ensure the hotfix does not regress other drivers or workloads with 7.2 being >> just a few days ahead. >> >> I suggest to not wait and get the reverts ready. > > Yes reverts would be safer. I have them in a branch at > people.freedesktop.org/~tursulin/drm-intel drm-sched-fair-reverts, > mostly straightforward apart from one easy conflict in > amdgpu_xcp_release_sched. Smoke tested on Steam Deck looks fine. > > Having said that, the fix for min_vruntime handling I provided earlier > in the thread also looks fine in my testing and is simple. No > regressions found with synthetic unit test workloads or messing around > on the Steam Deck. But we need to wait to hear from Luke since I haven't > been able to repro his report locally yet. > > Regards, > > Tvrtko > > ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT 2026-08-10 16:40 ` Luke.Wildhardt 2026-08-11 7:29 ` Philipp Stanner 2026-08-11 7:49 ` Tvrtko Ursulin @ 2026-08-11 17:59 ` Luke.Wildhardt 2026-08-12 11:37 ` Tvrtko Ursulin 2 siblings, 1 reply; 21+ messages in thread From: Luke.Wildhardt @ 2026-08-11 17:59 UTC (permalink / raw) To: Tvrtko Ursulin Cc: Danilo Krummrich, Philipp Stanner, phasta, matthew.brost, ckoenig.leichtzumerken, dri-devel, regressions, amd-gfx, linux-kernel I want to be very clear here, I was working with Claude Opus on debugging this issue. I am not a good programmer and work in on the hardware side of things, so this is only code I've tested and not of my creation, just hoping to be helpful. It appears the freezing (and maybe what I'll call micro freezing instead of stuttering) was resolved after an issue was potentially identified in drivers/gpu/drm/scheduler/sched_rq.c I did a 20m play session attempting to reproduce the issue but was not able to. Previously, I could reproduce this on-demand. I will copy the writeup it created and the testing we did below. I am only claiming that the freezing had been seemingly resolved, not a particular fitness or correct identification of the issue, that is for developers to judge. Quotations are an output of a summery I had Claude create. " THE INVARIANT drm_sched_entity_stats::vruntime is held absolute while the entity is linked in the run queue tree, and relative to min_vruntime while it is idle. save_vruntime() converts one way on the way out, restore_vruntime() the other on the way back in. drm_sched_rq_add_entity() calls restore unconditionally, which is only correct if the entity has definitely been through save. THE WINDOW drm_sched_entity_pop_job(): spsc_queue_pop(&entity->job_queue); /* queue now empty */ drm_sched_rq_pop_entity(entity); /* takes entity->lock inside */ No lock held across those two lines. A concurrent push seeing the queue empty gets first == true, calls add_entity, and restores a vruntime that has not yet been saved -- adding min_vruntime to an already-absolute value. If the entity is leftmost, get_min_vruntime() returns its own vruntime and it doubles. It persists: pop_entity then finds the pushed job, takes the next_job branch, and update_vruntime() carries the inflated value forward. The entity sits far right in the tree and is not selected until min_vruntime catches up, which under sustained load may be effectively never. EVIDENCE WARN_ON_ONCE(!RB_EMPTY_NODE(&entity->rb_tree_node)) at the top of add_entity(), after the rq lock, fires reproducibly. That condition should be unreachable. add_entity() is reached only from the "first job" branch of push_job(), and first is true only when the SPSC queue was empty. A linked entity implies its last pop_entity() peeked non-empty. Empty queue and still linked cannot both hold in any ordered execution. Fresh entities are excluded -- entity_init() calls RB_CLEAR_NODE() unconditionally. Instrumented as a counter instead, with ring and comm logged: idle 10s: 0 races / 5662 restores gameplay 60s: 8 races / 43315 restores [ 192.754604] vruntime race on ring gfx_0.0.0 (comm kwin_wayla:cs0) [ 226.882135] vruntime race on ring gfx_0.0.0 (comm kwin_wayla:cs0) [ 292.226032] vruntime race on ring gfx_0.0.0 (comm kwin_wayla:cs0) plus Xwayland:cs0 and vkd3d_queue on gfx, and many more on sdma0/sdma1. The gfx_0.0.0 / kwin_wayla:cs0 entries are the compositor submission thread on the graphics ring -- the process that dies when the desktop freezes. CHANGE The principled fix is presumably to close the window by holding entity->lock across the pop, which needs a _locked variant of drm_sched_rq_pop_entity(). I did not attempt that. What I tested checks the invariant instead: a still- linked entity never left, so its vruntime is already absolute and its tree position valid. --- a/drivers/gpu/drm/scheduler/sched_rq.c +++ b/drivers/gpu/drm/scheduler/sched_rq.c @@ drm_sched_rq_add_entity - ts = drm_sched_rq_get_min_vruntime(rq); - ts = drm_sched_entity_restore_vruntime(entity, ts, rq->head_prio); - drm_sched_rq_update_tree_locked(entity, rq, ts); + if (RB_EMPTY_NODE(&entity->rb_tree_node)) { + ts = drm_sched_rq_get_min_vruntime(rq); + ts = drm_sched_entity_restore_vruntime(entity, ts, + rq->head_prio); + drm_sched_rq_update_tree_locked(entity, rq, ts); + } On stock 7.2-rc7 with this applied, freezes are gone against a trigger that previously reproduced on demand, and perceptible hitching appears gone too. The race counter keeps incrementing (35 in the last session against 295744 restores), so the window still opens at the same rate and is now absorbed -- the timing has not merely shifted. Extended soak in progress. " On Monday, August 10th, 2026 at 9:40 AM, Luke.Wildhardt@proton.me <Luke.Wildhardt@proton.me> wrote: > Oops, forgot to reply all in my last message. > > I did a quick test of the patch before I left this morning, the freezing seems to gone, but when the GPU is under the same condition, the stuttering is more uniform if that makes sense. Granted this was only a 5 minute test but the freeze was rather reliable before, I can do more testing later. > > I know these things can be hard to describe, if that helps. If there's any profiling I could do or detailed logging that would help don't hesitate to ask. > > > > -------- Original Message -------- > On Monday, 08/10/26 at 06:48 Tvrtko Ursulin <tvrtko.ursulin@igalia.com> wrote: > > On 10/08/2026 13:45, Danilo Krummrich wrote: > > On Mon Aug 10, 2026 at 10:28 AM CEST, Philipp Stanner wrote: > >>> Reverted Commits: > >>> > >>> d09339388b77 drm/sched: Remove drm_sched_init_args->num_rqs > >>> 2833a0512b4c drm/sched: Remove drm_sched_init_args->num_rqs usage > >>> 16e7698bc04d drm/sched: Embed run queue singleton into the scheduler > >>> 77a6809f1dc3 drm/sched: Remove FIFO and RR and simplify to a single run queue > >>> 45c211ddf92a drm/sched: Switch default policy to fair > >>> 2462a0ce23b0 drm/amdgpu: Remove drm_sched_init_args->num_rqs usage > > > > [...] > > > >> That might be the wisest thing to do; but let's hear if Tvrtko has an idea for > >> a hotfix first. > > > > -rc7 was released yesterday; even with a working hotfix today it'd be tricky to > > ensure the hotfix does not regress other drivers or workloads with 7.2 being > > just a few days ahead. > > > > I suggest to not wait and get the reverts ready. > > Yes reverts would be safer. I have them in a branch at > people.freedesktop.org/~tursulin/drm-intel drm-sched-fair-reverts, > mostly straightforward apart from one easy conflict in > amdgpu_xcp_release_sched. Smoke tested on Steam Deck looks fine. > > Having said that, the fix for min_vruntime handling I provided earlier > in the thread also looks fine in my testing and is simple. No > regressions found with synthetic unit test workloads or messing around > on the Steam Deck. But we need to wait to hear from Luke since I haven't > been able to repro his report locally yet. > > Regards, > > Tvrtko > > ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT 2026-08-11 17:59 ` Luke.Wildhardt @ 2026-08-12 11:37 ` Tvrtko Ursulin 2026-08-12 11:58 ` Philipp Stanner 2026-08-12 15:12 ` Luke.Wildhardt 0 siblings, 2 replies; 21+ messages in thread From: Tvrtko Ursulin @ 2026-08-12 11:37 UTC (permalink / raw) To: Luke.Wildhardt Cc: Danilo Krummrich, Philipp Stanner, phasta, matthew.brost, ckoenig.leichtzumerken, dri-devel, regressions, amd-gfx, linux-kernel On 11/08/2026 18:59, Luke.Wildhardt@proton.me wrote: > I want to be very clear here, I was working with Claude Opus on debugging this issue. I am not a good programmer and work in on the hardware side of things, so this is only code I've tested and not of my creation, just hoping to be helpful. It is really helpful and appreciated! Both that you found the regression and that you are helping debug and fix it. > It appears the freezing (and maybe what I'll call micro freezing instead of stuttering) was resolved after an issue was potentially identified in drivers/gpu/drm/scheduler/sched_rq.c > > I did a 20m play session attempting to reproduce the issue but was not able to. Previously, I could reproduce this on-demand. > > I will copy the writeup it created and the testing we did below. I am only claiming that the freezing had been seemingly resolved, not a particular fitness or correct identification of the issue, that is for developers to judge. > > Quotations are an output of a summery I had Claude create. > > " > THE INVARIANT > > drm_sched_entity_stats::vruntime is held absolute while the entity is linked > in the run queue tree, and relative to min_vruntime while it is idle. > save_vruntime() converts one way on the way out, restore_vruntime() the other > on the way back in. drm_sched_rq_add_entity() calls restore unconditionally, > which is only correct if the entity has definitely been through save. > > > THE WINDOW > > drm_sched_entity_pop_job(): > > spsc_queue_pop(&entity->job_queue); /* queue now empty */ > > drm_sched_rq_pop_entity(entity); /* takes entity->lock inside */ > > No lock held across those two lines. A concurrent push seeing the queue empty > gets first == true, calls add_entity, and restores a vruntime that has not yet > been saved -- adding min_vruntime to an already-absolute value. If the entity > is leftmost, get_min_vruntime() returns its own vruntime and it doubles. > > It persists: pop_entity then finds the pushed job, takes the next_job branch, > and update_vruntime() carries the inflated value forward. The entity sits far > right in the tree and is not selected until min_vruntime catches up, which > under sustained load may be effectively never. > > > EVIDENCE > > WARN_ON_ONCE(!RB_EMPTY_NODE(&entity->rb_tree_node)) at the top of > add_entity(), after the rq lock, fires reproducibly. > > That condition should be unreachable. add_entity() is reached only from the > "first job" branch of push_job(), and first is true only when the SPSC queue > was empty. A linked entity implies its last pop_entity() peeked non-empty. > Empty queue and still linked cannot both hold in any ordered execution. Fresh > entities are excluded -- entity_init() calls RB_CLEAR_NODE() unconditionally. > > Instrumented as a counter instead, with ring and comm logged: > > idle 10s: 0 races / 5662 restores > gameplay 60s: 8 races / 43315 restores > > [ 192.754604] vruntime race on ring gfx_0.0.0 (comm kwin_wayla:cs0) > [ 226.882135] vruntime race on ring gfx_0.0.0 (comm kwin_wayla:cs0) > [ 292.226032] vruntime race on ring gfx_0.0.0 (comm kwin_wayla:cs0) > > plus Xwayland:cs0 and vkd3d_queue on gfx, and many more on sdma0/sdma1. The > gfx_0.0.0 / kwin_wayla:cs0 entries are the compositor submission thread on the > graphics ring -- the process that dies when the desktop freezes. > > > CHANGE > > The principled fix is presumably to close the window by holding entity->lock > across the pop, which needs a _locked variant of drm_sched_rq_pop_entity(). > I did not attempt that. What I tested checks the invariant instead: a still- > linked entity never left, so its vruntime is already absolute and its tree > position valid. > > --- a/drivers/gpu/drm/scheduler/sched_rq.c > +++ b/drivers/gpu/drm/scheduler/sched_rq.c > @@ drm_sched_rq_add_entity > - ts = drm_sched_rq_get_min_vruntime(rq); > - ts = drm_sched_entity_restore_vruntime(entity, ts, rq->head_prio); > - drm_sched_rq_update_tree_locked(entity, rq, ts); > + if (RB_EMPTY_NODE(&entity->rb_tree_node)) { > + ts = drm_sched_rq_get_min_vruntime(rq); > + ts = drm_sched_entity_restore_vruntime(entity, ts, > + rq->head_prio); > + drm_sched_rq_update_tree_locked(entity, rq, ts); > + } > > On stock 7.2-rc7 with this applied, freezes are gone against a trigger that > previously reproduced on demand, and perceptible hitching appears gone too. > The race counter keeps incrementing (35 in the last session against 295744 > restores), so the window still opens at the same rate and is now absorbed -- > the timing has not merely shifted. Extended soak in progress. > " Good find! I almost feel obsolete. Perhaps our new AI overlords should make a pension fund out of the proceeds obtained by training their models on the decades of our work so us old programmers can safely retire. Jokes aside, I agree with the above analysis that a better fix would be to pull things under the lock, and interestingly, Philipp had a lock widening patch not so long ago but I don't remember what happened it or how exactly did it look. It would possibly have fixed this problem. Anyway, I have prepared a branch with the two fixes if you would be kind enough to give it a spin: https://cgit.freedesktop.org/~tursulin/drm-intel/log/?h=drm-sched-fair-fixes In my testing it all looks good. Fingers crossed. Regards, Tvrtko > On Monday, August 10th, 2026 at 9:40 AM, Luke.Wildhardt@proton.me <Luke.Wildhardt@proton.me> wrote: > >> Oops, forgot to reply all in my last message. >> >> I did a quick test of the patch before I left this morning, the freezing seems to gone, but when the GPU is under the same condition, the stuttering is more uniform if that makes sense. Granted this was only a 5 minute test but the freeze was rather reliable before, I can do more testing later. >> >> I know these things can be hard to describe, if that helps. If there's any profiling I could do or detailed logging that would help don't hesitate to ask. >> >> >> >> -------- Original Message -------- >> On Monday, 08/10/26 at 06:48 Tvrtko Ursulin <tvrtko.ursulin@igalia.com> wrote: >> >> On 10/08/2026 13:45, Danilo Krummrich wrote: >>> On Mon Aug 10, 2026 at 10:28 AM CEST, Philipp Stanner wrote: >>>>> Reverted Commits: >>>>> >>>>> d09339388b77 drm/sched: Remove drm_sched_init_args->num_rqs >>>>> 2833a0512b4c drm/sched: Remove drm_sched_init_args->num_rqs usage >>>>> 16e7698bc04d drm/sched: Embed run queue singleton into the scheduler >>>>> 77a6809f1dc3 drm/sched: Remove FIFO and RR and simplify to a single run queue >>>>> 45c211ddf92a drm/sched: Switch default policy to fair >>>>> 2462a0ce23b0 drm/amdgpu: Remove drm_sched_init_args->num_rqs usage >>> >>> [...] >>> >>>> That might be the wisest thing to do; but let's hear if Tvrtko has an idea for >>>> a hotfix first. >>> >>> -rc7 was released yesterday; even with a working hotfix today it'd be tricky to >>> ensure the hotfix does not regress other drivers or workloads with 7.2 being >>> just a few days ahead. >>> >>> I suggest to not wait and get the reverts ready. >> >> Yes reverts would be safer. I have them in a branch at >> people.freedesktop.org/~tursulin/drm-intel drm-sched-fair-reverts, >> mostly straightforward apart from one easy conflict in >> amdgpu_xcp_release_sched. Smoke tested on Steam Deck looks fine. >> >> Having said that, the fix for min_vruntime handling I provided earlier >> in the thread also looks fine in my testing and is simple. No >> regressions found with synthetic unit test workloads or messing around >> on the Steam Deck. But we need to wait to hear from Luke since I haven't >> been able to repro his report locally yet. >> >> Regards, >> >> Tvrtko >> >> ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT 2026-08-12 11:37 ` Tvrtko Ursulin @ 2026-08-12 11:58 ` Philipp Stanner 2026-08-12 12:29 ` Tvrtko Ursulin 2026-08-12 15:12 ` Luke.Wildhardt 1 sibling, 1 reply; 21+ messages in thread From: Philipp Stanner @ 2026-08-12 11:58 UTC (permalink / raw) To: Tvrtko Ursulin, Luke.Wildhardt Cc: Danilo Krummrich, phasta, matthew.brost, ckoenig.leichtzumerken, dri-devel, regressions, amd-gfx, linux-kernel On Wed, 2026-08-12 at 12:37 +0100, Tvrtko Ursulin wrote: > Jokes aside, I agree with the above analysis that a better fix would be > to pull things under the lock, and interestingly, Philipp had a lock > widening patch not so long ago but I don't remember what happened it or > how exactly did it look. It would possibly have fixed this problem. https://lore.kernel.org/dri-devel/a0d969408a5a55dfb6d0b3e65906fd7bbf3eee1c.camel@mailbox.org/ TBH I got discouraged and increasingly dissatisfied with the state of our community, because whenever I try to fix things that are *obviously* broken and incorrect (and which were usually not broken by me), with relatively simple patches, I have to work uphill and only receive pushback. Reviewers and supporters are usually absent unless they want to have more features or performance-hacks added, though it is apparent that many indeed read what we are doing. Should things get better it's usually not acknowledged, and should things go wrong one runs danger of getting blamed with hindsight-bias. Had my patch landed and (speculatively) prevented this regression, no one would have noticed, but Phoronix would write some nice article about the FPS gain in video games, achieved by Tvrtko Ursulin :) Ah well. My vision for drm_sched has been that we absolutely need to start establishing strict computer science standards and have to prioritze formal correctness, simplicity and documentation over features and *especially* performance. Though I maybe should have tried to be more polite and diplomatic about it, the evidence that I see is that the DRM community does not actually want the above. At least I have to conclude that from the lack of encouraging participation and the presence of discouragement. So if I want to add a spinlock for something obviously broken, the burden of proof is on me, not on the one who broke it. If I want to prevent someone from adding a 4th drm_sched_start_x() function for performance-reasons, then I receive not-so-nice emails. Folks here seem to want to continue like they did for the last 15 years. If anyone here reads this who *doesn't* want to continue like we did for the last 15 years, then I would kindly ask for public participation and help with moving things towards the right direction. Simple emails with "I support this idea, +1", would almost be enough. It would seem that these recent event might proof that I was right about our fundamental issues. P. ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT 2026-08-12 11:58 ` Philipp Stanner @ 2026-08-12 12:29 ` Tvrtko Ursulin 2026-09-16 7:37 ` Philipp Stanner 0 siblings, 1 reply; 21+ messages in thread From: Tvrtko Ursulin @ 2026-08-12 12:29 UTC (permalink / raw) To: phasta, Luke.Wildhardt Cc: Danilo Krummrich, matthew.brost, ckoenig.leichtzumerken, dri-devel, regressions, amd-gfx, linux-kernel On 12/08/2026 12:58, Philipp Stanner wrote: > On Wed, 2026-08-12 at 12:37 +0100, Tvrtko Ursulin wrote: >> Jokes aside, I agree with the above analysis that a better fix would be >> to pull things under the lock, and interestingly, Philipp had a lock >> widening patch not so long ago but I don't remember what happened it or >> how exactly did it look. It would possibly have fixed this problem. > > https://lore.kernel.org/dri-devel/a0d969408a5a55dfb6d0b3e65906fd7bbf3eee1c.camel@mailbox.org/ > > TBH I got discouraged and increasingly dissatisfied with the state of > our community, because whenever I try to fix things that are > *obviously* broken and incorrect (and which were usually not broken by > me), with relatively simple patches, I have to work uphill and only > receive pushback. > > Reviewers and supporters are usually absent unless they want to have > more features or performance-hacks added, though it is apparent that > many indeed read what we are doing. > > Should things get better it's usually not acknowledged, and should > things go wrong one runs danger of getting blamed with hindsight-bias. > > Had my patch landed and (speculatively) prevented this regression, no > one would have noticed, but Phoronix would write some nice article > about the FPS gain in video games, achieved by Tvrtko Ursulin :) Not sure if you are calling me out here or not, since apart from a direct mention, the lore link above is also a reply to my email, so for the record, I did give my ack for those patches. Plus I spent time testing them. So if you were in fact calling me out, I don't see how that is warranted. And on the wider topic I also never shied away from the non-glamorous work. Going back to your linked patch series, now that I re-read it and reminded myself, it wouldn't have solved the regression from this thread since it left drm_sched_rq_pop_entity() outside the lock. But I did agree even then it was a step in the right direction. So if you respin it (last patch seemed buggy according to sashiko), maybe replace with my version of the completion removal (what happened with that one?), ideally fold the patch which removes the double re-lock cycle into the actual locking change (I think that's better), my ack still stands. But I also still think ack from AMD is needed as well since the changed code paths are by large used from amdgpu. Regards, Tvrtko > > Ah well. > > My vision for drm_sched has been that we absolutely need to start > establishing strict computer science standards and have to prioritze > formal correctness, simplicity and documentation over features and > *especially* performance. > > Though I maybe should have tried to be more polite and diplomatic about > it, the evidence that I see is that the DRM community does not actually > want the above. At least I have to conclude that from the lack of > encouraging participation and the presence of discouragement. > > So if I want to add a spinlock for something obviously broken, the > burden of proof is on me, not on the one who broke it. If I want to > prevent someone from adding a 4th drm_sched_start_x() function for > performance-reasons, then I receive not-so-nice emails. > > Folks here seem to want to continue like they did for the last 15 > years. > > If anyone here reads this who *doesn't* want to continue like we did > for the last 15 years, then I would kindly ask for public participation > and help with moving things towards the right direction. Simple emails > with "I support this idea, +1", would almost be enough. > > > It would seem that these recent event might proof that I was right > about our fundamental issues. > > > P. ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT 2026-08-12 12:29 ` Tvrtko Ursulin @ 2026-09-16 7:37 ` Philipp Stanner 0 siblings, 0 replies; 21+ messages in thread From: Philipp Stanner @ 2026-09-16 7:37 UTC (permalink / raw) To: Tvrtko Ursulin, phasta, Luke.Wildhardt Cc: Danilo Krummrich, matthew.brost, ckoenig.leichtzumerken, dri-devel, regressions, amd-gfx, linux-kernel On Wed, 2026-08-12 at 13:29 +0100, Tvrtko Ursulin wrote: > Not sure if you are calling me out here or not, since apart from a > direct mention, the lore link above is also a reply to my email, so for > the record, I did give my ack for those patches. Plus I spent time > testing them. So if you were in fact calling me out, I don't see how > that is warranted. And on the wider topic I also never shied away from > the non-glamorous work. I was addressing what I perceive to be the general problems within DRM – I still believe the general points are valid, but want to stress that it's a collective problem in which we all play a role, though no one is the protagonist. I mentioned the media outlet (which quoted your name) as part of the problem because there are external cheerleaders for performance (whose unconditional priority is *the* reason for the design mistakes in drm_sched) and features. Nevertheless, I apologize for the rough tone and want to emphasize that you're indeed the most engaged contributor and have improved many things. The unit tests alone are worth their weight in silver. The wider, problematic patterns in DRM might be worth addressing with the appropriate tenderness in a suitable format; I will try to collect my thoughts on that. Regards, Philipp ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT 2026-08-12 11:37 ` Tvrtko Ursulin 2026-08-12 11:58 ` Philipp Stanner @ 2026-08-12 15:12 ` Luke.Wildhardt 2026-08-13 6:42 ` Luke.Wildhardt 1 sibling, 1 reply; 21+ messages in thread From: Luke.Wildhardt @ 2026-08-12 15:12 UTC (permalink / raw) To: Tvrtko Ursulin Cc: Danilo Krummrich, Philipp Stanner, phasta, matthew.brost, ckoenig.leichtzumerken, dri-devel, regressions, amd-gfx, linux-kernel I will test this later tonight when I get home and report my findings. If all is well, I do have a 6900XT I can dig out and cross validate. Also, I attempted to find a synthetic way to induce the condition; the only slight way I found to do so was running furmark tail, but the desktop never froze, only the desktop seemed to get mildly choppy. I have not yet checked against the FIFO scheduler. I was trying to find a more objective way to trigger the issue, but it appears in my case I got rather (un)lucky with the specific game I was playing the reproduces it so well. No problem for helping, I reported the issue and since I am capable of helping at the very least test a fix and reporting on it, I feel I have an obligation to, especially since you haven't been able to reproduce the issue on your end. -------- Original Message -------- On Wednesday, 08/12/26 at 04:38 Tvrtko Ursulin <tvrtko.ursulin@igalia.com> wrote: On 11/08/2026 18:59, Luke.Wildhardt@proton.me wrote: > I want to be very clear here, I was working with Claude Opus on debugging this issue. I am not a good programmer and work in on the hardware side of things, so this is only code I've tested and not of my creation, just hoping to be helpful. It is really helpful and appreciated! Both that you found the regression and that you are helping debug and fix it. > It appears the freezing (and maybe what I'll call micro freezing instead of stuttering) was resolved after an issue was potentially identified in drivers/gpu/drm/scheduler/sched_rq.c > > I did a 20m play session attempting to reproduce the issue but was not able to. Previously, I could reproduce this on-demand. > > I will copy the writeup it created and the testing we did below. I am only claiming that the freezing had been seemingly resolved, not a particular fitness or correct identification of the issue, that is for developers to judge. > > Quotations are an output of a summery I had Claude create. > > " > THE INVARIANT > > drm_sched_entity_stats::vruntime is held absolute while the entity is linked > in the run queue tree, and relative to min_vruntime while it is idle. > save_vruntime() converts one way on the way out, restore_vruntime() the other > on the way back in. drm_sched_rq_add_entity() calls restore unconditionally, > which is only correct if the entity has definitely been through save. > > > THE WINDOW > > drm_sched_entity_pop_job(): > > spsc_queue_pop(&entity->job_queue); /* queue now empty */ > > drm_sched_rq_pop_entity(entity); /* takes entity->lock inside */ > > No lock held across those two lines. A concurrent push seeing the queue empty > gets first == true, calls add_entity, and restores a vruntime that has not yet > been saved -- adding min_vruntime to an already-absolute value. If the entity > is leftmost, get_min_vruntime() returns its own vruntime and it doubles. > > It persists: pop_entity then finds the pushed job, takes the next_job branch, > and update_vruntime() carries the inflated value forward. The entity sits far > right in the tree and is not selected until min_vruntime catches up, which > under sustained load may be effectively never. > > > EVIDENCE > > WARN_ON_ONCE(!RB_EMPTY_NODE(&entity->rb_tree_node)) at the top of > add_entity(), after the rq lock, fires reproducibly. > > That condition should be unreachable. add_entity() is reached only from the > "first job" branch of push_job(), and first is true only when the SPSC queue > was empty. A linked entity implies its last pop_entity() peeked non-empty. > Empty queue and still linked cannot both hold in any ordered execution. Fresh > entities are excluded -- entity_init() calls RB_CLEAR_NODE() unconditionally. > > Instrumented as a counter instead, with ring and comm logged: > > idle 10s: 0 races / 5662 restores > gameplay 60s: 8 races / 43315 restores > > [ 192.754604] vruntime race on ring gfx_0.0.0 (comm kwin_wayla:cs0) > [ 226.882135] vruntime race on ring gfx_0.0.0 (comm kwin_wayla:cs0) > [ 292.226032] vruntime race on ring gfx_0.0.0 (comm kwin_wayla:cs0) > > plus Xwayland:cs0 and vkd3d_queue on gfx, and many more on sdma0/sdma1. The > gfx_0.0.0 / kwin_wayla:cs0 entries are the compositor submission thread on the > graphics ring -- the process that dies when the desktop freezes. > > > CHANGE > > The principled fix is presumably to close the window by holding entity->lock > across the pop, which needs a _locked variant of drm_sched_rq_pop_entity(). > I did not attempt that. What I tested checks the invariant instead: a still- > linked entity never left, so its vruntime is already absolute and its tree > position valid. > > --- a/drivers/gpu/drm/scheduler/sched_rq.c > +++ b/drivers/gpu/drm/scheduler/sched_rq.c > @@ drm_sched_rq_add_entity > - ts = drm_sched_rq_get_min_vruntime(rq); > - ts = drm_sched_entity_restore_vruntime(entity, ts, rq->head_prio); > - drm_sched_rq_update_tree_locked(entity, rq, ts); > + if (RB_EMPTY_NODE(&entity->rb_tree_node)) { > + ts = drm_sched_rq_get_min_vruntime(rq); > + ts = drm_sched_entity_restore_vruntime(entity, ts, > + rq->head_prio); > + drm_sched_rq_update_tree_locked(entity, rq, ts); > + } > > On stock 7.2-rc7 with this applied, freezes are gone against a trigger that > previously reproduced on demand, and perceptible hitching appears gone too. > The race counter keeps incrementing (35 in the last session against 295744 > restores), so the window still opens at the same rate and is now absorbed -- > the timing has not merely shifted. Extended soak in progress. > " Good find! I almost feel obsolete. Perhaps our new AI overlords should make a pension fund out of the proceeds obtained by training their models on the decades of our work so us old programmers can safely retire. Jokes aside, I agree with the above analysis that a better fix would be to pull things under the lock, and interestingly, Philipp had a lock widening patch not so long ago but I don't remember what happened it or how exactly did it look. It would possibly have fixed this problem. Anyway, I have prepared a branch with the two fixes if you would be kind enough to give it a spin: https://cgit.freedesktop.org/~tursulin/drm-intel/log/?h=drm-sched-fair-fixes In my testing it all looks good. Fingers crossed. Regards, Tvrtko > On Monday, August 10th, 2026 at 9:40 AM, Luke.Wildhardt@proton.me <Luke.Wildhardt@proton.me> wrote: > >> Oops, forgot to reply all in my last message. >> >> I did a quick test of the patch before I left this morning, the freezing seems to gone, but when the GPU is under the same condition, the stuttering is more uniform if that makes sense. Granted this was only a 5 minute test but the freeze was rather reliable before, I can do more testing later. >> >> I know these things can be hard to describe, if that helps. If there's any profiling I could do or detailed logging that would help don't hesitate to ask. >> >> >> >> -------- Original Message -------- >> On Monday, 08/10/26 at 06:48 Tvrtko Ursulin <tvrtko.ursulin@igalia.com> wrote: >> >> On 10/08/2026 13:45, Danilo Krummrich wrote: >>> On Mon Aug 10, 2026 at 10:28 AM CEST, Philipp Stanner wrote: >>>>> Reverted Commits: >>>>> >>>>> d09339388b77 drm/sched: Remove drm_sched_init_args->num_rqs >>>>> 2833a0512b4c drm/sched: Remove drm_sched_init_args->num_rqs usage >>>>> 16e7698bc04d drm/sched: Embed run queue singleton into the scheduler >>>>> 77a6809f1dc3 drm/sched: Remove FIFO and RR and simplify to a single run queue >>>>> 45c211ddf92a drm/sched: Switch default policy to fair >>>>> 2462a0ce23b0 drm/amdgpu: Remove drm_sched_init_args->num_rqs usage >>> >>> [...] >>> >>>> That might be the wisest thing to do; but let's hear if Tvrtko has an idea for >>>> a hotfix first. >>> >>> -rc7 was released yesterday; even with a working hotfix today it'd be tricky to >>> ensure the hotfix does not regress other drivers or workloads with 7.2 being >>> just a few days ahead. >>> >>> I suggest to not wait and get the reverts ready. >> >> Yes reverts would be safer. I have them in a branch at >> people.freedesktop.org/~tursulin/drm-intel drm-sched-fair-reverts, >> mostly straightforward apart from one easy conflict in >> amdgpu_xcp_release_sched. Smoke tested on Steam Deck looks fine. >> >> Having said that, the fix for min_vruntime handling I provided earlier >> in the thread also looks fine in my testing and is simple. No >> regressions found with synthetic unit test workloads or messing around >> on the Steam Deck. But we need to wait to hear from Luke since I haven't >> been able to repro his report locally yet. >> >> Regards, >> >> Tvrtko >> >> ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT 2026-08-12 15:12 ` Luke.Wildhardt @ 2026-08-13 6:42 ` Luke.Wildhardt 2026-08-13 7:58 ` Tvrtko Ursulin 0 siblings, 1 reply; 21+ messages in thread From: Luke.Wildhardt @ 2026-08-13 6:42 UTC (permalink / raw) To: Tvrtko Ursulin Cc: Danilo Krummrich, Philipp Stanner, phasta, matthew.brost, ckoenig.leichtzumerken, dri-devel, regressions, amd-gfx, linux-kernel Ok, after a lot of A/B testing, it appears that trying your "drm-intel/drm-sched-fair-fixed" branch, the stuttering is still there. It's definitely paced differently though, I will attach a video. It's "choppier" when it happens. I tested in a few different areas / times of day, in case you notice the scenery isn't the same; this is representative of what I experienced in other places. I will note, that I tried playing 4 times, 3/4 times, shortly after boot. 1/4 times I left my desktop to idle for about 30 minutes while I did something else, and I didn't encounter the issue at all in 10 minutes. Not sure how significant that is, or if it's just noise. https://www.youtube.com/watch?v=GWyIVEuhooM To make sure I wasn't insane, I went back to the patch Claude gave me, applied against a clean 7.2.0-rc7. At least again in a few sessions after boot, no issue. I'll post that exact code below. If you need any other info from me let me know. --- a/drivers/gpu/drm/scheduler/sched_rq.c +++ b/drivers/gpu/drm/scheduler/sched_rq.c @@ -2,6 +2,7 @@ /* Copyright 2015 Advanced Micro Devices, Inc. */ /* Copyright (c) 2025 Valve Corporation */ +#include <linux/moduleparam.h> #include <linux/rbtree.h> #include <drm/drm_print.h> @@ -9,6 +10,32 @@ #include "sched_internal.h" +/* + * Diagnostic counters. Not for submission. + * + * dbg_pop_stayed - pops where the entity had another job queued and so stayed + * in the tree. No save/restore of vruntime occurs. + * dbg_pop_left - pops where the entity queue drained, so it left the tree + * and drm_sched_entity_save_vruntime() ran. Only this path + * arms a later restore. + * dbg_add_restore - calls to drm_sched_rq_add_entity(), i.e. restores. + * dbg_add_race - restores where the entity was still linked in the tree, + * meaning the vruntime being restored is still absolute. + * + * Incremented under rq->lock, so exact per scheduler and only mildly lossy + * when summed across rings. Writable so they can be reset between runs: + * echo 0 > /sys/module/gpu_sched/parameters/dbg_pop_left + */ +static unsigned long dbg_pop_stayed; +static unsigned long dbg_pop_left; +static unsigned long dbg_add_restore; +static unsigned long dbg_add_race; + +module_param(dbg_pop_stayed, ulong, 0644); +module_param(dbg_pop_left, ulong, 0644); +module_param(dbg_add_restore, ulong, 0644); +module_param(dbg_add_race, ulong, 0644); + static __always_inline bool drm_sched_entity_compare_before(struct rb_node *a, const struct rb_node *b) { @@ -266,14 +293,40 @@ sched = container_of(rq, typeof(*sched), rq); spin_lock(&rq->lock); + dbg_add_restore++; + if (!RB_EMPTY_NODE(&entity->rb_tree_node)) { + dbg_add_race++; + pr_warn_ratelimited("drm_sched: vruntime race on ring %s (comm %s)\n", + sched->name, current->comm); + } + if (list_empty(&entity->list)) { atomic_inc(sched->score); list_add_tail(&entity->list, &rq->entities); } - ts = drm_sched_rq_get_min_vruntime(rq); - ts = drm_sched_entity_restore_vruntime(entity, ts, rq->head_prio); - drm_sched_rq_update_tree_locked(entity, rq, ts); + /* + * Only restore the vruntime if the entity actually left the run queue. + * + * drm_sched_entity_pop_job() dequeues the last job and only afterwards + * calls drm_sched_rq_pop_entity(), which is where the entity is removed + * from the tree and drm_sched_entity_save_vruntime() converts its + * vruntime to min_vruntime-relative form. A push landing in that window + * sees an empty queue, takes the "first job" path to here, and would + * restore a vruntime which is still absolute -- adding min_vruntime to + * it a second time. If the entity is also the leftmost one, + * drm_sched_rq_get_min_vruntime() returns the entity's own vruntime and + * the value is doubled, placing it far to the right of the tree where it + * will not be selected again until the run queue catches up. + * + * A still-linked entity never left, so its vruntime is already absolute + * and its tree position is valid. The concurrent pop will update both. + */ + if (RB_EMPTY_NODE(&entity->rb_tree_node)) { + ts = drm_sched_rq_get_min_vruntime(rq); + ts = drm_sched_entity_restore_vruntime(entity, ts, rq->head_prio); + drm_sched_rq_update_tree_locked(entity, rq, ts); + } spin_unlock(&rq->lock); spin_unlock(&entity->lock); @@ -330,11 +383,13 @@ if (next_job) { ktime_t ts; + dbg_pop_stayed++; ts = drm_sched_entity_update_vruntime(entity); drm_sched_rq_update_tree_locked(entity, rq, ts); } else { ktime_t min_vruntime; + dbg_pop_left++; drm_sched_rq_remove_tree_locked(entity, rq); min_vruntime = drm_sched_rq_get_min_vruntime(rq); drm_sched_entity_save_vruntime(entity, min_vruntime); ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT 2026-08-13 6:42 ` Luke.Wildhardt @ 2026-08-13 7:58 ` Tvrtko Ursulin 2026-08-13 14:27 ` Luke.Wildhardt 0 siblings, 1 reply; 21+ messages in thread From: Tvrtko Ursulin @ 2026-08-13 7:58 UTC (permalink / raw) To: Luke.Wildhardt Cc: Danilo Krummrich, Philipp Stanner, phasta, matthew.brost, ckoenig.leichtzumerken, dri-devel, regressions, amd-gfx, linux-kernel On 13/08/2026 07:42, Luke.Wildhardt@proton.me wrote: > Ok, after a lot of A/B testing, it appears that trying your "drm-intel/drm-sched-fair-fixed" branch, the stuttering is still there. It's definitely paced differently though, I will attach a video. It's "choppier" when it happens. I tested in a few different areas / times of day, in case you notice the scenery isn't the same; this is representative of what I experienced in other places. I will note, that I tried playing 4 times, 3/4 times, shortly after boot. 1/4 times I left my desktop to idle for about 30 minutes while I did something else, and I didn't encounter the issue at all in 10 minutes. Not sure how significant that is, or if it's just noise. > > https://www.youtube.com/watch?v=GWyIVEuhooM > > To make sure I wasn't insane, I went back to the patch Claude gave me, applied against a clean 7.2.0-rc7. At least again in a few sessions after boot, no issue. I'll post that exact code below. Yes, I had a brain fart yesterday and had only pulled the lock out on the pop side. I have now pushed the updated branch, with both the add and pop side made symmetric in lock taking aspect. If you could pull and re-test once more that would be great. Regards, Tvrtko > If you need any other info from me let me know. > > > --- a/drivers/gpu/drm/scheduler/sched_rq.c > +++ b/drivers/gpu/drm/scheduler/sched_rq.c > @@ -2,6 +2,7 @@ > /* Copyright 2015 Advanced Micro Devices, Inc. */ > /* Copyright (c) 2025 Valve Corporation */ > > +#include <linux/moduleparam.h> > #include <linux/rbtree.h> > > #include <drm/drm_print.h> > @@ -9,6 +10,32 @@ > > #include "sched_internal.h" > > +/* > + * Diagnostic counters. Not for submission. > + * > + * dbg_pop_stayed - pops where the entity had another job queued and so stayed > + * in the tree. No save/restore of vruntime occurs. > + * dbg_pop_left - pops where the entity queue drained, so it left the tree > + * and drm_sched_entity_save_vruntime() ran. Only this path > + * arms a later restore. > + * dbg_add_restore - calls to drm_sched_rq_add_entity(), i.e. restores. > + * dbg_add_race - restores where the entity was still linked in the tree, > + * meaning the vruntime being restored is still absolute. > + * > + * Incremented under rq->lock, so exact per scheduler and only mildly lossy > + * when summed across rings. Writable so they can be reset between runs: > + * echo 0 > /sys/module/gpu_sched/parameters/dbg_pop_left > + */ > +static unsigned long dbg_pop_stayed; > +static unsigned long dbg_pop_left; > +static unsigned long dbg_add_restore; > +static unsigned long dbg_add_race; > + > +module_param(dbg_pop_stayed, ulong, 0644); > +module_param(dbg_pop_left, ulong, 0644); > +module_param(dbg_add_restore, ulong, 0644); > +module_param(dbg_add_race, ulong, 0644); > + > static __always_inline bool > drm_sched_entity_compare_before(struct rb_node *a, const struct rb_node *b) > { > @@ -266,14 +293,40 @@ > sched = container_of(rq, typeof(*sched), rq); > spin_lock(&rq->lock); > > + dbg_add_restore++; > + if (!RB_EMPTY_NODE(&entity->rb_tree_node)) { > + dbg_add_race++; > + pr_warn_ratelimited("drm_sched: vruntime race on ring %s (comm %s)\n", > + sched->name, current->comm); > + } > + > if (list_empty(&entity->list)) { > atomic_inc(sched->score); > list_add_tail(&entity->list, &rq->entities); > } > > - ts = drm_sched_rq_get_min_vruntime(rq); > - ts = drm_sched_entity_restore_vruntime(entity, ts, rq->head_prio); > - drm_sched_rq_update_tree_locked(entity, rq, ts); > + /* > + * Only restore the vruntime if the entity actually left the run queue. > + * > + * drm_sched_entity_pop_job() dequeues the last job and only afterwards > + * calls drm_sched_rq_pop_entity(), which is where the entity is removed > + * from the tree and drm_sched_entity_save_vruntime() converts its > + * vruntime to min_vruntime-relative form. A push landing in that window > + * sees an empty queue, takes the "first job" path to here, and would > + * restore a vruntime which is still absolute -- adding min_vruntime to > + * it a second time. If the entity is also the leftmost one, > + * drm_sched_rq_get_min_vruntime() returns the entity's own vruntime and > + * the value is doubled, placing it far to the right of the tree where it > + * will not be selected again until the run queue catches up. > + * > + * A still-linked entity never left, so its vruntime is already absolute > + * and its tree position is valid. The concurrent pop will update both. > + */ > + if (RB_EMPTY_NODE(&entity->rb_tree_node)) { > + ts = drm_sched_rq_get_min_vruntime(rq); > + ts = drm_sched_entity_restore_vruntime(entity, ts, rq->head_prio); > + drm_sched_rq_update_tree_locked(entity, rq, ts); > + } > > spin_unlock(&rq->lock); > spin_unlock(&entity->lock); > @@ -330,11 +383,13 @@ > if (next_job) { > ktime_t ts; > > + dbg_pop_stayed++; > ts = drm_sched_entity_update_vruntime(entity); > drm_sched_rq_update_tree_locked(entity, rq, ts); > } else { > ktime_t min_vruntime; > > + dbg_pop_left++; > drm_sched_rq_remove_tree_locked(entity, rq); > min_vruntime = drm_sched_rq_get_min_vruntime(rq); > drm_sched_entity_save_vruntime(entity, min_vruntime); > > > ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT 2026-08-13 7:58 ` Tvrtko Ursulin @ 2026-08-13 14:27 ` Luke.Wildhardt 2026-08-14 6:49 ` Luke.Wildhardt 2026-08-14 8:32 ` Feng, Kenneth 0 siblings, 2 replies; 21+ messages in thread From: Luke.Wildhardt @ 2026-08-13 14:27 UTC (permalink / raw) To: Tvrtko Ursulin Cc: Danilo Krummrich, Philipp Stanner, phasta, matthew.brost, ckoenig.leichtzumerken, dri-devel, regressions, amd-gfx, linux-kernel Looks good, in the 10 minutes I spent this morning doing a quick test, I wasn't able to reproduce the failure. I will do extended testing tonight. I will report back later with more details, but given the issue was typically reproducible on-demand, this is looking very promising. Thanks for all the hard work everyone! On Thursday, August 13th, 2026 at 12:58 AM, Tvrtko Ursulin <tvrtko.ursulin@igalia.com> wrote: > > On 13/08/2026 07:42, Luke.Wildhardt@proton.me wrote: > > Ok, after a lot of A/B testing, it appears that trying your "drm-intel/drm-sched-fair-fixed" branch, the stuttering is still there. It's definitely paced differently though, I will attach a video. It's "choppier" when it happens. I tested in a few different areas / times of day, in case you notice the scenery isn't the same; this is representative of what I experienced in other places. I will note, that I tried playing 4 times, 3/4 times, shortly after boot. 1/4 times I left my desktop to idle for about 30 minutes while I did something else, and I didn't encounter the issue at all in 10 minutes. Not sure how significant that is, or if it's just noise. > > > > https://www.youtube.com/watch?v=GWyIVEuhooM > > > > To make sure I wasn't insane, I went back to the patch Claude gave me, applied against a clean 7.2.0-rc7. At least again in a few sessions after boot, no issue. I'll post that exact code below. > > Yes, I had a brain fart yesterday and had only pulled the lock out on > the pop side. I have now pushed the updated branch, with both the add > and pop side made symmetric in lock taking aspect. > > If you could pull and re-test once more that would be great. > > Regards, > > Tvrtko > > > If you need any other info from me let me know. > > > > > > --- a/drivers/gpu/drm/scheduler/sched_rq.c > > +++ b/drivers/gpu/drm/scheduler/sched_rq.c > > @@ -2,6 +2,7 @@ > > /* Copyright 2015 Advanced Micro Devices, Inc. */ > > /* Copyright (c) 2025 Valve Corporation */ > > > > +#include <linux/moduleparam.h> > > #include <linux/rbtree.h> > > > > #include <drm/drm_print.h> > > @@ -9,6 +10,32 @@ > > > > #include "sched_internal.h" > > > > +/* > > + * Diagnostic counters. Not for submission. > > + * > > + * dbg_pop_stayed - pops where the entity had another job queued and so stayed > > + * in the tree. No save/restore of vruntime occurs. > > + * dbg_pop_left - pops where the entity queue drained, so it left the tree > > + * and drm_sched_entity_save_vruntime() ran. Only this path > > + * arms a later restore. > > + * dbg_add_restore - calls to drm_sched_rq_add_entity(), i.e. restores. > > + * dbg_add_race - restores where the entity was still linked in the tree, > > + * meaning the vruntime being restored is still absolute. > > + * > > + * Incremented under rq->lock, so exact per scheduler and only mildly lossy > > + * when summed across rings. Writable so they can be reset between runs: > > + * echo 0 > /sys/module/gpu_sched/parameters/dbg_pop_left > > + */ > > +static unsigned long dbg_pop_stayed; > > +static unsigned long dbg_pop_left; > > +static unsigned long dbg_add_restore; > > +static unsigned long dbg_add_race; > > + > > +module_param(dbg_pop_stayed, ulong, 0644); > > +module_param(dbg_pop_left, ulong, 0644); > > +module_param(dbg_add_restore, ulong, 0644); > > +module_param(dbg_add_race, ulong, 0644); > > + > > static __always_inline bool > > drm_sched_entity_compare_before(struct rb_node *a, const struct rb_node *b) > > { > > @@ -266,14 +293,40 @@ > > sched = container_of(rq, typeof(*sched), rq); > > spin_lock(&rq->lock); > > > > + dbg_add_restore++; > > + if (!RB_EMPTY_NODE(&entity->rb_tree_node)) { > > + dbg_add_race++; > > + pr_warn_ratelimited("drm_sched: vruntime race on ring %s (comm %s)\n", > > + sched->name, current->comm); > > + } > > + > > if (list_empty(&entity->list)) { > > atomic_inc(sched->score); > > list_add_tail(&entity->list, &rq->entities); > > } > > > > - ts = drm_sched_rq_get_min_vruntime(rq); > > - ts = drm_sched_entity_restore_vruntime(entity, ts, rq->head_prio); > > - drm_sched_rq_update_tree_locked(entity, rq, ts); > > + /* > > + * Only restore the vruntime if the entity actually left the run queue. > > + * > > + * drm_sched_entity_pop_job() dequeues the last job and only afterwards > > + * calls drm_sched_rq_pop_entity(), which is where the entity is removed > > + * from the tree and drm_sched_entity_save_vruntime() converts its > > + * vruntime to min_vruntime-relative form. A push landing in that window > > + * sees an empty queue, takes the "first job" path to here, and would > > + * restore a vruntime which is still absolute -- adding min_vruntime to > > + * it a second time. If the entity is also the leftmost one, > > + * drm_sched_rq_get_min_vruntime() returns the entity's own vruntime and > > + * the value is doubled, placing it far to the right of the tree where it > > + * will not be selected again until the run queue catches up. > > + * > > + * A still-linked entity never left, so its vruntime is already absolute > > + * and its tree position is valid. The concurrent pop will update both. > > + */ > > + if (RB_EMPTY_NODE(&entity->rb_tree_node)) { > > + ts = drm_sched_rq_get_min_vruntime(rq); > > + ts = drm_sched_entity_restore_vruntime(entity, ts, rq->head_prio); > > + drm_sched_rq_update_tree_locked(entity, rq, ts); > > + } > > > > spin_unlock(&rq->lock); > > spin_unlock(&entity->lock); > > @@ -330,11 +383,13 @@ > > if (next_job) { > > ktime_t ts; > > > > + dbg_pop_stayed++; > > ts = drm_sched_entity_update_vruntime(entity); > > drm_sched_rq_update_tree_locked(entity, rq, ts); > > } else { > > ktime_t min_vruntime; > > > > + dbg_pop_left++; > > drm_sched_rq_remove_tree_locked(entity, rq); > > min_vruntime = drm_sched_rq_get_min_vruntime(rq); > > drm_sched_entity_save_vruntime(entity, min_vruntime); > > > > > > > > ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT 2026-08-13 14:27 ` Luke.Wildhardt @ 2026-08-14 6:49 ` Luke.Wildhardt 2026-08-14 8:32 ` Feng, Kenneth 1 sibling, 0 replies; 21+ messages in thread From: Luke.Wildhardt @ 2026-08-14 6:49 UTC (permalink / raw) To: Tvrtko Ursulin Cc: Danilo Krummrich, Philipp Stanner, phasta, matthew.brost, ckoenig.leichtzumerken, dri-devel, regressions, amd-gfx, linux-kernel Did a 90 min play session across the problematic titles, and also let the game run in the background while I did other things on the desktop (Youtube, web browsing, Blender). No issues, seems to work as good as FIFO did on previous kernels. To reiterate, the issue at least for me was extremely easy to reproduce so I would be confident in calling this fixed (from my perspective). I want to say thanks to everyone for the assistance and responsiveness. I know these things can be stressful at the tail end of a RC. If you need other testing done or feedback on specific activities, other iterations, feel free to reach out, I don't mind. Sent with Proton Mail secure email. On Thursday, August 13th, 2026 at 7:27 AM, Luke.Wildhardt@proton.me <Luke.Wildhardt@proton.me> wrote: > Looks good, in the 10 minutes I spent this morning doing a quick test, I wasn't able to reproduce the failure. I will do extended testing tonight. > > I will report back later with more details, but given the issue was typically reproducible on-demand, this is looking very promising. > > Thanks for all the hard work everyone! > > On Thursday, August 13th, 2026 at 12:58 AM, Tvrtko Ursulin <tvrtko.ursulin@igalia.com> wrote: > > > > > On 13/08/2026 07:42, Luke.Wildhardt@proton.me wrote: > > > Ok, after a lot of A/B testing, it appears that trying your "drm-intel/drm-sched-fair-fixed" branch, the stuttering is still there. It's definitely paced differently though, I will attach a video. It's "choppier" when it happens. I tested in a few different areas / times of day, in case you notice the scenery isn't the same; this is representative of what I experienced in other places. I will note, that I tried playing 4 times, 3/4 times, shortly after boot. 1/4 times I left my desktop to idle for about 30 minutes while I did something else, and I didn't encounter the issue at all in 10 minutes. Not sure how significant that is, or if it's just noise. > > > > > > https://www.youtube.com/watch?v=GWyIVEuhooM > > > > > > To make sure I wasn't insane, I went back to the patch Claude gave me, applied against a clean 7.2.0-rc7. At least again in a few sessions after boot, no issue. I'll post that exact code below. > > > > Yes, I had a brain fart yesterday and had only pulled the lock out on > > the pop side. I have now pushed the updated branch, with both the add > > and pop side made symmetric in lock taking aspect. > > > > If you could pull and re-test once more that would be great. > > > > Regards, > > > > Tvrtko > > > > > If you need any other info from me let me know. > > > > > > > > > --- a/drivers/gpu/drm/scheduler/sched_rq.c > > > +++ b/drivers/gpu/drm/scheduler/sched_rq.c > > > @@ -2,6 +2,7 @@ > > > /* Copyright 2015 Advanced Micro Devices, Inc. */ > > > /* Copyright (c) 2025 Valve Corporation */ > > > > > > +#include <linux/moduleparam.h> > > > #include <linux/rbtree.h> > > > > > > #include <drm/drm_print.h> > > > @@ -9,6 +10,32 @@ > > > > > > #include "sched_internal.h" > > > > > > +/* > > > + * Diagnostic counters. Not for submission. > > > + * > > > + * dbg_pop_stayed - pops where the entity had another job queued and so stayed > > > + * in the tree. No save/restore of vruntime occurs. > > > + * dbg_pop_left - pops where the entity queue drained, so it left the tree > > > + * and drm_sched_entity_save_vruntime() ran. Only this path > > > + * arms a later restore. > > > + * dbg_add_restore - calls to drm_sched_rq_add_entity(), i.e. restores. > > > + * dbg_add_race - restores where the entity was still linked in the tree, > > > + * meaning the vruntime being restored is still absolute. > > > + * > > > + * Incremented under rq->lock, so exact per scheduler and only mildly lossy > > > + * when summed across rings. Writable so they can be reset between runs: > > > + * echo 0 > /sys/module/gpu_sched/parameters/dbg_pop_left > > > + */ > > > +static unsigned long dbg_pop_stayed; > > > +static unsigned long dbg_pop_left; > > > +static unsigned long dbg_add_restore; > > > +static unsigned long dbg_add_race; > > > + > > > +module_param(dbg_pop_stayed, ulong, 0644); > > > +module_param(dbg_pop_left, ulong, 0644); > > > +module_param(dbg_add_restore, ulong, 0644); > > > +module_param(dbg_add_race, ulong, 0644); > > > + > > > static __always_inline bool > > > drm_sched_entity_compare_before(struct rb_node *a, const struct rb_node *b) > > > { > > > @@ -266,14 +293,40 @@ > > > sched = container_of(rq, typeof(*sched), rq); > > > spin_lock(&rq->lock); > > > > > > + dbg_add_restore++; > > > + if (!RB_EMPTY_NODE(&entity->rb_tree_node)) { > > > + dbg_add_race++; > > > + pr_warn_ratelimited("drm_sched: vruntime race on ring %s (comm %s)\n", > > > + sched->name, current->comm); > > > + } > > > + > > > if (list_empty(&entity->list)) { > > > atomic_inc(sched->score); > > > list_add_tail(&entity->list, &rq->entities); > > > } > > > > > > - ts = drm_sched_rq_get_min_vruntime(rq); > > > - ts = drm_sched_entity_restore_vruntime(entity, ts, rq->head_prio); > > > - drm_sched_rq_update_tree_locked(entity, rq, ts); > > > + /* > > > + * Only restore the vruntime if the entity actually left the run queue. > > > + * > > > + * drm_sched_entity_pop_job() dequeues the last job and only afterwards > > > + * calls drm_sched_rq_pop_entity(), which is where the entity is removed > > > + * from the tree and drm_sched_entity_save_vruntime() converts its > > > + * vruntime to min_vruntime-relative form. A push landing in that window > > > + * sees an empty queue, takes the "first job" path to here, and would > > > + * restore a vruntime which is still absolute -- adding min_vruntime to > > > + * it a second time. If the entity is also the leftmost one, > > > + * drm_sched_rq_get_min_vruntime() returns the entity's own vruntime and > > > + * the value is doubled, placing it far to the right of the tree where it > > > + * will not be selected again until the run queue catches up. > > > + * > > > + * A still-linked entity never left, so its vruntime is already absolute > > > + * and its tree position is valid. The concurrent pop will update both. > > > + */ > > > + if (RB_EMPTY_NODE(&entity->rb_tree_node)) { > > > + ts = drm_sched_rq_get_min_vruntime(rq); > > > + ts = drm_sched_entity_restore_vruntime(entity, ts, rq->head_prio); > > > + drm_sched_rq_update_tree_locked(entity, rq, ts); > > > + } > > > > > > spin_unlock(&rq->lock); > > > spin_unlock(&entity->lock); > > > @@ -330,11 +383,13 @@ > > > if (next_job) { > > > ktime_t ts; > > > > > > + dbg_pop_stayed++; > > > ts = drm_sched_entity_update_vruntime(entity); > > > drm_sched_rq_update_tree_locked(entity, rq, ts); > > > } else { > > > ktime_t min_vruntime; > > > > > > + dbg_pop_left++; > > > drm_sched_rq_remove_tree_locked(entity, rq); > > > min_vruntime = drm_sched_rq_get_min_vruntime(rq); > > > drm_sched_entity_save_vruntime(entity, min_vruntime); > > > > > > > > > > > > > ^ permalink raw reply [flat|nested] 21+ messages in thread
* RE: [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT 2026-08-13 14:27 ` Luke.Wildhardt 2026-08-14 6:49 ` Luke.Wildhardt @ 2026-08-14 8:32 ` Feng, Kenneth 1 sibling, 0 replies; 21+ messages in thread From: Feng, Kenneth @ 2026-08-14 8:32 UTC (permalink / raw) To: Luke.Wildhardt, Tvrtko Ursulin Cc: Danilo Krummrich, Philipp Stanner, phasta, matthew.brost, ckoenig.leichtzumerken, dri-devel, regressions, amd-gfx, linux-kernel AMD General Hi Luke, I may have a dumb question, but I am curious to know. May I know if there is any performance difference between FAIR policy and FIFO policy on your cards for these test scenarios? Thanks. -----Original Message----- From: amd-gfx <amd-gfx-bounces@lists.freedesktop.org> On Behalf Of Luke.Wildhardt@proton.me Sent: Thursday, August 13, 2026 10:28 PM To: Tvrtko Ursulin <tvrtko.ursulin@igalia.com> Cc: Danilo Krummrich <dakr@kernel.org>; Philipp Stanner <phasta@mailbox.org>; phasta@kernel.org; matthew.brost@intel.com; ckoenig.leichtzumerken@gmail.com; dri-devel@lists.freedesktop.org; regressions@lists.linux.dev; amd-gfx@lists.freedesktop.org; linux-kernel@vger.kernel.org Subject: Re: [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT Looks good, in the 10 minutes I spent this morning doing a quick test, I wasn't able to reproduce the failure. I will do extended testing tonight. I will report back later with more details, but given the issue was typically reproducible on-demand, this is looking very promising. Thanks for all the hard work everyone! On Thursday, August 13th, 2026 at 12:58 AM, Tvrtko Ursulin <tvrtko.ursulin@igalia.com> wrote: > > On 13/08/2026 07:42, Luke.Wildhardt@proton.me wrote: > > Ok, after a lot of A/B testing, it appears that trying your "drm-intel/drm-sched-fair-fixed" branch, the stuttering is still there. It's definitely paced differently though, I will attach a video. It's "choppier" when it happens. I tested in a few different areas / times of day, in case you notice the scenery isn't the same; this is representative of what I experienced in other places. I will note, that I tried playing 4 times, 3/4 times, shortly after boot. 1/4 times I left my desktop to idle for about 30 minutes while I did something else, and I didn't encounter the issue at all in 10 minutes. Not sure how significant that is, or if it's just noise. > > > > https://www.youtube.com/watch?v=GWyIVEuhooM > > > > To make sure I wasn't insane, I went back to the patch Claude gave me, applied against a clean 7.2.0-rc7. At least again in a few sessions after boot, no issue. I'll post that exact code below. > > Yes, I had a brain fart yesterday and had only pulled the lock out on > the pop side. I have now pushed the updated branch, with both the add > and pop side made symmetric in lock taking aspect. > > If you could pull and re-test once more that would be great. > > Regards, > > Tvrtko > > > If you need any other info from me let me know. > > > > > > --- a/drivers/gpu/drm/scheduler/sched_rq.c > > +++ b/drivers/gpu/drm/scheduler/sched_rq.c > > @@ -2,6 +2,7 @@ > > /* Copyright 2015 Advanced Micro Devices, Inc. */ > > /* Copyright (c) 2025 Valve Corporation */ > > > > +#include <linux/moduleparam.h> > > #include <linux/rbtree.h> > > > > #include <drm/drm_print.h> > > @@ -9,6 +10,32 @@ > > > > #include "sched_internal.h" > > > > +/* > > + * Diagnostic counters. Not for submission. > > + * > > + * dbg_pop_stayed - pops where the entity had another job queued and so stayed > > + * in the tree. No save/restore of vruntime occurs. > > + * dbg_pop_left - pops where the entity queue drained, so it left the tree > > + * and drm_sched_entity_save_vruntime() ran. Only this path > > + * arms a later restore. > > + * dbg_add_restore - calls to drm_sched_rq_add_entity(), i.e. restores. > > + * dbg_add_race - restores where the entity was still linked in the tree, > > + * meaning the vruntime being restored is still absolute. > > + * > > + * Incremented under rq->lock, so exact per scheduler and only > > +mildly lossy > > + * when summed across rings. Writable so they can be reset between runs: > > + * echo 0 > /sys/module/gpu_sched/parameters/dbg_pop_left > > + */ > > +static unsigned long dbg_pop_stayed; static unsigned long > > +dbg_pop_left; static unsigned long dbg_add_restore; static unsigned > > +long dbg_add_race; > > + > > +module_param(dbg_pop_stayed, ulong, 0644); > > +module_param(dbg_pop_left, ulong, 0644); > > +module_param(dbg_add_restore, ulong, 0644); > > +module_param(dbg_add_race, ulong, 0644); > > + > > static __always_inline bool > > drm_sched_entity_compare_before(struct rb_node *a, const struct rb_node *b) > > { > > @@ -266,14 +293,40 @@ > > sched = container_of(rq, typeof(*sched), rq); > > spin_lock(&rq->lock); > > > > + dbg_add_restore++; > > + if (!RB_EMPTY_NODE(&entity->rb_tree_node)) { > > + dbg_add_race++; > > + pr_warn_ratelimited("drm_sched: vruntime race on ring %s (comm %s)\n", > > + sched->name, current->comm); > > + } > > + > > if (list_empty(&entity->list)) { > > atomic_inc(sched->score); > > list_add_tail(&entity->list, &rq->entities); > > } > > > > - ts = drm_sched_rq_get_min_vruntime(rq); > > - ts = drm_sched_entity_restore_vruntime(entity, ts, rq->head_prio); > > - drm_sched_rq_update_tree_locked(entity, rq, ts); > > + /* > > + * Only restore the vruntime if the entity actually left the run queue. > > + * > > + * drm_sched_entity_pop_job() dequeues the last job and only afterwards > > + * calls drm_sched_rq_pop_entity(), which is where the entity is removed > > + * from the tree and drm_sched_entity_save_vruntime() converts its > > + * vruntime to min_vruntime-relative form. A push landing in that window > > + * sees an empty queue, takes the "first job" path to here, and would > > + * restore a vruntime which is still absolute -- adding min_vruntime to > > + * it a second time. If the entity is also the leftmost one, > > + * drm_sched_rq_get_min_vruntime() returns the entity's own vruntime and > > + * the value is doubled, placing it far to the right of the tree where it > > + * will not be selected again until the run queue catches up. > > + * > > + * A still-linked entity never left, so its vruntime is already absolute > > + * and its tree position is valid. The concurrent pop will update both. > > + */ > > + if (RB_EMPTY_NODE(&entity->rb_tree_node)) { > > + ts = drm_sched_rq_get_min_vruntime(rq); > > + ts = drm_sched_entity_restore_vruntime(entity, ts, rq->head_prio); > > + drm_sched_rq_update_tree_locked(entity, rq, ts); > > + } > > > > spin_unlock(&rq->lock); > > spin_unlock(&entity->lock); > > @@ -330,11 +383,13 @@ > > if (next_job) { > > ktime_t ts; > > > > + dbg_pop_stayed++; > > ts = drm_sched_entity_update_vruntime(entity); > > drm_sched_rq_update_tree_locked(entity, rq, ts); > > } else { > > ktime_t min_vruntime; > > > > + dbg_pop_left++; > > drm_sched_rq_remove_tree_locked(entity, rq); > > min_vruntime = drm_sched_rq_get_min_vruntime(rq); > > drm_sched_entity_save_vruntime(entity, min_vruntime); > > > > > > > > ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT 2026-08-08 23:33 [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT Luke.Wildhardt 2026-08-10 8:28 ` Philipp Stanner @ 2026-08-10 11:36 ` Tvrtko Ursulin 2026-08-11 7:46 ` Tvrtko Ursulin 2 siblings, 0 replies; 21+ messages in thread From: Tvrtko Ursulin @ 2026-08-10 11:36 UTC (permalink / raw) To: Luke.Wildhardt, matthew.brost, dakr, phasta, ckoenig.leichtzumerken, dri-devel Cc: regressions, amd-gfx, linux-kernel Hi, On 09/08/2026 00:33, Luke.Wildhardt@proton.me wrote: > Since the DRM scheduler default policy was switched to FAIR in 7.2, sustained 100% GPU load in-game on my RX 9070 XT causes severe performance degradation. In the example title running through Proton, Project Silverfish, the foreground application degrades to roughly 10 fps or freezes outright while audio continues, and the KDE Plasma Wayland session can lock up entirely, requiring a reboot or killing the compositor to recover. The problem is not limited to games. Running the game in the background + alt tabbing, then playing a YouTube video can also cause the entire desktop to freeze. > > Then, if the game starts to stutter, either alt-tabbing, or opening up a menu in game (which reduced the load on the GPU) immediately stops the stuttering. > > The issue is reliably reproducible under sustained GPU saturation, although the time-to-failure varies. Increasing shadow load (which decreases in-game frame rate) substantially shortens the time-to-failure. > > In game, MangoHud does not reveal any frametime discrepancy. > > Bisection / test matrix > > Reproducer: Project Silverfish (Proton) at settings that saturate the GPU, with a second GPU client active (browser video). Other titles show the same symptom under sustained load; I have not yet re-tested those specific titles against the fix. > > Kernel Result > torvalds/master stock FAIR (default) FAIL > same as above + reverts below gpu_sched.sched_policy=2 (FAIR) FAIL > same as above + reverts below gpu_sched.sched_policy=1 (FIFO) PASS > 7.1.5 release PASS > drm-next FAIR (default) FAIL > > Reverted Commits: > > d09339388b77 drm/sched: Remove drm_sched_init_args->num_rqs > 2833a0512b4c drm/sched: Remove drm_sched_init_args->num_rqs usage > 16e7698bc04d drm/sched: Embed run queue singleton into the scheduler > 77a6809f1dc3 drm/sched: Remove FIFO and RR and simplify to a single run queue > 45c211ddf92a drm/sched: Switch default policy to fair > 2462a0ce23b0 drm/amdgpu: Remove drm_sched_init_args->num_rqs usage > > > > Kernel (failing): 7.2.0-rc6 (stock), also current drm-next > Kernel (working): 7.2.0-rc6-default-revert-00006-g7e5de9b07e3d with sched_policy=1 > 7.1 release > Distro: Arch Linux > Session: KDE Plasma 6.7.4 / Wayland (KWin), Qt 6.11.1, KF 6.28.0 > CPU: AMD Ryzen 9 9950X3D > RAM: 64 GiB > GPU: Sapphire Pulse Radeon RX 9070 XT > 03:00.0 [1002:7550] rev c0, subsys Sapphire 1478 > Motherboard: ASUS (AM5 / 800-series chipset) > Mesa: 26.1.6-arch1.1 (RADV, driverVersion 26.1.6) > Vulkan: instance 1.4.357 / device 1.4.354 > > > > > On the same reverted kernel build, changing only gpu_sched.sched_policy from 2 (FAIR) to 1 (FIFO) changes the result from reliably failing to passing extended stress testing. > > > > Possible mitigation options > > Given that 7.2 is late in the release cycle and 77a6809f1dc3 removes FIFO/RR as selectable policies, there is currently no runtime workaround for affected systems. > > > Revert 45c211ddf92a so FIFO remains the default while FAIR remains available for further testing. > > If necessary, also revert 77a6809f1dc3 and dependent changes to restore runtime policy selection. > > Alternatively, fix the underlying FAIR regression if a suitable fix can be identified in time for 7.2. > > > I mention the revert options because the justification for removing FIFO/RR included the absence of known regressions relative to those policies. This report appears to provide at least one counterexample. > > > I am willing to test patches responsively and collect data if necessary. Initially I thought it could be a missing wakeup but I couldn't spot any. Then I thought about the clue that the regressed/stuck state happens after some runtime, and is accelerated by the competing clients. So I thought could there be an accumulating error in vruntime handling somehow. The only problem I found so far is that I think the min_vruntime handling has a bug where if an entity never exists the run-queue it can get penalised by its' vruntime only growing, while the re-joining ones get the pull ahead of it. I think the fix is to make sure min_vruntime is strictly monotonic and strictly follows the last popped entity. On a re-read I also think that aligns with how the CPU scheduler does it. Could you please test the below diff against the tip (no reverts): diff --git a/drivers/gpu/drm/scheduler/sched_rq.c b/drivers/gpu/drm/scheduler/sched_rq.c index 044546bcb5f8..6b40f0c44fe5 100644 --- a/drivers/gpu/drm/scheduler/sched_rq.c +++ b/drivers/gpu/drm/scheduler/sched_rq.c @@ -95,6 +95,7 @@ void drm_sched_rq_init(struct drm_sched_rq *rq) INIT_LIST_HEAD(&rq->entities); rq->rb_tree_root = RB_ROOT_CACHED; rq->head_prio = DRM_SCHED_PRIORITY_INVALID; + rq->min_vruntime = 0; } /* @@ -117,28 +118,6 @@ static const unsigned int vruntime_shift[] = { [DRM_SCHED_PRIORITY_LOW] = 7, }; -static ktime_t -drm_sched_rq_get_min_vruntime(struct drm_sched_rq *rq) -{ - ktime_t vruntime = 0; - struct rb_node *rb; - - lockdep_assert_held(&rq->lock); - - rb = rb_first_cached(&rq->rb_tree_root); - if (rb) { - struct drm_sched_entity *entity = - rb_entry(rb, typeof(*entity), rb_tree_node); - struct drm_sched_entity_stats *stats = entity->stats; - - spin_lock(&stats->lock); - vruntime = stats->vruntime; - spin_unlock(&stats->lock); - } - - return vruntime; -} - static void drm_sched_entity_save_vruntime(struct drm_sched_entity *entity, ktime_t min_vruntime) @@ -148,7 +127,7 @@ drm_sched_entity_save_vruntime(struct drm_sched_entity *entity, spin_lock(&stats->lock); vruntime = stats->vruntime; - if (min_vruntime && vruntime > min_vruntime) + if (ktime_after(vruntime, min_vruntime)) vruntime = ktime_sub(vruntime, min_vruntime); else vruntime = 0; @@ -271,8 +250,8 @@ drm_sched_rq_add_entity(struct drm_sched_entity *entity) list_add_tail(&entity->list, &rq->entities); } - ts = drm_sched_rq_get_min_vruntime(rq); - ts = drm_sched_entity_restore_vruntime(entity, ts, rq->head_prio); + ts = drm_sched_entity_restore_vruntime(entity, rq->min_vruntime, + rq->head_prio); drm_sched_rq_update_tree_locked(entity, rq, ts); spin_unlock(&rq->lock); @@ -318,6 +297,7 @@ void drm_sched_rq_pop_entity(struct drm_sched_entity *entity) { struct drm_sched_job *next_job; struct drm_sched_rq *rq; + ktime_t ts; /* * Update the entity's location in the min heap according to @@ -326,18 +306,17 @@ void drm_sched_rq_pop_entity(struct drm_sched_entity *entity) spin_lock(&entity->lock); rq = entity->rq; spin_lock(&rq->lock); + + ts = drm_sched_entity_update_vruntime(entity); + if (ktime_after(ts, rq->min_vruntime)) + rq->min_vruntime = ts; + next_job = drm_sched_entity_queue_peek(entity); if (next_job) { - ktime_t ts; - - ts = drm_sched_entity_update_vruntime(entity); drm_sched_rq_update_tree_locked(entity, rq, ts); } else { - ktime_t min_vruntime; - drm_sched_rq_remove_tree_locked(entity, rq); - min_vruntime = drm_sched_rq_get_min_vruntime(rq); - drm_sched_entity_save_vruntime(entity, min_vruntime); + drm_sched_entity_save_vruntime(entity, rq->min_vruntime); } spin_unlock(&rq->lock); spin_unlock(&entity->lock); diff --git a/include/drm/gpu_scheduler.h b/include/drm/gpu_scheduler.h index 363d13fc929f..892ef474c5c1 100644 --- a/include/drm/gpu_scheduler.h +++ b/include/drm/gpu_scheduler.h @@ -252,6 +252,7 @@ struct drm_sched_entity { * @entities: list of the entities to be scheduled. * @rb_tree_root: root of time based priority queue of entities for FIFO scheduling * @head_prio: priority of the top tree element. + * @min_vruntime: Minimum virtual runtime for the run-queue. * * Run queue is a set of entities scheduling command submissions for * one specific ring. It implements the scheduling policy that selects @@ -263,6 +264,7 @@ struct drm_sched_rq { struct list_head entities; struct rb_root_cached rb_tree_root; enum drm_sched_priority head_prio; + ktime_t min_vruntime; }; /** Regards, Tvrtko ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT 2026-08-08 23:33 [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT Luke.Wildhardt 2026-08-10 8:28 ` Philipp Stanner 2026-08-10 11:36 ` Tvrtko Ursulin @ 2026-08-11 7:46 ` Tvrtko Ursulin 2026-08-12 4:52 ` Luke.Wildhardt 2 siblings, 1 reply; 21+ messages in thread From: Tvrtko Ursulin @ 2026-08-11 7:46 UTC (permalink / raw) To: Luke.Wildhardt, matthew.brost, dakr, phasta, ckoenig.leichtzumerken, dri-devel Cc: regressions, amd-gfx, linux-kernel On 09/08/2026 00:33, Luke.Wildhardt@proton.me wrote: > Since the DRM scheduler default policy was switched to FAIR in 7.2, sustained 100% GPU load in-game on my RX 9070 XT causes severe performance degradation. In the example title running through Proton, Project Silverfish, the foreground application degrades to roughly 10 fps or freezes outright while audio continues, and the KDE Plasma Wayland session can lock up entirely, requiring a reboot or killing the compositor to recover. The problem is not limited to games. Running the game in the background + alt tabbing, then playing a YouTube video can also cause the entire desktop to freeze. Was there anything interesting in dmesg then the total UI lockup happened? > > Then, if the game starts to stutter, either alt-tabbing, or opening up a menu in game (which reduced the load on the GPU) immediately stops the stuttering. > > The issue is reliably reproducible under sustained GPU saturation, although the time-to-failure varies. Increasing shadow load (which decreases in-game frame rate) substantially shortens the time-to-failure. > > In game, MangoHud does not reveal any frametime discrepancy. What exactly did you mean by not revealing discrepancy? Game is at 10fps and HUD shows 10fps? Or game appears to stutter but HUD shows "normal" fps? > Bisection / test matrix > > Reproducer: Project Silverfish (Proton) at settings that saturate the GPU, with a second GPU client active (browser video). Other titles show the same symptom under sustained load; I have not yet re-tested those specific titles against the fix. Could you list the other titles which showed the issue? Regards, Tvrtko > > Kernel Result > torvalds/master stock FAIR (default) FAIL > same as above + reverts below gpu_sched.sched_policy=2 (FAIR) FAIL > same as above + reverts below gpu_sched.sched_policy=1 (FIFO) PASS > 7.1.5 release PASS > drm-next FAIR (default) FAIL > > Reverted Commits: > > d09339388b77 drm/sched: Remove drm_sched_init_args->num_rqs > 2833a0512b4c drm/sched: Remove drm_sched_init_args->num_rqs usage > 16e7698bc04d drm/sched: Embed run queue singleton into the scheduler > 77a6809f1dc3 drm/sched: Remove FIFO and RR and simplify to a single run queue > 45c211ddf92a drm/sched: Switch default policy to fair > 2462a0ce23b0 drm/amdgpu: Remove drm_sched_init_args->num_rqs usage > > > > Kernel (failing): 7.2.0-rc6 (stock), also current drm-next > Kernel (working): 7.2.0-rc6-default-revert-00006-g7e5de9b07e3d with sched_policy=1 > 7.1 release > Distro: Arch Linux > Session: KDE Plasma 6.7.4 / Wayland (KWin), Qt 6.11.1, KF 6.28.0 > CPU: AMD Ryzen 9 9950X3D > RAM: 64 GiB > GPU: Sapphire Pulse Radeon RX 9070 XT > 03:00.0 [1002:7550] rev c0, subsys Sapphire 1478 > Motherboard: ASUS (AM5 / 800-series chipset) > Mesa: 26.1.6-arch1.1 (RADV, driverVersion 26.1.6) > Vulkan: instance 1.4.357 / device 1.4.354 > > > > > On the same reverted kernel build, changing only gpu_sched.sched_policy from 2 (FAIR) to 1 (FIFO) changes the result from reliably failing to passing extended stress testing. > > > > Possible mitigation options > > Given that 7.2 is late in the release cycle and 77a6809f1dc3 removes FIFO/RR as selectable policies, there is currently no runtime workaround for affected systems. > > > Revert 45c211ddf92a so FIFO remains the default while FAIR remains available for further testing. > > If necessary, also revert 77a6809f1dc3 and dependent changes to restore runtime policy selection. > > Alternatively, fix the underlying FAIR regression if a suitable fix can be identified in time for 7.2. > > > I mention the revert options because the justification for removing FIFO/RR included the absence of known regressions relative to those policies. This report appears to provide at least one counterexample. > > > I am willing to test patches responsively and collect data if necessary. > ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT 2026-08-11 7:46 ` Tvrtko Ursulin @ 2026-08-12 4:52 ` Luke.Wildhardt 0 siblings, 0 replies; 21+ messages in thread From: Luke.Wildhardt @ 2026-08-12 4:52 UTC (permalink / raw) To: Tvrtko Ursulin Cc: matthew.brost, dakr, phasta, ckoenig.leichtzumerken, dri-devel, regressions, amd-gfx, linux-kernel Here are some videos I took of what I encountered Stock 7.2.0-rc7 https://www.youtube.com/watch?v=DAmnZ25i1uA The first patch you sent me: https://www.youtube.com/watch?v=86kYUwGwv2c My initial quick observation with the patch you sent me might have just been a fluke, or things were just right for it not to freeze entirely. Additional games I saw this in was Spyro: Reignited Trilogy Hitman: Absolution Same sort of deal, opening a menu (which reduced GPU load) recovered the game. At least in the original game, Project Silverfish, and a small amount of time in Spyro, it appeared that playing with gamescope stopped the issue from appearing. Maybe gamescope does something under the hood that changes the scheduling? On Tuesday, August 11th, 2026 at 12:46 AM, Tvrtko Ursulin <tursulin@ursulin.net> wrote: > > On 09/08/2026 00:33, Luke.Wildhardt@proton.me wrote: > > Since the DRM scheduler default policy was switched to FAIR in 7.2, sustained 100% GPU load in-game on my RX 9070 XT causes severe performance degradation. In the example title running through Proton, Project Silverfish, the foreground application degrades to roughly 10 fps or freezes outright while audio continues, and the KDE Plasma Wayland session can lock up entirely, requiring a reboot or killing the compositor to recover. The problem is not limited to games. Running the game in the background + alt tabbing, then playing a YouTube video can also cause the entire desktop to freeze. > > Was there anything interesting in dmesg then the total UI lockup happened? > > > > > Then, if the game starts to stutter, either alt-tabbing, or opening up a menu in game (which reduced the load on the GPU) immediately stops the stuttering. > > > > The issue is reliably reproducible under sustained GPU saturation, although the time-to-failure varies. Increasing shadow load (which decreases in-game frame rate) substantially shortens the time-to-failure. > > > > In game, MangoHud does not reveal any frametime discrepancy. > > What exactly did you mean by not revealing discrepancy? Game is at 10fps > and HUD shows 10fps? Or game appears to stutter but HUD shows "normal" fps? > > > Bisection / test matrix > > > > Reproducer: Project Silverfish (Proton) at settings that saturate the GPU, with a second GPU client active (browser video). Other titles show the same symptom under sustained load; I have not yet re-tested those specific titles against the fix. > > Could you list the other titles which showed the issue? > > Regards, > > Tvrtko > > > > > Kernel Result > > torvalds/master stock FAIR (default) FAIL > > same as above + reverts below gpu_sched.sched_policy=2 (FAIR) FAIL > > same as above + reverts below gpu_sched.sched_policy=1 (FIFO) PASS > > 7.1.5 release PASS > > drm-next FAIR (default) FAIL > > > > Reverted Commits: > > > > d09339388b77 drm/sched: Remove drm_sched_init_args->num_rqs > > 2833a0512b4c drm/sched: Remove drm_sched_init_args->num_rqs usage > > 16e7698bc04d drm/sched: Embed run queue singleton into the scheduler > > 77a6809f1dc3 drm/sched: Remove FIFO and RR and simplify to a single run queue > > 45c211ddf92a drm/sched: Switch default policy to fair > > 2462a0ce23b0 drm/amdgpu: Remove drm_sched_init_args->num_rqs usage > > > > > > > > Kernel (failing): 7.2.0-rc6 (stock), also current drm-next > > Kernel (working): 7.2.0-rc6-default-revert-00006-g7e5de9b07e3d with sched_policy=1 > > 7.1 release > > Distro: Arch Linux > > Session: KDE Plasma 6.7.4 / Wayland (KWin), Qt 6.11.1, KF 6.28.0 > > CPU: AMD Ryzen 9 9950X3D > > RAM: 64 GiB > > GPU: Sapphire Pulse Radeon RX 9070 XT > > 03:00.0 [1002:7550] rev c0, subsys Sapphire 1478 > > Motherboard: ASUS (AM5 / 800-series chipset) > > Mesa: 26.1.6-arch1.1 (RADV, driverVersion 26.1.6) > > Vulkan: instance 1.4.357 / device 1.4.354 > > > > > > > > > > On the same reverted kernel build, changing only gpu_sched.sched_policy from 2 (FAIR) to 1 (FIFO) changes the result from reliably failing to passing extended stress testing. > > > > > > > > Possible mitigation options > > > > Given that 7.2 is late in the release cycle and 77a6809f1dc3 removes FIFO/RR as selectable policies, there is currently no runtime workaround for affected systems. > > > > > > Revert 45c211ddf92a so FIFO remains the default while FAIR remains available for further testing. > > > > If necessary, also revert 77a6809f1dc3 and dependent changes to restore runtime policy selection. > > > > Alternatively, fix the underlying FAIR regression if a suitable fix can be identified in time for 7.2. > > > > > > I mention the revert options because the justification for removing FIFO/RR included the absence of known regressions relative to those policies. This report appears to provide at least one counterexample. > > > > > > I am willing to test patches responsively and collect data if necessary. > > > > ^ permalink raw reply [flat|nested] 21+ messages in thread
end of thread, other threads:[~2026-09-16 7:37 UTC | newest] Thread overview: 21+ messages (download: mbox.gz / follow: Atom feed) -- links below jump to the message on this page -- 2026-08-08 23:33 [REGRESSION] drm/sched: FAIR policy causes serious performance degradation at max GPU load on 9070XT Luke.Wildhardt 2026-08-10 8:28 ` Philipp Stanner 2026-08-10 12:45 ` Danilo Krummrich 2026-08-10 13:48 ` Tvrtko Ursulin 2026-08-10 16:40 ` Luke.Wildhardt 2026-08-11 7:29 ` Philipp Stanner 2026-08-11 7:49 ` Tvrtko Ursulin 2026-08-11 17:59 ` Luke.Wildhardt 2026-08-12 11:37 ` Tvrtko Ursulin 2026-08-12 11:58 ` Philipp Stanner 2026-08-12 12:29 ` Tvrtko Ursulin 2026-09-16 7:37 ` Philipp Stanner 2026-08-12 15:12 ` Luke.Wildhardt 2026-08-13 6:42 ` Luke.Wildhardt 2026-08-13 7:58 ` Tvrtko Ursulin 2026-08-13 14:27 ` Luke.Wildhardt 2026-08-14 6:49 ` Luke.Wildhardt 2026-08-14 8:32 ` Feng, Kenneth 2026-08-10 11:36 ` Tvrtko Ursulin 2026-08-11 7:46 ` Tvrtko Ursulin 2026-08-12 4:52 ` Luke.Wildhardt
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®