From: Peter Zijlstra <peterz@infradead.org>
To: Tejun Heo <tj@kernel.org>
Cc: David Vernet <void@manifault.com>,
Andrea Righi <arighi@nvidia.com>,
Changwoo Min <changwoo@igalia.com>,
sched-ext@lists.linux.dev, Emil Tsalapatis <emil@etsalapatis.com>,
ElXreno <elxreno@gmail.com>,
linux-kernel@vger.kernel.org, stable@vger.kernel.org
Subject: Re: [PATCH 1/6] sched/core: Handle pick_task() releasing the rq lock
Date: Wed, 19 Aug 2026 14:24:38 +0200 [thread overview]
Message-ID: <20260819122438.GA1248307@noisy.programming.kicks-ass.net> (raw)
In-Reply-To: <20260819093751.GI1247881@noisy.programming.kicks-ass.net>
On Wed, Aug 19, 2026 at 11:37:51AM +0200, Peter Zijlstra wrote:
> On Fri, Aug 07, 2026 at 11:02:16AM -1000, Tejun Heo wrote:
> > Core scheduling's pick_next_task() breaks when a ->pick_task()
> > implementation can release the rq lock. The selection state derived on entry
> > is only valid while the lock is held continuously. Once a pick can drop the
> > lock, an interleaving selection can invalidate all of it: the single-CPU
> > fast path can commit an uncookied pick although the core went cookied during
> > the release, and forceidle committed by the interleaving selection skews the
> > restarted pass's accounting.
> >
> > Fix it by restarting the whole selection when a pick returns RETRY_TASK
> > after releasing the lock: a single restart point above the state derivation
> > replaces the per-loop restart labels, so a retry picks up state committed by
> > interleaving selections and accounts and resets forceidle like a fresh
> > selection would.
> >
> > need_sync and fi_before latch across retries. Clock validity can't be
> > re-derived - there is no program-ordered way to tell whether the own and
> > core rq clocks are still updated after the lock was released, as other
> > lockers' pin cycles may or may not have invalidated them. When restarting,
> > clear core_clock_updated so that the sibling loop re-updates the core rq,
> > and update the own rq clock if invalidated.
> >
> > Fixes: 4c95380701f5 ("sched/ext: Fold balance_scx() into pick_task_scx()")
> > Cc: stable@vger.kernel.org # v6.19+
> > Signed-off-by: Tejun Heo <tj@kernel.org>
>
> Suppose the SMT siblings CPU0 and CPU1; this core sched pick nonsense
> runs on CPU0 and does that multi pick thing.
>
> For CPU0 it pulls a task from the global DSQ, places it in the local
> DSQ, and returns that as the pick. No retry, all good.
>
> Then for CPU1 it does the same, but hits a RETRY, so it stuffs the task
> back on the global DSQ and return RETRY.
>
> Then on retry we find a FIFO task on CPU0, because lock-break and all
> that.
>
> Now we pick the FIFO task, but have not had an opportunity to put the
> CPU0 task back into the global DSQ.
>
> This is still possible, right?
Ah, I read the follow up patches more carefully, and I think you're
dealing with that here. That and I misremembered how the re-enqueue
worked. I though it was still done in pick_task_scx() on RETRY, but we
moved that to wakeup_preempt_scx().
OK, let me go over all this once more...
next prev parent reply other threads:[~2026-08-19 12:24 UTC|newest]
Thread overview: 18+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-07 21:02 [PATCHSET sched_ext/for-7.2-fixes] sched_ext: Fix core scheduling Tejun Heo
2026-08-07 21:02 ` [PATCH 1/6] sched/core: Handle pick_task() releasing the rq lock Tejun Heo
2026-08-10 11:00 ` Peter Zijlstra
2026-08-19 9:37 ` Peter Zijlstra
2026-08-19 12:24 ` Peter Zijlstra [this message]
2026-08-19 18:30 ` Tejun Heo
2026-08-07 21:02 ` [PATCH 2/6] sched/core: Make core-sched flips wait for in-flight selections Tejun Heo
2026-08-10 11:15 ` Peter Zijlstra
2026-08-10 22:10 ` Tejun Heo
2026-08-11 16:05 ` Peter Zijlstra
2026-08-07 21:02 ` [PATCH 3/6] sched_ext: Replace SCX_RQ_BAL_KEEP with a dispatch verdict return Tejun Heo
2026-08-11 7:43 ` Andrea Righi
2026-08-12 17:06 ` [PATCH v2 " Tejun Heo
2026-08-07 21:02 ` [PATCH 4/6] sched_ext: Fix this_rq() assumptions in dispatch kfuncs Tejun Heo
2026-08-07 21:02 ` [PATCH 5/6] sched_ext: Count rq lock releases in rq->scx.lock_drop_seq Tejun Heo
2026-08-07 21:02 ` [PATCH 6/6] sched_ext: Fix rq->core_pick corruption under core scheduling Tejun Heo
2026-08-12 17:07 ` [PATCHSET sched_ext/for-7.2-fixes] sched_ext: Fix " Tejun Heo
2026-08-12 20:25 ` Andrea Righi
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260819122438.GA1248307@noisy.programming.kicks-ass.net \
--to=peterz@infradead.org \
--cc=arighi@nvidia.com \
--cc=changwoo@igalia.com \
--cc=elxreno@gmail.com \
--cc=emil@etsalapatis.com \
--cc=linux-kernel@vger.kernel.org \
--cc=sched-ext@lists.linux.dev \
--cc=stable@vger.kernel.org \
--cc=tj@kernel.org \
--cc=void@manifault.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®