mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Kuba Piecuch <jpiecuch@google.com>
To: Tejun Heo <tj@kernel.org>
Cc: Kuba Piecuch <jpiecuch@google.com>,
	Andrea Righi <arighi@nvidia.com>,
	 David Vernet <void@manifault.com>,
	Changwoo Min <changwoo@igalia.com>,
	 Emil Tsalapatis <emil@etsalapatis.com>,
	sched-ext@lists.linux.dev,  linux-kernel@vger.kernel.org
Subject: Re: [PATCHSET sched_ext/for-7.3-fixes] sched_ext: Fix missing ops.dequeue() on remote local DSQ moves
Date: Wed, 30 Sep 2026 11:50:39 +0000	[thread overview]
Message-ID: <20260930115044.337291-1-jpiecuch@google.com> (raw)
In-Reply-To: <46f4b66249f016b675fa16446b2f2f99@kernel.org>

Hi Tejun,

On Tue, Sep 29, 2026 at 07:13:07AM -1000, Tejun Heo wrote:
> The fix looks good to me. It effectively reverts the enqueue_task_scx()
> half of b75aaea24c9f ("sched_ext: Properly mark SCX-internal migrations
> via sticky_cpu"), which as far as I can see only ever suppressed this
> ops.dequeue(). Andrea, can you confirm?

Thanks for the review. Following Andrea's suggestion, v2 clears
p->scx.sticky_cpu right after it's read, as before b75aaea24c9f, so the
fix is now a straight revert of the enqueue side.

> - dequeue_remote.c isn't built until 3/3, so 1/3 can't be built or run.
>   Can you put the fix first, followed by the test with its Makefile entry?

Sure, will do in v2.

> - 2/3: 7.1.y also needs 18d62044cda7 ("sched_ext: Preserve rq tracking
>   across local DSQ dispatch"). Without it, the nested ops.dequeue() trips
>   lockdep when ops.dispatch() uses scx_bpf_dsq_move() to another CPU's
>   local DSQ. It's tagged for stable too, but maybe note it as a
>   prerequisite?

Thanks, I missed that. v2 lists it as a prerequisite:

> - 2/3: With sub-scheds, scx_resolve_local_dsq() can divert the task to the
>   reject or rescue DSQ, so "inserted into the local DSQ" in the comments
>   isn't always accurate. Maybe "destination DSQ"? The new comment in
>   enqueue_task_scx() could be two lines, and the description could lead
>   with the late SCX_DEQ_CORE_SCHED_EXEC and be a lot shorter.

Agreed on all three, will address these in v2.

> - 1/3: A task can only be picked straight out of custody through
>   sched_core_find(), which only returns tasks with a core cookie. Checking
>   p->core_cookie on SCX_DEQ_CORE_SCHED_EXEC would be exact and would
>   remove core_sched_in_use() and the skip.

That's much better, thanks. v2 reads core_cookie through a CO-RE shadow
struct, so the test still builds and loads without CONFIG_SCHED_CORE,
and core_sched_in_use() and the skip are gone. I also ran the test with
the runner and all its workers sharing a core cookie on an SMT guest.
Legitimate core-sched picks out of custody do happen there, and the test
passes.

> - 1/3: _SC_NPROCESSORS_ONLN ignores affinity. With the runner confined to
>   one CPU, the test fails instead of skipping. sched_getaffinity() and
>   CPU_COUNT()?

Done.

> - 1/3: Nits. If the /proc scan stays, PR_SCHED_CORE_GET writes a u64, so
>   the cookie should be u64. ops.dispatch() pops one entry per call, so a
>   stale one idles the CPU until the next kick. Maybe loop a few times?
>   missed_dequeue_cnt and core_sched_exec_dequeue_cnt aren't printed
>   per-scenario like the other counters.

The /proc scan is gone in v2. ops.dispatch() now pops up to 8 entries until
it finds one that isn't stale. All counters are now reset before eac
scenario, so everything printed is per-scenario.

> - 1/3: The variants, error conditions and core-sched caveat are repeated
>   across the cover, description, file header and comments. Can you say
>   each once? Also, single-line comments are usually lowercase in
>   sched_ext.

Done.

Thanks,
Kuba

      parent reply	other threads:[~2026-09-30 11:51 UTC|newest]

Thread overview: 11+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-29 16:17 Kuba Piecuch
2026-09-29 16:17 ` [PATCH 1/3] selftests/sched_ext: Add a test for " Kuba Piecuch
2026-09-29 18:42   ` Andrea Righi
2026-09-30 11:50     ` Kuba Piecuch
2026-09-29 16:17 ` [PATCH 2/3] sched_ext: Call ops.dequeue() when a task arrives on a remote local DSQ Kuba Piecuch
2026-09-29 18:33   ` Andrea Righi
2026-09-30 11:50     ` Kuba Piecuch
2026-09-29 16:17 ` [PATCH 3/3] selftests/sched_ext: Enable the dequeue_remote test Kuba Piecuch
2026-09-29 17:13 ` [PATCHSET sched_ext/for-7.3-fixes] sched_ext: Fix missing ops.dequeue() on remote local DSQ moves Tejun Heo
2026-09-29 18:30   ` Andrea Righi
2026-09-30 11:50   ` Kuba Piecuch [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260930115044.337291-1-jpiecuch@google.com \
    --to=jpiecuch@google.com \
    --cc=arighi@nvidia.com \
    --cc=changwoo@igalia.com \
    --cc=emil@etsalapatis.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=sched-ext@lists.linux.dev \
    --cc=tj@kernel.org \
    --cc=void@manifault.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®