* [PATCHSET sched_ext/for-7.3-fixes] sched_ext: Fix missing ops.dequeue() on remote local DSQ moves
@ 2026-09-29 16:17 Kuba Piecuch
2026-09-29 16:17 ` [PATCH 1/3] selftests/sched_ext: Add a test for " Kuba Piecuch
` (3 more replies)
0 siblings, 4 replies; 8+ messages in thread
From: Kuba Piecuch @ 2026-09-29 16:17 UTC (permalink / raw)
To: Tejun Heo, David Vernet, Andrea Righi, Changwoo Min
Cc: Kuba Piecuch, Emil Tsalapatis, sched-ext, linux-kernel
Hi,
Since ebf1ccff79c4 ("sched_ext: Fix ops.dequeue() semantics"), every task
entering the BPF scheduler's custody gets exactly one ops.dequeue() when
it leaves it. A dispatch to a terminal DSQ ends custody with a flag-less
ops.dequeue() at insertion time.
That doesn't happen when the task is moved to the local DSQ of a CPU other
than the one whose rq it's on (SCX_DSQ_LOCAL_ON dispatch,
scx_bpf_dsq_move_to_local(), scx_bpf_dsq_move*()).
move_remote_task_to_local_dsq() sets p->scx.sticky_cpu to mark the
internal migration, but enqueue_task_scx() only clears it after the task
has been inserted into the destination local DSQ. task_leave_custody()
thus skips the custody exit on insertion, and the task sits on the local
DSQ with SCX_TASK_IN_CUSTODY still set. The BPF scheduler is only told
later, either:
- when the task is picked, via ops.dequeue(SCX_DEQ_CORE_SCHED_EXEC) from
set_next_task_scx(), even though core scheduling isn't involved, or
- on a property change while the task waits on the local DSQ, via
ops.dequeue(SCX_DEQ_SCHED_CHANGE) for a task already out of custody.
The existing dequeue selftest doesn't notice because it treats the late
SCX_DEQ_CORE_SCHED_EXEC dequeue like a regular dispatch dequeue, and it
still arrives before ops.running().
Patch 1 adds a dequeue_remote selftest that makes remote moves the common
case and fails on a SCX_DEQ_CORE_SCHED_EXEC dequeue, on a task running
without a preceding ops.dequeue(), or on unbalanced ops.enqueue() /
ops.dequeue() calls. Since core scheduling can legitimately pick tasks
straight out of custody, the test is skipped if any task has a core
scheduling cookie when it starts. It isn't added to auto-test-targets
since it fails on the current tree. Patch 2 fixes the bug by clearing
p->scx.sticky_cpu before scx_do_enqueue_task(). Patch 3 enables the test.
This is based on sched_ext/for-7.3-fixes (94480606a677) and also applies
cleanly to sched_ext/for-next. Testing was done on x86_64 in virtme-ng
with 4 vCPUs in an SMT topology (2 cores x 2 threads) and
CONFIG_SCHED_CORE=y, with no core scheduling cookies in use.
Without patch 2, dequeue_remote fails in every run (30/30), within the
first few custody enqueues after the scheduler is loaded:
sched_ext: dequeue_remote: dequeue_remote.bpf.c:143: 156 (runner): late ops.dequeue() with SCX_DEQ_CORE_SCHED_EXEC (enq_cpu=1 cpu=3 seq=1)
...
ops_dequeue+0x114/0x170
set_next_task_scx+0x104/0x1e0
and the full sched_ext selftest suite reports:
PASSED: 31
SKIPPED: 1 (nohz_tick)
FAILED: 1 (dequeue_remote)
With patch 2, dequeue_remote passes in every run (30/30). A typical run
does ~140k custody enqueues across both variants, each matched by exactly
one ops.dequeue(), with ~90k of them followed by the task running on a
CPU other than the one it was enqueued on. The full suite reports:
PASSED: 32
SKIPPED: 1 (nohz_tick)
FAILED: 0
With a core scheduling cookie in use, dequeue_remote is skipped as
expected.
Thanks,
Kuba
Assisted-by: Claude:claude-opus-5.5
Kuba Piecuch (3):
selftests/sched_ext: Add a test for ops.dequeue() on remote local DSQ
moves
sched_ext: Call ops.dequeue() when a task arrives on a remote local
DSQ
selftests/sched_ext: Enable the dequeue_remote test
kernel/sched/ext/ext.c | 16 +-
tools/testing/selftests/sched_ext/Makefile | 1 +
.../selftests/sched_ext/dequeue_remote.bpf.c | 272 ++++++++++++++++++
.../selftests/sched_ext/dequeue_remote.c | 243 ++++++++++++++++
4 files changed, 527 insertions(+), 5 deletions(-)
create mode 100644 tools/testing/selftests/sched_ext/dequeue_remote.bpf.c
create mode 100644 tools/testing/selftests/sched_ext/dequeue_remote.c
base-commit: 94480606a677deb68d5622cf0ded88626514b3f2
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH 1/3] selftests/sched_ext: Add a test for ops.dequeue() on remote local DSQ moves
2026-09-29 16:17 [PATCHSET sched_ext/for-7.3-fixes] sched_ext: Fix missing ops.dequeue() on remote local DSQ moves Kuba Piecuch
@ 2026-09-29 16:17 ` Kuba Piecuch
2026-09-29 18:42 ` Andrea Righi
2026-09-29 16:17 ` [PATCH 2/3] sched_ext: Call ops.dequeue() when a task arrives on a remote local DSQ Kuba Piecuch
` (2 subsequent siblings)
3 siblings, 1 reply; 8+ messages in thread
From: Kuba Piecuch @ 2026-09-29 16:17 UTC (permalink / raw)
To: Tejun Heo, David Vernet, Andrea Righi, Changwoo Min
Cc: Kuba Piecuch, Emil Tsalapatis, sched-ext, linux-kernel
When a BPF scheduler moves a task that is in its custody to the local DSQ
of a CPU other than the one whose rq the task is on, sched_ext migrates the
task with move_remote_task_to_local_dsq(). The task leaves the BPF
scheduler's custody when it is inserted into the destination local DSQ and
ops.dequeue() should be called at that point, just like for same-rq
dispatches.
Add a dequeue_remote test that exercises this path. All tasks are put
into custody from ops.enqueue() and moved to the local DSQ of whichever
CPU runs ops.dispatch(), so that most moves cross CPUs. Two variants are
covered, selected by test_use_move_to_local:
- false: tasks are queued in a BPF queue and dispatched with
scx_bpf_dsq_insert(SCX_DSQ_LOCAL_ON | cpu), going through
dispatch_to_local_dsq(),
- true: tasks are queued on a user DSQ and consumed with
scx_bpf_dsq_move_to_local(), going through consume_remote_task().
The BPF scheduler tracks each task's custody state and triggers
scx_bpf_error() if:
- ops.dequeue() is called with SCX_DEQ_CORE_SCHED_EXEC, meaning that
the custody exit was missed on insertion into the local DSQ and
deferred to set_next_task_scx(),
- a task starts running without ops.dequeue() having been called since
its last ops.enqueue(),
- ops.dequeue() is called for a task that isn't in custody, or
ops.enqueue() for a task that is.
Core scheduling can legitimately pick tasks straight out of custody, so
the test is skipped if any task has a core scheduling cookie when it
starts.
The test currently fails:
sched_ext: dequeue_remote: dequeue_remote.bpf.c:143: 156 (runner): late ops.dequeue() with SCX_DEQ_CORE_SCHED_EXEC (enq_cpu=1 cpu=3 seq=1)
...
ops_dequeue+0x114/0x170
set_next_task_scx+0x104/0x1e0
__pick_next_task+0xc7/0x180
__schedule+0x154/0x1870
so don't add it to auto-test-targets yet. It will be enabled once the
underlying bug is fixed.
Assisted-by: Claude:claude-opus-5.5
Signed-off-by: Kuba Piecuch <jpiecuch@google.com>
---
.../selftests/sched_ext/dequeue_remote.bpf.c | 272 ++++++++++++++++++
.../selftests/sched_ext/dequeue_remote.c | 243 ++++++++++++++++
2 files changed, 515 insertions(+)
create mode 100644 tools/testing/selftests/sched_ext/dequeue_remote.bpf.c
create mode 100644 tools/testing/selftests/sched_ext/dequeue_remote.c
diff --git a/tools/testing/selftests/sched_ext/dequeue_remote.bpf.c b/tools/testing/selftests/sched_ext/dequeue_remote.bpf.c
new file mode 100644
index 000000000000..4da428116ee4
--- /dev/null
+++ b/tools/testing/selftests/sched_ext/dequeue_remote.bpf.c
@@ -0,0 +1,272 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Verify that ops.dequeue() is called when a task leaves the BPF
+ * scheduler's custody through an SCX-internal cross-CPU migration, i.e.
+ * when the task is moved to the local DSQ of a CPU other than the one
+ * whose rq it is currently on (move_remote_task_to_local_dsq()).
+ *
+ * Every task is put into BPF custody from ops.enqueue() and later moved to
+ * a local DSQ from ops.dispatch() running on an arbitrary CPU, which makes
+ * remote moves very frequent. Depending on @test_use_move_to_local, tasks
+ * are either:
+ *
+ * - false: queued in a BPF queue and dispatched with
+ * scx_bpf_dsq_insert(SCX_DSQ_LOCAL_ON | cpu) (dispatch_to_local_dsq()),
+ * - true: queued on a user DSQ and consumed with
+ * scx_bpf_dsq_move_to_local() (consume_remote_task()).
+ *
+ * A task can only start running from a local DSQ, i.e. after it has left
+ * custody. So by the time ops.running() is invoked, ops.dequeue() must have
+ * been called for the preceding ops.enqueue(). Moreover, as core scheduling
+ * isn't in use (the test is skipped otherwise), that ops.dequeue() must
+ * happen when the task is inserted into the local DSQ, not later when it's
+ * picked for execution, which would be reported with
+ * %SCX_DEQ_CORE_SCHED_EXEC.
+ *
+ * Copyright (c) 2026 Google LLC.
+ */
+
+#include <scx/common.bpf.h>
+
+#define SHARED_DSQ 0
+
+char _license[] SEC("license") = "GPL";
+
+UEI_DEFINE(uei);
+
+struct {
+ __uint(type, BPF_MAP_TYPE_QUEUE);
+ __uint(max_entries, 32768);
+ __type(value, s32);
+} global_queue SEC(".maps");
+
+enum task_state {
+ TASK_NONE = 0,
+ TASK_ENQUEUED, /* in BPF custody, waiting for ops.dequeue() */
+ TASK_DISPATCHED, /* left custody */
+};
+
+struct task_ctx {
+ enum task_state state;
+ s32 enq_cpu; /* scx_bpf_task_cpu() at ops.enqueue() */
+ u64 enqueue_seq;
+};
+
+struct {
+ __uint(type, BPF_MAP_TYPE_TASK_STORAGE);
+ __uint(map_flags, BPF_F_NO_PREALLOC);
+ __type(key, int);
+ __type(value, struct task_ctx);
+} task_ctx_stor SEC(".maps");
+
+bool test_use_move_to_local;
+
+u64 enqueue_cnt, dequeue_cnt, dispatch_dequeue_cnt, change_dequeue_cnt;
+u64 remote_dispatch_cnt, remote_running_cnt, missed_dequeue_cnt;
+u64 core_sched_exec_dequeue_cnt;
+
+static struct task_ctx *lookup_task_ctx(struct task_struct *p)
+{
+ return bpf_task_storage_get(&task_ctx_stor, p, 0, 0);
+}
+
+s32 BPF_STRUCT_OPS(dequeue_remote_select_cpu, struct task_struct *p,
+ s32 prev_cpu, u64 wake_flags)
+{
+ /* No direct dispatch: always go through ops.enqueue() */
+ return prev_cpu;
+}
+
+void BPF_STRUCT_OPS(dequeue_remote_enqueue, struct task_struct *p, u64 enq_flags)
+{
+ struct task_ctx *tctx;
+ s32 pid = p->pid;
+
+ tctx = lookup_task_ctx(p);
+ if (!tctx) {
+ scx_bpf_dsq_insert(p, SCX_DSQ_GLOBAL, SCX_SLICE_DFL, enq_flags);
+ return;
+ }
+
+ /*
+ * Every task that entered custody must have received ops.dequeue()
+ * before it can be enqueued again.
+ */
+ if (tctx->state == TASK_ENQUEUED)
+ scx_bpf_error("%d (%s): enqueue while in ENQUEUED state seq=%llu",
+ p->pid, p->comm, tctx->enqueue_seq);
+
+ /*
+ * Mark @p as enqueued before making it visible to ops.dispatch() on
+ * other CPUs, which skips queue entries of tasks not in ENQUEUED
+ * state as stale.
+ */
+ tctx->state = TASK_ENQUEUED;
+ tctx->enq_cpu = scx_bpf_task_cpu(p);
+ tctx->enqueue_seq++;
+
+ if (test_use_move_to_local) {
+ scx_bpf_dsq_insert(p, SHARED_DSQ, SCX_SLICE_DFL, enq_flags);
+ } else if (bpf_map_push_elem(&global_queue, &pid, 0)) {
+ scx_bpf_dsq_insert(p, SCX_DSQ_GLOBAL, SCX_SLICE_DFL, enq_flags);
+ tctx->state = TASK_DISPATCHED;
+ tctx->enq_cpu = -1;
+ goto out;
+ }
+
+ __sync_fetch_and_add(&enqueue_cnt, 1);
+out:
+ scx_bpf_kick_cpu(scx_bpf_task_cpu(p), SCX_KICK_IDLE);
+}
+
+void BPF_STRUCT_OPS(dequeue_remote_dequeue, struct task_struct *p, u64 deq_flags)
+{
+ struct task_ctx *tctx;
+
+ __sync_fetch_and_add(&dequeue_cnt, 1);
+
+ tctx = lookup_task_ctx(p);
+ if (!tctx)
+ return;
+
+ /*
+ * The test doesn't run when core scheduling is in use, so no task can
+ * be picked for execution before being dispatched to a local DSQ. A
+ * %SCX_DEQ_CORE_SCHED_EXEC dequeue means that the task reached a
+ * local DSQ without leaving custody (no ops.dequeue() on arrival)
+ * and the custody exit was deferred to set_next_task_scx().
+ */
+ if (deq_flags & SCX_DEQ_CORE_SCHED_EXEC) {
+ __sync_fetch_and_add(&core_sched_exec_dequeue_cnt, 1);
+ scx_bpf_error("%d (%s): late ops.dequeue() with SCX_DEQ_CORE_SCHED_EXEC (enq_cpu=%d cpu=%d seq=%llu)",
+ p->pid, p->comm, tctx->enq_cpu,
+ scx_bpf_task_cpu(p), tctx->enqueue_seq);
+ }
+
+ /*
+ * ops.dequeue() is called exactly once per custody period, which
+ * starts with ops.enqueue(). Whatever the reason, @p must be in
+ * ENQUEUED state. In particular, NONE is only entered through a
+ * property change dequeue, which ends custody.
+ */
+ if (tctx->state != TASK_ENQUEUED)
+ scx_bpf_error("%d (%s): dequeue outside custody deq_flags=0x%llx state=%d seq=%llu",
+ p->pid, p->comm, deq_flags, tctx->state,
+ tctx->enqueue_seq);
+
+ if (deq_flags & SCX_DEQ_SCHED_CHANGE) {
+ __sync_fetch_and_add(&change_dequeue_cnt, 1);
+ tctx->state = TASK_NONE;
+ } else {
+ __sync_fetch_and_add(&dispatch_dequeue_cnt, 1);
+ tctx->state = TASK_DISPATCHED;
+ }
+}
+
+void BPF_STRUCT_OPS(dequeue_remote_dispatch, s32 cpu, struct task_struct *prev)
+{
+ struct task_ctx *tctx;
+ struct task_struct *p;
+ s32 pid;
+
+ if (test_use_move_to_local) {
+ scx_bpf_dsq_move_to_local(SHARED_DSQ, 0);
+ return;
+ }
+
+ if (bpf_map_pop_elem(&global_queue, &pid))
+ return;
+
+ p = bpf_task_from_pid(pid);
+ if (!p)
+ return;
+
+ /*
+ * Skip stale entries: tasks that left custody through a property
+ * change dequeue (state NONE) or that were already dispatched from
+ * a duplicate queue entry.
+ */
+ tctx = lookup_task_ctx(p);
+ if (!tctx || tctx->state != TASK_ENQUEUED) {
+ bpf_task_release(p);
+ return;
+ }
+
+ /*
+ * Move the task to this CPU's local DSQ. The task sits on the rq of
+ * the CPU it was enqueued on, so this is a remote move whenever
+ * that CPU differs from @cpu.
+ */
+ if (bpf_cpumask_test_cpu(cpu, p->cpus_ptr)) {
+ if (scx_bpf_task_cpu(p) != cpu)
+ __sync_fetch_and_add(&remote_dispatch_cnt, 1);
+ scx_bpf_dsq_insert(p, SCX_DSQ_LOCAL_ON | cpu, SCX_SLICE_DFL, 0);
+ } else {
+ scx_bpf_dsq_insert(p, SCX_DSQ_GLOBAL, SCX_SLICE_DFL, 0);
+ }
+
+ bpf_task_release(p);
+}
+
+void BPF_STRUCT_OPS(dequeue_remote_running, struct task_struct *p)
+{
+ struct task_ctx *tctx;
+
+ tctx = lookup_task_ctx(p);
+ if (!tctx)
+ return;
+
+ if (tctx->state == TASK_ENQUEUED) {
+ /* Running from a local DSQ without having left custody */
+ __sync_fetch_and_add(&missed_dequeue_cnt, 1);
+ scx_bpf_error("%d (%s): running without ops.dequeue() (enq_cpu=%d cpu=%d seq=%llu)",
+ p->pid, p->comm, tctx->enq_cpu,
+ scx_bpf_task_cpu(p), tctx->enqueue_seq);
+ return;
+ }
+
+ if (tctx->enq_cpu >= 0 && tctx->enq_cpu != scx_bpf_task_cpu(p))
+ __sync_fetch_and_add(&remote_running_cnt, 1);
+ tctx->enq_cpu = -1;
+}
+
+s32 BPF_STRUCT_OPS(dequeue_remote_init_task, struct task_struct *p,
+ struct scx_init_task_args *args)
+{
+ struct task_ctx *tctx;
+
+ tctx = bpf_task_storage_get(&task_ctx_stor, p, 0,
+ BPF_LOCAL_STORAGE_GET_F_CREATE);
+ if (!tctx)
+ return -ENOMEM;
+
+ /* task storage persists across attachments, start from scratch */
+ tctx->state = TASK_NONE;
+ tctx->enq_cpu = -1;
+
+ return 0;
+}
+
+s32 BPF_STRUCT_OPS_SLEEPABLE(dequeue_remote_init)
+{
+ return scx_bpf_create_dsq(SHARED_DSQ, -1);
+}
+
+void BPF_STRUCT_OPS(dequeue_remote_exit, struct scx_exit_info *ei)
+{
+ UEI_RECORD(uei, ei);
+}
+
+SEC(".struct_ops.link")
+struct sched_ext_ops dequeue_remote_ops = {
+ .select_cpu = (void *)dequeue_remote_select_cpu,
+ .enqueue = (void *)dequeue_remote_enqueue,
+ .dequeue = (void *)dequeue_remote_dequeue,
+ .dispatch = (void *)dequeue_remote_dispatch,
+ .running = (void *)dequeue_remote_running,
+ .init_task = (void *)dequeue_remote_init_task,
+ .init = (void *)dequeue_remote_init,
+ .exit = (void *)dequeue_remote_exit,
+ .flags = SCX_OPS_ENQ_LAST,
+ .name = "dequeue_remote",
+};
diff --git a/tools/testing/selftests/sched_ext/dequeue_remote.c b/tools/testing/selftests/sched_ext/dequeue_remote.c
new file mode 100644
index 000000000000..ab42ac694ac6
--- /dev/null
+++ b/tools/testing/selftests/sched_ext/dequeue_remote.c
@@ -0,0 +1,243 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Verify that ops.dequeue() is called for tasks leaving BPF custody through
+ * an SCX-internal cross-CPU migration (move_remote_task_to_local_dsq()).
+ *
+ * Copyright (c) 2026 Google LLC.
+ */
+#define _GNU_SOURCE
+#include <stdio.h>
+#include <stdlib.h>
+#include <unistd.h>
+#include <signal.h>
+#include <time.h>
+#include <sched.h>
+#include <ctype.h>
+#include <dirent.h>
+#include <sys/prctl.h>
+#include <bpf/bpf.h>
+#include <scx/common.h>
+#include <sys/wait.h>
+#include "scx_test.h"
+#include "dequeue_remote.bpf.skel.h"
+
+#define MAX_WORKERS 64
+#define RUN_MS 2000
+
+static long long now_ms(void)
+{
+ struct timespec ts;
+
+ clock_gettime(CLOCK_MONOTONIC, &ts);
+ return ts.tv_sec * 1000LL + ts.tv_nsec / 1000000;
+}
+
+static int get_core_cookie(pid_t tid, unsigned long *cookie)
+{
+ return prctl(PR_SCHED_CORE, PR_SCHED_CORE_GET, tid,
+ PR_SCHED_CORE_SCOPE_THREAD, (unsigned long)cookie);
+}
+
+/*
+ * Core scheduling is in use iff some task has a core scheduling cookie. If
+ * none does now, assume that none will get one for the duration of the test.
+ */
+static bool core_sched_in_use(void)
+{
+ unsigned long cookie;
+ struct dirent *pde, *tde;
+ bool in_use = false;
+ DIR *proc, *task;
+ char path[sizeof("/proc//task") + sizeof(pde->d_name)];
+
+ /* fails with EINVAL if !CONFIG_SCHED_CORE and ENODEV if SMT is off */
+ if (get_core_cookie(0, &cookie))
+ return false;
+
+ proc = opendir("/proc");
+ if (!proc)
+ return false;
+
+ while (!in_use && (pde = readdir(proc))) {
+ if (!isdigit(pde->d_name[0]))
+ continue;
+
+ snprintf(path, sizeof(path), "/proc/%s/task", pde->d_name);
+ task = opendir(path);
+ if (!task)
+ continue;
+
+ /* errors, e.g. ESRCH for exited tasks, are ignored */
+ while (!in_use && (tde = readdir(task))) {
+ if (isdigit(tde->d_name[0]) &&
+ !get_core_cookie(atoi(tde->d_name), &cookie))
+ in_use = cookie;
+ }
+
+ closedir(task);
+ }
+
+ closedir(proc);
+ return in_use;
+}
+
+/* Mix of short bursts and sleeps to generate lots of enqueues and wakeups */
+static void worker_fn(int id)
+{
+ long long end = now_ms() + RUN_MS;
+ volatile unsigned long sum = 0;
+
+ while (now_ms() < end) {
+ unsigned long j;
+
+ for (j = 0; j < 20000 + id * 1000; j++)
+ sum += j;
+ if (id & 1)
+ usleep(100);
+ else
+ sched_yield();
+ }
+
+ exit(0);
+}
+
+static enum scx_test_status run_scenario(struct dequeue_remote *skel,
+ bool use_move_to_local,
+ const char *name)
+{
+ struct bpf_link *link;
+ pid_t pids[MAX_WORKERS];
+ int nr_workers, i, status;
+ u64 enq, deq, dsp_deq, chg_deq, remote_dsp, remote_run;
+
+ nr_workers = 2 * sysconf(_SC_NPROCESSORS_ONLN);
+ if (nr_workers < 4)
+ nr_workers = 4;
+ if (nr_workers > MAX_WORKERS)
+ nr_workers = MAX_WORKERS;
+
+ skel->bss->test_use_move_to_local = use_move_to_local;
+ enq = skel->bss->enqueue_cnt;
+ deq = skel->bss->dequeue_cnt;
+ dsp_deq = skel->bss->dispatch_dequeue_cnt;
+ chg_deq = skel->bss->change_dequeue_cnt;
+ remote_dsp = skel->bss->remote_dispatch_cnt;
+ remote_run = skel->bss->remote_running_cnt;
+
+ link = bpf_map__attach_struct_ops(skel->maps.dequeue_remote_ops);
+ SCX_FAIL_IF(!link, "Failed to attach struct_ops for %s", name);
+
+ fflush(stdout);
+ fflush(stderr);
+
+ for (i = 0; i < nr_workers; i++) {
+ pids[i] = fork();
+ SCX_FAIL_IF(pids[i] < 0, "Failed to fork worker %d", i);
+ if (pids[i] == 0)
+ worker_fn(i);
+ }
+
+ for (i = 0; i < nr_workers; i++) {
+ SCX_FAIL_IF(waitpid(pids[i], &status, 0) != pids[i],
+ "Failed to wait for worker %d", i);
+ SCX_FAIL_IF(status != 0, "Worker %d exited with status %d", i, status);
+ }
+
+ bpf_link__destroy(link);
+
+ enq = skel->bss->enqueue_cnt - enq;
+ deq = skel->bss->dequeue_cnt - deq;
+ dsp_deq = skel->bss->dispatch_dequeue_cnt - dsp_deq;
+ chg_deq = skel->bss->change_dequeue_cnt - chg_deq;
+ remote_dsp = skel->bss->remote_dispatch_cnt - remote_dsp;
+ remote_run = skel->bss->remote_running_cnt - remote_run;
+
+ printf("%s:\n", name);
+ printf(" workers: %d\n", nr_workers);
+ printf(" enqueues: %lu\n", (unsigned long)enq);
+ printf(" dequeues: %lu (dispatch: %lu, property_change: %lu)\n",
+ (unsigned long)deq, (unsigned long)dsp_deq,
+ (unsigned long)chg_deq);
+ if (!use_move_to_local)
+ printf(" remote SCX_DSQ_LOCAL_ON dispatches: %lu\n",
+ (unsigned long)remote_dsp);
+ printf(" ran on a CPU other than the enqueue CPU: %lu\n",
+ (unsigned long)remote_run);
+ printf(" ran without ops.dequeue(): %lu\n",
+ (unsigned long)skel->bss->missed_dequeue_cnt);
+ printf(" late SCX_DEQ_CORE_SCHED_EXEC dequeues: %lu\n",
+ (unsigned long)skel->bss->core_sched_exec_dequeue_cnt);
+
+ if (skel->data->uei.kind != EXIT_KIND(SCX_EXIT_UNREG))
+ SCX_ERR("Scheduler exited with kind=%lld: %s",
+ (long long)skel->data->uei.kind, skel->data->uei.msg);
+ SCX_EQ(skel->data->uei.kind, EXIT_KIND(SCX_EXIT_UNREG));
+
+ /* the test is meaningless if no task was moved across CPUs */
+ SCX_GT(remote_run, 0);
+ SCX_EQ(enq, deq);
+
+ return SCX_TEST_PASS;
+}
+
+static enum scx_test_status setup(void **ctx)
+{
+ struct dequeue_remote *skel;
+
+ if (sysconf(_SC_NPROCESSORS_ONLN) < 2) {
+ fprintf(stderr, "Skipping: requires at least 2 CPUs\n");
+ return SCX_TEST_SKIP;
+ }
+
+ /*
+ * Core scheduling can legitimately pick tasks straight out of BPF
+ * custody, which the test would flag as late
+ * %SCX_DEQ_CORE_SCHED_EXEC dequeues.
+ */
+ if (core_sched_in_use()) {
+ fprintf(stderr, "Skipping: core scheduling is in use\n");
+ return SCX_TEST_SKIP;
+ }
+
+ skel = SCX_OPS_OPEN(dequeue_remote_ops, dequeue_remote);
+ SCX_OPS_LOAD(skel, dequeue_remote_ops, dequeue_remote, uei);
+
+ *ctx = skel;
+
+ return SCX_TEST_PASS;
+}
+
+static enum scx_test_status run(void *ctx)
+{
+ struct dequeue_remote *skel = ctx;
+ enum scx_test_status status;
+
+ status = run_scenario(skel, false,
+ "BPF queue -> SCX_DSQ_LOCAL_ON | cpu");
+ if (status != SCX_TEST_PASS)
+ return status;
+
+ status = run_scenario(skel, true,
+ "user DSQ -> scx_bpf_dsq_move_to_local()");
+ if (status != SCX_TEST_PASS)
+ return status;
+
+ return SCX_TEST_PASS;
+}
+
+static void cleanup(void *ctx)
+{
+ struct dequeue_remote *skel = ctx;
+
+ dequeue_remote__destroy(skel);
+}
+
+struct scx_test dequeue_remote_test = {
+ .name = "dequeue_remote",
+ .description = "Verify ops.dequeue() on SCX-internal cross-CPU migrations",
+ .setup = setup,
+ .run = run,
+ .cleanup = cleanup,
+};
+
+REGISTER_SCX_TEST(&dequeue_remote_test)
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH 2/3] sched_ext: Call ops.dequeue() when a task arrives on a remote local DSQ
2026-09-29 16:17 [PATCHSET sched_ext/for-7.3-fixes] sched_ext: Fix missing ops.dequeue() on remote local DSQ moves Kuba Piecuch
2026-09-29 16:17 ` [PATCH 1/3] selftests/sched_ext: Add a test for " Kuba Piecuch
@ 2026-09-29 16:17 ` Kuba Piecuch
2026-09-29 18:33 ` Andrea Righi
2026-09-29 16:17 ` [PATCH 3/3] selftests/sched_ext: Enable the dequeue_remote test Kuba Piecuch
2026-09-29 17:13 ` [PATCHSET sched_ext/for-7.3-fixes] sched_ext: Fix missing ops.dequeue() on remote local DSQ moves Tejun Heo
3 siblings, 1 reply; 8+ messages in thread
From: Kuba Piecuch @ 2026-09-29 16:17 UTC (permalink / raw)
To: Tejun Heo, David Vernet, Andrea Righi, Changwoo Min
Cc: Kuba Piecuch, Emil Tsalapatis, sched-ext, linux-kernel
When the BPF scheduler moves a task to the local DSQ of a CPU other than
the one whose rq the task is on, e.g. with SCX_DSQ_LOCAL_ON dispatch,
scx_bpf_dsq_move_to_local() or scx_bpf_dsq_move*(), the task is migrated
by move_remote_task_to_local_dsq(). p->scx.sticky_cpu is set across the
migration so that dequeueing the task from the source rq isn't treated as
the task leaving the BPF scheduler's custody.
However, enqueue_task_scx() only clears p->scx.sticky_cpu after
scx_do_enqueue_task() has inserted the task into the destination local
DSQ. task_leave_custody(), called from rq_owned_post_enq() on insertion,
still sees the migration in progress and skips the custody exit. The task
ends up on a terminal DSQ with SCX_TASK_IN_CUSTODY set and without
ops.dequeue() having been called.
The custody exit then happens only when the task is picked for execution,
in set_next_task_scx(), which reports it to the BPF scheduler as
ops.dequeue(SCX_DEQ_CORE_SCHED_EXEC) even though no core-sched pick took
place. If a scheduling property change hits the task while it's waiting
on the local DSQ, the BPF scheduler instead gets
ops.dequeue(SCX_DEQ_SCHED_CHANGE) for a task that has already left its
custody. Both break the ops.dequeue() semantics, under which a dispatch to
a terminal DSQ ends custody with an ops.dequeue() call without special
flags.
Clear p->scx.sticky_cpu before calling scx_do_enqueue_task(). The routing
decision in scx_do_enqueue_task() uses the local copy, and the departure
side is unaffected as p->scx.sticky_cpu is still set across
deactivate_task(). ops.dequeue() is now invoked on the destination rq when
the task is inserted into the local DSQ, as it already is for same-rq
dispatches.
Fixes: ebf1ccff79c4 ("sched_ext: Fix ops.dequeue() semantics")
Cc: stable@vger.kernel.org # v7.1+
Assisted-by: Claude:claude-opus-5.5
Signed-off-by: Kuba Piecuch <jpiecuch@google.com>
---
kernel/sched/ext/ext.c | 16 +++++++++++-----
1 file changed, 11 insertions(+), 5 deletions(-)
diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
index ad391a8cbd05..763de7797056 100644
--- a/kernel/sched/ext/ext.c
+++ b/kernel/sched/ext/ext.c
@@ -1501,8 +1501,9 @@ static inline bool task_scx_migrating(struct task_struct *p)
/*
* We only need to check sticky_cpu: it is set to the destination
* CPU in move_remote_task_to_local_dsq() before deactivate_task()
- * and cleared when the task is enqueued on the destination, so it
- * is only non-negative during an internal SCX migration.
+ * and cleared in enqueue_task_scx() on the destination before @p is
+ * inserted into the local DSQ, so it is only non-negative while @p
+ * is in transit between the two rqs.
*/
return p->scx.sticky_cpu >= 0;
}
@@ -2189,10 +2190,15 @@ static void enqueue_task_scx(struct rq *rq, struct task_struct *p, int core_enq_
if (rq->scx.nr_running == 1)
dl_server_start(&rq->ext_server);
- scx_do_enqueue_task(rq, p, enq_flags, sticky_cpu);
+ /*
+ * An SCX-internal migration ends once @p arrives on the destination
+ * rq. Clear sticky_cpu before enqueueing so that @p leaves the BPF
+ * scheduler's custody when inserted into the destination local DSQ.
+ * The local copy in @sticky_cpu is used for routing.
+ */
+ p->scx.sticky_cpu = -1;
- if (sticky_cpu >= 0)
- p->scx.sticky_cpu = -1;
+ scx_do_enqueue_task(rq, p, enq_flags, sticky_cpu);
out:
rq->scx.flags &= ~SCX_RQ_IN_WAKEUP;
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH 3/3] selftests/sched_ext: Enable the dequeue_remote test
2026-09-29 16:17 [PATCHSET sched_ext/for-7.3-fixes] sched_ext: Fix missing ops.dequeue() on remote local DSQ moves Kuba Piecuch
2026-09-29 16:17 ` [PATCH 1/3] selftests/sched_ext: Add a test for " Kuba Piecuch
2026-09-29 16:17 ` [PATCH 2/3] sched_ext: Call ops.dequeue() when a task arrives on a remote local DSQ Kuba Piecuch
@ 2026-09-29 16:17 ` Kuba Piecuch
2026-09-29 17:13 ` [PATCHSET sched_ext/for-7.3-fixes] sched_ext: Fix missing ops.dequeue() on remote local DSQ moves Tejun Heo
3 siblings, 0 replies; 8+ messages in thread
From: Kuba Piecuch @ 2026-09-29 16:17 UTC (permalink / raw)
To: Tejun Heo, David Vernet, Andrea Righi, Changwoo Min
Cc: Kuba Piecuch, Emil Tsalapatis, sched-ext, linux-kernel
Now that ops.dequeue() is called when a task arrives on a remote local
DSQ, the dequeue_remote test passes. Add it to auto-test-targets.
Assisted-by: Claude:claude-opus-5.5
Signed-off-by: Kuba Piecuch <jpiecuch@google.com>
---
tools/testing/selftests/sched_ext/Makefile | 1 +
1 file changed, 1 insertion(+)
diff --git a/tools/testing/selftests/sched_ext/Makefile b/tools/testing/selftests/sched_ext/Makefile
index 4e06d0baaeec..af919a56a4f4 100644
--- a/tools/testing/selftests/sched_ext/Makefile
+++ b/tools/testing/selftests/sched_ext/Makefile
@@ -165,6 +165,7 @@ auto-test-targets := \
create_dsq \
dequeue \
dequeue_iter \
+ dequeue_remote \
enq_last_no_enq_fails \
ddsp_bogus_dsq_fail \
ddsp_vtimelocal_fail \
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCHSET sched_ext/for-7.3-fixes] sched_ext: Fix missing ops.dequeue() on remote local DSQ moves
2026-09-29 16:17 [PATCHSET sched_ext/for-7.3-fixes] sched_ext: Fix missing ops.dequeue() on remote local DSQ moves Kuba Piecuch
` (2 preceding siblings ...)
2026-09-29 16:17 ` [PATCH 3/3] selftests/sched_ext: Enable the dequeue_remote test Kuba Piecuch
@ 2026-09-29 17:13 ` Tejun Heo
2026-09-29 18:30 ` Andrea Righi
3 siblings, 1 reply; 8+ messages in thread
From: Tejun Heo @ 2026-09-29 17:13 UTC (permalink / raw)
To: Kuba Piecuch, Andrea Righi
Cc: David Vernet, Changwoo Min, Emil Tsalapatis, sched-ext, linux-kernel
Hello,
On Tue, Sep 29, 2026 at 04:17:23PM +0000, Kuba Piecuch wrote:
> Since ebf1ccff79c4 ("sched_ext: Fix ops.dequeue() semantics"), every task
> entering the BPF scheduler's custody gets exactly one ops.dequeue() when
> it leaves it. A dispatch to a terminal DSQ ends custody with a flag-less
> ops.dequeue() at insertion time.
The fix looks good to me. It effectively reverts the enqueue_task_scx()
half of b75aaea24c9f ("sched_ext: Properly mark SCX-internal migrations
via sticky_cpu"), which as far as I can see only ever suppressed this
ops.dequeue(). Andrea, can you confirm?
- dequeue_remote.c isn't built until 3/3, so 1/3 can't be built or run.
Can you put the fix first, followed by the test with its Makefile entry?
- 2/3: 7.1.y also needs 18d62044cda7 ("sched_ext: Preserve rq tracking
across local DSQ dispatch"). Without it, the nested ops.dequeue() trips
lockdep when ops.dispatch() uses scx_bpf_dsq_move() to another CPU's
local DSQ. It's tagged for stable too, but maybe note it as a
prerequisite?
- 2/3: With sub-scheds, scx_resolve_local_dsq() can divert the task to the
reject or rescue DSQ, so "inserted into the local DSQ" in the comments
isn't always accurate. Maybe "destination DSQ"? The new comment in
enqueue_task_scx() could be two lines, and the description could lead
with the late SCX_DEQ_CORE_SCHED_EXEC and be a lot shorter.
- 1/3: A task can only be picked straight out of custody through
sched_core_find(), which only returns tasks with a core cookie. Checking
p->core_cookie on SCX_DEQ_CORE_SCHED_EXEC would be exact and would
remove core_sched_in_use() and the skip.
- 1/3: _SC_NPROCESSORS_ONLN ignores affinity. With the runner confined to
one CPU, the test fails instead of skipping. sched_getaffinity() and
CPU_COUNT()?
- 1/3: Nits. If the /proc scan stays, PR_SCHED_CORE_GET writes a u64, so
the cookie should be u64. ops.dispatch() pops one entry per call, so a
stale one idles the CPU until the next kick. Maybe loop a few times?
missed_dequeue_cnt and core_sched_exec_dequeue_cnt aren't printed
per-scenario like the other counters.
- 1/3: The variants, error conditions and core-sched caveat are repeated
across the cover, description, file header and comments. Can you say
each once? Also, single-line comments are usually lowercase in
sched_ext.
Thanks.
--
tejun
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCHSET sched_ext/for-7.3-fixes] sched_ext: Fix missing ops.dequeue() on remote local DSQ moves
2026-09-29 17:13 ` [PATCHSET sched_ext/for-7.3-fixes] sched_ext: Fix missing ops.dequeue() on remote local DSQ moves Tejun Heo
@ 2026-09-29 18:30 ` Andrea Righi
0 siblings, 0 replies; 8+ messages in thread
From: Andrea Righi @ 2026-09-29 18:30 UTC (permalink / raw)
To: Tejun Heo
Cc: Kuba Piecuch, David Vernet, Changwoo Min, Emil Tsalapatis,
sched-ext, linux-kernel
On Tue, Sep 29, 2026 at 07:13:07AM -1000, Tejun Heo wrote:
> Hello,
>
> On Tue, Sep 29, 2026 at 04:17:23PM +0000, Kuba Piecuch wrote:
> > Since ebf1ccff79c4 ("sched_ext: Fix ops.dequeue() semantics"), every task
> > entering the BPF scheduler's custody gets exactly one ops.dequeue() when
> > it leaves it. A dispatch to a terminal DSQ ends custody with a flag-less
> > ops.dequeue() at insertion time.
>
> The fix looks good to me. It effectively reverts the enqueue_task_scx()
> half of b75aaea24c9f ("sched_ext: Properly mark SCX-internal migrations
> via sticky_cpu"), which as far as I can see only ever suppressed this
> ops.dequeue(). Andrea, can you confirm?
Yes, I confirm. The source-side sticky_cpu assignment remains in place across
deactivate_task(), so the internal migration doesn't trigger ops.dequeue().
Kuba, thanks for catching this!
>
> - dequeue_remote.c isn't built until 3/3, so 1/3 can't be built or run.
> Can you put the fix first, followed by the test with its Makefile entry?
>
> - 2/3: 7.1.y also needs 18d62044cda7 ("sched_ext: Preserve rq tracking
> across local DSQ dispatch"). Without it, the nested ops.dequeue() trips
> lockdep when ops.dispatch() uses scx_bpf_dsq_move() to another CPU's
> local DSQ. It's tagged for stable too, but maybe note it as a
> prerequisite?
Agreed. Please mention 18d62044cda7 as a prerequisite for 7.1.y.
Thanks,
-Andrea
>
> - 2/3: With sub-scheds, scx_resolve_local_dsq() can divert the task to the
> reject or rescue DSQ, so "inserted into the local DSQ" in the comments
> isn't always accurate. Maybe "destination DSQ"? The new comment in
> enqueue_task_scx() could be two lines, and the description could lead
> with the late SCX_DEQ_CORE_SCHED_EXEC and be a lot shorter.
>
> - 1/3: A task can only be picked straight out of custody through
> sched_core_find(), which only returns tasks with a core cookie. Checking
> p->core_cookie on SCX_DEQ_CORE_SCHED_EXEC would be exact and would
> remove core_sched_in_use() and the skip.
>
> - 1/3: _SC_NPROCESSORS_ONLN ignores affinity. With the runner confined to
> one CPU, the test fails instead of skipping. sched_getaffinity() and
> CPU_COUNT()?
>
> - 1/3: Nits. If the /proc scan stays, PR_SCHED_CORE_GET writes a u64, so
> the cookie should be u64. ops.dispatch() pops one entry per call, so a
> stale one idles the CPU until the next kick. Maybe loop a few times?
> missed_dequeue_cnt and core_sched_exec_dequeue_cnt aren't printed
> per-scenario like the other counters.
>
> - 1/3: The variants, error conditions and core-sched caveat are repeated
> across the cover, description, file header and comments. Can you say
> each once? Also, single-line comments are usually lowercase in
> sched_ext.
>
> Thanks.
>
> --
> tejun
>
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH 2/3] sched_ext: Call ops.dequeue() when a task arrives on a remote local DSQ
2026-09-29 16:17 ` [PATCH 2/3] sched_ext: Call ops.dequeue() when a task arrives on a remote local DSQ Kuba Piecuch
@ 2026-09-29 18:33 ` Andrea Righi
0 siblings, 0 replies; 8+ messages in thread
From: Andrea Righi @ 2026-09-29 18:33 UTC (permalink / raw)
To: Kuba Piecuch
Cc: Tejun Heo, David Vernet, Changwoo Min, Emil Tsalapatis,
sched-ext, linux-kernel
Hi Kuba,
On Tue, Sep 29, 2026 at 04:17:25PM +0000, Kuba Piecuch wrote:
> When the BPF scheduler moves a task to the local DSQ of a CPU other than
> the one whose rq the task is on, e.g. with SCX_DSQ_LOCAL_ON dispatch,
> scx_bpf_dsq_move_to_local() or scx_bpf_dsq_move*(), the task is migrated
> by move_remote_task_to_local_dsq(). p->scx.sticky_cpu is set across the
> migration so that dequeueing the task from the source rq isn't treated as
> the task leaving the BPF scheduler's custody.
>
> However, enqueue_task_scx() only clears p->scx.sticky_cpu after
> scx_do_enqueue_task() has inserted the task into the destination local
> DSQ. task_leave_custody(), called from rq_owned_post_enq() on insertion,
> still sees the migration in progress and skips the custody exit. The task
> ends up on a terminal DSQ with SCX_TASK_IN_CUSTODY set and without
> ops.dequeue() having been called.
>
> The custody exit then happens only when the task is picked for execution,
> in set_next_task_scx(), which reports it to the BPF scheduler as
> ops.dequeue(SCX_DEQ_CORE_SCHED_EXEC) even though no core-sched pick took
> place. If a scheduling property change hits the task while it's waiting
> on the local DSQ, the BPF scheduler instead gets
> ops.dequeue(SCX_DEQ_SCHED_CHANGE) for a task that has already left its
> custody. Both break the ops.dequeue() semantics, under which a dispatch to
> a terminal DSQ ends custody with an ops.dequeue() call without special
> flags.
>
> Clear p->scx.sticky_cpu before calling scx_do_enqueue_task(). The routing
> decision in scx_do_enqueue_task() uses the local copy, and the departure
> side is unaffected as p->scx.sticky_cpu is still set across
> deactivate_task(). ops.dequeue() is now invoked on the destination rq when
> the task is inserted into the local DSQ, as it already is for same-rq
> dispatches.
>
> Fixes: ebf1ccff79c4 ("sched_ext: Fix ops.dequeue() semantics")
> Cc: stable@vger.kernel.org # v7.1+
> Assisted-by: Claude:claude-opus-5.5
> Signed-off-by: Kuba Piecuch <jpiecuch@google.com>
> ---
> kernel/sched/ext/ext.c | 16 +++++++++++-----
> 1 file changed, 11 insertions(+), 5 deletions(-)
>
> diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
> index ad391a8cbd05..763de7797056 100644
> --- a/kernel/sched/ext/ext.c
> +++ b/kernel/sched/ext/ext.c
> @@ -1501,8 +1501,9 @@ static inline bool task_scx_migrating(struct task_struct *p)
> /*
> * We only need to check sticky_cpu: it is set to the destination
> * CPU in move_remote_task_to_local_dsq() before deactivate_task()
> - * and cleared when the task is enqueued on the destination, so it
> - * is only non-negative during an internal SCX migration.
> + * and cleared in enqueue_task_scx() on the destination before @p is
> + * inserted into the local DSQ, so it is only non-negative while @p
> + * is in transit between the two rqs.
> */
> return p->scx.sticky_cpu >= 0;
> }
> @@ -2189,10 +2190,15 @@ static void enqueue_task_scx(struct rq *rq, struct task_struct *p, int core_enq_
> if (rq->scx.nr_running == 1)
> dl_server_start(&rq->ext_server);
>
> - scx_do_enqueue_task(rq, p, enq_flags, sticky_cpu);
> + /*
> + * An SCX-internal migration ends once @p arrives on the destination
> + * rq. Clear sticky_cpu before enqueueing so that @p leaves the BPF
> + * scheduler's custody when inserted into the destination local DSQ.
> + * The local copy in @sticky_cpu is used for routing.
> + */
> + p->scx.sticky_cpu = -1;
>
> - if (sticky_cpu >= 0)
> - p->scx.sticky_cpu = -1;
> + scx_do_enqueue_task(rq, p, enq_flags, sticky_cpu);
Nit: nothing between the top of enqueue_task_scx() and scx_do_enqueue_task()
looks at p->scx.sticky_cpu, so p->scx.sticky_cpu = -1 could go right after
reading it into the local variable, as it was before b75aaea24c9f. Same
behavior, but it also covers the SCX_TASK_QUEUED early exit, which currently
leaves p->scx.sticky_cpu set and essentially makes the fix a revert of the
enqueue side of b75aaea24c9f.
Either way:
Reviewed-by: Andrea Righi <arighi@nvidia.com>
Thanks,
-Andrea
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH 1/3] selftests/sched_ext: Add a test for ops.dequeue() on remote local DSQ moves
2026-09-29 16:17 ` [PATCH 1/3] selftests/sched_ext: Add a test for " Kuba Piecuch
@ 2026-09-29 18:42 ` Andrea Righi
0 siblings, 0 replies; 8+ messages in thread
From: Andrea Righi @ 2026-09-29 18:42 UTC (permalink / raw)
To: Kuba Piecuch
Cc: Tejun Heo, David Vernet, Changwoo Min, Emil Tsalapatis,
sched-ext, linux-kernel
Hi Kuba,
On Tue, Sep 29, 2026 at 04:17:24PM +0000, Kuba Piecuch wrote:
> When a BPF scheduler moves a task that is in its custody to the local DSQ
> of a CPU other than the one whose rq the task is on, sched_ext migrates the
> task with move_remote_task_to_local_dsq(). The task leaves the BPF
> scheduler's custody when it is inserted into the destination local DSQ and
> ops.dequeue() should be called at that point, just like for same-rq
> dispatches.
>
> Add a dequeue_remote test that exercises this path. All tasks are put
> into custody from ops.enqueue() and moved to the local DSQ of whichever
> CPU runs ops.dispatch(), so that most moves cross CPUs. Two variants are
> covered, selected by test_use_move_to_local:
>
> - false: tasks are queued in a BPF queue and dispatched with
> scx_bpf_dsq_insert(SCX_DSQ_LOCAL_ON | cpu), going through
> dispatch_to_local_dsq(),
>
> - true: tasks are queued on a user DSQ and consumed with
> scx_bpf_dsq_move_to_local(), going through consume_remote_task().
>
> The BPF scheduler tracks each task's custody state and triggers
> scx_bpf_error() if:
>
> - ops.dequeue() is called with SCX_DEQ_CORE_SCHED_EXEC, meaning that
> the custody exit was missed on insertion into the local DSQ and
> deferred to set_next_task_scx(),
>
> - a task starts running without ops.dequeue() having been called since
> its last ops.enqueue(),
>
> - ops.dequeue() is called for a task that isn't in custody, or
> ops.enqueue() for a task that is.
>
> Core scheduling can legitimately pick tasks straight out of custody, so
> the test is skipped if any task has a core scheduling cookie when it
> starts.
>
> The test currently fails:
>
> sched_ext: dequeue_remote: dequeue_remote.bpf.c:143: 156 (runner): late ops.dequeue() with SCX_DEQ_CORE_SCHED_EXEC (enq_cpu=1 cpu=3 seq=1)
> ...
> ops_dequeue+0x114/0x170
> set_next_task_scx+0x104/0x1e0
> __pick_next_task+0xc7/0x180
> __schedule+0x154/0x1870
>
> so don't add it to auto-test-targets yet. It will be enabled once the
> underlying bug is fixed.
>
> Assisted-by: Claude:claude-opus-5.5
> Signed-off-by: Kuba Piecuch <jpiecuch@google.com>
We can probably fold the this into PATCH 3.
Thanks,
-Andrea
^ permalink raw reply [flat|nested] 8+ messages in thread
end of thread, other threads:[~2026-09-29 18:42 UTC | newest]
Thread overview: 8+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-29 16:17 [PATCHSET sched_ext/for-7.3-fixes] sched_ext: Fix missing ops.dequeue() on remote local DSQ moves Kuba Piecuch
2026-09-29 16:17 ` [PATCH 1/3] selftests/sched_ext: Add a test for " Kuba Piecuch
2026-09-29 18:42 ` Andrea Righi
2026-09-29 16:17 ` [PATCH 2/3] sched_ext: Call ops.dequeue() when a task arrives on a remote local DSQ Kuba Piecuch
2026-09-29 18:33 ` Andrea Righi
2026-09-29 16:17 ` [PATCH 3/3] selftests/sched_ext: Enable the dequeue_remote test Kuba Piecuch
2026-09-29 17:13 ` [PATCHSET sched_ext/for-7.3-fixes] sched_ext: Fix missing ops.dequeue() on remote local DSQ moves Tejun Heo
2026-09-29 18:30 ` Andrea Righi
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®