From: Andrea Righi <arighi@nvidia.com>
To: Wanwu Li <liwanwu@kylinos.cn>
Cc: Tejun Heo <tj@kernel.org>, David Vernet <void@manifault.com>,
Changwoo Min <changwoo@igalia.com>,
sched-ext@lists.linux.dev, linux-kernel@vger.kernel.org,
stable@vger.kernel.org
Subject: Re: [PATCH] sched_ext: Fix NULL sched deref in select_cpu_and sub-sched error path
Date: Wed, 2 Sep 2026 18:14:18 +0200 [thread overview]
Message-ID: <aphLWtCVY1XVME9C@gpd4> (raw)
In-Reply-To: <20260902153640.144791-1-liwanwu@kylinos.cn>
Hi Wanwu,
On Wed, Sep 02, 2026 at 11:36:40PM +0800, Wanwu Li wrote:
> scx_bpf_select_cpu_and() errors out @p's scheduler when the root scheduler
> has sub-scheds attached:
>
> scx_error(scx_task_sched(p), "... must be used");
>
> scx_task_sched(p) is p->scx.sched, which is NULL for any task that is not
> on an scx scheduler: it is memset() by init_scx_entity() and cleared by
> scx_disable_and_exit_task() -- which sched_ext_dead() runs when a task
> exits. It is also an rcu_dereference_protected() that must be called with
> @p's pi_lock or rq lock held -- neither of which a BPF_PROG_TYPE_SYSCALL
> program holds.
>
> The wrapper is reachable from such a program -- scx_kfunc_context_filter()
> allows the select_cpu kfunc group for BPF_PROG_TYPE_SYSCALL -- and the
> program can pass any task, e.g. one obtained with bpf_task_from_pid() that
> exited in between. scx_error() then calls scx_vexit(), which dereferences
> sch->exit_info unconditionally, so passing NULL oopses the kernel.
>
> This was triggered live on a v7.2 based kernel with a
> BPF_PROG_TYPE_SYSCALL program calling the wrapper on an exited-but-not
> reaped task while a sub-scheduler was attached (faulting instruction is
> the scx_vexit() prologue "mov r15,[rdi+0x398]" with RDI=NULL and 0x398
> the offset of sch->exit_info):
>
> sched_ext: BPF scheduler "kfunc_subsched_null" enabled
> sched_ext: BPF sub-scheduler "kfunc_subsched_null" enabled
> sched_ext: Unassociated program run_select_cpu_ (id 76)
> BUG: kernel NULL pointer dereference, address: 0000000000000398
> #PF: supervisor read access in kernel mode
> #PF: error_code(0x0000) - not-present page
> Oops: Oops: 0000 [#1] SMP NOPTI
> CPU: 7 UID: 0 PID: 8201 Comm: kfunc_test_runn Tainted: G W
> RIP: 0010:scx_vexit+0x25/0xa0
> Code: ... <4c> 8b bf 98 03 00 00 ...
> CR2: 0000000000000398
> Call Trace:
> <TASK>
> __scx_exit+0x4f/0x70
> scx_bpf_select_cpu_and+0xab/0xb0
> bpf_prog_430ed61a7b66e03a_run_select_cpu_and+0x9c/0xe7
> ? __x64_sys_bpf+0x2c/0x40
> bpf_prog_test_run_syscall+0x130/0x2f0
> __sys_bpf+0x930/0x10d0
> ? __x64_sys_bpf+0x2c/0x40
> __x64_sys_bpf+0x2c/0x40
> do_syscall_64+0xbc/0x460
> ? rseq_set_ids_get_csaddr+0x81/0x140
> ? __rseq_handle_slowpath+0xd0/0x130
> ? switch_fpu_return+0x51/0xd0
> ? arch_exit_to_user_mode_prepare.constprop.0+0x87/0xb0
> ? do_syscall_64+0xf3/0x460
> ? irqentry_exit+0x48/0x740
> ? clear_bhb_loop+0x40/0x90
> ? do_syscall_64+0x35/0x460
> entry_SYSCALL_64_after_hwframe+0x76/0x7e
> </TASK>
>
> Keep the "error out @p's scheduler" attribution -- it is what every other
> kfunc error path does (select_cpu_from_kfunc()'s cross_task,
> scx_kf_arg_task_ok()) and it is correct for the callers that actually take
> this path: the only other callers are the struct_ops select_cpu/enqueue
> ops, where @p is the caller's own task and p->scx.sched is its scheduler.
> Just read it safely: use scx_task_sched_rcu() (valid under the guard(rcu)()
> the wrapper already holds, no @p lock required) and fall back to @sch -- the
> root scheduler, guaranteed non-NULL here -- when @p is not on an scx
> scheduler, which is precisely the case that used to be NULL.
>
> scx_bpf_dsq_insert_vtime() has the same error path but it is not reachable
> with a NULL @p: SYSCALL programs are rejected for its kfunc set and @p is
> always the calling scheduler's own task in the contexts where it runs, so
> it is left unchanged.
The reported crash looks valid to me, and using scx_task_sched_rcu() with the
root scheduler as fallback also looks correct.
However, as also pointed out by sashiko, the assumption above doesn't hold for
scx_bpf_dsq_insert_vtime(), although SYSCALL programs can't call it, STRUCT_OPS
programs can call it from ops.enqueue() and ops.dispatch().
Can you update scx_bpf_dsq_insert_vtime() as well with the same fallback?
Thanks,
-Andrea
>
> The proper endgame for these COMPAT wrappers is removal once the
> deprecation grace period is announced and elapsed; this fix only keeps
> the window from oopsing the kernel until that happens.
>
> Cc: <stable@vger.kernel.org>
> Fixes: a5fa0708cbfd ("sched_ext: Enforce scheduling authority in dispatch and select_cpu operations")
> Signed-off-by: Wanwu Li <liwanwu@kylinos.cn>
> ---
> kernel/sched/ext/idle.c | 9 +++++++--
> 1 file changed, 7 insertions(+), 2 deletions(-)
>
> diff --git a/kernel/sched/ext/idle.c b/kernel/sched/ext/idle.c
> index d2973fb3af6d..014599d82bb0 100644
> --- a/kernel/sched/ext/idle.c
> +++ b/kernel/sched/ext/idle.c
> @@ -1142,10 +1142,15 @@ __bpf_kfunc s32 scx_bpf_select_cpu_and(struct task_struct *p, s32 prev_cpu, u64
> #ifdef CONFIG_EXT_SUB_SCHED
> /*
> * Disallow if any sub-scheds are attached. There is no way to tell
> - * which scheduler called us, just error out @p's scheduler.
> + * which scheduler called us, so error out @p's scheduler -- but read
> + * it under RCU (@p's locks aren't held here) and fall back to @sch if
> + * @p isn't on an scx scheduler: a BPF_PROG_TYPE_SYSCALL prog can pass
> + * any task and p->scx.sched is NULL for one that has exited or is
> + * managed by another scheduler.
> */
> if (unlikely(!list_empty(&sch->children))) {
> - scx_error(scx_task_sched(p), "__scx_bpf_select_cpu_and() must be used");
> + scx_error(scx_task_sched_rcu(p) ?: sch,
> + "__scx_bpf_select_cpu_and() must be used");
> return -EINVAL;
> }
> #endif
> --
> 2.25.1
>
next prev parent reply other threads:[~2026-09-02 16:14 UTC|newest]
Thread overview: 9+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-02 15:36 Wanwu Li
2026-09-02 16:14 ` Andrea Righi [this message]
2026-09-02 16:34 ` liwanwu
2026-09-02 17:07 ` [PATCH v2] sched_ext: Fix NULL sched deref in kfunc sub-sched error paths Wanwu Li
2026-09-02 18:11 ` Andrea Righi
2026-09-02 19:37 ` Tejun Heo
2026-09-02 19:39 ` Tejun Heo
2026-09-03 6:06 ` [PATCH v3] " Wanwu Li
2026-09-03 17:56 ` Tejun Heo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aphLWtCVY1XVME9C@gpd4 \
--to=arighi@nvidia.com \
--cc=changwoo@igalia.com \
--cc=linux-kernel@vger.kernel.org \
--cc=liwanwu@kylinos.cn \
--cc=sched-ext@lists.linux.dev \
--cc=stable@vger.kernel.org \
--cc=tj@kernel.org \
--cc=void@manifault.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®