mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Andrea Righi <arighi@nvidia.com>
To: Wanwu Li <liwanwu@kylinos.cn>
Cc: Tejun Heo <tj@kernel.org>, David Vernet <void@manifault.com>,
	Changwoo Min <changwoo@igalia.com>,
	sched-ext@lists.linux.dev, linux-kernel@vger.kernel.org,
	stable@vger.kernel.org
Subject: Re: [PATCH v2] sched_ext: Fix NULL sched deref in kfunc sub-sched error paths
Date: Wed, 2 Sep 2026 20:11:14 +0200	[thread overview]
Message-ID: <aphmwjLkdS4tKpbv@gpd4> (raw)
In-Reply-To: <20260902170751.256434-1-liwanwu@kylinos.cn>

Hi Wanwu,

On Thu, Sep 03, 2026 at 01:07:51AM +0800, Wanwu Li wrote:
> The COMPAT kfunc wrappers scx_bpf_select_cpu_and() and
> scx_bpf_dsq_insert_vtime() error out @p's scheduler when the root
> scheduler has sub-scheds attached:
> 
> 	scx_error(scx_task_sched(p), "... must be used");
> 
> scx_task_sched(p) is p->scx.sched, which is NULL for any task that is not
> on an scx scheduler: it is memset() by init_scx_entity() and cleared by
> scx_disable_and_exit_task() -- which sched_ext_dead() runs when a task
> exits. It is also an rcu_dereference_protected() that must be called with
> @p's pi_lock or rq lock held, which neither wrapper does. Passing NULL to
> scx_error() reaches scx_vexit(), which dereferences sch->exit_info
> unconditionally, oopsing the kernel.
> 
> Both wrappers are reachable with such a @p. scx_bpf_select_cpu_and() is in
> the select_cpu kfunc group, which scx_kfunc_context_filter() opens to
> BPF_PROG_TYPE_SYSCALL programs. scx_bpf_dsq_insert_vtime() is in the
> enqueue_dispatch group, which ops.enqueue() and ops.dispatch() may call
> with any KF_RCU task: the group has no kf_tasks validation, and
> scx_dsq_insert_preamble() checks task ownership with scx_task_on_sched()
> precisely because @p may be an arbitrary task.
> 
> Neither wrapper requires a contrived @p. Tasks that are never enabled --
> kthreads and tasks of other classes under SCX_SWITCH_ALL=n -- keep
> p->scx.sched NULL indefinitely; and a task handed over from
> bpf_task_from_pid() can exit before the call lands, as its zombie stays
> visible to pid lookups until it is reaped.
> 
> One concrete trigger exercised for this changelog: a SYSCALL program
> calling the select_cpu_and wrapper on an exited-but-not-reaped task while
> a sub-scheduler is attached (faulting instruction is the scx_vexit()
> prologue "mov r15,[rdi+0x398]" with RDI=NULL and 0x398 the offset of
> sch->exit_info):
> 
>   sched_ext: BPF scheduler "kfunc_subsched_null" enabled
>   sched_ext: BPF sub-scheduler "kfunc_subsched_null" enabled
>   sched_ext: Unassociated program run_select_cpu_ (id 76)
>   BUG: kernel NULL pointer dereference, address: 0000000000000398
>   #PF: supervisor read access in kernel mode
>   #PF: error_code(0x0000) - not-present page
>   Oops: Oops: 0000 [#1] SMP NOPTI
>   CPU: 7 UID: 0 PID: 8201 Comm: kfunc_test_runn Tainted: G W
>   RIP: 0010:scx_vexit+0x25/0xa0
>   Code: ... <4c> 8b bf 98 03 00 00 ...
>   CR2: 0000000000000398
>   Call Trace:
>    <TASK>
>    __scx_exit+0x4f/0x70
>    scx_bpf_select_cpu_and+0xab/0xb0
>    bpf_prog_430ed61a7b66e03a_run_select_cpu_and+0x9c/0xe7
>    ? __x64_sys_bpf+0x2c/0x40
>    bpf_prog_test_run_syscall+0x130/0x2f0
>    __sys_bpf+0x930/0x10d0
>    ? __x64_sys_bpf+0x2c/0x40
>    __x64_sys_bpf+0x2c/0x40
>    do_syscall_64+0xbc/0x460
>    ? rseq_set_ids_get_csaddr+0x81/0x140
>    ? __rseq_handle_slowpath+0xd0/0x130
>    ? switch_fpu_return+0x51/0xd0
>    ? arch_exit_to_user_mode_prepare.constprop.0+0x87/0xb0
>    ? do_syscall_64+0xf3/0x460
>    ? irqentry_exit+0x48/0x740
>    ? clear_bhb_loop+0x40/0x90
>    ? do_syscall_64+0x35/0x460
>    entry_SYSCALL_64_after_hwframe+0x76/0x7e
>    </TASK>
> 
> Keep the "error out @p's scheduler" attribution -- it is what every other
> kfunc error path does (select_cpu_from_kfunc()'s cross_task,
> scx_kf_arg_task_ok()) and it is correct for the callers that take this
> path with a live task: p->scx.sched is their scheduler. Just read it
> safely: use scx_task_sched_rcu() (valid under the guard(rcu)() both
> wrappers already hold, no @p lock required) and fall back to @sch -- the
> root scheduler, guaranteed non-NULL here -- when @p is not on an scx
> scheduler, which is precisely the case that used to be NULL.
> 
> These COMPAT wrappers are scheduled for eventual removal once the
> deprecation grace period elapses, but until then -- and regardless of
> their removal timeline -- they must not oops the kernel on a task they
> are handed; this fix makes the error path safe.
> 
> Cc: <stable@vger.kernel.org>
> Fixes: a5fa0708cbfd ("sched_ext: Enforce scheduling authority in dispatch and select_cpu operations")
> Signed-off-by: Wanwu Li <liwanwu@kylinos.cn>

This looks good to me.

Reviewed-by: Andrea Righi <arighi@nvidia.com>

Thanks,
-Andrea

> ---
> 
> Changes v1 -> v2:
>  - Extend the same fallback to scx_bpf_dsq_insert_vtime() as suggested
>    by Andrea (and flagged by sashiko-bot); rewrite the reachability
>    argument in the commit message accordingly.
> 
>  kernel/sched/ext/ext.c  | 9 +++++++--
>  kernel/sched/ext/idle.c | 9 +++++++--
>  2 files changed, 14 insertions(+), 4 deletions(-)
> 
> diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
> index 10af28a9f2c0..fdfaa7e9c8f5 100644
> --- a/kernel/sched/ext/ext.c
> +++ b/kernel/sched/ext/ext.c
> @@ -8943,10 +8943,15 @@ __bpf_kfunc void scx_bpf_dsq_insert_vtime(struct task_struct *p, u64 dsq_id,
>  #ifdef CONFIG_EXT_SUB_SCHED
>  	/*
>  	 * Disallow if any sub-scheds are attached. There is no way to tell
> -	 * which scheduler called us, just error out @p's scheduler.
> +	 * which scheduler called us, so error out @p's scheduler -- but read
> +	 * it under RCU (@p's locks aren't necessarily held here) and fall
> +	 * back to @sch if @p isn't on an scx scheduler: enqueue/dispatch
> +	 * contexts may pass any KF_RCU task, and p->scx.sched is NULL for
> +	 * one that has exited or is managed by another scheduler.
>  	 */
>  	if (unlikely(!list_empty(&sch->children))) {
> -		scx_error(scx_task_sched(p), "__scx_bpf_dsq_insert_vtime() must be used");
> +		scx_error(scx_task_sched_rcu(p) ?: sch,
> +			  "__scx_bpf_dsq_insert_vtime() must be used");
>  		return;
>  	}
>  #endif
> diff --git a/kernel/sched/ext/idle.c b/kernel/sched/ext/idle.c
> index d2973fb3af6d..014599d82bb0 100644
> --- a/kernel/sched/ext/idle.c
> +++ b/kernel/sched/ext/idle.c
> @@ -1142,10 +1142,15 @@ __bpf_kfunc s32 scx_bpf_select_cpu_and(struct task_struct *p, s32 prev_cpu, u64
>  #ifdef CONFIG_EXT_SUB_SCHED
>  	/*
>  	 * Disallow if any sub-scheds are attached. There is no way to tell
> -	 * which scheduler called us, just error out @p's scheduler.
> +	 * which scheduler called us, so error out @p's scheduler -- but read
> +	 * it under RCU (@p's locks aren't held here) and fall back to @sch if
> +	 * @p isn't on an scx scheduler: a BPF_PROG_TYPE_SYSCALL prog can pass
> +	 * any task and p->scx.sched is NULL for one that has exited or is
> +	 * managed by another scheduler.
>  	 */
>  	if (unlikely(!list_empty(&sch->children))) {
> -		scx_error(scx_task_sched(p), "__scx_bpf_select_cpu_and() must be used");
> +		scx_error(scx_task_sched_rcu(p) ?: sch,
> +			  "__scx_bpf_select_cpu_and() must be used");
>  		return -EINVAL;
>  	}
>  #endif
> -- 
> 2.25.1
> 

  reply	other threads:[~2026-09-02 18:11 UTC|newest]

Thread overview: 9+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-02 15:36 [PATCH] sched_ext: Fix NULL sched deref in select_cpu_and sub-sched error path Wanwu Li
2026-09-02 16:14 ` Andrea Righi
2026-09-02 16:34   ` liwanwu
2026-09-02 17:07 ` [PATCH v2] sched_ext: Fix NULL sched deref in kfunc sub-sched error paths Wanwu Li
2026-09-02 18:11   ` Andrea Righi [this message]
2026-09-02 19:37   ` Tejun Heo
2026-09-02 19:39     ` Tejun Heo
2026-09-03  6:06 ` [PATCH v3] " Wanwu Li
2026-09-03 17:56   ` Tejun Heo

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=aphmwjLkdS4tKpbv@gpd4 \
    --to=arighi@nvidia.com \
    --cc=changwoo@igalia.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=liwanwu@kylinos.cn \
    --cc=sched-ext@lists.linux.dev \
    --cc=stable@vger.kernel.org \
    --cc=tj@kernel.org \
    --cc=void@manifault.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®