From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pj1-f51.google.com (mail-pj1-f51.google.com [209.85.216.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6F4A130C60D for ; Wed, 22 Apr 2026 06:33:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.51 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1776839627; cv=none; b=fYbN5k/4eB7gWzXLgxFub7u1BMrpswo3vmwkAB0wyogY9m93EC4HwTcLdhQD9r4t/z1PyWtJyK+PN8/Pqi+1dGLGSA3bBZX+PLqjt/6Ys/Yjy5nsKelc+/Q5WpmOZcGQOTx2IOHH+aAA4UZN7Nc3/gveIIamHljA/CMojqbvmdo= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1776839627; c=relaxed/simple; bh=gFeiypU5qclQlKyGSympAwDRn3JZ+YXx7QpsThI06dg=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=V9g8vd5b3VZQCwddz/VujV9ih+20bVRZAgfa6ttduGFmclN3tKqKaYB17sCYIiPluTggtJQc7DBqnCHsBO0/S+S1OW0jEQfT0/k/49hbloK3CMMO22QuKlLG7a40MxWWMTEZsigAGO5kMiF1cG2VY9CdC1INzzh3AH7Ll5S4o2I= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=BP5MdeqC; arc=none smtp.client-ip=209.85.216.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="BP5MdeqC" Received: by mail-pj1-f51.google.com with SMTP id 98e67ed59e1d1-35d95017a68so3360375a91.3 for ; Tue, 21 Apr 2026 23:33:46 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1776839626; x=1777444426; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=qVY0Ul7cq0NSPU/FuepEJhjqjn4ieUx/hBL/nitwkMs=; b=BP5MdeqCkAZfn1UGQ61+Q7cP3RyoRpLU6s5rhJEQ4+5VKxxdMWLOech75YSpNWqfld p2wcYhVMFPcghQ+N8/MeUuEWiKKwYiBgHCuLuPuLmUFTjWzW3Mlecl/VhhRSkWnbATP3 dcSszOmayocS+fvOCQXHJXSSgBS5fp+ZU2sARTN6RhsOCFAKDJr+SYidVqUlv7YzmdaN 9V0qERhk88NF9AogiOWAG34/XtgAZXcCXIsO6+G7VEwI4peL8dpL7gMnu1Rp2YSCgcwx FmIQxsJ3UjUBBQWuW4hgY1QKZ6jdCKlhKKw9ofTwvbHGcOkTgYlD8kchKnB/LD6rbBoj 1dtQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1776839626; x=1777444426; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-gg:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=qVY0Ul7cq0NSPU/FuepEJhjqjn4ieUx/hBL/nitwkMs=; b=AHB3Wx2RVJ3Qt3yOQLajlmFAsXJ9BmjgDgwZwrXnxaR3TEQN11kbUU4Q6CACUoofdW K2mwVssi+CaPI06JPkMdqNTIsxVU3sMhAGjf5xSmKAfg5KKlyeUVdjWUycp63ySAr3Bw CzH2EfonROcmXg4f+12u8cSLcz44irnqA5339TPEScgLdj/B6PvYUoSoVeJ1zInxIKdS 17uIrfTjsOYClVHxt+bPUc3FxBFuuvQlk6Zqu99a3zkH92xYBBaWhc8LLOCTSy89R5/W rtvlFWlWH/wHF2qe6HMUlc/YuyzB2qUmz5Jglm3UwRYgO0csaNWGE//wPzakdPDyY7a1 s/lw== X-Forwarded-Encrypted: i=1; AFNElJ9AxlTB2SCulC/AFtrV/XRq37l/MP+tNcLN3PQGQuuwypKTVx/RuKCqXOtKb/AZGBadU9OItYDfsGDusPI=@vger.kernel.org X-Gm-Message-State: AOJu0Yz1eKXRTJoHmgJ6K1FzlWgmrm+hyP3+r5FNRUc2MfH9PLxVK6ot VwcL+yWxdEwEsclsZm14lUCT4eKd3jROYAz0VcBilTBblTFLV1jSf11Z X-Gm-Gg: AeBDies1WJiOcbB3y2wnfHRhCyvVQr+HLKPJ+CNlQ0gAsPwkeWeeKrD5w08CrV59V5g M7z3VIiKF4zwEpoOaErHHc0gpAgZkc6r6DXQnywEM+ibKrvhXdXEpTm1Ake9lAQUQUNIqSjYOyd IbmScuDlK0iAK8RiVV0fIFBkXzFj2rRC+45r/9Ty5pccP7xQuH3J3u16+UbvimxCnOJpTbHxpNu dUIxgv2kuBMRzMNcon1OkYeBvcI2y5ZNC1MyxKpe+TDEeQ//uW8Wj558x4SQWeO8MIGwSxlyp5o wnuNRyH/7/CHMKizsAfGiCPniGJn5fC5q7an9iliTFQSMDnqfQ0uZJ6vS6KKInvd81bQKvoftaq I875zOkT+gqtwdDeZvBRb+WbXdys/05w/hN0hHQhADFMVsVYvxAazOYWSpWSlOSLQWInCmr87TR 41lCLFFdmfMgLSaLSjkWsoXv43rMKeogJrRsmJRnE0gDle18BebIghK6/BSCo8TeJ+/l4wZsS0a SWVInholYoh0weQ X-Received: by 2002:a17:903:3c70:b0:2b2:58c7:2ce1 with SMTP id d9443c01a7336-2b5fa02f3famr229855325ad.36.1776839625681; Tue, 21 Apr 2026 23:33:45 -0700 (PDT) Received: from cchengyang.duckdns.org (36-225-97-241.dynamic-ip.hinet.net. [36.225.97.241]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2b5fa9ff409sm193209615ad.14.2026.04.21.23.33.42 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 21 Apr 2026 23:33:45 -0700 (PDT) Date: Wed, 22 Apr 2026 14:33:40 +0800 From: Cheng-Yang Chou To: Tejun Heo Cc: Kuba Piecuch , Andrea Righi , David Vernet , Changwoo Min , Emil Tsalapatis , Christian Loehle , Daniel Hodges , sched-ext@lists.linux.dev, linux-kernel@vger.kernel.org, Ching-Chun Huang , Chia-Ping Tsai Subject: Re: [PATCH v2 sched_ext/for-7.1] sched_ext: Invalidate dispatch decisions on CPU affinity changes Message-ID: <20260422142633.G7180@cchengyang.duckdns.org> References: <20260319083518.94673-1-arighi@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: Hi Tejun, Andrea, and Kuba On Mon, Mar 23, 2026 at 01:13:20PM -1000, Tejun Heo wrote: > > The simple way to do this is to do scx_bpf_dsq_insert() at the very beginning, > > once we know which task we would like to dispatch, and cancel the pending > > dispatch via scx_bpf_dispatch_cancel() if any of the pre-dispatch checks fail > > on the BPF side. This way, the "critical section" includes BPF-side checks, and > > SCX will ignore the dispatch if there was a dequeue/enqueue racing with the > > critical section. > > > > With this solution, we can throw an error if task_can_run_on_remote_rq() is > > false, because we know that there was no racing cpumask change (if there was, > > it would have been caught earlier, in finish_dispatch()). > > Yeah, I think this makes more sense. qseq is already there to provide > protection against these events. It's just that the capturing of qseq is too > late. If insert/cancel is too ugly, we can introduce another kfunc to > capture the qseq - scx_bpf_dsq_insert_begin() or something like that - and > stash it in a per-cpu variable. That way, qseq would be cover the "current" > queued instance and the existing qseq mechanism would be able to reliably > ignore the ones that lost race to dequeue. Since this has been stale for a while, I prepared a patch to implement scx_bpf_dsq_insert_begin() as suggested. Is anyone else working on this? If not, I'm happy to send the formal patch to fix this. -- Cheers, Cheng-Yang diff --git a/kernel/sched/ext.c b/kernel/sched/ext.c index 0a53a0dd64bf..0215a21a02db 100644 --- a/kernel/sched/ext.c +++ b/kernel/sched/ext.c @@ -7933,6 +7933,7 @@ static void scx_dsq_insert_commit(struct scx_sched *sch, struct task_struct *p, { struct scx_dsp_ctx *dspc = &this_cpu_ptr(sch->pcpu)->dsp_ctx; struct task_struct *ddsp_task; + unsigned long qseq; ddsp_task = __this_cpu_read(direct_dispatch_task); if (ddsp_task) { @@ -7945,9 +7946,16 @@ static void scx_dsq_insert_commit(struct scx_sched *sch, struct task_struct *p, return; } + if (dspc->insert_begin_valid) { + qseq = dspc->insert_begin_qseq; + dspc->insert_begin_valid = false; + } else { + qseq = atomic_long_read(&p->scx.ops_state) & SCX_OPSS_QSEQ_MASK; + } + dspc->buf[dspc->cursor++] = (struct scx_dsp_buf_ent){ .task = p, - .qseq = atomic_long_read(&p->scx.ops_state) & SCX_OPSS_QSEQ_MASK, + .qseq = qseq, .dsq_id = dsq_id, .enq_flags = enq_flags, }; @@ -7955,6 +7963,39 @@ static void scx_dsq_insert_commit(struct scx_sched *sch, struct task_struct *p, __bpf_kfunc_start_defs(); +/** + * scx_bpf_dsq_insert_begin - Snapshot qseq before a dispatch decision + * @p: task_struct being considered for dispatch + * @aux: implicit BPF argument to access bpf_prog_aux hidden from BPF progs + * + * Capture @p's qseq before the BPF scheduler reads @p's properties (e.g. + * cpus_ptr) to make a dispatch decision. The snapshot is used by the + * subsequent scx_bpf_dsq_insert() call, extending the race detection window + * to cover any BPF-side checks between this call and the insert. If a + * concurrent dequeue/re-enqueue races within this window, finish_dispatch() + * detects the qseq mismatch and discards the stale dispatch. + */ +__bpf_kfunc void scx_bpf_dsq_insert_begin(struct task_struct *p, + const struct bpf_prog_aux *aux) +{ + struct scx_sched *sch; + struct scx_dsp_ctx *dspc; + + guard(rcu)(); + + sch = scx_prog_sched(aux); + if (unlikely(!sch)) + return; + + if (!scx_kf_allowed(sch, SCX_KF_ENQUEUE | SCX_KF_DISPATCH)) + return; + + dspc = &this_cpu_ptr(sch->pcpu)->dsp_ctx; + dspc->insert_begin_qseq = atomic_long_read(&p->scx.ops_state) & + SCX_OPSS_QSEQ_MASK; + dspc->insert_begin_valid = true; +} + /** * scx_bpf_dsq_insert - Insert a task into the FIFO queue of a DSQ * @p: task_struct to insert @@ -8134,6 +8175,7 @@ __bpf_kfunc void scx_bpf_dsq_insert_vtime(struct task_struct *p, u64 dsq_id, __bpf_kfunc_end_defs(); BTF_KFUNCS_START(scx_kfunc_ids_enqueue_dispatch) +BTF_ID_FLAGS(func, scx_bpf_dsq_insert_begin, KF_IMPLICIT_ARGS | KF_RCU) BTF_ID_FLAGS(func, scx_bpf_dsq_insert, KF_IMPLICIT_ARGS | KF_RCU) BTF_ID_FLAGS(func, scx_bpf_dsq_insert___v2, KF_IMPLICIT_ARGS | KF_RCU) BTF_ID_FLAGS(func, __scx_bpf_dsq_insert_vtime, KF_IMPLICIT_ARGS | KF_RCU) diff --git a/kernel/sched/ext_internal.h b/kernel/sched/ext_internal.h index 4a7ffc7f55d2..adc4f1c01b56 100644 --- a/kernel/sched/ext_internal.h +++ b/kernel/sched/ext_internal.h @@ -989,6 +989,8 @@ struct scx_dsp_ctx { struct rq *rq; u32 cursor; u32 nr_tasks; + unsigned long insert_begin_qseq; + bool insert_begin_valid; struct scx_dsp_buf_ent buf[]; }; diff --git a/tools/sched_ext/scx_central.bpf.c b/tools/sched_ext/scx_central.bpf.c index 64dd60b3e922..fb68a7d7e201 100644 --- a/tools/sched_ext/scx_central.bpf.c +++ b/tools/sched_ext/scx_central.bpf.c @@ -155,6 +155,8 @@ static bool dispatch_to_cpu(s32 cpu) * reflect the migration-disabled state yet if * migrate_disable_switch() hasn't run. */ + scx_bpf_dsq_insert_begin(p); + if (!bpf_cpumask_test_cpu(cpu, p->cpus_ptr) || (is_migration_disabled(p) && scx_bpf_task_cpu(p) != cpu)) { __sync_fetch_and_add(&nr_mismatches, 1); -- -- 2.48.1