From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mx0b-0031df01.pphosted.com (mx0b-0031df01.pphosted.com [205.220.180.131]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 132503E6DD8 for ; Mon, 15 Jun 2026 12:17:40 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=205.220.180.131 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781525861; cv=none; b=RFzm9Za6rVlellUiopr/IUaDjoo9pVPII95IR6fiRL0qWF6EkySXIYEs3Z6ETaqGMQGjmSDC0tcFSNBy/STNx/uRNjZtRHTnsGj8Qg/kxENWrvRdl1OUKKf+lKKFJl/zgmAnWsC9idVDQy5l49rkxSuDy4372fsB8ZEgeABfk3A= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781525861; c=relaxed/simple; bh=16fLVsgJwQw8RxGuNkUZmp/ZZRob8n63K+bKHNgIvhg=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=rJQUg2DKitp2AVAJrNal8E65L5W1x/iT6JLA+DtN9fFL+FjaZhdK2HxTLYgAcF12zZ5306HvL1XM1jn3y2HLgs6U2BmfMSrlE4zcpu4R88RZS4ZlKXU0zTsXbI9YaD65PE7CIbaG2wbiwSZhne+ydj2neg4p3yDwNVPADNtdn+Y= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=oss.qualcomm.com; spf=pass smtp.mailfrom=oss.qualcomm.com; dkim=pass (2048-bit key) header.d=qualcomm.com header.i=@qualcomm.com header.b=ovT7MLQi; dkim=pass (2048-bit key) header.d=oss.qualcomm.com header.i=@oss.qualcomm.com header.b=ZOdRS/6m; arc=none smtp.client-ip=205.220.180.131 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=oss.qualcomm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=oss.qualcomm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=qualcomm.com header.i=@qualcomm.com header.b="ovT7MLQi"; dkim=pass (2048-bit key) header.d=oss.qualcomm.com header.i=@oss.qualcomm.com header.b="ZOdRS/6m" Received: from pps.filterd (m0279868.ppops.net [127.0.0.1]) by mx0a-0031df01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 65FApdrE151819 for ; Mon, 15 Jun 2026 12:17:39 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=qualcomm.com; h= cc:content-transfer-encoding:content-type:date:from:in-reply-to :message-id:mime-version:references:subject:to; s=qcppdkim1; bh= n+cg/szuB5nopCfExxFx7dLfNUGtVuWLcGK9H/4gV+8=; b=ovT7MLQiBYVvJNiQ f1mkQX3XXijHtQ9LV9/YdyM2HLM0Ya8O+8QqcEk4bQ+wonLK6YDER/Cjj1CAdnjJ 2lzyCfTkbF45tebmzrXBEkrxLiWzhA/6mnYkPnUH5XAd+5b86L98Lo+2XL60tgxY RdJy+s+3g63XmlEZp5m3CU04rIuLTK2XB88OpzE0bN9P6b0FO1DPemlEc1t8jAQO okY2ZFa6ie8h0Vb/LEy8nX5rgFJ4Ie3rXszHOp6ygltkYNVHtgWRDnUbxVCKhnZ1 lmVk04fXm14YuAI66/yDpXZSrYkCMVukGlQe92h1bmzjcbzHO2+ZiOo6zk7HiyJ8 H1z/Lg== Received: from mail-pj1-f72.google.com (mail-pj1-f72.google.com [209.85.216.72]) by mx0a-0031df01.pphosted.com (PPS) with ESMTPS id 4etetf0kva-1 (version=TLSv1.3 cipher=TLS_AES_128_GCM_SHA256 bits=128 verify=NOT) for ; Mon, 15 Jun 2026 12:17:39 +0000 (GMT) Received: by mail-pj1-f72.google.com with SMTP id 98e67ed59e1d1-36bd282e47bso826823a91.2 for ; Mon, 15 Jun 2026 05:17:38 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=oss.qualcomm.com; s=google; t=1781525858; x=1782130658; darn=vger.kernel.org; h=content-transfer-encoding:in-reply-to:content-language:from :references:cc:to:subject:user-agent:mime-version:date:message-id :from:to:cc:subject:date:message-id:reply-to; bh=n+cg/szuB5nopCfExxFx7dLfNUGtVuWLcGK9H/4gV+8=; b=ZOdRS/6mgXveQajAWVlw3DxTlZ8yURO6NhYdEMIOX/NCybT+YpqsWU5dyqhbIhclVO myAUSasKwil6KELx6RQZhMX69c5QovScfwWUA2kbWwPdvlzUVGXuydnK2PO+LS0z0jZj xUWj7nVfsvWfnWw1z//s6DzwxStGMoImirtA5QYcFldrlMt3TndQA4TA5jGjz/+BTXsX T2FEB5OSMoc4ecCWY7SXgTUpGACwnwfwJ4Yc9QTssqweEFElZP8C2d7HJytb4V7uIU3x AtXrZyqOhXBBuToE/nNdRfMJZJPu2MWjJDSh746khfUIEcugwada7KB7qPFIT1tL0Uw7 0YWA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1781525858; x=1782130658; h=content-transfer-encoding:in-reply-to:content-language:from :references:cc:to:subject:user-agent:mime-version:date:message-id :x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to; bh=n+cg/szuB5nopCfExxFx7dLfNUGtVuWLcGK9H/4gV+8=; b=oLrRbU+XQfDcTz4VGfs4wS4Zqtcnwosal8fkNuS32J45twsyqHo8PpJF5brbovYH8f Buc67VmUVj56VpZrDjj9cTkxYPDeG+TqJrSX1xohwjwZ4AO4RFU7ZfwcJ6es5KJ0uMiN 1QYDZkuSBpQxdj1m/yWZFKNfCfD0Aka53j2CIvT3DzIg8UIrCyXLxNL27l9Ji0ncbRtH xXiendYeBvNgozF7kK3pTo+WpveI2Swn2KfFqkMVvncf31NlcSHM2oocZHB5XWd2x9t/ iO0Gk9j/rfKNUBkyrBLZycVGuB2b7UFxobEhx93G/5cLIPf/1hrg2GQ2XUKLbrqapntf Dmaw== X-Forwarded-Encrypted: i=1; AFNElJ9Ezt/8aZQcAmu2wr135vIk/4B3snj3BIYh8J21GOGfZo8+L2b1pptm4mRvOfz/Zn9k+tC9Qy81G7loPrI=@vger.kernel.org X-Gm-Message-State: AOJu0YwB6g2o+6LeJqv45PIQvhZKnolJ7Y0pLzxpCEOkP6Swl0GIvLW5 u/DSPnGcx3QRTSS6x/cbhtqriYf5LumPziyzWL9ap3t8/O3prXwurbtcLjkVHEHjr8rcixKPn8g 7GC2fCXtue2KHpuSxculPDvuZpnKchJEgjfJENM8zBHwXSSHm9VfOVFkdqu5dG1j49pw= X-Gm-Gg: Acq92OEX3x85ipfMjXJiiPQFmVVqcpVh1ctyfq4YTnSMpT+4q/Bm5S9Lj8zJjBqdG5U LqbYBpbG/F9bQVxAJrJzMITmMcjvxGZecNH7yLAyoEnoe5iUK9hSicl1LAvIZ7Bnai5yyXqKdQL P98nDli7jqsTM9q9iSy0TVPY0IxP0R44jiHrNoYw0lx8aqbdVnwTsOEGLGegF1lQRPXtkIU8vJr zoqgeebvODIDUSKcx+PBFHXX6Ws3jfCXEwIYGTojSIS5fLVvu1vp3TKsslwOD0iMSTq+bawDb1k ri2mWm5D+jyvoLad1tsRi6M6jvdr4AiW+SSwYlOgWxilhgEFjfOSEgbXd6BkMnDeyHQzbAGYJei +f//W0HBDGIacy1EfdTo7qNwxP+jjvXSTt1que0NCF0Kf1q8HRxDiQDlUiDMb4S6jsKujji42Y/ Kk9UFMh5Mkizk= X-Received: by 2002:a17:90b:3d47:b0:36b:9323:c726 with SMTP id 98e67ed59e1d1-37a030f8db0mr7226664a91.4.1781525858059; Mon, 15 Jun 2026 05:17:38 -0700 (PDT) X-Received: by 2002:a17:90b:3d47:b0:36b:9323:c726 with SMTP id 98e67ed59e1d1-37a030f8db0mr7226623a91.4.1781525857410; Mon, 15 Jun 2026 05:17:37 -0700 (PDT) Received: from [10.133.33.25] (tpe-colo-wan-fw-bordernet.qualcomm.com. [103.229.16.4]) by smtp.gmail.com with ESMTPSA id 41be03b00d2f7-c866518677dsm8855378a12.19.2026.06.15.05.17.32 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Mon, 15 Jun 2026 05:17:37 -0700 (PDT) Message-ID: <1cf3cf13-5f86-40d8-bdab-dd25a0c35bb7@oss.qualcomm.com> Date: Mon, 15 Jun 2026 20:17:31 +0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH 01/10] sched/core: Skip migration disabled tasks in proxy execution To: Andrea Righi , K Prateek Nayak Cc: John Stultz , Tejun Heo , David Vernet , Changwoo Min , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , Christian Loehle , Koba Ko , Joel Fernandes , sched-ext@lists.linux.dev, linux-kernel@vger.kernel.org References: <20260506174639.535232-1-arighi@nvidia.com> <20260506174639.535232-2-arighi@nvidia.com> <427e64df-2d3c-47a5-925f-ef9a751f1ca3@amd.com> <24ffc508-a806-4be0-9b33-fbe8c02d1742@amd.com> From: "Aiqun(Maria) Yu" Content-Language: en-US In-Reply-To: Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit X-Authority-Analysis: v=2.4 cv=adxRWxot c=1 sm=1 tr=0 ts=6a2fed63 cx=c_pps a=RP+M6JBNLl+fLTcSJhASfg==:117 a=nuhDOHQX5FNHPW3J6Bj6AA==:17 a=IkcTkHD0fZMA:10 a=FelO9ux0wxsA:10 a=s4-Qcg_JpJYA:10 a=VkNPw1HP01LnGYTKEx00:22 a=u7WPNUs3qKkmUXheDGA7:22 a=ZpdpYltYx_vBUK5n70dp:22 a=btlSXbVrFCAxagvMHvEA:9 a=3ZKOabzyN94A:10 a=QEXdDO2ut3YA:10 a=iS9zxrgQBfv6-_F4QbHw:22 X-Proofpoint-GUID: 2tlK2G5arEbEcPNrZjQxbOPTjhwLh9z4 X-Proofpoint-ORIG-GUID: 2tlK2G5arEbEcPNrZjQxbOPTjhwLh9z4 X-Proofpoint-Spam-Info: AW1haW4tMjYwNjE1MDEyOSBTYWx0ZWRfXze86QI4CUs3L Oya60bzYI+lEj0eLIQ+D6cF+XE65B79KAPYY1OQPCHuyMY8qyCtgHonjnQJ+ssp9DC89ONe3ZXL YtuiKoDH+UlVu7ZACvmKLgE8+UQkcho= X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwNjE1MDEyOSBTYWx0ZWRfX2TUyTPpc5sEj VTRIarNWch2Yd8parz1mPikP6E1ZAVQXQscwXIrEGo2Bsb4sVtaYcOJIAA5grhq3tn+Mjyx7M+s pt6IDi6x0rrx6f4yZWMbaVQ65LFtbwkCuYFWfPpXU1TZoq1Ikt3pGZ7q3pmeZQM8voK16LLgDpD OrX29q3bayUhLMx81BoixdNLIGjvAo5ZNCNOX7Et1So+vnyQGXEGceg5LOSPCN+yA9ZK352HUP2 OZRxB9uv/lhGl2Bq1UbSTNvwzQiUVj2M3L8TNyQGVl/IH6bMWrumnNQqXlVRkt2R9yqz03ljIq6 ChWGNPE/CKX0SuTPceZUeZHg575d812YXpMmvkgea18IkLdCqwpHvMa58qc0RIWOxEwcmiPQE9w tmDl8A1InaEJzsxeDStnbIeCytZkM+xjf3YCUy9azbrQQJrun24+uExOxLeD5R4npwPsqfWj1FU KFRN7JXRlPodc9sl2cQ== X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1143,Hydra:6.1.125,FMLib:17.12.100.49 definitions=2026-06-15_03,2026-06-15_01,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 priorityscore=1501 clxscore=1015 impostorscore=0 lowpriorityscore=0 malwarescore=0 suspectscore=0 spamscore=0 phishscore=0 adultscore=0 bulkscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606040000 definitions=main-2606150129 On 5/8/2026 3:40 PM, Andrea Righi wrote: > Hi Prateek, > > On Thu, May 07, 2026 at 09:17:34PM +0530, K Prateek Nayak wrote: >> Hello Andrea, >> >> On 5/7/2026 3:43 PM, Andrea Righi wrote: >>>>>> scx flow should look something like (please correct me if I'm >>>>>> wrong): >>>>>> >>>>>> CPU0: donor CPU1: owner >>>>>> =========== =========== >>>>>> >>>>>> /* Donor is retained on rq*/ >>>>>> put_prev_task_scx() >>>>>> ops.stopping() >>>>>> ops.dispatch() /* May be skipped if SCX_OPS_ENQ_LAST is not set */ >>>>>> do_pick_task_scx() >>>>>> next = donor; >>>>>> find_proxy_task() >>>>>> proxy_migrate_task() >>>>>> ops.dequeue() >>>>>> ======================> /* >>>> >>>> At this point I mean ^ >>>> >>>>>> * Moves to owner CPU (May be outside of affinity list) >>>>>> * ops.enqueue() still happens on CPU0 but I've shown it >>>>>> * here to depict the context has moved to owner's CPU. >>>>>> */ >>>>>> ops.enqueue() >>>>>> scx_bpf_dsq_insert() >>>>>> /* >>>>>> * !!! Cannot dispatch to local CPU; Outside affinity !!! >>>>>> * >>>>>> * We need to allow local dispatch outside affinity iff: >>>>>> * >>>>>> * p->is_blocked && cpu == task_cpu(p) >>>>>> * >>>>>> * Since enqueue_task_scx() hold's the task's rq_lock, the >>>>>> * is_blocked indicator should be stable during a dispatch. >>>>>> */ >>>>>> ops.dispatch() >>>>>> do_pick_task_scx() >>>>>> set_next_task_scx() >>>>>> ops.running(donor) >>>>>> find_proxy_task() >>>>>> next = owner >>>>>> /* >>>>>> * !!! Owner stats running without any notification. !!! >>>>>> * >>>>>> * If owner blocks, dequeue_task_scx() is executed first and >>>>>> * the sched-ext scheduler sees: >>>>>> * >>>>>> * ops.stopping(owner) >>>>>> * >>>>>> * which leads to some asymmetry. >>>>>> * >>>>>> * XXX: Below is how I imagine the flow should continue. >>>>>> */ >>>>>> ops.quiescent(owner) /* Core is taking back control of owner's running */ >>>>>> /* Runs owner */ >>>>>> ops.runnable(owner) /* Core is giving back control to ext layer */ >>>>>> ops.stopping(donor); /* Accounting symmetry for donor */ >>>>> >>>>> I think the order of operations should be the following: >>>>> >>>>> ops.runnable(donor) >>>>> -> ops.enqueue(donor) >>>>> -> donor becomes curr >>>>> -> ops.running(donor) /* set_next_task_scx(donor); !task_is_blocked(donor) */ >>>>> -> donor executes >>>>> -> donor blocks on mutex (proxy: stays on_rq; task_is_blocked(donor) true) >>>>> -> __schedule() >>>>> -> pick_next -> proxy-exec selects owner as next >>>>> -> put_prev_task_scx(donor) >>>>> -> ops.stopping(donor) >>>>> -> dispatch_enqueue(local_dsq) /* blocked donor: ext core parks on local DSQ */ >>>>> -> set_next_task_scx(owner) >>>>> -> ops.running(owner) >>>> >>>> So ext will just switch the context back to owner? But how does this >>>> happen with the changes in your series? >>>> >>>> Based on my understanding, this happens: >>>> >>>> -> pick_next -> sced-ext returns donor as next >>>> /* prev's context is put back */ >>>> -> set_next_task_scx(donor) >>>> -> ops.running(donor) >>>> >>>> /* In core.c */ >>>> >>>> /* next = donor */ >>>> if (next->blocked_on) /* true since we have blocked donor */ >>>> next = find_proxy_task(); /* Returns owner */ >>>> >>>> /* next = owner; */ >>>> /* Starts running owner */ >>>> >>>> How does ext core swap back the owner context here? Am I missing >>>> something? find_proxy_task() doesn't call put_prev_set_next_task() so >>>> I'm at a loss how we get to set_next_task_scx(owner). >>> >>> The sequence should be the following: >> >> Still a bit confused! Hope you can bear with me for just a little >> bit longer :-) > > No, thank you! This is super useful for me! I want to make sure I'm not > missing/misinterpreting anything obvious. :) > >> >>> >>> - pick_next_task(rq, rq->donor, &rf) returns donor (because we parked it on the local DSQ) >> >> So put_prev_set_next_task() happens as a part of pick_next_task(). >> >> When we pick the donor, we have already called set_next_task(donor) >> on it before returning it from pick_next_task(). >> >> "owner" is still not known at this point ... > > That seems correct. > >> >>> - in __schedule() (still holding rq->lock), proxy sees next->blocked_on and does: >>> - next = find_proxy_task(rq, next, &rf); -> returns owner (or triggers migration / retries) >>> - Only after that, __schedule() reaches the point where it performs the switch >>> (put_prev_set_next_task(rq, prev, next) via the pick path). Ao that point, >> >> ... and we don't do put_prev_set_next_task(donor, owner) after >> (or within) find_proxy_task() as far as I'm aware. The "donor" >> remains as the task on which we last called put_prev_task(). > > Also correct. > >> >> If you are referring to the bits in your Patch2, the calls to >> put_prev_task() and set_next_task() is done on the same "donor" >> task. It is purely for the sake of adding a balance callback if >> we had skipped migrating away the prev task due to proxy. >> >> AFAIC, nothing does a set_next_task(owner) after >> pick_next_task() in __schedule() unless I'm grossly mistaken. > > I think you're right. > > Let me try to recap what happens in two different scenarios: > > # donor and owner running on the same CPU > > Owner runs on CPU0, it expires its p->scx.slice, so it's de-scheduled and added > to a DSQ; donor is next, it runs and blocks on a mutex on CPU0, we park the > donor on CPU0's local DSQ, pick_next_task(rq, rq->donor, &rf) on CPU0 returns > next == donor, we see next->blocked_on == true, so we trigger find_proxy_task(), > inside find_proxy_task() we see owner_cpu == task_cpu, find_proxy_task() returns > owner, replacing next, set_next_task_scx(owner) triggers ops_dequeue() + > dispatch_dequeue(), removing the owner from the DSQ, then We've seen an issue which cpu1 do find_proxy_task trying to schedule the owner, and the cpu0 trying to migrate owner to consume remote. And finally cpu1 and cpu0 both do context switch to the same task. Is there any corner case here? > set_next_task_scx(owner) will trigger ops.running(onwer), then > ops.stopping(owner). And in this case we don't trigger ops.stopping(donor) + > ops.running(donor) during the proxy switch. > > # donor and owner running on different CPUs > > Owner runs on CPU0, it expires its p->scx.slice, so it's de-scheduled and added > to a DSQ; donor runs on CPU1, it blocks on a mutex on CPU1, we park donor on > CPU1's local DSQ, pick_next_task(rq, rq->donor, &rf) on CPU1 returns donor as > next, we see next->blocked_on == true, we trigger > find_proxy_task() on CPU1, find_proxy_task() sees owner_cpu != this_cpu, so it > triggers proxy_migrate_task() to migrate donor to CPU0, which triggers We've seen a crash that one cpu(cpu1 here for example) trying to do proxy_migrate_task to another cpu(cpu6 here for example), while cpu6 is trying to pause itself and trying to migrate tasks back to cpu1. And in a live lock. > deactivate_task(donor), unlinking it from CPU1's local DSQ, then > proxy_set_task_cpu(donor, CPU0). But at this point we're not adding donor to > CPU0's local DSQ. I think this is the part that is missing, if we add donor to > CPU0's local DSQ at this point we would effectively fall back to the "same CPU" > scenario and (in theory) everything should work. > > Something like the following (not tested yet - about to). > > Thanks, > -Andrea > > kernel/sched/ext.c | 16 ++++++++++++++++ > 1 file changed, 16 insertions(+) > > diff --git a/kernel/sched/ext.c b/kernel/sched/ext.c > index af9b10cd82c4a..6125c4cbd6d64 100644 > --- a/kernel/sched/ext.c > +++ b/kernel/sched/ext.c > @@ -1915,6 +1915,22 @@ static void do_enqueue_task(struct rq *rq, struct task_struct *p, u64 enq_flags, > > WARN_ON_ONCE(!(p->scx.flags & SCX_TASK_QUEUED)); > > + /* > + * Under proxy execution, mutex-blocked donors can be migrated to a > + * different rq (e.g., towards the mutex owner's CPU). For sched_ext, rq > + * association alone isn't sufficient for the donor to be picked again > + * and drive find_proxy_task(); make it immediately visible on the > + * destination rq by parking it on the built-in local DSQ. > + * > + * This task is a scheduling context token and isn't supposed to run as > + * itself while blocked. > + */ > + if (unlikely(task_is_blocked(p))) { > + clear_direct_dispatch(p); > + dispatch_enqueue(sch, rq, &rq->scx.local_dsq, p, 0); > + return; > + } Prefer online CPUs when selecting a target — there is a race window between a blocked task being enqueued onto an offline CPU's runqueue and getting a chance of __schedule to push it back out. > + > /* internal movements - rq migration / RESTORE */ > if (sticky_cpu == cpu_of(rq)) > goto local_norefill; -- Thx and BRs, Aiqun(Maria) Yu