From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from SN4PR2101CU001.outbound.protection.outlook.com (mail-southcentralusazon11012002.outbound.protection.outlook.com [40.93.195.2]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4A3C83A16B8 for ; Thu, 30 Apr 2026 07:26:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=40.93.195.2 ARC-Seal:i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1777533971; cv=fail; b=FUTv+L4hQrz2uK27cLWYK6Py5f+CIaa5VtoKzt7w5LsRm4A0TIKt2TQP4s7TKAc1aCcm+5N7D8RIPgPU3lW7fJtMd5bcaj0ubMT2nscrnRijIDo+QfScmFMRRuQc0EuDh6vsB0Vo6tN1OWN1j04IetTyndhNmjLRZ787am+A1Y4= ARC-Message-Signature:i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1777533971; c=relaxed/simple; bh=kwJIGEryVVGdnA7nG7zMr8YS0PKiQvpar9TfnYsRkpY=; h=Message-ID:Date:MIME-Version:Subject:To:CC:References:From: In-Reply-To:Content-Type; b=uWjs4jUw7p9bIzRIuLW9JccBUTnVgxuJwYsXh9f4QS6qxFqbHqv9+ANM849H6+JhQhnYaq2H19vYWLcFyrud3gXmIk1x38/n4MzEwv3LlRB0lD45sQdv3G/MDctjn/Ph2B61bdSREKvvDPwCxqwMKY/eAak+86ULtyaNHzlv2uI= ARC-Authentication-Results:i=2; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amd.com; spf=fail smtp.mailfrom=amd.com; dkim=pass (1024-bit key) header.d=amd.com header.i=@amd.com header.b=Uxkw/sse; arc=fail smtp.client-ip=40.93.195.2 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amd.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=amd.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=amd.com header.i=@amd.com header.b="Uxkw/sse" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=G75O5emS1Az4N3vm+rCNne5suxCYvzfC4cB3Tsu/g5VfpdjuuXOouqC9t5Lb43Zt/qRjiFHzgX2amBC2Mj5mcQwx+SrrvZaCyAupMGCXWzGoXNT0S5hi6oee/diSbEo9HPffL92WT+/+cZfrGlsc3FEmxRHd21NwQXvspYfJMsU6hyPmGlFTMLcRK6+k1frLOUSZR+SS4I/3b3/KNbBlmBKpLKpa7qPYGrej1ccJk+vLwDqVHbaFaVbZ4GSTJ2Jtf6j/IpRMMrO4rYCRmX/hC8mlP/qVkler9GKg1Qh9BlE9AOw6AwePSuhVL1MIq2io7+65im2/vYclvEDogujSGA== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=6VuBAj0o2KzWxfhR/utELt3pnAL6SjizFsfW6KHQfOo=; b=QgvBuJsYC5Z/5k8jDwpx4v/PQz6sgt3VdvqQzxO5LtEVkAx50lnFJ+Xh39YuzxffL9HkIGQ5HRNOymrju7MqcNXvGTlr/X20ljc6tefMtxceBOvJHifNugoMoZ55ighoKhAXi7Ib8AFKd/DRTD9+9uDDrOPxOaKiRe6LIq8rH8KqnUbTnrnoZUw9IZ/fOKQ3DqTu1LjZNEft1TIAt02oa0x1NKNDFhg22rT0XDHB9/QY8jPHBTXx+DnE+EhqOmga+9R9OOJG+6DatehVxuHn8s7wupKQx5JvnJegR8vzIhD+Kg5xVg+0nEAXXsQlPWuqO1dd+G5cDZsXPn1iuhZCHQ== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 165.204.84.17) smtp.rcpttodomain=google.com smtp.mailfrom=amd.com; dmarc=pass (p=quarantine sp=quarantine pct=100) action=none header.from=amd.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=amd.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=6VuBAj0o2KzWxfhR/utELt3pnAL6SjizFsfW6KHQfOo=; b=Uxkw/sseUveEergKpmzOomf9M3/c4xLoqvW+VjtaN9TKoXXYeuzcEtPDo25Tg+dEZtrW/D/+lutL9XdNCtTc6QGAGcbSR5SMnr/P7RUc9a6Uu5w7JYz7pxY9E9aX5965wGGgJPkPH3pFhO7rd73xhILfng2KXRZJWznWjO00/Ro= Received: from PH7PR03CA0030.namprd03.prod.outlook.com (2603:10b6:510:339::35) by LV3PR12MB9185.namprd12.prod.outlook.com (2603:10b6:408:199::13) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.20.9870.20; Thu, 30 Apr 2026 07:26:02 +0000 Received: from SA2PEPF000015CB.namprd03.prod.outlook.com (2603:10b6:510:339:cafe::e3) by PH7PR03CA0030.outlook.office365.com (2603:10b6:510:339::35) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.20.9870.21 via Frontend Transport; Thu, 30 Apr 2026 07:26:01 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 165.204.84.17) smtp.mailfrom=amd.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=amd.com; Received-SPF: Pass (protection.outlook.com: domain of amd.com designates 165.204.84.17 as permitted sender) receiver=protection.outlook.com; client-ip=165.204.84.17; helo=satlexmb08.amd.com; pr=C Received: from satlexmb08.amd.com (165.204.84.17) by SA2PEPF000015CB.mail.protection.outlook.com (10.167.241.201) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.20.9846.18 via Frontend Transport; Thu, 30 Apr 2026 07:26:01 +0000 Received: from satlexmb10.amd.com (10.181.42.219) by satlexmb08.amd.com (10.181.42.217) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.17; Thu, 30 Apr 2026 02:25:57 -0500 Received: from satlexmb08.amd.com (10.181.42.217) by satlexmb10.amd.com (10.181.42.219) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.17; Thu, 30 Apr 2026 02:25:57 -0500 Received: from [172.31.184.125] (10.180.168.240) by satlexmb08.amd.com (10.181.42.217) with Microsoft SMTP Server id 15.2.2562.17 via Frontend Transport; Thu, 30 Apr 2026 02:25:48 -0500 Message-ID: <9479df04-7351-4a69-88f4-17f6a90e13a5@amd.com> Date: Thu, 30 Apr 2026 12:55:47 +0530 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH 1/2] sched: proxy-exec: Close race causing workqueue work being delayed To: John Stultz CC: Peter Zijlstra , LKML , Vineeth Pillai , "Sonam Sanju" , Sean Christopherson , "Kunwu Chan" , Tejun Heo , Joel Fernandes , Qais Yousef , Ingo Molnar , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Valentin Schneider , Steven Rostedt , Will Deacon , Waiman Long , Boqun Feng , "Paul E. McKenney" , Metin Kaya , Xuewen Yan , Thomas Gleixner , "Daniel Lezcano" , Suleiman Souhlal , kuyo chang , hupu , References: <20260427183848.698551-1-jstultz@google.com> <20260427183848.698551-2-jstultz@google.com> <20260428094353.GB1026330@noisy.programming.kicks-ass.net> <20260428111833.GL3102924@noisy.programming.kicks-ass.net> <9cf9b433-cba5-4a8e-8dbf-6410239cffb6@amd.com> Content-Language: en-US From: K Prateek Nayak In-Reply-To: Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: 8bit X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: SA2PEPF000015CB:EE_|LV3PR12MB9185:EE_ X-MS-Office365-Filtering-Correlation-Id: c5ae40fc-e8a4-434e-46c4-08dea689bb43 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|82310400026|1800799024|36860700016|7416014|376014|22082099003|13003099007|56012099003|18002099003; X-Microsoft-Antispam-Message-Info: 0ie/FyaAY+b5oB4d6cCPhto3Bckqg1TxmrDD3jRqGhwfTdzEorGsy+UqCQwkT0eWjIkmhT2JNzlr0JwYWpWnasAMCcRiiFpjRWvk6ESyZudm+cH7UOPt3+dgJBO+1EyXwPJVhQQYdxeDMNMFlgPHkpSubnZyD8pPf6xpWUI0uKj/I3HfuN5yZhuKMYgUDdqZ3SziJpmmOx45T2lVUQUds5xT5X/kt0sBILTTNK5J5IC4vTLhq/sAuVh0/AovgyTLV0IdMvoKcRT2dTnCjAnqS81MakDNf96a4nWi6Iv3QVdggU5YQ0hM3IXFyT4Aq7ZUIJMMFTEO4Nn+fLKVOuEMpRBwfKLewJP1rObM/adlUCRnujZNhSjDi9Z2NucV+ZeAfxE26jxxv4gSczGbdyMwgDqT2qtRsKeFVAOWmVH+oLeRdKHIIR3j6TlFVggs1FHljinl18RHw4Ao+wEVwA2Fc4+4qCewjg9VTe0FpseXe4YJ3C1NHPVDRfNPxpa+0pGaKs/uvX3Bm/sUcFwFf1s3W3/NfMOjB/RXNn4y4HzhFueCYH68tXcQCXYprwPXwQW8mIUhxjIKe6ZoL/JWvKG0qAAz5Q2gBCo1DW9K2ieM2MuC/UoxDQlQI/qa1tMh7WOS7cRD0S5zmedTls/0uNxaZBVdzxZmo4El5S2g1wY3W2V1o4YVkklGXDPxt6zOKVxctQPS1h1pcBSaqYnJFo7GPdkYin7O5Dho/+iGkMCquDhtQWMTDVAVvZFdoHvmX9HFLPEdVpCNhW0dsJNqbPtO7g== X-Forefront-Antispam-Report: CIP:165.204.84.17;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:satlexmb08.amd.com;PTR:InfoDomainNonexistent;CAT:NONE;SFS:(13230040)(82310400026)(1800799024)(36860700016)(7416014)(376014)(22082099003)(13003099007)(56012099003)(18002099003);DIR:OUT;SFP:1101; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: 9/zUurlyi7f7HobIXnTuD+9byro+zAaqWM3KF4/ie4fRCT5OJ7P5NkZJaCGUHVlenrpWjBJSFqU4icb8idXhJJiecA+XaY0fXMLDtcRo3DQPUdDB9baL+bKJpxGyN9oFwJhFCdG/caQA9/MgaGB61kNDbh6TQRm+R3bxrVObLdYq3THakDDBtjNXiUQqiHdeJPPQme8DNjOeFGqd/XLGbzETs0A4eK19YPUfUKuBc0zge+bDk4eLGzbGcnf7pS8fQt1GZyleZ+9uv69kMgp3ShtQLLjKqEPjYA5ACUA6080jRvo25ibgZGzvIyrjQUAy0DhXFNRTbpdT8cgCRE4VR+OCGYdDY3p5QAHBoMFvefYYAZoGpzzckQvsNJBq0QR+M42Upbql+6CEH2shaYTxDrY1PJdQl3FzWDAetAL8eHUQ3ui65ywuiKkEGSVsKjEM X-OriginatorOrg: amd.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 30 Apr 2026 07:26:01.4914 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: c5ae40fc-e8a4-434e-46c4-08dea689bb43 X-MS-Exchange-CrossTenant-Id: 3dd8961f-e488-4e60-8e11-a82d994e183d X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=3dd8961f-e488-4e60-8e11-a82d994e183d;Ip=[165.204.84.17];Helo=[satlexmb08.amd.com] X-MS-Exchange-CrossTenant-AuthSource: SA2PEPF000015CB.namprd03.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: LV3PR12MB9185 Hello John, On 4/30/2026 11:14 AM, John Stultz wrote: > On Wed, Apr 29, 2026 at 2:00 AM K Prateek Nayak wrote: >> On 4/29/2026 7:57 AM, John Stultz wrote: >>>> diff --git a/include/linux/sched.h b/include/linux/sched.h >>>> index 8ec3b6d7d718b..6ea74aecc5fbd 100644 >>>> --- a/include/linux/sched.h >>>> +++ b/include/linux/sched.h >>>> @@ -6535,8 +6536,13 @@ static bool try_to_block_task(struct rq *rq, struct task_struct *p, >>>> * blocked on a mutex, and we want to keep it on the runqueue >>>> * to be selectable for proxy-execution. >>>> */ >>>> - if (!should_block) >>>> + if (!should_block) { >>>> + guard(raw_spinlock)(&p->blocked_lock); >>>> + /* Stable against race */ >>>> + if (task_is_blocked(p)) >>>> + WRITE_ONCE(p->se.sched_proxy, 1); >>>> return false; >>>> + } >>> >>> So if we double check and find the task isn't blocked anymore, we >>> probably shouldn't return early here, no? >>> >>> Let me take a stab at the bit flag approach and see how it goes. >> >> In case you want to peek at my homework ;-) >> > > Well, now it will just seem like I'm cheating! :) But I can't prove I didn't copy your idea first ;-) >> @@ -7043,8 +7053,16 @@ static void __sched notrace __schedule(int sched_mode) >> if (sched_proxy_exec()) { >> struct task_struct *prev_donor = rq->donor; >> >> + /* >> + * A wakeup raced with block_task(); >> + * Clear blocked_on before running the task >> + * again. >> + */ >> + if (unlikely(!prev_state && prev->blocked_on)) >> + clear_task_blocked_on(prev, NULL); >> + > > So similar to your previous change, I was hitting this same case, and > adding in your change here was bothering me a bit that it felt like > clearing things here was papering over an unexpected state change that > we didn't have before. > > Part of the issue is the blocked_on state machine is already a little complex: > NULL -> ptr -> PROXY_WAKING ->NULL > (with some limited shortcuting from ptr->NULL when we are dealing with current) > > So adding this BO_FLAG_PROXY/"latch" bit to the mix complicates things: > NULL -> ptr:unlatched -> ptr:latched -> PROXY_WAKING -> NULL > > With extra lines for: > ptr:unlatched -> NULL > ptr:latched -> NULL > ptr:unlatched -> PROXY_WAKING > > I kept on seeing these unexpected transitions where before we latched > the blocked_on state, we'd get a wakeup on another cpu from the owner > unlocking the lock, so we'd go ptr:unlatched -> PROXY_WAKING (along > with setting TASK_RUNNING). Then because we weren't latched, we might > be running midway through the lock logic. If we go into schedule() > we'll just come back out eventually (as we're TASK_RUNNING). Then when > then try again in the lock loop to set blocked_on to ptr:unlatched > again, we hit warnings because we are setting it to a ptr when it is > PROXY_WAKING and not NULL. Yes, it was precisely that which caused the splat in my case too for the initial suggestion on Peter's diff. I'll try to add a little bit more context in those tricky comments from next time onwards. Sorry about that! > > I'm sure this was obvious to you when you wrote the snippit above, but > I've been running a little slow, so let's not talk about how long it > took for me to actaully understand this. :P > > And yes, your snippit above avoids this, but it feels a little > incidental, cleaning up after the fact. So I think it might be better > to simplify the state machine a little to prevent it. > > My current thought is to cut out the ptr:unlatched -> PROXY_WAKING > transition. And instead in __set_task_blocked_on_waking(), instead of > checking blocked_on and doing an early return if it is null, we could > instead check if the latch bit was set, and if so, set blocked_on to > NULL. This bascially treats ptr:unlatched the same as NULL in most > cases for the state machine. That is a fantastic idea! Since the sched bits have not blocked it yet, there is no need to go through the return migration path. Clearing it when not latched should do the trick. > > Now, earlier you made some nice explanations earlier about how some of > the blocked_on checks done w/ the rq_lock and not the blocked_lock > were safe because of the limited situations where we clear blocked_on. > And this at first seems to put that assumption risk, but I think it > still holds if we always consider ptr:unlatched equivalent to NULL. So > once ptr:latched is set, the rule still holds. Yes, that is indeed true. Transitioning the task to latched state, and unlatching it, will both be done with rq_lock + blocked_on_lock so it should be okay to peek at the latched state when holding the task's rq lock in ttwu_runnable(), and in schedule(). > > That gives us: > NULL -> ptr:unlatched -> ptr:latched -> PROXY_WAKING -> NULL > > Where NULL and ptr:unlatched are functionally equivalent except for > the ability to transition to ptr:latched. > > With: > ptr:latched -> NULL // only done on current > ptr:unlatched -> NULL // only done on current or when trying to set waking > > So far my testing is doing ok with this. Let me know if you see any > holes though. This looks good to me! Cannot spot any holes with that. However, to completely get rid of that hunk in __schedule() we probably need to order setting the blocked_on wrt task's __state based on Peter's analysis in https://lore.kernel.org/lkml/20260403125424.GA2872@noisy.programming.kicks-ass.net/ As Peter suggested, I think we need: __set_task_blocked_on(current, lock); /* * Pairs with smp_rmb() after ttwu_state_match() in try_to_wake_up(). * Ensures ttwu always see the correct blocked_on state for wakeups. */ smp_wmb(); set_current_state(state); in __mutex_lock_common() for it to be completely safe with the proxy_needs_return() bits and that should be good enough to avoid that race but I'll let Peter comment if he thinks there is still a race window open. > > I'll try to cleanup my current changes (surely copying some of your > excellent work :), and send this out tomorrow (I'm shot for today) for > more concrete review. I'll be looking forward to it :-) -- Thanks and Regards, Prateek