From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from MW6PR02CU001.outbound.protection.outlook.com (mail-westus2azon11012047.outbound.protection.outlook.com [52.101.48.47]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id F328E337BBD for ; Thu, 2 Apr 2026 05:28:29 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.48.47 ARC-Seal:i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1775107711; cv=fail; b=iS0FUNTcHq4BahHtEpegqtyJCQQTWep2p1dRUgcXsLh1rk5CTxTPWBBUKvxKkRAcBn1jUATkutcL4vODCALzK1hfqrbdn9vw/ok3oJdirjYDa0/wR+taAjAxoH4qajAgpWUCFBkl0oAEPiOw8Q7Bdcyx9BO03Rg3eF9gGbqE1cc= ARC-Message-Signature:i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1775107711; c=relaxed/simple; bh=2BHaUQ6zqcG1pXZoiQDLizlgZ3VrjPPgy2nMTK+FV9Y=; h=Message-ID:Date:MIME-Version:Subject:From:To:CC:References: In-Reply-To:Content-Type; b=SYM8+E4d5YxuZCNrKpp8vMI04QLxtvNBkK/xRIY9pOI98RkNhp4+A1qnN6wjFSPMmpJDOuC2CFtsX8/4I4KAQNYljOUP9rIV6N9pgx7kf+yjbZ0KRkCxuONi0IVUzPVELCVVA6hfwYIpH+stZWSvoy95Puxyn3XPItN4DHcQkfc= ARC-Authentication-Results:i=2; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amd.com; spf=fail smtp.mailfrom=amd.com; dkim=pass (1024-bit key) header.d=amd.com header.i=@amd.com header.b=B06dN6ks; arc=fail smtp.client-ip=52.101.48.47 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amd.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=amd.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=amd.com header.i=@amd.com header.b="B06dN6ks" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=RF5JXukMYvr6KAPa2s2c7G/FYdncCdVQ87nvDqxU5AIqOlKrIno7LZXYWiVtlu/iqs7Q4zn5D0zkH/V28dRsvqlg31cMErVm3zDwI4ZI0j4Rm3GyX0on7SS6FJ2EP+o1jdUNb4e7VApzjyLnQ12NGHcIutQNAZs24yGsJSyMbudD1dqDZnMtam3Skaqukka3JJzOG/MbUe9w5HgPCTjFLHfHeGiU2QHiVzUQiFHeDMQJxrpLMfRC7RKUCScjr7z4DLZWrCzdXnbGoqENZOiJWJd6tn36NiXovEv2k9KccyAlWYImMTpeyIS7viMlY4//7T047Dfsujku6rEEn4CQRw== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=vENIgBmedmjlAseTtPV2KirZxhDYViFODH5ACbVDKrU=; b=xoNAb6wMDP8qVWMUuCH3yvWzjLtrN/KLYKTlSKi1uFtSeKmg1nk0VOo4uHspWD2LIq94WMoRIyT60AHnj8lNi++IoJGgcKprRivoh+FjAuDgtnYrLOfTHeLNlzD0SW+9Cjly3/yBuGQAE9nY2Tjdw6moboc5iiE1ADvnuwmaHLXunlk+1kvIg6+ySYYJm42VkLyqUovuD9hl5hYZ34sa0eHczDG3VLC69tC+32jgVheLV3UkqjalC7uEmznv1C0O6VPR0KiM1ZuMHxALOZe+moG45sdfz/pwMMs1DiPJIx/mS5Jr6QfFjhOfQZ/q1cKTbc0MssOkclqqZ/wmGB/oiQ== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 165.204.84.17) smtp.rcpttodomain=infradead.org smtp.mailfrom=amd.com; dmarc=pass (p=quarantine sp=quarantine pct=100) action=none header.from=amd.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=amd.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=vENIgBmedmjlAseTtPV2KirZxhDYViFODH5ACbVDKrU=; b=B06dN6ks4wx7APAHfDYNqX0PaIOKtZN6o+IAc/Ss94nHWOaYvCzwPFMfgeKXJPdTZxJtb8LC8lwds0VwQySEpgmXK7KW5s7EoSrJOfV29graFT9ega/0t5JbkIQDXCZD5A1iSpsVvLDS5CPX+mh16oVQ0AEuu22oexqyfJ07y/o= Received: from BN0PR04CA0084.namprd04.prod.outlook.com (2603:10b6:408:ea::29) by BN3PR12MB9595.namprd12.prod.outlook.com (2603:10b6:408:2cb::13) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.20.9769.18; Thu, 2 Apr 2026 05:28:25 +0000 Received: from BN1PEPF0000468C.namprd05.prod.outlook.com (2603:10b6:408:ea:cafe::2a) by BN0PR04CA0084.outlook.office365.com (2603:10b6:408:ea::29) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.20.9745.31 via Frontend Transport; Thu, 2 Apr 2026 05:28:24 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 165.204.84.17) smtp.mailfrom=amd.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=amd.com; Received-SPF: Pass (protection.outlook.com: domain of amd.com designates 165.204.84.17 as permitted sender) receiver=protection.outlook.com; client-ip=165.204.84.17; helo=satlexmb08.amd.com; pr=C Received: from satlexmb08.amd.com (165.204.84.17) by BN1PEPF0000468C.mail.protection.outlook.com (10.167.243.137) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.20.9769.17 via Frontend Transport; Thu, 2 Apr 2026 05:28:24 +0000 Received: from SATLEXMB04.amd.com (10.181.40.145) by satlexmb08.amd.com (10.181.42.217) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256) id 15.2.2562.17; Thu, 2 Apr 2026 00:28:24 -0500 Received: from satlexmb07.amd.com (10.181.42.216) by SATLEXMB04.amd.com (10.181.40.145) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256) id 15.1.2507.39; Thu, 2 Apr 2026 00:28:22 -0500 Received: from [10.136.42.52] (10.180.168.240) by satlexmb07.amd.com (10.181.42.216) with Microsoft SMTP Server id 15.2.2562.17 via Frontend Transport; Thu, 2 Apr 2026 00:28:19 -0500 Message-ID: <99fa12f9-71d3-4766-8742-a3adc9ce4271@amd.com> Date: Thu, 2 Apr 2026 10:58:18 +0530 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v2 5/7] sched/fair: Increase weight bits for avg_vruntime From: K Prateek Nayak To: Peter Zijlstra , Vincent Guittot CC: , , , , , , , , , , , References: <20260219075840.162631716@infradead.org> <20260219080624.942813440@infradead.org> <20260223115100.GI2995752@noisy.programming.kicks-ass.net> <0d3680c3-3e17-47b8-8fdb-0cc1f97ffce0@amd.com> Content-Language: en-US In-Reply-To: <0d3680c3-3e17-47b8-8fdb-0cc1f97ffce0@amd.com> Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: 7bit Received-SPF: None (SATLEXMB04.amd.com: kprateek.nayak@amd.com does not designate permitted sender hosts) X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: BN1PEPF0000468C:EE_|BN3PR12MB9595:EE_ X-MS-Office365-Filtering-Correlation-Id: 2592f71b-8caf-448d-afdc-08de9078a93e X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|36860700016|376014|7416014|82310400026|30052699003|1800799024|18002099003|56012099003|22082099003; X-Microsoft-Antispam-Message-Info: e0izzSMFE246tiJmcrK47pP1hZlMNMI0+tL3GKLDrKZnpz2MnnBgNmU4W8ZGBEXmFkdfKZ0fKe6W9jkQF4C1l6CP+nhgHHaaJVTjDHsXYe798PAFfCpV4Nhx2hZ2CsJ6+zk5OKDFVUDj/tuZt60UbdOptZucX9xjhAlnTgiL+CAeEk6P1EuM3X2P6bHLpcAWfxQgf4qv1DHm8sAfmtHxFBAhaGCk5qJGmwBNgutKRaRgH62c6wD3Sn9jFIZeQnx0lbqbhq8S2nxeMHcGgJt3FD2q6efhgrAgcLuaLOq7XwNYEL0nhihWNC6rQX2Mk0VQSAQ79ImsPFZXd4B4BBtcf8yqfhRfGbDph1oEN2rtpFMVJpV3bR1TH2/jW2FArpFG6oIyeIPTcGiOfarE6xfqeTuCQlvHcqMI9cv3BmmG1NK+CwbYV1Rb+19ELNhfNYBFaSj5dSe3fZZ6HLgVBRNLmXsIav5EMH4MX8mObXpdB8WnaKFtmywaeiUsuU398lnB8BaSWT9w1m3EZroPKaVLvLHJGpffUUIy8Ksucsj/et/bRPzAH3jJ49YAVOIyFxenICwgxJyroZx3Je1OGe0H3qISAVS6FPD6a5CwdiepKu/37XKs8Yxno8nI80y74nWSuaRQDgF1cffZJ3XZbhmymO452ZD246XA+3OPB79vMEDSsdJnWPxYYgGU7w2Jh8t9sManp19Sr7DDLNjSP8h/G6f7qHLzXVVueyS/dx+1SoQrFBM4nLA4zzLxF6pd9m+2VvAu5a0g5VFEWwVBtTwZRw== X-Forefront-Antispam-Report: CIP:165.204.84.17;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:satlexmb08.amd.com;PTR:InfoDomainNonexistent;CAT:NONE;SFS:(13230040)(36860700016)(376014)(7416014)(82310400026)(30052699003)(1800799024)(18002099003)(56012099003)(22082099003);DIR:OUT;SFP:1101; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: R1KyCE3vy1X87PZifZj4RQMAcI5LZ+ia1a1GdDwz+we07ziiz1flclT6zo4GIrnjc37MKI6CVweLAD6nEk2SrOC6yTvdI1czu9MquYXiAAc4g3bJB0xfZP7jCQUjhHjZprHLW2COUbZ3f02lj6r4rUNkQckP+o4O1pPh01mJErgm+++Qq3U4bEWFGm4Yax+3Oke0yNh2D29FqyceAhKVMerVERhg+zVFnFO4BjBgGZrKoqQ+R09QVBjSFR00Rp7PZT0qZcF9B+lSVIJOTcdsPZk5v84nInHV/kB4vl8JKrB3uIibivKruc9x5GOWDXF8md8kfbuxQVWSFgNvs/ogba4rod6LP8H+K4QcR7ANWCrhBUiIiLTasa4S6Ib1Ps7Lo2pNCIpWI2PqR73MW6CW0VyVCtyXOBCdc5TBuROEWrUyJlj/dsXIQ8svh6HSVum1 X-OriginatorOrg: amd.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 02 Apr 2026 05:28:24.2718 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: 2592f71b-8caf-448d-afdc-08de9078a93e X-MS-Exchange-CrossTenant-Id: 3dd8961f-e488-4e60-8e11-a82d994e183d X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=3dd8961f-e488-4e60-8e11-a82d994e183d;Ip=[165.204.84.17];Helo=[satlexmb08.amd.com] X-MS-Exchange-CrossTenant-AuthSource: BN1PEPF0000468C.namprd05.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: BN3PR12MB9595 On 3/30/2026 1:25 PM, K Prateek Nayak wrote: > ------------[ cut here ]------------ > (w_vruntime >> 63) != (w_vruntime >> 62) > WARNING: kernel/sched/fair.c:692 at __enqueue_entity+0x382/0x3a0, CPU#5: stress-ng/5062 Back to this: I still see this with latest set of changes on queue:sched/urgent but it doesn't go kaboom. Nonetheless, it suggests we are closing in on the s64 limitations of "sum_w_vruntime" which isn't very comforting. Here is one scenario where it was triggered when running: stress-ng --yield=32 -t 10000000s& while true; do perf bench sched messaging -p -t -l 100000 -g 16; done on a 256CPUs machine after about an hour into the run: __enqeue_entity: entity_key(-141245081754) weight(90891264) overflow_mul(5608800059305154560) vlag(57498) delayed?(0) cfs_rq: zero_vruntime(3809707759657809) sum_w_vruntime(0) sum_weight(0) nr_queued(1) cfs_rq->curr: entity_key(0) vruntime(3809707759657809) deadline(3809723966988476) weight(37) The above comes from __enqueue_entity() after a place_entity(). Breaking this down: vlag_initial = 57498 vlag = (57498 * (37 + 90891264)) / 37 = 141,245,081,754 vruntime = 3809707759657809 - 141245081754 = 3,809,566,514,576,055 entity_key(se, cfs_rq) = -141,245,081,754 Now, multiplying the entity_key with its own weight results to 5,608,800,059,305,154,560 (same as what overflow_mul() suggests) but in Python, without overflow, this would be: -1,2837,944,014,404,397,056 Now, the fact that it doesn't crash suggests to me the later avg_vruntime() calculation would restore normality and the sum_w_vruntime turns to -57498 (vlag_initial) * 90891264 (weight) = -5,226,065,897,472 (assuming curr's vruntime is still the same) which only requires 43 bits. I also added the following at the bottom of dequeue_entity(): WARN_ON_ONCE(!cfs_rq->nr_queued && cfs_rq->sum_w_vruntime) which was never triggered when the cfs_rq goes idle so it isn't like we didn't account sum_w_vruntime properly. There was just a momentary overflow so we are fine but will it always be that way? One way to avoid the warning entirely would be to pull the zero_vruntime close to avg_vruntime is we are enqueuing a very heavy entity. The correct way to do this would be to compute the actual avg_vruntime() and move the zero_vruntime to that point (but that requires at least one multiply + divide + update_zero_vruntime()). One seemingly cheap way by which I've been able to avoid the warning is with: diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 226509231e67..bc708bb8b5d0 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -5329,6 +5329,7 @@ static void place_entity(struct cfs_rq *cfs_rq, struct sched_entity *se, int flags) { u64 vslice, vruntime = avg_vruntime(cfs_rq); + bool update_zero = false; s64 lag = 0; if (!se->custom_slice) @@ -5406,6 +5407,17 @@ place_entity(struct cfs_rq *cfs_rq, struct sched_entity *se, int flags) load += avg_vruntime_weight(cfs_rq, curr->load.weight); lag *= load + avg_vruntime_weight(cfs_rq, se->load.weight); + /* + * If the entity_key() * sum_weight of all the enqueued entities + * is more than the sum_w_vruntime, move the zero_vruntime + * point to the vruntime of the entity which prevents using + * more bits than necessary for sum_w_vruntime until the + * next avg_vruntime(). + * + * XXX: Cheap enough check? + */ + if (abs(lag) > abs(cfs_rq->sum_w_vruntime)) + update_zero = true; if (WARN_ON_ONCE(!load)) load = 1; lag = div64_long(lag, load); @@ -5413,6 +5425,9 @@ place_entity(struct cfs_rq *cfs_rq, struct sched_entity *se, int flags) se->vruntime = vruntime - lag; + if (update_zero) + update_zero_vruntime(cfs_rq, -lag); + if (sched_feat(PLACE_REL_DEADLINE) && se->rel_deadline) { se->deadline += se->vruntime; se->rel_deadline = 0; --- But I'm sure it'll make people nervous since we basically move the zero_vruntime to se->vruntime. It isn't too bad if: abs(sum_w_vuntime - (lag * load)) < abs(lag * se->load.weight) but we already know that the latter overflows so is there any other cheaper indicator that we can use to detect the necessity to adjust the avg_vruntime beforehand at place_entity()? -- Thanks and Regards, Prateek