From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pf1-f171.google.com (mail-pf1-f171.google.com [209.85.210.171]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4147D2EDD40 for ; Tue, 10 Feb 2026 15:42:00 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.171 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1770738121; cv=none; b=MbMtbZDFrGAtowAW0BeBAPPcMeoHHSNd7Cp7F6g36GPE6xzDPcwdemkWgR0rMIeNmrzXZb7Yv4Fmr5m0uVJfYvc+eYdfoO878w1d8VAQ9Ud3TzbFP8pe+SkcRMA705UVeOyOsDe73U6w+PhxxTRmZ32sdcrt9DRehYevqWwl+Nw= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1770738121; c=relaxed/simple; bh=m0GSy5+i2h23pR6PMwPh4WWI2KmjKuMzV/lDEw6C6bc=; h=From:To:Cc:References:In-Reply-To:Subject:Date:Message-ID: MIME-Version:Content-Type; b=B6Bzk2J+jjmBgOGoG6hSlJ4G6oFFzXKVmzMZfXMaxmjDgHJmuVCVdBkw/zseWICMZlQwLnwHLBOzrItDIQEjsTNu/6aNHAtE0/0D+iZ+buBt2WPLt2Jva+EwSev+9y7nrGR3AoUf4L/KrJK3VN3Wrdl8/JBPMYlGqxtVYWOHR9w= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=telus.net; spf=pass smtp.mailfrom=telus.net; dkim=pass (2048-bit key) header.d=telus.net header.i=@telus.net header.b=bZobrof4; arc=none smtp.client-ip=209.85.210.171 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=telus.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=telus.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=telus.net header.i=@telus.net header.b="bZobrof4" Received: by mail-pf1-f171.google.com with SMTP id d2e1a72fcca58-81f4f4d4822so484295b3a.3 for ; Tue, 10 Feb 2026 07:42:00 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=telus.net; s=google; t=1770738119; x=1771342919; darn=vger.kernel.org; h=thread-index:content-language:content-transfer-encoding :mime-version:message-id:date:subject:in-reply-to:references:cc:to :from:from:to:cc:subject:date:message-id:reply-to; bh=7EizRQEyV/ZbNOA7+ZwGNd0OIC54xuTnpy3ovkEvfRI=; b=bZobrof4mhHI8QBTTLybWT3GoGySeHqEm5Dkp0ub38YOGWQ5dz68I6s3zWR+r1I7+e gVUIoF4skcEuI9Db0mOPTIiU1bUUiuHaF21qC9JU9b0VE3xUHLyJapLzXRQlzuNK0iIW F65/HbYnfFUu3mzmYf+rVbXqFXW21iS7fM7sC30WEwnDDIAYklEb+619+80afddHLzIj zDXDKS06a+oMHaf7DZlCXiXDtBonE2N2RnwZ7NoQomYaDCOmMjtv1mBLHrz0c6BdG7JV bGsOPAflQo5BNs4zWWqSTvIc5DrHxmVPD3PgIWS8PC+pMbS7WGdCpo6GcumTd/l5HjCn Fedw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1770738119; x=1771342919; h=thread-index:content-language:content-transfer-encoding :mime-version:message-id:date:subject:in-reply-to:references:cc:to :from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to; bh=7EizRQEyV/ZbNOA7+ZwGNd0OIC54xuTnpy3ovkEvfRI=; b=A7WTE68nnCQnC8OevQ0IDkCbqeWJLaFp0qfBPdPPFE/2NyK2lG9b9DbKWIitUxyLwx z2B909m/gsjzPU72r/DCGZYk+PP4PAYjBa82reGSGNO0bNGpBp/6mPP3MEkS6r3Z1/gs DHOagnplM6vumyVE1+HA9pTSerCA84XARWYppJMZdZzZmKI3qbR23JigpFrH7Zd8W9F/ PitZIfH7DLIK3mi89qwYqodKy36C1D/YUXT1PVql+8Vez2SdO5GCKG3MnGcBME0BzIEX Tlif0nRn8slsg3pl/ynMYo5VaB0eF4OdDgZ/B39VSQeZfN/Z+jaUjB73tydynyBwwEna RayQ== X-Forwarded-Encrypted: i=1; AJvYcCUB08lRITjjq5PJLgjcgiDWEBRm0R0Am5Qh1qb+ziWlKuZ0g6iPAQeKA27KeQpLCj/UMo1TFh0erDsaWIQ=@vger.kernel.org X-Gm-Message-State: AOJu0YwS1d0mpn7dDOcr9J7+ymmldlQh6VL8VjIAunf1n/MTSbMzb37Z huw5x/DXIx9mP5MqNy02TSLPBogzaRSeqfjwEEdk24tH5zsPsUbcAGH0nnBX3+BABps= X-Gm-Gg: AZuq6aKpCRtxr2vzDjVF8+R+Dtrw+Om/MkOsJynKXQfiR63D8n/8SYH5XLgqAq29uG9 wGKUO+mZNTpQHu+Ll9JQJoci/AiaJ/Xc8uJp5UdUcZCM2WCWoLNwqfZwe+5ytW8iC1deN4RynZO zJR0i4lL1JjFs7AUnIZlXunSCFgQtJaKtH1OGCLZWAUkZ6zD8lgZUYZAqw/ukL5wnjEGeUajFYH Ly/MQvGHbUmYdRxRRgE5zbBw7MIej6btMIpPA6VI0QbMN77h80qyiZ0ZPcRw1EezozaOX6L+kdU QgvkamHnuqf8VDfzTA/McBmhs7IKG7ueopLCKmXRT4ayMQ1eKafqFzhvgNci5lMA4lJIDdTiBB/ DAKmidPBIF9wgkzlk+4B6l2FmEadVD3TaURt3Wf9tYbyofCMRejvpe6v16RlhJrra852spxegpS VlvQmWeIbjmgG/6lBNbDwRQgusmuAeyPUA0a7B6KeX2Syi3sANBCpreI0M0ssTiHWcPl8= X-Received: by 2002:a05:6a21:3990:b0:371:53a7:a48a with SMTP id adf61e73a8af0-393acf8fa3fmr14246684637.1.1770738119601; Tue, 10 Feb 2026 07:41:59 -0800 (PST) Received: from DougS18 (s66-183-142-209.bc.hsia.telus.net. [66.183.142.209]) by smtp.gmail.com with ESMTPSA id 41be03b00d2f7-c6dcb5249a0sm12111472a12.11.2026.02.10.07.41.58 (version=TLS1_2 cipher=ECDHE-ECDSA-AES128-GCM-SHA256 bits=128/128); Tue, 10 Feb 2026 07:41:59 -0800 (PST) From: "Doug Smythies" To: "'Peter Zijlstra'" , "'K Prateek Nayak'" Cc: , , , , , , , , , , , , "Doug Smythies" References: <20260130093439.803225718@infradead.org> <4c48fe59-8ff3-41fb-83cb-869409f6fbc6@amd.com> <20260203111134.GL1282955@noisy.programming.kicks-ass.net> <38ef3462-4c4e-4f40-8d63-84dd71cbd043@amd.com> <20260209154718.GW1282955@noisy.programming.kicks-ass.net> In-Reply-To: <20260209154718.GW1282955@noisy.programming.kicks-ass.net> Subject: RE: [PATCH 0/4] sched: Various reweight_entity() fixes Date: Tue, 10 Feb 2026 07:41:58 -0800 Message-ID: <001901dc9aa3$cbad47f0$6307d7d0$@telus.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="us-ascii" Content-Transfer-Encoding: 7bit X-Mailer: Microsoft Outlook 16.0 Content-Language: en-ca Thread-Index: AQK5IwR5loYZhW90rhLC3Nnzvp6wgAGWA92PAlBHLF4B5k73QgEcsVRAAlcnOhazePX1wA== On 2026.02.09.07:47 Peter Zijlstra wrote: > On Wed, Feb 04, 2026 at 03:45:58PM +0530, K Prateek Nayak wrote: > >> # Overflow on enqueue >> >> <...>-102371 [255] ... : __enqueue_entity: Overflowed cfs_rq: >> <...>-102371 [255] ... : dump_h_overflow_cfs_rq: cfs_rq: depth(0) weight(90894772) nr_queued(2) sum_w_vruntime(0) sum_weight(0) zero_vruntime(701164930256050) sum_shift(0) avg_vruntime(701809615900788) >> <...>-102371 [255] ... : dump_h_overflow_entity: se: weight(3508) vruntime(701809615900788) slice(2800000) deadline(701810568648095) curr?(1) task?(1) <-------- cfs_rq->curr >> <...>-102371 [255] ... : __enqueue_entity: Overflowed se: >> <...>-102371 [255] ... : dump_h_overflow_entity: se: weight(90891264) vruntime(701808975077099) slice(2800000) deadline(701808975109401) curr?(0) task?(0) <-------- new se > > So I spend a whole time trying to reproduce the splat, but alas. > > That said, I did spot something 'funny' in the above, note that > zero_vruntime and avg_vruntime/curr->vruntime are significantly apart. > That is not something that should happen. zero_vruntime is supposed to > closely track avg_vruntime. > > That lead me to hypothesise that there is a problem tracking > zero_vruntime when there is but a single runnable task, and sure > enough, I could reproduce that, albeit not at such a scale as to lead to > such problems (probably too much noise on my machine). > > I ended up with the below; and I've already pushed out a fresh > queue/sched/core. Could you please test again? I tested this "V2". The CPU migration times test results are not good. We expect the sample time to not deviate from the nominal 1 second by more than 10 milliseconds for this test. The test ran for about 13 hours and 41 minutes (49,243 samples). Histogram of times: kernel: 6.19.0-rc8-pz-v2 gov: powersave HWP: enabled 1.000, 29206 1.001, 19598 1.002, 19 1.003, 15 1.004, 32 1.005, 25 1.006, 3 1.007, 3 1.008, 5 1.009, 5 1.010, 13 1.011, 14 1.012, 13 1.013, 6 1.014, 10 1.015, 16 1.016, 54 1.017, 116 1.018, 57 1.019, 7 1.020, 2 1.021, 1 1.023, 2 1.024, 4 1.025, 7 1.026, 1 1.027, 1 1.028, 1 1.029, 1 1.030, 2 1.037, 1 Total: 49240 : Total >= 10 mSec: 329 ( 0.67 percent) For reference previous test results are copied and pasted below. Step 1: Confirm where we left off a year ago: The exact same kernel from a year ago, that we ended up happy with, was used. doug@s19:~/tmp/peterz/6.19/turbo$ cat 613.his Kernel: 6.13.0-stock gov: powersave HWP: enabled 1.000000, 23195 1.001000, 10897 1.002000, 49 1.003000, 23 1.004000, 21 1.005000, 9 Total: 34194 : Total >= 10 mSec: 0 ( 0.00 percent) So, over 9 hours and never a nominal sample time exceeded by over 5 milliseconds. Very good. Step 2: Take a baseline sample before this patch set: Mainline kernel 6.19-rc1 was used: doug@s19:~/tmp/peterz/6.19/turbo$ cat rc1.his Kernel: 6.19.0-rc1-stock gov: powersave HWP: enabled 1.000000, 19509 1.001000, 10430 1.002000, 32 1.003000, 19 1.004000, 24 1.005000, 13 1.006000, 9 1.007000, 4 1.008000, 3 1.009000, 4 1.010000, 6 1.011000, 2 1.012000, 1 1.013000, 4 1.014000, 10 1.015000, 10 1.016000, 7 1.017000, 10 1.018000, 20 1.019000, 12 1.020000, 5 1.021000, 3 1.022000, 1 1.023000, 2 1.024000, 2 <<< Clamped. Actually 26 and 25 milliseconds Total: 30142 : Total >= 10 mSec: 95 ( 0.32 percent) What!!! Over 8 hours. It seems something has regressed over the last year. Our threshold of 10 milliseconds was rather arbitrary. Step 3: This patch set [V1] and from Peter's git tree: doug@s19:~/tmp/peterz/6.19/turbo$ cat 02.his kernel: 6.19.0-rc1-pz gov: powersave HWP: enabled 1.000000, 19139 1.001000, 9532 1.002000, 19 1.003000, 17 1.004000, 8 1.005000, 3 1.006000, 2 1.009000, 1 Total: 28721 : Total >= 10 mSec: 0 ( 0.00 percent) Just about 8 hours. Never a time >= our arbitrary threshold of 10 milliseconds. So, good. My test computer also hung under the heavy heavy load test, albeit at a higher load than before. There was no log information that I could find after the re-boot. References: https://lore.kernel.org/lkml/000d01dc939e$0fc99fe0$2f5cdfa0$@telus.net/ https://lore.kernel.org/lkml/005f01db5a44$3bb698e0$b323caa0$@telus.net/ https://lore.kernel.org/lkml/004a01dc952b$471c94a0$d555bde0$@telus.net/