From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-yw1-f174.google.com (mail-yw1-f174.google.com [209.85.128.174]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CD55627E05A for ; Sun, 18 Jan 2026 20:46:25 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.174 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1768769187; cv=none; b=PLGeMosz1KfRjfH4EtGSSf+tXTrAMDy2zyJNaV1+t7KDGi0yFIIrx7VvcEDF26W71jCyBLtbBM7nnEFBL+NevSNMdIAh474GHuOuW/uNxzpzOxclS446taKXWvxkV6llk0c7XQOo39mA/IGu4Q23t6gOwmvYckX9YELvGrpZJTY= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1768769187; c=relaxed/simple; bh=3cRJAK5SBkZVt7pq88jkiQ3NfjGdDbOUcSKw1UJD17A=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=tRcPktRb4Jo+xSscyLx3M9mAxzj4lg34Vm/gFuqhO5tVokXdrtSUA6D8uC6wKQZwAEYnp+zRQhOVVFmTaQYnSQVqzCKYjlSBMDkYxVQa8kpsIKFhq986vmsbKKL2Ov8VjnOl3bgiLSRwv0BjRES+SLlk/KIIlRXcS1rWUQsjNm0= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=A/xppa/u; arc=none smtp.client-ip=209.85.128.174 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="A/xppa/u" Received: by mail-yw1-f174.google.com with SMTP id 00721157ae682-78c6a53187dso34686257b3.2 for ; Sun, 18 Jan 2026 12:46:25 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1768769185; x=1769373985; darn=vger.kernel.org; h=content-transfer-encoding:in-reply-to:from:content-language :references:cc:to:subject:user-agent:mime-version:date:message-id :from:to:cc:subject:date:message-id:reply-to; bh=aQiuSNyy/2EW1fraj6QSwIXG6ZtbRh6MwJCn0UtRIVs=; b=A/xppa/ulIviKA2x6AbPj24FSioL3JmcN0jCBoQ3x/SAlNdfVA14i6yup+bL08nP+p j8eNUPMCCFZkiWSgylCVIwxLpNIkfLfHA0k04yaSgMs6Kjn9zHopz4jHvG5iTVQHr2l0 yOOtLoUISPTYF9gGZsdj/6ekkvs+cg11InUH4XrxRD9H5rZOU8LPOXhPYEN2xFvsAmYF uV/sV8X0y6iWxUhnLeYlH4HGMj3LdoSTwEnt9GHCS2DXDPLMOtXX1yn7hkC7jN0+bOBp fsY2dxq6b9Bf4u5pA+LoJHtsCd8J0wOYZyw9wimYJZtj/vB0Lbr3UkK/68JaOm0caEwz 0VEQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1768769185; x=1769373985; h=content-transfer-encoding:in-reply-to:from:content-language :references:cc:to:subject:user-agent:mime-version:date:message-id :x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to; bh=aQiuSNyy/2EW1fraj6QSwIXG6ZtbRh6MwJCn0UtRIVs=; b=pdHcav7oHmGf/XIQG7tC7/pTT7JyCmF5GWCWZjxdlWVReRTKJNkYkCdgIdXbyxkT4g oXa5cF52o2EWF4OWx4Aa5lfxTUDj5ae69v3WD1MHpmbdatQbxuMWxSCq3HOPdSkaX89Z zqgvA+DBAnZ8KeM/PlOlfNCmEQpfuSuRPyHy7xUrSCkpZ8KsecQL+N184k4hawhx+7k3 IiTpoYVu8U46ZD2qyQFTUWDCduL0MSi2h/NnOzrCc/VUISgco/Ax0cCxEgY0R0rO2E0a Tmm0m9PC2rrHB3p5NSKfIOXV/OIOBprry9E9nHyFfRQsK6aLILgs/sbp1xZlwfVf6Ms1 5NMQ== X-Forwarded-Encrypted: i=1; AJvYcCWyEXkC+jOBEZarLEjzw4Lyvf18/hqSFHrf8i3EfSbX966JACWvpwzsPl0gyvCNge8kMJy+iB8UZXdT5q0=@vger.kernel.org X-Gm-Message-State: AOJu0YxIlIqz4m3fig/fLfpTE6ved7+AdOlkeF/P5wZVK050e/065DH4 DoitPB21xFUlvSR9eYRhpo0yk0uFd3+hlbiUjmiq0T1oUSUHY24kRTx2 X-Gm-Gg: AY/fxX6FGoLYvf1oouws84+B5ME9APRoYj3/7uDxw02ahN++2lSOIihJdIj3ZH9xSLV Mn9brHA3Z4qmgTMyospi0Jq0sHVeV1WWN2YT4qWz+tF3mOaH4uS+5Mbe9qvte4ok4hNN7Qc34H2 xH4ts2nBChXpXffa2iLzPo/LtQ0FbvyX56/WvjIz3GNJ8nzxE3VqtL3VXbinfw1N6CAdv/bcdTh 0SxRr4UZDUs/Q3dq300ozOfwaeyy+Qo9eQ92FkUjewCJ8Lw9THD9e7Re80JdPBfUAwtV60Qd0cE 4kDvG5H3GriwyicfypsmjfJ2ZP7tEhu2GBWiiLfvvJyc/0mRvUGT+azHK8DE4FhhYii+Kbgbxum +I7ovYLQ0nZwJNSiYC9qHFbe0px+Y4Su4ifNdg3X3Kh7PM2p8inxzt2mrQP8Yei4EG+EHjHqzJ2 lYKGI8x6SD+lUQnw== X-Received: by 2002:a05:690c:6706:b0:793:bbb4:eadd with SMTP id 00721157ae682-793c671d0a3mr75179437b3.30.1768769184672; Sun, 18 Jan 2026 12:46:24 -0800 (PST) Received: from [192.168.1.64] ([173.92.131.131]) by smtp.gmail.com with ESMTPSA id 00721157ae682-793c66f326bsm33306147b3.19.2026.01.18.12.46.22 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Sun, 18 Jan 2026 12:46:23 -0800 (PST) Message-ID: <8760001e-0274-454c-a4e4-1f38a9695b88@gmail.com> Date: Sun, 18 Jan 2026 15:46:22 -0500 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH 4/4] sched/fair: Proportional newidle balance To: Peter Zijlstra , Chris Mason , Joseph Salisbury , Adam Li , Hazem Mohamed Abuelfotoh , Josh Don Cc: mingo@redhat.com, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, linux-kernel@vger.kernel.org References: <20251107160645.929564468@infradead.org> <20251107161739.770122091@infradead.org> Content-Language: en-US From: Mario Roy In-Reply-To: <20251107161739.770122091@infradead.org> Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 8bit The patch "Proportional newidle balance" introduced a regression with Linux 6.12.65 and 6.18.5. There is noticeable regression with easyWave testing. [1] The CPU is AMD Threadripper 9960X CPU (24/48). I followed the source to install easyWave [2]. That is fetching the two tar.gz archives. #!/bin/bash # CXXFLAGS="-O3 $CXXFLAGS" ./configure # make -j8 trap 'rm -f *.ssh *.idx *.log *.sshmax *.time' EXIT OMP_NUM_THREADS=48 ./src/easywave \   -grid examples/e2Asean.grd -source examples/BengkuluSept2007.flt \   -time 1200 Before results with CachyOS 6.12.63-2 and 6.18.3-2 kernels. easyWave ver.2013-04-11 Model time = 00:00:00,   elapsed: 0 msec Model time = 00:10:00,   elapsed: 5 msec Model time = 00:20:00,   elapsed: 10 msec Model time = 00:30:00,   elapsed: 19 msec ... Model time = 05:00:00,   elapsed: 2908 msec Model time = 05:10:00,   elapsed: 3079 msec Model time = 05:20:00,   elapsed: 3307 msec Model time = 05:30:00,   elapsed: 3503 msec ... After results with CachyOS 6.12.66-2 and 6.18.6-2 kernels. easyWave ver.2013-04-11 Model time = 00:00:00,   elapsed: 0 msec Model time = 00:10:00,   elapsed: 5 msec Model time = 00:20:00,   elapsed: 10 msec Model time = 00:30:00,   elapsed: 18 msec ... Model time = 05:00:00,   elapsed: 13057 msec  (normal is < 3.0s) Model time = 05:10:00,   elapsed: 13512 msec Model time = 05:20:00,   elapsed: 13833 msec Model time = 05:30:00,   elapsed: 14206 msec ... Reverting the patch "sched/fair: Proportional newidle balance" returns back to prior performance. [1] https://openbenchmarking.org/test/pts/easywave [2] https://openbenchmarking.org/innhold/da7f1cf159033fdfbb925102284aea8a83e8afdc On 11/7/25 11:06 AM, Peter Zijlstra wrote: > Add a randomized algorithm that runs newidle balancing proportional to > its success rate. > > This improves schbench significantly: > > 6.18-rc4: 2.22 Mrps/s > 6.18-rc4+revert: 2.04 Mrps/s > 6.18-rc4+revert+random: 2.18 Mrps/S > > Conversely, per Adam Li this affects SpecJBB slightly, reducing it by 1%: > > 6.17: -6% > 6.17+revert: 0% > 6.17+revert+random: -1% > > Signed-off-by: Peter Zijlstra (Intel) > --- > include/linux/sched/topology.h | 3 ++ > kernel/sched/core.c | 3 ++ > kernel/sched/fair.c | 43 +++++++++++++++++++++++++++++++++++++---- > kernel/sched/features.h | 5 ++++ > kernel/sched/sched.h | 7 ++++++ > kernel/sched/topology.c | 6 +++++ > 6 files changed, 63 insertions(+), 4 deletions(-) > > --- a/include/linux/sched/topology.h > +++ b/include/linux/sched/topology.h > @@ -92,6 +92,9 @@ struct sched_domain { > unsigned int nr_balance_failed; /* initialise to 0 */ > > /* idle_balance() stats */ > + unsigned int newidle_call; > + unsigned int newidle_success; > + unsigned int newidle_ratio; > u64 max_newidle_lb_cost; > unsigned long last_decay_max_lb_cost; > > --- a/kernel/sched/core.c > +++ b/kernel/sched/core.c > @@ -121,6 +121,7 @@ EXPORT_TRACEPOINT_SYMBOL_GPL(sched_updat > EXPORT_TRACEPOINT_SYMBOL_GPL(sched_compute_energy_tp); > > DEFINE_PER_CPU_SHARED_ALIGNED(struct rq, runqueues); > +DEFINE_PER_CPU(struct rnd_state, sched_rnd_state); > > #ifdef CONFIG_SCHED_PROXY_EXEC > DEFINE_STATIC_KEY_TRUE(__sched_proxy_exec); > @@ -8589,6 +8590,8 @@ void __init sched_init_smp(void) > { > sched_init_numa(NUMA_NO_NODE); > > + prandom_init_once(&sched_rnd_state); > + > /* > * There's no userspace yet to cause hotplug operations; hence all the > * CPU masks are stable and all blatant races in the below code cannot > --- a/kernel/sched/fair.c > +++ b/kernel/sched/fair.c > @@ -12146,11 +12146,26 @@ void update_max_interval(void) > max_load_balance_interval = HZ*num_online_cpus()/10; > } > > -static inline bool update_newidle_cost(struct sched_domain *sd, u64 cost) > +static inline void update_newidle_stats(struct sched_domain *sd, unsigned int success) > +{ > + sd->newidle_call++; > + sd->newidle_success += success; > + > + if (sd->newidle_call >= 1024) { > + sd->newidle_ratio = sd->newidle_success; > + sd->newidle_call /= 2; > + sd->newidle_success /= 2; > + } > +} > + > +static inline bool > +update_newidle_cost(struct sched_domain *sd, u64 cost, unsigned int success) > { > unsigned long next_decay = sd->last_decay_max_lb_cost + HZ; > unsigned long now = jiffies; > > + update_newidle_stats(sd, success); > + > if (cost > sd->max_newidle_lb_cost) { > /* > * Track max cost of a domain to make sure to not delay the > @@ -12198,7 +12213,7 @@ static void sched_balance_domains(struct > * Decay the newidle max times here because this is a regular > * visit to all the domains. > */ > - need_decay = update_newidle_cost(sd, 0); > + need_decay = update_newidle_cost(sd, 0, 0); > max_cost += sd->max_newidle_lb_cost; > > /* > @@ -12843,6 +12858,22 @@ static int sched_balance_newidle(struct > break; > > if (sd->flags & SD_BALANCE_NEWIDLE) { > + unsigned int weight = 1; > + > + if (sched_feat(NI_RANDOM)) { > + /* > + * Throw a 1k sided dice; and only run > + * newidle_balance according to the success > + * rate. > + */ > + u32 d1k = sched_rng() % 1024; > + weight = 1 + sd->newidle_ratio; > + if (d1k > weight) { > + update_newidle_stats(sd, 0); > + continue; > + } > + weight = (1024 + weight/2) / weight; > + } > > pulled_task = sched_balance_rq(this_cpu, this_rq, > sd, CPU_NEWLY_IDLE, > @@ -12850,10 +12881,14 @@ static int sched_balance_newidle(struct > > t1 = sched_clock_cpu(this_cpu); > domain_cost = t1 - t0; > - update_newidle_cost(sd, domain_cost); > - > curr_cost += domain_cost; > t0 = t1; > + > + /* > + * Track max cost of a domain to make sure to not delay the > + * next wakeup on the CPU. > + */ > + update_newidle_cost(sd, domain_cost, weight * !!pulled_task); > } > > /* > --- a/kernel/sched/features.h > +++ b/kernel/sched/features.h > @@ -121,3 +121,8 @@ SCHED_FEAT(WA_BIAS, true) > SCHED_FEAT(UTIL_EST, true) > > SCHED_FEAT(LATENCY_WARN, false) > + > +/* > + * Do newidle balancing proportional to its success rate using randomization. > + */ > +SCHED_FEAT(NI_RANDOM, true) > --- a/kernel/sched/sched.h > +++ b/kernel/sched/sched.h > @@ -5,6 +5,7 @@ > #ifndef _KERNEL_SCHED_SCHED_H > #define _KERNEL_SCHED_SCHED_H > > +#include > #include > #include > #include > @@ -1348,6 +1349,12 @@ static inline bool is_migration_disabled > } > > DECLARE_PER_CPU_SHARED_ALIGNED(struct rq, runqueues); > +DECLARE_PER_CPU(struct rnd_state, sched_rnd_state); > + > +static inline u32 sched_rng(void) > +{ > + return prandom_u32_state(this_cpu_ptr(&sched_rnd_state)); > +} > > #define cpu_rq(cpu) (&per_cpu(runqueues, (cpu))) > #define this_rq() this_cpu_ptr(&runqueues) > --- a/kernel/sched/topology.c > +++ b/kernel/sched/topology.c > @@ -1662,6 +1662,12 @@ sd_init(struct sched_domain_topology_lev > > .last_balance = jiffies, > .balance_interval = sd_weight, > + > + /* 50% success rate */ > + .newidle_call = 512, > + .newidle_success = 256, > + .newidle_ratio = 512, > + > .max_newidle_lb_cost = 0, > .last_decay_max_lb_cost = jiffies, > .child = child, > >