From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wr1-f52.google.com (mail-wr1-f52.google.com [209.85.221.52]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5CE8EEEB3 for ; Wed, 4 Mar 2026 03:03:10 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.221.52 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1772593392; cv=none; b=BMPRmAvbGL9iHQNen6ogZMzyJ5YYJng/kwyYyWvGGkkWfMGEevmkSI+nTzQaNurS97rQ289fzgOUwlNHjpol0RDw2Uv1bMbWxXGg/lZoN4l8RPt4Lax8kXO1bKstEIeQ0PjhzYjGQZD9XdF7dGBQwNfYtWYL2sYxAjeW23caR/I= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1772593392; c=relaxed/simple; bh=sLKtiEu4eRsjhOwCLIRF2NxkHfX9J/Og8vAkKA8YDk8=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=LMMC5+he1QscujzJoh6oBpTqXE3AJo1IywU7FpQ8KpydAueT3iYc1Z4Rk/IvCjTMUtTvexEb8XlIGOdofOewPpNVkGIwEU1GPjgpFX7BloT0bsjaj+d3tLrfqohPuyPZjpgPQHP04LNDrHn0AmJDNbI6Q23pAn1DStD2CbbLUJQ= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=layalina.io; spf=pass smtp.mailfrom=layalina.io; dkim=pass (2048-bit key) header.d=layalina-io.20230601.gappssmtp.com header.i=@layalina-io.20230601.gappssmtp.com header.b=SVsk2feO; arc=none smtp.client-ip=209.85.221.52 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=layalina.io Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=layalina.io Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=layalina-io.20230601.gappssmtp.com header.i=@layalina-io.20230601.gappssmtp.com header.b="SVsk2feO" Received: by mail-wr1-f52.google.com with SMTP id ffacd0b85a97d-439b790af67so1681319f8f.0 for ; Tue, 03 Mar 2026 19:03:10 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=layalina-io.20230601.gappssmtp.com; s=20230601; t=1772593389; x=1773198189; darn=vger.kernel.org; h=in-reply-to:content-transfer-encoding:content-disposition :mime-version:references:message-id:subject:cc:to:from:date:from:to :cc:subject:date:message-id:reply-to; bh=8pdt0vZxPQxIunJ8XshRd9zJ2sBYhyNq3Qv+7jGXvPc=; b=SVsk2feO5Cvj0mp8MCCPaMdCAvPr+QjnJG1wV6CT1QCvdtwkLhTPCQN7NgIgvE4L83 S7G1bRu7FkKt0av+tZkBVgXCJh/7B7VwaezploOE+mAbvwjVqjMgg/xPLtrHSwaRlfYb u3chRRyeSpEqxdDAMiTfDqVBWIsKPwWGONYQM+w54ZJSjQr5ivN+oy/7OrhbSmKgO1ar SaASPud+56KFYXpifW/orDmJCwnbaSblHelHeVEYuBEBoAy+hBcoB9tgnYoDrzrT4v6w FKUJKXi9+E5yrb6YG26mWflF61hVwyzK4W7yY+02VhtMRmfZntbnMQx4YmjmSQVW2Fqv 1ocA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1772593389; x=1773198189; h=in-reply-to:content-transfer-encoding:content-disposition :mime-version:references:message-id:subject:cc:to:from:date:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=8pdt0vZxPQxIunJ8XshRd9zJ2sBYhyNq3Qv+7jGXvPc=; b=ArML0Oy5+dCdFYs57Ebb/yuzaUnjA6VpHYlFI6O97KbiEolbedVw1I8+p2LADF/q0Z oSKLuDoBa1Iqj+t1l2iH/4N3iqLU3vnOH/hGInnEH3Cjy2B9VMiPuB68kAgGBrW5/H6v +/+DV93TQyq7mhBtF1IWqrtpnMkLL+iogrcLM36se/wykAFAMf4YnCLFFbX77CCB5W4i wA1RU5h49WkuwnHBhfyqcU92mu4CvndPlIfNGUqCrN/ayc6693IjhNt3VIMfgBl7j0rP Md/I+c+wNkW+fTOUqCeyz5tyncp7j2l9+sDeYO6gvE8Kk2dWAq0kw0F1LIEA3OmLq67M 3SbQ== X-Forwarded-Encrypted: i=1; AJvYcCUNu/uFtHTpVIIng/rhqjqbqCW/gECoKj6AjbkEWJMD+V5GaVFnIwZf1VFzcpz3j6O9uJAjUMTinZqvBwo=@vger.kernel.org X-Gm-Message-State: AOJu0YwZiAFWa41IDFNMH6TcjFy7Vsccf15xtAEa+4JmXlUPLeQm6njv 3KZO/OLGDTbJk2agN2N+wQuv+bAmtXKL6bZdymKFF3+OC2g7w55CMP/BXH8U2eElPKw= X-Gm-Gg: ATEYQzw/lQZa/azggJHmXBrY9usyvQo26IiB5A5ojgF/yY+t46JVQ6+fMXcigDQ0hSB L6433m53rSUDotzGJg7ir/tr8XyD/mgYF97DizIrSEa3ZtSQOmWIkfKLo73+SJt5IvnZEzTVLXv 7NMFvl6R9kvH17ySDKmh5y5GZkhE6UeEg17DBiH7402JpPDzWCj+5Xf4Cct2F8JoM03Hkjq13OT XCQLn1Cu6NweW1HT20iR105N4VmoyN6RWWBRtuvuoBURQa0B6hikXTgPA7zY9U7ua/YFj4YYhjk aFslr57x9jG59lni1SLtPUzwUN/FQ2zCn0aPb/eod6Sll9QtphzA7/tqtR9kWfmqTNcf2FO34Cr jxGMt9UclzMadXl5Kj0SxhXF0OnnG09GTH2Vn2iAh/RklF3su4lUp3VBcmGDq542vAds5tJEj6I zT1EroF9Cj7kqfvSRqI43dgnhzKpsncRf4qxD2EUuy5DJT149VCGULUU9rxf766nhlJwF5 X-Received: by 2002:a05:6000:40dc:b0:437:7719:ca82 with SMTP id ffacd0b85a97d-439c8a3d1c8mr737528f8f.3.1772593388170; Tue, 03 Mar 2026 19:03:08 -0800 (PST) Received: from airbuntu (host86-169-41-76.range86-169.btcentralplus.com. [86.169.41.76]) by smtp.gmail.com with ESMTPSA id ffacd0b85a97d-439ae0e7abasm25622800f8f.23.2026.03.03.19.03.07 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 03 Mar 2026 19:03:07 -0800 (PST) Date: Wed, 4 Mar 2026 03:03:06 +0000 From: Qais Yousef To: "Rafael J. Wysocki" Cc: Christian Loehle , Thomas Gleixner , LKML , Peter Zijlstra , Frederic Weisbecker Subject: Re: [patch 2/2] sched/idle: Make default_idle_call() NOHZ aware Message-ID: <20260304030306.uk5c63xw4oqvjffb@airbuntu> References: <20260301191959.406218221@kernel.org> <20260301192915.171574741@kernel.org> <7d3e7a4b-05c1-4bef-9450-29c49528cc9c@arm.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Disposition: inline Content-Transfer-Encoding: 8bit In-Reply-To: On 03/02/26 22:25, Rafael J. Wysocki wrote: > On Mon, Mar 2, 2026 at 12:04 PM Christian Loehle > wrote: > > > > On 3/1/26 19:30, Thomas Gleixner wrote: > > > Guests fall back to default_idle_call() as there is no cpuidle driver > > > available to them by default. That causes a problem in fully loaded > > > scenarios where CPUs go briefly idle for a couple of microseconds: > > > > > > tick_nohz_idle_stop_tick() is invoked unconditionally which means unless > > > there is timer pending in the next tick, the tick is stopped and a couple > > > of microseconds later when the idle condition goes away restarted. That > > > requires to program the clockevent device twice which implies a VM exit for > > > each reprogramming. > > > > > > It was suggested to remove the tick_nohz_idle_stop_tick() invocation from > > > the default idle code, but would be counterproductive. It would not allow > > > the host to go into deeper idle states when the guest CPU is fully idle as > > > it has to maintain the periodic tick. > > > > > > Cure this by implementing a trivial moving average filter which keeps track > > > of the recent idle recidency time and only stop the tick when the average > > > is larger than a tick. > > > > > > Signed-off-by: Thomas Gleixner > > > --- > > > kernel/sched/idle.c | 65 +++++++++++++++++++++++++++++++++++++++++++++------- > > > 1 file changed, 57 insertions(+), 8 deletions(-) > > > > > > --- a/kernel/sched/idle.c > > > +++ b/kernel/sched/idle.c > > > @@ -105,12 +105,7 @@ static inline void cond_tick_broadcast_e > > > static inline void cond_tick_broadcast_exit(void) { } > > > #endif /* !CONFIG_GENERIC_CLOCKEVENTS_BROADCAST_IDLE */ > > > > > > -/** > > > - * default_idle_call - Default CPU idle routine. > > > - * > > > - * To use when the cpuidle framework cannot be used. > > > - */ > > > -static void __cpuidle default_idle_call(void) > > > +static void __cpuidle __default_idle_call(void) > > > { > > > instrumentation_begin(); > > > if (!current_clr_polling_and_test()) { > > > @@ -130,6 +125,61 @@ static void __cpuidle default_idle_call( > > > instrumentation_end(); > > > } > > > > > > +#ifdef CONFIG_NO_HZ_COMMON > > > + > > > +/* Limit to 4 entries so it fits in a cache line */ > > > +#define IDLE_DUR_ENTRIES 4 > > > +#define IDLE_DUR_MASK (IDLE_DUR_ENTRIES - 1) > > > + > > > +struct idle_nohz_data { > > > + u64 duration[IDLE_DUR_ENTRIES]; > > > + u64 entry_time; > > > + u64 sum; > > > + unsigned int idx; > > > +}; > > > + > > > +static DEFINE_PER_CPU_ALIGNED(struct idle_nohz_data, nohz_data); > > > + > > > +/** > > > + * default_idle_call - Default CPU idle routine. > > > + * > > > + * To use when the cpuidle framework cannot be used. > > > + */ > > > +static void default_idle_call(void) > > > +{ > > > + struct idle_nohz_data *nd = this_cpu_ptr(&nohz_data); > > > + unsigned int idx = nd->idx; > > > + s64 delta; > > > + > > > + /* > > > + * If the CPU spends more than a tick on average in idle, try to stop > > > + * the tick. > > > + */ > > > + if (nd->sum > TICK_NSEC * IDLE_DUR_ENTRIES) > > > + tick_nohz_idle_stop_tick(); > > > + > > > + __default_idle_call(); > > > + > > > + /* > > > + * Build a moving average of the time spent in idle to prevent stopping > > > + * the tick on a loaded system which only goes idle briefly. > > > + */ > > > + delta = max(sched_clock() - nd->entry_time, 0); > > > + nd->sum += delta - nd->duration[idx]; > > > + nd->duration[idx] = delta; > > > + nd->idx = (idx + 1) & IDLE_DUR_MASK; > > > +} > > > + > > > +static void default_idle_enter(void) > > > +{ > > > + this_cpu_write(nohz_data.entry_time, sched_clock()); > > > +} > > > + > > > +#else /* CONFIG_NO_HZ_COMMON */ > > > +static inline void default_idle_call(void { __default_idle_call(); } > > > +static inline void default_idle_enter(void) { } > > > +#endif /* !CONFIG_NO_HZ_COMMON */ > > > + > > > static int call_cpuidle_s2idle(struct cpuidle_driver *drv, > > > struct cpuidle_device *dev, > > > u64 max_latency_ns) > > > @@ -186,8 +236,6 @@ static void cpuidle_idle_call(void) > > > } > > > > > > if (cpuidle_not_available(drv, dev)) { > > > - tick_nohz_idle_stop_tick(); > > > - > > > default_idle_call(); > > > goto exit_idle; > > > } > > > @@ -276,6 +324,7 @@ static void do_idle(void) > > > > > > __current_set_polling(); > > > tick_nohz_idle_enter(); > > > + default_idle_enter(); > > > > > > while (!need_resched()) { > > > > > > > > > > How does this work? We don't stop the tick until the average idle time is larger, > > but if we don't stop the tick how is that possible? > > > > Why don't we just require one or two consecutive tick wakeups before stopping? > > Exactly my thought and I think one should be sufficient. I concur. From our experience with TEO util threshold these averages can backfire. I think one tick is sufficient delay to not be obviously broken. But IMO the setup is broken too. No cpuidle driver and nohz is enabled but performance is important is not a good combination. Since this has proven to have both power and performance impact, ensuring there's a sensible cpuidle driver is the right thing to do IMHO.