From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from casper.infradead.org (casper.infradead.org [90.155.50.34]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D6BE5480DF3 for ; Wed, 29 Jul 2026 12:39:48 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=90.155.50.34 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785328796; cv=none; b=DOOnfqio8o4TKufETcV1lsRyuxHyj+ymnhDoUhSCisZOK7XCNsV/ekRr/gHPqVShmyrmHDee2NI8bgWSpILJHSCQK39XWgUcfxS4x33yycgUQ6EhGU86naLVCJUr0c6Sxdmc+MqkIU3ZEKnJSGHZdn1bfGA/E+C3K0p3fN2Zl34= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785328796; c=relaxed/simple; bh=t1B2mZywQ522aQyhpHF/7J5wRfZvu4ebphF3S3wXBHg=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=Btg5/uSams/gbCEwQF3U8sGSlp5hy3sLsU0hsB6ZAPEWbh9j1x+6HNy9m/etnRroFaI6aK3SYThuc/QFKTfqDxqqiznRACTcu8JWfw2JvMA/cPdgavoqoLjk1GS52VlxPTSizeUH6HQlK1xAWo57pvxmgLslE8rHJD3TJBUynRE= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=infradead.org; spf=pass smtp.mailfrom=infradead.org; dkim=pass (2048-bit key) header.d=infradead.org header.i=@infradead.org header.b=T09MddzN; arc=none smtp.client-ip=90.155.50.34 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=infradead.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=infradead.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=infradead.org header.i=@infradead.org header.b="T09MddzN" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=infradead.org; s=casper.20170209; h=In-Reply-To:Content-Type:MIME-Version: References:Message-ID:Subject:Cc:To:From:Date:Sender:Reply-To: Content-Transfer-Encoding:Content-ID:Content-Description; bh=T4zSjxEUq8pkO1UrtPsuN84VGah9W08WLtq4Nav1t20=; b=T09MddzNMgrsCEcVWIoTVw+gzN XCvmyArwlin/Mq0P16Tp22dacReqQjPL0taP4HCMX7KlxrmIllaK5rbGy6pN4VEC+Ut7jtShXgeVP 19SlVZjKCoNmJOjHaYgKUdwptG8va6GPpEdDHS1i0jc+Hnt+nCaSXjVpjBnZgcdS1Z9UHCXh44xSv tF6yLFxYAn1p/Vd9cy56SMekLBjAMgo5LevdDuneIhKI1fG2yRavg8mDG7So1synD6KuvF7PM9hks t1nK06KNSWvyZH/JK00HHzPzaR0lrd2Wh4AR5SNuV/pH2Tndgp1E8NVvzcrAfJM+YDCCJBgWP9Ydm r73N5Nig==; Received: from 77-249-17-252.cable.dynamic.v4.ziggo.nl ([77.249.17.252] helo=noisy.programming.kicks-ass.net) by casper.infradead.org with esmtpsa (Exim 4.99.1 #2 (Red Hat Linux)) id 1wp3Z3-00000003s3Y-3J4D; Wed, 29 Jul 2026 12:39:20 +0000 Received: by noisy.programming.kicks-ass.net (Postfix, from userid 1000) id C616B300E8B; Wed, 29 Jul 2026 14:39:12 +0200 (CEST) Date: Wed, 29 Jul 2026 14:39:12 +0200 From: Peter Zijlstra To: Jing Wu Cc: Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , x86@kernel.org, "H. Peter Anvin" , "Paul E. McKenney" , "Rafael J. Wysocki" , linux-kernel@vger.kernel.org, Qiliang Yuan , Jian Zhang , Frederic Weisbecker Subject: Re: [PATCH] x86/aperfmperf: Refresh stale sample via IPI for busy NOHZ_FULL CPUs Message-ID: <20260729123912.GZ751831@noisy.programming.kicks-ass.net> References: <20260728-bug-isolatecpu-cpufreq-v1-1-e95d34db8bcd@gmail.com> <20260728134434.GU751831@noisy.programming.kicks-ass.net> <20260728144224.GH651302@noisy.programming.kicks-ass.net> <20260729082225.1675233-1-realwujing@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260729082225.1675233-1-realwujing@gmail.com> On Wed, Jul 29, 2026 at 04:22:24PM +0800, Jing Wu wrote: > On Tue, Jul 28, 2026 at 04:42:24PM +0200, Peter Zijlstra wrote: > > Aside from the fact that sending IPIs to NOHZ_FULL is just plain wrong, > > this whole thing makes no sense. > > > > When the CPU is isolated, nothing should care about the ratio anyway. > > Just set the thing to '1' (1024) when the CPU enters NOHZ_FULL mode and > > ensure it isn't ever modified. > > Fair, understood. > > For context on why I went looking in the first place: stressing an > isolated, nohz_full CPU shows both /proc/cpuinfo's "cpu MHz" and > /sys/devices/system/cpu/cpuN/cpufreq/scaling_cur_freq stuck at the > P-state floor (e.g. 800MHz) for as long as the CPU stays busy and > isolated, while turbostat confirms the hardware is actually running > at full turbo (e.g. 3.2GHz) the whole time. Both interfaces go > through arch_freq_get_on_cpu(), so whatever affects one affects both. > > Getting the exact value would need an on-demand rdmsr on the target > CPU - which is what turbostat itself does via /dev/cpu/N/msr's > rdmsr_safe_regs_on_cpu(), i.e. the same smp_call_function_single() > IPI, just triggered manually by a human running a diagnostic tool > instead of sitting behind a commonly-polled sysfs file. > > I looked for a way around that: PCU mailbox telemetry can expose a > per-core P-state on some Xeon uncores without touching the target > CPU, and HFI publishes a shared table too, but that's a per-core > performance/efficiency class, not an achieved clock, and PCU access > is uncore/generation-specific, not a general mechanism. So as far as > I can tell there's no way to get the exact value for an isolated CPU > without an IPI of some form. Oh, you care about the silly sysfs files? I though this was about the scheduler use of aperf/mperf ratio. Both are driven from the same source, but the scheduler use makes no sense when isolated/NOHZ_FULL. And I would argue that keeping the CPU isolated is more important than having the silly number 'accurate'. Something like so perhaps? diff --git a/arch/x86/kernel/cpu/aperfmperf.c b/arch/x86/kernel/cpu/aperfmperf.c index 7ffc78d5ebf2..cedb40e6b5e6 100644 --- a/arch/x86/kernel/cpu/aperfmperf.c +++ b/arch/x86/kernel/cpu/aperfmperf.c @@ -438,7 +438,8 @@ static void scale_freq_tick(u64 acnt, u64 mcnt) { u64 freq_scale, freq_ratio; - if (!arch_scale_freq_invariant()) + if (!arch_scale_freq_invariant() || + !housekeeping_cpu(smp_processor_id(), HK_TYPE_TICK)) return; if (check_shl_overflow(acnt, 2*SCHED_CAPACITY_SHIFT, &acnt)) @@ -510,6 +511,9 @@ int arch_freq_get_on_cpu(int cpu) unsigned long last; u64 acnt, mcnt; + if (!housekeeping_cpu(cpu, HK_TYPE_TICK)) + return -EOPNOTSUPP; + if (!cpu_feature_enabled(X86_FEATURE_APERFMPERF)) goto fallback;