From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from casper.infradead.org (casper.infradead.org [90.155.50.34]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 69826199391 for ; Mon, 6 Jan 2025 11:57:36 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=90.155.50.34 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1736164658; cv=none; b=PILxgKQmi75F8nf3OH+t2CvCyz5R+CMPQ4jYT1A+h/VyjQQjJcwkIu6CUWeSKOqJgkFaVz8ywrxAzCXsOol7VP5jMlXgop4cPyaoVbKUHPxuE9WdOPwPNhJBkev33yJiP2JVdqgOskUDgyBobxsE/Yt8VbXIMHhEuHAWnzTD2cA= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1736164658; c=relaxed/simple; bh=FxRjsVqCRyLfcz2IxNo+ODXJcOMYQ2a7hygoSgwz+38=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=H6Dys6+AClEh90qFzVWXgtXZPWEHisqLmp3HcDEu5laTr2SPCX/TQkzu1TEyH377ZUYft6UN++JS1hJ6711WMA13V4aKE+3xZHmClOAHHGVGqUn1wVbiLKDUiFkpHXNCF6xv76tpEYevhDgy8B46NzkiUteAuzFFZzMhpdolPPs= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=infradead.org; spf=none smtp.mailfrom=infradead.org; dkim=pass (2048-bit key) header.d=infradead.org header.i=@infradead.org header.b=XYQU5NK5; arc=none smtp.client-ip=90.155.50.34 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=infradead.org Authentication-Results: smtp.subspace.kernel.org; spf=none smtp.mailfrom=infradead.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=infradead.org header.i=@infradead.org header.b="XYQU5NK5" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=infradead.org; s=casper.20170209; h=In-Reply-To:Content-Type:MIME-Version: References:Message-ID:Subject:Cc:To:From:Date:Sender:Reply-To: Content-Transfer-Encoding:Content-ID:Content-Description; bh=8vB6hXHTOkocY6aSx21j1sQsiNepparzwP4zjqTAum4=; b=XYQU5NK57zeovlaz/2udv9A+y0 8qb6JgMfjFsl5WjXAqbU6jEXiT27erTa5sQJE+HyZmW/SDlD79StWvJ0185FbDPpJYHlIdLJofP3G QvK5Ec20z+PuGByZZHegtSkf/LETnTXfr1olib+oN2vBKpNWWsHpqU6I+ARMwkp3neYkib5Agr/Vi 0GlReYRmjmUGNLPIJtg98iO28F8uselbea5uI+8bzP9CmdtrTZxnla/WIrqcIJmtyUCOqBxfzWXAW DY2M4oG+1Lgz5TKzDgiKkTYuFfz1MfeAbi+nXEfULcjgtw0AYssw1x0tFfKh2lXvu6bwPRFtNBcSm iSwWydWQ==; Received: from 77-249-17-89.cable.dynamic.v4.ziggo.nl ([77.249.17.89] helo=noisy.programming.kicks-ass.net) by casper.infradead.org with esmtpsa (Exim 4.98 #2 (Red Hat Linux)) id 1tUljh-000000098VS-31EF; Mon, 06 Jan 2025 11:57:34 +0000 Received: by noisy.programming.kicks-ass.net (Postfix, from userid 1000) id 2644E3005D6; Mon, 6 Jan 2025 12:57:33 +0100 (CET) Date: Mon, 6 Jan 2025 12:57:32 +0100 From: Peter Zijlstra To: Doug Smythies Cc: linux-kernel@vger.kernel.org, vincent.guittot@linaro.org Subject: Re: [REGRESSION] Re: [PATCH 00/24] Complete EEVDF Message-ID: <20250106115732.GE20870@noisy.programming.kicks-ass.net> References: <005f01db5a44$3bb698e0$b323caa0$@telus.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <005f01db5a44$3bb698e0$b323caa0$@telus.net> On Sun, Dec 29, 2024 at 02:51:43PM -0800, Doug Smythies wrote: > Hi Peter, > > I have been having trouble with turbostat reporting processor package power levels that can not possibly be true. > After eliminating the turbostat program itself as the source of the issue I bisected the kernel. > An edited summary (actual log attached): > > 82e9d0456e06 sched/fair: Avoid re-setting virtual deadline on 'migrations' > b10 bad fc1892becd56 sched/eevdf: Fixup PELT vs DELAYED_DEQUEUE > b13 bad 54a58a787791 sched/fair: Implement DELAY_ZERO > skip 152e11f6df29 sched/fair: Implement delayed dequeue > skip e1459a50ba31 sched: Teach dequeue_task() about special task states > skip a1c446611e31 sched,freezer: Mark TASK_FROZEN special > skip 781773e3b680 sched/fair: Implement ENQUEUE_DELAYED > skip f12e148892ed sched/fair: Prepare pick_next_task() for delayed dequeue > skip 2e0199df252a sched/fair: Prepare exit/cleanup paths for delayed_dequeue > b12 good e28b5f8bda01 sched/fair: Assert {set_next,put_prev}_entity() are properly balanced > dfa0a574cbc4 sched/uclamg: Handle delayed dequeue > b11 good abc158c82ae5 sched: Prepare generic code for delayed dequeue > e8901061ca0c sched: Split DEQUEUE_SLEEP from deactivate_task() > > Where "bN" is just my assigned kernel name for each bisection step. > > In the linux-kernel email archives I found a thread that isolated these same commits. > It was from late Novermebr / early December: > > https://lore.kernel.org/all/20240727105030.226163742@infradead.org/T/#m9aeb4d897e029cf7546513bb09499c320457c174 > > An example of the turbostat manifestation of the issue: > > doug@s19:~$ sudo ~/kernel/linux/tools/power/x86/turbostat/turbostat --quiet --Summary --show > Busy%,Bzy_MHz,IRQ,PkgWatt,PkgTmp,TSC_MHz --interval 1 > [sudo] password for doug: > Busy% Bzy_MHz TSC_MHz IRQ PkgTmp PkgWatt > 99.76 4800 4104 12304 73 80.08 > 99.76 4800 4104 12047 73 80.23 > 99.76 4800 879 12157 73 11.40 > 99.76 4800 26667 84214 72 557.23 > 99.76 4800 4104 12036 72 79.39 > > Where TSC_MHz was reported as 879, there was a big gap in time. > Like 4.7 seconds instead of 1. > Where TSC_MHz was reported as 26667, there was not a big gap in time. > > It happens for about 5% of the samples + or - a lot. > It only happens when the workload is almost exactly 100%. > More load, it doesn't occur. > Less load, it doesn't occur. Although, I did get this once: > > Busy% Bzy_MHz TSC_MHz IRQ PkgTmp PkgWatt > 91.46 4800 4104 11348 73 103.98 > 91.46 4800 4104 11353 73 103.89 > 91.50 4800 3903 11339 73 98.16 > 91.43 4800 4271 12001 73 108.52 > 91.45 4800 4148 11481 73 105.13 > 91.46 4800 4104 11341 73 103.96 > 91.46 4800 4104 11348 73 103.99 > > So, it might just be much less probable and less severe. > > It happens over many different types of workload that I have tried. In private email you've communicated it happens due to sched_setaffinity() sometimes taking multiple seconds. I'm trying to reproduce by starting a bash 'while ;: do :; done' spinner for each CPU, but so far am not able to reproduce. What is the easiest 100% load you're seeing this with?