From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from casper.infradead.org (casper.infradead.org [90.155.50.34]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7BA7F78F2B for ; Mon, 6 Jan 2025 17:14:05 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=90.155.50.34 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1736183647; cv=none; b=ZwLp7mPUrNjyoq2Ms8b5Khvdc5E7zBgT6zRhgdyVAFkUzgQhysrGVHQpTeyQmwm9xqn/pz4jKp6tTMDQ2fK2qG6jNbM+YQsba4W2lT7y39BA/W37pEDjWgks4BH99aHwxzzYFuVNdQj97s1ggIS5/x2ylHmEuj6LKDDkaXM8ejo= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1736183647; c=relaxed/simple; bh=XT59zmXXo8KEX2KY6U9zWRkYpnnpgSHXjYTZ0ciFi5Q=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=nP18e9GZP39n5Q0NQwBKxPou9VrpS3p9zuXR5rFXlnw5Q1g/WWuS2Z35a+u6Lem/BitjFoW/DWdpyVTRkEQ9oTuy1R/aV+lqGtaqQormFkIRlAuNOWRnUyRpTTJCjt7cJZj/ODPpCz6s6OCpRPz+k7GSX6+ZQxvqkaDt+aLri3c= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=infradead.org; spf=none smtp.mailfrom=infradead.org; dkim=pass (2048-bit key) header.d=infradead.org header.i=@infradead.org header.b=mg6AH11Z; arc=none smtp.client-ip=90.155.50.34 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=infradead.org Authentication-Results: smtp.subspace.kernel.org; spf=none smtp.mailfrom=infradead.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=infradead.org header.i=@infradead.org header.b="mg6AH11Z" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=infradead.org; s=casper.20170209; h=In-Reply-To:Content-Type:MIME-Version: References:Message-ID:Subject:Cc:To:From:Date:Sender:Reply-To: Content-Transfer-Encoding:Content-ID:Content-Description; bh=gseNizjSVzU2exF+Rr3INXnOERcgb6IvTHlHFJ3oQlw=; b=mg6AH11ZR4QfAmRi3No9yHkUow ZjKviGHARTDkm2VbcySF8FoM97QdA7VnFjXy7aze075UUljKsxulWoRtDDcfOuBn3G5fZ0652R/U/ q5wADZn05s8KQbCBU7xfQeHng74i9PAm63niP0p5qp+Jb+0xMtDwZWp3y9C4/GkE6Z7P0ZZbfIcmw fzD9EwV7xzYdiVOB98XplzfOkhsW/dEI//kGIzE0SuvuO+hnjnkCnhtY/XO2XZxeFv5MvrquVshgI qUkDP4CDMHmgSxAIbv7YqB+DHcCscxj8q4NabneRyWX0Y4MAcpkJTgUjEUhELjmoNG6fNpFFS/Zpk 3E15nOOg==; Received: from 77-249-17-89.cable.dynamic.v4.ziggo.nl ([77.249.17.89] helo=noisy.programming.kicks-ass.net) by casper.infradead.org with esmtpsa (Exim 4.98 #2 (Red Hat Linux)) id 1tUqfy-0000000GlCF-3gJb; Mon, 06 Jan 2025 17:14:03 +0000 Received: by noisy.programming.kicks-ass.net (Postfix, from userid 1000) id ADF5B3005D6; Mon, 6 Jan 2025 18:14:02 +0100 (CET) Date: Mon, 6 Jan 2025 18:14:02 +0100 From: Peter Zijlstra To: Doug Smythies Cc: linux-kernel@vger.kernel.org, vincent.guittot@linaro.org Subject: Re: [REGRESSION] Re: [PATCH 00/24] Complete EEVDF Message-ID: <20250106171402.GC22191@noisy.programming.kicks-ass.net> References: <005f01db5a44$3bb698e0$b323caa0$@telus.net> <20250106115732.GE20870@noisy.programming.kicks-ass.net> <000801db604b$e0f6b580$a2e42080$@telus.net> <20250106165932.GG20870@noisy.programming.kicks-ass.net> <20250106170455.GB22191@noisy.programming.kicks-ass.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20250106170455.GB22191@noisy.programming.kicks-ass.net> On Mon, Jan 06, 2025 at 06:04:55PM +0100, Peter Zijlstra wrote: > On Mon, Jan 06, 2025 at 05:59:32PM +0100, Peter Zijlstra wrote: > > On Mon, Jan 06, 2025 at 07:01:34AM -0800, Doug Smythies wrote: > > > > > > What is the easiest 100% load you're seeing this with? > > > > > > Lately, and specifically to be able to tell others, I have been using: > > > > > > yes > /dev/null & > > > > > > On my Intel i5-10600K, with 6 cores and 2 threads per core, 12 CPUs, > > > I run 12 of those work loads. > > > > On my headless ivb-ep 2 sockets, 10 cores each and 2 threads per core, I > > do: > > > > for ((i=0; i<40; i++)) ; do yes > /dev/null & done > > tools/power/x86/turbostat/turbostat --quiet --Summary --show Busy%,Bzy_MHz,IRQ,PkgWatt,PkgTmp,TSC_MHz --interval 1 > > > > But no so far, nada :-( I've tried with full preemption and voluntary, > > HZ=1000. > > > > And just as I send this, I see these happen: > > 100.00 3100 2793 40302 71 195.22 > 100.00 3100 2618 40459 72 183.58 > 100.00 3100 2993 46215 71 209.21 > 100.00 3100 2789 40467 71 195.19 > 99.92 3100 2798 40589 71 195.76 > 100.00 3100 2793 40397 72 195.46 > ... > 100.00 3100 2844 41906 71 199.43 > 100.00 3100 2779 40468 71 194.51 > 99.96 3100 2320 40933 71 163.23 > 100.00 3100 3529 61823 72 245.70 > 100.00 3100 2793 40493 72 195.45 > 100.00 3100 2793 40462 72 195.56 > > They look like funny little blips. Nowhere near as bad as you had > though. Anyway, given you've confirmed disabling DELAY_DEQUEUE fixes things, could you perhaps try the below hackery for me? Its a bit of a wild guess, but throw stuff at wall, see what sticks etc.. --- diff --git a/kernel/sched/core.c b/kernel/sched/core.c index 84902936a620..fa4b9891f93a 100644 --- a/kernel/sched/core.c +++ b/kernel/sched/core.c @@ -3019,7 +3019,7 @@ static int affine_move_task(struct rq *rq, struct task_struct *p, struct rq_flag } else { if (!is_migration_disabled(p)) { - if (task_on_rq_queued(p)) + if (task_on_rq_queued(p) && !p->se.sched_delayed) rq = move_queued_task(rq, rf, p, dest_cpu); if (!pending->stop_pending) { @@ -3776,28 +3776,30 @@ ttwu_do_activate(struct rq *rq, struct task_struct *p, int wake_flags, */ static int ttwu_runnable(struct task_struct *p, int wake_flags) { - struct rq_flags rf; - struct rq *rq; - int ret = 0; + CLASS(__task_rq_lock, rq_guard)(p); + struct rq *rq = rq_guard.rq; - rq = __task_rq_lock(p, &rf); - if (task_on_rq_queued(p)) { - update_rq_clock(rq); - if (p->se.sched_delayed) - enqueue_task(rq, p, ENQUEUE_NOCLOCK | ENQUEUE_DELAYED); - if (!task_on_cpu(rq, p)) { - /* - * When on_rq && !on_cpu the task is preempted, see if - * it should preempt the task that is current now. - */ - wakeup_preempt(rq, p, wake_flags); + if (!task_on_rq_queued(p)) + return 0; + + update_rq_clock(rq); + if (p->se.sched_delayed) { + int queue_flags = ENQUEUE_NOCLOCK | ENQUEUE_DELAYED; + if (!is_cpu_allowed(p, cpu_of(rq))) { + dequeue_task(rq, p, DEQUEUE_SLEEP | queue_flags); + return 0; } - ttwu_do_wakeup(p); - ret = 1; + enqueue_task(rq, p, queue_flags); } - __task_rq_unlock(rq, &rf); - - return ret; + if (!task_on_cpu(rq, p)) { + /* + * When on_rq && !on_cpu the task is preempted, see if + * it should preempt the task that is current now. + */ + wakeup_preempt(rq, p, wake_flags); + } + ttwu_do_wakeup(p); + return 1; } #ifdef CONFIG_SMP diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h index 65fa64845d9f..b4c1f6c06c18 100644 --- a/kernel/sched/sched.h +++ b/kernel/sched/sched.h @@ -1793,6 +1793,11 @@ task_rq_unlock(struct rq *rq, struct task_struct *p, struct rq_flags *rf) raw_spin_unlock_irqrestore(&p->pi_lock, rf->flags); } +DEFINE_LOCK_GUARD_1(__task_rq_lock, struct task_struct, + _T->rq = __task_rq_lock(_T->lock, &_T->rf), + __task_rq_unlock(_T->rq, &_T->rf), + struct rq *rq; struct rq_flags rf) + DEFINE_LOCK_GUARD_1(task_rq_lock, struct task_struct, _T->rq = task_rq_lock(_T->lock, &_T->rf), task_rq_unlock(_T->rq, _T->lock, &_T->rf),