From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mx0b-001b2d01.pphosted.com (mx0b-001b2d01.pphosted.com [148.163.158.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 601FD3C2799 for ; Wed, 8 Apr 2026 12:56:52 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.158.5 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1775653013; cv=none; b=If3yaFJfWc0FOFkqM0ghDxoPDy7eR2nkxxZmIDskjQHrRaRIWmzl0qDXEt5g5UIG0Cx8y84KYGygjb+BA4oQCaQkACtRWF5ybkQWOuMkIG603BWhLcPcXiEdaO1Cxl5yoaviq07oZRXbMvaAFy7F1VVgXqKOxp/pX0iAd+GhSDk= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1775653013; c=relaxed/simple; bh=K9fxMIMGhNsTIgPXxwGc5dnq09NeMZ0ZQp3OzOlYl0I=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=EHymx3ruWmnqdzJvI8+diq4Ut6GdFuXdUiWUUMoeb01pbSxHXOSTKjUhBzR8BKAtBbXO2EYOE5tXiVZW1m3UQgmhZlH4F+p47xSJPTUu1DzN33d/+dwB+Os06zDG9jReih5TK9nFctahXpHhUJp1+hY8UO/wttVRMc6wiUJadvs= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=IYQCtj9E; arc=none smtp.client-ip=148.163.158.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="IYQCtj9E" Received: from pps.filterd (m0356516.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 6386FJYo2326591; Wed, 8 Apr 2026 12:56:31 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:content-type:date:from:in-reply-to :message-id:mime-version:references:subject:to; s=pp1; bh=SnaO2Q sGVLnaDZEp00XdGaUyAyXNQ6raCaxrFv8ZkW8=; b=IYQCtj9EmcncRPZUj1D/8k i24LEEkgnVXaefGxoUSbn7rSz4mTEWzds09vwIQrpIuHOkvcwITYWj4as6hFB506 edtP284Bc68mLSaq+nwPnLkDWIXXrdAgPx6j+GMeGqR8HpA09CO9QubYcw6EMlZ0 Sdp8Pz4M1sXWScbk5b2NoIZHzJhCVpsVEZXpw3TKS7P15VlH9l4sSUBxXn8vJ3Lo OyA2KzSe7G/h/diBcF2cn0NwU7/wtT5aqQykQW2PBPLLfV6yYnTkvnVn4AtSSQ7W 2bBNWoQlvAbnuE9Y9q8yDh/z7qba3Vz+f2l141JF0NtfFAlci4HuHVpp4gHkkCIg == Received: from ppma21.wdc07v.mail.ibm.com (5b.69.3da9.ip4.static.sl-reverse.com [169.61.105.91]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4dcn2kfepm-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 08 Apr 2026 12:56:30 +0000 (GMT) Received: from pps.filterd (ppma21.wdc07v.mail.ibm.com [127.0.0.1]) by ppma21.wdc07v.mail.ibm.com (8.18.1.2/8.18.1.2) with ESMTP id 638BVsxg030030; Wed, 8 Apr 2026 12:56:30 GMT Received: from smtprelay02.fra02v.mail.ibm.com ([9.218.2.226]) by ppma21.wdc07v.mail.ibm.com (PPS) with ESMTPS id 4dcme7fe6h-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 08 Apr 2026 12:56:30 +0000 Received: from smtpav02.fra02v.mail.ibm.com (smtpav02.fra02v.mail.ibm.com [10.20.54.101]) by smtprelay02.fra02v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 638CuQCj40763708 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Wed, 8 Apr 2026 12:56:26 GMT Received: from smtpav02.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 1A88220043; Wed, 8 Apr 2026 12:56:26 +0000 (GMT) Received: from smtpav02.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 391E520040; Wed, 8 Apr 2026 12:56:21 +0000 (GMT) Received: from [9.124.212.125] (unknown [9.124.212.125]) by smtpav02.fra02v.mail.ibm.com (Postfix) with ESMTP; Wed, 8 Apr 2026 12:56:20 +0000 (GMT) Message-ID: <1430f070-a0cb-4c16-a96d-99acb5e53ab1@linux.ibm.com> Date: Wed, 8 Apr 2026 18:26:20 +0530 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v2 05/17] sched/core: allow only preferred CPUs in is_cpu_allowed To: Yury Norov Cc: linux-kernel@vger.kernel.org, mingo@kernel.org, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, tglx@linutronix.de, yury.norov@gmail.com, gregkh@linuxfoundation.org, pbonzini@redhat.com, seanjc@google.com, kprateek.nayak@amd.com, vschneid@redhat.com, iii@linux.ibm.com, huschle@linux.ibm.com, rostedt@goodmis.org, dietmar.eggemann@arm.com, mgorman@suse.de, bsegall@google.com, maddy@linux.ibm.com, srikar@linux.ibm.com, hdanton@sina.com, chleroy@kernel.org, vineeth@bitbyteword.org, joelagnelf@nvidia.com References: <20260407191950.643549-1-sshegde@linux.ibm.com> <20260407191950.643549-6-sshegde@linux.ibm.com> Content-Language: en-US From: Shrikanth Hegde In-Reply-To: Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 7bit X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwNDA4MDExOCBTYWx0ZWRfX7faXUCDewrXb EOLEb2FqeOtBahP3pMA0GCgif1sVLXsZcAFfgONVx8l928n2pFepDnH35CqyDTFHhsnHITRpcaJ iY2A2imMpTy7Z2R8p/UhHnq1cJO6HdzNyVHIZrzTDxQHQVcUlQ5p+p+YB/5kCkL51Kou94UmM2K DSiZ+/C34euZrAJ5ypZ8ebQ2Oj3boHqDdC8FA+EGtr3mD6Pz+By+51e4poeov2ss/f3gjOMVHbn zF4TB1Sd+YCOFLpWtySGDd2ImZUkCwWoMiBQhkEfvY4TxG+eypyLTlPN6UOwy9e9aC7AO3GPZ6m uGWWsGO4RkHoqFBmQxtM/doqXhF0ywEMhkx3ETq0yOTgtjR6R5TRuE9FD66bSjEvAVT/Y6mUXrj vxFCAv/xO9s2u7Ef5wUOKduipVi8cZvtSgjStuipXdNIEtW3Zrimg33rjQFnKhiDqIJKopCdcN2 yAuFam42r5snvVgsoGA== X-Proofpoint-ORIG-GUID: 9OPIx0mo3wbCDTteou8xBNvSLI9KMPmZ X-Authority-Analysis: v=2.4 cv=e9k2j6p/ c=1 sm=1 tr=0 ts=69d6507f cx=c_pps a=GFwsV6G8L6GxiO2Y/PsHdQ==:117 a=GFwsV6G8L6GxiO2Y/PsHdQ==:17 a=IkcTkHD0fZMA:10 a=A5OVakUREuEA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=Y2IxJ9c9Rs8Kov3niI8_:22 a=VnNF1IyMAAAA:8 a=8xJMZ2msQdqkUbr768EA:9 a=QEXdDO2ut3YA:10 X-Proofpoint-GUID: yP_n5hirE2dmxNFdps2LBzPAjoOOKrWS X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1143,Hydra:6.1.51,FMLib:17.12.100.49 definitions=2026-04-08_04,2026-04-08_01,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 clxscore=1015 impostorscore=0 malwarescore=0 suspectscore=0 spamscore=0 bulkscore=0 adultscore=0 priorityscore=1501 phishscore=0 lowpriorityscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2604010000 definitions=main-2604080118 Hi Yury. On 4/8/26 6:35 AM, Yury Norov wrote: > On Wed, Apr 08, 2026 at 12:49:38AM +0530, Shrikanth Hegde wrote: >> >> When possible, choose a preferred CPUs to pick. >> >> Push task mechanism uses stopper thread which going to call >> select_fallback_rq and use this mechanism to pick only a preferred CPU. >> >> When task is affined only to non-preferred CPUs it should continue to >> run there. Detect that by checking if cpus_ptr and cpu_preferred_mask >> interesect or not. >> >> Signed-off-by: Shrikanth Hegde >> --- >> kernel/sched/core.c | 17 ++++++++++++++--- >> kernel/sched/sched.h | 12 ++++++++++++ >> 2 files changed, 26 insertions(+), 3 deletions(-) >> >> diff --git a/kernel/sched/core.c b/kernel/sched/core.c >> index 7ea05a7a717b..336e7c694eb7 100644 >> --- a/kernel/sched/core.c >> +++ b/kernel/sched/core.c >> @@ -2463,9 +2463,16 @@ static inline bool is_cpu_allowed(struct task_struct *p, int cpu) >> if (is_migration_disabled(p)) >> return cpu_online(cpu); >> >> - /* Non kernel threads are not allowed during either online or offline. */ >> - if (!(p->flags & PF_KTHREAD)) >> - return cpu_active(cpu); >> + /* >> + * Non kernel threads are not allowed during either online or offline. >> + * Ensure it is a preferred CPU to avoid further contention >> + */ >> + if (!(p->flags & PF_KTHREAD)) { >> + if (!cpu_active(cpu)) >> + return false; >> + if (!cpu_preferred(cpu) && task_can_run_on_preferred_cpu(p)) >> + return false; >> + } >> >> /* KTHREAD_IS_PER_CPU is always allowed. */ >> if (kthread_is_per_cpu(p)) >> @@ -2475,6 +2482,10 @@ static inline bool is_cpu_allowed(struct task_struct *p, int cpu) >> if (cpu_dying(cpu)) >> return false; >> >> + /* Try on preferred CPU first */ >> + if (!cpu_preferred(cpu) && task_can_run_on_preferred_cpu(p)) >> + return false; > First one was regular tasks, this is for unbound kernel threads. Both will need. No? > You repeat this for the 2nd time. The cpu_preferred() call should go > inside task_can_run_on_preferred_cpu(). I want to keep this check for cpu_preferred() first. the reason being it is inexpensive since it is bit check. Only if it fails, then one should bother about task_can_run_on_preferred_cpu which is O(N) as you said. I am using the task_can_run_on_preferred_cpu in push task mechanism too. PATCH 10/17. I get there only on non-preferred CPU. So can I keep as is? > > And can you please pick some shorter name? > task_has_preferred_cpus? >> + >> /* But are allowed during online. */ >> return cpu_online(cpu); >> } >> diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h >> index 88e0c93b9e21..7271af2ca64f 100644 >> --- a/kernel/sched/sched.h >> +++ b/kernel/sched/sched.h >> @@ -4130,4 +4130,16 @@ DEFINE_CLASS_IS_UNCONDITIONAL(sched_change) >> >> #include "ext.h" >> >> +#ifdef CONFIG_PARAVIRT >> +static inline bool task_can_run_on_preferred_cpu(struct task_struct *p) >> +{ >> + return cpumask_intersects(p->cpus_ptr, cpu_preferred_mask); > > This makes is_cpu_allowed() O(N). Even if CONFIG_PARAVIRT is enabled, > I think some people would prefer to avoid this. Also, select_fallback_rq() > calls it in a loop, and this makes it O(N^2). > > > /* Any allowed, online CPU? */ > for_each_cpu(dest_cpu, p->cpus_ptr) { > if (!is_cpu_allowed(p, dest_cpu)) > continue; > > goto out; > } > > You can keep it O(N): > for_each_cpu_and(dest_cpu, p->cpus_ptr, cpu_preferred_mask) { > ... > } This would leave tasks which has affinity only on non-preferred CPUs without a CPU. That breaks below case, 600 CPUs, high steal time and hence preferred is 0-399. In that state, user does "taskset -c 500 ", that task ends up going to preferred CPUs since its affinity gets reset in the switch block later in select_fallback_rq. > > Not sure how critical that path is, but this looks suspicious. > Fair point. But we come here only if cpu_preferred(cpu) == false. When that happens, we expect the task running on will get preempted to a preferred CPUs. If it migrated to a preferred cpu, then on-wards it wont suffer the additional overhead of task_can_run_on_preferred_cpu case where it would be O(N^2) is for tasks which have explicit affinity on non-preferred CPUs. I see that as rare case. For the majority of the cases, it should still be O(N) since cpu_preferred(cpu) will be true. Here is the benchmark data on system where there is 0 steal time. It is dedicated LPAR(VM). | Test | Baseline | NO STEAL | %diff to base| STEAL | %diff to base | | | | MONITOR | | MONITOR | | |---------------------------------|----------|----------|--------------|-----------|---------------| | HackBench Process 20 groups | 2.70 | 2.67 | +1.11% | 2.66 | +1.48% | | HackBench Process 40 groups | 5.31 | 5.26 | +0.94% | 5.30 | +0.19% | | HackBench Process 60 groups | 7.95 | 7.82 | +1.64% | 7.90 | +0.63% | | HackBench thread 10 Time | 1.61 | 1.59 | +1.24% | 1.56 | +3.11% | | HackBench thread 20 Time | 2.82 | 2.83 | -0.35% | 2.80 | +0.71% | | HackBench Process(Pipe) 20 Time | 1.51 | 1.50 | +0.66% | 1.47 | +2.65% | | HackBench Process(Pipe) 40 Time | 2.78 | 2.69 | +3.24% | 2.69 | +3.24% | | HackBench Process(Pipe) 60 Time | 3.73 | 3.76 | -0.80% | 3.64 | +2.41% | | HackBench thread(Pipe) 10 Time | 0.94 | 0.91 | +3.19% | 0.91 | +3.19% | | HackBench thread(Pipe) 20 Time | 1.58 | 1.58 | -0.00% | 1.59 | -0.63% | >> +} >> +#else >> +static inline bool task_can_run_on_preferred_cpu(struct task_struct *p) >> +{ >> + return true; >> +} >> +#endif > > Same comment as in patch 3. I believe, it's worth to declare cpu_preferred_mask > unrelated to CONFIG_PARAVIRT, so that you'll not have to spread this > ifdefery around. > Yes. >> + >> #endif /* _KERNEL_SCHED_SCHED_H */ >> -- >> 2.47.3