From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mx0a-001b2d01.pphosted.com (mx0a-001b2d01.pphosted.com [148.163.156.1]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0286519DF55 for ; Thu, 18 Jun 2026 05:41:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.156.1 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781761274; cv=none; b=OX6LoBhTc8T3/5Ro40knDmF70zLRLrAfGUKnBbKe2Tq1jKpdrsqjD3WmRQTYE3Z0YSRHgz9BCaZzoIOQ4eBzXywqLDY7U27g0DYq/+0XnIR45Tb4LQKrZJp0qGjvHYbCv9w+Q7UdIfVf8aUDHwapFaFokrvefhB8Gau34P3KyKk= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781761274; c=relaxed/simple; bh=Yg4IBSZAFV+8eVbZKjSAj3NUno4OaWreqwJk5nC5vtQ=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=khf3BN0bWCkA80z1PtjjvicH3PCkicQEAWN1IVNGDF1CTo035h0VKhP8w3v8FmvBFwXfh7Bx8nSLar29q8MjfcdAqHj80/xTpsxby8HKNMu64cgVCtF3ExWmb8sdP3rzBRWErtEoX/8cPduaykUy3g65ysdln2Zw4jHa3s+hFDI= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=hH0cpc/m; arc=none smtp.client-ip=148.163.156.1 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="hH0cpc/m" Received: from pps.filterd (m0356517.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 65HHmDqR1082471; Thu, 18 Jun 2026 05:40:06 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:content-type:date:from:in-reply-to :message-id:mime-version:references:subject:to; s=pp1; bh=DfyxsS YQBe9lusmr+ynYMXd4GQbrJMJzj5z8cpBTSVs=; b=hH0cpc/mU/vAhMzlXWjDEU fKIlsAICTGgb9IU8XRUd7EXb7MqOp3Fk88Tv3uUCDulwOvkQqTcojHYF40HUmZbT oM72E5+53asaSDYbLbYltYJzbSBuV1dMnomxtQhU+81X4/6r4egswYUdnOsnDqUN WGzODRDAtq6eLLEdzA18BCUhUdj59gI1W/IK1nqsbfGCQLyH1ExTPetH8ObGbFuV J0RHB9VFRy1FKM9OmJskOdawlSUZeOcpiZLyhqsW0RbM0TTSUwq3LjpDd/i8D9Do 4SKzS8JrHjtXQSZrB9utcosu90XzYWmHrI8BJsE/mFTdvztnWIpD8I0OEZycc/fw == Received: from ppma23.wdc07v.mail.ibm.com (5d.69.3da9.ip4.static.sl-reverse.com [169.61.105.93]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4eueqvxgs2-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Thu, 18 Jun 2026 05:40:05 +0000 (GMT) Received: from pps.filterd (ppma23.wdc07v.mail.ibm.com [127.0.0.1]) by ppma23.wdc07v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 65I5YoTs032157; Thu, 18 Jun 2026 05:40:04 GMT Received: from smtprelay06.fra02v.mail.ibm.com ([9.218.2.230]) by ppma23.wdc07v.mail.ibm.com (PPS) with ESMTPS id 4ev172a5sh-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Thu, 18 Jun 2026 05:40:04 +0000 (GMT) Received: from smtpav02.fra02v.mail.ibm.com (smtpav02.fra02v.mail.ibm.com [10.20.54.101]) by smtprelay06.fra02v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 65I5e0ks27722116 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Thu, 18 Jun 2026 05:40:00 GMT Received: from smtpav02.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 6A6302004B; Thu, 18 Jun 2026 05:40:00 +0000 (GMT) Received: from smtpav02.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 0B77B20043; Thu, 18 Jun 2026 05:39:54 +0000 (GMT) Received: from [9.123.5.233] (unknown [9.123.5.233]) by smtpav02.fra02v.mail.ibm.com (Postfix) with ESMTP; Thu, 18 Jun 2026 05:39:53 +0000 (GMT) Message-ID: Date: Thu, 18 Jun 2026 11:09:52 +0530 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v4 15/20] sched/core: Compute steal values at regular intervals To: Yury Norov Cc: linux-kernel@vger.kernel.org, mingo@kernel.org, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, kprateek.nayak@amd.com, iii@linux.ibm.com, tglx@kernel.org, gregkh@linuxfoundation.org, pbonzini@redhat.com, seanjc@google.com, vschneid@redhat.com, huschle@linux.ibm.com, rostedt@goodmis.org, dietmar.eggemann@arm.com, mgorman@suse.de, bsegall@google.com, maddy@linux.ibm.com, srikar@linux.ibm.com, hdanton@sina.com, chleroy@kernel.org, vineeth@bitbyteword.org, frederic@kernel.org, arighi@nvidia.com, pauld@redhat.com, christian.loehle@arm.com, tj@kernel.org, tommaso.cucinotta@gmail.com, maz@kernel.org, rafael@kernel.org References: <20260617174139.155540-1-sshegde@linux.ibm.com> <20260617174139.155540-16-sshegde@linux.ibm.com> From: Shrikanth Hegde Content-Language: en-US In-Reply-To: Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 7bit X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Proofpoint-Spam-Info: AW1haW4tMjYwNjE4MDA0NSBTYWx0ZWRfXxXvuY5jCLbVG wfS3bxLqOGIJQdyG/XWehlOFmVJ1zNLIgIZUcLODKBliUdIM7izzxl3tG7xtd3sxQbVbgn3x4kd QN/2RXZKsQFzZkJxgsQY3+WcbhUOqr4= X-Proofpoint-GUID: PWK8b32thZXqd4lyEr7Dpm0_-f6gVQan X-Authority-Analysis: v=2.4 cv=bMgm5v+Z c=1 sm=1 tr=0 ts=6a3384b6 cx=c_pps a=3Bg1Hr4SwmMryq2xdFQyZA==:117 a=3Bg1Hr4SwmMryq2xdFQyZA==:17 a=IkcTkHD0fZMA:10 a=FelO9ux0wxsA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=U7nrCbtTmkRpXpFmAIza:22 a=VnNF1IyMAAAA:8 a=gvhWzT5KAuCHVE-r-poA:9 a=QEXdDO2ut3YA:10 X-Proofpoint-ORIG-GUID: 8KJsboS3Ewq9owWfKorcYPekmi2C_kIU X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwNjE4MDA0NSBTYWx0ZWRfX6yhPvwAZDBWb 6R5PCD6X6Nm/lXRMi+zaaOy9zGnr06ccACHBxwQlfIV79bV57Qu/ilUm1LJqSPxxGb2BpSu/ivK 7J/bIFIdJkUYImgB7nBZWaq5Ur7BrVoLYROtjJofEgehICAU9E7grNV15VMtId/LJZfazEzBnFP u0K9xGbT4YylCtazPYK4mnEQrFAiT4a1oK/c12I8Py7egsSzUGJeFjtwmhAS5EdF6b3K7kBTEqe jMIlfOLvPb0reaHP+itpcK5IIKqNMAxJ5DxAIKsGqa/rY+qGUe+sxi2sPKwUWqmeSpMqYEz/1rw vob1fknSEXhw4GIThHQx35Y751OmA5uAxy84doZN5jP4WLCx7js47mQYRLa9Mmo7cLXE5WPRDro Tqy1zDw1FJ8QhmH1Zl/F+iBNq0/qscd4qdJTGasApQ8EVImg/xRHix0dsfLnf2fP0TTG3rDh9PU bhyHSGvUgQ5xmsFLqfA== X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1143,Hydra:6.1.125,FMLib:17.12.100.49 definitions=2026-06-17_02,2026-06-17_03,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 suspectscore=0 impostorscore=0 malwarescore=0 adultscore=0 spamscore=0 bulkscore=0 priorityscore=1501 phishscore=0 clxscore=1015 lowpriorityscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2606180045 On 6/18/26 9:34 AM, Yury Norov wrote: > On Wed, Jun 17, 2026 at 11:11:34PM +0530, Shrikanth Hegde wrote: >> Kick off the work to compute the steal time at regular interval. >> Gated with steal monitor enabled static key check to avoid any overhead >> when its disabled. >> >> The sampling period can changed at runtime using steal_mon/sampling_period. >> By default is 1000 milliseconds. I.e. 1 second >> >> This work is done by first active housekeeping CPU only. Hence it won't >> need any complicated synchronization. >> >> Now, that sched_steal_mon_enabled() is available which is a static branch, >> add this to hotpath such as wakeup and load balance. >> This will make them effectively nop when the feature is disabled. >> >> Signed-off-by: Shrikanth Hegde >> --- >> v3->v4: >> - Add static key check in hotpaths. Could be split into a separate >> patch. Let me know if thats better. >> >> include/linux/sched.h | 2 ++ >> kernel/sched/core.c | 28 +++++++++++++++++++++++++++- >> kernel/sched/debug.c | 1 + >> kernel/sched/fair.c | 3 ++- >> kernel/sched/sched.h | 10 +++++++++- >> 5 files changed, 41 insertions(+), 3 deletions(-) >> >> diff --git a/include/linux/sched.h b/include/linux/sched.h >> index ce6bc8a22eb1..5b15353ed7ef 100644 >> --- a/include/linux/sched.h >> +++ b/include/linux/sched.h >> @@ -2527,5 +2527,7 @@ struct steal_monitor_t { >> unsigned int high_threshold; >> unsigned int sampling_period_ms; >> }; >> + >> +extern struct steal_monitor_t steal_mon; >> #endif >> #endif >> diff --git a/kernel/sched/core.c b/kernel/sched/core.c >> index cc48632dd42d..f1a91021e357 100644 >> --- a/kernel/sched/core.c >> +++ b/kernel/sched/core.c >> @@ -5793,7 +5793,7 @@ void sched_tick(void) >> unsigned long hw_pressure; >> u64 resched_latency; >> >> - if (!cpu_preferred(cpu)) >> + if (sched_steal_mon_enabled() && !cpu_preferred(cpu)) >> sched_push_current_non_preferred_cpu(rq); > > This looks like CPU can be non-preferred only if steal monitor is > enabled. To properly implement it, you need to mark all active CPUs > as preferred during the steal monitor disabling. That way you don't > need to complicate the condition. > That is done in disabling the feature [PATCH 13/20]. if (sched_sm_wr_enable && !orig) { static_branch_enable(&__sched_sm_enable); } else if (!sched_sm_wr_enable && orig) { static_branch_disable(&__sched_sm_enable); cpumask_copy(&__cpu_preferred_mask, cpu_active_mask); >> >> if (housekeeping_cpu(cpu, HK_TYPE_KERNEL_NOISE)) >> @@ -5834,6 +5834,9 @@ void sched_tick(void) >> rq->idle_balance = idle_cpu(cpu); >> sched_balance_trigger(rq); >> } >> + >> + if (sched_steal_mon_enabled()) >> + sched_trigger_steal_computation(cpu); >> } >> >> #ifdef CONFIG_NO_HZ_FULL >> @@ -11407,4 +11410,27 @@ void sched_steal_detection_work(struct work_struct *work) >> now = ktime_get(); >> sm->prev_time = now; >> } >> + >> +void sched_trigger_steal_computation(int cpu) >> +{ >> + int first_hk_cpu = cpumask_first_and(housekeeping_cpumask(HK_TYPE_KERNEL_NOISE), >> + cpu_active_mask); >> + ktime_t now; >> + >> + /* Done by first active housekeeping CPU only */ >> + if (likely(cpu != first_hk_cpu)) >> + return; >> + >> + /* >> + * Since everything is updated by first housekeeping CPU, >> + * There is no need for complex syncronization. >> + */ >> + now = ktime_get(); >> + >> + /* Default is once per second */ >> + if (likely(ktime_ms_delta(now, steal_mon.prev_time) < steal_mon.sampling_period_ms)) >> + return; >> + >> + schedule_work_on(first_hk_cpu, &steal_mon.work); > > I think, there should be a better way to schedule a work on regular > interval... > > Maybe steal_mon.work would schedule itself? So, the first time it's > scheduled on steal monitor enablement, and then just reschedules > itself. This way you'll avoid polluting sched_tick(). > That's a good idea too. All it needs a periodic call at the granularity of sampling period. maybe start hrtimer when enabling the feature. hrtimer kicks off the work. work function enables the hrtimer again. > >> +} >> #endif >> diff --git a/kernel/sched/debug.c b/kernel/sched/debug.c >> index 2d62858f9cc0..55b8beb42574 100644 >> --- a/kernel/sched/debug.c >> +++ b/kernel/sched/debug.c >> @@ -649,6 +649,7 @@ static ssize_t sched_sm_en_write(struct file *filp, const char __user *ubuf, >> static_branch_enable(&__sched_sm_enable); >> } else if (!sched_sm_wr_enable && orig) { >> static_branch_disable(&__sched_sm_enable); >> + cancel_work_sync(&steal_mon.work); >> cpumask_copy(&__cpu_preferred_mask, cpu_active_mask); >> } >> >> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c >> index 3f3c7f0ca489..b02a414ffaae 100644 >> --- a/kernel/sched/fair.c >> +++ b/kernel/sched/fair.c >> @@ -13292,7 +13292,8 @@ static int sched_balance_rq(int this_cpu, struct rq *this_rq, >> cpumask_and(cpus, sched_domain_span(sd), cpu_active_mask); >> >> /* Spread load among preferred CPUs */ >> - cpumask_and(cpus, cpus, cpu_preferred_mask); >> + if (sched_steal_mon_enabled()) >> + cpumask_and(cpus, cpus, cpu_preferred_mask); > > Again, if you mark do cpumask_copy(preferred, active) on the steal > monitor disablement, you don't need to complicate core logic here and > there. > >> >> schedstat_inc(sd->lb_count[idle]); >> >> diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h >> index 984da3827f19..f3814099cc0b 100644 >> --- a/kernel/sched/sched.h >> +++ b/kernel/sched/sched.h >> @@ -1060,6 +1060,7 @@ struct root_domain { >> struct perf_domain __rcu *pd; >> }; >> >> +static inline bool sched_steal_mon_enabled(void); >> extern void init_defrootdomain(void); >> extern int sched_init_domains(const struct cpumask *cpu_map); >> extern void rq_attach_root(struct rq *rq, struct root_domain *rd); >> @@ -1436,7 +1437,7 @@ static inline bool available_idle_cpu(int cpu) >> if (!idle_rq(cpu_rq(cpu))) >> return 0; >> >> - if (!cpu_preferred(cpu)) >> + if (sched_steal_mon_enabled() && !cpu_preferred(cpu)) >> return 0; >> >> if (vcpu_is_preempted(cpu)) >> @@ -4243,8 +4244,15 @@ DECLARE_STATIC_KEY_FALSE(__sched_sm_enable); >> void sched_init_steal_monitor(void); >> void sched_steal_detection_work(struct work_struct *work); >> void sched_push_current_non_preferred_cpu(struct rq *rq); >> +void sched_trigger_steal_computation(int cpu); >> +static inline bool sched_steal_mon_enabled(void) >> +{ >> + return static_branch_unlikely(&__sched_sm_enable); >> +} >> #else /* !CONFIG_PREFERRED_CPU */ >> static inline void sched_push_current_non_preferred_cpu(struct rq *rq) { } >> static inline void sched_init_steal_monitor(void) { } >> +static inline void sched_trigger_steal_computation(int cpu) { } >> +static inline bool sched_steal_mon_enabled(void) { return false; } >> #endif >> #endif /* _KERNEL_SCHED_SCHED_H */ >> -- >> 2.47.3