From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pl1-f173.google.com (mail-pl1-f173.google.com [209.85.214.173]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8AFE027FD76 for ; Sun, 1 Mar 2026 08:36:56 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.173 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1772354218; cv=none; b=grtfnaTlv3PyV4PaFTvSzp445pLCqQAvHkuSWlK9tcpVBwLph8Y0YM8e27kKCyy+hHEzrFWG1aKe8k1rDIZCrmRVfTzHJu0I8wJM3QALApw7g61T6Nkr+3YYG/HUctS3oaiv1nlxveRGB6og7fYC8xiXCscK7nm8T0ZuU4q9dew= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1772354218; c=relaxed/simple; bh=DNKCygubMvS/J6FYbEOQ+2BOZw99zB+m6S3j21z8ndI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=aksMoyc+Bqm4jk/i/PiCSUY4+PQuKa+CP8Djg0gKNavb2oCsbh1CfK08yeCLHKcAK1wBJRtSdRht8acfSYY3ftPeEyub0UD1XaxmOluxnnXGnDx764TIr7M6yMJRTMElrQ9/U63b6A6NzdscmP3FyGf+gF0KCIMkvdwqSpP9XGk= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=KIcdrvXl; arc=none smtp.client-ip=209.85.214.173 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="KIcdrvXl" Received: by mail-pl1-f173.google.com with SMTP id d9443c01a7336-2a7a9b8ed69so40642525ad.2 for ; Sun, 01 Mar 2026 00:36:56 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1772354216; x=1772959016; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to; bh=OL2KbCMFg/WPrKJuC3joywLs3WaKufzdMJwuprQDC+k=; b=KIcdrvXlAMIvFmjwgO2bsr4PyA4cDwy2LSGi4qeyfHf5i1565/KNa9QCIGfrds8ILF T4N5QokrqqckjbPY3EaQFAG+0NkZNZwT8ReyHA9yxEOZ2gl4v522UlRoElMAqubFjAy0 lZ6MZunbAXgGVOIV9D/Bo2mgsWf9jaerFnsIfinYO+GEAKwb6x1crTjcyVzmsLlWTkwp 8fFOsKPZsSoK0CafGX5dlEIBczHFwFjXrHmD107kOgDFBipx/TO9aAaSTct9gM3FOzKj jEJoAlWY0lxhmUgBhgnN1ab5zOIltglA+FJiXRfLwtXIx+34HxHxRImpSdt1R46eya74 6tyg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1772354216; x=1772959016; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to; bh=OL2KbCMFg/WPrKJuC3joywLs3WaKufzdMJwuprQDC+k=; b=ocpmA7jAQ6UpKi09Al+pExDDMxeYvdHV9F0jiJCoIYKdIlheh5P9cwU5ebZUKLYYqA 2Y3gsncld4pWdxQAjtz3go8ToJAsEoURKOwqWHelr7Z8Vnt/UvirtRELuSrTOXEN8eLm TCaiGJCDujkPtWDkhXNpVMqOGtT/HXIsdyU4zAIyjInlNtaLBVGpnnPL2bfqQC86btuU g1MzJxu5PjghoKkJJh8yarKc231Ew77mI4A0IoQ+W68sPv0P+K1qfEEKe/zve81iIJlH ByOdAJv27Z0AuWgKvX2+7SC5K5QCi2GFdWcYMEVq+2Zl/Z9rt86zfcvlGTT5Ld1xq9ey AT8Q== X-Forwarded-Encrypted: i=1; AJvYcCUk1pmQ4osJOLpLmrTCcMlORzA8JUXDP5ra9V6inEqcLGsvAiBV+W/RHD69lvcZX6IcPGVFTU43QKZppvE=@vger.kernel.org X-Gm-Message-State: AOJu0YyD1iHVG374atLGCjQi/aXebUZwhgKYdDFipvZd4rAjs9FAumrW YsIaA4NJE2ld1TRBs0Hq3YpPPwHNiSTp9yfzSpOcyG3n4q/Nf3WRuBJE X-Gm-Gg: ATEYQzzwMAM/0Ey1hZ+jsnZTIg64l2TrFKfYRNo0eCVQKMJs9k7dy3y22Q91eVouB3r 6wmMwxaRqPkKEyWGZcG6Fw/dLhGqdQlVUNzeJ3lB0HcS3pY5F3GvV0nxxd95gLCt1l4T7QR7BiK GP+H6YhChwguXDdGdzekfHi+6Clxs19AexPcBSkDbv2+umOzoH0SQxMxzQ/qMsgGnnYxhykkZJv 744TyfYFzcf6EthPY2fWPOaGfWhaU81p+VZxlaEVQG1VAXSjzuGxv1/7X8pBbrTPuLbBpnTCPft TtMyL367DwrfGtu0Gktloqfqjdg+SsTYsQqj+mRgb4/JNxXwNU6VT8VKcNMLjt51HTjOnOzu+b1 lMY8YBSTUMCwYNhcZ6a1gmZsJT4T55YYAeDzwPM3DG/4+1bKwuKwP+rIUaYWh+wo7130AGVhkVz VuD0cu6RQK2nR9EKgyhsbPWtjxlf/mj0YZTR4ieBJT45El+GvQf6Z3tg== X-Received: by 2002:a17:903:37d0:b0:2ad:dfb8:8ed3 with SMTP id d9443c01a7336-2ae2e3ecea7mr79127965ad.8.1772354215752; Sun, 01 Mar 2026 00:36:55 -0800 (PST) Received: from DESKTOP-3LEPQG8.localdomain ([119.28.20.50]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2adfb6b416bsm107799835ad.61.2026.03.01.00.36.42 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 01 Mar 2026 00:36:55 -0800 (PST) From: Xie Yuanbin To: peterz@infradead.org, tglx@kernel.org, dave.hansen@linux.intel.com, hpa@zytor.com, riel@surriel.com, david@kernel.org, segher@kernel.crashing.org, arnd@arndb.de, mingo@redhat.com, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, bp@alien8.de, luto@kernel.org, frederic@kernel.org, mingo@kernel.org, houwenlong.hwl@antgroup.com, will@kernel.org, akpm@linux-foundation.org, jgross@suse.com, baohua@kernel.org, ryan.roberts@arm.com, lorenzo.stoakes@oracle.com, nysal@linux.ibm.com, urezki@gmail.com, max.kellermann@ionos.com Cc: x86@kernel.org, linux-kernel@vger.kernel.org, Xie Yuanbin Subject: [PATCH v8 3/3] sched/core: Make finish_task_switch() and its subfunctions always inline Date: Sun, 1 Mar 2026 16:35:20 +0800 Message-ID: <20260301083520.110969-4-qq570070308@gmail.com> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260301083520.110969-1-qq570070308@gmail.com> References: <20260301083520.110969-1-qq570070308@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit finish_task_switch() is not inlined even in the O2 level optimization, performance testing indicates that this could lead to a significant performance degradation when certain Spectre vulnerability mitigations are enabled. In switch_mm_irq_off(), some mitigations may clear branch prediction history, or the instruction cache, like arm64_apply_bp_hardening() on arm64, BPIALL/ICIALLU on arm, and indirect_branch_prediction_barrier() on x86. finish_task_switch() is right after switch_mm_irqs_off(), the performance is greatly affected by function calls and branch jumps. __schedule() has a __sched attribute, which makes it be placed in '.sched.text' section, while finish_task_switch() does not. This makes they "far away from each other" in vmlinux, which aggravating the performance degradation. Make finish_task_switch() and its subfunctions always inline to optimize the performance. Performance test - time spent on calling finish_task_switch(): 1. x86-64: Intel i5-8300h@4Ghz, DDR4@2666mhz; unit: x86's tsc | test scenario | old | new | delta | | gcc 15.2 | 27.50 | 25.45 | -2.05 ( -7.5%) | | gcc 15.2 + spectre_v2_user=on | 46.75 | 25.96 | -20.79 (-44.5%) | | clang 21.1.7 | 27.25 | 25.45 | -1.80 ( -6.6%) | | clang 21.1.7 + spectre_v2_user=on | 39.50 | 26.00 | -13.50 (-34.2%) | 2. x86-64: AMD 9600x@5.45Ghz, DDR5@4800mhz; unit: x86's tsc | test scenario | old | new | delta | | gcc 15.2 | 27.51 | 27.51 | 0 ( 0%) | | gcc 15.2 + spectre_v2_user=on | 105.21 | 67.89 | -37.32 (-35.5%) | | clang 21.1.7 | 27.51 | 27.51 | 0 ( 0%) | | clang 21.1.7 + spectre_v2_user=on | 104.15 | 67.52 | -36.63 (-35.2%) | 3. arm64: Raspberry Pi 3b Rev 1.2, Cortex-A53@1.2Ghz; unit: cntvct_el0 | test scenario | old | new | delta | | gcc 15.2 | 1.453 | 1.115 | -0.338 (-23.3%) | | clang 21.1.7 | 1.532 | 1.123 | -0.409 (-26.7%) | 4. arm32: Raspberry Pi 3b Rev 1.2, Cortex-A53@1.2Ghz; unit: cntvct_el0 | test scenario | old | new | delta | | gcc 15.2 | 1.421 | 1.187 | -0.234 (-16.5%) | | clang 21.1.7 | 1.437 | 1.200 | -0.237 (-16.5%) | Cc: Peter Zijlstra Cc: Thomas Gleixner Cc: Rik van Riel Cc: Segher Boessenkool Cc: David Hildenbrand (Arm) Cc: H. Peter Anvin (Intel) Cc: Arnd Bergmann Signed-off-by: Xie Yuanbin --- More detailed information about the test can be found in the cover letter: Link: https://lore.kernel.org/20260301083520.110969-1-qq570070308@gmail.com arch/arm/include/asm/mmu_context.h | 2 +- arch/riscv/include/asm/sync_core.h | 2 +- arch/s390/include/asm/mmu_context.h | 2 +- arch/sparc/include/asm/mmu_context_64.h | 2 +- arch/x86/include/asm/sync_core.h | 2 +- include/linux/perf_event.h | 2 +- include/linux/sched/mm.h | 20 ++++++++++---------- include/linux/tick.h | 4 ++-- include/linux/vtime.h | 8 ++++---- kernel/sched/core.c | 12 ++++++------ kernel/sched/sched.h | 24 ++++++++++++------------ 11 files changed, 40 insertions(+), 40 deletions(-) diff --git a/arch/arm/include/asm/mmu_context.h b/arch/arm/include/asm/mmu_context.h index db2cb06aa8cf..bebde469f81a 100644 --- a/arch/arm/include/asm/mmu_context.h +++ b/arch/arm/include/asm/mmu_context.h @@ -80,7 +80,7 @@ static inline void check_and_switch_context(struct mm_struct *mm, #ifndef MODULE #define finish_arch_post_lock_switch \ finish_arch_post_lock_switch -static inline void finish_arch_post_lock_switch(void) +static __always_inline void finish_arch_post_lock_switch(void) { struct mm_struct *mm = current->mm; diff --git a/arch/riscv/include/asm/sync_core.h b/arch/riscv/include/asm/sync_core.h index 9153016da8f1..2fe6b7fe6b12 100644 --- a/arch/riscv/include/asm/sync_core.h +++ b/arch/riscv/include/asm/sync_core.h @@ -6,7 +6,7 @@ * RISC-V implements return to user-space through an xRET instruction, * which is not core serializing. */ -static inline void sync_core_before_usermode(void) +static __always_inline void sync_core_before_usermode(void) { asm volatile ("fence.i" ::: "memory"); } diff --git a/arch/s390/include/asm/mmu_context.h b/arch/s390/include/asm/mmu_context.h index bd1ef5e2d2eb..95d03be2ce45 100644 --- a/arch/s390/include/asm/mmu_context.h +++ b/arch/s390/include/asm/mmu_context.h @@ -93,7 +93,7 @@ static inline void switch_mm(struct mm_struct *prev, struct mm_struct *next, } #define finish_arch_post_lock_switch finish_arch_post_lock_switch -static inline void finish_arch_post_lock_switch(void) +static __always_inline void finish_arch_post_lock_switch(void) { struct task_struct *tsk = current; struct mm_struct *mm = tsk->mm; diff --git a/arch/sparc/include/asm/mmu_context_64.h b/arch/sparc/include/asm/mmu_context_64.h index 78bbacc14d2d..d1967214ef25 100644 --- a/arch/sparc/include/asm/mmu_context_64.h +++ b/arch/sparc/include/asm/mmu_context_64.h @@ -160,7 +160,7 @@ static inline void arch_start_context_switch(struct task_struct *prev) } #define finish_arch_post_lock_switch finish_arch_post_lock_switch -static inline void finish_arch_post_lock_switch(void) +static __always_inline void finish_arch_post_lock_switch(void) { /* Restore the state of MCDPER register for the new process * just switched to. diff --git a/arch/x86/include/asm/sync_core.h b/arch/x86/include/asm/sync_core.h index 96bda43538ee..4b55fa353bb5 100644 --- a/arch/x86/include/asm/sync_core.h +++ b/arch/x86/include/asm/sync_core.h @@ -93,7 +93,7 @@ static __always_inline void sync_core(void) * to user-mode. x86 implements return to user-space through sysexit, * sysrel, and sysretq, which are not core serializing. */ -static inline void sync_core_before_usermode(void) +static __always_inline void sync_core_before_usermode(void) { /* With PTI, we unconditionally serialize before running user code. */ if (static_cpu_has(X86_FEATURE_PTI)) diff --git a/include/linux/perf_event.h b/include/linux/perf_event.h index 48d851fbd8ea..7c1dac8da5e5 100644 --- a/include/linux/perf_event.h +++ b/include/linux/perf_event.h @@ -1632,7 +1632,7 @@ static inline void perf_event_task_migrate(struct task_struct *task) task->sched_migrated = 1; } -static inline void perf_event_task_sched_in(struct task_struct *prev, +static __always_inline void perf_event_task_sched_in(struct task_struct *prev, struct task_struct *task) { if (static_branch_unlikely(&perf_sched_events)) diff --git a/include/linux/sched/mm.h b/include/linux/sched/mm.h index 95d0040df584..a17730d63f88 100644 --- a/include/linux/sched/mm.h +++ b/include/linux/sched/mm.h @@ -32,7 +32,7 @@ extern struct mm_struct *mm_alloc(void); * See also for an in-depth explanation * of &mm_struct.mm_count vs &mm_struct.mm_users. */ -static inline void mmgrab(struct mm_struct *mm) +static __always_inline void mmgrab(struct mm_struct *mm) { atomic_inc(&mm->mm_count); } @@ -44,7 +44,7 @@ static inline void smp_mb__after_mmgrab(void) extern void __mmdrop(struct mm_struct *mm); -static inline void mmdrop(struct mm_struct *mm) +static __always_inline void mmdrop(struct mm_struct *mm) { /* * The implicit full barrier implied by atomic_dec_and_test() is @@ -71,27 +71,27 @@ static inline void __mmdrop_delayed(struct rcu_head *rhp) * Invoked from finish_task_switch(). Delegates the heavy lifting on RT * kernels via RCU. */ -static inline void mmdrop_sched(struct mm_struct *mm) +static __always_inline void mmdrop_sched(struct mm_struct *mm) { /* Provides a full memory barrier. See mmdrop() */ if (atomic_dec_and_test(&mm->mm_count)) call_rcu(&mm->delayed_drop, __mmdrop_delayed); } #else -static inline void mmdrop_sched(struct mm_struct *mm) +static __always_inline void mmdrop_sched(struct mm_struct *mm) { mmdrop(mm); } #endif /* Helpers for lazy TLB mm refcounting */ -static inline void mmgrab_lazy_tlb(struct mm_struct *mm) +static __always_inline void mmgrab_lazy_tlb(struct mm_struct *mm) { if (IS_ENABLED(CONFIG_MMU_LAZY_TLB_REFCOUNT)) mmgrab(mm); } -static inline void mmdrop_lazy_tlb(struct mm_struct *mm) +static __always_inline void mmdrop_lazy_tlb(struct mm_struct *mm) { if (IS_ENABLED(CONFIG_MMU_LAZY_TLB_REFCOUNT)) { mmdrop(mm); @@ -104,7 +104,7 @@ static inline void mmdrop_lazy_tlb(struct mm_struct *mm) } } -static inline void mmdrop_lazy_tlb_sched(struct mm_struct *mm) +static __always_inline void mmdrop_lazy_tlb_sched(struct mm_struct *mm) { if (IS_ENABLED(CONFIG_MMU_LAZY_TLB_REFCOUNT)) mmdrop_sched(mm); @@ -128,12 +128,12 @@ static inline void mmdrop_lazy_tlb_sched(struct mm_struct *mm) * See also for an in-depth explanation * of &mm_struct.mm_count vs &mm_struct.mm_users. */ -static inline void mmget(struct mm_struct *mm) +static __always_inline void mmget(struct mm_struct *mm) { atomic_inc(&mm->mm_users); } -static inline bool mmget_not_zero(struct mm_struct *mm) +static __always_inline bool mmget_not_zero(struct mm_struct *mm) { return atomic_inc_not_zero(&mm->mm_users); } @@ -532,7 +532,7 @@ enum { #include #endif -static inline void membarrier_mm_sync_core_before_usermode(struct mm_struct *mm) +static __always_inline void membarrier_mm_sync_core_before_usermode(struct mm_struct *mm) { /* * The atomic_read() below prevents CSE. The following should diff --git a/include/linux/tick.h b/include/linux/tick.h index 738007d6f577..df9934a60faf 100644 --- a/include/linux/tick.h +++ b/include/linux/tick.h @@ -177,7 +177,7 @@ extern cpumask_var_t tick_nohz_full_mask; #ifdef CONFIG_NO_HZ_FULL extern bool tick_nohz_full_running; -static inline bool tick_nohz_full_enabled(void) +static __always_inline bool tick_nohz_full_enabled(void) { if (!context_tracking_enabled()) return false; @@ -301,7 +301,7 @@ static inline void __tick_nohz_task_switch(void) { } static inline void tick_nohz_full_setup(cpumask_var_t cpumask) { } #endif -static inline void tick_nohz_task_switch(void) +static __always_inline void tick_nohz_task_switch(void) { if (tick_nohz_full_enabled()) __tick_nohz_task_switch(); diff --git a/include/linux/vtime.h b/include/linux/vtime.h index 29dd5b91dd7d..428464bb81b3 100644 --- a/include/linux/vtime.h +++ b/include/linux/vtime.h @@ -67,24 +67,24 @@ static __always_inline void vtime_account_guest_exit(void) * For now vtime state is tied to context tracking. We might want to decouple * those later if necessary. */ -static inline bool vtime_accounting_enabled(void) +static __always_inline bool vtime_accounting_enabled(void) { return context_tracking_enabled(); } -static inline bool vtime_accounting_enabled_cpu(int cpu) +static __always_inline bool vtime_accounting_enabled_cpu(int cpu) { return context_tracking_enabled_cpu(cpu); } -static inline bool vtime_accounting_enabled_this_cpu(void) +static __always_inline bool vtime_accounting_enabled_this_cpu(void) { return context_tracking_enabled_this_cpu(); } extern void vtime_task_switch_generic(struct task_struct *prev); -static inline void vtime_task_switch(struct task_struct *prev) +static __always_inline void vtime_task_switch(struct task_struct *prev) { if (vtime_accounting_enabled_this_cpu()) vtime_task_switch_generic(prev); diff --git a/kernel/sched/core.c b/kernel/sched/core.c index b59bab255e57..e394731434bd 100644 --- a/kernel/sched/core.c +++ b/kernel/sched/core.c @@ -4889,7 +4889,7 @@ static inline void prepare_task(struct task_struct *next) WRITE_ONCE(next->on_cpu, 1); } -static inline void finish_task(struct task_struct *prev) +static __always_inline void finish_task(struct task_struct *prev) { /* * This must be the very last reference to @prev from this CPU. After @@ -4905,7 +4905,7 @@ static inline void finish_task(struct task_struct *prev) smp_store_release(&prev->on_cpu, 0); } -static void do_balance_callbacks(struct rq *rq, struct balance_callback *head) +static __always_inline void do_balance_callbacks(struct rq *rq, struct balance_callback *head) { void (*func)(struct rq *rq); struct balance_callback *next; @@ -4940,7 +4940,7 @@ struct balance_callback balance_push_callback = { .func = balance_push, }; -static inline struct balance_callback * +static __always_inline struct balance_callback * __splice_balance_callbacks(struct rq *rq, bool split) { struct balance_callback *head = rq->balance_callback; @@ -5014,7 +5014,7 @@ prepare_lock_switch(struct rq *rq, struct task_struct *next, struct rq_flags *rf __acquire(__rq_lockp(this_rq())); } -static inline void finish_lock_switch(struct rq *rq) +static __always_inline void finish_lock_switch(struct rq *rq) __releases(__rq_lockp(rq)) { /* @@ -5047,7 +5047,7 @@ static inline void kmap_local_sched_out(void) #endif } -static inline void kmap_local_sched_in(void) +static __always_inline void kmap_local_sched_in(void) { #ifdef CONFIG_KMAP_LOCAL if (unlikely(current->kmap_ctrl.idx)) @@ -5101,7 +5101,7 @@ prepare_task_switch(struct rq *rq, struct task_struct *prev, * past. 'prev == current' is still correct but we need to recalculate this_rq * because prev may have moved to another CPU. */ -static struct rq *finish_task_switch(struct task_struct *prev) +static __always_inline struct rq *finish_task_switch(struct task_struct *prev) __releases(__rq_lockp(this_rq())) { struct rq *rq = this_rq(); diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h index 953d89d71804..b3a0d8174546 100644 --- a/kernel/sched/sched.h +++ b/kernel/sched/sched.h @@ -1430,12 +1430,12 @@ static inline struct cpumask *sched_group_span(struct sched_group *sg); DECLARE_STATIC_KEY_FALSE(__sched_core_enabled); -static inline bool sched_core_enabled(struct rq *rq) +static __always_inline bool sched_core_enabled(struct rq *rq) { return static_branch_unlikely(&__sched_core_enabled) && rq->core_enabled; } -static inline bool sched_core_disabled(void) +static __always_inline bool sched_core_disabled(void) { return !static_branch_unlikely(&__sched_core_enabled); } @@ -1444,7 +1444,7 @@ static inline bool sched_core_disabled(void) * Be careful with this function; not for general use. The return value isn't * stable unless you actually hold a relevant rq->__lock. */ -static inline raw_spinlock_t *rq_lockp(struct rq *rq) +static __always_inline raw_spinlock_t *rq_lockp(struct rq *rq) { if (sched_core_enabled(rq)) return &rq->core->__lock; @@ -1452,7 +1452,7 @@ static inline raw_spinlock_t *rq_lockp(struct rq *rq) return &rq->__lock; } -static inline raw_spinlock_t *__rq_lockp(struct rq *rq) +static __always_inline raw_spinlock_t *__rq_lockp(struct rq *rq) __returns_ctx_lock(rq_lockp(rq)) /* alias them */ { if (rq->core_enabled) @@ -1547,12 +1547,12 @@ static inline bool sched_core_disabled(void) return true; } -static inline raw_spinlock_t *rq_lockp(struct rq *rq) +static __always_inline raw_spinlock_t *rq_lockp(struct rq *rq) { return &rq->__lock; } -static inline raw_spinlock_t *__rq_lockp(struct rq *rq) +static __always_inline raw_spinlock_t *__rq_lockp(struct rq *rq) __returns_ctx_lock(rq_lockp(rq)) /* alias them */ { return &rq->__lock; @@ -1607,33 +1607,33 @@ extern void raw_spin_rq_lock_nested(struct rq *rq, int subclass) extern bool raw_spin_rq_trylock(struct rq *rq) __cond_acquires(true, __rq_lockp(rq)); -static inline void raw_spin_rq_lock(struct rq *rq) +static __always_inline void raw_spin_rq_lock(struct rq *rq) __acquires(__rq_lockp(rq)) { raw_spin_rq_lock_nested(rq, 0); } -static inline void raw_spin_rq_unlock(struct rq *rq) +static __always_inline void raw_spin_rq_unlock(struct rq *rq) __releases(__rq_lockp(rq)) { raw_spin_unlock(rq_lockp(rq)); } -static inline void raw_spin_rq_lock_irq(struct rq *rq) +static __always_inline void raw_spin_rq_lock_irq(struct rq *rq) __acquires(__rq_lockp(rq)) { local_irq_disable(); raw_spin_rq_lock(rq); } -static inline void raw_spin_rq_unlock_irq(struct rq *rq) +static __always_inline void raw_spin_rq_unlock_irq(struct rq *rq) __releases(__rq_lockp(rq)) { raw_spin_rq_unlock(rq); local_irq_enable(); } -static inline unsigned long _raw_spin_rq_lock_irqsave(struct rq *rq) +static __always_inline unsigned long _raw_spin_rq_lock_irqsave(struct rq *rq) __acquires(__rq_lockp(rq)) { unsigned long flags; @@ -1644,7 +1644,7 @@ static inline unsigned long _raw_spin_rq_lock_irqsave(struct rq *rq) return flags; } -static inline void raw_spin_rq_unlock_irqrestore(struct rq *rq, unsigned long flags) +static __always_inline void raw_spin_rq_unlock_irqrestore(struct rq *rq, unsigned long flags) __releases(__rq_lockp(rq)) { raw_spin_rq_unlock(rq); -- 2.51.0