From: Joel Fernandes <joelagnelf@nvidia.com>
To: Karl Mehltretter <kmehltretter@gmail.com>,
Peter Zijlstra <peterz@infradead.org>,
Thomas Gleixner <tglx@kernel.org>
Cc: Sebastian Andrzej Siewior <bigeasy@linutronix.de>,
Frederic Weisbecker <frederic@kernel.org>,
Clark Williams <clrkwllms@kernel.org>,
Steven Rostedt <rostedt@goodmis.org>,
Boqun Feng <boqun@kernel.org>, Lyude Paul <lyude@redhat.com>,
Alexander Potapenko <glider@google.com>,
Marco Elver <elver@google.com>, Jonathan Corbet <corbet@lwn.net>,
Bradley Morgan <brads@mainlining.org>,
linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org,
linux-rt-devel@lists.linux.dev
Subject: Re: [PATCH v5] softirq: Preserve interrupt context during IRQ exit
Date: Fri, 2 Oct 2026 20:50:39 -0400 [thread overview]
Message-ID: <fe301e2c-504b-42dd-954e-a133a54aa0ce@nvidia.com> (raw)
In-Reply-To: <20260930191432.62760-1-kmehltretter@gmail.com>
On 9/30/2026 3:14 PM, Karl Mehltretter wrote:
> On the return from interrupt path, __irq_exit_rcu() removes
> HARDIRQ_OFFSET from the preemption counter at the very top of the
> function. Everything after that reports the current context as task
> instead of hard interrupt. The code in the function itself, such as
> invoke_softirq(), is aware of this and does not rely on the counter.
> Everything else which derives the context from preempt_count gets it
> wrong in that window:
>
> - ftrace, perf and the ring buffer record task context and use the
> task recursion and context slots.
> - KCSAN attributes the accesses to the interrupted task, KMSAN uses
> and changes its state. KCOV and the printk caller id see a task.
> - On PREEMPT_RT can_spin_trylock() and local_trylock() reject hard
> interrupt context to avoid interfering with PI when the interrupted
> task is blocked on a lock. That check does not reject calls made in
> this window. BPF programs attached to sched_waking or sched_wakeup
> can reach it through kmalloc_nolock().
> - An oops kills the interrupted task instead of ending in "Fatal
> exception in interrupt".
>
> Tracing and the sanitizers see the wrong context in this window. No
> failure caused by this misclassification is known. The early removal of
> HARDIRQ_OFFSET predates git. lockdep is not affected because
> lockdep_hardirq_exit() is the last operation in irq_exit().
>
> Keep HARDIRQ_OFFSET until right before tick_irq_exit(), which needs
> in_hardirq() to be false for the outermost interrupt. Softirq handlers
> must not run with HARDIRQ_OFFSET set, so softirq_handle_begin() replaces
> it with SOFTIRQ_OFFSET and softirq_handle_end() reverts that, each in a
> single raw preempt_count update. The raw operations keep the preemption
> disable location recorded by irq_enter_rcu(), and lockdep is updated by
> hand. softirq_handle_begin() detects the case with in_hardirq() because
> __do_softirq() is reached through the stack switch in
> do_softirq_own_stack() and cannot take an argument.
>
> The checks run before HARDIRQ_OFFSET is removed. !in_interrupt() becomes
> irq_count() == HARDIRQ_OFFSET, as in irq_enter_rcu(). The timer thread
> check becomes (in_nmi() | hardirq_count()) == HARDIRQ_OFFSET. It does
> not test softirq_count(): the timer thread must also wake when the
> interrupt hit softirq processing or a section with BHs disabled.
> A softirq raised in the timer thread wakeup is handled by the timer
> thread, which handles all pending softirqs.
>
> A softirq raised from a tracepoint on the final preempt_count_sub()
> waits for the next interrupt exit and can trigger NOHZ tick-stop
> warnings meanwhile. That is not the normal path and does not justify a
> check on every interrupt exit. A tracepoint on tick_irq_exit() already
> behaves the same way.
>
> The number of preempt_count updates and the interrupt time accounting
> are unchanged. The preemptoff tracer now reports the interrupt and the
> softirq processing on top of it as one section, and function graph with
> nofuncgraph-irqs also skips the interrupt exit work, including the
> __do_softirq() frame.
>
> Suggested-by: Peter Zijlstra <peterz@infradead.org>
> Link: https://lore.kernel.org/r/20260813130826.GW687043@noisy.programming.kicks-ass.net
> Assisted-by: LLM
> Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com>
> Reviewed-by: Bradley Morgan <brads@mainlining.org>
> Reviewed-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
> ---
>
> Notes:
> Changes in v5:
> - Timer thread wakeup: merge the NMI test into the hardirq test,
> (in_nmi() | hardirq_count()) == HARDIRQ_OFFSET (Sebastian).
> - Add Sebastian's Reviewed-by, given on v4.
> - Rebase on v7.3-rc5. The two touched files are unchanged since rc4.
>
> Testing: v5 differs from v4 by that one expression. Both forms agree
> for all 2^32 preempt_count values and at the real site on every IRQ
> exit in four QEMU boots (arm64, arm32; plain and threadirqs; 1.1M
> evaluations). gcc 15 emits one conditional branch less on x86-64,
> arm64 and arm32; clang 22 on arm64, and 16 bytes less on x86-64. On
> v7.3-rc5, base against v5 in QEMU on arm64 (virt, SMP, lockdep) and
> arm32 (versatilepb, lockdep): boot and stress, nothing on v5 that the
> base does not show.
>
> v4: https://lore.kernel.org/r/20260926143505.66024-1-kmehltretter@gmail.com
>
> Documentation/core-api/entry.rst | 16 +++++---
> kernel/softirq.c | 65 ++++++++++++++++++++++++++------
> 2 files changed, 63 insertions(+), 18 deletions(-)
>
> diff --git a/Documentation/core-api/entry.rst b/Documentation/core-api/entry.rst
> index 79fdaed954d9..ff3df997b151 100644
> --- a/Documentation/core-api/entry.rst
> +++ b/Documentation/core-api/entry.rst
> @@ -197,8 +197,9 @@ return true, handles NOHZ tick state and interrupt time accounting. This
> means that up to the point where irq_enter_rcu() is invoked in_hardirq()
> returns false.
>
> -irq_exit_rcu() handles interrupt time accounting, undoes the preemption
> -count update and eventually handles soft interrupts and NOHZ tick state.
> +irq_exit_rcu() handles interrupt time accounting, handles soft interrupts if
> +possible, undoes the preemption count update and finally handles the NOHZ tick
> +state.
>
> In theory, the preemption count could be updated in irqentry_enter(). In
> practice, deferring this update to irq_enter_rcu() allows the preemption-count
> @@ -207,10 +208,13 @@ irqentry_exit(), which are described in the next paragraph. The only downside
> is that the early entry code up to irq_enter_rcu() must be aware that the
> preemption count has not yet been updated with the HARDIRQ_OFFSET state.
>
> -Note that irq_exit_rcu() must remove HARDIRQ_OFFSET from the preemption count
> -before it handles soft interrupts, whose handlers must run in BH context rather
> -than irq-disabled context. In addition, irqentry_exit() might schedule, which
> -also requires that HARDIRQ_OFFSET has been removed from the preemption count.
> +Note that soft interrupt handlers must run in BH context rather than in hard
> +interrupt context. irq_exit_rcu() therefore replaces HARDIRQ_OFFSET with
> +SOFTIRQ_OFFSET in the preemption count while it handles soft interrupts and
> +puts HARDIRQ_OFFSET back afterwards, so that the remaining interrupt exit work
> +is still attributed to the interrupt. HARDIRQ_OFFSET is removed before
> +irq_exit_rcu() returns because irqentry_exit() might schedule, which requires
> +that HARDIRQ_OFFSET has been removed from the preemption count.
>
> Even though interrupt handlers are expected to run with local interrupts
> disabled, interrupt nesting is common from an entry/exit perspective. For
> diff --git a/kernel/softirq.c b/kernel/softirq.c
> index 5d02c36c40e3..288e9e37b806 100644
> --- a/kernel/softirq.c
> +++ b/kernel/softirq.c
> @@ -350,8 +350,8 @@ static inline void ksoftirqd_run_end(void)
> local_irq_enable();
> }
>
> -static inline void softirq_handle_begin(void) { }
> -static inline void softirq_handle_end(void) { }
> +static inline bool softirq_handle_begin(void) { return false; }
> +static inline void softirq_handle_end(bool from_irq_exit) { }
>
> static inline bool should_wake_ksoftirqd(void)
> {
> @@ -481,15 +481,40 @@ void __local_bh_enable_ip(unsigned long ip, unsigned int cnt)
> }
> EXPORT_SYMBOL(__local_bh_enable_ip);
>
> -static inline void softirq_handle_begin(void)
> +static inline bool softirq_handle_begin(void)
> {
> - __local_bh_disable_ip(_RET_IP_, SOFTIRQ_OFFSET);
> + bool from_irq_exit = in_hardirq();
> +
> + if (!from_irq_exit) {
> + __local_bh_disable_ip(_RET_IP_, SOFTIRQ_OFFSET);
> + return false;
> + }
> +
> + /*
> + * Only reached from irq_exit(), with HARDIRQ_OFFSET still set.
> + * Replace it with SOFTIRQ_OFFSET before handle_softirqs() enables
> + * interrupts. Use the raw operation to preserve the preemption
> + * disable location recorded by irq_enter_rcu(), and update lockdep
> + * directly.
> + */
> + __preempt_count_sub(HARDIRQ_OFFSET - SOFTIRQ_OFFSET);
> + lockdep_softirqs_off(_RET_IP_);
> + WARN_ON_ONCE(irq_count() != SOFTIRQ_OFFSET);
> +
> + return true;
> }
>
> -static inline void softirq_handle_end(void)
> +static inline void softirq_handle_end(bool from_irq_exit)
> {
> - __local_bh_enable(SOFTIRQ_OFFSET);
> - WARN_ON_ONCE(in_interrupt());
> + if (!from_irq_exit) {
> + __local_bh_enable(SOFTIRQ_OFFSET);
> + WARN_ON_ONCE(in_interrupt());
> + return;
> + }
> +
> + lockdep_softirqs_on(_RET_IP_);
> + __preempt_count_add(HARDIRQ_OFFSET - SOFTIRQ_OFFSET);
> + WARN_ON_ONCE(irq_count() != HARDIRQ_OFFSET);
> }
>
> static inline void ksoftirqd_run_begin(void)
> @@ -605,6 +630,7 @@ static void handle_softirqs(bool ksirqd)
> unsigned long old_flags = current->flags;
> int max_restart = MAX_SOFTIRQ_RESTART;
> struct softirq_action *h;
> + bool from_irq_exit;
> bool in_hardirq;
> __u32 pending;
> int softirq_bit;
> @@ -618,7 +644,7 @@ static void handle_softirqs(bool ksirqd)
>
> pending = local_softirq_pending();
>
> - softirq_handle_begin();
> + from_irq_exit = softirq_handle_begin();
> in_hardirq = lockdep_softirq_start();
> account_softirq_enter(current);
>
> @@ -670,7 +696,7 @@ static void handle_softirqs(bool ksirqd)
>
> account_softirq_exit(current);
> lockdep_softirq_end(in_hardirq);
> - softirq_handle_end();
> + softirq_handle_end(from_irq_exit);
> current_restore_flags(old_flags, PF_MEMALLOC);
> }
>
> @@ -748,8 +774,12 @@ static inline void __irq_exit_rcu(void)
> lockdep_assert_irqs_disabled();
> #endif
> account_hardirq_exit(current);
> - preempt_count_sub(HARDIRQ_OFFSET);
> - if (!in_interrupt() && local_softirq_pending()) {
> +
> + /*
> + * HARDIRQ_OFFSET is still set. Only the outermost interrupt handles
> + * softirqs, and only if it did not hit a softirq or BH disabled section.
> + */
> + if (irq_count() == HARDIRQ_OFFSET && local_softirq_pending()) {
> /*
> * If we left hrtimers unarmed, make sure to arm them now,
> * before enabling interrupts to run softirq.
> @@ -758,10 +788,21 @@ static inline void __irq_exit_rcu(void)
> invoke_softirq();
> }
>
> + /*
> + * Wake the timer thread even if the interrupt hit a softirq or a
> + * section with BHs disabled. Only nested interrupts and NMIs are
> + * excluded.
> + */
> if (IS_ENABLED(CONFIG_IRQ_FORCED_THREADING) && force_irqthreads() &&
> - local_timers_pending_force_th() && !(in_nmi() | in_hardirq()))
> + local_timers_pending_force_th() &&
> + (in_nmi() | hardirq_count()) == HARDIRQ_OFFSET)
Heh, I reviewed the v4 and started typing the same comment as Sebastian's only
to realize its already changed to what I was also about to suggest, so great! :)
Reviewed-by: Joel Fernandes <joelagnelf@nvidia.com>
thanks,
--
Joel Fernandes
prev parent reply other threads:[~2026-10-03 0:50 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-30 19:14 Karl Mehltretter
2026-10-01 7:39 ` Sebastian Andrzej Siewior
2026-10-03 0:50 ` Joel Fernandes [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=fe301e2c-504b-42dd-954e-a133a54aa0ce@nvidia.com \
--to=joelagnelf@nvidia.com \
--cc=bigeasy@linutronix.de \
--cc=boqun@kernel.org \
--cc=brads@mainlining.org \
--cc=clrkwllms@kernel.org \
--cc=corbet@lwn.net \
--cc=elver@google.com \
--cc=frederic@kernel.org \
--cc=glider@google.com \
--cc=kmehltretter@gmail.com \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-rt-devel@lists.linux.dev \
--cc=lyude@redhat.com \
--cc=peterz@infradead.org \
--cc=rostedt@goodmis.org \
--cc=tglx@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®