mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Joel Fernandes <joelagnelf@nvidia.com>
To: Karl Mehltretter <kmehltretter@gmail.com>,
	Peter Zijlstra <peterz@infradead.org>,
	Thomas Gleixner <tglx@kernel.org>
Cc: Sebastian Andrzej Siewior <bigeasy@linutronix.de>,
	Frederic Weisbecker <frederic@kernel.org>,
	Clark Williams <clrkwllms@kernel.org>,
	Steven Rostedt <rostedt@goodmis.org>,
	Boqun Feng <boqun@kernel.org>, Lyude Paul <lyude@redhat.com>,
	Alexander Potapenko <glider@google.com>,
	Marco Elver <elver@google.com>, Jonathan Corbet <corbet@lwn.net>,
	Bradley Morgan <brads@mainlining.org>,
	linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org,
	linux-rt-devel@lists.linux.dev
Subject: Re: [PATCH v5] softirq: Preserve interrupt context during IRQ exit
Date: Fri, 2 Oct 2026 20:50:39 -0400	[thread overview]
Message-ID: <fe301e2c-504b-42dd-954e-a133a54aa0ce@nvidia.com> (raw)
In-Reply-To: <20260930191432.62760-1-kmehltretter@gmail.com>



On 9/30/2026 3:14 PM, Karl Mehltretter wrote:
> On the return from interrupt path, __irq_exit_rcu() removes
> HARDIRQ_OFFSET from the preemption counter at the very top of the
> function. Everything after that reports the current context as task
> instead of hard interrupt. The code in the function itself, such as
> invoke_softirq(), is aware of this and does not rely on the counter.
> Everything else which derives the context from preempt_count gets it
> wrong in that window:
> 
>   - ftrace, perf and the ring buffer record task context and use the
>     task recursion and context slots.
>   - KCSAN attributes the accesses to the interrupted task, KMSAN uses
>     and changes its state. KCOV and the printk caller id see a task.
>   - On PREEMPT_RT can_spin_trylock() and local_trylock() reject hard
>     interrupt context to avoid interfering with PI when the interrupted
>     task is blocked on a lock. That check does not reject calls made in
>     this window. BPF programs attached to sched_waking or sched_wakeup
>     can reach it through kmalloc_nolock().
>   - An oops kills the interrupted task instead of ending in "Fatal
>     exception in interrupt".
> 
> Tracing and the sanitizers see the wrong context in this window. No
> failure caused by this misclassification is known. The early removal of
> HARDIRQ_OFFSET predates git. lockdep is not affected because
> lockdep_hardirq_exit() is the last operation in irq_exit().
> 
> Keep HARDIRQ_OFFSET until right before tick_irq_exit(), which needs
> in_hardirq() to be false for the outermost interrupt. Softirq handlers
> must not run with HARDIRQ_OFFSET set, so softirq_handle_begin() replaces
> it with SOFTIRQ_OFFSET and softirq_handle_end() reverts that, each in a
> single raw preempt_count update. The raw operations keep the preemption
> disable location recorded by irq_enter_rcu(), and lockdep is updated by
> hand. softirq_handle_begin() detects the case with in_hardirq() because
> __do_softirq() is reached through the stack switch in
> do_softirq_own_stack() and cannot take an argument.
> 
> The checks run before HARDIRQ_OFFSET is removed. !in_interrupt() becomes
> irq_count() == HARDIRQ_OFFSET, as in irq_enter_rcu(). The timer thread
> check becomes (in_nmi() | hardirq_count()) == HARDIRQ_OFFSET. It does
> not test softirq_count(): the timer thread must also wake when the
> interrupt hit softirq processing or a section with BHs disabled.
> A softirq raised in the timer thread wakeup is handled by the timer
> thread, which handles all pending softirqs.
> 
> A softirq raised from a tracepoint on the final preempt_count_sub()
> waits for the next interrupt exit and can trigger NOHZ tick-stop
> warnings meanwhile. That is not the normal path and does not justify a
> check on every interrupt exit. A tracepoint on tick_irq_exit() already
> behaves the same way.
> 
> The number of preempt_count updates and the interrupt time accounting
> are unchanged. The preemptoff tracer now reports the interrupt and the
> softirq processing on top of it as one section, and function graph with
> nofuncgraph-irqs also skips the interrupt exit work, including the
> __do_softirq() frame.
> 
> Suggested-by: Peter Zijlstra <peterz@infradead.org>
> Link: https://lore.kernel.org/r/20260813130826.GW687043@noisy.programming.kicks-ass.net
> Assisted-by: LLM
> Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com>
> Reviewed-by: Bradley Morgan <brads@mainlining.org>
> Reviewed-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
> ---
> 
> Notes:
>     Changes in v5:
>     - Timer thread wakeup: merge the NMI test into the hardirq test,
>       (in_nmi() | hardirq_count()) == HARDIRQ_OFFSET (Sebastian).
>     - Add Sebastian's Reviewed-by, given on v4.
>     - Rebase on v7.3-rc5. The two touched files are unchanged since rc4.
>     
>     Testing: v5 differs from v4 by that one expression. Both forms agree
>     for all 2^32 preempt_count values and at the real site on every IRQ
>     exit in four QEMU boots (arm64, arm32; plain and threadirqs; 1.1M
>     evaluations). gcc 15 emits one conditional branch less on x86-64,
>     arm64 and arm32; clang 22 on arm64, and 16 bytes less on x86-64. On
>     v7.3-rc5, base against v5 in QEMU on arm64 (virt, SMP, lockdep) and
>     arm32 (versatilepb, lockdep): boot and stress, nothing on v5 that the
>     base does not show.
>     
>     v4: https://lore.kernel.org/r/20260926143505.66024-1-kmehltretter@gmail.com
> 
>  Documentation/core-api/entry.rst | 16 +++++---
>  kernel/softirq.c                 | 65 ++++++++++++++++++++++++++------
>  2 files changed, 63 insertions(+), 18 deletions(-)
> 
> diff --git a/Documentation/core-api/entry.rst b/Documentation/core-api/entry.rst
> index 79fdaed954d9..ff3df997b151 100644
> --- a/Documentation/core-api/entry.rst
> +++ b/Documentation/core-api/entry.rst
> @@ -197,8 +197,9 @@ return true, handles NOHZ tick state and interrupt time accounting. This
>  means that up to the point where irq_enter_rcu() is invoked in_hardirq()
>  returns false.
>  
> -irq_exit_rcu() handles interrupt time accounting, undoes the preemption
> -count update and eventually handles soft interrupts and NOHZ tick state.
> +irq_exit_rcu() handles interrupt time accounting, handles soft interrupts if
> +possible, undoes the preemption count update and finally handles the NOHZ tick
> +state.
>  
>  In theory, the preemption count could be updated in irqentry_enter(). In
>  practice, deferring this update to irq_enter_rcu() allows the preemption-count
> @@ -207,10 +208,13 @@ irqentry_exit(), which are described in the next paragraph. The only downside
>  is that the early entry code up to irq_enter_rcu() must be aware that the
>  preemption count has not yet been updated with the HARDIRQ_OFFSET state.
>  
> -Note that irq_exit_rcu() must remove HARDIRQ_OFFSET from the preemption count
> -before it handles soft interrupts, whose handlers must run in BH context rather
> -than irq-disabled context. In addition, irqentry_exit() might schedule, which
> -also requires that HARDIRQ_OFFSET has been removed from the preemption count.
> +Note that soft interrupt handlers must run in BH context rather than in hard
> +interrupt context. irq_exit_rcu() therefore replaces HARDIRQ_OFFSET with
> +SOFTIRQ_OFFSET in the preemption count while it handles soft interrupts and
> +puts HARDIRQ_OFFSET back afterwards, so that the remaining interrupt exit work
> +is still attributed to the interrupt. HARDIRQ_OFFSET is removed before
> +irq_exit_rcu() returns because irqentry_exit() might schedule, which requires
> +that HARDIRQ_OFFSET has been removed from the preemption count.
>  
>  Even though interrupt handlers are expected to run with local interrupts
>  disabled, interrupt nesting is common from an entry/exit perspective. For
> diff --git a/kernel/softirq.c b/kernel/softirq.c
> index 5d02c36c40e3..288e9e37b806 100644
> --- a/kernel/softirq.c
> +++ b/kernel/softirq.c
> @@ -350,8 +350,8 @@ static inline void ksoftirqd_run_end(void)
>  	local_irq_enable();
>  }
>  
> -static inline void softirq_handle_begin(void) { }
> -static inline void softirq_handle_end(void) { }
> +static inline bool softirq_handle_begin(void) { return false; }
> +static inline void softirq_handle_end(bool from_irq_exit) { }
>  
>  static inline bool should_wake_ksoftirqd(void)
>  {
> @@ -481,15 +481,40 @@ void __local_bh_enable_ip(unsigned long ip, unsigned int cnt)
>  }
>  EXPORT_SYMBOL(__local_bh_enable_ip);
>  
> -static inline void softirq_handle_begin(void)
> +static inline bool softirq_handle_begin(void)
>  {
> -	__local_bh_disable_ip(_RET_IP_, SOFTIRQ_OFFSET);
> +	bool from_irq_exit = in_hardirq();
> +
> +	if (!from_irq_exit) {
> +		__local_bh_disable_ip(_RET_IP_, SOFTIRQ_OFFSET);
> +		return false;
> +	}
> +
> +	/*
> +	 * Only reached from irq_exit(), with HARDIRQ_OFFSET still set.
> +	 * Replace it with SOFTIRQ_OFFSET before handle_softirqs() enables
> +	 * interrupts. Use the raw operation to preserve the preemption
> +	 * disable location recorded by irq_enter_rcu(), and update lockdep
> +	 * directly.
> +	 */
> +	__preempt_count_sub(HARDIRQ_OFFSET - SOFTIRQ_OFFSET);
> +	lockdep_softirqs_off(_RET_IP_);
> +	WARN_ON_ONCE(irq_count() != SOFTIRQ_OFFSET);
> +
> +	return true;
>  }
>  
> -static inline void softirq_handle_end(void)
> +static inline void softirq_handle_end(bool from_irq_exit)
>  {
> -	__local_bh_enable(SOFTIRQ_OFFSET);
> -	WARN_ON_ONCE(in_interrupt());
> +	if (!from_irq_exit) {
> +		__local_bh_enable(SOFTIRQ_OFFSET);
> +		WARN_ON_ONCE(in_interrupt());
> +		return;
> +	}
> +
> +	lockdep_softirqs_on(_RET_IP_);
> +	__preempt_count_add(HARDIRQ_OFFSET - SOFTIRQ_OFFSET);
> +	WARN_ON_ONCE(irq_count() != HARDIRQ_OFFSET);
>  }
>  
>  static inline void ksoftirqd_run_begin(void)
> @@ -605,6 +630,7 @@ static void handle_softirqs(bool ksirqd)
>  	unsigned long old_flags = current->flags;
>  	int max_restart = MAX_SOFTIRQ_RESTART;
>  	struct softirq_action *h;
> +	bool from_irq_exit;
>  	bool in_hardirq;
>  	__u32 pending;
>  	int softirq_bit;
> @@ -618,7 +644,7 @@ static void handle_softirqs(bool ksirqd)
>  
>  	pending = local_softirq_pending();
>  
> -	softirq_handle_begin();
> +	from_irq_exit = softirq_handle_begin();
>  	in_hardirq = lockdep_softirq_start();
>  	account_softirq_enter(current);
>  
> @@ -670,7 +696,7 @@ static void handle_softirqs(bool ksirqd)
>  
>  	account_softirq_exit(current);
>  	lockdep_softirq_end(in_hardirq);
> -	softirq_handle_end();
> +	softirq_handle_end(from_irq_exit);
>  	current_restore_flags(old_flags, PF_MEMALLOC);
>  }
>  
> @@ -748,8 +774,12 @@ static inline void __irq_exit_rcu(void)
>  	lockdep_assert_irqs_disabled();
>  #endif
>  	account_hardirq_exit(current);
> -	preempt_count_sub(HARDIRQ_OFFSET);
> -	if (!in_interrupt() && local_softirq_pending()) {
> +
> +	/*
> +	 * HARDIRQ_OFFSET is still set. Only the outermost interrupt handles
> +	 * softirqs, and only if it did not hit a softirq or BH disabled section.
> +	 */
> +	if (irq_count() == HARDIRQ_OFFSET && local_softirq_pending()) {
>  		/*
>  		 * If we left hrtimers unarmed, make sure to arm them now,
>  		 * before enabling interrupts to run softirq.
> @@ -758,10 +788,21 @@ static inline void __irq_exit_rcu(void)
>  		invoke_softirq();
>  	}
>  
> +	/*
> +	 * Wake the timer thread even if the interrupt hit a softirq or a
> +	 * section with BHs disabled. Only nested interrupts and NMIs are
> +	 * excluded.
> +	 */
>  	if (IS_ENABLED(CONFIG_IRQ_FORCED_THREADING) && force_irqthreads() &&
> -	    local_timers_pending_force_th() && !(in_nmi() | in_hardirq()))
> +	    local_timers_pending_force_th() &&
> +	    (in_nmi() | hardirq_count()) == HARDIRQ_OFFSET)

Heh, I reviewed the v4 and started typing the same comment as Sebastian's only
to realize its already changed to what I was also about to suggest, so great! :)

Reviewed-by: Joel Fernandes <joelagnelf@nvidia.com>

thanks,
-- 
Joel Fernandes


      parent reply	other threads:[~2026-10-03  0:50 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-30 19:14 Karl Mehltretter
2026-10-01  7:39 ` Sebastian Andrzej Siewior
2026-10-03  0:50 ` Joel Fernandes [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=fe301e2c-504b-42dd-954e-a133a54aa0ce@nvidia.com \
    --to=joelagnelf@nvidia.com \
    --cc=bigeasy@linutronix.de \
    --cc=boqun@kernel.org \
    --cc=brads@mainlining.org \
    --cc=clrkwllms@kernel.org \
    --cc=corbet@lwn.net \
    --cc=elver@google.com \
    --cc=frederic@kernel.org \
    --cc=glider@google.com \
    --cc=kmehltretter@gmail.com \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-rt-devel@lists.linux.dev \
    --cc=lyude@redhat.com \
    --cc=peterz@infradead.org \
    --cc=rostedt@goodmis.org \
    --cc=tglx@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®