Re: [PATCH v5] softirq: Preserve interrupt context during IRQ exit
From: Joel Fernandes
Date: Fri Oct 02 2026 - 20:51:07 EST
On 9/30/2026 3:14 PM, Karl Mehltretter wrote:
> On the return from interrupt path, __irq_exit_rcu() removes
> HARDIRQ_OFFSET from the preemption counter at the very top of the
> function. Everything after that reports the current context as task
> instead of hard interrupt. The code in the function itself, such as
> invoke_softirq(), is aware of this and does not rely on the counter.
> Everything else which derives the context from preempt_count gets it
> wrong in that window:
>
> - ftrace, perf and the ring buffer record task context and use the
> task recursion and context slots.
> - KCSAN attributes the accesses to the interrupted task, KMSAN uses
> and changes its state. KCOV and the printk caller id see a task.
> - On PREEMPT_RT can_spin_trylock() and local_trylock() reject hard
> interrupt context to avoid interfering with PI when the interrupted
> task is blocked on a lock. That check does not reject calls made in
> this window. BPF programs attached to sched_waking or sched_wakeup
> can reach it through kmalloc_nolock().
> - An oops kills the interrupted task instead of ending in "Fatal
> exception in interrupt".
>
> Tracing and the sanitizers see the wrong context in this window. No
> failure caused by this misclassification is known. The early removal of
> HARDIRQ_OFFSET predates git. lockdep is not affected because
> lockdep_hardirq_exit() is the last operation in irq_exit().
>
> Keep HARDIRQ_OFFSET until right before tick_irq_exit(), which needs
> in_hardirq() to be false for the outermost interrupt. Softirq handlers
> must not run with HARDIRQ_OFFSET set, so softirq_handle_begin() replaces
> it with SOFTIRQ_OFFSET and softirq_handle_end() reverts that, each in a
> single raw preempt_count update. The raw operations keep the preemption
> disable location recorded by irq_enter_rcu(), and lockdep is updated by
> hand. softirq_handle_begin() detects the case with in_hardirq() because
> __do_softirq() is reached through the stack switch in
> do_softirq_own_stack() and cannot take an argument.
>
> The checks run before HARDIRQ_OFFSET is removed. !in_interrupt() becomes
> irq_count() == HARDIRQ_OFFSET, as in irq_enter_rcu(). The timer thread
> check becomes (in_nmi() | hardirq_count()) == HARDIRQ_OFFSET. It does
> not test softirq_count(): the timer thread must also wake when the
> interrupt hit softirq processing or a section with BHs disabled.
> A softirq raised in the timer thread wakeup is handled by the timer
> thread, which handles all pending softirqs.
>
> A softirq raised from a tracepoint on the final preempt_count_sub()
> waits for the next interrupt exit and can trigger NOHZ tick-stop
> warnings meanwhile. That is not the normal path and does not justify a
> check on every interrupt exit. A tracepoint on tick_irq_exit() already
> behaves the same way.
>
> The number of preempt_count updates and the interrupt time accounting
> are unchanged. The preemptoff tracer now reports the interrupt and the
> softirq processing on top of it as one section, and function graph with
> nofuncgraph-irqs also skips the interrupt exit work, including the
> __do_softirq() frame.
>
> Suggested-by: Peter Zijlstra <peterz@xxxxxxxxxxxxx>
> Link: https://lore.kernel.org/r/20260813130826.GW687043@xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
> Assisted-by: LLM
> Signed-off-by: Karl Mehltretter <kmehltretter@xxxxxxxxx>
> Reviewed-by: Bradley Morgan <brads@xxxxxxxxxxxxxx>
> Reviewed-by: Sebastian Andrzej Siewior <bigeasy@xxxxxxxxxxxxx>
> ---
>
> Notes:
> Changes in v5:
> - Timer thread wakeup: merge the NMI test into the hardirq test,
> (in_nmi() | hardirq_count()) == HARDIRQ_OFFSET (Sebastian).
> - Add Sebastian's Reviewed-by, given on v4.
> - Rebase on v7.3-rc5. The two touched files are unchanged since rc4.
>
> Testing: v5 differs from v4 by that one expression. Both forms agree
> for all 2^32 preempt_count values and at the real site on every IRQ
> exit in four QEMU boots (arm64, arm32; plain and threadirqs; 1.1M
> evaluations). gcc 15 emits one conditional branch less on x86-64,
> arm64 and arm32; clang 22 on arm64, and 16 bytes less on x86-64. On
> v7.3-rc5, base against v5 in QEMU on arm64 (virt, SMP, lockdep) and
> arm32 (versatilepb, lockdep): boot and stress, nothing on v5 that the
> base does not show.
>
> v4: https://lore.kernel.org/r/20260926143505.66024-1-kmehltretter@xxxxxxxxx
>
> Documentation/core-api/entry.rst | 16 +++++---
> kernel/softirq.c | 65 ++++++++++++++++++++++++++------
> 2 files changed, 63 insertions(+), 18 deletions(-)
>
> diff --git a/Documentation/core-api/entry.rst b/Documentation/core-api/entry.rst
> index 79fdaed954d9..ff3df997b151 100644
> --- a/Documentation/core-api/entry.rst
> +++ b/Documentation/core-api/entry.rst
> @@ -197,8 +197,9 @@ return true, handles NOHZ tick state and interrupt time accounting. This
> means that up to the point where irq_enter_rcu() is invoked in_hardirq()
> returns false.
>
> -irq_exit_rcu() handles interrupt time accounting, undoes the preemption
> -count update and eventually handles soft interrupts and NOHZ tick state.
> +irq_exit_rcu() handles interrupt time accounting, handles soft interrupts if
> +possible, undoes the preemption count update and finally handles the NOHZ tick
> +state.
>
> In theory, the preemption count could be updated in irqentry_enter(). In
> practice, deferring this update to irq_enter_rcu() allows the preemption-count
> @@ -207,10 +208,13 @@ irqentry_exit(), which are described in the next paragraph. The only downside
> is that the early entry code up to irq_enter_rcu() must be aware that the
> preemption count has not yet been updated with the HARDIRQ_OFFSET state.
>
> -Note that irq_exit_rcu() must remove HARDIRQ_OFFSET from the preemption count
> -before it handles soft interrupts, whose handlers must run in BH context rather
> -than irq-disabled context. In addition, irqentry_exit() might schedule, which
> -also requires that HARDIRQ_OFFSET has been removed from the preemption count.
> +Note that soft interrupt handlers must run in BH context rather than in hard
> +interrupt context. irq_exit_rcu() therefore replaces HARDIRQ_OFFSET with
> +SOFTIRQ_OFFSET in the preemption count while it handles soft interrupts and
> +puts HARDIRQ_OFFSET back afterwards, so that the remaining interrupt exit work
> +is still attributed to the interrupt. HARDIRQ_OFFSET is removed before
> +irq_exit_rcu() returns because irqentry_exit() might schedule, which requires
> +that HARDIRQ_OFFSET has been removed from the preemption count.
>
> Even though interrupt handlers are expected to run with local interrupts
> disabled, interrupt nesting is common from an entry/exit perspective. For
> diff --git a/kernel/softirq.c b/kernel/softirq.c
> index 5d02c36c40e3..288e9e37b806 100644
> --- a/kernel/softirq.c
> +++ b/kernel/softirq.c
> @@ -350,8 +350,8 @@ static inline void ksoftirqd_run_end(void)
> local_irq_enable();
> }
>
> -static inline void softirq_handle_begin(void) { }
> -static inline void softirq_handle_end(void) { }
> +static inline bool softirq_handle_begin(void) { return false; }
> +static inline void softirq_handle_end(bool from_irq_exit) { }
>
> static inline bool should_wake_ksoftirqd(void)
> {
> @@ -481,15 +481,40 @@ void __local_bh_enable_ip(unsigned long ip, unsigned int cnt)
> }
> EXPORT_SYMBOL(__local_bh_enable_ip);
>
> -static inline void softirq_handle_begin(void)
> +static inline bool softirq_handle_begin(void)
> {
> - __local_bh_disable_ip(_RET_IP_, SOFTIRQ_OFFSET);
> + bool from_irq_exit = in_hardirq();
> +
> + if (!from_irq_exit) {
> + __local_bh_disable_ip(_RET_IP_, SOFTIRQ_OFFSET);
> + return false;
> + }
> +
> + /*
> + * Only reached from irq_exit(), with HARDIRQ_OFFSET still set.
> + * Replace it with SOFTIRQ_OFFSET before handle_softirqs() enables
> + * interrupts. Use the raw operation to preserve the preemption
> + * disable location recorded by irq_enter_rcu(), and update lockdep
> + * directly.
> + */
> + __preempt_count_sub(HARDIRQ_OFFSET - SOFTIRQ_OFFSET);
> + lockdep_softirqs_off(_RET_IP_);
> + WARN_ON_ONCE(irq_count() != SOFTIRQ_OFFSET);
> +
> + return true;
> }
>
> -static inline void softirq_handle_end(void)
> +static inline void softirq_handle_end(bool from_irq_exit)
> {
> - __local_bh_enable(SOFTIRQ_OFFSET);
> - WARN_ON_ONCE(in_interrupt());
> + if (!from_irq_exit) {
> + __local_bh_enable(SOFTIRQ_OFFSET);
> + WARN_ON_ONCE(in_interrupt());
> + return;
> + }
> +
> + lockdep_softirqs_on(_RET_IP_);
> + __preempt_count_add(HARDIRQ_OFFSET - SOFTIRQ_OFFSET);
> + WARN_ON_ONCE(irq_count() != HARDIRQ_OFFSET);
> }
>
> static inline void ksoftirqd_run_begin(void)
> @@ -605,6 +630,7 @@ static void handle_softirqs(bool ksirqd)
> unsigned long old_flags = current->flags;
> int max_restart = MAX_SOFTIRQ_RESTART;
> struct softirq_action *h;
> + bool from_irq_exit;
> bool in_hardirq;
> __u32 pending;
> int softirq_bit;
> @@ -618,7 +644,7 @@ static void handle_softirqs(bool ksirqd)
>
> pending = local_softirq_pending();
>
> - softirq_handle_begin();
> + from_irq_exit = softirq_handle_begin();
> in_hardirq = lockdep_softirq_start();
> account_softirq_enter(current);
>
> @@ -670,7 +696,7 @@ static void handle_softirqs(bool ksirqd)
>
> account_softirq_exit(current);
> lockdep_softirq_end(in_hardirq);
> - softirq_handle_end();
> + softirq_handle_end(from_irq_exit);
> current_restore_flags(old_flags, PF_MEMALLOC);
> }
>
> @@ -748,8 +774,12 @@ static inline void __irq_exit_rcu(void)
> lockdep_assert_irqs_disabled();
> #endif
> account_hardirq_exit(current);
> - preempt_count_sub(HARDIRQ_OFFSET);
> - if (!in_interrupt() && local_softirq_pending()) {
> +
> + /*
> + * HARDIRQ_OFFSET is still set. Only the outermost interrupt handles
> + * softirqs, and only if it did not hit a softirq or BH disabled section.
> + */
> + if (irq_count() == HARDIRQ_OFFSET && local_softirq_pending()) {
> /*
> * If we left hrtimers unarmed, make sure to arm them now,
> * before enabling interrupts to run softirq.
> @@ -758,10 +788,21 @@ static inline void __irq_exit_rcu(void)
> invoke_softirq();
> }
>
> + /*
> + * Wake the timer thread even if the interrupt hit a softirq or a
> + * section with BHs disabled. Only nested interrupts and NMIs are
> + * excluded.
> + */
> if (IS_ENABLED(CONFIG_IRQ_FORCED_THREADING) && force_irqthreads() &&
> - local_timers_pending_force_th() && !(in_nmi() | in_hardirq()))
> + local_timers_pending_force_th() &&
> + (in_nmi() | hardirq_count()) == HARDIRQ_OFFSET)
Heh, I reviewed the v4 and started typing the same comment as Sebastian's only
to realize its already changed to what I was also about to suggest, so great! :)
Reviewed-by: Joel Fernandes <joelagnelf@xxxxxxxxxx>
thanks,
--
Joel Fernandes