Re: [PATCH v4 05/17] irq & spin_lock: Add counted interrupt disabling/enabling

From: Boqun Feng

Date: Wed Aug 05 2026 - 13:40:48 EST


On Wed, Aug 05, 2026 at 10:23:20PM +0530, Shrikanth Hegde wrote:
> Hi Boqun.
>
> On 8/5/26 8:41 PM, Boqun Feng wrote:
> > On Wed, Aug 05, 2026 at 08:26:49PM +0530, Shrikanth Hegde wrote:
> > >
> > >
> > > On 8/5/26 7:50 PM, Boqun Feng wrote:
> > > > On Wed, Aug 05, 2026 at 07:40:18PM +0530, Shrikanth Hegde wrote:
> > > > >
> > > > > Hi Boqun,
> > > > >
> > > > > >
> > > > > > Something as below? Going to send it to kernel build bot and see if it
> > > > > > works for all configs.
> > > > > >
> > > > > > ----------------->8
> > > > > > diff --git a/include/linux/interrupt_rc.h b/include/linux/interrupt_rc.h
> > > > > > index b9a7f05ecf42..39f30bc65548 100644
> > > > > > --- a/include/linux/interrupt_rc.h
> > > > > > +++ b/include/linux/interrupt_rc.h
> > > > > > @@ -12,6 +12,7 @@
> > > > > > */
> > > > > >
> > > > > > #include <linux/irqflags.h>
> > > > > > +#include <linux/debug_locks.h>
> > > > > > #include <linux/preempt.h>
> > > > > > #include <linux/processor.h>
> > > > > > #include <linux/smp.h>
> > > > > > @@ -63,6 +64,12 @@ static inline void local_interrupt_disable(void)
> > > > > >
> > > > > > new_count = hardirq_disable_enter();
> > > > > >
> > > > > > + /* Is hardirq disable count overflow soon? */
> > > > > > + if (IS_ENABLED(CONFIG_DEBUG_PREEMPT))
> > > > > > + DEBUG_LOCKS_WARN_ON((new_count & HARDIRQ_DISABLE_MASK) +
> > > > > > + (10 << HARDIRQ_DISABLE_SHIFT) >
> > > > > > + HARDIRQ_DISABLE_MASK);
> > > > > > +
> > > > >
> > > > > This needs a return here right? Else we will see warning for 10 times
> > > > > and then overflow happens and we will call _local_interrupt_disable. No?
> >
> > Oh, seems I overlooked something... could you elaborate on this? What's
> > the scenario in your mind? You said we hit 10 times warning and *then*
> > overflow?
> >
>
>
> (the "10 times" wording was for the possible wraparound.
> I didn't know DEBUG_LOCKS_WARN_ON() will turn debug_locks
> off after the first warning, I thought it will print 10 times.)
>
> Main concern is the wraparound itself. If the HARDIRQ_DISABLE field reaches
> 0xff and we increment it once more, the field becomes zero after masking
> (ff + 1 -> 100 in the shifted field). Then local_interrupt_enable() can
> observe:
>
> (new_count & HARDIRQ_DISABLE_MASK) == 0
>
> and treat it as the outermost enable, potentially calling
> _local_interrupt_enable() while there are still outstanding logical
> local_interrupt_disable() users.
>
> So I think the useful debug check is one that catches the count before it
> gets close enough to wrap. Thing I wanted to avoid is silently making the
> HARDIRQ_DISABLE field look like zero after overflow.
>

Given that the detection is only on debug kernel, I think the damage of
keeping increment after warn triggered is low (this is assuming that we
can detect all the issues with the debug kernel, similar treatment as
the preempt_count nowadays). However, your suggestion does look better
for the corner cases and it's a better damage control, so I will add the
return part, thank you!

For the future, we need better detection and report for these counters
for sure.

Regards,
Boqun

> Whether it returns or just warns, I am not sure. As recovery may
> not be easy. It is meant more to be a damage contol than recovery.
>
>
> My line of thought is,
>
> If we return after debug checks in local_interrupt_disable(), we wont advance the preempt count
> further, so it wont wrap around. Now, there will be corresponding local_interrupt_enable(),
> at some point it will reach 0, we enable the interrupts. There will be some more
> local_interrupt_enable() still, but they will be caught with your debug check in
> local_interrupt_enable() and preempt count won't be decremented further.
> So if we put return we may have a damage control.
>
> > > > >
> > > >
> > > > DEBUG_LOCKS_WARN_ON() uses debug_locks_off() to avoid this, so we won't
> > > > see it 10 times. The reason not using return here, because we would
> > > > introduce unpaired local_interrupt_disable() if we returned:
> > >
> > > Yes, it could be a weird case if the overflow actually happens.
> > > So just warning maybe enough to catch such callers.
> > >
> > > Maybe your kunit test can actually help test the behavior with the loop
> > > count.
> > >
> > > >
> > > > // hardirq disable count is n
> > > > local_interrupt_disable(); // hardirq disable count is n + 1
> > > >
> > > > local_interrupt_disable(); <- trigger the warning, if we return
> > > > // hardirq disable count is n + 1
> > >
> > > Likely i am missing to understand.
> > >
> >
> > No, it was me who misunderstand ;-)
> >
> > > Isn't the count incremented earlier than return?
> > > I.e even if return happens it should be n + 2 right?
> > >
> >
> > Yeah, you're right, but then why do we want to return earlier? Since the
> > following if:
> >
> > if ((new_count & HARDIRQ_DISABLE_MASK) == HARDIRQ_DISABLE_OFFSET)
> >
> > will be false, and we will just return from the function, no?
> >
> > Regards,
> > Boqun
> >
> > > >
> > > > local_interrupt_enable(); // hardirq disable count is n
> > > >
> > > > local_interrupt_enable(); // hardirq disable count is n - 1
> > > >
> > > > Regards,
> > > > Boqun
> > > >
> > > > > Not sure, if below is any better? (Igore whitespace mangling)
> > > > >
> > > > > if (IS_ENABLED(CONFIG_DEBUG_PREEMPT) &&
> > > > > DEBUG_LOCKS_WARN_ON((preempt_count() & HARDIRQ_DISABLE_MASK) >=
> > > > > HARDIRQ_DISABLE_MASK - (10 << HARDIRQ_DISABLE_SHIFT)))
> > > > > return;
> > > > >
> > > > > > /* Interrupts can happen here, but it's OK, see __irq_exit_rcu(). */
> > > > > >
> > > > > > if ((new_count & HARDIRQ_DISABLE_MASK) == HARDIRQ_DISABLE_OFFSET)
> > > > > > @@ -73,6 +80,11 @@ static inline void local_interrupt_enable(void)
> > > > > > {
> > > > > > int new_count;
> > > > > >
> > > > > > + /* Unpaired local_interrupt_enable()? Warn and abort. */
> > > > > > + if (IS_ENABLED(CONFIG_DEBUG_PREEMPT) &&
> > > > > > + DEBUG_LOCKS_WARN_ON((preempt_count() & HARDIRQ_DISABLE_MASK) == 0))
> > > > > > + return;
> > > > > > +
> > > > > > new_count = hardirq_disable_exit();
> > > > > >
> > > > > > if ((new_count & HARDIRQ_DISABLE_MASK) == 0) >
> > > > >
> > >
>