Re: [PATCH] mm/mmu_notifier: Remove non_block_start/end() from notifier invocation

From: David Woodhouse

Date: Tue Aug 11 2026 - 10:30:38 EST


On Tue, 2026-08-11 at 10:55 -0300, Jason Gunthorpe wrote:
> On Tue, Aug 11, 2026 at 09:58:44AM +0100, David Woodhouse wrote:
> > As Michal Hocko put it when the
> > annotation was first proposed (quoted in commit 312364f3534c
> > ("kernel.h: Add non_block_start/end()")), the OOM reaper "shouldn't
> > depend on any locks or sleepable conditionals" and checking for
> > sleepable context was "the best thing we could come up with that would
> > describe these demands at least partially".
>
> Sure, but we translated this into the notifier must run in an atomic
> context and everyone has been happy with this.
>
> > That distinction now matters in both directions:
> >
> >  - On PREEMPT_RT, spinning locks become sleeping locks, and perfectly
> >    legitimate spinlock/rwlock usage in notifier implementations (e.g.
> >    KVM's mn_invalidate_lock and gfn_to_pfn_cache locks) triggers the
> >    splat despite having no allocator dependency whatsoever. This is
> >    reproducible today on a PREEMPT_RT kernel: KVM takes
> >    kvm->mn_invalidate_lock in kvm_mmu_notifier_invalidate_range_start(),
> >    and if the OOM reaper reaps a KVM process the result is a "BUG:
> >    sleeping function called from invalid context" from
> >    rt_spin_lock().
>
> I don't know anything about PREEEMPT_RT, but this seems like an issue
> with RT if a traditionally atomic safe functions are now triggering
> might sleep failures?

I can sympathise with that point of view. In fact I've spent the last
couple of years mostly ignoring this "problem" and just blaming RT for
doing exactly that, but I don't think we can really get away with it
any more.

cf. https://lore.kernel.org/all/787aa26cf62dfd361eea8ed19f384fc517892501.camel@xxxxxxxxxxxxx/

> >  - A notifier implementation may legitimately need to wait for an RCU
> >    grace period before allowing the caller to proceed with unmapping
>
> That's not allowed. We really want to forbid that, it is not an
> acceptable way to implement a driver using these APIs due to
> performance.

Speak for yourself. For the KVM gfn-to-pfn-cache the performance scales
*much* better with RCU than with explicit locking:
https://lore.kernel.org/all/8f41cb82b7c99d5a3d1dda016e4841326b4d8a52.camel@xxxxxxxxxxxxx/

Perhaps we could find a way to push down an *accurate* sanity check
into the code paths where what you say is *true*? I guess it could be
done with a flag on each notifier? Or *into* the notifier callback
function(s)?

Attachment: smime.p7s
Description: S/MIME cryptographic signature