Re: [PATCH] mm/mmu_notifier: Remove non_block_start/end() from notifier invocation
From: David Woodhouse
Date: Wed Aug 12 2026 - 17:57:15 EST
On Wed, 2026-08-12 at 14:38 -0700, Paul E. McKenney wrote:
> On Wed, Aug 12, 2026 at 06:17:50PM +0100, David Woodhouse wrote:
> > On 12 August 2026 17:20:51 BST, Sean Christopherson <seanjc@xxxxxxxxxx> wrote:
> > > On Wed, Aug 12, 2026, David Woodhouse wrote:
> > > > On Wed, 2026-08-12 at 16:03 +0100, David Woodhouse wrote:
> > > > > On Wed, 2026-08-12 at 15:34 +0100, David Woodhouse wrote:
> > > > > > On Wed, 2026-08-12 at 15:05 +0100, David Woodhouse wrote:
> > > > > > > I'll rephrase that for my own understanding:
> > > > > > >
> > > > > > > *If* we go all the way to building a whole SRCU flavour for this *and*
> > > > > > > implementing a spin-only variant of srcu_synchronize() which is
> > > > > > > tailored to the atomic-reader use case, *then* we don't need to remove
> > > > > > > the non_block_{start,end} guards around the MMU notifiers, which are
> > > > > > > basically never being called anyway and don't actually seem to protect
> > > > > > > against any real bugs.
> > > > > > >
> > > > > > > Yes?
> > > > > >
> > > > > > FWIW it looks something like this. I'll throw it into my torture and
> > > > > > latency tests, and we can see what Paul thinks of it. I'm still utterly
> > > > > > unconvinced it's needed, but I concede it has its good points.
> > > > >
> > > > > This slightly refactored version is the one that's actually going into
> > > > > my torture tests...
> > > >
> > > > Well, it survived first contact, and it's doing the soak testing now.
> > > >
> > > > The average is basically no better than the try_synchronize_srcu()
> > > > case, unsurprisingly — as *both* of them just observe that there are no
> > > > readers and proceed immediately, in at least 99% of cases.
> > > >
> > > > Like the existing rwlock case, it still manages double-digit p100
> > > > latency even when though *doesn't* actually sleep.
> > > >
> > > > I don't *hate* it, but I do question the benefit of it over try-first.
> > >
> > > FWIW, the max latency and >8ms numbers are very appealing to me, as my concerns
> > > with using SRCU are all about the tail latencies.
> > >
> > > But I'm obviously not the one who'd be saddled with maintaining the code, so I'm
> > > more than a little biased towards choosing the more complex version.
> >
> > Ack. Well, it isn't even *that* bad in terms of complexity. It's
> > mostly just a while loop around the core of the
> > try_synchronize_srcu() I posted before, and some syntactic sugar
> > and safeguards.
> >
> > Let's see what Paul thinks.
>
> A single call to try_synchronize_srcu() that, upon failure, invokes
> synchronize_srcu_expedited() would work fine. And a limited loop on
> try_synchronize_srcu() that invokes synchronize_srcu_expedited() upon
> repeated failure return would also work fine.
>
> But if you had a large system, an unlimited loop on try_synchronize_srcu()
> could wait forever on a system that was large enough that there was at
> least one reader at all times. To guarantee loop termination, it is
> necessary to flip the index. That way, new readers would use the index
> that was already seen to be clear, and old readers would eventually
> drain from the old index.
Right. The 'while loop around try_synchronize_srcu()' as I described it
above was playing on my mind a little as I was out hunting the solar
eclipse this evening, and you *just* beat me to pointing it out.
> So is your use case adddressed by the single call to
> try_synchronize_srcu() that, upon failure, invokes
> synchronize_srcu_expedited()? If so, we have a solution!
I was angling for that... but given Sean's latest response above, I
don't think I was winning :)
Even if it wasn't a single call, but spin for a while before falling
back, that still doesn't address Sean's concern about the tail
latencies.
> If not, please let me know, and I will put together that does the job.
I think it would be good to take a look at that option, yes please if
it's a reasonable thing to do.
> Thanx, Paul
>
> > > > Again, I'll defer to Paul, but personally I'd want to see a more
> > > > compelling use case for it.
> > > >
> > > > ┌───────────────┬─────────────────────┬───────────────────┬─────────────────────┐
> > > > │ │ expedited │ try-first │ atomic │
> > > > ├───────────────┼─────────────────────┼───────────────────┼─────────────────────┤
> > > > │ median drain │ 32-128µs │ 4-16µs │ 4-16µs │
> > > > ├───────────────┼─────────────────────┼───────────────────┼─────────────────────┤
> > > > │ avg │ 118µs │ 13.8µs │ 12.3µs │
> > > > ├───────────────┼─────────────────────┼───────────────────┼─────────────────────┤
> > > > │ >1ms │ ~950ppm │ ~990ppm │ 838ppm │
> > > > ├───────────────┼─────────────────────┼───────────────────┼─────────────────────┤
> > > > │ >8ms │ 42ppm │ 4.4ppm │ 0.10ppm │
> > > > ├───────────────┼─────────────────────┼───────────────────┼─────────────────────┤
> > > > │ max │ 33.6ms │ 17.6ms │ 10.25ms │
> > > > ├───────────────┼─────────────────────┼───────────────────┼─────────────────────┤
> > > > │ fallback rate │ — │ 1.2% │ 0% │
> > > > ├───────────────┼─────────────────────┼───────────────────┼─────────────────────┤
> > > > │ sample │ 32.6M drains, 10min │ 41M drains, 10min │ 40.5M drains, 10min │
> > > > └───────────────┴─────────────────────┴───────────────────┴─────────────────────┘
> >
Attachment:
smime.p7s
Description: S/MIME cryptographic signature