Re: [RFC PATCH v2 0/3] locking, mm: Add atomic allocator trylocks on RT
From: Karl Mehltretter
Date: Fri Oct 09 2026 - 06:04:53 EST
On Wed, Oct 07, 2026 at 07:15:29PM +0100, Harry Yoo wrote:
> > This RFC instead adds an atomic owner state for bounded PREEMPT_RT spinlock
> > trylocks. Atomic acquisition succeeds only from the completely free state.
>
> Please don't skip discussion stage and jump into submitting a very
> intrusive change across locking and mm, assisted by LLM? :/
>
> I don't think attempts to tackle this problem will make progress without
> discussing and getting buy-in from (PREEMPT_RT) locking folks.
>
Sorry, I should have followed up on the additional v1 replies before
posting this v2.
On the alternatives in [1], Alexei [2] preferred detecting
scheduler-lock contexts and returning NULL. He confirmed that NULL
returns are allowed, rejected restoring the local-storage allocator, and
raised the RT unlock/wakeup objection.
I initially tried that detection, tracking pi_lock, rtmutex wait_lock
and rq locks. With RT, lockdep and slab_debug enabled, task-storage
allocation under an independently held hrtimer base lock reported this
cycle:
hrtimer_bases.lock -> rtmutex wait_lock -> pi_lock -> rq->__lock
-> hrtimer_bases.lock
I don't know whether my changes introduced or exposed the dependency.
The timer lock was not covered by that guard. This was a lockdep warning
and the workload completed.
I then tried atomic ownership in v2 to avoid that unlock path, at
the cost of IRQ-disabled spinning without PI boosting of the owner.
One possible interpretation of "raw_local_trylock_t" would be preserving
the non-RT local_trylock semantics on RT for short per-CPU cache
operations, with allocation failing when it needs an rtmutex-backed
fallback.
Would that be worth exploring? Sebastian suggested
consolidating the existing local-lock variants first. What should that
cover?
For arena user faults, backing-page allocation failure returns
VM_FAULT_SIGSEGV on the RFC base. next-20261007 preallocates outside the
raw lock. Failed fallback allocation returns VM_FAULT_SIGBUS. If
restricting that fallback increases fault failures, is that acceptable?
Thanks,
Karl
[1] https://lore.kernel.org/r/arIY0wMQjCCsqidI@xxxxxxxxx
[2] https://lore.kernel.org/r/DLO3RV3IQUOV.21BNLF4QS9T4W@xxxxxxxxx