Re: [PATCH v2] mm: vmalloc: fix vmap_purge_lock livelock under memory pressure

From: Dev Jain

Date: Wed Sep 02 2026 - 00:31:49 EST




On 01/09/26 10:29 pm, Uladzislau Rezki wrote:
> On Tue, Sep 01, 2026 at 06:43:43PM +0200, Uladzislau Rezki wrote:
>> On Tue, Sep 01, 2026 at 05:07:30PM +0800, Ye Liu wrote:
>>>
>>>
>>> 在 2026/9/1 14:22, Dev Jain 写道:
>>>>
>>>>
>>>> On 31/08/26 3:24 pm, Uladzislau Rezki wrote:
>>>>> On Mon, Aug 31, 2026 at 11:39:14AM +0530, Dev Jain wrote:
>>>>>>
>>>>>>
>>>>>> On 28/08/26 11:35 pm, Andrew Morton wrote:
>>>>>>> On Fri, 28 Aug 2026 17:17:53 +0800 Ye Liu <ye.liu@xxxxxxxxx> wrote:
>>>>>>>
>>>>>>>> From: Ye Liu <liuye@xxxxxxxxxx>
>>>>>>>>
>>>>>>>> The vmap_purge_lock mutex can be held for an extended period by
>>>>>>>> __purge_vmap_area_lazy() which calls flush_work() to wait for
>>>>>>>> purge_vmap_node workers while holding the lock. Under memory
>>>>>>>> pressure, those workers may themselves be blocked in direct
>>>>>>>> reclaim trying to acquire the same lock via the
>>>>>>>> vmap_node_shrink_scan() shrinker callback, creating a circular
>>>>>>>> dependency that deadlocks the entire system.
>>>>>>>>
>>>>>>>> Two places acquire vmap_purge_lock from paths that can be reached
>>>>>>>> during direct reclaim:
>>>>>>>>
>>>>>>>> 1. vmap_node_shrink_scan(): replace blocking guard(mutex) with
>>>>>>>> mutex_trylock(). This is a shrinker that only decays the vmap
>>>>>>>> pool and returns SHRINK_STOP without freeing memory; skipping a
>>>>>>>> decay cycle when the lock is contended is harmless and prevents
>>>>>>>> tasks from piling up on the mutex in the direct reclaim path.
>>>>>>>>
>>>>>>>> 2. reclaim_and_purge_vmap_areas(): replace mutex_lock() with
>>>>>>>> mutex_trylock(). This is called from the vmalloc allocation
>>>>>>>> overflow path; if trylock fails, another thread is already
>>>>>>>> purging and the allocator's retry will find freed space. The
>>>>>>>> notifier chain provides a fallback if the retry still fails.
>>>>>>>>
>>>>>>>> Both trylock failures break the circular dependency: the lock
>>>>>>>> holder's flush_work() can complete because workers are no longer
>>>>>>>> blocked on vmap_purge_lock in the direct reclaim path.
>>>>>>>
>>>>>>> Thanks. AI review expressed a couple of concerns:
>>>>>>> https://sashiko.dev/#/patchset/20260828091753.299295-1-ye.liu@xxxxxxxxx
>>>>>>
>>>>>>
>>>>>> Sounds legit to me. Now there is no guarantee of purge being successful, and we
>>>>>> can get a spurious failure.
>>>>>>
>>>>>> How about using a WQ_RECLAIM workqueue:
>>>>>>
>>>>>> vmap_purge_wq = alloc_workqueue("vmap_purge",
>>>>>> WQ_MEM_RECLAIM | WQ_PERCPU, 0);
>>>>>>
>>>>>> I see the same pattern in __lru_add_drain_all() and kmem_cache_init_late().
>>>>>>
>>>>> WQ_MEM_RECLAIM makes sense but this is another patch.
>>>>>
>>>>> I copied here AI comment:
>>>>>>
>>>>>> Does replacing this blocking lock with a trylock break synchronization for
>>>>>> the callers?
>>>>>>
>>>>> No it does not. If someone is doing reclaim we do not wait and do not try
>>>>> to do it again thus fail allocation.
>>>>>
>>>>>> When vmalloc space is exhausted, alloc_vmap_area() calls
>>>>>> reclaim_and_purge_vmap_areas() and relies on its blocking behavior to ensure
>>>>>> that free space has actually been reclaimed before looping back to retry:
>>>>>> mm/vmalloc.c:alloc_vmap_area() {
>>>>>> ...
>>>>>> overflow:
>>>>>> if (!purged) {
>>>>>> reclaim_and_purge_vmap_areas();
>>>>>> purged = 1;
>>>>>> goto retry;
>>>>>> }
>>>>>> ...
>>>>>> }
>>>>>> With this patch, if another thread holds vmap_purge_lock, mutex_trylock()
>>>>>> fails and the function returns immediately.
>>>>>>
>>>>> If reclaim is in progress and trylock fails a caller repeats only one
>>>>> time to retry an allocation. There is no any infinite loop.
>>>>>
>>>>>>
>>>>>> The allocator then retries
>>>>>> instantly without waiting for the concurrent purge to complete.
>>>>>> Because the retry fails and purged is already 1, could this cause the
>>>>>> allocation to abort and return a spurious vmalloc allocation failure
>>>>>> (-EBUSY or -ENOMEM)?
>>>>>>
>>>>> vmap space can be fragmented and not avail for 32-bit systems. For
>>>>> 64-bit system it is likely impossible.
>>>>>
>>>>> But, i think we can overt mutex_lock() into mutex_trylock() just only
>>>>> in the:
>>>>>
>>>>> static unsigned long
>>>>> vmap_node_shrink_scan(struct shrinker *shrink, struct shrink_control *sc)
>>>>> {
>>>>> struct vmap_node *vn;
>>>>>
>>>>> guard(mutex)(&vmap_purge_lock);
>>>>> for_each_vmap_node(vn)
>>>>> decay_va_pool_node(vn, true);
>>>>>
>>>>> return SHRINK_STOP;
>>>>> }
>>>>>
>>>>> so the reclaim path is not blocked. It should also address an issue
>>>>> reported by the Ye Liu <ye.liu@xxxxxxxxx>.
>>>>
>>>> IIUC you are suggesting mutex_trylock() only in the shrinker path. But
>>>> then, the following is possible no: take purge lock, try to get a
>>>> worker thread, worker thread is stuck in vmalloc -> alloc_vmap_area
>>>> -> reclaim_and_purge_vmap_areas -> take purge lock?
>>>>
>>>>
>>>
>>> Personally, I lean toward the WQ_MEM_RECLAIM workqueue solution.
>>> After taking a closer look at the code, I noticed a subtle but
>>> potentially problematic scenario:
>>>
>>> When drain_vmap_area_work acquires vmap_purge_lock and calls into
>>> __purge_vmap_area_lazy, it may subsequently invoke queue_work/queue_work_on
>>> on the same CPU's system_wq. If the newly queued work ends up waiting
>>> for an available worker on that same CPU, while the current worker is
>>> blocked waiting for that very work to complete (via flush_work),
>>> we could end up with a self-deadlock on a single CPU.
>>>
>>> Theoretically, this seems possible. I suspect the reason we don't
>>> see widespread reports of such deadlocks is that the nr_purge_helpers
>>> logic limits the number of asynchronous workers; when resources are tight,
>>> it falls back to synchronous execution (purge_vmap_node directly),
>>> which avoids queuing additional work.
>>>
>>> Using a dedicated workqueue with WQ_MEM_RECLAIM would provide a clean,
>>> explicit isolation—ensuring forward progress under memory pressure and
>>> eliminating the risk of interfering with other subsystems' workqueues.
>>> I believe this approach is more robust in the long run.
>>>
>>> Perhaps like the code below:
>>> the dedicated queue eliminates the self‑deadlock risk, and the trylock
>>> in the shrinker prevents recursive lock attempts from reclaim contexts.
>>>
>>> diff --git a/mm/vmalloc.c b/mm/vmalloc.c
>>> index bea9f76ed7e7..68fc1f5acb2f 100644
>>> --- a/mm/vmalloc.c
>>> +++ b/mm/vmalloc.c
>>> @@ -2218,6 +2218,9 @@ static unsigned long lazy_max_pages(void)
>>> */
>>> static DEFINE_MUTEX(vmap_purge_lock);
>>>
>>> +/* Workqueue for lazy vmap purging; WQ_MEM_RECLAIM guarantees progress. */
>>> +static struct workqueue_struct *vmap_purge_wq;
>>> +
>>> /* for per-CPU blocks */
>>> static void purge_fragmented_blocks_allcpus(void);
>>>
>>> @@ -2408,9 +2411,9 @@ static bool __purge_vmap_area_lazy(unsigned long start, unsigned long end,
>>> INIT_WORK(&vn->purge_work, purge_vmap_node);
>>>
>>> if (cpumask_test_cpu(i, cpu_online_mask))
>>> - schedule_work_on(i, &vn->purge_work);
>>> + queue_work_on(i, vmap_purge_wq, &vn->purge_work);
>>> else
>>> - schedule_work(&vn->purge_work);
>>> + queue_work(vmap_purge_wq, &vn->purge_work);
>>>
>>> nr_purge_helpers--;
>>> } else {
>>> @@ -5519,10 +5522,14 @@ vmap_node_shrink_scan(struct shrinker *shrink, struct shrink_control *sc)
>>> {
>>> struct vmap_node *vn;
>>>
>>> - guard(mutex)(&vmap_purge_lock);
>>> + if (!mutex_trylock(&vmap_purge_lock))
>>> + return SHRINK_STOP;
>>> +
>>> for_each_vmap_node(vn)
>>> decay_va_pool_node(vn, true);
>>>
>>> + mutex_unlock(&vmap_purge_lock);
>>> +
>>> return SHRINK_STOP;
>>> }
>>>
>>> @@ -5575,6 +5582,17 @@ void __init vmalloc_init(void)
>>> * Now we can initialize a free vmap space.
>>> */
>>> vmap_init_free_space();
>>> +
>>> + /*
>>> + * A dedicated workqueue for lazy vmap purging. WQ_MEM_RECLAIM
>>> + * reserves a rescue worker so queued purge work items are executed
>>> + * even under memory pressure, when workers of the system workqueue
>>> + * may be stuck in direct reclaim.
>>> + */
>>> + vmap_purge_wq = alloc_workqueue("vmap_purge",
>>> + WQ_MEM_RECLAIM | WQ_PERCPU, 0);
>>> + WARN_ON(!vmap_purge_wq);
>>> +
>>> vmap_initialized = true;
>>>
>> I agree. We should have it and it should be as separate patch, i.e.
>> split vmap_node_shrink_scan() and dedicated per-cpu WQs per vmap drain.
>>
> And i sent out already the WQ_UNBOUND | WQ_MEM_RECLAIM and separate WQ
> for vmap drain logic. It looks like Andrew/me forgot about it:
>
> https://lore.kernel.org/all/20260331202352.879718-1-urezki@xxxxxxxxx/

Great! Perhaps resend it and we can review it?


>
> --
> Uladzislau Rezki