Re: [PATCH] mm: add cond_resched() to free_pud_range()
From: Leon Hwang
Date: Tue Aug 18 2026 - 21:55:29 EST
On 18/8/26 22:09, Mike Rapoport wrote:
> On Tue, Aug 18, 2026 at 09:49:34PM +0800, Leon Hwang wrote:
>> Packet receive timeouts were seen in production. Tracing showed that an
>> exiting process with a sparse 2.5 TiB mapping could remain in kernel
>> context for over 20 ms without reaching a scheduling point while
>> freeing PTE page-table pages. Hard IRQs could still be handled, but the
>> per-CPU ksoftirqd thread and other runnable tasks could not run during
>> that interval, delaying NET_RX softirq work queued to ksoftirqd.
>>
>> Like zap_pud_range(), add cond_resched() to free_pud_range() so
>> ksoftirqd and other runnable tasks can run between PUD entries. Testing
>> with PREEMPT_NONE showed that the maximum interval between scheduling
>> points fell from over 20 ms to below 2 ms.
>>
>> Reported-by: Lance Yang <lance.yang@xxxxxxxxx>
>> Tested-by: Lance Yang <lance.yang@xxxxxxxxx>
>> Signed-off-by: Leon Hwang <leon.hwang@xxxxxxxxx>
>> ---
>> mm/memory.c | 2 +-
>> 1 file changed, 1 insertion(+), 1 deletion(-)
>>
>> diff --git a/mm/memory.c b/mm/memory.c
>> index 4134ac607ee0..68c15449de07 100644
>> --- a/mm/memory.c
>> +++ b/mm/memory.c
>> @@ -231,7 +231,7 @@ static inline void free_pud_range(struct mmu_gather *tlb, p4d_t *p4d,
>> if (pud_none_or_clear_bad(pud))
>> continue;
>> free_pmd_range(tlb, pud, addr, next, floor, ceiling);
>> - } while (pud++, addr = next, addr != end);
>> + } while (pud++, cond_resched(), addr = next, addr != end);
>
> This is really obscure, can't it be in the loop body?
Agreed. It should be in the loop body, after free_pmd_range().
Thanks,
Leon
>>
>> start &= P4D_MASK;
>> if (start < floor)
>> --
>> 2.55.0
>>
>>
>