Re: [PATCH] mm/oom_kill: fix hung tasks queued on mmap_lock behind a long reap
From: Jiayuan Chen
Date: Tue Sep 15 2026 - 04:23:42 EST
On 9/15/26 2:57 PM, Michal Hocko wrote:
On Mon 14-09-26 20:35:33, Andrew Morton wrote:
On Mon, 14 Sep 2026 20:36:16 +0200 Michal Hocko <mhocko@xxxxxxxx> wrote:
Call Trace:Why is this a practical problem we need to care about? It is kind of
<TASK>
__schedule+0x487/0x1870
schedule+0x28/0xb0
schedule_preempt_disabled+0x16/0x30
rwsem_down_write_slowpath+0x1d4/0x750
down_write+0x60/0x70
__ksm_exit+0xb4/0x230
__mmput+0x12c/0x150
mmput+0x1e/0x30
do_exit+0x283/0xa30
do_group_exit+0x34/0x90
get_signal+0x952/0x960
arch_do_signal_or_restart+0x41/0x250
exit_to_user_mode_loop+0xd3/0x560
do_syscall_64+0x385/0x470
</TASK>
KSM is just the one LTP happened to hit: __khugepaged_exit() has the
same write lock cycle ahead of exit_mmap().
natural that the oom victim exit path might race with the oom reaper. They
share the same lock that is mutualy exclusive. The whole point of the
reaper is to ensure there is a forward progress achieved. So before we
start modifying this let's talk about any practical/real life problems.
Hi Michal, Andrew
Agreed, the hung task warning itself is harmless, especially for a dying
task. The real problems are what sits behind it.
With a 500G swapped-out victim the reap takes ~600s, and for all of it:
1. The victim cannot exit. __ksm_exit() needs mmap_lock for write and
queues behind the reaper, so the process stays alive in D state for
10 minutes and whoever waits for it (parent, container runtime)
waits too.
2. ksmd and khugepaged stall. Once that writer is queued, their
mmap_read_lock() on this mm queues as well, so both daemons stop
for the whole system for the same ~600s.
With the patch they wait for one 1G chunk at most, and since exit_mmap()
now frees alongside the reaper, the victim's memory is gone in 264s
instead of ~600s. The reaper still never blocks on mmap_lock and still
makes progress on every pass, so forward progress is not changed.
If this situation is expected, unavoidable etc then perhaps the bestRight. Reaping 10s of GBs worth of VMAs might take some time indeed and
change is to periodically poke the hung-task detector?
that could trigger the hung task detector.