Re: [PATCH] mm/oom_kill: fix hung tasks queued on mmap_lock behind a long reap
From: Michal Hocko
Date: Tue Sep 15 2026 - 04:38:05 EST
On Tue 15-09-26 16:13:11, Jiayuan Chen wrote:
>
> On 9/15/26 2:57 PM, Michal Hocko wrote:
> > On Mon 14-09-26 20:35:33, Andrew Morton wrote:
> > > On Mon, 14 Sep 2026 20:36:16 +0200 Michal Hocko <mhocko@xxxxxxxx> wrote:
> > >
> > > > > Call Trace:
> > > > > <TASK>
> > > > > __schedule+0x487/0x1870
> > > > > schedule+0x28/0xb0
> > > > > schedule_preempt_disabled+0x16/0x30
> > > > > rwsem_down_write_slowpath+0x1d4/0x750
> > > > > down_write+0x60/0x70
> > > > > __ksm_exit+0xb4/0x230
> > > > > __mmput+0x12c/0x150
> > > > > mmput+0x1e/0x30
> > > > > do_exit+0x283/0xa30
> > > > > do_group_exit+0x34/0x90
> > > > > get_signal+0x952/0x960
> > > > > arch_do_signal_or_restart+0x41/0x250
> > > > > exit_to_user_mode_loop+0xd3/0x560
> > > > > do_syscall_64+0x385/0x470
> > > > > </TASK>
> > > > >
> > > > > KSM is just the one LTP happened to hit: __khugepaged_exit() has the
> > > > > same write lock cycle ahead of exit_mmap().
> > > > Why is this a practical problem we need to care about? It is kind of
> > > > natural that the oom victim exit path might race with the oom reaper. They
> > > > share the same lock that is mutualy exclusive. The whole point of the
> > > > reaper is to ensure there is a forward progress achieved. So before we
> > > > start modifying this let's talk about any practical/real life problems.
>
> Hi Michal, Andrew
>
> Agreed, the hung task warning itself is harmless, especially for a dying
> task. The real problems are what sits behind it.
>
> With a 500G swapped-out victim the reap takes ~600s, and for all of it:
>
> 1. The victim cannot exit. __ksm_exit() needs mmap_lock for write and
> queues behind the reaper, so the process stays alive in D state for
> 10 minutes and whoever waits for it (parent, container runtime)
> waits too.
>
> 2. ksmd and khugepaged stall. Once that writer is queued, their
> mmap_read_lock() on this mm queues as well, so both daemons stop
> for the whole system for the same ~600s.
Right. But why is that a problem we need to fix? OOM reaper is taking a
prortion of the exit time by doing the leg work of tearing down the
address space. Exiting task would need to do the same so it is unlikely
to terminate much faster.
--
Michal Hocko
SUSE Labs