Re: [PATCH] mm/oom_kill: fix hung tasks queued on mmap_lock behind a long reap

From: Andrew Morton

Date: Mon Sep 14 2026 - 23:37:19 EST


On Mon, 14 Sep 2026 20:36:16 +0200 Michal Hocko <mhocko@xxxxxxxx> wrote:

> > Call Trace:
> > <TASK>
> > __schedule+0x487/0x1870
> > schedule+0x28/0xb0
> > schedule_preempt_disabled+0x16/0x30
> > rwsem_down_write_slowpath+0x1d4/0x750
> > down_write+0x60/0x70
> > __ksm_exit+0xb4/0x230
> > __mmput+0x12c/0x150
> > mmput+0x1e/0x30
> > do_exit+0x283/0xa30
> > do_group_exit+0x34/0x90
> > get_signal+0x952/0x960
> > arch_do_signal_or_restart+0x41/0x250
> > exit_to_user_mode_loop+0xd3/0x560
> > do_syscall_64+0x385/0x470
> > </TASK>
> >
> > KSM is just the one LTP happened to hit: __khugepaged_exit() has the
> > same write lock cycle ahead of exit_mmap().
>
> Why is this a practical problem we need to care about? It is kind of
> natural that the oom victim exit path might race with the oom reaper. They
> share the same lock that is mutualy exclusive. The whole point of the
> reaper is to ensure there is a forward progress achieved. So before we
> start modifying this let's talk about any practical/real life problems.

If this situation is expected, unavoidable etc then perhaps the best
change is to periodically poke the hung-task detector?