Re: [PATCH] mm: memcg: stop reclaim when a limit update is superseded
From: Guopeng Zhang
Date: Wed Jul 29 2026 - 02:16:56 EST
在 2026/7/27 21:58, Michal Hocko 写道:
> On Mon 27-07-26 20:59:23, Guopeng Zhang wrote:
>>
>>
>> 在 2026/7/27 16:16, Michal Hocko 写道:
>>> On Fri 24-07-26 10:18:05, Guopeng Zhang wrote:
>>>> From: Guopeng Zhang <zhangguopeng@xxxxxxxxxx>
>>>>
>>>> kernfs serializes file operations only per open file, so separate open
>>>> files can update the same memory.high or memory.max file concurrently.
>>>> Both handlers store the new limit before synchronous reclaim, but
>>>> continue to use the writer's local target in the reclaim loop. If another
>>>> writer raises or removes the limit, the first writer can continue
>>>> reclaiming toward a stale target.
>>>>
>>>> For memory.max, this can leave the writer looping indefinitely once
>>>> reclaim retries are exhausted. The OOM path sees sufficient margin under
>>>> the current limit and returns true without killing, while the writer
>>>> still compares usage against its stale target and records another OOM
>>>> event.
>>>
>>> The current behavior is deliberate as described in b6e6edcfa4056.
>>> What is an actual problem you are trying to fix?
>>
>> Thanks for raising this. I agree that the behavior introduced by
>> b6e6edcfa405 is deliberate while the limit installed by the writer
>> remains current. The problem occurs when that limit is overwritten by a
>> later write through another open file.
>>
>> Writer A stores a low memory.max and enters synchronous reclaim. Writer
>> B then restores memory.max to "max". A still compares usage against its
>> local old target, while mem_cgroup_out_of_memory() checks the current
>> memory.max. Since 1378b37d03e8, the current-margin check in
>> mem_cgroup_out_of_memory() sees sufficient margin and returns true
>> without selecting a victim. Once the reclaim retries are exhausted, A
>> therefore loops indefinitely and increments the oom counter in
>> memory.events on every iteration.
>>
>> I reproduced this with a cgroup holding 128 MiB of anonymous memory and
>> with swapping disabled for the cgroup. Writer A lowered memory.max to
>> 32 MiB. After that value became visible, writer B restored memory.max
>> to "max" through another open file.
>
> Is this trying to replicate any real workload? One would expect that
> writers to limit do some sort of coordination otherwise the exact
> behavior is not really well defined.
>
No, this was not motivated by a reported production workload. We found
it through automated randomized testing for our cgroup observability
work and reduced it to the reproducer above.
>> On the unpatched kernel, A remained blocked after B's write, and the oom
>> counter in memory.events increased from 37333 to 13512861 during the
>> reproducer's one-second sampling interval. With the patch, the same
>> reproducer observed A return after B superseded A's limit.
>>
>> The new check does not change the behavior introduced by b6e6edcfa405
>> while the writer's target remains the active limit. For the indefinite
>> loop described above, 1378b37d03e8 appears to be the more precise
>> Fixes: target for the memory.max hunk. Does that match your reading of
>> the history?
>
> Well, to be really honest I am not really convinced this needs fixing.
> And if yes, your patch changes a well established behavior existing
> userspace might already depend on. While your described case doesn't
> look great it doesn't seem really harmful and the looping task is
> killable.
That makes sense. Without a concrete workload showing practical impact,
there is not enough justification to change the established behavior.
We can revisit this if such a workload turns up.
Andrew, please drop this patch from your queue.
Thanks,
Guopeng