Re: [PATCH v2 2/2] mm, memcg: fix memory.peak reset clobbering other fds' watermark

From: Johannes Weiner

Date: Wed Aug 12 2026 - 12:49:01 EST


On Fri, Aug 07, 2026 at 05:00:00PM +0800, Ridong wrote:
> From: Ridong Chen <chenridong@xxxxxxxxxx>
>
> Writing to memory.peak resets the peak for that fd only. Each fd is a
> watcher and reads back max(its own value, the shared local_watermark).
>
> peak_write() resets by lowering local_watermark to the current usage.
> To keep the other watchers' peaks it then walks the watcher list, but it
> stores the current usage into them instead of the old watermark. So once
> usage has dropped from a peak, a reset on one fd wrongly drags every
> other fd's peak down too, even fds that never reset.
>
> Reproduced on 7.2.0-rc5-next under QEMU, two fds A and B on one cgroup:
> B sees the peak (410624 KB), usage drops, then A resets -- and B's peak
> collapses to 1060 KB although B never reset. With this patch B keeps
> reading 410624 KB.
>
> Fix: save the old watermark before lowering it and use that to floor the
> other watchers, so a reset only affects the fd that issued it.
>
> Fixes: c6f53ed8f213 ("mm, memcg: cg2 memory{.swap,}.peak write handlers")
> Assisted-by: Claude:claude-opus-4-8
> Signed-off-by: Ridong Chen <chenridong@xxxxxxxxxx>
> ---
> mm/memcontrol.c | 7 ++++---
> 1 file changed, 4 insertions(+), 3 deletions(-)
>
> diff --git a/mm/memcontrol.c b/mm/memcontrol.c
> index 2da55b778ae3..28577beeb3d0 100644
> --- a/mm/memcontrol.c
> +++ b/mm/memcontrol.c
> @@ -4746,7 +4746,7 @@ static ssize_t peak_write(struct kernfs_open_file *of, char *buf, size_t nbytes,
> loff_t off, struct page_counter *pc,
> struct list_head *watchers)
> {
> - unsigned long usage;
> + unsigned long usage, peer_watermark;
> struct cgroup_of_peak *peer_ctx;
> struct mem_cgroup *memcg = mem_cgroup_from_css(of_css(of));
> struct cgroup_of_peak *ofp = of_peak(of);
> @@ -4754,11 +4754,12 @@ static ssize_t peak_write(struct kernfs_open_file *of, char *buf, size_t nbytes,
> spin_lock(&memcg->peaks_lock);
>
> usage = page_counter_read(pc);
> + peer_watermark = max(usage, READ_ONCE(pc->local_watermark));
> WRITE_ONCE(pc->local_watermark, usage);
>
> list_for_each_entry(peer_ctx, watchers, list)
> - if (usage > peer_ctx->value)
> - WRITE_ONCE(peer_ctx->value, usage);
> + if (peer_ctx != ofp && peer_watermark > peer_ctx->value)
> + WRITE_ONCE(peer_ctx->value, peer_watermark);

Sorry for letting your previous reply sit unanswered. You made a good
point on the peer_watermark = max(usage, local_watermark) being
pointless because that's how local_watermark moves to begin with.

So what you had before was indeed better. It was just me missing that
detail. Could you please go back to your original? Feel free to
include:

Acked-by: Johannes Weiner <hannes@xxxxxxxxxxx>