Re: [PATCH v2] writeback: let foreign flushes reach dying cgwbs

From: Andrew Morton

Date: Mon Sep 28 2026 - 19:56:55 EST


On Mon, 28 Sep 2026 22:12:59 +0000 (UTC) Liz Fong-Jones <lizf@xxxxxxxxxxxx> wrote:

> ---
> At Honeycomb, a container that reads from Kafka and writes columnar
> files to a host volume stalls for 30-60s after each deploy replaces it
> (v6.18). We're increasingly confident this is the cause: the stall
> looks the same as in the reproducer (little CPU use, lag growing
> linearly and then recovering), and the mitigation it predicts, moving
> our final syncfs after the last write, worked in production (below).
> We haven't caught it with probes in production yet.
>
> ...
>
> Trigger: a cgroup dirties files and is removed while they are still
> dirty; a sibling keeps appending to the same files under a parent memory
> limit. Symptom: cgroup_writeback_by_id() returns -ENOENT and the sibling
> stalls in balance_dirty_pages(). Minimal recipe below; the harness that
> produced the numbers (paced writer, lag per second, MODE=alive control)
> is at https://gist.github.com/lizthegrey/2209d831930588f63076bdc0ac7b78e2.

IMO the above two paragraphs are the most important part of the patch
description yet they're in the throw-away section. They should be right at
the start of everything. Thanks for at least including them - many do not.

This is what people want to know! What problem does this solve? What
benefit is this to our users? Downstream people want to know "what
benefit is this to me?". Maintainers want to know "why should I spend
time on this person's patch rather than the billion others"?

But I keep saying that, to no observable effect, sigh.