Re: [PATCH] mm/fadvise: skip remote LRU drains for ineligible folios

From: Andrew Morton

Date: Wed Oct 07 2026 - 16:23:15 EST


On Tue, 06 Oct 2026 00:27:19 +0200 Serapheim Dimitropoulos <serapheimd@xxxxxxxxx> wrote:

> From: Serapheim Dimitropoulos <sdimitropoulos@xxxxxxxxxxxxx>
>
> POSIX_FADV_DONTNEED retries invalidation after a global LRU drain whenever
> mapping_try_invalidate() reports a failed eviction. This includes failures
> for mapped, dirty or writeback folios, which dropping LRU batch references
> cannot make evictable while those conditions persist.
>
> Filter those failures in mapping_try_invalidate(), while the folio is still
> locked and before deactivation can enqueue another batch reference. Leave
> mapping_evict_folio() and its eviction safety checks unchanged.
>
> Use a boolean retry flag instead of a failure count: generic_fadvise() only
> needs to decide whether to drain and retry once. This remains a heuristic,
> not a test for remote LRU references. Preserve the retry for other failures
> on clean, unmapped folios, including failures from filemap_release_folio()
> and remove_mapping(), rather than limiting it to the early refcount check.
>
> This follows the problem identified in fujunjie's earlier proposal, with
> the filtering kept in mapping_try_invalidate() and a conservative fallback
> for other eviction failures.

It's not particularly clear from the above, but this appears to be a
performance optimization. No increase in POSIX_FADV_DONTNEED's success
rate is expected?

> Link: https://lkml.rescloud.iu.edu/2605.0/07547.html
> Link: https://lkml.iu.edu/2605.1/03284.html
> Link: https://lkml.iu.edu/2605.1/03732.html
> Signed-off-by: Serapheim Dimitropoulos <sdimitropoulos@xxxxxxxxxxxxx>
> ---
> Tested baseline and patched kernels in QEMU on ext4, XFS and OverlayFS.
> For clean mapped files on each filesystem, 32 POSIX_FADV_DONTNEED calls
> produced 32 lru_add_drain_all() calls before the patch and none afterwards.
> A separate dirty, unmapped test using ext4 data=journal showed the same
> reduction.

The whole point of the patch is a performance optimization, so it lives
or dies by measurements. This vital info shouldn't be below the ---
throwaway line!

Also, fujunjie's original had timing measurements, which are nice to
see.

Anyway, I totally believe that this makes things faster so there's no
need to do much work on this. Just saying.

We're in lockdown mode for this -rc cycle so please await review and
plan to respin/resend after next -rc1, thanks.