Re: [PATCH v2] mm/khugepaged: flush deferred unmaps before dropping a failed folio
From: Hugh Dickins
Date: Fri Oct 09 2026 - 02:30:37 EST
On Tue, 6 Oct 2026, Kyle Zeng wrote:
> collapse_file() can fail its reference-count or dirty-folio check after
> unmapping with TTU_BATCH_FLUSH. Both paths put back the isolated folio,
> then unlock it and drop the lookup reference before reaching the common
> try_to_unmap_flush().
>
> The page-cache reference does not keep the folio stable once the lock is
> released. Another collapse can replace and free it. With its PTEs
> already gone, that collapse cannot flush the first task's per-task TLB
> batch, and retract_page_tables() skips short or unaligned VMAs. A CPU
> can therefore retain a user translation to the freed folio. This has
> been reproduced with unprivileged MADV_COLLAPSE on a memfd.
>
> Flush at out_unlock while the lookup reference and folio lock are still
> held. The common flush continues to cover the accumulated pagelist on
> both success and rollback, preserving batching on successful collapses.
>
> Fixes: 6d9df8a5889c ("mm/thp: collapse_file() do try_to_unmap(TTU_BATCH_FLUSH)")
> Cc: stable@xxxxxxxxxxxxxxx
> Assisted-by: LLM
> Signed-off-by: Kyle Zeng <kylebot@xxxxxxxxxx>
Good catch, yes, thanks: the folios already on the pagelist were correctly
flushed before reference dropped, but the one failing folio had its
reference dropped too soon.
I might have chosen to fix it differently (holding the final reference),
rather than duplicating the try_to_unmap_flush(); but you've put a nice
comment on its no-op when already flushed (thanks to Zi Yan), so I don't
think it's worth messing around with this good fix you've already tested.
Acked-by: Hugh Dickins <hughd@xxxxxxxxxx>
> ---
> Changes in v2:
> - Use Assisted-by: LLM.
> - Explain that the common flush is a no-op after out_unlock flushes.
>
> mm/khugepaged.c | 10 +++++++---
> 1 file changed, 7 insertions(+), 3 deletions(-)
>
> diff --git a/mm/khugepaged.c b/mm/khugepaged.c
> index 75639298efc2..e1a5890818ad 100644
> --- a/mm/khugepaged.c
> +++ b/mm/khugepaged.c
> @@ -2478,6 +2478,11 @@ static enum scan_result collapse_file(struct mm_struct *mm, unsigned long addr,
> index += folio_nr_pages(folio);
> continue;
> out_unlock:
> + /*
> + * The folio may have been unmapped with TTU_BATCH_FLUSH.
> + * Flush before releasing the lock and our last reference.
> + */
> + try_to_unmap_flush();
> folio_unlock(folio);
> folio_put(folio);
> goto xa_unlocked;
> @@ -2488,9 +2493,8 @@ static enum scan_result collapse_file(struct mm_struct *mm, unsigned long addr,
> xa_unlocked:
>
> /*
> - * If collapse is successful, flush must be done now before copying.
> - * If collapse is unsuccessful, does flush actually need to be done?
> - * Do it anyway, to clear the state.
> + * Flush before copying the folios, or releasing them in rollback.
> + * This is a no-op if out_unlock already flushed the batch.
> */
> try_to_unmap_flush();
>