Re: [PATCH 1/2] fuse: zero the partial EOF page when extending a file
From: Miklos Szeredi
Date: Wed Aug 19 2026 - 05:33:58 EST
[Cc: Jan + Joanne]
On Fri, 31 Jul 2026 at 22:39, Jimmy Zuber <jamz@xxxxxxxxxx> wrote:
>
> Extending a fuse file past a non-page-aligned EOF does not zero the tail of
> the old last page. If that page is cached and was dirtied beyond the old
> EOF -- e.g. an application mmap()ed the EOF page and stored into the region
> past EOF, which is undefined until the file grows -- the now in-bounds tail
> is exposed to subsequent reads as stale data instead of zeros, in violation
> of POSIX file-extension semantics.
>
> Other filesystems zero this via pagecache_isize_extended(), but that helper
> is a no-op for fuse: it returns early when i_blocksize() >= PAGE_SIZE, and
> a non-fuseblk fuse mount has s_blocksize == PAGE_SIZE (the server-supplied
> st_blksize only sets fi->cached_i_blkbits, not i_blkbits). The NFS client
> hit the same problem and open-codes the zeroing in
> nfs_truncate_last_folio(); add the equivalent fuse_zero_partial_eof_folio()
> and call it from the three paths that extend a file: a buffered write, a
> size-extending setattr/truncate, and a size-extending fallocate
> (fuse_write_update_attr(), fuse_do_setattr() and fuse_file_fallocate()).
>
> writeback_cache connections are unaffected, as their writes go through
> iomap_file_buffered_write(), which zeroes post-EOF folios. The bug is
> observable on a non-writeback_cache server that returns FOPEN_KEEP_CACHE on
> writable files (without FOPEN_DIRECT_IO), and is caught by the new
> write_extend_eof fuse selftest.
>
> Signed-off-by: Jimmy Zuber <jamz@xxxxxxxxxx>
> ---
> fs/fuse/dir.c | 3 +++
> fs/fuse/file.c | 56 ++++++++++++++++++++++++++++++++++++++++++++++++
> fs/fuse/fuse_i.h | 1 +
> 3 files changed, 60 insertions(+)
>
> diff --git a/fs/fuse/dir.c b/fs/fuse/dir.c
> index 795e92037ce7..f6614ccef186 100644
> --- a/fs/fuse/dir.c
> +++ b/fs/fuse/dir.c
> @@ -2282,6 +2282,9 @@ int fuse_do_setattr(struct mnt_idmap *idmap, struct dentry *dentry,
> */
> if ((is_truncate || !is_wb) &&
> S_ISREG(inode->i_mode) && oldsize != outarg.attr.size) {
> + if (outarg.attr.size > oldsize)
> + fuse_zero_partial_eof_folio(inode, oldsize,
> + outarg.attr.size);
> truncate_pagecache(inode, outarg.attr.size);
> invalidate_inode_pages2(mapping);
> }
> diff --git a/fs/fuse/file.c b/fs/fuse/file.c
> index cb8da4c06d17..a9063b4e9217 100644
> --- a/fs/fuse/file.c
> +++ b/fs/fuse/file.c
> @@ -21,6 +21,8 @@
> #include <linux/splice.h>
> #include <linux/task_io_accounting_ops.h>
> #include <linux/iomap.h>
> +#include <linux/highmem.h>
> +#include <linux/rmap.h>
>
> static int fuse_send_open(struct fuse_mount *fm, u64 nodeid,
> unsigned int open_flags, int opcode,
> @@ -1200,20 +1202,64 @@ static ssize_t fuse_send_write(struct fuse_io_args *ia, loff_t pos,
> return err ?: ia->write.out.size;
> }
>
> +/*
> + * An operation extended i_size past a non-folio-aligned old EOF at @from,
> + * turning [@from, @to) into a hole that must read back as zero. If the old
> + * last folio is cached and was dirtied beyond the old EOF (e.g. mmap stores
> + * into the post-EOF region, which are undefined until the file grows), zero
> + * that tail so it is not exposed as stale data (xfstests generic/363).
> + *
> + * pagecache_isize_extended() cannot be used: it bails out for
> + * i_blocksize() >= PAGE_SIZE, and a non-fuseblk mount has
> + * s_blocksize == PAGE_SIZE, so the zeroing has to be done here.
> + * Callers hold i_rwsem, serialising this against concurrent writes and
> + * truncates; it must not run under fi->lock, as it locks the folio.
> + */
> +void fuse_zero_partial_eof_folio(struct inode *inode, loff_t from, loff_t to)
> +{
> + struct folio *folio;
> + size_t offset, end;
> +
> + if (from >= to)
> + return;
> +
> + folio = filemap_lock_folio(inode->i_mapping, from >> PAGE_SHIFT);
> + if (IS_ERR(folio))
> + return;
> +
> + if (folio_mkclean(folio))
> + folio_mark_dirty(folio);
> +
> + if (folio_test_dirty(folio)) {
> + offset = offset_in_folio(folio, from);
> + end = min_t(loff_t, to - folio_pos(folio), folio_size(folio));
> + folio_zero_segment(folio, offset, end);
> + }
> +
> + folio_unlock(folio);
> + folio_put(folio);
> +}
> +
While the fix is probably correct, I'm not happy with it.
Why does pagecache_isize_extended() not handle blocksize >= PAGE_SIZE?
Where does straddling folio zeroing happen for other filesystems?
Is it done by the buffer head code? Is it done by iomap? What
about cases that use neither?
Thanks,
Miklos