Re: [PATCH v2] fuse: add inode generation number support

From: Russell Harmon

Date: Tue Sep 29 2026 - 16:56:04 EST


On Mon, Sep 28, 2026 at 9:54 AM Amir Goldstein <amir73il@xxxxxxxxx> wrote:
>
> On Mon, Sep 28, 2026 at 6:04 PM Russell Harmon <russ@xxxxxx> wrote:
> >
> > (re-sending in plaintext)
> >
> > Thanks for the review! Comments inline.
> >
> > On Mon, Sep 28, 2026 at 2:38 AM Amir Goldstein <amir73il@xxxxxxxxx> wrote:
> > >
> > > On Mon, Sep 28, 2026 at 3:44 AM Russell Harmon <russ@xxxxxx> wrote:
> > > >
> > > > This patch adds support for propagating the inode generation number from
> > > > the FUSE server to the kernel.
> > >
> > > Incomplete statement.
> > > Without the mention of GETATTR/SETATTR this is a misleading statement.
> > >
> > > > This is useful for exporting FUSE
> > > > filesystems over NFS, where the generation number is used to detect
> > > > stale file handles (ESTALE) when inodes are recycled.
> > >
> > > How exactly does it help?
> > > I am not trying to troll you, I am really curious. how?
> > >
> > > Context: I have been trying to improve FUSE NFS export support for a while
> > > I have built a library that provides reliable NFS export for FUSE passthrough fs
> > > for specific backing file system types [1].
> > >
> > > [1] https://github.com/amir73il/libfuse/tree/libfuse_passthrough/passthrough
> > >
> > > It is broadly understood that real NFS export support requires extending the
> > > FUSE protocol to identify objects using file handles and Luis has
> > > already started
> > > with this work [2]
> > >
> > > [2] https://lore.kernel.org/linux-fsdevel/20260225112439.27276-1-luis@xxxxxxxxxx/
> > >
> > > So my question is, what does adding generation id to GETATTR/SETATTR
> > > improve for FUSE filesystem writers that wish to export their filesystem to NFS?
> >
> > Long answer is https://russ.har.mn/blog/2026-04-09/fuse-loopback-is-incomplete
> >
> > Short answer is that inode on its own doesn't uniquely identify a
> > file. The unique identifier is inode + generation. Reason being that a
> > user can delete a file and create a new one, and the filesystem is
> > allowed to reuse the inode (but if it does it must use a different
> > generation).
>
> This is obvious. That's the reason LOOKUP already returns a generation.
>
> >
> > As for why it needs to be a part of GETATTR/SETATTR, consider that,
> > given the above, without a generation number such calls are actually
> > ambiguous. They don't uniquely identify a filesystem object. e.g.
> > "what are you asked to GETATTR/SETATTR _on_?"
> >
>
> You phrased the question correctly, but the same question
> applies to any other command, i.e.
> "What are you asked to OPEN/TRUNCATE _on_?"
>
> So you see, what's needed is to change the input argument
> of ALL the fuse requests from nodeid to nodeid+generation, or
> in the more generic form of NFS handles, a variable size object id blob.

Yes, fair, we'll need to change all the fuse requests.

> This is from the latest proposal for FUSEX protocol [3]
>
> struct fusex_id {
> u64 nodeid;
> /* will extend with file handle */
> };
>
> [3] https://lore.kernel.org/fuse-devel/20260429102058.1362965-1-mszeredi@xxxxxxxxxx/
>
> > > >
> > > > Key changes:
> > > > - Bump FUSE protocol version to 7.47.
> > > > - Repurpose the unused `dummy` field in `struct fuse_attr_out` as
> > > > `generation`.
> > > > - Add a `FUSE_ATTR_GENERATION` INIT flag with which the filesystem opts
> > > > into the kernel consuming that field. Gating on the protocol minor
> > > > version alone would break existing filesystems: the minor version only
> > > > reflects the library, not whether the individual filesystem fills the
> > > > field, and a zero there would look like a generation change for any
> > > > filesystem that reports nonzero generations in LOOKUP.
> > > > - Update `fuse_change_attributes` and related functions to accept and set
> > > > `inode->i_generation`.
> > > > - Populate `i_generation` from `LOOKUP`, `GETATTR`, and `READDIRPLUS`
> > > > responses.
> > > > - Detect nodeid recycling on `GETATTR` and `SETATTR` responses via
> > > > `fuse_stale_inode()`, the same check already used by the `LOOKUP` and
> > > > `READDIRPLUS` paths, and mark the inode bad (EIO) instead of merging
> > > > the recycled file's attributes into the existing, possibly still-open,
> > > > inode.
> > > > - Update `fuse_get_dentry` to validate the generation number against the
> > > > file handle, returning ESTALE on mismatch.
> > > > - Maintain backward compatibility: without `FUSE_ATTR_GENERATION` the
> > > > generation field in attr replies is ignored.
> > > >
> > > > Verification:
> > > > Tested with a QEMU harness in fuse-generation-qemu against a patched
> > > > libfuse (FUSE_CAP_ATTR_GENERATION, fuse_reply_attr_with_generation) and
> > > > its passthrough_ll example reporting real backing-filesystem generation
> > > > numbers. The suite verifies that:
> > > > 1. The generation from `LOOKUP` reaches `name_to_handle_at()` file
> > > > handles and matches the backing filesystem's FS_IOC_GETVERSION.
> > > > 2. `open_by_handle_at()` succeeds for a valid handle and fails with
> > > > ESTALE for a handle whose generation does not match, both while the
> > > > inode is cached and after cache eviction.
> > >
> > > 1 and 2 should work on upstream FUSE right?
> > > This is something worth mentioning.
> > >
> > > > 3. When a `GETATTR` reply reports a new generation for a cached inode
> > > > (inode recycling), the kernel marks the inode bad: fstat() on an
> > > > open fd fails with EIO, while a fresh path lookup recovers and
> > > > pre-recycling file handles fail with ESTALE.
> > >
> > > How did the inode recycle happen with passtrhough_ll which keeps
> > > open fds for fuse inodes?
> > > Something is missing from this test report.
> > > If you just used a mock filesystem which makes no sense in the real world
> > > then the value of this change is questionable.
> >
> > Ack, I'll dig more into this. Sending responses to your other
> > questions now and will reply again once I've got an answer.
> >
> >
> > > > Signed-off-by: Russell Harmon <russ@xxxxxx>
> > > > Assisted-by: Gemini:gemini-3.1
> > > > ---
> > > > v2:
> > > > - Bump FUSE_KERNEL_MINOR_VERSION to 47 to match the new 7.47 changelog
> > > > entry (v1 added the entry but left the minor at 46).
> > > >
> > > > v1: https://lore.kernel.org/all/20260927141437.1432584-1-russ@xxxxxx/
> > > >
> > > > Documentation/filesystems/fuse/fuse.rst | 32 +++++++++++++++++++++++++
> > > > fs/fuse/dir.c | 21 +++++++++++-----
> > > > fs/fuse/fuse_i.h | 10 ++++++--
> > > > fs/fuse/inode.c | 23 ++++++++++++------
> > > > fs/fuse/readdir.c | 2 +-
> > > > include/uapi/linux/fuse.h | 11 +++++++--
> > > > 6 files changed, 81 insertions(+), 18 deletions(-)
> > > >
> > > > diff --git a/Documentation/filesystems/fuse/fuse.rst b/Documentation/filesystems/fuse/fuse.rst
> > > > index 0fbd5a03fdc9..f67bc9fc6316 100644
> > > > --- a/Documentation/filesystems/fuse/fuse.rst
> > > > +++ b/Documentation/filesystems/fuse/fuse.rst
> > > > @@ -49,6 +49,38 @@ using the sftp protocol.
> > > > The userspace library and utilities are available from the
> > > > `FUSE homepage: <https://github.com/libfuse/>`_
> > > >
> > > > +NFS export support
> > > > +==================
> > > > +
> > > > +FUSE filesystems can be exported via NFS if the filesystem daemon supports it.
> > > > +For reliable NFS export, the filesystem should provide a unique inode
> > > > +generation number for each inode. This generation number is used by the
> > > > +NFS server to distinguish between different file instances that may
> > > > +share the same inode number (e.g. after an inode number is reused).
> > > > +
> > > > +The inode generation number is provided by the filesystem daemon in the
> > > > +following messages:
> > > > +
> > > > +- `FUSE_LOOKUP`
> > > > +- `FUSE_GETATTR` (see below)
> > > > +- `FUSE_SETATTR` (see below)
> > > > +- `FUSE_READDIRPLUS`
> > > > +- `FUSE_CREATE` / `FUSE_TMPFILE` / `FUSE_MKNOD` / `FUSE_MKDIR` / `FUSE_SYMLINK` / `FUSE_LINK`
> > > > +
> > > > +A daemon that keeps the generation number in its `FUSE_GETATTR` and
> > > > +`FUSE_SETATTR` replies (the `generation` field of `fuse_attr_out`,
> > > > +protocol 7.46) must announce this by setting `FUSE_ATTR_GENERATION` in
> > > > +its `FUSE_INIT` reply flags. When the flag is negotiated, the kernel
> > > > +compares the generation in every getattr/setattr reply against the
> > > > +cached inode: a mismatch means the daemon has reused the node ID for a
> > > > +different file, and the cached inode is marked bad (subsequent
> > > > +operations on it fail with EIO). Without the flag, the field is ignored
> > > > +and the generation is only taken from lookup-type replies, preserving
> > > > +the behavior of existing filesystems.
> > > > +
> > > > +If the filesystem daemon does not provide a generation number, the kernel
> > > > +will use a default value of 0.
> > > > +
> > >
> > > On the one hand, I still need to understand the value of reporting generation
> > > in GETATTR/SETATTR.
> > >
> > > On the other hand, I do see the value in the server negotiating at init time the
> > > fact that "Generation values are reliable".
> > >
> > > What happens today is that NFS exporting is allowed for all FUSE filesystems
> > > regardless of the reliability of generation id, so after inode evict
> > > and recycle,
> > > an NFSv3 client that had access to inode X.Y may get access to a completely
> > > different file or even a directory with inode X.Y, where X is the recycle nodeid
> > > and Y is an unreliable generation provided by the server.
> > >
> > > The problem is that FUSE does not require opt-in for NFS export, it only
> > > allows servers to opt-out of NFS export (FUSE_NO_EXPORT_SUPPORT).
> > >
> > > So what can be done given a declaration of the server that generation
> > > is reliable?
> > >
> > > One option is to set a non-zero uuid/fsid to the fuse filesystem.
> > > This will allow exporting the fuse filesystem without the opt-in uuid/fsid=
> > > in /etc/exports.
> > > This will also allow setting fanotify FAN_MARK_FILESYSTEM watches
> > > on this fuse filesystem, whose file handles could be trusted to be a genuine
> > > unique identity of the filesystem objects.
> > >
> > > But if we take this route, it is better to take it one step further
> > > and allow the
> > > server to determine the filesystem uuid/fsid during negotiation.
> > >
> > > In any case, I am not convinced there is value in doing all this
> > > without extending
> > > the protocol with LOOKUP_HANDLE/LOOKUPX lookup by file handle, so if you
> > > have compelling use cases, please spell them out.
> >
> > I'm trying to implement a caching filesystem which uses a sqlite
> > database (stored on an SSD) to cache all file and directory metadata,
> > ultimately to avoid HDD spinups (and the associated electricity cost)
> > for operations other than file reads. I'm also exporting this
> > filesystem over NFS, hence the need for NFS support.
>
> OK, so you are implementing a passthrough filesystem (to HDD)
> with a fast caching layer, I assume to cache results of GETATTR/
> LOOKUP/READDIR/READDIRPLUS? Anything else?
> Are you caching a specific type of backing filesystem?
>
> I advise you to look at libfuse_passthrough [1].
> You can either see if it fits you or learn from my experience with NFS
> export of a fuse passthrough filesystem.
> I am giving a talk on the subject in Linux Plumbers next week [4].
>
> >
> > If I remember correctly, name_to_handle_at on a FUSE filesystem
> > contains just the inode+generation, and open_by_handle_at results in a
> > LOOKUP of that inode (and now inode+generation). But I'll
> > double-check. Assuming my memory is correct, I think that's
> > sufficient, assuming that we consider inode+generation as a unique
> > file object identifier.
>
> The problem is, that what you assume is not an API definition, this is
> a specific
> filesystem implementation, which happens to be correct for some commonly
> used filesystems (e.g. ext4, xfs), but is not true in general.
> In general the API {name_to,open_by}_handle_at() the file handle is an
> opaque blob.

Maybe I'm thinking about this backwards, but we're not talking about
"any" filesystem here. We're talking about the FUSE filesystem where
the definition of a handle _is_ inode+generation. All I'm proposing we
do is to expose that fact to the FUSE userspace daemon. That daemon
can then turn around and call open on another filesystem, or make a
network call to AWS or whatever. In my case I want to
open_by_handle_at on an underlying filesystem, but I can solve that by
storing a mapping from fuse_ino+fuse_generation_num to
underlying_fs_file_handle in my database.


> The way that libfuse_passthrough gates against recycled inode numbers
> is that the file handle is stored in the server's inode table at lookup time
> (generation is recorded) and all the commands that follow, lookup in the
> server's inode table (could be in sqlite in your case) and use
> open_by_handle_at() with the recorded file handle from lookup time
> to get an O_PATH fd to the object on the backing filesystem (on HDD).
>
> This is called InodeRef in libfuse_passthrough and any request will return
> ESTALE to the kernel if it cannot get an InodeRef.
>
> Thanks,
> Amir.
>
> [4] https://lpc.events/event/20/contributions/2367