Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
From: NeilBrown
Date: Wed Sep 30 2026 - 20:36:57 EST
On Thu, 01 Oct 2026, Jori Koolstra wrote:
> > Op 29-09-2026 18:15 EDT schreef NeilBrown <neilb@xxxxxxxxxxx>:
> >
> >
> > On Tue, 29 Sep 2026, Jori Koolstra wrote:
> > > > Op 19-09-2026 02:01 CEST schreef NeilBrown <neilb@xxxxxxxxxxx>:
> > > >
> > > > >
> > > > > nfsd_create_locked() used to do that before vfs_mkdir() could return a
> > > > > dentry, but it doesn't any more. The reason was because
> > > > > d_splice_alias() on might return a different dentry.
> > > > > In this case we want the same dentry, but we need to do a lookup on it.
> > > > >
> > > > > I'd rather fix this in kernfs, but maybe that is a longer-term goal.
> > > > >
> > > > > The comment in kernfs_dop_revalidate() suggests the we should d_drop()
> > > > > the negative dentry and d_alloc_parallel() a new one and ->lookup that.
> > > > > I'm not certain that is needed if we keep the parent locked, but we
> > > > > would need to be certain.
> > > > > We at least need to d_drop() the dentry before ->lookup as ->lookup
> > > > > cannot handle hashed dentries and a hashed-negative dentry is passed
> > > > > to ->mkdir.
> > > > >
> > > > > I wonder if we could just disable O_CREATE|O_DIRECTORY on kernfs ....
> > > > > probably not.
> > > > >
> > > > > Summary: I think that if vfs_mkdir() returns NULL (success) but the
> > > > > dentry is negative, we need to d_drop() and call ->lookup with a big
> > > > > comment about kernfs. But we need to double-check that this will do the
> > > > > right thing with ->d_time (I think it will).
> > > > > We also need to think carefully about races with
> > > > > kernfs_dop_revalidate(), which could happen concurrently with the
> > > > > ->lookup.
> > > >
> > > > I've thought a bit more about this ... I think that doing a lookup after
> > > > the vfs_mkdir() results in a negative is a bit ugly. It assumes things
> > > > about the fs that I would rather not assume.
> > > >
> > > > I would rather have the current proposed code check for a negative
> > > > dentry, and fail with -EIO or similar.
> > > >
> > >
> > > I just noticed that there's precedent for this in overlayfs in super.c:
> > >
> > > /* Weird filesystem returning with hashed negative (kernfs)? */
> > > err = -EINVAL;
> > > if (d_really_is_negative(work))
> > > goto out_dput;
> > >
> > > Shall we just do this for current kernel release, then we can add support
> > > later if wanted.
> > >
> > > (But let's do EOPNOTSUPP instead of EINVAL)
> > >
> > > What do you think?
> >
> > The problem with this approach is that open(.., O_CREAT|O_DIRECTORY)
> > might create the directory, then return -EOPNOTSUPP. This is weird and
> > I'd rather it not be visible.
> >
>
> Err, *derp*, what a stupid suggestion of mine.
>
> > Currently O_DIRECTORY|O_CREAT results in -EINVAL. I would rather it
> > remain a -EINVAL on any filesystem which doesn't completely support
> > the functionality.
> >
>
> I don't think that works for the reason I just wrote in my email to Amir:
> it would make lookup dependent on the dentry cache. If it's in-cache, you get
> your dir, otherwise suddenly -EINVAL.
>
> > To do that we need some way to detect kernfs and tracefs. I think
> > the only way we can do that is to make some change to those two
> > filesystems.
> > Maybe a new SB_I_ flag in sb->s_iflags would be ok in the short term.
> >
>
> We can just implement atomic_open() for kernfs/tracefs, do a lookup there,
> and if negative with O_CREAT return maybe -ENOENT (or really we need a new
> error that says "the requested create could not be serviced," like -ENOCREATE,
> or whatever). And if it is positive we do finish_no_open().
>
> It's a bit of a hack because it does not really have anything to do with
> atomicity, but it does short-circuit the mkdir call in lookup_open(). I guess
> that would work. Maybe I am confused, but wasn't that what you proposed here
> earlier?
Yes, it is what I proposed earlier. But I think it would require more
review and probably make it unrealistic to land this cycle. But I'm not
thinking it is unlikely to be ready this cycle any way.
I'm now wondering if we should keep ->atomic_open out of the loop and
always use ->mkdir to create a directory.
Based on your justification you probably always want O_EXCL and I would
be inclined to require that.
So if the dentry is in-lookup we call ->atomic_open(O_DIRECTORY). If
that succeeds - good. If it reports ENOENT or a negative dentry, then
we cal ->mkdir. If that succeeds with a positive dentry, we call
through to call ->open.
If ->mkdir succeeds with a negative dentry - we have the problem of
kernfs and tracefs. I'm leaning towards fixing those to do the lookup.
I don't think any filesystems *can* combine mkdir with open, so not
using ->atomic_open for the mkdir doesn't actually lose anything.
NeilBrown