Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)
From: Jori Koolstra
Date: Wed Sep 30 2026 - 18:51:34 EST
> Op 29-09-2026 18:15 EDT schreef NeilBrown <neilb@xxxxxxxxxxx>:
>
>
> On Tue, 29 Sep 2026, Jori Koolstra wrote:
> > > Op 19-09-2026 02:01 CEST schreef NeilBrown <neilb@xxxxxxxxxxx>:
> > >
> > > >
> > > > nfsd_create_locked() used to do that before vfs_mkdir() could return a
> > > > dentry, but it doesn't any more. The reason was because
> > > > d_splice_alias() on might return a different dentry.
> > > > In this case we want the same dentry, but we need to do a lookup on it.
> > > >
> > > > I'd rather fix this in kernfs, but maybe that is a longer-term goal.
> > > >
> > > > The comment in kernfs_dop_revalidate() suggests the we should d_drop()
> > > > the negative dentry and d_alloc_parallel() a new one and ->lookup that.
> > > > I'm not certain that is needed if we keep the parent locked, but we
> > > > would need to be certain.
> > > > We at least need to d_drop() the dentry before ->lookup as ->lookup
> > > > cannot handle hashed dentries and a hashed-negative dentry is passed
> > > > to ->mkdir.
> > > >
> > > > I wonder if we could just disable O_CREATE|O_DIRECTORY on kernfs ....
> > > > probably not.
> > > >
> > > > Summary: I think that if vfs_mkdir() returns NULL (success) but the
> > > > dentry is negative, we need to d_drop() and call ->lookup with a big
> > > > comment about kernfs. But we need to double-check that this will do the
> > > > right thing with ->d_time (I think it will).
> > > > We also need to think carefully about races with
> > > > kernfs_dop_revalidate(), which could happen concurrently with the
> > > > ->lookup.
> > >
> > > I've thought a bit more about this ... I think that doing a lookup after
> > > the vfs_mkdir() results in a negative is a bit ugly. It assumes things
> > > about the fs that I would rather not assume.
> > >
> > > I would rather have the current proposed code check for a negative
> > > dentry, and fail with -EIO or similar.
> > >
> >
> > I just noticed that there's precedent for this in overlayfs in super.c:
> >
> > /* Weird filesystem returning with hashed negative (kernfs)? */
> > err = -EINVAL;
> > if (d_really_is_negative(work))
> > goto out_dput;
> >
> > Shall we just do this for current kernel release, then we can add support
> > later if wanted.
> >
> > (But let's do EOPNOTSUPP instead of EINVAL)
> >
> > What do you think?
>
> The problem with this approach is that open(.., O_CREAT|O_DIRECTORY)
> might create the directory, then return -EOPNOTSUPP. This is weird and
> I'd rather it not be visible.
>
Err, *derp*, what a stupid suggestion of mine.
> Currently O_DIRECTORY|O_CREAT results in -EINVAL. I would rather it
> remain a -EINVAL on any filesystem which doesn't completely support
> the functionality.
>
I don't think that works for the reason I just wrote in my email to Amir:
it would make lookup dependent on the dentry cache. If it's in-cache, you get
your dir, otherwise suddenly -EINVAL.
> To do that we need some way to detect kernfs and tracefs. I think
> the only way we can do that is to make some change to those two
> filesystems.
> Maybe a new SB_I_ flag in sb->s_iflags would be ok in the short term.
>
We can just implement atomic_open() for kernfs/tracefs, do a lookup there,
and if negative with O_CREAT return maybe -ENOENT (or really we need a new
error that says "the requested create could not be serviced," like -ENOCREATE,
or whatever). And if it is positive we do finish_no_open().
It's a bit of a hack because it does not really have anything to do with
atomicity, but it does short-circuit the mkdir call in lookup_open(). I guess
that would work. Maybe I am confused, but wasn't that what you proposed here
earlier?
Best,
Jori.