Re: [PATCH v6 07/12] vfs: add O_CREAT|O_DIRECTORY to open*(2)

From: NeilBrown

Date: Tue Sep 29 2026 - 18:16:34 EST


On Tue, 29 Sep 2026, Jori Koolstra wrote:
> > Op 19-09-2026 02:01 CEST schreef NeilBrown <neilb@xxxxxxxxxxx>:
> >
> > >
> > > nfsd_create_locked() used to do that before vfs_mkdir() could return a
> > > dentry, but it doesn't any more. The reason was because
> > > d_splice_alias() on might return a different dentry.
> > > In this case we want the same dentry, but we need to do a lookup on it.
> > >
> > > I'd rather fix this in kernfs, but maybe that is a longer-term goal.
> > >
> > > The comment in kernfs_dop_revalidate() suggests the we should d_drop()
> > > the negative dentry and d_alloc_parallel() a new one and ->lookup that.
> > > I'm not certain that is needed if we keep the parent locked, but we
> > > would need to be certain.
> > > We at least need to d_drop() the dentry before ->lookup as ->lookup
> > > cannot handle hashed dentries and a hashed-negative dentry is passed
> > > to ->mkdir.
> > >
> > > I wonder if we could just disable O_CREATE|O_DIRECTORY on kernfs ....
> > > probably not.
> > >
> > > Summary: I think that if vfs_mkdir() returns NULL (success) but the
> > > dentry is negative, we need to d_drop() and call ->lookup with a big
> > > comment about kernfs. But we need to double-check that this will do the
> > > right thing with ->d_time (I think it will).
> > > We also need to think carefully about races with
> > > kernfs_dop_revalidate(), which could happen concurrently with the
> > > ->lookup.
> >
> > I've thought a bit more about this ... I think that doing a lookup after
> > the vfs_mkdir() results in a negative is a bit ugly. It assumes things
> > about the fs that I would rather not assume.
> >
> > I would rather have the current proposed code check for a negative
> > dentry, and fail with -EIO or similar.
> >
>
> I just noticed that there's precedent for this in overlayfs in super.c:
>
> /* Weird filesystem returning with hashed negative (kernfs)? */
> err = -EINVAL;
> if (d_really_is_negative(work))
> goto out_dput;
>
> Shall we just do this for current kernel release, then we can add support
> later if wanted.
>
> (But let's do EOPNOTSUPP instead of EINVAL)
>
> What do you think?

The problem with this approach is that open(.., O_CREAT|O_DIRECTORY)
might create the directory, then return -EOPNOTSUPP. This is weird and
I'd rather it not be visible.

Currently O_DIRECTORY|O_CREAT results in -EINVAL. I would rather it
remain a -EINVAL on any filesystem which doesn't completely support
the functionality.

To do that we need some way to detect kernfs and tracefs. I think
the only way we can do that is to make some change to those two
filesystems.
Maybe a new SB_I_ flag in sb->s_iflags would be ok in the short term.

NeilBrown


>
> Thanks,
> Jori.
>