Re: [PATCH v2] ext4: use fsdata to track inline data write state and fix race
From: Aditya Prakash Srivastava
Date: Thu Jul 02 2026 - 07:57:35 EST
Thank You Jan for your valuable inputs. As you pointed out,
CONVERT_INLINE_DATA indeed becomes unused. I will be doing
that cleanup once we sort this one out.
Coming back to this patch, I will incorporate the suggestions, test it
and will send the next revision of this patch. I agree that
interactions with FALL_BACK_TO_NONDELALLOC introduce
unnecessary cross-over of two independent paths and should be
avoided altogether.
Thanks,
Aditya
On Thu, Jul 2, 2026 at 5:04 PM Jan Kara <jack@xxxxxxx> wrote:
>
> Hi Aditya!
>
> On Thu 02-07-26 16:19:56, Aditya Prakash Srivastava wrote:
> > I have been thinking about the VFS retry case where if we do
> > not reset the fsdata, we might end up in an infinite loop.
>
> Indeed, good spotting. I have also realized that after your patch
> CONVERT_INLINE_DATA is effectively unused and can be deleted along with
> passing fsdata to ext4_generic_write_inline_data() and
> ext4_da_convert_inline_data_to_extent(). So that would be a good cleanup
> patch on top of your fix.
>
> > The proposed fix from my side is to reintroduce setting
> > fsdata = NULL in the beginning of ext4_write_begin,
> > of course we should be careful not to overwrite it if it
> > has some other value. Sample code like below
> >
> > if (unlikely(ret))
> > return ret;
> >
> > /*
> > * Reset fsdata for a vfs retry case
> > */
> > if (*fsdata != (void *)FALL_BACK_TO_NONDELALLOC)
> > *fsdata = NULL;
>
> Hum, thinking more about this I think we should use fsdata more as flags
> than enum. So ext4_write_begin() should be doing
>
> *fsdata = (void *)((unsigned long)*fsdata | EXT4_WRITE_DATA_INLINE);
>
> and ext4_write_end() should be doing (unconditionally on exit):
>
> *fsdata = (void *)((unsigned long)*fsdata & ~EXT4_WRITE_DATA_INLINE);
>
> to unroll the state back. Otherwise the interactions with
> FALL_BACK_TO_NONDELALLOC and delalloc path get too tedious to verify and
> they *are* in fact independent states of the write path.
>
> Honza
>
> >
> > trace_ext4_write_begin(inode, pos, len);
> > /*
> > * Reserve one block more for addition to orphan list in case
> >
> > This should take care of the retry case not getting into
> > an infinite loop.
> > What do you think? I shall send you a revised version if we
> > agree on this.
> >
> > Thanks,
> > Aditya
> >
> > On Thu, Jul 2, 2026 at 4:05 PM Jan Kara <jack@xxxxxxx> wrote:
> > >
> > > On Thu 02-07-26 05:07:24, Aditya Srivastava wrote:
> > > > From: Aditya Prakash Srivastava <aditya.ansh182@xxxxxxxxx>
> > > >
> > > > Instead of checking the live inode state (ext4_has_inline_data(inode)
> > > > and ext4_test_inode_state(inode, EXT4_STATE_MAY_INLINE_DATA)) in the
> > > > write_end handlers, use the fsdata parameter of the address space
> > > > operations to explicitly pass down the state in which write_begin
> > > > prepared the write.
> > > >
> > > > A concurrent thread (such as ext4_page_mkwrite()) can convert the
> > > > inline data to an extent between write_begin and write_end. If this
> > > > happens, the write_end handlers would previously miss the inline
> > > > write_end path and fall through to extent-based write_end logic.
> > > > However, since block buffers were never allocated in write_begin,
> > > > this resulted in NULL pointer dereferences or data loss because
> > > > folio_buffers(folio) was NULL.
> > > >
> > > > Define EXT4_WRITE_DATA_INLINE (3) and communicate this state via
> > > > fsdata:
> > > > 1) ext4_write_begin() and ext4_da_write_begin() explicitly set
> > > > *fsdata to EXT4_WRITE_DATA_INLINE when an inline write is
> > > > successfully prepared.
> > > > 2) ext4_write_end(), ext4_journalled_write_end(), and
> > > > ext4_da_write_end() rely solely on fsdata / write_mode to
> > > > invoke ext4_write_inline_data_end().
> > > >
> > > > Furthermore, during a buffered write, ext4_write_inline_data_end()
> > > > acquires the xattr lock after preparing the write. If a concurrent
> > > > page fault (ext4_page_mkwrite()) converts the inline data to an extent
> > > > after the write_end handlers check the state but before
> > > > ext4_write_inline_data_end() acquires the xattr write lock, the
> > > > subsequent check will trigger a kernel panic via
> > > > BUG_ON(!ext4_has_inline_data(inode)).
> > > >
> > > > Replace the BUG_ON check in ext4_write_inline_data_end() with a graceful
> > > > error-handling retry path. If the inline data is cleared after locking
> > > > the xattr, we safely release all resources (releasing iloc.bh,
> > > > unlocking/putting the folio, stopping the active journal transaction
> > > > handle) and return 0 (VFS retry) to let the generic write path retry
> > > > the operation safely.
> > > >
> > > > Reported-by: syzbot+0c89d865531d053abb2d@xxxxxxxxxxxxxxxxxxxxxxxxx
> > > > Closes: https://syzkaller.appspot.com/bug?extid=0c89d865531d053abb2d
> > > > Fixes: 3fdcfb668fd7 ("ext4: add journalled write support for inline data")
> > > > Suggested-by: Jan Kara <jack@xxxxxxx>
> > > > Signed-off-by: Aditya Prakash Srivastava <aditya.ansh182@xxxxxxxxx>
> > >
> > > Looks good! Feel free to add:
> > >
> > > Reviewed-by: Jan Kara <jack@xxxxxxx>
> > >
> > > Honza
> > >
> > > > ---
> > > > Changes in v2:
> > > > - Folded the BUG_ON fix from the second patch into the first one to
> > > > ensure bisectability across git history, as suggested by Jan Kara.
> > > > - Removed the pointless initialization `*fsdata = NULL` on entry to
> > > > `ext4_write_begin()`.
> > > > - Removed the redundant check `if (fsdata)` in `ext4_write_begin()`.
> > > >
> > > > fs/ext4/ext4.h | 1 +
> > > > fs/ext4/inline.c | 14 +++++++++++++-
> > > > fs/ext4/inode.c | 18 +++++++++---------
> > > > 3 files changed, 23 insertions(+), 10 deletions(-)
> > > >
> > > > diff --git a/fs/ext4/ext4.h b/fs/ext4/ext4.h
> > > > index b37c136ea3ab..521bd5d6321c 100644
> > > > --- a/fs/ext4/ext4.h
> > > > +++ b/fs/ext4/ext4.h
> > > > @@ -3138,6 +3138,7 @@ int do_journal_get_write_access(handle_t *handle, struct inode *inode,
> > > > void ext4_set_inode_mapping_order(struct inode *inode);
> > > > #define FALL_BACK_TO_NONDELALLOC 1
> > > > #define CONVERT_INLINE_DATA 2
> > > > +#define EXT4_WRITE_DATA_INLINE 3
> > > >
> > > > typedef enum {
> > > > EXT4_IGET_NORMAL = 0,
> > > > diff --git a/fs/ext4/inline.c b/fs/ext4/inline.c
> > > > index 8045e4ff270c..cfd591dc1d9c 100644
> > > > --- a/fs/ext4/inline.c
> > > > +++ b/fs/ext4/inline.c
> > > > @@ -812,7 +812,19 @@ int ext4_write_inline_data_end(struct inode *inode, loff_t pos, unsigned len,
> > > > goto out;
> > > > }
> > > > ext4_write_lock_xattr(inode, &no_expand);
> > > > - BUG_ON(!ext4_has_inline_data(inode));
> > > > + /*
> > > > + * We could have raced with ext4_page_mkwrite() converting
> > > > + * the inode and clearing the inline data flag, so we just
> > > > + * release resources and retry the whole write.
> > > > + */
> > > > + if (unlikely(!ext4_has_inline_data(inode))) {
> > > > + ext4_write_unlock_xattr(inode, &no_expand);
> > > > + brelse(iloc.bh);
> > > > + folio_unlock(folio);
> > > > + folio_put(folio);
> > > > + ext4_journal_stop(handle);
> > > > + return 0;
> > > > + }
> > > >
> > > > /*
> > > > * ei->i_inline_off may have changed since
> > > > diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c
> > > > index ce99807c5f5b..d3d9e1999670 100644
> > > > --- a/fs/ext4/inode.c
> > > > +++ b/fs/ext4/inode.c
> > > > @@ -1316,8 +1316,10 @@ static int ext4_write_begin(const struct kiocb *iocb,
> > > > foliop);
> > > > if (ret < 0)
> > > > return ret;
> > > > - if (ret == 1)
> > > > + if (ret == 1) {
> > > > + *fsdata = (void *)EXT4_WRITE_DATA_INLINE;
> > > > return 0;
> > > > + }
> > > > }
> > > >
> > > > /*
> > > > @@ -1450,8 +1452,7 @@ static int ext4_write_end(const struct kiocb *iocb,
> > > >
> > > > trace_ext4_write_end(inode, pos, len, copied);
> > > >
> > > > - if (ext4_has_inline_data(inode) &&
> > > > - ext4_test_inode_state(inode, EXT4_STATE_MAY_INLINE_DATA))
> > > > + if (fsdata == (void *)EXT4_WRITE_DATA_INLINE)
> > > > return ext4_write_inline_data_end(inode, pos, len, copied,
> > > > folio);
> > > >
> > > > @@ -1560,8 +1561,7 @@ static int ext4_journalled_write_end(const struct kiocb *iocb,
> > > >
> > > > BUG_ON(!ext4_handle_valid(handle));
> > > >
> > > > - if (ext4_has_inline_data(inode) &&
> > > > - ext4_test_inode_state(inode, EXT4_STATE_MAY_INLINE_DATA))
> > > > + if (fsdata == (void *)EXT4_WRITE_DATA_INLINE)
> > > > return ext4_write_inline_data_end(inode, pos, len, copied,
> > > > folio);
> > > >
> > > > @@ -3161,8 +3161,10 @@ static int ext4_da_write_begin(const struct kiocb *iocb,
> > > > foliop, fsdata, true);
> > > > if (ret < 0)
> > > > return ret;
> > > > - if (ret == 1)
> > > > + if (ret == 1) {
> > > > + *fsdata = (void *)EXT4_WRITE_DATA_INLINE;
> > > > return 0;
> > > > + }
> > > > }
> > > >
> > > > retry:
> > > > @@ -3299,9 +3301,7 @@ static int ext4_da_write_end(const struct kiocb *iocb,
> > > >
> > > > trace_ext4_da_write_end(inode, pos, len, copied);
> > > >
> > > > - if (write_mode != CONVERT_INLINE_DATA &&
> > > > - ext4_test_inode_state(inode, EXT4_STATE_MAY_INLINE_DATA) &&
> > > > - ext4_has_inline_data(inode))
> > > > + if (write_mode == EXT4_WRITE_DATA_INLINE)
> > > > return ext4_write_inline_data_end(inode, pos, len, copied,
> > > > folio);
> > > >
> > > > --
> > > > 2.47.3
> > > >
> > > --
> > > Jan Kara <jack@xxxxxxxx>
> > > SUSE Labs, CR
> --
> Jan Kara <jack@xxxxxxxx>
> SUSE Labs, CR