[PATCH v3 3/6] ntfs: fail the mount when $MFT needs its own extent records

From: Matthias Goergens

Date: Tue Sep 29 2026 - 23:51:24 EST


Mounting a crafted image hangs the mount process forever, unkillable at
0% CPU. Only the hung-task detector reports it:

INFO: task mount:74 blocked in I/O wait for more than 30 seconds.

folio_wait_bit_common <- waits forever
filemap_read_folio
map_mft_record_folio
map_mft_record
ntfs_map_runlist_nolock
ntfs_attr_vcn_to_rl
__ntfs_read_iomap_begin
iomap_read_folio
ntfs_read_folio <- already holds that folio's lock
ntfs_read_inode_mount
ntfs_fill_super

ntfs_read_inode_mount() assembles $MFT's runlist one $DATA extent at a
time, hoping, as its comment says, that it never needs a part of $MFT it
has not decoded yet. If $MFT's attribute list puts one of its
attributes or later $DATA extents in an extent record outside the
runlist decoded so far, reading that record re-enters
ntfs_map_runlist_nolock() for $MFT. Depending on the layout, that waits
on the folio lock it already holds, as above, blocks on $MFT's runlist
lock, or dereferences NULL in map_extent_mft_record().

Mark the volume for the whole $DATA enumeration and have
ntfs_map_runlist_nolock() refuse $MFT with -EIO while the mark is set.
The enumeration decodes each extent itself with
ntfs_mapping_pairs_decompress(), so it does not need that path.

A validly placed record can trip the check too. The driver puts an
extent record for $MFT's own mapping pairs before the first vcn it
describes, but with clusters smaller than a page, the folio holding it
can still run past the decoded runlist. An unpatched kernel hangs on
that layout as well; with the check the mount fails. The check then
refuses only the folio's tail, so this patch needs "ntfs: do not map an
unmappable runlist fragment as a hole": without it the tail is
zero-filled and the mount serves an all-zero mft record. The mount now
fails instead of hanging:

ntfs: (device vda): ntfs_map_runlist_nolock(): $MFT needs its own
extent records to describe itself; cannot mount.

Tested under qemu with KASAN, PROVE_LOCKING and the hung-task detector
on ntfs-next, with and without the whole series: six crafted images
that hang or crash an unpatched kernel fail to mount with the series,
and nine images that mount without it still mount, three of them with
the same file listing and contents, among them a volume written by
ntfs-3g whose $MFT has its $DATA in three extent records.

Fixes: b041ca562526 ("ntfs: update iomap and address space operations")
Suggested-by: Hyunchul Lee <hyc.lee@xxxxxxxxx>
Cc: stable@xxxxxxxxxxxxxxx
Signed-off-by: Matthias Goergens <matthias.goergens@xxxxxxxxx>
---
fs/ntfs/attrib.c | 11 +++++++++++
fs/ntfs/inode.c | 9 +++++++++
fs/ntfs/volume.h | 3 +++
3 files changed, 23 insertions(+)

diff --git a/fs/ntfs/attrib.c b/fs/ntfs/attrib.c
index 5f5a91265af5e..2e72cd816d04f 100644
--- a/fs/ntfs/attrib.c
+++ b/fs/ntfs/attrib.c
@@ -103,6 +103,17 @@ int ntfs_map_runlist_nolock(struct ntfs_inode *ni, s64 vcn, struct ntfs_attr_sea
base_ni = ni;
else
base_ni = ni->ext.base_ntfs_ino;
+ /*
+ * ntfs_read_inode_mount() builds $MFT's runlist itself, so nothing
+ * should reach here for $MFT. A crafted image can: the read that
+ * gets here already holds the $MFT folio lock it would wait on.
+ */
+ if (unlikely(NVolMftBootstrap(ni->vol) &&
+ base_ni == NTFS_I(ni->vol->mft_ino))) {
+ ntfs_error(ni->vol->sb,
+ "$MFT needs its own extent records to describe itself; cannot mount.");
+ return -EIO;
+ }
if (!ctx) {
ctx_is_temporary = ctx_needs_reset = true;
m = map_mft_record(base_ni);
diff --git a/fs/ntfs/inode.c b/fs/ntfs/inode.c
index 9583b2c6c7a26..a61f1519549cb 100644
--- a/fs/ntfs/inode.c
+++ b/fs/ntfs/inode.c
@@ -2080,6 +2080,11 @@ int ntfs_read_inode_mount(struct inode *vi)
/* Now load all attribute extents. */
a = NULL;
next_vcn = last_vcn = highest_vcn = 0;
+ /*
+ * Reading one of $MFT's own extent records in this loop can re-enter
+ * ntfs_map_runlist_nolock() for $MFT; see the check there.
+ */
+ NVolSetMftBootstrap(vol);
while (!(err = ntfs_attr_lookup(AT_DATA, NULL, 0, 0, next_vcn, NULL, 0,
ctx))) {
struct runlist_element *nrl;
@@ -2162,6 +2167,7 @@ int ntfs_read_inode_mount(struct inode *vi)
err = ntfs_read_locked_inode(vi);
if (err) {
ntfs_error(sb, "ntfs_read_inode() of $MFT failed.\n");
+ NVolClearMftBootstrap(vol);
ntfs_attr_put_search_ctx(ctx);
/* Revert to the safe super operations. */
kfree(m);
@@ -2195,6 +2201,7 @@ int ntfs_read_inode_mount(struct inode *vi)
goto put_err_out;
}
}
+ NVolClearMftBootstrap(vol);
if (err != -ENOENT) {
ntfs_error(sb, "Failed to lookup $MFT/$DATA attribute extent. Run chkdsk.\n");
goto put_err_out;
@@ -2229,6 +2236,8 @@ int ntfs_read_inode_mount(struct inode *vi)
put_err_out:
ntfs_attr_put_search_ctx(ctx);
err_out:
+ /* Also reached from inside the $DATA loop. */
+ NVolClearMftBootstrap(vol);
ntfs_error(sb, "Failed. Marking inode as bad.");
kfree(m);
return -1;
diff --git a/fs/ntfs/volume.h b/fs/ntfs/volume.h
index fdb57279de84c..0473b602084c5 100644
--- a/fs/ntfs/volume.h
+++ b/fs/ntfs/volume.h
@@ -194,6 +194,7 @@ struct ntfs_volume {
* NV_Discard Issue discard/TRIM commands for freed clusters.
* NV_DisableSparse Disable creation of sparse regions.
* NV_NativeSymlinkRel Translate absolute Windows reparse targets (native_symlink=rel).
+ * NV_MftBootstrap Mount is still assembling $MFT's own runlist.
*/
enum {
NV_Errors,
@@ -214,6 +215,7 @@ enum {
NV_DisableSparse,
NV_NativeSymlinkRel,
NV_SymlinkNative,
+ NV_MftBootstrap,
};

/*
@@ -253,6 +255,7 @@ DEFINE_NVOL_BIT_OPS(Discard)
DEFINE_NVOL_BIT_OPS(DisableSparse)
DEFINE_NVOL_BIT_OPS(NativeSymlinkRel)
DEFINE_NVOL_BIT_OPS(SymlinkNative)
+DEFINE_NVOL_BIT_OPS(MftBootstrap)

static inline void ntfs_inc_free_clusters(struct ntfs_volume *vol, s64 nr)
{
--
2.55.0