[PATCH v3 4/4] ntfs: do not give a contiguous cluster allocation a second run

From: Matthias Goergens

Date: Sun Oct 04 2026 - 05:57:04 EST


ntfs_attr_map_cluster() asks ntfs_cluster_alloc() for contiguous
clusters and merges the whole result into the runlist, but reports only
its first run, so its callers zero only that run. Within a bitmap page
the allocator keeps such a request to one run, because a cluster in use
ends it. The zone search can also move on without testing the clusters
after the run, though: it skips a bitmap page that has no free clusters,
it goes from the end of the zone to the start of pass 2, and once every
zone has been searched it shrinks the MFT zone by half and goes on at
its new end. After each of these jumps the first cluster tested is
still taken as the next one of the run, and if it is free it starts a
second run elsewhere on the volume.

fallocate() over a hole then leaves the clusters of that second run
holding whatever they held before: on a nearly full volume, 16 of 20
fallocated clusters read back a deleted file's data, in each of the
three cases.

Delayed allocation gets a second run without punching any holes: filling
the same volume with write() reaches the last case when the MFT zone is
shrunk. The allocator then takes only the first run off the
dirty-cluster count, and after the fill statfs() reported no free space
while 939 clusters were still free, until the next mount.

Before setting the bit of a cluster for a contiguous request, check that
the cluster comes right after the run, and return the run as it is if
not; no bit has to be cleared. ntfs_attr_map_cluster() already handles
a short run, which a cluster in use produces in the same way: its
callers ask again for the rest.

Fixes: 11ccc9107dc4 ("ntfs: update runlist handling and cluster allocator")
Suggested-by: Namjae Jeon <linkinjeon@xxxxxxxxxx>
Link: https://lore.kernel.org/all/CAKYAXd-agPFvDj2-HvjP3SLtG7AxkqUeNPqYyrLt-BN7fT8w6g@xxxxxxxxxxxxxx/
Signed-off-by: Matthias Goergens <matthias.goergens@xxxxxxxxx>
---
New in v3, after Namjae Jeon's question on v2 2/3. Patch 2's early
return for the run from the hint stays: without it the restarted search
would read the bitmap until it found a free cluster just to return the
run, and if there was none it would fail the whole request with ENOSPC.

The other contiguous callers are the attribute list allocations, which
free a result that is not one run of the full length, and
ntfs_write_cb() in compress.c, which writes the whole compression block
from the first cluster without looking at the length. That was already
wrong for a short run before this patch and is not changed here.

contig-paths.c and its guest init script init-paths in
https://github.com/matthiasgoergens/linux/tree/reproducer/2026-10-04-ntfs-nearfull-v3
set up each case; the README there has the results.
---
fs/ntfs/lcnalloc.c | 10 ++++++++++
1 file changed, 10 insertions(+)

diff --git a/fs/ntfs/lcnalloc.c b/fs/ntfs/lcnalloc.c
index c58e689fb585..3ebcac4c2864 100644
--- a/fs/ntfs/lcnalloc.c
+++ b/fs/ntfs/lcnalloc.c
@@ -375,6 +375,16 @@ struct runlist_element *ntfs_cluster_alloc(struct ntfs_volume *vol, const s64 st
has_guess = 1;
continue;
}
+ /*
+ * A contiguous request gets a single run. The scan can
+ * move on without testing the clusters after the run
+ * (past a bitmap page that is full, to pass 2 or to
+ * another zone), so return the run when this cluster
+ * does not follow it.
+ */
+ if (is_contig && rlpos &&
+ lcn + bmp_pos != prev_lcn + prev_run_len)
+ goto out;
/*
* Allocate more memory if needed, including space for
* the terminator element.
--
2.56.0