mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v3 0/6] ntfs: fix the $MFT bootstrap hang and reads of unmapped runlist ranges
@ 2026-09-30  3:48 Matthias Goergens
  2026-09-30  3:48 ` [PATCH v3 1/6] ntfs: do not map an unmappable runlist fragment as a hole Matthias Goergens
                   ` (7 more replies)
  0 siblings, 8 replies; 9+ messages in thread
From: Matthias Goergens @ 2026-09-30  3:48 UTC (permalink / raw)
  To: Namjae Jeon, Hyunchul Lee; +Cc: Baolin Liu, ntfs, linux-kernel

Hi Hyunchul,

This is v3, with your review of v2 and Baolin's applied:

- Patch 1: a failed retry for a vcn at or beyond allocated_size returns
  the runlist end again, as in ntfs-next; only lookups below it fail
  with -EIO.  On its own, v2's patch 1 also failed lookups past EOF,
  for instance when reading a file's last folio with 512-byte clusters
  (Baolin).  Baolin suggested moving patch 4's allocated_size check
  before the retry instead, but that skips a retry ntfs-next makes: on
  a corrupt volume whose allocated_size is cut to where a file's last
  extent record starts, the rest of the file was then read as zeros,
  where ntfs-next and v3's patch 1 read the data (with the whole series,
  patch 6 rejects that file).
- Patch 4: the if statement you asked me to merge is gone, as the
  allocated_size check now sits in patch 1, and in
  ntfs_attr_vcn_to_rl() patch 4 only extends patch 1's -EIO to
  LCN_ENOENT.  The expansion rollback restores allocated_size under
  size_lock.
- Patch 6: the comment is gone, and $MFT's own non-resident attribute
  list is checked too (Baolin); a volume where that list claims more
  data than its allocation used to mount and now fails to.

Patches 2, 3 and 5 are unchanged.

The series fixes a hang at mount when $MFT needs its own extent records
(patch 3, which needs patch 1).  Patches 1 and 2 fix reads and writes
of a runlist range held in an extent record that cannot be read, which
returned zeros with no error or silently lost buffered writes, and
patch 4 does the same for a runlist that ends before allocated_size.
Patches 5 and 6 check the sizes of $MFT's data and of every non-resident
attribute against their allocation, as fs/ntfs3 does.

Tested under qemu with KASAN and the hung-task detector on ntfs-next,
with and without the series.  Nine crafted images that hang or crash
the mount without it fail to mount with it, and six that read zeros
that are not on disk return errors instead.  Eighteen images that mount
without the series, among them eight public test volumes written by
Windows or mkntfs, read every file the same with it, and six small
mkntfs images still mount.  v3 gives the same results as v2 on all of
them.  The images and the scripts that generate them are at

  https://github.com/matthiasgoergens/linux/tree/reproducer/2026-09-26-ntfs-mft-runlist

and, for the ordinary file of patch 4, the unaligned allocated_size of
patch 6 and three of the eighteen volumes that mount, at

  https://github.com/matthiasgoergens/linux/tree/reproducer/2026-09-26-ntfs-v2-extra

and, for v3's changes, at

  https://github.com/matthiasgoergens/linux/tree/reproducer/2026-09-30-ntfs-v3

v2: https://lore.kernel.org/all/cover.1790417653.git.matthias.goergens@gmail.com/
v1: https://lore.kernel.org/all/20260922153931.1976405-1-matthias.goergens@gmail.com/

Thanks,
Matthias

Matthias Goergens (6):
  ntfs: do not map an unmappable runlist fragment as a hole
  ntfs: do not turn an unmappable runlist fragment into delalloc on
    write
  ntfs: fail the mount when $MFT needs its own extent records
  ntfs: do not map a vcn as a hole when its runlist lookup failed
  ntfs: fail the mount when $MFT's data size exceeds its allocation
  ntfs: reject non-resident attributes whose sizes exceed their
    allocation

 fs/ntfs/attrib.c   | 51 +++++++++++++++++++++++++++++++++++++++++-----
 fs/ntfs/attrlist.c |  8 ++++++--
 fs/ntfs/inode.c    | 46 +++++++++++++++++++++++++++++++++++++++++
 fs/ntfs/iomap.c    | 16 +++++++++++++--
 fs/ntfs/layout.h   | 13 +++++++-----
 fs/ntfs/volume.h   |  3 +++
 6 files changed, 123 insertions(+), 14 deletions(-)


base-commit: 259abb551e2944998cad4214c201954ab1ac5c8d
-- 
2.55.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [PATCH v3 1/6] ntfs: do not map an unmappable runlist fragment as a hole
  2026-09-30  3:48 [PATCH v3 0/6] ntfs: fix the $MFT bootstrap hang and reads of unmapped runlist ranges Matthias Goergens
@ 2026-09-30  3:48 ` Matthias Goergens
  2026-09-30  3:48 ` [PATCH v3 2/6] ntfs: do not turn an unmappable runlist fragment into delalloc on write Matthias Goergens
                   ` (6 subsequent siblings)
  7 siblings, 0 replies; 9+ messages in thread
From: Matthias Goergens @ 2026-09-30  3:48 UTC (permalink / raw)
  To: Namjae Jeon, Hyunchul Lee; +Cc: Baolin Liu, ntfs, linux-kernel

When the extent mft record holding part of a file's runlist cannot be
read, ntfs_attr_vcn_to_rl() ignores the failed ntfs_map_runlist_nolock()
retry and returns with *lcn == LCN_RL_NOT_MAPPED.  The iomap read path
only rejects lcn < LCN_ENOENT, so it maps the range as a hole and read()
returns zeros with no error.  The first read already does this:
ntfs_attr_map_whole_runlist() keeps the fragments it could read,
readahead drops its error, and the next lookup finds the unmapped tail.

On a file whose runlist is split between its base record (vcn 0-214) and
one extent record (vcn 215-1499), breaking only the extent record's FILE
magic makes the kernel log "Failed to map extent mft record", yet read()
returns 1285 clusters of zeros for vcn 215-1499.  A damaged attribute
list entry, which makes the retry fail with -ENOENT, gives the same
zeros.

Return -ENOMEM if the retry ran out of memory and -EIO otherwise.  Both
callers already handle an ERR_PTR; other negative lcns, including
LCN_ENOENT, are returned as before, and so is LCN_RL_NOT_MAPPED at or
beyond allocated_size, where nothing is mapped: the runlist ends there
with LCN_RL_NOT_MAPPED when only the last extent is mapped, as after a
write into it, and with clusters smaller than a page every read of a
file's last folio looks up such vcns.

Reads of vcn 215-1499 now fail with -EIO.  An intact volume exercised
with buffered, mmap and O_DIRECT I/O, fallocate, truncate, sparse and
compressed files behaves as before.

Fixes: 495e90fa3348 ("ntfs: update attrib operations")
Cc: stable@vger.kernel.org
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
---
 fs/ntfs/attrib.c | 24 ++++++++++++++++++++++--
 1 file changed, 22 insertions(+), 2 deletions(-)

diff --git a/fs/ntfs/attrib.c b/fs/ntfs/attrib.c
index 333b3371acb47..5f5a91265af5e 100644
--- a/fs/ntfs/attrib.c
+++ b/fs/ntfs/attrib.c
@@ -317,7 +317,7 @@ int ntfs_map_runlist(struct ntfs_inode *ni, s64 vcn)
 struct runlist_element *ntfs_attr_vcn_to_rl(struct ntfs_inode *ni, s64 vcn, s64 *lcn)
 {
 	struct runlist_element *rl = ni->runlist.rl;
-	int err;
+	int err = 0;
 	bool is_retry = false;
 
 	if (!rl) {
@@ -335,12 +335,32 @@ struct runlist_element *ntfs_attr_vcn_to_rl(struct ntfs_inode *ni, s64 vcn, s64
 
 	if (*lcn <= LCN_RL_NOT_MAPPED && is_retry == false) {
 		is_retry = true;
-		if (!ntfs_map_runlist_nolock(ni, vcn, NULL)) {
+		err = ntfs_map_runlist_nolock(ni, vcn, NULL);
+		if (!err) {
 			rl = ni->runlist.rl;
 			goto remap_rl;
 		}
 	}
 
+	/*
+	 * The runlist fragment containing @vcn could not be mapped, e.g.
+	 * because the extent mft record holding it is corrupt.  Do not hand
+	 * LCN_RL_NOT_MAPPED back to callers, which would treat it as a hole.
+	 * At or beyond the allocated size nothing is mapped, and the runlist
+	 * ends there with LCN_RL_NOT_MAPPED if only a later extent has been
+	 * mapped, so return that end as it is.
+	 */
+	if (*lcn == LCN_RL_NOT_MAPPED) {
+		unsigned long flags;
+		s64 allocated_size;
+
+		read_lock_irqsave(&ni->size_lock, flags);
+		allocated_size = ni->allocated_size;
+		read_unlock_irqrestore(&ni->size_lock, flags);
+		if ((s64)ntfs_cluster_to_bytes(ni->vol, vcn) < allocated_size)
+			return ERR_PTR(err == -ENOMEM ? -ENOMEM : -EIO);
+	}
+
 	return rl;
 }
 
-- 
2.55.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [PATCH v3 2/6] ntfs: do not turn an unmappable runlist fragment into delalloc on write
  2026-09-30  3:48 [PATCH v3 0/6] ntfs: fix the $MFT bootstrap hang and reads of unmapped runlist ranges Matthias Goergens
  2026-09-30  3:48 ` [PATCH v3 1/6] ntfs: do not map an unmappable runlist fragment as a hole Matthias Goergens
@ 2026-09-30  3:48 ` Matthias Goergens
  2026-09-30  3:48 ` [PATCH v3 3/6] ntfs: fail the mount when $MFT needs its own extent records Matthias Goergens
                   ` (5 subsequent siblings)
  7 siblings, 0 replies; 9+ messages in thread
From: Matthias Goergens @ 2026-09-30  3:48 UTC (permalink / raw)
  To: Namjae Jeon, Hyunchul Lee; +Cc: Baolin Liu, ntfs, linux-kernel

ntfs_write_simple_iomap_begin_non_resident() has the same unchecked
retry.  LCN_RL_NOT_MAPPED passes its "lcn <= LCN_HOLE" test, so a
buffered write into a range whose extent record cannot be read is
merged into the runlist as LCN_DELALLOC over clusters that are already
allocated on disk, and write() succeeds.  If the record stays
unreadable, writeback fails too and the data never reaches the disk;
the error only shows up at a later fsync() or a synchronous write.

On the volume from the previous patch, after a read of the file, an
8 KiB pwrite() at vcn 1000 returns 8192, fsync() returns -EIO, and
after remount the range is unchanged.

Fail the lookup here too, with -ENOMEM or -EIO.  That pwrite() now
fails with -EIO.  Writes whose range can be mapped are unaffected.

Fixes: b041ca562526 ("ntfs: update iomap and address space operations")
Cc: stable@vger.kernel.org
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
---
 fs/ntfs/iomap.c | 16 ++++++++++++++--
 1 file changed, 14 insertions(+), 2 deletions(-)

diff --git a/fs/ntfs/iomap.c b/fs/ntfs/iomap.c
index c812d7f19b360..b4475963e57a7 100644
--- a/fs/ntfs/iomap.c
+++ b/fs/ntfs/iomap.c
@@ -394,7 +394,7 @@ static int ntfs_write_simple_iomap_begin_non_resident(struct inode *inode, loff_
 	loff_t vcn_ofs, rl_length;
 	struct runlist_element *rl, *rlc;
 	bool is_retry = false;
-	int err = 0;
+	int err = 0, map_err = 0;
 	s64 vcn, lcn;
 	s64 max_clu_count =
 		ntfs_bytes_to_cluster(vol, round_up(length, vol->cluster_size));
@@ -429,12 +429,24 @@ static int ntfs_write_simple_iomap_begin_non_resident(struct inode *inode, loff_
 
 	if (lcn <= LCN_RL_NOT_MAPPED && is_retry == false) {
 		is_retry = true;
-		if (!ntfs_map_runlist_nolock(ni, vcn, NULL)) {
+		map_err = ntfs_map_runlist_nolock(ni, vcn, NULL);
+		if (!map_err) {
 			rl = ni->runlist.rl;
 			goto remap_rl;
 		}
 	}
 
+	/*
+	 * As in ntfs_attr_vcn_to_rl(): a runlist fragment that could not be
+	 * mapped is not a hole.  Treating it as one would put a delalloc
+	 * extent over clusters that are allocated on disk but unknown to us.
+	 */
+	if (lcn == LCN_RL_NOT_MAPPED) {
+		up_write(&ni->runlist.lock);
+		mutex_unlock(&ni->mrec_lock);
+		return map_err == -ENOMEM ? -ENOMEM : -EIO;
+	}
+
 	max_clu_count = min(max_clu_count, rl->length - (vcn - rl->vcn));
 	if (max_clu_count == 0) {
 		ntfs_error(inode->i_sb,
-- 
2.55.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [PATCH v3 3/6] ntfs: fail the mount when $MFT needs its own extent records
  2026-09-30  3:48 [PATCH v3 0/6] ntfs: fix the $MFT bootstrap hang and reads of unmapped runlist ranges Matthias Goergens
  2026-09-30  3:48 ` [PATCH v3 1/6] ntfs: do not map an unmappable runlist fragment as a hole Matthias Goergens
  2026-09-30  3:48 ` [PATCH v3 2/6] ntfs: do not turn an unmappable runlist fragment into delalloc on write Matthias Goergens
@ 2026-09-30  3:48 ` Matthias Goergens
  2026-09-30  3:48 ` [PATCH v3 4/6] ntfs: do not map a vcn as a hole when its runlist lookup failed Matthias Goergens
                   ` (4 subsequent siblings)
  7 siblings, 0 replies; 9+ messages in thread
From: Matthias Goergens @ 2026-09-30  3:48 UTC (permalink / raw)
  To: Namjae Jeon, Hyunchul Lee; +Cc: Baolin Liu, ntfs, linux-kernel

Mounting a crafted image hangs the mount process forever, unkillable at
0% CPU.  Only the hung-task detector reports it:

  INFO: task mount:74 blocked in I/O wait for more than 30 seconds.

  folio_wait_bit_common      <- waits forever
  filemap_read_folio
  map_mft_record_folio
  map_mft_record
  ntfs_map_runlist_nolock
  ntfs_attr_vcn_to_rl
  __ntfs_read_iomap_begin
  iomap_read_folio
  ntfs_read_folio            <- already holds that folio's lock
  ntfs_read_inode_mount
  ntfs_fill_super

ntfs_read_inode_mount() assembles $MFT's runlist one $DATA extent at a
time, hoping, as its comment says, that it never needs a part of $MFT it
has not decoded yet.  If $MFT's attribute list puts one of its
attributes or later $DATA extents in an extent record outside the
runlist decoded so far, reading that record re-enters
ntfs_map_runlist_nolock() for $MFT.  Depending on the layout, that waits
on the folio lock it already holds, as above, blocks on $MFT's runlist
lock, or dereferences NULL in map_extent_mft_record().

Mark the volume for the whole $DATA enumeration and have
ntfs_map_runlist_nolock() refuse $MFT with -EIO while the mark is set.
The enumeration decodes each extent itself with
ntfs_mapping_pairs_decompress(), so it does not need that path.

A validly placed record can trip the check too.  The driver puts an
extent record for $MFT's own mapping pairs before the first vcn it
describes, but with clusters smaller than a page, the folio holding it
can still run past the decoded runlist.  An unpatched kernel hangs on
that layout as well; with the check the mount fails.  The check then
refuses only the folio's tail, so this patch needs "ntfs: do not map an
unmappable runlist fragment as a hole": without it the tail is
zero-filled and the mount serves an all-zero mft record.  The mount now
fails instead of hanging:

  ntfs: (device vda): ntfs_map_runlist_nolock(): $MFT needs its own
    extent records to describe itself; cannot mount.

Tested under qemu with KASAN, PROVE_LOCKING and the hung-task detector
on ntfs-next, with and without the whole series: six crafted images
that hang or crash an unpatched kernel fail to mount with the series,
and nine images that mount without it still mount, three of them with
the same file listing and contents, among them a volume written by
ntfs-3g whose $MFT has its $DATA in three extent records.

Fixes: b041ca562526 ("ntfs: update iomap and address space operations")
Suggested-by: Hyunchul Lee <hyc.lee@gmail.com>
Cc: stable@vger.kernel.org
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
---
 fs/ntfs/attrib.c | 11 +++++++++++
 fs/ntfs/inode.c  |  9 +++++++++
 fs/ntfs/volume.h |  3 +++
 3 files changed, 23 insertions(+)

diff --git a/fs/ntfs/attrib.c b/fs/ntfs/attrib.c
index 5f5a91265af5e..2e72cd816d04f 100644
--- a/fs/ntfs/attrib.c
+++ b/fs/ntfs/attrib.c
@@ -103,6 +103,17 @@ int ntfs_map_runlist_nolock(struct ntfs_inode *ni, s64 vcn, struct ntfs_attr_sea
 		base_ni = ni;
 	else
 		base_ni = ni->ext.base_ntfs_ino;
+	/*
+	 * ntfs_read_inode_mount() builds $MFT's runlist itself, so nothing
+	 * should reach here for $MFT.  A crafted image can: the read that
+	 * gets here already holds the $MFT folio lock it would wait on.
+	 */
+	if (unlikely(NVolMftBootstrap(ni->vol) &&
+		     base_ni == NTFS_I(ni->vol->mft_ino))) {
+		ntfs_error(ni->vol->sb,
+			   "$MFT needs its own extent records to describe itself; cannot mount.");
+		return -EIO;
+	}
 	if (!ctx) {
 		ctx_is_temporary = ctx_needs_reset = true;
 		m = map_mft_record(base_ni);
diff --git a/fs/ntfs/inode.c b/fs/ntfs/inode.c
index 9583b2c6c7a26..a61f1519549cb 100644
--- a/fs/ntfs/inode.c
+++ b/fs/ntfs/inode.c
@@ -2080,6 +2080,11 @@ int ntfs_read_inode_mount(struct inode *vi)
 	/* Now load all attribute extents. */
 	a = NULL;
 	next_vcn = last_vcn = highest_vcn = 0;
+	/*
+	 * Reading one of $MFT's own extent records in this loop can re-enter
+	 * ntfs_map_runlist_nolock() for $MFT; see the check there.
+	 */
+	NVolSetMftBootstrap(vol);
 	while (!(err = ntfs_attr_lookup(AT_DATA, NULL, 0, 0, next_vcn, NULL, 0,
 			ctx))) {
 		struct runlist_element *nrl;
@@ -2162,6 +2167,7 @@ int ntfs_read_inode_mount(struct inode *vi)
 			err = ntfs_read_locked_inode(vi);
 			if (err) {
 				ntfs_error(sb, "ntfs_read_inode() of $MFT failed.\n");
+				NVolClearMftBootstrap(vol);
 				ntfs_attr_put_search_ctx(ctx);
 				/* Revert to the safe super operations. */
 				kfree(m);
@@ -2195,6 +2201,7 @@ int ntfs_read_inode_mount(struct inode *vi)
 			goto put_err_out;
 		}
 	}
+	NVolClearMftBootstrap(vol);
 	if (err != -ENOENT) {
 		ntfs_error(sb, "Failed to lookup $MFT/$DATA attribute extent. Run chkdsk.\n");
 		goto put_err_out;
@@ -2229,6 +2236,8 @@ int ntfs_read_inode_mount(struct inode *vi)
 put_err_out:
 	ntfs_attr_put_search_ctx(ctx);
 err_out:
+	/* Also reached from inside the $DATA loop. */
+	NVolClearMftBootstrap(vol);
 	ntfs_error(sb, "Failed. Marking inode as bad.");
 	kfree(m);
 	return -1;
diff --git a/fs/ntfs/volume.h b/fs/ntfs/volume.h
index fdb57279de84c..0473b602084c5 100644
--- a/fs/ntfs/volume.h
+++ b/fs/ntfs/volume.h
@@ -194,6 +194,7 @@ struct ntfs_volume {
  * NV_Discard			Issue discard/TRIM commands for freed clusters.
  * NV_DisableSparse		Disable creation of sparse regions.
  * NV_NativeSymlinkRel		Translate absolute Windows reparse targets (native_symlink=rel).
+ * NV_MftBootstrap		Mount is still assembling $MFT's own runlist.
  */
 enum {
 	NV_Errors,
@@ -214,6 +215,7 @@ enum {
 	NV_DisableSparse,
 	NV_NativeSymlinkRel,
 	NV_SymlinkNative,
+	NV_MftBootstrap,
 };
 
 /*
@@ -253,6 +255,7 @@ DEFINE_NVOL_BIT_OPS(Discard)
 DEFINE_NVOL_BIT_OPS(DisableSparse)
 DEFINE_NVOL_BIT_OPS(NativeSymlinkRel)
 DEFINE_NVOL_BIT_OPS(SymlinkNative)
+DEFINE_NVOL_BIT_OPS(MftBootstrap)
 
 static inline void ntfs_inc_free_clusters(struct ntfs_volume *vol, s64 nr)
 {
-- 
2.55.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [PATCH v3 4/6] ntfs: do not map a vcn as a hole when its runlist lookup failed
  2026-09-30  3:48 [PATCH v3 0/6] ntfs: fix the $MFT bootstrap hang and reads of unmapped runlist ranges Matthias Goergens
                   ` (2 preceding siblings ...)
  2026-09-30  3:48 ` [PATCH v3 3/6] ntfs: fail the mount when $MFT needs its own extent records Matthias Goergens
@ 2026-09-30  3:48 ` Matthias Goergens
  2026-09-30  3:48 ` [PATCH v3 5/6] ntfs: fail the mount when $MFT's data size exceeds its allocation Matthias Goergens
                   ` (3 subsequent siblings)
  7 siblings, 0 replies; 9+ messages in thread
From: Matthias Goergens @ 2026-09-30  3:48 UTC (permalink / raw)
  To: Namjae Jeon, Hyunchul Lee; +Cc: Baolin Liu, ntfs, linux-kernel

ntfs_attr_vcn_to_rl() retries ntfs_map_runlist_nolock() for any lcn up
to LCN_RL_NOT_MAPPED, which includes LCN_ENOENT, but turns a failed
retry into an error only for LCN_RL_NOT_MAPPED.  For LCN_ENOENT the
error is dropped and the read path maps the range as a hole.

An LCN_ENOENT below allocated_size comes from a base extent with a
highest_vcn of 0, which ntfs_mapping_pairs_decompress() takes to map the
whole attribute, so the runlist ends after its last mapping pair.  If
the pairs end early, the retry finds the same extent and fails with
-ENOENT.  On a crafted volume with 4 KiB clusters, a 64-cluster file
whose mapping pairs stop after 16 clusters reads 48 clusters of zeros,
with no error.

The same layout gets a crafted $MFT past the check from "ntfs: fail the
mount when $MFT needs its own extent records".  With 512-byte clusters
and $MFT's mapping pairs ending at vcn 4, an unpatched kernel hangs on
the folio lock reading records 0-3.  With the check alone, the -EIO is
dropped, records 2 and 3 read as zeros and the mount carries on until
check_mft_mirror() finds the zeroed record 2.

Fail the lookup whenever the retry leaves @vcn unmapped below
allocated_size, -ENOENT included.  At or beyond allocated_size the end
of the runlist is still returned as it is.

A failed expansion in ntfs_non_resident_attr_expand() or
ntfs_attrlist_repack() truncates the runlist under the runlist lock but
restores allocated_size only after dropping it.  A lookup in between
would now fail, so restore allocated_size before dropping the lock in
both, under size_lock, which ntfs_attr_vcn_to_rl() takes to read it.

The crafted file now fails from vcn 16 on with -EIO, and the crafted
volume fails to mount with the check's message.

Fixes: 495e90fa3348 ("ntfs: update attrib operations")
Cc: stable@vger.kernel.org
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
---
 fs/ntfs/attrib.c   | 30 ++++++++++++++++++++----------
 fs/ntfs/attrlist.c |  8 ++++++--
 2 files changed, 26 insertions(+), 12 deletions(-)

diff --git a/fs/ntfs/attrib.c b/fs/ntfs/attrib.c
index 2e72cd816d04f..5439c12f92808 100644
--- a/fs/ntfs/attrib.c
+++ b/fs/ntfs/attrib.c
@@ -354,14 +354,16 @@ struct runlist_element *ntfs_attr_vcn_to_rl(struct ntfs_inode *ni, s64 vcn, s64
 	}
 
 	/*
-	 * The runlist fragment containing @vcn could not be mapped, e.g.
-	 * because the extent mft record holding it is corrupt.  Do not hand
-	 * LCN_RL_NOT_MAPPED back to callers, which would treat it as a hole.
-	 * At or beyond the allocated size nothing is mapped, and the runlist
-	 * ends there with LCN_RL_NOT_MAPPED if only a later extent has been
-	 * mapped, so return that end as it is.
+	 * Neither the runlist nor the retry mapped @vcn, e.g. because the
+	 * extent mft record holding it is corrupt or because the mapping
+	 * pairs end too soon.  ntfs_map_runlist_nolock() reports the latter
+	 * as -ENOENT, as @vcn lies past the extent it found.  Below the
+	 * allocated size, callers would treat LCN_RL_NOT_MAPPED or LCN_ENOENT
+	 * as a hole, so fail instead.  At or beyond it nothing is mapped: the
+	 * runlist ends there with LCN_ENOENT, or with LCN_RL_NOT_MAPPED if
+	 * only a later extent has been mapped, so return that end as it is.
 	 */
-	if (*lcn == LCN_RL_NOT_MAPPED) {
+	if (*lcn <= LCN_RL_NOT_MAPPED) {
 		unsigned long flags;
 		s64 allocated_size;
 
@@ -4484,6 +4486,7 @@ static int ntfs_non_resident_attr_expand(struct ntfs_inode *ni, const s64 newsiz
 	struct ntfs_inode *base_ni;
 	struct super_block *sb = ni->vol->sb;
 	size_t new_rl_count;
+	unsigned long flags;
 
 	ntfs_debug("Inode 0x%llx, attr 0x%x, new size %lld old size %lld\n",
 			(unsigned long long)ni->mft_no, ni->type,
@@ -4714,11 +4717,20 @@ static int ntfs_non_resident_attr_expand(struct ntfs_inode *ni, const s64 newsiz
 	if (err2)
 		ntfs_debug("Leaking clusters");
 
-	/* Now, truncate the runlist itself. */
+	/*
+	 * Now, truncate the runlist itself.  Restore allocated_size before
+	 * dropping the lock: ntfs_attr_vcn_to_rl() fails a lookup below the
+	 * allocated size that falls past the end of the runlist.
+	 */
 	if (ni != locked_ni)
 		down_write(&ni->runlist.lock);
 	err2 = ntfs_rl_truncate_nolock(vol, &ni->runlist,
 			ntfs_bytes_to_cluster(vol, org_alloc_size));
+	if (!err2) {
+		write_lock_irqsave(&ni->size_lock, flags);
+		ni->allocated_size = org_alloc_size;
+		write_unlock_irqrestore(&ni->size_lock, flags);
+	}
 	if (ni != locked_ni)
 		up_write(&ni->runlist.lock);
 	if (err2) {
@@ -4730,8 +4742,6 @@ static int ntfs_non_resident_attr_expand(struct ntfs_inode *ni, const s64 newsiz
 		ni->runlist.rl = NULL;
 		ntfs_error(sb, "Couldn't truncate runlist. Rollback failed");
 	} else {
-		/* Prepare to mapping pairs update. */
-		ni->allocated_size = org_alloc_size;
 		/* Restore mapping pairs. */
 		if (ni != locked_ni)
 			down_read(&ni->runlist.lock);
diff --git a/fs/ntfs/attrlist.c b/fs/ntfs/attrlist.c
index 1bbd2bc62c582..3660e7fd24b13 100644
--- a/fs/ntfs/attrlist.c
+++ b/fs/ntfs/attrlist.c
@@ -168,14 +168,18 @@ static int ntfs_attrlist_repack(struct inode *attr_vi,
 	return 0;
 
 restore_old_runlist:
+	/*
+	 * Restore allocated_size before dropping the runlist lock:
+	 * ntfs_attr_vcn_to_rl() fails a lookup below the allocated size that
+	 * falls past the end of the runlist.
+	 */
 	down_write(&attr_ni->runlist.lock);
 	attr_ni->runlist.rl = old_rl;
 	attr_ni->runlist.count = old_rl_count;
-	up_write(&attr_ni->runlist.lock);
-
 	write_lock_irqsave(&attr_ni->size_lock, flags);
 	attr_ni->allocated_size = old_alloc_size;
 	write_unlock_irqrestore(&attr_ni->size_lock, flags);
+	up_write(&attr_ni->runlist.lock);
 
 	restore_err = ntfs_attr_update_mapping_pairs_locked(
 			attr_ni, 0, locked_ni);
-- 
2.55.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [PATCH v3 5/6] ntfs: fail the mount when $MFT's data size exceeds its allocation
  2026-09-30  3:48 [PATCH v3 0/6] ntfs: fix the $MFT bootstrap hang and reads of unmapped runlist ranges Matthias Goergens
                   ` (3 preceding siblings ...)
  2026-09-30  3:48 ` [PATCH v3 4/6] ntfs: do not map a vcn as a hole when its runlist lookup failed Matthias Goergens
@ 2026-09-30  3:48 ` Matthias Goergens
  2026-09-30  3:48 ` [PATCH v3 6/6] ntfs: reject non-resident attributes whose sizes exceed their allocation Matthias Goergens
                   ` (2 subsequent siblings)
  7 siblings, 0 replies; 9+ messages in thread
From: Matthias Goergens @ 2026-09-30  3:48 UTC (permalink / raw)
  To: Namjae Jeon, Hyunchul Lee; +Cc: Baolin Liu, ntfs, linux-kernel

ntfs_read_inode_mount() takes $MFT's i_size from data_size without
checking it against allocated_size, unlike the resident case in
ntfs_read_locked_inode().  mft records between the two are not on disk.

On a crafted volume with 512-byte clusters whose $MFT is cut to 4
clusters while data_size still says 27 records, reading the folio
holding records 0-3 finds the end of the runlist at vcn 4.  An unpatched
kernel hangs there on the folio lock, as in "ntfs: fail the mount when
$MFT needs its own extent records".  With the previous patch vcn 4-7 are
past the allocation and so read as a hole: records 2 and 3 come back as
zeros and the mount carries on until check_mft_mirror() finds the zeroed
record 2.

Refuse such a $MFT before anything is read through it.  The mount now
fails with:

  ntfs: (device vda): ntfs_read_inode_mount(): $MFT data size 27648
    exceeds its allocated size 2048. $MFT is corrupt. Run chkdsk.

fs/ntfs3 rejects any non-resident attribute whose data_size exceeds its
allocated size.  This patch checks only $MFT; the next one covers the
other non-resident attributes.

Fixes: b041ca562526 ("ntfs: update iomap and address space operations")
Cc: stable@vger.kernel.org
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
---
 fs/ntfs/inode.c | 10 ++++++++++
 1 file changed, 10 insertions(+)

diff --git a/fs/ntfs/inode.c b/fs/ntfs/inode.c
index a61f1519549cb..9c97fc3f9e5ff 100644
--- a/fs/ntfs/inode.c
+++ b/fs/ntfs/inode.c
@@ -2136,6 +2136,16 @@ int ntfs_read_inode_mount(struct inode *vi)
 			vi->i_size = le64_to_cpu(a->data.non_resident.data_size);
 			ni->initialized_size = le64_to_cpu(a->data.non_resident.initialized_size);
 			ni->allocated_size = le64_to_cpu(a->data.non_resident.allocated_size);
+			/*
+			 * Records between allocated_size and data_size are not
+			 * on disk, and would be read as zeros.
+			 */
+			if (vi->i_size > ni->allocated_size) {
+				ntfs_error(sb,
+					   "$MFT data size %lld exceeds its allocated size %lld. $MFT is corrupt. Run chkdsk.",
+					   vi->i_size, ni->allocated_size);
+				goto put_err_out;
+			}
 			/*
 			 * Verify the number of mft records does not exceed
 			 * 2^32 - 1.
-- 
2.55.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [PATCH v3 6/6] ntfs: reject non-resident attributes whose sizes exceed their allocation
  2026-09-30  3:48 [PATCH v3 0/6] ntfs: fix the $MFT bootstrap hang and reads of unmapped runlist ranges Matthias Goergens
                   ` (4 preceding siblings ...)
  2026-09-30  3:48 ` [PATCH v3 5/6] ntfs: fail the mount when $MFT's data size exceeds its allocation Matthias Goergens
@ 2026-09-30  3:48 ` Matthias Goergens
  2026-09-30  9:03 ` [PATCH v3 0/6] ntfs: fix the $MFT bootstrap hang and reads of unmapped runlist ranges Hyunchul Lee
  2026-09-30 10:16 ` Namjae Jeon
  7 siblings, 0 replies; 9+ messages in thread
From: Matthias Goergens @ 2026-09-30  3:48 UTC (permalink / raw)
  To: Namjae Jeon, Hyunchul Lee; +Cc: Baolin Liu, ntfs, linux-kernel

ntfs_read_locked_inode(), ntfs_read_locked_attr_inode() and
ntfs_read_locked_index_inode() take a non-resident attribute's data_size
and initialized_size from disk without checking them against its
allocated_size.  The runlist ends at allocated_size and the read path
maps what lies beyond it as a hole, so a data_size larger than the
allocation makes reads return zeros that are not on disk, with no error.

On a crafted volume with 4 KiB clusters where a 6144000-byte file claims
a data_size and initialized_size 64 KiB beyond its allocation, reading
the file returns 16 clusters of such zeros.  A sparse file behaves the
same.  A compressed file instead fails with -EIO after a burst of "Still
have pages left!" errors from ntfs_read_compressed_block().

Require 0 <= initialized_size <= data_size <= allocated_size, with
allocated_size a multiple of the cluster size, when the inode is read,
as fs/ntfs3 does in mi_enum_attr(), and fail with -EIO otherwise, which
marks the inode bad and the volume as having errors.  initialized_size
above data_size is rejected too because an extending write zeroes the
gap it opens only from initialized_size onwards.  An unaligned
allocated_size ends inside a cluster that the read path treats as past
the end: a crafted 66440-byte file with 4 KiB clusters reads its last
904 bytes as zeros.

Sparse and compressed attributes need no exception.  Every non-resident
attribute has a cluster-aligned allocated_size >= data_size, holes
included, on eight public test volumes (seven written by Windows, with
LZNT1-compressed files, a sparse $UsnJrnl:$J and a OneDrive placeholder
whose unnamed $DATA is one hole, and one by mkntfs), on a volume written
by ntfs-3g and on one written by this driver, and the patched driver
reads every file on them as before.  fs/ntfs3 has enforced the same
checks since v6.6, also without exceptions.  The layout.h comment saying
that data_size can exceed allocated_size for compressed and sparse
attributes is corrected.

This also covers a non-resident attribute list, which
load_attribute_list() reads through ntfs_attr_iget(), and $MFT, whose
inode goes through ntfs_read_locked_inode() after the previous patch's
check.  $MFT's own non-resident attribute list goes through neither:
ntfs_read_inode_mount() loads it with load_attribute_list_mount() and
ntfs_read_locked_inode() skips it for $MFT, so the check is added there
too.  The check was already missing in the classic driver; the Fixes
tag names the commit that brought that code back.

Fixes: 1e9ea7e04472 ("Revert "fs: Remove NTFS classic"")
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
---
 fs/ntfs/inode.c  | 27 +++++++++++++++++++++++++++
 fs/ntfs/layout.h | 13 ++++++++-----
 2 files changed, 35 insertions(+), 5 deletions(-)

diff --git a/fs/ntfs/inode.c b/fs/ntfs/inode.c
index 9c97fc3f9e5ff..7e475d339e920 100644
--- a/fs/ntfs/inode.c
+++ b/fs/ntfs/inode.c
@@ -651,6 +651,25 @@ void ntfs_set_vfs_operations(struct inode *inode, mode_t mode, dev_t dev)
 	}
 }
 
+static bool ntfs_non_resident_sizes_inconsistent(struct inode *vi,
+						 const struct attr_record *a)
+{
+	s64 allocated_size = le64_to_cpu(a->data.non_resident.allocated_size);
+	s64 data_size = le64_to_cpu(a->data.non_resident.data_size);
+	s64 initialized_size = le64_to_cpu(a->data.non_resident.initialized_size);
+
+	if (initialized_size >= 0 && initialized_size <= data_size &&
+	    data_size <= allocated_size &&
+	    !ntfs_bytes_to_cluster_off(NTFS_I(vi)->vol, allocated_size))
+		return false;
+
+	ntfs_error(vi->i_sb,
+		   "Attribute 0x%x of inode 0x%llx is corrupt (initialized size %lld, data size %lld, allocated size %lld).",
+		   le32_to_cpu(a->type), NTFS_I(vi)->mft_no, initialized_size,
+		   data_size, allocated_size);
+	return true;
+}
+
 /*
  * ntfs_read_locked_inode - read an inode from its device
  * @vi:		inode to read
@@ -1184,6 +1203,8 @@ static int ntfs_read_locked_inode(struct inode *vi)
 					"First extent of $DATA attribute has non zero lowest_vcn.");
 				goto unm_err_out;
 			}
+			if (ntfs_non_resident_sizes_inconsistent(vi, a))
+				goto unm_err_out;
 			vi->i_size = ni->data_size = le64_to_cpu(a->data.non_resident.data_size);
 			ni->initialized_size = le64_to_cpu(a->data.non_resident.initialized_size);
 			ni->allocated_size = le64_to_cpu(a->data.non_resident.allocated_size);
@@ -1446,6 +1467,8 @@ static int ntfs_read_locked_attr_inode(struct inode *base_vi, struct inode *vi)
 			ntfs_error(vi->i_sb, "First extent of attribute has non-zero lowest_vcn.");
 			goto unm_err_out;
 		}
+		if (ntfs_non_resident_sizes_inconsistent(vi, a))
+			goto unm_err_out;
 		vi->i_size = ni->data_size = le64_to_cpu(a->data.non_resident.data_size);
 		ni->initialized_size = le64_to_cpu(a->data.non_resident.initialized_size);
 		ni->allocated_size = le64_to_cpu(a->data.non_resident.allocated_size);
@@ -1675,6 +1698,8 @@ static int ntfs_read_locked_index_inode(struct inode *base_vi, struct inode *vi)
 			"First extent of $INDEX_ALLOCATION attribute has non zero lowest_vcn.");
 		goto unm_err_out;
 	}
+	if (ntfs_non_resident_sizes_inconsistent(vi, a))
+		goto unm_err_out;
 	vi->i_size = ni->data_size = le64_to_cpu(a->data.non_resident.data_size);
 	ni->initialized_size = le64_to_cpu(a->data.non_resident.initialized_size);
 	ni->allocated_size = le64_to_cpu(a->data.non_resident.allocated_size);
@@ -2008,6 +2033,8 @@ int ntfs_read_inode_mount(struct inode *vi)
 					"Attribute list has non zero lowest_vcn. $MFT is corrupt. You should run chkdsk.");
 				goto put_err_out;
 			}
+			if (ntfs_non_resident_sizes_inconsistent(vi, a))
+				goto put_err_out;
 
 			rl = ntfs_mapping_pairs_decompress(vol, a, NULL, &new_rl_count);
 			if (IS_ERR(rl)) {
diff --git a/fs/ntfs/layout.h b/fs/ntfs/layout.h
index 8f5792139d719..2de83d4ca40b9 100644
--- a/fs/ntfs/layout.h
+++ b/fs/ntfs/layout.h
@@ -811,14 +811,17 @@ enum {
  *                                  on XP SP2+.
  * @data.non_resident.reserved:     5 bytes for 8-byte alignment.
  * @data.non_resident.allocated_size:
- *                                  Allocated disk space in bytes.
- *                                  For compressed: logical allocated size.
+ *                                  Allocated size in bytes, a multiple of
+ *                                  the cluster size.  For compressed and
+ *                                  sparse attributes holes count as
+ *                                  allocated; the clusters actually in use
+ *                                  are in compressed_size.
  * @data.non_resident.data_size:    Logical attribute value size in bytes.
- *                                  Can be larger than allocated_size if
- *                                  compressed/sparse.
+ *                                  Never larger than allocated_size, also
+ *                                  when compressed/sparse.
  * @data.non_resident.initialized_size:
  *                                  Initialized portion size in bytes.
- *                                  Usually equals data_size.
+ *                                  Usually equals data_size, never larger.
  * @data.non_resident.compressed_size:
  *                                  Compressed on-disk size in bytes.
  *                                  Only present when compressed or sparse.
-- 
2.55.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH v3 0/6] ntfs: fix the $MFT bootstrap hang and reads of unmapped runlist ranges
  2026-09-30  3:48 [PATCH v3 0/6] ntfs: fix the $MFT bootstrap hang and reads of unmapped runlist ranges Matthias Goergens
                   ` (5 preceding siblings ...)
  2026-09-30  3:48 ` [PATCH v3 6/6] ntfs: reject non-resident attributes whose sizes exceed their allocation Matthias Goergens
@ 2026-09-30  9:03 ` Hyunchul Lee
  2026-09-30 10:16 ` Namjae Jeon
  7 siblings, 0 replies; 9+ messages in thread
From: Hyunchul Lee @ 2026-09-30  9:03 UTC (permalink / raw)
  To: Matthias Goergens; +Cc: Namjae Jeon, Baolin Liu, ntfs, linux-kernel


The whole series looks good to me.

Reviewed-by: Hyunchul Lee <hyc.lee@gmail.com>

-- 
Thanks,
Hyunchul

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH v3 0/6] ntfs: fix the $MFT bootstrap hang and reads of unmapped runlist ranges
  2026-09-30  3:48 [PATCH v3 0/6] ntfs: fix the $MFT bootstrap hang and reads of unmapped runlist ranges Matthias Goergens
                   ` (6 preceding siblings ...)
  2026-09-30  9:03 ` [PATCH v3 0/6] ntfs: fix the $MFT bootstrap hang and reads of unmapped runlist ranges Hyunchul Lee
@ 2026-09-30 10:16 ` Namjae Jeon
  7 siblings, 0 replies; 9+ messages in thread
From: Namjae Jeon @ 2026-09-30 10:16 UTC (permalink / raw)
  To: Matthias Goergens; +Cc: Hyunchul Lee, Baolin Liu, ntfs, linux-kernel

On Wed, Sep 30, 2026 at 12:48 PM Matthias Goergens
<matthias.goergens@gmail.com> wrote:
>
> Hi Hyunchul,
>
> This is v3, with your review of v2 and Baolin's applied:
>
> - Patch 1: a failed retry for a vcn at or beyond allocated_size returns
>   the runlist end again, as in ntfs-next; only lookups below it fail
>   with -EIO.  On its own, v2's patch 1 also failed lookups past EOF,
>   for instance when reading a file's last folio with 512-byte clusters
>   (Baolin).  Baolin suggested moving patch 4's allocated_size check
>   before the retry instead, but that skips a retry ntfs-next makes: on
>   a corrupt volume whose allocated_size is cut to where a file's last
>   extent record starts, the rest of the file was then read as zeros,
>   where ntfs-next and v3's patch 1 read the data (with the whole series,
>   patch 6 rejects that file).
> - Patch 4: the if statement you asked me to merge is gone, as the
>   allocated_size check now sits in patch 1, and in
>   ntfs_attr_vcn_to_rl() patch 4 only extends patch 1's -EIO to
>   LCN_ENOENT.  The expansion rollback restores allocated_size under
>   size_lock.
> - Patch 6: the comment is gone, and $MFT's own non-resident attribute
>   list is checked too (Baolin); a volume where that list claims more
>   data than its allocation used to mount and now fails to.
>
> Patches 2, 3 and 5 are unchanged.
>
> The series fixes a hang at mount when $MFT needs its own extent records
> (patch 3, which needs patch 1).  Patches 1 and 2 fix reads and writes
> of a runlist range held in an extent record that cannot be read, which
> returned zeros with no error or silently lost buffered writes, and
> patch 4 does the same for a runlist that ends before allocated_size.
> Patches 5 and 6 check the sizes of $MFT's data and of every non-resident
> attribute against their allocation, as fs/ntfs3 does.
>
> Tested under qemu with KASAN and the hung-task detector on ntfs-next,
> with and without the series.  Nine crafted images that hang or crash
> the mount without it fail to mount with it, and six that read zeros
> that are not on disk return errors instead.  Eighteen images that mount
> without the series, among them eight public test volumes written by
> Windows or mkntfs, read every file the same with it, and six small
> mkntfs images still mount.  v3 gives the same results as v2 on all of
> them.  The images and the scripts that generate them are at
>
>   https://github.com/matthiasgoergens/linux/tree/reproducer/2026-09-26-ntfs-mft-runlist
>
> and, for the ordinary file of patch 4, the unaligned allocated_size of
> patch 6 and three of the eighteen volumes that mount, at
>
>   https://github.com/matthiasgoergens/linux/tree/reproducer/2026-09-26-ntfs-v2-extra
>
> and, for v3's changes, at
>
>   https://github.com/matthiasgoergens/linux/tree/reproducer/2026-09-30-ntfs-v3
>
> v2: https://lore.kernel.org/all/cover.1790417653.git.matthias.goergens@gmail.com/
> v1: https://lore.kernel.org/all/20260922153931.1976405-1-matthias.goergens@gmail.com/
>
> Thanks,
> Matthias
>
> Matthias Goergens (6):
>   ntfs: do not map an unmappable runlist fragment as a hole
>   ntfs: do not turn an unmappable runlist fragment into delalloc on
>     write
>   ntfs: fail the mount when $MFT needs its own extent records
>   ntfs: do not map a vcn as a hole when its runlist lookup failed
>   ntfs: fail the mount when $MFT's data size exceeds its allocation
>   ntfs: reject non-resident attributes whose sizes exceed their
>     allocation
Applied them to #ntfs-next.
Thanks!

^ permalink raw reply	[flat|nested] 9+ messages in thread

end of thread, other threads:[~2026-09-30 10:16 UTC | newest]

Thread overview: 9+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-30  3:48 [PATCH v3 0/6] ntfs: fix the $MFT bootstrap hang and reads of unmapped runlist ranges Matthias Goergens
2026-09-30  3:48 ` [PATCH v3 1/6] ntfs: do not map an unmappable runlist fragment as a hole Matthias Goergens
2026-09-30  3:48 ` [PATCH v3 2/6] ntfs: do not turn an unmappable runlist fragment into delalloc on write Matthias Goergens
2026-09-30  3:48 ` [PATCH v3 3/6] ntfs: fail the mount when $MFT needs its own extent records Matthias Goergens
2026-09-30  3:48 ` [PATCH v3 4/6] ntfs: do not map a vcn as a hole when its runlist lookup failed Matthias Goergens
2026-09-30  3:48 ` [PATCH v3 5/6] ntfs: fail the mount when $MFT's data size exceeds its allocation Matthias Goergens
2026-09-30  3:48 ` [PATCH v3 6/6] ntfs: reject non-resident attributes whose sizes exceed their allocation Matthias Goergens
2026-09-30  9:03 ` [PATCH v3 0/6] ntfs: fix the $MFT bootstrap hang and reads of unmapped runlist ranges Hyunchul Lee
2026-09-30 10:16 ` Namjae Jeon

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®