From: Matthias Goergens <matthias.goergens@gmail.com>
To: Namjae Jeon <linkinjeon@kernel.org>, Hyunchul Lee <hyc.lee@gmail.com>
Cc: ntfs@lists.linux.dev, linux-kernel@vger.kernel.org
Subject: [PATCH v2 3/6] ntfs: fail the mount when $MFT needs its own extent records
Date: Sun, 27 Sep 2026 13:08:22 +0800 [thread overview]
Message-ID: <20260927050831.2739166-3-matthias.goergens@gmail.com> (raw)
In-Reply-To: <cover.1790417653.git.matthias.goergens@gmail.com>
Mounting a crafted image hangs the mount process forever, unkillable at
0% CPU. Only the hung-task detector reports it:
INFO: task mount:74 blocked in I/O wait for more than 30 seconds.
folio_wait_bit_common <- waits forever
filemap_read_folio
map_mft_record_folio
map_mft_record
ntfs_map_runlist_nolock
ntfs_attr_vcn_to_rl
__ntfs_read_iomap_begin
iomap_read_folio
ntfs_read_folio <- already holds that folio's lock
ntfs_read_inode_mount
ntfs_fill_super
ntfs_read_inode_mount() assembles $MFT's runlist one $DATA extent at a
time, hoping, as its comment says, that it never needs a part of $MFT it
has not decoded yet. If $MFT's attribute list puts one of its
attributes or later $DATA extents in an extent record outside the
runlist decoded so far, reading that record re-enters
ntfs_map_runlist_nolock() for $MFT. Depending on the layout, that waits
on the folio lock it already holds, as above, blocks on $MFT's runlist
lock, or dereferences NULL in map_extent_mft_record().
Mark the volume for the whole $DATA enumeration and have
ntfs_map_runlist_nolock() refuse $MFT with -EIO while the mark is set.
The enumeration decodes each extent itself with
ntfs_mapping_pairs_decompress(), so it does not need that path.
A validly placed record can trip the check too. The driver puts an
extent record for $MFT's own mapping pairs before the first vcn it
describes, but with clusters smaller than a page, the folio holding it
can still run past the decoded runlist. An unpatched kernel hangs on
that layout as well; with the check the mount fails. The check then
refuses only the folio's tail, so this patch needs "ntfs: do not map an
unmappable runlist fragment as a hole": without it the tail is
zero-filled and the mount serves an all-zero mft record. The mount now
fails instead of hanging:
ntfs: (device vda): ntfs_map_runlist_nolock(): $MFT needs its own
extent records to describe itself; cannot mount.
Tested under qemu with KASAN, PROVE_LOCKING and the hung-task detector
on ntfs-next, with and without the whole series: six crafted images
that hang or crash an unpatched kernel fail to mount with the series,
and nine images that mount without it still mount, three of them with
the same file listing and contents, among them a volume written by
ntfs-3g whose $MFT has its $DATA in three extent records.
Fixes: b041ca562526 ("ntfs: update iomap and address space operations")
Suggested-by: Hyunchul Lee <hyc.lee@gmail.com>
Cc: stable@vger.kernel.org
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
---
fs/ntfs/attrib.c | 11 +++++++++++
fs/ntfs/inode.c | 9 +++++++++
fs/ntfs/volume.h | 3 +++
3 files changed, 23 insertions(+)
diff --git a/fs/ntfs/attrib.c b/fs/ntfs/attrib.c
index a337a3429b401..eab4d8d32132f 100644
--- a/fs/ntfs/attrib.c
+++ b/fs/ntfs/attrib.c
@@ -103,6 +103,17 @@ int ntfs_map_runlist_nolock(struct ntfs_inode *ni, s64 vcn, struct ntfs_attr_sea
base_ni = ni;
else
base_ni = ni->ext.base_ntfs_ino;
+ /*
+ * ntfs_read_inode_mount() builds $MFT's runlist itself, so nothing
+ * should reach here for $MFT. A crafted image can: the read that
+ * gets here already holds the $MFT folio lock it would wait on.
+ */
+ if (unlikely(NVolMftBootstrap(ni->vol) &&
+ base_ni == NTFS_I(ni->vol->mft_ino))) {
+ ntfs_error(ni->vol->sb,
+ "$MFT needs its own extent records to describe itself; cannot mount.");
+ return -EIO;
+ }
if (!ctx) {
ctx_is_temporary = ctx_needs_reset = true;
m = map_mft_record(base_ni);
diff --git a/fs/ntfs/inode.c b/fs/ntfs/inode.c
index 9583b2c6c7a26..a61f1519549cb 100644
--- a/fs/ntfs/inode.c
+++ b/fs/ntfs/inode.c
@@ -2080,6 +2080,11 @@ int ntfs_read_inode_mount(struct inode *vi)
/* Now load all attribute extents. */
a = NULL;
next_vcn = last_vcn = highest_vcn = 0;
+ /*
+ * Reading one of $MFT's own extent records in this loop can re-enter
+ * ntfs_map_runlist_nolock() for $MFT; see the check there.
+ */
+ NVolSetMftBootstrap(vol);
while (!(err = ntfs_attr_lookup(AT_DATA, NULL, 0, 0, next_vcn, NULL, 0,
ctx))) {
struct runlist_element *nrl;
@@ -2162,6 +2167,7 @@ int ntfs_read_inode_mount(struct inode *vi)
err = ntfs_read_locked_inode(vi);
if (err) {
ntfs_error(sb, "ntfs_read_inode() of $MFT failed.\n");
+ NVolClearMftBootstrap(vol);
ntfs_attr_put_search_ctx(ctx);
/* Revert to the safe super operations. */
kfree(m);
@@ -2195,6 +2201,7 @@ int ntfs_read_inode_mount(struct inode *vi)
goto put_err_out;
}
}
+ NVolClearMftBootstrap(vol);
if (err != -ENOENT) {
ntfs_error(sb, "Failed to lookup $MFT/$DATA attribute extent. Run chkdsk.\n");
goto put_err_out;
@@ -2229,6 +2236,8 @@ int ntfs_read_inode_mount(struct inode *vi)
put_err_out:
ntfs_attr_put_search_ctx(ctx);
err_out:
+ /* Also reached from inside the $DATA loop. */
+ NVolClearMftBootstrap(vol);
ntfs_error(sb, "Failed. Marking inode as bad.");
kfree(m);
return -1;
diff --git a/fs/ntfs/volume.h b/fs/ntfs/volume.h
index fdb57279de84c..0473b602084c5 100644
--- a/fs/ntfs/volume.h
+++ b/fs/ntfs/volume.h
@@ -194,6 +194,7 @@ struct ntfs_volume {
* NV_Discard Issue discard/TRIM commands for freed clusters.
* NV_DisableSparse Disable creation of sparse regions.
* NV_NativeSymlinkRel Translate absolute Windows reparse targets (native_symlink=rel).
+ * NV_MftBootstrap Mount is still assembling $MFT's own runlist.
*/
enum {
NV_Errors,
@@ -214,6 +215,7 @@ enum {
NV_DisableSparse,
NV_NativeSymlinkRel,
NV_SymlinkNative,
+ NV_MftBootstrap,
};
/*
@@ -253,6 +255,7 @@ DEFINE_NVOL_BIT_OPS(Discard)
DEFINE_NVOL_BIT_OPS(DisableSparse)
DEFINE_NVOL_BIT_OPS(NativeSymlinkRel)
DEFINE_NVOL_BIT_OPS(SymlinkNative)
+DEFINE_NVOL_BIT_OPS(MftBootstrap)
static inline void ntfs_inc_free_clusters(struct ntfs_volume *vol, s64 nr)
{
--
2.55.0
next prev parent reply other threads:[~2026-09-27 5:08 UTC|newest]
Thread overview: 14+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-22 15:39 [PATCH] " Matthias Goergens
2026-09-23 1:16 ` Hyunchul Lee
2026-09-27 5:08 ` [PATCH v2 0/6] ntfs: fix the $MFT bootstrap hang and reads of unmapped runlist ranges Matthias Goergens
2026-09-27 5:08 ` [PATCH v2 1/6] ntfs: do not map an unmappable runlist fragment as a hole Matthias Goergens
2026-09-27 23:03 ` liubaolin
2026-09-27 5:08 ` [PATCH v2 2/6] ntfs: do not turn an unmappable runlist fragment into delalloc on write Matthias Goergens
2026-09-28 0:21 ` Hyunchul Lee
2026-09-28 1:46 ` Hyunchul Lee
2026-09-27 5:08 ` Matthias Goergens [this message]
2026-09-27 5:08 ` [PATCH v2 4/6] ntfs: do not map a vcn as a hole when its runlist lookup failed Matthias Goergens
2026-09-28 1:09 ` Hyunchul Lee
2026-09-27 5:08 ` [PATCH v2 5/6] ntfs: fail the mount when $MFT's data size exceeds its allocation Matthias Goergens
2026-09-27 5:08 ` [PATCH v2 6/6] ntfs: reject non-resident attributes whose sizes exceed their allocation Matthias Goergens
2026-09-28 1:42 ` Hyunchul Lee
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260927050831.2739166-3-matthias.goergens@gmail.com \
--to=matthias.goergens@gmail.com \
--cc=hyc.lee@gmail.com \
--cc=linkinjeon@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=ntfs@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®