mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v3 0/4] ntfs: fix spurious EIO and ENOSPC on nearly full volumes
@ 2026-10-04  9:56 Matthias Goergens
  2026-10-04  9:56 ` [PATCH v3 1/4] ntfs: set the attribute list size before reserving space for it Matthias Goergens
                   ` (4 more replies)
  0 siblings, 5 replies; 6+ messages in thread
From: Matthias Goergens @ 2026-10-04  9:56 UTC (permalink / raw)
  To: Namjae Jeon, Hyunchul Lee; +Cc: Baolin Liu, ntfs, linux-kernel

These fix three errors returned while free space is left, and a
contiguous cluster allocation that comes back in two runs.  Patch 1
fixes file creation failing with EIO on a nearly full volume; it now
goes on until the volume is full and then fails with ENOSPC.  Patch 2
fixes the cluster allocator missing free clusters when its start hint
lies at or past the end of the volume, which empties the MFT zone, and
patch 3 removes the fallocate() check that then fails every later call.
Together the two bugs make fallocate() refuse space that is free until
the next mount.  Patch 4 keeps a contiguous request to a single run:
fallocate() zeroes only the first run, so the clusters of a second one
kept the data of a deleted file.  Reproducers, run with KASAN and
lockdep:
https://github.com/matthiasgoergens/linux/tree/reproducer/2026-10-04-ntfs-nearfull-v3

Changes in v3: Namjae Jeon asked whether a contiguous request can still
get a second run once the search has left the hint, for instance when
the run reaches the end of a bitmap page and the next page is skipped as
full.  It can, on that path and on two more: where the search wraps to
pass 2, and where it shrinks the MFT zone after trying every zone.  On
ntfs-next each left 16 of 20 fallocated clusters reading the data of a
deleted file.  New patch 4 does what he suggested: before setting a
cluster's bit, it checks that the cluster follows the run, and returns
the run otherwise.  Patches 1 to 3 are unchanged apart from Hyunchul
Lee's Reviewed-by.  Rebased on ntfs-next.

v2: https://lore.kernel.org/all/cover.1790853718.git.matthias.goergens@gmail.com/
v1: https://lore.kernel.org/all/cover.1790503811.git.matthias.goergens@gmail.com/

Matthias Goergens (4):
  ntfs: set the attribute list size before reserving space for it
  ntfs: restart the zone search when the allocation hint fails
  ntfs: do not refuse fallocate() when the MFT zone is empty
  ntfs: do not give a contiguous cluster allocation a second run

 fs/ntfs/attrlist.c |  5 +++--
 fs/ntfs/file.c     |  3 ---
 fs/ntfs/lcnalloc.c | 35 ++++++++++++++++++++++++++++++++---
 3 files changed, 35 insertions(+), 8 deletions(-)


base-commit: b2a3b6b6c7b09040e3f44ade6e20a7689e4acd28
-- 
2.56.0


^ permalink raw reply	[flat|nested] 6+ messages in thread

* [PATCH v3 1/4] ntfs: set the attribute list size before reserving space for it
  2026-10-04  9:56 [PATCH v3 0/4] ntfs: fix spurious EIO and ENOSPC on nearly full volumes Matthias Goergens
@ 2026-10-04  9:56 ` Matthias Goergens
  2026-10-04  9:56 ` [PATCH v3 2/4] ntfs: restart the zone search when the allocation hint fails Matthias Goergens
                   ` (3 subsequent siblings)
  4 siblings, 0 replies; 6+ messages in thread
From: Matthias Goergens @ 2026-10-04  9:56 UTC (permalink / raw)
  To: Namjae Jeon, Hyunchul Lee; +Cc: Baolin Liu, ntfs, linux-kernel

ntfs_attrlist_update_locked() resizes the attribute list and then, for
$MFT, tries to reserve the maximum list size.  When that allocation
fails with -ENOSPC it falls back to ntfs_attrlist_repack(), which reads
the list through ntfs_inode_attr_pread().  The read stops at i_size,
but i_size is only updated after the reserve, so when the list has just
grown the repack gets a short read and returns -EIO.  Only -ENOSPC from
the reserve is ignored, so the -EIO fails the $MFT extension and with
it the file creation that needed a new mft record.

On a 32 MiB volume with 512-byte clusters whose $MFT has a non-resident
attribute list, 591 of 1500 creations of empty files succeed and the
rest fail with EIO while 742 KiB is still free:

  ntfs: (device vda): ntfs_attrlist_update_locked(): Failed to reserve attribute list space
  ntfs: Failed add attr entry to attrlist
  ntfs: MP update failed
  ntfs: (device vda): ntfs_mft_record_alloc(): Failed to extend mft data allocation.

Update i_size as soon as the list has been resized.  File creation then
goes on until the volume is full and fails with ENOSPC.

Fixes: b1d732e62a5b ("ntfs: repack $MFT/$ATTRIBUTE LIST")
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
Reviewed-by: Baolin Liu <liubaolin@kylinos.cn>
Reviewed-by: Hyunchul Lee <hyc.lee@gmail.com>
---
The volume is frag.img from
https://github.com/matthiasgoergens/linux/tree/reproducer/2026-09-26-ntfs-mft-runlist
(mft-bootstrap/frag.img.xz); creating empty files on it until one fails
is enough.  Tested with KASAN and lockdep: after the patch the files
written before the volume filled read back intact after a remount, and
ntfsfix and ntfsresize --info find nothing wrong with the image.
---
 fs/ntfs/attrlist.c | 5 +++--
 1 file changed, 3 insertions(+), 2 deletions(-)

diff --git a/fs/ntfs/attrlist.c b/fs/ntfs/attrlist.c
index 3660e7fd24b1..6edd5d2d9bf2 100644
--- a/fs/ntfs/attrlist.c
+++ b/fs/ntfs/attrlist.c
@@ -256,6 +256,9 @@ int ntfs_attrlist_update_locked(struct ntfs_inode *base_ni,
 		return err;
 	}
 
+	/* ntfs_attrlist_repack() below reads the list up to i_size. */
+	i_size_write(attr_vi, base_ni->attr_list_size);
+
 	/*
 	 * Reserve the maximum legal list size while the MFT metadata area is
 	 * still easy to allocate contiguously. This prevents a later list entry
@@ -283,8 +286,6 @@ int ntfs_attrlist_update_locked(struct ntfs_inode *base_ni,
 		}
 	}
 
-	i_size_write(attr_vi, base_ni->attr_list_size);
-
 	if (NInoNonResident(attr_ni) && !NInoAttrListNonResident(base_ni))
 		NInoSetAttrListNonResident(base_ni);
 
-- 
2.56.0


^ permalink raw reply	[flat|nested] 6+ messages in thread

* [PATCH v3 2/4] ntfs: restart the zone search when the allocation hint fails
  2026-10-04  9:56 [PATCH v3 0/4] ntfs: fix spurious EIO and ENOSPC on nearly full volumes Matthias Goergens
  2026-10-04  9:56 ` [PATCH v3 1/4] ntfs: set the attribute list size before reserving space for it Matthias Goergens
@ 2026-10-04  9:56 ` Matthias Goergens
  2026-10-04  9:56 ` [PATCH v3 3/4] ntfs: do not refuse fallocate() when the MFT zone is empty Matthias Goergens
                   ` (2 subsequent siblings)
  4 siblings, 0 replies; 6+ messages in thread
From: Matthias Goergens @ 2026-10-04  9:56 UTC (permalink / raw)
  To: Namjae Jeon, Hyunchul Lee; +Cc: Baolin Liu, ntfs, linux-kernel

When ntfs_cluster_alloc() is given a start_lcn, it first tries the
clusters from there on.  If that does not satisfy the request, it moves
on to the zone's current position, but keeps the pass, the zone_end and
the has_guess state it had.  That loses free clusters in two ways.  If
the search had already moved on to pass 2, whose range ends at the hint,
it scans only from the zone position to the hint and misses every free
cluster between the start of the zone and the zone position.  If no
cluster had been tested yet, has_guess is still set, so the cluster at
the zone position is tried as if it had been the hint, and when that
cluster is in use the rest of the bitmap buffer is skipped.

Both happen when the hint lies at or past the end of the volume, which
ntfs_attr_map_cluster() produces when it extrapolates from the last
allocated run across a hole.  The allocator then finds nothing, shrinks
the MFT zone to nothing trying to satisfy the request, and fails with
-ENOSPC.

To reproduce on a 128 MiB volume with 4 KiB clusters, extend two files
in turn by one cluster at a time, alternating between fallocate(),
write() past EOF and truncate() up, so that the runs of each file are
separated by holes.  When the volume is full, truncate both files back
to 1 MiB and start again.  During the second round a one-cluster
fallocate() fails with ENOSPC while 40 MiB is free.

Start the zone over as if no hint had been given.  A contiguous request
can get there with the run from the hint already allocated, when that
run reached the end of a bitmap buffer or of the zone.  Return that run
instead, as when a cluster in use ends it; going on at the zone position
gives the request a second run that its caller does not expect.  That
could happen before this patch too, whenever the cluster at the zone
position was free.  ntfs_attr_map_cluster() reports and zeroes only the
first run, so after a fallocate() of 20 clusters over a hole, with the
hint 4 clusters before the end of the volume, 16 of them read back the
data of a deleted file.

Fixes: 11ccc9107dc4 ("ntfs: update runlist handling and cluster allocator")
Suggested-by: Baolin Liu <liubaolin@kylinos.cn>
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
Reviewed-by: Hyunchul Lee <hyc.lee@gmail.com>
---
v2: for a contiguous request, return the run from the hint instead of
going on at the zone position (Baolin Liu).  This also fixes the stale
data after fallocate() described in the last paragraph.

v3: unchanged.  The other ways to a second run that Namjae Jeon asked
about, after the search has left the hint, are fixed in patch 4.
---
 fs/ntfs/lcnalloc.c | 25 ++++++++++++++++++++++---
 1 file changed, 22 insertions(+), 3 deletions(-)

diff --git a/fs/ntfs/lcnalloc.c b/fs/ntfs/lcnalloc.c
index 0d6cd08ee2e7..c58e689fb585 100644
--- a/fs/ntfs/lcnalloc.c
+++ b/fs/ntfs/lcnalloc.c
@@ -506,13 +506,32 @@ struct runlist_element *ntfs_cluster_alloc(struct ntfs_volume *vol, const s64 st
 		}
 
 		if (!used_zone_pos) {
+			/*
+			 * The run at @start_lcn reached the end of the buffer
+			 * or the zone.  A contiguous request gets that run
+			 * alone, not a second one from the zone position.
+			 */
+			if (is_contig && rlpos)
+				goto out;
+			/*
+			 * Leaving @start_lcn for the zone position starts the
+			 * zone over as if no hint had been given, even if the
+			 * search had already reached pass 2, whose range ends
+			 * at @start_lcn.
+			 */
 			used_zone_pos = 1;
-			if (search_zone == 1)
+			has_guess = 0;
+			pass = 1;
+			if (search_zone == 1) {
 				zone_start = vol->mft_zone_pos;
-			else if (search_zone == 2)
+				zone_end = vol->mft_zone_end;
+			} else if (search_zone == 2) {
 				zone_start = vol->data1_zone_pos;
-			else
+				zone_end = vol->nr_clusters;
+			} else {
 				zone_start = vol->data2_zone_pos;
+				zone_end = vol->mft_zone_start;
+			}
 
 			if (!zone_start || zone_start == vol->mft_zone_start ||
 			    zone_start == vol->mft_zone_end)
-- 
2.56.0


^ permalink raw reply	[flat|nested] 6+ messages in thread

* [PATCH v3 3/4] ntfs: do not refuse fallocate() when the MFT zone is empty
  2026-10-04  9:56 [PATCH v3 0/4] ntfs: fix spurious EIO and ENOSPC on nearly full volumes Matthias Goergens
  2026-10-04  9:56 ` [PATCH v3 1/4] ntfs: set the attribute list size before reserving space for it Matthias Goergens
  2026-10-04  9:56 ` [PATCH v3 2/4] ntfs: restart the zone search when the allocation hint fails Matthias Goergens
@ 2026-10-04  9:56 ` Matthias Goergens
  2026-10-04  9:56 ` [PATCH v3 4/4] ntfs: do not give a contiguous cluster allocation a second run Matthias Goergens
  2026-10-05  2:26 ` [PATCH v3 0/4] ntfs: fix spurious EIO and ENOSPC on nearly full volumes Namjae Jeon
  4 siblings, 0 replies; 6+ messages in thread
From: Matthias Goergens @ 2026-10-04  9:56 UTC (permalink / raw)
  To: Namjae Jeon, Hyunchul Lee; +Cc: Baolin Liu, ntfs, linux-kernel

ntfs_fallocate() fails with -ENOSPC whenever the MFT zone has shrunk to
nothing.  The cluster allocator gives the MFT zone up when it cannot
find free clusters elsewhere, and nothing restores it before the next
mount, so after one such failure every fallocate() on the volume fails
with ENOSPC, punching holes included, however much space is freed.  The
allocator failure fixed by the previous patch was enough to get there:
afterwards a one-cluster fallocate() failed with ENOSPC while 123 MiB of
the 128 MiB volume were free.

write() and truncate() do not look at the MFT zone, and
ntfs_allocate_range() already fails with -ENOSPC when too few clusters
are free.  Remove the check.

Fixes: 9c87959601e8 ("ntfs: update file operations")
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
Reviewed-by: Baolin Liu <liubaolin@kylinos.cn>
Reviewed-by: Hyunchul Lee <hyc.lee@gmail.com>
---
 fs/ntfs/file.c | 3 ---
 1 file changed, 3 deletions(-)

diff --git a/fs/ntfs/file.c b/fs/ntfs/file.c
index 3ec82715a588..762045d26296 100644
--- a/fs/ntfs/file.c
+++ b/fs/ntfs/file.c
@@ -1163,9 +1163,6 @@ static long ntfs_fallocate(struct file *file, int mode, loff_t offset, loff_t le
 	if (!NVolFreeClusterKnown(vol))
 		wait_event(vol->free_waitq, NVolFreeClusterKnown(vol));
 
-	if ((ni->vol->mft_zone_end - ni->vol->mft_zone_start) == 0)
-		return -ENOSPC;
-
 	if (NInoNonResident(ni) && !NInoFullyMapped(ni)) {
 		down_write(&ni->runlist.lock);
 		err = ntfs_attr_map_whole_runlist(ni);
-- 
2.56.0


^ permalink raw reply	[flat|nested] 6+ messages in thread

* [PATCH v3 4/4] ntfs: do not give a contiguous cluster allocation a second run
  2026-10-04  9:56 [PATCH v3 0/4] ntfs: fix spurious EIO and ENOSPC on nearly full volumes Matthias Goergens
                   ` (2 preceding siblings ...)
  2026-10-04  9:56 ` [PATCH v3 3/4] ntfs: do not refuse fallocate() when the MFT zone is empty Matthias Goergens
@ 2026-10-04  9:56 ` Matthias Goergens
  2026-10-05  2:26 ` [PATCH v3 0/4] ntfs: fix spurious EIO and ENOSPC on nearly full volumes Namjae Jeon
  4 siblings, 0 replies; 6+ messages in thread
From: Matthias Goergens @ 2026-10-04  9:56 UTC (permalink / raw)
  To: Namjae Jeon, Hyunchul Lee; +Cc: Baolin Liu, ntfs, linux-kernel

ntfs_attr_map_cluster() asks ntfs_cluster_alloc() for contiguous
clusters and merges the whole result into the runlist, but reports only
its first run, so its callers zero only that run.  Within a bitmap page
the allocator keeps such a request to one run, because a cluster in use
ends it.  The zone search can also move on without testing the clusters
after the run, though: it skips a bitmap page that has no free clusters,
it goes from the end of the zone to the start of pass 2, and once every
zone has been searched it shrinks the MFT zone by half and goes on at
its new end.  After each of these jumps the first cluster tested is
still taken as the next one of the run, and if it is free it starts a
second run elsewhere on the volume.

fallocate() over a hole then leaves the clusters of that second run
holding whatever they held before: on a nearly full volume, 16 of 20
fallocated clusters read back a deleted file's data, in each of the
three cases.

Delayed allocation gets a second run without punching any holes: filling
the same volume with write() reaches the last case when the MFT zone is
shrunk.  The allocator then takes only the first run off the
dirty-cluster count, and after the fill statfs() reported no free space
while 939 clusters were still free, until the next mount.

Before setting the bit of a cluster for a contiguous request, check that
the cluster comes right after the run, and return the run as it is if
not; no bit has to be cleared.  ntfs_attr_map_cluster() already handles
a short run, which a cluster in use produces in the same way: its
callers ask again for the rest.

Fixes: 11ccc9107dc4 ("ntfs: update runlist handling and cluster allocator")
Suggested-by: Namjae Jeon <linkinjeon@kernel.org>
Link: https://lore.kernel.org/all/CAKYAXd-agPFvDj2-HvjP3SLtG7AxkqUeNPqYyrLt-BN7fT8w6g@mail.gmail.com/
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
---
New in v3, after Namjae Jeon's question on v2 2/3.  Patch 2's early
return for the run from the hint stays: without it the restarted search
would read the bitmap until it found a free cluster just to return the
run, and if there was none it would fail the whole request with ENOSPC.

The other contiguous callers are the attribute list allocations, which
free a result that is not one run of the full length, and
ntfs_write_cb() in compress.c, which writes the whole compression block
from the first cluster without looking at the length.  That was already
wrong for a short run before this patch and is not changed here.

contig-paths.c and its guest init script init-paths in
https://github.com/matthiasgoergens/linux/tree/reproducer/2026-10-04-ntfs-nearfull-v3
set up each case; the README there has the results.
---
 fs/ntfs/lcnalloc.c | 10 ++++++++++
 1 file changed, 10 insertions(+)

diff --git a/fs/ntfs/lcnalloc.c b/fs/ntfs/lcnalloc.c
index c58e689fb585..3ebcac4c2864 100644
--- a/fs/ntfs/lcnalloc.c
+++ b/fs/ntfs/lcnalloc.c
@@ -375,6 +375,16 @@ struct runlist_element *ntfs_cluster_alloc(struct ntfs_volume *vol, const s64 st
 				has_guess = 1;
 				continue;
 			}
+			/*
+			 * A contiguous request gets a single run.  The scan can
+			 * move on without testing the clusters after the run
+			 * (past a bitmap page that is full, to pass 2 or to
+			 * another zone), so return the run when this cluster
+			 * does not follow it.
+			 */
+			if (is_contig && rlpos &&
+			    lcn + bmp_pos != prev_lcn + prev_run_len)
+				goto out;
 			/*
 			 * Allocate more memory if needed, including space for
 			 * the terminator element.
-- 
2.56.0


^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: [PATCH v3 0/4] ntfs: fix spurious EIO and ENOSPC on nearly full volumes
  2026-10-04  9:56 [PATCH v3 0/4] ntfs: fix spurious EIO and ENOSPC on nearly full volumes Matthias Goergens
                   ` (3 preceding siblings ...)
  2026-10-04  9:56 ` [PATCH v3 4/4] ntfs: do not give a contiguous cluster allocation a second run Matthias Goergens
@ 2026-10-05  2:26 ` Namjae Jeon
  4 siblings, 0 replies; 6+ messages in thread
From: Namjae Jeon @ 2026-10-05  2:26 UTC (permalink / raw)
  To: Matthias Goergens; +Cc: Hyunchul Lee, Baolin Liu, ntfs, linux-kernel

On Sun, Oct 4, 2026 at 6:56 PM Matthias Goergens
<matthias.goergens@gmail.com> wrote:
>
> These fix three errors returned while free space is left, and a
> contiguous cluster allocation that comes back in two runs.  Patch 1
> fixes file creation failing with EIO on a nearly full volume; it now
> goes on until the volume is full and then fails with ENOSPC.  Patch 2
> fixes the cluster allocator missing free clusters when its start hint
> lies at or past the end of the volume, which empties the MFT zone, and
> patch 3 removes the fallocate() check that then fails every later call.
> Together the two bugs make fallocate() refuse space that is free until
> the next mount.  Patch 4 keeps a contiguous request to a single run:
> fallocate() zeroes only the first run, so the clusters of a second one
> kept the data of a deleted file.  Reproducers, run with KASAN and
> lockdep:
> https://github.com/matthiasgoergens/linux/tree/reproducer/2026-10-04-ntfs-nearfull-v3
>
> Changes in v3: Namjae Jeon asked whether a contiguous request can still
> get a second run once the search has left the hint, for instance when
> the run reaches the end of a bitmap page and the next page is skipped as
> full.  It can, on that path and on two more: where the search wraps to
> pass 2, and where it shrinks the MFT zone after trying every zone.  On
> ntfs-next each left 16 of 20 fallocated clusters reading the data of a
> deleted file.  New patch 4 does what he suggested: before setting a
> cluster's bit, it checks that the cluster follows the run, and returns
> the run otherwise.  Patches 1 to 3 are unchanged apart from Hyunchul
> Lee's Reviewed-by.  Rebased on ntfs-next.
>
> v2: https://lore.kernel.org/all/cover.1790853718.git.matthias.goergens@gmail.com/
> v1: https://lore.kernel.org/all/cover.1790503811.git.matthias.goergens@gmail.com/
>
> Matthias Goergens (4):
>   ntfs: set the attribute list size before reserving space for it
>   ntfs: restart the zone search when the allocation hint fails
>   ntfs: do not refuse fallocate() when the MFT zone is empty
>   ntfs: do not give a contiguous cluster allocation a second run
Applied them to #ntfs-next.
Thanks!

^ permalink raw reply	[flat|nested] 6+ messages in thread

end of thread, other threads:[~2026-10-05  2:26 UTC | newest]

Thread overview: 6+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-04  9:56 [PATCH v3 0/4] ntfs: fix spurious EIO and ENOSPC on nearly full volumes Matthias Goergens
2026-10-04  9:56 ` [PATCH v3 1/4] ntfs: set the attribute list size before reserving space for it Matthias Goergens
2026-10-04  9:56 ` [PATCH v3 2/4] ntfs: restart the zone search when the allocation hint fails Matthias Goergens
2026-10-04  9:56 ` [PATCH v3 3/4] ntfs: do not refuse fallocate() when the MFT zone is empty Matthias Goergens
2026-10-04  9:56 ` [PATCH v3 4/4] ntfs: do not give a contiguous cluster allocation a second run Matthias Goergens
2026-10-05  2:26 ` [PATCH v3 0/4] ntfs: fix spurious EIO and ENOSPC on nearly full volumes Namjae Jeon

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®