mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH 0/3] ntfs: fix spurious EIO and ENOSPC on nearly full volumes
@ 2026-09-27 10:57 Matthias Goergens
  2026-09-27 10:57 ` [PATCH 1/3] ntfs: set the attribute list size before reserving space for it Matthias Goergens
                   ` (2 more replies)
  0 siblings, 3 replies; 7+ messages in thread
From: Matthias Goergens @ 2026-09-27 10:57 UTC (permalink / raw)
  To: Namjae Jeon, Hyunchul Lee; +Cc: ntfs, linux-kernel

These fix three errors returned while free space is left.  Patch 1
fixes file creation failing with EIO on a nearly full volume; it now
goes on until the volume is full and then fails with ENOSPC.  Patch 2
fixes the cluster allocator missing free clusters when its start hint
lies at or past the end of the volume, which empties the MFT zone, and
patch 3 removes the fallocate() check that then fails every later call.
Together the two bugs make fallocate() refuse space that is free until
the next mount.  Reproducers, run with KASAN and lockdep:
https://github.com/matthiasgoergens/linux/tree/reproducer/2026-09-27-ntfs-nearfull

Matthias Goergens (3):
  ntfs: set the attribute list size before reserving space for it
  ntfs: restart the zone search when the allocation hint fails
  ntfs: do not refuse fallocate() when the MFT zone is empty

 fs/ntfs/attrlist.c |  5 +++--
 fs/ntfs/file.c     |  3 ---
 fs/ntfs/lcnalloc.c | 18 +++++++++++++++---
 3 files changed, 18 insertions(+), 8 deletions(-)

base-commit: 259abb551e2944998cad4214c201954ab1ac5c8d

-- 
2.55.0


^ permalink raw reply	[flat|nested] 7+ messages in thread

* [PATCH 1/3] ntfs: set the attribute list size before reserving space for it
  2026-09-27 10:57 [PATCH 0/3] ntfs: fix spurious EIO and ENOSPC on nearly full volumes Matthias Goergens
@ 2026-09-27 10:57 ` Matthias Goergens
  2026-09-27 18:11   ` liubaolin
  2026-09-27 10:57 ` [PATCH 2/3] ntfs: restart the zone search when the allocation hint fails Matthias Goergens
  2026-09-27 10:57 ` [PATCH 3/3] ntfs: do not refuse fallocate() when the MFT zone is empty Matthias Goergens
  2 siblings, 1 reply; 7+ messages in thread
From: Matthias Goergens @ 2026-09-27 10:57 UTC (permalink / raw)
  To: Namjae Jeon, Hyunchul Lee; +Cc: ntfs, linux-kernel

ntfs_attrlist_update_locked() resizes the attribute list and then, for
$MFT, tries to reserve the maximum list size.  When that allocation
fails with -ENOSPC it falls back to ntfs_attrlist_repack(), which reads
the list through ntfs_inode_attr_pread().  The read stops at i_size,
but i_size is only updated after the reserve, so when the list has just
grown the repack gets a short read and returns -EIO.  Only -ENOSPC from
the reserve is ignored, so the -EIO fails the $MFT extension and with
it the file creation that needed a new mft record.

On a 32 MiB volume with 512-byte clusters whose $MFT has a non-resident
attribute list, 591 of 1500 creations of empty files succeed and the
rest fail with EIO while 742 KiB is still free:

  ntfs: (device vda): ntfs_attrlist_update_locked(): Failed to reserve attribute list space
  ntfs: Failed add attr entry to attrlist
  ntfs: MP update failed
  ntfs: (device vda): ntfs_mft_record_alloc(): Failed to extend mft data allocation.

Update i_size as soon as the list has been resized.  File creation then
goes on until the volume is full and fails with ENOSPC.

Fixes: b1d732e62a5b ("ntfs: repack $MFT/$ATTRIBUTE LIST")
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
---
The volume is frag.img from
https://github.com/matthiasgoergens/linux/tree/reproducer/2026-09-26-ntfs-mft-runlist
(mft-bootstrap/frag.img.xz); creating empty files on it until one fails
is enough.  Tested with KASAN and lockdep: after the patch the files
written before the volume filled read back intact after a remount, and
ntfsfix and ntfsresize --info find nothing wrong with the image.
---
 fs/ntfs/attrlist.c | 5 +++--
 1 file changed, 3 insertions(+), 2 deletions(-)

diff --git a/fs/ntfs/attrlist.c b/fs/ntfs/attrlist.c
index 1bbd2bc62c582..62dfb96e31a3c 100644
--- a/fs/ntfs/attrlist.c
+++ b/fs/ntfs/attrlist.c
@@ -252,6 +252,9 @@ int ntfs_attrlist_update_locked(struct ntfs_inode *base_ni,
 		return err;
 	}
 
+	/* ntfs_attrlist_repack() below reads the list up to i_size. */
+	i_size_write(attr_vi, base_ni->attr_list_size);
+
 	/*
 	 * Reserve the maximum legal list size while the MFT metadata area is
 	 * still easy to allocate contiguously. This prevents a later list entry
@@ -279,8 +282,6 @@ int ntfs_attrlist_update_locked(struct ntfs_inode *base_ni,
 		}
 	}
 
-	i_size_write(attr_vi, base_ni->attr_list_size);
-
 	if (NInoNonResident(attr_ni) && !NInoAttrListNonResident(base_ni))
 		NInoSetAttrListNonResident(base_ni);
 
-- 
2.55.0


^ permalink raw reply	[flat|nested] 7+ messages in thread

* [PATCH 2/3] ntfs: restart the zone search when the allocation hint fails
  2026-09-27 10:57 [PATCH 0/3] ntfs: fix spurious EIO and ENOSPC on nearly full volumes Matthias Goergens
  2026-09-27 10:57 ` [PATCH 1/3] ntfs: set the attribute list size before reserving space for it Matthias Goergens
@ 2026-09-27 10:57 ` Matthias Goergens
  2026-09-27 18:06   ` liubaolin
  2026-09-27 10:57 ` [PATCH 3/3] ntfs: do not refuse fallocate() when the MFT zone is empty Matthias Goergens
  2 siblings, 1 reply; 7+ messages in thread
From: Matthias Goergens @ 2026-09-27 10:57 UTC (permalink / raw)
  To: Namjae Jeon, Hyunchul Lee; +Cc: ntfs, linux-kernel

When ntfs_cluster_alloc() is given a start_lcn, it first tries the
clusters from there on.  If that does not satisfy the request, it moves
on to the zone's current position, but keeps the pass, the zone_end and
the has_guess state it had.  That loses free clusters in two ways.  If
the search had already moved on to pass 2, whose range ends at the hint,
it scans only from the zone position to the hint and misses every free
cluster between the start of the zone and the zone position.  If no
cluster had been tested yet, has_guess is still set, so the cluster at
the zone position is tried as if it had been the hint, and when that
cluster is in use the rest of the bitmap buffer is skipped.

Both happen when the hint lies at or past the end of the volume, which
ntfs_attr_map_cluster() produces when it extrapolates from the last
allocated run across a hole.  The allocator then finds nothing, shrinks
the MFT zone to nothing trying to satisfy the request, and fails with
-ENOSPC.

To reproduce on a 128 MiB volume with 4 KiB clusters, extend two files
in turn by one cluster at a time, alternating between fallocate(),
write() past EOF and truncate() up, so that the runs of each file are
separated by holes.  When the volume is full, truncate both files back
to 1 MiB and start again.  During the second round a one-cluster
fallocate() fails with ENOSPC while 40 MiB is free.

Start the zone over as if no hint had been given.

Fixes: 11ccc9107dc4 ("ntfs: update runlist handling and cluster allocator")
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
---
 fs/ntfs/lcnalloc.c | 18 +++++++++++++++---
 1 file changed, 15 insertions(+), 3 deletions(-)

diff --git a/fs/ntfs/lcnalloc.c b/fs/ntfs/lcnalloc.c
index 0d6cd08ee2e76..30faf422a5424 100644
--- a/fs/ntfs/lcnalloc.c
+++ b/fs/ntfs/lcnalloc.c
@@ -506,13 +506,25 @@ struct runlist_element *ntfs_cluster_alloc(struct ntfs_volume *vol, const s64 st
 		}
 
 		if (!used_zone_pos) {
+			/*
+			 * Leaving @start_lcn for the zone position starts the
+			 * zone over as if no hint had been given, even if the
+			 * search had already reached pass 2, whose range ends
+			 * at @start_lcn.
+			 */
 			used_zone_pos = 1;
-			if (search_zone == 1)
+			has_guess = 0;
+			pass = 1;
+			if (search_zone == 1) {
 				zone_start = vol->mft_zone_pos;
-			else if (search_zone == 2)
+				zone_end = vol->mft_zone_end;
+			} else if (search_zone == 2) {
 				zone_start = vol->data1_zone_pos;
-			else
+				zone_end = vol->nr_clusters;
+			} else {
 				zone_start = vol->data2_zone_pos;
+				zone_end = vol->mft_zone_start;
+			}
 
 			if (!zone_start || zone_start == vol->mft_zone_start ||
 			    zone_start == vol->mft_zone_end)
-- 
2.55.0


^ permalink raw reply	[flat|nested] 7+ messages in thread

* [PATCH 3/3] ntfs: do not refuse fallocate() when the MFT zone is empty
  2026-09-27 10:57 [PATCH 0/3] ntfs: fix spurious EIO and ENOSPC on nearly full volumes Matthias Goergens
  2026-09-27 10:57 ` [PATCH 1/3] ntfs: set the attribute list size before reserving space for it Matthias Goergens
  2026-09-27 10:57 ` [PATCH 2/3] ntfs: restart the zone search when the allocation hint fails Matthias Goergens
@ 2026-09-27 10:57 ` Matthias Goergens
  2026-09-27 18:11   ` liubaolin
  2 siblings, 1 reply; 7+ messages in thread
From: Matthias Goergens @ 2026-09-27 10:57 UTC (permalink / raw)
  To: Namjae Jeon, Hyunchul Lee; +Cc: ntfs, linux-kernel

ntfs_fallocate() fails with -ENOSPC whenever the MFT zone has shrunk to
nothing.  The cluster allocator gives the MFT zone up when it cannot
find free clusters elsewhere, and nothing restores it before the next
mount, so after one such failure every fallocate() on the volume fails
with ENOSPC, punching holes included, however much space is freed.  The
allocator failure fixed by the previous patch was enough to get there:
afterwards a one-cluster fallocate() failed with ENOSPC while 123 MiB of
the 128 MiB volume were free.

write() and truncate() do not look at the MFT zone, and
ntfs_allocate_range() already fails with -ENOSPC when too few clusters
are free.  Remove the check.

Fixes: 9c87959601e8 ("ntfs: update file operations")
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
---
 fs/ntfs/file.c | 3 ---
 1 file changed, 3 deletions(-)

diff --git a/fs/ntfs/file.c b/fs/ntfs/file.c
index 3ec82715a5882..762045d262963 100644
--- a/fs/ntfs/file.c
+++ b/fs/ntfs/file.c
@@ -1163,9 +1163,6 @@ static long ntfs_fallocate(struct file *file, int mode, loff_t offset, loff_t le
 	if (!NVolFreeClusterKnown(vol))
 		wait_event(vol->free_waitq, NVolFreeClusterKnown(vol));
 
-	if ((ni->vol->mft_zone_end - ni->vol->mft_zone_start) == 0)
-		return -ENOSPC;
-
 	if (NInoNonResident(ni) && !NInoFullyMapped(ni)) {
 		down_write(&ni->runlist.lock);
 		err = ntfs_attr_map_whole_runlist(ni);
-- 
2.55.0


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH 2/3] ntfs: restart the zone search when the allocation hint fails
  2026-09-27 10:57 ` [PATCH 2/3] ntfs: restart the zone search when the allocation hint fails Matthias Goergens
@ 2026-09-27 18:06   ` liubaolin
  0 siblings, 0 replies; 7+ messages in thread
From: liubaolin @ 2026-09-27 18:06 UTC (permalink / raw)
  To: Matthias Goergens, Namjae Jeon, Hyunchul Lee; +Cc: ntfs, linux-kernel



在 2026/9/27 18:57, Matthias Goergens 写道:
> When ntfs_cluster_alloc() is given a start_lcn, it first tries the
> clusters from there on.  If that does not satisfy the request, it moves
> on to the zone's current position, but keeps the pass, the zone_end and
> the has_guess state it had.  That loses free clusters in two ways.  If
> the search had already moved on to pass 2, whose range ends at the hint,
> it scans only from the zone position to the hint and misses every free
> cluster between the start of the zone and the zone position.  If no
> cluster had been tested yet, has_guess is still set, so the cluster at
> the zone position is tried as if it had been the hint, and when that
> cluster is in use the rest of the bitmap buffer is skipped.
> 
> Both happen when the hint lies at or past the end of the volume, which
> ntfs_attr_map_cluster() produces when it extrapolates from the last
> allocated run across a hole.  The allocator then finds nothing, shrinks
> the MFT zone to nothing trying to satisfy the request, and fails with
> -ENOSPC.
> 
> To reproduce on a 128 MiB volume with 4 KiB clusters, extend two files
> in turn by one cluster at a time, alternating between fallocate(),
> write() past EOF and truncate() up, so that the runs of each file are
> separated by holes.  When the volume is full, truncate both files back
> to 1 MiB and start again.  During the second round a one-cluster
> fallocate() fails with ENOSPC while 40 MiB is free.
> 
> Start the zone over as if no hint had been given.
> 
> Fixes: 11ccc9107dc4 ("ntfs: update runlist handling and cluster allocator")
> Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
> ---
>   fs/ntfs/lcnalloc.c | 18 +++++++++++++++---
>   1 file changed, 15 insertions(+), 3 deletions(-)
> 
> diff --git a/fs/ntfs/lcnalloc.c b/fs/ntfs/lcnalloc.c
> index 0d6cd08ee2e76..30faf422a5424 100644
> --- a/fs/ntfs/lcnalloc.c
> +++ b/fs/ntfs/lcnalloc.c
> @@ -506,13 +506,25 @@ struct runlist_element *ntfs_cluster_alloc(struct ntfs_volume *vol, const s64 st
>   		}
>   
>   		if (!used_zone_pos) {
> +			/*
> +			 * Leaving @start_lcn for the zone position starts the
> +			 * zone over as if no hint had been given, even if the
> +			 * search had already reached pass 2, whose range ends
> +			 * at @start_lcn.
> +			 */
>   			used_zone_pos = 1;
> -			if (search_zone == 1)
> +			has_guess = 0;
> +			pass = 1;

Hi Matthias,
   Resetting pass, zone_end and has_guess here fixes the missed free 
clusters when falling back from the hint to the zone position.

   However, !used_zone_pos only means that the zone position has not 
been used yet; it does not mean that no clusters have been allocated.
   A partial hint allocation can also reach this block at the end of the 
bitmap buffer. Clearing has_guess then allows a nonadjacent run to be 
appended even when is_contig is true.

   Could we return the existing partial run before resetting the search 
state, like this?

   	if (!used_zone_pos) {
   		if (is_contig && rlpos)
   			goto out;

   		used_zone_pos = 1;
   		has_guess = 0;
   		pass = 1;
   		...
   	}

Thanks,
Baolin.

> +			if (search_zone == 1) {
>   				zone_start = vol->mft_zone_pos;
> -			else if (search_zone == 2)
> +				zone_end = vol->mft_zone_end;
> +			} else if (search_zone == 2) {
>   				zone_start = vol->data1_zone_pos;
> -			else
> +				zone_end = vol->nr_clusters;
> +			} else {
>   				zone_start = vol->data2_zone_pos;
> +				zone_end = vol->mft_zone_start;
> +			}
>   
>   			if (!zone_start || zone_start == vol->mft_zone_start ||
>   			    zone_start == vol->mft_zone_end)


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH 1/3] ntfs: set the attribute list size before reserving space for it
  2026-09-27 10:57 ` [PATCH 1/3] ntfs: set the attribute list size before reserving space for it Matthias Goergens
@ 2026-09-27 18:11   ` liubaolin
  0 siblings, 0 replies; 7+ messages in thread
From: liubaolin @ 2026-09-27 18:11 UTC (permalink / raw)
  To: Matthias Goergens, Namjae Jeon, Hyunchul Lee; +Cc: ntfs, linux-kernel



在 2026/9/27 18:57, Matthias Goergens 写道:
> ntfs_attrlist_update_locked() resizes the attribute list and then, for
> $MFT, tries to reserve the maximum list size.  When that allocation
> fails with -ENOSPC it falls back to ntfs_attrlist_repack(), which reads
> the list through ntfs_inode_attr_pread().  The read stops at i_size,
> but i_size is only updated after the reserve, so when the list has just
> grown the repack gets a short read and returns -EIO.  Only -ENOSPC from
> the reserve is ignored, so the -EIO fails the $MFT extension and with
> it the file creation that needed a new mft record.
> 
> On a 32 MiB volume with 512-byte clusters whose $MFT has a non-resident
> attribute list, 591 of 1500 creations of empty files succeed and the
> rest fail with EIO while 742 KiB is still free:
> 
>    ntfs: (device vda): ntfs_attrlist_update_locked(): Failed to reserve attribute list space
>    ntfs: Failed add attr entry to attrlist
>    ntfs: MP update failed
>    ntfs: (device vda): ntfs_mft_record_alloc(): Failed to extend mft data allocation.
> 
> Update i_size as soon as the list has been resized.  File creation then
> goes on until the volume is full and fails with ENOSPC.
> 
> Fixes: b1d732e62a5b ("ntfs: repack $MFT/$ATTRIBUTE LIST")
> Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
> ---
> The volume is frag.img from
> https://github.com/matthiasgoergens/linux/tree/reproducer/2026-09-26-ntfs-mft-runlist
> (mft-bootstrap/frag.img.xz); creating empty files on it until one fails
> is enough.  Tested with KASAN and lockdep: after the patch the files
> written before the volume filled read back intact after a remount, and
> ntfsfix and ntfsresize --info find nothing wrong with the image.
> ---
>   fs/ntfs/attrlist.c | 5 +++--
>   1 file changed, 3 insertions(+), 2 deletions(-)
> 
> diff --git a/fs/ntfs/attrlist.c b/fs/ntfs/attrlist.c
> index 1bbd2bc62c582..62dfb96e31a3c 100644
> --- a/fs/ntfs/attrlist.c
> +++ b/fs/ntfs/attrlist.c
> @@ -252,6 +252,9 @@ int ntfs_attrlist_update_locked(struct ntfs_inode *base_ni,
>   		return err;
>   	}
>   
> +	/* ntfs_attrlist_repack() below reads the list up to i_size. */
> +	i_size_write(attr_vi, base_ni->attr_list_size);
> +
>   	/*
>   	 * Reserve the maximum legal list size while the MFT metadata area is
>   	 * still easy to allocate contiguously. This prevents a later list entry
> @@ -279,8 +282,6 @@ int ntfs_attrlist_update_locked(struct ntfs_inode *base_ni,
>   		}
>   	}
>   
> -	i_size_write(attr_vi, base_ni->attr_list_size);
> -
>   	if (NInoNonResident(attr_ni) && !NInoAttrListNonResident(base_ni))
>   		NInoSetAttrListNonResident(base_ni);
>   

Looks good to me.

Reviewed-by: Baolin Liu <liubaolin@kylinos.cn>


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH 3/3] ntfs: do not refuse fallocate() when the MFT zone is empty
  2026-09-27 10:57 ` [PATCH 3/3] ntfs: do not refuse fallocate() when the MFT zone is empty Matthias Goergens
@ 2026-09-27 18:11   ` liubaolin
  0 siblings, 0 replies; 7+ messages in thread
From: liubaolin @ 2026-09-27 18:11 UTC (permalink / raw)
  To: Matthias Goergens, Namjae Jeon, Hyunchul Lee; +Cc: ntfs, linux-kernel



在 2026/9/27 18:57, Matthias Goergens 写道:
> ntfs_fallocate() fails with -ENOSPC whenever the MFT zone has shrunk to
> nothing.  The cluster allocator gives the MFT zone up when it cannot
> find free clusters elsewhere, and nothing restores it before the next
> mount, so after one such failure every fallocate() on the volume fails
> with ENOSPC, punching holes included, however much space is freed.  The
> allocator failure fixed by the previous patch was enough to get there:
> afterwards a one-cluster fallocate() failed with ENOSPC while 123 MiB of
> the 128 MiB volume were free.
> 
> write() and truncate() do not look at the MFT zone, and
> ntfs_allocate_range() already fails with -ENOSPC when too few clusters
> are free.  Remove the check.
> 
> Fixes: 9c87959601e8 ("ntfs: update file operations")
> Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
> ---
>   fs/ntfs/file.c | 3 ---
>   1 file changed, 3 deletions(-)
> 
> diff --git a/fs/ntfs/file.c b/fs/ntfs/file.c
> index 3ec82715a5882..762045d262963 100644
> --- a/fs/ntfs/file.c
> +++ b/fs/ntfs/file.c
> @@ -1163,9 +1163,6 @@ static long ntfs_fallocate(struct file *file, int mode, loff_t offset, loff_t le
>   	if (!NVolFreeClusterKnown(vol))
>   		wait_event(vol->free_waitq, NVolFreeClusterKnown(vol));
>   
> -	if ((ni->vol->mft_zone_end - ni->vol->mft_zone_start) == 0)
> -		return -ENOSPC;
> -
>   	if (NInoNonResident(ni) && !NInoFullyMapped(ni)) {
>   		down_write(&ni->runlist.lock);
>   		err = ntfs_attr_map_whole_runlist(ni);

Looks good to me.

Reviewed-by: Baolin Liu <liubaolin@kylinos.cn>


^ permalink raw reply	[flat|nested] 7+ messages in thread

end of thread, other threads:[~2026-09-27 18:11 UTC | newest]

Thread overview: 7+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-27 10:57 [PATCH 0/3] ntfs: fix spurious EIO and ENOSPC on nearly full volumes Matthias Goergens
2026-09-27 10:57 ` [PATCH 1/3] ntfs: set the attribute list size before reserving space for it Matthias Goergens
2026-09-27 18:11   ` liubaolin
2026-09-27 10:57 ` [PATCH 2/3] ntfs: restart the zone search when the allocation hint fails Matthias Goergens
2026-09-27 18:06   ` liubaolin
2026-09-27 10:57 ` [PATCH 3/3] ntfs: do not refuse fallocate() when the MFT zone is empty Matthias Goergens
2026-09-27 18:11   ` liubaolin

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®