mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* ftruncate() after FICLONERANGE costs 300-600us/file, ~10x more than the clone itself
@ 2026-07-06 21:45 Matteo Croce
  2026-07-07  5:56 ` Darrick J. Wong
  0 siblings, 1 reply; 5+ messages in thread
From: Matteo Croce @ 2026-07-06 21:45 UTC (permalink / raw)
  To: linux-xfs; +Cc: LKML

Hi,

while adding reflink support to GNU tar I found that on XFS a
FICLONERANGE followed by a shrinking ftruncate() is pathologically
expensive: 300-600us for the truncate, plus another ~100us at close()
for the post-EOF cleanup, around 10x the cost of the clone itself.
The file has no dirty pages at that point, so this is not writeback:
it looks like synchronous per-file metadata work. BtrFS (the only
other mainstream filesystem with reflinks, so my only comparison
point) does the same shrink in a flat ~15us.

Any program cloning a non-block-aligned range out of a larger file
hits this (in my case, tar extracting a member from an archive):
FICLONERANGE requires block-aligned lengths (unless the range ends at
the source EOF), so the only option is to clone the length rounded up
to the fs block and truncate to the real size. Whole-file cloners
like cp --reflink use FICLONE and are not affected, which is probably
why nobody noticed so far. Extracting a linux source tree (~93k
files) this way with tar --reflink on XFS is 3.4x SLOWER than a
regular copy (68s vs 20s); the same binary on BtrFS, on another
partition of the same spinning disk, is 1.7x faster than the copy.

The core of the issue reproduces with xfs_io alone (on a freshly
created filesystem with mkfs defaults, rmapbt=1):

  $ dd if=/dev/urandom of=src.dat bs=1M count=1; sync
  $ strace -T -e trace=ioctl,ftruncate xfs_io -f \
        -c "reflink src.dat 0 0 90112" -c "truncate 90000" dest.dat
  ioctl(3, FICLONERANGE, {src_fd=4, src_offset=0, src_length=90112,
        dest_offset=0}) = 0 <0.000089>
  ftruncate(3, 90000) = 0 <0.000446>        <-- shrink by 112 bytes

and the control, a no-op truncate at the exact cloned size:

  $ xfs_io -f -c "reflink src.dat 0 0 90112" -c "truncate 90112" ...
  ioctl(3, FICLONERANGE, ...) = 0 <0.000088>
  ftruncate(3, 90112) = 0 <0.000030>

For the statistics over many files, the reproducer below creates
20000 files of SIZE bytes from a common (fsync'ed) source file,
timing each syscall. Modes:

  write        plain write() of the content
  write-trunc  write() rounded up to the block, then ftruncate()
  clone-tail   FICLONERANGE of the whole blocks, then write() of
               the final partial block (no truncate)
  clone-trunc  FICLONERANGE rounded up to the block, then
               ftruncate() to the real size

XFS with mkfs defaults (rmapbt=1), per-file cost of the ftruncate in
clone-trunc, plus the neighboring syscalls for context (us):

  size      clone   trunc   close     whole loop
  1000       40.2   294.2    82.4        9.10s
  5000       60.8   285.5    95.5        9.61s
  20000      61.8   416.0    99.4       12.31s
  65536      50.1     7.3     2.7        1.86s   <-- block-aligned,
  90000      62.9   306.7    94.1       10.06s       truncate is a no-op

The 65536 control row (exact multiple of the block size) shows the
whole overhead vanishing when the truncate does not shrink into the
cloned extent; at 90000 (21 blocks + 3984 bytes) it is back.

To rule out the reverse-mapping btree, I recreated the fs on the same
partition with -m rmapbt=0 and reran the matrix:

  size      clone   trunc   close     whole loop
  1000       21.8   295.8    52.0        8.14s
  5000       27.1   332.8    56.2        9.09s
  20000      27.1   366.6    58.4        9.81s
  65536      19.7     6.7     2.6        1.21s
  90000      27.1   574.3    57.0       13.94s

The clone and the close get cheaper without rmapbt, but the truncate
does not improve at all; if anything it grows with the number of
cloned blocks (~13us per block from 1 to 22 blocks). So the rmap
btree is not where the cost is.

For reference, on the same mkfs-default filesystem:

  - clone-tail (no truncate): clone ~50us, tail write ~12us, close
    ~7us, at every size. A prototype of tar using this scheme turns
    the 3.4x slowdown into a 1.6x speedup on XFS, so a userspace
    workaround exists, but it should not be needed.
  - The clone path is otherwise excellent: at size=90000, plain
    write pays 11.6s of writeback at syncfs time, the clone modes
    ~0.1s.
  - BtrFS clone-trunc for comparison: trunc ~15us and close ~2.3us,
    flat from 1k to 90k.

Unrelated, maybe: the write-trunc mode (not shown in the tables
above) also shows that on XFS an ftruncate on a file with dirty
delalloc data synchronously flushes it (the deferred writeback cost
moves into the truncate, sync time drops to ~0.1s), even when the
truncate does not change the size: at size=65536 the no-op ftruncate
costs 514us on the freshly written file, versus the 7.3us that the
same no-op costs on a clean cloned file in the table above. BtrFS
only pays this for an actual shrink, not for the no-op (19.6us).
This does not apply to the clone case above, where the cloned file
has nothing dirty, which is why the 300-600us look like pure
metadata work.

Environment:
  kernel 7.1.2 vanilla
  All filesystems were freshly created on the same partition of a
  spinning HDD, with mkfs defaults.
  meta-data: isize=512 agcount=4
             sectsz=4096 crc=1 finobt=1, sparse=1, rmapbt=1
             reflink=1 bigtime=1 inobtcount=1 nrext64=1
  mount: rw,noatime,inode64,logbufs=8,logbsize=32k,noquota
  data:  bsize=4096

The benchmark and its driver script are here:

  https://gist.github.com/teknoraver/67b50ce366d7cf430bb7f82a8018acfa

(gcc -O2 -o clone_bench clone_bench.c; the script runs the whole
mode/size matrix in fresh subdirectories of the given target dir.)

Happy to run further tests (tracepoints, perf, other mkfs/mount
options) if useful.

Regards,
-- 
Matteo Croce

per aspera ad upstream

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: ftruncate() after FICLONERANGE costs 300-600us/file, ~10x more than the clone itself
  2026-07-06 21:45 ftruncate() after FICLONERANGE costs 300-600us/file, ~10x more than the clone itself Matteo Croce
@ 2026-07-07  5:56 ` Darrick J. Wong
  2026-07-07  6:20   ` Christoph Hellwig
  2026-07-07 15:07   ` Matteo Croce
  0 siblings, 2 replies; 5+ messages in thread
From: Darrick J. Wong @ 2026-07-07  5:56 UTC (permalink / raw)
  To: Matteo Croce; +Cc: linux-xfs, LKML

On Mon, Jul 06, 2026 at 11:45:07PM +0200, Matteo Croce wrote:
> Hi,
> 
> while adding reflink support to GNU tar I found that on XFS a
> FICLONERANGE followed by a shrinking ftruncate() is pathologically
> expensive: 300-600us for the truncate, plus another ~100us at close()
> for the post-EOF cleanup, around 10x the cost of the clone itself.
> The file has no dirty pages at that point, so this is not writeback:
> it looks like synchronous per-file metadata work. BtrFS (the only
> other mainstream filesystem with reflinks, so my only comparison
> point) does the same shrink in a flat ~15us.
> 
> Any program cloning a non-block-aligned range out of a larger file
> hits this (in my case, tar extracting a member from an archive):
> FICLONERANGE requires block-aligned lengths (unless the range ends at
> the source EOF), so the only option is to clone the length rounded up
> to the fs block and truncate to the real size. Whole-file cloners
> like cp --reflink use FICLONE and are not affected, which is probably
> why nobody noticed so far. Extracting a linux source tree (~93k
> files) this way with tar --reflink on XFS is 3.4x SLOWER than a
> regular copy (68s vs 20s); the same binary on BtrFS, on another
> partition of the same spinning disk, is 1.7x faster than the copy.
> 
> The core of the issue reproduces with xfs_io alone (on a freshly
> created filesystem with mkfs defaults, rmapbt=1):
> 
>   $ dd if=/dev/urandom of=src.dat bs=1M count=1; sync
>   $ strace -T -e trace=ioctl,ftruncate xfs_io -f \
>         -c "reflink src.dat 0 0 90112" -c "truncate 90000" dest.dat
>   ioctl(3, FICLONERANGE, {src_fd=4, src_offset=0, src_length=90112,
>         dest_offset=0}) = 0 <0.000089>
>   ftruncate(3, 90000) = 0 <0.000446>        <-- shrink by 112 bytes

XFS zeroes the tail block when you truncate down, which causes an out of
place write.

Assuming src.dat is the fully written 90000 byte file, can you

$ xfs_io -c "reflink src.dat 86016 86016 0" dest.dat

to link only the eof-block into dest.dat and keep its file size at
90000?

--D

> and the control, a no-op truncate at the exact cloned size:
> 
>   $ xfs_io -f -c "reflink src.dat 0 0 90112" -c "truncate 90112" ...
>   ioctl(3, FICLONERANGE, ...) = 0 <0.000088>
>   ftruncate(3, 90112) = 0 <0.000030>
> 
> For the statistics over many files, the reproducer below creates
> 20000 files of SIZE bytes from a common (fsync'ed) source file,
> timing each syscall. Modes:
> 
>   write        plain write() of the content
>   write-trunc  write() rounded up to the block, then ftruncate()
>   clone-tail   FICLONERANGE of the whole blocks, then write() of
>                the final partial block (no truncate)
>   clone-trunc  FICLONERANGE rounded up to the block, then
>                ftruncate() to the real size
> 
> XFS with mkfs defaults (rmapbt=1), per-file cost of the ftruncate in
> clone-trunc, plus the neighboring syscalls for context (us):
> 
>   size      clone   trunc   close     whole loop
>   1000       40.2   294.2    82.4        9.10s
>   5000       60.8   285.5    95.5        9.61s
>   20000      61.8   416.0    99.4       12.31s
>   65536      50.1     7.3     2.7        1.86s   <-- block-aligned,
>   90000      62.9   306.7    94.1       10.06s       truncate is a no-op
> 
> The 65536 control row (exact multiple of the block size) shows the
> whole overhead vanishing when the truncate does not shrink into the
> cloned extent; at 90000 (21 blocks + 3984 bytes) it is back.
> 
> To rule out the reverse-mapping btree, I recreated the fs on the same
> partition with -m rmapbt=0 and reran the matrix:
> 
>   size      clone   trunc   close     whole loop
>   1000       21.8   295.8    52.0        8.14s
>   5000       27.1   332.8    56.2        9.09s
>   20000      27.1   366.6    58.4        9.81s
>   65536      19.7     6.7     2.6        1.21s
>   90000      27.1   574.3    57.0       13.94s
> 
> The clone and the close get cheaper without rmapbt, but the truncate
> does not improve at all; if anything it grows with the number of
> cloned blocks (~13us per block from 1 to 22 blocks). So the rmap
> btree is not where the cost is.
> 
> For reference, on the same mkfs-default filesystem:
> 
>   - clone-tail (no truncate): clone ~50us, tail write ~12us, close
>     ~7us, at every size. A prototype of tar using this scheme turns
>     the 3.4x slowdown into a 1.6x speedup on XFS, so a userspace
>     workaround exists, but it should not be needed.
>   - The clone path is otherwise excellent: at size=90000, plain
>     write pays 11.6s of writeback at syncfs time, the clone modes
>     ~0.1s.
>   - BtrFS clone-trunc for comparison: trunc ~15us and close ~2.3us,
>     flat from 1k to 90k.
> 
> Unrelated, maybe: the write-trunc mode (not shown in the tables
> above) also shows that on XFS an ftruncate on a file with dirty
> delalloc data synchronously flushes it (the deferred writeback cost
> moves into the truncate, sync time drops to ~0.1s), even when the
> truncate does not change the size: at size=65536 the no-op ftruncate
> costs 514us on the freshly written file, versus the 7.3us that the
> same no-op costs on a clean cloned file in the table above. BtrFS
> only pays this for an actual shrink, not for the no-op (19.6us).
> This does not apply to the clone case above, where the cloned file
> has nothing dirty, which is why the 300-600us look like pure
> metadata work.
> 
> Environment:
>   kernel 7.1.2 vanilla
>   All filesystems were freshly created on the same partition of a
>   spinning HDD, with mkfs defaults.
>   meta-data: isize=512 agcount=4
>              sectsz=4096 crc=1 finobt=1, sparse=1, rmapbt=1
>              reflink=1 bigtime=1 inobtcount=1 nrext64=1
>   mount: rw,noatime,inode64,logbufs=8,logbsize=32k,noquota
>   data:  bsize=4096
> 
> The benchmark and its driver script are here:
> 
>   https://gist.github.com/teknoraver/67b50ce366d7cf430bb7f82a8018acfa
> 
> (gcc -O2 -o clone_bench clone_bench.c; the script runs the whole
> mode/size matrix in fresh subdirectories of the given target dir.)
> 
> Happy to run further tests (tracepoints, perf, other mkfs/mount
> options) if useful.
> 
> Regards,
> -- 
> Matteo Croce
> 
> per aspera ad upstream
> 

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: ftruncate() after FICLONERANGE costs 300-600us/file, ~10x more than the clone itself
  2026-07-07  5:56 ` Darrick J. Wong
@ 2026-07-07  6:20   ` Christoph Hellwig
  2026-07-07 15:09     ` Matteo Croce
  2026-07-07 15:07   ` Matteo Croce
  1 sibling, 1 reply; 5+ messages in thread
From: Christoph Hellwig @ 2026-07-07  6:20 UTC (permalink / raw)
  To: Darrick J. Wong; +Cc: Matteo Croce, linux-xfs, LKML

On Mon, Jul 06, 2026 at 10:56:51PM -0700, Darrick J. Wong wrote:
> XFS zeroes the tail block when you truncate down, which causes an out of
> place write.

All the file systems do it (or at least should).  But somehow we
manage to be slower.

I wonder if we hit the filemap_write_and_wait_range case in
xfs_vn_setattr_size due to a non-uptodate i_disk_size for this
workload somehow?  Although reflink should update i_disk_size
properly.

> 
> Assuming src.dat is the fully written 90000 byte file, can you
> 
> $ xfs_io -c "reflink src.dat 86016 86016 0" dest.dat
> 
> to link only the eof-block into dest.dat and keep its file size at
> 90000?

Only doing the reflink block aligned and copying data for partial
blocks should indeed always be faster.  But I wonder if we have
some dragons lurking in the truncate down path.


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: ftruncate() after FICLONERANGE costs 300-600us/file, ~10x more than the clone itself
  2026-07-07  5:56 ` Darrick J. Wong
  2026-07-07  6:20   ` Christoph Hellwig
@ 2026-07-07 15:07   ` Matteo Croce
  1 sibling, 0 replies; 5+ messages in thread
From: Matteo Croce @ 2026-07-07 15:07 UTC (permalink / raw)
  To: Darrick J. Wong; +Cc: linux-xfs, LKML, Christoph Hellwig

Il giorno mar 7 lug 2026 alle ore 07:56 Darrick J. Wong
<djwong@kernel.org> ha scritto:
> Assuming src.dat is the fully written 90000 byte file, can you
>
> $ xfs_io -c "reflink src.dat 86016 86016 0" dest.dat
>
> to link only the eof-block into dest.dat and keep its file size at
> 90000?
>

Yes, that works nicely:

  $ dd if=/dev/urandom of=src.dat bs=90000 count=1; sync
  $ strace -T -e trace=ioctl,ftruncate xfs_io -f \
        -c "reflink src.dat 86016 86016 0" dest.dat
  ioctl(3, FICLONERANGE, {src_fd=4, src_offset=86016, src_length=0,
        dest_offset=86016}) = 0 <0.000093>
  $ stat -c %s dest.dat
  90000
  $ cmp -i 86016 src.dat dest.dat && echo tail OK
  tail OK
  $ filefrag -v dest.dat
  Filesystem type is: 58465342
  File size of dest.dat is 90000 (22 blocks of 4096 bytes)
   ext:     logical_offset:        physical_offset: length:   expected: flags:
     0:       21..      21:         45..        45:      1:
21: last,shared,eof
  dest.dat: 1 extent found

The clone of the unaligned EOF block costs the same ~90us as any
other clone, it sets the destination size without any truncate, and
the block is really shared. So the whole overhead I measured is in
the truncate-down zeroing of the shared tail block, as you said.

The out of place write is also nicely visible in filefrag after the
clone+truncate sequence (src.dat is a 1 MiB file here, so the cloned
range lies inside it): only the tail block loses the sharing, and it
lands out of sequence with respect to the shared run:

  $ xfs_io -f -c "reflink src.dat 0 0 90112" -c "truncate 90000" dest.dat
  $ filefrag -v dest.dat
  File size of dest.dat is 90000 (22 blocks of 4096 bytes)
   ext:     logical_offset:        physical_offset: length:   expected: flags:
     0:        0..      20:         46..        66:     21:             shared
     1:       21..      21:         45..        45:      1:         67: last,eof
  dest.dat: 2 extents found

Sadly, tar cannot use the eof-block shape for the general case: the
member data does not end at the archive EOF (the end-of-archive
marker follows), so there is no source EOF to link from. The plan
for tar on the affected filesystems is to clone the whole blocks and
write the final partial block, which avoids the truncate entirely
and in my prototype turns the 3.4x slowdown into a 1.6x speedup.

Regards,
-- 
per aspera ad upstream

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: ftruncate() after FICLONERANGE costs 300-600us/file, ~10x more than the clone itself
  2026-07-07  6:20   ` Christoph Hellwig
@ 2026-07-07 15:09     ` Matteo Croce
  0 siblings, 0 replies; 5+ messages in thread
From: Matteo Croce @ 2026-07-07 15:09 UTC (permalink / raw)
  To: Christoph Hellwig; +Cc: Darrick J. Wong, linux-xfs, LKML

Il giorno mar 7 lug 2026 alle ore 08:20 Christoph Hellwig
<hch@infradead.org> ha scritto:
> I wonder if we hit the filemap_write_and_wait_range case in
> xfs_vn_setattr_size due to a non-uptodate i_disk_size for this
> workload somehow?  Although reflink should update i_disk_size
> properly.

Possibly related datapoint, from the write-trunc mode of my
benchmark (write of the size rounded up to the block, then
ftruncate to the real size): the truncate synchronously flushes the
delalloc data even when it does not change the file size at all. A
no-op ftruncate at a block-aligned 65536 costs 514us on a freshly
written file, versus 7us for the same no-op on a freshly cloned
file, and the syncfs time of the run drops accordingly, so the
deferred writeback cost moves into the truncate. If a tracepoint or
counter would help pinpointing where the time goes, I am happy to
run it.

> >
> > Assuming src.dat is the fully written 90000 byte file, can you
> >
> > $ xfs_io -c "reflink src.dat 86016 86016 0" dest.dat
> >
> > to link only the eof-block into dest.dat and keep its file size at
> > 90000?
>
> Only doing the reflink block aligned and copying data for partial
> blocks should indeed always be faster.  But I wonder if we have
> some dragons lurking in the truncate down path.
>

That is the plan for tar: clone the whole blocks and write the
final partial block, avoiding the truncate entirely. In my
prototype this turns the 3.4x slowdown into a 1.6x speedup on XFS,
with the small cost that the tail block of each member is no longer
shared with the archive.

Regards,
-- 
per aspera ad upstream

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-07-07 15:09 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-07-06 21:45 ftruncate() after FICLONERANGE costs 300-600us/file, ~10x more than the clone itself Matteo Croce
2026-07-07  5:56 ` Darrick J. Wong
2026-07-07  6:20   ` Christoph Hellwig
2026-07-07 15:09     ` Matteo Croce
2026-07-07 15:07   ` Matteo Croce

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®