mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* ftruncate() after FICLONERANGE costs 300-600us/file, ~10x more than the clone itself
@ 2026-07-06 21:45 Matteo Croce
  2026-07-07  5:56 ` Darrick J. Wong
  0 siblings, 1 reply; 5+ messages in thread
From: Matteo Croce @ 2026-07-06 21:45 UTC (permalink / raw)
  To: linux-xfs; +Cc: LKML

Hi,

while adding reflink support to GNU tar I found that on XFS a
FICLONERANGE followed by a shrinking ftruncate() is pathologically
expensive: 300-600us for the truncate, plus another ~100us at close()
for the post-EOF cleanup, around 10x the cost of the clone itself.
The file has no dirty pages at that point, so this is not writeback:
it looks like synchronous per-file metadata work. BtrFS (the only
other mainstream filesystem with reflinks, so my only comparison
point) does the same shrink in a flat ~15us.

Any program cloning a non-block-aligned range out of a larger file
hits this (in my case, tar extracting a member from an archive):
FICLONERANGE requires block-aligned lengths (unless the range ends at
the source EOF), so the only option is to clone the length rounded up
to the fs block and truncate to the real size. Whole-file cloners
like cp --reflink use FICLONE and are not affected, which is probably
why nobody noticed so far. Extracting a linux source tree (~93k
files) this way with tar --reflink on XFS is 3.4x SLOWER than a
regular copy (68s vs 20s); the same binary on BtrFS, on another
partition of the same spinning disk, is 1.7x faster than the copy.

The core of the issue reproduces with xfs_io alone (on a freshly
created filesystem with mkfs defaults, rmapbt=1):

  $ dd if=/dev/urandom of=src.dat bs=1M count=1; sync
  $ strace -T -e trace=ioctl,ftruncate xfs_io -f \
        -c "reflink src.dat 0 0 90112" -c "truncate 90000" dest.dat
  ioctl(3, FICLONERANGE, {src_fd=4, src_offset=0, src_length=90112,
        dest_offset=0}) = 0 <0.000089>
  ftruncate(3, 90000) = 0 <0.000446>        <-- shrink by 112 bytes

and the control, a no-op truncate at the exact cloned size:

  $ xfs_io -f -c "reflink src.dat 0 0 90112" -c "truncate 90112" ...
  ioctl(3, FICLONERANGE, ...) = 0 <0.000088>
  ftruncate(3, 90112) = 0 <0.000030>

For the statistics over many files, the reproducer below creates
20000 files of SIZE bytes from a common (fsync'ed) source file,
timing each syscall. Modes:

  write        plain write() of the content
  write-trunc  write() rounded up to the block, then ftruncate()
  clone-tail   FICLONERANGE of the whole blocks, then write() of
               the final partial block (no truncate)
  clone-trunc  FICLONERANGE rounded up to the block, then
               ftruncate() to the real size

XFS with mkfs defaults (rmapbt=1), per-file cost of the ftruncate in
clone-trunc, plus the neighboring syscalls for context (us):

  size      clone   trunc   close     whole loop
  1000       40.2   294.2    82.4        9.10s
  5000       60.8   285.5    95.5        9.61s
  20000      61.8   416.0    99.4       12.31s
  65536      50.1     7.3     2.7        1.86s   <-- block-aligned,
  90000      62.9   306.7    94.1       10.06s       truncate is a no-op

The 65536 control row (exact multiple of the block size) shows the
whole overhead vanishing when the truncate does not shrink into the
cloned extent; at 90000 (21 blocks + 3984 bytes) it is back.

To rule out the reverse-mapping btree, I recreated the fs on the same
partition with -m rmapbt=0 and reran the matrix:

  size      clone   trunc   close     whole loop
  1000       21.8   295.8    52.0        8.14s
  5000       27.1   332.8    56.2        9.09s
  20000      27.1   366.6    58.4        9.81s
  65536      19.7     6.7     2.6        1.21s
  90000      27.1   574.3    57.0       13.94s

The clone and the close get cheaper without rmapbt, but the truncate
does not improve at all; if anything it grows with the number of
cloned blocks (~13us per block from 1 to 22 blocks). So the rmap
btree is not where the cost is.

For reference, on the same mkfs-default filesystem:

  - clone-tail (no truncate): clone ~50us, tail write ~12us, close
    ~7us, at every size. A prototype of tar using this scheme turns
    the 3.4x slowdown into a 1.6x speedup on XFS, so a userspace
    workaround exists, but it should not be needed.
  - The clone path is otherwise excellent: at size=90000, plain
    write pays 11.6s of writeback at syncfs time, the clone modes
    ~0.1s.
  - BtrFS clone-trunc for comparison: trunc ~15us and close ~2.3us,
    flat from 1k to 90k.

Unrelated, maybe: the write-trunc mode (not shown in the tables
above) also shows that on XFS an ftruncate on a file with dirty
delalloc data synchronously flushes it (the deferred writeback cost
moves into the truncate, sync time drops to ~0.1s), even when the
truncate does not change the size: at size=65536 the no-op ftruncate
costs 514us on the freshly written file, versus the 7.3us that the
same no-op costs on a clean cloned file in the table above. BtrFS
only pays this for an actual shrink, not for the no-op (19.6us).
This does not apply to the clone case above, where the cloned file
has nothing dirty, which is why the 300-600us look like pure
metadata work.

Environment:
  kernel 7.1.2 vanilla
  All filesystems were freshly created on the same partition of a
  spinning HDD, with mkfs defaults.
  meta-data: isize=512 agcount=4
             sectsz=4096 crc=1 finobt=1, sparse=1, rmapbt=1
             reflink=1 bigtime=1 inobtcount=1 nrext64=1
  mount: rw,noatime,inode64,logbufs=8,logbsize=32k,noquota
  data:  bsize=4096

The benchmark and its driver script are here:

  https://gist.github.com/teknoraver/67b50ce366d7cf430bb7f82a8018acfa

(gcc -O2 -o clone_bench clone_bench.c; the script runs the whole
mode/size matrix in fresh subdirectories of the given target dir.)

Happy to run further tests (tracepoints, perf, other mkfs/mount
options) if useful.

Regards,
-- 
Matteo Croce

per aspera ad upstream

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-07-07 15:09 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-07-06 21:45 ftruncate() after FICLONERANGE costs 300-600us/file, ~10x more than the clone itself Matteo Croce
2026-07-07  5:56 ` Darrick J. Wong
2026-07-07  6:20   ` Christoph Hellwig
2026-07-07 15:09     ` Matteo Croce
2026-07-07 15:07   ` Matteo Croce

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®