* [BUG] btrfs: degraded RAID6 preallocated writes return EIO from fdatasync
@ 2026-09-29 10:34 Bartosz Fenski
2026-09-29 10:41 ` Qu Wenruo
0 siblings, 1 reply; 2+ messages in thread
From: Bartosz Fenski @ 2026-09-29 10:34 UTC (permalink / raw)
To: David Sterba; +Cc: linux-btrfs, linux-kernel, Qu Wenruo, Chris Mason
Hello,
A fresh four-device Btrfs RAID6 filesystem returns EIO from fdatasync()
when writing to a file preallocated by fio after one device is removed.
The filesystem is clean before the device is removed, one missing device
is within RAID6 tolerance, and all surviving-device error counters remain
zero.
I reproduced this twice on unmodified upstream Linux v7.3-rc5. Both
runs failed after exactly the same number of writes and at the same
fdatasync offset. Disabling fio's preallocation makes the same workload
complete successfully.
Test environment
----------------
Kernel:
Linux v7.3-rc5
commit 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e
arm64 defconfig with CONFIG_BTRFS_FS=y and CONFIG_BLK_DEV_LOOP=y
kernel taint before and after test: 0
Userspace:
btrfs-progs v6.17.1
fio 3.41
Ubuntu 26.04 userspace in an arm64 virtual machine
Filesystem:
four 16 GiB loop devices
data profile: RAID6
metadata profile: RAID1C3
checksum: crc32c
sector and page size: 4096 bytes
Reproducer
----------
The complete reproducer is available here:
https://github.com/fenio/modern-fs-benchmark/blob/main/scripts/repro-btrfs-raid6-degraded-write.sh
The failing invocation is:
sudo env \
PRELOAD_SIZE=0 \
FALLOCATE_MODE=default \
SYNC_MODE=fdatasync16 \
MISSING_INDEX=1 \
RUNTIME=10 \
./scripts/repro-btrfs-raid6-degraded-write.sh
It performs these operations:
1. Creates four sparse 16 GiB files and attaches them to loop devices.
2. Creates Btrfs with:
mkfs.btrfs -f -d raid6 -m raid1c3 <four loop devices>
3. Mounts the filesystem and creates a subvolume.
4. Unmounts it and detaches device 2.
5. Mounts it with degraded,noatime.
6. Verifies that device 2 is reported as MISSING.
7. Runs fio random 4 KiB writes to a 1 GiB file with:
--rw=randwrite
--bs=4k
--size=1G
--runtime=10
--time_based
--randrepeat=1
--fdatasync=16
Actual result
-------------
Both v7.3-rc5 runs failed with:
fio: io_u error on file degraded-write.0.0:
Input/output error: datasync offset=368435200, buflen=0
The fio JSON output was identical in both runs:
error: 5 (EIO)
bytes written: 5439488
completed writes: 1328
successful syncs: 82
After the failure, Btrfs still reported exactly one missing device:
devid 2 size 0 used 0 path <missing disk> MISSING
All Btrfs device error counters, including those for the missing device,
were zero:
write_io_errs 0
read_io_errs 0
flush_io_errs 0
corruption_errs 0
generation_errs 0
There was no kernel warning, Oops, transaction abort, checksum error, or
physical I/O error in dmesg.
Control
-------
Running the same test with fio preallocation disabled succeeds:
sudo env \
PRELOAD_SIZE=0 \
FALLOCATE_MODE=none \
SYNC_MODE=fdatasync16 \
MISSING_INDEX=1 \
RUNTIME=10 \
./scripts/repro-btrfs-raid6-degraded-write.sh
The control completed the full runtime with:
error: 0
bytes written: 409731072
completed writes: 100032
successful syncs: 6252
The filesystem still had the same missing member and all device error
counters remained zero.
Independent x86-64 reproductions
---------------------------------
The same result was independently reproduced on GitHub-hosted Ubuntu
26.04 x86-64 runners using Linux 7.0.0-1012-azure:
https://github.com/fenio/modern-fs-benchmark/actions/runs/36482348497
https://github.com/fenio/modern-fs-benchmark/actions/runs/36482737600
In both runs, the fresh-filesystem/default-preallocation variant failed
at exactly the same point as v7.3-rc5:
datasync offset: 368435200
bytes written: 5439488
completed writes: 1328
successful syncs: 82
The corresponding fallocate=none control succeeded in both GitHub runs.
Thus the failure has reproduced on both arm64 and x86-64, with two
different kernel builds and virtualization environments.
Possible cause
--------------
This appears to involve checksum verification of reconstructed sectors
which do not themselves have a checksum.
fill_data_csums() allocates zeroed csum_buf and csum_bitmap storage for
the whole data stripe. If any sector has a checksum, those buffers are
retained even though some bits in csum_bitmap can remain clear.
verify_one_sector() currently checks only whether csum_bitmap and
csum_buf exist:
if (!rbio->csum_bitmap || !rbio->csum_buf)
return 0;
It does not check whether the bit for the specific reconstructed sector
is set before comparing the reconstructed data against its entry in the
zeroed checksum buffer.
fio's default preallocation creates PREALLOC extents containing a mixture
of written sectors with checksums and unwritten sectors without
checksums. filefrag after the failure confirms that written blocks are
interleaved with unwritten extents.
If a checksum-less sector is reconstructed during degraded RAID6 RMW,
verify_one_sector() appears able to compare it against a zero-filled
checksum slot and return -EIO.
A guard similar to the following before accessing csum_buf may be needed:
if (!test_bit(rbio_sector_index(rbio, stripe_nr, sector_nr),
rbio->csum_bitmap))
return 0;
The relevant checksum verification was introduced by commit:
7a3150723061 ("btrfs: raid56: do data csum verification during RMW
cycle")
Current Btrfs for-next still appears to lack the per-sector bitmap check.
I searched the linux-btrfs archive and did not find an existing report
with this combination of a clean filesystem, one missing RAID6 member,
preallocated extents, and sync-time EIO.
I can test a proposed patch on the same v7.3-rc5 VM.
Regards,
Bartosz Fenski
^ permalink raw reply [flat|nested] 2+ messages in thread
* Re: [BUG] btrfs: degraded RAID6 preallocated writes return EIO from fdatasync
2026-09-29 10:34 [BUG] btrfs: degraded RAID6 preallocated writes return EIO from fdatasync Bartosz Fenski
@ 2026-09-29 10:41 ` Qu Wenruo
0 siblings, 0 replies; 2+ messages in thread
From: Qu Wenruo @ 2026-09-29 10:41 UTC (permalink / raw)
To: Bartosz Fenski, David Sterba; +Cc: linux-btrfs, linux-kernel, Chris Mason
在 2026/9/29 20:04, Bartosz Fenski 写道:
> Hello,
>
> A fresh four-device Btrfs RAID6 filesystem returns EIO from fdatasync()
> when writing to a file preallocated by fio after one device is removed.
>
> The filesystem is clean before the device is removed, one missing device
> is within RAID6 tolerance, and all surviving-device error counters remain
> zero.
>
> I reproduced this twice on unmodified upstream Linux v7.3-rc5. Both
> runs failed after exactly the same number of writes and at the same
> fdatasync offset. Disabling fio's preallocation makes the same workload
> complete successfully.
I'll take a look.
[...]
> A guard similar to the following before accessing csum_buf may be needed:
>
> if (!test_bit(rbio_sector_index(rbio, stripe_nr, sector_nr),
> rbio->csum_bitmap))
> return 0;
It's already there, check verify_bio_data_sectors().
Anyway I'll try to reproduce and debug it here.
Thanks for the report,
Qu
>
> The relevant checksum verification was introduced by commit:
>
> 7a3150723061 ("btrfs: raid56: do data csum verification during RMW
> cycle")
>
> Current Btrfs for-next still appears to lack the per-sector bitmap check.
>
> I searched the linux-btrfs archive and did not find an existing report
> with this combination of a clean filesystem, one missing RAID6 member,
> preallocated extents, and sync-time EIO.
>
> I can test a proposed patch on the same v7.3-rc5 VM.
>
> Regards,
> Bartosz Fenski
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-09-29 10:41 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-29 10:34 [BUG] btrfs: degraded RAID6 preallocated writes return EIO from fdatasync Bartosz Fenski
2026-09-29 10:41 ` Qu Wenruo
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®