From: Chao Yu <chao@kernel.org>
To: jaegeuk@kernel.org
Cc: linux-f2fs-devel@lists.sourceforge.net,
linux-kernel@vger.kernel.org, Chao Yu <chao@kernel.org>,
stable@vger.kernel.org
Subject: [PATCH] f2fs: zone: fix to avoid bio split on sequential zone
Date: Wed, 7 Oct 2026 10:18:12 +0000 [thread overview]
Message-ID: <20261007101812.2348510-1-chao@kernel.org> (raw)
From: Chao Yu <chao@kernel.org>
In the block layer, bio_split_to_limits() splits a bio if it exceeds
any underlying queue limits:
1. queue_max_sectors: Splits bios larger than the transfer limit. This
limit is set by the device driver/DMA controller (max_hw_sectors) and
device-mapper targets. When flushing dirty pages (governed by VFS
writeback_chunk_size() and f2fs nr_pages_to_skip()), f2fs merges
contiguous pages into bios up to 512KB (1024 sectors) or more without
checking queue_max_sectors, which will be split if the device or DM
target enforces a smaller limit.
2. queue_max_segments: Splits bios exceeding scatter-gather segment
limits. f2fs allocates bios with up to 256 pages, exceeding the
128-segment limit of DM targets (dm-default-key/dm-crypt) if pages
are discontiguous.
3. chunk_sectors: Splits bios crossing chunk/zone boundaries. In f2fs,
section_size strictly matches zone_size, and block allocation never
spans across section boundaries, so bios never cross chunk_sectors.
Crucially, when bio_split_to_limits() splits a bio, bio_submit_split()
submits the tail half via submit_bio_noacct() before returning the head
half to the caller. On a sequential zoned device, this causes the tail
bio to arrive at the zone write plug before the head:
dm-3: zone 7: prepare_bio wp mismatch: bio sector 30720 offset 2048
!= wp_offset 248, len 248
The write plug rejects the out-of-order tail with BLK_STS_IOERR, causing
f2fs_stop_checkpoint() and remounting the filesystem read-only:
2026-09-30T22:09:23.725678+08:00 localhost kernel: CPU: 9 UID: 0 PID: 243 Comm: kworker/9:1H Not tainted 6.12.81-gb3ef518fb26d #5
2026-09-30T22:09:23.725750+08:00 localhost kernel: Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS 1.17.0-debian-1.17.0-1 04/01/2014
2026-09-30T22:09:23.725753+08:00 localhost kernel: Workqueue: dm-3_zwplugs blk_zone_wplug_bio_work
2026-09-30T22:09:23.725756+08:00 localhost kernel: Call Trace:
2026-09-30T22:09:23.725758+08:00 localhost kernel: <TASK>
2026-09-30T22:09:23.725758+08:00 localhost kernel: dump_stack_lvl+0x4d/0x70
2026-09-30T22:09:23.725759+08:00 localhost kernel: f2fs_stop_checkpoint+0x1df/0x5d0
2026-09-30T22:09:23.725759+08:00 localhost kernel: ? blk_zone_wplug_bio_work+0x4eb/0x740
2026-09-30T22:09:23.725760+08:00 localhost kernel: f2fs_cache_write_end_io+0x545/0x6e0
2026-09-30T22:09:23.725761+08:00 localhost kernel: blk_zone_wplug_bio_work+0x4eb/0x740
2026-09-30T22:09:23.725762+08:00 localhost kernel: process_one_work+0x5c6/0xde0
2026-09-30T22:09:23.725764+08:00 localhost kernel: ? assign_work+0x124/0x4e0
2026-09-30T22:09:23.725765+08:00 localhost kernel: worker_thread+0x403/0xb70
2026-09-30T22:09:23.725765+08:00 localhost kernel: ? __kthread_parkme+0x82/0x140
2026-09-30T22:09:23.725766+08:00 localhost kernel: ? __pfx_worker_thread+0x10/0x10
2026-09-30T22:09:23.725766+08:00 localhost kernel: kthread+0x235/0x2e0
2026-09-30T22:09:23.725767+08:00 localhost kernel: ? recalc_sigpending+0x128/0x1d0
2026-09-30T22:09:23.725767+08:00 localhost kernel: ? __pfx_kthread+0x10/0x10
2026-09-30T22:09:23.725769+08:00 localhost kernel: ret_from_fork+0x2f/0x70
2026-09-30T22:09:23.725769+08:00 localhost kernel: ? __pfx_kthread+0x10/0x10
2026-09-30T22:09:23.725770+08:00 localhost kernel: ret_from_fork_asm+0x1a/0x30
2026-09-30T22:09:23.725771+08:00 localhost kernel: </TASK>
2026-09-30T22:09:23.725771+08:00 localhost kernel: F2FS-fs (dm-0): Stopped filesystem due to reason: 3
To prevent write bios destined for sequential zones from being split:
1. In __bio_alloc(), bound npages to bdev_max_segments(bdev) for writes to
sequential zone areas.
2. In page_is_mergeable(), check queue_max_segments() and queue_max_sectors()
for sequential zone areas, so f2fs will not merge beyond the queue limits.
Fixes: 52763a4b7a21 ("f2fs: detect host-managed SMR by feature flag")
Fixes: 664ba972df9b ("f2fs: use BIO_MAX_PAGES for bio allocation")
Cc: stable@vger.kernel.org
Signed-off-by: Chao Yu <chao@kernel.org>
---
fs/f2fs/data.c | 18 +++++++++++++++++-
1 file changed, 17 insertions(+), 1 deletion(-)
diff --git a/fs/f2fs/data.c b/fs/f2fs/data.c
index 4bfa301a5df6..15c6af8515f3 100644
--- a/fs/f2fs/data.c
+++ b/fs/f2fs/data.c
@@ -536,6 +536,10 @@ static struct bio *__bio_alloc(struct f2fs_io_info *fio, int npages)
struct bio *bio;
bdev = f2fs_target_device(sbi, fio->new_blkaddr, §or);
+ /* avoid reversed IO if bio is split due to exceeding max_segments */
+ if (npages > 1 && !is_read_io(fio->op) &&
+ f2fs_is_sequential_zone_area(sbi, fio->new_blkaddr))
+ npages = min_t(int, npages, bdev_max_segments(bdev));
bio = bio_alloc_bioset(bdev, npages,
fio->op | fio->op_flags | f2fs_io_flags(fio),
GFP_NOIO, &f2fs_bioset);
@@ -892,7 +896,19 @@ static bool page_is_mergeable(struct f2fs_sb_info *sbi, struct bio *bio,
return false;
if (last_blkaddr + 1 != cur_blkaddr)
return false;
- return bio->bi_bdev == f2fs_target_device(sbi, cur_blkaddr, NULL);
+ if (bio->bi_bdev != f2fs_target_device(sbi, cur_blkaddr, NULL))
+ return false;
+ /* avoid reversed IO if bio is split due to exceeding max_segments/sectors */
+ if (f2fs_is_sequential_zone_area(sbi, cur_blkaddr)) {
+ struct request_queue *q = bdev_get_queue(bio->bi_bdev);
+
+ if (bio->bi_vcnt >= queue_max_segments(q))
+ return false;
+ if (bio_sectors(bio) + (F2FS_BLKSIZE(sbi) >> SECTOR_SHIFT) >
+ queue_max_sectors(q))
+ return false;
+ }
+ return true;
}
static bool io_type_is_mergeable(struct f2fs_bio_info *io,
--
2.49.0
reply other threads:[~2026-10-07 10:18 UTC|newest]
Thread overview: [no followups] expand[flat|nested] mbox.gz Atom feed
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261007101812.2348510-1-chao@kernel.org \
--to=chao@kernel.org \
--cc=jaegeuk@kernel.org \
--cc=linux-f2fs-devel@lists.sourceforge.net \
--cc=linux-kernel@vger.kernel.org \
--cc=stable@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®