* [PATCH 0/3] dm zoned: fix two idle reclaim problems
@ 2026-10-10 10:51 illofspeed
2026-10-10 10:51 ` [PATCH 1/3] dm zoned: back off when reclaim finds no destination zone illofspeed
` (2 more replies)
0 siblings, 3 replies; 4+ messages in thread
From: illofspeed @ 2026-10-10 10:51 UTC (permalink / raw)
To: dm-devel; +Cc: agk, snitzer, mpatocka, bmarzins, hare, dlemoal, linux-kernel
These patches fix two problems in dm-zoned's reclaim worker, found while
running md RAID5 on three dm-zoned targets on 27 TB host-managed SMR drives
(WD Ultrastar DC HC680), each target with a regular cache device.
1. On a target with cache zones and no free sequential zone, an idle
target keeps copying random zones into other random zones, forever
(patch 2; patch 1 makes the resulting -ENOSPC back off instead of
spinning). After a full md rebuild every chunk is mapped, so this
happens right after a rebuild: the idle member read and wrote about
72 MB/s nonstop for hours.
2. After a reclaim pass that ends while the target is busy, the idle poll
is never re-armed, so zones filled under load are not reclaimed when the
target goes idle (patch 3). Users notice it as "the buffer never drains";
the workaround has been a timer sending "dmsetup message <dev> 0 reclaim".
Testing: the three patches were built as an out-of-tree dm-zoned module for
7.2.6 (the changed code is identical in current master) and tested
- on emulated host-managed drives (tcmu-runner ZBC handler) in a VM: the
idle copy loop reproduced with the stock module (340 MB/s on an idle
target, zone counters unchanged), 0 MB/s with the patches;
- on the three real drives since 2026-10-05: crash test, three 1.5 TiB
write benchmarks, and eight conversions between single-device and
cache-device layouts; emptied zones were reclaimed within minutes of the
targets going idle without the reclaim timer.
- with a reproducer that needs no special hardware (scsi_debug zbc=managed,
64 MiB zones, plus a 1 GiB loop cache device): stock module ~600 MB/s of
read+write on the idle target with unchanged zone counters, patched
module 0 MB/s:
https://github.com/illofspeed/synology-host-managed-smr/tree/main/repro
Not done: a test on current master itself (only 7.2.6), and a rebuild onto a
cached member with the patches (the first problem was seen after a rebuild
with the stock module).
Related, not addressed here: the double-free in dmz_load_sb() reported on
dm-devel on 2026-05-30.
The problems were found and the patches written with the help of an AI
coding assistant (hence Assisted-by); the analysis, the tests on real
hardware and this description were reviewed by me.
illofspeed (3):
dm zoned: back off when reclaim finds no destination zone
dm zoned: do not reclaim a random zone into another random zone
dm zoned: keep polling for idle reclaim after a reclaim pass
drivers/md/dm-zoned-reclaim.c | 28 ++++++++++++++++++++++++++--
1 file changed, 26 insertions(+), 2 deletions(-)
--
2.47.3
^ permalink raw reply [flat|nested] 4+ messages in thread
* [PATCH 1/3] dm zoned: back off when reclaim finds no destination zone
2026-10-10 10:51 [PATCH 0/3] dm zoned: fix two idle reclaim problems illofspeed
@ 2026-10-10 10:51 ` illofspeed
2026-10-10 10:51 ` [PATCH 2/3] dm zoned: do not reclaim a random zone into another random zone illofspeed
2026-10-10 10:51 ` [PATCH 3/3] dm zoned: keep polling for idle reclaim after a reclaim pass illofspeed
2 siblings, 0 replies; 4+ messages in thread
From: illofspeed @ 2026-10-10 10:51 UTC (permalink / raw)
To: dm-devel; +Cc: agk, snitzer, mpatocka, bmarzins, hare, dlemoal, linux-kernel
When dmz_do_reclaim() fails with -ENOSPC because no free zone is
available to copy the data into, dmz_reclaim_work() still ends with
dmz_schedule_reclaim(). If the target is idle, dmz_should_reclaim()
keeps returning true, the work is queued again with no delay and the
worker spins on the same failing pass.
Re-arm the idle poll instead, so that reclaim is retried after
DMZ_IDLE_PERIOD. The next patch makes -ENOSPC the expected result when
an idle target has random zones but no free sequential zone.
Tested on a 7.2.6 kernel as an out-of-tree dm-zoned module, together
with the next two patches, on emulated host-managed drives (tcmu-runner
ZBC handler) and on three 27 TB host-managed drives used as md RAID5
members.
Assisted-by: LLM
Signed-off-by: illofspeed <volvo.mail@gmail.com>
---
drivers/md/dm-zoned-reclaim.c | 8 ++++++++
1 file changed, 8 insertions(+)
diff --git a/drivers/md/dm-zoned-reclaim.c b/drivers/md/dm-zoned-reclaim.c
index c041413c729..315f0da99f3 100644
--- a/drivers/md/dm-zoned-reclaim.c
+++ b/drivers/md/dm-zoned-reclaim.c
@@ -540,6 +540,14 @@ static void dmz_reclaim_work(struct work_struct *work)
if (ret && ret != -EINTR) {
if (!dmz_check_dev(zmd))
return;
+ if (ret == -ENOSPC) {
+ /*
+ * No zone to copy into: nothing can progress now.
+ * Rescheduling immediately would spin; poll instead.
+ */
+ mod_delayed_work(zrc->wq, &zrc->work, DMZ_IDLE_PERIOD);
+ return;
+ }
}
dmz_schedule_reclaim(zrc);
--
2.43.0
^ permalink raw reply [flat|nested] 4+ messages in thread
* [PATCH 2/3] dm zoned: do not reclaim a random zone into another random zone
2026-10-10 10:51 [PATCH 0/3] dm zoned: fix two idle reclaim problems illofspeed
2026-10-10 10:51 ` [PATCH 1/3] dm zoned: back off when reclaim finds no destination zone illofspeed
@ 2026-10-10 10:51 ` illofspeed
2026-10-10 10:51 ` [PATCH 3/3] dm zoned: keep polling for idle reclaim after a reclaim pass illofspeed
2 siblings, 0 replies; 4+ messages in thread
From: illofspeed @ 2026-10-10 10:51 UTC (permalink / raw)
To: dm-devel; +Cc: agk, snitzer, mpatocka, bmarzins, hare, dlemoal, linux-kernel
dmz_reclaim_rnd_data() falls back to a random destination zone when no
free sequential zone is available and the target has cache zones. When
the target is idle and the cache zone list is empty,
dmz_get_rnd_zone_for_reclaim() selects random zones of the zoned device
as well, so a random zone is then copied into another random zone. That
frees nothing: the next pass picks the next random zone and copies it
again, for as long as the target stays idle.
Seen on a two-device target (256 GiB regular cache device plus a 27 TB
host-managed drive) used as an md RAID5 member: after a full rebuild
every chunk was mapped and no sequential zone was free, and the idle
drive read and wrote about 72 MB/s each without pause for hours while
the free zone counters in "dmsetup status" did not change. Reproduced on
emulated host-managed drives (tcmu-runner ZBC handler): 340 MB/s of
copying on an idle target, 0 MB/s with this patch.
Only allow the random-zone fallback when the zone being reclaimed is a
cache zone: moving it to the zoned device still frees cache space.
Fixes: c5c788595292 ("dm zoned: start reclaim with sequential zones")
Assisted-by: LLM
Signed-off-by: illofspeed <volvo.mail@gmail.com>
---
drivers/md/dm-zoned-reclaim.c | 9 ++++++++-
1 file changed, 8 insertions(+), 1 deletion(-)
diff --git a/drivers/md/dm-zoned-reclaim.c b/drivers/md/dm-zoned-reclaim.c
index 315f0da99f3..2bcc163d209 100644
--- a/drivers/md/dm-zoned-reclaim.c
+++ b/drivers/md/dm-zoned-reclaim.c
@@ -288,7 +288,14 @@ static int dmz_reclaim_rnd_data(struct dmz_reclaim *zrc, struct dm_zone *dzone)
again:
szone = dmz_alloc_zone(zmd, zrc->dev_idx,
alloc_flags | DMZ_ALLOC_RECLAIM);
- if (!szone && alloc_flags == DMZ_ALLOC_SEQ && dmz_nr_cache_zones(zmd)) {
+ /*
+ * Without a free sequential zone, moving a cache zone into a random
+ * zone of the zoned device still frees cache space. Moving a random
+ * zone into another random zone gains nothing, and when the target
+ * is idle it repeats forever. Only fall back for cache zones.
+ */
+ if (!szone && alloc_flags == DMZ_ALLOC_SEQ && dmz_nr_cache_zones(zmd) &&
+ dmz_is_cache(dzone)) {
alloc_flags = DMZ_ALLOC_RND;
goto again;
}
--
2.43.0
^ permalink raw reply [flat|nested] 4+ messages in thread
* [PATCH 3/3] dm zoned: keep polling for idle reclaim after a reclaim pass
2026-10-10 10:51 [PATCH 0/3] dm zoned: fix two idle reclaim problems illofspeed
2026-10-10 10:51 ` [PATCH 1/3] dm zoned: back off when reclaim finds no destination zone illofspeed
2026-10-10 10:51 ` [PATCH 2/3] dm zoned: do not reclaim a random zone into another random zone illofspeed
@ 2026-10-10 10:51 ` illofspeed
2 siblings, 0 replies; 4+ messages in thread
From: illofspeed @ 2026-10-10 10:51 UTC (permalink / raw)
To: dm-devel; +Cc: agk, snitzer, mpatocka, bmarzins, hare, dlemoal, linux-kernel
dmz_reclaim_work() re-arms its DMZ_IDLE_PERIOD poll only when
dmz_should_reclaim() returns false at the start of a pass. After a pass
it calls dmz_schedule_reclaim(), which queues the work only if reclaim
is needed right now. When a pass ends while the target is still busy and
has enough free zones, nothing is queued at all. Once the target goes
idle, nothing re-evaluates dmz_should_reclaim() until new I/O arrives or
user space sends a "reclaim" message, so the random or cache zones that
filled up under load stay full.
Seen on md RAID5 members on 27 TB host-managed drives: after heavy
writes the random zones stayed full for hours on idle targets; the
workaround was a timer sending "dmsetup message <dev> 0 reclaim". With
this patch the emptied and filled zones were reclaimed within minutes of
the targets going idle, without the timer.
After a pass, requeue the work immediately if reclaim is still needed,
otherwise re-arm the idle poll.
Fixes: 3b1a94c88b79 ("dm zoned: drive-managed zoned block device target")
Assisted-by: LLM
Signed-off-by: illofspeed <volvo.mail@gmail.com>
---
drivers/md/dm-zoned-reclaim.c | 11 ++++++++++-
1 file changed, 10 insertions(+), 1 deletion(-)
diff --git a/drivers/md/dm-zoned-reclaim.c b/drivers/md/dm-zoned-reclaim.c
index 2bcc163d209..7feb2628bd0 100644
--- a/drivers/md/dm-zoned-reclaim.c
+++ b/drivers/md/dm-zoned-reclaim.c
@@ -557,7 +557,16 @@ static void dmz_reclaim_work(struct work_struct *work)
}
}
- dmz_schedule_reclaim(zrc);
+ /*
+ * Keep the worker alive: when a pass ends while the target is busy
+ * with enough free zones, dmz_schedule_reclaim() queues nothing and
+ * the idle poll is never re-armed, so zones filled under load are not
+ * reclaimed once the target goes idle.
+ */
+ if (dmz_should_reclaim(zrc, dmz_reclaim_percentage(zrc)))
+ mod_delayed_work(zrc->wq, &zrc->work, 0);
+ else
+ mod_delayed_work(zrc->wq, &zrc->work, DMZ_IDLE_PERIOD);
}
/*
--
2.43.0
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-10-10 10:51 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-10 10:51 [PATCH 0/3] dm zoned: fix two idle reclaim problems illofspeed
2026-10-10 10:51 ` [PATCH 1/3] dm zoned: back off when reclaim finds no destination zone illofspeed
2026-10-10 10:51 ` [PATCH 2/3] dm zoned: do not reclaim a random zone into another random zone illofspeed
2026-10-10 10:51 ` [PATCH 3/3] dm zoned: keep polling for idle reclaim after a reclaim pass illofspeed
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®