mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v2 0/3] dm zoned: fix two idle reclaim problems
@ 2026-10-10 16:16 illofspeed
  2026-10-10 16:16 ` [PATCH v2 1/3] dm zoned: back off when reclaim finds no destination zone illofspeed
                   ` (4 more replies)
  0 siblings, 5 replies; 14+ messages in thread
From: illofspeed @ 2026-10-10 16:16 UTC (permalink / raw)
  To: dm-devel; +Cc: agk, snitzer, mpatocka, bmarzins, hare, dlemoal, linux-kernel

v2: no code changes. The only change is my author address: v1 went out on
2026-10-10 from my previous address. Please apply this version instead of v1.

These patches fix two problems in dm-zoned's reclaim worker, found while
running md RAID5 on three dm-zoned targets on 27 TB host-managed SMR drives
(WD Ultrastar DC HC680), each target with a regular cache device.

1. On a target with cache zones and no free sequential zone, an idle
   target keeps copying random zones into other random zones, forever
   (patch 2; patch 1 makes the resulting -ENOSPC back off instead of
   spinning). After a full md rebuild every chunk is mapped, so this
   happens right after a rebuild: the idle member read and wrote about
   72 MB/s nonstop for hours.

2. After a reclaim pass that ends while the target is busy, the idle poll
   is never re-armed, so zones filled under load are not reclaimed when the
   target goes idle (patch 3). Users notice it as "the buffer never drains";
   the workaround has been a timer sending "dmsetup message <dev> 0 reclaim".

Testing: the three patches were built as an out-of-tree dm-zoned module for
7.2.6 (the changed code is identical in current master) and tested
- on emulated host-managed drives (tcmu-runner ZBC handler) in a VM: the
  idle copy loop reproduced with the stock module (340 MB/s on an idle
  target, zone counters unchanged), 0 MB/s with the patches;
- on the three real drives since 2026-10-05: crash test, three 1.5 TiB
  write benchmarks, and eight conversions between single-device and
  cache-device layouts; emptied zones were reclaimed within minutes of the
  targets going idle without the reclaim timer.
- with a reproducer that needs no special hardware (scsi_debug zbc=managed,
  64 MiB zones, plus a 1 GiB loop cache device): stock module ~600 MB/s of
  read+write on the idle target with unchanged zone counters, patched
  module 0 MB/s:
  https://github.com/illofspeed/synology-host-managed-smr/tree/main/repro
Not done: a test on current master itself (only 7.2.6), and a rebuild onto a
cached member with the patches (the first problem was seen after a rebuild
with the stock module).

Related, not addressed here: the double-free in dmz_load_sb() reported on
dm-devel on 2026-05-30.

The problems were found and the patches written with the help of an AI
coding assistant (hence Assisted-by); the analysis, the tests on real
hardware and this description were reviewed by me.

illofspeed (3):
  dm zoned: back off when reclaim finds no destination zone
  dm zoned: do not reclaim a random zone into another random zone
  dm zoned: keep polling for idle reclaim after a reclaim pass

 drivers/md/dm-zoned-reclaim.c | 28 ++++++++++++++++++++++++++--
 1 file changed, 26 insertions(+), 2 deletions(-)

-- 
2.47.3


^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH v2 1/3] dm zoned: back off when reclaim finds no destination zone
  2026-10-10 16:16 [PATCH v2 0/3] dm zoned: fix two idle reclaim problems illofspeed
@ 2026-10-10 16:16 ` illofspeed
  2026-10-11 14:24   ` Damien Le Moal
  2026-10-10 16:16 ` [PATCH v2 2/3] dm zoned: do not reclaim a random zone into another random zone illofspeed
                   ` (3 subsequent siblings)
  4 siblings, 1 reply; 14+ messages in thread
From: illofspeed @ 2026-10-10 16:16 UTC (permalink / raw)
  To: dm-devel; +Cc: agk, snitzer, mpatocka, bmarzins, hare, dlemoal, linux-kernel

When dmz_do_reclaim() fails with -ENOSPC because no free zone is
available to copy the data into, dmz_reclaim_work() still ends with
dmz_schedule_reclaim(). If the target is idle, dmz_should_reclaim()
keeps returning true, the work is queued again with no delay and the
worker spins on the same failing pass.

Re-arm the idle poll instead, so that reclaim is retried after
DMZ_IDLE_PERIOD. The next patch makes -ENOSPC the expected result when
an idle target has random zones but no free sequential zone.

Tested on a 7.2.6 kernel as an out-of-tree dm-zoned module, together
with the next two patches, on emulated host-managed drives (tcmu-runner
ZBC handler) and on three 27 TB host-managed drives used as md RAID5
members.

Assisted-by: LLM
Signed-off-by: illofspeed <illofspeed@gmail.com>
---
 drivers/md/dm-zoned-reclaim.c | 8 ++++++++
 1 file changed, 8 insertions(+)

diff --git a/drivers/md/dm-zoned-reclaim.c b/drivers/md/dm-zoned-reclaim.c
index c041413c729..315f0da99f3 100644
--- a/drivers/md/dm-zoned-reclaim.c
+++ b/drivers/md/dm-zoned-reclaim.c
@@ -540,6 +540,14 @@ static void dmz_reclaim_work(struct work_struct *work)
 	if (ret && ret != -EINTR) {
 		if (!dmz_check_dev(zmd))
 			return;
+		if (ret == -ENOSPC) {
+			/*
+			 * No zone to copy into: nothing can progress now.
+			 * Rescheduling immediately would spin; poll instead.
+			 */
+			mod_delayed_work(zrc->wq, &zrc->work, DMZ_IDLE_PERIOD);
+			return;
+		}
 	}
 
 	dmz_schedule_reclaim(zrc);
-- 
2.43.0


^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH v2 2/3] dm zoned: do not reclaim a random zone into another random zone
  2026-10-10 16:16 [PATCH v2 0/3] dm zoned: fix two idle reclaim problems illofspeed
  2026-10-10 16:16 ` [PATCH v2 1/3] dm zoned: back off when reclaim finds no destination zone illofspeed
@ 2026-10-10 16:16 ` illofspeed
  2026-10-11 14:30   ` Damien Le Moal
  2026-10-10 16:16 ` [PATCH v2 3/3] dm zoned: keep polling for idle reclaim after a reclaim pass illofspeed
                   ` (2 subsequent siblings)
  4 siblings, 1 reply; 14+ messages in thread
From: illofspeed @ 2026-10-10 16:16 UTC (permalink / raw)
  To: dm-devel; +Cc: agk, snitzer, mpatocka, bmarzins, hare, dlemoal, linux-kernel

dmz_reclaim_rnd_data() falls back to a random destination zone when no
free sequential zone is available and the target has cache zones. When
the target is idle and the cache zone list is empty,
dmz_get_rnd_zone_for_reclaim() selects random zones of the zoned device
as well, so a random zone is then copied into another random zone. That
frees nothing: the next pass picks the next random zone and copies it
again, for as long as the target stays idle.

Seen on a two-device target (256 GiB regular cache device plus a 27 TB
host-managed drive) used as an md RAID5 member: after a full rebuild
every chunk was mapped and no sequential zone was free, and the idle
drive read and wrote about 72 MB/s each without pause for hours while
the free zone counters in "dmsetup status" did not change. Reproduced on
emulated host-managed drives (tcmu-runner ZBC handler): 340 MB/s of
copying on an idle target, 0 MB/s with this patch.

Only allow the random-zone fallback when the zone being reclaimed is a
cache zone: moving it to the zoned device still frees cache space.

Fixes: c5c788595292 ("dm zoned: start reclaim with sequential zones")
Assisted-by: LLM
Signed-off-by: illofspeed <illofspeed@gmail.com>
---
 drivers/md/dm-zoned-reclaim.c | 9 ++++++++-
 1 file changed, 8 insertions(+), 1 deletion(-)

diff --git a/drivers/md/dm-zoned-reclaim.c b/drivers/md/dm-zoned-reclaim.c
index 315f0da99f3..2bcc163d209 100644
--- a/drivers/md/dm-zoned-reclaim.c
+++ b/drivers/md/dm-zoned-reclaim.c
@@ -288,7 +288,14 @@ static int dmz_reclaim_rnd_data(struct dmz_reclaim *zrc, struct dm_zone *dzone)
 again:
 	szone = dmz_alloc_zone(zmd, zrc->dev_idx,
 			       alloc_flags | DMZ_ALLOC_RECLAIM);
-	if (!szone && alloc_flags == DMZ_ALLOC_SEQ && dmz_nr_cache_zones(zmd)) {
+	/*
+	 * Without a free sequential zone, moving a cache zone into a random
+	 * zone of the zoned device still frees cache space. Moving a random
+	 * zone into another random zone gains nothing, and when the target
+	 * is idle it repeats forever. Only fall back for cache zones.
+	 */
+	if (!szone && alloc_flags == DMZ_ALLOC_SEQ && dmz_nr_cache_zones(zmd) &&
+	    dmz_is_cache(dzone)) {
 		alloc_flags = DMZ_ALLOC_RND;
 		goto again;
 	}
-- 
2.43.0


^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH v2 3/3] dm zoned: keep polling for idle reclaim after a reclaim pass
  2026-10-10 16:16 [PATCH v2 0/3] dm zoned: fix two idle reclaim problems illofspeed
  2026-10-10 16:16 ` [PATCH v2 1/3] dm zoned: back off when reclaim finds no destination zone illofspeed
  2026-10-10 16:16 ` [PATCH v2 2/3] dm zoned: do not reclaim a random zone into another random zone illofspeed
@ 2026-10-10 16:16 ` illofspeed
  2026-10-11 14:31   ` Damien Le Moal
  2026-10-11 14:23 ` [PATCH v2 0/3] dm zoned: fix two idle reclaim problems Damien Le Moal
  2026-10-11 14:33 ` Damien Le Moal
  4 siblings, 1 reply; 14+ messages in thread
From: illofspeed @ 2026-10-10 16:16 UTC (permalink / raw)
  To: dm-devel; +Cc: agk, snitzer, mpatocka, bmarzins, hare, dlemoal, linux-kernel

dmz_reclaim_work() re-arms its DMZ_IDLE_PERIOD poll only when
dmz_should_reclaim() returns false at the start of a pass. After a pass
it calls dmz_schedule_reclaim(), which queues the work only if reclaim
is needed right now. When a pass ends while the target is still busy and
has enough free zones, nothing is queued at all. Once the target goes
idle, nothing re-evaluates dmz_should_reclaim() until new I/O arrives or
user space sends a "reclaim" message, so the random or cache zones that
filled up under load stay full.

Seen on md RAID5 members on 27 TB host-managed drives: after heavy
writes the random zones stayed full for hours on idle targets; the
workaround was a timer sending "dmsetup message <dev> 0 reclaim". With
this patch the emptied and filled zones were reclaimed within minutes of
the targets going idle, without the timer.

After a pass, requeue the work immediately if reclaim is still needed,
otherwise re-arm the idle poll.

Fixes: 3b1a94c88b79 ("dm zoned: drive-managed zoned block device target")
Assisted-by: LLM
Signed-off-by: illofspeed <illofspeed@gmail.com>
---
 drivers/md/dm-zoned-reclaim.c | 11 ++++++++++-
 1 file changed, 10 insertions(+), 1 deletion(-)

diff --git a/drivers/md/dm-zoned-reclaim.c b/drivers/md/dm-zoned-reclaim.c
index 2bcc163d209..7feb2628bd0 100644
--- a/drivers/md/dm-zoned-reclaim.c
+++ b/drivers/md/dm-zoned-reclaim.c
@@ -557,7 +557,16 @@ static void dmz_reclaim_work(struct work_struct *work)
 		}
 	}
 
-	dmz_schedule_reclaim(zrc);
+	/*
+	 * Keep the worker alive: when a pass ends while the target is busy
+	 * with enough free zones, dmz_schedule_reclaim() queues nothing and
+	 * the idle poll is never re-armed, so zones filled under load are not
+	 * reclaimed once the target goes idle.
+	 */
+	if (dmz_should_reclaim(zrc, dmz_reclaim_percentage(zrc)))
+		mod_delayed_work(zrc->wq, &zrc->work, 0);
+	else
+		mod_delayed_work(zrc->wq, &zrc->work, DMZ_IDLE_PERIOD);
 }
 
 /*
-- 
2.43.0


^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH v2 0/3] dm zoned: fix two idle reclaim problems
  2026-10-10 16:16 [PATCH v2 0/3] dm zoned: fix two idle reclaim problems illofspeed
                   ` (2 preceding siblings ...)
  2026-10-10 16:16 ` [PATCH v2 3/3] dm zoned: keep polling for idle reclaim after a reclaim pass illofspeed
@ 2026-10-11 14:23 ` Damien Le Moal
  2026-10-11 14:33 ` Damien Le Moal
  4 siblings, 0 replies; 14+ messages in thread
From: Damien Le Moal @ 2026-10-11 14:23 UTC (permalink / raw)
  To: illofspeed, dm-devel; +Cc: agk, snitzer, mpatocka, bmarzins, hare, linux-kernel

On 2026/10/10 18:16, illofspeed wrote:
> v2: no code changes. The only change is my author address: v1 went out on
> 2026-10-10 from my previous address. Please apply this version instead of v1.
> 
> These patches fix two problems in dm-zoned's reclaim worker, found while
> running md RAID5 on three dm-zoned targets on 27 TB host-managed SMR drives
> (WD Ultrastar DC HC680), each target with a regular cache device.
> 
> 1. On a target with cache zones and no free sequential zone, an idle
>    target keeps copying random zones into other random zones, forever
>    (patch 2; patch 1 makes the resulting -ENOSPC back off instead of
>    spinning). After a full md rebuild every chunk is mapped, so this
>    happens right after a rebuild: the idle member read and wrote about
>    72 MB/s nonstop for hours.
> 
> 2. After a reclaim pass that ends while the target is busy, the idle poll
>    is never re-armed, so zones filled under load are not reclaimed when the
>    target goes idle (patch 3). Users notice it as "the buffer never drains";
>    the workaround has been a timer sending "dmsetup message <dev> 0 reclaim".
> 
> Testing: the three patches were built as an out-of-tree dm-zoned module for
> 7.2.6 (the changed code is identical in current master) and tested
> - on emulated host-managed drives (tcmu-runner ZBC handler) in a VM: the
>   idle copy loop reproduced with the stock module (340 MB/s on an idle
>   target, zone counters unchanged), 0 MB/s with the patches;
> - on the three real drives since 2026-10-05: crash test, three 1.5 TiB
>   write benchmarks, and eight conversions between single-device and
>   cache-device layouts; emptied zones were reclaimed within minutes of the
>   targets going idle without the reclaim timer.
> - with a reproducer that needs no special hardware (scsi_debug zbc=managed,
>   64 MiB zones, plus a 1 GiB loop cache device): stock module ~600 MB/s of
>   read+write on the idle target with unchanged zone counters, patched
>   module 0 MB/s:
>   https://github.com/illofspeed/synology-host-managed-smr/tree/main/repro
> Not done: a test on current master itself (only 7.2.6), and a rebuild onto a
> cached member with the patches (the first problem was seen after a rebuild
> with the stock module).
> 
> Related, not addressed here: the double-free in dmz_load_sb() reported on
> dm-devel on 2026-05-30.
> 
> The problems were found and the patches written with the help of an AI
> coding assistant (hence Assisted-by); the analysis, the tests on real
> hardware and this description were reviewed by me.

What changes from v1?

> 
> illofspeed (3):
>   dm zoned: back off when reclaim finds no destination zone
>   dm zoned: do not reclaim a random zone into another random zone
>   dm zoned: keep polling for idle reclaim after a reclaim pass
> 
>  drivers/md/dm-zoned-reclaim.c | 28 ++++++++++++++++++++++++++--
>  1 file changed, 26 insertions(+), 2 deletions(-)
> 


-- 
Damien Le Moal
Western Digital Research

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH v2 1/3] dm zoned: back off when reclaim finds no destination zone
  2026-10-10 16:16 ` [PATCH v2 1/3] dm zoned: back off when reclaim finds no destination zone illofspeed
@ 2026-10-11 14:24   ` Damien Le Moal
  2026-10-11 16:17     ` illofspeed
  0 siblings, 1 reply; 14+ messages in thread
From: Damien Le Moal @ 2026-10-11 14:24 UTC (permalink / raw)
  To: illofspeed, dm-devel; +Cc: agk, snitzer, mpatocka, bmarzins, hare, linux-kernel

On 2026/10/10 18:16, illofspeed wrote:
> When dmz_do_reclaim() fails with -ENOSPC because no free zone is
> available to copy the data into, dmz_reclaim_work() still ends with
> dmz_schedule_reclaim(). If the target is idle, dmz_should_reclaim()
> keeps returning true, the work is queued again with no delay and the
> worker spins on the same failing pass.
> 
> Re-arm the idle poll instead, so that reclaim is retried after
> DMZ_IDLE_PERIOD. The next patch makes -ENOSPC the expected result when
> an idle target has random zones but no free sequential zone.
> 
> Tested on a 7.2.6 kernel as an out-of-tree dm-zoned module, together
> with the next two patches, on emulated host-managed drives (tcmu-runner
> ZBC handler) and on three 27 TB host-managed drives used as md RAID5
> members.
> 
> Assisted-by: LLM
> Signed-off-by: illofspeed <illofspeed@gmail.com>

Patch looks OK, but is "ilofspeed" your full name?

> ---
>  drivers/md/dm-zoned-reclaim.c | 8 ++++++++
>  1 file changed, 8 insertions(+)
> 
> diff --git a/drivers/md/dm-zoned-reclaim.c b/drivers/md/dm-zoned-reclaim.c
> index c041413c729..315f0da99f3 100644
> --- a/drivers/md/dm-zoned-reclaim.c
> +++ b/drivers/md/dm-zoned-reclaim.c
> @@ -540,6 +540,14 @@ static void dmz_reclaim_work(struct work_struct *work)
>  	if (ret && ret != -EINTR) {
>  		if (!dmz_check_dev(zmd))
>  			return;
> +		if (ret == -ENOSPC) {
> +			/*
> +			 * No zone to copy into: nothing can progress now.
> +			 * Rescheduling immediately would spin; poll instead.
> +			 */
> +			mod_delayed_work(zrc->wq, &zrc->work, DMZ_IDLE_PERIOD);
> +			return;
> +		}
>  	}
>  
>  	dmz_schedule_reclaim(zrc);


-- 
Damien Le Moal
Western Digital Research

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH v2 2/3] dm zoned: do not reclaim a random zone into another random zone
  2026-10-10 16:16 ` [PATCH v2 2/3] dm zoned: do not reclaim a random zone into another random zone illofspeed
@ 2026-10-11 14:30   ` Damien Le Moal
  2026-10-11 16:17     ` illofspeed
  0 siblings, 1 reply; 14+ messages in thread
From: Damien Le Moal @ 2026-10-11 14:30 UTC (permalink / raw)
  To: illofspeed, dm-devel; +Cc: agk, snitzer, mpatocka, bmarzins, hare, linux-kernel

On 2026/10/10 18:16, illofspeed wrote:
> dmz_reclaim_rnd_data() falls back to a random destination zone when no
> free sequential zone is available and the target has cache zones. When
> the target is idle and the cache zone list is empty,
> dmz_get_rnd_zone_for_reclaim() selects random zones of the zoned device
> as well, so a random zone is then copied into another random zone. That
> frees nothing: the next pass picks the next random zone and copies it
> again, for as long as the target stays idle.
> 
> Seen on a two-device target (256 GiB regular cache device plus a 27 TB
> host-managed drive) used as an md RAID5 member: after a full rebuild
> every chunk was mapped and no sequential zone was free, and the idle
> drive read and wrote about 72 MB/s each without pause for hours while
> the free zone counters in "dmsetup status" did not change. Reproduced on
> emulated host-managed drives (tcmu-runner ZBC handler): 340 MB/s of
> copying on an idle target, 0 MB/s with this patch.
> 
> Only allow the random-zone fallback when the zone being reclaimed is a
> cache zone: moving it to the zoned device still frees cache space.
> 
> Fixes: c5c788595292 ("dm zoned: start reclaim with sequential zones")
> Assisted-by: LLM
> Signed-off-by: illofspeed <illofspeed@gmail.com>
> ---
>  drivers/md/dm-zoned-reclaim.c | 9 ++++++++-
>  1 file changed, 8 insertions(+), 1 deletion(-)
> 
> diff --git a/drivers/md/dm-zoned-reclaim.c b/drivers/md/dm-zoned-reclaim.c
> index 315f0da99f3..2bcc163d209 100644
> --- a/drivers/md/dm-zoned-reclaim.c
> +++ b/drivers/md/dm-zoned-reclaim.c
> @@ -288,7 +288,14 @@ static int dmz_reclaim_rnd_data(struct dmz_reclaim *zrc, struct dm_zone *dzone)
>  again:
>  	szone = dmz_alloc_zone(zmd, zrc->dev_idx,
>  			       alloc_flags | DMZ_ALLOC_RECLAIM);
> -	if (!szone && alloc_flags == DMZ_ALLOC_SEQ && dmz_nr_cache_zones(zmd)) {
> +	/*
> +	 * Without a free sequential zone, moving a cache zone into a random
> +	 * zone of the zoned device still frees cache space. Moving a random
> +	 * zone into another random zone gains nothing, and when the target
> +	 * is idle it repeats forever. Only fall back for cache zones.

"fall back" to what? This comment is not clear.

> +	 */
> +	if (!szone && alloc_flags == DMZ_ALLOC_SEQ && dmz_nr_cache_zones(zmd) &&

Long line. Please split after "DMZ_ALLOC_SEQ &&".
Also, is the test for dmz_nr_cache_zones() needed? dmz_is_cache(dzone) will
never be true if there are no cache zones.

> +	    dmz_is_cache(dzone)) {
>  		alloc_flags = DMZ_ALLOC_RND;
>  		goto again;
>  	}


-- 
Damien Le Moal
Western Digital Research

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH v2 3/3] dm zoned: keep polling for idle reclaim after a reclaim pass
  2026-10-10 16:16 ` [PATCH v2 3/3] dm zoned: keep polling for idle reclaim after a reclaim pass illofspeed
@ 2026-10-11 14:31   ` Damien Le Moal
  2026-10-11 16:17     ` illofspeed
  0 siblings, 1 reply; 14+ messages in thread
From: Damien Le Moal @ 2026-10-11 14:31 UTC (permalink / raw)
  To: illofspeed, dm-devel; +Cc: agk, snitzer, mpatocka, bmarzins, hare, linux-kernel

On 2026/10/10 18:16, illofspeed wrote:
> dmz_reclaim_work() re-arms its DMZ_IDLE_PERIOD poll only when
> dmz_should_reclaim() returns false at the start of a pass. After a pass
> it calls dmz_schedule_reclaim(), which queues the work only if reclaim
> is needed right now. When a pass ends while the target is still busy and
> has enough free zones, nothing is queued at all. Once the target goes
> idle, nothing re-evaluates dmz_should_reclaim() until new I/O arrives or
> user space sends a "reclaim" message, so the random or cache zones that
> filled up under load stay full.
> 
> Seen on md RAID5 members on 27 TB host-managed drives: after heavy
> writes the random zones stayed full for hours on idle targets; the
> workaround was a timer sending "dmsetup message <dev> 0 reclaim". With
> this patch the emptied and filled zones were reclaimed within minutes of
> the targets going idle, without the timer.
> 
> After a pass, requeue the work immediately if reclaim is still needed,
> otherwise re-arm the idle poll.
> 
> Fixes: 3b1a94c88b79 ("dm zoned: drive-managed zoned block device target")
> Assisted-by: LLM
> Signed-off-by: illofspeed <illofspeed@gmail.com>

Patch looks OK to me, but same comment as patch 1 about your SoB name. Please
confirm.

> ---
>  drivers/md/dm-zoned-reclaim.c | 11 ++++++++++-
>  1 file changed, 10 insertions(+), 1 deletion(-)
> 
> diff --git a/drivers/md/dm-zoned-reclaim.c b/drivers/md/dm-zoned-reclaim.c
> index 2bcc163d209..7feb2628bd0 100644
> --- a/drivers/md/dm-zoned-reclaim.c
> +++ b/drivers/md/dm-zoned-reclaim.c
> @@ -557,7 +557,16 @@ static void dmz_reclaim_work(struct work_struct *work)
>  		}
>  	}
>  
> -	dmz_schedule_reclaim(zrc);
> +	/*
> +	 * Keep the worker alive: when a pass ends while the target is busy
> +	 * with enough free zones, dmz_schedule_reclaim() queues nothing and
> +	 * the idle poll is never re-armed, so zones filled under load are not
> +	 * reclaimed once the target goes idle.
> +	 */
> +	if (dmz_should_reclaim(zrc, dmz_reclaim_percentage(zrc)))
> +		mod_delayed_work(zrc->wq, &zrc->work, 0);
> +	else
> +		mod_delayed_work(zrc->wq, &zrc->work, DMZ_IDLE_PERIOD);
>  }
>  
>  /*


-- 
Damien Le Moal
Western Digital Research

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH v2 0/3] dm zoned: fix two idle reclaim problems
  2026-10-10 16:16 [PATCH v2 0/3] dm zoned: fix two idle reclaim problems illofspeed
                   ` (3 preceding siblings ...)
  2026-10-11 14:23 ` [PATCH v2 0/3] dm zoned: fix two idle reclaim problems Damien Le Moal
@ 2026-10-11 14:33 ` Damien Le Moal
  4 siblings, 0 replies; 14+ messages in thread
From: Damien Le Moal @ 2026-10-11 14:33 UTC (permalink / raw)
  To: illofspeed, dm-devel; +Cc: agk, snitzer, mpatocka, bmarzins, hare, linux-kernel

On 2026/10/10 18:16, illofspeed wrote:
> v2: no code changes. The only change is my author address: v1 went out on
> 2026-10-10 from my previous address. Please apply this version instead of v1.

OK. Missed this. So this answer my question about the changes.

> 
> These patches fix two problems in dm-zoned's reclaim worker, found while
> running md RAID5 on three dm-zoned targets on 27 TB host-managed SMR drives
> (WD Ultrastar DC HC680), each target with a regular cache device.
> 
> 1. On a target with cache zones and no free sequential zone, an idle
>    target keeps copying random zones into other random zones, forever
>    (patch 2; patch 1 makes the resulting -ENOSPC back off instead of
>    spinning). After a full md rebuild every chunk is mapped, so this
>    happens right after a rebuild: the idle member read and wrote about
>    72 MB/s nonstop for hours.
> 
> 2. After a reclaim pass that ends while the target is busy, the idle poll
>    is never re-armed, so zones filled under load are not reclaimed when the
>    target goes idle (patch 3). Users notice it as "the buffer never drains";
>    the workaround has been a timer sending "dmsetup message <dev> 0 reclaim".
> 
> Testing: the three patches were built as an out-of-tree dm-zoned module for
> 7.2.6 (the changed code is identical in current master) and tested
> - on emulated host-managed drives (tcmu-runner ZBC handler) in a VM: the
>   idle copy loop reproduced with the stock module (340 MB/s on an idle
>   target, zone counters unchanged), 0 MB/s with the patches;
> - on the three real drives since 2026-10-05: crash test, three 1.5 TiB
>   write benchmarks, and eight conversions between single-device and
>   cache-device layouts; emptied zones were reclaimed within minutes of the
>   targets going idle without the reclaim timer.
> - with a reproducer that needs no special hardware (scsi_debug zbc=managed,
>   64 MiB zones, plus a 1 GiB loop cache device): stock module ~600 MB/s of
>   read+write on the idle target with unchanged zone counters, patched
>   module 0 MB/s:
>   https://github.com/illofspeed/synology-host-managed-smr/tree/main/repro
> Not done: a test on current master itself (only 7.2.6), and a rebuild onto a
> cached member with the patches (the first problem was seen after a rebuild
> with the stock module).
> 
> Related, not addressed here: the double-free in dmz_load_sb() reported on
> dm-devel on 2026-05-30.
> 
> The problems were found and the patches written with the help of an AI
> coding assistant (hence Assisted-by); the analysis, the tests on real
> hardware and this description were reviewed by me.
> 
> illofspeed (3):
>   dm zoned: back off when reclaim finds no destination zone
>   dm zoned: do not reclaim a random zone into another random zone
>   dm zoned: keep polling for idle reclaim after a reclaim pass
> 
>  drivers/md/dm-zoned-reclaim.c | 28 ++++++++++++++++++++++++++--
>  1 file changed, 26 insertions(+), 2 deletions(-)
> 


-- 
Damien Le Moal
Western Digital Research

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH v2 1/3] dm zoned: back off when reclaim finds no destination zone
  2026-10-11 14:24   ` Damien Le Moal
@ 2026-10-11 16:17     ` illofspeed
  2026-10-11 16:32       ` Damien Le Moal
  2026-10-11 16:47       ` Damien Le Moal
  0 siblings, 2 replies; 14+ messages in thread
From: illofspeed @ 2026-10-11 16:17 UTC (permalink / raw)
  To: dlemoal; +Cc: dm-devel, agk, snitzer, mpatocka, bmarzins, hare, linux-kernel

Damien Le Moal wrote:
> Patch looks OK, but is "ilofspeed" your full name?

illofspeed is not my legal name, but it is the name I use for all my
public work (GitHub and this series), and I understood
Documentation/process/submitting-patches.rst ("using a known identity")
to allow that. If you need a legal name in the Signed-off-by, please let
me know and I will resend.

(Resent as plain text: my first reply had an HTML part and did not reach
the lists.)

illofspeed

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH v2 2/3] dm zoned: do not reclaim a random zone into another random zone
  2026-10-11 14:30   ` Damien Le Moal
@ 2026-10-11 16:17     ` illofspeed
  0 siblings, 0 replies; 14+ messages in thread
From: illofspeed @ 2026-10-11 16:17 UTC (permalink / raw)
  To: dlemoal; +Cc: dm-devel, agk, snitzer, mpatocka, bmarzins, hare, linux-kernel

Damien Le Moal wrote:
>> +	 * Without a free sequential zone, moving a cache zone into a random
>> +	 * zone of the zoned device still frees cache space. Moving a random
>> +	 * zone into another random zone gains nothing, and when the target
>> +	 * is idle it repeats forever. Only fall back for cache zones.
>
> "fall back" to what? This comment is not clear.

To a random destination zone: the switch from DMZ_ALLOC_SEQ to
DMZ_ALLOC_RND just below. v3 rewords it:

	/*
	 * If no free sequential zone is available, a cache zone can still be
	 * moved into a free random zone of the zoned device, which frees cache
	 * space. Do not do this for a random zone: moving it into another
	 * random zone frees nothing and, on an idle target, repeats forever.
	 */

>> +	if (!szone && alloc_flags == DMZ_ALLOC_SEQ && dmz_nr_cache_zones(zmd) &&
>
> Long line. Please split after "DMZ_ALLOC_SEQ &&".
> Also, is the test for dmz_nr_cache_zones() needed? dmz_is_cache(dzone) will
> never be true if there are no cache zones.

You are right: DMZ_CACHE is only set on zones of the regular device, so
the test is redundant. Without it the condition fits on one line:

	if (!szone && alloc_flags == DMZ_ALLOC_SEQ && dmz_is_cache(dzone)) {

With this change the reproducer from the cover letter still shows 0 MB/s
on the idle target. I will send v3 once the Signed-off-by question on
1/3 is settled; patches 1 and 3 stay unchanged.

illofspeed

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH v2 3/3] dm zoned: keep polling for idle reclaim after a reclaim pass
  2026-10-11 14:31   ` Damien Le Moal
@ 2026-10-11 16:17     ` illofspeed
  0 siblings, 0 replies; 14+ messages in thread
From: illofspeed @ 2026-10-11 16:17 UTC (permalink / raw)
  To: dlemoal; +Cc: dm-devel, agk, snitzer, mpatocka, bmarzins, hare, linux-kernel

Damien Le Moal wrote:
> Patch looks OK to me, but same comment as patch 1 about your SoB name. Please
> confirm.

Same answer as for 1/3: illofspeed is the name I use for all my public
work. If a legal name is required, I will resend.

(Resent as plain text: my first reply had an HTML part and did not reach
the lists.)

illofspeed

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH v2 1/3] dm zoned: back off when reclaim finds no destination zone
  2026-10-11 16:17     ` illofspeed
@ 2026-10-11 16:32       ` Damien Le Moal
  2026-10-11 16:47       ` Damien Le Moal
  1 sibling, 0 replies; 14+ messages in thread
From: Damien Le Moal @ 2026-10-11 16:32 UTC (permalink / raw)
  To: illofspeed; +Cc: dm-devel, agk, snitzer, mpatocka, bmarzins, hare, linux-kernel

On 2026/10/11 18:17, illofspeed wrote:
> Damien Le Moal wrote:
>> Patch looks OK, but is "ilofspeed" your full name?
> 
> illofspeed is not my legal name, but it is the name I use for all my
> public work (GitHub and this series), and I understood
> Documentation/process/submitting-patches.rst ("using a known identity")
> to allow that. If you need a legal name in the Signed-off-by, please let
> me know and I will resend.

Hmmm... I have not read that doc in a while, but the point here would be how
"identity" is defined. I do not mind you using your github handle, as long as
the rules allow it, that is, "identity" includes a github handle as OK.

-- 
Damien Le Moal
Western Digital Research

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH v2 1/3] dm zoned: back off when reclaim finds no destination zone
  2026-10-11 16:17     ` illofspeed
  2026-10-11 16:32       ` Damien Le Moal
@ 2026-10-11 16:47       ` Damien Le Moal
  1 sibling, 0 replies; 14+ messages in thread
From: Damien Le Moal @ 2026-10-11 16:47 UTC (permalink / raw)
  To: illofspeed; +Cc: dm-devel, agk, snitzer, mpatocka, bmarzins, hare, linux-kernel

On 2026/10/11 18:17, illofspeed wrote:
> Damien Le Moal wrote:
>> Patch looks OK, but is "ilofspeed" your full name?
> 
> illofspeed is not my legal name, but it is the name I use for all my
> public work (GitHub and this series), and I understood
> Documentation/process/submitting-patches.rst ("using a known identity")
> to allow that. If you need a legal name in the Signed-off-by, please let
> me know and I will resend.

Re-reading the doc, "Sign your work - the Developer's Certificate of Origin"
states that:

...
using a known identity (sorry, no anonymous contributions.)

The key here is "sorry, no anonymous contributions.)", but your github handle is
anonymous, in a way. I am not a lawyer, but that is how I interpret it.

Nevertheless, some googling is telling me that pseudonyms like a github user
name are OK. Personally, I do prefer and appreciate peoples giving their full
name, but if you prefer your github handle, that is fine.


-- 
Damien Le Moal
Western Digital Research

^ permalink raw reply	[flat|nested] 14+ messages in thread

end of thread, other threads:[~2026-10-11 16:47 UTC | newest]

Thread overview: 14+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-10 16:16 [PATCH v2 0/3] dm zoned: fix two idle reclaim problems illofspeed
2026-10-10 16:16 ` [PATCH v2 1/3] dm zoned: back off when reclaim finds no destination zone illofspeed
2026-10-11 14:24   ` Damien Le Moal
2026-10-11 16:17     ` illofspeed
2026-10-11 16:32       ` Damien Le Moal
2026-10-11 16:47       ` Damien Le Moal
2026-10-10 16:16 ` [PATCH v2 2/3] dm zoned: do not reclaim a random zone into another random zone illofspeed
2026-10-11 14:30   ` Damien Le Moal
2026-10-11 16:17     ` illofspeed
2026-10-10 16:16 ` [PATCH v2 3/3] dm zoned: keep polling for idle reclaim after a reclaim pass illofspeed
2026-10-11 14:31   ` Damien Le Moal
2026-10-11 16:17     ` illofspeed
2026-10-11 14:23 ` [PATCH v2 0/3] dm zoned: fix two idle reclaim problems Damien Le Moal
2026-10-11 14:33 ` Damien Le Moal

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®