mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v3 0/2] writeback: let foreign flushes reach dying cgwbs
@ 2026-09-29  1:40 Liz Fong-Jones
  2026-09-29  1:40 ` [PATCH v3 2/2] writeback: kick writeback on a cgwb replaced in cgwb_create() Liz Fong-Jones
                   ` (2 more replies)
  0 siblings, 3 replies; 6+ messages in thread
From: Liz Fong-Jones @ 2026-09-29  1:40 UTC (permalink / raw)
  To: Christian Brauner, Jan Kara, Tejun Heo
  Cc: Alexander Viro, Jens Axboe, Andrew Morton, Johannes Weiner,
	Roman Gushchin, Shakeel Butt, Xin Yin, linux-fsdevel, linux-mm,
	cgroups, linux-kernel, ian, Liz Fong-Jones

At Honeycomb, a container that reads from Kafka and writes columnar
files to a host volume stalls for 30-60s after each deploy replaces it
(v6.18). We're increasingly confident this is the cause: the stall
looks the same as in the reproducer (little CPU use, lag growing
linearly and then recovering), and the mitigation it predicts, moving
our final syncfs after the last write, worked in production (below).
We haven't caught it with probes in production yet.

Patch 1 keeps killed wbs in bdi->cgwb_tree until they are released, so
foreign flushes can still reach them. Patch 2 kicks writeback on a
killed wb that a live memcg's new wb replaces after a blkcg association
change, which patch 1 leaves out of reach.

Based on vfs-7.4.misc, next to Xin Yin's 168a8c13159c, which fixes the
sizing but not the lookup of a removed cgroup's wb. Each helps without
the other; this series also applies to mainline. On mainline
(fd179f8a05be), without 168a8c13159c, the test from the changelog gave
52.8s, 73.3s and 5.2s worst lag before this series and 1.4s, 1.0s and
1.0s with it.

Syncing before the removal isn't enough: cleanup_offline_cgwbs_workfn()
skips a dying wb with any dirty inode and only retries on the next memcg
offline. On ext4, syncfs can leave inodes I_DIRTY_SYNC: at 5000 files
one syncfs still stalled in 7 of 9 runs (19-26s worst lag), a second
syncfs avoided it in 5 of 5, and with this series one syncfs lagged at
most 1.3s in 2 of 2. On XFS one syncfs left the wb clean in 5 of 5 runs,
but a single 4KiB append from the old cgroup after it stranded the wb
again (15-28s worst lag, 5 of 5; 1.0s in 2 of 2 with this series). In
production, our writer synced after closing its data files but before
writing its final Kafka offset to a small shutdown file, and moving the
syncfs after that write mitigated the stall for us on XFS. In the
reproducer, a 14-byte file written after the syncfs stranded the wb the
same way, whether created directly or via a temp file and rename (16-29s
worst lag, 7 of 7; with this series, 1.0-1.2s in 2 of 2 created
directly); an empty file did not.

The harness that produced the numbers (paced writer, lag per second,
MODE=alive control) is at
https://gist.github.com/lizthegrey/2209d831930588f63076bdc0ac7b78e2, and
a minimal recipe is below. Settings for all numbers: MEM_MAX=max
POD_MAX=4G RATE_MBPS=250, plus NFILES=5000 SYNC_OLD=1 (or 2) and
DIRTY_AFTER=1 or SHUTDOWN_FILE=offset|offset_rename|empty for the syncfs
runs and NFILES=500000 OLD_SECS=120 SYNC_OLD=1 POD_HOP=10 for the
500k-file runs.

Also reproduced unpatched on bare-metal arm64 with Ubuntu's 7.0 kernel
(ext4 on a loop device): 12.9s worst lag with the old cgroup removed,
2.5s with it kept.

Tested: arm64 KVM guests (virtme-ng), ext4 and XFS on virtio. No KASAN,
lockdep, PROVE_RCU, DEBUG_LIST, DEBUG_OBJECTS_WORK, DEBUG_ATOMIC_SLEEP
or DEBUG_PAGEALLOC reports, including a stress run toggling io on the
parent of two writers 40 times so their wbs are killed and replaced in
the same slot (160 takeovers, each followed by patch 2's kick, counted
with bpftrace). W=1 and checkpatch --strict clean, no new sparse
warnings. Patch 1 is under CONFIG_CGROUP_WRITEBACK; patch 2's
wb_wakeup() and wb_start_writeback() changes are not, so v3 was also
built with it disabled (arm64 defconfig without BLK_CGROUP, W=1 on the
touched files, no new warnings).

Not tested: x86 runtime, this series in production, linux-next.

Workaround until this lands: syncfs after the old container's last
write (twice on ext4). Lowering vm.dirty_expire_centisecs from 3000 to
500 cut the worst lag on unpatched mainline to 4.8s, 4.6s and 14.3s, but
did not remove the stall.

Related: without 168a8c13159c the flush is sized from the dead memcg's
own dirty pages, so this only helps while it has some. f6988c90671e
(mainline) also matters: on XFS with 500k synced inodes and the pod
cgroup removed 10s later, the handover was slow enough without it that
the replacement fell 21-22s behind (vfs-7.4.misc), against 1.0-1.2s
with it (mainline).

Patch 2 was tested with a live cgroup whose io controller is disabled on
its parent (its wb is killed and replaced) while a sibling keeps
appending to its files, on XFS with 5000 files. With the kick, the
replaced wb's inodes start moving over right at the takeover; without
it, they started 12s late in one of three instrumented runs. The sibling
still stalled in most of six runs either way (worst lag 20.3s, 1.2s,
10.8s, 28.5s, 3.8s and 17.0s without patch 2; 2.8s, 22.3s, 25.7s, 17.3s,
1.5s and 19.5s with it), because writing back the old wb's backlog took
up to 27s and the sibling's foreign flushes reached the new wb in the
meantime. Should foreign flushes also reach a replaced wb until it is
clean, or is that not worth it for this case?

Developed with Claude Opus 5.5, which wrote the code, the reproducers
and first drafts of this text, and ran the builds and VM tests. Claude
Fable 5.1 reviewed the code and the claims in this text before v2 and v3
were posted. I drove the investigation from the production symptoms,
designed the experiments and controls (repeated runs, parent-commit
baselines, testing this patch on its own), and reviewed the analysis,
code and results.

Minimal recipe:
	#!/bin/bash
	# As root, cgroup v2, in a directory on ext4/xfs/btrfs ($DIR).
	cg=/sys/fs/cgroup/wbmini
	mkdir $cg && echo +memory > $cg/cgroup.subtree_control
	echo 4G > $cg/memory.max
	mkdir $cg/old $cg/new
	append() {  # append $1 blocks of $2 to each of 1000 files
		for i in $(seq 1000); do
			dd if=/dev/zero of=$DIR/f$i bs=$2 count=$1 \
			   oflag=append conv=notrunc status=none
		done
	}
	(echo $BASHPID > $cg/old/cgroup.procs
	 append 8 1M            # old owner writes the files...
	 append 1 64k)          # ...and exits with all of them dirty
	rmdir $cg/old           # its cgroup goes away while they are dirty
	bpftrace -e 'kretprobe:cgroup_writeback_by_id { @ret[(int32)retval] = count(); }' &
	sleep 3
	time (echo $BASHPID > $cg/new/cgroup.procs
	      append 4 1M)      # replacement keeps appending to the same files
	kill -INT $!; wait
	rmdir $cg/new $cg
	# Unpatched: many @ret[-2] (-ENOENT) and the append stalls.
	# Patched: @ret[0] only.

---
Changes in v3:
- Use wb_dying(), scoped_guard() with the offline_node list_del moved
  into it, and {}; filter dying wbs right after the lookup in
  cgwb_create(); comment wording (Tejun)
- New patch 2: kick writeback on a wb replaced in cgwb_create() (Tejun)
- Lead patch 1's changelog with the problem and its trigger (Andrew)
- Link to v2: https://patch.msgid.link/20260928-wb-dying-cgwb-flush-v2-1-56b54cda74f2@honeycomb.io

Changes in v2:
- Keep killed wbs in bdi->cgwb_tree until release instead of walking
  bdi->wb_list (Tejun)
- Drop Fixes: and Cc: stable (Tejun)
- Link to v1: https://patch.msgid.link/20260926-wb-dying-cgwb-flush-v1-1-a8d898085a3a@honeycomb.io

---
Liz Fong-Jones (2):
      writeback: let foreign flushes reach dying cgwbs
      writeback: kick writeback on a cgwb replaced in cgwb_create()

 fs/fs-writeback.c                |   8 +--
 include/linux/backing-dev-defs.h |  13 ++++-
 include/linux/backing-dev.h      |   3 +-
 mm/backing-dev.c                 | 102 +++++++++++++++++++++++++++++++--------
 4 files changed, 101 insertions(+), 25 deletions(-)
---
base-commit: 168a8c13159c6e3f0f08da6f8fa2f633a91ba9fd
change-id: 20260926-wb-dying-cgwb-flush-06e260e2abad

Best regards,
--  
Liz Fong-Jones <lizf@honeycomb.io>


^ permalink raw reply	[flat|nested] 6+ messages in thread

end of thread, other threads:[~2026-09-30 15:47 UTC | newest]

Thread overview: 6+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-29  1:40 [PATCH v3 0/2] writeback: let foreign flushes reach dying cgwbs Liz Fong-Jones
2026-09-29  1:40 ` [PATCH v3 2/2] writeback: kick writeback on a cgwb replaced in cgwb_create() Liz Fong-Jones
2026-09-29  1:40 ` [PATCH v3 1/2] writeback: let foreign flushes reach dying cgwbs Liz Fong-Jones
2026-09-29 17:58   ` Tejun Heo
2026-09-29 19:32 ` [PATCH v3 0/2] " Tejun Heo
2026-09-30 15:47   ` Jan Kara

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®