mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Peter Zijlstra <peterz@infradead.org>
To: Shakeel Butt <shakeel.butt@linux.dev>
Cc: "Tejun Heo" <tj@kernel.org>,
	"Johannes Weiner" <hannes@cmpxchg.org>,
	"Michal Koutný" <mkoutny@suse.com>,
	"Michal Hocko" <mhocko@kernel.org>,
	"Roman Gushchin" <roman.gushchin@linux.dev>,
	"Muchun Song" <muchun.song@linux.dev>,
	"Andrew Morton" <akpm@linux-foundation.org>,
	"Ingo Molnar" <mingo@redhat.com>,
	"Juri Lelli" <juri.lelli@redhat.com>,
	"Vincent Guittot" <vincent.guittot@linaro.org>,
	"Dietmar Eggemann" <dietmar.eggemann@arm.com>,
	"Steven Rostedt" <rostedt@goodmis.org>,
	"Ben Segall" <bsegall@google.com>, "Mel Gorman" <mgorman@suse.de>,
	"Valentin Schneider" <vschneid@redhat.com>,
	"K Prateek Nayak" <kprateek.nayak@amd.com>,
	"Suren Baghdasaryan" <surenb@google.com>,
	"Kumar Kartikeya Dwivedi" <memxor@gmail.com>,
	"David Dai" <david.dai@linux.dev>,
	"JP Kobryn" <jp.kobryn@linux.dev>,
	"Frederic Weisbecker" <frederic@kernel.org>,
	"Aaron Lu" <ziqianlu@bytedance.com>,
	"Daniel Jordan" <daniel.m.jordan@oracle.com>,
	"Hao Lee" <haolee.swjtu@gmail.com>,
	kernel-team@meta.com, cgroups@vger.kernel.org,
	bpf@vger.kernel.org, linux-mm@kvack.org,
	linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [RFC PATCH 0/7] cgroup: charge kernel work to the cgroup it is done for
Date: Thu, 1 Oct 2026 12:54:24 +0200	[thread overview]
Message-ID: <20261001105424.GK4121339@noisy.programming.kicks-ass.net> (raw)
In-Reply-To: <20260924184714.912181-1-shakeel.butt@linux.dev>

On Thu, Sep 24, 2026 at 11:47:04AM -0700, Shakeel Butt wrote:

> Options we looked at
> ====================
> 
> A. Keep the shared kworker, and only charge the cgroup.
> 
>   A1. Measure how long the reclaim ran, then charge that time to the
>       memcg's cgroup.
> 
>   A2. Let the kworker say "charge this cgroup" while it works, the
>       same way set_active_memcg() works for memory.
> 
> B. Also make the cgroup's CPU limits apply.
> 
>   B1. Run the kworker in the cgroup's scheduling group while it works,
>       so cpu.weight and cpu.max apply to it.
> 
>   B2. Run the reclaim at full speed, then take its CPU time out of the
>       cgroup's cpu.max quota afterwards (back-charging).
> 
> C. Do the reclaim inside the cgroup.
> 
>   C1. Record the overage as a debt on the memcg, and let the memcg's
>       own tasks pay it the next time they charge memory or return to
>       user space.
> 
>   C2. Give each memcg its own reclaim thread that lives in the cgroup.
> 
>   C3. Use a cgroup-aware workqueue, as in [1], or per-cgroup worker
>       pools.

> What we chose and why
> =====================
> 
> We chose A2, together with B2 for cpu.max.
> 
> We do not throttle the reclaim itself. Adding limits to it is more
> complicated and most probably unneeded, as we envision that we will
> need concurrent background reclaimers instead of throttling in a
> real-world environment. We are working on developing a system to
> balance the rate of allocations/charges with the rate of reclaim, to
> keep the system always running effectively.
> 
> In addition, reclaim takes sleeping locks, like i_mmap_rwsem and the
> anon_vma lock in rmap walks, and fs locks in shrinkers. With a low
> cgroup weight, the kworker can be preempted while it holds them, and
> then tasks in other cgroups wait on it. It also hurts the workqueue. A
> kworker that is runnable but not running still counts as running, and
> the workqueue only spots CPU hogs by the CPU time they use. So other
> work queued on that CPU's system_wq waits too.

With proxy execution, this should be alleviated. Then all tasks waiting
on a lock acquisition will contribute to the runnability of the lock
holder.

> Still, the reclaim should not be free CPU time on top of the cgroup's
> limit. With A2 alone, the reclaim shows up in cpu.stat, but the
> cgroup's own tasks still get their full cpu.max quota. B2 fixes that
> without slowing the reclaim down. The kworker runs at full speed, and
> afterwards its time is taken out of the cgroup's cpu.max quota, so the
> cgroup's own tasks get less CPU time instead. This is the
> back-charging Tejun described [2]. It only covers cpu.max.
> Back-charging cpu.weight is future work.

Fiddling with weight sounds like horrible garbage. If you find you need
that, I would really rather you went C[23].

> We did not take C2 or C3. A high_work run asks for only 64 pages,
> which is far too little to pay for moving a thread into a cgroup
> through the global cgroup locks. A thread per memcg means thousands of
> threads. cgroup v2 does not allow tasks in a non-leaf domain cgroup.
> And a kernel thread in a cgroup shows up in cgroup.procs and blocks
> rmdir.

You don't move the threads, you spawn them on group creation and leave
them there. I'm sure you can fudge the rmdir thing with less ugly than
you're proposing here and in the other series.



  parent reply	other threads:[~2026-10-01 10:54 UTC|newest]

Thread overview: 19+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-24 18:47 Shakeel Butt
2026-09-24 18:47 ` [RFC PATCH 1/7] cgroup: add cgroup_account_system_time() Shakeel Butt
2026-09-24 18:47 ` [RFC PATCH 2/7] cgroup: add set_active_cgroup() to charge CPU time to a cgroup Shakeel Butt
2026-09-24 18:47 ` [RFC PATCH 3/7] psi: charge pressure to the task's active cgroup Shakeel Butt
2026-09-24 18:47 ` [RFC PATCH 4/7] sched/fair: add cfs_bandwidth_charge() for kernel work done for a cgroup Shakeel Butt
2026-09-24 18:47 ` [RFC PATCH 5/7] cgroup: take set_active_cgroup() time out of the cgroup's cpu.max Shakeel Butt
2026-09-24 18:47 ` [RFC PATCH 6/7] memcg: charge high_work reclaim to the memcg Shakeel Butt
2026-09-24 18:47 ` [RFC PATCH 7/7] selftests: cgroup: check that high_work reclaim is charged " Shakeel Butt
2026-09-24 20:28 ` [RFC PATCH 0/7] cgroup: charge kernel work to the cgroup it is done for Tejun Heo
2026-09-24 21:11   ` Shakeel Butt
2026-09-28 22:39     ` Tejun Heo
2026-10-01 10:59   ` Peter Zijlstra
2026-10-01 12:18     ` Peter Zijlstra
2026-10-01 16:00       ` Shakeel Butt
2026-10-01 12:40     ` Sebastian Andrzej Siewior
2026-10-01 18:59     ` Tejun Heo
2026-10-01 10:54 ` Peter Zijlstra [this message]
  -- strict thread matches above, loose matches on Subject: below --
2026-09-24 18:45 Shakeel Butt
2026-09-24 18:50 ` Shakeel Butt

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261001105424.GK4121339@noisy.programming.kicks-ass.net \
    --to=peterz@infradead.org \
    --cc=akpm@linux-foundation.org \
    --cc=bpf@vger.kernel.org \
    --cc=bsegall@google.com \
    --cc=cgroups@vger.kernel.org \
    --cc=daniel.m.jordan@oracle.com \
    --cc=david.dai@linux.dev \
    --cc=dietmar.eggemann@arm.com \
    --cc=frederic@kernel.org \
    --cc=hannes@cmpxchg.org \
    --cc=haolee.swjtu@gmail.com \
    --cc=jp.kobryn@linux.dev \
    --cc=juri.lelli@redhat.com \
    --cc=kernel-team@meta.com \
    --cc=kprateek.nayak@amd.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=memxor@gmail.com \
    --cc=mgorman@suse.de \
    --cc=mhocko@kernel.org \
    --cc=mingo@redhat.com \
    --cc=mkoutny@suse.com \
    --cc=muchun.song@linux.dev \
    --cc=roman.gushchin@linux.dev \
    --cc=rostedt@goodmis.org \
    --cc=shakeel.butt@linux.dev \
    --cc=surenb@google.com \
    --cc=tj@kernel.org \
    --cc=vincent.guittot@linaro.org \
    --cc=vschneid@redhat.com \
    --cc=ziqianlu@bytedance.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®