From: Shakeel Butt <shakeel.butt@linux.dev>
To: Tejun Heo <tj@kernel.org>
Cc: Andrew Morton <akpm@linux-foundation.org>,
Alexei Starovoitov <ast@kernel.org>,
Johannes Weiner <hannes@cmpxchg.org>,
Michal Hocko <mhocko@kernel.org>,
Roman Gushchin <roman.gushchin@linux.dev>,
JP Kobryn <jp.kobryn@linux.dev>,
Muchun Song <muchun.song@linux.dev>,
Michal Koutny <mkoutny@suse.com>,
Amery Hung <ameryhung@gmail.com>,
Daniel Borkmann <daniel@iogearbox.net>,
Andrii Nakryiko <andrii@kernel.org>,
Eduard Zingerman <eddyz87@gmail.com>,
Kumar Kartikeya Dwivedi <memxor@gmail.com>,
Martin KaFai Lau <martin.lau@linux.dev>,
Song Liu <song@kernel.org>,
Yonghong Song <yonghong.song@linux.dev>,
Emil Tsalapatis <emil@etsalapatis.com>,
Jiri Olsa <jolsa@kernel.org>,
Ihor Solodrai <ihor.solodrai@linux.dev>,
John Fastabend <john.fastabend@gmail.com>,
Jiayuan Chen <jiayuan.chen@linux.dev>,
hui.zhu@linux.dev, Donet Tom <donettom@linux.ibm.com>,
Greg Thelen <gthelen@google.com>,
Meta kernel team <kernel-team@meta.com>,
linux-mm@kvack.org, bpf@vger.kernel.org, cgroups@vger.kernel.org,
linux-kernel@vger.kernel.org
Subject: Re: [RFC PATCH 0/4] memcg_ext: memcg policy through cgroup-attached struct_ops
Date: Wed, 30 Sep 2026 06:28:21 -0700 [thread overview]
Message-ID: <arw4rjVkm2ob9ioL@linux.dev> (raw)
In-Reply-To: <bda525506b3fb3e61da28a112f28abba@kernel.org>
On Mon, Sep 28, 2026 at 10:40:20AM -1000, Tejun Heo wrote:
> Hello, Shakeel.
>
> On Mon, Sep 21, 2026 at 12:25:55PM -0700, Shakeel Butt wrote:
> > try_charge_memcg() calls __mem_cgroup_handle_over_high() before it returns,
> > which reclaims and can throttle the task. That happens wherever the charge
> > happens, so a task holding a kernel lock can be stuck there, and everything
> > waiting on that lock is stuck behind it.
>
> Why not just raise the lazy bound high enough that most charges never
> enforce inline, and maybe annotate the specific paths that can allocate a
> lot so that they do? Inline enforcement should be the exception, not the
> rule. Flipping that and then trying to reverse it with custom BPF policies
> doesn't make a lot of sense.
>
Please correct me if I misunderstood you. Mainly, you are saying that we
should have a sane default behavior for memory.high. At the moment, if the
current charging process accumulates charges totaling more than
MEMCG_CHARGE_BATCH pages and the target memcg is over its high limit,
memory.high is enforced synchronously. You are suggesting that we should
increase the threshold from MEMCG_CHARGE_BATCH to some arbitrarily large
number. In that case, synchronous enforcement of memory.high will be very
rare.
I am fine with changing the default behavior. Actually, I have been
contemplating whether I should propose a revert of commit c9afe31ec443e
("memcg: synchronously enforce memory.high for large overcharges") because
it has introduced more problems than it has solved, but that is a separate
topic. The initial commit already mentioned that MEMCG_CHARGE_BATCH was used
arbitrarily, so replacing it with something big might be acceptable. I want
to keep that decision separate.
I am not sure about annotating specific paths. I think it would impose a
greater maintenance burden as the kernel evolves, since the annotations
might become stale. Also, people might object to adding memcg-internal hooks
in non-memcg code paths. In any case, this can be explored separately.
Returning to the actual proposal, my plan was to start small with a narrow,
specific use case. However, my long-term plan is to provide a mechanism to
change the default behavior for custom use cases. For example, for
memory.high, I will provide a way for users to specify what behavior they
want, i.e., whether or not they want more synchronous throttling. Second, I
will introduce a mechanism to trigger async reclaim workers. I also plan to
extend this functionality to memory.max. I just wanted to convey that I will
keep pushing this proposal, with the use cases adjusted a bit.
> > One concrete scenario which can be resolved by this new feature is the
> > kernfs notify worker. It delivers notifications with the cgroup2
> > kernfs_rwsem held for read, and the charge for the delivery allocation goes
> > to the cgroup that set the watch, usually one already under pressure. So
> > the worker reclaims while holding the lock, a waiting writer blocks every
> > later reader, and anything touching cgroupfs stalls for seconds.
>
> Slowing down the kworker inline doesn't make sense. It's charging on behalf
> of the watcher through set_active_memcg(), which already tells us whose debt
> it is. Wouldn't it make more sense to defer the debt to that cgroup instead
> of slowing down the kernel thread?
>
Yes, that makes sense. When a non-task-context charge exceeds memory.high,
the kernel already schedules high_work for the charged memcg. I plan to do
the same for kthreads so they need not reclaim or throttle inline. This is
best-effort reclaim rather than exact debt accounting; I still need to
examine userspace tasks that charge another memcg through
set_active_memcg(). I am addressing the reclaim worker's CPU accounting
separately.
Thanks for taking a look and providing feedback.
next prev parent reply other threads:[~2026-09-30 13:28 UTC|newest]
Thread overview: 19+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-21 19:25 Shakeel Butt
2026-09-21 19:25 ` [RFC PATCH 1/4] bpf, cgroup: fix cgroup struct_ops query for a second attach type Shakeel Butt
2026-09-21 20:19 ` bot+bpf-ci
2026-09-29 11:49 ` Yafang Shao
2026-09-21 19:25 ` [RFC PATCH 2/4] memcg_ext: add cgroup-attached bpf_memcg_ops Shakeel Butt
2026-09-21 19:25 ` [RFC PATCH 3/4] memcg_ext: allow BPF to defer memory.high enforcement Shakeel Butt
2026-09-24 20:14 ` JP Kobryn
2026-09-24 21:19 ` Shakeel Butt
2026-09-21 19:25 ` [RFC PATCH 4/4] selftests/bpf: add a cgroupfs lock-holder bpf_memcg_ops sample Shakeel Butt
2026-09-23 13:07 ` [RFC PATCH 0/4] memcg_ext: memcg policy through cgroup-attached struct_ops Yafang Shao
2026-09-23 15:47 ` Shakeel Butt
2026-09-24 10:01 ` Yafang Shao
2026-09-24 20:42 ` Shakeel Butt
2026-09-28 3:23 ` Yafang Shao
2026-09-28 20:40 ` Tejun Heo
2026-09-30 13:28 ` Shakeel Butt [this message]
2026-09-30 23:35 ` Tejun Heo
2026-10-01 1:00 ` Shakeel Butt
2026-10-01 18:42 ` Tejun Heo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=arw4rjVkm2ob9ioL@linux.dev \
--to=shakeel.butt@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=ameryhung@gmail.com \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=cgroups@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=donettom@linux.ibm.com \
--cc=eddyz87@gmail.com \
--cc=emil@etsalapatis.com \
--cc=gthelen@google.com \
--cc=hannes@cmpxchg.org \
--cc=hui.zhu@linux.dev \
--cc=ihor.solodrai@linux.dev \
--cc=jiayuan.chen@linux.dev \
--cc=john.fastabend@gmail.com \
--cc=jolsa@kernel.org \
--cc=jp.kobryn@linux.dev \
--cc=kernel-team@meta.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=martin.lau@linux.dev \
--cc=memxor@gmail.com \
--cc=mhocko@kernel.org \
--cc=mkoutny@suse.com \
--cc=muchun.song@linux.dev \
--cc=roman.gushchin@linux.dev \
--cc=song@kernel.org \
--cc=tj@kernel.org \
--cc=yonghong.song@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®