From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mta1.migadu.com (out-4.mta1.migadu.com [95.215.58.4]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7ABBF3C988E for ; Wed, 19 Aug 2026 06:35:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=95.215.58.4 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787121360; cv=none; b=mbvgCcqaw79OoXebvQtCiLaOF7n+dIol8uJqNB/ohU5HNpcqJLX6cNg6rn/ZMgeJgaNu/BZZc4ZWi4HNMOpc0qHhI9tMaRkAAPrfj2fjp2n/487CFrai3jvK+/fXDlC1Ickqutp0RxlcjneHd6VnfbKKU/zMJ7UjsrvBbZecbYw= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787121360; c=relaxed/simple; bh=Ru/OP1lG0ST1wZXA7RhW3VLfH8WpvYmZ6YLYeLoQvMo=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=ZDXtOO4EALZ59CcWRi4rvFzdREkrItHJvPBULq50yZsHXwjW2h2NSnHZq84BuYSWx1Qmx6tlYKwn6o7YuPQXQjHhTvlX43rk2XjiHpAD2+XYKLBW40vTOHu6qDYFhAYCZ265iRPw0MLzXl5e8pCiHZ4xiHHesiEhY8ACgLvoIDU= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=tCGKGeUZ; arc=none smtp.client-ip=95.215.58.4 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="tCGKGeUZ" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=Ru/OP1lG0ST1wZXA7RhW3VLfH8WpvYmZ6YLYeLoQvMo=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1787121356; v=1; x=1787726156; b=tCGKGeUZAIc9/dTmIsJSO3St8/r/1zzT+lXdqKSujyXo5txJx13fAMmOkehJPQG1BRm73N3r fh55Ks+td0A4fcW/nsp2UQk64kPVCeDf5cIUWNmBL+SOyAvxh74ONb6Vq1npM0KMFTLX77ZDiRM RVp+MifnSRgSpha3HUZ3RLMw= X-Envelope-To: linux-kernel@vger.kernel.org Received: from teawater-KVM-Virtual-Machine (39.156.73.13) by smtp.migadu.com with ESMTPS id 6237c860b5f0182a; Wed, 19 Aug 2026 06:35:55 +0000 X-Migadu-Flow: FLOW_OUT From: "Hui Zhu" To: Roman Gushchin , JP Kobryn , Shakeel Butt , Andrew Morton , Andrii Nakryiko , Eduard Zingerman , Ihor Solodrai , Alexei Starovoitov , Daniel Borkmann , Kumar Kartikeya Dwivedi , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Shuah Khan , Barry Song , Geliang Tang , linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-mm@kvack.org, linux-kselftest@vger.kernel.org Cc: Hui Zhu Subject: [PATCH bpf-next v3 0/2] bpf: BPF-driven proactive memcg reclaim Date: Wed, 19 Aug 2026 14:35:42 +0800 Message-ID: X-Mailer: git-send-email 2.53.0 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit From: Hui Zhu This series lets a BPF program decide when to trigger memcg reclaim and how aggressively to do it, based on whatever runtime signal it chooses to observe -- rather than reclaim only being triggered once a cgroup's usage crosses a fixed threshold. The core idea is a pair of new kfuncs, bpf_proactive_reclaim() and bpf_proactive_reclaim_swappiness(), which give BPF direct access to the proactive reclaim path so this decision can be made in BPF policy rather than hard-coded threshold logic. This was originally part of a larger series posted here [1]. That series also adds a memcg BPF struct_ops (memcg_charged, memcg_uncharged, below_low, below_min) for synchronous, in-line memory protection decisions. That mechanism and this one solve different problems -- struct_ops hooks run inline on the charge/reclaim path, while the kfuncs here are for asynchronous, out-of-band reclaim decided independently by a BPF program -- so they are reviewed as separate series. This series carries only the async reclaim piece. Compared to v1, the kfunc interface has been reworked based on review feedback: instead of a thin wrapper around try_to_free_mem_cgroup_pages() exposing raw gfp/reclaim-option knobs, the series now provides use-case-driven kfuncs that perform one proactive reclaim pass with the same parameters memory.reclaim uses. The bpf_thread_wq patches from v1 (old patches 2-3) are dropped from this series: following the discussion in [2], the cgroup-aware workqueue is being superseded by a disaggregated set of async primitives (bpf_kthread/bpf_waitq) that will be developed separately (discussion in [3]), and the selftest now queues its reclaim work through bpf_wq. Patch 1 adds bpf_proactive_reclaim() and bpf_proactive_reclaim_swappiness(), sleepable kfuncs that perform one reclaim pass on a target memcg, like a write to memory.reclaim: swap is allowed, and the anon/file balance follows the cgroup's swappiness or an explicit override in [MIN_SWAPPINESS, MAX_SWAPPINESS] plus SWAPPINESS_ANON_ONLY. Both delegate to try_to_free_mem_cgroup_pages() with GFP_KERNEL and MEMCG_RECLAIM_MAY_SWAP | MEMCG_RECLAIM_PROACTIVE, the same parameters user_proactive_reclaim() uses, and unlike memory.reclaim they do not retry until the requested size is reached. Both refuse to run when the caller already holds PF_MEMALLOC, since a nested try_to_free_mem_cgroup_pages() would clobber the outer reclaim's current->reclaim_state (e.g. MGLRU dereferences current->reclaim_state->mm_walk). Patch 2 (selftests/bpf: add memcg async reclaim test) ties the kfuncs into a worked example: it watches the WORKINGSET_REFAULT_FILE counter of a high-priority cgroup as a proxy for memory-pressure impact, and once it starts climbing, proactively reclaims pages from a low-priority cgroup via bpf_proactive_reclaim(), with the reclaim work queued asynchronously through bpf_wq. The test asserts that the monitored cgroup's workload finishes faster once async reclaim kicks in. This demonstrates the end-to-end use case: BPF observes pressure on the cgroup it wants to protect, and reclaims from the cgroup it wants to reclaim from, in one self-contained mechanism. Note that, without bpf_thread_wq, the CPU cost of the reclaim work is not yet attributed to a chosen cgroup; that part waits for the async primitives work mentioned above. Changelog: v3: According to the comments of bot+bpf-ci, add a shared helper bpf_proactive_reclaim_pages() that is called by bpf_proactive_reclaim and bpf_proactive_reclaim_swappiness. According to the comments of sashiko and bot+bpf-ci, fix the issues of selftests. v2: According to the comments of Shakeel Butt, replace bpf_try_to_free_mem_cgroup_pages() with bpf_proactive_reclaim(memcg, size) and bpf_proactive_reclaim_swappiness(memcg, size, swappiness). According to the comments of Kumar Kartikeya Dwivedi, drop patch 2 and patch 3. Remove bpf_thread_wq code in patch 4. According to the comments of sashiko-bot, fix the issues of selftests. [1] https://sashiko.dev/#/message/cover.1779760876.git.zhuhui%40kylinos.cn [2] https://sashiko.dev/#/message/1b58d56976202f26818d31dbd0da2ecb2e2460f5%40linux.dev [3] https://sashiko.dev/#/message/DKNHV09PBQZP.IRQL20BY574I%40gmail.com Hui Zhu (2): mm/bpf: Add bpf_proactive_reclaim kfuncs selftests/bpf: add memcg async reclaim test mm/bpf_memcontrol.c | 101 ++++ .../bpf/prog_tests/memcg_async_reclaim.c | 479 ++++++++++++++++++ .../selftests/bpf/progs/memcg_async_reclaim.c | 180 +++++++ 3 files changed, 760 insertions(+) create mode 100644 tools/testing/selftests/bpf/prog_tests/memcg_async_reclaim.c create mode 100644 tools/testing/selftests/bpf/progs/memcg_async_reclaim.c -- 2.53.0