From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm2-f9.google.com (mail-wm2-f9.google.com [74.125.225.137]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 72821319617 for ; Thu, 13 Aug 2026 03:26:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.225.137 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786591593; cv=none; b=bJxskziP1mopJXiMJ2Cj5s2mfc2n7PpJFvl0EGYNPMP98N+m25CCuSAxg5iIiUJMEf/0SZqgF+/KmQ6BYrKhzlUriCRgIjxTZ8dIlNCuWJIGqDxAfnHkm9kftyTfvhrgvprDX0y+NQxxUQ1NlddNBDKSBAkK1QPWaWko0zyso20= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786591593; c=relaxed/simple; bh=X0pHP4S2OsRNcX4/HbKprcP2h56QDiwvZLsJATuwvS4=; h=Mime-Version:Content-Type:Date:Message-Id:Cc:Subject:From:To: References:In-Reply-To; b=eEHtnJ4lG7ZDJElUC1LOjaAviPAXVQI5BYNWrW39IYlKpcgfRKQrWIa4spxTrMNZC3cpR2Zp2YotkQr3JfXoHkviPutPjpnXT9+Sn0F79EGtNc0TZuK+cwwayopNk74sUab5mZGyDK9Kj1PSOKgPWGcjv/vb2IGMjQzoaHf10Uw= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=OyGzc+CZ; arc=none smtp.client-ip=74.125.225.137 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="OyGzc+CZ" Received: by mail-wm2-f9.google.com with SMTP id 5b1f17b1804b1-49556ce3549so6598085e9.0 for ; Wed, 12 Aug 2026 20:26:30 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786591589; x=1787196389; darn=vger.kernel.org; h=in-reply-to:references:to:from:subject:cc:message-id:date :content-type:content-transfer-encoding:mime-version:from:to:cc :subject:date:message-id:reply-to:content-type; bh=glE18fTWGFGOtieGZ/t9cUibr9Ip5396IdYy4itNJLM=; b=OyGzc+CZYyaigbVJVAELJ3hPfH3bFkcydqDPyzVpDTLspUl9dMB92tfymlwohqKazP Ug4Yd+yHsVtYxlyNDZHc5WvTq8zmOtRitcivDH++cQDCss766gKdtsbTOExoUDZ+Gp2i Mus2HK92AsvmSLPg0i0A1PEhW6M5oJnqFfO/pSc6E68EoWe8mNaQPiyfRMgEeXWYbHoA JfZAkzOb3JMI9H+VUUvKNWw63ss7uMtCTNMu4vxzwDvCweoJFD0Wn5s2wUGn7idjv9Vk BWWwX2VqK/c+Jyv/qvVUWxj38c5KILmrhF4Q2QiqbyQDZMn9oWY4TsvgO+9wqT7lDTbz KD6A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786591589; x=1787196389; h=in-reply-to:references:to:from:subject:cc:message-id:date :content-type:content-transfer-encoding:mime-version:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=glE18fTWGFGOtieGZ/t9cUibr9Ip5396IdYy4itNJLM=; b=GVodOcxeBh3BdIZvHlGMf9/Is+o+dZYhS0PVa8uSamPp7xXnPfNWgmWeHvldDe9P2v jXzFFq9uu3aBmzf7UspKAN8foM1NgSHYM0P2Emo8sWgkHskoDdBr0lrj0Ip8RqjlpC6W YQqxSr4c4rLaXXHCqWTtcjQ9xq9tGSgRoYV60k09d4sTJlfs9BxJAAktRuur/XtJgF4d XF6yviFpZ+uLHBo3iVdZk30yvl0nSOhsYoD/QOl2HHMiWbrtsT85eAgL2LVXU6IuTZUR Xr7004bdJO81+/BssKSfZscNhjOILIuzIPR/lFZN0GJCcTihwdQfXYIhoYBP2bWwDOrV gX1g== X-Forwarded-Encrypted: i=1; AHgh+Ro29c2jrN8P6kkH8IK8A+kVd+OTNVzL+bEP3InjR7+GuQaeXhNjDdKBz+dKhwDNVYT2FNQg+em95TQz4lA=@vger.kernel.org X-Gm-Message-State: AOJu0YzTBtGwHE3D+28YtlS1HIeJtQbLWl+3T4oIoKGoM8QmZ2Lp1EQT bWZSuNsYjlEJd+BaNYc9YowMMnv6POfR0hHQkC59n3iAmInQsXMxGROp X-Gm-Gg: AR+sD12un02cSCMd0mtQ8UuTpMnG9Z3HIC6Ai3FSXJqUcCv03N9uTFkLVJs4LkzQ2A/ Hcjz2BqDw1WU7vObh6CARcTNhetAjRBf6eDQ4FxmcRHyZONV840dU45TjJpi786XJY3s2AOjOt/ rw3ctnUyjNKT1iI3tcf2bhHJ+e5xrEsK6PdGnfXvCv2PzKMUebf2//fAbUNO/CLMxDVkMrEF4Xi iE5Cel52TrxeTs6d6M+EUOUR1o5pIQ7ysPQcf8K9MBB/lONKg+IniLOvFZEkk5hUkhzFDCuX1MH FRv5SlqEOTOIbVytiFy623tgl2MYii/a7xG4NU3q3fRg4ammIknU1ivsktC8WOf5jOkmSt0uFGQ zmBmEUEzuDUE2fs5ToJjn2TQB4xbs42euR2LG6Fhyai+qccm6IDfxoqpef/A1DKzvdL+8Tg/gHy I0rm9MrnFCh1mgHVTTY+p6a+nQ+lDtdwTVZ+ZPXyvycocGZLBVgVydpPprC/KHRqUjYgPPDDK/e mXzB5wlRz/fh9Ypn5euZbkMXRCLbs0QtAXAB/zUkBEVDEwrltwlA+4WXYqGWVu2Xx5lpwzPIVCr TX+QKEICFm/EabKrHmEjfJ4q5xYejgXNvWbmDQ== X-Received: by 2002:adf:e008:0:10b0:481:5167:d526 with SMTP id ffacd0b85a97d-48159c8d9f8mr2408652f8f.6.1786591588609; Wed, 12 Aug 2026 20:26:28 -0700 (PDT) Received: from localhost (nat-icclus-192-26-29-3.epfl.ch. [192.26.29.3]) by smtp.gmail.com with ESMTPSA id ffacd0b85a97d-4815a569ab3sm2609540f8f.11.2026.08.12.20.26.26 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Wed, 12 Aug 2026 20:26:28 -0700 (PDT) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset=UTF-8 Date: Thu, 13 Aug 2026 05:26:26 +0200 Message-Id: Cc: "Hui Zhu" Subject: Re: [PATCH bpf-next 2/4] bpf: add bpf_thread_wq kthread-backed workqueue with cgroup placement From: "Kumar Kartikeya Dwivedi" To: "Hui Zhu" , "Alexei Starovoitov" , "Daniel Borkmann" , "John Fastabend" , "Andrii Nakryiko" , "Martin KaFai Lau" , "Eduard Zingerman" , "Song Liu" , "Yonghong Song" , "Jiri Olsa" , "Johannes Weiner" , "Michal Hocko" , "Roman Gushchin" , "Shakeel Butt" , "Muchun Song" , "JP Kobryn" , "Andrew Morton" , "Shuah Khan" , , "Jakub Kicinski" , "Jesper Dangaard Brouer" , "Stanislav Fomichev" , "KP Singh" , "Tao Chen" , "Mykyta Yatsenko" , "Leon Hwang" , "Anton Protopopov" , "Amery Hung" , "Tobias Klauser" , "Eyal Birger" , "Rong Tao" , "Hao Luo" , "Peter Zijlstra" , "Miguel Ojeda" , "Nathan Chancellor" , "Kees Cook" , "Tejun Heo" , "Jeff Xu" , , "Jan Hendrik Farr" , "Christian Brauner" , "Randy Dunlap" , "Brian Gerst" , "Masahiro Yamada" , "Willem de Bruijn" , "Jason Xing" , "Paul Chaignon" , "Lance Yang" , "Jiayuan Chen" , "Emil Tsalapatis" , "Ihor Solodrai" , "Barry Song" , "Geliang Tang" , , , , , , X-Mailer: aerc 0.21.0 References: In-Reply-To: On Fri Aug 7, 2026 at 9:01 AM CEST, Hui Zhu wrote: > From: Hui Zhu > > Introduce bpf_thread_wq, a new BPF embedded map field similar to > bpf_wq but backed by a dedicated kthread_worker instead of a system > workqueue. The worker kthread can be attached to a specific cgroup at > init time so BPF-deferred callbacks run under the resource limits of > the target cgroup. > > Three kfuncs are exposed: > bpf_thread_wq_init(twq, map, cgroup_id, flags) [KF_SLEEPABLE] > bpf_thread_wq_set_callback(twq, cb, flags, aux) > bpf_thread_wq_start(twq, flags) > > bpf_thread_wq_init() is registered only for BPF_PROG_TYPE_SYSCALL > programs. It creates a kthread worker and may attach it to a cgroup; > those paths can sleep and acquire kthread and cgroup locks. Restricting > init to syscall programs prevents it from running in BPF contexts that > may already hold locks which could deadlock with those paths. > > bpf_thread_wq intentionally avoids the bpf_async infrastructure used by > bpf_timer and bpf_wq. That infrastructure drives cleanup from irq_work > in hardirq context, while bpf_thread_wq cancellation and final teardown > may need to sleep through kthread_cancel_work_sync(), > kthread_destroy_worker() and a final cgroup_put(). > bpf_thread_wq_cancel_and_free() therefore cancels work synchronously and > drops the context reference; the last put waits for tasks-trace RCU > readers and then schedules process-context work to run bpf_prog_put(), > cgroup_put(), kthread_destroy_worker() and kfree(). > > Add BTF/map support for bpf_thread_wq fields, map teardown hooks, > verifier handling for the callback kfunc, and cgroup_kthread_attach() to > move the worker into the requested cgroup. > > Supported map types are BPF_MAP_TYPE_HASH, BPF_MAP_TYPE_LRU_HASH, and > BPF_MAP_TYPE_ARRAY, consistent with bpf_wq and bpf_task_work. > > Signed-off-by: Hui Zhu > --- Hi Hui, Thanks for sharing the patches. I think Sashiko and Mykyta already pointed = out a couple of issues with the current implementation, but I would like to comme= nt on the higher-level approach. If I understood the past discussions and current set correctly, the reason = for your choice to move from bpf_wq to bpf_thread_wq was primarily to enable co= rrect CPU accounting of the work done by threads to specific cgroups. I think this is a step in the right direction, but looking at the bigger picture, I feel we need a more flexible solution. In practice, users deciding to do async reclaim through such a BPF interfac= e would want to scale and compact the number of threads doing reclaim-related= work dynamically, based on available idle resources, but also application-specif= ic metrics, and I do not think the bpf_thread_wq abstraction provides enough flexibility in managing work scheduling related aspects precisely. What we probably should expose is the ability for programs to manage their = own wait queues and subscribe kthreads managed by BPF programs dynamically to t= hem. User-defined policies can dictate how many threads remain active, whether t= hey busy poll, and when they go to sleep, completely under the program's contro= l. Looking beyond this particular example, there are cases where work is stash= ed in queues, and threads draw items from the pool and process them. In such case= s having flexibility in deciding the mapping between threads and queues is al= so important. We want to thus allow management of the work item queues from th= e program itself, and not hide it behind the API to offer the desired level o= f control. The order in which items are ranked in individual queues, and the order in which queues are processed horizontally by threads is also a desir= able property. Correctly associating the BPF-managed kthread to a cgroup should still be possible in a manner similar to what you did in this set. But the amount of control programs can exert over how work is scheduled and the level of concurrency will be much higher if we disaggregate and generalize each part= of the picture (wait queues, kthreads, and work item queues, which can be implemented in BPF itself). I am not aware of ways to back charge time spent to remote cgroups, but if desired we could also explore that option when a single thread does work on behalf of multiple cgroups. It is something to be explored. I've been working on related patches, and will post RFC set for bpf_kthread= and bpf_waitq management in due time. Until then I recommend that you continue experimenting with the existing async execution primitives for now. > [...]