From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mta0.migadu.com (out-75.mta0.migadu.com [91.218.175.75]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 989C645198A for ; Fri, 2 Oct 2026 22:19:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.75 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790979590; cv=none; b=OnroTCO/tpNGDckd8+0WsPEuAEi1sHhhBLkvXZIcggwyDN4U+C8sBGuqV0Ik+fYTB9oivoiqP7+sDsOViJQs6QrnqjcnCQewo8yo6Yg5GPuSZss4HxPSGmAUAy9JVpdxAAhRAuNKYblX8N95/dfhj1WCuA/oZlzvh3iAvZJ1MV0= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790979590; c=relaxed/simple; bh=LIeypzTLxZFUaXt5ugGV2tt9HvkLhl45i2GZQQOzRtA=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=Nevp9DLgn18It4ID3aKr3YPmUOJjZQ/6IY9R+l58Afelfc322Qx8DSnIggTCPH2AOQXltkf8OC/sFC/0tq9jG8CNG5iCWcy+RZeQltRa8XLHO31FPRSJ2OoldDo2ephgkx6xaqLD454l2ZyVjKuqS9ZUP6PpiArkVzlRcWJUZEs= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=BhepoCzX; arc=none smtp.client-ip=91.218.175.75 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="BhepoCzX" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=LIeypzTLxZFUaXt5ugGV2tt9HvkLhl45i2GZQQOzRtA=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1790979584; v=1; x=1791584384; b=BhepoCzXqCV6+2aFXT/5CmOqdAoSYoHbqDxEnkF4m0Xvdqb6AWvH8+ChwpvQPvS343Ccuu3o J42qH3XqO24w3NmnS8ix4Pe+F7ROpe3BO3JCMsVAwMHUneXcpT9E/0xfeCr/FHmKGgh1yIX2Baa kARXwtyunWysy7L0cbgpXQKM= X-Envelope-To: linux-kernel@vger.kernel.org Received: by smtp.migadu.com with ESMTPS id 2a1b46a32a9d1e0a; Fri, 02 Oct 2026 22:19:43 +0000 X-Mizu-Trace-ID: 2a1b46a32a9d1e0a X-Migadu-Flow: FLOW_OUT Date: Fri, 2 Oct 2026 15:19:38 -0700 From: Shakeel Butt To: Tejun Heo Cc: Andrew Morton , Alexei Starovoitov , Johannes Weiner , Michal Hocko , Roman Gushchin , JP Kobryn , Muchun Song , Michal Koutny , Amery Hung , Daniel Borkmann , Andrii Nakryiko , Eduard Zingerman , Kumar Kartikeya Dwivedi , Martin KaFai Lau , Song Liu , Yonghong Song , Emil Tsalapatis , Jiri Olsa , Ihor Solodrai , John Fastabend , Jiayuan Chen , hui.zhu@linux.dev, Donet Tom , Greg Thelen , Meta kernel team , linux-mm@kvack.org, bpf@vger.kernel.org, cgroups@vger.kernel.org, linux-kernel@vger.kernel.org Subject: Re: [RFC PATCH 0/4] memcg_ext: memcg policy through cgroup-attached struct_ops Message-ID: References: <20260921192559.2619635-1-shakeel.butt@linux.dev> <31a871a07fccc99e953a2d633fc51edc@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: On Fri, Oct 02, 2026 at 06:25:08AM -1000, Tejun Heo wrote: > Hello, Shakeel. > > On Thu, Oct 01, 2026 at 03:56:47PM -0700, Shakeel Butt wrote: > > We are in agreement on (1) and (2) completely. For (3), I am fine with removing > > the sync enforcement, but for throttling points for bulk operation sites, > > I think we should only add them when there is an actual use case for that > > or someone complains about overrun from those sites. > > On (3), if removing synchronous enforcement wouldn't regress anything, > that's fine, but why was it added in the first place? > I added the sync enforcement to replace a Google internal feature which, on memcg OOM, allows node controller couple of seconds to either increase the max limit or let the memcg die. With sync enforcement, memory.high helped in simple benchmarks. However later testing on some realistic Google workloads, I found out that several thousand threads are very normal of typical Google workload and memory.high sync enforcement is not effective on applications with large amount of threads. In addition, there were workloads which on noticing blocked threads, keep forking more threads. At the end implementing that feature using memory.high didn't pan out. > > Now, setting aside the default behavior of memory.high, I want to provide > > additional flexibility to users for (3) specifically. One specific case is > > letting users opt in to async reclaim instead of the other forms of memory.high > > enforcement. Basically, users can specify that instead of having their > > application threads throttled, they would prefer async reclaim to bring their > > usage back below memory.high. > > As for flexibility, we already have a gradient of enforcement around > memory.high. Is the need here to make the shape of that gradient > configurable? Can you give specific examples where this is needed? > The concrete example I have is the kswapd like async reclaimers (plural) per memcg. Kswapd is woken up on free pages falling below low watermark and then when free pages fall below min watermark, allocators get throttled (enter direct reclaim). I want to apply similar concept to memcg (but with right cpu accounting and more concurrency).