From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtpout.efficios.com (smtpout.efficios.com [158.69.130.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4168C23817F for ; Sat, 13 Dec 2025 19:04:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=158.69.130.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1765652663; cv=none; b=IVXThQmGh1ogeH2S56QkB8ZgKwjhHQeo0NpSGEoKHZW+dFK3TjosQpY1UIAghfPRtDGyZQq7tFD833zP3hwOSl40I/e7MdnX6eWXT602WABRZoN9Spyp67Lvax0AalltBEcpYQzOosV7mxPh9FOPj243znHsZGTozP4i38Ml4OI= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1765652663; c=relaxed/simple; bh=GETXBNJ/LSDNSEBhKYXtfbtjTjT8hoQqPituucnfiUE=; h=From:To:Cc:Subject:Date:Message-Id:MIME-Version; b=cnKCGpDag3mEINZ8PUkQi67ptIQ4xg0iv8nbVu+NBve9x7ki5KVgVIHbwCw6FDTD622k15me16xTjPjuJPAhxsxjD51ummVINf6mcXVLCGx2z+sI9oWUDh82AE+XS8IJlYOGuaTBvLol9GL/2hi1w84pZHr1txjZcrhsWFTVe18= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=efficios.com; spf=pass smtp.mailfrom=efficios.com; dkim=pass (2048-bit key) header.d=efficios.com header.i=@efficios.com header.b=Ge9Q+6l6; arc=none smtp.client-ip=158.69.130.18 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=efficios.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=efficios.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=efficios.com header.i=@efficios.com header.b="Ge9Q+6l6" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=efficios.com; s=smtpout1; t=1765652174; bh=XnWXxYgPlTX2nSNMqtPdMpxR5WcKB7bf0h7dJyPySv4=; h=From:To:Cc:Subject:Date:From; b=Ge9Q+6l6VnDaMS3disfHY1Hc0/VgRuCHOpiChO0uZBE+EyBQxuTx5tbrOeDCn7Nn0 G38I0FMTYQWBCRehovjp4nyDli/R0dueY9b2rYo8AaX3wD4SPUwhvqHfnyYCJ5GaX/ 7S9+tMUrKHidWTkS0kw3tqK0XpMFub1zltk/eRD0UbCuL7JaVIobFlryV2gD38BPUM sEZEQ7YWk015OkoT24w8jk15KRacqk/SMGvMxgygGxtXdLtr7yj6larse18WU3USeV LqWggS5SvyB6J0zZO530o/XQa1L3WLTrbR08qC7dcxkpe588+OaEQq4BbnUCUf93ui a6fUtZs0OghtA== Received: from thinkos.internal.efficios.com (mtl.efficios.com [216.120.195.104]) by smtpout.efficios.com (Postfix) with ESMTPSA id 4dTFsV3fkvzZXc; Sat, 13 Dec 2025 13:56:14 -0500 (EST) From: Mathieu Desnoyers To: Andrew Morton Cc: linux-kernel@vger.kernel.org, Mathieu Desnoyers , "Paul E. McKenney" , Steven Rostedt , Masami Hiramatsu , Dennis Zhou , Tejun Heo , Christoph Lameter , Martin Liu , David Rientjes , christian.koenig@amd.com, Shakeel Butt , SeongJae Park , Michal Hocko , Johannes Weiner , Sweet Tea Dorminy , Lorenzo Stoakes , "Liam R . Howlett" , Mike Rapoport , Suren Baghdasaryan , Vlastimil Babka , Christian Brauner Subject: [PATCH v10 0/3] mm: Fix OOM killer inaccuracy on large many-core systems Date: Sat, 13 Dec 2025 13:56:05 -0500 Message-Id: <20251213185608.3418096-1-mathieu.desnoyers@efficios.com> X-Mailer: git-send-email 2.39.5 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Introduce hierarchical per-cpu counters and use them for RSS tracking to fix the per-mm RSS tracking which has become too inaccurate for OOM killer purposes on large many-core systems. The following rss tracking issues were noted by Sweet Tea Dorminy [1], which lead to picking wrong tasks as OOM kill target: Recently, several internal services had an RSS usage regression as part of a kernel upgrade. Previously, they were on a pre-6.2 kernel and were able to read RSS statistics in a backup watchdog process to monitor and decide if they'd overrun their memory budget. Now, however, a representative service with five threads, expected to use about a hundred MB of memory, on a 250-cpu machine had memory usage tens of megabytes different from the expected amount -- this constituted a significant percentage of inaccuracy, causing the watchdog to act. This was a result of commit f1a7941243c1 ("mm: convert mm's rss stats into percpu_counter") [1]. Previously, the memory error was bounded by 64*nr_threads pages, a very livable megabyte. Now, however, as a result of scheduler decisions moving the threads around the CPUs, the memory error could be as large as a gigabyte. This is a really tremendous inaccuracy for any few-threaded program on a large machine and impedes monitoring significantly. These stat counters are also used to make OOM killing decisions, so this additional inaccuracy could make a big difference in OOM situations -- either resulting in the wrong process being killed, or in less memory being returned from an OOM-kill than expected. The approach proposed here is to replace this by the hierarchical per-cpu counters, which bounds the inaccuracy based on the system topology with O(N*logN). Notable change for v10: The new patch 3/3 changes the implementation of the oom killer task selection to a 2-pass algorithm, where the first pass uses the fast approximation provided by the hierarchical percpu counters, and the second pass does a precise sum for all tasks which have badness values within the range of the approximation accuracy. I've done moderate testing of this series on a 256-core VM with 128GB RAM. Figuring out whether this indeed helps solve issues with real-life workloads will require broader feedback from the community. The one request I did not have time to fulfill yet is to port the tests from the librseq feature branch implementation (userspace) to the kernel selftests. This series is based on v6.18. Andrew, are you interested to try this out in mm-new ? Thanks, Mathieu Link: https://lore.kernel.org/lkml/20250331223516.7810-2-sweettea-kernel@dorminy.me/ # [1] To: Andrew Morton Cc: "Paul E. McKenney" Cc: Steven Rostedt Cc: Masami Hiramatsu Cc: Mathieu Desnoyers Cc: Dennis Zhou Cc: Tejun Heo Cc: Christoph Lameter Cc: Martin Liu Cc: David Rientjes Cc: christian.koenig@amd.com Cc: Shakeel Butt Cc: SeongJae Park Cc: Michal Hocko Cc: Johannes Weiner Cc: Sweet Tea Dorminy Cc: Lorenzo Stoakes Cc: "Liam R . Howlett" Cc: Mike Rapoport Cc: Suren Baghdasaryan Cc: Vlastimil Babka Cc: Christian Brauner Mathieu Desnoyers (3): lib: Introduce hierarchical per-cpu counters mm: Fix OOM killer inaccuracy on large many-core systems mm: Implement precise OOM killer task selection fs/proc/base.c | 2 +- include/linux/mm.h | 58 ++- include/linux/mm_types.h | 4 +- include/linux/oom.h | 12 +- include/linux/percpu_counter_tree.h | 242 ++++++++++ include/trace/events/kmem.h | 2 +- init/main.c | 2 + kernel/fork.c | 24 +- lib/Makefile | 1 + lib/percpu_counter_tree.c | 705 ++++++++++++++++++++++++++++ mm/oom_kill.c | 72 ++- 11 files changed, 1089 insertions(+), 35 deletions(-) create mode 100644 include/linux/percpu_counter_tree.h create mode 100644 lib/percpu_counter_tree.c -- 2.39.5