From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm1-f49.google.com (mail-wm1-f49.google.com [209.85.128.49]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1A01B33D4F5 for ; Mon, 12 Jan 2026 08:42:15 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.49 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1768207337; cv=none; b=dTABrkOZqBL9jj3HotodEYUrT8mxiX53vKKGOH/zv0lTGCNLldwdx7b++LHy5mpJvCda+2yjE7LqUzAeB0libaO9AXllKxyEVpAXxpoTtWsSCsFhbSzexWVVF/sbq8REPkPw3wU5a30tvRnXnZLmuULN2kx53TEN0c/CBRcN3ac= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1768207337; c=relaxed/simple; bh=rjITqn12+D2tvtJTf7/BRP0AALPZnd8r3PQY4MECiIw=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=MFLrhuhVDoACuoq/FKvyZBMsr588WhZ2AoCqMNEJU4lO9OqlBTblXq0GLDPX1G4AOZh2iREzmwJ7dBXO7LrFh9ngm0HKxuc84DB3XpESJ0MOf1lkEYsWcyxYvZ8d9t3cmMtdDWEOIbAppi2GNvdRCBv0SiUVJSlPoMxF5aYJqCI= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=suse.com; spf=pass smtp.mailfrom=suse.com; dkim=pass (2048-bit key) header.d=suse.com header.i=@suse.com header.b=ECMsPALb; arc=none smtp.client-ip=209.85.128.49 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=suse.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=suse.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=suse.com header.i=@suse.com header.b="ECMsPALb" Received: by mail-wm1-f49.google.com with SMTP id 5b1f17b1804b1-4779a4fc95aso25633575e9.1 for ; Mon, 12 Jan 2026 00:42:15 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=suse.com; s=google; t=1768207334; x=1768812134; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=CDfFNoiSuTKA5KfAUoXISxuSK51bp7lujfs6vADIdxE=; b=ECMsPALbh38oKqxWkweA8gysOOon0feqAtyuXePZlNC8xXrr2CugQ/ykGFDLCYIo+I xv2m7iBI0t9mvHg2JGY2rVLY1wUuwbPHXMX/pME2zoOh/n/GH5I4HRqi/uasbl6t9ljU 1sjCWoTfadNh/PbwUetBw99/0XLPAiX/lVtj7+91xnRazzkSxbhTrcdhyrCYCIEDHvvU UIkRrhQpE+YRErb80gOnrwsOTfTG7tSH/7cyPCiRkgsElZmsMs5FZ1G1v8c1XbhudRpq wVeC5U89KChbpVBHtiCh3f8ro2sWFlGJJ/P8xNdwXWcNM6fnxx/S+ssy3nVvazNdGU// X+AQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1768207334; x=1768812134; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-gg:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=CDfFNoiSuTKA5KfAUoXISxuSK51bp7lujfs6vADIdxE=; b=WZ5yTMSPdqAVjcGoxhc4tqp8q+nmEs0OzZjaygAG9d9G5n0sgUGkjuWMY9r383DPj/ GAKsCY6OpZ2+Ho0pC6N7h2YDv/CbtPYQhXjxTXDuW0B7dPTNfKR+ZTAJ04pjI5FeaOh8 F+7DkLddFJhwY8VsVCD44CLMztFvKO+pXOnK8K5x5hbEwvpA9KWurjDwzxTewvhoL4zo p/z/KX1bMfCsQNBQX1I+wwzNgPKEP+9mOlMUZz1S1c6n2pGCEEL8zAwjo4NGNrr5S7mV aUBke32UoInfnrq+U7Yh5kkExVoHHupCqJLsaMIjdYYdTD8beaIliyIX/0DCCEBBtwEX hqjg== X-Forwarded-Encrypted: i=1; AJvYcCVdmO3qDqqrHOHAKoT9IKQ3V93xWlivzJNiKV0z46anCyA26QeWIJ+B0j59IWPdyNwSAakWZlSt4r3IS6Q=@vger.kernel.org X-Gm-Message-State: AOJu0YyQXNpn42Um1/kB4cYsFoW0s9QfMDGUSdRLlW6gwsf4FVNI959L 14LjFJ6nOfXis6V2RGPTuwCBzJo9peixpId6uPF4UR0ragVlce+ozO9EZBtO2PDRRyU= X-Gm-Gg: AY/fxX76JzXH0Efst26AFADfH8+X0sRQViLkoAp7V5K3fbff8PkuBg506h7N8QuDwUj bQsab+OB9iWMCKl3SM7BPsniD8JHB8EfD/DLROafGwCWaC045VT5bOU8XB2yuZbdTatFA0Oqf5o dEduNG1Sa1EPEbeZUOHk0fBgh9XQ/V0hYHFD8Ke+rqFKxpOY9pqJfo+OPuIUbCZZ3iZd/kXDlT3 RaRGhzITKhUxVeq5sHh+s3TWF3C4L0GSpM9HgZmQ34CnTbOaOxX9kyMSvlMTnKJiRTOXrBvC/lf N8S65bPGRduYPWGh+miCQ/wVzIW1uG/AflPpxna3+09A5SasNsm9aS7xgUM18t8pIAZuEBdwWjE 6Bt1DB9sjI9dgqVyflVWDbc9cLRo1XhcBwGg8ULQJ7+re7qcvzldeKEATFNYwMg24+93sy/NPOf uA433PDtZSqGagXRC/G2THTWc1Dnhj6K7q5xc= X-Google-Smtp-Source: AGHT+IGOZ9lkMoCzrUzWNxChXoAXhL7/lCG+k3HeWbbjlLCJz2d0Tg02k9SjW7Z9t8X4D9e1/0CcXQ== X-Received: by 2002:a05:600c:3b05:b0:475:ddad:c3a9 with SMTP id 5b1f17b1804b1-47d84877e51mr207488795e9.13.1768207334460; Mon, 12 Jan 2026 00:42:14 -0800 (PST) Received: from localhost (109-81-19-111.rct.o2.cz. [109.81.19.111]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-47d865f84besm132725385e9.1.2026.01.12.00.42.13 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 12 Jan 2026 00:42:13 -0800 (PST) Date: Mon, 12 Jan 2026 09:42:12 +0100 From: Michal Hocko To: Mathieu Desnoyers Cc: Andrew Morton , linux-kernel@vger.kernel.org, "Paul E. McKenney" , Steven Rostedt , Masami Hiramatsu , Dennis Zhou , Tejun Heo , Christoph Lameter , Martin Liu , David Rientjes , christian.koenig@amd.com, Shakeel Butt , SeongJae Park , Johannes Weiner , Sweet Tea Dorminy , Lorenzo Stoakes , "Liam R . Howlett" , Mike Rapoport , Suren Baghdasaryan , Vlastimil Babka , Christian Brauner , Wei Yang , David Hildenbrand , Miaohe Lin , Al Viro , linux-mm@kvack.org, linux-trace-kernel@vger.kernel.org, Yu Zhao , Roman Gushchin , Mateusz Guzik , Matthew Wilcox , Baolin Wang , Aboorva Devarajan Subject: Re: [PATCH v13 2/3] mm: Fix OOM killer inaccuracy on large many-core systems Message-ID: References: <20260111194958.1231477-1-mathieu.desnoyers@efficios.com> <20260111194958.1231477-3-mathieu.desnoyers@efficios.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260111194958.1231477-3-mathieu.desnoyers@efficios.com> Hi, sorry to jump in this late but the timing of previous versions didn't really work well for me. On Sun 11-01-26 14:49:57, Mathieu Desnoyers wrote: [...] > Here is a (possibly incomplete) list of the prior approaches that were > used or proposed, along with their downside: > > 1) Per-thread rss tracking: large error on many-thread processes. > > 2) Per-CPU counters: up to 12% slower for short-lived processes and 9% > increased system time in make test workloads [1]. Moreover, the > inaccuracy increases with O(n^2) with the number of CPUs. > > 3) Per-NUMA-node counters: requires atomics on fast-path (overhead), > error is high with systems that have lots of NUMA nodes (32 times > the number of NUMA nodes). > > The approach proposed here is to replace this by the hierarchical > per-cpu counters, which bounds the inaccuracy based on the system > topology with O(N*logN). The concept of hierarchical pcp counter is interesting and I am definitely not opposed if there are more users that would benefit. >From the OOM POV, IIUC the primary problem is that get_mm_counter (percpu_counter_read_positive) is too imprecise on systems when the task is moving around a large number of cpus. In the list of alternative solutions I do not see percpu_counter_sum_positive to be mentioned. oom_badness() is a really slow path and taking the slow path to calculate a much more precise value seems acceptable. Have you considered that option? -- Michal Hocko SUSE Labs