From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-qk1-f180.google.com (mail-qk1-f180.google.com [209.85.222.180]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E80115477E for ; Sun, 28 Jun 2026 00:31:25 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.222.180 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782606687; cv=none; b=G5jiJmWMgxesvfFK1nqV/CVxC9IheMAiJbyFMeqUJ/EOeEhQ7Uv95Y51tfSg6sDbvFb6S0WFdaruhYcmmyFPyf61uQ0HgOa/GntIa0ui9hWaV13ObjT1BijdQSUUfb6F2Fp/FTuIAj+T60KwBzNly69WrudqmJ1XcKTCbQMb/m4= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782606687; c=relaxed/simple; bh=s+Y9WU5DlLmpv9cAOgVfG446Lpq8V5LIwHPuEJ5FUkY=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=RX75hshBU9tpj7Rnl19iRDYNHL2lFfXILDs0BPH2TpIbQn0Gc7fsseGVOGfSUUlvHx8C5jcLeFoqqtX9nq6uwjaCPJ/VaVvRXZuwItKOg9Oz9c73zXzNJpZ3G9O/Vd5tkwzZ7J0DmCo3DqYO/sPNvKQTGzb9a7XlKHiJQQezLPs= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=awWgUufr; arc=none smtp.client-ip=209.85.222.180 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="awWgUufr" Received: by mail-qk1-f180.google.com with SMTP id af79cd13be357-9225cc6fbb8so177857685a.2 for ; Sat, 27 Jun 2026 17:31:25 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1782606685; x=1783211485; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=wkN+srl6wjGtNu7iwo0FYJ4HSQUHdkfF1SLjXCC02H4=; b=awWgUufrjZ7rMswnM+3waZLxOEPEGf9h4MGeLZ732IDx/hoxL8hBL0L/S6ssH+UeKq 3ESUQ2s5gRUqkVSDxaxxBQC/2zAQcKy5XtSnVHuED/T+1nGuGbtdBpHlL1Z92uQeS4jn eXRlBsJ//AOWfjZL0ZXc0zeb+5g/UnNoyp2Ko8UI9xlmLHs7mgPgJFbbEs6oeBM/JOvF nphmdGct+eNkPvhnBAbAE4YHNqPr97OMxL5TYrli2bx57cmV0+C9fUr7znorPYjuuNrb ZfuWGtkzXd2cfRpTN4kGmpbV0aBCQ+IKtyaFYUgkeZIu59wa2H7PiRMXrkPRtNIJwzqL kApw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1782606685; x=1783211485; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-gg:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=wkN+srl6wjGtNu7iwo0FYJ4HSQUHdkfF1SLjXCC02H4=; b=nT5D/3OcAc+Dtpz70BHFBaNIMrzhDe24Fyhi0pKZmEjtKyInlwsM5ozgjL6aQo6ODF KcuQkXxaBEsGEC/iDDB34M3f8Cgi/h8XqtecAWqK8aWeV3XLkZcZ+GgSth1lR308THqR e2SBqkTzVyubYmmKQKF7X91SvwVzJHdVdTkisYFsDz+ra7kOXuRDg+zrU2gJXkI0zwz/ W4ehiEDCXAWCZnDmMbMvAF/dn4koLliiLf4rOFOhkRJBQ7TL3UzGJXiTGe/gBUkFrQSx z51Hm4uPd1xjtn6hzGNZB6AjPxYKygvUHcA2O6CHNDlAatWjuhkAyl4gjyL9S8O9A7gd a6lQ== X-Forwarded-Encrypted: i=1; AFNElJ8wAesS9jJZZWWIaJmXt4WjmT2gdcKupllpYPmhVwh8YbRSixDfrzsLoH+Qwb1cemr7sWLfG28zbJPiVkA=@vger.kernel.org X-Gm-Message-State: AOJu0Yy6yQ9PgXkDG0GjjV3XE4W0j74VCO1H6sMek8vsp/8v/wvUPhdH L9z+dtaODInvCxabkeW9DrEKGUQIjRs7/fKJ/Uor0qrUQ+6wl6XMpNSYnzJXWbClwsXI0oRaTr/ jk4hE X-Gm-Gg: AfdE7cnor1SNQ771oyhCCro9+llYOg4rnYNNhngAKnahBYPliUJt6Og5xtFp/5l5D6Z tcJppPCGqiksYrsFT38lgEYLfPHDB7oFS4/y5JKQDsDyFIktd8UdZAMAT0uzo5YpKqyalqzSm+i LCOlnuSYLwLVCSd4PyCRnTKZm1CssEpWiPZ4WKYd8fOjALuwrJH4/ZxH1AWqoxEoBJxaZT4KVHZ jsszvpvm5VQg6ffdNgGVmI5ztS+00ydd4yNbmenfvbDjST4x9ecIpD1AVUTVQkZQrCHZFREtCT/ 5gwetlqnZqqp2z15zKEmxDrgf2dS/xJtiqVjCg8xmsYfMhBZB7k0Xbdda0ak5j8dCZfuT38Qp1/ UuQebhBEhqqA6C7cVWu7mZX4V4Hakb4dERHNeHfQUgaXvE3zct3Ud3PvEyC9KtWilh4GhLCGXDC VGBLpy2CAbr3y99GxPinaMv8flZgc+lEITabrkG6BB0aOa2Ljzw39EFOvC+l3xL5iN6iEm X-Received: by 2002:a05:620a:2b90:b0:915:7fe4:cac5 with SMTP id af79cd13be357-92b3e96cd68mr914439085a.49.1782606684710; Sat, 27 Jun 2026 17:31:24 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-926000c31b4sm1637191585a.23.2026.06.27.17.31.22 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 27 Jun 2026 17:31:24 -0700 (PDT) Date: Sat, 27 Jun 2026 20:31:18 -0400 From: Gregory Price To: Andrew Morton Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, kernel-team@meta.com, rppt@kernel.org, vbabka@kernel.org, mgorman@techsingularity.net, hannes@cmpxchg.org, stable@vger.kernel.org Subject: Re: [PATCH v2] mm/vmstat: fold stranded per-cpu node stats when a node comes online Message-ID: References: <20260627202243.758289-1-gourry@gourry.net> <20260627161007.81e4533ce561c2951a69f927@linux-foundation.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260627161007.81e4533ce561c2951a69f927@linux-foundation.org> On Sat, Jun 27, 2026 at 04:10:07PM -0700, Andrew Morton wrote: > > > The existing code zeroes the per-cpu counters and causes a permanent > > skew. Fold the stranded deltas instead, before the node rejoins the > > online set. The node is not online yet and the hotplug lock is held, > > so the remote access to per-cpu values is safe. > > Oh. Shouldn't we be doing this during offlining? > I tried this first. I was unable to convince myself there was a safe way to accomplish this. 1) sashiko pointed out we can't schedule_on_each_cpu while holding the hotplug lock because we'll re-take cpus_read_lock and cause a deadlock condition with cpu-hotplug. 2) I'm not sure we can do it after the hotplug lock as been dropped, at least not safely. At the very least another hot-plug re-adding the node could start. That just seemed like a bad path. 3) foreign cpu access to the per-cpu values are not atomic with respect to in-flight folds on the target cpu. this_cpu_xchg and this_cpu_add are (i believe) only atomic wrt the cpu itself (can't be interrupted mid-exchange). doing it before node_offline() has problems (in-flight folds), doing it after node_offline() still *technically* carries the same in-flight fold risk - just narrower (fold has to have started already). I couldn't convince myself there wasn't still a race, so here we are. > > + for_each_possible_cpu(cpu) { > > That's a lot of CPUs > Unfortunately - cpus may have gone offline while the node was offline, so we legitimately have to visit every *possible* cpu :[ > > + struct per_cpu_nodestat *p = per_cpu_ptr(pgdat->per_cpu_nodestats, cpu); > > > > - p = per_cpu_ptr(pgdat->per_cpu_nodestats, cpu); > > + for (i = 0; i < NR_VM_NODE_STAT_ITEMS; i++) > > and that's a lot of items. I am aware :[. I suppose we could vectorize the collection here on some archs, but I try to avoid being clever where I can. > > I guess the overall loop count won't be large enough to cause issues, > but it's large! > > Perhaps there's some simple test we can do on the per_cpu_nodestat to > avoid the inner loop? Perhaps might need to add a field for this? Hadn't considered this, but maybe. Will take a look. > > btw, "for(int i..." is allowed nowadays. It'll make this code nicer, IMO. > aye aye o7 > And... Sashiko seems to have found a pre-existing issue: > https://sashiko.dev/#/patchset/20260627202243.758289-1-gourry@gourry.net > Will take a look, thanks! ~Gregory