From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-yw1-f174.google.com (mail-yw1-f174.google.com [209.85.128.174]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A2E761A0B05 for ; Thu, 19 Dec 2024 18:27:02 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.174 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1734632824; cv=none; b=uQ6DNmhcQyRpRMImUo/7olLFW0G/5gl84xUcsiQdaSWsjX8/1uz0l9aC+LAwSVjkjZ3dEHhMdwF/LDy159Sa0GGRq9csh7uL6aWjOBbuCWw+fTSnIW8wNL8Z5rLZ57lhed4d/+c+d62HsXSgflYAz4mvhoZA8peB4K5llTceM4o= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1734632824; c=relaxed/simple; bh=9U46bY47k6B/jyHiZW+9dWA0OnKjOFGQh69IOLDDOD4=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=q+/9AhhF1XXrLw/tF9GKj1hh7wvb7gWQtjkkIJWNuDSCd6MsJHCmtSxeU8oqqapqU9QP76DjcfaY6DVcUDq0h2xiiNax07+ezqKw1v4S989nsPa1WBnScln/vt9SMtECm/8q1iMIts0YBMgcOIGw9FHsB8nCfCI2PlMj53Vhhes= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=Tpipb3hH; arc=none smtp.client-ip=209.85.128.174 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="Tpipb3hH" Received: by mail-yw1-f174.google.com with SMTP id 00721157ae682-6eff4f0d627so10371127b3.1 for ; Thu, 19 Dec 2024 10:27:02 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1734632821; x=1735237621; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=SygkKPVnlRnoo/IT+7wKW1HPev+nB5KffD41QdoD2l0=; b=Tpipb3hH6fMauYCItr8+ppmBCA3BpY0389D+pPQiMqAmtnl7eoul7/BFPlUiW3UEhc 51j26x899f9OgIJv3L3wfb4BqxQWulU7hso+Enw4IMW+tzpwbfJSE5zQU7CZpj8xu6Sw /jnKVqHpshBmINC3YbH8PWiNmtd5ZQ+Ae0Rjkz35GyK3r0YuqXOfNY0K1kIiqXw//nvr e00n9Br7jWQU/YvN78y092OBLMRWdJ2ek87RlhrSK6e1Tlf8IuLjJogSJvwnmRjvASxz Ll89hoU2ZFn3aiydjIlkzv19XQIfdV+Abg2+kWSJBa5Hm1ElNYE1sJrquS7mYiGRm/rn 3hcQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1734632821; x=1735237621; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=SygkKPVnlRnoo/IT+7wKW1HPev+nB5KffD41QdoD2l0=; b=n/2h97L/hkQ7PkiiYayKl+Al8P0Q9sfDWDTAXpvCKn3s+cbsKR/G0MP5CJi2edhrbi mVMvvLaSt5M2zS6Q0rWxY1xw3XwTRTlRIabXZTlIvGlhFi1Fq1oeWocr5cmqLFAzkjfv roEOrQaouOqE3PK/tyt6wyA0XJICtzEHdqqnu1w/UMWzKN3a9ABbxG5eo5YgBgs2BeeB rd6nAIqxtShQmD7TY3gzhr/lQhLIIDU96lp3M5WU4ppeYPDXHAtjwHnhgap9GpADAKog ClpoopkVE0IijMT5Kn+Mgbm6/IkmeroReCz+YJYymYOb4g9GO0+e3ABIZyUzswlwDLrr DyJQ== X-Forwarded-Encrypted: i=1; AJvYcCU+JN+hLwxUQGHFYGc+Pe8cISYDDq2zwfOtankAZOr4wOoreOMc24TYBXoA+oeJAIVgi4fNuFhgzDOl16s=@vger.kernel.org X-Gm-Message-State: AOJu0YzdpEXRGXVIIwo79XrPlKmwketN+IjCaDk7AilEwThWMP1uHywq pNqSVpDhzA9XwL4R/ik6q253u//8tjrhLqtTxaeQYw3INMboGloi X-Gm-Gg: ASbGnctNqKVcR5jGjUyC4UJcBamLA0uP9YjXEdBacaHvibyNc2K1AGcdl/GqPGByT9P 8y2cG1BbZt/pwUH3074GxFHJT7XEVJ2pSL0iy/kvOATsjEOWzZH6+hvgpAqxMD0upCOroLa6X7U mVrfdhrptlrYCBaLlhhEyXWoOPQMQh3dMCdOKG7jw0qpUNHZxPMY2syve/jN/AKNNDYSbnetTZV 2PcVFu3K/UWCSg1WHhj47+XVAn71LdijKWjdjSmZ8jJCzVCpb/5NB36pvvfumU3AlsPeUir2JAV SvwAsD5DuvjPcDL0 X-Google-Smtp-Source: AGHT+IFUWwtaLoXNIgr0ysjjvnlQAr+xeKEh3Mtxx/e1YcNpP/53+ymF4IZulprdDduACaU+jNVaSw== X-Received: by 2002:a05:690c:d86:b0:6ef:522a:1c41 with SMTP id 00721157ae682-6f3e1b4062amr47911237b3.13.1734632821448; Thu, 19 Dec 2024 10:27:01 -0800 (PST) Received: from localhost (c-24-129-28-254.hsd1.fl.comcast.net. [24.129.28.254]) by smtp.gmail.com with ESMTPSA id 00721157ae682-6f3e743ff00sm4033727b3.32.2024.12.19.10.27.00 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 19 Dec 2024 10:27:00 -0800 (PST) Date: Thu, 19 Dec 2024 10:26:59 -0800 From: Yury Norov To: Tejun Heo Cc: Andrea Righi , David Vernet , Changwoo Min , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , linux-kernel@vger.kernel.org Subject: Re: [PATCH 1/6] sched/topology: introduce for_each_numa_hop_node() / sched_numa_hop_node() Message-ID: References: <20241217094156.577262-1-arighi@nvidia.com> <20241217094156.577262-2-arighi@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: On Wed, Dec 18, 2024 at 06:04:53AM -1000, Tejun Heo wrote: > Hello, > > On Wed, Dec 18, 2024 at 11:23:40AM +0100, Andrea Righi wrote: > ... > > > So, this would work but given that there is nothing dynamic about this > > > ordering, would it make more sense to build the ordering and store it > > > per-node? Then, the iteration just becomes walking that array. > > > > I've also considered doing that. I don't know if it'd work with offline > > nodes, but maybe we can just check node_online(node) at each iteration and > > skip those that are not online. for_each_numa_hop_mask() only traverses N_CPU nodes, and N_CPU nodes have proper distances. I think that for_each_numa_hop_node() should match for_each_numa_hop_mask(). It would be good to cross-test them to ensure that they generate the same order at least for N_CPU nodes. If you think that for_each_numa_hop_node() should traverse non-N_CPU nodes, you need a 'node_state' parameter. This will allow to make sure that at least N_CPU portion works correctly. > Yeah, there can be e.g. for_each_possible_node_by_dist() wheke nodes with > unknown distances (offline ones?) are put at the end and then there's also > for_each_online_node_by_dist() which filters out offline ones, and the > ordering can be updated from a CPU hotplug callback. We can assign UINT_MAX for those nodes I guess? > The ordering can be > probably put in an rcu protected array? I'm not sure what's the > synchronization convention around node on/offlining. Is that protected > together with CPU on/offlining? The machinery is already there, we just need another array of nodemasks - sched_domains_numa_nodes in addition to sched_domains_numa_nodes. The last one is already protected by RCU, and we need to update new array every time when sched_domains_numa_nodes updated. > Given that there usually aren't that many nodes, the current implementation > is probably fine too, so please feel free to ignore this suggestion for now > too. I agree. The number of nodes on typical system is 1 or 2. Even if it's 8, the Andrea's bubble sort will be still acceptable. So, I'm OK with O(N^2) if you guys OK with it. I only would like to have this choice explained in commit message. Thanks, Yury