From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wr1-f54.google.com (mail-wr1-f54.google.com [209.85.221.54]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4E324376BD3 for ; Wed, 1 Apr 2026 15:41:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.221.54 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1775058107; cv=none; b=INU1gB6I7D7vFxOXowAWXa8LCcKHik55DnCAJi87XOHmYO5Ko+M08G9vVSy2PxjSww4EmDAVJ0NES8fua7Tl7ZYqjJM9pSIdZjOLnlk38oSpCnsGY7m9eLZHj47GrbY+lV1a0AZ0PrdqJNNsvLgp05VlSvKpwuK5AE+GyP9kUWU= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1775058107; c=relaxed/simple; bh=VKBGCTQ4MKdm2hSyQT5lYZZPRod8DhE8ixCvt/fdYY4=; h=Message-ID:Date:MIME-Version:From:Subject:To:Cc:References: In-Reply-To:Content-Type; b=LD78BDSUyLyMuVB+30FKQaTi3uQUYB5NLFfomV1JnvtIW6pH8yZ9+EcqoT09uj2yh4FUlslm9dRbxQgxJvzpVgvsjZuhBgc+p6Y6vIX9QT9BFJ6IaRF2QvVeWXcCiR4PoaKhhDm+MomhnBYtRI6gZtj4/fSMePrX8DTpfaEElSY= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=EUJGW7zY; arc=none smtp.client-ip=209.85.221.54 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="EUJGW7zY" Received: by mail-wr1-f54.google.com with SMTP id ffacd0b85a97d-43d02a71526so1901940f8f.3 for ; Wed, 01 Apr 2026 08:41:46 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1775058105; x=1775662905; darn=vger.kernel.org; h=content-transfer-encoding:in-reply-to:content-language:references :cc:to:subject:from:user-agent:mime-version:date:message-id:from:to :cc:subject:date:message-id:reply-to; bh=y3mEKAopnQO6A/xQG6kbs0MVhm944oh12TNeKFEVQ5I=; b=EUJGW7zYPCtM8PvfZttcEXbvkoGjEO+TWxWNUT38eKXaODQu6T6sUhZGakOKCUxyl0 EPTBNkQfdSVqGlvDoFw0GPBib5UjKZVAvdTSpdyn6TYa94AaAHYHeiboG6XUMSBhsrUj Zhdt7Lm3qzsgVfmfXeacWY+Cv9LvIIOGr/lrTdcX6tCQVjPE5kN8DZUGWGWJOa6kYXuf NOmpf8KbZryZPP+6y7L1t9efUSFYLMGxodNsDhOd9ySqePO7NehiBeLdwbB5CL21Gk4T HSAYJdttAuDKsAYWITQRI+1h5A+puY0ZH/9kyrbi0HI9FqPjHgKY0fDCFc4cLfcQQ0W9 oDvA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1775058105; x=1775662905; h=content-transfer-encoding:in-reply-to:content-language:references :cc:to:subject:from:user-agent:mime-version:date:message-id:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=y3mEKAopnQO6A/xQG6kbs0MVhm944oh12TNeKFEVQ5I=; b=SiXrX8RkXT34/Xzf8LiwuUrdTrgCcQg8BhcSUzWedMajJ6Vbjr+PIXhGWSIco/53Qw T39xWf5hnk9J75xAMg6fjhGxPcyBOTQUD0utKJudoiQ/ZdIbLFjuVW0k0Nvjy9I7D36q IsbHmNCBgF6jKE0snUrhVnnnLhl1qgmIFw4gT9BPG8u7SsVKuRT0tme97Ma8A57QZUc9 77AYsDwCZkXnfuvOSjDfRvNP8cL0k1oRraK6whWqjfVuX0EPHXXNfpLc7QcjUTzlAcK6 HevVZV8ZETqjpIH4DPud0wVdSlW2hPuNYX2CSA6mZjTNiaJedah14YyRh+27p/A3Yh7Z sKAg== X-Gm-Message-State: AOJu0Yz5ZskWqmHqC5hxLNnG36xNMCUggQ23ypejq/eprXhhuJaQC77S +tqA3C7s3ZSsSXIiXWWAwhyqcXHmAjKXPmggvO4ZtuTulHJTg4TQy9Su X-Gm-Gg: ATEYQzxnBBz0KugyiQB6DQLSX6KrkSXJlUBgmeOrch8InmEYhwbiB5ByRjnlyr4JQD9 VGKoD8k4jbPpmGTs2Nyxv4jYZvePV8rs5gLT2XLl7uNctWUvtied7BikK8uwdLhz97rggpFDPL+ jFfu5Ll8GWFr3WIsGkAc4V8tFz6dcP71f8mq2EQNQyB63ydzmO9O5/xNCyeB+oz9/Dsa+0cDhOD 75gcnl3RyjSXaA9LDgbtsGutPpPPP7HnZKGKKvSTpafd8mFh72rNIuygcZfsvsdU/NUhCHUjnS+ V5XqYvccQpI/ndY3HN1orxbasIUiH2970Rqu3zR8bsMgfv7GxUFmZdtJYRyIeZOzmIYMQFayuuu aqyWt7uz/k1F1zw2pVrNNC/Asz+SnNmxnwF7Y3nwYgiu8121B8QkI8Q1lQ/4jK1iHd7tmfKWzqx eAdFCm5EPMvrzpZEjU2uoyrDJWzBgwkq1Q/63UYN+Zo0WViciQsbgGPZmOPmkI1MgLJz7oGijMV Rvr X-Received: by 2002:a05:6000:4201:b0:43b:8f38:3b88 with SMTP id ffacd0b85a97d-43d150dd7ccmr8285546f8f.25.1775058104316; Wed, 01 Apr 2026 08:41:44 -0700 (PDT) Received: from ?IPV6:2a01:5a8:304:153c:3ab3:7a9:6529:7104? ([2a01:5a8:304:153c:3ab3:7a9:6529:7104]) by smtp.gmail.com with ESMTPSA id ffacd0b85a97d-43d1e2c5253sm698396f8f.9.2026.04.01.08.41.42 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Wed, 01 Apr 2026 08:41:43 -0700 (PDT) Message-ID: Date: Wed, 1 Apr 2026 18:41:40 +0300 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird From: "Nikola Z. Ivanov" Subject: Re: [RFC PATCH] x86/topo: Unify srat_detect_node among amd/intel/hygon To: K Prateek Nayak , tglx@kernel.org, mingo@redhat.com, bp@alien8.de, dave.hansen@linux.intel.com, hpa@zytor.com, puwen@hygon.cn, peterz@infradead.org, mario.limonciello@amd.com, yazen.ghannam@amd.com, andrew.cooper3@citrix.com, kai.huang@intel.com, i@rong.moe, pawan.kumar.gupta@linux.intel.com, xin@zytor.com, darwi@linutronix.de, sohil.mehta@intel.com, suchitkarunakaran@gmail.com, sshegde@linux.ibm.com, yury.norov@gmail.com, ricardo.neri-calderon@linux.intel.com Cc: linux-kernel@vger.kernel.org, x86@kernel.org References: <20260329120841.2118684-1-zlatistiv@gmail.com> <4dde81e8-98cc-4666-a4e7-b5083210fbd0@amd.com> Content-Language: en-US In-Reply-To: <4dde81e8-98cc-4666-a4e7-b5083210fbd0@amd.com> Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 8bit On 3/30/26 7:57 AM, K Prateek Nayak wrote: > Hello Nikola, > > On 3/29/2026 5:38 PM, Nikola Z. Ivanov wrote: >> This change is provoked by an observed warning after >> commit 717b64d58cff ("x86/topo: Replace x86_has_numa_in_package") >> when faking numa nodes on intel. >> >> For example: >> >> qemu-system-x86_64 \ >> -kernel arch/x86/boot/bzImage \ >> -append "console=ttyS0 root=/dev/sda debug numa=fake=2" \ >> -hda $IMAGES/unstable.img \ >> -cpu qemu64,vendor=GenuineIntel \ >> -nographic \ >> -m 2G \ >> -smp 2 \ > You can also say: > > -smp 2,sockets=2 > > and that fixes the warning but that is not a valid solution? Why? > >> Will trigger: >> >> [ 0.066755][ T0] ------------[ cut here ]------------ >> [ 0.066755][ T0] WARNING: arch/x86/kernel/smpboot.c:698 at >> set_cpu_sibling_map+0xe41/0x1f90, CPU#1: swapper/1/0 >> [ 0.066755][ T0] Call Trace: >> [ 0.066755][ T0] >> [ 0.066755][ T0] ap_starting+0x9e/0x140 >> [ 0.066755][ T0] ? __pfx_ap_starting+0x10/0x10 >> [ 0.066755][ T0] ? fpu__init_cpu_xstate+0x5c/0x320 >> [ 0.066755][ T0] start_secondary+0x66/0x110 >> [ 0.066755][ T0] common_startup_64+0x13e/0x147 >> [ 0.066755][ T0] >> >> smpboot.c suggests that the topology is invalid as >> the CPUs are in the same package but different nodes. > To me, that looks like a broken topology from a virtualization use case > and the user can easily go fix their QEMU cmdlline if they care. I'm > pretty sure folks using NUMA emulation in production know what they are > doing. > >> Fix this by unifying the srat_detect_node function >> among amd/intel/hygon and taking the amd/hygon approach >> of falling back to LLC when SRAT is not detected. > As far as the AMD, Hygon unification goes, I don't mind that but > someone has to confirm if nearby_node() holds for all APICID > distribution on Intel. > >> Place the function inside common.c and expose it in topology.h > There is no need to make it visible out of arch/x86/kernel/cpu/ > Perhaps arch/x86/kernel/cpu/cpu/cpu.h? Yes, I have made a pretty bad mistake here also because topology.h is not x86 specific and will cause build warnings. >> The hygon code is already basically identical to amd >> except for the way it obtains the LLC ID. >> We can reuse that from the hygon code since we >> already have the struct cpuinfo_x86 passed to us. >> >> Signed-off-by: Nikola Z. Ivanov >> --- >> This is marked RFC as I lack the context for the reason >> why the intel code looks the way it does. I can see >> it went through a few changes in the 2008-2010 year range, >> which makes be believe that the comment regarding >> "not doing AMD heuristics for now" is long overdue. > So prior to you patch, If I launch: > > -smp 4,sockets=2,cores=2 > > and "numa=fake=2", the srat_detect_node() for an Intel VM maps: > > CPU#0 -> Node#0 > CPU#1 -> Node#1 > CPU#2 -> Node#0 > CPU#3 -> Node#1 This is not necessarily correct either. Those qemu parameters do not produce an interleaved topology, but the cpu to node map will end up interleaved, this ties back to the warning in smpboot.c This is what it looks like pre-patch: cpu to node mapping (interleaved): # ls -l /sys/devices/system/cpu/cpu*/node* lrwxrwxrwx    1 root     root             0 Apr  1 13:57 /sys/devices/system/cpu/cpu0/node0 -> ../../node/node0 lrwxrwxrwx    1 root     root             0 Apr  1 13:57 /sys/devices/system/cpu/cpu1/node1 -> ../../node/node1 lrwxrwxrwx    1 root     root             0 Apr  1 13:57 /sys/devices/system/cpu/cpu2/node0 -> ../../node/node0 lrwxrwxrwx    1 root     root             0 Apr  1 13:57 /sys/devices/system/cpu/cpu3/node1 -> ../../node/node1 # cpu to socket (not interleaved): # lscpu -e CPU NODE SOCKET CORE L1d:L1i:L2:L3 ONLINE   0    0      0    0 0:0:0:0          yes   1    0      0    1 1:1:1:0          yes   2    0      1    2 2:2:2:1          yes   3    0      1    3 3:3:3:1          yes # The NODE output of "lscpu -e" is kind of bogus, as it doesn't look up the cpu to node mapping, but instead reads the node to cpu map and figures out the reverse, which caused me to make a lot of false assumptions earlier... However, this is an unrelated matter and the SOCKET is correct. The early initialization code first assigns nodes in round-robin fashion, which is later overridden on AMD by srat_detect_node, but persists for Intel. > > Which resembles Intel baremetal node assignments where the CPUs > are interleaved. After your patch, it does: > > CPU#0 -> Node#0 > CPU#1 -> Node#0 > CPU#2 -> Node#0 > CPU#3 -> Node#0 This happens because we take an interesting path in srat_detect_node when the patch is applied. The cpuid to llc_id map looks like this: 0 -> 0 1 -> 0 2 -> 2 3 -> 2 Since node 2 does not exist, we enter the if(!node_online(node)) path and our final mapping ends up like this: 0 -> 0 1 -> 0 2 -> 0 3 -> 0 > So despite there being 2 LLCs, the Node assignments all go to Node#0 > which may have other unintended consequences. > > The statement "falling back to LLC when SRAT is not detected." isn't > accurate right? We have 2 LLCs and 2 Nodes but topology bits associate > both LLCs to the same node. > > I'm all for unifying the AMD and Hygon's srat_detect_node() but unifying > all three for an obviously broken use-case isn't a good motivation. > > I'll let others comment since they are more familiar with the NUMA > emulation bits and maybe all this is acceptable. > Hi Prateek, Thank you for the feedback! I have tried to dig a bit deeper and left my findings in response. I suppose I will wait to see if someone chimes in with some more details, if not I will do as you suggest and try to come up with a good way to unify amd/hygon and leave intel as is.