From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-yw1-f172.google.com (mail-yw1-f172.google.com [209.85.128.172]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E63CF4279FE for ; Wed, 29 Jul 2026 08:34:25 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.172 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785314071; cv=none; b=SqT0Chmo92MGw1IC3/RmoXl3ircjis4ceeugB3I57mY0mnTBpaHmmePNse9KA9LV7z3HO8wKsvVA3jMdDp4A18M8GJ59AKUIXqxoynA2Hsv76xbLxKFyAnnZj5f3EvLMubzkNEjB9vmUlto3nuGv3SP+hqFzxw2sJVRsp1lzx78= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785314071; c=relaxed/simple; bh=kR/9sK9mcB7hctDuy/2H8HtGbmrt9byIGI+UPwtztOA=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=OrNgMWSpmI7GHhsB6pDHjhdHiS7Te2VDrKvIN8dmEG8RzUZgmDtPEsPl8ysq48oshRwlDhtmZQX/v2mBh1CSU3eGKeeGQ6cQFtJIalFchcydH4gFr8k9O+1E7Ah288DuGxcwcmPe1aR7HsZL+fJf0CYHa/3+6BnDpR2A6jxDrfQ= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=D3TaQ2K6; arc=none smtp.client-ip=209.85.128.172 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="D3TaQ2K6" Received: by mail-yw1-f172.google.com with SMTP id 00721157ae682-81f64e8dfbcso11171927b3.2 for ; Wed, 29 Jul 2026 01:34:25 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1785314062; x=1785918862; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=WCfZWs2qBzfBTMw3/b4qNRHgwpbTts4mgogpiMUra6w=; b=D3TaQ2K6IffDJ7kXXEtpopMtUGZk0E3sA44jl3URVZ+HVNd7oM6CanxJAhXYk9NC86 nuvqDBNAr3E1cc9uzf+B3XgGP6VU3CsJDKnI+/TC9RRaeDbW6i4LpkAMIF/VTLT2Xi0K sdAURKNJ+W+aMdwJVrmnMeCAagttFtUw+1hkRp2TFR6gm5IXMePMuMM4fO1BRah0CvVu B/YfyTq3483+8ABMxmIbOWiHeZN/42cAlTGJiUDKjW+2v9v3A/mAWI2LmJtYDeP26wCh gDpIz1hPRjTQH5K6K1euOMDlahK3HIzqh9n9RJmmzlgsjmhbQnok1sLzcU6xjhLZOsbS vHbg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785314062; x=1785918862; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=WCfZWs2qBzfBTMw3/b4qNRHgwpbTts4mgogpiMUra6w=; b=MgdbV1+RMiogMi8nv+355scoaQC10sF6kD1LVC79YvxYNypSjHmwEMXamJLd1eE4V+ CL9vJz/CXbBwJ/onPqCPHzFMOZLpw+BGzTVAn8xMCx2ZyikToktZ4umig99AxzeIGoIs IW2O90MBtWKOjfEGvGUxISJpd/UIM3mfwM6DX+x9buGUdTsoPbIR7l2Q1YzG9K2RVegF xAdDKOkBanJi4pnyOU7SI6EghvfyEIa6A+gYOSM+WMUkofjxV4T3WhLRzhFMHhWzxRNF +jzY/vrK29qMKsQjn/tF1+L9VzsYPaNpgh9+MLnKoJ8hQLOAoUQ5Q5KgFjmVfAUURZ9p y1mw== X-Forwarded-Encrypted: i=1; AHgh+RqZBwPBQM5YnaJP4debJniW96Dp79on0TC1mPegJ2D1gZhlPN3sqq/QXcBBBZQb+xujWK2XEPEa+Hjgmy4=@vger.kernel.org X-Gm-Message-State: AOJu0YxogdSoX1/JnA1zZgo09o2UAFOjw0vV/3JxvhBcJDeYyyj4RqoC ehY5go0LkuT5ajXB17HVBEQSIAJIZl/OBoDUcGP7snZK7jV4bYpc8ksi X-Gm-Gg: AR+sD12zCOYqf3y86rnSt+mgS2DLMwWY8LLZBkQrzXjHWZZKL8LN6mC6Dfkzz8sb80j 2YJ/BfUMnxzoGnolx8q7jl5wcZ79TO9nRV5J3QuTz8pt9ZIlg76lnJvn2CGGMWEakDnfVnRUQOk JFg9DRecPBJd5lwL6Rnocogl1D957KZpeM4LZXYaebz1jq+BbOWKwRroFBM9B6xT5+PIz/HHXkr 16TeFy82XmF3Na2uZQLSL6S7poOhNkV7XYaK6V0EPFAFPINAq63ppOL5OPmYSKYEbgCNYbS9nxI oLI6jIUNAPZK9T2jXOR9s8xDpO/xvkFGqA3mgCfu2nVW0RyDNovyG2J+vF7QnOgxEaIaTUtC6NA GZj+p1hj6z9IPVw3R1N95zj/4OjYC02wQR+4gGju6HgzhcCpeRIrR+dW80IbULAFvK/H9CBocbU H05gPpYKIztmFqw0tvKaeHUE18euA+Ep/zz0jJe8Xo8m2sN9W5SK6W9swTKP1KBUtnxQjvkunNc gkaeV0qtGIxshGazfhik3SEsgj20uZHlhd4SRYMAA== X-Received: by 2002:a05:690c:3685:b0:81e:9f17:8050 with SMTP id 00721157ae682-81f99240917mr28751857b3.21.1785314062516; Wed, 29 Jul 2026 01:34:22 -0700 (PDT) Received: from localhost.localdomain ([107.198.84.185]) by smtp.gmail.com with ESMTPSA id 00721157ae682-81fa2957751sm14029097b3.35.2026.07.29.01.34.20 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 29 Jul 2026 01:34:21 -0700 (PDT) From: Paul Sherman To: palmer@dabbelt.com, pjw@kernel.org, aou@eecs.berkeley.edu Cc: alex@ghiti.fr, linux-riscv@lists.infradead.org, linux-kernel@vger.kernel.org, Paul Sherman Subject: [PATCH] riscv: smp: fix non-SPINWAIT secondary hart rendezvous for fw_dynamic platforms Date: Wed, 29 Jul 2026 01:34:14 -0700 Message-ID: <20260729083414.39339-1-shermanpauldylan@gmail.com> X-Mailer: git-send-email 2.53.0 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit On platforms where firmware (e.g. fw_dynamic) releases all harts to the same Linux entry point simultaneously, the hart selected by OpenSBI as Domain0 Boot HART is already in SBI_HSM_STATE_STARTED when Linux later attempts to bring it online as a secondary CPU via SBI HART_START. OpenSBI correctly returns SBI_ERR_ALREADY_STARTED, but Linux has no recovery path: with CONFIG_RISCV_BOOT_SPINWAIT=n, there is no secondary wait path in _start_kernel for harts that entered Linux directly from firmware, and sbi_cpu_start() has no handler for SBI_ERR_ALREADY_STARTED. This causes one CPU to be permanently dropped per boot. The missing CPU is always the OpenSBI Domain0 Boot HART, which varies between boots on Sophgo SG2042 (hart 1, 2, or 3), explaining the apparent 'moving victim'. Fix with three cooperating changes: 1. Initialize boot_cpu_hartid to INVALID_HARTID instead of relying on BSS zero-initialization. Without this, boot_cpu_hartid aliases with hart 0, causing hart 0 to always appear to win the boot CPU race regardless of which hart actually stored its hartid first. 2. Add a non-SPINWAIT secondary wait path in _start_kernel. When firmware releases multiple harts to the same entry point, non-primary harts divert into the existing spinwait rendezvous arrays (previously used only by CONFIG_RISCV_BOOT_SPINWAIT) and wait for cpu_start() to provide boot data before proceeding to secondary startup. 3. Handle SBI_ERR_ALREADY_STARTED (-EALREADY) in sbi_cpu_start(). When HART_START returns -EALREADY, the hart is already executing in Linux and spinning in .Lwait_for_cpu_up_sbi. Write the spinwait rendezvous arrays to release it into secondary startup, matching the approach used by cpu_ops_spinwait.c. The arrays __cpu_spinwait_stack_pointer and __cpu_spinwait_task_pointer are defined unconditionally in cpu_ops_spinwait.c but their extern declarations in head.h were guarded by CONFIG_RISCV_BOOT_SPINWAIT. Move the declarations outside the guard since the arrays are always present and now used by both boot paths. Note: The NR_CPUS bound check mirrors the identical pattern in cpu_ops_spinwait.c:32 which guards the same arrays against out-of-range hartids on platforms with discontiguous hart numbering. Link: https://lore.kernel.org/linux-riscv/20260727221508.5179-1-shermanpauldylan@gmail.com/ Tested-on: Milk-V Pioneer (Sophgo SG2042, 64-hart RISC-V, 4-NUMA nodes, 128GB DDR4, OpenSBI v1.5, Linux v7.2-rc5) Result: boot_cpu_hartid correctly reflects Domain0 Boot HART, all 64 CPUs online in 2.7 seconds (was 63 CPUs, boot always on hart 0) Signed-off-by: Paul Sherman --- arch/riscv/kernel/cpu_ops_sbi.c | 41 ++++++++++++++++++++++++++++++++- arch/riscv/kernel/head.S | 25 ++++++++++++++++++++ arch/riscv/kernel/head.h | 2 -- arch/riscv/kernel/setup.c | 9 +++++++- 4 files changed, 73 insertions(+), 4 deletions(-) diff --git a/arch/riscv/kernel/cpu_ops_sbi.c b/arch/riscv/kernel/cpu_ops_sbi.c index ee6e4b5cc39e9..3268fda5be35f 100644 --- a/arch/riscv/kernel/cpu_ops_sbi.c +++ b/arch/riscv/kernel/cpu_ops_sbi.c @@ -12,6 +12,7 @@ #include #include #include +#include "head.h" extern char secondary_start_sbi[]; const struct cpu_operations cpu_ops_sbi; @@ -23,6 +24,18 @@ const struct cpu_operations cpu_ops_sbi; */ static struct sbi_hart_boot_data boot_data[NR_CPUS]; +#ifndef CONFIG_RISCV_BOOT_SPINWAIT +/* + * Secondary hart rendezvous arrays, shared with head.S. + * These arrays are named for historical reasons after the spinwait + * boot protocol, but serve a generic purpose: holding per-hart boot + * data until a secondary hart is ready to proceed. Defined here when + * CONFIG_RISCV_BOOT_SPINWAIT=n; otherwise defined in cpu_ops_spinwait.c. + */ +void *__cpu_spinwait_stack_pointer[NR_CPUS] __section(".data"); +void *__cpu_spinwait_task_pointer[NR_CPUS] __section(".data"); +#endif + static int sbi_hsm_hart_start(unsigned long hartid, unsigned long saddr, unsigned long priv) { @@ -62,6 +75,7 @@ static int sbi_cpu_start(unsigned int cpuid, struct task_struct *tidle) unsigned long boot_addr = __pa_symbol(secondary_start_sbi); unsigned long hartid = cpuid_to_hartid_map(cpuid); unsigned long hsm_data; + int ret; struct sbi_hart_boot_data *bdata = &boot_data[cpuid]; /* Make sure tidle is updated */ @@ -71,7 +85,32 @@ static int sbi_cpu_start(unsigned int cpuid, struct task_struct *tidle) /* Make sure boot data is updated */ smp_mb(); hsm_data = __pa(bdata); - return sbi_hsm_hart_start(hartid, boot_addr, hsm_data); + + ret = sbi_hsm_hart_start(hartid, boot_addr, hsm_data); + + /* + * The firmware boot hart enters Linux directly from the bootloader + * and is already in SBI_HSM_STATE_STARTED when Linux attempts to + * bring it online as a secondary CPU. HART_START correctly returns + * SBI_ERR_ALREADY_STARTED in this case. The hart is spinning in + * .Lwait_for_cpu_up_sbi waiting for boot data - write the spinwait + * rendezvous arrays to release it into secondary startup. + * + * Guard against invalid or out-of-range hartids, matching the + * same constraint enforced in cpu_ops_spinwait.c. + */ + if (ret == -EALREADY) { + if (hartid != INVALID_HARTID && + hartid < (unsigned long)NR_CPUS) { /* array bound */ + /* Ensure bdata writes visible before spinwait arrays */ + smp_wmb(); + WRITE_ONCE(__cpu_spinwait_stack_pointer[hartid], + task_pt_regs(tidle)); + WRITE_ONCE(__cpu_spinwait_task_pointer[hartid], tidle); + } + ret = 0; + } + return ret; } #ifdef CONFIG_HOTPLUG_CPU diff --git a/arch/riscv/kernel/head.S b/arch/riscv/kernel/head.S index f6a8ca49e6277..da48f5efd928b 100644 --- a/arch/riscv/kernel/head.S +++ b/arch/riscv/kernel/head.S @@ -281,6 +281,31 @@ SYM_CODE_START(_start_kernel) la a2, boot_cpu_hartid REG_S a0, (a2) +#ifndef CONFIG_RISCV_BOOT_SPINWAIT + /* + * On platforms where firmware releases all harts to the same + * entry point (e.g. fw_dynamic), non-primary harts must divert + * here before MMU setup. Wait in the spinwait rendezvous arrays + * until cpu_start() provides boot data. a0 = hartid. + */ + REG_L a3, (a2) + beq a0, a3, .Lprimary_hart + slli a3, a0, LGREG + la a1, __cpu_spinwait_stack_pointer + la a2, __cpu_spinwait_task_pointer + add a1, a3, a1 + add a2, a3, a2 +.Lwait_for_cpu_up_sbi: + fence r, r + REG_L sp, (a1) + REG_L tp, (a2) + beqz sp, .Lwait_for_cpu_up_sbi + beqz tp, .Lwait_for_cpu_up_sbi + fence + tail .Lsecondary_start_common +.Lprimary_hart: +#endif /* !CONFIG_RISCV_BOOT_SPINWAIT */ + /* Initialize page tables and relocate to virtual addresses */ la tp, init_task la sp, init_thread_union + THREAD_SIZE diff --git a/arch/riscv/kernel/head.h b/arch/riscv/kernel/head.h index 05a04bef442b1..1b34f8e655581 100644 --- a/arch/riscv/kernel/head.h +++ b/arch/riscv/kernel/head.h @@ -12,9 +12,7 @@ extern atomic_t hart_lottery; asmlinkage void __init setup_vm(uintptr_t dtb_pa); -#ifdef CONFIG_RISCV_BOOT_SPINWAIT extern void *__cpu_spinwait_stack_pointer[]; extern void *__cpu_spinwait_task_pointer[]; -#endif #endif /* __ASM_HEAD_H */ diff --git a/arch/riscv/kernel/setup.c b/arch/riscv/kernel/setup.c index 52d1d2b8f338b..d2eecb7ad4965 100644 --- a/arch/riscv/kernel/setup.c +++ b/arch/riscv/kernel/setup.c @@ -47,7 +47,14 @@ * BSS. */ atomic_t hart_lottery __section(".sdata"); -unsigned long boot_cpu_hartid; +/* + * Initialize to INVALID_HARTID so that the first hart to store its + * hartid in head.S is unambiguously the boot CPU. Without this, + * boot_cpu_hartid starts as 0 (BSS), which aliases with hart 0 and + * causes hart 0 to always appear to win the boot CPU race regardless + * of which hart actually wrote first. + */ +unsigned long boot_cpu_hartid = INVALID_HARTID; EXPORT_SYMBOL_GPL(boot_cpu_hartid); /* -- 2.53.0