From: salil.mehta@opnsrc.net
To: Thomas Gleixner <tglx@kernel.org>,
Catalin Marinas <catalin.marinas@arm.com>,
Will Deacon <will@kernel.org>
Cc: Jonathan Cameron <jic23@kernel.org>,
James Morse <james.morse@arm.com>,
Peter Zijlstra <peterz@infradead.org>,
Mark Rutland <mark.rutland@arm.com>,
Jinjie Ruan <ruanjinjie@huawei.com>,
Yicong Yang <yangyicong@hisilicon.com>,
Jonathan Corbet <corbet@lwn.net>, Gavin Shan <gshan@redhat.com>,
linux-kernel@vger.kernel.org,
linux-arm-kernel@lists.infradead.org, linux-doc@vger.kernel.org,
Salil Mehta <salil.mehta@opnsrc.net>
Subject: [RFC PATCH 0/2] cpu/hotplug: use cpu_enabled_mask for SMT bringup
Date: Tue, 29 Sep 2026 19:41:28 +0000 [thread overview]
Message-ID: <cover.1790692295.git.salil.mehta@opnsrc.net> (raw)
From: Salil Mehta <salil.mehta@opnsrc.net>
Hi,
For context, I had been away from the QEMU/kernel mailing-list work for an
extended period due to exceptional personal circumstances. I have only
recently started catching up with the Arm vCPU hotplug work and the related
changes that have landed upstream. While doing so, I came across the change
discussed below.
============
I. The Story
============
This is an RFC and, in particular, an open question about the interaction
between cpu_present_mask, cpu_enabled_mask and cpuhp_smt_enable(). I may have
missed a later constraint or discussion while catching up, so I would very
much appreciate a sanity check on the direction proposed here.
Commit f9a82544c717 ("cpu/hotplug: Fix NULL kobject warning in
cpuhp_smt_enable()") fixed a real warning on arm64 by changing the arm64
present-mask semantics: an ACPI Online-Capable CPU which is not MADT
Enabled is no longer initially marked present, and acpi_map_cpu() /
acpi_unmap_cpu() now add and remove it from cpu_present_mask.
I wondered whether the generic enabled mask gives us a narrower way to fix
the original problem.
cpu_enabled_mask was introduced by 4e1a7df45480 ("cpumask: Add enabled
cpumask for present CPUs that can be brought online") specifically for the
case where a CPU can be present but is not currently allowed to be brought
online. cpuhp_smt_enable() is itself trying to bring offline CPUs online,
so patch 1 simply skips a CPU when !cpu_enabled(cpu).
If that is the intended meaning of cpu_enabled_mask, should
cpuhp_smt_enable() consume it rather than require arm64 to collapse the
present/not-enabled distinction? Or is there another reason why changing
the arm64 present-mask semantics is preferred here? I may be missing a
constraint in another architecture or in the SMT hotplug path, hence this
RFC.
Patch 2 reverts f9a82544c717 so that the question can be evaluated with the
original arm64 virtual CPU hotplug model restored. The ordering is
intentional: the generic enabled check lands first, so reverting the arm64
present-mask change does not reintroduce the NULL-kobject warning.
===========
II. Testing
===========
The series is based on current upstream master, after Linux v7.3-rc5.
An arm64 Image build with CONFIG_ACPI=y, CONFIG_HOTPLUG_CPU=y and
CONFIG_HOTPLUG_SMT=y passes.
The runtime tests were performed on an NVIDIA Jetson Orin Nano system.
The host does not provide hardware SMT; QEMU's virtual SMT topology was
used to exercise the guest HOTPLUG_SMT paths described below.
===============
III. Reproducer
===============
The warning can be reproduced with an arm64 ACPI guest using a virtual SMT
topology, for example:
qemu-system-aarch64 \
-machine virt,gic-version=3,acpi=on \
-accel tcg,thread=multi -cpu cortex-a57 \
-smp cpus=3,maxcpus=6,sockets=1,clusters=1,cores=3,threads=2 \
-m 1024M \
-bios QEMU_EFI.fd \
-kernel Image \
-initrd rootfs.cpio.gz \
-append "console=ttyAMA0 root=/dev/ram rdinit=/init acpi=force" \
-nographic
After boot:
cat /sys/devices/system/cpu/present
cat /sys/devices/system/cpu/enabled
cat /sys/devices/system/cpu/online
With the pre-f9a arm64 semantics these report:
present=0-5
enabled=0-2
online=0-2
Then exercise SMT disable/enable:
echo off > /sys/devices/system/cpu/smt/control
cat /sys/devices/system/cpu/online
echo on > /sys/devices/system/cpu/smt/control
cat /sys/devices/system/cpu/online
With the original cpuhp_smt_enable() implementation, the second write
produces the following warning:
WARNING: fs/sysfs/group.c:137 at internal_create_group+0x418/0x550,
CPU#2: bash/1
with the relevant call trace:
internal_create_group+0x418/0x550
sysfs_create_group+0x20/0x38
topology_add_dev+0x24/0x38
cpuhp_invoke_callback+0x174/0x2c0
__cpuhp_invoke_callback_range+0x98/0x128
_cpu_up+0x150/0x280
cpuhp_smt_enable+0xc4/0x128
control_store+0xf8/0x1f0
With patch 1 followed by patch 2, the same sequence gives:
SMT off: online=0,2
SMT on: online=0-2
and:
dmesg | grep -E "WARNING:|internal_create_group|cpuhp_smt_enable"
produces no output.
=========================
IV. Additional Comparison
=========================
I also checked cpus=4,maxcpus=6: all four requested initial CPUs are brought
up with the restored present semantics. By comparison, applying the
f9a82544c717 changes in isolation to the same firmware model resulted in
only the boot CPU being brought up during early SMP initialization, with
the remaining enabled CPUs appearing later through ACPI enumeration.
Both patches pass scripts/checkpatch.pl --strict with no warnings or
errors.
I also have an up-to-date QEMU branch containing the corresponding Arm vCPU
hotplug work used for the tests above. If that would be useful for reproducing
or testing this RFC, please shout and I can share the branch.
=================
V. 'The' Question
=================
The main question for review is therefore whether cpu_enabled_mask should be
the generic eligibility check in cpuhp_smt_enable(), preserving the
present/enabled distinction that motivated the mask, or whether there is a
reason that distinction should instead be removed from arm64.
Many thanks in anticipation!
Salil
Salil Mehta (2):
cpu/hotplug: Skip disabled CPUs in cpuhp_smt_enable
Revert "cpu/hotplug: Fix NULL kobject warning in cpuhp_smt_enable()"
Documentation/arch/arm64/cpu-hotplug.rst | 28 ++++++++++--------------
arch/arm64/kernel/acpi.c | 2 --
arch/arm64/kernel/smp.c | 12 +---------
kernel/cpu.c | 5 +++--
4 files changed, 16 insertions(+), 31 deletions(-)
--
2.34.1
next reply other threads:[~2026-09-29 19:41 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-29 19:41 salil.mehta [this message]
2026-09-29 19:41 ` [RFC PATCH 1/2] cpu/hotplug: Skip disabled CPUs in cpuhp_smt_enable salil.mehta
2026-09-29 20:16 ` Bradley Morgan
2026-09-30 8:42 ` Salil Mehta
2026-09-29 19:41 ` [RFC PATCH 2/2] Revert "cpu/hotplug: Fix NULL kobject warning in cpuhp_smt_enable()" salil.mehta
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=cover.1790692295.git.salil.mehta@opnsrc.net \
--to=salil.mehta@opnsrc.net \
--cc=catalin.marinas@arm.com \
--cc=corbet@lwn.net \
--cc=gshan@redhat.com \
--cc=james.morse@arm.com \
--cc=jic23@kernel.org \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=mark.rutland@arm.com \
--cc=peterz@infradead.org \
--cc=ruanjinjie@huawei.com \
--cc=tglx@kernel.org \
--cc=will@kernel.org \
--cc=yangyicong@hisilicon.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®