From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id ACACD2C21DF; Fri, 2 Oct 2026 12:26:39 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790944001; cv=none; b=e4gC30egtaEi2g3MYdzbzPhhH2+p/Lyu/9WsIGygpr2aE9O1CNqBlFkSZF/tbNpQIYyedJlIihmuVfeD5B8hovWwk42mR9s2OugjsiOjRNH27nkUOYL5W13AA7iJTjCoD160SNFrE/bjywfQCxXI6K5F+o6G1uwQ7c4T8RaedNU= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790944001; c=relaxed/simple; bh=zCerN7gWVzPgDqXrohsaZl5dB9fo6RbehtnbPNz3A/g=; h=Date:Message-ID:From:To:Cc:Subject:In-Reply-To:References: MIME-Version:Content-Type; b=UfhTFJy2rOrHryJdCDeC+AsL/tXlKPGbIdEqF3l1aSJALwVF7nZHdweJlgU7hXKJB8cOxs8NCPrrhHgqryqX3kcaRzTJf/D5qaTXVPuic2q54AAMAm4075tjFvsP3+74q8rO7rIBUW91IL0RqYVXXDbv70DF3OCtHu+HbC6k31s= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=j7eGrxGo; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="j7eGrxGo" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 147F61F00893; Fri, 2 Oct 2026 12:26:39 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1790943999; bh=CLh+ENzud1W7jlMwtzalp6l/ROxXZyI85Y4rCwOcmlM=; h=Date:From:To:Cc:Subject:In-Reply-To:References; b=j7eGrxGoMElYV4W9Lz4fI2ZW4MqRmu+5mNpqqKCzc/Dby0H12nTiNsAm0tuZic02q jpLGsjnOS0AQPYWo1JSS4pR6QnP7JKldDGLWQyG278rGB7vghxFUIUdxJ62oYuVAjo 1Fc5ozTZQg91N6QyeQtboasYsYG7gPi+oRhH8eYE7dXmmf5T3vd+qMGU01fe+//6y6 aIzpCSzgXlaFJWsLankN9iqkW9sRcRYLWcJNizGYfQVaP/hp2R01LuP0ZrqN4sDQfi +8UbBAP+Q/XamlwFJ2fnhXthNVpCcwgsnqZoVN2dwaBXtrqwlrkZZRYPlvKIQ2x5GB 8u8p8cTo/Q0ug== Received: from sofa.misterjones.org ([185.219.108.64] helo=goblin-girl.misterjones.org) by disco-boy.misterjones.org with esmtpsa (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.98.2) (envelope-from ) id 1xCcLU-0000000GDNO-42kB; Fri, 02 Oct 2026 12:26:37 +0000 Date: Fri, 02 Oct 2026 13:26:36 +0100 Message-ID: <86tsn42sqr.wl-maz@kernel.org> From: Marc Zyngier To: Tomohiro Misono Cc: Catalin Marinas , Will Deacon , Mark Rutland , Jonathan Corbet , Shuah Khan , Randy Dunlap , Thomas Gleixner , Radu Rendec , Kohei Enju , linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-perf-users@vger.kernel.org Subject: Re: [PATCH 4/4] irqchip/gicv3: Add workaround for FUJITSU-MONAKA erratum E#030003 In-Reply-To: <20261002-monaka-fix-for-upstream-v1-4-4aec0b0cbe34@fujitsu.com> References: <20261002-monaka-fix-for-upstream-v1-0-4aec0b0cbe34@fujitsu.com> <20261002-monaka-fix-for-upstream-v1-4-4aec0b0cbe34@fujitsu.com> User-Agent: Wanderlust/2.15.9 (Almost Unreal) SEMI-EPG/1.14.7 (Harue) FLIM-LB/1.14.9 (=?UTF-8?B?R29qxY0=?=) APEL-LB/10.8 EasyPG/1.0.0 Emacs/30.1 (aarch64-unknown-linux-gnu) MULE/6.0 (HANACHIRUSATO) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 (generated by SEMI-EPG 1.14.7 - "Harue") Content-Type: text/plain; charset=US-ASCII X-SA-Exim-Connect-IP: 185.219.108.64 X-SA-Exim-Rcpt-To: misono.tomohiro@fujitsu.com, catalin.marinas@arm.com, will@kernel.org, mark.rutland@arm.com, corbet@lwn.net, skhan@linuxfoundation.org, rdunlap@infradead.org, tglx@kernel.org, radu@rendec.net, enju.kohei@fujitsu.com, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-perf-users@vger.kernel.org X-SA-Exim-Mail-From: maz@kernel.org X-SA-Exim-Scanned: No (on disco-boy.misterjones.org); SAEximRunCond expanded to false On Fri, 02 Oct 2026 11:26:52 +0100, Tomohiro Misono wrote: > > From: Kohei Enju > > On affected FUJITSU-MONAKA CPUs, an SGI generated by writing to > ICC_SGI0R_EL1, ICC_SGI1R_EL1, or ICC_ASGI1R_EL1 may be lost if the > operation races with CPU interface processing triggered by the arrival > of a higher-priority interrupt, an update to a pending interrupt, or a > transition of the PE to the Sleep state. When this occurs, the system > register write does not complete, causing the issuing core to hang. > Is the SGI lost? Or is the sender core hanging? I expect that virtualised accesses to these registers are not affected, since they trap, but it'd be good to document this. > Work around the erratum by writing 1 to the bit corresponding to the > target SGI INTID in GICR_ISPENDR0 of each target PE's GIC Redistributor, > instead of writing to the affected system registers. > > The affinity topology of affected systems has Aff0 == 0 for every PE, > with PEs distinguished by higher affinity levels. Consequently, an > ICC_SGI1R_EL1 write can target only one PE, so the workaround does not > increase the number of writes required to send an SGI to multiple PEs. > > Accessing a target PE's GICR_ISPENDR0 requires its Redistributor base to > have been discovered. Affected platforms conform to SBBR, which requires > PSCI for secondary CPU boot. They therefore do not use the ACPI Parking > protocol, which sends an IPI before the secondary CPU has initialized > its Redistributor. With PSCI, a secondary CPU discovers its > Redistributor before becoming an IPI target. > > Signed-off-by: Kohei Enju > --- > Documentation/arch/arm64/silicon-errata.rst | 2 ++ > drivers/irqchip/irq-gic-v3.c | 48 +++++++++++++++++++++++++++++ > 2 files changed, 50 insertions(+) > > diff --git a/Documentation/arch/arm64/silicon-errata.rst b/Documentation/arch/arm64/silicon-errata.rst > index 429bdf9a3d9d..13328a30df59 100644 > --- a/Documentation/arch/arm64/silicon-errata.rst > +++ b/Documentation/arch/arm64/silicon-errata.rst > @@ -377,6 +377,8 @@ stable kernels. > +----------------+-----------------+-----------------+-----------------------------+ > | Fujitsu | MONAKA | E#030002 | FUJITSU_ERRATUM_030002 | > +----------------+-----------------+-----------------+-----------------------------+ > +| Fujitsu | MONAKA GICv3/v4 | E#030003 | N/A | > ++----------------+-----------------+-----------------+-----------------------------+ > +----------------+-----------------+-----------------+-----------------------------+ > | ASR | ASR8601 | #8601001 | N/A | > +----------------+-----------------+-----------------+-----------------------------+ > diff --git a/drivers/irqchip/irq-gic-v3.c b/drivers/irqchip/irq-gic-v3.c > index 6e1fa5b247fc..e73fea8a0e27 100644 > --- a/drivers/irqchip/irq-gic-v3.c > +++ b/drivers/irqchip/irq-gic-v3.c > @@ -80,6 +80,8 @@ static DEFINE_STATIC_KEY_FALSE(gic_nvidia_t241_erratum); > > static DEFINE_STATIC_KEY_FALSE(gic_arm64_2941627_erratum); > > +static DEFINE_STATIC_KEY_FALSE(gic_fujitsu_030003_erratum); > + > static struct gic_chip_data gic_data __read_mostly; > static DEFINE_STATIC_KEY_TRUE(supports_deactivate_key); > > @@ -238,6 +240,9 @@ static DEFINE_PER_CPU(bool, has_rss); > #define gic_data_rdist() (this_cpu_ptr(gic_data.rdists.rdist)) > #define gic_data_rdist_rd_base() (gic_data_rdist()->rd_base) > #define gic_data_rdist_sgi_base() (gic_data_rdist_rd_base() + SZ_64K) > +#define gic_data_rdist_cpu(cpu) (per_cpu_ptr(gic_data.rdists.rdist, cpu)) > +#define gic_data_rdist_rd_base_cpu(cpu) (gic_data_rdist_cpu(cpu)->rd_base) > +#define gic_data_rdist_sgi_base_cpu(cpu) (gic_data_rdist_rd_base_cpu(cpu) + SZ_64K) > > /* Our default, arbitrary priority value. Linux only uses one anyway. */ > #define DEFAULT_PMR_VALUE 0xf0 > @@ -1373,6 +1378,13 @@ static void gic_send_sgi(u64 cluster_id, u16 tlist, unsigned int irq) > gic_write_sgi1r(val); > } > > +static void gic_send_sgi_via_rdist(int cpu, unsigned int irq) > +{ > + void __iomem *base = gic_data_rdist_sgi_base_cpu(cpu); > + > + writel_relaxed(BIT(irq), base + GICR_ISPENDR0); > +} > + I don't think there is any need for a helper, given that there is a single caller. > static void gic_ipi_send_mask(struct irq_data *d, const struct cpumask *mask) > { > int cpu; > @@ -1386,6 +1398,22 @@ static void gic_ipi_send_mask(struct irq_data *d, const struct cpumask *mask) > */ > dsb(ishst); > > + if (static_branch_unlikely(&gic_fujitsu_030003_erratum)) { > + /* > + * The affinity topology of affected systems has Aff0 == 0 for > + * every PE; PEs are distinguished by higher affinity levels. > + * ICC_SGI1R_EL1 therefore targets only one PE per write, so > + * using GICR_ISPENDR0 does not increase the number of writes > + * required to send an SGI to multiple PEs. This is not about the number of writes, but the cost of individual writes. A sysreg access is almost free (at least it is on decent implementations), while an MMIO access is probably one of the worst offenders. Therefore implying that there is no extra overhead is likely to be misleading. The commit message has the same problem. In any case, I don't think you need to justify anything here, as the choice between costly SGIs and a dead CPU is pretty moot. > + */ > + for_each_cpu(cpu, mask) > + gic_send_sgi_via_rdist(cpu, d->hwirq); > + > + /* Force the above writes to GICR_ISPENDR0 to be executed */ > + dsb(st); This doesn't force things to be executed. This is about completion of the access, and with an nGnRE mapping, it doesn't enforce that the stores actually reach the RDs, only an arbitrary point in the memory subsystem. The only way to guarantee this is to perform a read-back. If this is relying on some additional properties that are implementation specific, then this requires to be documented (for example, if the implementation treats nGnRE as nGnRnE). > + return; > + } > + > for_each_cpu(cpu, mask) { > u64 cluster_id = MPIDR_TO_SGI_CLUSTER_ID(gic_cpu_to_affinity(cpu)); > u16 tlist; > @@ -1871,6 +1899,20 @@ static bool gic_enable_quirk_rk3399(void *data) > return false; > } > > +#define SMCCC_SOC_ID_FUJITSU_MONAKA 0x00040003 > + > +static bool gic_enable_quirk_fujitsu_030003(void *data) > +{ > + s32 soc_id = arm_smccc_get_soc_id_version(); > + > + /* Check JEP106 code for FUJITSU-MONAKA chip (0004:0003) */ > + if (soc_id != SMCCC_SOC_ID_FUJITSU_MONAKA) > + return false; Why is this keyed on some firmware interface, while it is the CPU interface that is at fault? I'd expect that looking at the MIDR would be more reliable. > + > + static_branch_enable(&gic_fujitsu_030003_erratum); > + return true; > +} > + > static bool rd_set_non_coherent(void *data) > { > struct gic_chip_data *d = data; > @@ -1951,6 +1993,12 @@ static const struct gic_quirk gic_quirks[] = { > .mask = 0xff000fff, > .init = gic_enable_quirk_rk3399, > }, > + { > + .desc = "GICv3: Fujitsu erratum 030003", > + .iidr = 0x0403043b, > + .mask = 0xffffffff, > + .init = gic_enable_quirk_fujitsu_030003, Same problem. This is looking that the distributor instead of the CPU. It's OK to use it as a proxy for further filtering, but the final decision should probably rest on the MIDR. Thanks, M. -- Without deviation from the norm, progress is not possible.