* [PATCH RFC 0/1] Support unknown NMI based on SSE
@ 2025-09-19 7:00 Yunhui Cui
2025-09-19 7:00 ` [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support Yunhui Cui
0 siblings, 1 reply; 11+ messages in thread
From: Yunhui Cui @ 2025-09-19 7:00 UTC (permalink / raw)
To: paul.walmsley, palmer, aou, alex, conor, atishp, cleger, ajones,
apatel, cuiyunhui, mchitale, linux-riscv, linux-kernel
This patch is based on the SSE features by Clément Léger:
https://lore.kernel.org/all/20250908181717.1997461-1-cleger@rivosinc.com/
It mainly adds support for unknown NMI functionality. Since extension
development has not yet been implemented in SBI, an RFC version is
submitted first.
Yunhui Cui (1):
drivers: firmware: riscv: add unknown NMI support
arch/riscv/include/asm/sbi.h | 1 +
drivers/firmware/riscv/Kconfig | 10 +++++
drivers/firmware/riscv/Makefile | 1 +
drivers/firmware/riscv/sse_nmi.c | 77 ++++++++++++++++++++++++++++++++
4 files changed, 89 insertions(+)
create mode 100644 drivers/firmware/riscv/sse_nmi.c
--
2.39.5
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support
2025-09-19 7:00 [PATCH RFC 0/1] Support unknown NMI based on SSE Yunhui Cui
@ 2025-09-19 7:00 ` Yunhui Cui
2025-09-19 7:17 ` Clément Léger
0 siblings, 1 reply; 11+ messages in thread
From: Yunhui Cui @ 2025-09-19 7:00 UTC (permalink / raw)
To: paul.walmsley, palmer, aou, alex, conor, atishp, cleger, ajones,
apatel, cuiyunhui, mchitale, linux-riscv, linux-kernel
Unknown NMI can force the kernel to respond (e.g., panic) when the
system encounters unrecognized critical hardware events, aiding in
troubleshooting system faults. This is implemented via the Supervisor
Software Events (SSE) framework.
Signed-off-by: Yunhui Cui <cuiyunhui@bytedance.com>
---
arch/riscv/include/asm/sbi.h | 1 +
drivers/firmware/riscv/Kconfig | 10 +++++
drivers/firmware/riscv/Makefile | 1 +
drivers/firmware/riscv/sse_nmi.c | 77 ++++++++++++++++++++++++++++++++
4 files changed, 89 insertions(+)
create mode 100644 drivers/firmware/riscv/sse_nmi.c
diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h
index 874cc1d7603a5..5801f90a88f62 100644
--- a/arch/riscv/include/asm/sbi.h
+++ b/arch/riscv/include/asm/sbi.h
@@ -481,6 +481,7 @@ enum sbi_sse_attr_id {
#define SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS 0x00000000
#define SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP 0x00000001
+#define SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI 0x00000002
#define SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS 0x00008000
#define SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW 0x00010000
#define SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS 0x00100000
diff --git a/drivers/firmware/riscv/Kconfig b/drivers/firmware/riscv/Kconfig
index ed5b663ac5f91..746bac862ac46 100644
--- a/drivers/firmware/riscv/Kconfig
+++ b/drivers/firmware/riscv/Kconfig
@@ -12,4 +12,14 @@ config RISCV_SBI_SSE
this option provides support to register callbacks on specific SSE
events.
+config RISCV_SSE_UNKNOWN_NMI
+ bool "Enable SBI Supervisor Software Events unknown NMI support"
+ depends on RISCV_SBI_SSE
+ default y
+ help
+ This option enables support for delivering unknown Non-Maskable Interrupt (NMI)
+ notifications via the Supervisor Software Events (SSE) framework. When enabled,
+ unknown NMIs can trigger kernel responses (e.g., panic) for unrecognized critical
+ hardware events, aiding in system fault diagnosis.
+
endmenu
diff --git a/drivers/firmware/riscv/Makefile b/drivers/firmware/riscv/Makefile
index c8795d4bbb2ea..9242c6cd5e3e9 100644
--- a/drivers/firmware/riscv/Makefile
+++ b/drivers/firmware/riscv/Makefile
@@ -1,3 +1,4 @@
# SPDX-License-Identifier: GPL-2.0
obj-$(CONFIG_RISCV_SBI_SSE) += riscv_sbi_sse.o
+obj-$(CONFIG_RISCV_SSE_UNKNOWN_NMI) += sse_nmi.o
diff --git a/drivers/firmware/riscv/sse_nmi.c b/drivers/firmware/riscv/sse_nmi.c
new file mode 100644
index 0000000000000..43063f42efff0
--- /dev/null
+++ b/drivers/firmware/riscv/sse_nmi.c
@@ -0,0 +1,77 @@
+// SPDX-License-Identifier: GPL-2.0+
+
+#include <linux/mm.h>
+#include <linux/nmi.h>
+#include <linux/riscv_sbi_sse.h>
+#include <linux/sched/debug.h>
+#include <linux/sysctl.h>
+
+#include <asm/irq_regs.h>
+#include <asm/sbi.h>
+
+int panic_on_unknown_nmi = 1;
+struct sse_event *evt;
+static struct ctl_table_header *unknown_nmi_sysctl_header;
+
+const struct ctl_table unknown_nmi_table[] = {
+ {
+ .procname = "panic_enable",
+ .data = &panic_on_unknown_nmi,
+ .maxlen = sizeof(int),
+ .mode = 0644,
+ .proc_handler = proc_dointvec_minmax,
+ .extra1 = SYSCTL_ZERO,
+ .extra2 = SYSCTL_ONE,
+ },
+};
+
+static void nmi_handler(struct pt_regs *regs)
+{
+ pr_emerg("NMI received for unknown on CPU %d.\n", smp_processor_id());
+
+ if (panic_on_unknown_nmi)
+ nmi_panic(regs, "NMI: Not continuing");
+
+ pr_emerg("Dazed and confused, but trying to continue\n");
+}
+
+static int nmi_sse_handler(u32 evt, void *arg, struct pt_regs *regs)
+{
+ nmi_handler(regs);
+
+ return 0;
+}
+
+static int sse_nmi_init(void)
+{
+ int ret;
+
+ evt = sse_event_register(SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI, 0,
+ nmi_sse_handler, NULL);
+ if (IS_ERR(evt))
+ return PTR_ERR(evt);
+
+ ret = sse_event_enable(evt);
+ if (ret) {
+ sse_event_unregister(evt);
+ return ret;
+ }
+
+ unknown_nmi_sysctl_header = register_sysctl("kernel", unknown_nmi_table);
+ if (!unknown_nmi_sysctl_header) {
+ sse_event_disable(evt);
+ sse_event_unregister(evt);
+ return -ENOMEM;
+ }
+
+ pr_info("Using SSE for NMI event delivery\n");
+
+ return 0;
+}
+
+static int __init unknow_nmi_init(void)
+{
+ return sse_nmi_init();
+}
+
+late_initcall(unknow_nmi_init);
--
2.39.5
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support
2025-09-19 7:00 ` [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support Yunhui Cui
@ 2025-09-19 7:17 ` Clément Léger
2025-09-19 7:52 ` [External] " yunhui cui
0 siblings, 1 reply; 11+ messages in thread
From: Clément Léger @ 2025-09-19 7:17 UTC (permalink / raw)
To: Yunhui Cui, paul.walmsley, palmer, aou, alex, conor, atishp,
ajones, apatel, mchitale, linux-riscv, linux-kernel
On 19/09/2025 09:00, Yunhui Cui wrote:
> Unknown NMI can force the kernel to respond (e.g., panic) when the
> system encounters unrecognized critical hardware events, aiding in
> troubleshooting system faults. This is implemented via the Supervisor
> Software Events (SSE) framework.
>
> Signed-off-by: Yunhui Cui <cuiyunhui@bytedance.com>
> ---
> arch/riscv/include/asm/sbi.h | 1 +
> drivers/firmware/riscv/Kconfig | 10 +++++
> drivers/firmware/riscv/Makefile | 1 +
> drivers/firmware/riscv/sse_nmi.c | 77 ++++++++++++++++++++++++++++++++
> 4 files changed, 89 insertions(+)
> create mode 100644 drivers/firmware/riscv/sse_nmi.c
>
> diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h
> index 874cc1d7603a5..5801f90a88f62 100644
> --- a/arch/riscv/include/asm/sbi.h
> +++ b/arch/riscv/include/asm/sbi.h
> @@ -481,6 +481,7 @@ enum sbi_sse_attr_id {
>
> #define SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS 0x00000000
> #define SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP 0x00000001
> +#define SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI 0x00000002
Was this submitted to the PRS WG ? This a specification modification so
it should go through the usual process.
> #define SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS 0x00008000
> #define SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW 0x00010000
> #define SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS 0x00100000
> diff --git a/drivers/firmware/riscv/Kconfig b/drivers/firmware/riscv/Kconfig
> index ed5b663ac5f91..746bac862ac46 100644
> --- a/drivers/firmware/riscv/Kconfig
> +++ b/drivers/firmware/riscv/Kconfig
> @@ -12,4 +12,14 @@ config RISCV_SBI_SSE
> this option provides support to register callbacks on specific SSE
> events.
>
> +config RISCV_SSE_UNKNOWN_NMI
> + bool "Enable SBI Supervisor Software Events unknown NMI support"
> + depends on RISCV_SBI_SSE
> + default y
> + help
> + This option enables support for delivering unknown Non-Maskable Interrupt (NMI)
> + notifications via the Supervisor Software Events (SSE) framework. When enabled,
> + unknown NMIs can trigger kernel responses (e.g., panic) for unrecognized critical
> + hardware events, aiding in system fault diagnosis.
> +
> endmenu
> diff --git a/drivers/firmware/riscv/Makefile b/drivers/firmware/riscv/Makefile
> index c8795d4bbb2ea..9242c6cd5e3e9 100644
> --- a/drivers/firmware/riscv/Makefile
> +++ b/drivers/firmware/riscv/Makefile
> @@ -1,3 +1,4 @@
> # SPDX-License-Identifier: GPL-2.0
>
> obj-$(CONFIG_RISCV_SBI_SSE) += riscv_sbi_sse.o
> +obj-$(CONFIG_RISCV_SSE_UNKNOWN_NMI) += sse_nmi.o
> diff --git a/drivers/firmware/riscv/sse_nmi.c b/drivers/firmware/riscv/sse_nmi.c
> new file mode 100644
> index 0000000000000..43063f42efff0
> --- /dev/null
> +++ b/drivers/firmware/riscv/sse_nmi.c
> @@ -0,0 +1,77 @@
> +// SPDX-License-Identifier: GPL-2.0+
> +
> +#include <linux/mm.h>
> +#include <linux/nmi.h>
> +#include <linux/riscv_sbi_sse.h>
> +#include <linux/sched/debug.h>
> +#include <linux/sysctl.h>
> +
> +#include <asm/irq_regs.h>
> +#include <asm/sbi.h>
> +
> +int panic_on_unknown_nmi = 1;
> +struct sse_event *evt;
> +static struct ctl_table_header *unknown_nmi_sysctl_header;
> +
> +const struct ctl_table unknown_nmi_table[] = {
> + {
> + .procname = "panic_enable",
> + .data = &panic_on_unknown_nmi,
> + .maxlen = sizeof(int),
> + .mode = 0644,
> + .proc_handler = proc_dointvec_minmax,
> + .extra1 = SYSCTL_ZERO,
> + .extra2 = SYSCTL_ONE,
> + },
> +};
> +
> +static void nmi_handler(struct pt_regs *regs)
> +{
> + pr_emerg("NMI received for unknown on CPU %d.\n", smp_processor_id());
> +
> + if (panic_on_unknown_nmi)
> + nmi_panic(regs, "NMI: Not continuing");
> +
> + pr_emerg("Dazed and confused, but trying to continue\n");
> +}
I'm dazed and confused as well ;) What's the point of this except
interrupting the kernel with a panic ? It seems like it's a better idea
to let the firmware handle that properly and display whatever
information are needed. Was your idea to actually force the kernel to
enter in some debug mode ?
Thanks,
Clément
> +
> +static int nmi_sse_handler(u32 evt, void *arg, struct pt_regs *regs)
> +{
> + nmi_handler(regs);
> +
> + return 0;
> +}
> +
> +static int sse_nmi_init(void)
> +{
> + int ret;
> +
> + evt = sse_event_register(SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI, 0,
> + nmi_sse_handler, NULL);
> + if (IS_ERR(evt))
> + return PTR_ERR(evt);
> +
> + ret = sse_event_enable(evt);
> + if (ret) {
> + sse_event_unregister(evt);
> + return ret;
> + }
> +
> + unknown_nmi_sysctl_header = register_sysctl("kernel", unknown_nmi_table);
> + if (!unknown_nmi_sysctl_header) {
> + sse_event_disable(evt);
> + sse_event_unregister(evt);
> + return -ENOMEM;
> + }
> +
> + pr_info("Using SSE for NMI event delivery\n");
> +
> + return 0;
> +}
> +
> +static int __init unknow_nmi_init(void)
> +{
> + return sse_nmi_init();
> +}
> +
> +late_initcall(unknow_nmi_init);
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [External] Re: [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support
2025-09-19 7:17 ` Clément Léger
@ 2025-09-19 7:52 ` yunhui cui
2025-09-22 8:11 ` yunhui cui
0 siblings, 1 reply; 11+ messages in thread
From: yunhui cui @ 2025-09-19 7:52 UTC (permalink / raw)
To: Clément Léger
Cc: paul.walmsley, palmer, aou, alex, conor, atishp, ajones, apatel,
mchitale, linux-riscv, linux-kernel
Hi Clément,
On Fri, Sep 19, 2025 at 3:18 PM Clément Léger <cleger@rivosinc.com> wrote:
>
>
>
> On 19/09/2025 09:00, Yunhui Cui wrote:
> > Unknown NMI can force the kernel to respond (e.g., panic) when the
> > system encounters unrecognized critical hardware events, aiding in
> > troubleshooting system faults. This is implemented via the Supervisor
> > Software Events (SSE) framework.
> >
> > Signed-off-by: Yunhui Cui <cuiyunhui@bytedance.com>
> > ---
> > arch/riscv/include/asm/sbi.h | 1 +
> > drivers/firmware/riscv/Kconfig | 10 +++++
> > drivers/firmware/riscv/Makefile | 1 +
> > drivers/firmware/riscv/sse_nmi.c | 77 ++++++++++++++++++++++++++++++++
> > 4 files changed, 89 insertions(+)
> > create mode 100644 drivers/firmware/riscv/sse_nmi.c
> >
> > diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h
> > index 874cc1d7603a5..5801f90a88f62 100644
> > --- a/arch/riscv/include/asm/sbi.h
> > +++ b/arch/riscv/include/asm/sbi.h
> > @@ -481,6 +481,7 @@ enum sbi_sse_attr_id {
> >
> > #define SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS 0x00000000
> > #define SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP 0x00000001
> > +#define SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI 0x00000002
>
> Was this submitted to the PRS WG ? This a specification modification so
> it should go through the usual process.
>
> > #define SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS 0x00008000
> > #define SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW 0x00010000
> > #define SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS 0x00100000
> > diff --git a/drivers/firmware/riscv/Kconfig b/drivers/firmware/riscv/Kconfig
> > index ed5b663ac5f91..746bac862ac46 100644
> > --- a/drivers/firmware/riscv/Kconfig
> > +++ b/drivers/firmware/riscv/Kconfig
> > @@ -12,4 +12,14 @@ config RISCV_SBI_SSE
> > this option provides support to register callbacks on specific SSE
> > events.
> >
> > +config RISCV_SSE_UNKNOWN_NMI
> > + bool "Enable SBI Supervisor Software Events unknown NMI support"
> > + depends on RISCV_SBI_SSE
> > + default y
> > + help
> > + This option enables support for delivering unknown Non-Maskable Interrupt (NMI)
> > + notifications via the Supervisor Software Events (SSE) framework. When enabled,
> > + unknown NMIs can trigger kernel responses (e.g., panic) for unrecognized critical
> > + hardware events, aiding in system fault diagnosis.
> > +
> > endmenu
> > diff --git a/drivers/firmware/riscv/Makefile b/drivers/firmware/riscv/Makefile
> > index c8795d4bbb2ea..9242c6cd5e3e9 100644
> > --- a/drivers/firmware/riscv/Makefile
> > +++ b/drivers/firmware/riscv/Makefile
> > @@ -1,3 +1,4 @@
> > # SPDX-License-Identifier: GPL-2.0
> >
> > obj-$(CONFIG_RISCV_SBI_SSE) += riscv_sbi_sse.o
> > +obj-$(CONFIG_RISCV_SSE_UNKNOWN_NMI) += sse_nmi.o
> > diff --git a/drivers/firmware/riscv/sse_nmi.c b/drivers/firmware/riscv/sse_nmi.c
> > new file mode 100644
> > index 0000000000000..43063f42efff0
> > --- /dev/null
> > +++ b/drivers/firmware/riscv/sse_nmi.c
> > @@ -0,0 +1,77 @@
> > +// SPDX-License-Identifier: GPL-2.0+
> > +
> > +#include <linux/mm.h>
> > +#include <linux/nmi.h>
> > +#include <linux/riscv_sbi_sse.h>
> > +#include <linux/sched/debug.h>
> > +#include <linux/sysctl.h>
> > +
> > +#include <asm/irq_regs.h>
> > +#include <asm/sbi.h>
> > +
> > +int panic_on_unknown_nmi = 1;
> > +struct sse_event *evt;
> > +static struct ctl_table_header *unknown_nmi_sysctl_header;
> > +
> > +const struct ctl_table unknown_nmi_table[] = {
> > + {
> > + .procname = "panic_enable",
> > + .data = &panic_on_unknown_nmi,
> > + .maxlen = sizeof(int),
> > + .mode = 0644,
> > + .proc_handler = proc_dointvec_minmax,
> > + .extra1 = SYSCTL_ZERO,
> > + .extra2 = SYSCTL_ONE,
> > + },
> > +};
> > +
> > +static void nmi_handler(struct pt_regs *regs)
> > +{
> > + pr_emerg("NMI received for unknown on CPU %d.\n", smp_processor_id());
> > +
> > + if (panic_on_unknown_nmi)
> > + nmi_panic(regs, "NMI: Not continuing");
> > +
> > + pr_emerg("Dazed and confused, but trying to continue\n");
> > +}
>
> I'm dazed and confused as well ;) What's the point of this except
> interrupting the kernel with a panic ? It seems like it's a better idea
> to let the firmware handle that properly and display whatever
> information are needed. Was your idea to actually force the kernel to
> enter in some debug mode ?
There is an important scenario: when the kernel becomes unresponsive,
we need to trigger an unknown NMI to cause the system to panic() and
then collect the vmcore, and such a requirement is common on x86
servers.
>
> Thanks,
>
> Clément
>
> > +
> > +static int nmi_sse_handler(u32 evt, void *arg, struct pt_regs *regs)
> > +{
> > + nmi_handler(regs);
> > +
> > + return 0;
> > +}
> > +
> > +static int sse_nmi_init(void)
> > +{
> > + int ret;
> > +
> > + evt = sse_event_register(SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI, 0,
> > + nmi_sse_handler, NULL);
> > + if (IS_ERR(evt))
> > + return PTR_ERR(evt);
> > +
> > + ret = sse_event_enable(evt);
> > + if (ret) {
> > + sse_event_unregister(evt);
> > + return ret;
> > + }
> > +
> > + unknown_nmi_sysctl_header = register_sysctl("kernel", unknown_nmi_table);
> > + if (!unknown_nmi_sysctl_header) {
> > + sse_event_disable(evt);
> > + sse_event_unregister(evt);
> > + return -ENOMEM;
> > + }
> > +
> > + pr_info("Using SSE for NMI event delivery\n");
> > +
> > + return 0;
> > +}
> > +
> > +static int __init unknow_nmi_init(void)
> > +{
> > + return sse_nmi_init();
> > +}
> > +
> > +late_initcall(unknow_nmi_init);
>
Thanks,
Yunhui
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [External] Re: [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support
2025-09-19 7:52 ` [External] " yunhui cui
@ 2025-09-22 8:11 ` yunhui cui
2025-09-22 8:31 ` Clément Léger
2025-09-22 10:39 ` Anup Patel
0 siblings, 2 replies; 11+ messages in thread
From: yunhui cui @ 2025-09-22 8:11 UTC (permalink / raw)
To: Clément Léger
Cc: paul.walmsley, palmer, aou, alex, conor, atishp, ajones, apatel,
mchitale, linux-riscv, linux-kernel
Hi Clément,
On Fri, Sep 19, 2025 at 3:52 PM yunhui cui <cuiyunhui@bytedance.com> wrote:
>
> Hi Clément,
>
> On Fri, Sep 19, 2025 at 3:18 PM Clément Léger <cleger@rivosinc.com> wrote:
> >
> >
> >
> > On 19/09/2025 09:00, Yunhui Cui wrote:
> > > Unknown NMI can force the kernel to respond (e.g., panic) when the
> > > system encounters unrecognized critical hardware events, aiding in
> > > troubleshooting system faults. This is implemented via the Supervisor
> > > Software Events (SSE) framework.
> > >
> > > Signed-off-by: Yunhui Cui <cuiyunhui@bytedance.com>
> > > ---
> > > arch/riscv/include/asm/sbi.h | 1 +
> > > drivers/firmware/riscv/Kconfig | 10 +++++
> > > drivers/firmware/riscv/Makefile | 1 +
> > > drivers/firmware/riscv/sse_nmi.c | 77 ++++++++++++++++++++++++++++++++
> > > 4 files changed, 89 insertions(+)
> > > create mode 100644 drivers/firmware/riscv/sse_nmi.c
> > >
> > > diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h
> > > index 874cc1d7603a5..5801f90a88f62 100644
> > > --- a/arch/riscv/include/asm/sbi.h
> > > +++ b/arch/riscv/include/asm/sbi.h
> > > @@ -481,6 +481,7 @@ enum sbi_sse_attr_id {
> > >
> > > #define SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS 0x00000000
> > > #define SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP 0x00000001
> > > +#define SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI 0x00000002
> >
> > Was this submitted to the PRS WG ? This a specification modification so
> > it should go through the usual process.
> >
> > > #define SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS 0x00008000
> > > #define SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW 0x00010000
> > > #define SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS 0x00100000
> > > diff --git a/drivers/firmware/riscv/Kconfig b/drivers/firmware/riscv/Kconfig
> > > index ed5b663ac5f91..746bac862ac46 100644
> > > --- a/drivers/firmware/riscv/Kconfig
> > > +++ b/drivers/firmware/riscv/Kconfig
> > > @@ -12,4 +12,14 @@ config RISCV_SBI_SSE
> > > this option provides support to register callbacks on specific SSE
> > > events.
> > >
> > > +config RISCV_SSE_UNKNOWN_NMI
> > > + bool "Enable SBI Supervisor Software Events unknown NMI support"
> > > + depends on RISCV_SBI_SSE
> > > + default y
> > > + help
> > > + This option enables support for delivering unknown Non-Maskable Interrupt (NMI)
> > > + notifications via the Supervisor Software Events (SSE) framework. When enabled,
> > > + unknown NMIs can trigger kernel responses (e.g., panic) for unrecognized critical
> > > + hardware events, aiding in system fault diagnosis.
> > > +
> > > endmenu
> > > diff --git a/drivers/firmware/riscv/Makefile b/drivers/firmware/riscv/Makefile
> > > index c8795d4bbb2ea..9242c6cd5e3e9 100644
> > > --- a/drivers/firmware/riscv/Makefile
> > > +++ b/drivers/firmware/riscv/Makefile
> > > @@ -1,3 +1,4 @@
> > > # SPDX-License-Identifier: GPL-2.0
> > >
> > > obj-$(CONFIG_RISCV_SBI_SSE) += riscv_sbi_sse.o
> > > +obj-$(CONFIG_RISCV_SSE_UNKNOWN_NMI) += sse_nmi.o
> > > diff --git a/drivers/firmware/riscv/sse_nmi.c b/drivers/firmware/riscv/sse_nmi.c
> > > new file mode 100644
> > > index 0000000000000..43063f42efff0
> > > --- /dev/null
> > > +++ b/drivers/firmware/riscv/sse_nmi.c
> > > @@ -0,0 +1,77 @@
> > > +// SPDX-License-Identifier: GPL-2.0+
> > > +
> > > +#include <linux/mm.h>
> > > +#include <linux/nmi.h>
> > > +#include <linux/riscv_sbi_sse.h>
> > > +#include <linux/sched/debug.h>
> > > +#include <linux/sysctl.h>
> > > +
> > > +#include <asm/irq_regs.h>
> > > +#include <asm/sbi.h>
> > > +
> > > +int panic_on_unknown_nmi = 1;
> > > +struct sse_event *evt;
> > > +static struct ctl_table_header *unknown_nmi_sysctl_header;
> > > +
> > > +const struct ctl_table unknown_nmi_table[] = {
> > > + {
> > > + .procname = "panic_enable",
> > > + .data = &panic_on_unknown_nmi,
> > > + .maxlen = sizeof(int),
> > > + .mode = 0644,
> > > + .proc_handler = proc_dointvec_minmax,
> > > + .extra1 = SYSCTL_ZERO,
> > > + .extra2 = SYSCTL_ONE,
> > > + },
> > > +};
> > > +
> > > +static void nmi_handler(struct pt_regs *regs)
> > > +{
> > > + pr_emerg("NMI received for unknown on CPU %d.\n", smp_processor_id());
> > > +
> > > + if (panic_on_unknown_nmi)
> > > + nmi_panic(regs, "NMI: Not continuing");
> > > +
> > > + pr_emerg("Dazed and confused, but trying to continue\n");
> > > +}
> >
> > I'm dazed and confused as well ;) What's the point of this except
> > interrupting the kernel with a panic ? It seems like it's a better idea
> > to let the firmware handle that properly and display whatever
> > information are needed. Was your idea to actually force the kernel to
> > enter in some debug mode ?
>
> There is an important scenario: when the kernel becomes unresponsive,
> we need to trigger an unknown NMI to cause the system to panic() and
> then collect the vmcore, and such a requirement is common on x86
> servers.
>
> >
> > Thanks,
> >
> > Clément
> >
> > > +
> > > +static int nmi_sse_handler(u32 evt, void *arg, struct pt_regs *regs)
> > > +{
> > > + nmi_handler(regs);
> > > +
> > > + return 0;
> > > +}
> > > +
> > > +static int sse_nmi_init(void)
> > > +{
> > > + int ret;
> > > +
> > > + evt = sse_event_register(SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI, 0,
> > > + nmi_sse_handler, NULL);
Should we add this UNKNOWN_NMI event ID in Chapter 17 of the SBI spec?
> > > + if (IS_ERR(evt))
> > > + return PTR_ERR(evt);
> > > +
> > > + ret = sse_event_enable(evt);
> > > + if (ret) {
> > > + sse_event_unregister(evt);
> > > + return ret;
> > > + }
> > > +
> > > + unknown_nmi_sysctl_header = register_sysctl("kernel", unknown_nmi_table);
> > > + if (!unknown_nmi_sysctl_header) {
> > > + sse_event_disable(evt);
> > > + sse_event_unregister(evt);
> > > + return -ENOMEM;
> > > + }
> > > +
> > > + pr_info("Using SSE for NMI event delivery\n");
> > > +
> > > + return 0;
> > > +}
> > > +
> > > +static int __init unknow_nmi_init(void)
> > > +{
> > > + return sse_nmi_init();
> > > +}
> > > +
> > > +late_initcall(unknow_nmi_init);
> >
>
> Thanks,
> Yunhui
Thanks,
Yunhui
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [External] Re: [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support
2025-09-22 8:11 ` yunhui cui
@ 2025-09-22 8:31 ` Clément Léger
2025-09-22 10:39 ` Anup Patel
1 sibling, 0 replies; 11+ messages in thread
From: Clément Léger @ 2025-09-22 8:31 UTC (permalink / raw)
To: yunhui cui
Cc: paul.walmsley, palmer, aou, alex, conor, atishp, ajones, apatel,
mchitale, linux-riscv, linux-kernel
On 22/09/2025 10:11, yunhui cui wrote:
> Hi Clément,
>
> On Fri, Sep 19, 2025 at 3:52 PM yunhui cui <cuiyunhui@bytedance.com> wrote:
>>
>> Hi Clément,
>>
>> On Fri, Sep 19, 2025 at 3:18 PM Clément Léger <cleger@rivosinc.com> wrote:
>>>
>>>
>>>
>>> On 19/09/2025 09:00, Yunhui Cui wrote:
>>>> Unknown NMI can force the kernel to respond (e.g., panic) when the
>>>> system encounters unrecognized critical hardware events, aiding in
>>>> troubleshooting system faults. This is implemented via the Supervisor
>>>> Software Events (SSE) framework.
>>>>
>>>> Signed-off-by: Yunhui Cui <cuiyunhui@bytedance.com>
>>>> ---
>>>> arch/riscv/include/asm/sbi.h | 1 +
>>>> drivers/firmware/riscv/Kconfig | 10 +++++
>>>> drivers/firmware/riscv/Makefile | 1 +
>>>> drivers/firmware/riscv/sse_nmi.c | 77 ++++++++++++++++++++++++++++++++
>>>> 4 files changed, 89 insertions(+)
>>>> create mode 100644 drivers/firmware/riscv/sse_nmi.c
>>>>
>>>> diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h
>>>> index 874cc1d7603a5..5801f90a88f62 100644
>>>> --- a/arch/riscv/include/asm/sbi.h
>>>> +++ b/arch/riscv/include/asm/sbi.h
>>>> @@ -481,6 +481,7 @@ enum sbi_sse_attr_id {
>>>>
>>>> #define SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS 0x00000000
>>>> #define SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP 0x00000001
>>>> +#define SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI 0x00000002
>>>
>>> Was this submitted to the PRS WG ? This a specification modification so
>>> it should go through the usual process.
>>>
>>>> #define SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS 0x00008000
>>>> #define SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW 0x00010000
>>>> #define SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS 0x00100000
>>>> diff --git a/drivers/firmware/riscv/Kconfig b/drivers/firmware/riscv/Kconfig
>>>> index ed5b663ac5f91..746bac862ac46 100644
>>>> --- a/drivers/firmware/riscv/Kconfig
>>>> +++ b/drivers/firmware/riscv/Kconfig
>>>> @@ -12,4 +12,14 @@ config RISCV_SBI_SSE
>>>> this option provides support to register callbacks on specific SSE
>>>> events.
>>>>
>>>> +config RISCV_SSE_UNKNOWN_NMI
>>>> + bool "Enable SBI Supervisor Software Events unknown NMI support"
>>>> + depends on RISCV_SBI_SSE
>>>> + default y
>>>> + help
>>>> + This option enables support for delivering unknown Non-Maskable Interrupt (NMI)
>>>> + notifications via the Supervisor Software Events (SSE) framework. When enabled,
>>>> + unknown NMIs can trigger kernel responses (e.g., panic) for unrecognized critical
>>>> + hardware events, aiding in system fault diagnosis.
>>>> +
>>>> endmenu
>>>> diff --git a/drivers/firmware/riscv/Makefile b/drivers/firmware/riscv/Makefile
>>>> index c8795d4bbb2ea..9242c6cd5e3e9 100644
>>>> --- a/drivers/firmware/riscv/Makefile
>>>> +++ b/drivers/firmware/riscv/Makefile
>>>> @@ -1,3 +1,4 @@
>>>> # SPDX-License-Identifier: GPL-2.0
>>>>
>>>> obj-$(CONFIG_RISCV_SBI_SSE) += riscv_sbi_sse.o
>>>> +obj-$(CONFIG_RISCV_SSE_UNKNOWN_NMI) += sse_nmi.o
>>>> diff --git a/drivers/firmware/riscv/sse_nmi.c b/drivers/firmware/riscv/sse_nmi.c
>>>> new file mode 100644
>>>> index 0000000000000..43063f42efff0
>>>> --- /dev/null
>>>> +++ b/drivers/firmware/riscv/sse_nmi.c
>>>> @@ -0,0 +1,77 @@
>>>> +// SPDX-License-Identifier: GPL-2.0+
>>>> +
>>>> +#include <linux/mm.h>
>>>> +#include <linux/nmi.h>
>>>> +#include <linux/riscv_sbi_sse.h>
>>>> +#include <linux/sched/debug.h>
>>>> +#include <linux/sysctl.h>
>>>> +
>>>> +#include <asm/irq_regs.h>
>>>> +#include <asm/sbi.h>
>>>> +
>>>> +int panic_on_unknown_nmi = 1;
>>>> +struct sse_event *evt;
>>>> +static struct ctl_table_header *unknown_nmi_sysctl_header;
>>>> +
>>>> +const struct ctl_table unknown_nmi_table[] = {
>>>> + {
>>>> + .procname = "panic_enable",
>>>> + .data = &panic_on_unknown_nmi,
>>>> + .maxlen = sizeof(int),
>>>> + .mode = 0644,
>>>> + .proc_handler = proc_dointvec_minmax,
>>>> + .extra1 = SYSCTL_ZERO,
>>>> + .extra2 = SYSCTL_ONE,
>>>> + },
>>>> +};
>>>> +
>>>> +static void nmi_handler(struct pt_regs *regs)
>>>> +{
>>>> + pr_emerg("NMI received for unknown on CPU %d.\n", smp_processor_id());
>>>> +
>>>> + if (panic_on_unknown_nmi)
>>>> + nmi_panic(regs, "NMI: Not continuing");
>>>> +
>>>> + pr_emerg("Dazed and confused, but trying to continue\n");
>>>> +}
>>>
>>> I'm dazed and confused as well ;) What's the point of this except
>>> interrupting the kernel with a panic ? It seems like it's a better idea
>>> to let the firmware handle that properly and display whatever
>>> information are needed. Was your idea to actually force the kernel to
>>> enter in some debug mode ?
>>
>> There is an important scenario: when the kernel becomes unresponsive,
>> we need to trigger an unknown NMI to cause the system to panic() and
>> then collect the vmcore, and such a requirement is common on x86
>> servers.
>>
>>>
>>> Thanks,
>>>
>>> Clément
>>>
>>>> +
>>>> +static int nmi_sse_handler(u32 evt, void *arg, struct pt_regs *regs)
>>>> +{
>>>> + nmi_handler(regs);
>>>> +
>>>> + return 0;
>>>> +}
>>>> +
>>>> +static int sse_nmi_init(void)
>>>> +{
>>>> + int ret;
>>>> +
>>>> + evt = sse_event_register(SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI, 0,
>>>> + nmi_sse_handler, NULL);
>
> Should we add this UNKNOWN_NMI event ID in Chapter 17 of the SBI spec?
Hi Yunhui,
If you want it to be part of the spec, that should indeed be submitted
for review.
Thanks,
Clément
>
>
>>>> + if (IS_ERR(evt))
>>>> + return PTR_ERR(evt);
>>>> +
>>>> + ret = sse_event_enable(evt);
>>>> + if (ret) {
>>>> + sse_event_unregister(evt);
>>>> + return ret;
>>>> + }
>>>> +
>>>> + unknown_nmi_sysctl_header = register_sysctl("kernel", unknown_nmi_table);
>>>> + if (!unknown_nmi_sysctl_header) {
>>>> + sse_event_disable(evt);
>>>> + sse_event_unregister(evt);
>>>> + return -ENOMEM;
>>>> + }
>>>> +
>>>> + pr_info("Using SSE for NMI event delivery\n");
>>>> +
>>>> + return 0;
>>>> +}
>>>> +
>>>> +static int __init unknow_nmi_init(void)
>>>> +{
>>>> + return sse_nmi_init();
>>>> +}
>>>> +
>>>> +late_initcall(unknow_nmi_init);
>>>
>>
>> Thanks,
>> Yunhui
>
> Thanks,
> Yunhui
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [External] Re: [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support
2025-09-22 8:11 ` yunhui cui
2025-09-22 8:31 ` Clément Léger
@ 2025-09-22 10:39 ` Anup Patel
2025-09-22 11:58 ` yunhui cui
1 sibling, 1 reply; 11+ messages in thread
From: Anup Patel @ 2025-09-22 10:39 UTC (permalink / raw)
To: yunhui cui
Cc: Clément Léger, paul.walmsley, palmer, aou, alex, conor,
atishp, ajones, apatel, mchitale, linux-riscv, linux-kernel
On Mon, Sep 22, 2025 at 1:42 PM yunhui cui <cuiyunhui@bytedance.com> wrote:
>
> Hi Clément,
>
> On Fri, Sep 19, 2025 at 3:52 PM yunhui cui <cuiyunhui@bytedance.com> wrote:
> >
> > Hi Clément,
> >
> > On Fri, Sep 19, 2025 at 3:18 PM Clément Léger <cleger@rivosinc.com> wrote:
> > >
> > >
> > >
> > > On 19/09/2025 09:00, Yunhui Cui wrote:
> > > > Unknown NMI can force the kernel to respond (e.g., panic) when the
> > > > system encounters unrecognized critical hardware events, aiding in
> > > > troubleshooting system faults. This is implemented via the Supervisor
> > > > Software Events (SSE) framework.
> > > >
> > > > Signed-off-by: Yunhui Cui <cuiyunhui@bytedance.com>
> > > > ---
> > > > arch/riscv/include/asm/sbi.h | 1 +
> > > > drivers/firmware/riscv/Kconfig | 10 +++++
> > > > drivers/firmware/riscv/Makefile | 1 +
> > > > drivers/firmware/riscv/sse_nmi.c | 77 ++++++++++++++++++++++++++++++++
> > > > 4 files changed, 89 insertions(+)
> > > > create mode 100644 drivers/firmware/riscv/sse_nmi.c
> > > >
> > > > diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h
> > > > index 874cc1d7603a5..5801f90a88f62 100644
> > > > --- a/arch/riscv/include/asm/sbi.h
> > > > +++ b/arch/riscv/include/asm/sbi.h
> > > > @@ -481,6 +481,7 @@ enum sbi_sse_attr_id {
> > > >
> > > > #define SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS 0x00000000
> > > > #define SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP 0x00000001
> > > > +#define SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI 0x00000002
> > >
> > > Was this submitted to the PRS WG ? This a specification modification so
> > > it should go through the usual process.
> > >
> > > > #define SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS 0x00008000
> > > > #define SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW 0x00010000
> > > > #define SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS 0x00100000
> > > > diff --git a/drivers/firmware/riscv/Kconfig b/drivers/firmware/riscv/Kconfig
> > > > index ed5b663ac5f91..746bac862ac46 100644
> > > > --- a/drivers/firmware/riscv/Kconfig
> > > > +++ b/drivers/firmware/riscv/Kconfig
> > > > @@ -12,4 +12,14 @@ config RISCV_SBI_SSE
> > > > this option provides support to register callbacks on specific SSE
> > > > events.
> > > >
> > > > +config RISCV_SSE_UNKNOWN_NMI
> > > > + bool "Enable SBI Supervisor Software Events unknown NMI support"
> > > > + depends on RISCV_SBI_SSE
> > > > + default y
> > > > + help
> > > > + This option enables support for delivering unknown Non-Maskable Interrupt (NMI)
> > > > + notifications via the Supervisor Software Events (SSE) framework. When enabled,
> > > > + unknown NMIs can trigger kernel responses (e.g., panic) for unrecognized critical
> > > > + hardware events, aiding in system fault diagnosis.
> > > > +
> > > > endmenu
> > > > diff --git a/drivers/firmware/riscv/Makefile b/drivers/firmware/riscv/Makefile
> > > > index c8795d4bbb2ea..9242c6cd5e3e9 100644
> > > > --- a/drivers/firmware/riscv/Makefile
> > > > +++ b/drivers/firmware/riscv/Makefile
> > > > @@ -1,3 +1,4 @@
> > > > # SPDX-License-Identifier: GPL-2.0
> > > >
> > > > obj-$(CONFIG_RISCV_SBI_SSE) += riscv_sbi_sse.o
> > > > +obj-$(CONFIG_RISCV_SSE_UNKNOWN_NMI) += sse_nmi.o
> > > > diff --git a/drivers/firmware/riscv/sse_nmi.c b/drivers/firmware/riscv/sse_nmi.c
> > > > new file mode 100644
> > > > index 0000000000000..43063f42efff0
> > > > --- /dev/null
> > > > +++ b/drivers/firmware/riscv/sse_nmi.c
> > > > @@ -0,0 +1,77 @@
> > > > +// SPDX-License-Identifier: GPL-2.0+
> > > > +
> > > > +#include <linux/mm.h>
> > > > +#include <linux/nmi.h>
> > > > +#include <linux/riscv_sbi_sse.h>
> > > > +#include <linux/sched/debug.h>
> > > > +#include <linux/sysctl.h>
> > > > +
> > > > +#include <asm/irq_regs.h>
> > > > +#include <asm/sbi.h>
> > > > +
> > > > +int panic_on_unknown_nmi = 1;
> > > > +struct sse_event *evt;
> > > > +static struct ctl_table_header *unknown_nmi_sysctl_header;
> > > > +
> > > > +const struct ctl_table unknown_nmi_table[] = {
> > > > + {
> > > > + .procname = "panic_enable",
> > > > + .data = &panic_on_unknown_nmi,
> > > > + .maxlen = sizeof(int),
> > > > + .mode = 0644,
> > > > + .proc_handler = proc_dointvec_minmax,
> > > > + .extra1 = SYSCTL_ZERO,
> > > > + .extra2 = SYSCTL_ONE,
> > > > + },
> > > > +};
> > > > +
> > > > +static void nmi_handler(struct pt_regs *regs)
> > > > +{
> > > > + pr_emerg("NMI received for unknown on CPU %d.\n", smp_processor_id());
> > > > +
> > > > + if (panic_on_unknown_nmi)
> > > > + nmi_panic(regs, "NMI: Not continuing");
> > > > +
> > > > + pr_emerg("Dazed and confused, but trying to continue\n");
> > > > +}
> > >
> > > I'm dazed and confused as well ;) What's the point of this except
> > > interrupting the kernel with a panic ? It seems like it's a better idea
> > > to let the firmware handle that properly and display whatever
> > > information are needed. Was your idea to actually force the kernel to
> > > enter in some debug mode ?
> >
> > There is an important scenario: when the kernel becomes unresponsive,
> > we need to trigger an unknown NMI to cause the system to panic() and
> > then collect the vmcore, and such a requirement is common on x86
> > servers.
> >
> > >
> > > Thanks,
> > >
> > > Clément
> > >
> > > > +
> > > > +static int nmi_sse_handler(u32 evt, void *arg, struct pt_regs *regs)
> > > > +{
> > > > + nmi_handler(regs);
> > > > +
> > > > + return 0;
> > > > +}
> > > > +
> > > > +static int sse_nmi_init(void)
> > > > +{
> > > > + int ret;
> > > > +
> > > > + evt = sse_event_register(SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI, 0,
> > > > + nmi_sse_handler, NULL);
>
> Should we add this UNKNOWN_NMI event ID in Chapter 17 of the SBI spec?
The ratified SBI v3.0 defines a "Software injected local event" (ID 0xffff0000)
which can be used to inject panic() on a particular HART from another HART.
Regards,
Anup
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [External] Re: [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support
2025-09-22 10:39 ` Anup Patel
@ 2025-09-22 11:58 ` yunhui cui
2025-09-22 17:20 ` Atish Kumar Patra
2025-09-23 5:05 ` Anup Patel
0 siblings, 2 replies; 11+ messages in thread
From: yunhui cui @ 2025-09-22 11:58 UTC (permalink / raw)
To: Anup Patel
Cc: Clément Léger, paul.walmsley, palmer, aou, alex, conor,
atishp, ajones, apatel, mchitale, linux-riscv, linux-kernel
Hi Anup,
On Mon, Sep 22, 2025 at 6:40 PM Anup Patel <anup@brainfault.org> wrote:
>
> On Mon, Sep 22, 2025 at 1:42 PM yunhui cui <cuiyunhui@bytedance.com> wrote:
> >
> > Hi Clément,
> >
> > On Fri, Sep 19, 2025 at 3:52 PM yunhui cui <cuiyunhui@bytedance.com> wrote:
> > >
> > > Hi Clément,
> > >
> > > On Fri, Sep 19, 2025 at 3:18 PM Clément Léger <cleger@rivosinc.com> wrote:
> > > >
> > > >
> > > >
> > > > On 19/09/2025 09:00, Yunhui Cui wrote:
> > > > > Unknown NMI can force the kernel to respond (e.g., panic) when the
> > > > > system encounters unrecognized critical hardware events, aiding in
> > > > > troubleshooting system faults. This is implemented via the Supervisor
> > > > > Software Events (SSE) framework.
> > > > >
> > > > > Signed-off-by: Yunhui Cui <cuiyunhui@bytedance.com>
> > > > > ---
> > > > > arch/riscv/include/asm/sbi.h | 1 +
> > > > > drivers/firmware/riscv/Kconfig | 10 +++++
> > > > > drivers/firmware/riscv/Makefile | 1 +
> > > > > drivers/firmware/riscv/sse_nmi.c | 77 ++++++++++++++++++++++++++++++++
> > > > > 4 files changed, 89 insertions(+)
> > > > > create mode 100644 drivers/firmware/riscv/sse_nmi.c
> > > > >
> > > > > diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h
> > > > > index 874cc1d7603a5..5801f90a88f62 100644
> > > > > --- a/arch/riscv/include/asm/sbi.h
> > > > > +++ b/arch/riscv/include/asm/sbi.h
> > > > > @@ -481,6 +481,7 @@ enum sbi_sse_attr_id {
> > > > >
> > > > > #define SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS 0x00000000
> > > > > #define SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP 0x00000001
> > > > > +#define SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI 0x00000002
> > > >
> > > > Was this submitted to the PRS WG ? This a specification modification so
> > > > it should go through the usual process.
> > > >
> > > > > #define SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS 0x00008000
> > > > > #define SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW 0x00010000
> > > > > #define SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS 0x00100000
> > > > > diff --git a/drivers/firmware/riscv/Kconfig b/drivers/firmware/riscv/Kconfig
> > > > > index ed5b663ac5f91..746bac862ac46 100644
> > > > > --- a/drivers/firmware/riscv/Kconfig
> > > > > +++ b/drivers/firmware/riscv/Kconfig
> > > > > @@ -12,4 +12,14 @@ config RISCV_SBI_SSE
> > > > > this option provides support to register callbacks on specific SSE
> > > > > events.
> > > > >
> > > > > +config RISCV_SSE_UNKNOWN_NMI
> > > > > + bool "Enable SBI Supervisor Software Events unknown NMI support"
> > > > > + depends on RISCV_SBI_SSE
> > > > > + default y
> > > > > + help
> > > > > + This option enables support for delivering unknown Non-Maskable Interrupt (NMI)
> > > > > + notifications via the Supervisor Software Events (SSE) framework. When enabled,
> > > > > + unknown NMIs can trigger kernel responses (e.g., panic) for unrecognized critical
> > > > > + hardware events, aiding in system fault diagnosis.
> > > > > +
> > > > > endmenu
> > > > > diff --git a/drivers/firmware/riscv/Makefile b/drivers/firmware/riscv/Makefile
> > > > > index c8795d4bbb2ea..9242c6cd5e3e9 100644
> > > > > --- a/drivers/firmware/riscv/Makefile
> > > > > +++ b/drivers/firmware/riscv/Makefile
> > > > > @@ -1,3 +1,4 @@
> > > > > # SPDX-License-Identifier: GPL-2.0
> > > > >
> > > > > obj-$(CONFIG_RISCV_SBI_SSE) += riscv_sbi_sse.o
> > > > > +obj-$(CONFIG_RISCV_SSE_UNKNOWN_NMI) += sse_nmi.o
> > > > > diff --git a/drivers/firmware/riscv/sse_nmi.c b/drivers/firmware/riscv/sse_nmi.c
> > > > > new file mode 100644
> > > > > index 0000000000000..43063f42efff0
> > > > > --- /dev/null
> > > > > +++ b/drivers/firmware/riscv/sse_nmi.c
> > > > > @@ -0,0 +1,77 @@
> > > > > +// SPDX-License-Identifier: GPL-2.0+
> > > > > +
> > > > > +#include <linux/mm.h>
> > > > > +#include <linux/nmi.h>
> > > > > +#include <linux/riscv_sbi_sse.h>
> > > > > +#include <linux/sched/debug.h>
> > > > > +#include <linux/sysctl.h>
> > > > > +
> > > > > +#include <asm/irq_regs.h>
> > > > > +#include <asm/sbi.h>
> > > > > +
> > > > > +int panic_on_unknown_nmi = 1;
> > > > > +struct sse_event *evt;
> > > > > +static struct ctl_table_header *unknown_nmi_sysctl_header;
> > > > > +
> > > > > +const struct ctl_table unknown_nmi_table[] = {
> > > > > + {
> > > > > + .procname = "panic_enable",
> > > > > + .data = &panic_on_unknown_nmi,
> > > > > + .maxlen = sizeof(int),
> > > > > + .mode = 0644,
> > > > > + .proc_handler = proc_dointvec_minmax,
> > > > > + .extra1 = SYSCTL_ZERO,
> > > > > + .extra2 = SYSCTL_ONE,
> > > > > + },
> > > > > +};
> > > > > +
> > > > > +static void nmi_handler(struct pt_regs *regs)
> > > > > +{
> > > > > + pr_emerg("NMI received for unknown on CPU %d.\n", smp_processor_id());
> > > > > +
> > > > > + if (panic_on_unknown_nmi)
> > > > > + nmi_panic(regs, "NMI: Not continuing");
> > > > > +
> > > > > + pr_emerg("Dazed and confused, but trying to continue\n");
> > > > > +}
> > > >
> > > > I'm dazed and confused as well ;) What's the point of this except
> > > > interrupting the kernel with a panic ? It seems like it's a better idea
> > > > to let the firmware handle that properly and display whatever
> > > > information are needed. Was your idea to actually force the kernel to
> > > > enter in some debug mode ?
> > >
> > > There is an important scenario: when the kernel becomes unresponsive,
> > > we need to trigger an unknown NMI to cause the system to panic() and
> > > then collect the vmcore, and such a requirement is common on x86
> > > servers.
> > >
> > > >
> > > > Thanks,
> > > >
> > > > Clément
> > > >
> > > > > +
> > > > > +static int nmi_sse_handler(u32 evt, void *arg, struct pt_regs *regs)
> > > > > +{
> > > > > + nmi_handler(regs);
> > > > > +
> > > > > + return 0;
> > > > > +}
> > > > > +
> > > > > +static int sse_nmi_init(void)
> > > > > +{
> > > > > + int ret;
> > > > > +
> > > > > + evt = sse_event_register(SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI, 0,
> > > > > + nmi_sse_handler, NULL);
> >
> > Should we add this UNKNOWN_NMI event ID in Chapter 17 of the SBI spec?
>
> The ratified SBI v3.0 defines a "Software injected local event" (ID 0xffff0000)
> which can be used to inject panic() on a particular HART from another HART.
Has SBI_SSE_EVENT_LOCAL_SOFTWARE (0xffff0000) been given meaning
currently? The scenario of the unknown NMI is that after a HART
receives an interrupt from the platform, M-mode responds to this
interrupt and then enters S-mode to execute the registered handler.
Can SBI_SSE_EVENT_LOCAL_SOFTWARE be used directly?
>
> Regards,
> Anup
Thanks,
Yunhui
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [External] Re: [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support
2025-09-22 11:58 ` yunhui cui
@ 2025-09-22 17:20 ` Atish Kumar Patra
2025-09-23 2:15 ` yunhui cui
2025-09-23 5:05 ` Anup Patel
1 sibling, 1 reply; 11+ messages in thread
From: Atish Kumar Patra @ 2025-09-22 17:20 UTC (permalink / raw)
To: yunhui cui
Cc: Anup Patel, Clément Léger, paul.walmsley, palmer, aou,
alex, conor, ajones, apatel, mchitale, linux-riscv, linux-kernel
On Mon, Sep 22, 2025 at 4:58 AM yunhui cui <cuiyunhui@bytedance.com> wrote:
>
> Hi Anup,
>
> On Mon, Sep 22, 2025 at 6:40 PM Anup Patel <anup@brainfault.org> wrote:
> >
> > On Mon, Sep 22, 2025 at 1:42 PM yunhui cui <cuiyunhui@bytedance.com> wrote:
> > >
> > > Hi Clément,
> > >
> > > On Fri, Sep 19, 2025 at 3:52 PM yunhui cui <cuiyunhui@bytedance.com> wrote:
> > > >
> > > > Hi Clément,
> > > >
> > > > On Fri, Sep 19, 2025 at 3:18 PM Clément Léger <cleger@rivosinc.com> wrote:
> > > > >
> > > > >
> > > > >
> > > > > On 19/09/2025 09:00, Yunhui Cui wrote:
> > > > > > Unknown NMI can force the kernel to respond (e.g., panic) when the
> > > > > > system encounters unrecognized critical hardware events, aiding in
> > > > > > troubleshooting system faults. This is implemented via the Supervisor
> > > > > > Software Events (SSE) framework.
> > > > > >
> > > > > > Signed-off-by: Yunhui Cui <cuiyunhui@bytedance.com>
> > > > > > ---
> > > > > > arch/riscv/include/asm/sbi.h | 1 +
> > > > > > drivers/firmware/riscv/Kconfig | 10 +++++
> > > > > > drivers/firmware/riscv/Makefile | 1 +
> > > > > > drivers/firmware/riscv/sse_nmi.c | 77 ++++++++++++++++++++++++++++++++
> > > > > > 4 files changed, 89 insertions(+)
> > > > > > create mode 100644 drivers/firmware/riscv/sse_nmi.c
> > > > > >
> > > > > > diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h
> > > > > > index 874cc1d7603a5..5801f90a88f62 100644
> > > > > > --- a/arch/riscv/include/asm/sbi.h
> > > > > > +++ b/arch/riscv/include/asm/sbi.h
> > > > > > @@ -481,6 +481,7 @@ enum sbi_sse_attr_id {
> > > > > >
> > > > > > #define SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS 0x00000000
> > > > > > #define SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP 0x00000001
> > > > > > +#define SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI 0x00000002
> > > > >
> > > > > Was this submitted to the PRS WG ? This a specification modification so
> > > > > it should go through the usual process.
> > > > >
> > > > > > #define SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS 0x00008000
> > > > > > #define SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW 0x00010000
> > > > > > #define SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS 0x00100000
> > > > > > diff --git a/drivers/firmware/riscv/Kconfig b/drivers/firmware/riscv/Kconfig
> > > > > > index ed5b663ac5f91..746bac862ac46 100644
> > > > > > --- a/drivers/firmware/riscv/Kconfig
> > > > > > +++ b/drivers/firmware/riscv/Kconfig
> > > > > > @@ -12,4 +12,14 @@ config RISCV_SBI_SSE
> > > > > > this option provides support to register callbacks on specific SSE
> > > > > > events.
> > > > > >
> > > > > > +config RISCV_SSE_UNKNOWN_NMI
> > > > > > + bool "Enable SBI Supervisor Software Events unknown NMI support"
> > > > > > + depends on RISCV_SBI_SSE
> > > > > > + default y
> > > > > > + help
> > > > > > + This option enables support for delivering unknown Non-Maskable Interrupt (NMI)
> > > > > > + notifications via the Supervisor Software Events (SSE) framework. When enabled,
> > > > > > + unknown NMIs can trigger kernel responses (e.g., panic) for unrecognized critical
> > > > > > + hardware events, aiding in system fault diagnosis.
> > > > > > +
> > > > > > endmenu
> > > > > > diff --git a/drivers/firmware/riscv/Makefile b/drivers/firmware/riscv/Makefile
> > > > > > index c8795d4bbb2ea..9242c6cd5e3e9 100644
> > > > > > --- a/drivers/firmware/riscv/Makefile
> > > > > > +++ b/drivers/firmware/riscv/Makefile
> > > > > > @@ -1,3 +1,4 @@
> > > > > > # SPDX-License-Identifier: GPL-2.0
> > > > > >
> > > > > > obj-$(CONFIG_RISCV_SBI_SSE) += riscv_sbi_sse.o
> > > > > > +obj-$(CONFIG_RISCV_SSE_UNKNOWN_NMI) += sse_nmi.o
> > > > > > diff --git a/drivers/firmware/riscv/sse_nmi.c b/drivers/firmware/riscv/sse_nmi.c
> > > > > > new file mode 100644
> > > > > > index 0000000000000..43063f42efff0
> > > > > > --- /dev/null
> > > > > > +++ b/drivers/firmware/riscv/sse_nmi.c
> > > > > > @@ -0,0 +1,77 @@
> > > > > > +// SPDX-License-Identifier: GPL-2.0+
> > > > > > +
> > > > > > +#include <linux/mm.h>
> > > > > > +#include <linux/nmi.h>
> > > > > > +#include <linux/riscv_sbi_sse.h>
> > > > > > +#include <linux/sched/debug.h>
> > > > > > +#include <linux/sysctl.h>
> > > > > > +
> > > > > > +#include <asm/irq_regs.h>
> > > > > > +#include <asm/sbi.h>
> > > > > > +
> > > > > > +int panic_on_unknown_nmi = 1;
> > > > > > +struct sse_event *evt;
> > > > > > +static struct ctl_table_header *unknown_nmi_sysctl_header;
> > > > > > +
> > > > > > +const struct ctl_table unknown_nmi_table[] = {
> > > > > > + {
> > > > > > + .procname = "panic_enable",
> > > > > > + .data = &panic_on_unknown_nmi,
> > > > > > + .maxlen = sizeof(int),
> > > > > > + .mode = 0644,
> > > > > > + .proc_handler = proc_dointvec_minmax,
> > > > > > + .extra1 = SYSCTL_ZERO,
> > > > > > + .extra2 = SYSCTL_ONE,
> > > > > > + },
> > > > > > +};
> > > > > > +
> > > > > > +static void nmi_handler(struct pt_regs *regs)
> > > > > > +{
> > > > > > + pr_emerg("NMI received for unknown on CPU %d.\n", smp_processor_id());
> > > > > > +
> > > > > > + if (panic_on_unknown_nmi)
> > > > > > + nmi_panic(regs, "NMI: Not continuing");
> > > > > > +
> > > > > > + pr_emerg("Dazed and confused, but trying to continue\n");
> > > > > > +}
> > > > >
> > > > > I'm dazed and confused as well ;) What's the point of this except
> > > > > interrupting the kernel with a panic ? It seems like it's a better idea
> > > > > to let the firmware handle that properly and display whatever
> > > > > information are needed. Was your idea to actually force the kernel to
> > > > > enter in some debug mode ?
> > > >
> > > > There is an important scenario: when the kernel becomes unresponsive,
> > > > we need to trigger an unknown NMI to cause the system to panic() and
> > > > then collect the vmcore, and such a requirement is common on x86
> > > > servers.
> > > >
> > > > >
> > > > > Thanks,
> > > > >
> > > > > Clément
> > > > >
> > > > > > +
> > > > > > +static int nmi_sse_handler(u32 evt, void *arg, struct pt_regs *regs)
> > > > > > +{
> > > > > > + nmi_handler(regs);
> > > > > > +
> > > > > > + return 0;
> > > > > > +}
> > > > > > +
> > > > > > +static int sse_nmi_init(void)
> > > > > > +{
> > > > > > + int ret;
> > > > > > +
> > > > > > + evt = sse_event_register(SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI, 0,
> > > > > > + nmi_sse_handler, NULL);
> > >
> > > Should we add this UNKNOWN_NMI event ID in Chapter 17 of the SBI spec?
> >
> > The ratified SBI v3.0 defines a "Software injected local event" (ID 0xffff0000)
> > which can be used to inject panic() on a particular HART from another HART.
>
> Has SBI_SSE_EVENT_LOCAL_SOFTWARE (0xffff0000) been given meaning
> currently? The scenario of the unknown NMI is that after a HART
> receives an interrupt from the platform, M-mode responds to this
Which interrupt ? A standard one or platform specific one ?
> interrupt and then enters S-mode to execute the registered handler.
> Can SBI_SSE_EVENT_LOCAL_SOFTWARE be used directly?
>
>
> >
> > Regards,
> > Anup
>
> Thanks,
> Yunhui
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [External] Re: [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support
2025-09-22 17:20 ` Atish Kumar Patra
@ 2025-09-23 2:15 ` yunhui cui
0 siblings, 0 replies; 11+ messages in thread
From: yunhui cui @ 2025-09-23 2:15 UTC (permalink / raw)
To: Atish Kumar Patra
Cc: Anup Patel, Clément Léger, paul.walmsley, palmer, aou,
alex, conor, ajones, apatel, mchitale, linux-riscv, linux-kernel
Hi Atish,
On Tue, Sep 23, 2025 at 1:20 AM Atish Kumar Patra <atishp@rivosinc.com> wrote:
>
> On Mon, Sep 22, 2025 at 4:58 AM yunhui cui <cuiyunhui@bytedance.com> wrote:
> >
> > Hi Anup,
> >
> > On Mon, Sep 22, 2025 at 6:40 PM Anup Patel <anup@brainfault.org> wrote:
> > >
> > > On Mon, Sep 22, 2025 at 1:42 PM yunhui cui <cuiyunhui@bytedance.com> wrote:
> > > >
> > > > Hi Clément,
> > > >
> > > > On Fri, Sep 19, 2025 at 3:52 PM yunhui cui <cuiyunhui@bytedance.com> wrote:
> > > > >
> > > > > Hi Clément,
> > > > >
> > > > > On Fri, Sep 19, 2025 at 3:18 PM Clément Léger <cleger@rivosinc.com> wrote:
> > > > > >
> > > > > >
> > > > > >
> > > > > > On 19/09/2025 09:00, Yunhui Cui wrote:
> > > > > > > Unknown NMI can force the kernel to respond (e.g., panic) when the
> > > > > > > system encounters unrecognized critical hardware events, aiding in
> > > > > > > troubleshooting system faults. This is implemented via the Supervisor
> > > > > > > Software Events (SSE) framework.
> > > > > > >
> > > > > > > Signed-off-by: Yunhui Cui <cuiyunhui@bytedance.com>
> > > > > > > ---
> > > > > > > arch/riscv/include/asm/sbi.h | 1 +
> > > > > > > drivers/firmware/riscv/Kconfig | 10 +++++
> > > > > > > drivers/firmware/riscv/Makefile | 1 +
> > > > > > > drivers/firmware/riscv/sse_nmi.c | 77 ++++++++++++++++++++++++++++++++
> > > > > > > 4 files changed, 89 insertions(+)
> > > > > > > create mode 100644 drivers/firmware/riscv/sse_nmi.c
> > > > > > >
> > > > > > > diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h
> > > > > > > index 874cc1d7603a5..5801f90a88f62 100644
> > > > > > > --- a/arch/riscv/include/asm/sbi.h
> > > > > > > +++ b/arch/riscv/include/asm/sbi.h
> > > > > > > @@ -481,6 +481,7 @@ enum sbi_sse_attr_id {
> > > > > > >
> > > > > > > #define SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS 0x00000000
> > > > > > > #define SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP 0x00000001
> > > > > > > +#define SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI 0x00000002
> > > > > >
> > > > > > Was this submitted to the PRS WG ? This a specification modification so
> > > > > > it should go through the usual process.
> > > > > >
> > > > > > > #define SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS 0x00008000
> > > > > > > #define SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW 0x00010000
> > > > > > > #define SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS 0x00100000
> > > > > > > diff --git a/drivers/firmware/riscv/Kconfig b/drivers/firmware/riscv/Kconfig
> > > > > > > index ed5b663ac5f91..746bac862ac46 100644
> > > > > > > --- a/drivers/firmware/riscv/Kconfig
> > > > > > > +++ b/drivers/firmware/riscv/Kconfig
> > > > > > > @@ -12,4 +12,14 @@ config RISCV_SBI_SSE
> > > > > > > this option provides support to register callbacks on specific SSE
> > > > > > > events.
> > > > > > >
> > > > > > > +config RISCV_SSE_UNKNOWN_NMI
> > > > > > > + bool "Enable SBI Supervisor Software Events unknown NMI support"
> > > > > > > + depends on RISCV_SBI_SSE
> > > > > > > + default y
> > > > > > > + help
> > > > > > > + This option enables support for delivering unknown Non-Maskable Interrupt (NMI)
> > > > > > > + notifications via the Supervisor Software Events (SSE) framework. When enabled,
> > > > > > > + unknown NMIs can trigger kernel responses (e.g., panic) for unrecognized critical
> > > > > > > + hardware events, aiding in system fault diagnosis.
> > > > > > > +
> > > > > > > endmenu
> > > > > > > diff --git a/drivers/firmware/riscv/Makefile b/drivers/firmware/riscv/Makefile
> > > > > > > index c8795d4bbb2ea..9242c6cd5e3e9 100644
> > > > > > > --- a/drivers/firmware/riscv/Makefile
> > > > > > > +++ b/drivers/firmware/riscv/Makefile
> > > > > > > @@ -1,3 +1,4 @@
> > > > > > > # SPDX-License-Identifier: GPL-2.0
> > > > > > >
> > > > > > > obj-$(CONFIG_RISCV_SBI_SSE) += riscv_sbi_sse.o
> > > > > > > +obj-$(CONFIG_RISCV_SSE_UNKNOWN_NMI) += sse_nmi.o
> > > > > > > diff --git a/drivers/firmware/riscv/sse_nmi.c b/drivers/firmware/riscv/sse_nmi.c
> > > > > > > new file mode 100644
> > > > > > > index 0000000000000..43063f42efff0
> > > > > > > --- /dev/null
> > > > > > > +++ b/drivers/firmware/riscv/sse_nmi.c
> > > > > > > @@ -0,0 +1,77 @@
> > > > > > > +// SPDX-License-Identifier: GPL-2.0+
> > > > > > > +
> > > > > > > +#include <linux/mm.h>
> > > > > > > +#include <linux/nmi.h>
> > > > > > > +#include <linux/riscv_sbi_sse.h>
> > > > > > > +#include <linux/sched/debug.h>
> > > > > > > +#include <linux/sysctl.h>
> > > > > > > +
> > > > > > > +#include <asm/irq_regs.h>
> > > > > > > +#include <asm/sbi.h>
> > > > > > > +
> > > > > > > +int panic_on_unknown_nmi = 1;
> > > > > > > +struct sse_event *evt;
> > > > > > > +static struct ctl_table_header *unknown_nmi_sysctl_header;
> > > > > > > +
> > > > > > > +const struct ctl_table unknown_nmi_table[] = {
> > > > > > > + {
> > > > > > > + .procname = "panic_enable",
> > > > > > > + .data = &panic_on_unknown_nmi,
> > > > > > > + .maxlen = sizeof(int),
> > > > > > > + .mode = 0644,
> > > > > > > + .proc_handler = proc_dointvec_minmax,
> > > > > > > + .extra1 = SYSCTL_ZERO,
> > > > > > > + .extra2 = SYSCTL_ONE,
> > > > > > > + },
> > > > > > > +};
> > > > > > > +
> > > > > > > +static void nmi_handler(struct pt_regs *regs)
> > > > > > > +{
> > > > > > > + pr_emerg("NMI received for unknown on CPU %d.\n", smp_processor_id());
> > > > > > > +
> > > > > > > + if (panic_on_unknown_nmi)
> > > > > > > + nmi_panic(regs, "NMI: Not continuing");
> > > > > > > +
> > > > > > > + pr_emerg("Dazed and confused, but trying to continue\n");
> > > > > > > +}
> > > > > >
> > > > > > I'm dazed and confused as well ;) What's the point of this except
> > > > > > interrupting the kernel with a panic ? It seems like it's a better idea
> > > > > > to let the firmware handle that properly and display whatever
> > > > > > information are needed. Was your idea to actually force the kernel to
> > > > > > enter in some debug mode ?
> > > > >
> > > > > There is an important scenario: when the kernel becomes unresponsive,
> > > > > we need to trigger an unknown NMI to cause the system to panic() and
> > > > > then collect the vmcore, and such a requirement is common on x86
> > > > > servers.
> > > > >
> > > > > >
> > > > > > Thanks,
> > > > > >
> > > > > > Clément
> > > > > >
> > > > > > > +
> > > > > > > +static int nmi_sse_handler(u32 evt, void *arg, struct pt_regs *regs)
> > > > > > > +{
> > > > > > > + nmi_handler(regs);
> > > > > > > +
> > > > > > > + return 0;
> > > > > > > +}
> > > > > > > +
> > > > > > > +static int sse_nmi_init(void)
> > > > > > > +{
> > > > > > > + int ret;
> > > > > > > +
> > > > > > > + evt = sse_event_register(SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI, 0,
> > > > > > > + nmi_sse_handler, NULL);
> > > >
> > > > Should we add this UNKNOWN_NMI event ID in Chapter 17 of the SBI spec?
> > >
> > > The ratified SBI v3.0 defines a "Software injected local event" (ID 0xffff0000)
> > > which can be used to inject panic() on a particular HART from another HART.
> >
> > Has SBI_SSE_EVENT_LOCAL_SOFTWARE (0xffff0000) been given meaning
> > currently? The scenario of the unknown NMI is that after a HART
> > receives an interrupt from the platform, M-mode responds to this
>
> Which interrupt ? A standard one or platform specific one ?
Platform specific one.
Are there any recommended event IDs? Or should a new UNKNOWN NMI event
ID be added to the SBI spec?
>
> > interrupt and then enters S-mode to execute the registered handler.
> > Can SBI_SSE_EVENT_LOCAL_SOFTWARE be used directly?
> >
> >
> > >
> > > Regards,
> > > Anup
> >
> > Thanks,
> > Yunhui
Thanks,
Yunhui
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [External] Re: [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support
2025-09-22 11:58 ` yunhui cui
2025-09-22 17:20 ` Atish Kumar Patra
@ 2025-09-23 5:05 ` Anup Patel
1 sibling, 0 replies; 11+ messages in thread
From: Anup Patel @ 2025-09-23 5:05 UTC (permalink / raw)
To: yunhui cui
Cc: Clément Léger, paul.walmsley, palmer, aou, alex, conor,
atishp, ajones, apatel, mchitale, linux-riscv, linux-kernel
On Mon, Sep 22, 2025 at 5:28 PM yunhui cui <cuiyunhui@bytedance.com> wrote:
>
> Hi Anup,
>
> On Mon, Sep 22, 2025 at 6:40 PM Anup Patel <anup@brainfault.org> wrote:
> >
> > On Mon, Sep 22, 2025 at 1:42 PM yunhui cui <cuiyunhui@bytedance.com> wrote:
> > >
> > > Hi Clément,
> > >
> > > On Fri, Sep 19, 2025 at 3:52 PM yunhui cui <cuiyunhui@bytedance.com> wrote:
> > > >
> > > > Hi Clément,
> > > >
> > > > On Fri, Sep 19, 2025 at 3:18 PM Clément Léger <cleger@rivosinc.com> wrote:
> > > > >
> > > > >
> > > > >
> > > > > On 19/09/2025 09:00, Yunhui Cui wrote:
> > > > > > Unknown NMI can force the kernel to respond (e.g., panic) when the
> > > > > > system encounters unrecognized critical hardware events, aiding in
> > > > > > troubleshooting system faults. This is implemented via the Supervisor
> > > > > > Software Events (SSE) framework.
> > > > > >
> > > > > > Signed-off-by: Yunhui Cui <cuiyunhui@bytedance.com>
> > > > > > ---
> > > > > > arch/riscv/include/asm/sbi.h | 1 +
> > > > > > drivers/firmware/riscv/Kconfig | 10 +++++
> > > > > > drivers/firmware/riscv/Makefile | 1 +
> > > > > > drivers/firmware/riscv/sse_nmi.c | 77 ++++++++++++++++++++++++++++++++
> > > > > > 4 files changed, 89 insertions(+)
> > > > > > create mode 100644 drivers/firmware/riscv/sse_nmi.c
> > > > > >
> > > > > > diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h
> > > > > > index 874cc1d7603a5..5801f90a88f62 100644
> > > > > > --- a/arch/riscv/include/asm/sbi.h
> > > > > > +++ b/arch/riscv/include/asm/sbi.h
> > > > > > @@ -481,6 +481,7 @@ enum sbi_sse_attr_id {
> > > > > >
> > > > > > #define SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS 0x00000000
> > > > > > #define SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP 0x00000001
> > > > > > +#define SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI 0x00000002
> > > > >
> > > > > Was this submitted to the PRS WG ? This a specification modification so
> > > > > it should go through the usual process.
> > > > >
> > > > > > #define SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS 0x00008000
> > > > > > #define SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW 0x00010000
> > > > > > #define SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS 0x00100000
> > > > > > diff --git a/drivers/firmware/riscv/Kconfig b/drivers/firmware/riscv/Kconfig
> > > > > > index ed5b663ac5f91..746bac862ac46 100644
> > > > > > --- a/drivers/firmware/riscv/Kconfig
> > > > > > +++ b/drivers/firmware/riscv/Kconfig
> > > > > > @@ -12,4 +12,14 @@ config RISCV_SBI_SSE
> > > > > > this option provides support to register callbacks on specific SSE
> > > > > > events.
> > > > > >
> > > > > > +config RISCV_SSE_UNKNOWN_NMI
> > > > > > + bool "Enable SBI Supervisor Software Events unknown NMI support"
> > > > > > + depends on RISCV_SBI_SSE
> > > > > > + default y
> > > > > > + help
> > > > > > + This option enables support for delivering unknown Non-Maskable Interrupt (NMI)
> > > > > > + notifications via the Supervisor Software Events (SSE) framework. When enabled,
> > > > > > + unknown NMIs can trigger kernel responses (e.g., panic) for unrecognized critical
> > > > > > + hardware events, aiding in system fault diagnosis.
> > > > > > +
> > > > > > endmenu
> > > > > > diff --git a/drivers/firmware/riscv/Makefile b/drivers/firmware/riscv/Makefile
> > > > > > index c8795d4bbb2ea..9242c6cd5e3e9 100644
> > > > > > --- a/drivers/firmware/riscv/Makefile
> > > > > > +++ b/drivers/firmware/riscv/Makefile
> > > > > > @@ -1,3 +1,4 @@
> > > > > > # SPDX-License-Identifier: GPL-2.0
> > > > > >
> > > > > > obj-$(CONFIG_RISCV_SBI_SSE) += riscv_sbi_sse.o
> > > > > > +obj-$(CONFIG_RISCV_SSE_UNKNOWN_NMI) += sse_nmi.o
> > > > > > diff --git a/drivers/firmware/riscv/sse_nmi.c b/drivers/firmware/riscv/sse_nmi.c
> > > > > > new file mode 100644
> > > > > > index 0000000000000..43063f42efff0
> > > > > > --- /dev/null
> > > > > > +++ b/drivers/firmware/riscv/sse_nmi.c
> > > > > > @@ -0,0 +1,77 @@
> > > > > > +// SPDX-License-Identifier: GPL-2.0+
> > > > > > +
> > > > > > +#include <linux/mm.h>
> > > > > > +#include <linux/nmi.h>
> > > > > > +#include <linux/riscv_sbi_sse.h>
> > > > > > +#include <linux/sched/debug.h>
> > > > > > +#include <linux/sysctl.h>
> > > > > > +
> > > > > > +#include <asm/irq_regs.h>
> > > > > > +#include <asm/sbi.h>
> > > > > > +
> > > > > > +int panic_on_unknown_nmi = 1;
> > > > > > +struct sse_event *evt;
> > > > > > +static struct ctl_table_header *unknown_nmi_sysctl_header;
> > > > > > +
> > > > > > +const struct ctl_table unknown_nmi_table[] = {
> > > > > > + {
> > > > > > + .procname = "panic_enable",
> > > > > > + .data = &panic_on_unknown_nmi,
> > > > > > + .maxlen = sizeof(int),
> > > > > > + .mode = 0644,
> > > > > > + .proc_handler = proc_dointvec_minmax,
> > > > > > + .extra1 = SYSCTL_ZERO,
> > > > > > + .extra2 = SYSCTL_ONE,
> > > > > > + },
> > > > > > +};
> > > > > > +
> > > > > > +static void nmi_handler(struct pt_regs *regs)
> > > > > > +{
> > > > > > + pr_emerg("NMI received for unknown on CPU %d.\n", smp_processor_id());
> > > > > > +
> > > > > > + if (panic_on_unknown_nmi)
> > > > > > + nmi_panic(regs, "NMI: Not continuing");
> > > > > > +
> > > > > > + pr_emerg("Dazed and confused, but trying to continue\n");
> > > > > > +}
> > > > >
> > > > > I'm dazed and confused as well ;) What's the point of this except
> > > > > interrupting the kernel with a panic ? It seems like it's a better idea
> > > > > to let the firmware handle that properly and display whatever
> > > > > information are needed. Was your idea to actually force the kernel to
> > > > > enter in some debug mode ?
> > > >
> > > > There is an important scenario: when the kernel becomes unresponsive,
> > > > we need to trigger an unknown NMI to cause the system to panic() and
> > > > then collect the vmcore, and such a requirement is common on x86
> > > > servers.
> > > >
> > > > >
> > > > > Thanks,
> > > > >
> > > > > Clément
> > > > >
> > > > > > +
> > > > > > +static int nmi_sse_handler(u32 evt, void *arg, struct pt_regs *regs)
> > > > > > +{
> > > > > > + nmi_handler(regs);
> > > > > > +
> > > > > > + return 0;
> > > > > > +}
> > > > > > +
> > > > > > +static int sse_nmi_init(void)
> > > > > > +{
> > > > > > + int ret;
> > > > > > +
> > > > > > + evt = sse_event_register(SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI, 0,
> > > > > > + nmi_sse_handler, NULL);
> > >
> > > Should we add this UNKNOWN_NMI event ID in Chapter 17 of the SBI spec?
> >
> > The ratified SBI v3.0 defines a "Software injected local event" (ID 0xffff0000)
> > which can be used to inject panic() on a particular HART from another HART.
>
> Has SBI_SSE_EVENT_LOCAL_SOFTWARE (0xffff0000) been given meaning
> currently? The scenario of the unknown NMI is that after a HART
> receives an interrupt from the platform, M-mode responds to this
> interrupt and then enters S-mode to execute the registered handler.
> Can SBI_SSE_EVENT_LOCAL_SOFTWARE be used directly?
It's a software injected event with lower priority which can be used by
supervisor software on a HART to inject an event to another HART.
If you need a platform specific event then there are multiple ranges
with different priority-levels.
Regards,
Anup
^ permalink raw reply [flat|nested] 11+ messages in thread
end of thread, other threads:[~2025-09-23 5:06 UTC | newest]
Thread overview: 11+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2025-09-19 7:00 [PATCH RFC 0/1] Support unknown NMI based on SSE Yunhui Cui
2025-09-19 7:00 ` [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support Yunhui Cui
2025-09-19 7:17 ` Clément Léger
2025-09-19 7:52 ` [External] " yunhui cui
2025-09-22 8:11 ` yunhui cui
2025-09-22 8:31 ` Clément Léger
2025-09-22 10:39 ` Anup Patel
2025-09-22 11:58 ` yunhui cui
2025-09-22 17:20 ` Atish Kumar Patra
2025-09-23 2:15 ` yunhui cui
2025-09-23 5:05 ` Anup Patel
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®