* [PATCH RFC 0/1] Support unknown NMI based on SSE @ 2025-09-19 7:00 Yunhui Cui 2025-09-19 7:00 ` [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support Yunhui Cui 0 siblings, 1 reply; 11+ messages in thread From: Yunhui Cui @ 2025-09-19 7:00 UTC (permalink / raw) To: paul.walmsley, palmer, aou, alex, conor, atishp, cleger, ajones, apatel, cuiyunhui, mchitale, linux-riscv, linux-kernel This patch is based on the SSE features by Clément Léger: https://lore.kernel.org/all/20250908181717.1997461-1-cleger@rivosinc.com/ It mainly adds support for unknown NMI functionality. Since extension development has not yet been implemented in SBI, an RFC version is submitted first. Yunhui Cui (1): drivers: firmware: riscv: add unknown NMI support arch/riscv/include/asm/sbi.h | 1 + drivers/firmware/riscv/Kconfig | 10 +++++ drivers/firmware/riscv/Makefile | 1 + drivers/firmware/riscv/sse_nmi.c | 77 ++++++++++++++++++++++++++++++++ 4 files changed, 89 insertions(+) create mode 100644 drivers/firmware/riscv/sse_nmi.c -- 2.39.5 ^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support 2025-09-19 7:00 [PATCH RFC 0/1] Support unknown NMI based on SSE Yunhui Cui @ 2025-09-19 7:00 ` Yunhui Cui 2025-09-19 7:17 ` Clément Léger 0 siblings, 1 reply; 11+ messages in thread From: Yunhui Cui @ 2025-09-19 7:00 UTC (permalink / raw) To: paul.walmsley, palmer, aou, alex, conor, atishp, cleger, ajones, apatel, cuiyunhui, mchitale, linux-riscv, linux-kernel Unknown NMI can force the kernel to respond (e.g., panic) when the system encounters unrecognized critical hardware events, aiding in troubleshooting system faults. This is implemented via the Supervisor Software Events (SSE) framework. Signed-off-by: Yunhui Cui <cuiyunhui@bytedance.com> --- arch/riscv/include/asm/sbi.h | 1 + drivers/firmware/riscv/Kconfig | 10 +++++ drivers/firmware/riscv/Makefile | 1 + drivers/firmware/riscv/sse_nmi.c | 77 ++++++++++++++++++++++++++++++++ 4 files changed, 89 insertions(+) create mode 100644 drivers/firmware/riscv/sse_nmi.c diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h index 874cc1d7603a5..5801f90a88f62 100644 --- a/arch/riscv/include/asm/sbi.h +++ b/arch/riscv/include/asm/sbi.h @@ -481,6 +481,7 @@ enum sbi_sse_attr_id { #define SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS 0x00000000 #define SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP 0x00000001 +#define SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI 0x00000002 #define SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS 0x00008000 #define SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW 0x00010000 #define SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS 0x00100000 diff --git a/drivers/firmware/riscv/Kconfig b/drivers/firmware/riscv/Kconfig index ed5b663ac5f91..746bac862ac46 100644 --- a/drivers/firmware/riscv/Kconfig +++ b/drivers/firmware/riscv/Kconfig @@ -12,4 +12,14 @@ config RISCV_SBI_SSE this option provides support to register callbacks on specific SSE events. +config RISCV_SSE_UNKNOWN_NMI + bool "Enable SBI Supervisor Software Events unknown NMI support" + depends on RISCV_SBI_SSE + default y + help + This option enables support for delivering unknown Non-Maskable Interrupt (NMI) + notifications via the Supervisor Software Events (SSE) framework. When enabled, + unknown NMIs can trigger kernel responses (e.g., panic) for unrecognized critical + hardware events, aiding in system fault diagnosis. + endmenu diff --git a/drivers/firmware/riscv/Makefile b/drivers/firmware/riscv/Makefile index c8795d4bbb2ea..9242c6cd5e3e9 100644 --- a/drivers/firmware/riscv/Makefile +++ b/drivers/firmware/riscv/Makefile @@ -1,3 +1,4 @@ # SPDX-License-Identifier: GPL-2.0 obj-$(CONFIG_RISCV_SBI_SSE) += riscv_sbi_sse.o +obj-$(CONFIG_RISCV_SSE_UNKNOWN_NMI) += sse_nmi.o diff --git a/drivers/firmware/riscv/sse_nmi.c b/drivers/firmware/riscv/sse_nmi.c new file mode 100644 index 0000000000000..43063f42efff0 --- /dev/null +++ b/drivers/firmware/riscv/sse_nmi.c @@ -0,0 +1,77 @@ +// SPDX-License-Identifier: GPL-2.0+ + +#include <linux/mm.h> +#include <linux/nmi.h> +#include <linux/riscv_sbi_sse.h> +#include <linux/sched/debug.h> +#include <linux/sysctl.h> + +#include <asm/irq_regs.h> +#include <asm/sbi.h> + +int panic_on_unknown_nmi = 1; +struct sse_event *evt; +static struct ctl_table_header *unknown_nmi_sysctl_header; + +const struct ctl_table unknown_nmi_table[] = { + { + .procname = "panic_enable", + .data = &panic_on_unknown_nmi, + .maxlen = sizeof(int), + .mode = 0644, + .proc_handler = proc_dointvec_minmax, + .extra1 = SYSCTL_ZERO, + .extra2 = SYSCTL_ONE, + }, +}; + +static void nmi_handler(struct pt_regs *regs) +{ + pr_emerg("NMI received for unknown on CPU %d.\n", smp_processor_id()); + + if (panic_on_unknown_nmi) + nmi_panic(regs, "NMI: Not continuing"); + + pr_emerg("Dazed and confused, but trying to continue\n"); +} + +static int nmi_sse_handler(u32 evt, void *arg, struct pt_regs *regs) +{ + nmi_handler(regs); + + return 0; +} + +static int sse_nmi_init(void) +{ + int ret; + + evt = sse_event_register(SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI, 0, + nmi_sse_handler, NULL); + if (IS_ERR(evt)) + return PTR_ERR(evt); + + ret = sse_event_enable(evt); + if (ret) { + sse_event_unregister(evt); + return ret; + } + + unknown_nmi_sysctl_header = register_sysctl("kernel", unknown_nmi_table); + if (!unknown_nmi_sysctl_header) { + sse_event_disable(evt); + sse_event_unregister(evt); + return -ENOMEM; + } + + pr_info("Using SSE for NMI event delivery\n"); + + return 0; +} + +static int __init unknow_nmi_init(void) +{ + return sse_nmi_init(); +} + +late_initcall(unknow_nmi_init); -- 2.39.5 ^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support 2025-09-19 7:00 ` [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support Yunhui Cui @ 2025-09-19 7:17 ` Clément Léger 2025-09-19 7:52 ` [External] " yunhui cui 0 siblings, 1 reply; 11+ messages in thread From: Clément Léger @ 2025-09-19 7:17 UTC (permalink / raw) To: Yunhui Cui, paul.walmsley, palmer, aou, alex, conor, atishp, ajones, apatel, mchitale, linux-riscv, linux-kernel On 19/09/2025 09:00, Yunhui Cui wrote: > Unknown NMI can force the kernel to respond (e.g., panic) when the > system encounters unrecognized critical hardware events, aiding in > troubleshooting system faults. This is implemented via the Supervisor > Software Events (SSE) framework. > > Signed-off-by: Yunhui Cui <cuiyunhui@bytedance.com> > --- > arch/riscv/include/asm/sbi.h | 1 + > drivers/firmware/riscv/Kconfig | 10 +++++ > drivers/firmware/riscv/Makefile | 1 + > drivers/firmware/riscv/sse_nmi.c | 77 ++++++++++++++++++++++++++++++++ > 4 files changed, 89 insertions(+) > create mode 100644 drivers/firmware/riscv/sse_nmi.c > > diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h > index 874cc1d7603a5..5801f90a88f62 100644 > --- a/arch/riscv/include/asm/sbi.h > +++ b/arch/riscv/include/asm/sbi.h > @@ -481,6 +481,7 @@ enum sbi_sse_attr_id { > > #define SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS 0x00000000 > #define SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP 0x00000001 > +#define SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI 0x00000002 Was this submitted to the PRS WG ? This a specification modification so it should go through the usual process. > #define SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS 0x00008000 > #define SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW 0x00010000 > #define SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS 0x00100000 > diff --git a/drivers/firmware/riscv/Kconfig b/drivers/firmware/riscv/Kconfig > index ed5b663ac5f91..746bac862ac46 100644 > --- a/drivers/firmware/riscv/Kconfig > +++ b/drivers/firmware/riscv/Kconfig > @@ -12,4 +12,14 @@ config RISCV_SBI_SSE > this option provides support to register callbacks on specific SSE > events. > > +config RISCV_SSE_UNKNOWN_NMI > + bool "Enable SBI Supervisor Software Events unknown NMI support" > + depends on RISCV_SBI_SSE > + default y > + help > + This option enables support for delivering unknown Non-Maskable Interrupt (NMI) > + notifications via the Supervisor Software Events (SSE) framework. When enabled, > + unknown NMIs can trigger kernel responses (e.g., panic) for unrecognized critical > + hardware events, aiding in system fault diagnosis. > + > endmenu > diff --git a/drivers/firmware/riscv/Makefile b/drivers/firmware/riscv/Makefile > index c8795d4bbb2ea..9242c6cd5e3e9 100644 > --- a/drivers/firmware/riscv/Makefile > +++ b/drivers/firmware/riscv/Makefile > @@ -1,3 +1,4 @@ > # SPDX-License-Identifier: GPL-2.0 > > obj-$(CONFIG_RISCV_SBI_SSE) += riscv_sbi_sse.o > +obj-$(CONFIG_RISCV_SSE_UNKNOWN_NMI) += sse_nmi.o > diff --git a/drivers/firmware/riscv/sse_nmi.c b/drivers/firmware/riscv/sse_nmi.c > new file mode 100644 > index 0000000000000..43063f42efff0 > --- /dev/null > +++ b/drivers/firmware/riscv/sse_nmi.c > @@ -0,0 +1,77 @@ > +// SPDX-License-Identifier: GPL-2.0+ > + > +#include <linux/mm.h> > +#include <linux/nmi.h> > +#include <linux/riscv_sbi_sse.h> > +#include <linux/sched/debug.h> > +#include <linux/sysctl.h> > + > +#include <asm/irq_regs.h> > +#include <asm/sbi.h> > + > +int panic_on_unknown_nmi = 1; > +struct sse_event *evt; > +static struct ctl_table_header *unknown_nmi_sysctl_header; > + > +const struct ctl_table unknown_nmi_table[] = { > + { > + .procname = "panic_enable", > + .data = &panic_on_unknown_nmi, > + .maxlen = sizeof(int), > + .mode = 0644, > + .proc_handler = proc_dointvec_minmax, > + .extra1 = SYSCTL_ZERO, > + .extra2 = SYSCTL_ONE, > + }, > +}; > + > +static void nmi_handler(struct pt_regs *regs) > +{ > + pr_emerg("NMI received for unknown on CPU %d.\n", smp_processor_id()); > + > + if (panic_on_unknown_nmi) > + nmi_panic(regs, "NMI: Not continuing"); > + > + pr_emerg("Dazed and confused, but trying to continue\n"); > +} I'm dazed and confused as well ;) What's the point of this except interrupting the kernel with a panic ? It seems like it's a better idea to let the firmware handle that properly and display whatever information are needed. Was your idea to actually force the kernel to enter in some debug mode ? Thanks, Clément > + > +static int nmi_sse_handler(u32 evt, void *arg, struct pt_regs *regs) > +{ > + nmi_handler(regs); > + > + return 0; > +} > + > +static int sse_nmi_init(void) > +{ > + int ret; > + > + evt = sse_event_register(SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI, 0, > + nmi_sse_handler, NULL); > + if (IS_ERR(evt)) > + return PTR_ERR(evt); > + > + ret = sse_event_enable(evt); > + if (ret) { > + sse_event_unregister(evt); > + return ret; > + } > + > + unknown_nmi_sysctl_header = register_sysctl("kernel", unknown_nmi_table); > + if (!unknown_nmi_sysctl_header) { > + sse_event_disable(evt); > + sse_event_unregister(evt); > + return -ENOMEM; > + } > + > + pr_info("Using SSE for NMI event delivery\n"); > + > + return 0; > +} > + > +static int __init unknow_nmi_init(void) > +{ > + return sse_nmi_init(); > +} > + > +late_initcall(unknow_nmi_init); ^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [External] Re: [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support 2025-09-19 7:17 ` Clément Léger @ 2025-09-19 7:52 ` yunhui cui 2025-09-22 8:11 ` yunhui cui 0 siblings, 1 reply; 11+ messages in thread From: yunhui cui @ 2025-09-19 7:52 UTC (permalink / raw) To: Clément Léger Cc: paul.walmsley, palmer, aou, alex, conor, atishp, ajones, apatel, mchitale, linux-riscv, linux-kernel Hi Clément, On Fri, Sep 19, 2025 at 3:18 PM Clément Léger <cleger@rivosinc.com> wrote: > > > > On 19/09/2025 09:00, Yunhui Cui wrote: > > Unknown NMI can force the kernel to respond (e.g., panic) when the > > system encounters unrecognized critical hardware events, aiding in > > troubleshooting system faults. This is implemented via the Supervisor > > Software Events (SSE) framework. > > > > Signed-off-by: Yunhui Cui <cuiyunhui@bytedance.com> > > --- > > arch/riscv/include/asm/sbi.h | 1 + > > drivers/firmware/riscv/Kconfig | 10 +++++ > > drivers/firmware/riscv/Makefile | 1 + > > drivers/firmware/riscv/sse_nmi.c | 77 ++++++++++++++++++++++++++++++++ > > 4 files changed, 89 insertions(+) > > create mode 100644 drivers/firmware/riscv/sse_nmi.c > > > > diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h > > index 874cc1d7603a5..5801f90a88f62 100644 > > --- a/arch/riscv/include/asm/sbi.h > > +++ b/arch/riscv/include/asm/sbi.h > > @@ -481,6 +481,7 @@ enum sbi_sse_attr_id { > > > > #define SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS 0x00000000 > > #define SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP 0x00000001 > > +#define SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI 0x00000002 > > Was this submitted to the PRS WG ? This a specification modification so > it should go through the usual process. > > > #define SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS 0x00008000 > > #define SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW 0x00010000 > > #define SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS 0x00100000 > > diff --git a/drivers/firmware/riscv/Kconfig b/drivers/firmware/riscv/Kconfig > > index ed5b663ac5f91..746bac862ac46 100644 > > --- a/drivers/firmware/riscv/Kconfig > > +++ b/drivers/firmware/riscv/Kconfig > > @@ -12,4 +12,14 @@ config RISCV_SBI_SSE > > this option provides support to register callbacks on specific SSE > > events. > > > > +config RISCV_SSE_UNKNOWN_NMI > > + bool "Enable SBI Supervisor Software Events unknown NMI support" > > + depends on RISCV_SBI_SSE > > + default y > > + help > > + This option enables support for delivering unknown Non-Maskable Interrupt (NMI) > > + notifications via the Supervisor Software Events (SSE) framework. When enabled, > > + unknown NMIs can trigger kernel responses (e.g., panic) for unrecognized critical > > + hardware events, aiding in system fault diagnosis. > > + > > endmenu > > diff --git a/drivers/firmware/riscv/Makefile b/drivers/firmware/riscv/Makefile > > index c8795d4bbb2ea..9242c6cd5e3e9 100644 > > --- a/drivers/firmware/riscv/Makefile > > +++ b/drivers/firmware/riscv/Makefile > > @@ -1,3 +1,4 @@ > > # SPDX-License-Identifier: GPL-2.0 > > > > obj-$(CONFIG_RISCV_SBI_SSE) += riscv_sbi_sse.o > > +obj-$(CONFIG_RISCV_SSE_UNKNOWN_NMI) += sse_nmi.o > > diff --git a/drivers/firmware/riscv/sse_nmi.c b/drivers/firmware/riscv/sse_nmi.c > > new file mode 100644 > > index 0000000000000..43063f42efff0 > > --- /dev/null > > +++ b/drivers/firmware/riscv/sse_nmi.c > > @@ -0,0 +1,77 @@ > > +// SPDX-License-Identifier: GPL-2.0+ > > + > > +#include <linux/mm.h> > > +#include <linux/nmi.h> > > +#include <linux/riscv_sbi_sse.h> > > +#include <linux/sched/debug.h> > > +#include <linux/sysctl.h> > > + > > +#include <asm/irq_regs.h> > > +#include <asm/sbi.h> > > + > > +int panic_on_unknown_nmi = 1; > > +struct sse_event *evt; > > +static struct ctl_table_header *unknown_nmi_sysctl_header; > > + > > +const struct ctl_table unknown_nmi_table[] = { > > + { > > + .procname = "panic_enable", > > + .data = &panic_on_unknown_nmi, > > + .maxlen = sizeof(int), > > + .mode = 0644, > > + .proc_handler = proc_dointvec_minmax, > > + .extra1 = SYSCTL_ZERO, > > + .extra2 = SYSCTL_ONE, > > + }, > > +}; > > + > > +static void nmi_handler(struct pt_regs *regs) > > +{ > > + pr_emerg("NMI received for unknown on CPU %d.\n", smp_processor_id()); > > + > > + if (panic_on_unknown_nmi) > > + nmi_panic(regs, "NMI: Not continuing"); > > + > > + pr_emerg("Dazed and confused, but trying to continue\n"); > > +} > > I'm dazed and confused as well ;) What's the point of this except > interrupting the kernel with a panic ? It seems like it's a better idea > to let the firmware handle that properly and display whatever > information are needed. Was your idea to actually force the kernel to > enter in some debug mode ? There is an important scenario: when the kernel becomes unresponsive, we need to trigger an unknown NMI to cause the system to panic() and then collect the vmcore, and such a requirement is common on x86 servers. > > Thanks, > > Clément > > > + > > +static int nmi_sse_handler(u32 evt, void *arg, struct pt_regs *regs) > > +{ > > + nmi_handler(regs); > > + > > + return 0; > > +} > > + > > +static int sse_nmi_init(void) > > +{ > > + int ret; > > + > > + evt = sse_event_register(SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI, 0, > > + nmi_sse_handler, NULL); > > + if (IS_ERR(evt)) > > + return PTR_ERR(evt); > > + > > + ret = sse_event_enable(evt); > > + if (ret) { > > + sse_event_unregister(evt); > > + return ret; > > + } > > + > > + unknown_nmi_sysctl_header = register_sysctl("kernel", unknown_nmi_table); > > + if (!unknown_nmi_sysctl_header) { > > + sse_event_disable(evt); > > + sse_event_unregister(evt); > > + return -ENOMEM; > > + } > > + > > + pr_info("Using SSE for NMI event delivery\n"); > > + > > + return 0; > > +} > > + > > +static int __init unknow_nmi_init(void) > > +{ > > + return sse_nmi_init(); > > +} > > + > > +late_initcall(unknow_nmi_init); > Thanks, Yunhui ^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [External] Re: [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support 2025-09-19 7:52 ` [External] " yunhui cui @ 2025-09-22 8:11 ` yunhui cui 2025-09-22 8:31 ` Clément Léger 2025-09-22 10:39 ` Anup Patel 0 siblings, 2 replies; 11+ messages in thread From: yunhui cui @ 2025-09-22 8:11 UTC (permalink / raw) To: Clément Léger Cc: paul.walmsley, palmer, aou, alex, conor, atishp, ajones, apatel, mchitale, linux-riscv, linux-kernel Hi Clément, On Fri, Sep 19, 2025 at 3:52 PM yunhui cui <cuiyunhui@bytedance.com> wrote: > > Hi Clément, > > On Fri, Sep 19, 2025 at 3:18 PM Clément Léger <cleger@rivosinc.com> wrote: > > > > > > > > On 19/09/2025 09:00, Yunhui Cui wrote: > > > Unknown NMI can force the kernel to respond (e.g., panic) when the > > > system encounters unrecognized critical hardware events, aiding in > > > troubleshooting system faults. This is implemented via the Supervisor > > > Software Events (SSE) framework. > > > > > > Signed-off-by: Yunhui Cui <cuiyunhui@bytedance.com> > > > --- > > > arch/riscv/include/asm/sbi.h | 1 + > > > drivers/firmware/riscv/Kconfig | 10 +++++ > > > drivers/firmware/riscv/Makefile | 1 + > > > drivers/firmware/riscv/sse_nmi.c | 77 ++++++++++++++++++++++++++++++++ > > > 4 files changed, 89 insertions(+) > > > create mode 100644 drivers/firmware/riscv/sse_nmi.c > > > > > > diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h > > > index 874cc1d7603a5..5801f90a88f62 100644 > > > --- a/arch/riscv/include/asm/sbi.h > > > +++ b/arch/riscv/include/asm/sbi.h > > > @@ -481,6 +481,7 @@ enum sbi_sse_attr_id { > > > > > > #define SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS 0x00000000 > > > #define SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP 0x00000001 > > > +#define SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI 0x00000002 > > > > Was this submitted to the PRS WG ? This a specification modification so > > it should go through the usual process. > > > > > #define SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS 0x00008000 > > > #define SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW 0x00010000 > > > #define SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS 0x00100000 > > > diff --git a/drivers/firmware/riscv/Kconfig b/drivers/firmware/riscv/Kconfig > > > index ed5b663ac5f91..746bac862ac46 100644 > > > --- a/drivers/firmware/riscv/Kconfig > > > +++ b/drivers/firmware/riscv/Kconfig > > > @@ -12,4 +12,14 @@ config RISCV_SBI_SSE > > > this option provides support to register callbacks on specific SSE > > > events. > > > > > > +config RISCV_SSE_UNKNOWN_NMI > > > + bool "Enable SBI Supervisor Software Events unknown NMI support" > > > + depends on RISCV_SBI_SSE > > > + default y > > > + help > > > + This option enables support for delivering unknown Non-Maskable Interrupt (NMI) > > > + notifications via the Supervisor Software Events (SSE) framework. When enabled, > > > + unknown NMIs can trigger kernel responses (e.g., panic) for unrecognized critical > > > + hardware events, aiding in system fault diagnosis. > > > + > > > endmenu > > > diff --git a/drivers/firmware/riscv/Makefile b/drivers/firmware/riscv/Makefile > > > index c8795d4bbb2ea..9242c6cd5e3e9 100644 > > > --- a/drivers/firmware/riscv/Makefile > > > +++ b/drivers/firmware/riscv/Makefile > > > @@ -1,3 +1,4 @@ > > > # SPDX-License-Identifier: GPL-2.0 > > > > > > obj-$(CONFIG_RISCV_SBI_SSE) += riscv_sbi_sse.o > > > +obj-$(CONFIG_RISCV_SSE_UNKNOWN_NMI) += sse_nmi.o > > > diff --git a/drivers/firmware/riscv/sse_nmi.c b/drivers/firmware/riscv/sse_nmi.c > > > new file mode 100644 > > > index 0000000000000..43063f42efff0 > > > --- /dev/null > > > +++ b/drivers/firmware/riscv/sse_nmi.c > > > @@ -0,0 +1,77 @@ > > > +// SPDX-License-Identifier: GPL-2.0+ > > > + > > > +#include <linux/mm.h> > > > +#include <linux/nmi.h> > > > +#include <linux/riscv_sbi_sse.h> > > > +#include <linux/sched/debug.h> > > > +#include <linux/sysctl.h> > > > + > > > +#include <asm/irq_regs.h> > > > +#include <asm/sbi.h> > > > + > > > +int panic_on_unknown_nmi = 1; > > > +struct sse_event *evt; > > > +static struct ctl_table_header *unknown_nmi_sysctl_header; > > > + > > > +const struct ctl_table unknown_nmi_table[] = { > > > + { > > > + .procname = "panic_enable", > > > + .data = &panic_on_unknown_nmi, > > > + .maxlen = sizeof(int), > > > + .mode = 0644, > > > + .proc_handler = proc_dointvec_minmax, > > > + .extra1 = SYSCTL_ZERO, > > > + .extra2 = SYSCTL_ONE, > > > + }, > > > +}; > > > + > > > +static void nmi_handler(struct pt_regs *regs) > > > +{ > > > + pr_emerg("NMI received for unknown on CPU %d.\n", smp_processor_id()); > > > + > > > + if (panic_on_unknown_nmi) > > > + nmi_panic(regs, "NMI: Not continuing"); > > > + > > > + pr_emerg("Dazed and confused, but trying to continue\n"); > > > +} > > > > I'm dazed and confused as well ;) What's the point of this except > > interrupting the kernel with a panic ? It seems like it's a better idea > > to let the firmware handle that properly and display whatever > > information are needed. Was your idea to actually force the kernel to > > enter in some debug mode ? > > There is an important scenario: when the kernel becomes unresponsive, > we need to trigger an unknown NMI to cause the system to panic() and > then collect the vmcore, and such a requirement is common on x86 > servers. > > > > > Thanks, > > > > Clément > > > > > + > > > +static int nmi_sse_handler(u32 evt, void *arg, struct pt_regs *regs) > > > +{ > > > + nmi_handler(regs); > > > + > > > + return 0; > > > +} > > > + > > > +static int sse_nmi_init(void) > > > +{ > > > + int ret; > > > + > > > + evt = sse_event_register(SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI, 0, > > > + nmi_sse_handler, NULL); Should we add this UNKNOWN_NMI event ID in Chapter 17 of the SBI spec? > > > + if (IS_ERR(evt)) > > > + return PTR_ERR(evt); > > > + > > > + ret = sse_event_enable(evt); > > > + if (ret) { > > > + sse_event_unregister(evt); > > > + return ret; > > > + } > > > + > > > + unknown_nmi_sysctl_header = register_sysctl("kernel", unknown_nmi_table); > > > + if (!unknown_nmi_sysctl_header) { > > > + sse_event_disable(evt); > > > + sse_event_unregister(evt); > > > + return -ENOMEM; > > > + } > > > + > > > + pr_info("Using SSE for NMI event delivery\n"); > > > + > > > + return 0; > > > +} > > > + > > > +static int __init unknow_nmi_init(void) > > > +{ > > > + return sse_nmi_init(); > > > +} > > > + > > > +late_initcall(unknow_nmi_init); > > > > Thanks, > Yunhui Thanks, Yunhui ^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [External] Re: [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support 2025-09-22 8:11 ` yunhui cui @ 2025-09-22 8:31 ` Clément Léger 2025-09-22 10:39 ` Anup Patel 1 sibling, 0 replies; 11+ messages in thread From: Clément Léger @ 2025-09-22 8:31 UTC (permalink / raw) To: yunhui cui Cc: paul.walmsley, palmer, aou, alex, conor, atishp, ajones, apatel, mchitale, linux-riscv, linux-kernel On 22/09/2025 10:11, yunhui cui wrote: > Hi Clément, > > On Fri, Sep 19, 2025 at 3:52 PM yunhui cui <cuiyunhui@bytedance.com> wrote: >> >> Hi Clément, >> >> On Fri, Sep 19, 2025 at 3:18 PM Clément Léger <cleger@rivosinc.com> wrote: >>> >>> >>> >>> On 19/09/2025 09:00, Yunhui Cui wrote: >>>> Unknown NMI can force the kernel to respond (e.g., panic) when the >>>> system encounters unrecognized critical hardware events, aiding in >>>> troubleshooting system faults. This is implemented via the Supervisor >>>> Software Events (SSE) framework. >>>> >>>> Signed-off-by: Yunhui Cui <cuiyunhui@bytedance.com> >>>> --- >>>> arch/riscv/include/asm/sbi.h | 1 + >>>> drivers/firmware/riscv/Kconfig | 10 +++++ >>>> drivers/firmware/riscv/Makefile | 1 + >>>> drivers/firmware/riscv/sse_nmi.c | 77 ++++++++++++++++++++++++++++++++ >>>> 4 files changed, 89 insertions(+) >>>> create mode 100644 drivers/firmware/riscv/sse_nmi.c >>>> >>>> diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h >>>> index 874cc1d7603a5..5801f90a88f62 100644 >>>> --- a/arch/riscv/include/asm/sbi.h >>>> +++ b/arch/riscv/include/asm/sbi.h >>>> @@ -481,6 +481,7 @@ enum sbi_sse_attr_id { >>>> >>>> #define SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS 0x00000000 >>>> #define SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP 0x00000001 >>>> +#define SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI 0x00000002 >>> >>> Was this submitted to the PRS WG ? This a specification modification so >>> it should go through the usual process. >>> >>>> #define SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS 0x00008000 >>>> #define SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW 0x00010000 >>>> #define SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS 0x00100000 >>>> diff --git a/drivers/firmware/riscv/Kconfig b/drivers/firmware/riscv/Kconfig >>>> index ed5b663ac5f91..746bac862ac46 100644 >>>> --- a/drivers/firmware/riscv/Kconfig >>>> +++ b/drivers/firmware/riscv/Kconfig >>>> @@ -12,4 +12,14 @@ config RISCV_SBI_SSE >>>> this option provides support to register callbacks on specific SSE >>>> events. >>>> >>>> +config RISCV_SSE_UNKNOWN_NMI >>>> + bool "Enable SBI Supervisor Software Events unknown NMI support" >>>> + depends on RISCV_SBI_SSE >>>> + default y >>>> + help >>>> + This option enables support for delivering unknown Non-Maskable Interrupt (NMI) >>>> + notifications via the Supervisor Software Events (SSE) framework. When enabled, >>>> + unknown NMIs can trigger kernel responses (e.g., panic) for unrecognized critical >>>> + hardware events, aiding in system fault diagnosis. >>>> + >>>> endmenu >>>> diff --git a/drivers/firmware/riscv/Makefile b/drivers/firmware/riscv/Makefile >>>> index c8795d4bbb2ea..9242c6cd5e3e9 100644 >>>> --- a/drivers/firmware/riscv/Makefile >>>> +++ b/drivers/firmware/riscv/Makefile >>>> @@ -1,3 +1,4 @@ >>>> # SPDX-License-Identifier: GPL-2.0 >>>> >>>> obj-$(CONFIG_RISCV_SBI_SSE) += riscv_sbi_sse.o >>>> +obj-$(CONFIG_RISCV_SSE_UNKNOWN_NMI) += sse_nmi.o >>>> diff --git a/drivers/firmware/riscv/sse_nmi.c b/drivers/firmware/riscv/sse_nmi.c >>>> new file mode 100644 >>>> index 0000000000000..43063f42efff0 >>>> --- /dev/null >>>> +++ b/drivers/firmware/riscv/sse_nmi.c >>>> @@ -0,0 +1,77 @@ >>>> +// SPDX-License-Identifier: GPL-2.0+ >>>> + >>>> +#include <linux/mm.h> >>>> +#include <linux/nmi.h> >>>> +#include <linux/riscv_sbi_sse.h> >>>> +#include <linux/sched/debug.h> >>>> +#include <linux/sysctl.h> >>>> + >>>> +#include <asm/irq_regs.h> >>>> +#include <asm/sbi.h> >>>> + >>>> +int panic_on_unknown_nmi = 1; >>>> +struct sse_event *evt; >>>> +static struct ctl_table_header *unknown_nmi_sysctl_header; >>>> + >>>> +const struct ctl_table unknown_nmi_table[] = { >>>> + { >>>> + .procname = "panic_enable", >>>> + .data = &panic_on_unknown_nmi, >>>> + .maxlen = sizeof(int), >>>> + .mode = 0644, >>>> + .proc_handler = proc_dointvec_minmax, >>>> + .extra1 = SYSCTL_ZERO, >>>> + .extra2 = SYSCTL_ONE, >>>> + }, >>>> +}; >>>> + >>>> +static void nmi_handler(struct pt_regs *regs) >>>> +{ >>>> + pr_emerg("NMI received for unknown on CPU %d.\n", smp_processor_id()); >>>> + >>>> + if (panic_on_unknown_nmi) >>>> + nmi_panic(regs, "NMI: Not continuing"); >>>> + >>>> + pr_emerg("Dazed and confused, but trying to continue\n"); >>>> +} >>> >>> I'm dazed and confused as well ;) What's the point of this except >>> interrupting the kernel with a panic ? It seems like it's a better idea >>> to let the firmware handle that properly and display whatever >>> information are needed. Was your idea to actually force the kernel to >>> enter in some debug mode ? >> >> There is an important scenario: when the kernel becomes unresponsive, >> we need to trigger an unknown NMI to cause the system to panic() and >> then collect the vmcore, and such a requirement is common on x86 >> servers. >> >>> >>> Thanks, >>> >>> Clément >>> >>>> + >>>> +static int nmi_sse_handler(u32 evt, void *arg, struct pt_regs *regs) >>>> +{ >>>> + nmi_handler(regs); >>>> + >>>> + return 0; >>>> +} >>>> + >>>> +static int sse_nmi_init(void) >>>> +{ >>>> + int ret; >>>> + >>>> + evt = sse_event_register(SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI, 0, >>>> + nmi_sse_handler, NULL); > > Should we add this UNKNOWN_NMI event ID in Chapter 17 of the SBI spec? Hi Yunhui, If you want it to be part of the spec, that should indeed be submitted for review. Thanks, Clément > > >>>> + if (IS_ERR(evt)) >>>> + return PTR_ERR(evt); >>>> + >>>> + ret = sse_event_enable(evt); >>>> + if (ret) { >>>> + sse_event_unregister(evt); >>>> + return ret; >>>> + } >>>> + >>>> + unknown_nmi_sysctl_header = register_sysctl("kernel", unknown_nmi_table); >>>> + if (!unknown_nmi_sysctl_header) { >>>> + sse_event_disable(evt); >>>> + sse_event_unregister(evt); >>>> + return -ENOMEM; >>>> + } >>>> + >>>> + pr_info("Using SSE for NMI event delivery\n"); >>>> + >>>> + return 0; >>>> +} >>>> + >>>> +static int __init unknow_nmi_init(void) >>>> +{ >>>> + return sse_nmi_init(); >>>> +} >>>> + >>>> +late_initcall(unknow_nmi_init); >>> >> >> Thanks, >> Yunhui > > Thanks, > Yunhui ^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [External] Re: [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support 2025-09-22 8:11 ` yunhui cui 2025-09-22 8:31 ` Clément Léger @ 2025-09-22 10:39 ` Anup Patel 2025-09-22 11:58 ` yunhui cui 1 sibling, 1 reply; 11+ messages in thread From: Anup Patel @ 2025-09-22 10:39 UTC (permalink / raw) To: yunhui cui Cc: Clément Léger, paul.walmsley, palmer, aou, alex, conor, atishp, ajones, apatel, mchitale, linux-riscv, linux-kernel On Mon, Sep 22, 2025 at 1:42 PM yunhui cui <cuiyunhui@bytedance.com> wrote: > > Hi Clément, > > On Fri, Sep 19, 2025 at 3:52 PM yunhui cui <cuiyunhui@bytedance.com> wrote: > > > > Hi Clément, > > > > On Fri, Sep 19, 2025 at 3:18 PM Clément Léger <cleger@rivosinc.com> wrote: > > > > > > > > > > > > On 19/09/2025 09:00, Yunhui Cui wrote: > > > > Unknown NMI can force the kernel to respond (e.g., panic) when the > > > > system encounters unrecognized critical hardware events, aiding in > > > > troubleshooting system faults. This is implemented via the Supervisor > > > > Software Events (SSE) framework. > > > > > > > > Signed-off-by: Yunhui Cui <cuiyunhui@bytedance.com> > > > > --- > > > > arch/riscv/include/asm/sbi.h | 1 + > > > > drivers/firmware/riscv/Kconfig | 10 +++++ > > > > drivers/firmware/riscv/Makefile | 1 + > > > > drivers/firmware/riscv/sse_nmi.c | 77 ++++++++++++++++++++++++++++++++ > > > > 4 files changed, 89 insertions(+) > > > > create mode 100644 drivers/firmware/riscv/sse_nmi.c > > > > > > > > diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h > > > > index 874cc1d7603a5..5801f90a88f62 100644 > > > > --- a/arch/riscv/include/asm/sbi.h > > > > +++ b/arch/riscv/include/asm/sbi.h > > > > @@ -481,6 +481,7 @@ enum sbi_sse_attr_id { > > > > > > > > #define SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS 0x00000000 > > > > #define SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP 0x00000001 > > > > +#define SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI 0x00000002 > > > > > > Was this submitted to the PRS WG ? This a specification modification so > > > it should go through the usual process. > > > > > > > #define SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS 0x00008000 > > > > #define SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW 0x00010000 > > > > #define SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS 0x00100000 > > > > diff --git a/drivers/firmware/riscv/Kconfig b/drivers/firmware/riscv/Kconfig > > > > index ed5b663ac5f91..746bac862ac46 100644 > > > > --- a/drivers/firmware/riscv/Kconfig > > > > +++ b/drivers/firmware/riscv/Kconfig > > > > @@ -12,4 +12,14 @@ config RISCV_SBI_SSE > > > > this option provides support to register callbacks on specific SSE > > > > events. > > > > > > > > +config RISCV_SSE_UNKNOWN_NMI > > > > + bool "Enable SBI Supervisor Software Events unknown NMI support" > > > > + depends on RISCV_SBI_SSE > > > > + default y > > > > + help > > > > + This option enables support for delivering unknown Non-Maskable Interrupt (NMI) > > > > + notifications via the Supervisor Software Events (SSE) framework. When enabled, > > > > + unknown NMIs can trigger kernel responses (e.g., panic) for unrecognized critical > > > > + hardware events, aiding in system fault diagnosis. > > > > + > > > > endmenu > > > > diff --git a/drivers/firmware/riscv/Makefile b/drivers/firmware/riscv/Makefile > > > > index c8795d4bbb2ea..9242c6cd5e3e9 100644 > > > > --- a/drivers/firmware/riscv/Makefile > > > > +++ b/drivers/firmware/riscv/Makefile > > > > @@ -1,3 +1,4 @@ > > > > # SPDX-License-Identifier: GPL-2.0 > > > > > > > > obj-$(CONFIG_RISCV_SBI_SSE) += riscv_sbi_sse.o > > > > +obj-$(CONFIG_RISCV_SSE_UNKNOWN_NMI) += sse_nmi.o > > > > diff --git a/drivers/firmware/riscv/sse_nmi.c b/drivers/firmware/riscv/sse_nmi.c > > > > new file mode 100644 > > > > index 0000000000000..43063f42efff0 > > > > --- /dev/null > > > > +++ b/drivers/firmware/riscv/sse_nmi.c > > > > @@ -0,0 +1,77 @@ > > > > +// SPDX-License-Identifier: GPL-2.0+ > > > > + > > > > +#include <linux/mm.h> > > > > +#include <linux/nmi.h> > > > > +#include <linux/riscv_sbi_sse.h> > > > > +#include <linux/sched/debug.h> > > > > +#include <linux/sysctl.h> > > > > + > > > > +#include <asm/irq_regs.h> > > > > +#include <asm/sbi.h> > > > > + > > > > +int panic_on_unknown_nmi = 1; > > > > +struct sse_event *evt; > > > > +static struct ctl_table_header *unknown_nmi_sysctl_header; > > > > + > > > > +const struct ctl_table unknown_nmi_table[] = { > > > > + { > > > > + .procname = "panic_enable", > > > > + .data = &panic_on_unknown_nmi, > > > > + .maxlen = sizeof(int), > > > > + .mode = 0644, > > > > + .proc_handler = proc_dointvec_minmax, > > > > + .extra1 = SYSCTL_ZERO, > > > > + .extra2 = SYSCTL_ONE, > > > > + }, > > > > +}; > > > > + > > > > +static void nmi_handler(struct pt_regs *regs) > > > > +{ > > > > + pr_emerg("NMI received for unknown on CPU %d.\n", smp_processor_id()); > > > > + > > > > + if (panic_on_unknown_nmi) > > > > + nmi_panic(regs, "NMI: Not continuing"); > > > > + > > > > + pr_emerg("Dazed and confused, but trying to continue\n"); > > > > +} > > > > > > I'm dazed and confused as well ;) What's the point of this except > > > interrupting the kernel with a panic ? It seems like it's a better idea > > > to let the firmware handle that properly and display whatever > > > information are needed. Was your idea to actually force the kernel to > > > enter in some debug mode ? > > > > There is an important scenario: when the kernel becomes unresponsive, > > we need to trigger an unknown NMI to cause the system to panic() and > > then collect the vmcore, and such a requirement is common on x86 > > servers. > > > > > > > > Thanks, > > > > > > Clément > > > > > > > + > > > > +static int nmi_sse_handler(u32 evt, void *arg, struct pt_regs *regs) > > > > +{ > > > > + nmi_handler(regs); > > > > + > > > > + return 0; > > > > +} > > > > + > > > > +static int sse_nmi_init(void) > > > > +{ > > > > + int ret; > > > > + > > > > + evt = sse_event_register(SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI, 0, > > > > + nmi_sse_handler, NULL); > > Should we add this UNKNOWN_NMI event ID in Chapter 17 of the SBI spec? The ratified SBI v3.0 defines a "Software injected local event" (ID 0xffff0000) which can be used to inject panic() on a particular HART from another HART. Regards, Anup ^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [External] Re: [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support 2025-09-22 10:39 ` Anup Patel @ 2025-09-22 11:58 ` yunhui cui 2025-09-22 17:20 ` Atish Kumar Patra 2025-09-23 5:05 ` Anup Patel 0 siblings, 2 replies; 11+ messages in thread From: yunhui cui @ 2025-09-22 11:58 UTC (permalink / raw) To: Anup Patel Cc: Clément Léger, paul.walmsley, palmer, aou, alex, conor, atishp, ajones, apatel, mchitale, linux-riscv, linux-kernel Hi Anup, On Mon, Sep 22, 2025 at 6:40 PM Anup Patel <anup@brainfault.org> wrote: > > On Mon, Sep 22, 2025 at 1:42 PM yunhui cui <cuiyunhui@bytedance.com> wrote: > > > > Hi Clément, > > > > On Fri, Sep 19, 2025 at 3:52 PM yunhui cui <cuiyunhui@bytedance.com> wrote: > > > > > > Hi Clément, > > > > > > On Fri, Sep 19, 2025 at 3:18 PM Clément Léger <cleger@rivosinc.com> wrote: > > > > > > > > > > > > > > > > On 19/09/2025 09:00, Yunhui Cui wrote: > > > > > Unknown NMI can force the kernel to respond (e.g., panic) when the > > > > > system encounters unrecognized critical hardware events, aiding in > > > > > troubleshooting system faults. This is implemented via the Supervisor > > > > > Software Events (SSE) framework. > > > > > > > > > > Signed-off-by: Yunhui Cui <cuiyunhui@bytedance.com> > > > > > --- > > > > > arch/riscv/include/asm/sbi.h | 1 + > > > > > drivers/firmware/riscv/Kconfig | 10 +++++ > > > > > drivers/firmware/riscv/Makefile | 1 + > > > > > drivers/firmware/riscv/sse_nmi.c | 77 ++++++++++++++++++++++++++++++++ > > > > > 4 files changed, 89 insertions(+) > > > > > create mode 100644 drivers/firmware/riscv/sse_nmi.c > > > > > > > > > > diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h > > > > > index 874cc1d7603a5..5801f90a88f62 100644 > > > > > --- a/arch/riscv/include/asm/sbi.h > > > > > +++ b/arch/riscv/include/asm/sbi.h > > > > > @@ -481,6 +481,7 @@ enum sbi_sse_attr_id { > > > > > > > > > > #define SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS 0x00000000 > > > > > #define SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP 0x00000001 > > > > > +#define SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI 0x00000002 > > > > > > > > Was this submitted to the PRS WG ? This a specification modification so > > > > it should go through the usual process. > > > > > > > > > #define SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS 0x00008000 > > > > > #define SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW 0x00010000 > > > > > #define SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS 0x00100000 > > > > > diff --git a/drivers/firmware/riscv/Kconfig b/drivers/firmware/riscv/Kconfig > > > > > index ed5b663ac5f91..746bac862ac46 100644 > > > > > --- a/drivers/firmware/riscv/Kconfig > > > > > +++ b/drivers/firmware/riscv/Kconfig > > > > > @@ -12,4 +12,14 @@ config RISCV_SBI_SSE > > > > > this option provides support to register callbacks on specific SSE > > > > > events. > > > > > > > > > > +config RISCV_SSE_UNKNOWN_NMI > > > > > + bool "Enable SBI Supervisor Software Events unknown NMI support" > > > > > + depends on RISCV_SBI_SSE > > > > > + default y > > > > > + help > > > > > + This option enables support for delivering unknown Non-Maskable Interrupt (NMI) > > > > > + notifications via the Supervisor Software Events (SSE) framework. When enabled, > > > > > + unknown NMIs can trigger kernel responses (e.g., panic) for unrecognized critical > > > > > + hardware events, aiding in system fault diagnosis. > > > > > + > > > > > endmenu > > > > > diff --git a/drivers/firmware/riscv/Makefile b/drivers/firmware/riscv/Makefile > > > > > index c8795d4bbb2ea..9242c6cd5e3e9 100644 > > > > > --- a/drivers/firmware/riscv/Makefile > > > > > +++ b/drivers/firmware/riscv/Makefile > > > > > @@ -1,3 +1,4 @@ > > > > > # SPDX-License-Identifier: GPL-2.0 > > > > > > > > > > obj-$(CONFIG_RISCV_SBI_SSE) += riscv_sbi_sse.o > > > > > +obj-$(CONFIG_RISCV_SSE_UNKNOWN_NMI) += sse_nmi.o > > > > > diff --git a/drivers/firmware/riscv/sse_nmi.c b/drivers/firmware/riscv/sse_nmi.c > > > > > new file mode 100644 > > > > > index 0000000000000..43063f42efff0 > > > > > --- /dev/null > > > > > +++ b/drivers/firmware/riscv/sse_nmi.c > > > > > @@ -0,0 +1,77 @@ > > > > > +// SPDX-License-Identifier: GPL-2.0+ > > > > > + > > > > > +#include <linux/mm.h> > > > > > +#include <linux/nmi.h> > > > > > +#include <linux/riscv_sbi_sse.h> > > > > > +#include <linux/sched/debug.h> > > > > > +#include <linux/sysctl.h> > > > > > + > > > > > +#include <asm/irq_regs.h> > > > > > +#include <asm/sbi.h> > > > > > + > > > > > +int panic_on_unknown_nmi = 1; > > > > > +struct sse_event *evt; > > > > > +static struct ctl_table_header *unknown_nmi_sysctl_header; > > > > > + > > > > > +const struct ctl_table unknown_nmi_table[] = { > > > > > + { > > > > > + .procname = "panic_enable", > > > > > + .data = &panic_on_unknown_nmi, > > > > > + .maxlen = sizeof(int), > > > > > + .mode = 0644, > > > > > + .proc_handler = proc_dointvec_minmax, > > > > > + .extra1 = SYSCTL_ZERO, > > > > > + .extra2 = SYSCTL_ONE, > > > > > + }, > > > > > +}; > > > > > + > > > > > +static void nmi_handler(struct pt_regs *regs) > > > > > +{ > > > > > + pr_emerg("NMI received for unknown on CPU %d.\n", smp_processor_id()); > > > > > + > > > > > + if (panic_on_unknown_nmi) > > > > > + nmi_panic(regs, "NMI: Not continuing"); > > > > > + > > > > > + pr_emerg("Dazed and confused, but trying to continue\n"); > > > > > +} > > > > > > > > I'm dazed and confused as well ;) What's the point of this except > > > > interrupting the kernel with a panic ? It seems like it's a better idea > > > > to let the firmware handle that properly and display whatever > > > > information are needed. Was your idea to actually force the kernel to > > > > enter in some debug mode ? > > > > > > There is an important scenario: when the kernel becomes unresponsive, > > > we need to trigger an unknown NMI to cause the system to panic() and > > > then collect the vmcore, and such a requirement is common on x86 > > > servers. > > > > > > > > > > > Thanks, > > > > > > > > Clément > > > > > > > > > + > > > > > +static int nmi_sse_handler(u32 evt, void *arg, struct pt_regs *regs) > > > > > +{ > > > > > + nmi_handler(regs); > > > > > + > > > > > + return 0; > > > > > +} > > > > > + > > > > > +static int sse_nmi_init(void) > > > > > +{ > > > > > + int ret; > > > > > + > > > > > + evt = sse_event_register(SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI, 0, > > > > > + nmi_sse_handler, NULL); > > > > Should we add this UNKNOWN_NMI event ID in Chapter 17 of the SBI spec? > > The ratified SBI v3.0 defines a "Software injected local event" (ID 0xffff0000) > which can be used to inject panic() on a particular HART from another HART. Has SBI_SSE_EVENT_LOCAL_SOFTWARE (0xffff0000) been given meaning currently? The scenario of the unknown NMI is that after a HART receives an interrupt from the platform, M-mode responds to this interrupt and then enters S-mode to execute the registered handler. Can SBI_SSE_EVENT_LOCAL_SOFTWARE be used directly? > > Regards, > Anup Thanks, Yunhui ^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [External] Re: [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support 2025-09-22 11:58 ` yunhui cui @ 2025-09-22 17:20 ` Atish Kumar Patra 2025-09-23 2:15 ` yunhui cui 2025-09-23 5:05 ` Anup Patel 1 sibling, 1 reply; 11+ messages in thread From: Atish Kumar Patra @ 2025-09-22 17:20 UTC (permalink / raw) To: yunhui cui Cc: Anup Patel, Clément Léger, paul.walmsley, palmer, aou, alex, conor, ajones, apatel, mchitale, linux-riscv, linux-kernel On Mon, Sep 22, 2025 at 4:58 AM yunhui cui <cuiyunhui@bytedance.com> wrote: > > Hi Anup, > > On Mon, Sep 22, 2025 at 6:40 PM Anup Patel <anup@brainfault.org> wrote: > > > > On Mon, Sep 22, 2025 at 1:42 PM yunhui cui <cuiyunhui@bytedance.com> wrote: > > > > > > Hi Clément, > > > > > > On Fri, Sep 19, 2025 at 3:52 PM yunhui cui <cuiyunhui@bytedance.com> wrote: > > > > > > > > Hi Clément, > > > > > > > > On Fri, Sep 19, 2025 at 3:18 PM Clément Léger <cleger@rivosinc.com> wrote: > > > > > > > > > > > > > > > > > > > > On 19/09/2025 09:00, Yunhui Cui wrote: > > > > > > Unknown NMI can force the kernel to respond (e.g., panic) when the > > > > > > system encounters unrecognized critical hardware events, aiding in > > > > > > troubleshooting system faults. This is implemented via the Supervisor > > > > > > Software Events (SSE) framework. > > > > > > > > > > > > Signed-off-by: Yunhui Cui <cuiyunhui@bytedance.com> > > > > > > --- > > > > > > arch/riscv/include/asm/sbi.h | 1 + > > > > > > drivers/firmware/riscv/Kconfig | 10 +++++ > > > > > > drivers/firmware/riscv/Makefile | 1 + > > > > > > drivers/firmware/riscv/sse_nmi.c | 77 ++++++++++++++++++++++++++++++++ > > > > > > 4 files changed, 89 insertions(+) > > > > > > create mode 100644 drivers/firmware/riscv/sse_nmi.c > > > > > > > > > > > > diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h > > > > > > index 874cc1d7603a5..5801f90a88f62 100644 > > > > > > --- a/arch/riscv/include/asm/sbi.h > > > > > > +++ b/arch/riscv/include/asm/sbi.h > > > > > > @@ -481,6 +481,7 @@ enum sbi_sse_attr_id { > > > > > > > > > > > > #define SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS 0x00000000 > > > > > > #define SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP 0x00000001 > > > > > > +#define SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI 0x00000002 > > > > > > > > > > Was this submitted to the PRS WG ? This a specification modification so > > > > > it should go through the usual process. > > > > > > > > > > > #define SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS 0x00008000 > > > > > > #define SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW 0x00010000 > > > > > > #define SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS 0x00100000 > > > > > > diff --git a/drivers/firmware/riscv/Kconfig b/drivers/firmware/riscv/Kconfig > > > > > > index ed5b663ac5f91..746bac862ac46 100644 > > > > > > --- a/drivers/firmware/riscv/Kconfig > > > > > > +++ b/drivers/firmware/riscv/Kconfig > > > > > > @@ -12,4 +12,14 @@ config RISCV_SBI_SSE > > > > > > this option provides support to register callbacks on specific SSE > > > > > > events. > > > > > > > > > > > > +config RISCV_SSE_UNKNOWN_NMI > > > > > > + bool "Enable SBI Supervisor Software Events unknown NMI support" > > > > > > + depends on RISCV_SBI_SSE > > > > > > + default y > > > > > > + help > > > > > > + This option enables support for delivering unknown Non-Maskable Interrupt (NMI) > > > > > > + notifications via the Supervisor Software Events (SSE) framework. When enabled, > > > > > > + unknown NMIs can trigger kernel responses (e.g., panic) for unrecognized critical > > > > > > + hardware events, aiding in system fault diagnosis. > > > > > > + > > > > > > endmenu > > > > > > diff --git a/drivers/firmware/riscv/Makefile b/drivers/firmware/riscv/Makefile > > > > > > index c8795d4bbb2ea..9242c6cd5e3e9 100644 > > > > > > --- a/drivers/firmware/riscv/Makefile > > > > > > +++ b/drivers/firmware/riscv/Makefile > > > > > > @@ -1,3 +1,4 @@ > > > > > > # SPDX-License-Identifier: GPL-2.0 > > > > > > > > > > > > obj-$(CONFIG_RISCV_SBI_SSE) += riscv_sbi_sse.o > > > > > > +obj-$(CONFIG_RISCV_SSE_UNKNOWN_NMI) += sse_nmi.o > > > > > > diff --git a/drivers/firmware/riscv/sse_nmi.c b/drivers/firmware/riscv/sse_nmi.c > > > > > > new file mode 100644 > > > > > > index 0000000000000..43063f42efff0 > > > > > > --- /dev/null > > > > > > +++ b/drivers/firmware/riscv/sse_nmi.c > > > > > > @@ -0,0 +1,77 @@ > > > > > > +// SPDX-License-Identifier: GPL-2.0+ > > > > > > + > > > > > > +#include <linux/mm.h> > > > > > > +#include <linux/nmi.h> > > > > > > +#include <linux/riscv_sbi_sse.h> > > > > > > +#include <linux/sched/debug.h> > > > > > > +#include <linux/sysctl.h> > > > > > > + > > > > > > +#include <asm/irq_regs.h> > > > > > > +#include <asm/sbi.h> > > > > > > + > > > > > > +int panic_on_unknown_nmi = 1; > > > > > > +struct sse_event *evt; > > > > > > +static struct ctl_table_header *unknown_nmi_sysctl_header; > > > > > > + > > > > > > +const struct ctl_table unknown_nmi_table[] = { > > > > > > + { > > > > > > + .procname = "panic_enable", > > > > > > + .data = &panic_on_unknown_nmi, > > > > > > + .maxlen = sizeof(int), > > > > > > + .mode = 0644, > > > > > > + .proc_handler = proc_dointvec_minmax, > > > > > > + .extra1 = SYSCTL_ZERO, > > > > > > + .extra2 = SYSCTL_ONE, > > > > > > + }, > > > > > > +}; > > > > > > + > > > > > > +static void nmi_handler(struct pt_regs *regs) > > > > > > +{ > > > > > > + pr_emerg("NMI received for unknown on CPU %d.\n", smp_processor_id()); > > > > > > + > > > > > > + if (panic_on_unknown_nmi) > > > > > > + nmi_panic(regs, "NMI: Not continuing"); > > > > > > + > > > > > > + pr_emerg("Dazed and confused, but trying to continue\n"); > > > > > > +} > > > > > > > > > > I'm dazed and confused as well ;) What's the point of this except > > > > > interrupting the kernel with a panic ? It seems like it's a better idea > > > > > to let the firmware handle that properly and display whatever > > > > > information are needed. Was your idea to actually force the kernel to > > > > > enter in some debug mode ? > > > > > > > > There is an important scenario: when the kernel becomes unresponsive, > > > > we need to trigger an unknown NMI to cause the system to panic() and > > > > then collect the vmcore, and such a requirement is common on x86 > > > > servers. > > > > > > > > > > > > > > Thanks, > > > > > > > > > > Clément > > > > > > > > > > > + > > > > > > +static int nmi_sse_handler(u32 evt, void *arg, struct pt_regs *regs) > > > > > > +{ > > > > > > + nmi_handler(regs); > > > > > > + > > > > > > + return 0; > > > > > > +} > > > > > > + > > > > > > +static int sse_nmi_init(void) > > > > > > +{ > > > > > > + int ret; > > > > > > + > > > > > > + evt = sse_event_register(SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI, 0, > > > > > > + nmi_sse_handler, NULL); > > > > > > Should we add this UNKNOWN_NMI event ID in Chapter 17 of the SBI spec? > > > > The ratified SBI v3.0 defines a "Software injected local event" (ID 0xffff0000) > > which can be used to inject panic() on a particular HART from another HART. > > Has SBI_SSE_EVENT_LOCAL_SOFTWARE (0xffff0000) been given meaning > currently? The scenario of the unknown NMI is that after a HART > receives an interrupt from the platform, M-mode responds to this Which interrupt ? A standard one or platform specific one ? > interrupt and then enters S-mode to execute the registered handler. > Can SBI_SSE_EVENT_LOCAL_SOFTWARE be used directly? > > > > > > Regards, > > Anup > > Thanks, > Yunhui ^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [External] Re: [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support 2025-09-22 17:20 ` Atish Kumar Patra @ 2025-09-23 2:15 ` yunhui cui 0 siblings, 0 replies; 11+ messages in thread From: yunhui cui @ 2025-09-23 2:15 UTC (permalink / raw) To: Atish Kumar Patra Cc: Anup Patel, Clément Léger, paul.walmsley, palmer, aou, alex, conor, ajones, apatel, mchitale, linux-riscv, linux-kernel Hi Atish, On Tue, Sep 23, 2025 at 1:20 AM Atish Kumar Patra <atishp@rivosinc.com> wrote: > > On Mon, Sep 22, 2025 at 4:58 AM yunhui cui <cuiyunhui@bytedance.com> wrote: > > > > Hi Anup, > > > > On Mon, Sep 22, 2025 at 6:40 PM Anup Patel <anup@brainfault.org> wrote: > > > > > > On Mon, Sep 22, 2025 at 1:42 PM yunhui cui <cuiyunhui@bytedance.com> wrote: > > > > > > > > Hi Clément, > > > > > > > > On Fri, Sep 19, 2025 at 3:52 PM yunhui cui <cuiyunhui@bytedance.com> wrote: > > > > > > > > > > Hi Clément, > > > > > > > > > > On Fri, Sep 19, 2025 at 3:18 PM Clément Léger <cleger@rivosinc.com> wrote: > > > > > > > > > > > > > > > > > > > > > > > > On 19/09/2025 09:00, Yunhui Cui wrote: > > > > > > > Unknown NMI can force the kernel to respond (e.g., panic) when the > > > > > > > system encounters unrecognized critical hardware events, aiding in > > > > > > > troubleshooting system faults. This is implemented via the Supervisor > > > > > > > Software Events (SSE) framework. > > > > > > > > > > > > > > Signed-off-by: Yunhui Cui <cuiyunhui@bytedance.com> > > > > > > > --- > > > > > > > arch/riscv/include/asm/sbi.h | 1 + > > > > > > > drivers/firmware/riscv/Kconfig | 10 +++++ > > > > > > > drivers/firmware/riscv/Makefile | 1 + > > > > > > > drivers/firmware/riscv/sse_nmi.c | 77 ++++++++++++++++++++++++++++++++ > > > > > > > 4 files changed, 89 insertions(+) > > > > > > > create mode 100644 drivers/firmware/riscv/sse_nmi.c > > > > > > > > > > > > > > diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h > > > > > > > index 874cc1d7603a5..5801f90a88f62 100644 > > > > > > > --- a/arch/riscv/include/asm/sbi.h > > > > > > > +++ b/arch/riscv/include/asm/sbi.h > > > > > > > @@ -481,6 +481,7 @@ enum sbi_sse_attr_id { > > > > > > > > > > > > > > #define SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS 0x00000000 > > > > > > > #define SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP 0x00000001 > > > > > > > +#define SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI 0x00000002 > > > > > > > > > > > > Was this submitted to the PRS WG ? This a specification modification so > > > > > > it should go through the usual process. > > > > > > > > > > > > > #define SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS 0x00008000 > > > > > > > #define SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW 0x00010000 > > > > > > > #define SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS 0x00100000 > > > > > > > diff --git a/drivers/firmware/riscv/Kconfig b/drivers/firmware/riscv/Kconfig > > > > > > > index ed5b663ac5f91..746bac862ac46 100644 > > > > > > > --- a/drivers/firmware/riscv/Kconfig > > > > > > > +++ b/drivers/firmware/riscv/Kconfig > > > > > > > @@ -12,4 +12,14 @@ config RISCV_SBI_SSE > > > > > > > this option provides support to register callbacks on specific SSE > > > > > > > events. > > > > > > > > > > > > > > +config RISCV_SSE_UNKNOWN_NMI > > > > > > > + bool "Enable SBI Supervisor Software Events unknown NMI support" > > > > > > > + depends on RISCV_SBI_SSE > > > > > > > + default y > > > > > > > + help > > > > > > > + This option enables support for delivering unknown Non-Maskable Interrupt (NMI) > > > > > > > + notifications via the Supervisor Software Events (SSE) framework. When enabled, > > > > > > > + unknown NMIs can trigger kernel responses (e.g., panic) for unrecognized critical > > > > > > > + hardware events, aiding in system fault diagnosis. > > > > > > > + > > > > > > > endmenu > > > > > > > diff --git a/drivers/firmware/riscv/Makefile b/drivers/firmware/riscv/Makefile > > > > > > > index c8795d4bbb2ea..9242c6cd5e3e9 100644 > > > > > > > --- a/drivers/firmware/riscv/Makefile > > > > > > > +++ b/drivers/firmware/riscv/Makefile > > > > > > > @@ -1,3 +1,4 @@ > > > > > > > # SPDX-License-Identifier: GPL-2.0 > > > > > > > > > > > > > > obj-$(CONFIG_RISCV_SBI_SSE) += riscv_sbi_sse.o > > > > > > > +obj-$(CONFIG_RISCV_SSE_UNKNOWN_NMI) += sse_nmi.o > > > > > > > diff --git a/drivers/firmware/riscv/sse_nmi.c b/drivers/firmware/riscv/sse_nmi.c > > > > > > > new file mode 100644 > > > > > > > index 0000000000000..43063f42efff0 > > > > > > > --- /dev/null > > > > > > > +++ b/drivers/firmware/riscv/sse_nmi.c > > > > > > > @@ -0,0 +1,77 @@ > > > > > > > +// SPDX-License-Identifier: GPL-2.0+ > > > > > > > + > > > > > > > +#include <linux/mm.h> > > > > > > > +#include <linux/nmi.h> > > > > > > > +#include <linux/riscv_sbi_sse.h> > > > > > > > +#include <linux/sched/debug.h> > > > > > > > +#include <linux/sysctl.h> > > > > > > > + > > > > > > > +#include <asm/irq_regs.h> > > > > > > > +#include <asm/sbi.h> > > > > > > > + > > > > > > > +int panic_on_unknown_nmi = 1; > > > > > > > +struct sse_event *evt; > > > > > > > +static struct ctl_table_header *unknown_nmi_sysctl_header; > > > > > > > + > > > > > > > +const struct ctl_table unknown_nmi_table[] = { > > > > > > > + { > > > > > > > + .procname = "panic_enable", > > > > > > > + .data = &panic_on_unknown_nmi, > > > > > > > + .maxlen = sizeof(int), > > > > > > > + .mode = 0644, > > > > > > > + .proc_handler = proc_dointvec_minmax, > > > > > > > + .extra1 = SYSCTL_ZERO, > > > > > > > + .extra2 = SYSCTL_ONE, > > > > > > > + }, > > > > > > > +}; > > > > > > > + > > > > > > > +static void nmi_handler(struct pt_regs *regs) > > > > > > > +{ > > > > > > > + pr_emerg("NMI received for unknown on CPU %d.\n", smp_processor_id()); > > > > > > > + > > > > > > > + if (panic_on_unknown_nmi) > > > > > > > + nmi_panic(regs, "NMI: Not continuing"); > > > > > > > + > > > > > > > + pr_emerg("Dazed and confused, but trying to continue\n"); > > > > > > > +} > > > > > > > > > > > > I'm dazed and confused as well ;) What's the point of this except > > > > > > interrupting the kernel with a panic ? It seems like it's a better idea > > > > > > to let the firmware handle that properly and display whatever > > > > > > information are needed. Was your idea to actually force the kernel to > > > > > > enter in some debug mode ? > > > > > > > > > > There is an important scenario: when the kernel becomes unresponsive, > > > > > we need to trigger an unknown NMI to cause the system to panic() and > > > > > then collect the vmcore, and such a requirement is common on x86 > > > > > servers. > > > > > > > > > > > > > > > > > Thanks, > > > > > > > > > > > > Clément > > > > > > > > > > > > > + > > > > > > > +static int nmi_sse_handler(u32 evt, void *arg, struct pt_regs *regs) > > > > > > > +{ > > > > > > > + nmi_handler(regs); > > > > > > > + > > > > > > > + return 0; > > > > > > > +} > > > > > > > + > > > > > > > +static int sse_nmi_init(void) > > > > > > > +{ > > > > > > > + int ret; > > > > > > > + > > > > > > > + evt = sse_event_register(SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI, 0, > > > > > > > + nmi_sse_handler, NULL); > > > > > > > > Should we add this UNKNOWN_NMI event ID in Chapter 17 of the SBI spec? > > > > > > The ratified SBI v3.0 defines a "Software injected local event" (ID 0xffff0000) > > > which can be used to inject panic() on a particular HART from another HART. > > > > Has SBI_SSE_EVENT_LOCAL_SOFTWARE (0xffff0000) been given meaning > > currently? The scenario of the unknown NMI is that after a HART > > receives an interrupt from the platform, M-mode responds to this > > Which interrupt ? A standard one or platform specific one ? Platform specific one. Are there any recommended event IDs? Or should a new UNKNOWN NMI event ID be added to the SBI spec? > > > interrupt and then enters S-mode to execute the registered handler. > > Can SBI_SSE_EVENT_LOCAL_SOFTWARE be used directly? > > > > > > > > > > Regards, > > > Anup > > > > Thanks, > > Yunhui Thanks, Yunhui ^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [External] Re: [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support 2025-09-22 11:58 ` yunhui cui 2025-09-22 17:20 ` Atish Kumar Patra @ 2025-09-23 5:05 ` Anup Patel 1 sibling, 0 replies; 11+ messages in thread From: Anup Patel @ 2025-09-23 5:05 UTC (permalink / raw) To: yunhui cui Cc: Clément Léger, paul.walmsley, palmer, aou, alex, conor, atishp, ajones, apatel, mchitale, linux-riscv, linux-kernel On Mon, Sep 22, 2025 at 5:28 PM yunhui cui <cuiyunhui@bytedance.com> wrote: > > Hi Anup, > > On Mon, Sep 22, 2025 at 6:40 PM Anup Patel <anup@brainfault.org> wrote: > > > > On Mon, Sep 22, 2025 at 1:42 PM yunhui cui <cuiyunhui@bytedance.com> wrote: > > > > > > Hi Clément, > > > > > > On Fri, Sep 19, 2025 at 3:52 PM yunhui cui <cuiyunhui@bytedance.com> wrote: > > > > > > > > Hi Clément, > > > > > > > > On Fri, Sep 19, 2025 at 3:18 PM Clément Léger <cleger@rivosinc.com> wrote: > > > > > > > > > > > > > > > > > > > > On 19/09/2025 09:00, Yunhui Cui wrote: > > > > > > Unknown NMI can force the kernel to respond (e.g., panic) when the > > > > > > system encounters unrecognized critical hardware events, aiding in > > > > > > troubleshooting system faults. This is implemented via the Supervisor > > > > > > Software Events (SSE) framework. > > > > > > > > > > > > Signed-off-by: Yunhui Cui <cuiyunhui@bytedance.com> > > > > > > --- > > > > > > arch/riscv/include/asm/sbi.h | 1 + > > > > > > drivers/firmware/riscv/Kconfig | 10 +++++ > > > > > > drivers/firmware/riscv/Makefile | 1 + > > > > > > drivers/firmware/riscv/sse_nmi.c | 77 ++++++++++++++++++++++++++++++++ > > > > > > 4 files changed, 89 insertions(+) > > > > > > create mode 100644 drivers/firmware/riscv/sse_nmi.c > > > > > > > > > > > > diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h > > > > > > index 874cc1d7603a5..5801f90a88f62 100644 > > > > > > --- a/arch/riscv/include/asm/sbi.h > > > > > > +++ b/arch/riscv/include/asm/sbi.h > > > > > > @@ -481,6 +481,7 @@ enum sbi_sse_attr_id { > > > > > > > > > > > > #define SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS 0x00000000 > > > > > > #define SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP 0x00000001 > > > > > > +#define SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI 0x00000002 > > > > > > > > > > Was this submitted to the PRS WG ? This a specification modification so > > > > > it should go through the usual process. > > > > > > > > > > > #define SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS 0x00008000 > > > > > > #define SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW 0x00010000 > > > > > > #define SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS 0x00100000 > > > > > > diff --git a/drivers/firmware/riscv/Kconfig b/drivers/firmware/riscv/Kconfig > > > > > > index ed5b663ac5f91..746bac862ac46 100644 > > > > > > --- a/drivers/firmware/riscv/Kconfig > > > > > > +++ b/drivers/firmware/riscv/Kconfig > > > > > > @@ -12,4 +12,14 @@ config RISCV_SBI_SSE > > > > > > this option provides support to register callbacks on specific SSE > > > > > > events. > > > > > > > > > > > > +config RISCV_SSE_UNKNOWN_NMI > > > > > > + bool "Enable SBI Supervisor Software Events unknown NMI support" > > > > > > + depends on RISCV_SBI_SSE > > > > > > + default y > > > > > > + help > > > > > > + This option enables support for delivering unknown Non-Maskable Interrupt (NMI) > > > > > > + notifications via the Supervisor Software Events (SSE) framework. When enabled, > > > > > > + unknown NMIs can trigger kernel responses (e.g., panic) for unrecognized critical > > > > > > + hardware events, aiding in system fault diagnosis. > > > > > > + > > > > > > endmenu > > > > > > diff --git a/drivers/firmware/riscv/Makefile b/drivers/firmware/riscv/Makefile > > > > > > index c8795d4bbb2ea..9242c6cd5e3e9 100644 > > > > > > --- a/drivers/firmware/riscv/Makefile > > > > > > +++ b/drivers/firmware/riscv/Makefile > > > > > > @@ -1,3 +1,4 @@ > > > > > > # SPDX-License-Identifier: GPL-2.0 > > > > > > > > > > > > obj-$(CONFIG_RISCV_SBI_SSE) += riscv_sbi_sse.o > > > > > > +obj-$(CONFIG_RISCV_SSE_UNKNOWN_NMI) += sse_nmi.o > > > > > > diff --git a/drivers/firmware/riscv/sse_nmi.c b/drivers/firmware/riscv/sse_nmi.c > > > > > > new file mode 100644 > > > > > > index 0000000000000..43063f42efff0 > > > > > > --- /dev/null > > > > > > +++ b/drivers/firmware/riscv/sse_nmi.c > > > > > > @@ -0,0 +1,77 @@ > > > > > > +// SPDX-License-Identifier: GPL-2.0+ > > > > > > + > > > > > > +#include <linux/mm.h> > > > > > > +#include <linux/nmi.h> > > > > > > +#include <linux/riscv_sbi_sse.h> > > > > > > +#include <linux/sched/debug.h> > > > > > > +#include <linux/sysctl.h> > > > > > > + > > > > > > +#include <asm/irq_regs.h> > > > > > > +#include <asm/sbi.h> > > > > > > + > > > > > > +int panic_on_unknown_nmi = 1; > > > > > > +struct sse_event *evt; > > > > > > +static struct ctl_table_header *unknown_nmi_sysctl_header; > > > > > > + > > > > > > +const struct ctl_table unknown_nmi_table[] = { > > > > > > + { > > > > > > + .procname = "panic_enable", > > > > > > + .data = &panic_on_unknown_nmi, > > > > > > + .maxlen = sizeof(int), > > > > > > + .mode = 0644, > > > > > > + .proc_handler = proc_dointvec_minmax, > > > > > > + .extra1 = SYSCTL_ZERO, > > > > > > + .extra2 = SYSCTL_ONE, > > > > > > + }, > > > > > > +}; > > > > > > + > > > > > > +static void nmi_handler(struct pt_regs *regs) > > > > > > +{ > > > > > > + pr_emerg("NMI received for unknown on CPU %d.\n", smp_processor_id()); > > > > > > + > > > > > > + if (panic_on_unknown_nmi) > > > > > > + nmi_panic(regs, "NMI: Not continuing"); > > > > > > + > > > > > > + pr_emerg("Dazed and confused, but trying to continue\n"); > > > > > > +} > > > > > > > > > > I'm dazed and confused as well ;) What's the point of this except > > > > > interrupting the kernel with a panic ? It seems like it's a better idea > > > > > to let the firmware handle that properly and display whatever > > > > > information are needed. Was your idea to actually force the kernel to > > > > > enter in some debug mode ? > > > > > > > > There is an important scenario: when the kernel becomes unresponsive, > > > > we need to trigger an unknown NMI to cause the system to panic() and > > > > then collect the vmcore, and such a requirement is common on x86 > > > > servers. > > > > > > > > > > > > > > Thanks, > > > > > > > > > > Clément > > > > > > > > > > > + > > > > > > +static int nmi_sse_handler(u32 evt, void *arg, struct pt_regs *regs) > > > > > > +{ > > > > > > + nmi_handler(regs); > > > > > > + > > > > > > + return 0; > > > > > > +} > > > > > > + > > > > > > +static int sse_nmi_init(void) > > > > > > +{ > > > > > > + int ret; > > > > > > + > > > > > > + evt = sse_event_register(SBI_SSE_EVENT_LOCAL_UNKNOWN_NMI, 0, > > > > > > + nmi_sse_handler, NULL); > > > > > > Should we add this UNKNOWN_NMI event ID in Chapter 17 of the SBI spec? > > > > The ratified SBI v3.0 defines a "Software injected local event" (ID 0xffff0000) > > which can be used to inject panic() on a particular HART from another HART. > > Has SBI_SSE_EVENT_LOCAL_SOFTWARE (0xffff0000) been given meaning > currently? The scenario of the unknown NMI is that after a HART > receives an interrupt from the platform, M-mode responds to this > interrupt and then enters S-mode to execute the registered handler. > Can SBI_SSE_EVENT_LOCAL_SOFTWARE be used directly? It's a software injected event with lower priority which can be used by supervisor software on a HART to inject an event to another HART. If you need a platform specific event then there are multiple ranges with different priority-levels. Regards, Anup ^ permalink raw reply [flat|nested] 11+ messages in thread
end of thread, other threads:[~2025-09-23 5:06 UTC | newest] Thread overview: 11+ messages (download: mbox.gz / follow: Atom feed) -- links below jump to the message on this page -- 2025-09-19 7:00 [PATCH RFC 0/1] Support unknown NMI based on SSE Yunhui Cui 2025-09-19 7:00 ` [PATCH RFC 1/1] drivers: firmware: riscv: add unknown NMI support Yunhui Cui 2025-09-19 7:17 ` Clément Léger 2025-09-19 7:52 ` [External] " yunhui cui 2025-09-22 8:11 ` yunhui cui 2025-09-22 8:31 ` Clément Léger 2025-09-22 10:39 ` Anup Patel 2025-09-22 11:58 ` yunhui cui 2025-09-22 17:20 ` Atish Kumar Patra 2025-09-23 2:15 ` yunhui cui 2025-09-23 5:05 ` Anup Patel
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®