* [PATCH v20 1/9] firmware: arm_rmm: Add SMC definitions for calling the RMM
2026-09-29 22:16 [PATCH v20 0/9] firmware: arm_rmm: Add RMM v2.0 base RMI support Suzuki K Poulose
@ 2026-09-29 22:16 ` Suzuki K Poulose
2026-09-30 9:51 ` Catalin Marinas
2026-09-29 22:16 ` [PATCH v20 2/9] firmware: arm_rmm: Check for RMI support at init Suzuki K Poulose
` (7 subsequent siblings)
8 siblings, 1 reply; 23+ messages in thread
From: Suzuki K Poulose @ 2026-09-29 22:16 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, sudeep.holla,
jonathan.cameron, Suzuki K Poulose
From: Steven Price <steven.price@arm.com>
The RMM (Realm Management Monitor) provides functionality that can be
accessed by SMC calls from the host.
The SMC definitions are based on DEN0137[1] version 2.0-bet3
[1] https://developer.arm.com/documentation/den0137/2-0bet3/
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Reviewed-by: Gavin Shan <gshan@redhat.com>
Tested-by: Gavin Shan <gshan@redhat.com>
Signed-off-by: Steven Price <steven.price@arm.com>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Changes since v19:
* Fold the 'padding' only union in rec_exit to the previous group of fields
* Add static_assert to make sure the struct field offsets match the spec
* Rename 'padding' fields to 'sizer' which they really are
* Add MAINTAINERS entry for the new header file under ARM64
Changes since v18:
* Use _ULL version for masks spanning beyond 32bits
* Add GENMASK, FIELD_GET, FIELD_PREP for RMI_ABI_VERSION_*
* Consistent naming for masks, order the remaing ones by MSB to LSB
* Rename RMI_RETURN_ => RMI_RESULT and add relevant Rmi types for
clear indication on what they represent.
Changes since v17:
* Use GENMASK()/BIT() for masks consistently
* Rename RMI_{ADDR_RANGE, DONATE}_SIZE => RMI_{*}_BLOCK_SIZE
* Reorder the definitions for MSB to LSB
* Add definions for RMI_OP_MEM_*CONTIG and RMI_OP_CAN*_CANCEL
Changes since v16:
* Updated definitions to RMM specification v2.0-bet3.
---
MAINTAINERS | 1 +
include/linux/arm-smccc-rmi.h | 516 ++++++++++++++++++++++++++++++++++
2 files changed, 517 insertions(+)
create mode 100644 include/linux/arm-smccc-rmi.h
diff --git a/MAINTAINERS b/MAINTAINERS
index 6215fcb077705..ed82c55bf566f 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -3953,6 +3953,7 @@ F: Documentation/arch/arm64/
F: arch/arm64/
F: drivers/virt/coco/arm-cca-guest/
F: drivers/virt/coco/pkvm-guest/
+F: include/linux/arm-smccc-rmi.h
F: tools/testing/selftests/arm64/
X: arch/arm64/boot/dts/
X: arch/arm64/configs/defconfig
diff --git a/include/linux/arm-smccc-rmi.h b/include/linux/arm-smccc-rmi.h
new file mode 100644
index 0000000000000..655726b374624
--- /dev/null
+++ b/include/linux/arm-smccc-rmi.h
@@ -0,0 +1,516 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (C) 2023-2026 ARM Ltd.
+ *
+ * The values and structures in this file are from the Realm Management Monitor
+ * specification (DEN0137) version 2.0-bet3:
+ * https://developer.arm.com/documentation/den0137/2-0bet3/
+ */
+
+#ifndef __LINUX_ARM_SMCCC_RMI_H_
+#define __LINUX_ARM_SMCCC_RMI_H_
+
+#include <linux/arm-smccc.h>
+#include <linux/bitfield.h>
+#include <linux/bits.h>
+#include <linux/build_bug.h>
+#include <linux/sizes.h>
+#include <linux/stddef.h>
+
+#include <asm/page.h>
+
+#define SMC_RMI_CALL(func) \
+ ARM_SMCCC_CALL_VAL(ARM_SMCCC_FAST_CALL, \
+ ARM_SMCCC_SMC_64, \
+ ARM_SMCCC_OWNER_STANDARD, \
+ (func))
+
+#define SMC_RMI_VERSION SMC_RMI_CALL(0x0150)
+
+#define SMC_RMI_RTT_DATA_MAP_INIT SMC_RMI_CALL(0x0153)
+
+#define SMC_RMI_REALM_ACTIVATE SMC_RMI_CALL(0x0157)
+#define SMC_RMI_REALM_CREATE SMC_RMI_CALL(0x0158)
+#define SMC_RMI_REALM_DESTROY SMC_RMI_CALL(0x0159)
+#define SMC_RMI_REC_CREATE SMC_RMI_CALL(0x015a)
+#define SMC_RMI_REC_DESTROY SMC_RMI_CALL(0x015b)
+#define SMC_RMI_REC_ENTER SMC_RMI_CALL(0x015c)
+#define SMC_RMI_RTT_CREATE SMC_RMI_CALL(0x015d)
+#define SMC_RMI_RTT_DESTROY SMC_RMI_CALL(0x015e)
+
+#define SMC_RMI_RTT_READ_ENTRY SMC_RMI_CALL(0x0161)
+
+#define SMC_RMI_RTT_DEV_VALIDATE SMC_RMI_CALL(0x0163)
+#define SMC_RMI_PSCI_COMPLETE SMC_RMI_CALL(0x0164)
+#define SMC_RMI_FEATURES SMC_RMI_CALL(0x0165)
+#define SMC_RMI_RTT_FOLD SMC_RMI_CALL(0x0166)
+
+#define SMC_RMI_RTT_INIT_RIPAS SMC_RMI_CALL(0x0168)
+#define SMC_RMI_RTT_SET_RIPAS SMC_RMI_CALL(0x0169)
+#define SMC_RMI_VSMMU_CREATE SMC_RMI_CALL(0x016a)
+#define SMC_RMI_VSMMU_DESTROY SMC_RMI_CALL(0x016b)
+
+#define SMC_RMI_RMM_CONFIG_SET SMC_RMI_CALL(0x016e)
+#define SMC_RMI_PSMMU_IRQ_NOTIFY SMC_RMI_CALL(0x016f)
+#define SMC_RMI_ATTEST_PLAT_TOKEN_REFRESH SMC_RMI_CALL(0x0170)
+
+#define SMC_RMI_PDEV_ABORT SMC_RMI_CALL(0x0174)
+#define SMC_RMI_PDEV_COMMUNICATE SMC_RMI_CALL(0x0175)
+#define SMC_RMI_PDEV_CREATE SMC_RMI_CALL(0x0176)
+#define SMC_RMI_PDEV_DESTROY SMC_RMI_CALL(0x0177)
+#define SMC_RMI_PDEV_GET_STATE SMC_RMI_CALL(0x0178)
+
+#define SMC_RMI_PDEV_STREAM_KEY_REFRESH SMC_RMI_CALL(0x017a)
+#define SMC_RMI_PDEV_SET_PUBKEY SMC_RMI_CALL(0x017b)
+#define SMC_RMI_PDEV_STOP SMC_RMI_CALL(0x017c)
+#define SMC_RMI_RTT_AUX_CREATE SMC_RMI_CALL(0x017d)
+#define SMC_RMI_RTT_AUX_DESTROY SMC_RMI_CALL(0x017e)
+#define SMC_RMI_RTT_AUX_FOLD SMC_RMI_CALL(0x017f)
+
+#define SMC_RMI_VDEV_ABORT SMC_RMI_CALL(0x0185)
+#define SMC_RMI_VDEV_COMMUNICATE SMC_RMI_CALL(0x0186)
+#define SMC_RMI_VDEV_CREATE SMC_RMI_CALL(0x0187)
+#define SMC_RMI_VDEV_DESTROY SMC_RMI_CALL(0x0188)
+#define SMC_RMI_VDEV_GET_STATE SMC_RMI_CALL(0x0189)
+#define SMC_RMI_VDEV_UNLOCK SMC_RMI_CALL(0x018a)
+#define SMC_RMI_RTT_SET_S2AP SMC_RMI_CALL(0x018b)
+
+#define SMC_RMI_VDEV_GET_INTERFACE_REPORT SMC_RMI_CALL(0x01d0)
+#define SMC_RMI_VDEV_GET_MEASUREMENTS SMC_RMI_CALL(0x01d1)
+#define SMC_RMI_VDEV_LOCK SMC_RMI_CALL(0x01d2)
+#define SMC_RMI_VDEV_START SMC_RMI_CALL(0x01d3)
+
+#define SMC_RMI_VSMMU_EVENT_HANDLE SMC_RMI_CALL(0x01d6)
+#define SMC_RMI_PSMMU_ACTIVATE SMC_RMI_CALL(0x01d7)
+#define SMC_RMI_PSMMU_DEACTIVATE SMC_RMI_CALL(0x01d8)
+
+#define SMC_RMI_PSMMU_ST_L2_CREATE SMC_RMI_CALL(0x01db)
+#define SMC_RMI_PSMMU_ST_L2_DESTROY SMC_RMI_CALL(0x01dc)
+#define SMC_RMI_DPT_L0_CREATE SMC_RMI_CALL(0x01dd)
+#define SMC_RMI_DPT_L0_DESTROY SMC_RMI_CALL(0x01de)
+#define SMC_RMI_DPT_L1_CREATE SMC_RMI_CALL(0x01df)
+#define SMC_RMI_DPT_L1_DESTROY SMC_RMI_CALL(0x01e0)
+#define SMC_RMI_GRANULE_TRACKING_GET SMC_RMI_CALL(0x01e1)
+
+#define SMC_RMI_GRANULE_TRACKING_SET SMC_RMI_CALL(0x01e3)
+
+#define SMC_RMI_RMM_CONFIG_GET SMC_RMI_CALL(0x01ec)
+
+#define SMC_RMI_RMM_STATE_GET SMC_RMI_CALL(0x01ee)
+
+#define SMC_RMI_PSMMU_EVENT_CONSUME SMC_RMI_CALL(0x01f0)
+#define SMC_RMI_GRANULE_RANGE_DELEGATE SMC_RMI_CALL(0x01f1)
+#define SMC_RMI_GRANULE_RANGE_UNDELEGATE SMC_RMI_CALL(0x01f2)
+#define SMC_RMI_GPT_L1_CREATE SMC_RMI_CALL(0x01f3)
+#define SMC_RMI_GPT_L1_DESTROY SMC_RMI_CALL(0x01f4)
+#define SMC_RMI_RTT_DATA_MAP SMC_RMI_CALL(0x01f5)
+#define SMC_RMI_RTT_DATA_UNMAP SMC_RMI_CALL(0x01f6)
+#define SMC_RMI_RTT_DEV_MAP SMC_RMI_CALL(0x01f7)
+#define SMC_RMI_RTT_DEV_UNMAP SMC_RMI_CALL(0x01f8)
+#define SMC_RMI_RTT_ARCH_DEV_MAP SMC_RMI_CALL(0x01f9)
+#define SMC_RMI_RTT_ARCH_DEV_UNMAP SMC_RMI_CALL(0x01fa)
+#define SMC_RMI_RTT_UNPROT_MAP SMC_RMI_CALL(0x01fb)
+#define SMC_RMI_RTT_UNPROT_UNMAP SMC_RMI_CALL(0x01fc)
+#define SMC_RMI_RTT_AUX_PROT_MAP SMC_RMI_CALL(0x01fd)
+#define SMC_RMI_RTT_AUX_PROT_UNMAP SMC_RMI_CALL(0x01fe)
+#define SMC_RMI_RTT_AUX_UNPROT_MAP SMC_RMI_CALL(0x01ff)
+#define SMC_RMI_RTT_AUX_UNPROT_UNMAP SMC_RMI_CALL(0x0200)
+#define SMC_RMI_REALM_TERMINATE SMC_RMI_CALL(0x0201)
+#define SMC_RMI_RMM_ACTIVATE SMC_RMI_CALL(0x0202)
+#define SMC_RMI_OP_CONTINUE SMC_RMI_CALL(0x0203)
+#define SMC_RMI_PDEV_STREAM_CONNECT SMC_RMI_CALL(0x0204)
+#define SMC_RMI_PDEV_STREAM_DISCONNECT SMC_RMI_CALL(0x0205)
+#define SMC_RMI_PDEV_STREAM_COMPLETE SMC_RMI_CALL(0x0206)
+#define SMC_RMI_PDEV_STREAM_KEY_PURGE SMC_RMI_CALL(0x0207)
+#define SMC_RMI_OP_MEM_DONATE SMC_RMI_CALL(0x0208)
+#define SMC_RMI_OP_MEM_RECLAIM SMC_RMI_CALL(0x0209)
+#define SMC_RMI_OP_CANCEL SMC_RMI_CALL(0x020a)
+#define SMC_RMI_VSMMU_FEATURES SMC_RMI_CALL(0x020b)
+#define SMC_RMI_VSMMU_CMD_GET SMC_RMI_CALL(0x020c)
+#define SMC_RMI_VSMMU_CMD_COMPLETE SMC_RMI_CALL(0x020d)
+#define SMC_RMI_PSMMU_INFO SMC_RMI_CALL(0x020e)
+#define SMC_RMI_RMM_DEACTIVATE SMC_RMI_CALL(0x020f)
+#define SMC_RMI_PDEV_STREAM_INFO SMC_RMI_CALL(0x0210)
+#define SMC_RMI_GPT_INFO SMC_RMI_CALL(0x0211)
+
+#define RMI_ABI_MAJOR_VERSION 2
+#define RMI_ABI_MINOR_VERSION 0
+
+#define RMI_ABI_VERSION_MAJOR_MASK GENMASK(30, 16)
+#define RMI_ABI_VERSION_MINOR_MASK GENMASK(15, 0)
+
+#define RMI_ABI_VERSION_GET_MAJOR(v) FIELD_GET(RMI_ABI_VERSION_MAJOR_MASK, (v))
+#define RMI_ABI_VERSION_GET_MINOR(v) FIELD_GET(RMI_ABI_VERSION_MINOR_MASK, (v))
+#define RMI_ABI_VERSION(major, minor) \
+ (FIELD_PREP(RMI_ABI_VERSION_MAJOR_MASK, (major)) | \
+ FIELD_PREP(RMI_ABI_VERSION_MINOR_MASK, (minor)))
+
+/* RmiResult Type definitions */
+#define RMI_RESULT_STATUS_MASK GENMASK(7, 0)
+/* RmiResultDataLevel field */
+#define RMI_RESULT_DATA_LEVEL_MASK GENMASK(15, 8)
+/* RmiResultDataIncomplete fields */
+#define RMI_RESULT_MEMREQ_MASK GENMASK(9, 8)
+#define RMI_RESULT_CAN_CANCEL_MASK BIT(10)
+
+#define RMI_RESULT_STATUS(ret) FIELD_GET(RMI_RESULT_STATUS_MASK, ret)
+#define RMI_RESULT_DATA_LEVEL(ret) FIELD_GET(RMI_RESULT_DATA_LEVEL_MASK, ret)
+#define RMI_RESULT_MEMREQ(ret) FIELD_GET(RMI_RESULT_MEMREQ_MASK, ret)
+#define RMI_RESULT_CAN_CANCEL(ret) FIELD_GET(RMI_RESULT_CAN_CANCEL_MASK, ret)
+
+#define RMI_SUCCESS 0
+#define RMI_ERROR_INPUT 1
+#define RMI_ERROR_REALM 2
+#define RMI_ERROR_REC 3
+#define RMI_ERROR_RTT 4
+#define RMI_ERROR_NOT_SUPPORTED 5
+#define RMI_ERROR_DEVICE 6
+#define RMI_ERROR_RTT_AUX 7
+#define RMI_ERROR_PSMMU_ST 8
+#define RMI_ERROR_DPT 9
+#define RMI_BUSY 10
+#define RMI_ERROR_GLOBAL 11
+#define RMI_ERROR_TRACKING 12
+#define RMI_INCOMPLETE 13
+#define RMI_BLOCKED 14
+#define RMI_ERROR_GPT 15
+#define RMI_ERROR_GRANULE 16
+
+#define RMI_CONTINUE_KEEP_GOING 0
+#define RMI_CONTINUE_STOP 1
+
+#define RMI_OP_MEM_REQ_NONE 0
+#define RMI_OP_MEM_REQ_DONATE 1
+#define RMI_OP_MEM_REQ_RECLAIM 2
+
+#define RMI_OP_CANNOT_CANCEL 0
+#define RMI_OP_CAN_CANCEL 1
+
+#define RMI_DONATE_STATE_MASK GENMASK(18, 17)
+#define RMI_DONATE_CONTIG_MASK BIT(16)
+#define RMI_DONATE_COUNT_MASK GENMASK(15, 2)
+#define RMI_DONATE_BLOCK_SIZE_MASK GENMASK(1, 0)
+
+#define RMI_DONATE_STATE(req) FIELD_GET(RMI_DONATE_STATE_MASK, req)
+#define RMI_DONATE_CONTIG(req) FIELD_GET(RMI_DONATE_CONTIG_MASK, req)
+#define RMI_DONATE_COUNT(req) FIELD_GET(RMI_DONATE_COUNT_MASK, req)
+#define RMI_DONATE_BLOCK_SIZE(req) FIELD_GET(RMI_DONATE_BLOCK_SIZE_MASK, req)
+
+#define RMI_OP_MEM_DELEGATED 0
+#define RMI_OP_MEM_UNDELEGATED 1
+#define RMI_OP_MEM_CONDITIONAL 2
+
+#define RMI_OP_MEM_NON_CONTIG 0
+#define RMI_OP_MEM_CONTIG 1
+
+#define RMI_ADDR_TYPE_NONE 0
+#define RMI_ADDR_TYPE_SINGLE 1
+#define RMI_ADDR_TYPE_LIST 2
+
+#define RMI_ADDR_RANGE_STATE_MASK GENMASK_ULL(63, 62)
+#define RMI_ADDR_RANGE_ADDR_MASK GENMASK_ULL(51, PAGE_SHIFT)
+#define RMI_ADDR_RANGE_COUNT_MASK GENMASK(PAGE_SHIFT - 1, 2)
+#define RMI_ADDR_RANGE_BLOCK_SIZE_MASK GENMASK(1, 0)
+
+#define RMI_ADDR_RANGE_STATE(r) FIELD_GET(RMI_ADDR_RANGE_STATE_MASK, (r))
+#define RMI_ADDR_RANGE_ADDR(r) ((r) & RMI_ADDR_RANGE_ADDR_MASK)
+#define RMI_ADDR_RANGE_COUNT(r) FIELD_GET(RMI_ADDR_RANGE_COUNT_MASK, (r))
+#define RMI_ADDR_RANGE_BLOCK_SIZE(r) FIELD_GET(RMI_ADDR_RANGE_BLOCK_SIZE_MASK, (r))
+
+enum rmi_ripas {
+ RMI_EMPTY = 0,
+ RMI_RAM = 1,
+ RMI_DESTROYED = 2,
+ RMI_DEV = 3,
+};
+
+#define RMI_NO_MEASURE_CONTENT 0
+#define RMI_MEASURE_CONTENT 1
+
+#define RMI_FEATURE_REGISTER_0_S2OASZ GENMASK_ULL(40, 33)
+#define RMI_FEATURE_REGISTER_0_L0GPT_BLOCK_DELEGATE BIT_ULL(32)
+#define RMI_FEATURE_REGISTER_0_PMU_NUM_CTRS GENMASK(31, 27)
+#define RMI_FEATURE_REGISTER_0_PMU BIT(26)
+#define RMI_FEATURE_REGISTER_0_NUM_WPS GENMASK(25, 20)
+#define RMI_FEATURE_REGISTER_0_NUM_BPS GENMASK(19, 14)
+#define RMI_FEATURE_REGISTER_0_SVE_VL GENMASK(13, 10)
+#define RMI_FEATURE_REGISTER_0_SVE BIT(9)
+#define RMI_FEATURE_REGISTER_0_LPA2 BIT(8)
+#define RMI_FEATURE_REGISTER_0_S2SZ GENMASK(7, 0)
+
+#define RMI_FEATURE_REGISTER_1_PPS GENMASK(16, 14)
+#define RMI_FEATURE_REGISTER_1_L0GPTSZ GENMASK(13, 10)
+#define RMI_FEATURE_REGISTER_1_MAX_RECS_ORDER GENMASK(9, 6)
+#define RMI_FEATURE_REGISTER_1_HASH_SHA_512 BIT(5)
+#define RMI_FEATURE_REGISTER_1_HASH_SHA_384 BIT(4)
+#define RMI_FEATURE_REGISTER_1_HASH_SHA_256 BIT(3)
+#define RMI_FEATURE_REGISTER_1_RMI_GRAN_SZ_64KB BIT(2)
+#define RMI_FEATURE_REGISTER_1_RMI_GRAN_SZ_16KB BIT(1)
+#define RMI_FEATURE_REGISTER_1_RMI_GRAN_SZ_4KB BIT(0)
+
+#define RMI_FEATURE_REGISTER_2_REALM_MAX_VDEVS_ORDER GENMASK(14, 10)
+#define RMI_FEATURE_REGISTER_2_NON_TEE_STREAM BIT(9)
+#define RMI_FEATURE_REGISTER_2_VDEV_KROU BIT(8)
+#define RMI_FEATURE_REGISTER_2_PDEV_MAX_VDEVS_ORDER GENMASK(7, 4)
+#define RMI_FEATURE_REGISTER_2_ATS BIT(3)
+#define RMI_FEATURE_REGISTER_2_VSMMU BIT(2)
+#define RMI_FEATURE_REGISTER_2_DA_COH BIT(1)
+#define RMI_FEATURE_REGISTER_2_DA BIT(0)
+
+#define RMI_FEATURE_REGISTER_3_RTT_S2AP_INDIRECT BIT(6)
+#define RMI_FEATURE_REGISTER_3_RTT_PLANE GENMASK(5, 4)
+#define RMI_FEATURE_REGISTER_3_MAX_NUM_AUX_PLANES GENMASK(3, 0)
+
+#define RMI_FEATURE_REGISTER_4_MEC_COUNT GENMASK_ULL(63, 0)
+
+#define RMI_MEM_CATEGORY_CONVENTIONAL 0
+#define RMI_MEM_CATEGORY_DEV_NCOH 1
+#define RMI_MEM_CATEGORY_DEV_COH 2
+#define RMI_MEM_CATEGORY_NONE 3
+
+#define RMI_TRACKING_RESERVED 0
+#define RMI_TRACKING_NONE 1
+#define RMI_TRACKING_FINE 2
+#define RMI_TRACKING_COARSE 3
+#define RMI_TRACKING_INTERMEDIATE 4
+
+#define RMI_GRANULE_SIZE_4KB 0
+#define RMI_GRANULE_SIZE_16KB 1
+#define RMI_GRANULE_SIZE_64KB 2
+
+#define RMI_GPT_PAR_RESERVED 0U
+#define RMI_GPT_PAR_PLAT 1U
+#define RMI_GPT_PAR_HOST_NOT_CREATED 2U
+#define RMI_GPT_PAR_HOST_CREATED 3U
+
+/*
+ * Note many of these fields are smaller than u64 but all fields have u64
+ * alignment, so use u64 to ensure correct alignment.
+ */
+struct rmm_config {
+ union { /* 0x0 */
+ struct {
+ u64 tracking_region_size;
+ u64 rmi_granule_size;
+ };
+ u8 sizer[SZ_4K];
+ };
+};
+
+static_assert(sizeof(struct rmm_config) == SZ_4K);
+
+#define RMI_REALM_PARAM_FLAG_MEC_POLICY_MASK GENMASK(8, 7)
+#define RMI_REALM_PARAM_FLAG_LFA_POLICY_MASK GENMASK(6, 5)
+#define RMI_REALM_PARAM_FLAG_DA BIT(3)
+#define RMI_REALM_PARAM_FLAG_PMU BIT(2)
+#define RMI_REALM_PARAM_FLAG_SVE BIT(1)
+
+#define RMI_HASH_SHA_256 0
+#define RMI_HASH_SHA_512 1
+#define RMI_HASH_SHA_384 2
+
+struct realm_params {
+ union { /* 0x0 */
+ struct {
+ u64 flags0;
+ u64 s2sz;
+ u64 sve_vl;
+ u64 num_bps;
+ u64 num_wps;
+ u64 pmu_num_ctrs;
+ u64 hash_algo;
+ u64 num_aux_planes;
+ };
+ u8 sizer0[0x400];
+ };
+ union { /* 0x400 */
+ struct {
+ u8 rpv[64];
+ u64 ats_plane;
+ };
+ u8 sizer1[0x400];
+ };
+ union { /* 0x800 */
+ struct {
+ u64 sizer2;
+ u64 rtt_base;
+ s64 rtt_level_start;
+ u64 rtt_num_start;
+ u64 flags1;
+ u64 max_num_vdevs;
+ };
+ u8 sizer3[0x700];
+ };
+ union { /* 0xf00 */
+ struct {
+ u8 sizer4[0x80];
+ u64 aux_rtt_base[3];
+ };
+ u8 sizer5[0x100];
+ };
+};
+
+static_assert(offsetof(struct realm_params, rpv) == 0x400);
+static_assert(offsetof(struct realm_params, rtt_base) == 0x808);
+static_assert(offsetof(struct realm_params, aux_rtt_base) == 0xf80);
+static_assert(sizeof(struct realm_params) == SZ_4K);
+
+/*
+ * The number of GPRs (starting from X0) that are configured by the host when
+ * a REC is created.
+ */
+#define REC_CREATE_NR_GPRS 8
+
+#define REC_PARAMS_FLAG_RUNNABLE BIT(0)
+
+struct rec_params {
+ union { /* 0x0 */
+ u64 flags;
+ u8 sizer0[0x100];
+ };
+ union { /* 0x100 */
+ u64 mpidr;
+ u8 sizer1[0x100];
+ };
+ union { /* 0x200 */
+ u64 pc;
+ u8 sizer2[0x100];
+ };
+ union { /* 0x300 */
+ u64 gprs[REC_CREATE_NR_GPRS];
+ u8 sizer3[0xd00];
+ };
+};
+
+static_assert(offsetof(struct rec_params, mpidr) == 0x100);
+static_assert(offsetof(struct rec_params, pc) == 0x200);
+static_assert(offsetof(struct rec_params, gprs) == 0x300);
+static_assert(sizeof(struct rec_params) == SZ_4K);
+
+#define REC_ENTER_FLAG_FORCE_P0 BIT(7)
+#define REC_ENTER_FLAG_DEV_MEM_RESPONSE BIT(6)
+#define REC_ENTER_FLAG_S2AP_RESPONSE BIT(5)
+#define REC_ENTER_FLAG_RIPAS_RESPONSE BIT(4)
+#define REC_ENTER_FLAG_TRAP_WFE BIT(3)
+#define REC_ENTER_FLAG_TRAP_WFI BIT(2)
+#define REC_ENTER_FLAG_INJECT_SEA BIT(1)
+#define REC_ENTER_FLAG_EMULATED_MMIO BIT(0)
+
+#define REC_RUN_GPRS 31
+
+struct rec_enter {
+ union { /* 0x000 */
+ u64 flags;
+ u8 sizer0[0x200];
+ };
+ union { /* 0x200 */
+ u64 gprs[REC_RUN_GPRS];
+ u8 sizer1[0x600];
+ };
+};
+
+static_assert(offsetof(struct rec_enter, gprs) == 0x200);
+static_assert(sizeof(struct rec_enter) == SZ_2K);
+
+#define RMI_EXIT_SYNC 0x00
+#define RMI_EXIT_IRQ 0x01
+#define RMI_EXIT_FIQ 0x02
+#define RMI_EXIT_PSCI 0x03
+#define RMI_EXIT_RIPAS_CHANGE 0x04
+#define RMI_EXIT_HOST_CALL 0x05
+#define RMI_EXIT_SERROR 0x06
+#define RMI_EXIT_S2AP_CHANGE 0x07
+#define RMI_EXIT_VDEV_VALIDATE_MAPPING 0x08
+#define RMI_EXIT_VSMMU_COMMAND 0x0a
+
+struct rec_exit {
+ union { /* 0x000 */
+ u8 exit_reason;
+ u8 sizer0[0x100];
+ };
+ union { /* 0x100 */
+ struct {
+ u64 esr;
+ u64 far;
+ u64 hpfar;
+ u64 rtt_tree;
+ };
+ u8 sizer1[0x100];
+ };
+ union { /* 0x200 */
+ u64 gprs[REC_RUN_GPRS];
+ u8 sizer2[0x200];
+ };
+ union { /* 0x400 */
+ struct {
+ u64 cntp_ctl;
+ u64 cntp_cval;
+ u64 cntv_ctl;
+ u64 cntv_cval;
+ };
+ u8 sizer4[0x100];
+ };
+ union { /* 0x500 */
+ struct {
+ u64 ripas_base;
+ u64 ripas_top;
+ u8 ripas_value;
+ u8 sizer5[0xf];
+ u64 s2ap_base;
+ u64 s2ap_top;
+ u64 vdev_id_1;
+ u64 vdev_id_2;
+ u64 dev_mem_base;
+ u64 dev_mem_top;
+ u64 dev_mem_pa;
+ };
+ u8 sizer6[0x100];
+ };
+ union { /* 0x600 */
+ struct {
+ u16 imm;
+ u8 sizer7[0x6];
+ u64 plane;
+ };
+ u8 sizer8[0x100];
+ };
+ union { /* 0x700 */
+ struct {
+ u8 pmu_ovf_status;
+ u8 sizer9[0xf];
+ u64 vsmmu;
+ };
+ u8 sizer10[0x100];
+ };
+};
+
+static_assert(sizeof(struct rec_exit) == SZ_2K);
+static_assert(offsetof(struct rec_exit, esr) == 0x100);
+static_assert(offsetof(struct rec_exit, gprs) == 0x200);
+static_assert(offsetof(struct rec_exit, cntp_ctl) == 0x400);
+static_assert(offsetof(struct rec_exit, ripas_base) == 0x500);
+static_assert(offsetof(struct rec_exit, imm) == 0x600);
+static_assert(offsetof(struct rec_exit, pmu_ovf_status) == 0x700);
+
+struct rec_run {
+ struct rec_enter enter;
+ struct rec_exit exit;
+};
+
+static_assert(sizeof(struct rec_run) == SZ_4K);
+
+/* RMI_RTT_UNPROT_MAP_FLAGS definitions */
+#define RMI_RTT_UNPROT_MAP_FLAGS_S2AP_MASK GENMASK(22, 19)
+#define RMI_RTT_UNPROT_MAP_FLAGS_MEMATTR_MASK GENMASK(18, 16)
+#define RMI_RTT_UNPROT_MAP_FLAGS_LIST_COUNT_MASK GENMASK(15, 2)
+#define RMI_RTT_UNPROT_MAP_FLAGS_OADDR_TYPE_MASK GENMASK(1, 0)
+
+/* RMI_RTT_PROT_MAP_FLAGS definitions */
+#define RMI_RTT_PROT_MAP_FLAGS_LIST_COUNT_MASK GENMASK(15, 2)
+#define RMI_RTT_PROT_MAP_FLAGS_OADDR_TYPE_MASK GENMASK(1, 0)
+
+/* S2AP Direct Encodings, used in RMI_RTT_UNPROT_MAP_FLAGS_S2AP */
+#define RMI_S2AP_DIRECT_READ BIT(1)
+#define RMI_S2AP_DIRECT_WRITE BIT(0)
+
+#endif /* __LINUX_ARM_SMCCC_RMI_H_ */
--
2.43.0
^ permalink raw reply [flat|nested] 23+ messages in thread* Re: [PATCH v20 1/9] firmware: arm_rmm: Add SMC definitions for calling the RMM
2026-09-29 22:16 ` [PATCH v20 1/9] firmware: arm_rmm: Add SMC definitions for calling the RMM Suzuki K Poulose
@ 2026-09-30 9:51 ` Catalin Marinas
0 siblings, 0 replies; 23+ messages in thread
From: Catalin Marinas @ 2026-09-30 9:51 UTC (permalink / raw)
To: Suzuki K Poulose
Cc: kvm, kvmarm, maz, will, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, sudeep.holla,
jonathan.cameron
On Tue, Sep 29, 2026 at 11:16:15PM +0100, Suzuki K Poulose wrote:
> From: Steven Price <steven.price@arm.com>
>
> The RMM (Realm Management Monitor) provides functionality that can be
> accessed by SMC calls from the host.
>
> The SMC definitions are based on DEN0137[1] version 2.0-bet3
>
> [1] https://developer.arm.com/documentation/den0137/2-0bet3/
>
> Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
> Reviewed-by: Gavin Shan <gshan@redhat.com>
> Tested-by: Gavin Shan <gshan@redhat.com>
> Signed-off-by: Steven Price <steven.price@arm.com>
> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
Acked-by: Catalin Marinas <catalin.marinas@arm.com>
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v20 2/9] firmware: arm_rmm: Check for RMI support at init
2026-09-29 22:16 [PATCH v20 0/9] firmware: arm_rmm: Add RMM v2.0 base RMI support Suzuki K Poulose
2026-09-29 22:16 ` [PATCH v20 1/9] firmware: arm_rmm: Add SMC definitions for calling the RMM Suzuki K Poulose
@ 2026-09-29 22:16 ` Suzuki K Poulose
2026-09-29 22:16 ` [PATCH v20 3/9] firmware: arm_rmm: Configure the RMM with the host's page size Suzuki K Poulose
` (6 subsequent siblings)
8 siblings, 0 replies; 23+ messages in thread
From: Suzuki K Poulose @ 2026-09-29 22:16 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, sudeep.holla,
jonathan.cameron, Suzuki K Poulose
From: Steven Price <steven.price@arm.com>
Query the RMI version number and check if it is a compatible version.
The first two feature registers are read and exposed for future code to
use.
We only support this for Little Endian kernels, the Big Endian kernel
support is anyway marked BROKEN and is being removed.
Cc: Marc Zyngier <maz@kernel.org>
Cc: Oliver Upton <oupton@kernel.org>
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Reviewed-by: Gavin Shan <gshan@redhat.com>
Tested-by: Gavin Shan <gshan@redhat.com>
Signed-off-by: Steven Price <steven.price@arm.com>
Co-developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
v20:
* Fix comment : s/RmiFeatureRegister5/RmiFeatureRegister4/
* rmi_feat_reg() s/unsigned long/unsigned int for index
* linux/arm-rmi-cmds.h -> Mark the "#endif"
* Use += rmi.o in Makefile
* Mask the ID_AA64PFR0_EL1_RME field for KVM guests
* Add MAINTAINERS entry for header file and arm_rmm directory
(Potential conflict with Aneesh's smccc series)
v19:
* Read all implemented RmiFeatureRegisters - 5
* Use ARRAY_SIZE(rmi_feat_reg_cache) for the loop in rmi_read_features()
* Fold rmi_features() into rmi_read_features
* Fix comment for rmi_smccc_invoke()
* Drop default y
* Add retry for RMI_BLOCKED and return to caller
v18:
* Always use arm_smccc_1_2_invoke() for all RMIs making sure the unsused
parameters are 0 - Sashiko
* Move rmi_features() calls away from the arm-rmi-cmds.h to rmi.c - Gavin
v17:
* Rename ARM_RMM to ARM_RMM_RMI to make it easier to add Guest facing RSI
support, which is also in progress
v16:
* Update Kconfig text to include PCIe TDISP.
* Export rmi_feat_reg() here rather than in a later commit.
v15:
* The code is moved again, this time into the 'firmware' directory.
v14:
* This moves the basic RMI setup into the 'kernel' directory. This is
because RMI will be used for some features outside of KVM so should
be available even if KVM isn't compiled in.
---
MAINTAINERS | 2 +
arch/arm64/Kconfig | 1 +
arch/arm64/kernel/cpufeature.c | 1 +
arch/arm64/kvm/sys_regs.c | 2 +
drivers/firmware/Kconfig | 1 +
drivers/firmware/Makefile | 1 +
drivers/firmware/arm_rmm/Kconfig | 25 +++++++
drivers/firmware/arm_rmm/Makefile | 1 +
drivers/firmware/arm_rmm/rmi.c | 109 ++++++++++++++++++++++++++++++
include/linux/arm-rmi-cmds.h | 48 +++++++++++++
10 files changed, 191 insertions(+)
create mode 100644 drivers/firmware/arm_rmm/Kconfig
create mode 100644 drivers/firmware/arm_rmm/Makefile
create mode 100644 drivers/firmware/arm_rmm/rmi.c
create mode 100644 include/linux/arm-rmi-cmds.h
diff --git a/MAINTAINERS b/MAINTAINERS
index ed82c55bf566f..1646012706823 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -3951,8 +3951,10 @@ S: Maintained
T: git git://git.kernel.org/pub/scm/linux/kernel/git/arm64/linux.git
F: Documentation/arch/arm64/
F: arch/arm64/
+F: drivers/firmware/arm_rmm/
F: drivers/virt/coco/arm-cca-guest/
F: drivers/virt/coco/pkvm-guest/
+F: include/linux/arm-rmi-cmds.h
F: include/linux/arm-smccc-rmi.h
F: tools/testing/selftests/arm64/
X: arch/arm64/boot/dts/
diff --git a/arch/arm64/Kconfig b/arch/arm64/Kconfig
index b5a51b0ef9440..ff9565d3ffa59 100644
--- a/arch/arm64/Kconfig
+++ b/arch/arm64/Kconfig
@@ -38,6 +38,7 @@ config ARM64
select ARCH_HAS_MEMBARRIER_SYNC_CORE
select ARCH_HAS_MEM_ENCRYPT
select ARCH_SUPPORTS_MSEAL_SYSTEM_MAPPINGS
+ select ARCH_SUPPORTS_RMM
select ARCH_HAS_NMI_SAFE_THIS_CPU_OPS
select ARCH_HAS_NON_OVERLAPPING_ADDRESS_SPACE
select ARCH_HAS_NONLEAF_PMD_YOUNG if ARM64_HAFT
diff --git a/arch/arm64/kernel/cpufeature.c b/arch/arm64/kernel/cpufeature.c
index 32102c3912fa7..e8b29983b0021 100644
--- a/arch/arm64/kernel/cpufeature.c
+++ b/arch/arm64/kernel/cpufeature.c
@@ -293,6 +293,7 @@ static const struct arm64_ftr_bits ftr_id_aa64isar3[] = {
static const struct arm64_ftr_bits ftr_id_aa64pfr0[] = {
ARM64_FTR_BITS(FTR_HIDDEN, FTR_NONSTRICT, FTR_LOWER_SAFE, ID_AA64PFR0_EL1_CSV3_SHIFT, 4, 0),
ARM64_FTR_BITS(FTR_HIDDEN, FTR_NONSTRICT, FTR_LOWER_SAFE, ID_AA64PFR0_EL1_CSV2_SHIFT, 4, 0),
+ ARM64_FTR_BITS(FTR_HIDDEN, FTR_NONSTRICT, FTR_LOWER_SAFE, ID_AA64PFR0_EL1_RME_SHIFT, 4, 0),
ARM64_FTR_BITS(FTR_VISIBLE, FTR_STRICT, FTR_LOWER_SAFE, ID_AA64PFR0_EL1_DIT_SHIFT, 4, 0),
ARM64_FTR_BITS(FTR_HIDDEN, FTR_NONSTRICT, FTR_LOWER_SAFE, ID_AA64PFR0_EL1_AMU_SHIFT, 4, 0),
ARM64_FTR_BITS(FTR_HIDDEN, FTR_STRICT, FTR_LOWER_SAFE, ID_AA64PFR0_EL1_MPAM_SHIFT, 4, 0),
diff --git a/arch/arm64/kvm/sys_regs.c b/arch/arm64/kvm/sys_regs.c
index 44aae52c473d7..f0ef348f725b6 100644
--- a/arch/arm64/kvm/sys_regs.c
+++ b/arch/arm64/kvm/sys_regs.c
@@ -2163,6 +2163,8 @@ static u64 sanitise_id_aa64pfr0_el1(const struct kvm_vcpu *vcpu, u64 val)
*/
val &= ~ID_AA64PFR0_EL1_MPAM_MASK;
+ val &= ~ID_AA64PFR0_EL1_RME_MASK;
+
return val;
}
diff --git a/drivers/firmware/Kconfig b/drivers/firmware/Kconfig
index b7cc11e4fbfa6..62660bf520a8d 100644
--- a/drivers/firmware/Kconfig
+++ b/drivers/firmware/Kconfig
@@ -310,5 +310,6 @@ source "drivers/firmware/samsung/Kconfig"
source "drivers/firmware/smccc/Kconfig"
source "drivers/firmware/tegra/Kconfig"
source "drivers/firmware/xilinx/Kconfig"
+source "drivers/firmware/arm_rmm/Kconfig"
endmenu
diff --git a/drivers/firmware/Makefile b/drivers/firmware/Makefile
index be46f1e1dc77f..196a650ccf025 100644
--- a/drivers/firmware/Makefile
+++ b/drivers/firmware/Makefile
@@ -39,3 +39,4 @@ obj-y += samsung/
obj-y += smccc/
obj-y += tegra/
obj-y += xilinx/
+obj-y += arm_rmm/
diff --git a/drivers/firmware/arm_rmm/Kconfig b/drivers/firmware/arm_rmm/Kconfig
new file mode 100644
index 0000000000000..a37cba6647360
--- /dev/null
+++ b/drivers/firmware/arm_rmm/Kconfig
@@ -0,0 +1,25 @@
+
+config ARCH_SUPPORTS_RMM
+ bool
+
+config ARM_RMM_RMI
+ bool "Realm Management Interface (RMI) Support"
+ depends on ARCH_SUPPORTS_RMM
+ help
+ Support the Realm Management Monitor (RMM) on Arm systems that
+ implement the Realm Management Extension (RME), as defined by the
+ Arm Confidential Compute Architecture.
+
+ The RMM runs at EL2 in the Realm world and provides the Realm
+ Management Interface (RMI) used by a Normal World host to create,
+ manage and run protected virtual machines called Realms. The RMM can
+ also act as a TSM, as defined by the PCIe TDISP and can manage the
+ PCI IDE setup for securing the PCIe links.
+
+ This option builds the host-side RMI support used by KVM to detect a
+ compatible RMM, configure it, manage delegated memory and enable
+ Realm guests.
+
+ Selecting this option does not by itself make Realm guests available:
+ the system must also provide RME-capable hardware and firmware with a
+ compatible RMM implementation.
diff --git a/drivers/firmware/arm_rmm/Makefile b/drivers/firmware/arm_rmm/Makefile
new file mode 100644
index 0000000000000..25e6ad80d3a19
--- /dev/null
+++ b/drivers/firmware/arm_rmm/Makefile
@@ -0,0 +1 @@
+obj-$(CONFIG_ARM_RMM_RMI) += rmi.o
diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/arm_rmm/rmi.c
new file mode 100644
index 0000000000000..90706a16c8836
--- /dev/null
+++ b/drivers/firmware/arm_rmm/rmi.c
@@ -0,0 +1,109 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (C) 2023-2026 ARM Ltd.
+ */
+
+#include <linux/arm-rmi-cmds.h>
+#include <linux/cpufeature.h>
+#include <linux/memblock.h>
+#include <linux/slab.h>
+
+#include <asm/memory.h>
+#include <asm/pgtable-hwdef.h>
+
+/* RMM v2.0 defines RmiFeatureRegister0 to RmiFeatureRegister4. */
+static unsigned long rmi_feat_reg_cache[5] __ro_after_init;
+
+static int rmi_check_version(void)
+{
+ unsigned short version_major, version_minor;
+ unsigned long host_version = RMI_ABI_VERSION(RMI_ABI_MAJOR_VERSION,
+ RMI_ABI_MINOR_VERSION);
+ unsigned long aa64pfr0 = read_sanitised_ftr_reg(SYS_ID_AA64PFR0_EL1);
+ struct arm_smccc_1_2_regs res = {
+ SMC_RMI_VERSION, host_version,
+ };
+
+ /* If RME isn't supported, then RMI can't be */
+ if (cpuid_feature_extract_unsigned_field(aa64pfr0, ID_AA64PFR0_EL1_RME_SHIFT) == 0)
+ return -ENXIO;
+
+ rmi_smccc_invoke(&res);
+ if (res.a0 == SMCCC_RET_NOT_SUPPORTED)
+ return -ENXIO;
+
+ version_major = RMI_ABI_VERSION_GET_MAJOR(res.a1);
+ version_minor = RMI_ABI_VERSION_GET_MINOR(res.a1);
+
+ if (res.a0 != RMI_SUCCESS) {
+ unsigned short high_version_major, high_version_minor;
+
+ high_version_major = RMI_ABI_VERSION_GET_MAJOR(res.a2);
+ high_version_minor = RMI_ABI_VERSION_GET_MINOR(res.a2);
+
+ pr_err("Unsupported RMI ABI (v%d.%d - v%d.%d) we want v%d.%d\n",
+ version_major, version_minor,
+ high_version_major, high_version_minor,
+ RMI_ABI_MAJOR_VERSION,
+ RMI_ABI_MINOR_VERSION);
+ return -ENXIO;
+ }
+
+ pr_info("RMI ABI version %d.%d\n", version_major, version_minor);
+
+ return 0;
+}
+
+static int rmi_read_features(void)
+{
+ /*
+ * Since we've negotiated a compatible version these feature registers
+ * should always be available
+ */
+ for (int i = 0; i < ARRAY_SIZE(rmi_feat_reg_cache); i++) {
+ struct arm_smccc_1_2_regs args = {
+ SMC_RMI_FEATURES, i,
+ };
+
+ rmi_smccc_invoke(&args);
+ if (WARN_ON(args.a0 != RMI_SUCCESS))
+ return -EINVAL;
+
+ rmi_feat_reg_cache[i] = args.a1;
+ }
+
+ return 0;
+}
+
+unsigned long rmi_feat_reg(unsigned int index)
+{
+ if (WARN_ON(index >= ARRAY_SIZE(rmi_feat_reg_cache)))
+ return 0;
+
+ return rmi_feat_reg_cache[index];
+}
+EXPORT_SYMBOL_GPL(rmi_feat_reg);
+
+
+static int __init arm64_init_rmi(void)
+{
+ int ret;
+
+ /* If we can't agree on the RMI ABI version, don't proceed further */
+ ret = rmi_check_version();
+ if (ret)
+ return ret;
+
+ ret = rmi_read_features();
+ if (ret)
+ return ret;
+
+ return 0;
+}
+
+/*
+ * Note arm64_init_rmi() must be called before kvm_init_rmi() otherwise KVM
+ * will not support realm guests. subsys_initcall() is called before
+ * module_init() (used for KVM) so this is OK.
+ */
+subsys_initcall(arm64_init_rmi);
diff --git a/include/linux/arm-rmi-cmds.h b/include/linux/arm-rmi-cmds.h
new file mode 100644
index 0000000000000..911489b636522
--- /dev/null
+++ b/include/linux/arm-rmi-cmds.h
@@ -0,0 +1,48 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (C) 2026 ARM Ltd.
+ */
+
+#ifndef __LINUX_ARM_RMI_CMDS_H_
+#define __LINUX_ARM_RMI_CMDS_H_
+
+#include <linux/arm-smccc-rmi.h>
+#include <linux/bug.h>
+#include <linux/processor.h>
+#include <linux/types.h>
+
+#define RMM_BLOCKED_RETRY_COUNT 2
+/*
+ * rmi_smccc_invoke: Invoke the RMI call and return the results, retrying the
+ * command when status is RMI_BUSY. If we encounter RMI_BLOCKED, we retry
+ * it one more time before we give up. The caller is supposed to handle the
+ * result and reissue if required.
+ *
+ * We don't expect to see RMI_BLOCKED on a practical system, except when
+ * there are parallel requests that results in long standing operation,
+ * with one blocking the other.
+ *
+ * @regs: Input parameters filled in. Updated with the ouptput results
+ * after the call.
+ */
+static inline void rmi_smccc_invoke(struct arm_smccc_1_2_regs *regs)
+{
+ struct arm_smccc_1_2_regs args = *regs;
+ long status;
+ int i = 0;
+
+ while (i < RMM_BLOCKED_RETRY_COUNT) {
+ arm_smccc_1_2_invoke(&args, regs);
+
+ status = RMI_RESULT_STATUS(regs->a0);
+ if (status != RMI_BUSY && status != RMI_BLOCKED)
+ break;
+ if (status == RMI_BLOCKED)
+ i++;
+ cpu_relax();
+ }
+}
+
+unsigned long rmi_feat_reg(unsigned int index);
+
+#endif /* __LINUX_ARM_RMI_CMDS_H_ */
--
2.43.0
^ permalink raw reply [flat|nested] 23+ messages in thread* [PATCH v20 3/9] firmware: arm_rmm: Configure the RMM with the host's page size
2026-09-29 22:16 [PATCH v20 0/9] firmware: arm_rmm: Add RMM v2.0 base RMI support Suzuki K Poulose
2026-09-29 22:16 ` [PATCH v20 1/9] firmware: arm_rmm: Add SMC definitions for calling the RMM Suzuki K Poulose
2026-09-29 22:16 ` [PATCH v20 2/9] firmware: arm_rmm: Check for RMI support at init Suzuki K Poulose
@ 2026-09-29 22:16 ` Suzuki K Poulose
2026-09-30 11:10 ` Catalin Marinas
2026-09-29 22:16 ` [PATCH v20 4/9] firmware: arm_rmm: Add support for SRO Suzuki K Poulose
` (5 subsequent siblings)
8 siblings, 1 reply; 23+ messages in thread
From: Suzuki K Poulose @ 2026-09-29 22:16 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, sudeep.holla,
jonathan.cameron, Suzuki K Poulose
From: Steven Price <steven.price@arm.com>
RMM v2.0 brings the ability to set the RMM's granule size. Check the
feature registers and configure the RMM so that it matches the host's
page size. This means that operations can be done with a granularity
equal to PAGE_SIZE.
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Reviewed-by: Gavin Shan <gshan@redhat.com>
Tested-by: Gavin Shan <gshan@redhat.com>
Signed-off-by: Steven Price <steven.price@arm.com>
Co-developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Changes since v19:
* Don't initialize 'ret' in rmi_configure, add explicit return 0/-EINVAL
Changes since v18:
* Move the rmi_config_set closer to rmi_configure, drop the comment
* Use __free() magic for rmm_config object
Changes since v17:
* Move rmi_config_set() out of the header file.
* Print the error message for rmi_config_set if it fails
Changes since v15:
* Actually check the feature register for the host's page-size support.
Changes since v14:
* Move the implementation into drivers/firmware/arm_rmm.
Changes since v13:
* Moved out of KVM.
---
drivers/firmware/arm_rmm/rmi.c | 70 ++++++++++++++++++++++++++++++++++
1 file changed, 70 insertions(+)
diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/arm_rmm/rmi.c
index 90706a16c8836..b558965a05a2c 100644
--- a/drivers/firmware/arm_rmm/rmi.c
+++ b/drivers/firmware/arm_rmm/rmi.c
@@ -84,6 +84,72 @@ unsigned long rmi_feat_reg(unsigned int index)
}
EXPORT_SYMBOL_GPL(rmi_feat_reg);
+static int rmi_rmm_config_set(unsigned long cfg_ptr)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_RMM_CONFIG_SET, cfg_ptr,
+ };
+
+ rmi_smccc_invoke(®s);
+
+ return regs.a0;
+}
+
+static int rmi_configure(void)
+{
+ unsigned long granule_feature;
+ unsigned long granule_size;
+ int ret;
+
+ switch (PAGE_SIZE) {
+ case SZ_4K:
+ granule_size = RMI_GRANULE_SIZE_4KB;
+ granule_feature = RMI_FEATURE_REGISTER_1_RMI_GRAN_SZ_4KB;
+ break;
+ case SZ_16K:
+ granule_size = RMI_GRANULE_SIZE_16KB;
+ granule_feature = RMI_FEATURE_REGISTER_1_RMI_GRAN_SZ_16KB;
+ break;
+ case SZ_64K:
+ granule_size = RMI_GRANULE_SIZE_64KB;
+ granule_feature = RMI_FEATURE_REGISTER_1_RMI_GRAN_SZ_64KB;
+ break;
+ default:
+ BUILD_BUG();
+ }
+
+ if (!(rmi_feat_reg(1) & granule_feature)) {
+ pr_err("RMM does not support %luKB granules\n",
+ PAGE_SIZE >> 10);
+ return -ENXIO;
+ }
+
+ struct rmm_config *config __free(free_page) =
+ (struct rmm_config *)get_zeroed_page(GFP_KERNEL);
+
+ if (!config) {
+ pr_err("Unable to allocate memory for RMM config\n");
+ return -ENOMEM;
+ }
+
+ config->rmi_granule_size = granule_size;
+
+ /*
+ * For now we set the tracking_region_size to 0 which is the only option
+ * for 4KB PAGE_SIZE (1GB for 4KB PAGE_SIZE, 32MB/512MB for 16KB/64KB).
+ * TODO: Support other tracking sizes via Kconfig option for other
+ * PAGE_SIZES
+ */
+ config->tracking_region_size = 0;
+
+ ret = rmi_rmm_config_set(virt_to_phys(config));
+ if (ret) {
+ pr_err("RMM config set failed (%d)\n", ret);
+ return -EINVAL;
+ }
+
+ return 0;
+}
static int __init arm64_init_rmi(void)
{
@@ -98,6 +164,10 @@ static int __init arm64_init_rmi(void)
if (ret)
return ret;
+ ret = rmi_configure();
+ if (ret)
+ return ret;
+
return 0;
}
--
2.43.0
^ permalink raw reply [flat|nested] 23+ messages in thread* Re: [PATCH v20 3/9] firmware: arm_rmm: Configure the RMM with the host's page size
2026-09-29 22:16 ` [PATCH v20 3/9] firmware: arm_rmm: Configure the RMM with the host's page size Suzuki K Poulose
@ 2026-09-30 11:10 ` Catalin Marinas
0 siblings, 0 replies; 23+ messages in thread
From: Catalin Marinas @ 2026-09-30 11:10 UTC (permalink / raw)
To: Suzuki K Poulose
Cc: kvm, kvmarm, maz, will, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, sudeep.holla,
jonathan.cameron
On Tue, Sep 29, 2026 at 11:16:17PM +0100, Suzuki K Poulose wrote:
> From: Steven Price <steven.price@arm.com>
>
> RMM v2.0 brings the ability to set the RMM's granule size. Check the
> feature registers and configure the RMM so that it matches the host's
> page size. This means that operations can be done with a granularity
> equal to PAGE_SIZE.
>
> Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
> Reviewed-by: Gavin Shan <gshan@redhat.com>
> Tested-by: Gavin Shan <gshan@redhat.com>
> Signed-off-by: Steven Price <steven.price@arm.com>
> Co-developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
Reviewed-by: Catalin Marinas <catalin.marinas@arm.com>
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v20 4/9] firmware: arm_rmm: Add support for SRO
2026-09-29 22:16 [PATCH v20 0/9] firmware: arm_rmm: Add RMM v2.0 base RMI support Suzuki K Poulose
` (2 preceding siblings ...)
2026-09-29 22:16 ` [PATCH v20 3/9] firmware: arm_rmm: Configure the RMM with the host's page size Suzuki K Poulose
@ 2026-09-29 22:16 ` Suzuki K Poulose
2026-09-30 13:15 ` Catalin Marinas
2026-09-29 22:16 ` [PATCH v20 5/9] firmware: arm_rmm: Activate the RMM Suzuki K Poulose
` (4 subsequent siblings)
8 siblings, 1 reply; 23+ messages in thread
From: Suzuki K Poulose @ 2026-09-29 22:16 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, sudeep.holla,
jonathan.cameron, Suzuki K Poulose
From: Steven Price <steven.price@arm.com>
RMM v2.0 introduces the concept of "Stateful RMI Operations" (SRO). This
means that an SMC can return with an operation still in progress. The
host is expected to continue the operation until it reaches a conclusion
(either success or failure). During this process the RMM can request
additional memory ('donate') or hand memory back to the host
('reclaim'). The host can request an in progress operation is cancelled,
but still continue the operation until it has completed (otherwise the
incomplete operation may cause future RMM operations to fail).
All memory donate operations are cancellable and if the RMM request cannot
be satisfied, we cancel the request and fail the operation. (e.g., if the
allocation size > MAX_PAGE_ORDER)
The SRO is tracked using a struct rmi_sro_state object which keeps track
of any memory which has been allocated but not yet consumed by the RMM
or reclaimed from the RMM. This allows the memory to be reused in a
future request within the same operation. It will also permit an
operation to be done in a context where memory allocation may be
difficult (e.g. atomic context) with the option to abort the operation
and retry the memory allocation outside of the atomic context. The
memory stored in the struct rmi_sro_state object can then be reused on
the subsequent attempt.
Wrappers for SRO RMI commands are also provided here because they depend
on the rmi_sro_execute() implementation added by this patch.
Delegate/undelegate handles are also added here because they now use the
SRO/stateful command infrastructure and are also used for the memory
DONATE/RECLAIM flows.
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Reviewed-by: Gavin Shan <gshan@redhat.com>
Tested-by: Gavin Shan <gshan@redhat.com>
Signed-off-by: Steven Price <steven.price@arm.com>
Co-developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
v20:
* Fix typos s/-ECANCELLED/ECANCELED/ rmi_sro_execute()
* Drop duplicate, incorrect assignment of sro_handle in rmi_sro_execute()
* Drop int i, redeclartion in for() for rmi_sro_donate_noncontig()
* Drop donated_granules for rmi_sro_donate_contig()
* Drop "inline" and be consistent with the rest of the code
* Avoid nesting if..else in rmi_*delegate_range by breaking out on error
* Consistent styling in comments, wrap to 80space, spelling corrections
* s/\*\*/\*/ for comment above functions
* Reorder the functions to appear as rmi_*_delegate* followed by rmi_*_undelegate*
* WARN_ON if the RMM MemDonateReq:count exceeds the RmiAddrRangeDesc:count
* Reject requests larger than MAX_PAGE_ORDER
* Reject the reserved state and RMI_OP_MEM_CONDITIONAL. Flip to
check the states we support
v19:
* Clamp the mem donate request count to RMI_MAX_ADDR_LIST to prevent overflow
for non_contig requests
* Donate gathered memory when sro list runs out of space for non_contig case
* Add comments for rmi_delegate_range(), rmi_free_delegated_page()
* Fix handling of buggy RMM for rmi_delegate_range()
* Move rmi_granule_*delegate_range() closer to their callers.
* Clarify the requirements for rmi_granule_delegate_range()
* Fix return result to -ENXIO if the MEM_OP is unknown
* Handle unsupported RMI_OP_MEM_CONDITIONAL
* Rename "out" label to "mem_donate"
* Switch to use while loop for rmi_sro_donate_noncontig(), add comments
where we gather the cached entries
* If we don't have capacity in the SRO object, try with what we managed to
collect for non-contiguous requests
* Switch to for() loop for allocation of granules
* Add comment, make code reader friendly for the caching the remaining
entries after the memory donate
* Add documentation for rmi_sro_execute()
v18:
* Prevent overflow for donated_granules output from buggy RMM
* Handle corrupted addr_count in the sro
* Avoid spilling literal pools on stack with sro initialisation
* Handle buggy RMM when the out_top is not changed with RMI_SUCCESS
for delegat/undelegate range calls
* Rename free_delegated_page => rmi_free_delegated_page
* Rename donate_req_to_unit_size => donate_req_to_block_size
* Introduce rmi_addr_block_size_to_bytes() helper to convert a RmiAddrBlockSize
encoding used in RMI_DONATE_REQ and RMI_ADDR_RANGE Descriptors, replaces
donate_req_to_unit_size()
* Rename unit_size => block_size_fld, unit_size_bytes => block_size etc.
* Explicitly check for MEM_CONTIG/CAN_CANCEL fields to match the RMM spec values.
* Rename free_delegated_page => rmi_free_delegated_page()
* Drop RMI_BUSY, RMI_BLOCKED checks from rmi_*delegate_range as they are already
handled by the rmi_smccc_invoke() used by the SRO.
* Add a helper to free an address range entry, which may be partially consumed.
* Ensure RMI_OP_RECLAIM output is valid before consumption
v17:
* Handle buggy RMM firmware to avoid looping forever for non-cancellable SROs.
* Add comment (for the AI agents) to clarify that all memory donating SROs are
cancellable.
v16:
* Wrappers for realm guests split into a separate patch.
* Better support for cancellation - previously a cancelled operation
could be treated as successful.
* Consistently use a signed type for wrapper return values so that
Linux error codes can be returned as well as RMI return values.
v15:
* Wrappers for SRO RMI functions are provided in this patch due to
their dependency on the SRO infrastructure.
* Fold the range delegate/undelegate wrappers into this patch because
they depend on the stateful command infrastructure.
* Add cpu_relax() calls when RMI_BUSY/RMI_BLOCKED is returned.
* Various fixes.
v14:
* SRO support has improved although is still not fully complete. The
infrastructure has been moved out of KVM.
---
drivers/firmware/arm_rmm/rmi.c | 682 +++++++++++++++++++++++++++++++++
include/linux/arm-rmi-cmds.h | 41 ++
2 files changed, 723 insertions(+)
diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/arm_rmm/rmi.c
index b558965a05a2c..8d839e9e61424 100644
--- a/drivers/firmware/arm_rmm/rmi.c
+++ b/drivers/firmware/arm_rmm/rmi.c
@@ -6,6 +6,7 @@
#include <linux/arm-rmi-cmds.h>
#include <linux/cpufeature.h>
#include <linux/memblock.h>
+#include <linux/mmzone.h>
#include <linux/slab.h>
#include <asm/memory.h>
@@ -14,6 +15,687 @@
/* RMM v2.0 defines RmiFeatureRegister0 to RmiFeatureRegister4. */
static unsigned long rmi_feat_reg_cache[5] __ro_after_init;
+/*
+ * rmi_granule_range_delegate() - Delegate granules
+ * @base: PA of the first granule of the range
+ * @top: PA of the first granule after the range
+ * @out_top: PA of the first granule not delegated
+ *
+ * Delegate a range of granule for use by the realm world. If the entire range
+ * was delegated then @out_top == @top, otherwise the function should be called
+ * again with @base == @out_top.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static long rmi_granule_range_delegate(unsigned long base,
+ unsigned long top,
+ unsigned long *out_top)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_GRANULE_RANGE_DELEGATE, base, top
+ };
+ long ret = rmi_sro_execute(®s);
+
+ if (ret == RMI_SUCCESS && out_top)
+ *out_top = regs.a1;
+
+ return ret;
+}
+
+/*
+ * rmi_delegate_range: Delegate a physically contiguous range.
+ * We iterate over the range until we hit an error. So we may
+ * return an error, but with a partially delegated range. The
+ * caller must always look at the @out_phys to figure out, how
+ * much progress was made.
+ *
+ * @phys: Base of the physical address range
+ * @size: Size of the physical address range
+ * @out_phys: Top of the range that was completed. This is always
+ * valid, irrespective of the result.
+ *
+ * Returns RMI_SUCCESS on successful completion. Otherwise, returns
+ * the Linux error number or the RMI status code as described
+ * by the RMM spec for RMI_GRANULE_DELEGATE_RANGE or RMI_BLOCKED.
+ */
+int rmi_delegate_range(phys_addr_t phys,
+ unsigned long size,
+ phys_addr_t *out_phys)
+{
+ long ret = 0;
+ unsigned long top = phys + size;
+ unsigned long out_top;
+
+ while (phys < top) {
+ ret = rmi_granule_range_delegate(phys, top, &out_top);
+
+ if (ret != RMI_SUCCESS)
+ break;
+ /*
+ * Buggy RMM ? Let the caller handle the failure.
+ * We can't know how far the RMM delegated in this
+ * iteration, so we return the best known good limit.
+ * RMM can deal with granules already in "undelegated"
+ * in a given range. So, it is fine for the caller to
+ * try the range we return.
+ */
+ if (WARN_ON(out_top <= phys)) {
+ ret = -ENXIO;
+ break;
+ }
+ phys = out_top;
+ }
+
+ if (out_phys)
+ *out_phys = phys;
+
+ return ret;
+}
+EXPORT_SYMBOL_GPL(rmi_delegate_range);
+
+/*
+ * rmi_granule_range_undelegate() - Undelegate a range of granules
+ * @base: Base PA of the target range
+ * @top: Top PA of the target range
+ * @out_top: Returns the top PA of range whose state is undelegated
+ *
+ * Undelegate a range of granules to allow use by the normal world. Will fail
+ * if the granules are in use by RMM. RMM can ignore granules that are already
+ * undelegated and thus is safe to be called on a range with a mix of delegated
+ * and undelegated granules.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static long rmi_granule_range_undelegate(unsigned long base,
+ unsigned long top,
+ unsigned long *out_top)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_GRANULE_RANGE_UNDELEGATE, base, top
+ };
+ long ret = rmi_sro_execute(®s);
+
+ if (ret == RMI_SUCCESS && out_top)
+ *out_top = regs.a1;
+
+ return ret;
+}
+
+int rmi_undelegate_range(phys_addr_t phys,
+ unsigned long size)
+{
+ long ret = 0;
+ unsigned long top = phys + size;
+ unsigned long out_top;
+
+ while (phys < top) {
+ ret = rmi_granule_range_undelegate(phys, top, &out_top);
+
+ if (ret != RMI_SUCCESS)
+ break;
+ /* Buggy RMM ? Let the caller leak the pages */
+ if (WARN_ON(out_top <= phys))
+ return -ENXIO;
+ phys = out_top;
+ }
+
+ return ret;
+}
+EXPORT_SYMBOL_GPL(rmi_undelegate_range);
+
+/*
+ * rmi_free_delegated_page: Undelegate and free a page that has been previously
+ * delegated to the Realm world. If we are unable to undelegate it, the page is
+ * leaked.
+ * NOTE: Do not use this helper if the page could be concurrently operated by
+ * another thread, as it may get leaked if the undelegation fails due to RMI_BLOCKED
+ */
+int rmi_free_delegated_page(phys_addr_t phys)
+{
+ if (WARN_ON_ONCE(rmi_undelegate_page(phys))) {
+ /* Undelegate failed: leak the page */
+ return -EBUSY;
+ }
+
+ free_page((unsigned long)phys_to_virt(phys));
+
+ return 0;
+}
+EXPORT_SYMBOL_GPL(rmi_free_delegated_page);
+
+/*
+ * Convert the RmiAddrBlockSize to actual size. This is used in RmiDonateReq
+ * and RmiAddrRangeDesc*.
+ */
+static unsigned long rmi_addr_block_size_to_bytes(unsigned long block_size_fld)
+{
+ return BIT(ARM64_HW_PGTABLE_LEVEL_SHIFT(3 - block_size_fld));
+}
+
+/*
+ * free_addr_range: Free memory described by the address range entry, which may
+ * be partially consumed by RMM.
+ *
+ * @entry: RMI_ADDR_RANGE descriptor
+ * @consumed_size: Page aligned size consumed by the RMM from the address range.
+ *
+ * If the state of the address is DELEGATED, undelegate it back, before freeing.
+ * Leaks the memory if we cannot undelegate the range.
+ */
+static void free_addr_range(unsigned long entry, unsigned long consumed_size)
+{
+ unsigned long phys = RMI_ADDR_RANGE_ADDR(entry);
+ unsigned long block_size_fld = RMI_ADDR_RANGE_BLOCK_SIZE(entry);
+ unsigned long count = RMI_ADDR_RANGE_COUNT(entry);
+ unsigned long state = RMI_ADDR_RANGE_STATE(entry);
+ unsigned long size = rmi_addr_block_size_to_bytes(block_size_fld) * count;
+
+ WARN_ON(!PAGE_ALIGNED(phys) || !PAGE_ALIGNED(consumed_size));
+
+ /* We shouldn't see this in reclaim path, leak it for now */
+ if (WARN_ON((state != RMI_OP_MEM_DELEGATED) &&
+ (state != RMI_OP_MEM_UNDELEGATED)))
+ return;
+
+ /* Adjust the address and size for partially consumed entry */
+ phys += consumed_size;
+ size -= consumed_size;
+ /*
+ * Undelegate the pages back if required. If we can't
+ * change them back, leak the pages.
+ */
+ if (state == RMI_OP_MEM_DELEGATED &&
+ WARN_ON(rmi_undelegate_range(phys, size)))
+ return;
+ free_pages_exact(phys_to_virt(phys), size);
+}
+
+static void rmi_op_continue(unsigned long sro_handle, unsigned long flags,
+ struct arm_smccc_1_2_regs *out_regs)
+{
+ *out_regs = (struct arm_smccc_1_2_regs) {
+ SMC_RMI_OP_CONTINUE, sro_handle, flags
+ };
+
+ rmi_smccc_invoke(out_regs);
+}
+
+static void rmi_op_cancel(unsigned long sro_handle,
+ struct arm_smccc_1_2_regs *out_regs)
+{
+ *out_regs = (struct arm_smccc_1_2_regs) {
+ SMC_RMI_OP_CANCEL, sro_handle
+ };
+
+ rmi_smccc_invoke(out_regs);
+}
+
+static void rmi_op_mem_donate(unsigned long sro_handle, unsigned long list_addr,
+ unsigned long list_count, unsigned long flags,
+ struct arm_smccc_1_2_regs *out_regs)
+{
+ *out_regs = (struct arm_smccc_1_2_regs) {
+ SMC_RMI_OP_MEM_DONATE, sro_handle, list_addr, list_count, flags
+ };
+
+ /*
+ * The output donated count (a1) is always valid, irrespective
+ * of the return result. i.e., 0 if there was an error
+ */
+ rmi_smccc_invoke(out_regs);
+}
+
+static void rmi_op_mem_reclaim(unsigned long sro_handle,
+ unsigned long list_addr,
+ unsigned long list_count,
+ struct arm_smccc_1_2_regs *out_regs)
+{
+ *out_regs = (struct arm_smccc_1_2_regs) {
+ SMC_RMI_OP_MEM_RECLAIM, sro_handle, list_addr, list_count
+ };
+
+ rmi_smccc_invoke(out_regs);
+}
+
+static int rmi_sro_ensure_capacity(struct rmi_sro_state *sro,
+ unsigned long count)
+{
+ if (WARN_ON_ONCE(sro->addr_count > RMI_MAX_ADDR_LIST))
+ return -EOVERFLOW;
+
+ if (count > RMI_MAX_ADDR_LIST - sro->addr_count)
+ return -ENOSPC;
+
+ return 0;
+}
+
+static int rmi_sro_donate_contig(struct rmi_sro_state *sro,
+ unsigned long sro_handle,
+ unsigned long donatereq,
+ struct arm_smccc_1_2_regs *out_regs,
+ gfp_t gfp)
+{
+ unsigned long block_size_fld = RMI_DONATE_BLOCK_SIZE(donatereq);
+ unsigned long block_size = rmi_addr_block_size_to_bytes(block_size_fld);
+ unsigned long count = RMI_DONATE_COUNT(donatereq);
+ unsigned long state = RMI_DONATE_STATE(donatereq);
+ unsigned long size = block_size * count;
+ unsigned long addr_range;
+ unsigned long donated_size;
+ int ret;
+ void *virt;
+ phys_addr_t phys;
+
+ /*
+ * The RMM specification requires contiguous allocations are always a
+ * power of 2
+ */
+ if (WARN_ON_ONCE(!is_power_of_2(size)))
+ return -EINVAL;
+ /*
+ * RMM clamps the Maximum value of RmiOpMemDonateReq:count to prevent
+ * overflow in the RMI_ADDR_RANGE_COUNT field.
+ */
+ if (WARN_ON_ONCE(count > (BIT(PAGE_SHIFT - 2) - 1)))
+ return -EINVAL;
+
+ /*
+ * We can't allocate blocks more than what the buddy allocator can
+ * satisfy. TODO: Add support for larger allocations.
+ */
+ if (get_order(size) > MAX_PAGE_ORDER)
+ return -ENOMEM;
+
+ /* Reuse the cached address range if we have one */
+ for (int i = 0; i < sro->addr_count; i++) {
+ unsigned long entry = sro->addr_list[i];
+
+ if (RMI_ADDR_RANGE_BLOCK_SIZE(entry) == block_size_fld &&
+ RMI_ADDR_RANGE_COUNT(entry) == count &&
+ RMI_ADDR_RANGE_STATE(entry) == state &&
+ IS_ALIGNED(RMI_ADDR_RANGE_ADDR(entry), size)) {
+ sro->addr_count--;
+ swap(sro->addr_list[sro->addr_count],
+ sro->addr_list[i]);
+
+ goto mem_donate;
+ }
+ }
+
+ ret = rmi_sro_ensure_capacity(sro, 1);
+ if (ret)
+ return ret;
+
+ virt = alloc_pages_exact(size, gfp);
+ if (!virt)
+ return -ENOMEM;
+ phys = virt_to_phys(virt);
+
+ if (state == RMI_OP_MEM_DELEGATED) {
+ phys_addr_t delegated_phys;
+
+ if (rmi_delegate_range(phys, size, &delegated_phys)) {
+ if (!rmi_undelegate_range(phys, delegated_phys - phys))
+ free_pages_exact(virt, size);
+ return -ENXIO;
+ }
+ }
+
+ addr_range = phys & RMI_ADDR_RANGE_ADDR_MASK;
+ FIELD_MODIFY(RMI_ADDR_RANGE_BLOCK_SIZE_MASK, &addr_range, block_size_fld);
+ FIELD_MODIFY(RMI_ADDR_RANGE_COUNT_MASK, &addr_range, count);
+ FIELD_MODIFY(RMI_ADDR_RANGE_STATE_MASK, &addr_range, state);
+
+ sro->addr_list[sro->addr_count] = addr_range;
+
+mem_donate:
+ rmi_op_mem_donate(sro_handle,
+ virt_to_phys(&sro->addr_list[sro->addr_count]), 1,
+ 0, out_regs);
+ donated_size = out_regs->a1 << PAGE_SHIFT;
+
+ if (WARN_ON(donated_size > size))
+ donated_size = size;
+
+ /* All granules consumed by the RMM */
+ if (donated_size == size)
+ return 0;
+ /* No granules were consumed by the RMM, cache them */
+ if (donated_size == 0) {
+ sro->addr_count++;
+ return 0;
+ }
+
+ /* The granules were partially consumed, reclaim the unused ones. */
+ free_addr_range(sro->addr_list[sro->addr_count], donated_size);
+
+ return 0;
+}
+
+static int rmi_sro_donate_noncontig(struct rmi_sro_state *sro,
+ unsigned long sro_handle,
+ unsigned long donatereq,
+ struct arm_smccc_1_2_regs *out_regs,
+ gfp_t gfp)
+{
+ unsigned long block_size_fld = RMI_DONATE_BLOCK_SIZE(donatereq);
+ unsigned long block_size = rmi_addr_block_size_to_bytes(block_size_fld);
+ unsigned long count = RMI_DONATE_COUNT(donatereq);
+ unsigned long state = RMI_DONATE_STATE(donatereq);
+ unsigned long found = 0;
+ unsigned long donated_granules;
+ unsigned long granules_per_block = block_size >> PAGE_SHIFT;
+ unsigned long consumed_blocks;
+ int addr_list_start = sro->addr_count;
+ int ret, i, src;
+
+ /*
+ * We can't allocate blocks more than what the buddy allocator can
+ * satisfy. TODO: Add support for larger allocations.
+ */
+ if (get_order(block_size) > MAX_PAGE_ORDER)
+ return -ENOMEM;
+ /*
+ * Maximum value of RmiOpMemDonateReq:count is 2^(PAGE_SHIFT-2) - 1.
+ * But we further clamp it down by the number of entries we can do
+ * in one go. The RMM can request the remaining in the next iteration.
+ */
+ if (count > RMI_MAX_ADDR_LIST)
+ count = RMI_MAX_ADDR_LIST;
+
+ /* Gather the suitable entries to the end of the list */
+ i = 0;
+ while (i < addr_list_start && found < count) {
+ unsigned long entry = sro->addr_list[i];
+
+ if (RMI_ADDR_RANGE_BLOCK_SIZE(entry) == block_size_fld &&
+ RMI_ADDR_RANGE_COUNT(entry) == 1 &&
+ RMI_ADDR_RANGE_STATE(entry) == state) {
+ addr_list_start--;
+ swap(sro->addr_list[addr_list_start],
+ sro->addr_list[i]);
+ found++;
+ /* Continue from the swapped in entry */
+ continue;
+ }
+ /* Skip past the entry */
+ i++;
+ }
+
+ ret = rmi_sro_ensure_capacity(sro, count - found);
+ if (ret) {
+ /* If we have found some entries, donate them and try again */
+ if (found)
+ goto mem_donate;
+ /* Otherwise free up the list and start again */
+ rmi_sro_free(sro);
+ /* Reset the addr_list_start to match sro->addr_count */
+ addr_list_start = 0;
+ }
+
+ for (; found < count; found++) {
+ unsigned long addr_range;
+ void *virt = alloc_pages_exact(block_size, gfp);
+ phys_addr_t phys;
+
+ if (!virt)
+ return -ENOMEM;
+
+ phys = virt_to_phys(virt);
+
+ if (state == RMI_OP_MEM_DELEGATED) {
+ phys_addr_t delegated_phys;
+
+ if (rmi_delegate_range(phys, block_size, &delegated_phys)) {
+ if (!rmi_undelegate_range(phys, delegated_phys - phys))
+ free_pages_exact(virt, block_size);
+ return -ENXIO;
+ }
+ }
+
+ addr_range = phys & RMI_ADDR_RANGE_ADDR_MASK;
+ FIELD_MODIFY(RMI_ADDR_RANGE_BLOCK_SIZE_MASK, &addr_range, block_size_fld);
+ FIELD_MODIFY(RMI_ADDR_RANGE_COUNT_MASK, &addr_range, 1);
+ FIELD_MODIFY(RMI_ADDR_RANGE_STATE_MASK, &addr_range, state);
+
+ sro->addr_list[sro->addr_count++] = addr_range;
+ }
+
+mem_donate:
+ rmi_op_mem_donate(sro_handle,
+ virt_to_phys(&sro->addr_list[addr_list_start]),
+ found, 0, out_regs);
+
+ donated_granules = out_regs->a1;
+ /*
+ * The RMM shouldn't report more granules than we provided, but clamp
+ * just in case.
+ */
+ if (WARN_ON_ONCE(donated_granules > found * granules_per_block))
+ donated_granules = found * granules_per_block;
+
+ /*
+ * The RMM reports the consumed memory in terms of granules, but we
+ * track in the address lists in block-sized ranges. So divide to get
+ * the number of (complete) consumed blocks.
+ */
+ consumed_blocks = donated_granules / granules_per_block;
+ if (donated_granules % granules_per_block) {
+ /*
+ * A block has been partially consumed, the start is owned by
+ * the RMM, the tail is owned by the host
+ */
+ unsigned long entry =
+ sro->addr_list[addr_list_start + consumed_blocks];
+ unsigned long donated_size =
+ (donated_granules % granules_per_block) << PAGE_SHIFT;
+
+ free_addr_range(entry, donated_size);
+ /*
+ * This block is now fully 'consumed' (either held by the RMM or
+ * freed)
+ */
+ consumed_blocks++;
+ }
+
+ /*
+ * Keep just the blocks the RMM didn't use in addr_list
+ * RMM claimed consumed_blocks entries from addr_list_start.
+ * Move the entries left out at the end i.e.,
+ * [ addr_list_start + consumed_blocks, addr_list_start + found)
+ * to the rest of the valid entries and adjust the addr_count to
+ * reflect the available entries.
+ */
+ for (i = 0, src = addr_list_start + consumed_blocks;
+ i < found - consumed_blocks; i++)
+ sro->addr_list[addr_list_start + i] = sro->addr_list[src + i];
+
+ sro->addr_count -= consumed_blocks;
+
+ return 0;
+}
+
+static int rmi_sro_donate(struct rmi_sro_state *sro,
+ unsigned long sro_handle,
+ unsigned long donatereq,
+ struct arm_smccc_1_2_regs *regs,
+ gfp_t gfp)
+{
+ unsigned long state = RMI_DONATE_STATE(donatereq);
+
+ if (WARN_ON_ONCE(!RMI_DONATE_COUNT(donatereq)))
+ return -EINVAL;
+
+ /*
+ * We do not support RMI_OP_MEM_CONDITIONAL yet. This is only required
+ * for use in RMI_GRANULE_TRACKING_SET, which we don't support yet.
+ */
+ if (WARN_ON_ONCE((state != RMI_OP_MEM_DELEGATED) &&
+ (state != RMI_OP_MEM_UNDELEGATED)))
+ return -EINVAL;
+
+ if (RMI_DONATE_CONTIG(donatereq) == RMI_OP_MEM_CONTIG)
+ return rmi_sro_donate_contig(sro, sro_handle, donatereq, regs, gfp);
+ else
+ return rmi_sro_donate_noncontig(sro, sro_handle, donatereq, regs, gfp);
+}
+
+static int rmi_sro_reclaim(struct rmi_sro_state *sro,
+ unsigned long sro_handle,
+ struct arm_smccc_1_2_regs *out_regs)
+{
+ unsigned long capacity;
+
+ /*
+ * We don't do a partial free of the entries. So for now free the
+ * entire address list as we prepare to reclaim more from the RMM.
+ */
+ if (rmi_sro_ensure_capacity(sro, 1))
+ rmi_sro_free(sro);
+
+ capacity = RMI_MAX_ADDR_LIST - sro->addr_count;
+
+ rmi_op_mem_reclaim(sro_handle,
+ virt_to_phys(&sro->addr_list[sro->addr_count]),
+ capacity, out_regs);
+
+ /*
+ * RMI_OP_MEM_RECLAIM always return RMI_INCOMPLETE, except when the
+ * input parameters were invalid.
+ */
+ if (WARN_ON_ONCE(RMI_RESULT_STATUS(out_regs->a0) != RMI_INCOMPLETE))
+ return -EINVAL;
+ if (WARN_ON_ONCE(out_regs->a1 > capacity))
+ out_regs->a1 = capacity;
+
+ sro->addr_count += out_regs->a1;
+
+ return 0;
+}
+
+void rmi_sro_free(struct rmi_sro_state *sro)
+{
+ /* Handle the worse */
+ if (WARN_ON(sro->addr_count < 0))
+ return;
+
+ if (WARN_ON(sro->addr_count > RMI_MAX_ADDR_LIST))
+ sro->addr_count = RMI_MAX_ADDR_LIST;
+
+ for (int i = 0; i < sro->addr_count; i++)
+ free_addr_range(sro->addr_list[i], 0);
+
+ sro->addr_count = 0;
+}
+EXPORT_SYMBOL_GPL(rmi_sro_free);
+
+long rmi_sro_memxfer_execute(struct rmi_sro_state *sro, gfp_t gfp)
+{
+ struct arm_smccc_1_2_regs *regs = &sro->regs;
+ bool cancelled = false;
+ unsigned long sro_handle;
+
+ rmi_smccc_invoke(regs);
+
+ sro_handle = regs->a1;
+ while (RMI_RESULT_STATUS(regs->a0) == RMI_INCOMPLETE) {
+ bool can_cancel = RMI_RESULT_CAN_CANCEL(regs->a0) == RMI_OP_CAN_CANCEL;
+ int ret = 0;
+
+ switch (RMI_RESULT_MEMREQ(regs->a0)) {
+ case RMI_OP_MEM_REQ_NONE:
+ rmi_op_continue(sro_handle, RMI_CONTINUE_KEEP_GOING,
+ regs);
+ break;
+ case RMI_OP_MEM_REQ_DONATE:
+ ret = rmi_sro_donate(sro, sro_handle, regs->a2, regs,
+ gfp);
+ break;
+ case RMI_OP_MEM_REQ_RECLAIM:
+ ret = rmi_sro_reclaim(sro, sro_handle, regs);
+ break;
+ default:
+ WARN_ON_ONCE(1);
+ ret = -ENXIO;
+ }
+
+ if (ret) {
+ /*
+ * All memory donating SROs must be cancellable. So a
+ * failure in memory allocation shouldn't be an issue.
+ * However, if we encounter a random failure (e.g.,
+ * buggy RMM), don't loop forever, just give up.
+ */
+ if (WARN_ON_ONCE(!can_cancel))
+ return ret;
+ /*
+ * If we have already cancelled, and came back here due
+ * to an error in MEMREQ, then there is no point
+ * in going in loops.
+ */
+ if (WARN_ON_ONCE(cancelled))
+ break;
+ rmi_op_cancel(sro_handle, regs);
+ cancelled = true;
+
+ if (WARN_ON_ONCE(RMI_RESULT_STATUS(regs->a0) != RMI_INCOMPLETE))
+ return ret;
+ }
+ }
+
+ if (cancelled)
+ return -ECANCELED;
+
+ return regs->a0;
+}
+EXPORT_SYMBOL_GPL(rmi_sro_memxfer_execute);
+
+/*
+ * rmi_sro_execute: Execute an RMI command that is stateful but not memory
+ * transferring. Takes regs, filled with the FIDs and the arguments in place.
+ *
+ * Returns :
+ * -ECANCELED - If the operation had to be aborted and SRO was cancellable.
+ * Otherwise, returns the result of the RMI command.
+ */
+long rmi_sro_execute(struct arm_smccc_1_2_regs *regs)
+{
+ bool cancelled = false;
+ unsigned long sro_handle;
+
+ rmi_smccc_invoke(regs);
+
+ sro_handle = regs->a1;
+ while (RMI_RESULT_STATUS(regs->a0) == RMI_INCOMPLETE) {
+ bool can_cancel = RMI_RESULT_CAN_CANCEL(regs->a0) == RMI_OP_CAN_CANCEL;
+
+ switch (RMI_RESULT_MEMREQ(regs->a0)) {
+ case RMI_OP_MEM_REQ_NONE:
+ rmi_op_continue(sro_handle, RMI_CONTINUE_KEEP_GOING,
+ regs);
+ break;
+ default:
+ WARN_ON_ONCE(1);
+ if (!can_cancel)
+ return regs->a0;
+ /*
+ * We can't get here normally, but handle this anyway
+ * for a buggy RMM implementation.
+ */
+ if (cancelled)
+ return -ECANCELED;
+ rmi_op_cancel(sro_handle, regs);
+ cancelled = true;
+ }
+ }
+
+ if (cancelled)
+ return -ECANCELED;
+
+ return regs->a0;
+}
+EXPORT_SYMBOL_GPL(rmi_sro_execute);
+
static int rmi_check_version(void)
{
unsigned short version_major, version_minor;
diff --git a/include/linux/arm-rmi-cmds.h b/include/linux/arm-rmi-cmds.h
index 911489b636522..5354839d76077 100644
--- a/include/linux/arm-rmi-cmds.h
+++ b/include/linux/arm-rmi-cmds.h
@@ -8,9 +8,19 @@
#include <linux/arm-smccc-rmi.h>
#include <linux/bug.h>
+#include <linux/gfp.h>
#include <linux/processor.h>
+#include <linux/string.h>
#include <linux/types.h>
+#define RMI_MAX_ADDR_LIST 256
+
+struct rmi_sro_state {
+ struct arm_smccc_1_2_regs regs;
+ int addr_count;
+ unsigned long addr_list[RMI_MAX_ADDR_LIST];
+};
+
#define RMM_BLOCKED_RETRY_COUNT 2
/*
* rmi_smccc_invoke: Invoke the RMI call and return the results, retrying the
@@ -45,4 +55,35 @@ static inline void rmi_smccc_invoke(struct arm_smccc_1_2_regs *regs)
unsigned long rmi_feat_reg(unsigned int index);
+int rmi_delegate_range(phys_addr_t phys, unsigned long size,
+ phys_addr_t *out_phys);
+int rmi_undelegate_range(phys_addr_t phys, unsigned long size);
+int rmi_free_delegated_page(phys_addr_t phys);
+
+static inline int rmi_delegate_page(phys_addr_t phys)
+{
+ return rmi_delegate_range(phys, PAGE_SIZE, NULL);
+}
+
+static inline int rmi_undelegate_page(phys_addr_t phys)
+{
+ return rmi_undelegate_range(phys, PAGE_SIZE);
+}
+
+long rmi_sro_memxfer_execute(struct rmi_sro_state *sro, gfp_t gfp);
+void rmi_sro_free(struct rmi_sro_state *sro);
+long rmi_sro_execute(struct arm_smccc_1_2_regs *regs);
+
+/*
+ * Resetting the addr_count is sufficient to ignore the addr_list contents.
+ */
+#define rmi_sro_memxfer_cmd(sro, gfp, ...) ({ \
+ struct rmi_sro_state *__sro = (sro); \
+ __sro->addr_count = 0; \
+ __sro->regs = (struct arm_smccc_1_2_regs){ __VA_ARGS__ }; \
+ long __ret = rmi_sro_memxfer_execute(__sro, gfp); \
+ rmi_sro_free(__sro); \
+ __ret; \
+})
+
#endif /* __LINUX_ARM_RMI_CMDS_H_ */
--
2.43.0
^ permalink raw reply [flat|nested] 23+ messages in thread* Re: [PATCH v20 4/9] firmware: arm_rmm: Add support for SRO
2026-09-29 22:16 ` [PATCH v20 4/9] firmware: arm_rmm: Add support for SRO Suzuki K Poulose
@ 2026-09-30 13:15 ` Catalin Marinas
2026-09-30 13:48 ` Suzuki K Poulose
0 siblings, 1 reply; 23+ messages in thread
From: Catalin Marinas @ 2026-09-30 13:15 UTC (permalink / raw)
To: Suzuki K Poulose
Cc: kvm, kvmarm, maz, will, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, sudeep.holla,
jonathan.cameron
On Tue, Sep 29, 2026 at 11:16:18PM +0100, Suzuki K Poulose wrote:
> +static int rmi_sro_donate_contig(struct rmi_sro_state *sro,
> + unsigned long sro_handle,
> + unsigned long donatereq,
> + struct arm_smccc_1_2_regs *out_regs,
> + gfp_t gfp)
> +{
> + unsigned long block_size_fld = RMI_DONATE_BLOCK_SIZE(donatereq);
> + unsigned long block_size = rmi_addr_block_size_to_bytes(block_size_fld);
> + unsigned long count = RMI_DONATE_COUNT(donatereq);
> + unsigned long state = RMI_DONATE_STATE(donatereq);
> + unsigned long size = block_size * count;
> + unsigned long addr_range;
> + unsigned long donated_size;
> + int ret;
> + void *virt;
> + phys_addr_t phys;
> +
> + /*
> + * The RMM specification requires contiguous allocations are always a
> + * power of 2
> + */
> + if (WARN_ON_ONCE(!is_power_of_2(size)))
> + return -EINVAL;
> + /*
> + * RMM clamps the Maximum value of RmiOpMemDonateReq:count to prevent
> + * overflow in the RMI_ADDR_RANGE_COUNT field.
> + */
> + if (WARN_ON_ONCE(count > (BIT(PAGE_SHIFT - 2) - 1)))
> + return -EINVAL;
Nit: it might be easier to read as FIELD_MAX(RMI_ADDR_RANGE_COUNT_MASK)
as that's what we want to limit it to.
Reviewed-by: Catalin Marinas <catalin.marinas@arm.com>
(Sashiko seems to have more findings but the rest looks alright to me)
^ permalink raw reply [flat|nested] 23+ messages in thread* Re: [PATCH v20 4/9] firmware: arm_rmm: Add support for SRO
2026-09-30 13:15 ` Catalin Marinas
@ 2026-09-30 13:48 ` Suzuki K Poulose
0 siblings, 0 replies; 23+ messages in thread
From: Suzuki K Poulose @ 2026-09-30 13:48 UTC (permalink / raw)
To: Catalin Marinas
Cc: kvm, kvmarm, maz, will, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, sudeep.holla,
jonathan.cameron
On 30/09/2026 14:15, Catalin Marinas wrote:
> On Tue, Sep 29, 2026 at 11:16:18PM +0100, Suzuki K Poulose wrote:
>> +static int rmi_sro_donate_contig(struct rmi_sro_state *sro,
>> + unsigned long sro_handle,
>> + unsigned long donatereq,
>> + struct arm_smccc_1_2_regs *out_regs,
>> + gfp_t gfp)
>> +{
>> + unsigned long block_size_fld = RMI_DONATE_BLOCK_SIZE(donatereq);
>> + unsigned long block_size = rmi_addr_block_size_to_bytes(block_size_fld);
>> + unsigned long count = RMI_DONATE_COUNT(donatereq);
>> + unsigned long state = RMI_DONATE_STATE(donatereq);
>> + unsigned long size = block_size * count;
>> + unsigned long addr_range;
>> + unsigned long donated_size;
>> + int ret;
>> + void *virt;
>> + phys_addr_t phys;
>> +
>> + /*
>> + * The RMM specification requires contiguous allocations are always a
>> + * power of 2
>> + */
>> + if (WARN_ON_ONCE(!is_power_of_2(size)))
>> + return -EINVAL;
>> + /*
>> + * RMM clamps the Maximum value of RmiOpMemDonateReq:count to prevent
>> + * overflow in the RMI_ADDR_RANGE_COUNT field.
>> + */
>> + if (WARN_ON_ONCE(count > (BIT(PAGE_SHIFT - 2) - 1)))
>> + return -EINVAL;
>
> Nit: it might be easier to read as FIELD_MAX(RMI_ADDR_RANGE_COUNT_MASK)
> as that's what we want to limit it to.
Thanks, that is neat. I will update this.
>
> Reviewed-by: Catalin Marinas <catalin.marinas@arm.com>
>
> (Sashiko seems to have more findings but the rest looks alright to me)
I have updated the code to address the issues.
Thanks Catalin.
Suzuki
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v20 5/9] firmware: arm_rmm: Activate the RMM
2026-09-29 22:16 [PATCH v20 0/9] firmware: arm_rmm: Add RMM v2.0 base RMI support Suzuki K Poulose
` (3 preceding siblings ...)
2026-09-29 22:16 ` [PATCH v20 4/9] firmware: arm_rmm: Add support for SRO Suzuki K Poulose
@ 2026-09-29 22:16 ` Suzuki K Poulose
2026-09-30 13:20 ` Catalin Marinas
2026-09-29 22:16 ` [PATCH v20 6/9] firmware: arm_rmm: Ensure the RMM has GPT entries for memory Suzuki K Poulose
` (3 subsequent siblings)
8 siblings, 1 reply; 23+ messages in thread
From: Suzuki K Poulose @ 2026-09-29 22:16 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, sudeep.holla,
jonathan.cameron, Suzuki K Poulose
From: Steven Price <steven.price@arm.com>
Activate the RMM after the basic configuration. This is a memory
transferring stateful operation.
Reviewed-by: Gavin Shan <gshan@redhat.com>
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Tested-by: Gavin Shan <gshan@redhat.com>
Signed-off-by: Steven Price <steven.price@arm.com>
Co-developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Changes since v17:
* Inline RMM_ACTIVATE command and remove the definitions from arm-rmi-cmds.h
* Use scope-based cleanup to free sro object
Changes since v16:
* Split into a new patch
---
drivers/firmware/arm_rmm/rmi.c | 13 ++++++++++++-
1 file changed, 12 insertions(+), 1 deletion(-)
diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/arm_rmm/rmi.c
index 8d839e9e61424..d982cf158f8da 100644
--- a/drivers/firmware/arm_rmm/rmi.c
+++ b/drivers/firmware/arm_rmm/rmi.c
@@ -850,7 +850,18 @@ static int __init arm64_init_rmi(void)
if (ret)
return ret;
- return 0;
+ /* Activate the RMM */
+ struct rmi_sro_state *sro __free(kfree) = kmalloc_obj(*sro);
+ if (!sro)
+ return -ENOMEM;
+
+ ret = rmi_sro_memxfer_cmd(sro, GFP_KERNEL, SMC_RMI_RMM_ACTIVATE);
+ if (ret) {
+ pr_err("RMM activate failed (%d)\n", ret);
+ ret = ret < 0 ? ret : -ENXIO;
+ }
+
+ return ret;
}
/*
--
2.43.0
^ permalink raw reply [flat|nested] 23+ messages in thread* Re: [PATCH v20 5/9] firmware: arm_rmm: Activate the RMM
2026-09-29 22:16 ` [PATCH v20 5/9] firmware: arm_rmm: Activate the RMM Suzuki K Poulose
@ 2026-09-30 13:20 ` Catalin Marinas
0 siblings, 0 replies; 23+ messages in thread
From: Catalin Marinas @ 2026-09-30 13:20 UTC (permalink / raw)
To: Suzuki K Poulose
Cc: kvm, kvmarm, maz, will, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, sudeep.holla,
jonathan.cameron
On Tue, Sep 29, 2026 at 11:16:19PM +0100, Suzuki K Poulose wrote:
> From: Steven Price <steven.price@arm.com>
>
> Activate the RMM after the basic configuration. This is a memory
> transferring stateful operation.
>
> Reviewed-by: Gavin Shan <gshan@redhat.com>
> Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
> Tested-by: Gavin Shan <gshan@redhat.com>
> Signed-off-by: Steven Price <steven.price@arm.com>
> Co-developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
Reviewed-by: Catalin Marinas <catalin.marinas@arm.com>
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v20 6/9] firmware: arm_rmm: Ensure the RMM has GPT entries for memory
2026-09-29 22:16 [PATCH v20 0/9] firmware: arm_rmm: Add RMM v2.0 base RMI support Suzuki K Poulose
` (4 preceding siblings ...)
2026-09-29 22:16 ` [PATCH v20 5/9] firmware: arm_rmm: Activate the RMM Suzuki K Poulose
@ 2026-09-29 22:16 ` Suzuki K Poulose
2026-09-30 13:39 ` Catalin Marinas
2026-09-30 14:44 ` Sudeep Holla
2026-09-29 22:16 ` [PATCH v20 7/9] arm64: Block hibernate and kexec while RMM is active Suzuki K Poulose
` (2 subsequent siblings)
8 siblings, 2 replies; 23+ messages in thread
From: Suzuki K Poulose @ 2026-09-29 22:16 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, sudeep.holla,
jonathan.cameron, Suzuki K Poulose
From: Steven Price <steven.price@arm.com>
The RMM maintains the state of all the granules in the system to make
sure that the host is abiding by the rules. This state can be maintained
at different granularity, per page (TRACKING_FINE) or per region
(TRACKING_COARSE or TRACKING_INTERMEDIATE). The region size depends on the
underlying "RMI_GRANULE_SIZE". For a "coarse"/"intermediate" region,
all pages in the region must be of the same state, this implies we need to
have "fine" tracking for DRAM, so that we can delegate individual pages.
For now we only support a statically carved out memory for tracking
granules for the "fine" regions. This can be extended in the future to
allow modifying the tracking granularity and remove the need for a
static allocation by the firmware.
Similarly, the firmware may create L0 GPT entries describing the total
address space. But if we change the "PAS" (Physical Address Space) of a
granule, then the firmware may need to create L1 tables to track the PAS
at a finer granularity. Linux therefore checks if the platform firmware
manages the PAR region. i.e., the firmware is in charge of managing the
L1 GPTs (creation and the required memory for the GPT tables - via static
carveouts) without host intervention. Support for dynamic GPT creation by
the host will be added later.
If the firmware requires us to manage the tracking or GPT memory,
deactivate the RMM and reclaim any memory donated at RMM activation.
Apply the same checks when hotplugged memory is brought online.
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Reviewed-by: Gavin Shan <gshan@redhat.com>
Tested-by: Gavin Shan <gshan@redhat.com>
Signed-off-by: Steven Price <steven.price@arm.com>
Co-developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Changes since v19:
* Wrap comments to 80spaces.
* Add a comment on ALIGN_DOWN the end of memory range.
* Add is_rmm_active() to indicate if the RMM is active
* Switch errors to -ENODEV from -ENOMEM when RMM doesn't manage the region
* Print error messages when RMI_GPT_INFO/RMI_GRANULE_TRACKIN_GET fails
Changes since v18:
* Avoid mixing gotos with __free cleanups for arm64_init_rmi()
* Handle buggy RMM to make forward progress for RMI_GPT_INFO and
RMI_GRANULE_TRACKING_GET
* Make sure the memory ranges are not inverted and reject such ranges.
* Move granule_tracking_get/gpt_info wrappers closer to the callers
Drop "inline", let the compiler do its job
Changes since v17:
* Move wrappers that may not be used elsewhere, out of arm-rmi-cmds.h
Changes since v16:
* Check fine tracking and create L1 GPTs for hotplug-added memory.
* Clarify the L1 GPT setup and move the explanatory comment.
* Switch to using RMI_GPT_INFO command for checking the GPTs.
* Deactivate the RMM and reclaim the memory if we can't proceed.
Changes since v15:
* Skip firmware-reserved NOMAP memory in rmi_init_metadata()
* Handle negative error codes from wrappers.
Changes since v14:
* Move the implementation into drivers/firmware/arm_rmm.
Changes since v13:
* Moved out of KVM
---
drivers/firmware/arm_rmm/rmi.c | 247 ++++++++++++++++++++++++++++++++-
include/linux/arm-rmi-cmds.h | 18 +++
2 files changed, 264 insertions(+), 1 deletion(-)
diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/arm_rmm/rmi.c
index d982cf158f8da..d4d098448e46e 100644
--- a/drivers/firmware/arm_rmm/rmi.c
+++ b/drivers/firmware/arm_rmm/rmi.c
@@ -6,6 +6,7 @@
#include <linux/arm-rmi-cmds.h>
#include <linux/cpufeature.h>
#include <linux/memblock.h>
+#include <linux/memory.h>
#include <linux/mmzone.h>
#include <linux/slab.h>
@@ -14,6 +15,25 @@
/* RMM v2.0 defines RmiFeatureRegister0 to RmiFeatureRegister4. */
static unsigned long rmi_feat_reg_cache[5] __ro_after_init;
+static bool arm64_rmm_active;
+static bool arm64_rmi_is_available;
+
+/* is_rmm_active: Returns if the RMM is ACTIVE */
+bool is_rmm_active(void)
+{
+ return READ_ONCE(arm64_rmm_active);
+}
+EXPORT_SYMBOL_GPL(is_rmm_active);
+
+/*
+ * is_rmi_available: Returns if the RMM is ACTIVE and is configured with static
+ * carveout for metadata nad L1GPTs.
+ */
+bool is_rmi_available(void)
+{
+ return READ_ONCE(arm64_rmi_is_available);
+}
+EXPORT_SYMBOL_GPL(is_rmi_available);
/*
* rmi_granule_range_delegate() - Delegate granules
@@ -833,6 +853,216 @@ static int rmi_configure(void)
return 0;
}
+/**
+ * rmi_granule_tracking_get() - Get configuration of a Granule tracking region
+ * @start: Base PA of the tracking region
+ * @end: End of the PA region
+ * @out_category: Memory category
+ * @out_state: Tracking region state
+ * @out_top: Top of the memory region
+ *
+ * Return: RMI return code
+ */
+static int rmi_granule_tracking_get(unsigned long start,
+ unsigned long end,
+ unsigned long *out_category,
+ unsigned long *out_state,
+ unsigned long *out_top)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_GRANULE_TRACKING_GET, start, end,
+ };
+
+ rmi_smccc_invoke(®s);
+
+ if (regs.a0 != RMI_SUCCESS)
+ return regs.a0;
+
+ if (out_category)
+ *out_category = regs.a1;
+ if (out_state)
+ *out_state = regs.a2;
+ if (out_top)
+ *out_top = regs.a3;
+
+ return RMI_SUCCESS;
+}
+
+/*
+ * Make sure the area is tracked by RMM at FINE granularity.
+ * We do not support changing the tracking yet.
+ */
+static int rmi_verify_memory_tracking(phys_addr_t start, phys_addr_t end)
+{
+ while (start < end) {
+ unsigned long ret, category, state, next;
+
+ ret = rmi_granule_tracking_get(start, end, &category, &state, &next);
+ if (ret != RMI_SUCCESS) {
+ pr_err("RMI_GRANULE_TRACKING_GET failed for %llx-%llx: (%lu)\n",
+ start, end, ret);
+ return -ENXIO;
+ }
+
+ if (WARN_ON(next <= start))
+ return -ENXIO;
+
+ if (state != RMI_TRACKING_FINE ||
+ category != RMI_MEM_CATEGORY_CONVENTIONAL) {
+ /* TODO: Set granule tracking in this case */
+ pr_err("Granule tracking for region isn't fine/conventional: %llx-%lx\n",
+ start, next);
+ return -ENODEV;
+ }
+ start = next;
+ }
+
+ return 0;
+}
+
+/*
+ * rmi_gpt_info - Query the GPT info for the given PAR.
+ * @start: Base of the physical address region
+ * @end: Top of the physical address region
+ * @out_top: Top of the physical address region for which the GPT @out_gpt_par_state
+ * is valid
+ * @out_gpt_par_state: State of the GPT covered by [start, out_top)
+ */
+static long rmi_gpt_info(unsigned long start, unsigned long end,
+ unsigned long *out_top,
+ unsigned long *out_gpt_par_state)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_GPT_INFO, start, end,
+ };
+
+ rmi_smccc_invoke(®s);
+ if (regs.a0 != RMI_SUCCESS)
+ return regs.a0;
+
+ if (out_top)
+ *out_top = regs.a1;
+ if (out_gpt_par_state)
+ *out_gpt_par_state = regs.a2;
+
+ return RMI_SUCCESS;
+}
+
+/*
+ * We do not support creating L1 GPTs yet. So, make sure that all the regions
+ * are managed by the firmware.
+ */
+static int rmi_verify_gpt_firmware_managed(phys_addr_t start, phys_addr_t end)
+{
+ unsigned long l0gpt_sz;
+ unsigned long next, par_state;
+
+ l0gpt_sz = 1UL << (30 + FIELD_GET(RMI_FEATURE_REGISTER_1_L0GPTSZ,
+ rmi_feat_reg(1)));
+ start = ALIGN_DOWN(start, l0gpt_sz);
+ end = ALIGN(end, l0gpt_sz);
+
+ while (start < end) {
+ long ret = rmi_gpt_info(start, end, &next, &par_state);
+
+ if (ret != RMI_SUCCESS) {
+ pr_err("RMI_GPT_INFO failed for %llx-%llx: (%ld)\n",
+ start, end, ret);
+ return -ENXIO;
+ }
+
+ if (WARN_ON(next <= start))
+ return -ENXIO;
+
+ if (par_state != RMI_GPT_PAR_PLAT) {
+ pr_err("GPT for the region is not managed by firmware %llx-%lx\n",
+ start, next);
+ return -ENODEV;
+ }
+ start = next;
+ }
+
+ return 0;
+}
+
+static int rmi_prepare_memory(phys_addr_t start, phys_addr_t end)
+{
+ int ret;
+
+ if (start >= end)
+ return -EINVAL;
+
+ ret = rmi_verify_memory_tracking(start, end);
+ if (ret)
+ return ret;
+
+ return rmi_verify_gpt_firmware_managed(start, end);
+}
+
+static int rmi_init_metadata(void)
+{
+ phys_addr_t start, end;
+ struct memblock_region *r;
+
+ for_each_mem_region(r) {
+ int ret;
+
+ /* Firmware-reserved NOMAP regions are not usable system RAM */
+ if (memblock_is_nomap(r))
+ continue;
+
+ start = PAGE_ALIGN(r->base);
+ /*
+ * We always deal with PAGE_SIZE and if in the odd case this
+ * region boundary is not PAGE aligned, we stick to the page
+ * that we can use from the region.
+ */
+ end = PAGE_ALIGN_DOWN(r->base + r->size);
+ /* Too small ? */
+ if (start >= end)
+ continue;
+
+ ret = rmi_prepare_memory(start, end);
+ if (ret)
+ return ret;
+ }
+
+ return 0;
+}
+
+static int rmi_memory_notifier(struct notifier_block *nb,
+ unsigned long action, void *data)
+{
+ struct memory_notify *arg = data;
+ phys_addr_t start, end;
+ int ret;
+
+ if (action != MEM_GOING_ONLINE)
+ return NOTIFY_DONE;
+
+ start = PFN_PHYS(arg->start_pfn);
+ end = PFN_PHYS(arg->start_pfn + arg->nr_pages);
+
+ ret = rmi_prepare_memory(start, end);
+
+ return notifier_from_errno(ret);
+}
+
+static struct notifier_block rmi_memory_nb = {
+ .notifier_call = rmi_memory_notifier,
+};
+
+static int rmi_init_memory(void)
+{
+ int ret;
+
+ ret = rmi_init_metadata();
+ if (ret)
+ return ret;
+
+ return register_memory_notifier(&rmi_memory_nb);
+}
+
static int __init arm64_init_rmi(void)
{
int ret;
@@ -859,9 +1089,24 @@ static int __init arm64_init_rmi(void)
if (ret) {
pr_err("RMM activate failed (%d)\n", ret);
ret = ret < 0 ? ret : -ENXIO;
+ return ret;
}
- return ret;
+ WRITE_ONCE(arm64_rmm_active, true);
+
+ ret = rmi_init_memory();
+ if (ret) {
+ /* Deactivate the RMM */
+ if (!WARN_ON(rmi_sro_memxfer_cmd(sro, GFP_KERNEL, SMC_RMI_RMM_DEACTIVATE)))
+ WRITE_ONCE(arm64_rmm_active, false);
+ return ret;
+ }
+
+ /* Mark the RMM is usable */
+ WRITE_ONCE(arm64_rmi_is_available, true);
+ pr_info("RMI configured\n");
+
+ return 0;
}
/*
diff --git a/include/linux/arm-rmi-cmds.h b/include/linux/arm-rmi-cmds.h
index 5354839d76077..5d005054f3c6c 100644
--- a/include/linux/arm-rmi-cmds.h
+++ b/include/linux/arm-rmi-cmds.h
@@ -86,4 +86,22 @@ long rmi_sro_execute(struct arm_smccc_1_2_regs *regs);
__ret; \
})
+#ifdef CONFIG_ARM_RMM_RMI
+
+bool is_rmm_active(void);
+bool is_rmi_available(void);
+
+#else
+
+static inline bool is_rmm_active(void)
+{
+ return false;
+}
+
+static inline bool is_rmi_available(void)
+{
+ return false;
+}
+#endif /* CONFIG_ARM_RMM_RMI */
+
#endif /* __LINUX_ARM_RMI_CMDS_H_ */
--
2.43.0
^ permalink raw reply [flat|nested] 23+ messages in thread* Re: [PATCH v20 6/9] firmware: arm_rmm: Ensure the RMM has GPT entries for memory
2026-09-29 22:16 ` [PATCH v20 6/9] firmware: arm_rmm: Ensure the RMM has GPT entries for memory Suzuki K Poulose
@ 2026-09-30 13:39 ` Catalin Marinas
2026-09-30 14:44 ` Sudeep Holla
1 sibling, 0 replies; 23+ messages in thread
From: Catalin Marinas @ 2026-09-30 13:39 UTC (permalink / raw)
To: Suzuki K Poulose
Cc: kvm, kvmarm, maz, will, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, sudeep.holla,
jonathan.cameron
On Tue, Sep 29, 2026 at 11:16:20PM +0100, Suzuki K Poulose wrote:
> From: Steven Price <steven.price@arm.com>
>
> The RMM maintains the state of all the granules in the system to make
> sure that the host is abiding by the rules. This state can be maintained
> at different granularity, per page (TRACKING_FINE) or per region
> (TRACKING_COARSE or TRACKING_INTERMEDIATE). The region size depends on the
> underlying "RMI_GRANULE_SIZE". For a "coarse"/"intermediate" region,
> all pages in the region must be of the same state, this implies we need to
> have "fine" tracking for DRAM, so that we can delegate individual pages.
>
> For now we only support a statically carved out memory for tracking
> granules for the "fine" regions. This can be extended in the future to
> allow modifying the tracking granularity and remove the need for a
> static allocation by the firmware.
>
> Similarly, the firmware may create L0 GPT entries describing the total
> address space. But if we change the "PAS" (Physical Address Space) of a
> granule, then the firmware may need to create L1 tables to track the PAS
> at a finer granularity. Linux therefore checks if the platform firmware
> manages the PAR region. i.e., the firmware is in charge of managing the
> L1 GPTs (creation and the required memory for the GPT tables - via static
> carveouts) without host intervention. Support for dynamic GPT creation by
> the host will be added later.
>
> If the firmware requires us to manage the tracking or GPT memory,
> deactivate the RMM and reclaim any memory donated at RMM activation.
>
> Apply the same checks when hotplugged memory is brought online.
>
> Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
> Reviewed-by: Gavin Shan <gshan@redhat.com>
> Tested-by: Gavin Shan <gshan@redhat.com>
> Signed-off-by: Steven Price <steven.price@arm.com>
> Co-developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
Reviewed-by: Catalin Marinas <catalin.marinas@arm.com>
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH v20 6/9] firmware: arm_rmm: Ensure the RMM has GPT entries for memory
2026-09-29 22:16 ` [PATCH v20 6/9] firmware: arm_rmm: Ensure the RMM has GPT entries for memory Suzuki K Poulose
2026-09-30 13:39 ` Catalin Marinas
@ 2026-09-30 14:44 ` Sudeep Holla
2026-09-30 15:55 ` Suzuki K Poulose
1 sibling, 1 reply; 23+ messages in thread
From: Sudeep Holla @ 2026-09-30 14:44 UTC (permalink / raw)
To: Suzuki K Poulose
Cc: kvm, kvmarm, maz, Sudeep Holla, will, catalin.marinas,
linux-kernel, linux-arm-kernel, steven.price, aneesh.kumar,
oupton, gshan, joey.gouly, tabba, yuzenghui, linux-coco,
gankulkarni, sdonthineni, alpergun, fj0570is, WeiLin.Chang,
lpieralisi, enju.kohei, jonathan.cameron
On Tue, Sep 29, 2026 at 11:16:20PM +0100, Suzuki K Poulose wrote:
> From: Steven Price <steven.price@arm.com>
>
> The RMM maintains the state of all the granules in the system to make
> sure that the host is abiding by the rules. This state can be maintained
> at different granularity, per page (TRACKING_FINE) or per region
> (TRACKING_COARSE or TRACKING_INTERMEDIATE). The region size depends on the
> underlying "RMI_GRANULE_SIZE". For a "coarse"/"intermediate" region,
> all pages in the region must be of the same state, this implies we need to
> have "fine" tracking for DRAM, so that we can delegate individual pages.
>
> For now we only support a statically carved out memory for tracking
> granules for the "fine" regions. This can be extended in the future to
> allow modifying the tracking granularity and remove the need for a
> static allocation by the firmware.
>
> Similarly, the firmware may create L0 GPT entries describing the total
> address space. But if we change the "PAS" (Physical Address Space) of a
> granule, then the firmware may need to create L1 tables to track the PAS
> at a finer granularity. Linux therefore checks if the platform firmware
> manages the PAR region. i.e., the firmware is in charge of managing the
> L1 GPTs (creation and the required memory for the GPT tables - via static
> carveouts) without host intervention. Support for dynamic GPT creation by
> the host will be added later.
>
> If the firmware requires us to manage the tracking or GPT memory,
> deactivate the RMM and reclaim any memory donated at RMM activation.
>
> Apply the same checks when hotplugged memory is brought online.
>
> Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
> Reviewed-by: Gavin Shan <gshan@redhat.com>
> Tested-by: Gavin Shan <gshan@redhat.com>
> Signed-off-by: Steven Price <steven.price@arm.com>
> Co-developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
> ---
> Changes since v19:
> * Wrap comments to 80spaces.
> * Add a comment on ALIGN_DOWN the end of memory range.
> * Add is_rmm_active() to indicate if the RMM is active
> * Switch errors to -ENODEV from -ENOMEM when RMM doesn't manage the region
> * Print error messages when RMI_GPT_INFO/RMI_GRANULE_TRACKIN_GET fails
> Changes since v18:
> * Avoid mixing gotos with __free cleanups for arm64_init_rmi()
> * Handle buggy RMM to make forward progress for RMI_GPT_INFO and
> RMI_GRANULE_TRACKING_GET
> * Make sure the memory ranges are not inverted and reject such ranges.
> * Move granule_tracking_get/gpt_info wrappers closer to the callers
> Drop "inline", let the compiler do its job
> Changes since v17:
> * Move wrappers that may not be used elsewhere, out of arm-rmi-cmds.h
> Changes since v16:
> * Check fine tracking and create L1 GPTs for hotplug-added memory.
> * Clarify the L1 GPT setup and move the explanatory comment.
> * Switch to using RMI_GPT_INFO command for checking the GPTs.
> * Deactivate the RMM and reclaim the memory if we can't proceed.
> Changes since v15:
> * Skip firmware-reserved NOMAP memory in rmi_init_metadata()
> * Handle negative error codes from wrappers.
> Changes since v14:
> * Move the implementation into drivers/firmware/arm_rmm.
> Changes since v13:
> * Moved out of KVM
> ---
> drivers/firmware/arm_rmm/rmi.c | 247 ++++++++++++++++++++++++++++++++-
> include/linux/arm-rmi-cmds.h | 18 +++
> 2 files changed, 264 insertions(+), 1 deletion(-)
>
> diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/arm_rmm/rmi.c
> index d982cf158f8da..d4d098448e46e 100644
> --- a/drivers/firmware/arm_rmm/rmi.c
> +++ b/drivers/firmware/arm_rmm/rmi.c
[...]
> +static int rmi_verify_gpt_firmware_managed(phys_addr_t start, phys_addr_t end)
> +{
> + unsigned long l0gpt_sz;
> + unsigned long next, par_state;
> +
> + l0gpt_sz = 1UL << (30 + FIELD_GET(RMI_FEATURE_REGISTER_1_L0GPTSZ,
> + rmi_feat_reg(1)));
> + start = ALIGN_DOWN(start, l0gpt_sz);
> + end = ALIGN(end, l0gpt_sz);
> +
Is it guaranteed that end with not be = PASZ ? The spec says RMI_GPT_INFO
will return RMI_GPT_INFO when top >= pasz. Unlike start, end for good
reason is aligned up(ceil) instead of aligned down, but that also means
end will become PASZ when we reach end of memory, no ?
Sorry if I am missing something, trying to understand bits and piece of
RME still, just getting started.
> + while (start < end) {
> + long ret = rmi_gpt_info(start, end, &next, &par_state);
> +
> + if (ret != RMI_SUCCESS) {
> + pr_err("RMI_GPT_INFO failed for %llx-%llx: (%ld)\n",
> + start, end, ret);
> + return -ENXIO;
> + }
> +
> + if (WARN_ON(next <= start))
> + return -ENXIO;
> +
> + if (par_state != RMI_GPT_PAR_PLAT) {
> + pr_err("GPT for the region is not managed by firmware %llx-%lx\n",
> + start, next);
> + return -ENODEV;
> + }
> + start = next;
> + }
> +
> + return 0;
> +}
> +
--
Regards,
Sudeep
^ permalink raw reply [flat|nested] 23+ messages in thread* Re: [PATCH v20 6/9] firmware: arm_rmm: Ensure the RMM has GPT entries for memory
2026-09-30 14:44 ` Sudeep Holla
@ 2026-09-30 15:55 ` Suzuki K Poulose
0 siblings, 0 replies; 23+ messages in thread
From: Suzuki K Poulose @ 2026-09-30 15:55 UTC (permalink / raw)
To: Sudeep Holla
Cc: kvm, kvmarm, maz, will, catalin.marinas, linux-kernel,
linux-arm-kernel, steven.price, aneesh.kumar, oupton, gshan,
joey.gouly, tabba, yuzenghui, linux-coco, gankulkarni,
sdonthineni, alpergun, fj0570is, WeiLin.Chang, lpieralisi,
enju.kohei, jonathan.cameron
Hi Sudeep
On 30/09/2026 15:44, Sudeep Holla wrote:
> On Tue, Sep 29, 2026 at 11:16:20PM +0100, Suzuki K Poulose wrote:
>> From: Steven Price <steven.price@arm.com>
>>
>> The RMM maintains the state of all the granules in the system to make
>> sure that the host is abiding by the rules. This state can be maintained
>> at different granularity, per page (TRACKING_FINE) or per region
>> (TRACKING_COARSE or TRACKING_INTERMEDIATE). The region size depends on the
>> underlying "RMI_GRANULE_SIZE". For a "coarse"/"intermediate" region,
>> all pages in the region must be of the same state, this implies we need to
>> have "fine" tracking for DRAM, so that we can delegate individual pages.
>>
>> For now we only support a statically carved out memory for tracking
>> granules for the "fine" regions. This can be extended in the future to
>> allow modifying the tracking granularity and remove the need for a
>> static allocation by the firmware.
>>
>> Similarly, the firmware may create L0 GPT entries describing the total
>> address space. But if we change the "PAS" (Physical Address Space) of a
>> granule, then the firmware may need to create L1 tables to track the PAS
>> at a finer granularity. Linux therefore checks if the platform firmware
>> manages the PAR region. i.e., the firmware is in charge of managing the
>> L1 GPTs (creation and the required memory for the GPT tables - via static
>> carveouts) without host intervention. Support for dynamic GPT creation by
>> the host will be added later.
>>
>> If the firmware requires us to manage the tracking or GPT memory,
>> deactivate the RMM and reclaim any memory donated at RMM activation.
>>
>> Apply the same checks when hotplugged memory is brought online.
>>
...
>> ---
>> drivers/firmware/arm_rmm/rmi.c | 247 ++++++++++++++++++++++++++++++++-
>> include/linux/arm-rmi-cmds.h | 18 +++
>> 2 files changed, 264 insertions(+), 1 deletion(-)
>>
>> diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/arm_rmm/rmi.c
>> index d982cf158f8da..d4d098448e46e 100644
>> --- a/drivers/firmware/arm_rmm/rmi.c
>> +++ b/drivers/firmware/arm_rmm/rmi.c
>
> [...]
>
>> +static int rmi_verify_gpt_firmware_managed(phys_addr_t start, phys_addr_t end)
>> +{
>> + unsigned long l0gpt_sz;
>> + unsigned long next, par_state;
>> +
>> + l0gpt_sz = 1UL << (30 + FIELD_GET(RMI_FEATURE_REGISTER_1_L0GPTSZ,
>> + rmi_feat_reg(1)));
>> + start = ALIGN_DOWN(start, l0gpt_sz);
>> + end = ALIGN(end, l0gpt_sz);
>> +
>
> Is it guaranteed that end with not be = PASZ ? The spec says RMI_GPT_INFO
> will return RMI_GPT_INFO when top >= pasz. Unlike start, end for good
> reason is aligned up(ceil) instead of aligned down, but that also means
> end will become PASZ when we reach end of memory, no ?
Good point. The RMM spec needs to be fixed to make the condition
top > pasz
as the top, is really the top the region and is not inclusive.
This is already a known issue, tracked by FENIMORE-1807. The proposed
fix is :
pre: UInt(top) > rmm.static.pasz
>
> Sorry if I am missing something, trying to understand bits and piece of
> RME still, just getting started.
Thanks for looking !
Cheers
Suzuki
>
>> + while (start < end) {
>> + long ret = rmi_gpt_info(start, end, &next, &par_state);
>> +
>> + if (ret != RMI_SUCCESS) {
>> + pr_err("RMI_GPT_INFO failed for %llx-%llx: (%ld)\n",
>> + start, end, ret);
>> + return -ENXIO;
>> + }
>> +
>> + if (WARN_ON(next <= start))
>> + return -ENXIO;
>> +
>> + if (par_state != RMI_GPT_PAR_PLAT) {
>> + pr_err("GPT for the region is not managed by firmware %llx-%lx\n",
>> + start, next);
>> + return -ENODEV;
>> + }
>> + start = next;
>> + }
>> +
>> + return 0;
>> +}
>> +
>
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v20 7/9] arm64: Block hibernate and kexec while RMM is active
2026-09-29 22:16 [PATCH v20 0/9] firmware: arm_rmm: Add RMM v2.0 base RMI support Suzuki K Poulose
` (5 preceding siblings ...)
2026-09-29 22:16 ` [PATCH v20 6/9] firmware: arm_rmm: Ensure the RMM has GPT entries for memory Suzuki K Poulose
@ 2026-09-29 22:16 ` Suzuki K Poulose
2026-09-30 16:06 ` Jonathan Cameron
2026-09-29 22:16 ` [PATCH v20 8/9] firmware: arm_rmm: Add wrappers for Realm related RMI commands Suzuki K Poulose
2026-09-29 22:16 ` [PATCH v20 9/9] firmware: arm_rmm: hotplug: Skip memory added to ZONE_MOVABLE Suzuki K Poulose
8 siblings, 1 reply; 23+ messages in thread
From: Suzuki K Poulose @ 2026-09-29 22:16 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, sudeep.holla,
jonathan.cameron, Suzuki K Poulose
RMM can be deactivated only after all delegated granules have been
reclaimed. If a new kernel is entered while any granules remain in the
Realm PAS, accesses to that memory can raise a Granule Protection Fault
and be fatal to the new kernel.
Crash kexec/kdump needs separate handling. It can be supported only once
the crash kernel can tolerate delegated memory inherited from the primary
kernel. i.e., be able to read the pages safely and fixup the GPF. Until
then disable the kexec completely.
Hibernate has a similar problem. The image cannot be safely saved for
delegated pages, as the RMM doesn't support exporting the pages.
Disable both kexec and hiberation while the RMM is active.
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Changes since v19:
- New patch to disable kexec and hibernation with RMM
---
arch/arm64/kernel/hibernate.c | 6 ++++++
arch/arm64/kernel/machine_kexec.c | 11 +++++++++++
2 files changed, 17 insertions(+)
diff --git a/arch/arm64/kernel/hibernate.c b/arch/arm64/kernel/hibernate.c
index 7bf1174277772..d02f01c17febf 100644
--- a/arch/arm64/kernel/hibernate.c
+++ b/arch/arm64/kernel/hibernate.c
@@ -10,6 +10,7 @@
* Copyright (C) 2006 Rafael J. Wysocki <rjw@sisk.pl>
*/
#define pr_fmt(x) "hibernate: " x
+#include <linux/arm-rmi-cmds.h>
#include <linux/cpu.h>
#include <linux/kvm_host.h>
#include <linux/pm.h>
@@ -341,6 +342,11 @@ int swsusp_arch_suspend(void)
return -EBUSY;
}
+ if (is_rmm_active()) {
+ pr_err("Can't hibernate: RMM is active.\n");
+ return -EBUSY;
+ }
+
flags = local_daif_save();
if (__cpu_suspend_enter(&state)) {
diff --git a/arch/arm64/kernel/machine_kexec.c b/arch/arm64/kernel/machine_kexec.c
index 8f9bc2327dc85..48f343704cb54 100644
--- a/arch/arm64/kernel/machine_kexec.c
+++ b/arch/arm64/kernel/machine_kexec.c
@@ -6,6 +6,7 @@
* Copyright (C) Huawei Futurewei Technologies.
*/
+#include <linux/arm-rmi-cmds.h>
#include <linux/interrupt.h>
#include <linux/irq.h>
#include <linux/kernel.h>
@@ -59,6 +60,16 @@ int machine_kexec_prepare(struct kimage *kimage)
return -EBUSY;
}
+ /*
+ * We will be able to allow kdump to proceed, once we have the support
+ * for handling GPF from vmcore accesses to delegated pages. Until then
+ * block kexec completely.
+ */
+ if (is_rmm_active()) {
+ pr_err("Can't kexec: RMM is active.\n");
+ return -EBUSY;
+ }
+
return 0;
}
--
2.43.0
^ permalink raw reply [flat|nested] 23+ messages in thread* Re: [PATCH v20 7/9] arm64: Block hibernate and kexec while RMM is active
2026-09-29 22:16 ` [PATCH v20 7/9] arm64: Block hibernate and kexec while RMM is active Suzuki K Poulose
@ 2026-09-30 16:06 ` Jonathan Cameron
0 siblings, 0 replies; 23+ messages in thread
From: Jonathan Cameron @ 2026-09-30 16:06 UTC (permalink / raw)
To: Suzuki K Poulose
Cc: kvm, kvmarm, maz, will, catalin.marinas, linux-kernel,
linux-arm-kernel, steven.price, aneesh.kumar, oupton, gshan,
joey.gouly, tabba, yuzenghui, linux-coco, gankulkarni,
sdonthineni, alpergun, fj0570is, WeiLin.Chang, lpieralisi,
enju.kohei, sudeep.holla
On Tue, 29 Sep 2026 23:16:21 +0100
Suzuki K Poulose <suzuki.poulose@arm.com> wrote:
> RMM can be deactivated only after all delegated granules have been
> reclaimed. If a new kernel is entered while any granules remain in the
> Realm PAS, accesses to that memory can raise a Granule Protection Fault
> and be fatal to the new kernel.
>
> Crash kexec/kdump needs separate handling. It can be supported only once
> the crash kernel can tolerate delegated memory inherited from the primary
> kernel. i.e., be able to read the pages safely and fixup the GPF. Until
> then disable the kexec completely.
>
> Hibernate has a similar problem. The image cannot be safely saved for
> delegated pages, as the RMM doesn't support exporting the pages.
>
> Disable both kexec and hiberation while the RMM is active.
>
> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
Not my area of expertise but follows similar patterns to blocking these
for other reasons.
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v20 8/9] firmware: arm_rmm: Add wrappers for Realm related RMI commands
2026-09-29 22:16 [PATCH v20 0/9] firmware: arm_rmm: Add RMM v2.0 base RMI support Suzuki K Poulose
` (6 preceding siblings ...)
2026-09-29 22:16 ` [PATCH v20 7/9] arm64: Block hibernate and kexec while RMM is active Suzuki K Poulose
@ 2026-09-29 22:16 ` Suzuki K Poulose
2026-09-30 15:51 ` Catalin Marinas
2026-09-29 22:16 ` [PATCH v20 9/9] firmware: arm_rmm: hotplug: Skip memory added to ZONE_MOVABLE Suzuki K Poulose
8 siblings, 1 reply; 23+ messages in thread
From: Suzuki K Poulose @ 2026-09-29 22:16 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, sudeep.holla,
jonathan.cameron, Suzuki K Poulose
From: Steven Price <steven.price@arm.com>
Introduce wrappers for the RMI functions needed for creating and
managing realm guests. This will be used by the KVM to manage the
Realms
Reviewed-by: Gavin Shan <gshan@redhat.com>
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Signed-off-by: Steven Price <steven.price@arm.com>
Co-developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Changes since v19:
* Use rmi_sro_memxfr_command() for rec_create - Catalin
Changes since v18:
* Convert the last man standing nesting if - Gavin
* Pass out_top for RMI_ERROR_RTT for rmi_rtt_destroy()
Changes since v17:
* Avoid nesting if conditions for output populating RMI calls
* Clean up comments
* Always use arm_smccc_1_2_invoke() for RMI calls.
Changes since v16:
* Split into a separate patch and move away from arch/arm64 to
include/linux/.
* Also moved into the firmware_rmm series from the KVM CCA support.
This is done in a hope to reduce the merge conflicts and make
the KVM CCA upstreaming in independent parallel chunks
---
include/linux/arm-rmi-cmds.h | 460 +++++++++++++++++++++++++++++++++++
1 file changed, 460 insertions(+)
diff --git a/include/linux/arm-rmi-cmds.h b/include/linux/arm-rmi-cmds.h
index 5d005054f3c6c..4f0d6a8006a4a 100644
--- a/include/linux/arm-rmi-cmds.h
+++ b/include/linux/arm-rmi-cmds.h
@@ -104,4 +104,464 @@ static inline bool is_rmi_available(void)
}
#endif /* CONFIG_ARM_RMM_RMI */
+
+/**
+ * rmi_rtt_data_map_init() - Create a mapping at protected IPA, copying contents
+ * from a given non-secure source granule.
+ * @rd: PA of the RD
+ * @data: PA of the target granule mapped in the guest
+ * @ipa: IPA at which the granule @data will be mapped in the guest
+ * @src: PA of the source granule with contents
+ * @flags: RMI_MEASURE_CONTENT if the contents should be measured
+ *
+ * Create a mapping from Protected IPA space to conventional memory, copying
+ * contents from a Non-secure Granule provided by the caller.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_rtt_data_map_init(unsigned long rd, unsigned long data,
+ unsigned long ipa, unsigned long src,
+ unsigned long flags)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_RTT_DATA_MAP_INIT, rd, data, ipa, src, flags
+ };
+
+ return rmi_sro_execute(®s);
+}
+
+/**
+ * rmi_rtt_data_map() - Create mappings in protected IPA range with unknown contents
+ * @rd: PA of the RD
+ * @base: Base of the target IPA range
+ * @top: Top of the target IPA range
+ * @flags: Flags
+ * @oaddr: Output address set descriptor
+ * @out_top: Top address of range which was processed.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_rtt_data_map(unsigned long rd,
+ unsigned long base,
+ unsigned long top,
+ unsigned long flags,
+ unsigned long oaddr,
+ unsigned long *out_top)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_RTT_DATA_MAP, rd, base, top, flags, oaddr
+ };
+ long ret;
+
+ ret = rmi_sro_execute(®s);
+
+ if (ret == RMI_SUCCESS && out_top)
+ *out_top = regs.a1;
+
+ return ret;
+}
+
+/**
+ * rmi_rtt_data_unmap() - Remove mappings to conventional memory at a protected
+ * IPA range
+ * @rd: PA of the RD
+ * @base: Base of the target IPA range
+ * @top: Top of the target IPA range
+ * @flags: Flags
+ * @oaddr: Output address set descriptor
+ * @out_top: Returns top IPA of range which has been unmapped
+ * @out_range: Output address range
+ * @out_count: Number of entries in output address list
+ *
+ * Removes mappings to convention memory with a target Protected IPA range.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_rtt_data_unmap(unsigned long rd,
+ unsigned long base,
+ unsigned long top,
+ unsigned long flags,
+ unsigned long oaddr,
+ unsigned long *out_top,
+ unsigned long *out_range,
+ unsigned long *out_count)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_RTT_DATA_UNMAP, rd, base, top, flags, oaddr
+ };
+ long ret;
+
+ ret = rmi_sro_execute(®s);
+
+ if (ret != RMI_SUCCESS)
+ return ret;
+
+ if (out_top)
+ *out_top = regs.a1;
+ if (out_range)
+ *out_range = regs.a2;
+ if (out_count)
+ *out_count = regs.a3;
+
+ return RMI_SUCCESS;
+}
+
+/**
+ * rmi_psci_complete() - Complete pending PSCI command
+ * @calling_rec: PA of the calling REC
+ * @status: Status of the PSCI request
+ *
+ * Completes a pending PSCI command.
+ *
+ * Return: RMI return code
+ */
+static inline long rmi_psci_complete(unsigned long calling_rec,
+ unsigned long status)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_PSCI_COMPLETE, calling_rec, status,
+ };
+
+ rmi_smccc_invoke(®s);
+ return regs.a0;
+}
+
+/**
+ * rmi_realm_activate() - Activate a realm
+ * @rd: PA of the RD
+ *
+ * Mark a realm as Active, signalling that creation is completed, allowing
+ * execution of the realm.
+ *
+ * Return: RMI return code
+ */
+static inline long rmi_realm_activate(unsigned long rd)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_REALM_ACTIVATE, rd,
+ };
+
+ rmi_smccc_invoke(®s);
+ return regs.a0;
+}
+
+/**
+ * rmi_realm_create() - Create a realm
+ * @rd: PA of the RD
+ * @params: PA of realm parameters
+ * @sro: Preallocated SRO context
+ *
+ * Create a new realm using the given parameters.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_realm_create(unsigned long rd, unsigned long params,
+ struct rmi_sro_state *sro)
+{
+ return rmi_sro_memxfer_cmd(sro, GFP_KERNEL,
+ SMC_RMI_REALM_CREATE, rd, params);
+}
+
+/**
+ * rmi_realm_terminate() - Terminate a realm
+ * @rd: PA of the RD
+ * @sro: Preallocated SRO context
+ *
+ * Terminates a realm, moving it into a ZOMBIE state
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_realm_terminate(unsigned long rd,
+ struct rmi_sro_state *sro)
+{
+ return rmi_sro_memxfer_cmd(sro, GFP_KERNEL,
+ SMC_RMI_REALM_TERMINATE, rd);
+}
+
+/**
+ * rmi_realm_destroy() - Destroy a realm
+ * @rd: PA of the RD
+ * @sro: Preallocated SRO context
+ *
+ * Destroys a realm, all objects belonging to the realm must be destroyed first.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_realm_destroy(unsigned long rd,
+ struct rmi_sro_state *sro)
+{
+ return rmi_sro_memxfer_cmd(sro, GFP_KERNEL,
+ SMC_RMI_REALM_DESTROY, rd);
+}
+
+/**
+ * rmi_rec_create() - Create a REC
+ * @rd: PA of the RD
+ * @rec: PA of the target REC
+ * @params: PA of REC parameters
+ * @sro: Allocated SRO context to be used
+ *
+ * Create a REC using the parameters specified in the struct rec_params pointed
+ * to by @params.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_rec_create(unsigned long rd,
+ unsigned long rec,
+ unsigned long params,
+ struct rmi_sro_state *sro)
+{
+ return rmi_sro_memxfer_cmd(sro, GFP_KERNEL,
+ SMC_RMI_REC_CREATE, rd, rec, params);
+}
+
+/**
+ * rmi_rec_destroy() - Destroy a REC
+ * @rec: PA of the target REC
+ * @sro: Allocated SRO context to be used
+ *
+ * Destroys a REC. The REC must not be running.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_rec_destroy(unsigned long rec,
+ struct rmi_sro_state *sro)
+{
+ return rmi_sro_memxfer_cmd(sro, GFP_KERNEL, SMC_RMI_REC_DESTROY, rec);
+}
+
+/**
+ * rmi_rec_enter() - Enter a REC
+ * @rec: PA of the target REC
+ * @run_ptr: PA of RecRun structure
+ *
+ * Starts (or continues) execution within a REC.
+ *
+ * Return: RMI return code
+ */
+static inline long rmi_rec_enter(unsigned long rec, unsigned long run_ptr)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_REC_ENTER, rec, run_ptr,
+ };
+
+ rmi_smccc_invoke(®s);
+ return regs.a0;
+}
+
+/**
+ * rmi_rtt_create() - Creates an RTT
+ * @rd: PA of the RD
+ * @rtt: PA of the target RTT
+ * @ipa: Base of the IPA range described by the RTT
+ * @level: Depth of the RTT within the tree
+ *
+ * Creates an RTT (Realm Translation Table) at the specified level for the
+ * translation of the specified address within the realm.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_rtt_create(unsigned long rd, unsigned long rtt,
+ unsigned long ipa, long level)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_RTT_CREATE, rd, rtt, ipa, level
+ };
+
+ return rmi_sro_execute(®s);
+}
+
+/**
+ * rmi_rtt_destroy() - Destroy an RTT
+ * @rd: PA of the RD
+ * @ipa: Base of the IPA range described by the RTT
+ * @level: RTT level
+ * @out_rtt: Pointer to write the PA of the RTT which was destroyed
+ * @out_top: Pointer to write the top IPA of non-live RTT entries, from entry
+ * at which the RTT walk terminated.
+ *
+ * Destroys an RTT. The RTT must be non-live, i.e. none of the entries in the
+ * table are in ASSIGNED or TABLE state.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code.
+ */
+static inline long rmi_rtt_destroy(unsigned long rd,
+ unsigned long ipa,
+ long level,
+ unsigned long *out_rtt,
+ unsigned long *out_top)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_RTT_DESTROY, rd, ipa, level
+ };
+ long ret = rmi_sro_execute(®s);
+
+ switch (RMI_RESULT_STATUS(ret)) {
+ case RMI_SUCCESS:
+ if (out_rtt)
+ *out_rtt = regs.a1;
+ fallthrough;
+ case RMI_ERROR_RTT:
+ if (out_top)
+ *out_top = regs.a2;
+ break;
+ default:
+ break;
+ }
+
+ return ret;
+}
+
+/**
+ * rmi_rtt_fold() - Fold an RTT
+ * @rd: PA of the RD
+ * @ipa: Base of the IPA range described by the RTT
+ * @level: Depth of the RTT within the tree
+ * @out_rtt: Pointer to write the PA of the RTT which was destroyed
+ *
+ * Folds an RTT. If all entries with the RTT are 'homogeneous' the RTT can be
+ * folded into the parent and the RTT destroyed.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_rtt_fold(unsigned long rd, unsigned long ipa,
+ long level, unsigned long *out_rtt)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_RTT_FOLD, rd, ipa, level
+ };
+ long ret = rmi_sro_execute(®s);
+
+ if (ret == RMI_SUCCESS && out_rtt)
+ *out_rtt = regs.a1;
+
+ return ret;
+}
+
+/**
+ * rmi_rtt_init_ripas() - Set RIPAS for new realm
+ * @rd: PA of the RD
+ * @base: Base of target IPA region
+ * @top: Top of target IPA region
+ * @out_top: Top IPA of range whose RIPAS was modified
+ *
+ * Sets the RIPAS of a target IPA range to RAM, for a realm in the NEW state.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_rtt_init_ripas(unsigned long rd, unsigned long base,
+ unsigned long top, unsigned long *out_top)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_RTT_INIT_RIPAS, rd, base, top
+ };
+ long ret = rmi_sro_execute(®s);
+
+ if (ret == RMI_SUCCESS && out_top)
+ *out_top = regs.a1;
+
+ return ret;
+}
+
+/**
+ * rmi_rtt_unprot_map() - Map unprotected granules into a realm
+ * @rd: PA of the RD
+ * @base: Base IPA of the mapping
+ * @top: Top of the target IPA range
+ * @flags: Flags
+ * @oaddr: Output address set descriptor
+ * @out_top: Top IPA of range which has been mapped
+ *
+ * Create mappings to memory within a target unprotected IPA range.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_rtt_unprot_map(unsigned long rd,
+ unsigned long base,
+ unsigned long top,
+ unsigned long flags,
+ unsigned long oaddr,
+ unsigned long *out_top)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_RTT_UNPROT_MAP, rd, base, top, flags, oaddr
+ };
+ long ret = rmi_sro_execute(®s);
+
+ if (ret == RMI_SUCCESS && out_top)
+ *out_top = regs.a1;
+
+ return ret;
+}
+
+/**
+ * rmi_rtt_set_ripas() - Set RIPAS for an running realm
+ * @rd: PA of the RD
+ * @rec: PA of the REC making the request
+ * @base: Base of target IPA region
+ * @top: Top of target IPA region
+ * @out_top: Pointer to write top IPA of range whose RIPAS was modified
+ *
+ * Completes a request made by the realm to change the RIPAS of a target IPA
+ * range.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_rtt_set_ripas(unsigned long rd, unsigned long rec,
+ unsigned long base, unsigned long top,
+ unsigned long *out_top)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_RTT_SET_RIPAS, rd, rec, base, top
+ };
+ long ret = rmi_sro_execute(®s);
+
+ if (ret == RMI_SUCCESS && out_top)
+ *out_top = regs.a1;
+
+ return ret;
+}
+
+/**
+ * rmi_rtt_unprot_unmap() - Remove mappings within an unprotected IPA range
+ * @rd: PA of the RD
+ * @base: Base IPA of the mapping
+ * @top: Top of the target IPA range
+ * @flags: Flags
+ * @oaddr: Output address set descriptor
+ * @out_top: Top IPA which has been unmapped
+ * @out_range: Output address range
+ * @out_count: Number of entries in output address list
+ *
+ * Removes mappings to memory within a target unprotected IPA range.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_rtt_unprot_unmap(unsigned long rd,
+ unsigned long base,
+ unsigned long top,
+ unsigned long flags,
+ unsigned long oaddr,
+ unsigned long *out_top,
+ unsigned long *out_range,
+ unsigned long *out_count)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_RTT_UNPROT_UNMAP, rd, base, top, flags, oaddr
+ };
+ long ret = rmi_sro_execute(®s);
+
+ if (ret != RMI_SUCCESS)
+ return ret;
+
+ if (out_top)
+ *out_top = regs.a1;
+ if (out_range)
+ *out_range = regs.a2;
+ if (out_count)
+ *out_count = regs.a3;
+
+ return RMI_SUCCESS;
+}
+
#endif /* __LINUX_ARM_RMI_CMDS_H_ */
--
2.43.0
^ permalink raw reply [flat|nested] 23+ messages in thread* Re: [PATCH v20 8/9] firmware: arm_rmm: Add wrappers for Realm related RMI commands
2026-09-29 22:16 ` [PATCH v20 8/9] firmware: arm_rmm: Add wrappers for Realm related RMI commands Suzuki K Poulose
@ 2026-09-30 15:51 ` Catalin Marinas
0 siblings, 0 replies; 23+ messages in thread
From: Catalin Marinas @ 2026-09-30 15:51 UTC (permalink / raw)
To: Suzuki K Poulose
Cc: kvm, kvmarm, maz, will, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, sudeep.holla,
jonathan.cameron
On Tue, Sep 29, 2026 at 11:16:22PM +0100, Suzuki K Poulose wrote:
> From: Steven Price <steven.price@arm.com>
>
> Introduce wrappers for the RMI functions needed for creating and
> managing realm guests. This will be used by the KVM to manage the
> Realms
>
> Reviewed-by: Gavin Shan <gshan@redhat.com>
> Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
> Signed-off-by: Steven Price <steven.price@arm.com>
> Co-developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
Reviewed-by: Catalin Marinas <catalin.marinas@arm.com>
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v20 9/9] firmware: arm_rmm: hotplug: Skip memory added to ZONE_MOVABLE
2026-09-29 22:16 [PATCH v20 0/9] firmware: arm_rmm: Add RMM v2.0 base RMI support Suzuki K Poulose
` (7 preceding siblings ...)
2026-09-29 22:16 ` [PATCH v20 8/9] firmware: arm_rmm: Add wrappers for Realm related RMI commands Suzuki K Poulose
@ 2026-09-29 22:16 ` Suzuki K Poulose
2026-09-30 12:22 ` David Hildenbrand (Arm)
2026-09-30 15:53 ` Catalin Marinas
8 siblings, 2 replies; 23+ messages in thread
From: Suzuki K Poulose @ 2026-09-29 22:16 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, sudeep.holla,
jonathan.cameron, Suzuki K Poulose, David Hildenbrand
If the hotadded memory is in ZONE_MOVABLE, we could allow that to proceed
without the RMM having fine grained tracking, as we won't allocate pages
from there for guest-memfd or kernel allocation
Cc: Catalin Marinas <catalin.marinas@arm.com>
Cc: David Hildenbrand (Arm) <david@kernel.org>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Changes since v19:
- New patch in v20, keeping this at the tail to allow for the rest of the
series to proceed, while this patch gets discussed.
---
drivers/firmware/arm_rmm/rmi.c | 12 ++++++++++++
1 file changed, 12 insertions(+)
diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/arm_rmm/rmi.c
index d4d098448e46e..ef0a171ba9677 100644
--- a/drivers/firmware/arm_rmm/rmi.c
+++ b/drivers/firmware/arm_rmm/rmi.c
@@ -1043,6 +1043,18 @@ static int rmi_memory_notifier(struct notifier_block *nb,
start = PFN_PHYS(arg->start_pfn);
end = PFN_PHYS(arg->start_pfn + arg->nr_pages);
+ /*
+ * If the hotplugged memory region is in ZONE_MOVABLE, neither the
+ * guest-memfd nor the kernel allocations can come from there. So we are
+ * fine to skip this.
+ * For boot time memory, we don't allow this discount, as the
+ * decision to move boot memory pages to ZONE_MOVABLE is purely
+ * a kernel choice and the firmware must have decided what it covers.
+ */
+ if (page_zonenum(pfn_to_page(arg->start_pfn)) == ZONE_MOVABLE &&
+ page_zonenum(pfn_to_page(arg->start_pfn + arg->nr_pages - 1)) == ZONE_MOVABLE)
+ return NOTIFY_DONE;
+
ret = rmi_prepare_memory(start, end);
return notifier_from_errno(ret);
--
2.43.0
^ permalink raw reply [flat|nested] 23+ messages in thread* Re: [PATCH v20 9/9] firmware: arm_rmm: hotplug: Skip memory added to ZONE_MOVABLE
2026-09-29 22:16 ` [PATCH v20 9/9] firmware: arm_rmm: hotplug: Skip memory added to ZONE_MOVABLE Suzuki K Poulose
@ 2026-09-30 12:22 ` David Hildenbrand (Arm)
2026-09-30 12:47 ` Suzuki K Poulose
2026-09-30 15:53 ` Catalin Marinas
1 sibling, 1 reply; 23+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-30 12:22 UTC (permalink / raw)
To: Suzuki K Poulose, kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, sudeep.holla,
jonathan.cameron
On 9/30/26 00:16, Suzuki K Poulose wrote:
> If the hotadded memory is in ZONE_MOVABLE, we could allow that to proceed
> without the RMM having fine grained tracking, as we won't allocate pages
> from there for guest-memfd or kernel allocation
>
> Cc: Catalin Marinas <catalin.marinas@arm.com>
> Cc: David Hildenbrand (Arm) <david@kernel.org>
> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
> ---
> Changes since v19:
> - New patch in v20, keeping this at the tail to allow for the rest of the
> series to proceed, while this patch gets discussed.
> ---
> drivers/firmware/arm_rmm/rmi.c | 12 ++++++++++++
> 1 file changed, 12 insertions(+)
>
> diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/arm_rmm/rmi.c
> index d4d098448e46e..ef0a171ba9677 100644
> --- a/drivers/firmware/arm_rmm/rmi.c
> +++ b/drivers/firmware/arm_rmm/rmi.c
> @@ -1043,6 +1043,18 @@ static int rmi_memory_notifier(struct notifier_block *nb,
> start = PFN_PHYS(arg->start_pfn);
> end = PFN_PHYS(arg->start_pfn + arg->nr_pages);
>
> + /*
> + * If the hotplugged memory region is in ZONE_MOVABLE, neither the
> + * guest-memfd nor the kernel allocations can come from there. So we are
> + * fine to skip this.
> + * For boot time memory, we don't allow this discount, as the
> + * decision to move boot memory pages to ZONE_MOVABLE is purely
> + * a kernel choice and the firmware must have decided what it covers.
> + */
> + if (page_zonenum(pfn_to_page(arg->start_pfn)) == ZONE_MOVABLE &&
> + page_zonenum(pfn_to_page(arg->start_pfn + arg->nr_pages - 1)) == ZONE_MOVABLE)
> + return NOTIFY_DONE;
> +
is_zone_movable_page(pfn_to_page(start_pfn)
as used by virtio-mem should be sufficient. Checking only the start pfn is
sufficient.
--
Cheers,
David
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH v20 9/9] firmware: arm_rmm: hotplug: Skip memory added to ZONE_MOVABLE
2026-09-30 12:22 ` David Hildenbrand (Arm)
@ 2026-09-30 12:47 ` Suzuki K Poulose
0 siblings, 0 replies; 23+ messages in thread
From: Suzuki K Poulose @ 2026-09-30 12:47 UTC (permalink / raw)
To: David Hildenbrand (Arm), kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, sudeep.holla,
jonathan.cameron
On 30/09/2026 13:22, David Hildenbrand (Arm) wrote:
> On 9/30/26 00:16, Suzuki K Poulose wrote:
>> If the hotadded memory is in ZONE_MOVABLE, we could allow that to proceed
>> without the RMM having fine grained tracking, as we won't allocate pages
>> from there for guest-memfd or kernel allocation
>>
>> Cc: Catalin Marinas <catalin.marinas@arm.com>
>> Cc: David Hildenbrand (Arm) <david@kernel.org>
>> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
>> ---
>> Changes since v19:
>> - New patch in v20, keeping this at the tail to allow for the rest of the
>> series to proceed, while this patch gets discussed.
>> ---
>> drivers/firmware/arm_rmm/rmi.c | 12 ++++++++++++
>> 1 file changed, 12 insertions(+)
>>
>> diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/arm_rmm/rmi.c
>> index d4d098448e46e..ef0a171ba9677 100644
>> --- a/drivers/firmware/arm_rmm/rmi.c
>> +++ b/drivers/firmware/arm_rmm/rmi.c
>> @@ -1043,6 +1043,18 @@ static int rmi_memory_notifier(struct notifier_block *nb,
>> start = PFN_PHYS(arg->start_pfn);
>> end = PFN_PHYS(arg->start_pfn + arg->nr_pages);
>>
>> + /*
>> + * If the hotplugged memory region is in ZONE_MOVABLE, neither the
>> + * guest-memfd nor the kernel allocations can come from there. So we are
>> + * fine to skip this.
>> + * For boot time memory, we don't allow this discount, as the
>> + * decision to move boot memory pages to ZONE_MOVABLE is purely
>> + * a kernel choice and the firmware must have decided what it covers.
>> + */
>> + if (page_zonenum(pfn_to_page(arg->start_pfn)) == ZONE_MOVABLE &&
>> + page_zonenum(pfn_to_page(arg->start_pfn + arg->nr_pages - 1)) == ZONE_MOVABLE)
>> + return NOTIFY_DONE;
>> +
>
> is_zone_movable_page(pfn_to_page(start_pfn)
>
> as used by virtio-mem should be sufficient. Checking only the start pfn is
> sufficient.
Thanks David, I will update it.
Cheers
Suzuki
>
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH v20 9/9] firmware: arm_rmm: hotplug: Skip memory added to ZONE_MOVABLE
2026-09-29 22:16 ` [PATCH v20 9/9] firmware: arm_rmm: hotplug: Skip memory added to ZONE_MOVABLE Suzuki K Poulose
2026-09-30 12:22 ` David Hildenbrand (Arm)
@ 2026-09-30 15:53 ` Catalin Marinas
1 sibling, 0 replies; 23+ messages in thread
From: Catalin Marinas @ 2026-09-30 15:53 UTC (permalink / raw)
To: Suzuki K Poulose
Cc: kvm, kvmarm, maz, will, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, sudeep.holla,
jonathan.cameron, David Hildenbrand
On Tue, Sep 29, 2026 at 11:16:23PM +0100, Suzuki K Poulose wrote:
> If the hotadded memory is in ZONE_MOVABLE, we could allow that to proceed
> without the RMM having fine grained tracking, as we won't allocate pages
> from there for guest-memfd or kernel allocation
>
> Cc: Catalin Marinas <catalin.marinas@arm.com>
> Cc: David Hildenbrand (Arm) <david@kernel.org>
> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
With David's suggestion:
Reviewed-by: Catalin Marinas <catalin.marinas@arm.com>
^ permalink raw reply [flat|nested] 23+ messages in thread