* [PATCH v19 01/20] KVM: arm64: protected VM: Handle user writes to CNTVCT_EL0/CNTPCT_EL0
2026-09-20 21:28 [PATCH v19 00/20] KVM: arm64: CCA: Add basic plumbing for Realms Suzuki K Poulose
@ 2026-09-20 21:28 ` Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 02/20] KVM: arm64: Disable Steal time accounting for protected guests Suzuki K Poulose
` (18 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: Suzuki K Poulose @ 2026-09-20 21:28 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose
Protected VMs doesn't allow setting offsets for virtual and phyiscal
counters, as the offset is always fixed to 0. The VM ioclt is filtered
out based on the cap. However we don't prevent the userspace from trying
to write to the CNTVCT/CNTPCT registers. This would lead to KVM triggering
a WARN() in timer_set_offset() as the vm_offset pointer is set to NULL.
Fix this by always "fixing" the timer offsets to 0 and marking that the
timer offset is set in the kvm->arch.flags at KVM init time for protected
VMs. A userspace writing to the CNT*CT_EL0 would observe success, without
any real effect. This was chosen over preventing the writes to these
registers and returning -EPERM.
Reported by Sashiko
Link: https://lore.kernel.org/all/20260908164641.416911F00A3A@smtp.kernel.org
Fixes: f7d05ee84a6a ("KVM: arm64: Prevent host from managing timer offsets for protected VMs")
Suggested-by: Marc Zyngier <maz@kernel.org>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Changes since v18:
- Retain NULL vm_offset for protected VMs to avoid host tampering with the
offset.
- Moved the flag setting into kvm_timer_init_vm(), where it should have been
in the first place
---
arch/arm64/kvm/arch_timer.c | 12 ++++++++++--
1 file changed, 10 insertions(+), 2 deletions(-)
diff --git a/arch/arm64/kvm/arch_timer.c b/arch/arm64/kvm/arch_timer.c
index 6ac3321f4c575..226cd5a495c8b 100644
--- a/arch/arm64/kvm/arch_timer.c
+++ b/arch/arm64/kvm/arch_timer.c
@@ -1110,8 +1110,7 @@ void kvm_timer_vcpu_init(struct kvm_vcpu *vcpu)
timer_context_init(vcpu, i);
/* Synchronize offsets across timers of a VM if not already provided */
- if (!vcpu_is_protected(vcpu) &&
- !test_bit(KVM_ARCH_FLAG_VM_COUNTER_OFFSET, &vcpu->kvm->arch.flags)) {
+ if (!test_bit(KVM_ARCH_FLAG_VM_COUNTER_OFFSET, &vcpu->kvm->arch.flags)) {
timer_set_offset(vcpu_vtimer(vcpu), kvm_phys_timer_read());
timer_set_offset(vcpu_ptimer(vcpu), 0);
}
@@ -1133,6 +1132,15 @@ void kvm_timer_init_vm(struct kvm *kvm)
*/
for (int i = 0; i < NR_KVM_TIMERS; i++)
kvm->arch.timer_data.ppi[i] = get_vgic_ppi(kvm, default_ppi[i]);
+
+ /*
+ * Protected VMs don't allow any offset being set from userspace,
+ * either set via writes to the counters or using the dedicated
+ * ioctl. Pretend the offset has already been set and rely on the
+ * default offset being 0.
+ */
+ if (kvm_vm_is_protected(kvm))
+ set_bit(KVM_ARCH_FLAG_VM_COUNTER_OFFSET, &kvm->arch.flags);
}
void kvm_timer_cpu_up(void)
--
2.43.0
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v19 02/20] KVM: arm64: Disable Steal time accounting for protected guests
2026-09-20 21:28 [PATCH v19 00/20] KVM: arm64: CCA: Add basic plumbing for Realms Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 01/20] KVM: arm64: protected VM: Handle user writes to CNTVCT_EL0/CNTPCT_EL0 Suzuki K Poulose
@ 2026-09-20 21:28 ` Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 03/20] KVM: arm64: Include kvm_emulate.h in kvm/arm_psci.h Suzuki K Poulose
` (17 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: Suzuki K Poulose @ 2026-09-20 21:28 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose,
Fuad Tabba
PVTIME support is advertised by KVM_CAP_STEAL_TIME, which doesn't take into
account the kvm instance. Even with that, a VMM could skip the CAP check
and proceed to configure the PVTIME as we don't do further check on the
DEVICE_CTRL. Tighten this up by passing the KVM instance around wherever
possible and catch things early.
Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
arch/arm64/include/asm/kvm_host.h | 2 +-
arch/arm64/kvm/arm.c | 2 +-
arch/arm64/kvm/pvtime.c | 14 +++++++-------
3 files changed, 9 insertions(+), 9 deletions(-)
diff --git a/arch/arm64/include/asm/kvm_host.h b/arch/arm64/include/asm/kvm_host.h
index 27fe0cd5b2d7a..286489a69dff5 100644
--- a/arch/arm64/include/asm/kvm_host.h
+++ b/arch/arm64/include/asm/kvm_host.h
@@ -1346,7 +1346,7 @@ long kvm_hypercall_pv_features(struct kvm_vcpu *vcpu);
gpa_t kvm_init_stolen_time(struct kvm_vcpu *vcpu);
void kvm_update_stolen_time(struct kvm_vcpu *vcpu);
-bool kvm_arm_pvtime_supported(void);
+bool kvm_arm_pvtime_supported(struct kvm *kvm);
int kvm_arm_pvtime_set_attr(struct kvm_vcpu *vcpu,
struct kvm_device_attr *attr);
int kvm_arm_pvtime_get_attr(struct kvm_vcpu *vcpu,
diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c
index 8b080804bc90b..db36815630790 100644
--- a/arch/arm64/kvm/arm.c
+++ b/arch/arm64/kvm/arm.c
@@ -447,7 +447,7 @@ int kvm_vm_ioctl_check_extension(struct kvm *kvm, long ext)
r = system_supports_mte();
break;
case KVM_CAP_STEAL_TIME:
- r = kvm_arm_pvtime_supported();
+ r = kvm_arm_pvtime_supported(kvm);
break;
case KVM_CAP_ARM_EL1_32BIT:
r = cpus_have_final_cap(ARM64_HAS_32BIT_EL1);
diff --git a/arch/arm64/kvm/pvtime.c b/arch/arm64/kvm/pvtime.c
index 4ceabaa4c30bd..579e0a4720ad2 100644
--- a/arch/arm64/kvm/pvtime.c
+++ b/arch/arm64/kvm/pvtime.c
@@ -67,9 +67,9 @@ gpa_t kvm_init_stolen_time(struct kvm_vcpu *vcpu)
return base;
}
-bool kvm_arm_pvtime_supported(void)
+bool kvm_arm_pvtime_supported(struct kvm *kvm)
{
- return !!sched_info_on();
+ return !!sched_info_on() && (!kvm || !kvm_vm_is_protected(kvm));
}
int kvm_arm_pvtime_set_attr(struct kvm_vcpu *vcpu,
@@ -81,8 +81,8 @@ int kvm_arm_pvtime_set_attr(struct kvm_vcpu *vcpu,
int ret = 0;
int idx;
- if (!kvm_arm_pvtime_supported() ||
- attr->attr != KVM_ARM_VCPU_PVTIME_IPA)
+ if (!kvm_arm_pvtime_supported(kvm) ||
+ (attr->attr != KVM_ARM_VCPU_PVTIME_IPA))
return -ENXIO;
if (get_user(ipa, user))
@@ -110,8 +110,8 @@ int kvm_arm_pvtime_get_attr(struct kvm_vcpu *vcpu,
u64 __user *user = (u64 __user *)attr->addr;
u64 ipa;
- if (!kvm_arm_pvtime_supported() ||
- attr->attr != KVM_ARM_VCPU_PVTIME_IPA)
+ if (!kvm_arm_pvtime_supported(vcpu->kvm) ||
+ (attr->attr != KVM_ARM_VCPU_PVTIME_IPA))
return -ENXIO;
ipa = vcpu->arch.steal.base;
@@ -126,7 +126,7 @@ int kvm_arm_pvtime_has_attr(struct kvm_vcpu *vcpu,
{
switch (attr->attr) {
case KVM_ARM_VCPU_PVTIME_IPA:
- if (kvm_arm_pvtime_supported())
+ if (kvm_arm_pvtime_supported(vcpu->kvm))
return 0;
}
return -ENXIO;
--
2.43.0
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v19 03/20] KVM: arm64: Include kvm_emulate.h in kvm/arm_psci.h
2026-09-20 21:28 [PATCH v19 00/20] KVM: arm64: CCA: Add basic plumbing for Realms Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 01/20] KVM: arm64: protected VM: Handle user writes to CNTVCT_EL0/CNTPCT_EL0 Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 02/20] KVM: arm64: Disable Steal time accounting for protected guests Suzuki K Poulose
@ 2026-09-20 21:28 ` Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 04/20] KVM: arm64: Avoid including linux/kvm_host.h in kvm_pgtable.h Suzuki K Poulose
` (16 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: Suzuki K Poulose @ 2026-09-20 21:28 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose,
Fuad Tabba
Fix a potential build error (like below, when asm/kvm_emulate.h gets
included after the kvm/arm_psci.h) by including the missing header file
in kvm/arm_psci.h:
./include/kvm/arm_psci.h: In function ‘kvm_psci_version’:
./include/kvm/arm_psci.h:29:13: error: implicit declaration of function
‘vcpu_has_feature’; did you mean ‘cpu_have_feature’? [-Werror=implicit-function-declaration]
29 | if (vcpu_has_feature(vcpu, KVM_ARM_VCPU_PSCI_0_2)) {
| ^~~~~~~~~~~~~~~~
| cpu_have_feature
Reviewed-by: Gavin Shan <gshan@redhat.com>
Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
include/kvm/arm_psci.h | 2 ++
1 file changed, 2 insertions(+)
diff --git a/include/kvm/arm_psci.h b/include/kvm/arm_psci.h
index f86a006d67136..06c20612e9e7d 100644
--- a/include/kvm/arm_psci.h
+++ b/include/kvm/arm_psci.h
@@ -10,6 +10,8 @@
#include <linux/kvm_host.h>
#include <uapi/linux/psci.h>
+#include <asm/kvm_emulate.h>
+
#define KVM_ARM_PSCI_0_1 PSCI_VERSION(0, 1)
#define KVM_ARM_PSCI_0_2 PSCI_VERSION(0, 2)
#define KVM_ARM_PSCI_1_0 PSCI_VERSION(1, 0)
--
2.43.0
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v19 04/20] KVM: arm64: Avoid including linux/kvm_host.h in kvm_pgtable.h
2026-09-20 21:28 [PATCH v19 00/20] KVM: arm64: CCA: Add basic plumbing for Realms Suzuki K Poulose
` (2 preceding siblings ...)
2026-09-20 21:28 ` [PATCH v19 03/20] KVM: arm64: Include kvm_emulate.h in kvm/arm_psci.h Suzuki K Poulose
@ 2026-09-20 21:28 ` Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 05/20] KVM: arm64: Track the type of VM in kvm_arch Suzuki K Poulose
` (15 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: Suzuki K Poulose @ 2026-09-20 21:28 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Fuad Tabba,
Suzuki K Poulose
From: Steven Price <steven.price@arm.com>
To avoid future include cycles, drop the linux/kvm_host.h include in
kvm_pgtable.h and include the lightweight headers required for the types
and inline helpers used there. Additionally provide a forward
declaration for struct kvm_s2_mmu as it's only used as a pointer in this
file.
Both pgtable.c and kvm_pkvm.h relied on the indirect inclusion of
kvm_host.h, so make that explicit.
Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev>
Reviewed-by: Gavin Shan <gshan@redhat.com>
Signed-off-by: Steven Price <steven.price@arm.com>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
arch/arm64/include/asm/kvm_pgtable.h | 6 +++++-
arch/arm64/include/asm/kvm_pkvm.h | 2 +-
arch/arm64/kvm/hyp/pgtable.c | 1 +
3 files changed, 7 insertions(+), 2 deletions(-)
diff --git a/arch/arm64/include/asm/kvm_pgtable.h b/arch/arm64/include/asm/kvm_pgtable.h
index 41a8687938eb6..c2e4b29e605fc 100644
--- a/arch/arm64/include/asm/kvm_pgtable.h
+++ b/arch/arm64/include/asm/kvm_pgtable.h
@@ -8,9 +8,13 @@
#define __ARM64_KVM_PGTABLE_H__
#include <linux/bits.h>
-#include <linux/kvm_host.h>
+#include <linux/kvm_types.h>
+#include <linux/rbtree_types.h>
+#include <linux/rcupdate.h>
#include <linux/types.h>
+struct kvm_s2_mmu;
+
#define KVM_PGTABLE_FIRST_LEVEL -1
#define KVM_PGTABLE_LAST_LEVEL 3
diff --git a/arch/arm64/include/asm/kvm_pkvm.h b/arch/arm64/include/asm/kvm_pkvm.h
index beea00e693a0a..54a618d887fa4 100644
--- a/arch/arm64/include/asm/kvm_pkvm.h
+++ b/arch/arm64/include/asm/kvm_pkvm.h
@@ -7,9 +7,9 @@
#define __ARM64_KVM_PKVM_H__
#include <linux/arm_ffa.h>
+#include <linux/kvm_host.h>
#include <linux/memblock.h>
#include <linux/scatterlist.h>
-#include <asm/kvm_host.h>
#include <asm/kvm_pgtable.h>
/* Maximum number of VMs that can co-exist under pKVM. */
diff --git a/arch/arm64/kvm/hyp/pgtable.c b/arch/arm64/kvm/hyp/pgtable.c
index b74dd5ce1efd3..f48253b9d88b5 100644
--- a/arch/arm64/kvm/hyp/pgtable.c
+++ b/arch/arm64/kvm/hyp/pgtable.c
@@ -8,6 +8,7 @@
*/
#include <linux/bitfield.h>
+#include <linux/kvm_host.h>
#include <asm/kvm_pgtable.h>
#include <asm/stage2_pgtable.h>
--
2.43.0
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v19 05/20] KVM: arm64: Track the type of VM in kvm_arch
2026-09-20 21:28 [PATCH v19 00/20] KVM: arm64: CCA: Add basic plumbing for Realms Suzuki K Poulose
` (3 preceding siblings ...)
2026-09-20 21:28 ` [PATCH v19 04/20] KVM: arm64: Avoid including linux/kvm_host.h in kvm_pgtable.h Suzuki K Poulose
@ 2026-09-20 21:28 ` Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 06/20] KVM: arm64: Refactor the vcpu_load to allow for VM specific callbacks Suzuki K Poulose
` (14 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: Suzuki K Poulose @ 2026-09-20 21:28 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose
KVM arm64 has different types of VMs with all the different modes in which
the hypervisor code can be run. e.g., VHE, nVHE, pKVM etc. Then there is
protected VM and normal VMs with pKVM. We might soon add other types,
e.g., Arm CCA Realm. So in an effort to make the handling of these
different types of VMs a bit more friendly to the eyes, add a VM flavor to
the kvm_arch and we could then add handlers for different operations based
on the VM type.
Keep the flavor initialisation at the beginning to allow for the detection
early enough and fail out on any unsupported requests.
With that, add wrappers for checking the "type" of a VM and replace the
existing users with the new wrappers.
Given we already have the construct of "kvm_vm_is_protected" in the core
KVM code, use that for all confidential compute guests including Realms
that we are about to add.
Adds __VM_PROTECTED marker vm flavor to generalize kvm_vm_is_protected()
to predicate all confidential guests running on KVM. In later patches, we
would add Realm VMs, which would also be classified as protected.
Add explicit helper to detect if a given VM is a "protected" VM under pKVM.
Change the existing users that precisely want to check the VM type. These
include :
- kvm_arch_prepare_memory_region - For preventing memslot changes after
pVM creation.
All the others are retained as a wider check for confidential guest VMs.
These are:
- kvm_vm_ioctl_set_counter_offset - For disallowing timer offset
configuration
- io_mem_abort for dabt handling without valid syndrome information
Both of which are true for Realms too.
Realms support is restricted to VHE host and thus "kvm_vm_is_protected()"
checks in the pkvm hyp specific code doesn't need to change, as the only
protected guests it deals with is "protected pKVM" guests. To tighten this
init_pkvm_hyp_vm() restricts the hyp copy of the vm_flavor to the ones it
supports.
Suggested-by: Marc Zyngier <maz@kernel.org>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Changes since v18:
- Merge the __VM_PROTECTED marker and the widening of kvm_vm_is_protected()
to this patch.
- Merge the use of kvm_vm_is_unprotected_pkvm() for !kvm_vm_is_protected()
given the scope changes here.
- Drop Fuad's review tag, as this patch has multiple merges
- Restrict the VM flavors to the supported types in init_pkvm_hyp_vm().
- Drop kvm_vm_hyp_is_pkvm() and revert to is_protected_kvm_enabled()
- Use is_protected_kvm_enabled() to make the pKVM guest flavor checks.
- s/PKVM/pKVM for commit descriptions too
Changes since v17:
* s/PKVM/pKVM for the comments
* Drop type argument for pkvm_init_host_vm and also drop protected variable.
* Add helpers for checking if the VM is running on pKVM (kvm_vm_hyp_is_pkvm())
* Use kvm_vm_hyp_is_pkvm() to replace is_protected_kvm_enabled() with valid
kvm instance
---
arch/arm64/include/asm/kvm_host.h | 22 +++++++++++++++++++---
arch/arm64/include/asm/kvm_pkvm.h | 4 ++--
arch/arm64/kvm/arm.c | 31 ++++++++++++++++++++++++++-----
arch/arm64/kvm/handle_exit.c | 2 +-
arch/arm64/kvm/hyp/nvhe/pkvm.c | 6 +++++-
arch/arm64/kvm/mmu.c | 2 +-
arch/arm64/kvm/pkvm.c | 6 ++----
7 files changed, 56 insertions(+), 17 deletions(-)
diff --git a/arch/arm64/include/asm/kvm_host.h b/arch/arm64/include/asm/kvm_host.h
index 286489a69dff5..9b1cf9c59e81f 100644
--- a/arch/arm64/include/asm/kvm_host.h
+++ b/arch/arm64/include/asm/kvm_host.h
@@ -257,7 +257,6 @@ struct kvm_protected_vm {
pkvm_handle_t handle;
struct kvm_hyp_memcache teardown_mc;
struct kvm_hyp_memcache stage2_teardown_mc;
- bool is_protected;
bool is_created;
/*
@@ -306,9 +305,19 @@ enum fgt_group_id {
__NR_FGT_GROUP_IDS__
};
+enum kvm_arm_vm_flavor {
+ VM_NVHE,
+ VM_VHE,
+ VM_PKVM, /* Normal guests on pKVM */
+ MARKER(__VM_PROTECTED),
+ VM_PROTECTED_PKVM, /* Protected VM */
+ VM_FLAVOR_MAX,
+};
+
struct kvm_arch {
struct kvm_s2_mmu mmu;
+ enum kvm_arm_vm_flavor vm_flavor;
/*
* Fine-Grained UNDEF, mimicking the FGT layout defined by the
* architecture. We track them globally, as we present the
@@ -1504,10 +1513,17 @@ struct kvm *kvm_arch_alloc_vm(void);
#define __KVM_HAVE_ARCH_FLUSH_REMOTE_TLBS_RANGE
-#define kvm_vm_is_protected(kvm) (is_protected_kvm_enabled() && (kvm)->arch.pkvm.is_protected)
-
+#define kvm_vm_is_protected(kvm) ((kvm)->arch.vm_flavor >= __VM_PROTECTED)
#define vcpu_is_protected(vcpu) kvm_vm_is_protected((vcpu)->kvm)
+#define kvm_vm_is_protected_pkvm(kvm) \
+ (is_protected_kvm_enabled() && ((kvm)->arch.vm_flavor == VM_PROTECTED_PKVM))
+#define vcpu_is_protected_pkvm(vcpu) kvm_vm_is_protected_pkvm(vcpu->kvm)
+
+#define kvm_vm_is_unprotected_pkvm(kvm) \
+ (is_protected_kvm_enabled() && ((kvm)->arch.vm_flavor == VM_PKVM))
+
+
int kvm_arm_vcpu_finalize(struct kvm_vcpu *vcpu, int feature);
bool kvm_arm_vcpu_is_finalized(struct kvm_vcpu *vcpu);
diff --git a/arch/arm64/include/asm/kvm_pkvm.h b/arch/arm64/include/asm/kvm_pkvm.h
index 54a618d887fa4..e4ea80711bec6 100644
--- a/arch/arm64/include/asm/kvm_pkvm.h
+++ b/arch/arm64/include/asm/kvm_pkvm.h
@@ -17,7 +17,7 @@
#define HYP_MEMBLOCK_REGIONS 128
-int pkvm_init_host_vm(struct kvm *kvm, unsigned long type);
+int pkvm_init_host_vm(struct kvm *kvm);
int pkvm_create_hyp_vm(struct kvm *kvm);
bool pkvm_hyp_vm_is_created(struct kvm *kvm);
void pkvm_destroy_hyp_vm(struct kvm *kvm);
@@ -49,7 +49,7 @@ static inline bool kvm_pkvm_ext_allowed(struct kvm *kvm, long ext)
case KVM_CAP_ARM_SUPPORTED_BLOCK_SIZES:
return false;
default:
- return !kvm || !kvm_vm_is_protected(kvm);
+ return !kvm || kvm_vm_is_unprotected_pkvm(kvm);
}
}
diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c
index db36815630790..8c784b266a8e8 100644
--- a/arch/arm64/kvm/arm.c
+++ b/arch/arm64/kvm/arm.c
@@ -214,6 +214,26 @@ static int kvm_arm_default_max_vcpus(void)
return vgic_present ? kvm_vgic_get_max_vcpus() : KVM_MAX_VCPUS;
}
+static int kvm_init_vm_flavor(struct kvm *kvm, unsigned long type)
+{
+ bool protected = type & KVM_VM_TYPE_ARM_PROTECTED;
+
+ if (is_protected_kvm_enabled()) {
+ if (protected)
+ kvm->arch.vm_flavor = VM_PROTECTED_PKVM;
+ else
+ kvm->arch.vm_flavor = VM_PKVM;
+ } else if (protected) {
+ return -EINVAL;
+ } else if (has_vhe()) {
+ kvm->arch.vm_flavor = VM_VHE;
+ } else {
+ kvm->arch.vm_flavor = VM_NVHE;
+ }
+
+ return 0;
+}
+
/**
* kvm_arch_init_vm - initializes a VM data structure
* @kvm: pointer to the KVM struct
@@ -236,6 +256,10 @@ int kvm_arch_init_vm(struct kvm *kvm, unsigned long type)
mutex_unlock(&kvm->lock);
#endif
+ ret = kvm_init_vm_flavor(kvm, type);
+ if (ret)
+ return ret;
+
kvm_init_nested(kvm);
ret = kvm_share_hyp(kvm, kvm + 1);
@@ -257,12 +281,9 @@ int kvm_arch_init_vm(struct kvm *kvm, unsigned long type)
* If any failures occur after this is successful, make sure to
* call __pkvm_unreserve_vm to unreserve the VM in hyp.
*/
- ret = pkvm_init_host_vm(kvm, type);
+ ret = pkvm_init_host_vm(kvm);
if (ret)
goto err_uninit_mmu;
- } else if (type & KVM_VM_TYPE_ARM_PROTECTED) {
- ret = -EINVAL;
- goto err_uninit_mmu;
}
kvm_vgic_early_init(kvm);
@@ -985,7 +1006,7 @@ int kvm_arch_vcpu_run_pid_change(struct kvm_vcpu *vcpu)
if (is_protected_kvm_enabled()) {
/* Start with the vcpu in a dirty state */
- if (!kvm_vm_is_protected(vcpu->kvm))
+ if (kvm_vm_is_unprotected_pkvm(vcpu->kvm))
vcpu_set_flag(vcpu, PKVM_HOST_STATE_DIRTY);
ret = pkvm_create_hyp_vm(kvm);
if (ret)
diff --git a/arch/arm64/kvm/handle_exit.c b/arch/arm64/kvm/handle_exit.c
index db37678dcb05c..384c5d258c7f8 100644
--- a/arch/arm64/kvm/handle_exit.c
+++ b/arch/arm64/kvm/handle_exit.c
@@ -490,7 +490,7 @@ static void handle_exit_pkvm_state(struct kvm_vcpu *vcpu, int exception_index)
{
int exception_code = ARM_EXCEPTION_CODE(exception_index);
- if (!is_protected_kvm_enabled() || kvm_vm_is_protected(vcpu->kvm))
+ if (!kvm_vm_is_unprotected_pkvm(vcpu->kvm))
return;
/*
diff --git a/arch/arm64/kvm/hyp/nvhe/pkvm.c b/arch/arm64/kvm/hyp/nvhe/pkvm.c
index 459bd9eb7e4bc..57e2eef6d7426 100644
--- a/arch/arm64/kvm/hyp/nvhe/pkvm.c
+++ b/arch/arm64/kvm/hyp/nvhe/pkvm.c
@@ -432,7 +432,11 @@ static void init_pkvm_hyp_vm(struct kvm *host_kvm, struct pkvm_hyp_vm *hyp_vm,
hyp_vm->host_kvm = host_kvm;
hyp_vm->kvm.created_vcpus = nr_vcpus;
- hyp_vm->kvm.arch.pkvm.is_protected = READ_ONCE(host_kvm->arch.pkvm.is_protected);
+ if (kvm_vm_is_protected(host_kvm))
+ hyp_vm->kvm.arch.vm_flavor = VM_PROTECTED_PKVM;
+ else
+ hyp_vm->kvm.arch.vm_flavor = VM_PKVM;
+
hyp_vm->kvm.arch.flags = 0;
pkvm_init_features_from_host(hyp_vm, host_kvm);
diff --git a/arch/arm64/kvm/mmu.c b/arch/arm64/kvm/mmu.c
index 9ba86450fe4af..0f4e8b71fa85d 100644
--- a/arch/arm64/kvm/mmu.c
+++ b/arch/arm64/kvm/mmu.c
@@ -2624,7 +2624,7 @@ int kvm_arch_prepare_memory_region(struct kvm *kvm,
hva_t hva, reg_end;
int ret = 0;
- if (kvm_vm_is_protected(kvm)) {
+ if (kvm_vm_is_protected_pkvm(kvm)) {
/* Cannot modify memslots once a pVM has run. */
if (pkvm_hyp_vm_is_created(kvm) &&
(change == KVM_MR_DELETE || change == KVM_MR_MOVE)) {
diff --git a/arch/arm64/kvm/pkvm.c b/arch/arm64/kvm/pkvm.c
index 8e4c6e4bec123..8e9176a700926 100644
--- a/arch/arm64/kvm/pkvm.c
+++ b/arch/arm64/kvm/pkvm.c
@@ -229,10 +229,9 @@ void pkvm_destroy_hyp_vm(struct kvm *kvm)
mutex_unlock(&kvm->arch.config_lock);
}
-int pkvm_init_host_vm(struct kvm *kvm, unsigned long type)
+int pkvm_init_host_vm(struct kvm *kvm)
{
int ret;
- bool protected = type & KVM_VM_TYPE_ARM_PROTECTED;
/* Reserve the VM in hyp and obtain a hyp handle for the VM. */
ret = kvm_call_hyp_nvhe(__pkvm_reserve_vm);
@@ -240,8 +239,7 @@ int pkvm_init_host_vm(struct kvm *kvm, unsigned long type)
return ret;
kvm->arch.pkvm.handle = ret;
- kvm->arch.pkvm.is_protected = protected;
- if (protected) {
+ if (kvm_vm_is_protected(kvm)) {
pr_warn_once("kvm: protected VMs are experimental and for development only, tainting kernel\n");
add_taint(TAINT_USER, LOCKDEP_STILL_OK);
}
--
2.43.0
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v19 06/20] KVM: arm64: Refactor the vcpu_load to allow for VM specific callbacks
2026-09-20 21:28 [PATCH v19 00/20] KVM: arm64: CCA: Add basic plumbing for Realms Suzuki K Poulose
` (4 preceding siblings ...)
2026-09-20 21:28 ` [PATCH v19 05/20] KVM: arm64: Track the type of VM in kvm_arch Suzuki K Poulose
@ 2026-09-20 21:28 ` Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 07/20] KVM: arm64: Add vcpu load/put call backs for flavors Suzuki K Poulose
` (13 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: Suzuki K Poulose @ 2026-09-20 21:28 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose
To keep the VCPU load/put handling cleaner with the different kinds of VM
types, we are about to introduce VM specific callbacks to do just the right
thing. In preparation for that, make some refactoring to add the change
easier.
No functional changes intended. Based on a work by Marc Zyngier.
Reviewed-by: Gavin Shan <gshan@redhat.com>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
arch/arm64/kvm/arm.c | 48 ++++++++++++++++++++++++++------------------
1 file changed, 29 insertions(+), 19 deletions(-)
diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c
index 8c784b266a8e8..c74706ed9a531 100644
--- a/arch/arm64/kvm/arm.c
+++ b/arch/arm64/kvm/arm.c
@@ -683,14 +683,11 @@ static bool kvm_vcpu_should_clear_twe(struct kvm_vcpu *vcpu)
return single_task_running();
}
-void kvm_arch_vcpu_load(struct kvm_vcpu *vcpu, int cpu)
+static void vcpu_prepare_mmu(struct kvm_vcpu *vcpu)
{
struct kvm_s2_mmu *mmu;
int *last_ran;
- if (is_protected_kvm_enabled())
- goto nommu;
-
if (vcpu_has_nv(vcpu))
kvm_vcpu_load_hw_mmu(vcpu);
@@ -720,10 +717,33 @@ void kvm_arch_vcpu_load(struct kvm_vcpu *vcpu, int cpu)
kvm_call_hyp(__kvm_flush_cpu_context, mmu);
*last_ran = vcpu->vcpu_idx;
}
+}
+
+static void vcpu_set_wfx_traps(struct kvm_vcpu *vcpu)
+{
+ if (kvm_vcpu_should_clear_twe(vcpu))
+ vcpu->arch.hcr_el2 &= ~HCR_TWE;
+ else
+ vcpu->arch.hcr_el2 |= HCR_TWE;
+
+ if (kvm_vcpu_should_clear_twi(vcpu))
+ vcpu->arch.hcr_el2 &= ~HCR_TWI;
+ else
+ vcpu->arch.hcr_el2 |= HCR_TWI;
+}
+
+static void vcpu_load_pvtime(struct kvm_vcpu *vcpu)
+{
+ if (kvm_arm_is_pvtime_enabled(&vcpu->arch))
+ kvm_make_request(KVM_REQ_RECORD_STEAL, vcpu);
+}
+
+void kvm_arch_vcpu_load(struct kvm_vcpu *vcpu, int cpu)
+{
+ if (!is_protected_kvm_enabled())
+ vcpu_prepare_mmu(vcpu);
-nommu:
vcpu->cpu = cpu;
-
/*
* The timer must be loaded before the vgic to correctly set up physical
* interrupt deactivation in nested state (e.g. timer interrupt).
@@ -736,19 +756,9 @@ void kvm_arch_vcpu_load(struct kvm_vcpu *vcpu, int cpu)
kvm_vcpu_load_vhe(vcpu);
kvm_arch_vcpu_load_fp(vcpu);
kvm_vcpu_pmu_restore_guest(vcpu);
- if (kvm_arm_is_pvtime_enabled(&vcpu->arch))
- kvm_make_request(KVM_REQ_RECORD_STEAL, vcpu);
-
- if (kvm_vcpu_should_clear_twe(vcpu))
- vcpu->arch.hcr_el2 &= ~HCR_TWE;
- else
- vcpu->arch.hcr_el2 |= HCR_TWE;
-
- if (kvm_vcpu_should_clear_twi(vcpu))
- vcpu->arch.hcr_el2 &= ~HCR_TWI;
- else
- vcpu->arch.hcr_el2 |= HCR_TWI;
+ vcpu_load_pvtime(vcpu);
+ vcpu_set_wfx_traps(vcpu);
vcpu_set_pauth_traps(vcpu);
if (is_protected_kvm_enabled()) {
@@ -772,7 +782,7 @@ void kvm_arch_vcpu_put(struct kvm_vcpu *vcpu)
kvm_call_hyp_nvhe(__pkvm_vcpu_put);
/* __pkvm_vcpu_put implies a sync of the state */
- if (!kvm_vm_is_protected(vcpu->kvm))
+ if (kvm_vm_is_unprotected_pkvm(vcpu->kvm))
vcpu_set_flag(vcpu, PKVM_HOST_STATE_DIRTY);
}
--
2.43.0
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v19 07/20] KVM: arm64: Add vcpu load/put call backs for flavors
2026-09-20 21:28 [PATCH v19 00/20] KVM: arm64: CCA: Add basic plumbing for Realms Suzuki K Poulose
` (5 preceding siblings ...)
2026-09-20 21:28 ` [PATCH v19 06/20] KVM: arm64: Refactor the vcpu_load to allow for VM specific callbacks Suzuki K Poulose
@ 2026-09-20 21:28 ` Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 08/20] KVM: arm64: Reuse kvm_stage2_unmap_range in kvm_unmap_gfn_range Suzuki K Poulose
` (12 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: Suzuki K Poulose @ 2026-09-20 21:28 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose
Add VM flavor specific handlers for VCPU load/put, in an effort to make it
easier to follow the code. pauth traps were removed from VMs running PKVM
as it is a no-op for them.
Based on a patch by Marc Zyngier
Suggested-by: Marc Zyngier <maz@kernel.org>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Changes since v17:
- Use macro to initialize the per-flavor ops
- Add a wrapper to initialise ops in the vcpu structure.
- Add BUILD_BUG_ON for the array size
- Remove irrelevant comment about the order of timer loading for !VHE
- Use the explicti kvm_call_hyp_nvhe for nVHE flavor
- Don't call nvhe_vcpu_put from pkvm_vcpu_put, open code them
- Drop cpu argument for vcpu_load() callback. We set the cpu
before the callbacks are invoked
---
arch/arm64/include/asm/kvm_host.h | 6 ++
arch/arm64/kvm/arm.c | 134 ++++++++++++++++++++++++------
2 files changed, 114 insertions(+), 26 deletions(-)
diff --git a/arch/arm64/include/asm/kvm_host.h b/arch/arm64/include/asm/kvm_host.h
index 9b1cf9c59e81f..149f4582c8b6a 100644
--- a/arch/arm64/include/asm/kvm_host.h
+++ b/arch/arm64/include/asm/kvm_host.h
@@ -150,6 +150,11 @@ struct kvm_vmid {
atomic64_t id;
};
+struct kvm_vcpu_ops {
+ void (*vcpu_load)(struct kvm_vcpu *vcpu);
+ void (*vcpu_put)(struct kvm_vcpu *vcpu);
+};
+
struct kvm_s2_mmu {
struct kvm_vmid vmid;
@@ -855,6 +860,7 @@ struct vncr_tlb;
struct kvm_vcpu_arch {
struct kvm_cpu_context ctxt;
+ const struct kvm_vcpu_ops *vcpu_ops;
/*
* Guest floating point state
diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c
index c74706ed9a531..9b977dc734220 100644
--- a/arch/arm64/kvm/arm.c
+++ b/arch/arm64/kvm/arm.c
@@ -93,6 +93,7 @@ static const struct kvm_ioctl_cap_map vm_ioctl_caps[] = {
{ KVM_ARM_PREFERRED_TARGET, KVM_CAP_ARM_BASIC },
};
+static void kvm_init_vcpu_ops(struct kvm_vcpu *vcpu);
/*
* Set *ext to the capability.
* Return 0 if found, or -EINVAL if no IOCTL matches.
@@ -569,6 +570,8 @@ int kvm_arch_vcpu_create(struct kvm_vcpu *vcpu)
mutex_unlock(&vcpu->mutex);
#endif
+ kvm_init_vcpu_ops(vcpu);
+
/* Force users to call KVM_ARM_VCPU_INIT */
vcpu_clear_flag(vcpu, VCPU_INITIALIZED);
@@ -738,12 +741,9 @@ static void vcpu_load_pvtime(struct kvm_vcpu *vcpu)
kvm_make_request(KVM_REQ_RECORD_STEAL, vcpu);
}
-void kvm_arch_vcpu_load(struct kvm_vcpu *vcpu, int cpu)
+static void vhe_vcpu_load(struct kvm_vcpu *vcpu)
{
- if (!is_protected_kvm_enabled())
- vcpu_prepare_mmu(vcpu);
-
- vcpu->cpu = cpu;
+ vcpu_prepare_mmu(vcpu);
/*
* The timer must be loaded before the vgic to correctly set up physical
* interrupt deactivation in nested state (e.g. timer interrupt).
@@ -752,22 +752,53 @@ void kvm_arch_vcpu_load(struct kvm_vcpu *vcpu, int cpu)
kvm_vgic_load(vcpu);
kvm_vcpu_load_debug(vcpu);
kvm_vcpu_load_fgt(vcpu);
- if (has_vhe())
- kvm_vcpu_load_vhe(vcpu);
+ kvm_vcpu_load_vhe(vcpu);
kvm_arch_vcpu_load_fp(vcpu);
kvm_vcpu_pmu_restore_guest(vcpu);
vcpu_load_pvtime(vcpu);
vcpu_set_wfx_traps(vcpu);
vcpu_set_pauth_traps(vcpu);
+}
- if (is_protected_kvm_enabled()) {
- kvm_call_hyp_nvhe(__pkvm_vcpu_load,
- vcpu->kvm->arch.pkvm.handle,
- vcpu->vcpu_idx, vcpu->arch.hcr_el2);
- kvm_call_hyp(__vgic_v3_restore_vmcr_aprs,
- &vcpu->arch.vgic_cpu.vgic_v3);
- }
+static void nvhe_vcpu_load(struct kvm_vcpu *vcpu)
+{
+ vcpu_prepare_mmu(vcpu);
+ kvm_timer_vcpu_load(vcpu);
+ kvm_vgic_load(vcpu);
+ kvm_vcpu_load_debug(vcpu);
+ kvm_vcpu_load_fgt(vcpu);
+ kvm_arch_vcpu_load_fp(vcpu);
+ kvm_vcpu_pmu_restore_guest(vcpu);
+
+ vcpu_load_pvtime(vcpu);
+ vcpu_set_wfx_traps(vcpu);
+ vcpu_set_pauth_traps(vcpu);
+}
+
+static void pkvm_vcpu_load(struct kvm_vcpu *vcpu)
+{
+ kvm_timer_vcpu_load(vcpu);
+ kvm_vgic_load(vcpu);
+ kvm_vcpu_load_debug(vcpu);
+ kvm_vcpu_load_fgt(vcpu);
+ kvm_arch_vcpu_load_fp(vcpu);
+ kvm_vcpu_pmu_restore_guest(vcpu);
+
+ vcpu_load_pvtime(vcpu);
+ vcpu_set_wfx_traps(vcpu);
+
+ kvm_call_hyp_nvhe(__pkvm_vcpu_load,
+ vcpu->kvm->arch.pkvm.handle,
+ vcpu->vcpu_idx, vcpu->arch.hcr_el2);
+ kvm_call_hyp_nvhe(__vgic_v3_restore_vmcr_aprs,
+ &vcpu->arch.vgic_cpu.vgic_v3);
+}
+
+void kvm_arch_vcpu_load(struct kvm_vcpu *vcpu, int cpu)
+{
+ vcpu->cpu = cpu;
+ vcpu->arch.vcpu_ops->vcpu_load(vcpu);
if (!cpumask_test_cpu(cpu, vcpu->kvm->arch.supported_cpus))
vcpu_set_on_unsupported_cpu(vcpu);
@@ -775,28 +806,48 @@ void kvm_arch_vcpu_load(struct kvm_vcpu *vcpu, int cpu)
vcpu->arch.pid = pid_nr(vcpu->pid);
}
-void kvm_arch_vcpu_put(struct kvm_vcpu *vcpu)
+static void vhe_vcpu_put(struct kvm_vcpu *vcpu)
{
- if (is_protected_kvm_enabled()) {
- kvm_call_hyp(__vgic_v3_save_aprs, &vcpu->arch.vgic_cpu.vgic_v3);
- kvm_call_hyp_nvhe(__pkvm_vcpu_put);
-
- /* __pkvm_vcpu_put implies a sync of the state */
- if (kvm_vm_is_unprotected_pkvm(vcpu->kvm))
- vcpu_set_flag(vcpu, PKVM_HOST_STATE_DIRTY);
- }
-
kvm_vcpu_put_debug(vcpu);
kvm_arch_vcpu_put_fp(vcpu);
- if (has_vhe())
- kvm_vcpu_put_vhe(vcpu);
+ kvm_vcpu_put_vhe(vcpu);
kvm_timer_vcpu_put(vcpu);
kvm_vgic_put(vcpu);
kvm_vcpu_pmu_restore_host(vcpu);
if (vcpu_has_nv(vcpu))
kvm_vcpu_put_hw_mmu(vcpu);
kvm_arm_vmid_clear_active();
+}
+static void nvhe_vcpu_put(struct kvm_vcpu *vcpu)
+{
+ kvm_vcpu_put_debug(vcpu);
+ kvm_arch_vcpu_put_fp(vcpu);
+ kvm_timer_vcpu_put(vcpu);
+ kvm_vgic_put(vcpu);
+ kvm_vcpu_pmu_restore_host(vcpu);
+ kvm_arm_vmid_clear_active();
+}
+
+static void pkvm_vcpu_put(struct kvm_vcpu *vcpu)
+{
+ kvm_call_hyp_nvhe(__vgic_v3_save_aprs, &vcpu->arch.vgic_cpu.vgic_v3);
+ kvm_call_hyp_nvhe(__pkvm_vcpu_put);
+
+ /* __pkvm_vcpu_put implies a sync of the state */
+ if (kvm_vm_is_unprotected_pkvm(vcpu->kvm))
+ vcpu_set_flag(vcpu, PKVM_HOST_STATE_DIRTY);
+
+ kvm_vcpu_put_debug(vcpu);
+ kvm_arch_vcpu_put_fp(vcpu);
+ kvm_timer_vcpu_put(vcpu);
+ kvm_vgic_put(vcpu);
+ kvm_vcpu_pmu_restore_host(vcpu);
+}
+
+void kvm_arch_vcpu_put(struct kvm_vcpu *vcpu)
+{
+ vcpu->arch.vcpu_ops->vcpu_put(vcpu);
vcpu_clear_on_unsupported_cpu(vcpu);
vcpu->cpu = -1;
}
@@ -2136,6 +2187,37 @@ int kvm_arch_vm_ioctl(struct file *filp, unsigned int ioctl, unsigned long arg)
}
}
+static const struct kvm_vcpu_ops vhe_vcpu_ops = {
+ .vcpu_load = vhe_vcpu_load,
+ .vcpu_put = vhe_vcpu_put,
+};
+
+static const struct kvm_vcpu_ops nvhe_vcpu_ops = {
+ .vcpu_load = nvhe_vcpu_load,
+ .vcpu_put = nvhe_vcpu_put,
+};
+
+static const struct kvm_vcpu_ops pkvm_vcpu_ops = {
+ .vcpu_load = pkvm_vcpu_load,
+ .vcpu_put = pkvm_vcpu_put,
+};
+
+#define KVM_VCPU_OPS(flavor, ops) \
+ [(flavor)] = (ops)
+
+static const struct kvm_vcpu_ops *arm64_vcpu_ops[] = {
+ KVM_VCPU_OPS(VM_VHE, &vhe_vcpu_ops),
+ KVM_VCPU_OPS(VM_NVHE, &nvhe_vcpu_ops),
+ KVM_VCPU_OPS(VM_PKVM, &pkvm_vcpu_ops),
+ KVM_VCPU_OPS(VM_PROTECTED_PKVM, &pkvm_vcpu_ops),
+};
+
+static void kvm_init_vcpu_ops(struct kvm_vcpu *vcpu)
+{
+ BUILD_BUG_ON(ARRAY_SIZE(arm64_vcpu_ops) != VM_FLAVOR_MAX);
+ vcpu->arch.vcpu_ops = arm64_vcpu_ops[vcpu->kvm->arch.vm_flavor];
+}
+
static unsigned long nvhe_percpu_size(void)
{
return (unsigned long)CHOOSE_NVHE_SYM(__per_cpu_end) -
--
2.43.0
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v19 08/20] KVM: arm64: Reuse kvm_stage2_unmap_range in kvm_unmap_gfn_range
2026-09-20 21:28 [PATCH v19 00/20] KVM: arm64: CCA: Add basic plumbing for Realms Suzuki K Poulose
` (6 preceding siblings ...)
2026-09-20 21:28 ` [PATCH v19 07/20] KVM: arm64: Add vcpu load/put call backs for flavors Suzuki K Poulose
@ 2026-09-20 21:28 ` Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 09/20] KVM: arm64: Add VM specific callback for S2 MMU operations Suzuki K Poulose
` (11 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: Suzuki K Poulose @ 2026-09-20 21:28 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose,
Fuad Tabba
In preparation for adding VM specific backends for stage2 operation,
switch to kvm_stage2_unmap_range() instead of __unmap_stage2_range()
from the kvm_unmap_gfn_range(). Drop the bail out check for protected
VMs and defer that to the one in kvm_stage2_unmap_range(). Later we
would replace the logic in kvm_stage2_unmap_range() with VM specific
backends.
No functional changes intended.
Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
arch/arm64/kvm/mmu.c | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
diff --git a/arch/arm64/kvm/mmu.c b/arch/arm64/kvm/mmu.c
index 0f4e8b71fa85d..03f2017a7404a 100644
--- a/arch/arm64/kvm/mmu.c
+++ b/arch/arm64/kvm/mmu.c
@@ -2436,12 +2436,12 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
bool kvm_unmap_gfn_range(struct kvm *kvm, struct kvm_gfn_range *range)
{
- if (!kvm->arch.mmu.pgt || kvm_vm_is_protected(kvm))
+ if (!kvm->arch.mmu.pgt)
return false;
- __unmap_stage2_range(&kvm->arch.mmu, range->start << PAGE_SHIFT,
- (range->end - range->start) << PAGE_SHIFT,
- range->may_block);
+ kvm_stage2_unmap_range(&kvm->arch.mmu, range->start << PAGE_SHIFT,
+ (range->end - range->start) << PAGE_SHIFT,
+ range->may_block);
kvm_nested_s2_unmap(kvm, range->may_block);
return false;
--
2.43.0
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v19 09/20] KVM: arm64: Add VM specific callback for S2 MMU operations
2026-09-20 21:28 [PATCH v19 00/20] KVM: arm64: CCA: Add basic plumbing for Realms Suzuki K Poulose
` (7 preceding siblings ...)
2026-09-20 21:28 ` [PATCH v19 08/20] KVM: arm64: Reuse kvm_stage2_unmap_range in kvm_unmap_gfn_range Suzuki K Poulose
@ 2026-09-20 21:28 ` Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 10/20] KVM: arm64: Abstract out memory abort handling Suzuki K Poulose
` (10 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: Suzuki K Poulose @ 2026-09-20 21:28 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose
Add VM type specific S2 MMU operation backends which can be initialized per
VM flavor, to keep the handling cleaner.
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
arch/arm64/include/asm/kvm_host.h | 15 ++++
arch/arm64/kvm/mmu.c | 137 +++++++++++++++++++++++++-----
2 files changed, 131 insertions(+), 21 deletions(-)
diff --git a/arch/arm64/include/asm/kvm_host.h b/arch/arm64/include/asm/kvm_host.h
index 149f4582c8b6a..7664d8b8cce5a 100644
--- a/arch/arm64/include/asm/kvm_host.h
+++ b/arch/arm64/include/asm/kvm_host.h
@@ -155,6 +155,19 @@ struct kvm_vcpu_ops {
void (*vcpu_put)(struct kvm_vcpu *vcpu);
};
+struct kvm_gfn_range;
+
+struct kvm_vm_s2_ops {
+ bool (*vm_age_gfn)(struct kvm *kvm, struct kvm_gfn_range *range);
+ bool (*vm_test_age_gfn)(struct kvm *kvm, struct kvm_gfn_range *range);
+ int (*vm_flush_remote_tlbs)(struct kvm *kvm);
+ int (*vm_flush_remote_tlbs_range)(struct kvm *kvm, gfn_t gfn,
+ u64 nr_pages);
+ void (*vm_stage2_unmap_range)(struct kvm_s2_mmu *mmu,
+ phys_addr_t start, u64 size,
+ bool may_block);
+};
+
struct kvm_s2_mmu {
struct kvm_vmid vmid;
@@ -332,6 +345,8 @@ struct kvm_arch {
*/
u64 fgu[__NR_FGT_GROUP_IDS__];
+ const struct kvm_vm_s2_ops *vm_s2_ops;
+
/*
* Stage 2 paging state for VMs with nested S2 using a virtual
* VMID.
diff --git a/arch/arm64/kvm/mmu.c b/arch/arm64/kvm/mmu.c
index 03f2017a7404a..d97a4a1bca23f 100644
--- a/arch/arm64/kvm/mmu.c
+++ b/arch/arm64/kvm/mmu.c
@@ -37,6 +37,8 @@ static unsigned long __ro_after_init io_map_base;
#define KVM_PGT_FN(fn) (!is_protected_kvm_enabled() ? fn : p ## fn)
+static int kvm_vm_init_vm_s2_ops(struct kvm *kvm);
+
static phys_addr_t __stage2_range_addr_end(phys_addr_t addr, phys_addr_t end,
phys_addr_t size)
{
@@ -166,6 +168,18 @@ static bool memslot_is_logging(struct kvm_memory_slot *memslot)
return memslot->dirty_bitmap && !(memslot->flags & KVM_MEM_READONLY);
}
+static int pkvm_flush_remote_tlbs(struct kvm *kvm)
+{
+ kvm_call_hyp_nvhe(__pkvm_tlb_flush_vmid, kvm->arch.pkvm.handle);
+ return 0;
+}
+
+static int kvm_vm_flush_remote_tlbs(struct kvm *kvm)
+{
+ kvm_call_hyp(__kvm_tlb_flush_vmid, &kvm->arch.mmu);
+ return 0;
+}
+
/**
* kvm_arch_flush_remote_tlbs() - flush all VM TLB entries for v7/8
* @kvm: pointer to kvm structure.
@@ -174,26 +188,36 @@ static bool memslot_is_logging(struct kvm_memory_slot *memslot)
*/
int kvm_arch_flush_remote_tlbs(struct kvm *kvm)
{
- if (is_protected_kvm_enabled())
- kvm_call_hyp_nvhe(__pkvm_tlb_flush_vmid, kvm->arch.pkvm.handle);
- else
- kvm_call_hyp(__kvm_tlb_flush_vmid, &kvm->arch.mmu);
- return 0;
+ if (!kvm->arch.vm_s2_ops->vm_flush_remote_tlbs)
+ return 1;
+ return kvm->arch.vm_s2_ops->vm_flush_remote_tlbs(kvm);
}
-int kvm_arch_flush_remote_tlbs_range(struct kvm *kvm,
- gfn_t gfn, u64 nr_pages)
+static int pkvm_flush_remote_tlbs_range(struct kvm *kvm,
+ gfn_t gfn, u64 nr_pages)
+{
+ return pkvm_flush_remote_tlbs(kvm);
+}
+
+static int kvm_vm_flush_remote_tlbs_range(struct kvm *kvm,
+ gfn_t gfn, u64 nr_pages)
{
u64 size = nr_pages << PAGE_SHIFT;
u64 addr = gfn << PAGE_SHIFT;
- if (is_protected_kvm_enabled())
- kvm_call_hyp_nvhe(__pkvm_tlb_flush_vmid, kvm->arch.pkvm.handle);
- else
- kvm_tlb_flush_vmid_range(&kvm->arch.mmu, addr, size);
+ kvm_tlb_flush_vmid_range(&kvm->arch.mmu, addr, size);
return 0;
}
+int kvm_arch_flush_remote_tlbs_range(struct kvm *kvm,
+ gfn_t gfn, u64 nr_pages)
+{
+ if (!kvm->arch.vm_s2_ops->vm_flush_remote_tlbs_range)
+ return 1;
+
+ return kvm->arch.vm_s2_ops->vm_flush_remote_tlbs_range(kvm, gfn, nr_pages);
+}
+
static void *stage2_memcache_zalloc_page(void *arg)
{
struct kvm_mmu_memory_cache *mc = arg;
@@ -337,13 +361,20 @@ static void __unmap_stage2_range(struct kvm_s2_mmu *mmu, phys_addr_t start, u64
may_block));
}
+static void kvm_vm_stage2_unmap_range(struct kvm_s2_mmu *mmu,
+ phys_addr_t start,
+ u64 size, bool may_block)
+{
+ __unmap_stage2_range(mmu, start, size, may_block);
+}
+
void kvm_stage2_unmap_range(struct kvm_s2_mmu *mmu, phys_addr_t start,
u64 size, bool may_block)
{
- if (kvm_vm_is_protected(kvm_s2_mmu_to_kvm(mmu)))
- return;
+ struct kvm *kvm = kvm_s2_mmu_to_kvm(mmu);
- __unmap_stage2_range(mmu, start, size, may_block);
+ if (kvm->arch.vm_s2_ops->vm_stage2_unmap_range)
+ kvm->arch.vm_s2_ops->vm_stage2_unmap_range(mmu, start, size, may_block);
}
void kvm_stage2_flush_range(struct kvm_s2_mmu *mmu, phys_addr_t addr, phys_addr_t end)
@@ -983,6 +1014,12 @@ int kvm_init_stage2_mmu(struct kvm *kvm, struct kvm_s2_mmu *mmu, unsigned long t
int cpu, err;
struct kvm_pgtable *pgt;
+ /* Initialize the VM ops for the VM instance for the first time */
+ if (mmu == &kvm->arch.mmu) {
+ err = kvm_vm_init_vm_s2_ops(kvm);
+ if (err)
+ return err;
+ }
/*
* If we already have our page tables in place, and that the
* MMU context is the canonical one, we have a bug somewhere,
@@ -2447,34 +2484,46 @@ bool kvm_unmap_gfn_range(struct kvm *kvm, struct kvm_gfn_range *range)
return false;
}
-bool kvm_age_gfn(struct kvm *kvm, struct kvm_gfn_range *range)
+static bool kvm_vm_age_gfn(struct kvm *kvm, struct kvm_gfn_range *range)
{
u64 size = (range->end - range->start) << PAGE_SHIFT;
- if (!kvm->arch.mmu.pgt || kvm_vm_is_protected(kvm))
- return false;
-
return KVM_PGT_FN(kvm_pgtable_stage2_test_clear_young)(kvm->arch.mmu.pgt,
range->start << PAGE_SHIFT,
size, true);
+}
+
+bool kvm_age_gfn(struct kvm *kvm, struct kvm_gfn_range *range)
+{
+ if (!kvm->arch.mmu.pgt || !kvm->arch.vm_s2_ops->vm_age_gfn)
+ return false;
+
+ return kvm->arch.vm_s2_ops->vm_age_gfn(kvm, range);
/*
* TODO: Handle nested_mmu structures here using the reverse mapping in
* a later version of patch series.
*/
}
-bool kvm_test_age_gfn(struct kvm *kvm, struct kvm_gfn_range *range)
+static bool kvm_vm_test_age_gfn(struct kvm *kvm, struct kvm_gfn_range *range)
{
u64 size = (range->end - range->start) << PAGE_SHIFT;
- if (!kvm->arch.mmu.pgt || kvm_vm_is_protected(kvm))
- return false;
return KVM_PGT_FN(kvm_pgtable_stage2_test_clear_young)(kvm->arch.mmu.pgt,
range->start << PAGE_SHIFT,
size, false);
}
+bool kvm_test_age_gfn(struct kvm *kvm, struct kvm_gfn_range *range)
+{
+
+ if (!kvm->arch.mmu.pgt || !kvm->arch.vm_s2_ops->vm_test_age_gfn)
+ return false;
+
+ return kvm->arch.vm_s2_ops->vm_test_age_gfn(kvm, range);
+}
+
phys_addr_t kvm_mmu_get_httbr(void)
{
return __pa(hyp_pgtable->pgd);
@@ -2796,3 +2845,49 @@ void kvm_toggle_cache(struct kvm_vcpu *vcpu, bool was_enabled)
trace_kvm_toggle_cache(*vcpu_pc(vcpu), was_enabled, now_enabled);
}
+
+static const struct kvm_vm_s2_ops protected_pkvm_vm_s2_ops = {
+ .vm_flush_remote_tlbs = pkvm_flush_remote_tlbs,
+ .vm_flush_remote_tlbs_range = pkvm_flush_remote_tlbs_range,
+ /*
+ * Not supported for Protected VMs under pKVM
+ * .vm_age_gfn
+ * .vm_test_age_gfn
+ * .vm_stage2_unmap_range
+ */
+};
+
+static const struct kvm_vm_s2_ops pkvm_vm_s2_ops = {
+ .vm_flush_remote_tlbs = pkvm_flush_remote_tlbs,
+ .vm_flush_remote_tlbs_range = pkvm_flush_remote_tlbs_range,
+ .vm_age_gfn = kvm_vm_age_gfn,
+ .vm_test_age_gfn = kvm_vm_test_age_gfn,
+ .vm_stage2_unmap_range = kvm_vm_stage2_unmap_range,
+};
+
+static const struct kvm_vm_s2_ops kvm_default_vm_s2_ops = {
+ .vm_flush_remote_tlbs = kvm_vm_flush_remote_tlbs,
+ .vm_flush_remote_tlbs_range = kvm_vm_flush_remote_tlbs_range,
+ .vm_age_gfn = kvm_vm_age_gfn,
+ .vm_test_age_gfn = kvm_vm_test_age_gfn,
+ .vm_stage2_unmap_range = kvm_vm_stage2_unmap_range,
+};
+
+#define KVM_VM_S2_OPS(flavor, ops) \
+ [flavor] = ops
+static const struct kvm_vm_s2_ops *arm64_vm_s2_ops[] = {
+ KVM_VM_S2_OPS(VM_VHE, &kvm_default_vm_s2_ops),
+ KVM_VM_S2_OPS(VM_NVHE, &kvm_default_vm_s2_ops),
+ KVM_VM_S2_OPS(VM_PKVM, &pkvm_vm_s2_ops),
+ KVM_VM_S2_OPS(VM_PROTECTED_PKVM, &protected_pkvm_vm_s2_ops),
+};
+
+static int kvm_vm_init_vm_s2_ops(struct kvm *kvm)
+{
+ BUILD_BUG_ON(ARRAY_SIZE(arm64_vm_s2_ops) != VM_FLAVOR_MAX);
+
+ kvm->arch.vm_s2_ops = arm64_vm_s2_ops[kvm->arch.vm_flavor];
+ if (WARN_ON(!kvm->arch.vm_s2_ops))
+ return -EINVAL;
+ return 0;
+}
--
2.43.0
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v19 10/20] KVM: arm64: Abstract out memory abort handling
2026-09-20 21:28 [PATCH v19 00/20] KVM: arm64: CCA: Add basic plumbing for Realms Suzuki K Poulose
` (8 preceding siblings ...)
2026-09-20 21:28 ` [PATCH v19 09/20] KVM: arm64: Add VM specific callback for S2 MMU operations Suzuki K Poulose
@ 2026-09-20 21:28 ` Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 11/20] KVM: arm64: Mandate VGIC v3 for for VMs running on hyp that don't trust the host Suzuki K Poulose
` (9 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: Suzuki K Poulose @ 2026-09-20 21:28 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose
Move the memory abort handling under VM specific s2 operation.
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Changes since v18:
- Ensure vm_mem_abort handler is always !NULL
---
arch/arm64/include/asm/kvm_host.h | 2 ++
arch/arm64/kvm/mmu.c | 38 +++++++++++++++++++------------
2 files changed, 25 insertions(+), 15 deletions(-)
diff --git a/arch/arm64/include/asm/kvm_host.h b/arch/arm64/include/asm/kvm_host.h
index 7664d8b8cce5a..3211543a85b38 100644
--- a/arch/arm64/include/asm/kvm_host.h
+++ b/arch/arm64/include/asm/kvm_host.h
@@ -156,6 +156,7 @@ struct kvm_vcpu_ops {
};
struct kvm_gfn_range;
+struct kvm_s2_fault_desc;
struct kvm_vm_s2_ops {
bool (*vm_age_gfn)(struct kvm *kvm, struct kvm_gfn_range *range);
@@ -166,6 +167,7 @@ struct kvm_vm_s2_ops {
void (*vm_stage2_unmap_range)(struct kvm_s2_mmu *mmu,
phys_addr_t start, u64 size,
bool may_block);
+ int (*vm_mem_abort)(const struct kvm_s2_fault_desc *s2fd);
};
struct kvm_s2_mmu {
diff --git a/arch/arm64/kvm/mmu.c b/arch/arm64/kvm/mmu.c
index d97a4a1bca23f..cf293d09e940a 100644
--- a/arch/arm64/kvm/mmu.c
+++ b/arch/arm64/kvm/mmu.c
@@ -1742,7 +1742,7 @@ struct kvm_s2_fault_vma_info {
bool map_non_cacheable;
};
-static int pkvm_mem_abort(const struct kvm_s2_fault_desc *s2fd)
+static int protected_vm_mem_abort(const struct kvm_s2_fault_desc *s2fd)
{
unsigned int flags = FOLL_HWPOISON | FOLL_LONGTERM | FOLL_WRITE;
struct kvm_vcpu *vcpu = s2fd->vcpu;
@@ -2180,6 +2180,22 @@ static int user_mem_abort(const struct kvm_s2_fault_desc *s2fd)
return kvm_s2_fault_map(s2fd, &s2vi, prot, memcache);
}
+static int kvm_vm_mem_abort(const struct kvm_s2_fault_desc *s2fd)
+{
+ int ret;
+ struct kvm_vcpu *vcpu = s2fd->vcpu;
+
+ VM_WARN_ON_ONCE(kvm_vcpu_trap_is_permission_fault(vcpu) &&
+ !kvm_is_write_fault(vcpu) &&
+ !kvm_vcpu_trap_is_exec_fault(vcpu));
+
+ if (kvm_slot_has_gmem(s2fd->memslot))
+ ret = gmem_abort(s2fd);
+ else
+ ret = user_mem_abort(s2fd);
+ return ret;
+}
+
/* Resolve the access fault by making the page young again. */
static void handle_access_fault(struct kvm_vcpu *vcpu, phys_addr_t fault_ipa)
{
@@ -2287,6 +2303,7 @@ int kvm_handle_guest_sea(struct kvm_vcpu *vcpu)
int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
{
struct kvm_s2_trans nested_trans, *nested = NULL;
+ struct kvm *kvm = vcpu->kvm;
unsigned long esr;
phys_addr_t fault_ipa; /* The address we faulted on */
phys_addr_t ipa; /* Always the IPA in the L1 guest phys space */
@@ -2448,19 +2465,7 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
.hva = hva,
};
- if (kvm_vm_is_protected(vcpu->kvm)) {
- ret = pkvm_mem_abort(&s2fd);
- } else {
- VM_WARN_ON_ONCE(kvm_vcpu_trap_is_permission_fault(vcpu) &&
- !write_fault &&
- !kvm_vcpu_trap_is_exec_fault(vcpu));
-
- if (kvm_slot_has_gmem(memslot))
- ret = gmem_abort(&s2fd);
- else
- ret = user_mem_abort(&s2fd);
- }
-
+ ret = kvm->arch.vm_s2_ops->vm_mem_abort(&s2fd);
if (ret == 0)
ret = 1;
out:
@@ -2855,6 +2860,7 @@ static const struct kvm_vm_s2_ops protected_pkvm_vm_s2_ops = {
* .vm_test_age_gfn
* .vm_stage2_unmap_range
*/
+ .vm_mem_abort = protected_vm_mem_abort,
};
static const struct kvm_vm_s2_ops pkvm_vm_s2_ops = {
@@ -2863,6 +2869,7 @@ static const struct kvm_vm_s2_ops pkvm_vm_s2_ops = {
.vm_age_gfn = kvm_vm_age_gfn,
.vm_test_age_gfn = kvm_vm_test_age_gfn,
.vm_stage2_unmap_range = kvm_vm_stage2_unmap_range,
+ .vm_mem_abort = kvm_vm_mem_abort,
};
static const struct kvm_vm_s2_ops kvm_default_vm_s2_ops = {
@@ -2871,6 +2878,7 @@ static const struct kvm_vm_s2_ops kvm_default_vm_s2_ops = {
.vm_age_gfn = kvm_vm_age_gfn,
.vm_test_age_gfn = kvm_vm_test_age_gfn,
.vm_stage2_unmap_range = kvm_vm_stage2_unmap_range,
+ .vm_mem_abort = kvm_vm_mem_abort,
};
#define KVM_VM_S2_OPS(flavor, ops) \
@@ -2887,7 +2895,7 @@ static int kvm_vm_init_vm_s2_ops(struct kvm *kvm)
BUILD_BUG_ON(ARRAY_SIZE(arm64_vm_s2_ops) != VM_FLAVOR_MAX);
kvm->arch.vm_s2_ops = arm64_vm_s2_ops[kvm->arch.vm_flavor];
- if (WARN_ON(!kvm->arch.vm_s2_ops))
+ if (WARN_ON(!kvm->arch.vm_s2_ops || !kvm->arch.vm_s2_ops->vm_mem_abort))
return -EINVAL;
return 0;
}
--
2.43.0
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v19 11/20] KVM: arm64: Mandate VGIC v3 for for VMs running on hyp that don't trust the host
2026-09-20 21:28 [PATCH v19 00/20] KVM: arm64: CCA: Add basic plumbing for Realms Suzuki K Poulose
` (9 preceding siblings ...)
2026-09-20 21:28 ` [PATCH v19 10/20] KVM: arm64: Abstract out memory abort handling Suzuki K Poulose
@ 2026-09-20 21:28 ` Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 12/20] KVM: arm64: CCA: Add a new mode for supporting Realm guests Suzuki K Poulose
` (8 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: Suzuki K Poulose @ 2026-09-20 21:28 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose
pKVM and RMM, both do not trust the host. Add a helper to detect the VMs
that have "distrusting" hyp. Use this for blocking ioremap of vgic-v2 into
stage2 and prevent creation of VGIC other than v3.
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Changes since v18:
- Cover pKVM guests for VGIC v3 mandate.
- Merge the VGIC mandate check in here.
---
arch/arm64/include/asm/kvm_host.h | 3 +++
arch/arm64/kvm/mmu.c | 2 +-
arch/arm64/kvm/vgic/vgic-init.c | 2 ++
3 files changed, 6 insertions(+), 1 deletion(-)
diff --git a/arch/arm64/include/asm/kvm_host.h b/arch/arm64/include/asm/kvm_host.h
index 3211543a85b38..bb156a633ab75 100644
--- a/arch/arm64/include/asm/kvm_host.h
+++ b/arch/arm64/include/asm/kvm_host.h
@@ -328,6 +328,8 @@ enum fgt_group_id {
enum kvm_arm_vm_flavor {
VM_NVHE,
VM_VHE,
+ /* VMs running on a hyp that doesn't trust */
+ MARKER(__VM_DISTRUSTING_HYP),
VM_PKVM, /* Normal guests on pKVM */
MARKER(__VM_PROTECTED),
VM_PROTECTED_PKVM, /* Protected VM */
@@ -1546,6 +1548,7 @@ struct kvm *kvm_arch_alloc_vm(void);
#define kvm_vm_is_unprotected_pkvm(kvm) \
(is_protected_kvm_enabled() && ((kvm)->arch.vm_flavor == VM_PKVM))
+#define kvm_vm_hyp_is_distrusting(kvm) ((kvm)->arch.vm_flavor >= __VM_DISTRUSTING_HYP)
int kvm_arm_vcpu_finalize(struct kvm_vcpu *vcpu, int feature);
bool kvm_arm_vcpu_is_finalized(struct kvm_vcpu *vcpu);
diff --git a/arch/arm64/kvm/mmu.c b/arch/arm64/kvm/mmu.c
index cf293d09e940a..413c4b114d75d 100644
--- a/arch/arm64/kvm/mmu.c
+++ b/arch/arm64/kvm/mmu.c
@@ -1251,7 +1251,7 @@ int kvm_phys_addr_ioremap(struct kvm *kvm, phys_addr_t guest_ipa,
KVM_PGTABLE_PROT_R |
(writable ? KVM_PGTABLE_PROT_W : 0);
- if (is_protected_kvm_enabled())
+ if (kvm_vm_hyp_is_distrusting(kvm))
return -EPERM;
size += offset_in_page(guest_ipa);
diff --git a/arch/arm64/kvm/vgic/vgic-init.c b/arch/arm64/kvm/vgic/vgic-init.c
index 4012df6002ea6..874025513afcc 100644
--- a/arch/arm64/kvm/vgic/vgic-init.c
+++ b/arch/arm64/kvm/vgic/vgic-init.c
@@ -84,6 +84,8 @@ int kvm_vgic_create(struct kvm *kvm, u32 type)
!kvm_vgic_global_state.can_emulate_gicv2)
return -ENODEV;
+ if (kvm_vm_hyp_is_distrusting(kvm) && type != KVM_DEV_TYPE_ARM_VGIC_V3)
+ return -ENODEV;
/*
* Ensure mutual exclusion with vCPU creation and any vCPU ioctls by:
*
--
2.43.0
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v19 12/20] KVM: arm64: CCA: Add a new mode for supporting Realm guests
2026-09-20 21:28 [PATCH v19 00/20] KVM: arm64: CCA: Add basic plumbing for Realms Suzuki K Poulose
` (10 preceding siblings ...)
2026-09-20 21:28 ` [PATCH v19 11/20] KVM: arm64: Mandate VGIC v3 for for VMs running on hyp that don't trust the host Suzuki K Poulose
@ 2026-09-20 21:28 ` Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 13/20] KVM: arm64: CCA: Add VCPU load/put for Realms Suzuki K Poulose
` (7 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: Suzuki K Poulose @ 2026-09-20 21:28 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose
Add an explicit mode to support Arm CCA guests.
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Documentation/admin-guide/kernel-parameters.txt | 3 +++
arch/arm64/include/asm/kvm_host.h | 1 +
arch/arm64/kvm/arm.c | 5 +++++
3 files changed, 9 insertions(+)
diff --git a/Documentation/admin-guide/kernel-parameters.txt b/Documentation/admin-guide/kernel-parameters.txt
index 68647ff4bdd24..1afe3df3b923e 100644
--- a/Documentation/admin-guide/kernel-parameters.txt
+++ b/Documentation/admin-guide/kernel-parameters.txt
@@ -3256,6 +3256,9 @@ Kernel parameters
nested: VHE-based mode with support for nested
virtualization. Requires at least ARMv8.4
hardware (with FEAT_NV2).
+ rmm: Support for running confidential guests in Realm
+ world using RMM, as defined by Arm Confidential
+ Compute Architecture (CCA)
Defaults to VHE/nVHE based on hardware support. Setting
mode to "protected" will disable kexec and hibernation
diff --git a/arch/arm64/include/asm/kvm_host.h b/arch/arm64/include/asm/kvm_host.h
index bb156a633ab75..ecbf5a35cd59e 100644
--- a/arch/arm64/include/asm/kvm_host.h
+++ b/arch/arm64/include/asm/kvm_host.h
@@ -69,6 +69,7 @@ enum kvm_mode {
KVM_MODE_DEFAULT,
KVM_MODE_PROTECTED,
KVM_MODE_NV,
+ KVM_MODE_RMM,
KVM_MODE_NONE,
};
#ifdef CONFIG_KVM
diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c
index 9b977dc734220..0e2ab5311e4d8 100644
--- a/arch/arm64/kvm/arm.c
+++ b/arch/arm64/kvm/arm.c
@@ -3268,6 +3268,11 @@ static int __init early_kvm_mode_cfg(char *arg)
return 0;
}
+ if (strcmp(arg, "rmm") == 0 && !WARN_ON(!is_kernel_in_hyp_mode())) {
+ kvm_mode = KVM_MODE_RMM;
+ return 0;
+ }
+
return -EINVAL;
}
early_param("kvm-arm.mode", early_kvm_mode_cfg);
--
2.43.0
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v19 13/20] KVM: arm64: CCA: Add VCPU load/put for Realms
2026-09-20 21:28 [PATCH v19 00/20] KVM: arm64: CCA: Add basic plumbing for Realms Suzuki K Poulose
` (11 preceding siblings ...)
2026-09-20 21:28 ` [PATCH v19 12/20] KVM: arm64: CCA: Add a new mode for supporting Realm guests Suzuki K Poulose
@ 2026-09-20 21:28 ` Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 14/20] KVM: arm64: CCA: Add bare minimal S2 operations for Realm Suzuki K Poulose
` (6 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: Suzuki K Poulose @ 2026-09-20 21:28 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose
RMM controls the VCPU settings and most are hidden from the KVM, except
for the VGIC and timer bits.
A later patch would add syncing the VCPU state into the Realm REC related
SMC parameters.
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
arch/arm64/kvm/arm.c | 18 ++++++++++++++++++
1 file changed, 18 insertions(+)
diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c
index 0e2ab5311e4d8..bcc274df2a915 100644
--- a/arch/arm64/kvm/arm.c
+++ b/arch/arm64/kvm/arm.c
@@ -795,6 +795,13 @@ static void pkvm_vcpu_load(struct kvm_vcpu *vcpu)
&vcpu->arch.vgic_cpu.vgic_v3);
}
+static void realm_vcpu_load(struct kvm_vcpu *vcpu)
+{
+ kvm_timer_vcpu_load(vcpu);
+ kvm_vgic_load(vcpu);
+ vcpu_set_wfx_traps(vcpu);
+}
+
void kvm_arch_vcpu_load(struct kvm_vcpu *vcpu, int cpu)
{
vcpu->cpu = cpu;
@@ -845,6 +852,12 @@ static void pkvm_vcpu_put(struct kvm_vcpu *vcpu)
kvm_vcpu_pmu_restore_host(vcpu);
}
+static void realm_vcpu_put(struct kvm_vcpu *vcpu)
+{
+ kvm_timer_vcpu_put(vcpu);
+ kvm_vgic_put(vcpu);
+}
+
void kvm_arch_vcpu_put(struct kvm_vcpu *vcpu)
{
vcpu->arch.vcpu_ops->vcpu_put(vcpu);
@@ -2202,6 +2215,11 @@ static const struct kvm_vcpu_ops pkvm_vcpu_ops = {
.vcpu_put = pkvm_vcpu_put,
};
+static const struct kvm_vcpu_ops realm_vcpu_ops = {
+ .vcpu_load = realm_vcpu_load,
+ .vcpu_put = realm_vcpu_put,
+};
+
#define KVM_VCPU_OPS(flavor, ops) \
[(flavor)] = (ops)
--
2.43.0
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v19 14/20] KVM: arm64: CCA: Add bare minimal S2 operations for Realm
2026-09-20 21:28 [PATCH v19 00/20] KVM: arm64: CCA: Add basic plumbing for Realms Suzuki K Poulose
` (12 preceding siblings ...)
2026-09-20 21:28 ` [PATCH v19 13/20] KVM: arm64: CCA: Add VCPU load/put for Realms Suzuki K Poulose
@ 2026-09-20 21:28 ` Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 15/20] KVM: arm64: CCA: Introduce Realms Suzuki K Poulose
` (5 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: Suzuki K Poulose @ 2026-09-20 21:28 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose
Add bare minimal MMU operation hooks for Realms. The mem_abort handling
is chosen as the default KVM variant. However this cannot be reached for
Realms yet and we would need real RMI command support to make it fully
functional.
RMM takes care of the TLB flushing as required, when the Stage2 is
modified. So host doesn't need to do anything explicitly. RMM doesn't
support access flags for the stage2, even for the shared IPA.
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
arch/arm64/kvm/mmu.c | 24 ++++++++++++++++++++++++
1 file changed, 24 insertions(+)
diff --git a/arch/arm64/kvm/mmu.c b/arch/arm64/kvm/mmu.c
index 413c4b114d75d..a38459f6b456e 100644
--- a/arch/arm64/kvm/mmu.c
+++ b/arch/arm64/kvm/mmu.c
@@ -180,6 +180,12 @@ static int kvm_vm_flush_remote_tlbs(struct kvm *kvm)
return 0;
}
+static int realm_vm_flush_remote_tlbs(struct kvm *kvm)
+{
+ /* Nothing to do here, RMM takes care of this */
+ return 0;
+}
+
/**
* kvm_arch_flush_remote_tlbs() - flush all VM TLB entries for v7/8
* @kvm: pointer to kvm structure.
@@ -209,6 +215,13 @@ static int kvm_vm_flush_remote_tlbs_range(struct kvm *kvm,
return 0;
}
+static int realm_vm_flush_remote_tlbs_range(struct kvm *kvm,
+ gfn_t gfn, u64 nr_pages)
+{
+ /* Nothing to do here, RMM takes care of this */
+ return 0;
+}
+
int kvm_arch_flush_remote_tlbs_range(struct kvm *kvm,
gfn_t gfn, u64 nr_pages)
{
@@ -2881,6 +2894,17 @@ static const struct kvm_vm_s2_ops kvm_default_vm_s2_ops = {
.vm_mem_abort = kvm_vm_mem_abort,
};
+static const struct kvm_vm_s2_ops realm_vm_s2_ops = {
+ .vm_flush_remote_tlbs = realm_vm_flush_remote_tlbs,
+ .vm_flush_remote_tlbs_range = realm_vm_flush_remote_tlbs_range,
+ .vm_mem_abort = kvm_vm_mem_abort,
+ /*
+ * Not supported for Realms
+ * .vm_age_gfn = realm_vm_age_gfn,
+ * .vm_test_age_gfn = realm_vm_test_age_gfn,
+ */
+};
+
#define KVM_VM_S2_OPS(flavor, ops) \
[flavor] = ops
static const struct kvm_vm_s2_ops *arm64_vm_s2_ops[] = {
--
2.43.0
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v19 15/20] KVM: arm64: CCA: Introduce Realms
2026-09-20 21:28 [PATCH v19 00/20] KVM: arm64: CCA: Add basic plumbing for Realms Suzuki K Poulose
` (13 preceding siblings ...)
2026-09-20 21:28 ` [PATCH v19 14/20] KVM: arm64: CCA: Add bare minimal S2 operations for Realm Suzuki K Poulose
@ 2026-09-20 21:28 ` Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 16/20] KVM: arm64: CCA: Don't expose unsupported capabilities for realm guests Suzuki K Poulose
` (4 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: Suzuki K Poulose @ 2026-09-20 21:28 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose
From: Steven Price <steven.price@arm.com>
Add foundational work for supporting Realms.
- Add a new VM flavor.
- At KVM init, check if the KVM can support Realms (though not functional yet)
and will be advertised by static key kvm_rmi_is_available. This will be
turned on in a later patches, once we have all the bits and pieces ready.
For now check if we are blessed with KVM_MODE_RMM.
- Add realm specific tracking in kvm_arch. Since Realm and protected pKVM
states are mutually exclusive, move them into a union.
Please note that we cannot create Realm VMs yet. This requires further changes
to the UABI and core RMI driver support, which will come later.
Signed-off-by: Steven Price <steven.price@arm.com>
Co-developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Changes since v16:
* Share mutually exclusive pKVM and Realm per-VM storage in a union.
* Move to the new VM flavor infrastructure, split bits out. Trimmed down
* Move in Realm state and basic boiler plates
Changes since v13:
* Most of the init has been moved out of the 'kvm' directory so this is
much more basic now.
Changes since v12:
* Drop check for 4k page size.
Changes since v11:
* Reword slightly the comments on the realm states.
Changes since v10:
* kvm_is_realm() no longer has a NULL check.
* Rename from "rme" to "rmi" when referring to the RMM interface.
* Check for RME (hardware) support before probing for RMI support.
Changes since v8:
* No need to guard kvm_init_rme() behind 'in_hyp_mode'.
Changes since v6:
* Improved message for an unsupported RMI ABI version.
Changes since v5:
* Reword "unsupported" message from "host supports" to "we want" to
clarify that 'we' are the 'host'.
Changes since v2:
* Drop return value from kvm_init_rme(), it was always 0.
* Rely on the RMM return value to identify whether the RSI ABI is
compatible.
---
arch/arm64/include/asm/kvm_emulate.h | 16 ++++++++
arch/arm64/include/asm/kvm_host.h | 18 +++++---
arch/arm64/include/asm/kvm_rmi.h | 61 ++++++++++++++++++++++++++++
arch/arm64/include/asm/virt.h | 1 +
arch/arm64/kvm/Makefile | 2 +-
arch/arm64/kvm/arm.c | 6 +++
arch/arm64/kvm/mmu.c | 1 +
arch/arm64/kvm/rmi.c | 18 ++++++++
8 files changed, 117 insertions(+), 6 deletions(-)
create mode 100644 arch/arm64/include/asm/kvm_rmi.h
create mode 100644 arch/arm64/kvm/rmi.c
diff --git a/arch/arm64/include/asm/kvm_emulate.h b/arch/arm64/include/asm/kvm_emulate.h
index a3c1928bdf743..d360a8b05b8bf 100644
--- a/arch/arm64/include/asm/kvm_emulate.h
+++ b/arch/arm64/include/asm/kvm_emulate.h
@@ -793,4 +793,20 @@ static inline void kvm_reset_vcpu_psci(struct kvm_vcpu *vcpu,
vcpu_set_reg(vcpu, 0, reset_state->r0);
}
+static inline enum realm_state kvm_realm_state(struct kvm *kvm)
+{
+ return READ_ONCE(kvm->arch.realm.state);
+}
+
+static inline void kvm_set_realm_state(struct kvm *kvm,
+ enum realm_state new_state)
+{
+ WRITE_ONCE(kvm->arch.realm.state, new_state);
+}
+
+static inline bool kvm_realm_is_created(struct kvm *kvm)
+{
+ return kvm_vm_is_realm(kvm) && kvm_realm_state(kvm) != REALM_STATE_NONE;
+}
+
#endif /* __ARM64_KVM_EMULATE_H__ */
diff --git a/arch/arm64/include/asm/kvm_host.h b/arch/arm64/include/asm/kvm_host.h
index ecbf5a35cd59e..6d12e2c360af6 100644
--- a/arch/arm64/include/asm/kvm_host.h
+++ b/arch/arm64/include/asm/kvm_host.h
@@ -27,6 +27,7 @@
#include <asm/fpsimd.h>
#include <asm/kvm.h>
#include <asm/kvm_asm.h>
+#include <asm/kvm_rmi.h>
#include <asm/vncr_mapping.h>
#define __KVM_HAVE_ARCH_INTC_INITIALIZED
@@ -334,6 +335,7 @@ enum kvm_arm_vm_flavor {
VM_PKVM, /* Normal guests on pKVM */
MARKER(__VM_PROTECTED),
VM_PROTECTED_PKVM, /* Protected VM */
+ VM_REALM, /* CCA */
VM_FLAVOR_MAX,
};
@@ -451,11 +453,14 @@ struct kvm_arch {
/* Count the number of VNCR_EL2 TLBs */
atomic_t vncr_tlb_count;
- /*
- * For an untrusted host VM, 'pkvm.handle' is used to lookup
- * the associated pKVM instance in the hypervisor.
- */
- struct kvm_protected_vm pkvm;
+ union {
+ /*
+ * For an untrusted host VM, 'pkvm.handle' is used to lookup
+ * the associated pKVM instance in the hypervisor.
+ */
+ struct kvm_protected_vm pkvm;
+ struct realm realm;
+ };
#ifdef CONFIG_PTDUMP_STAGE2_DEBUGFS
/* Nested virtualization info */
@@ -1549,6 +1554,9 @@ struct kvm *kvm_arch_alloc_vm(void);
#define kvm_vm_is_unprotected_pkvm(kvm) \
(is_protected_kvm_enabled() && ((kvm)->arch.vm_flavor == VM_PKVM))
+#define kvm_vm_is_realm(kvm) ((kvm)->arch.vm_flavor == VM_REALM)
+#define vcpu_is_rec(vcpu) kvm_vm_is_realm((vcpu)->kvm)
+
#define kvm_vm_hyp_is_distrusting(kvm) ((kvm)->arch.vm_flavor >= __VM_DISTRUSTING_HYP)
int kvm_arm_vcpu_finalize(struct kvm_vcpu *vcpu, int feature);
diff --git a/arch/arm64/include/asm/kvm_rmi.h b/arch/arm64/include/asm/kvm_rmi.h
new file mode 100644
index 0000000000000..44f5c75a27b5b
--- /dev/null
+++ b/arch/arm64/include/asm/kvm_rmi.h
@@ -0,0 +1,61 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (C) 2023-2026 ARM Ltd.
+ */
+
+#ifndef __ASM_KVM_RMI_H
+#define __ASM_KVM_RMI_H
+
+/**
+ * enum realm_state - State of a Realm
+ *
+ * Mirrors the RMM's Realm lifecycle states where they are meaningful to KVM,
+ * with REALM_STATE_DYING being a KVM-internal state used to prevent further
+ * requests while teardown is in progress. KVM does not track REALM_SYSTEM_OFF
+ * or REALM_ZOMBIE separately as they naturally lead to teardown.
+ */
+enum realm_state {
+ /**
+ * @REALM_STATE_NONE:
+ * Realm has not yet been created. rmi_realm_create() has not
+ * yet been called.
+ */
+ REALM_STATE_NONE,
+ /**
+ * @REALM_STATE_NEW:
+ * Realm is under construction, rmi_realm_create() has been
+ * called, but it is not yet activated. Pages may be populated.
+ */
+ REALM_STATE_NEW,
+ /**
+ * @REALM_STATE_ACTIVE:
+ * Realm has been created and is eligible for execution with
+ * rmi_rec_enter(). Pages may no longer be populated with
+ * rmi_data_create().
+ */
+ REALM_STATE_ACTIVE,
+ /**
+ * @REALM_STATE_DYING:
+ * Realm is in the process of being destroyed or has already been
+ * destroyed.
+ */
+ REALM_STATE_DYING,
+ /**
+ * @REALM_STATE_DEAD:
+ * Realm has been destroyed.
+ */
+ REALM_STATE_DEAD
+};
+
+/**
+ * struct realm - Additional per VM data for a Realm
+ *
+ * @state: The lifetime state machine for the realm
+ */
+struct realm {
+ enum realm_state state;
+};
+
+void kvm_init_rmi(void);
+
+#endif /* __ASM_KVM_RMI_H */
diff --git a/arch/arm64/include/asm/virt.h b/arch/arm64/include/asm/virt.h
index b546703c3ab9a..92cec42952f42 100644
--- a/arch/arm64/include/asm/virt.h
+++ b/arch/arm64/include/asm/virt.h
@@ -87,6 +87,7 @@ void __hyp_reset_vectors(void);
bool is_kvm_arm_initialised(void);
DECLARE_STATIC_KEY_FALSE(kvm_protected_mode_initialized);
+DECLARE_STATIC_KEY_FALSE(kvm_rmi_is_available);
static inline bool is_pkvm_initialized(void)
{
diff --git a/arch/arm64/kvm/Makefile b/arch/arm64/kvm/Makefile
index 59612d2f277c1..ed3cf30eb06e7 100644
--- a/arch/arm64/kvm/Makefile
+++ b/arch/arm64/kvm/Makefile
@@ -16,7 +16,7 @@ CFLAGS_handle_exit.o += -Wno-override-init
kvm-y += arm.o mmu.o mmio.o psci.o hypercalls.o pvtime.o \
inject_fault.o va_layout.o handle_exit.o config.o \
guest.o debug.o reset.o sys_regs.o stacktrace.o \
- vgic-sys-reg-v3.o fpsimd.o pkvm.o \
+ vgic-sys-reg-v3.o fpsimd.o pkvm.o rmi.o \
arch_timer.o trng.o vmid.o emulate-nested.o nested.o at.o \
vgic/vgic.o vgic/vgic-init.o \
vgic/vgic-irqfd.o vgic/vgic-v2.o \
diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c
index bcc274df2a915..86e705330d7bd 100644
--- a/arch/arm64/kvm/arm.c
+++ b/arch/arm64/kvm/arm.c
@@ -42,6 +42,7 @@
#include <asm/kvm_nested.h>
#include <asm/kvm_pkvm.h>
#include <asm/kvm_ptrauth.h>
+#include <asm/kvm_rmi.h>
#include <asm/sections.h>
#include <asm/stacktrace/nvhe.h>
@@ -112,6 +113,8 @@ long kvm_get_cap_for_kvm_ioctl(unsigned int ioctl, long *ext)
return -EINVAL;
}
+DEFINE_STATIC_KEY_FALSE(kvm_rmi_is_available);
+
DECLARE_KVM_HYP_PER_CPU(unsigned long, kvm_hyp_vector);
DEFINE_PER_CPU(unsigned long, kvm_arm_hyp_stack_base);
@@ -2228,6 +2231,7 @@ static const struct kvm_vcpu_ops *arm64_vcpu_ops[] = {
KVM_VCPU_OPS(VM_NVHE, &nvhe_vcpu_ops),
KVM_VCPU_OPS(VM_PKVM, &pkvm_vcpu_ops),
KVM_VCPU_OPS(VM_PROTECTED_PKVM, &pkvm_vcpu_ops),
+ KVM_VCPU_OPS(VM_REALM, &realm_vcpu_ops),
};
static void kvm_init_vcpu_ops(struct kvm_vcpu *vcpu)
@@ -3181,6 +3185,8 @@ static __init int kvm_arm_init(void)
in_hyp_mode = is_kernel_in_hyp_mode();
+ kvm_init_rmi();
+
if (cpus_have_final_cap(ARM64_WORKAROUND_DEVICE_LOAD_ACQUIRE) ||
cpus_have_final_cap(ARM64_WORKAROUND_1508412))
kvm_info("Guests without required CPU erratum workarounds can deadlock system!\n" \
diff --git a/arch/arm64/kvm/mmu.c b/arch/arm64/kvm/mmu.c
index a38459f6b456e..39ee27ba9d493 100644
--- a/arch/arm64/kvm/mmu.c
+++ b/arch/arm64/kvm/mmu.c
@@ -2912,6 +2912,7 @@ static const struct kvm_vm_s2_ops *arm64_vm_s2_ops[] = {
KVM_VM_S2_OPS(VM_NVHE, &kvm_default_vm_s2_ops),
KVM_VM_S2_OPS(VM_PKVM, &pkvm_vm_s2_ops),
KVM_VM_S2_OPS(VM_PROTECTED_PKVM, &protected_pkvm_vm_s2_ops),
+ KVM_VM_S2_OPS(VM_REALM, &realm_vm_s2_ops),
};
static int kvm_vm_init_vm_s2_ops(struct kvm *kvm)
diff --git a/arch/arm64/kvm/rmi.c b/arch/arm64/kvm/rmi.c
new file mode 100644
index 0000000000000..5ecc8b3498698
--- /dev/null
+++ b/arch/arm64/kvm/rmi.c
@@ -0,0 +1,18 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (C) 2023-2026 ARM Ltd.
+ */
+
+#include <linux/kvm_host.h>
+
+#include <asm/virt.h>
+
+void kvm_init_rmi(void)
+{
+ if (kvm_get_mode() != KVM_MODE_RMM)
+ return;
+
+ /* TODO: Check if the RMI is available */
+
+ /* Future patch will enable static branch kvm_rmi_is_available */
+}
--
2.43.0
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v19 16/20] KVM: arm64: CCA: Don't expose unsupported capabilities for realm guests
2026-09-20 21:28 [PATCH v19 00/20] KVM: arm64: CCA: Add basic plumbing for Realms Suzuki K Poulose
` (14 preceding siblings ...)
2026-09-20 21:28 ` [PATCH v19 15/20] KVM: arm64: CCA: Introduce Realms Suzuki K Poulose
@ 2026-09-20 21:28 ` Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 17/20] KVM: arm64: CCA: WARN on injected undef exceptions Suzuki K Poulose
` (3 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: Suzuki K Poulose @ 2026-09-20 21:28 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose
Limit the capabilities that are allowed for Realm VMs. Similarly block
the vm_ioctls backed by the capabilities.
Repurpose the kvm_pkvm_ioctl_allowed() to support both pKVM and Realm
ioctls. Rename the helper to kvm_vm_ioctl_allowed() and move it
into arch/arm64/kvm/arm.c. Also add a generic kvm_vm_ext_allowed()
to handle pKVM and Realm capability filtering and route them accordingly.
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Changes since v18:
* Bail out early for !pKVM && !Realm vms. Use kvm_vm_hyp_is_distrusting()
* WARN_ON_ONCE(!kvm), we are only called from kvm_arch_vm_ioctl() with a
valid kvm instance.
* Move the kvm_realm_ext_allowed() to asm/kvm_rmi.h - Fuad
* Rename kvm_arch_vm_*allowed => kvm_vm_*_allowed - Fuad
Changes since v17:
* Drop superfluous !kvm check from kvm_vm_ioctl_enable_cap() - Sashiko
* Drop KVM_CAP_CREATE_IRQCHIP, as we don't support VGIC_V2 for Realms
* Filter out the vm_ioctls that are based on blocked cap.
* Repurpose the pkvm plumbing for filtering the caps and ioctl to generic
and plumb the Realm support in
Changes since v13:
* Add missing check in kvm_vm_ioctl_enable_cap().
Changes since v10:
* Add a kvm_realm_ext_allowed() function which limits which extensions
are exposed to an allowlist. This removes the need for special casing
various extensions.
Changes since v7:
* Remove the helper functions and inline the kvm_is_realm() check with
a ternary operator.
* Rewrite the commit message to explain this patch.
---
arch/arm64/include/asm/kvm_pkvm.h | 19 ------------
arch/arm64/include/asm/kvm_rmi.h | 23 +++++++++++++++
arch/arm64/kvm/arm.c | 49 +++++++++++++++++++++++++++++--
3 files changed, 69 insertions(+), 22 deletions(-)
diff --git a/arch/arm64/include/asm/kvm_pkvm.h b/arch/arm64/include/asm/kvm_pkvm.h
index e4ea80711bec6..1bc4fe2726e9b 100644
--- a/arch/arm64/include/asm/kvm_pkvm.h
+++ b/arch/arm64/include/asm/kvm_pkvm.h
@@ -53,25 +53,6 @@ static inline bool kvm_pkvm_ext_allowed(struct kvm *kvm, long ext)
}
}
-/*
- * Check whether the KVM VM IOCTL is allowed in pKVM.
- *
- * Certain features are allowed only for non-protected VMs in pKVM, which is why
- * this takes the VM (kvm) as a parameter.
- */
-static inline bool kvm_pkvm_ioctl_allowed(struct kvm *kvm, unsigned int ioctl)
-{
- long ext;
- int r;
-
- r = kvm_get_cap_for_kvm_ioctl(ioctl, &ext);
-
- if (WARN_ON_ONCE(r < 0))
- return false;
-
- return kvm_pkvm_ext_allowed(kvm, ext);
-}
-
extern struct memblock_region kvm_nvhe_sym(hyp_memory)[];
extern unsigned int kvm_nvhe_sym(hyp_memblock_nr);
diff --git a/arch/arm64/include/asm/kvm_rmi.h b/arch/arm64/include/asm/kvm_rmi.h
index 44f5c75a27b5b..6b8b9ee9ea245 100644
--- a/arch/arm64/include/asm/kvm_rmi.h
+++ b/arch/arm64/include/asm/kvm_rmi.h
@@ -6,6 +6,8 @@
#ifndef __ASM_KVM_RMI_H
#define __ASM_KVM_RMI_H
+#include <linux/kvm.h>
+
/**
* enum realm_state - State of a Realm
*
@@ -58,4 +60,25 @@ struct realm {
void kvm_init_rmi(void);
+static inline bool kvm_realm_ext_allowed(long ext)
+{
+ switch (ext) {
+ case KVM_CAP_IRQCHIP:
+ case KVM_CAP_ARM_PSCI:
+ case KVM_CAP_ARM_PSCI_0_2:
+ case KVM_CAP_NR_VCPUS:
+ case KVM_CAP_MAX_VCPUS:
+ case KVM_CAP_MAX_VCPU_ID:
+ case KVM_CAP_MSI_DEVID:
+ case KVM_CAP_ARM_VM_IPA_SIZE:
+ case KVM_CAP_ARM_SVE:
+ case KVM_CAP_ONE_REG:
+ case KVM_CAP_ARM_PTRAUTH_ADDRESS:
+ case KVM_CAP_ARM_PTRAUTH_GENERIC:
+ case KVM_CAP_SYNC_MMU:
+ return true;
+ }
+ return false;
+}
+
#endif /* __ASM_KVM_RMI_H */
diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c
index 86e705330d7bd..eab8543a4194d 100644
--- a/arch/arm64/kvm/arm.c
+++ b/arch/arm64/kvm/arm.c
@@ -136,6 +136,49 @@ int kvm_arch_vcpu_should_kick(struct kvm_vcpu *vcpu)
return kvm_vcpu_exiting_guest_mode(vcpu) == IN_GUEST_MODE;
}
+static inline bool kvm_vm_ext_allowed(struct kvm *kvm, long ext)
+{
+ /*
+ * We could be called with kvm as NULL, so can't use kvm_vm_* for pKVM
+ * flavors
+ */
+ if (is_protected_kvm_enabled())
+ return kvm_pkvm_ext_allowed(kvm, ext);
+ else if (kvm && kvm_vm_is_realm(kvm))
+ return kvm_realm_ext_allowed(ext);
+ else
+ return true;
+}
+
+/*
+ * Check whether the KVM VM IOCTL is allowed. For pKVM and Realm VMs, certain
+ * ioctls are not allowed. Further, certain features are allowed only for
+ * non-protected VMs in pKVM.
+ */
+static inline bool kvm_vm_ioctl_allowed(struct kvm *kvm, unsigned int ioctl)
+{
+ long ext;
+ int r;
+
+ /*
+ * We are guaranteed to be called with a valid kvm instance, as the
+ * only caller is kvm_arch_vm_ioctl(). Catch any deviations, as we
+ * rely on the kvm instance below.
+ */
+ if (WARN_ON_ONCE(!kvm))
+ return false;
+
+ /* Cover both pKVM host and Realm VMs */
+ if (!kvm_vm_hyp_is_distrusting(kvm))
+ return true;
+
+ r = kvm_get_cap_for_kvm_ioctl(ioctl, &ext);
+ if (WARN_ON_ONCE(r < 0))
+ return false;
+
+ return kvm_vm_ext_allowed(kvm, ext);
+}
+
int kvm_vm_ioctl_enable_cap(struct kvm *kvm,
struct kvm_enable_cap *cap)
{
@@ -144,7 +187,7 @@ int kvm_vm_ioctl_enable_cap(struct kvm *kvm,
if (cap->flags)
return -EINVAL;
- if (is_protected_kvm_enabled() && !kvm_pkvm_ext_allowed(kvm, cap->cap))
+ if (!kvm_vm_ext_allowed(kvm, cap->cap))
return -EINVAL;
switch (cap->cap) {
@@ -403,7 +446,7 @@ int kvm_vm_ioctl_check_extension(struct kvm *kvm, long ext)
{
int r;
- if (is_protected_kvm_enabled() && !kvm_pkvm_ext_allowed(kvm, ext))
+ if (!kvm_vm_ext_allowed(kvm, ext))
return 0;
switch (ext) {
@@ -2135,7 +2178,7 @@ int kvm_arch_vm_ioctl(struct file *filp, unsigned int ioctl, unsigned long arg)
void __user *argp = (void __user *)arg;
struct kvm_device_attr attr;
- if (is_protected_kvm_enabled() && !kvm_pkvm_ioctl_allowed(kvm, ioctl))
+ if (!kvm_vm_ioctl_allowed(kvm, ioctl))
return -EINVAL;
switch (ioctl) {
--
2.43.0
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v19 17/20] KVM: arm64: CCA: WARN on injected undef exceptions
2026-09-20 21:28 [PATCH v19 00/20] KVM: arm64: CCA: Add basic plumbing for Realms Suzuki K Poulose
` (15 preceding siblings ...)
2026-09-20 21:28 ` [PATCH v19 16/20] KVM: arm64: CCA: Don't expose unsupported capabilities for realm guests Suzuki K Poulose
@ 2026-09-20 21:28 ` Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 18/20] KVM: arm64: CCA: Support timers in realm RECs Suzuki K Poulose
` (2 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: Suzuki K Poulose @ 2026-09-20 21:28 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei
From: Steven Price <steven.price@arm.com>
The RMM doesn't allow injection of a undefined exception into a realm
guest. Add a WARN to catch if this ever happens.
Signed-off-by: Steven Price <steven.price@arm.com>
---
Changes since v15:
* Switch to KVM_BUG() to mark the VM as bugged as well.
Changes since v6:
* if (x) WARN(1, ...) makes no sense, just WARN(x, ...)!
---
arch/arm64/kvm/inject_fault.c | 1 +
1 file changed, 1 insertion(+)
diff --git a/arch/arm64/kvm/inject_fault.c b/arch/arm64/kvm/inject_fault.c
index d6c4fc16f8795..d61a3ff04fabb 100644
--- a/arch/arm64/kvm/inject_fault.c
+++ b/arch/arm64/kvm/inject_fault.c
@@ -317,6 +317,7 @@ void kvm_inject_size_fault(struct kvm_vcpu *vcpu)
*/
void kvm_inject_undefined(struct kvm_vcpu *vcpu)
{
+ KVM_BUG(vcpu_is_rec(vcpu), vcpu->kvm, "Unexpected undefined exception injection to REC");
if (vcpu_el1_is_32bit(vcpu))
inject_undef32(vcpu);
else
--
2.43.0
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v19 18/20] KVM: arm64: CCA: Support timers in realm RECs
2026-09-20 21:28 [PATCH v19 00/20] KVM: arm64: CCA: Add basic plumbing for Realms Suzuki K Poulose
` (16 preceding siblings ...)
2026-09-20 21:28 ` [PATCH v19 17/20] KVM: arm64: CCA: WARN on injected undef exceptions Suzuki K Poulose
@ 2026-09-20 21:28 ` Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 19/20] KVM: arm64: CCA: Expose SVE VL register before VCPU finalization Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 20/20] KVM: arm64: CCA: Control user register access for Realms Suzuki K Poulose
19 siblings, 0 replies; 21+ messages in thread
From: Suzuki K Poulose @ 2026-09-20 21:28 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose
From: Steven Price <steven.price@arm.com>
The RMM keeps track of the timer while the realm REC is running, but on
exit to the normal world KVM is responsible for handling the timers.
A later patch adds the support for propagating the timer values from the
exit data structure and making sure the values are in sync for KVM.
Also, RMM doesn't support injecting virtual interrupts backed by Physical
interrupts. So, use the existing software resampling mechanims for Realm
timer interrupts.
Signed-off-by: Steven Price <steven.price@arm.com>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
arch/arm64/kvm/arch_timer.c | 22 ++++++++++++++++++++--
1 file changed, 20 insertions(+), 2 deletions(-)
diff --git a/arch/arm64/kvm/arch_timer.c b/arch/arm64/kvm/arch_timer.c
index 226cd5a495c8b..83c13bbdcc777 100644
--- a/arch/arm64/kvm/arch_timer.c
+++ b/arch/arm64/kvm/arch_timer.c
@@ -56,11 +56,25 @@ static unsigned long kvm_arch_timer_get_irq_flags(void)
return kvm_vgic_global_state.no_hw_deactivation ? VGIC_IRQ_SW_RESAMPLE : 0;
}
+static unsigned long kvm_realm_timer_get_irq_flags(void)
+{
+ /*
+ * RMI_REC_ENTER rejects LRs with the HW bit set, so use the existing
+ * software resampling mechanism for Realm timer interrupts.
+ */
+ return VGIC_IRQ_SW_RESAMPLE;
+}
+
static const struct irq_ops arch_timer_irq_ops = {
.get_flags = kvm_arch_timer_get_irq_flags,
.get_input_level = kvm_arch_timer_get_input_level,
};
+static const struct irq_ops realm_timer_irq_ops = {
+ .get_flags = kvm_realm_timer_get_irq_flags,
+ .get_input_level = kvm_arch_timer_get_input_level,
+};
+
static const struct irq_ops arch_timer_irq_ops_vgic_v5 = {
.get_input_level = kvm_arch_timer_get_input_level,
.queue_irq_unlock = vgic_v5_ppi_queue_irq_unlock,
@@ -1617,8 +1631,12 @@ int kvm_timer_enable(struct kvm_vcpu *vcpu)
get_timer_map(vcpu, &map);
- ops = vgic_is_v5(vcpu->kvm) ? &arch_timer_irq_ops_vgic_v5 :
- &arch_timer_irq_ops;
+ if (vcpu_is_rec(vcpu))
+ ops = &realm_timer_irq_ops;
+ else if (vgic_is_v5(vcpu->kvm))
+ ops = &arch_timer_irq_ops_vgic_v5;
+ else
+ ops = &arch_timer_irq_ops;
for (int i = 0; i < nr_timers(vcpu); i++)
kvm_vgic_set_irq_ops(vcpu, timer_irq(vcpu_get_timer(vcpu, i)), ops);
--
2.43.0
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v19 19/20] KVM: arm64: CCA: Expose SVE VL register before VCPU finalization
2026-09-20 21:28 [PATCH v19 00/20] KVM: arm64: CCA: Add basic plumbing for Realms Suzuki K Poulose
` (17 preceding siblings ...)
2026-09-20 21:28 ` [PATCH v19 18/20] KVM: arm64: CCA: Support timers in realm RECs Suzuki K Poulose
@ 2026-09-20 21:28 ` Suzuki K Poulose
2026-09-20 21:28 ` [PATCH v19 20/20] KVM: arm64: CCA: Control user register access for Realms Suzuki K Poulose
19 siblings, 0 replies; 21+ messages in thread
From: Suzuki K Poulose @ 2026-09-20 21:28 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei,
Jean-Philippe Brucker, Suzuki K Poulose
From: Jean-Philippe Brucker <jean-philippe@linaro.org>
Userspace must configure the SVE vector length before the Realm is created
(as it is part of the parameter for Realm creation), but the Realm VCPUs
cannot be finalized until after the Realm Descriptor has been created.
KVM_GET_REG_LIST currently rejects the unfinalized VCPUs, which prevents
the userspace from discovering and configuring the VLs for the Realm.
Allow KVM_GET_REG_LIST for unfinalized RECs and make the SVE register
enumeration handle the unfinalized case explicitly. i.e., only expose
KVM_REG_ARM64_SVE_VLS before SVE is finalized.
One adverse side effect of this change is that a KVM_GET_REG_LIST call that
only probes for the array size will now succeed even if SVE is not
finalized, but that seems harmless since the following KVM_GET_REG_LIST
with the full array will fail.
Signed-off-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
Signed-off-by: Steven Price <steven.price@arm.com>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Changes since v17:
- Rewrite the commit description to clearly describe the purpose
---
arch/arm64/kvm/arm.c | 15 ++++++++++++++-
arch/arm64/kvm/guest.c | 10 +++++-----
2 files changed, 19 insertions(+), 6 deletions(-)
diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c
index eab8543a4194d..6a52b102d0ce4 100644
--- a/arch/arm64/kvm/arm.c
+++ b/arch/arm64/kvm/arm.c
@@ -1993,6 +1993,19 @@ static int kvm_arm_vcpu_set_events(struct kvm_vcpu *vcpu,
return __kvm_arm_vcpu_set_events(vcpu, events);
}
+/*
+ * Realm VCPUs can be finalized only after the Realm descriptor is created.
+ * But in order to seal the SVE VL, we need to allow the userspace to read/write
+ * to the SVE_VL, before everything is finalized.
+ * Allow the register list for RECs before the VCPUs are finalized.
+ */
+static bool kvm_arm_vcpu_reg_list_allowed(struct kvm_vcpu *vcpu)
+{
+ if (kvm_arm_vcpu_is_finalized(vcpu))
+ return true;
+ return vcpu_is_rec(vcpu);
+}
+
long kvm_arch_vcpu_ioctl(struct file *filp,
unsigned int ioctl, unsigned long arg)
{
@@ -2048,7 +2061,7 @@ long kvm_arch_vcpu_ioctl(struct file *filp,
break;
r = -EPERM;
- if (!kvm_arm_vcpu_is_finalized(vcpu))
+ if (!kvm_arm_vcpu_reg_list_allowed(vcpu))
break;
r = -EFAULT;
diff --git a/arch/arm64/kvm/guest.c b/arch/arm64/kvm/guest.c
index b01d6622b8720..c3ca369882273 100644
--- a/arch/arm64/kvm/guest.c
+++ b/arch/arm64/kvm/guest.c
@@ -598,8 +598,8 @@ static unsigned long num_sve_regs(const struct kvm_vcpu *vcpu)
if (!vcpu_has_sve(vcpu))
return 0;
- /* Policed by KVM_GET_REG_LIST: */
- WARN_ON(!kvm_arm_vcpu_sve_finalized(vcpu));
+ if (!kvm_arm_vcpu_sve_finalized(vcpu))
+ return 1; /* KVM_REG_ARM64_SVE_VLS */
return slices * (SVE_NUM_PREGS + SVE_NUM_ZREGS + 1 /* FFR */)
+ 1; /* KVM_REG_ARM64_SVE_VLS */
@@ -616,9 +616,6 @@ static int copy_sve_reg_indices(const struct kvm_vcpu *vcpu,
if (!vcpu_has_sve(vcpu))
return 0;
- /* Policed by KVM_GET_REG_LIST: */
- WARN_ON(!kvm_arm_vcpu_sve_finalized(vcpu));
-
/*
* Enumerate this first, so that userspace can save/restore in
* the order reported by KVM_GET_REG_LIST:
@@ -628,6 +625,9 @@ static int copy_sve_reg_indices(const struct kvm_vcpu *vcpu,
return -EFAULT;
++num_regs;
+ if (!kvm_arm_vcpu_sve_finalized(vcpu))
+ return num_regs;
+
for (i = 0; i < slices; i++) {
for (n = 0; n < SVE_NUM_ZREGS; n++) {
reg = KVM_REG_ARM64_SVE_ZREG(n, i);
--
2.43.0
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v19 20/20] KVM: arm64: CCA: Control user register access for Realms
2026-09-20 21:28 [PATCH v19 00/20] KVM: arm64: CCA: Add basic plumbing for Realms Suzuki K Poulose
` (18 preceding siblings ...)
2026-09-20 21:28 ` [PATCH v19 19/20] KVM: arm64: CCA: Expose SVE VL register before VCPU finalization Suzuki K Poulose
@ 2026-09-20 21:28 ` Suzuki K Poulose
19 siblings, 0 replies; 21+ messages in thread
From: Suzuki K Poulose @ 2026-09-20 21:28 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei,
Jean-Philippe Brucker, Suzuki K Poulose
From: Jean-Philippe Brucker <jean-philippe@linaro.org>
The RMM restricts the access to the register states that the host can
read/modify for a given Realm.
e.g., At VCPU creation, can modify GPRS (x0-x30) and PC.
While servicing SMCCC calls via RSI_HOST_CALL or servicing PSCI
requests.
MMIO emulation in the unprotected space.
Additionally we use the sysreg configuration to advertise/configure the
following Realm parameters, which are required before the Realm Descriptor
:w
is created:
- SVE Vector Length
- Number of HW Breakpoints/Watchpoints
- PMU Counters.
Thus KVM also additionally allows access to ID_AA64DFR0_EL1 and SVE_VLS for
the configuration of Realm creation parameters. We don't support PMUs for
the Realm VMs yet, so PMCR is not exposed.
The RMM makes similar restrictions for reading of the guest's registers
(this is *confidential* compute after all), however we don't impose the
restriction here. This allows the VMM to read (stale) values from the
registers which might be useful to read back the initial values even if
the RMM doesn't provide the latest version. For migration of a realm VM,
a new interface will be needed so that the VMM can receive an
(encrypted) blob of the VM's state.
Reflect the above in KVM_GET_REG_LIST, KVM_SET_ONE_REG calls.
Signed-off-by: Jean-Philippe Brucker <jean-philippe@linaro.org>
Co-developed-by: Steven Price <steven.price@arm.com>
Signed-off-by: Steven Price <steven.price@arm.com>
Co-developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Changes since v18:
- Don't expose PMCR_EL0 to the userspace now, we could do that when we
support PMU
- Fix check patch warnings
Changes since v17:
- Merge related changes into one single patch for the user set/get registers
I have retained the Review tags, as the code hasn't changed, just the patches
were merged into a single one with the same tags.
- Limit KVM_GET_REG_LIST to the allowed CORE registers.
---
arch/arm64/kvm/guest.c | 63 +++++++++++++++++++++++++++++++++++++
arch/arm64/kvm/hypercalls.c | 4 +--
arch/arm64/kvm/sys_regs.c | 28 +++++++++++++----
3 files changed, 87 insertions(+), 8 deletions(-)
diff --git a/arch/arm64/kvm/guest.c b/arch/arm64/kvm/guest.c
index c3ca369882273..ffe3f5ce4b3cf 100644
--- a/arch/arm64/kvm/guest.c
+++ b/arch/arm64/kvm/guest.c
@@ -73,6 +73,25 @@ static u64 core_reg_offset_from_id(u64 id)
return id & ~(KVM_REG_ARCH_MASK | KVM_REG_SIZE_MASK | KVM_REG_ARM_CORE);
}
+static bool kvm_realm_validate_core_reg(u64 off)
+{
+ /*
+ * Note that GPRs can only sometimes be controlled by the VMM.
+ * For PSCI only X0-X6 are used, higher registers are ignored (restored
+ * from the REC).
+ * For HOST_CALL all of X0-X30 are copied to the RsiHostCall structure.
+ * For emulated MMIO X0 is always used.
+ * PC can only be set before the realm is activated.
+ */
+ switch (off) {
+ case KVM_REG_ARM_CORE_REG(regs.regs[0]) ...
+ KVM_REG_ARM_CORE_REG(regs.regs[30]):
+ case KVM_REG_ARM_CORE_REG(regs.pc):
+ return true;
+ }
+ return false;
+}
+
static int core_reg_size_from_offset(const struct kvm_vcpu *vcpu, u64 off)
{
int size;
@@ -553,6 +572,9 @@ static int copy_core_reg_indices(const struct kvm_vcpu *vcpu,
u64 reg = KVM_REG_ARM64 | KVM_REG_ARM_CORE | i;
int size = core_reg_size_from_offset(vcpu, i);
+ if (vcpu_is_rec(vcpu) && !kvm_realm_validate_core_reg(i))
+ continue;
+
if (size < 0)
continue;
@@ -598,6 +620,9 @@ static unsigned long num_sve_regs(const struct kvm_vcpu *vcpu)
if (!vcpu_has_sve(vcpu))
return 0;
+ if (kvm_vm_is_realm(vcpu->kvm))
+ return 1; /* KVM_REG_ARM64_SVE_VLS */
+
if (!kvm_arm_vcpu_sve_finalized(vcpu))
return 1; /* KVM_REG_ARM64_SVE_VLS */
@@ -625,6 +650,10 @@ static int copy_sve_reg_indices(const struct kvm_vcpu *vcpu,
return -EFAULT;
++num_regs;
+ /* For Realms only support SVE_VLS */
+ if (kvm_vm_is_realm(vcpu->kvm))
+ return num_regs;
+
if (!kvm_arm_vcpu_sve_finalized(vcpu))
return num_regs;
@@ -705,6 +734,11 @@ int kvm_arm_get_reg(struct kvm_vcpu *vcpu, const struct kvm_one_reg *reg)
if ((reg->id & ~KVM_REG_SIZE_MASK) >> 32 != KVM_REG_ARM64 >> 32)
return -EINVAL;
+ /*
+ * We don't filter out the register reads for Realms, like we do for
+ * the user writes. We expose junk data for the VMM instead of
+ * denying the requests.
+ */
switch (reg->id & KVM_REG_ARM_COPROC_MASK) {
case KVM_REG_ARM_CORE: return get_core_reg(vcpu, reg);
case KVM_REG_ARM_FW:
@@ -716,12 +750,41 @@ int kvm_arm_get_reg(struct kvm_vcpu *vcpu, const struct kvm_one_reg *reg)
return kvm_arm_sys_reg_get_reg(vcpu, reg);
}
+#define KVM_REG_ARM_ID_AA64DFR0_EL1 ARM64_SYS_REG(3, 0, 0, 5, 0)
+/*
+ * The RMI ABI only enables setting some GPRs and PC. The selection of GPRs
+ * that are available depends on the Realm state and the reason for the last
+ * exit. All other registers are reset to architectural or otherwise defined
+ * reset values by the RMM, except for a few configuration fields that
+ * correspond to Realm parameters.
+ */
+static bool validate_realm_set_reg(struct kvm_vcpu *vcpu,
+ const struct kvm_one_reg *reg)
+{
+ if ((reg->id & KVM_REG_ARM_COPROC_MASK) == KVM_REG_ARM_CORE) {
+ u64 off = core_reg_offset_from_id(reg->id);
+
+ return kvm_realm_validate_core_reg(off);
+ }
+
+ switch (reg->id) {
+ case KVM_REG_ARM_ID_AA64DFR0_EL1:
+ case KVM_REG_ARM64_SVE_VLS:
+ return true;
+ }
+
+ return false;
+}
+
int kvm_arm_set_reg(struct kvm_vcpu *vcpu, const struct kvm_one_reg *reg)
{
/* We currently use nothing arch-specific in upper 32 bits */
if ((reg->id & ~KVM_REG_SIZE_MASK) >> 32 != KVM_REG_ARM64 >> 32)
return -EINVAL;
+ if (kvm_vm_is_realm(vcpu->kvm) && !validate_realm_set_reg(vcpu, reg))
+ return -EINVAL;
+
switch (reg->id & KVM_REG_ARM_COPROC_MASK) {
case KVM_REG_ARM_CORE: return set_core_reg(vcpu, reg);
case KVM_REG_ARM_FW:
diff --git a/arch/arm64/kvm/hypercalls.c b/arch/arm64/kvm/hypercalls.c
index b11b8821c9fbc..2b1e6fdeb4d5c 100644
--- a/arch/arm64/kvm/hypercalls.c
+++ b/arch/arm64/kvm/hypercalls.c
@@ -414,14 +414,14 @@ void kvm_arm_teardown_hypercalls(struct kvm *kvm)
int kvm_arm_get_fw_num_regs(struct kvm_vcpu *vcpu)
{
- return ARRAY_SIZE(kvm_arm_fw_reg_ids);
+ return vcpu_is_rec(vcpu) ? 0 : ARRAY_SIZE(kvm_arm_fw_reg_ids);
}
int kvm_arm_copy_fw_reg_indices(struct kvm_vcpu *vcpu, u64 __user *uindices)
{
int i;
- for (i = 0; i < ARRAY_SIZE(kvm_arm_fw_reg_ids); i++) {
+ for (i = 0; i < kvm_arm_get_fw_num_regs(vcpu); i++) {
if (put_user(kvm_arm_fw_reg_ids[i], uindices++))
return -EFAULT;
}
diff --git a/arch/arm64/kvm/sys_regs.c b/arch/arm64/kvm/sys_regs.c
index 44aae52c473d7..47ebde943a09e 100644
--- a/arch/arm64/kvm/sys_regs.c
+++ b/arch/arm64/kvm/sys_regs.c
@@ -5638,18 +5638,18 @@ int kvm_arm_sys_reg_set_reg(struct kvm_vcpu *vcpu, const struct kvm_one_reg *reg
sys_reg_descs, ARRAY_SIZE(sys_reg_descs));
}
-static unsigned int num_demux_regs(void)
+static inline unsigned int num_demux_regs(struct kvm_vcpu *vcpu)
{
- return CSSELR_MAX;
+ return vcpu_is_rec(vcpu) ? 0 : CSSELR_MAX;
}
-static int write_demux_regids(u64 __user *uindices)
+static int write_demux_regids(struct kvm_vcpu *vcpu, u64 __user *uindices)
{
u64 val = KVM_REG_ARM64 | KVM_REG_SIZE_U32 | KVM_REG_ARM_DEMUX;
unsigned int i;
val |= KVM_REG_ARM_DEMUX_ID_CCSIDR;
- for (i = 0; i < CSSELR_MAX; i++) {
+ for (i = 0; i < num_demux_regs(vcpu); i++) {
if (put_user(val | i, uindices))
return -EFAULT;
uindices++;
@@ -5693,11 +5693,27 @@ static bool copy_reg_to_user(const struct sys_reg_desc *reg, u64 __user **uind)
return true;
}
+static inline bool kvm_realm_sys_reg_hidden_user(const struct kvm_vcpu *vcpu,
+ u64 reg)
+{
+ if (!vcpu_is_rec(vcpu))
+ return false;
+
+ switch (reg) {
+ case SYS_ID_AA64DFR0_EL1:
+ return false;
+ }
+ return true;
+}
+
static int walk_one_sys_reg(const struct kvm_vcpu *vcpu,
const struct sys_reg_desc *rd,
u64 __user **uind,
unsigned int *total)
{
+ if (kvm_realm_sys_reg_hidden_user(vcpu, reg_to_encoding(rd)))
+ return 0;
+
/*
* Ignore registers we trap but don't save,
* and for which no custom user accessor is provided.
@@ -5735,7 +5751,7 @@ static int walk_sys_regs(struct kvm_vcpu *vcpu, u64 __user *uind)
unsigned long kvm_arm_num_sys_reg_descs(struct kvm_vcpu *vcpu)
{
- return num_demux_regs()
+ return num_demux_regs(vcpu)
+ walk_sys_regs(vcpu, (u64 __user *)NULL);
}
@@ -5748,7 +5764,7 @@ int kvm_arm_copy_sys_reg_indices(struct kvm_vcpu *vcpu, u64 __user *uindices)
return err;
uindices += err;
- return write_demux_regids(uindices);
+ return write_demux_regids(vcpu, uindices);
}
#define KVM_ARM_FEATURE_ID_RANGE_INDEX(r) \
--
2.43.0
^ permalink raw reply [flat|nested] 21+ messages in thread