mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Marc Zyngier <maz@kernel.org>
To: Fuad Tabba <fuad.tabba@linux.dev>
Cc: oupton@kernel.org, kvmarm@lists.linux.dev,
	linux-arm-kernel@lists.infradead.org,
	linux-kernel@vger.kernel.org, catalin.marinas@arm.com,
	will@kernel.org, joey.gouly@arm.com, seiden@linux.ibm.com,
	suzuki.poulose@arm.com, yuzenghui@huawei.com,
	mark.rutland@arm.com, steven.price@arm.com,
	vdonnefort@google.com, qperret@google.com, tabba@google.com
Subject: Re: [PATCH v3 14/18] KVM: arm64: Add per-EC entry/exit state marshalling for protected guests
Date: Wed, 16 Sep 2026 17:27:04 +0100	[thread overview]
Message-ID: <86bj9x5fhj.wl-maz@kernel.org> (raw)
In-Reply-To: <20260914113338.159227-15-fuad.tabba@linux.dev>

On Mon, 14 Sep 2026 12:33:34 +0100,
Fuad Tabba <fuad.tabba@linux.dev> wrote:
> 
> Move a protected guest's state between the hyp vCPU and the host per
> exception class instead of copying the whole context. Add
> entry_hyp_pvm_handlers[] and exit_hyp_pvm_handlers[] for WFx, SYS64,
> IABT, DABT and HVC64, and route protected guests through them: on exit
> each handler copies out only what its class needs, and on re-entry
> only what the host may have changed: the PC update or exception it
> requested (PC_UPDATE_REQ), the value of a read it emulated, and a
> forwarded PSCI call's return value. entry_hyp_vm_handlers[] is
> removed: a non-protected vCPU's iflags are copied wholesale.
> 
> The host's fault view is EL2's own syndrome with the guest register
> index withheld, the deferred SError syndrome (DISR_EL1), and the
> addresses each class needs. The value of a written register is passed
> in r0. MMIO data is clamped to the access width and, for a load,
> sign-extended at EL2 from EL2's syndrome. Endianness stays with the
> host. A data abort's PC update is taken for a completed MMIO access,
> and for a cache maintenance operation the host skips on unbacked
> memory, as it does for any guest.
> 
> The host's copy of a protected vCPU's PSTATE is a view the guest never
> runs from: its reset value has PSTATE.A set, and EL2 sets PSTATE.A in
> the mode it copies out regardless of the guest's, so, unless the VMM
> wrote the host copy's PSTATE or SCTLR2_EL1 before the first run,
> serror_is_masked() is true and kvm_inject_serror_esr() pends a
> host-injected SError through HCR_EL2.VSE, masked by the guest's own
> PSTATE.A, instead of emulating the entry on that copy.
> 
> Neither dispatch runs for a trap taken with an SError pending: EL2
> doesn't handle it, and the guest replays it once the host has
> injected the SError. The exit handlers would otherwise marshal a trap
> EL2 never handled, and handle_pvm_exit_hvc64() would panic on an
> unfiltered function id.
> 
> For HVC64, only the PSCI calls EL2 forwards reach the host: the exit
> handler passes the function id and the arguments each call needs, and
> the entry handler returns the host's result. A CPU_ON the host failed
> is rolled back to OFF and returned as INTERNAL_FAILURE, or as
> ALREADY_ON when that's what the host returned: PSCI defines it as the
> retry signal for a CPU_ON that reaches the implementation before the
> target's CPU_OFF has been processed (DEN0022 section 6.6), which a
> guest that doesn't poll AFFINITY_INFO first can do. A target that
> already ran returns SUCCESS. The rollback leaves the published reset
> state in place: clearing it races a fresh CPU_ON's publication and
> wedges the target at ON_PENDING, and the entry point left behind is one
> the guest supplied. The rollback's cmpxchg carries no generation, so
> one that lands after a later CPU_ON has republished ON_PENDING cancels
> that cycle too. The target is then OFF at EL2 while the host has it
> runnable, and its CPU_ONs return ALREADY_ON until the VMM stops it
> again.
> 
> Signed-off-by: Fuad Tabba <fuad.tabba@linux.dev>
> ---
>  arch/arm64/kvm/hyp/nvhe/hyp-main.c | 430 +++++++++++++++++++++++++++--
>  1 file changed, 407 insertions(+), 23 deletions(-)
> 
> diff --git a/arch/arm64/kvm/hyp/nvhe/hyp-main.c b/arch/arm64/kvm/hyp/nvhe/hyp-main.c
> index 1dcc75261dc08..3d59c4827c42a 100644
> --- a/arch/arm64/kvm/hyp/nvhe/hyp-main.c
> +++ b/arch/arm64/kvm/hyp/nvhe/hyp-main.c
> @@ -4,6 +4,8 @@
>   * Author: Andrew Scull <ascull@google.com>
>   */
>  
> +#include <kvm/arm_hypercalls.h>
> +
>  #include <hyp/adjust_pc.h>
>  #include <hyp/switch.h>
>  
> @@ -34,13 +36,366 @@ void __kvm_hyp_host_forward_smc(struct kvm_cpu_context *host_ctxt);
>  
>  typedef void (*hyp_entry_exit_handler_fn)(struct pkvm_hyp_vcpu *);
>  
> -static void handle_vm_entry_generic(struct pkvm_hyp_vcpu *hyp_vcpu)
> +static bool pvm_sys64_is_write(u64 esr)
>  {
> -	vcpu_copy_flag(&hyp_vcpu->vcpu, hyp_vcpu->host_vcpu, PC_UPDATE_REQ);
> +	return (esr & ESR_ELx_SYS64_ISS_DIR_MASK) == ESR_ELx_SYS64_ISS_DIR_WRITE;
>  }
>  
> -static const hyp_entry_exit_handler_fn entry_hyp_vm_handlers[] = {
> -	[0 ... ESR_ELx_EC_MAX]		= handle_vm_entry_generic,
> +static void handle_pvm_entry_wfx(struct pkvm_hyp_vcpu *hyp_vcpu)
> +{
> +	struct kvm_vcpu *host_vcpu = hyp_vcpu->host_vcpu;
> +
> +	/* Exceptions have priority; the host injects none on WFx. */
> +	if (vcpu_get_flag(host_vcpu, PENDING_EXCEPTION))
> +		return;
> +
> +	if (vcpu_get_flag(host_vcpu, INCREMENT_PC)) {
> +		vcpu_clear_flag(&hyp_vcpu->vcpu, PC_UPDATE_REQ);
> +		kvm_incr_pc(&hyp_vcpu->vcpu);
> +	}
> +}
> +
> +static void handle_pvm_entry_sys64(struct pkvm_hyp_vcpu *hyp_vcpu)
> +{
> +	struct kvm_vcpu *host_vcpu = hyp_vcpu->host_vcpu;
> +	bool pc_update;
> +
> +	/* Exceptions have priority over anything else */
> +	if (vcpu_get_flag(host_vcpu, PENDING_EXCEPTION)) {
> +		/* A host-requested exception on SYS64 is always an UNDEF. */
> +		u32 esr = (ESR_ELx_EC_UNKNOWN << ESR_ELx_EC_SHIFT) | ESR_ELx_IL;
> +
> +		__vcpu_assign_sys_reg(&hyp_vcpu->vcpu, ESR_EL1, esr);
> +		kvm_pend_exception(&hyp_vcpu->vcpu, EXCEPT_AA64_EL1_SYNC);
> +		return;
> +	}
> +
> +	/* Handle PC increment on a host-emulated access */
> +	pc_update = vcpu_get_flag(host_vcpu, INCREMENT_PC);
> +	if (pc_update) {
> +		vcpu_clear_flag(&hyp_vcpu->vcpu, PC_UPDATE_REQ);
> +		kvm_incr_pc(&hyp_vcpu->vcpu);
> +	}
> +
> +	/* If the host emulated a read access, update the register */
> +	if (pc_update &&
> +	    !pvm_sys64_is_write(hyp_vcpu->vcpu.arch.fault.esr_el2)) {
> +		/* r0 as transfer register between the guest and the host. */
> +		u64 rt_val = READ_ONCE(host_vcpu->arch.ctxt.regs.regs[0]);

Why isn't this

		rt_val = READ_ONCE(vcpu_gp_regs(host_vcpu)->regs[0]);

and similarly everywhere else?

[...]

> +static void handle_pvm_exit_sys64(struct pkvm_hyp_vcpu *hyp_vcpu)
> +{
> +	struct kvm_vcpu *host_vcpu = hyp_vcpu->host_vcpu;
> +	u32 esr_el2 = hyp_vcpu->vcpu.arch.fault.esr_el2;
> +
> +	/* The mode is required for the host to emulate some sysregs */
> +	host_vcpu->arch.ctxt.regs.pstate =
> +		pvm_host_pstate(hyp_vcpu->vcpu.arch.ctxt.regs.pstate);
> +
> +	/* r0 as transfer register between the guest and the host. */
> +	if (pvm_sys64_is_write(esr_el2)) {
> +		int rt = kvm_vcpu_sys_get_rt(&hyp_vcpu->vcpu);
> +		u64 rt_val = vcpu_get_reg(&hyp_vcpu->vcpu, rt);
> +
> +		host_vcpu->arch.ctxt.regs.regs[0] = rt_val;

and this should be

		vcpu_set_reg(host_vcpu, 0, rt_val);

assuming you don't need a WRITE_ONCE() to match the READ_ONCE() in the
other direction.

> +	}
> +}
> +
> +static void handle_pvm_exit_iabt(struct pkvm_hyp_vcpu *hyp_vcpu)
> +{
> +	hyp_vcpu->host_vcpu->arch.fault.hpfar_el2 =
> +		hyp_vcpu->vcpu.arch.fault.hpfar_el2;

Please keep assignments on a single line (everywhere).

Thanks,

	M.

-- 
Without deviation from the norm, progress is not possible.

  reply	other threads:[~2026-09-16 16:27 UTC|newest]

Thread overview: 29+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-14 11:33 [PATCH v3 00/18] KVM: arm64: Confine protected VM vCPU state to EL2 Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 01/18] KVM: arm64: Sync HCR_EL2.VSE back to the host vCPU under pKVM Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 02/18] KVM: arm64: Validate the host vCPU's VM before reading it " Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 03/18] KVM: arm64: Pin the host vCPU before adjusting its PC " Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 04/18] KVM: arm64: Disable steal time for protected VMs Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 05/18] KVM: arm64: Introduce per-EC entry handlers for pKVM Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 06/18] KVM: arm64: Skip fixed-feature state flush for protected vCPUs Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 07/18] KVM: arm64: Add {flush,sync}_hyp_timer_state() primitives Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 08/18] KVM: arm64: Add system register reset framework for protected VMs Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 09/18] KVM: arm64: Implement HVC handling for protected guests at EL2 Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 10/18] KVM: arm64: Handle PSCI calls for protected VMs " Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 11/18] KVM: arm64: Restrict KVM_ARM_VCPU_INIT and PSCI version for protected VMs Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 12/18] KVM: arm64: Prevent host PC adjustments for protected vCPUs Fuad Tabba
2026-09-14 13:42   ` Marc Zyngier
2026-09-14 14:43     ` Fuad Tabba
2026-09-15 11:02       ` Marc Zyngier
2026-09-15 11:19         ` Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 13/18] KVM: arm64: Inject an UNDEF at EL2 for unhandled protected guest exits Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 14/18] KVM: arm64: Add per-EC entry/exit state marshalling for protected guests Fuad Tabba
2026-09-16 16:27   ` Marc Zyngier [this message]
2026-09-16 19:05     ` Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 15/18] KVM: arm64: Reject host access to protected VM private state Fuad Tabba
2026-09-16 16:30   ` Marc Zyngier
2026-09-16 19:07     ` Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 16/18] KVM: arm64: Reject host power-on of a vCPU that EL2 holds powered off Fuad Tabba
2026-09-16 16:43   ` Marc Zyngier
2026-09-16 19:08     ` Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 17/18] KVM: arm64: Advertise the capabilities that protected VMs support Fuad Tabba
2026-09-14 11:33 ` [PATCH v3 18/18] KVM: arm64: Document the protected VM userspace API Fuad Tabba

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=86bj9x5fhj.wl-maz@kernel.org \
    --to=maz@kernel.org \
    --cc=catalin.marinas@arm.com \
    --cc=fuad.tabba@linux.dev \
    --cc=joey.gouly@arm.com \
    --cc=kvmarm@lists.linux.dev \
    --cc=linux-arm-kernel@lists.infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mark.rutland@arm.com \
    --cc=oupton@kernel.org \
    --cc=qperret@google.com \
    --cc=seiden@linux.ibm.com \
    --cc=steven.price@arm.com \
    --cc=suzuki.poulose@arm.com \
    --cc=tabba@google.com \
    --cc=vdonnefort@google.com \
    --cc=will@kernel.org \
    --cc=yuzenghui@huawei.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®