mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Leonardo Bras <leo.bras@arm.com>
To: Oliver Upton <oupton@kernel.org>
Cc: Leonardo Bras <leo.bras@arm.com>, Marc Zyngier <maz@kernel.org>,
	Fuad Tabba <fuad.tabba@linux.dev>,
	Joey Gouly <joey.gouly@arm.com>,
	Steffen Eiden <seiden@linux.ibm.com>,
	Suzuki K Poulose <suzuki.poulose@arm.com>,
	Zenghui Yu <yuzenghui@huawei.com>,
	Catalin Marinas <catalin.marinas@arm.com>,
	Will Deacon <will@kernel.org>,
	Mark Rutland <mark.rutland@arm.com>,
	Raghavendra Rao Ananta <rananta@google.com>,
	Tian Zheng <zhengtian10@huawei.com>,
	linux-arm-kernel@lists.infradead.org, kvmarm@lists.linux.dev,
	linux-kernel@vger.kernel.org
Subject: Re: [RFC PATCH 5/5] KVM: arm64: Enable HAFDBS for guests not on migration
Date: Thu, 17 Sep 2026 14:40:42 +0100	[thread overview]
Message-ID: <aqvt2gnlBZeNsaf4@LeoBrasDK> (raw)
In-Reply-To: <aqsl1_IqnV2r8kUL@kernel.org>

On Wed, Sep 16, 2026 at 04:27:19PM -0700, Oliver Upton wrote:
> On Wed, Sep 16, 2026 at 03:00:46PM +0100, Leonardo Bras wrote:
> > On Tue, Sep 15, 2026 at 05:10:02PM -0700, Oliver Upton wrote:
> > > Hi,
> > > 
> > > On Tue, Sep 01, 2026 at 06:15:56PM +0100, Leonardo Bras wrote:
> > > > When dirty-logging is disabled, even non-write faults make a page dirty,
> > > > which avoids a second fault when the page is actually written to.
> > > > 
> > > > On dirty-logging enable, this approach causes all (writable) pages on the
> > > > memslot to be marked clean, even if they were not written to, which can
> > > > take a lot of time, while holding the MMU lock, doing atomic writes to
> > > > PTEs.
> > > 
> > > Do you have any performance numbers for this? Enabling HAFDBS seems a
> > > bit involved to avoid some stores on the first pass.
> > 
> > Not yet, but if the idea does not look too crazy I can find hardware and 
> > collect some data :)
> 
> TBH this looks like a micro-optimization so I'm not expecting the
> performance gains to justify the behavior change.
> 

Okay then, I will get a machine to figure the performance gains :)

> > > > 						      has_vhe() &&
> > > 
> > > I don't see a reason why this needs to be constrained to VHE-only.
> > > 
> > 
> > Humm, in nVHE would not the host kernel run in EL1?
> > I thought that this being a feature that depends on EL2 registers host 
> > would need to be in EL2 to make use of it.
> > 
> > That being said, I understand very little of how this works, so I 
> > constrained to VHE only at the start.
> > 
> > Would this work in nVHE?
> 
> We already pass a stage-2 MMU configuration to EL2 from EL1 in nVHE and
> hVHE. How is this any different?
> 

Ah, so we can just ask the hypervisor in EL2 to enable that behavior in 
those cpus? I have dig deeper in code to get a better understanding on how 
that would work.

Would that make sense in pKVM as well?

> I'm not opposed to making features VHE-only, but there needs to be some
> amount of reasoning to justify it.

Likewise, I am not opposed to make it work outside VHE as well, I just 
don't understand enough of nVHE/hVHE/pKVM yet, so I thought starting with 
VHE would be simpler.

> 
> > > > +		!kvm_vcpu_has_nv(kvm) && cpus_have_final_cap(ARM64_HW_DBM);
> > > 
> > > Same thing goes for nested... KVM can make use of HAFDBS in the
> > > canonical stage-2 MMU (or even a shadow stage-2) independent of the
> > > guest hypervisor.
> > > 
> > 
> > Humm, I remember reaching the conclusion that it could not be used if the 
> > guest supported NV. Let's say:
> > 
> > L0 - Host	- Has HAFDBS enabled
> > L1 - Hypervisor - Has HAFDBS disabled
> > L2 - Guest	-
> > 
> > Let's say guest writes to a page, and the shadow S2 has DBM=1, so it's 
> > marked as WD by HAFDBS. Since no fault was taken, how would the L1 be able 
> > to update it's S2 pagetables to mark the page dirty?
> > 
> > (We would have to transverse the Shadow S2 Pagetable updating the original 
> > S2 pagetable)

(*)

> > 
> > I was wondering, thought, that we could emulate it in the last level 
> > hypervisor, if it's guest does not support nested guests. That would mean 
> > we can have the last-1 level hypervisor to update the S2 pagetable on the 
> > last level hypervisor without it having to fault. Ex:
> > 
> > L0 Host - HAFDBS disabled
> > [...]
> > Ln-1 Hypervisor - HAFDBS disabled
> > Ln   Hypervisor - HAFDBS enabled
> > Ln+1 Guest - No E2H feature
> > 
> > When the guest writes to a page, the host should receive a fault, that IIUC 
> > have to propagate down up to Ln Hyp. If Ln Hyp has HAFDBS, we could skip 
> > injecting a fault in Ln Hyp, as  Ln-1 Hyp could emulate HAFDBS and write 
> > the dirty bit to S2 pagetagle of Ln+1 guest, that resides in Ln memory.
> > 
> > Not sure if the troulbe would be worth, though.
> > Does it make sense?
> 
> I'm not following your reasoning here.

Oh, sorry, that got confusing up there. Everything above (*) was my 
perception of why would not that work. Lines below (*) are a new idea on 
how we could use it in the last level hypervisor.

> Treat the shadow stage-2 MMU as a
> TLB; that TLB is filled with a writable translation when S2AP[1]=1 in the
> L1 translation.

L1 considering that L1 is the last level hypervisor, right? So the example 
above the (*). 

If so, yeah it makes sense. Shadow S2 acs as a TLB so we don't have to 
walk all intermediate pagetables to figure out the translation every time 
we need to access it and the entry is not in the HW TLB.

Thanks for sharing this... thinking about the shadow S2 just became easier 
:)

> 
> The dirty state of the pseudo-TLB is completely internal. You could then
> layer HAFDBS for the L1 translation on top of this (which we don't
> support) by potentially relaxing the descriptor _before_ evaluating the
> resulting permissions. You'd then take write permission faults to set
> S2AP[1] in the L1 translation.
> 

Would relaxing here mean mark the shadow S2 PTE as RO?
I think you got what I meant on the idea below (*), but exemplified with 
the example above (*).

What I meant in the idea below the (*) is that the level (n-1) hypervisor 
could just flip the dirty-bit for it's level n guest (which is the last 
level  hypervisor), emulating a hardware that supports HAFDBS. That would 
avoid injecting the fault in the level n guest and just flip the bit 
directly from level (n-1) viewpoint. For that, only the Ln hypervisor 
should have HAFDBS enabled. (this idea does not depend on the 
underlying hardware actually supporting HAFDBS).

Does it make sense, then?

> > > The name would suggest this thing takes a vcpu pointer...
> > > 
> > 
> > Ah, that name was based on
> > #define kvm_vcpu_has_feature(k, f)  __vcpu_has_feature(&(k)->arch, #(f))
> > 
> > That takes a kvm struct to check the kvm_arch one, instead of looking into 
> > the vcpu. I did it like this because there were some scenarios it was not 
> > quite straightforward to get the vcpu to use vcpu_has_nv(), which takes a 
> > vcpu.
> 
> This thing probably should've been "kvm_has_vcpu_feature()" or similar to
> massage the expected typing.

Humm, makes sense.
Would you like me to change the new one to kvm_has_nv() in the next 
version?

Thanks!
Leo

  reply	other threads:[~2026-09-17 13:41 UTC|newest]

Thread overview: 24+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-01 17:15 [RFC PATCH 0/5] KVM: arm64: New PTE dirty-page encoding, HAFDBS new usage Leonardo Bras
2026-09-01 17:15 ` [RFC PATCH 1/5] KVM: arm64: pgtables: Change write bit from S2AP_W to DBM Leonardo Bras
2026-09-13  9:00   ` Marc Zyngier
2026-09-15 17:12     ` Leonardo Bras
2026-09-16  0:37       ` Oliver Upton
2026-09-16 11:22         ` Leonardo Bras
2026-09-16 12:20           ` Marc Zyngier
2026-09-16 13:25             ` Leonardo Bras
2026-09-16  8:30       ` Marc Zyngier
2026-09-16 13:03         ` Leonardo Bras
2026-09-01 17:15 ` [RFC PATCH 2/5] KVM: arm64: Add KVM_PGTABLE_PROT_DIRTY Leonardo Bras
2026-09-13  9:09   ` Marc Zyngier
2026-09-15 17:33     ` Leonardo Bras
2026-09-01 17:15 ` [RFC PATCH 3/5] KVM: arm64: Introduce a dedicated walker for stage2 write-protect Leonardo Bras
2026-09-01 17:15 ` [RFC PATCH 4/5] KVM: arm64: Add KVM_REQ_RELOAD_STAGE2 Leonardo Bras
2026-09-02  3:41   ` Tian Zheng
2026-09-02 10:53     ` Leonardo Bras
2026-09-01 17:15 ` [RFC PATCH 5/5] KVM: arm64: Enable HAFDBS for guests not on migration Leonardo Bras
2026-09-16  0:10   ` Oliver Upton
2026-09-16 14:00     ` Leonardo Bras
2026-09-16 23:27       ` Oliver Upton
2026-09-17 13:40         ` Leonardo Bras [this message]
2026-09-12 12:24 ` [RFC PATCH 0/5] KVM: arm64: New PTE dirty-page encoding, HAFDBS new usage Marc Zyngier
2026-09-15 15:31   ` Leonardo Bras

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=aqvt2gnlBZeNsaf4@LeoBrasDK \
    --to=leo.bras@arm.com \
    --cc=catalin.marinas@arm.com \
    --cc=fuad.tabba@linux.dev \
    --cc=joey.gouly@arm.com \
    --cc=kvmarm@lists.linux.dev \
    --cc=linux-arm-kernel@lists.infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mark.rutland@arm.com \
    --cc=maz@kernel.org \
    --cc=oupton@kernel.org \
    --cc=rananta@google.com \
    --cc=seiden@linux.ibm.com \
    --cc=suzuki.poulose@arm.com \
    --cc=will@kernel.org \
    --cc=yuzenghui@huawei.com \
    --cc=zhengtian10@huawei.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®