From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pg1-f200.google.com (mail-pg1-f200.google.com [209.85.215.200]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 04248403150 for ; Tue, 11 Aug 2026 17:12:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.200 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786468331; cv=none; b=Ssuv8EdTCJSWGxTbCnNbJEP9uEU95sgpcL2YD+FI4qyEi19SU3GCoZHObcoYzaT6seKpLqvtJe3dT+hAxt3feQcE0G8q+F1nhwJBCyIf5zLPIya8NXSocQJHUya9LDfFZvXViXQOx+2ynuWME5pHktsRK3sYa8E5A9FEYN+GF2k= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786468331; c=relaxed/simple; bh=rRw18ZT/4KdoRFnzYNEvxKXNcU6Gsa6GEcOC2wP6zCA=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=jez8LWTvS47YiCSmo5cJ1OOtIlr+TCNIRziORH3bj5bjzmVtEQ7Z7twhQ7Px+fBdV0CjVK82JERgu6RR31YBW+H6/43fVCmcx+6z76AIllcy/lOD81dKsXTkejiMO1NvCDXqpf2nOKPURW1GcHLSFezyICgvtkSkMiZoh1YQNDg= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=G18KGCS7; arc=none smtp.client-ip=209.85.215.200 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="G18KGCS7" Received: by mail-pg1-f200.google.com with SMTP id 41be03b00d2f7-c96b4f58ddcso2703243a12.3 for ; Tue, 11 Aug 2026 10:12:09 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1786468329; x=1787073129; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=hDmw0Bd7Mgi6QHj5D0rabOEXrT9zZ2xMuzlP2EIUA30=; b=G18KGCS7J32T4BHEdRlyG29h9HkMK556Iu4OxYJd/lp1uIidcFj4RpnnduvvaiXxIR VMFSxyD+lW0CGsj/+yG3lsepXr0BDgYdNv8Gyjd2llcTIr7Vy6huDFopvV/RKpFlW436 u8OF1doMvrfSIuQOD/rLvjD1vRzP6zHr0UnSeTLZ9qfRCLAf/1w7OuWC3CTjQ6h3E1is 9deyptkpeu6A5DDCMkHSXsyVvs5Ji/ezNM2IJR/kzRxniiCNtvEgjIJceM+w5HOAgtpe lBSbEQ+QJbd7LcwFgDVLlGKqUOAxU7NAekFGc/9psR65J7gOWxt3RmilUWcc70iNsVNO EJsQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786468329; x=1787073129; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=hDmw0Bd7Mgi6QHj5D0rabOEXrT9zZ2xMuzlP2EIUA30=; b=mOOovylbvBQQ057ERTl+bqimNm102gZs9UsUwD2jrS0xc960udMaTQO7l1WPj0P+US eu5WviDlJANLjMKqDNgSe4Mtm/Nb4Uh8KAv84fXrCqBbBUA34P6cBNJHEAX8zwXG6KnP KH7nj42NdQLvWaKoWsbWZwadF0JqPX6/xqPT+pwH5AAF7YbsZmJbnz62ncRql/FoDUKF CbUNvtGcX1EC38HS4YDOnCI4GIbmGM5unQdf/EDsXL0rc49cxoecJCjIr3ChBDToRUcf 5G1gAfU/li57MeSiyLYpN8krXCxm6ckZFSK/GYjW6/YBnaHEIvoU3dkH0+3X0IMvbOv5 Zy2A== X-Forwarded-Encrypted: i=1; AHgh+Rqkmz1x9VXEBwZeGGtcZNGB50QBoamNPRoKqkRD1yDsj9hqjU0sCWoVCrn7yUPbYBQ4M/dpaWxaoWS4Ru0=@vger.kernel.org X-Gm-Message-State: AOJu0YwprhqKIwZ0MrbO/bnGDkAMqTt2VohdV3oZc5Z/J1I3mr8s9Tpe Wts6wjcvWsIvcNYjA4qc3+o1pyCgoYefhUDuzs87PnzGv9Fn7ZsXKP5s98L/bULcf89Z5EILbUQ ZPesgUA== X-Received: from pgne19.prod.google.com ([2002:a63:7453:0:b0:cbe:dff5:6872]) (user=seanjc job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a21:6003:b0:3b4:71c1:ab29 with SMTP id adf61e73a8af0-3cc2ba22d6amr7324765637.16.1786468329105; Tue, 11 Aug 2026 10:12:09 -0700 (PDT) Date: Tue, 11 Aug 2026 10:12:08 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260806214050.78058-1-seanjc@google.com> <20260806214050.78058-5-seanjc@google.com> Message-ID: Subject: Re: [PATCH 4/4] KVM: x86/mmu: Add sanity check to detect stale page faults in "map private PFN" From: Sean Christopherson To: Yan Zhao Cc: Paolo Bonzini , kvm@vger.kernel.org, linux-kernel@vger.kernel.org, Kai Huang , Rick Edgecombe , Sashiko Bot Content-Type: text/plain; charset="us-ascii" On Tue, Aug 11, 2026, Yan Zhao wrote: > On Thu, Aug 06, 2026 at 02:40:50PM -0700, Sean Christopherson wrote: > > Harden the "map private PFN" flow against potentially-fatal bugs or future > > KVM changes by checking for a stale "fault" prior to actually mapping the > > PFN into the guest. While it should be impossible for the "page fault" to > > become stale, the sanity check is cheap, whereas a broken assumption would > > have a high probability of leading to a guest-expoitable use-after-free. > > > > Snapshot the invalidation sequence after acquiring mmu_lock to avoid false > > positives, even though doing so completely voids anys and all protection > > against unexpected invalidations. Pretty much the entire point of > > kvm_tdp_mmu_map_private_pfn() is that it allows mapping a PFN that was > > gifted by the caller, i.e. the caller would have to mess up its one and > > only responsibility. > > > > Signed-off-by: Sean Christopherson > > --- > > arch/x86/kvm/mmu/mmu.c | 10 ++++++++++ > > 1 file changed, 10 insertions(+) > > > > diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c > > index 379f570ef04f..76e3cd717324 100644 > > --- a/arch/x86/kvm/mmu/mmu.c > > +++ b/arch/x86/kvm/mmu/mmu.c > > @@ -5210,6 +5210,16 @@ int kvm_tdp_mmu_map_private_pfn(struct kvm_vcpu *vcpu, gfn_t gfn, kvm_pfn_t pfn) > > */ > > WARN_ON_ONCE(kvm_test_request(KVM_REQ_MMU_FREE_OBSOLETE_ROOTS, vcpu)); > > > > + /* > > + * Snapshot the invalidation sequence counter after acquiring > > + * mmu_lock, as guest_memfd guarantees the validity of the pfn, > > + * i.e. any concurrent invalidations are guaranteed to be > > + * irrelevant. > > + */ > > + fault.mmu_seq = vcpu->kvm->mmu_invalidate_seq; > Could you explain more about the conditions under which a fault is stale while > guest_memfd guarantees the validity of the pfn? > > Given that kvm_tdp_mmu_map_private_pfn() already asserts holding slots_lock and > invalidate_lock, I can't think of one. A non-guest_memfd mmu_notifier invalidation bumps mmu_invalidate_seq. And because the range-based invalidation checks are deliberately coarse, in-flight invalidations could also trigger a false positive if GFNs N and N+2 are being invalidate, while kvm_tdp_mmu_map_private_pfn() is trying to map N+1. > If we add the sanity check because it's cheap, why don't we save > fault.mmu_seq before getting the pfn to guard against stale pfn as well? Because as the changelog says, the whole point of this API is to provide a stable PFN. I want to add an is_page_fault_stale() check as defense-in-depth against bugs, and against future changes, e.g. if we extend is_page_fault_stale() to cover more reason why a page fault can become stale. The downside is that is_page_fault_stale() is susceptible to false positives; the funky code here is to minimize the chances of a false positives, while maintaining a reasonable level of safety.