From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pl1-f199.google.com (mail-pl1-f199.google.com [209.85.214.199]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E902D3BBFC2 for ; Wed, 26 Aug 2026 09:18:14 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.199 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787735897; cv=none; b=XdMusU7nDFZRUPhikTi5Ny9Uo8mY+Ph6m3wu+7JtxUOccFRgCpgVIJQs7B1P26hDJKsQRBVEsGZ3PIZyAZcfqrlEpGOb89H81Htw+ipZHipZo4VaalzF7QJ1hYYs9+uIY3WGisN0d3W7moKmf2ZDalcasUp8Y5p9HLWLJGoGK7o= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787735897; c=relaxed/simple; bh=3j4HP/pcwLl+LC7YJJLl650dY7sJQ8xV6gpIzCkDtX4=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=r0QDOB2gUwgqt3OaR8rBv3i9HooZIvjf8bfigdzX4pK2iiLEl2j1vH1OiyG0hbge9l90aMC+GeEbXtFqmDrT+Viz9AZXIP7YWtDZ7YA1HJ09O7krGQMKwJAiHjvUV783gH7HB5GoTg97XQ9FgXyf+WzRhNdM4PNPPrCC+ujaM04= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--ackerleytng.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=wf6T039b; arc=none smtp.client-ip=209.85.214.199 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--ackerleytng.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="wf6T039b" Received: by mail-pl1-f199.google.com with SMTP id d9443c01a7336-2d6df0a1e18so9975395ad.1 for ; Wed, 26 Aug 2026 02:18:14 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787735893; x=1788340693; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=jW6qehcQroMLAfqimWC+/98gYPjKxQ/Xm+gQk2AWdxQ=; b=wf6T039bATc0puYn++erLl+8bxvURYMR0S1aig8HPEFtbxjk4Br7o2b6ZumvYjFTOo JYOJrD0zggbKP5ZzxRlQdgQlG9AgM0amqW0IHSxT7phkhSLw1SuAS9a2fInEJzMUGxCY QgnYIrhwuVoOkd5ZmJR7pE0JnUiO0LlGdpIK7AOpw97785ijEOP7CsOIZxl7h0DnLmPL H8Z+SItjdYIsw3Zazo7vEllD5F/SsqfkF/XWR/9K3AkQBDtZHIKeGVpeZLuBFujFME+l 3ylTKiHdvqMRcyQ/kCmo0eC4RS6wRywJUQTTKbuSei0HkQLBEJnAUkHQW12XAsTIt8n8 Kt0w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787735893; x=1788340693; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=jW6qehcQroMLAfqimWC+/98gYPjKxQ/Xm+gQk2AWdxQ=; b=GNo55cA6uFzPEFDFmKvIJLQ0ZjUvvZRCACwjxWUNHCAIt7O7BF4uL+v0DfWzD20wLs hb0igydQvUP8Al9bxz2awhKYH01w4zLbUuHdQ8Wfd1Gn9sjX2iEG7HrRPEKNdZngxTCN G7XCvBa961ldrcEKYqhIQLSS8Y8zShATwGmrHzK4VOn6IQM/oS5JcZVCG2yQ7GfHh4Ky niDK+5mH4BHbku3Ekal1QkC9aYbZSqP4Fxu8SYZbnKPJ8+f+nUIGe9xIfl1m+B4qobKf hOMvGs1GUU+QLKjRigGalncE/wmCAO9dsQ6sqWq9J3eMaTHjKZvMr7UwP2QQLP1MiaF+ hFpw== X-Forwarded-Encrypted: i=1; AHgh+RpIm04VDb6xjcqnXBN394Xy+Q3rx10iJSc3GsbLfmoaHBXydQdpOUS5d95Bw1fGLtQhW23RsMz8o1JmMCA=@vger.kernel.org X-Gm-Message-State: AFuF++mwCPzdh5QAL2qHLwzrmVyF/PMQven8DdVTEH3JzZ9J7UOZep1w D2B55U21IMfO2hAzX5jYJXVXx61s6NB0dsakUrUb01Ue0B9E/wJJtPAyvnVqZ2z3n0qaHHsWjPB 8WuqJhdy7dwEIiPyQDszac4EE/Q== X-Received: from plbkr3.prod.google.com ([2002:a17:903:803:b0:2d7:1ff:d1ee]) (user=ackerleytng job=prod-delivery.src-stubby-dispatcher) by 2002:a17:903:11c3:b0:2cc:f5aa:9513 with SMTP id d9443c01a7336-2d707bab981mr79455195ad.10.1787735892152; Wed, 26 Aug 2026 02:18:12 -0700 (PDT) Date: Wed, 26 Aug 2026 09:18:01 +0000 In-Reply-To: <20260826-gmem-inplace-conversion-v11-0-0a15d8a799aa@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260826-gmem-inplace-conversion-v11-0-0a15d8a799aa@google.com> X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Developer-Signature: v=1; a=ed25519-sha256; t=1787735885; l=8901; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=SEnzznEA5NUsL5Y4N4re64mzaHrxr3QSDUbuTg+NomA=; b=GhCAeUsftl/MAjjcKvIK0iPyIYi5/ZehgDnv1weoDN6O1ArKkXecL34ytH/RnfWRQ8sPaegrd FSm/zS6foADBWKsZuwKgbLzpHYu20tO75OkT63E6a1AUG8HUylRHSUC X-Mailer: b4 0.16.0 Message-ID: <20260826-gmem-inplace-conversion-v11-3-0a15d8a799aa@google.com> Subject: [PATCH v11 03/46] KVM: guest_memfd: Introduce per-gmem attributes, use to guard user mappings From: Ackerley Tng To: aik@amd.com, andrew.jones@linux.dev, binbin.wu@linux.intel.com, brauner@kernel.org, chao.p.peng@linux.intel.com, david@kernel.org, jmattson@google.com, jthoughton@google.com, michael.roth@amd.com, oupton@kernel.org, pankaj.gupta@amd.com, qperret@google.com, rick.p.edgecombe@intel.com, rientjes@google.com, shivankg@amd.com, steven.price@arm.com, willy@infradead.org, wyihan@google.com, yan.y.zhao@intel.com, forkloop@google.com, pratyush@kernel.org, suzuki.poulose@arm.com, aneesh.kumar@kernel.org, liam@infradead.org, Paolo Bonzini , Sean Christopherson , Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , x86@kernel.org, "H. Peter Anvin" , Steven Rostedt , Masami Hiramatsu , Mathieu Desnoyers , Jonathan Corbet , Shuah Khan , Shuah Khan , Vishal Annapurve , Andrew Morton , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Youngjun Park , Qi Zheng , Shakeel Butt , Kiryl Shutsemau , Baoquan He , Jason Gunthorpe , John Hubbard , Peter Xu , tarunsahu@google.com, Fuad Tabba , Vlastimil Babka Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-mm@kvack.org, linux-coco@lists.linux.dev, Ackerley Tng , Xiaoyao Li Content-Type: text/plain; charset="utf-8" From: Sean Christopherson Start plumbing in guest_memfd support for in-place private<=>shared conversions by tracking attributes via a maple tree. KVM currently tracks private vs. shared attributes on a per-VM basis, which made sense when a guest_memfd _only_ supported private memory, but tracking per-VM simply can't work for in-place conversions as the shared/private status of a given page needs to be per-gmem_inode, not per-VM. Use the filemap invalidation lock to protect the maple tree, as taking the lock for read when faulting in memory (for userspace or the guest) isn't expected to result in meaningful contention, and using a separate lock would add significant complexity (avoiding deadlock is quite difficult). In kvm_gmem_get_pfn(), drop the folio refcount before releasing filemap_invalidate_lock(). This ensures that a competing conversion request from userspace (to be added in a later patch), which also takes the filemap_invalidate_lock(), will never see an elevated refcount due to kvm_gmem_get_pfn(). Co-developed-by: Vishal Annapurve Signed-off-by: Vishal Annapurve Co-developed-by: Fuad Tabba Signed-off-by: Fuad Tabba Signed-off-by: Sean Christopherson Tested-by: Shivank Garg Reviewed-by: Xiaoyao Li Reviewed-by: Binbin Wu Co-developed-by: Ackerley Tng Signed-off-by: Ackerley Tng --- virt/kvm/guest_memfd.c | 136 ++++++++++++++++++++++++++++++++++++++++++------- 1 file changed, 119 insertions(+), 17 deletions(-) diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c index 2c5857a73ec94..154fc9622640d 100644 --- a/virt/kvm/guest_memfd.c +++ b/virt/kvm/guest_memfd.c @@ -4,6 +4,7 @@ #include #include #include +#include #include #include #include @@ -34,6 +35,13 @@ struct gmem_inode { struct list_head gmem_file_list; u64 flags; + /* + * Every index in this inode, whether memory is populated or + * not, is tracked in attributes. The entire range of indices, + * corresponding to the size of this inode, is represented in + * this maple tree. + */ + struct maple_tree attributes; }; static __always_inline struct gmem_inode *GMEM_I(struct inode *inode) @@ -61,9 +69,28 @@ static pgoff_t kvm_gmem_get_index(struct kvm_memory_slot *slot, gfn_t gfn) return gfn - slot->base_gfn + slot->gmem.pgoff; } +static u64 kvm_gmem_get_default_attributes(struct inode *inode) +{ + bool init_shared = GMEM_I(inode)->flags & GUEST_MEMFD_FLAG_INIT_SHARED; + + return init_shared ? 0 : KVM_MEMORY_ATTRIBUTE_PRIVATE; +} + +static u64 kvm_gmem_get_attributes(struct inode *inode, void *entry) +{ + if (WARN_ON_ONCE(!entry)) + return kvm_gmem_get_default_attributes(inode); + + return xa_to_value(entry); +} + static bool kvm_gmem_is_private_mem(struct inode *inode, pgoff_t index) { - return !(GMEM_I(inode)->flags & GUEST_MEMFD_FLAG_INIT_SHARED); + struct maple_tree *mt = &GMEM_I(inode)->attributes; + void *entry = mtree_load(mt, index); + + return kvm_gmem_get_attributes(inode, entry) & + KVM_MEMORY_ATTRIBUTE_PRIVATE; } static bool kvm_gmem_is_shared_mem(struct inode *inode, pgoff_t index) @@ -364,10 +391,13 @@ static vm_fault_t kvm_gmem_fault_user_mapping(struct vm_fault *vmf) if (((loff_t)vmf->pgoff << PAGE_SHIFT) >= i_size_read(inode)) return VM_FAULT_SIGBUS; - if (!kvm_gmem_is_shared_mem(inode, vmf->pgoff)) - return VM_FAULT_SIGBUS; + filemap_invalidate_lock_shared(inode->i_mapping); + if (kvm_gmem_is_shared_mem(inode, vmf->pgoff)) + folio = kvm_gmem_get_folio(inode, vmf->pgoff); + else + folio = ERR_PTR(-EACCES); + filemap_invalidate_unlock_shared(inode->i_mapping); - folio = kvm_gmem_get_folio(inode, vmf->pgoff); if (IS_ERR(folio)) { if (PTR_ERR(folio) == -EAGAIN) return VM_FAULT_RETRY; @@ -520,6 +550,51 @@ bool __weak kvm_arch_supports_gmem_init_shared(struct kvm *kvm) return true; } +static int kvm_gmem_init_inode(struct inode *inode, loff_t size, u64 flags) +{ + struct gmem_inode *gi = GMEM_I(inode); + MA_STATE(mas, &gi->attributes, 0, (size >> PAGE_SHIFT) - 1); + u64 attrs; + int r; + + inode->i_op = &kvm_gmem_iops; + inode->i_mapping->a_ops = &kvm_gmem_aops; + inode->i_mode |= S_IFREG; + inode->i_size = size; + mapping_set_gfp_mask(inode->i_mapping, GFP_HIGHUSER); + + /* + * guest_memfd memory is neither migratable nor swappable: set + * inaccessible to gate off both. + */ + mapping_set_inaccessible(inode->i_mapping); + WARN_ON_ONCE(!mapping_unevictable(inode->i_mapping)); + + gi->flags = flags; + + mt_set_external_lock(&gi->attributes, + &inode->i_mapping->invalidate_lock); + + /* + * Store default attributes for the entire gmem instance. Ensuring every + * index is represented in the maple tree at all times simplifies the + * conversion and merging logic. + */ + attrs = kvm_gmem_get_default_attributes(inode); + + /* + * Acquire the invalidation lock purely to make lockdep happy. The + * maple tree library expects all stores to be protected via the lock, + * and the library can't know when the tree is reachable only by the + * caller, as is the case here. + */ + filemap_invalidate_lock(inode->i_mapping); + r = mas_store_gfp(&mas, xa_mk_value(attrs), GFP_KERNEL); + filemap_invalidate_unlock(inode->i_mapping); + + return r; +} + static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags) { static const char *name = "[kvm-gmem]"; @@ -550,16 +625,9 @@ static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags) goto err_fops; } - inode->i_op = &kvm_gmem_iops; - inode->i_mapping->a_ops = &kvm_gmem_aops; - inode->i_mode |= S_IFREG; - inode->i_size = size; - mapping_set_gfp_mask(inode->i_mapping, GFP_HIGHUSER); - mapping_set_inaccessible(inode->i_mapping); - /* Unmovable mappings are supposed to be marked unevictable as well. */ - WARN_ON_ONCE(!mapping_unevictable(inode->i_mapping)); - - GMEM_I(inode)->flags = flags; + err = kvm_gmem_init_inode(inode, size, flags); + if (err) + goto err_inode; file = alloc_file_pseudo(inode, kvm_gmem_mnt, name, O_RDWR, &kvm_gmem_fops); if (IS_ERR(file)) { @@ -763,9 +831,13 @@ int kvm_gmem_get_pfn(struct kvm *kvm, struct kvm_memory_slot *slot, if (!file) return -EFAULT; + filemap_invalidate_lock_shared(file_inode(file)->i_mapping); + folio = __kvm_gmem_get_pfn(file, slot, index, pfn, max_order); - if (IS_ERR(folio)) - return PTR_ERR(folio); + if (IS_ERR(folio)) { + r = PTR_ERR(folio); + goto out; + } if (!folio_test_uptodate(folio)) { clear_highpage(folio_page(folio, 0)); @@ -780,6 +852,8 @@ int kvm_gmem_get_pfn(struct kvm *kvm, struct kvm_memory_slot *slot, folio_unlock(folio); folio_put(folio); +out: + filemap_invalidate_unlock_shared(file_inode(file)->i_mapping); return r; } EXPORT_SYMBOL_FOR_KVM_INTERNAL(kvm_gmem_get_pfn); @@ -909,6 +983,15 @@ static struct inode *kvm_gmem_alloc_inode(struct super_block *sb) mpol_shared_policy_init(&gi->policy, NULL); + /* + * Memory attributes are protected by the filemap invalidation lock, but + * the lock structure isn't available at this time. Immediately mark + * maple tree as using external locking so that accessing the tree + * before it's fully initialized results in NULL pointer dereferences + * and not more subtle bugs. + */ + mt_init_flags(&gi->attributes, MT_FLAGS_LOCK_EXTERN | MT_FLAGS_USE_RCU); + gi->flags = 0; INIT_LIST_HEAD(&gi->gmem_file_list); return &gi->vfs_inode; @@ -916,7 +999,26 @@ static struct inode *kvm_gmem_alloc_inode(struct super_block *sb) static void kvm_gmem_destroy_inode(struct inode *inode) { - mpol_free_shared_policy(&GMEM_I(inode)->policy); + struct gmem_inode *gi = GMEM_I(inode); + + mpol_free_shared_policy(&gi->policy); + + /* + * Note! Checking for an empty tree is functionally necessary + * to avoid explosions if the tree hasn't been fully + * initialized, i.e. if the inode is being destroyed before + * guest_memfd can set the external lock, lockdep would find + * that the tree's internal ma_lock was not held. + */ + if (!mtree_empty(&gi->attributes)) { + /* + * Acquire the invalidation lock purely to make lockdep happy, + * the inode is unreachable at this point. + */ + filemap_invalidate_lock(inode->i_mapping); + __mt_destroy(&gi->attributes); + filemap_invalidate_unlock(inode->i_mapping); + } } static void kvm_gmem_free_inode(struct inode *inode) -- 2.55.0.887.g758fc8c411-goog