* [PATCH v4 0/5] KVM: guest_memfd: Fix binding bugs
@ 2026-09-21 21:06 Sean Christopherson
2026-09-21 21:06 ` [PATCH v4 1/5] KVM: guest_memfd: Gracefully handle xarray errors when binding a memslot Sean Christopherson
` (4 more replies)
0 siblings, 5 replies; 6+ messages in thread
From: Sean Christopherson @ 2026-09-21 21:06 UTC (permalink / raw)
To: Paolo Bonzini, Sean Christopherson
Cc: David Hildenbrand, kvm, linux-kernel, Stefan Teodorescu,
Dennis Tighe, Sashiko Bot, Ackerley Tng, Yan Zhao
Fix guest_memfd bugs related to binding to a memslot:
- Handle errors when inserting into guest_memfd's binding xarray, e.g. to
do the right thing on ENOMEM.
- Bind a memslot only once the memslot is fully prepared (because it becomes
reachable/visible once its inserted into gmem's bindings arraxy).
Patch 5 is a related cleanup to remove a superflous WRITE_ONCE() (unwinding
the slot update on insertion failure isn't an option if the slot is observable,
i.e. if the WRITE_ONCE() is actually necessary).
v4:
- Drop intermediate xar variable. [David, Ackerley]
- Collect reviews. [David, Ackerley]
- Restrict kvm_gmem_bind() to CREATE as calling kvm_arch_free_memslot() on
FLAGS_ONLY changes is unsafe, and doing the right thing for dirty bitmaps
is tricky for similar reasons. As a bonus, this eliminates the ugly
almost-duplicate code that David pointed out.
- Explain why KVM must deal with the unwind during bind(). [Ackerley]
v3:
- https://lore.kernel.org/all/20260904004342.3162959-1-seanjc@google.com
- Nullify bindings on error before dropping invalidat lock. [Sashiko x3]
v2:
- https://lore.kernel.org/all/20260902182020.2615443-1-seanjc@google.com
- Fix the binding-too-early bug. [Sashiko]
- Explicitly zero the bindings entry on failure to ensure there are no
partial entries. [Sashiko]
v1: https://lore.kernel.org/all/20260826165154.766699-2-seanjc@google.com
Sean Christopherson (5):
KVM: guest_memfd: Gracefully handle xarray errors when binding a
memslot
KVM: Use goto to handle errors during memslot preparation
KVM: Only bind memslot to guest_memfd instance for CREATE operations
KVM: guest_memfd: Establish memslot<=>guest_memfd bindings *after*
memslot is ready
KVM: guest_memfd: Drop superfluous WRITE_ONCE() when binding a memslot
virt/kvm/guest_memfd.c | 14 ++++++++--
virt/kvm/kvm_main.c | 62 ++++++++++++++++++++++--------------------
2 files changed, 44 insertions(+), 32 deletions(-)
base-commit: 70c944caf570fda2d79baa71435589a8db39f048
--
2.55.0.1082.g2b9226bbc0-goog
^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH v4 1/5] KVM: guest_memfd: Gracefully handle xarray errors when binding a memslot
2026-09-21 21:06 [PATCH v4 0/5] KVM: guest_memfd: Fix binding bugs Sean Christopherson
@ 2026-09-21 21:06 ` Sean Christopherson
2026-09-21 21:06 ` [PATCH v4 2/5] KVM: Use goto to handle errors during memslot preparation Sean Christopherson
` (3 subsequent siblings)
4 siblings, 0 replies; 6+ messages in thread
From: Sean Christopherson @ 2026-09-21 21:06 UTC (permalink / raw)
To: Paolo Bonzini, Sean Christopherson
Cc: David Hildenbrand, kvm, linux-kernel, Stefan Teodorescu,
Dennis Tighe, Sashiko Bot, Ackerley Tng, Yan Zhao
If inserting a memslot into a guest_memfd's bindings xarray fails,
propagate the error back to the caller, i.e. fail memslot creation as well.
Signalling success and continuing on with memslot creation results in
use-after-free, as the guest_memfd instance will remain reachable via the
memslot after the file is freed (kvm_gmem_release() won't nullify the file
pointer due to lack of a valid binding).
Opportunistically WARN and reject binding if KVM_MEMSLOT_GMEM_ONLY is
already set, partly to guard against goofs elsewhere, but mostly so that
KVM doesn't need to worry about clobbering flags when unwinding on failure.
Regarding the unwind, the slot must be fully prepared before inserting it
into the bindings, at which point the slot becomes reachable. I.e. waiting
to update the slot in order to avoid the ugly unwind isn't an option. And
as part of the unwind, explicitly nullify the relevant bindings, as xarray
can store a subset of entries when populating a range.
Fixes: a7800aa80ea4 ("KVM: Add KVM_CREATE_GUEST_MEMFD ioctl() for guest-specific backing memory")
Cc: stable@vger.kernel.org
Reported-by: Stefan Teodorescu <fane@google.com>
Reported-by: Dennis Tighe <dtighe@google.com>
Reported-by: Sashiko Bot <sashiko-bot@kernel.org>
Closes: https://lore.kernel.org/all/20260823135031.4F6DC1F000E9%40smtp.kernel.org
Reviewed-by: David Hildenbrand (Arm) <david@kernel.org>
Reviewed-by: Ackerley Tng <ackerleytng@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
---
virt/kvm/guest_memfd.c | 12 ++++++++++--
1 file changed, 10 insertions(+), 2 deletions(-)
diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
index 63943aa253d4..c094611f7c7a 100644
--- a/virt/kvm/guest_memfd.c
+++ b/virt/kvm/guest_memfd.c
@@ -654,6 +654,9 @@ int kvm_gmem_bind(struct kvm *kvm, struct kvm_memory_slot *slot,
BUILD_BUG_ON(sizeof(gpa_t) != sizeof(offset));
BUILD_BUG_ON(sizeof(gfn_t) != sizeof(slot->gmem.pgoff));
+ if (WARN_ON_ONCE(slot->flags & KVM_MEMSLOT_GMEM_ONLY))
+ return -EINVAL;
+
file = fget(fd);
if (!file)
return -EBADF;
@@ -692,7 +695,13 @@ int kvm_gmem_bind(struct kvm *kvm, struct kvm_memory_slot *slot,
if (kvm_gmem_supports_mmap(inode))
slot->flags |= KVM_MEMSLOT_GMEM_ONLY;
- xa_store_range(&f->bindings, start, end - 1, slot, GFP_KERNEL);
+ r = xa_err(xa_store_range(&f->bindings, start, end - 1, slot, GFP_KERNEL));
+ if (r) {
+ xa_store_range(&f->bindings, start, end - 1, NULL, GFP_KERNEL);
+ slot->gmem.file = NULL;
+ slot->gmem.pgoff = 0;
+ slot->flags &= ~KVM_MEMSLOT_GMEM_ONLY;
+ }
filemap_invalidate_unlock(inode->i_mapping);
/*
@@ -700,7 +709,6 @@ int kvm_gmem_bind(struct kvm *kvm, struct kvm_memory_slot *slot,
* not the other way 'round. Active bindings are invalidated if the
* file is closed before memslots are destroyed.
*/
- r = 0;
err:
fput(file);
return r;
--
2.55.0.1082.g2b9226bbc0-goog
^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH v4 2/5] KVM: Use goto to handle errors during memslot preparation
2026-09-21 21:06 [PATCH v4 0/5] KVM: guest_memfd: Fix binding bugs Sean Christopherson
2026-09-21 21:06 ` [PATCH v4 1/5] KVM: guest_memfd: Gracefully handle xarray errors when binding a memslot Sean Christopherson
@ 2026-09-21 21:06 ` Sean Christopherson
2026-09-21 21:06 ` [PATCH v4 3/5] KVM: Only bind memslot to guest_memfd instance for CREATE operations Sean Christopherson
` (2 subsequent siblings)
4 siblings, 0 replies; 6+ messages in thread
From: Sean Christopherson @ 2026-09-21 21:06 UTC (permalink / raw)
To: Paolo Bonzini, Sean Christopherson
Cc: David Hildenbrand, kvm, linux-kernel, Stefan Teodorescu,
Dennis Tighe, Sashiko Bot, Ackerley Tng, Yan Zhao
Use a goto to unwind early memslot changes if preparing for a memslot
operation fails. This will allow moving the creation of guest_memfd
bindings into kvm_set_memslot() without needing to copy+paste the unwind
logic.
No functional change intended.
Cc: stable@vger.kernel.org
Reviewed-by: David Hildenbrand (Arm) <david@kernel.org>
Reviewed-by: Ackerley Tng <ackerleytng@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
---
virt/kvm/kvm_main.c | 31 ++++++++++++++++---------------
1 file changed, 16 insertions(+), 15 deletions(-)
diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c
index d9da8b51614a..24cf96840827 100644
--- a/virt/kvm/kvm_main.c
+++ b/virt/kvm/kvm_main.c
@@ -1941,21 +1941,8 @@ static int kvm_set_memslot(struct kvm *kvm,
}
r = kvm_prepare_memory_region(kvm, old, new, change);
- if (r) {
- /*
- * For DELETE/MOVE, revert the above INVALID change. No
- * modifications required since the original slot was preserved
- * in the inactive slots. Changing the active memslots also
- * release slots_arch_lock.
- */
- if (change == KVM_MR_DELETE || change == KVM_MR_MOVE) {
- kvm_activate_memslot(kvm, invalid_slot, old);
- kfree(invalid_slot);
- } else {
- mutex_unlock(&kvm->slots_arch_lock);
- }
- return r;
- }
+ if (r)
+ goto err;
/*
* For DELETE and MOVE, the working slot is now active as the INVALID
@@ -1987,6 +1974,20 @@ static int kvm_set_memslot(struct kvm *kvm,
kvm_commit_memory_region(kvm, old, new, change);
return 0;
+
+err:
+ /*
+ * For DELETE/MOVE, revert the above INVALID change. No modifications
+ * required since the original slot was preserved in the inactive slots.
+ * Changing the active memslots also release slots_arch_lock.
+ */
+ if (change == KVM_MR_DELETE || change == KVM_MR_MOVE) {
+ kvm_activate_memslot(kvm, invalid_slot, old);
+ kfree(invalid_slot);
+ } else {
+ mutex_unlock(&kvm->slots_arch_lock);
+ }
+ return r;
}
static bool kvm_check_memslot_overlap(struct kvm_memslots *slots, int id,
--
2.55.0.1082.g2b9226bbc0-goog
^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH v4 3/5] KVM: Only bind memslot to guest_memfd instance for CREATE operations
2026-09-21 21:06 [PATCH v4 0/5] KVM: guest_memfd: Fix binding bugs Sean Christopherson
2026-09-21 21:06 ` [PATCH v4 1/5] KVM: guest_memfd: Gracefully handle xarray errors when binding a memslot Sean Christopherson
2026-09-21 21:06 ` [PATCH v4 2/5] KVM: Use goto to handle errors during memslot preparation Sean Christopherson
@ 2026-09-21 21:06 ` Sean Christopherson
2026-09-21 21:06 ` [PATCH v4 4/5] KVM: guest_memfd: Establish memslot<=>guest_memfd bindings *after* memslot is ready Sean Christopherson
2026-09-21 21:06 ` [PATCH v4 5/5] KVM: guest_memfd: Drop superfluous WRITE_ONCE() when binding a memslot Sean Christopherson
4 siblings, 0 replies; 6+ messages in thread
From: Sean Christopherson @ 2026-09-21 21:06 UTC (permalink / raw)
To: Paolo Bonzini, Sean Christopherson
Cc: David Hildenbrand, kvm, linux-kernel, Stefan Teodorescu,
Dennis Tighe, Sashiko Bot, Ackerley Tng, Yan Zhao
For additional defense-in-depth, and to avoid having to handle impossible
unwind scenarios when binding to a memslot fails, bind a memslot to a gmem
instance only when for CREATE operations, i.e. don't attempt to establish a
binding for MOVE and FLAGS_ONLY operations. And when FLAGS_ONLY operations
are eventually supported (this is currently all dead code), creating a new
binding would be incorrect; KVM instead needs to do a 1:1 replacement of
the existing binding, i.e. FLAGS_ONLY will need its own dedicated handling.
Update the relevant TODO to make a better guess as to what needs to be done
to support toggling dirty logging for guest_memfd memslots.
Because it's dead code, no functional change intended.
Cc: stable@vger.kernel.org
Signed-off-by: Sean Christopherson <seanjc@google.com>
---
virt/kvm/kvm_main.c | 10 +++++-----
1 file changed, 5 insertions(+), 5 deletions(-)
diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c
index 24cf96840827..b417b1f7095f 100644
--- a/virt/kvm/kvm_main.c
+++ b/virt/kvm/kvm_main.c
@@ -1752,10 +1752,10 @@ static void kvm_commit_memory_region(struct kvm *kvm,
kvm_destroy_dirty_bitmap(old);
/*
- * Unbind the guest_memfd instance as needed; the @new slot has
- * already created its own binding. TODO: Drop the WARN when
- * dirty logging guest_memfd memslots is supported. Until then,
- * flags-only changes on guest_memfd slots should be impossible.
+ * TODO: Drop the WARN and do the unbind() call only for MOVE
+ * when dirty logging guest_memfd memslots is supported. Until
+ * then, flags-only changes on guest_memfd slots should also be
+ * impossible; unbind the old memslot for defense-in-depth.
*/
if (WARN_ON_ONCE(old->flags & KVM_MEM_GUEST_MEMFD))
kvm_gmem_unbind(old);
@@ -2116,7 +2116,7 @@ static int kvm_set_memory_region(struct kvm *kvm,
new->npages = npages;
new->flags = mem->flags;
new->userspace_addr = mem->userspace_addr;
- if (mem->flags & KVM_MEM_GUEST_MEMFD) {
+ if (change == KVM_MR_CREATE && (mem->flags & KVM_MEM_GUEST_MEMFD)) {
r = kvm_gmem_bind(kvm, new, mem->guest_memfd, mem->guest_memfd_offset);
if (r)
goto out;
--
2.55.0.1082.g2b9226bbc0-goog
^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH v4 4/5] KVM: guest_memfd: Establish memslot<=>guest_memfd bindings *after* memslot is ready
2026-09-21 21:06 [PATCH v4 0/5] KVM: guest_memfd: Fix binding bugs Sean Christopherson
` (2 preceding siblings ...)
2026-09-21 21:06 ` [PATCH v4 3/5] KVM: Only bind memslot to guest_memfd instance for CREATE operations Sean Christopherson
@ 2026-09-21 21:06 ` Sean Christopherson
2026-09-21 21:06 ` [PATCH v4 5/5] KVM: guest_memfd: Drop superfluous WRITE_ONCE() when binding a memslot Sean Christopherson
4 siblings, 0 replies; 6+ messages in thread
From: Sean Christopherson @ 2026-09-21 21:06 UTC (permalink / raw)
To: Paolo Bonzini, Sean Christopherson
Cc: David Hildenbrand, kvm, linux-kernel, Stefan Teodorescu,
Dennis Tighe, Sashiko Bot, Ackerley Tng, Yan Zhao
Wait to bind a memslot to a guest_memfd instance until *after* the memslot
is fully prepared, as creating the binding in guest_memfd will effectively
expose the memslot to readers. As pointed out by Sashiko, binding the
memslot before it's ready to be exposed to the rest of the world can break
various memslot assumption and rules. E.g. x86 could observe a NULL rmap
pointer if a PUNCH_HOLE hit the guest_memfd after the binding was created,
but before KVM made it through kvm_prepare_memory_region().
Begrudgingly resort to passing in the guest_memfd fd+offset pair to
kvm_set_memslot(), as creating the binding really does need to happen in
the middle of setting the new memslot. Alternatively, to preserve the
aesthetically pleasing function prototype, "struct kvm_memory_slot" could
be expanded to track the fd and the file, but that would create the
possibility for TOCTOU bugs on the fd vs. file, and would add zero value
beyond making kvm_set_memslot() look pretty.
Fixes:a7800aa80ea4 ("KVM: Add KVM_CREATE_GUEST_MEMFD ioctl() for guest-specific backing memory")
Cc: stable@vger.kernel.org
Reported-by: Sashiko Bot <sashiko-bot@kernel.org>
Closes: https://lore.kernel.org/all/20260826170551.BEF801F000E9@smtp.kernel.org
Signed-off-by: Sean Christopherson <seanjc@google.com>
---
virt/kvm/kvm_main.c | 27 +++++++++++++++------------
1 file changed, 15 insertions(+), 12 deletions(-)
diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c
index b417b1f7095f..cc79a33a7d39 100644
--- a/virt/kvm/kvm_main.c
+++ b/virt/kvm/kvm_main.c
@@ -1897,7 +1897,8 @@ static void kvm_update_flags_memslot(struct kvm *kvm,
static int kvm_set_memslot(struct kvm *kvm,
struct kvm_memory_slot *old,
struct kvm_memory_slot *new,
- enum kvm_mr_change change)
+ enum kvm_mr_change change,
+ unsigned int gmem_fd, uoff_t gmem_offset)
{
struct kvm_memory_slot *invalid_slot;
int r;
@@ -1944,6 +1945,15 @@ static int kvm_set_memslot(struct kvm *kvm,
if (r)
goto err;
+ if (change == KVM_MR_CREATE && (new->flags & KVM_MEM_GUEST_MEMFD)) {
+ r = kvm_gmem_bind(kvm, new, gmem_fd, gmem_offset);
+ if (r) {
+ kvm_arch_free_memslot(kvm, new);
+ kvm_destroy_dirty_bitmap(new);
+ goto err;
+ }
+ }
+
/*
* For DELETE and MOVE, the working slot is now active as the INVALID
* version of the old slot. MOVE is particularly special as it reuses
@@ -2069,7 +2079,7 @@ static int kvm_set_memory_region(struct kvm *kvm,
if (WARN_ON_ONCE(kvm->nr_memslot_pages < old->npages))
return -EIO;
- return kvm_set_memslot(kvm, old, NULL, KVM_MR_DELETE);
+ return kvm_set_memslot(kvm, old, NULL, KVM_MR_DELETE, -1, 0);
}
base_gfn = (mem->guest_phys_addr >> PAGE_SHIFT);
@@ -2116,21 +2126,14 @@ static int kvm_set_memory_region(struct kvm *kvm,
new->npages = npages;
new->flags = mem->flags;
new->userspace_addr = mem->userspace_addr;
- if (change == KVM_MR_CREATE && (mem->flags & KVM_MEM_GUEST_MEMFD)) {
- r = kvm_gmem_bind(kvm, new, mem->guest_memfd, mem->guest_memfd_offset);
- if (r)
- goto out;
- }
- r = kvm_set_memslot(kvm, old, new, change);
+ r = kvm_set_memslot(kvm, old, new, change,
+ mem->guest_memfd, mem->guest_memfd_offset);
if (r)
- goto out_unbind;
+ goto out;
return 0;
-out_unbind:
- if (mem->flags & KVM_MEM_GUEST_MEMFD)
- kvm_gmem_unbind(new);
out:
kfree(new);
return r;
--
2.55.0.1082.g2b9226bbc0-goog
^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH v4 5/5] KVM: guest_memfd: Drop superfluous WRITE_ONCE() when binding a memslot
2026-09-21 21:06 [PATCH v4 0/5] KVM: guest_memfd: Fix binding bugs Sean Christopherson
` (3 preceding siblings ...)
2026-09-21 21:06 ` [PATCH v4 4/5] KVM: guest_memfd: Establish memslot<=>guest_memfd bindings *after* memslot is ready Sean Christopherson
@ 2026-09-21 21:06 ` Sean Christopherson
4 siblings, 0 replies; 6+ messages in thread
From: Sean Christopherson @ 2026-09-21 21:06 UTC (permalink / raw)
To: Paolo Bonzini, Sean Christopherson
Cc: David Hildenbrand, kvm, linux-kernel, Stefan Teodorescu,
Dennis Tighe, Sashiko Bot, Ackerley Tng, Yan Zhao
Drop the superfluous WRITE_ONCE() when setting a memslot's guest_memfd file
during initial binding, as the memslot *must* be inactive and unreachable.
The superfluous WRITE_ONCE() was added by commit 67b43038ce14 ("KVM:
guest_memfd: Remove RCU-protected attribute from slot->gmem.file") to
maintain rough "parity" with the existing rcu_assign_pointer(), not
realizing that the only reason rcu_assign_pointer() was used was to make
sparse and other checkers happy.
Cc: Yan Zhao <yan.y.zhao@intel.com>
Reviewed-by: David Hildenbrand (Arm) <david@kernel.org>
Reviewed-by: Ackerley Tng <ackerleytng@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
---
virt/kvm/guest_memfd.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
index c094611f7c7a..11795ffe5830 100644
--- a/virt/kvm/guest_memfd.c
+++ b/virt/kvm/guest_memfd.c
@@ -690,7 +690,7 @@ int kvm_gmem_bind(struct kvm *kvm, struct kvm_memory_slot *slot,
* kvm_gmem_bind() must occur on a new memslot. Because the memslot
* is not visible yet, kvm_gmem_get_pfn() is guaranteed to see the file.
*/
- WRITE_ONCE(slot->gmem.file, file);
+ slot->gmem.file = file;
slot->gmem.pgoff = start;
if (kvm_gmem_supports_mmap(inode))
slot->flags |= KVM_MEMSLOT_GMEM_ONLY;
--
2.55.0.1082.g2b9226bbc0-goog
^ permalink raw reply [flat|nested] 6+ messages in thread
end of thread, other threads:[~2026-09-21 21:06 UTC | newest]
Thread overview: 6+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-21 21:06 [PATCH v4 0/5] KVM: guest_memfd: Fix binding bugs Sean Christopherson
2026-09-21 21:06 ` [PATCH v4 1/5] KVM: guest_memfd: Gracefully handle xarray errors when binding a memslot Sean Christopherson
2026-09-21 21:06 ` [PATCH v4 2/5] KVM: Use goto to handle errors during memslot preparation Sean Christopherson
2026-09-21 21:06 ` [PATCH v4 3/5] KVM: Only bind memslot to guest_memfd instance for CREATE operations Sean Christopherson
2026-09-21 21:06 ` [PATCH v4 4/5] KVM: guest_memfd: Establish memslot<=>guest_memfd bindings *after* memslot is ready Sean Christopherson
2026-09-21 21:06 ` [PATCH v4 5/5] KVM: guest_memfd: Drop superfluous WRITE_ONCE() when binding a memslot Sean Christopherson
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®