* [PATCH v4 1/3] KVM: Release memory-attribute reservations abandoned on ENOMEM
2026-09-15 17:53 [PATCH v4 0/3] KVM: Fix and account mem_attr_array reservations abandoned on ENOMEM David Ballesteros
@ 2026-09-15 17:53 ` David Ballesteros
2026-09-15 17:53 ` [PATCH v4 2/3] KVM: Make kvm_range_has_memory_attributes() consistent about reservations David Ballesteros
` (2 subsequent siblings)
3 siblings, 0 replies; 7+ messages in thread
From: David Ballesteros @ 2026-09-15 17:53 UTC (permalink / raw)
To: pbonzini, seanjc; +Cc: kvm, linux-kernel
kvm_vm_set_mem_attributes() reserves an xarray entry for every GFN in the
range before storing anything, so that the store phase cannot fail partway
through. When a reservation fails, the loop jumps to out_unlock and the
reservations already made are abandoned: nothing in the call releases them.
A later request that clears a range covering them does erase them, as the
clear stores NULL over the reservation, but nothing obliges userspace to
issue one and a caller exploiting this will not; absent such a call the
entries live until kvm_destroy_vm(). An unprivileged user with /dev/kvm on
a VM with private-memory support can therefore leak kernel memory across
calls (a 576-byte xa_node per 64 GFNs) for the life of the VM fd.
The abandoned entries are not inert. A bare reservation is an
XA_ZERO_ENTRY. kvm_range_has_memory_attributes() is inconsistent about it:
the end == start + 1 path and the general loop treat it as absent (matching
kvm_get_memory_attributes(), which maps it to NULL via xa_load()), but the
!attrs fast path calls xas_find() directly, which returns the zero entry as
present. Via hugepage_has_attrs(), that makes
kvm_arch_post_set_memory_attributes() mark a straddling head/tail hugepage
"mixed" for a range whose attributes are in fact uniform, so KVM stops
using a hugepage there until a later request re-covers it.
Release the reservations this call made on the failure path. xa_release()
erases an entry only while it is still a reservation, so value entries that
predate this call are left untouched; it takes no gfp and cannot fail.
Only [start, i) is walked, i being the index whose reservation failed (no
entry was created at or beyond it).
Runtime-verified on v6.18.48 (isolated sw-protected VM, no KASAN, no fault
injection; the reservations are left behind by real memcg pressure via
clone(CLONE_VM), not by fault injection). Unpatched, a failed request
retains on the order of 450000 xa_nodes (~250 MiB), and with one of them
inside a 2 MiB region a clear of the region's head page leaves pages_2m
unchanged (the reservation is invisible to xa_load), a clear of two pages
drops pages_2m by one and raises pages_4k by 512 (the hugepage is
degraded), and a clear covering the whole 2 MiB restores it. With this
patch the same run retains under 2000 nodes -- three orders of magnitude
less, at the level of run-to-run noise -- and pages_2m stays at 16 across
all three clears: the reservations are released, so the hugepage is never
degraded. Both arms used the same kernel config and the same test binary.
The bug reproduces with XA_FLAGS_ACCOUNT applied (patch 3/3), i.e.
accounting alone does not fix it.
Found by an AI-assisted security audit.
Fixes: 5a475554db1e ("KVM: Introduce per-page memory attributes")
Cc: stable@vger.kernel.org
Assisted-by: Claude-Code:claude-opus-5
Signed-off-by: David Ballesteros <davimaba.v@proton.me>
---
virt/kvm/kvm_main.c | 24 +++++++++++++++++++++++-
1 file changed, 23 insertions(+), 1 deletion(-)
diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c
--- a/virt/kvm/kvm_main.c
+++ b/virt/kvm/kvm_main.c
@@ -2575,7 +2575,7 @@ static int kvm_vm_set_mem_attributes(struct kvm *kvm, gfn_t start, gfn_t end,
for (i = start; i < end; i++) {
r = xa_reserve(&kvm->mem_attr_array, i, GFP_KERNEL_ACCOUNT);
if (r)
- goto out_unlock;
+ goto out_release;
cond_resched();
}
@@ -2594,6 +2594,28 @@ static int kvm_vm_set_mem_attributes(struct kvm *kvm, gfn_t start, gfn_t end,
out_unlock:
mutex_unlock(&kvm->slots_lock);
+ return r;
+
+out_release:
+ /*
+ * The reservation loop failed at @i; the entries in [start, i) were
+ * reserved by this call and, without releasing them here, would be
+ * retained until userspace happens to clear a range covering them, or
+ * until the VM is destroyed. The retained entries are not inert:
+ * a bare reservation is an XA_ZERO_ENTRY, which the !attrs fast path of
+ * kvm_range_has_memory_attributes() counts as present (it calls
+ * xas_find() directly) even though kvm_get_memory_attributes() reports
+ * it as absent, so a straddling hugepage over such an entry gets marked
+ * mixed and KVM stops using a hugepage for a range whose attributes are
+ * uniform. xa_release() erases an entry only while it is still a
+ * reservation, so value entries that predate this call are untouched.
+ */
+ while (i-- > start) {
+ xa_release(&kvm->mem_attr_array, i);
+ cond_resched();
+ }
+ mutex_unlock(&kvm->slots_lock);
+
return r;
}
static int kvm_vm_ioctl_set_mem_attributes(struct kvm *kvm,
^ permalink raw reply [flat|nested] 7+ messages in thread* [PATCH v4 2/3] KVM: Make kvm_range_has_memory_attributes() consistent about reservations
2026-09-15 17:53 [PATCH v4 0/3] KVM: Fix and account mem_attr_array reservations abandoned on ENOMEM David Ballesteros
2026-09-15 17:53 ` [PATCH v4 1/3] KVM: Release memory-attribute " David Ballesteros
@ 2026-09-15 17:53 ` David Ballesteros
2026-09-15 17:53 ` [PATCH v4 3/3] KVM: Account mem_attr_array nodes to the caller's memcg David Ballesteros
2026-09-24 21:51 ` [PATCH v4 0/3] KVM: Fix and account mem_attr_array reservations abandoned on ENOMEM Sean Christopherson
3 siblings, 0 replies; 7+ messages in thread
From: David Ballesteros @ 2026-09-15 17:53 UTC (permalink / raw)
To: pbonzini, seanjc; +Cc: kvm, linux-kernel
Make the !attrs fast path of kvm_range_has_memory_attributes() skip bare
reservations, so that all three of the function's query paths agree on
what an XA_ZERO_ENTRY means.
A reservation carries no attributes, and two of the three paths already
treat it as absent: the end == start + 1 path reads it through
kvm_get_memory_attributes(), which maps it to NULL via xa_load(), and the
general loop skips it via xas_retry(). Only the !attrs fast path calls
xas_find() directly, which returns the reservation as a present entry, so
it reports a range that is in fact all-shared as not-all-shared. Through
hugepage_has_attrs(), that marks a straddling hugepage mixed for a range
whose attributes are uniform.
This is a consistency fix rather than a fix for a reachable bug, and is
not tagged for stable. Every caller of kvm_range_has_memory_attributes()
holds kvm->slots_lock -- the idempotency check in
kvm_vm_set_mem_attributes(), hugepage_has_attrs() from both
kvm_arch_post_set_memory_attributes() and
kvm_mmu_init_memslot_memory_attributes(), and __kvm_gmem_populate() --
and slots_lock excludes the only writer, so with patch 1/3 applied no
caller can observe a reservation. The function should not have to depend
on that to answer consistently.
Note this patch depends on 1/3 and must not be applied without it. Today a
clear over a range that holds only reservations does not take the
idempotency early-out in kvm_vm_set_mem_attributes(), because the fast path
reports the range as not-all-clear; the clear therefore proceeds and its
xa_store(NULL) loop erases the reservations as a side effect. Teaching the
fast path to skip reservations makes that early-out fire and removes the
accidental cleanup, so the patch that stops the reservations from being
abandoned in the first place has to come first.
Found by an AI-assisted security audit.
Fixes: 5a475554db1e ("KVM: Introduce per-page memory attributes")
Assisted-by: Claude-Code:claude-opus-5
Signed-off-by: David Ballesteros <davimaba.v@proton.me>
---
virt/kvm/kvm_main.c | 13 +++++++++++--
1 file changed, 11 insertions(+), 2 deletions(-)
diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c
index 45e7844..4e5e497 100644
--- a/virt/kvm/kvm_main.c
+++ b/virt/kvm/kvm_main.c
@@ -2446,8 +2446,17 @@ bool kvm_range_has_memory_attributes(struct kvm *kvm, gfn_t start, gfn_t end,
return (kvm_get_memory_attributes(kvm, start) & mask) == attrs;
guard(rcu)();
- if (!attrs)
- return !xas_find(&xas, end - 1);
+ if (!attrs) {
+ /*
+ * Skip reservations: a bare XA_ZERO_ENTRY carries no
+ * attributes, but xas_find() returns it raw.
+ */
+ do {
+ entry = xas_find(&xas, end - 1);
+ } while (xas_retry(&xas, entry));
+
+ return !entry;
+ }
for (index = start; index < end; index++) {
do {
^ permalink raw reply [flat|nested] 7+ messages in thread* [PATCH v4 3/3] KVM: Account mem_attr_array nodes to the caller's memcg
2026-09-15 17:53 [PATCH v4 0/3] KVM: Fix and account mem_attr_array reservations abandoned on ENOMEM David Ballesteros
2026-09-15 17:53 ` [PATCH v4 1/3] KVM: Release memory-attribute " David Ballesteros
2026-09-15 17:53 ` [PATCH v4 2/3] KVM: Make kvm_range_has_memory_attributes() consistent about reservations David Ballesteros
@ 2026-09-15 17:53 ` David Ballesteros
2026-09-24 21:51 ` [PATCH v4 0/3] KVM: Fix and account mem_attr_array reservations abandoned on ENOMEM Sean Christopherson
3 siblings, 0 replies; 7+ messages in thread
From: David Ballesteros @ 2026-09-15 17:53 UTC (permalink / raw)
To: pbonzini, seanjc; +Cc: kvm, linux-kernel
kvm_vm_set_mem_attributes() passes GFP_KERNEL_ACCOUNT when reserving xarray
entries, but the nodes are allocated by xas_alloc(), which hardcodes
GFP_NOWAIT and only adds __GFP_ACCOUNT when the xarray carries
XA_FLAGS_ACCOUNT. mem_attr_array is initialized with plain xa_init(), so
the flag is never set and nodes taken from that fast path -- the
overwhelming majority -- are not charged to the caller; only the rare
__xas_nomem() slow path is, because it receives the caller's gfp. Measured
on v6.18.48: a process in a cgroup limited to 256 MiB grew
radix_tree_node slab by ~512 MiB while its memory.current stayed near 0.
Per-tenant memcg limits therefore do not contain the growth.
Set XA_FLAGS_ACCOUNT so the intended accounting takes effect.
Note this is a change in reachability, not in contract.
KVM_SET_MEMORY_ATTRIBUTES could already return -ENOMEM, but only under
global memory pressure, since the fast path allocated with plain
GFP_NOWAIT. With the nodes accounted, a tenant under a memory.max limit
can now hit it from a cgroup-local condition, i.e. conversions that
previously succeeded may fail. That is intended and matches every other
GFP_KERNEL_ACCOUNT allocation in KVM; the alternative is letting the tenant
grow host memory that is never charged to it. Userspace driving
conversions from guest KVM_HC_MAP_GPA_RANGE hypercalls surfaces the failure
on that path.
Runtime-verified on v6.18.48 (isolated VM, no KASAN): without the flag a
process in a 256 MiB cgroup materializes 512 MiB of radix_tree_node slab
with memory.current flat (the memcg is inert); with the flag the same
process is contained by the cgroup -- the memcg OOM killer selects the
attacker inside its own slice (CONSTRAINT_MEMCG) instead of exhausting
global memory.
Found by an AI-assisted security audit.
Not tagged for stable: unlike 1/3 and 2/3, which are pure corrections, this
one changes observable behaviour, and new -ENOMEM returns for
cgroup-limited tenants are a poor fit for stable's regression-risk bar.
Happy to send it to stable separately if maintainers disagree.
Fixes: 5a475554db1e ("KVM: Introduce per-page memory attributes")
Assisted-by: Claude-Code:claude-opus-5
Signed-off-by: David Ballesteros <davimaba.v@proton.me>
---
virt/kvm/kvm_main.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c
--- a/virt/kvm/kvm_main.c
+++ b/virt/kvm/kvm_main.c
@@ -1116,7 +1116,7 @@ static struct kvm *kvm_create_vm(unsigned long type, const char *fdname)
rcuwait_init(&kvm->mn_memslots_update_rcuwait);
xa_init(&kvm->vcpu_array);
#ifdef CONFIG_KVM_GENERIC_MEMORY_ATTRIBUTES
- xa_init(&kvm->mem_attr_array);
+ xa_init_flags(&kvm->mem_attr_array, XA_FLAGS_ACCOUNT);
#endif
INIT_LIST_HEAD(&kvm->gpc_list);
^ permalink raw reply [flat|nested] 7+ messages in thread* Re: [PATCH v4 0/3] KVM: Fix and account mem_attr_array reservations abandoned on ENOMEM
2026-09-15 17:53 [PATCH v4 0/3] KVM: Fix and account mem_attr_array reservations abandoned on ENOMEM David Ballesteros
` (2 preceding siblings ...)
2026-09-15 17:53 ` [PATCH v4 3/3] KVM: Account mem_attr_array nodes to the caller's memcg David Ballesteros
@ 2026-09-24 21:51 ` Sean Christopherson
2026-09-27 6:13 ` David Ballesteros
3 siblings, 1 reply; 7+ messages in thread
From: Sean Christopherson @ 2026-09-24 21:51 UTC (permalink / raw)
To: Sean Christopherson, pbonzini, David Ballesteros; +Cc: kvm, linux-kernel
On Tue, 15 Sep 2026 17:53:40 +0000, David Ballesteros wrote:
> Three small fixes to KVM's per-page memory attributes: release the xarray
> reservations that KVM_SET_MEMORY_ATTRIBUTES abandons when it fails partway
> through (1/3), make kvm_range_has_memory_attributes() agree with itself
> about what such a reservation means (2/3), and charge the xa_nodes to the
> caller's memcg as the code already intended (3/3).
>
> 1/3 Release the reservations abandoned on ENOMEM. This is a plain bug:
> xa_reserve() materializes entries GFN-by-GFN before the store phase,
> and on failure the loop bails without releasing what it reserved. A
> later clear covering them does erase them, but nothing obliges
> userspace to issue one; absent that, the reclaim path is
> kvm_destroy_vm(). The retained entries are not inert -- an
> abandoned reservation is an XA_ZERO_ENTRY, which
> kvm_range_has_memory_attributes()'s !attrs fast path counts as
> present (raw xas_find()) while kvm_get_memory_attributes() treats
> it as absent, so a straddling hugepage over such an entry is marked
> mixed and KVM stops using a hugepage for a range whose attributes
> are uniform. xa_release() erases only entries still reserved,
> leaving pre-existing value entries untouched.
>
> [...]
Applied patch 3, with a heavily modified changelog, to kvm-x86 fixes. For the
reservation behavior, I went with Zeng Chi's fix to have KVM treat ZERO values
as "no attributes". Having dangling reservations is a-ok, the memcg accounting
really needs to do the right thing there.
In the future, please don't have AI directly write changelogs. It's fine to
let AI generate a rough draft, for me at least, AI tends to be far too verbose
and uses terminology that isn't common in Linux/upstream. In other words, AI
tends to write changelogs (and bug reports) that require far too much effort
to understand.
I apologize in advance if you wrote the changelogs, i.e. if I am falsely
accusing you of being a robot. If AI didn't write the changelogs, well, you
do one heck of a job of imitating some of my newfound "friends" :-)
Gripes about AI aside, than you very much for the fixes!
[1/3] KVM: Release memory-attribute reservations abandoned on ENOMEM
[SKIP]
[2/3] KVM: Make kvm_range_has_memory_attributes() consistent about reservations
[SKIP]
[3/3] KVM: Account mem_attr_array nodes to the caller's memcg
https://github.com/kvm-x86/linux/commit/382e5d514b6f
--
https://github.com/kvm-x86/linux/tree/next
^ permalink raw reply [flat|nested] 7+ messages in thread* Re: [PATCH v4 0/3] KVM: Fix and account mem_attr_array reservations abandoned on ENOMEM
2026-09-24 21:51 ` [PATCH v4 0/3] KVM: Fix and account mem_attr_array reservations abandoned on ENOMEM Sean Christopherson
@ 2026-09-27 6:13 ` David Ballesteros
2026-09-28 16:30 ` Sean Christopherson
0 siblings, 1 reply; 7+ messages in thread
From: David Ballesteros @ 2026-09-27 6:13 UTC (permalink / raw)
To: Sean Christopherson; +Cc: pbonzini, kvm, linux-kernel
Haha, I didn't even realize using AI would make me sound like a bot
until now...
Jokes aside, I truly appreciate your kindness. I have to admit that
even though I tried to review everything as thoroughly as possible,
this was my very first patch submission, which is why I found it a bit tricky to follow all the conventions (though that's no excuse). I definitely leaned too heavily on AI for the write-up. With that said, I understand the extra work it causes and I am truly sorry. If there's a next time, I'll write it myself -- AI at most for a rough first pass, but the words that reach the list in general will be mine. Best regards, and thanks for the guidance and patience.
David Ballesteros
On 24/09/26 16:54, Sean Christopherson wrote:
> In the future, please don't have AI directly write changelogs. It's
> fine to let AI generate a rough draft, for me at least, AI tends to
> be far too verbose and uses terminology that isn't common in
> Linux/upstream. In other words, AI tends to write changelogs (and
> bug reports) that require far too much effort to understand.
>
> Gripes about AI aside, than you very much for the fixes!
^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH v4 0/3] KVM: Fix and account mem_attr_array reservations abandoned on ENOMEM
2026-09-27 6:13 ` David Ballesteros
@ 2026-09-28 16:30 ` Sean Christopherson
0 siblings, 0 replies; 7+ messages in thread
From: Sean Christopherson @ 2026-09-28 16:30 UTC (permalink / raw)
To: David Ballesteros; +Cc: pbonzini, kvm, linux-kernel
On Sun, Sep 27, 2026, David Ballesteros wrote:
>
> Haha, I didn't even realize using AI would make me sound like a bot
> until now...
>
> Jokes aside, I truly appreciate your kindness. I have to admit that even
> though I tried to review everything as thoroughly as possible, this was my
> very first patch submission, which is why I found it a bit tricky to follow
> all the conventions (though that's no excuse).
No worries, contributing upstream is definitely a gauntlet run. Calling it "a
bit tricky" is being *very* generous.
> I definitely leaned too heavily on AI for the write-up. With that said, I
> understand the extra work it causes and I am truly sorry. If there's a next
> time,
Hopefully there is a next time, but selfishly, I hope the next wave of security
fixes is to a different subsystem. :-D
> I'll write it myself -- AI at most for a rough first pass, but the
> words that reach the list in general will be mine. Best regards, and thanks
> for the guidance and patience.
In the future, please also wait a reasonable amount of time between versions,
e.g. 1-2 business days for relatively urgent fixes, more for non-urgent things
(though with so many of the guidelines, there are exceptions; use common sense).
With Sashiko and syzbot giving such quick feedback, it's tempting and all too
easy to blast versions back-to-back(-to-back), but that's often counter-productive.
E.g. there's almost no opportunity for a human to step in and provide guidance
(e.g. to point out the existing on-list fixes), the sheer volume of mail adds to
the load of maintainers, and if we overload the bots, then we risk losing some
of that early/prompt feedback.
Thanks again!
^ permalink raw reply [flat|nested] 7+ messages in thread