From: Zeng Chi <zeng_chi911@163.com>
To: Sean Christopherson <seanjc@google.com>
Cc: pbonzini@redhat.com, chao.p.peng@linux.intel.com,
kvm@vger.kernel.org, linux-kernel@vger.kernel.org,
zengchi@kylinos.cn
Subject: Re: [PATCH] KVM: Release reserved xarray entries if reserving memory attributes fails
Date: Fri, 28 Aug 2026 18:51:06 +0800 [thread overview]
Message-ID: <7d00f8a3-0d35-49ce-83ab-29b94a3490ed@163.com> (raw)
In-Reply-To: <apCD3F3oOyC6E5oN@google.com>
On 2026/8/28 02:37, Sean Christopherson wrote:
> On Thu, Aug 27, 2026, Zeng Chi wrote:
>> From: Zeng Chi <zengchi@kylinos.cn>
>>
>> kvm_vm_set_mem_attributes() reserves an xarray entry for every gfn in
>> the range before modifying any attributes, so that the actual updates
>> can't fail partway through. But if one of the reservations fails, e.g.
>> due to -ENOMEM, the entries that were already reserved are left behind,
>> as the error path bails without releasing them.
>>
>> A reserved entry is XA_ZERO_ENTRY, not NULL.
>
> Lovely.
>
>> xa_load() hides the difference, but kvm_range_has_memory_attributes() uses
>> xas_find() to check whether a range has no attributes at all, and xas_find()
>> returns zero entries as-is. As a result, a leaked reservation makes KVM
>> think the range has attributes set even though kvm_get_memory_attributes()
>> reports none. On x86, the next time mixed-attribute tracking is recomputed
>> for the range (memslot creation, or a later attribute change that straddles
>> the 2MiB page), hugepage_has_attrs() treats a fully shared 2MiB range as
On 2026/8/28 02:37, Sean Christopherson wrote:
> I'm inclined to fix kvm_range_has_memory_attributes() instead of unwinding the
> reservation. Because this isn't a memory leak per se, e.g. if it weren't for
> the false negative in kvm_range_has_memory_attributes(), I would say this is a
> complete non-issue (there's no leak, just a maybe-unused reservation).
> working as intended.
>
Agreed, fixing the reader is better. Reframed that way in v2.
> I think it would be this?
>
> diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c
> index 65eb26a0520d..a01b2af1cb17 100644
> --- a/virt/kvm/kvm_main.c
> +++ b/virt/kvm/kvm_main.c
> @@ -2447,8 +2447,9 @@ bool kvm_range_has_memory_attributes(struct kvm *kvm, gfn_t start, gfn_t end,
> return (kvm_get_memory_attributes(kvm, start) & mask) == attrs;
>
> guard(rcu)();
> - if (!attrs)
> - return !xas_find(&xas, end - 1);
> +
> + if (!attrs && !xas_find(&xas, end - 1))
> + return true;
>
> for (index = start; index < end; index++) {
> do {
>
I tried that first, but it doesn't fix the false positive. Falling through to
the generic loop for the !attrs case still returns false for a range that only
contains reserved (zero) entries: the loop does
do {
entry = xas_next(&xas);
} while (xas_retry(&xas, entry));
and xas_retry() returns true for zero entries (xa_is_zero()), so it skips the
reserved entries, then "xas.xa_index != index" trips and the function returns
false. So the leaked-reservation range is still reported as having attributes.
I confirmed it with the tools/testing/radix-tree harness (reserve [0, 512),
then query attrs == 0): the original code, the sketch above, and the sketch
with the xas cursor reset all return false, where absent is expected.
What does work is to walk the range and ignore the zero entries explicitly:
guard(rcu)();
if (!attrs) {
xas_for_each(&xas, entry, end - 1)
if (!xa_is_zero(entry))
return false;
return true;
}
That gives the right answer for {empty, only reservations, a real value present,
reservations + a real value, all set}.
> Side topic, does storing NULL even require an entry? Based on the above behavior,
> I assume not. So can't we also do? This feels like deja vu though...
>
> @@ -2573,7 +2574,7 @@ static int kvm_vm_set_mem_attributes(struct kvm *kvm, gfn_t start, gfn_t end,
> * Reserve memory ahead of time to avoid having to deal with failures
> * partway through setting the new attributes.
> */
> - for (i = start; i < end; i++) {
> + for (i = start; entry && i < end; i++) {
> r = xa_reserve(&kvm->mem_attr_array, i, GFP_KERNEL_ACCOUNT);
> if (r)
> goto out_unlock;
Right, storing NULL just erases and never allocates, so no reservation is
needed when clearing. I folded that into the same patch (the reservation loop
becomes "for (i = start; entry && i < end; i++)").
Thanks,
Zeng Chi
prev parent reply other threads:[~2026-08-28 10:51 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-27 11:05 Zeng Chi
2026-08-27 18:37 ` Sean Christopherson
2026-08-28 10:27 ` [PATCH v2] KVM: Don't treat reserved xarray entries as having memory attributes Zeng Chi
2026-08-28 17:15 ` Sean Christopherson
2026-08-28 18:17 ` Sean Christopherson
2026-08-28 10:51 ` Zeng Chi [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=7d00f8a3-0d35-49ce-83ab-29b94a3490ed@163.com \
--to=zeng_chi911@163.com \
--cc=chao.p.peng@linux.intel.com \
--cc=kvm@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=pbonzini@redhat.com \
--cc=seanjc@google.com \
--cc=zengchi@kylinos.cn \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®