* [RFC PATCH v2 00/11] KVM: Allow alternative providers of guest_memfd backed by PFNMAP memory
@ 2026-07-20 11:03 David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 01/11] KVM: selftests: sev_smoke_test: Only run VM types the host offers David Woodhouse
` (3 more replies)
0 siblings, 4 replies; 28+ messages in thread
From: David Woodhouse @ 2026-07-20 11:03 UTC (permalink / raw)
To: Sean Christopherson, Paolo Bonzini, kvm
Cc: linux-kernel, iommu, linux-kselftest, linux-media, dri-devel,
linaro-mm-sig, Thomas Gleixner, Ingo Molnar, Borislav Petkov,
Dave Hansen, H . Peter Anvin, x86, Jason Gunthorpe, Kevin Tian,
Joerg Roedel, Will Deacon, Robin Murphy, Shuah Khan,
Sumit Semwal, Christian König, Tom Lendacky, Tycho Andersen,
Kees Cook, Ashish Kalra, Connor Williamson, Andrew Morton,
Zhou Yuhang, Zi Li, Dan Williams, Eric Biggers, Colin Ian King,
connordw, graf, fgriffo
For dedicated hosting environments, there are a few reasons to prefer
using external PFNMAP memory for guests, rather than kernel-managed
memory:
• No struct page overheads
• Easily managed in 1GiB pages (or larger chunks); no fragmentation/THP
• Faster kexec/KHO live update (every millisecond spent faffing around
with memory management is an extra millisecond of steal time taken
from guests)
Today, there is only one implementation of guest_memfd: the internal
shmem-based page cache in virt/kvm/guest_memfd.c. This series allows
for other code to provide "a guest_memfd". This is an early path-
finding proof of concept tested with AMD SEV-SNP as well as non-CoCo
guests on x86 and arm64. I'm planning to build a memory-based file
system which gives files for both guest_memfd as well as memory-based
storage to allow things like userspace VM and orchestrator state
to be passed over KHO/kexec.
This series lifts guest_memfd's operations into a small ops struct and
converts KVM's own gmem to use it, so it becomes one implementation
among peers rather than the only game in town. It also adapts the
SEV-SNP fast paths to page-less operation, plumbs iommufd to map a
guest_memfd-backing dma-buf as CPU RAM (not MMIO), and adds the
provider->KVM revocation primitive that makes overcommit-style reclaim
possible.
It's largely AI-built at the moment. My main design concerns are around
how we tell that a given file is "a guest_memfd", and how we plug into
IOMMUFD. Presenting a dma-buf was the less intrusive choice because that
path already has support for taking pages away, but I'm less convinced
that it's the cleaner choice in the long term.
I'll continue to bikeshed it myself, but this is at least functional
enough to show the direction and solicit further opinions...
https://git.infradead.org/?p=users/dwmw2/linux.git;a=shortlog;h=gmem-provider-v2
Changes since v1 (mostly Sashiko feedback):
Core (KVM / SEV):
- populate: drop the transient page reference after ->post_populate()
rather than before, so the callback still sees the ref it relies on
- kvm_gmem_invalidate_range(): walk every address space and all
memslots in the range, not just address space 0 / a single slot
- SEV: the guest_memfd cache flush hit a WARN_ON_ONCE on a failed
temporary mapping; that path is reachable from userspace (hole-punch),
so downgrade to pr_warn_ratelimited to avoid a panic_on_warn DoS
- Export file_is_kvm() (EXPORT_SYMBOL_GPL) so an external provider can
confirm the fd handed to it at setup really is a KVM instance
Sample provider (samples/kvm/gmem_provider.c):
- Serialize per-fd state (info->kvm, bound_slot, present bitmap) under a
new info->lock; v1 accessed info->kvm without locking, racing
bind/unbind against release
- Validate setup.kvm_fd via file_is_kvm() before using it
- Fix bitmap_zalloc(unsigned int) truncation -> kvzalloc with
BITS_TO_LONGS
- Fix integer underflow in SET_PRESENT when start_index < pgoff
- gmem_release: zap active guest mappings + clear slot->gmem.file before
freeing the backing
- Pin THIS_MODULE while a provider fd is open
- gmem_mmap: refuse if the fd was not created mmap-capable
- gmem_max_order: clamp by the absent bitmap
- Bind/unbind: symmetric psmash + rmp_make_shared with proper 2MiB
alignment
Selftests:
- Revoke test: assert the revoked access faults unconditionally
- iommufd test: mark the range KVM_MEMORY_ATTRIBUTE_PRIVATE so the guest
actually exercises the provider path rather than the shared HVA mmap
- NVMe DMA test: read the vendor ID from PCI config and add compiler
barriers before the doorbell writes
Connor Williamson (1):
KVM: SEV: Remove struct page dependency from SNP gmem paths
David Woodhouse (10):
KVM: selftests: sev_smoke_test: Only run VM types the host offers
KVM: selftests: sev_init2_tests: Derive SEV availability from KVM
KVM: guest_memfd: Introduce guest memory ops and route native gmem through them
iommufd: Look up private-interconnect phys via exporter symbols
iommufd: Plumb dma-buf memory-type (RAM vs MMIO) through the phys map
KVM: guest_memfd: Add ops-driven page revocation
samples/kvm: Add guest_memfd backing sample
selftests/kvm: gmem_provider KVM-only tests
selftests/kvm: gmem_provider iommufd tests
samples/kvm, selftests/kvm: Allow the gmem_provider NVMe DMA test on arm64
arch/x86/kvm/Makefile | 5 +-
arch/x86/kvm/svm/sev.c | 93 ++-
arch/x86/virt/svm/sev.c | 27 +-
drivers/iommu/iommufd/io_pagetable.h | 7 +
drivers/iommu/iommufd/pages.c | 49 +-
include/linux/kvm_host.h | 91 +++
samples/Kconfig | 18 +
samples/Makefile | 1 +
samples/kvm/Makefile | 2 +
samples/kvm/gmem_provider.c | 772 +++++++++++++++++++++
samples/kvm/gmem_provider.h | 53 ++
tools/testing/selftests/kvm/Makefile.kvm | 6 +
.../selftests/kvm/gmem_provider_nvme_dma_test.c | 376 ++++++++++
.../kvm/x86/gmem_provider_hugepage_test.c | 130 ++++
.../selftests/kvm/x86/gmem_provider_iommufd_test.c | 158 +++++
.../selftests/kvm/x86/gmem_provider_revoke_test.c | 135 ++++
.../testing/selftests/kvm/x86/gmem_provider_test.c | 195 ++++++
.../selftests/kvm/x86/gmem_provider_vfio_test.c | 134 ++++
tools/testing/selftests/kvm/x86/sev_init2_tests.c | 16 +-
tools/testing/selftests/kvm/x86/sev_smoke_test.c | 9 +-
virt/kvm/guest_memfd.c | 391 +++++++++--
virt/kvm/kvm_main.c | 8 +-
virt/kvm/kvm_mm.h | 4 +-
23 files changed, 2577 insertions(+), 103 deletions(-)
^ permalink raw reply [flat|nested] 28+ messages in thread* [RFC PATCH v2 01/11] KVM: selftests: sev_smoke_test: Only run VM types the host offers 2026-07-20 11:03 [RFC PATCH v2 00/11] KVM: Allow alternative providers of guest_memfd backed by PFNMAP memory David Woodhouse @ 2026-07-20 11:03 ` David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 02/11] KVM: selftests: sev_init2_tests: Derive SEV availability from KVM David Woodhouse ` (9 more replies) 2026-07-20 15:11 ` [RFC PATCH v2 00/11] KVM: Allow alternative providers of guest_memfd backed by PFNMAP memory Paolo Bonzini ` (2 subsequent siblings) 3 siblings, 10 replies; 28+ messages in thread From: David Woodhouse @ 2026-07-20 11:03 UTC (permalink / raw) To: Sean Christopherson, Paolo Bonzini, kvm Cc: linux-kernel, iommu, linux-kselftest, linux-media, dri-devel, linaro-mm-sig, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, x86, Jason Gunthorpe, Kevin Tian, Joerg Roedel, Will Deacon, Robin Murphy, Shuah Khan, Sumit Semwal, Christian König, Tom Lendacky, Tycho Andersen, Kees Cook, Ashish Kalra, Connor Williamson, Andrew Morton, Zhou Yuhang, Zi Li, Dan Williams, Eric Biggers, Colin Ian King, connordw, graf, fgriffo, David Woodhouse From: David Woodhouse <dwmw@amazon.co.uk> sev_smoke_test ran the plain SEV subtest unconditionally, gated only on the X86_FEATURE_SEV CPUID bit, while gating SEV-ES and SNP on the KVM_CAP_VM_TYPES bits. CPUID reporting SEV does not mean KVM offers the SEV VM type: when all SEV ASIDs are assigned to SEV-SNP, KVM_X86_SEV_VM is unavailable even though X86_FEATURE_SEV is set. On such a host the test aborts in KVM_CREATE_VM instead of exercising the available modes. Gate the SEV subtest on KVM_CAP_VM_TYPES like the others, so the test runs the VM types the host actually offers. Reviewed-by: Tycho Andersen (AMD) <tycho@kernel.org> Signed-off-by: David Woodhouse <dwmw@amazon.co.uk> --- tools/testing/selftests/kvm/x86/sev_smoke_test.c | 9 ++++++++- 1 file changed, 8 insertions(+), 1 deletion(-) diff --git a/tools/testing/selftests/kvm/x86/sev_smoke_test.c b/tools/testing/selftests/kvm/x86/sev_smoke_test.c index 6b2cbe2a90b7..bf27b6187afa 100644 --- a/tools/testing/selftests/kvm/x86/sev_smoke_test.c +++ b/tools/testing/selftests/kvm/x86/sev_smoke_test.c @@ -247,7 +247,14 @@ int main(int argc, char *argv[]) { TEST_REQUIRE(kvm_cpu_has(X86_FEATURE_SEV)); - test_sev_smoke(guest_sev_code, KVM_X86_SEV_VM, 0); + /* + * Only exercise VM types the host actually offers. CPUID reporting + * SEV does not guarantee KVM offers the SEV VM type: when all SEV + * ASIDs are assigned to SEV-SNP, KVM_X86_SEV_VM is unavailable even + * though X86_FEATURE_SEV is set. Gate every type on KVM_CAP_VM_TYPES. + */ + if (kvm_check_cap(KVM_CAP_VM_TYPES) & BIT(KVM_X86_SEV_VM)) + test_sev_smoke(guest_sev_code, KVM_X86_SEV_VM, 0); if (kvm_check_cap(KVM_CAP_VM_TYPES) & BIT(KVM_X86_SEV_ES_VM)) test_sev_smoke(guest_sev_es_code, KVM_X86_SEV_ES_VM, SEV_POLICY_ES); base-commit: 0e35b9b6ec0ffcc5e23cbdec09f5c622ad532b53 -- 2.55.0 ^ permalink raw reply [flat|nested] 28+ messages in thread
* [RFC PATCH v2 02/11] KVM: selftests: sev_init2_tests: Derive SEV availability from KVM 2026-07-20 11:03 ` [RFC PATCH v2 01/11] KVM: selftests: sev_smoke_test: Only run VM types the host offers David Woodhouse @ 2026-07-20 11:03 ` David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 03/11] KVM: SEV: Remove struct page dependency from SNP gmem paths David Woodhouse ` (8 subsequent siblings) 9 siblings, 0 replies; 28+ messages in thread From: David Woodhouse @ 2026-07-20 11:03 UTC (permalink / raw) To: Sean Christopherson, Paolo Bonzini, kvm Cc: linux-kernel, iommu, linux-kselftest, linux-media, dri-devel, linaro-mm-sig, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, x86, Jason Gunthorpe, Kevin Tian, Joerg Roedel, Will Deacon, Robin Murphy, Shuah Khan, Sumit Semwal, Christian König, Tom Lendacky, Tycho Andersen, Kees Cook, Ashish Kalra, Connor Williamson, Andrew Morton, Zhou Yuhang, Zi Li, Dan Williams, Eric Biggers, Colin Ian King, connordw, graf, fgriffo, David Woodhouse From: David Woodhouse <dwmw@amazon.co.uk> The test asserted that the X86_FEATURE_SEV CPUID bit exactly matches whether KVM offers KVM_X86_SEV_VM. That is not an invariant: when all SEV ASIDs are assigned to SEV-SNP, KVM does not offer the SEV VM type even though CPUID reports SEV, so the test aborts on an SNP-only host. Derive SEV availability from KVM_CAP_VM_TYPES (as already done for SEV-ES and SNP), assert only the one-way implication that a type offered by KVM is also reported in CPUID, and TEST_REQUIRE() the SEV VM type so the test skips cleanly when it is unavailable. Reviewed-by: Tycho Andersen (AMD) <tycho@kernel.org> Signed-off-by: David Woodhouse <dwmw@amazon.co.uk> --- .../testing/selftests/kvm/x86/sev_init2_tests.c | 16 +++++++++++----- 1 file changed, 11 insertions(+), 5 deletions(-) diff --git a/tools/testing/selftests/kvm/x86/sev_init2_tests.c b/tools/testing/selftests/kvm/x86/sev_init2_tests.c index 8db88c355f16..689390c10f7c 100644 --- a/tools/testing/selftests/kvm/x86/sev_init2_tests.c +++ b/tools/testing/selftests/kvm/x86/sev_init2_tests.c @@ -130,12 +130,18 @@ int main(int argc, char *argv[]) KVM_X86_SEV_VMSA_FEATURES, &supported_vmsa_features); - have_sev = kvm_cpu_has(X86_FEATURE_SEV); - TEST_ASSERT(have_sev == !!(kvm_check_cap(KVM_CAP_VM_TYPES) & BIT(KVM_X86_SEV_VM)), - "sev: KVM_CAP_VM_TYPES (%x) does not match cpuid (checking %x)", - kvm_check_cap(KVM_CAP_VM_TYPES), 1 << KVM_X86_SEV_VM); + /* + * Whether a VM type is available depends on KVM, not just CPUID: e.g. + * when all SEV ASIDs are assigned to SEV-SNP, KVM does not offer the + * SEV VM type even though X86_FEATURE_SEV is set. Derive availability + * from KVM_CAP_VM_TYPES and only assert the one-way implication that a + * type offered by KVM must also be reported in CPUID. + */ + have_sev = kvm_check_cap(KVM_CAP_VM_TYPES) & BIT(KVM_X86_SEV_VM); + TEST_ASSERT(!have_sev || kvm_cpu_has(X86_FEATURE_SEV), + "sev: SEV_VM supported without SEV in CPUID"); - TEST_REQUIRE(kvm_check_cap(KVM_CAP_VM_TYPES) & BIT(KVM_X86_SEV_VM)); + TEST_REQUIRE(have_sev); have_sev_es = kvm_check_cap(KVM_CAP_VM_TYPES) & BIT(KVM_X86_SEV_ES_VM); TEST_ASSERT(!have_sev_es || kvm_cpu_has(X86_FEATURE_SEV_ES), -- 2.55.0 ^ permalink raw reply [flat|nested] 28+ messages in thread
* [RFC PATCH v2 03/11] KVM: SEV: Remove struct page dependency from SNP gmem paths 2026-07-20 11:03 ` [RFC PATCH v2 01/11] KVM: selftests: sev_smoke_test: Only run VM types the host offers David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 02/11] KVM: selftests: sev_init2_tests: Derive SEV availability from KVM David Woodhouse @ 2026-07-20 11:03 ` David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 04/11] KVM: guest_memfd: Introduce guest memory ops and route native gmem through them David Woodhouse ` (7 subsequent siblings) 9 siblings, 0 replies; 28+ messages in thread From: David Woodhouse @ 2026-07-20 11:03 UTC (permalink / raw) To: Sean Christopherson, Paolo Bonzini, kvm Cc: linux-kernel, iommu, linux-kselftest, linux-media, dri-devel, linaro-mm-sig, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, x86, Jason Gunthorpe, Kevin Tian, Joerg Roedel, Will Deacon, Robin Murphy, Shuah Khan, Sumit Semwal, Christian König, Tom Lendacky, Tycho Andersen, Kees Cook, Ashish Kalra, Connor Williamson, Andrew Morton, Zhou Yuhang, Zi Li, Dan Williams, Eric Biggers, Colin Ian King, connordw, graf, fgriffo, David Woodhouse From: Connor Williamson <connordw@amazon.com> The SNP paths assume every guest_memfd PFN has a struct page and a direct map alias. That is not true for memory handed out by an external guest_memfd provider that backs guest RAM with device or carved-out physical memory (e.g. via mem=), which has no struct page. Relax the RMP helpers in arch/x86/virt/svm/sev.c: - psmash(): drop the pfn_valid() gate. PSMASH operates on the RMP entry for a physical address and touches neither the direct map nor struct page. - adjust_direct_map(): return 0 instead of -EINVAL for a page-less PFN, which has no direct map alias to split. Checked before the 2M range test, which would otherwise reject the span. - __snp_leak_pages(): compute the struct page per iteration and skip list insertion when it is NULL; a page-less PFN cannot be chained onto the leaked pages list. And the KVM SNP gmem paths in arch/x86/kvm/svm/sev.c: - sev_gmem_map_pfn()/sev_gmem_unmap_pfn() fall back to memremap() for the launch-update copy and the CPUID error read-back in sev_gmem_post_populate(). The destination PFN is mapped before the atomic kmap_local_page() of the source, since memremap() may sleep. - sev_clflush_pfn() falls back to memremap() for the cache flush in sev_gmem_invalidate(). No functional change for page-backed PFNs. Signed-off-by: Connor Williamson <connordw@amazon.com> Signed-off-by: David Woodhouse <dwmw@amazon.co.uk> --- arch/x86/kvm/svm/sev.c | 93 +++++++++++++++++++++++++++++++++++------ arch/x86/virt/svm/sev.c | 27 ++++++++---- 2 files changed, 100 insertions(+), 20 deletions(-) diff --git a/arch/x86/kvm/svm/sev.c b/arch/x86/kvm/svm/sev.c index 427229347876..125779c82bc4 100644 --- a/arch/x86/kvm/svm/sev.c +++ b/arch/x86/kvm/svm/sev.c @@ -12,6 +12,7 @@ #include <linux/kvm_host.h> #include <linux/kernel.h> #include <linux/highmem.h> +#include <linux/io.h> #include <linux/psp.h> #include <linux/psp-sev.h> #include <linux/pagemap.h> @@ -2321,6 +2322,30 @@ struct sev_gmem_populate_args { int fw_error; }; +/* + * Map a guest_memfd PFN for CPU access. A PFN provided by an external + * guest_memfd provider may have no struct page, so fall back to memremap() + * for those. Pair each call with sev_gmem_unmap_pfn(). + */ +static void *sev_gmem_map_pfn(kvm_pfn_t pfn) +{ + if (pfn_valid(pfn)) + return kmap_local_pfn(pfn); + + return memremap(pfn_to_hpa(pfn), PAGE_SIZE, MEMREMAP_WB); +} + +static void sev_gmem_unmap_pfn(kvm_pfn_t pfn, void *vaddr) +{ + if (!vaddr) + return; + + if (pfn_valid(pfn)) + kunmap_local(vaddr); + else + memunmap(vaddr); +} + static int sev_gmem_post_populate(struct kvm *kvm, gfn_t gfn, kvm_pfn_t pfn, struct page *src_page, void *opaque) { @@ -2343,13 +2368,23 @@ static int sev_gmem_post_populate(struct kvm *kvm, gfn_t gfn, kvm_pfn_t pfn, } if (src_page) { - void *src_vaddr = kmap_local_page(src_page); - void *dst_vaddr = kmap_local_pfn(pfn); + void *dst_vaddr = sev_gmem_map_pfn(pfn); + void *src_vaddr; - memcpy(dst_vaddr, src_vaddr, PAGE_SIZE); + if (!dst_vaddr) { + ret = -ENOMEM; + goto out; + } - kunmap_local(dst_vaddr); + /* + * Map the destination (which may be page-less and thus sleep in + * memremap()) before the atomic kmap_local_page() of the source. + */ + src_vaddr = kmap_local_page(src_page); + memcpy(dst_vaddr, src_vaddr, PAGE_SIZE); kunmap_local(src_vaddr); + + sev_gmem_unmap_pfn(pfn, dst_vaddr); } ret = rmp_make_private(pfn, gfn << PAGE_SHIFT, PG_LEVEL_4K, @@ -2379,14 +2414,17 @@ static int sev_gmem_post_populate(struct kvm *kvm, gfn_t gfn, kvm_pfn_t pfn, if (ret && !snp_page_reclaim(kvm, pfn) && sev_populate_args->type == KVM_SEV_SNP_PAGE_TYPE_CPUID && sev_populate_args->fw_error == SEV_RET_INVALID_PARAM) { - void *src_vaddr = kmap_local_page(src_page); - void *dst_vaddr = kmap_local_pfn(pfn); + void *dst_vaddr = sev_gmem_map_pfn(pfn); + void *src_vaddr; - memcpy(src_vaddr, dst_vaddr, PAGE_SIZE); - set_page_dirty(src_page); + if (dst_vaddr) { + src_vaddr = kmap_local_page(src_page); + memcpy(src_vaddr, dst_vaddr, PAGE_SIZE); + set_page_dirty(src_page); + kunmap_local(src_vaddr); - kunmap_local(dst_vaddr); - kunmap_local(src_vaddr); + sev_gmem_unmap_pfn(pfn, dst_vaddr); + } } out: @@ -5117,6 +5155,38 @@ int sev_gmem_prepare(struct kvm *kvm, kvm_pfn_t pfn, gfn_t gfn, int max_order) return 0; } +/* + * Flush the CPU caches for a guest_memfd PFN. A PFN from an external + * guest_memfd provider may have no direct map alias, so fall back to a + * temporary memremap() mapping for the cache flush. + */ +static void sev_clflush_pfn(kvm_pfn_t pfn, size_t size) +{ + void *va; + + if (pfn_valid(pfn)) { + clflush_cache_range(__va(pfn_to_hpa(pfn)), size); + return; + } + + va = memremap(pfn_to_hpa(pfn), size, MEMREMAP_WB); + if (!va) { + /* + * memremap() can fail under memory pressure (vmalloc space + * exhaustion). Callers reach us during guest_memfd hole- + * punching, so this is user-triggerable; ratelimit rather + * than WARN, otherwise panic_on_warn hosts turn it into a + * denial of service. + */ + pr_warn_ratelimited("SEV: memremap failed for cache flush pfn=%llx size=%zu\n", + pfn, size); + return; + } + + clflush_cache_range(va, size); + memunmap(va); +} + void sev_gmem_invalidate(kvm_pfn_t start, kvm_pfn_t end) { kvm_pfn_t pfn; @@ -5172,8 +5242,7 @@ void sev_gmem_invalidate(kvm_pfn_t start, kvm_pfn_t end) * cache entries for these pages before free'ing them back to * the host. */ - clflush_cache_range(__va(pfn_to_hpa(pfn)), - use_2m_update ? PMD_SIZE : PAGE_SIZE); + sev_clflush_pfn(pfn, use_2m_update ? PMD_SIZE : PAGE_SIZE); next_pfn: pfn += use_2m_update ? PTRS_PER_PMD : 1; cond_resched(); diff --git a/arch/x86/virt/svm/sev.c b/arch/x86/virt/svm/sev.c index 8bcdce98f6dc..3623b52fbb11 100644 --- a/arch/x86/virt/svm/sev.c +++ b/arch/x86/virt/svm/sev.c @@ -897,8 +897,11 @@ int psmash(u64 pfn) if (!cc_platform_has(CC_ATTR_HOST_SEV_SNP)) return -ENODEV; - if (!pfn_valid(pfn)) - return -EINVAL; + /* + * Do not require a struct page. PSMASH operates on the RMP entry for + * a physical address; a valid PFN with no struct page backing (e.g. + * memory handed out by an external guest_memfd provider) is fine. + */ /* Binutils version 2.36 supports the PSMASH mnemonic. */ asm volatile(".byte 0xF3, 0x0F, 0x01, 0xFF" @@ -953,8 +956,13 @@ static int adjust_direct_map(u64 pfn, int rmp_level) if (WARN_ON_ONCE(rmp_level > PG_LEVEL_2M)) return -EINVAL; + /* + * A PFN with no struct page has no direct map alias to split, so + * there is nothing to do. This is checked before the 2M range test + * below, which would otherwise reject the whole span. + */ if (!pfn_valid(pfn)) - return -EINVAL; + return 0; if (rmp_level == PG_LEVEL_2M && (!IS_ALIGNED(pfn, PTRS_PER_PMD) || !pfn_valid(pfn + PTRS_PER_PMD - 1))) @@ -1060,32 +1068,35 @@ EXPORT_SYMBOL_GPL(rmp_make_shared); void __snp_leak_pages(u64 pfn, unsigned int npages, bool dump_rmp) { - struct page *page = pfn_to_page(pfn); - pr_warn("Leaking PFN range 0x%llx-0x%llx\n", pfn, pfn + npages); spin_lock(&snp_leaked_pages_list_lock); while (npages--) { + struct page *page = pfn_valid(pfn) ? pfn_to_page(pfn) : NULL; /* * Reuse the page's buddy list for chaining into the leaked * pages list. This page should not be on a free list currently * and is also unsafe to be added to a free list. + * + * A page-less PFN (e.g. memory backed by an external + * guest_memfd provider) has no struct page to chain, so it is + * only accounted and, optionally, its RMP entry dumped. */ - if (likely(!PageCompound(page)) || + if (page && + (likely(!PageCompound(page)) || /* * Skip inserting tail pages of compound page as * page->buddy_list of tail pages is not usable. */ - (PageHead(page) && compound_nr(page) <= npages)) + (PageHead(page) && compound_nr(page) <= npages))) list_add_tail(&page->buddy_list, &snp_leaked_pages_list); if (dump_rmp) dump_rmpentry(pfn); snp_nr_leaked_pages++; pfn++; - page++; } spin_unlock(&snp_leaked_pages_list_lock); } -- 2.55.0 ^ permalink raw reply [flat|nested] 28+ messages in thread
* [RFC PATCH v2 04/11] KVM: guest_memfd: Introduce guest memory ops and route native gmem through them 2026-07-20 11:03 ` [RFC PATCH v2 01/11] KVM: selftests: sev_smoke_test: Only run VM types the host offers David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 02/11] KVM: selftests: sev_init2_tests: Derive SEV availability from KVM David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 03/11] KVM: SEV: Remove struct page dependency from SNP gmem paths David Woodhouse @ 2026-07-20 11:03 ` David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 05/11] iommufd: Look up private-interconnect phys via exporter symbols David Woodhouse ` (6 subsequent siblings) 9 siblings, 0 replies; 28+ messages in thread From: David Woodhouse @ 2026-07-20 11:03 UTC (permalink / raw) To: Sean Christopherson, Paolo Bonzini, kvm Cc: linux-kernel, iommu, linux-kselftest, linux-media, dri-devel, linaro-mm-sig, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, x86, Jason Gunthorpe, Kevin Tian, Joerg Roedel, Will Deacon, Robin Murphy, Shuah Khan, Sumit Semwal, Christian König, Tom Lendacky, Tycho Andersen, Kees Cook, Ashish Kalra, Connor Williamson, Andrew Morton, Zhou Yuhang, Zi Li, Dan Williams, Eric Biggers, Colin Ian King, connordw, graf, fgriffo, David Woodhouse From: David Woodhouse <dwmw@amazon.co.uk> Guest_memfd today has exactly one backing implementation, hard-wired into virt/kvm/guest_memfd.c. To let other subsystems supply guest memory from an alternate source -- for example a physically contiguous range without struct page, or memory shared with a peer device -- lift the operations into a small vtable and route the built-in gmem through it. Introduce struct kvm_gmem_ops in <linux/kvm_host.h>: bind / unbind / get_pfn / populate / invalidate for the memory-side operations, plus release / mmap / fallocate for the file-level ones. Every file that can back guest memory is created with one shared struct file_operations, kvm_gmem_fops, and file->private_data points at a struct kvm_gmem_backing whose ->ops field is the per-instance vtable. This mirrors the dma-buf / socket / iommufd pattern: a canonical fops identifies the file (is_kvm_gmem_file() is a plain fops-equality check), and per-instance behaviour is reached via private_data. Every callback on kvm_gmem_fops -- release, mmap, fallocate, and the memory-side dispatchers already present -- reaches through the ops table, so foreign implementations get a private hook for each and KVM code never has to distinguish "native vs foreign". Native gmem becomes one implementation of the ops (kvm_gmem_native_ops), with the current release/mmap/fallocate bodies moved into kvm_gmem_native_release/_mmap/_fallocate. struct gmem_file gains a struct kvm_gmem_backing as its first member so file->private_data can be reinterpreted uniformly across implementations. .populate is deliberately left NULL for the built-in path; kvm_gmem_populate() keeps an inline fast path because the current folio-locked + attribute-checked shape of post_populate() does not fit the ops signature cleanly. Wiring that flow through ops is a separate cleanup. .invalidate is likewise NULL for the built-in path, because its invalidations run through inode->i_mapping (truncate, hwpoison, free_folio) rather than an out-of-band notification. kvm_gmem_unbind() gains a struct kvm * parameter so the dispatcher can pass it to ops->unbind() without keeping a duplicate copy anywhere. The three existing callers in kvm_main.c already have @kvm in scope. Based on a proof-of-concept by Connor Williamson <connordw@amazon.com>. Signed-off-by: David Woodhouse (Kiro) <dwmw@amazon.co.uk> --- arch/x86/kvm/Makefile | 5 +- include/linux/kvm_host.h | 82 ++++++++++ virt/kvm/guest_memfd.c | 339 +++++++++++++++++++++++++++++++-------- virt/kvm/kvm_main.c | 8 +- virt/kvm/kvm_mm.h | 4 +- 5 files changed, 365 insertions(+), 73 deletions(-) diff --git a/arch/x86/kvm/Makefile b/arch/x86/kvm/Makefile index 77337c37324b..36a319f382b7 100644 --- a/arch/x86/kvm/Makefile +++ b/arch/x86/kvm/Makefile @@ -63,7 +63,10 @@ exports_grep_trailer := --include='*.[ch]' -nrw $(srctree)/virt/kvm $(srctree)/a -e kvm_write_track_remove_gfn \ -e kvm_get_kvm \ -e kvm_get_kvm_safe \ - -e kvm_put_kvm + -e kvm_put_kvm \ + -e kvm_gmem_invalidate_range \ + -e kvm_gmem_fops \ + -e file_is_kvm # Force grep to emit a goofy group separator that can in turn be replaced with # the above newline macro (newlines in Make are a nightmare). Note, grep only diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h index ab8cfaec82d3..b33d2f475042 100644 --- a/include/linux/kvm_host.h +++ b/include/linux/kvm_host.h @@ -621,6 +621,88 @@ static inline bool kvm_slot_has_gmem(const struct kvm_memory_slot *slot) return slot && (slot->flags & KVM_MEM_GUEST_MEMFD); } +#ifdef CONFIG_KVM_GUEST_MEMFD +/* + * struct kvm_gmem_ops - vtable a guest_memfd backing implementation exposes. + * + * Every file that can back KVM guest memory -- KVM's own guest_memfd (see + * kvm_gmem_create()) and any other subsystem opting in -- is created with + * the shared kvm_gmem_fops (see is_kvm_gmem_file()). file->private_data on + * such a file points at a struct kvm_gmem_backing whose ->ops field is the + * per-instance vtable below. KVM never looks any further; the rest of the + * per-file state (allocator, refcount, private xarray, ...) is up to the + * implementation and is reached via container_of() from the backing. + * + * The memory-side callbacks (bind, unbind, get_pfn, populate, invalidate) + * are the reason the interface exists. The file-level callbacks (release, + * mmap, fallocate, ...) are dispatched by kvm_gmem_fops so that every + * implementation gets a private hook -- see the dispatchers in + * virt/kvm/guest_memfd.c. + * + * get_pfn() contract: + * - *pfn is the host PFN backing @gfn. It need not have a struct page. + * - *max_order is the largest order KVM may use to map starting at @gfn. + * The implementation MUST clamp it to the physical contiguity and the + * HPA/GPA alignment of the backing, so KVM never builds a huge SPTE that + * spans discontiguous physical memory. + * - If *page is non-NULL on return the caller owns that page reference and + * is responsible for put_page() after use. If *page is left NULL the PFN + * is treated as non-refcounted, and its lifetime is owned by the + * implementation across bind()/unbind(). + * + * Memory intended to back guest RAM MUST be reported as E820_TYPE_RAM by the + * host so KVM maps it write-back (and applies the memory-encryption bit on + * encrypted hosts) rather than treating it as MMIO. + */ +struct kvm_gmem_ops { + /* Memory operations. */ + int (*bind)(struct file *file, struct kvm *kvm, + struct kvm_memory_slot *slot, loff_t offset); + void (*unbind)(struct file *file, struct kvm *kvm, + struct kvm_memory_slot *slot); + int (*get_pfn)(struct file *file, struct kvm *kvm, + struct kvm_memory_slot *slot, gfn_t gfn, + kvm_pfn_t *pfn, struct page **page, int *max_order); + int (*populate)(struct file *file, struct kvm *kvm, + struct kvm_memory_slot *slot, gfn_t gfn, + kvm_pfn_t *pfn, struct page *src_page, int order); + void (*invalidate)(struct file *file, struct kvm *kvm, + struct kvm_memory_slot *slot, + gfn_t start, gfn_t end); + + /* + * File-level operations, dispatched by kvm_gmem_fops. These are + * per-instance analogues of the like-named file_operations members, + * with the same semantics. A NULL entry causes the dispatcher to + * fail the file-level operation with an appropriate errno. + */ + void (*release)(struct file *file); + int (*mmap)(struct file *file, struct vm_area_struct *vma); + long (*fallocate)(struct file *file, int mode, loff_t offset, + loff_t len); + long (*ioctl)(struct file *file, unsigned int cmd, unsigned long arg); +}; + +/* + * Anchor placed at file->private_data of every guest_memfd-backing file. + * KVM's own gmem embeds it as the first member of struct gmem_file; other + * implementations do the same in their own per-fd state. is_kvm_gmem_file() + * (fops-identity check) proves the reinterpretation is safe. + */ +struct kvm_gmem_backing { + const struct kvm_gmem_ops *ops; +}; + +/* Defined in guest_memfd.c; single canonical fops for every gmem file. */ +extern struct file_operations kvm_gmem_fops; + +static inline bool is_kvm_gmem_file(struct file *file) +{ + return file && file->f_op == &kvm_gmem_fops; +} + +#endif /* CONFIG_KVM_GUEST_MEMFD */ + static inline bool kvm_slot_dirty_track_enabled(const struct kvm_memory_slot *slot) { return slot->flags & KVM_MEM_LOG_DIRTY_PAGES; diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c index db57c5766ab6..057f90eb835a 100644 --- a/virt/kvm/guest_memfd.c +++ b/virt/kvm/guest_memfd.c @@ -22,11 +22,22 @@ static struct vfsmount *kvm_gmem_mnt; * specific to its associated VM, e.g. memslots=>gmem bindings. */ struct gmem_file { + /* + * MUST be first: file->private_data points here. + * is_kvm_gmem_file(file) proves this reinterpretation is safe. + */ + struct kvm_gmem_backing backing; + struct kvm *kvm; struct xarray bindings; struct list_head entry; }; +static inline struct gmem_file *gmem_file_of(struct file *file) +{ + return container_of(file->private_data, struct gmem_file, backing); +} + struct gmem_inode { struct shared_policy policy; struct inode vfs_inode; @@ -296,7 +307,7 @@ static long kvm_gmem_allocate(struct inode *inode, loff_t offset, loff_t len) return r; } -static long kvm_gmem_fallocate(struct file *file, int mode, loff_t offset, +static long kvm_gmem_native_fallocate(struct file *file, int mode, loff_t offset, loff_t len) { int ret; @@ -320,9 +331,10 @@ static long kvm_gmem_fallocate(struct file *file, int mode, loff_t offset, return ret; } -static int kvm_gmem_release(struct inode *inode, struct file *file) +static void kvm_gmem_native_release(struct file *file) { - struct gmem_file *f = file->private_data; + struct inode *inode = file_inode(file); + struct gmem_file *f = gmem_file_of(file); struct kvm_memory_slot *slot; struct kvm *kvm = f->kvm; unsigned long index; @@ -367,7 +379,6 @@ static int kvm_gmem_release(struct inode *inode, struct file *file) kvm_put_kvm(kvm); - return 0; } static inline struct file *kvm_gmem_get_file(struct kvm_memory_slot *slot) @@ -466,7 +477,7 @@ static const struct vm_operations_struct kvm_gmem_vm_ops = { #endif }; -static int kvm_gmem_mmap(struct file *file, struct vm_area_struct *vma) +static int kvm_gmem_native_mmap(struct file *file, struct vm_area_struct *vma) { if (!kvm_gmem_supports_mmap(file_inode(file))) return -ENODEV; @@ -481,12 +492,61 @@ static int kvm_gmem_mmap(struct file *file, struct vm_area_struct *vma) return 0; } -static struct file_operations kvm_gmem_fops = { - .mmap = kvm_gmem_mmap, +/* File-op dispatchers — thin: they all go through the backing's ops table. */ + +static int kvm_gmem_fops_release(struct inode *inode, struct file *file) +{ + struct kvm_gmem_backing *b = file->private_data; + + if (b && b->ops && b->ops->release) + b->ops->release(file); + return 0; +} + +static int kvm_gmem_fops_mmap(struct file *file, struct vm_area_struct *vma) +{ + struct kvm_gmem_backing *b = file->private_data; + + if (!b || !b->ops || !b->ops->mmap) + return -ENODEV; + return b->ops->mmap(file, vma); +} + +static long kvm_gmem_fops_fallocate(struct file *file, int mode, + loff_t offset, loff_t len) +{ + struct kvm_gmem_backing *b = file->private_data; + + if (!b || !b->ops || !b->ops->fallocate) + return -EOPNOTSUPP; + return b->ops->fallocate(file, mode, offset, len); +} + +static long kvm_gmem_fops_ioctl(struct file *file, unsigned int cmd, + unsigned long arg) +{ + struct kvm_gmem_backing *b = file->private_data; + + if (!b || !b->ops || !b->ops->ioctl) + return -ENOTTY; + return b->ops->ioctl(file, cmd, arg); +} + +/* + * A single struct file_operations for every guest_memfd-backing file -- + * KVM's built-in gmem and any subsystem-supplied backing alike. All + * callbacks that can vary per implementation dispatch through the + * kvm_gmem_backing at file->private_data. + */ +struct file_operations kvm_gmem_fops = { .open = generic_file_open, - .release = kvm_gmem_release, - .fallocate = kvm_gmem_fallocate, + .release = kvm_gmem_fops_release, + .mmap = kvm_gmem_fops_mmap, + .fallocate = kvm_gmem_fops_fallocate, + .unlocked_ioctl = kvm_gmem_fops_ioctl, + .compat_ioctl = kvm_gmem_fops_ioctl, }; +EXPORT_SYMBOL_FOR_KVM_INTERNAL(kvm_gmem_fops); static int kvm_gmem_migrate_folio(struct address_space *mapping, struct folio *dst, struct folio *src, @@ -557,6 +617,43 @@ bool __weak kvm_arch_supports_gmem_init_shared(struct kvm *kvm) return true; } +static int kvm_gmem_native_bind(struct file *file, struct kvm *kvm, + struct kvm_memory_slot *slot, loff_t offset); +static void kvm_gmem_native_unbind(struct file *slot_file, struct kvm *kvm, + struct kvm_memory_slot *slot); +static int kvm_gmem_native_get_pfn(struct file *file, struct kvm *kvm, + struct kvm_memory_slot *slot, gfn_t gfn, + kvm_pfn_t *pfn, struct page **page, + int *max_order); +static void kvm_gmem_native_release(struct file *file); +static int kvm_gmem_native_mmap(struct file *file, + struct vm_area_struct *vma); +static long kvm_gmem_native_fallocate(struct file *file, int mode, + loff_t offset, loff_t len); + +static const struct kvm_gmem_ops kvm_gmem_native_ops = { + /* Memory operations. */ + .bind = kvm_gmem_native_bind, + .unbind = kvm_gmem_native_unbind, + .get_pfn = kvm_gmem_native_get_pfn, + /* + * .populate is deliberately NULL: kvm_gmem_populate() has an inline + * fast path that drives post_populate() with the folio locked and + * the memory-attribute check in the right place; that flow doesn't + * fit the current ops signature cleanly. Wiring it through ops + * is a separate cleanup. + * + * .invalidate is deliberately NULL: invalidations for the built-in + * implementation flow through inode->i_mapping (truncate, hwpoison, + * free_folio) and don't need an out-of-band callback from KVM. + */ + + /* File-level operations (dispatched by kvm_gmem_fops). */ + .release = kvm_gmem_native_release, + .mmap = kvm_gmem_native_mmap, + .fallocate = kvm_gmem_native_fallocate, +}; + static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags) { static const char *name = "[kvm-gmem]"; @@ -605,7 +702,8 @@ static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags) } file->f_flags |= O_LARGEFILE; - file->private_data = f; + file->private_data = &f->backing; + f->backing.ops = &kvm_gmem_native_ops; kvm_get_kvm(kvm); f->kvm = kvm; @@ -640,34 +738,27 @@ int kvm_gmem_create(struct kvm *kvm, struct kvm_create_guest_memfd *args) return __kvm_gmem_create(kvm, size, flags); } -int kvm_gmem_bind(struct kvm *kvm, struct kvm_memory_slot *slot, - unsigned int fd, uoff_t offset) +static const struct kvm_gmem_ops *kvm_gmem_get_ops(struct file *file) +{ + if (!is_kvm_gmem_file(file)) + return NULL; + return ((struct kvm_gmem_backing *)file->private_data)->ops; +} + +static int kvm_gmem_native_bind(struct file *file, struct kvm *kvm, + struct kvm_memory_slot *slot, loff_t offset) { uoff_t size = slot->npages << PAGE_SHIFT; + struct gmem_file *f = gmem_file_of(file); + struct inode *inode = file_inode(file); unsigned long start, end; - struct gmem_file *f; - struct inode *inode; - struct file *file; int r = -EINVAL; - BUILD_BUG_ON(sizeof(gpa_t) != sizeof(offset)); - BUILD_BUG_ON(sizeof(gfn_t) != sizeof(slot->gmem.pgoff)); - - file = fget(fd); - if (!file) - return -EBADF; - - if (file->f_op != &kvm_gmem_fops) - goto err; - - f = file->private_data; if (f->kvm != kvm) - goto err; - - inode = file_inode(file); + return -EINVAL; - if (!PAGE_ALIGNED(offset) || offset + size > i_size_read(inode)) - goto err; + if (offset < 0 || !PAGE_ALIGNED(offset) || offset + size > i_size_read(inode)) + return -EINVAL; filemap_invalidate_lock(inode->i_mapping); @@ -677,14 +768,13 @@ int kvm_gmem_bind(struct kvm *kvm, struct kvm_memory_slot *slot, if (!xa_empty(&f->bindings) && xa_find(&f->bindings, &start, end - 1, XA_PRESENT)) { r = -EEXIST; - filemap_invalidate_unlock(inode->i_mapping); - goto err; + goto out; } /* - * memslots of flag KVM_MEM_GUEST_MEMFD are immutable to change, so - * kvm_gmem_bind() must occur on a new memslot. Because the memslot - * is not visible yet, kvm_gmem_get_pfn() is guaranteed to see the file. + * memslots of flag KVM_MEM_GUEST_MEMFD are immutable, so bind must + * occur on a new memslot. Because the memslot is not visible yet, + * kvm_gmem_get_pfn() is guaranteed to see the file. */ WRITE_ONCE(slot->gmem.file, file); slot->gmem.pgoff = start; @@ -692,15 +782,51 @@ int kvm_gmem_bind(struct kvm *kvm, struct kvm_memory_slot *slot, slot->flags |= KVM_MEMSLOT_GMEM_ONLY; xa_store_range(&f->bindings, start, end - 1, slot, GFP_KERNEL); + r = 0; +out: filemap_invalidate_unlock(inode->i_mapping); + return r; +} + +int kvm_gmem_bind(struct kvm *kvm, struct kvm_memory_slot *slot, + unsigned int fd, uoff_t offset) +{ + const struct kvm_gmem_ops *ops; + struct file *file; + int r; + + BUILD_BUG_ON(sizeof(gpa_t) != sizeof(offset)); + BUILD_BUG_ON(sizeof(gfn_t) != sizeof(slot->gmem.pgoff)); + + file = fget(fd); + if (!file) + return -EBADF; + + ops = kvm_gmem_get_ops(file); + if (!ops || !ops->bind || !ops->get_pfn) { + r = -EINVAL; + goto out; + } + + if (!PAGE_ALIGNED(offset)) { + r = -EINVAL; + goto out; + } /* - * Drop the reference to the file, even on success. The file pins KVM, - * not the other way 'round. Active bindings are invalidated if the - * file is closed before memslots are destroyed. + * On success .bind() associates the fd with @kvm and (if the + * implementation needs it) records the reference that makes the fd + * pin the VM. A .bind() that returns success is trusted to have + * done so; there is no separate ownership check here. + */ + r = ops->bind(file, kvm, slot, offset); +out: + /* + * Drop the reference to the file even on success. The file pins + * KVM, not the other way 'round: it is the implementation's job to + * invalidate its bindings if the fd is closed before the memslots + * are destroyed. */ - r = 0; -err: fput(file); return r; } @@ -719,21 +845,15 @@ static void __kvm_gmem_unbind(struct kvm_memory_slot *slot, struct gmem_file *f) WRITE_ONCE(slot->gmem.file, NULL); } -void kvm_gmem_unbind(struct kvm_memory_slot *slot) +static void kvm_gmem_native_unbind(struct file *slot_file, struct kvm *kvm, + struct kvm_memory_slot *slot) { - /* - * Nothing to do if the underlying file was _already_ closed, as - * kvm_gmem_release() invalidates and nullifies all bindings. - */ - if (!slot->gmem.file) - return; - CLASS(gmem_get_file, file)(slot); /* - * However, if the file is _being_ closed, then the bindings need to be - * removed as kvm_gmem_release() might not run until after the memslot - * is freed. Note, modifying the bindings is safe even though the file + * If the file is _being_ closed, then the bindings need to be removed + * as kvm_gmem_release() might not run until after the memslot is + * freed. Note, modifying the bindings is safe even though the file * is dying as kvm_gmem_release() nullifies slot->gmem.file under * slots_lock, and only puts its reference to KVM after destroying all * bindings. I.e. reaching this point means kvm_gmem_release() hasn't @@ -741,15 +861,41 @@ void kvm_gmem_unbind(struct kvm_memory_slot *slot) * until the caller drops slots_lock. */ if (!file) { - __kvm_gmem_unbind(slot, slot->gmem.file->private_data); + __kvm_gmem_unbind(slot, gmem_file_of(slot_file)); return; } filemap_invalidate_lock(file->f_mapping); - __kvm_gmem_unbind(slot, file->private_data); + __kvm_gmem_unbind(slot, gmem_file_of(file)); filemap_invalidate_unlock(file->f_mapping); } +void kvm_gmem_unbind(struct kvm *kvm, struct kvm_memory_slot *slot) +{ + const struct kvm_gmem_ops *ops; + struct file *file; + + /* + * Nothing to do if the underlying file was _already_ closed, as + * kvm_gmem_release() invalidates and nullifies all bindings. + */ + file = READ_ONCE(slot->gmem.file); + if (!file) + return; + + ops = kvm_gmem_get_ops(file); + if (ops && ops->unbind) + ops->unbind(file, kvm, slot); + + /* + * .unbind() may clear slot->gmem.file itself (as the built-in gmem + * ops do, under filemap_invalidate_lock, to serialise against + * release/invalidate); catch the ones that don't. + */ + if (READ_ONCE(slot->gmem.file)) + WRITE_ONCE(slot->gmem.file, NULL); +} + /* Returns a locked folio on success. */ static struct folio *__kvm_gmem_get_pfn(struct file *file, struct kvm_memory_slot *slot, @@ -757,7 +903,7 @@ static struct folio *__kvm_gmem_get_pfn(struct file *file, int *max_order) { struct file *slot_file = READ_ONCE(slot->gmem.file); - struct gmem_file *f = file->private_data; + struct gmem_file *f = gmem_file_of(file); struct folio *folio; if (file != slot_file) { @@ -787,17 +933,14 @@ static struct folio *__kvm_gmem_get_pfn(struct file *file, return folio; } -int kvm_gmem_get_pfn(struct kvm *kvm, struct kvm_memory_slot *slot, - gfn_t gfn, kvm_pfn_t *pfn, struct page **page, - int *max_order) +static int kvm_gmem_native_get_pfn(struct file *file, struct kvm *kvm, + struct kvm_memory_slot *slot, gfn_t gfn, + kvm_pfn_t *pfn, struct page **page, + int *max_order) { pgoff_t index = kvm_gmem_get_index(slot, gfn); struct folio *folio; - int r = 0; - - CLASS(gmem_get_file, file)(slot); - if (!file) - return -EFAULT; + int r; folio = __kvm_gmem_get_pfn(file, slot, index, pfn, max_order); if (IS_ERR(folio)) @@ -809,7 +952,6 @@ int kvm_gmem_get_pfn(struct kvm *kvm, struct kvm_memory_slot *slot, } r = kvm_gmem_prepare_folio(kvm, slot, gfn, folio); - folio_unlock(folio); if (!r) @@ -819,6 +961,24 @@ int kvm_gmem_get_pfn(struct kvm *kvm, struct kvm_memory_slot *slot, return r; } + +int kvm_gmem_get_pfn(struct kvm *kvm, struct kvm_memory_slot *slot, + gfn_t gfn, kvm_pfn_t *pfn, struct page **page, + int *max_order) +{ + const struct kvm_gmem_ops *ops; + + CLASS(gmem_get_file, file)(slot); + if (!file) + return -EFAULT; + + ops = kvm_gmem_get_ops(file); + if (!ops || !ops->get_pfn) + return -EFAULT; + + *page = NULL; + return ops->get_pfn(file, kvm, slot, gfn, pfn, page, max_order); +} EXPORT_SYMBOL_FOR_KVM_INTERNAL(kvm_gmem_get_pfn); #ifdef CONFIG_HAVE_KVM_ARCH_GMEM_POPULATE @@ -860,11 +1020,51 @@ static long __kvm_gmem_populate(struct kvm *kvm, struct kvm_memory_slot *slot, return ret; } +static int kvm_gmem_populate_one(const struct kvm_gmem_ops *ops, + struct file *file, struct kvm *kvm, + struct kvm_memory_slot *slot, gfn_t gfn, + struct page *src_page, + kvm_gmem_populate_cb post_populate, + void *opaque) +{ + struct page *ignored_page = NULL; + kvm_pfn_t pfn; + int ret; + + /* + * Resolve the PFN via the ops. A populate() op may also copy + * @src_page into the backing before the arch post_populate() step + * (e.g. SNP LAUNCH_UPDATE); otherwise fall back to get_pfn(). + */ + if (ops->populate) + ret = ops->populate(file, kvm, slot, gfn, &pfn, + src_page, 0); + else + ret = ops->get_pfn(file, kvm, slot, gfn, &pfn, + &ignored_page, NULL); + if (ret) + return ret; + + ret = post_populate(kvm, gfn, pfn, src_page, opaque); + + /* + * Populate does not consume a refcount on the destination page. + * Drop @page (if any) AFTER post_populate() so the PFN cannot be + * freed and reallocated behind an arch hook (e.g. RMPUPDATE) that + * is still operating on it. + */ + if (ignored_page) + put_page(ignored_page); + + return ret; +} + long kvm_gmem_populate(struct kvm *kvm, gfn_t start_gfn, void __user *src, long npages, bool may_writeback_src, kvm_gmem_populate_cb post_populate, void *opaque) { struct kvm_memory_slot *slot; + const struct kvm_gmem_ops *ops; int ret = 0; long i; @@ -884,6 +1084,8 @@ long kvm_gmem_populate(struct kvm *kvm, gfn_t start_gfn, void __user *src, if (!file) return -EFAULT; + ops = kvm_gmem_get_ops(file); + npages = min_t(ulong, slot->npages - (start_gfn - slot->base_gfn), npages); for (i = 0; i < npages; i++) { struct page *src_page = NULL; @@ -906,8 +1108,13 @@ long kvm_gmem_populate(struct kvm *kvm, gfn_t start_gfn, void __user *src, } } - ret = __kvm_gmem_populate(kvm, slot, file, start_gfn + i, src_page, - post_populate, opaque); + if (ops && (ops->populate || ops->get_pfn != kvm_gmem_native_get_pfn)) + ret = kvm_gmem_populate_one(ops, file, kvm, slot, + start_gfn + i, src_page, + post_populate, opaque); + else + ret = __kvm_gmem_populate(kvm, slot, file, start_gfn + i, + src_page, post_populate, opaque); if (src_page) put_page(src_page); diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c index e44c20c04961..3af5270d2a90 100644 --- a/virt/kvm/kvm_main.c +++ b/virt/kvm/kvm_main.c @@ -935,7 +935,7 @@ static void kvm_destroy_dirty_bitmap(struct kvm_memory_slot *memslot) static void kvm_free_memslot(struct kvm *kvm, struct kvm_memory_slot *slot) { if (slot->flags & KVM_MEM_GUEST_MEMFD) - kvm_gmem_unbind(slot); + kvm_gmem_unbind(kvm, slot); kvm_destroy_dirty_bitmap(slot); @@ -1747,7 +1747,7 @@ static void kvm_commit_memory_region(struct kvm *kvm, * flags-only changes on guest_memfd slots should be impossible. */ if (WARN_ON_ONCE(old->flags & KVM_MEM_GUEST_MEMFD)) - kvm_gmem_unbind(old); + kvm_gmem_unbind(kvm, old); /* * The final quirk. Free the detached, old slot, but only its @@ -2118,7 +2118,7 @@ static int kvm_set_memory_region(struct kvm *kvm, out_unbind: if (mem->flags & KVM_MEM_GUEST_MEMFD) - kvm_gmem_unbind(new); + kvm_gmem_unbind(kvm, new); out: kfree(new); return r; @@ -5475,7 +5475,7 @@ bool file_is_kvm(struct file *file) { return file && file->f_op == &kvm_vm_fops; } -EXPORT_SYMBOL_FOR_KVM_INTERNAL(file_is_kvm); +EXPORT_SYMBOL_GPL(file_is_kvm); static int kvm_dev_ioctl_create_vm(unsigned long type) { diff --git a/virt/kvm/kvm_mm.h b/virt/kvm/kvm_mm.h index 7510ca915dd1..ea0fe74e89d4 100644 --- a/virt/kvm/kvm_mm.h +++ b/virt/kvm/kvm_mm.h @@ -76,7 +76,7 @@ void kvm_gmem_exit(void); int kvm_gmem_create(struct kvm *kvm, struct kvm_create_guest_memfd *args); int kvm_gmem_bind(struct kvm *kvm, struct kvm_memory_slot *slot, unsigned int fd, uoff_t offset); -void kvm_gmem_unbind(struct kvm_memory_slot *slot); +void kvm_gmem_unbind(struct kvm *kvm, struct kvm_memory_slot *slot); #else static inline int kvm_gmem_init(struct module *module) { @@ -91,7 +91,7 @@ static inline int kvm_gmem_bind(struct kvm *kvm, return -EIO; } -static inline void kvm_gmem_unbind(struct kvm_memory_slot *slot) +static inline void kvm_gmem_unbind(struct kvm *kvm, struct kvm_memory_slot *slot) { WARN_ON_ONCE(1); } -- 2.55.0 ^ permalink raw reply [flat|nested] 28+ messages in thread
* [RFC PATCH v2 05/11] iommufd: Look up private-interconnect phys via exporter symbols 2026-07-20 11:03 ` [RFC PATCH v2 01/11] KVM: selftests: sev_smoke_test: Only run VM types the host offers David Woodhouse ` (2 preceding siblings ...) 2026-07-20 11:03 ` [RFC PATCH v2 04/11] KVM: guest_memfd: Introduce guest memory ops and route native gmem through them David Woodhouse @ 2026-07-20 11:03 ` David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 06/11] iommufd: Plumb dma-buf memory-type (RAM vs MMIO) through the phys map David Woodhouse ` (5 subsequent siblings) 9 siblings, 0 replies; 28+ messages in thread From: David Woodhouse @ 2026-07-20 11:03 UTC (permalink / raw) To: Sean Christopherson, Paolo Bonzini, kvm Cc: linux-kernel, iommu, linux-kselftest, linux-media, dri-devel, linaro-mm-sig, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, x86, Jason Gunthorpe, Kevin Tian, Joerg Roedel, Will Deacon, Robin Murphy, Shuah Khan, Sumit Semwal, Christian König, Tom Lendacky, Tycho Andersen, Kees Cook, Ashish Kalra, Connor Williamson, Andrew Morton, Zhou Yuhang, Zi Li, Dan Williams, Eric Biggers, Colin Ian King, connordw, graf, fgriffo, David Woodhouse From: David Woodhouse <dwmw@amazon.co.uk> The dma-buf phys-map dispatch used by iommufd's IOMMU_IOAS_MAP_FILE path currently knows only about the VFIO PCI exporter (looked up via symbol_get(vfio_pci_dma_buf_iommufd_map)). Widen the dispatch so additional exporters can advertise the same well-known symbol name convention and have their dma-bufs mapped without an explicit iommufd dependency in each exporter's module. For now the dispatch tries the sample guest_memfd provider's exporter first, then falls back to VFIO PCI, both via symbol_get. This mirrors the pattern the in-tree comment already flags: "iommufd and vfio have a circular dependency. Future work for a phys based private interconnect will remove this." The eventual formal negotiated dma_buf_ops op that returns phys plus a memory-type descriptor will subsume both of these lookups. The next patch plumbs the memory type through iommufd; a follow-up will convert the symbol lookups to the negotiated op. Based on a proof-of-concept by Connor Williamson <connordw@amazon.com> Signed-off-by: David Woodhouse (Kiro) <dwmw@amazon.co.uk> --- drivers/iommu/iommufd/pages.c | 20 ++++++++++++++++++++ 1 file changed, 20 insertions(+) diff --git a/drivers/iommu/iommufd/pages.c b/drivers/iommu/iommufd/pages.c index 03c8379bbc34..2d4ea41460fd 100644 --- a/drivers/iommu/iommufd/pages.c +++ b/drivers/iommu/iommufd/pages.c @@ -1470,6 +1470,26 @@ sym_vfio_pci_dma_buf_iommufd_map(struct dma_buf_attachment *attachment, if (rc != -EOPNOTSUPP) return rc; + /* + * Prototype: try the sample gmem provider's dma-buf exporter. This + * mirrors the vfio-pci private-interconnect hook, and (like it) is + * meant to be replaced by a formal negotiated exporter op returning + * phys for iommufd. + */ + { + extern int gmem_provider_dma_buf_iommufd_map( + struct dma_buf_attachment *, struct phys_vec *); + typeof(&gmem_provider_dma_buf_iommufd_map) gfn; + + gfn = symbol_get(gmem_provider_dma_buf_iommufd_map); + if (gfn) { + rc = gfn(attachment, phys); + symbol_put(gmem_provider_dma_buf_iommufd_map); + if (rc != -EOPNOTSUPP) + return rc; + } + } + if (!IS_ENABLED(CONFIG_VFIO_PCI_DMABUF)) return -EOPNOTSUPP; -- 2.55.0 ^ permalink raw reply [flat|nested] 28+ messages in thread
* [RFC PATCH v2 06/11] iommufd: Plumb dma-buf memory-type (RAM vs MMIO) through the phys map 2026-07-20 11:03 ` [RFC PATCH v2 01/11] KVM: selftests: sev_smoke_test: Only run VM types the host offers David Woodhouse ` (3 preceding siblings ...) 2026-07-20 11:03 ` [RFC PATCH v2 05/11] iommufd: Look up private-interconnect phys via exporter symbols David Woodhouse @ 2026-07-20 11:03 ` David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 07/11] KVM: guest_memfd: Add ops-driven page revocation David Woodhouse ` (4 subsequent siblings) 9 siblings, 0 replies; 28+ messages in thread From: David Woodhouse @ 2026-07-20 11:03 UTC (permalink / raw) To: Sean Christopherson, Paolo Bonzini, kvm Cc: linux-kernel, iommu, linux-kselftest, linux-media, dri-devel, linaro-mm-sig, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, x86, Jason Gunthorpe, Kevin Tian, Joerg Roedel, Will Deacon, Robin Murphy, Shuah Khan, Sumit Semwal, Christian König, Tom Lendacky, Tycho Andersen, Kees Cook, Ashish Kalra, Connor Williamson, Andrew Morton, Zhou Yuhang, Zi Li, Dan Williams, Eric Biggers, Colin Ian King, connordw, graf, fgriffo, David Woodhouse From: David Woodhouse <dwmw@amazon.co.uk> iommufd's dma-buf import path programmed the IOMMU unconditionally with BATCH_MMIO (IOMMU_MMIO set, IOMMU_CACHE cleared). This is correct for the VFIO PCI dma-buf exporter -- its phys is PCI BAR (MMIO) memory -- but it is the wrong prot for a dma-buf whose phys is RAM: on AMD-Vi the device DMA silently misroutes (Identify Controller against a gmem-provider-backed target completes with status=0 but the target region is not written). Carry the exporter's memory type through to the fill routine: - Add is_cpu_ram to iopt_pages_dmabuf and pfn_reader_dmabuf. - Make sym_vfio_pci_dma_buf_iommufd_map() also fill *is_cpu_ram. VFIO PCI's exporter is MMIO (false); the sample gmem provider's exporter is CPU RAM (true). - pfn_reader_fill_dmabuf() picks BATCH_CPU_MEMORY vs BATCH_MMIO based on the captured kind. The eventual formal negotiated exporter op should return phys and memory type together in one call, replacing both symbol_get() lookups. Signed-off-by: David Woodhouse (Kiro) <dwmw@amazon.co.uk> --- drivers/iommu/iommufd/io_pagetable.h | 7 ++++++ drivers/iommu/iommufd/pages.c | 33 +++++++++++++++++++++++----- 2 files changed, 34 insertions(+), 6 deletions(-) diff --git a/drivers/iommu/iommufd/io_pagetable.h b/drivers/iommu/iommufd/io_pagetable.h index 27e3e311d395..887ed94474c7 100644 --- a/drivers/iommu/iommufd/io_pagetable.h +++ b/drivers/iommu/iommufd/io_pagetable.h @@ -206,6 +206,13 @@ struct iopt_pages_dmabuf { /* Always PAGE_SIZE aligned */ unsigned long start; struct list_head tracker; + /* + * true if the exporter's phys is CPU RAM (map with IOMMU_CACHE, no + * IOMMU_MMIO); false for MMIO/BAR memory (map with IOMMU_MMIO). Set + * by the exporter dispatch at map time; consulted by + * pfn_reader_fill_dmabuf() to pick BATCH_CPU_MEMORY vs BATCH_MMIO. + */ + bool is_cpu_ram; }; /* diff --git a/drivers/iommu/iommufd/pages.c b/drivers/iommu/iommufd/pages.c index 2d4ea41460fd..f9b2ae6d7e96 100644 --- a/drivers/iommu/iommufd/pages.c +++ b/drivers/iommu/iommufd/pages.c @@ -1080,6 +1080,7 @@ static int pfn_reader_user_update_pinned(struct pfn_reader_user *user, struct pfn_reader_dmabuf { struct phys_vec phys; unsigned long start_offset; + bool is_cpu_ram; }; static int pfn_reader_dmabuf_init(struct pfn_reader_dmabuf *dmabuf, @@ -1091,6 +1092,7 @@ static int pfn_reader_dmabuf_init(struct pfn_reader_dmabuf *dmabuf, dmabuf->phys = pages->dmabuf.phys; dmabuf->start_offset = pages->dmabuf.start; + dmabuf->is_cpu_ram = pages->dmabuf.is_cpu_ram; return 0; } @@ -1106,9 +1108,15 @@ static int pfn_reader_fill_dmabuf(struct pfn_reader_dmabuf *dmabuf, * always filled using page size aligned PFNs just like the other types. * If the dmabuf has been sliced on a sub page offset then the common * batch to domain code will adjust it before mapping to the domain. + * + * The exporter's memory type (CPU RAM vs MMIO/BAR) selects the batch + * kind so downstream iommu_map sets IOMMU_CACHE for cache-coherent RAM + * or IOMMU_MMIO for BAR memory. The kind was captured at map time by + * the exporter dispatch. */ batch_add_pfn_num(batch, PHYS_PFN(dmabuf->phys.paddr + start), - last_index - start_index + 1, BATCH_MMIO); + last_index - start_index + 1, + dmabuf->is_cpu_ram ? BATCH_CPU_MEMORY : BATCH_MMIO); return 0; } @@ -1459,22 +1467,31 @@ static const struct dma_buf_attach_ops iopt_dmabuf_attach_revoke_ops = { * iommufd and vfio have a circular dependency. Future work for a phys * based private interconnect will remove this. */ +/* + * Look up the exporter's phys accessor for iommufd's private-interconnect + * path. Also fills *is_cpu_ram: true if the exporter's memory is normal + * cache-coherent RAM (needs BATCH_CPU_MEMORY / IOMMU_CACHE), false for MMIO + * (needs BATCH_MMIO / IOMMU_MMIO). This will be replaced by a formal + * exporter op that returns phys + memory type together. + */ static int sym_vfio_pci_dma_buf_iommufd_map(struct dma_buf_attachment *attachment, - struct phys_vec *phys) + struct phys_vec *phys, bool *is_cpu_ram) { typeof(&vfio_pci_dma_buf_iommufd_map) fn; int rc; rc = iommufd_test_dma_buf_iommufd_map(attachment, phys); - if (rc != -EOPNOTSUPP) + if (rc != -EOPNOTSUPP) { + *is_cpu_ram = false; /* test hook mimics VFIO MMIO */ return rc; + } /* * Prototype: try the sample gmem provider's dma-buf exporter. This * mirrors the vfio-pci private-interconnect hook, and (like it) is * meant to be replaced by a formal negotiated exporter op returning - * phys for iommufd. + * phys + memory type. The provider serves RAM, so mark it CPU_RAM. */ { extern int gmem_provider_dma_buf_iommufd_map( @@ -1485,8 +1502,10 @@ sym_vfio_pci_dma_buf_iommufd_map(struct dma_buf_attachment *attachment, if (gfn) { rc = gfn(attachment, phys); symbol_put(gmem_provider_dma_buf_iommufd_map); - if (rc != -EOPNOTSUPP) + if (rc != -EOPNOTSUPP) { + *is_cpu_ram = true; return rc; + } } } @@ -1498,6 +1517,7 @@ sym_vfio_pci_dma_buf_iommufd_map(struct dma_buf_attachment *attachment, return -EOPNOTSUPP; rc = fn(attachment, phys); symbol_put(vfio_pci_dma_buf_iommufd_map); + *is_cpu_ram = false; /* VFIO PCI dma-buf carries BAR (MMIO) memory */ return rc; } @@ -1526,7 +1546,8 @@ static int iopt_map_dmabuf(struct iommufd_ctx *ictx, struct iopt_pages *pages, if (rc) goto err_detach; - rc = sym_vfio_pci_dma_buf_iommufd_map(attach, &pages->dmabuf.phys); + rc = sym_vfio_pci_dma_buf_iommufd_map(attach, &pages->dmabuf.phys, + &pages->dmabuf.is_cpu_ram); if (rc) goto err_unpin; -- 2.55.0 ^ permalink raw reply [flat|nested] 28+ messages in thread
* [RFC PATCH v2 07/11] KVM: guest_memfd: Add ops-driven page revocation 2026-07-20 11:03 ` [RFC PATCH v2 01/11] KVM: selftests: sev_smoke_test: Only run VM types the host offers David Woodhouse ` (4 preceding siblings ...) 2026-07-20 11:03 ` [RFC PATCH v2 06/11] iommufd: Plumb dma-buf memory-type (RAM vs MMIO) through the phys map David Woodhouse @ 2026-07-20 11:03 ` David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 08/11] samples/kvm: Add guest_memfd backing sample David Woodhouse ` (3 subsequent siblings) 9 siblings, 0 replies; 28+ messages in thread From: David Woodhouse @ 2026-07-20 11:03 UTC (permalink / raw) To: Sean Christopherson, Paolo Bonzini, kvm Cc: linux-kernel, iommu, linux-kselftest, linux-media, dri-devel, linaro-mm-sig, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, x86, Jason Gunthorpe, Kevin Tian, Joerg Roedel, Will Deacon, Robin Murphy, Shuah Khan, Sumit Semwal, Christian König, Tom Lendacky, Tycho Andersen, Kees Cook, Ashish Kalra, Connor Williamson, Andrew Morton, Zhou Yuhang, Zi Li, Dan Williams, Eric Biggers, Colin Ian King, connordw, graf, fgriffo, David Woodhouse From: David Woodhouse <dwmw@amazon.co.uk> Add kvm_gmem_invalidate_range(): a guest_memfd backing calls it to zap the guest secondary-MMU mapping for a gfn range so the next access re-faults through ops->get_pfn(). Mirrors the existing internal __kvm_gmem_invalidate_start() for a single bound slot, bracketing the unmap with the mmu_invalidate window so a racing vCPU fault retries rather than installing a stale mapping. The existing ops->invalidate callback is KVM -> implementation (cleanup); this new symbol is the implementation -> KVM path, useful for an overcommit manager or any subsystem that reclaims pages from a running guest. Export it for module use and add it to the KVM external-export allowlist. Signed-off-by: David Woodhouse (Kiro) <dwmw@amazon.co.uk> --- include/linux/kvm_host.h | 9 +++++++ virt/kvm/guest_memfd.c | 52 ++++++++++++++++++++++++++++++++++++++++ 2 files changed, 61 insertions(+) diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h index b33d2f475042..04fa0cb126f6 100644 --- a/include/linux/kvm_host.h +++ b/include/linux/kvm_host.h @@ -701,6 +701,15 @@ static inline bool is_kvm_gmem_file(struct file *file) return file && file->f_op == &kvm_gmem_fops; } +/* + * Outbound (implementation -> KVM) revocation: zap the guest's secondary MMU + * (NPT/EPT) mapping for a gfn range so the next guest access re-faults + * through the ops' get_pfn(). This is the reverse of ->invalidate (which is + * KVM -> implementation) and is what a backing that manages its own memory + * (e.g. an overcommit manager) uses to reclaim a page from a running guest. + */ +void kvm_gmem_invalidate_range(struct kvm *kvm, gfn_t start, gfn_t end); + #endif /* CONFIG_KVM_GUEST_MEMFD */ static inline bool kvm_slot_dirty_track_enabled(const struct kvm_memory_slot *slot) diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c index 057f90eb835a..1da22da0a6ec 100644 --- a/virt/kvm/guest_memfd.c +++ b/virt/kvm/guest_memfd.c @@ -738,6 +738,58 @@ int kvm_gmem_create(struct kvm *kvm, struct kvm_create_guest_memfd *args) return __kvm_gmem_create(kvm, size, flags); } +/* + * Outbound (implementation -> KVM) revocation: an implementation zaps the + * guest's secondary MMU mapping for a gfn range so the next guest access + * re-faults through ops->get_pfn(). Mirrors __kvm_gmem_invalidate_start() + * for a single bound slot, bracketing the unmap with the mmu_invalidate + * window so a racing vCPU fault retries rather than installing a stale + * mapping. + */ +void kvm_gmem_invalidate_range(struct kvm *kvm, gfn_t start, gfn_t end) +{ + struct kvm_memory_slot *slot; + struct kvm_memslot_iter iter; + bool flush = false; + int as_id; + int idx; + + idx = srcu_read_lock(&kvm->srcu); + + /* + * Walk every address space (SMM has its own on x86) and every + * intersecting memslot. A single invalidation may span several + * slots and holes: skip the holes, clamp to each slot's range. + */ + KVM_MMU_LOCK(kvm); + kvm_mmu_invalidate_start(kvm); + + for (as_id = 0; as_id < kvm_arch_nr_memslot_as_ids(kvm); as_id++) { + struct kvm_memslots *slots = __kvm_memslots(kvm, as_id); + + kvm_for_each_memslot_in_gfn_range(&iter, slots, start, end) { + struct kvm_gfn_range gfn_range; + + slot = iter.slot; + gfn_range.slot = slot; + gfn_range.start = max(start, slot->base_gfn); + gfn_range.end = min(end, slot->base_gfn + slot->npages); + gfn_range.may_block = true; + gfn_range.attr_filter = KVM_FILTER_SHARED | KVM_FILTER_PRIVATE; + + flush |= kvm_mmu_unmap_gfn_range(kvm, &gfn_range); + } + } + if (flush) + kvm_flush_remote_tlbs(kvm); + + kvm_mmu_invalidate_end(kvm); + KVM_MMU_UNLOCK(kvm); + + srcu_read_unlock(&kvm->srcu, idx); +} +EXPORT_SYMBOL_GPL(kvm_gmem_invalidate_range); + static const struct kvm_gmem_ops *kvm_gmem_get_ops(struct file *file) { if (!is_kvm_gmem_file(file)) -- 2.55.0 ^ permalink raw reply [flat|nested] 28+ messages in thread
* [RFC PATCH v2 08/11] samples/kvm: Add guest_memfd backing sample 2026-07-20 11:03 ` [RFC PATCH v2 01/11] KVM: selftests: sev_smoke_test: Only run VM types the host offers David Woodhouse ` (5 preceding siblings ...) 2026-07-20 11:03 ` [RFC PATCH v2 07/11] KVM: guest_memfd: Add ops-driven page revocation David Woodhouse @ 2026-07-20 11:03 ` David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 09/11] selftests/kvm: gmem_provider KVM-only tests David Woodhouse ` (2 subsequent siblings) 9 siblings, 0 replies; 28+ messages in thread From: David Woodhouse @ 2026-07-20 11:03 UTC (permalink / raw) To: Sean Christopherson, Paolo Bonzini, kvm Cc: linux-kernel, iommu, linux-kselftest, linux-media, dri-devel, linaro-mm-sig, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, x86, Jason Gunthorpe, Kevin Tian, Joerg Roedel, Will Deacon, Robin Murphy, Shuah Khan, Sumit Semwal, Christian König, Tom Lendacky, Tycho Andersen, Kees Cook, Ashish Kalra, Connor Williamson, Andrew Morton, Zhou Yuhang, Zi Li, Dan Williams, Eric Biggers, Colin Ian King, connordw, graf, fgriffo, David Woodhouse From: David Woodhouse <dwmw@amazon.co.uk> A minimal in-tree implementation of the guest_memfd ops interface, useful both as reference code and as the substrate for the selftests that follow. Features: - Backs a slot from a fixed page-less physical range specified by the module parameters addr= and len=, or falls back to a CMA-allocated contiguous region when neither is given. The addr=/len= mode lets a reserved DRAM carve-out (mem=... or an explicit e820/DT reservation) become KVM slot backing without any struct page. - Per-fd mmap capability via GMEM_PROVIDER_FLAG_MMAP_CAPABLE at GMEM_PROVIDER_SETUP time. An mmap-capable fd is bound as a gmem-only slot (KVM_MEMSLOT_GMEM_ONLY) so the host mmap and the guest fault path both route through this implementation. The mmap-capable fd is refused for hardware-encrypted VM types (SEV, SEV-ES, SEV-SNP, TDX) where a host view onto private-marked pages can only fault; the sample enforces this at bind time. KVM_X86_SW_PROTECTED_VM has no hardware encryption and is allowed. - Page revocation via GMEM_PROVIDER_SET_PRESENT. Marking a range absent tracks it in a bitmap, invokes kvm_gmem_invalidate_range() to zap the guest NPT/EPT, and fans dma_buf_invalidate_mappings() out to any exported dma-buf so iommufd tears down its mapping. Subsequent get_pfn() on the range returns -EFAULT until it is restored with present=1. Models overcommit-style reclaim without needing the platform 's real reclaim machinery. - A dma-buf exporter (GMEM_PROVIDER_GET_DMABUF) that yields an fd suitable for IOMMU_IOAS_MAP_FILE. The exporter mirrors the shape of drivers/vfio/pci/vfio_pci_dmabuf.c: dma_buf_attach_revocable, ->map_dma_buf via dma_buf_phys_vec_to_sgt, a single-range phys_vec, and the "well-known symbol" hook for iommufd's private-interconnect phys-map lookup. The sample is layered on top of the shared kvm_gmem_fops: its fd is created with those fops and file->private_data points at a kvm_gmem_backing embedded at the head of its per-instance state. Release / mmap / ioctl for the fd all run through the ops table, so the sample owns no file_operations of its own. Exposing kvm_gmem_fops for module use required promoting its export from EXPORT_SYMBOL_FOR_KVM_INTERNAL to EXPORT_SYMBOL_GPL and adding it to the KVM external-export allowlist. SEV-SNP bind takes a psmash + rmp_make_shared sweep over the range so a new SNP VM can transition its pages private and re-encrypt them -- harmless on non-SNP hosts, where the calls return -ENODEV. Based on a proof-of-concept by Connor Williamson <connordw@amazon.com>. Signed-off-by: David Woodhouse (Kiro) <dwmw@amazon.co.uk> --- samples/Kconfig | 18 + samples/Makefile | 1 + samples/kvm/Makefile | 2 + samples/kvm/gmem_provider.c | 772 ++++++++++++++++++++++++++++++++++++ samples/kvm/gmem_provider.h | 53 +++ virt/kvm/guest_memfd.c | 2 +- 6 files changed, 847 insertions(+), 1 deletion(-) create mode 100644 samples/kvm/Makefile create mode 100644 samples/kvm/gmem_provider.c create mode 100644 samples/kvm/gmem_provider.h diff --git a/samples/Kconfig b/samples/Kconfig index a75e8e78330d..3482659638b5 100644 --- a/samples/Kconfig +++ b/samples/Kconfig @@ -326,6 +326,24 @@ source "samples/rust/Kconfig" source "samples/damon/Kconfig" +config SAMPLE_KVM_GMEM_PROVIDER + tristate "Build sample guest_memfd provider -- loadable module only" + depends on KVM_GUEST_MEMFD && X86_64 && m + help + This builds a sample external guest_memfd provider module with two + backing modes. Loaded with addr=/len= it backs guest memory with a + fixed physical range that has no struct page -- for example memory + carved out of the kernel with mem= on the command line. Without + those params it falls back to a physically contiguous, page-backed + region from alloc_contig_pages(), sized by the setup ioctl (easier to + run, no boot param). + + It demonstrates the guest_memfd provider ABI including 2M and 1G + guest mappings and, on SEV-SNP hosts, the RMP reset that lets a + page-less range be re-bound to a new VM across a live update. + + If unsure, say N. + endif # SAMPLES config HAVE_SAMPLE_FTRACE_DIRECT diff --git a/samples/Makefile b/samples/Makefile index 07641e177bd8..e85397e5e34f 100644 --- a/samples/Makefile +++ b/samples/Makefile @@ -37,6 +37,7 @@ obj-$(CONFIG_SAMPLE_TPS6594_PFSM) += pfsm/ subdir-$(CONFIG_SAMPLE_WATCHDOG) += watchdog subdir-$(CONFIG_SAMPLE_WATCH_QUEUE) += watch_queue obj-$(CONFIG_SAMPLE_KMEMLEAK) += kmemleak/ +obj-$(CONFIG_SAMPLE_KVM_GMEM_PROVIDER) += kvm/ obj-$(CONFIG_SAMPLE_CORESIGHT_SYSCFG) += coresight/ obj-$(CONFIG_SAMPLE_FPROBE) += fprobe/ obj-$(CONFIG_SAMPLES_RUST) += rust/ diff --git a/samples/kvm/Makefile b/samples/kvm/Makefile new file mode 100644 index 000000000000..dcad6e53ea78 --- /dev/null +++ b/samples/kvm/Makefile @@ -0,0 +1,2 @@ +# SPDX-License-Identifier: GPL-2.0 +obj-$(CONFIG_SAMPLE_KVM_GMEM_PROVIDER) += gmem_provider.o diff --git a/samples/kvm/gmem_provider.c b/samples/kvm/gmem_provider.c new file mode 100644 index 000000000000..9728f5a8029b --- /dev/null +++ b/samples/kvm/gmem_provider.c @@ -0,0 +1,772 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * gmem_provider - sample external guest_memfd provider. + * + * Demonstrates the KVM guest_memfd provider ABI with two backing modes: + * + * - External (page-less): loaded with addr=/len=, backs guest memory with a + * fixed physical range that has no struct page -- e.g. memory carved out of + * the kernel with mem= on the command line. This is the case the provider + * ABI exists for; get_pfn() returns bare PFNs KVM treats as non-refcounted. + * + * - CMA fallback (page-backed): when addr=/len= are not given, allocates a + * physically contiguous region via alloc_contig_pages() sized by the setup + * ioctl. Easier to run (no mem= boot param), and still exercises 2M/1G + * mappings. Not preserved across kexec/live update. + * + * On SEV-SNP hosts, bind() resets the range's RMP entries to 4K shared so a + * (new) SNP VM can re-encrypt it, which is what allows re-binding the range to + * a fresh VM across a live update. + * + * Usage: + * # page-less external range: + * insmod gmem_provider.ko addr=0x5D40000000 len=0x1000000 + * # or CMA fallback (no params); size comes from the ioctl + * insmod gmem_provider.ko + * fd = open("/dev/gmem_provider"); ioctl(fd, GMEM_PROVIDER_SETUP, {kvm_fd, size}); + * pass the returned fd + KVM_MEM_GUEST_MEMFD to KVM_SET_USER_MEMORY_REGION2. + */ + +#include <linux/anon_inodes.h> +#include <linux/bitmap.h> +#include <linux/dma-buf.h> +#include <linux/dma-buf-mapping.h> +#include <linux/dma-resv.h> +#include <linux/file.h> +#include <linux/fs.h> +#include <linux/gfp.h> +#include <linux/highmem.h> +#include <linux/io.h> +#include <linux/kvm_host.h> +#include <linux/miscdevice.h> +#include <linux/mm.h> +#include <linux/module.h> +#include <linux/pgtable.h> +#include <linux/slab.h> + +#if IS_ENABLED(CONFIG_AMD_MEM_ENCRYPT) +#include <asm/sev.h> +#endif + +#include "gmem_provider.h" + +MODULE_IMPORT_NS("DMA_BUF"); + +static unsigned long long addr; +module_param(addr, ullong, 0444); +MODULE_PARM_DESC(addr, "physical base of an external page-less backing region (optional)"); + +static unsigned long long len; +module_param(len, ullong, 0444); +MODULE_PARM_DESC(len, "size in bytes of the external backing region (optional)"); + + +struct gmem_info { + /* + * MUST be first: file->private_data points here. is_kvm_gmem_file() + * on the KVM side proves the reinterpretation is safe. + */ + struct kvm_gmem_backing backing; + + /* + * Protects everything below (except the immutable base_pfn/npages/ + * cma_pages fields set at setup). Ordering: info->lock is a leaf; + * do not acquire other locks under it. KVM's slots_lock is already + * held on the .bind/.unbind paths so info->lock is only needed to + * serialise those against ioctl(SET_PRESENT). + */ + struct mutex lock; + + /* Kept locally: kvm_gmem_ops does not carry a kvm pointer. */ + struct kvm *kvm; + + /* + * Single active memslot binding. Used at release time to zap any + * still-active guest mappings before the memory disappears. NULL + * when no memslot is bound (or the last one has been unbound). + */ + struct kvm_memory_slot *bound_slot; + + bool mmap_capable; /* -> KVM_MEMSLOT_GMEM_ONLY at bind */ + + unsigned long base_pfn; + unsigned long npages; + struct page *cma_pages; /* non-NULL if CMA-allocated */ + gfn_t base_gfn; /* recorded at bind, for revoke */ + pgoff_t pgoff; /* provider offset (pages) of the slot */ + unsigned long *absent; /* bitmap of currently-revoked pages */ + struct list_head dmabufs; /* struct gmem_dmabuf entries */ + struct mutex dmabufs_lock; +}; + +static struct gmem_info *to_gmem_info(struct file *file) +{ + return container_of(file->private_data, struct gmem_info, backing); +} + +/* Map a backing PFN for CPU access: page-backed via kmap, page-less via memremap. */ +static void *gmem_map_pfn(kvm_pfn_t pfn) +{ + if (pfn_valid(pfn)) + return kmap_local_pfn(pfn); + return memremap(PFN_PHYS(pfn), PAGE_SIZE, MEMREMAP_WB); +} + +static void gmem_unmap_pfn(kvm_pfn_t pfn, void *vaddr) +{ + if (!vaddr) + return; + if (pfn_valid(pfn)) + kunmap_local(vaddr); + else + memunmap(vaddr); +} + +/* Largest order KVM may map at @gfn, snapped to 4K/2M/1G. */ +static int gmem_max_order(struct gmem_info *info, gfn_t gfn, unsigned long index) +{ + unsigned long pfn = info->base_pfn + index; + unsigned long remaining = info->npages - index; + unsigned int pud_order = PUD_SHIFT - PAGE_SHIFT; + unsigned int pmd_order = PMD_SHIFT - PAGE_SHIFT; + unsigned long absent_next; + + /* + * A hugepage may not span any revoked page. Clamp by the distance to + * the next absent bit; scanning is cheap because @absent is a plain + * bitmap and the caller already checked test_bit(index). + */ + if (info->absent) { + absent_next = find_next_bit(info->absent, info->npages, + index + 1); + remaining = min(remaining, absent_next - index); + } + + if (IS_ALIGNED(pfn, 1UL << pud_order) && + IS_ALIGNED(gfn, 1UL << pud_order) && + remaining >= (1UL << pud_order)) + return pud_order; + + if (IS_ALIGNED(pfn, 1UL << pmd_order) && + IS_ALIGNED(gfn, 1UL << pmd_order) && + remaining >= (1UL << pmd_order)) + return pmd_order; + + return 0; +} + +static int gmem_get_pfn(struct file *file, struct kvm *kvm, + struct kvm_memory_slot *slot, gfn_t gfn, + kvm_pfn_t *pfn, struct page **page, int *max_order) +{ + struct gmem_info *info = to_gmem_info(file); + pgoff_t index = gfn - slot->base_gfn + slot->gmem.pgoff; + + if (index >= info->npages) + return -EINVAL; + + /* Revoked (absent) page: behave like not-present so the fault fails. */ + if (info->absent && test_bit(index, info->absent)) + return -EFAULT; + + *pfn = info->base_pfn + index; + if (max_order) + *max_order = gmem_max_order(info, gfn, index); + return 0; +} + +static int gmem_populate(struct file *file, struct kvm *kvm, + struct kvm_memory_slot *slot, gfn_t gfn, + kvm_pfn_t *pfn, struct page *src_page, int order) +{ + struct gmem_info *info = to_gmem_info(file); + pgoff_t index = gfn - slot->base_gfn + slot->gmem.pgoff; + + if (index >= info->npages) + return -EINVAL; + + *pfn = info->base_pfn + index; + + if (src_page) { + void *dst, *src; + + /* Map dst first: memremap() may sleep, kmap_local_page() must not. */ + dst = gmem_map_pfn(*pfn); + if (!dst) + return -ENOMEM; + src = kmap_local_page(src_page); + memcpy(dst, src, PAGE_SIZE); + kunmap_local(src); + gmem_unmap_pfn(*pfn, dst); + } + return 0; +} + +static int gmem_bind(struct file *file, struct kvm *kvm, + struct kvm_memory_slot *slot, loff_t offset) +{ + struct gmem_info *info = to_gmem_info(file); + struct kvm *old_kvm = NULL; + unsigned long start = offset >> PAGE_SHIFT; + + if (offset < 0 || !PAGE_ALIGNED(offset) || + start + slot->npages > info->npages) + return -EINVAL; + + /* + * An mmap-capable backing hands the VMM a host mapping onto pages + * that a hardware-encrypted VM (SEV, SEV-ES, SEV-SNP, TDX) will mark + * private in the RMP/EPT; the mmap can only ever fault on those + * pages. Refuse rather than hand the VMM a useless (and misleading) + * shared view. SW_PROTECTED_VM has no hardware encryption and is + * fine. + */ +#ifdef CONFIG_X86 + if (info->mmap_capable && + (kvm->arch.vm_type == KVM_X86_SEV_VM || + kvm->arch.vm_type == KVM_X86_SEV_ES_VM || + kvm->arch.vm_type == KVM_X86_SNP_VM || + kvm->arch.vm_type == KVM_X86_TDX_VM)) + return -EACCES; +#endif + + /* Record the binding so the revoke ioctl can translate offset -> gfn. */ + mutex_lock(&info->lock); + info->base_gfn = slot->base_gfn; + info->pgoff = start; + info->bound_slot = slot; + mutex_unlock(&info->lock); + +#if IS_ENABLED(CONFIG_AMD_MEM_ENCRYPT) + /* + * Reset the RMP for the range to 4K shared so a (new) SEV-SNP VM can + * transition it to private and re-encrypt it. PSMASH any 2M entries + * first. Harmless on non-SNP hosts, where these return -ENODEV. + */ + { + unsigned long i; + unsigned long first_pmd_pfn = ALIGN(info->base_pfn + start, + PTRS_PER_PMD); + + for (i = first_pmd_pfn - (info->base_pfn + start); + i + PTRS_PER_PMD <= slot->npages; + i += PTRS_PER_PMD) + psmash(info->base_pfn + start + i); + + for (i = 0; i < slot->npages; i++) { + unsigned long pfn = info->base_pfn + start + i; + int ret = rmp_make_shared(pfn, PG_LEVEL_4K); + + if (ret && ret != -ENODEV) + pr_info_once("gmem_provider: rmp_make_shared(0x%lx) = %d\n", + pfn, ret); + } + } +#endif + + /* + * Claim (or, on re-bind, transfer) VM ownership; pins the VM. + * kvm_put_kvm(old_kvm) MUST run outside info->lock: if the put + * drops the last ref, kvm_destroy_vm() runs inline and calls back + * into our gmem_unbind() (via kvm_gmem_unbind() on each memslot), + * which needs info->lock -- taking it here would self-deadlock. + */ + mutex_lock(&info->lock); + if (info->kvm != kvm) { + old_kvm = info->kvm; + kvm_get_kvm(kvm); + info->kvm = kvm; + } + mutex_unlock(&info->lock); + if (old_kvm) + kvm_put_kvm(old_kvm); + + /* + * Record the file on the slot (KVM's outer bind no longer does this + * for us) and mark the slot gmem-only if this backing serves host + * accesses through its own mmap. + */ + WRITE_ONCE(slot->gmem.file, file); + slot->gmem.pgoff = start; + if (info->mmap_capable) + slot->flags |= KVM_MEMSLOT_GMEM_ONLY; + + return 0; +} + +static void gmem_unbind(struct file *file, struct kvm *kvm, + struct kvm_memory_slot *slot) +{ + struct gmem_info *info = to_gmem_info(file); + + mutex_lock(&info->lock); + if (info->bound_slot == slot) + info->bound_slot = NULL; + mutex_unlock(&info->lock); + +#if IS_ENABLED(CONFIG_AMD_MEM_ENCRYPT) + { + unsigned long start = slot->gmem.pgoff; + unsigned long first_pmd_pfn = ALIGN(info->base_pfn + start, + PTRS_PER_PMD); + unsigned long i; + + /* + * Symmetric with gmem_bind(): PSMASH any 2M RMP entries first + * so rmp_make_shared(PG_LEVEL_4K) can succeed on SNP hosts. + * Harmless on non-SNP: psmash() returns -ENODEV. + */ + for (i = first_pmd_pfn - (info->base_pfn + start); + i + PTRS_PER_PMD <= slot->npages; + i += PTRS_PER_PMD) + psmash(info->base_pfn + start + i); + + for (i = 0; i < slot->npages; i++) { + unsigned long pfn = info->base_pfn + start + i; + + rmp_make_shared(pfn, PG_LEVEL_4K); + } + } +#endif +} + +static void gmem_release(struct file *file); +static int gmem_mmap(struct file *file, struct vm_area_struct *vma); +static long gmem_fd_ioctl(struct file *file, unsigned int cmd, unsigned long arg); + +static const struct kvm_gmem_ops gmem_ops = { + .bind = gmem_bind, + .unbind = gmem_unbind, + .get_pfn = gmem_get_pfn, + .populate = gmem_populate, + .release = gmem_release, + .mmap = gmem_mmap, + .ioctl = gmem_fd_ioctl, +}; + +static void gmem_release(struct file *file) +{ + struct gmem_info *info = to_gmem_info(file); + struct kvm_memory_slot *slot; + + /* + * If a memslot is still bound at close time, KVM has not yet had a + * chance to call ops->unbind. Zap the guest mappings for the range + * and clear slot->gmem.file so the eventual unbind is a no-op. This + * matches native gmem's kvm_gmem_release() and prevents the guest + * from continuing to hit backing memory after we free it below. + */ + mutex_lock(&info->lock); + slot = info->bound_slot; + if (slot && info->kvm) { + kvm_gmem_invalidate_range(info->kvm, slot->base_gfn, + slot->base_gfn + slot->npages); + WRITE_ONCE(slot->gmem.file, NULL); + info->bound_slot = NULL; + } + mutex_unlock(&info->lock); + + if (info->kvm) + kvm_put_kvm(info->kvm); + if (info->cma_pages) + free_contig_range(info->base_pfn, info->npages); + kvfree(info->absent); + kfree(info); + module_put(THIS_MODULE); +} + +static int gmem_mmap(struct file *file, struct vm_area_struct *vma) +{ + struct gmem_info *info = to_gmem_info(file); + unsigned long npages = vma_pages(vma); + + /* + * gmem_bind() refuses to bind an mmap-capable fd to a coco VM. A fd + * that was created without GMEM_PROVIDER_FLAG_MMAP_CAPABLE has no + * such gate at bind time, so its mmap must not succeed at any point. + */ + if (!info->mmap_capable) + return -EPERM; + + if (vma->vm_pgoff + npages > info->npages) + return -EINVAL; + + /* Page-less backing: map raw PFNs, not folios. */ + vm_flags_set(vma, VM_PFNMAP | VM_IO | VM_DONTEXPAND | VM_DONTDUMP); + return remap_pfn_range(vma, vma->vm_start, info->base_pfn + vma->vm_pgoff, + npages << PAGE_SHIFT, vma->vm_page_prot); +} + +/* + * Dynamic dma-buf exporter over the provider's backing. + * + * Follows the same shape as drivers/vfio/pci/vfio_pci_dmabuf.c: a per-dmabuf + * priv holding a phys_vec, a revocable dynamic attach, and a "private + * interconnect" symbol iommufd looks up to fetch phys directly (instead of + * mapping through the DMA API). + * + * A revoke on the provider (SET_PRESENT present=0) fans out to every exported + * dma-buf via dma_buf_invalidate_mappings(), so iommufd (which registered a + * revocable importer) tears down the IOMMU mapping alongside KVM's NPT zap. + */ +struct gmem_dmabuf { + struct dma_buf *dmabuf; + struct gmem_info *info; + struct file *provider_file; /* holds info alive */ + struct list_head list; /* info->dmabufs */ + struct phys_vec phys; /* single contiguous range */ + struct kref kref; + struct completion comp; + bool revoked; +}; + +static int gmem_dma_buf_attach(struct dma_buf *dmabuf, + struct dma_buf_attachment *attach) +{ + struct gmem_dmabuf *priv = dmabuf->priv; + + if (!attach->peer2peer) + return -EOPNOTSUPP; + if (priv->revoked) + return -ENODEV; + if (!dma_buf_attach_revocable(attach)) + return -EOPNOTSUPP; + return 0; +} + +static void gmem_dma_buf_done(struct kref *kref) +{ + struct gmem_dmabuf *priv = container_of(kref, struct gmem_dmabuf, kref); + + complete(&priv->comp); +} + +static struct sg_table *gmem_dma_buf_map(struct dma_buf_attachment *attach, + enum dma_data_direction dir) +{ + struct gmem_dmabuf *priv = attach->dmabuf->priv; + struct sg_table *sgt; + + dma_resv_assert_held(priv->dmabuf->resv); + if (priv->revoked) + return ERR_PTR(-ENODEV); + + /* RAM, not P2P MMIO: no p2pdma_provider. */ + sgt = dma_buf_phys_vec_to_sgt(attach, NULL, &priv->phys, 1, + priv->phys.len, dir); + if (IS_ERR(sgt)) + return sgt; + + kref_get(&priv->kref); + return sgt; +} + +static void gmem_dma_buf_unmap(struct dma_buf_attachment *attach, + struct sg_table *sgt, + enum dma_data_direction dir) +{ + struct gmem_dmabuf *priv = attach->dmabuf->priv; + + dma_resv_assert_held(priv->dmabuf->resv); + dma_buf_free_sgt(attach, sgt, dir); + kref_put(&priv->kref, gmem_dma_buf_done); +} + +static void gmem_dma_buf_release(struct dma_buf *dmabuf) +{ + struct gmem_dmabuf *priv = dmabuf->priv; + + if (priv->info) { + mutex_lock(&priv->info->dmabufs_lock); + list_del_init(&priv->list); + mutex_unlock(&priv->info->dmabufs_lock); + } + if (priv->provider_file) + fput(priv->provider_file); + kfree(priv); +} + +static const struct dma_buf_ops gmem_dma_buf_ops = { + .attach = gmem_dma_buf_attach, + .map_dma_buf = gmem_dma_buf_map, + .unmap_dma_buf = gmem_dma_buf_unmap, + .release = gmem_dma_buf_release, +}; + +/* + * Private interconnect for iommufd (mirrors vfio_pci_dma_buf_iommufd_map). + * Returns the single contiguous phys range for the exported region so iommufd + * can program the IOMMU directly, bypassing the DMA API. + */ +int gmem_provider_dma_buf_iommufd_map(struct dma_buf_attachment *attach, + struct phys_vec *phys); +int gmem_provider_dma_buf_iommufd_map(struct dma_buf_attachment *attach, + struct phys_vec *phys) +{ + struct gmem_dmabuf *priv; + + dma_resv_assert_held(attach->dmabuf->resv); + if (attach->dmabuf->ops != &gmem_dma_buf_ops) + return -EOPNOTSUPP; + priv = attach->dmabuf->priv; + if (priv->revoked) + return -ENODEV; + *phys = priv->phys; + return 0; +} +EXPORT_SYMBOL_FOR_MODULES(gmem_provider_dma_buf_iommufd_map, "iommufd"); + +/* Called with info->dmabufs_lock held on the revoke path. */ +static void gmem_dma_buf_revoke_all(struct gmem_info *info) +{ + struct gmem_dmabuf *priv; + + list_for_each_entry(priv, &info->dmabufs, list) { + dma_resv_lock(priv->dmabuf->resv, NULL); + if (!priv->revoked) { + priv->revoked = true; + dma_buf_invalidate_mappings(priv->dmabuf); + } + dma_resv_unlock(priv->dmabuf->resv); + } +} + +static int gmem_provider_get_dmabuf(struct file *file) +{ + struct gmem_info *info = to_gmem_info(file); + DEFINE_DMA_BUF_EXPORT_INFO(exp_info); + struct gmem_dmabuf *priv; + int fd; + + priv = kzalloc(sizeof(*priv), GFP_KERNEL); + if (!priv) + return -ENOMEM; + + priv->info = info; + priv->provider_file = get_file(file); + priv->phys.paddr = (u64)info->base_pfn << PAGE_SHIFT; + priv->phys.len = (u64)info->npages << PAGE_SHIFT; + kref_init(&priv->kref); + init_completion(&priv->comp); + INIT_LIST_HEAD(&priv->list); + + exp_info.ops = &gmem_dma_buf_ops; + exp_info.size = priv->phys.len; + exp_info.flags = O_RDWR; + exp_info.priv = priv; + + priv->dmabuf = dma_buf_export(&exp_info); + if (IS_ERR(priv->dmabuf)) { + fd = PTR_ERR(priv->dmabuf); + kfree(priv); + return fd; + } + + mutex_lock(&info->dmabufs_lock); + list_add(&priv->list, &info->dmabufs); + mutex_unlock(&info->dmabufs_lock); + + fd = dma_buf_fd(priv->dmabuf, O_CLOEXEC); + if (fd < 0) + dma_buf_put(priv->dmabuf); + return fd; +} + +static long gmem_fd_ioctl(struct file *file, unsigned int cmd, unsigned long arg) +{ + struct gmem_info *info = to_gmem_info(file); + struct gmem_provider_present p; + unsigned long start_index, end_index; + + if (cmd == GMEM_PROVIDER_GET_DMABUF) + return gmem_provider_get_dmabuf(file); + + if (cmd != GMEM_PROVIDER_SET_PRESENT) + return -ENOTTY; + if (copy_from_user(&p, (void __user *)arg, sizeof(p))) + return -EFAULT; + if (!p.len || !PAGE_ALIGNED(p.offset) || !PAGE_ALIGNED(p.len)) + return -EINVAL; + + start_index = p.offset >> PAGE_SHIFT; + end_index = start_index + (p.len >> PAGE_SHIFT); + if (end_index > info->npages || end_index < start_index) + return -EINVAL; + + if (p.present) { + /* Restore: next guest fault calls get_pfn() and re-maps. */ + mutex_lock(&info->lock); + bitmap_clear(info->absent, start_index, end_index - start_index); + mutex_unlock(&info->lock); + } else { + unsigned long clamped_start; + + /* Revoke: mark absent, then zap the guest NPT/EPT for the range. */ + mutex_lock(&info->lock); + bitmap_set(info->absent, start_index, end_index - start_index); + + /* + * Translate provider offset -> guest gfn. Skip any part of + * the range that falls outside the currently bound slot; a + * naive subtraction would underflow. + */ + clamped_start = max_t(unsigned long, start_index, info->pgoff); + if (info->kvm && end_index > clamped_start) + kvm_gmem_invalidate_range(info->kvm, + info->base_gfn + clamped_start - info->pgoff, + info->base_gfn + end_index - info->pgoff); + mutex_unlock(&info->lock); + + /* + * Fan out to iommufd (and any other dma-buf importer): mark the + * exported dma-buf(s) revoked and invalidate any active mappings. + * The provider stays ignorant of scratch-page policy; that lives + * in the importer. + */ + mutex_lock(&info->dmabufs_lock); + gmem_dma_buf_revoke_all(info); + mutex_unlock(&info->dmabufs_lock); + } + return 0; +} + +static long gmem_ctl_ioctl(struct file *file, unsigned int cmd, unsigned long arg) +{ + struct gmem_provider_setup setup; + struct gmem_info *info; + struct file *kvm_file; + struct page *pages = NULL; + struct kvm *kvm; + unsigned long npages; + int fd, ret; + + if (cmd != GMEM_PROVIDER_SETUP) + return -ENOTTY; + + if (copy_from_user(&setup, (void __user *)arg, sizeof(setup))) + return -EFAULT; + if (setup.flags & ~GMEM_PROVIDER_FLAG_MMAP_CAPABLE) + return -EINVAL; + + kvm_file = fget(setup.kvm_fd); + if (!kvm_file) + return -EBADF; + if (!file_is_kvm(kvm_file)) { + fput(kvm_file); + return -EINVAL; + } + kvm = kvm_file->private_data; + if (!kvm) { + fput(kvm_file); + return -EINVAL; + } + kvm_get_kvm(kvm); + fput(kvm_file); + + info = kzalloc(sizeof(*info), GFP_KERNEL); + if (!info) { + ret = -ENOMEM; + goto err_put_kvm; + } + + if (addr && len) { + /* External page-less range from module params. */ + info->base_pfn = addr >> PAGE_SHIFT; + info->npages = len >> PAGE_SHIFT; + } else { + /* CMA fallback: allocate a contiguous, page-backed region. */ + if (!setup.size || !PAGE_ALIGNED(setup.size)) { + ret = -EINVAL; + goto err_free_info; + } + npages = setup.size >> PAGE_SHIFT; + pages = alloc_contig_pages(npages, GFP_KERNEL, numa_node_id(), NULL); + if (!pages) { + ret = -ENOMEM; + goto err_free_info; + } + /* Provider path skips KVM's folio-clear; zero to avoid data leak. */ + memset(page_to_virt(pages), 0, (size_t)npages << PAGE_SHIFT); + info->base_pfn = page_to_pfn(pages); + info->npages = npages; + info->cma_pages = pages; + } + + info->absent = kvzalloc(BITS_TO_LONGS(info->npages) * sizeof(unsigned long), + GFP_KERNEL); + if (!info->absent) { + ret = -ENOMEM; + goto err_free_pages; + } + + INIT_LIST_HEAD(&info->dmabufs); + mutex_init(&info->dmabufs_lock); + mutex_init(&info->lock); + + /* + * Pin THIS_MODULE while the provider fd is alive. The fd is created + * with kvm_gmem_fops (owned by kvm.ko), which does not pin us, so + * rmmod of gmem_provider is otherwise free to run behind our ops. + */ + if (!try_module_get(THIS_MODULE)) { + ret = -ENODEV; + goto err_free_pages; + } + + info->backing.ops = &gmem_ops; + info->kvm = kvm; + info->mmap_capable = !!(setup.flags & GMEM_PROVIDER_FLAG_MMAP_CAPABLE); + + fd = anon_inode_getfd("[gmem-provider]", &kvm_gmem_fops, + &info->backing, O_RDWR | O_CLOEXEC); + if (fd < 0) { + ret = fd; + goto err_module_put; + } + return fd; + +err_module_put: + module_put(THIS_MODULE); + +err_free_pages: + if (pages) + free_contig_range(page_to_pfn(pages), npages); +err_free_info: + kvfree(info->absent); + kfree(info); +err_put_kvm: + kvm_put_kvm(kvm); + return ret; +} + +static const struct file_operations gmem_ctl_fops = { + .owner = THIS_MODULE, + .unlocked_ioctl = gmem_ctl_ioctl, + .compat_ioctl = gmem_ctl_ioctl, +}; + +static struct miscdevice gmem_dev = { + .minor = MISC_DYNAMIC_MINOR, + .name = "gmem_provider", + .fops = &gmem_ctl_fops, +}; + +static int __init gmem_provider_init(void) +{ + if ((addr || len) && + (!addr || !len || !PAGE_ALIGNED(addr) || !PAGE_ALIGNED(len))) { + pr_err("gmem_provider: addr= and len= must both be set and page aligned\n"); + return -EINVAL; + } + return misc_register(&gmem_dev); +} +module_init(gmem_provider_init); + +static void __exit gmem_provider_exit(void) +{ + misc_deregister(&gmem_dev); +} +module_exit(gmem_provider_exit); + +MODULE_LICENSE("GPL"); +MODULE_DESCRIPTION("Sample guest_memfd provider (external page-less range or CMA fallback)"); diff --git a/samples/kvm/gmem_provider.h b/samples/kvm/gmem_provider.h new file mode 100644 index 000000000000..45f1b8257f60 --- /dev/null +++ b/samples/kvm/gmem_provider.h @@ -0,0 +1,53 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +#ifndef _SAMPLES_KVM_GMEM_PROVIDER_H +#define _SAMPLES_KVM_GMEM_PROVIDER_H + +#include <linux/ioctl.h> +#include <linux/types.h> + +/* + * ioctl on /dev/gmem_provider: create a guest_memfd provider fd and return it + * for use with KVM_SET_USER_MEMORY_REGION2. + * + * If the module was loaded with addr=/len=, the fd is backed by that fixed + * page-less physical range and @size is ignored. Otherwise the fd is backed + * by a @size-byte physically contiguous region from alloc_contig_pages(). + */ +/* Flags for struct gmem_provider_setup.flags */ +#define GMEM_PROVIDER_FLAG_MMAP_CAPABLE (1u << 0) /* fd is mmap()-able; slot becomes gmem-only. + Refused for coco VMs (SEV-SNP/TDX). */ + +struct gmem_provider_setup { + __s32 kvm_fd; /* an open KVM VM fd */ + __u32 flags; /* GMEM_PROVIDER_FLAG_* */ + __u64 size; /* CMA fallback size in bytes, page aligned */ +}; + +#define GMEM_PROVIDER_IOCTL_BASE 'G' +#define GMEM_PROVIDER_SETUP _IOW(GMEM_PROVIDER_IOCTL_BASE, 1, struct gmem_provider_setup) + +/* + * ioctl on a provider fd (returned by SETUP): flip a byte range of the backing + * between present and absent. Revoking (present=0) marks the range absent and + * zaps the guest's NPT/EPT so the next access re-faults; get_pfn() then refuses + * the range until it is restored (present=1). Models overcommit page reclaim. + */ +struct gmem_provider_present { + __u64 offset; /* byte offset into the provider region, page aligned */ + __u64 len; /* byte length, page aligned */ + __u32 present; /* 0 = revoke (absent), 1 = restore (present) */ + __u32 pad; +}; + +#define GMEM_PROVIDER_SET_PRESENT _IOW(GMEM_PROVIDER_IOCTL_BASE, 2, struct gmem_provider_present) + +/* + * ioctl on a provider fd (returned by SETUP): export the backing region as a + * dynamic dma-buf and return an fd for it, suitable for + * IOMMU_IOAS_MAP_FILE. Revoking the region (SET_PRESENT present=0) fans out + * to the exported dma-buf via dma_buf_invalidate_mappings(), causing iommufd + * to tear down the IOMMU mapping so DMA to the reclaimed range faults. + */ +#define GMEM_PROVIDER_GET_DMABUF _IO(GMEM_PROVIDER_IOCTL_BASE, 3) + +#endif /* _SAMPLES_KVM_GMEM_PROVIDER_H */ diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c index 1da22da0a6ec..d284bb70fe05 100644 --- a/virt/kvm/guest_memfd.c +++ b/virt/kvm/guest_memfd.c @@ -546,7 +546,7 @@ struct file_operations kvm_gmem_fops = { .unlocked_ioctl = kvm_gmem_fops_ioctl, .compat_ioctl = kvm_gmem_fops_ioctl, }; -EXPORT_SYMBOL_FOR_KVM_INTERNAL(kvm_gmem_fops); +EXPORT_SYMBOL_GPL(kvm_gmem_fops); static int kvm_gmem_migrate_folio(struct address_space *mapping, struct folio *dst, struct folio *src, -- 2.55.0 ^ permalink raw reply [flat|nested] 28+ messages in thread
* [RFC PATCH v2 09/11] selftests/kvm: gmem_provider KVM-only tests 2026-07-20 11:03 ` [RFC PATCH v2 01/11] KVM: selftests: sev_smoke_test: Only run VM types the host offers David Woodhouse ` (6 preceding siblings ...) 2026-07-20 11:03 ` [RFC PATCH v2 08/11] samples/kvm: Add guest_memfd backing sample David Woodhouse @ 2026-07-20 11:03 ` David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 10/11] selftests/kvm: gmem_provider iommufd tests David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 11/11] samples/kvm, selftests/kvm: Allow the gmem_provider NVMe DMA test on arm64 David Woodhouse 9 siblings, 0 replies; 28+ messages in thread From: David Woodhouse @ 2026-07-20 11:03 UTC (permalink / raw) To: Sean Christopherson, Paolo Bonzini, kvm Cc: linux-kernel, iommu, linux-kselftest, linux-media, dri-devel, linaro-mm-sig, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, x86, Jason Gunthorpe, Kevin Tian, Joerg Roedel, Will Deacon, Robin Murphy, Shuah Khan, Sumit Semwal, Christian König, Tom Lendacky, Tycho Andersen, Kees Cook, Ashish Kalra, Connor Williamson, Andrew Morton, Zhou Yuhang, Zi Li, Dan Williams, Eric Biggers, Colin Ian King, connordw, graf, fgriffo, David Woodhouse From: David Woodhouse <dwmw@amazon.co.uk> Three selftests that exercise the guest_memfd provider ABI purely from the KVM side (no iommufd, no assigned devices): gmem_provider_test SEV-SNP launch + re-bind against a page-less provider region. Loads a 16 MiB guest at physical addr= from the module parameters, runs SNP_LAUNCH_START -> SNP_LAUNCH_UPDATE -> SNP_LAUNCH_FINISH -> KVM_CREATE_VCPU, tears the VM down, and re-binds the same provider fd to a fresh SNP VM. Proves the LU-survivable path (fd persists, KVM refs transfer, RMP transitions clean). gmem_provider_hugepage_test Non-CoCo (SW_PROTECTED_VM) guest whose slot is backed by a mmap-capable provider fd. Verifies via KVM_GET_STATS that the NPT for a 1 GiB, 1G-aligned slot contains at least one 1 GiB and one 2 MiB mapping. Exercises kvm_gmem_max_mapping_level and the mmap-capable = gmem-only path. gmem_provider_revoke_test GMEM_PROVIDER_SET_PRESENT round trip. Reads a page from the guest (populates NPT), revokes the range, re-enters the guest to observe the fault (memory attribute goes absent, vcpu_run fails EFAULT), restores and re-enters, confirms the guest reads the original value. Exercises the provider-driven kvm_gmem_provider_invalidate helper end-to-end. All three assume the sample provider is loaded with an addr=/len= page-less region; hugepage and revoke additionally require the fd to be set up with GMEM_PROVIDER_FLAG_MMAP_CAPABLE. Signed-off-by: David Woodhouse (Kiro) <dwmw@amazon.co.uk> --- tools/testing/selftests/kvm/Makefile.kvm | 3 + .../kvm/x86/gmem_provider_hugepage_test.c | 130 ++++++++++++ .../kvm/x86/gmem_provider_revoke_test.c | 135 ++++++++++++ .../selftests/kvm/x86/gmem_provider_test.c | 195 ++++++++++++++++++ 4 files changed, 463 insertions(+) create mode 100644 tools/testing/selftests/kvm/x86/gmem_provider_hugepage_test.c create mode 100644 tools/testing/selftests/kvm/x86/gmem_provider_revoke_test.c create mode 100644 tools/testing/selftests/kvm/x86/gmem_provider_test.c diff --git a/tools/testing/selftests/kvm/Makefile.kvm b/tools/testing/selftests/kvm/Makefile.kvm index d28a057fa6c2..e6e9ac9da45a 100644 --- a/tools/testing/selftests/kvm/Makefile.kvm +++ b/tools/testing/selftests/kvm/Makefile.kvm @@ -161,6 +161,9 @@ TEST_GEN_PROGS_x86 += rseq_test TEST_GEN_PROGS_x86 += steal_time TEST_GEN_PROGS_x86 += system_counter_offset_test TEST_GEN_PROGS_x86 += pre_fault_memory_test +TEST_GEN_PROGS_x86 += x86/gmem_provider_test +TEST_GEN_PROGS_x86 += x86/gmem_provider_hugepage_test +TEST_GEN_PROGS_x86 += x86/gmem_provider_revoke_test # Compiled outputs used by test targets TEST_GEN_PROGS_EXTENDED_x86 += x86/nx_huge_pages_test diff --git a/tools/testing/selftests/kvm/x86/gmem_provider_hugepage_test.c b/tools/testing/selftests/kvm/x86/gmem_provider_hugepage_test.c new file mode 100644 index 000000000000..fbbef9761e64 --- /dev/null +++ b/tools/testing/selftests/kvm/x86/gmem_provider_hugepage_test.c @@ -0,0 +1,130 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * gmem_provider_hugepage_test - verify KVM installs 2M and 1G NPT/EPT mappings + * for a non-confidential guest backed entirely by the samples/kvm gmem_provider + * in its gmem-only (mmap) mode. + * + * The test opens the provider with GMEM_PROVIDER_FLAG_MMAP_CAPABLE at SETUP time. + * Load the module with a >=1G, 1G-aligned page-less range: + * insmod gmem_provider.ko addr=0x5D40000000 len=0x42000000 + * + * The provider fd is mmap-capable, so its slot is gmem-only: all guest faults + * route to the provider (no private attributes needed) and the mapping level + * comes from the provider's max_order. The slot's userspace_addr is a + * 1G-aligned mmap of the provider, so the lpage HVA-alignment check passes and + * 1G is permitted. Skips if the module or SW_PROTECTED VM support is absent. + */ +#include <fcntl.h> +#include <errno.h> +#include <stdint.h> +#include <stdio.h> +#include <string.h> +#include <unistd.h> +#include <sys/ioctl.h> +#include <sys/mman.h> + +#include <linux/sizes.h> + +#include "test_util.h" +#include "kvm_util.h" +#include "processor.h" + +/* Mirrors samples/kvm/gmem_provider.h */ +struct gmem_provider_setup { + __s32 kvm_fd; + __u32 flags; + __u64 size; +}; +#define GMEM_PROVIDER_SETUP _IOW('G', 1, struct gmem_provider_setup) +#define GMEM_PROVIDER_FLAG_MMAP_CAPABLE (1u << 0) + +#define DATA_SLOT 10 +#define DATA_GPA (1ULL << 32) /* 4G: 1G-aligned */ +#define DATA_SIZE ((uint64_t)SZ_1G + SZ_2M) + +static void guest_code(void) +{ + *(volatile uint64_t *)DATA_GPA = 0xdead; /* -> 1G mapping */ + *(volatile uint64_t *)(DATA_GPA + SZ_1G) = 0xbeef; /* -> 2M mapping */ + GUEST_DONE(); +} + +int main(void) +{ + struct vm_shape shape = { + .mode = VM_MODE_DEFAULT, + .type = KVM_X86_SW_PROTECTED_VM, + }; + struct gmem_provider_setup setup = { .flags = GMEM_PROVIDER_FLAG_MMAP_CAPABLE }; + struct kvm_vcpu *vcpu; + struct kvm_vm *vm; + struct ucall uc; + int gmem_ctl, gmem_fd, r; + void *resv, *hva; + uint64_t p4k, p2m, p1g; + + TEST_REQUIRE(kvm_check_cap(KVM_CAP_VM_TYPES) & BIT(KVM_X86_SW_PROTECTED_VM)); + + gmem_ctl = open("/dev/gmem_provider", O_RDWR); + __TEST_REQUIRE(gmem_ctl >= 0, + "gmem_provider module not loaded (/dev/gmem_provider absent)"); + + vm = vm_create_shape_with_one_vcpu(shape, &vcpu, guest_code); + + setup.kvm_fd = vm->fd; + setup.size = DATA_SIZE; + gmem_fd = ioctl(gmem_ctl, GMEM_PROVIDER_SETUP, &setup); + TEST_ASSERT(gmem_fd >= 0, "GMEM_PROVIDER_SETUP failed, errno %d", errno); + + /* + * Map the provider at a 1G-aligned host VA so the slot's userspace_addr + * is 1G-congruent with the (1G-aligned) GPA and KVM permits 1G lpages. + * A non-mmap-capable provider (the coco path) would fail this mmap. + */ + resv = mmap(NULL, DATA_SIZE + SZ_1G, PROT_NONE, + MAP_PRIVATE | MAP_ANONYMOUS | MAP_NORESERVE, -1, 0); + TEST_ASSERT(resv != MAP_FAILED, "reserve VA failed, errno %d", errno); + hva = (void *)(((uintptr_t)resv + SZ_1G - 1) & ~((uintptr_t)SZ_1G - 1)); + hva = mmap(hva, DATA_SIZE, PROT_READ | PROT_WRITE, + MAP_SHARED | MAP_FIXED, gmem_fd, 0); + TEST_ASSERT(hva != MAP_FAILED, + "mmap(provider) failed, errno %d", errno); + + /* gmem-only slot: the provider is the whole backing; host uses @hva. */ + r = __vm_set_user_memory_region2(vm, DATA_SLOT, KVM_MEM_GUEST_MEMFD, + DATA_GPA, DATA_SIZE, hva, gmem_fd, 0); + TEST_ASSERT(!r, "KVM_SET_USER_MEMORY_REGION2 failed: %d errno %d", r, errno); + + /* Map only the two pages the guest touches into its page tables. */ + virt_map(vm, DATA_GPA, DATA_GPA, 1); + virt_map(vm, DATA_GPA + SZ_1G, DATA_GPA + SZ_1G, 1); + + vcpu_run(vcpu); + switch (get_ucall(vcpu, &uc)) { + case UCALL_DONE: + break; + case UCALL_ABORT: + REPORT_GUEST_ASSERT(uc); + default: + TEST_FAIL("Unexpected exit: %s", + exit_reason_str(vcpu->run->exit_reason)); + } + + p4k = vm_get_stat(vm, pages_4k); + p2m = vm_get_stat(vm, pages_2m); + p1g = vm_get_stat(vm, pages_1g); + pr_info("gmem-only provider NPT mappings: 4K=%lu 2M=%lu 1G=%lu\n", + p4k, p2m, p1g); + /* Informational: the host mapping and guest share the same memory. */ + pr_info("host reads at DATA_GPA via mmap: 0x%llx\n", + (unsigned long long)*(volatile uint64_t *)hva); + + TEST_ASSERT(p1g > 0, "expected at least one 1G NPT mapping, got 0"); + TEST_ASSERT(p2m > 0, "expected at least one 2M NPT mapping, got 0"); + + kvm_vm_free(vm); + munmap(hva, DATA_SIZE); + close(gmem_fd); + close(gmem_ctl); + return 0; +} diff --git a/tools/testing/selftests/kvm/x86/gmem_provider_revoke_test.c b/tools/testing/selftests/kvm/x86/gmem_provider_revoke_test.c new file mode 100644 index 000000000000..415972aa8a7e --- /dev/null +++ b/tools/testing/selftests/kvm/x86/gmem_provider_revoke_test.c @@ -0,0 +1,135 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * gmem_provider_revoke_test - exercise provider-driven page revocation. + * + * Flips a provider-backed page between present and absent via an ioctl on the + * provider fd and checks the provider->KVM path: revoking zaps the guest's NPT + * (so the guest re-faults into an absent get_pfn and cannot read the page), + * restoring lets the next fault re-map it. This is the mechanism a memory + * overcommit manager would use to reclaim a page from a running guest. + * + * The test opens the provider with GMEM_PROVIDER_FLAG_MMAP_CAPABLE at SETUP time (gmem-only). Load the module with a + * backing region of at least DATA_SIZE, e.g.: + * insmod gmem_provider.ko addr=0x5D40000000 len=0x1000000 + */ +#include <fcntl.h> +#include <errno.h> +#include <stdint.h> +#include <stdio.h> +#include <string.h> +#include <unistd.h> +#include <sys/ioctl.h> +#include <sys/mman.h> + +#include "test_util.h" +#include "kvm_util.h" +#include "processor.h" + +/* Mirrors samples/kvm/gmem_provider.h */ +struct gmem_provider_setup { + __s32 kvm_fd; + __u32 flags; + __u64 size; +}; +struct gmem_provider_present { + __u64 offset; + __u64 len; + __u32 present; + __u32 pad; +}; +#define GMEM_PROVIDER_SETUP _IOW('G', 1, struct gmem_provider_setup) +#define GMEM_PROVIDER_FLAG_MMAP_CAPABLE (1u << 0) +#define GMEM_PROVIDER_SET_PRESENT _IOW('G', 2, struct gmem_provider_present) + +#define DATA_SLOT 10 +#define DATA_GPA (1ULL << 32) +#define DATA_SIZE 0x200000ULL /* 2 MiB region */ +#define MAGIC 0x1234abcdULL + +static void guest_code(void) +{ + *(volatile uint64_t *)DATA_GPA = MAGIC; + for (;;) + GUEST_SYNC(*(volatile uint64_t *)DATA_GPA); +} + +int main(void) +{ + struct vm_shape shape = { + .mode = VM_MODE_DEFAULT, + .type = KVM_X86_SW_PROTECTED_VM, + }; + struct gmem_provider_setup setup = { .flags = GMEM_PROVIDER_FLAG_MMAP_CAPABLE }; + struct gmem_provider_present req; + struct kvm_vcpu *vcpu; + struct kvm_vm *vm; + struct ucall uc; + int gmem_ctl, gmem_fd, r; + void *hva; + + TEST_REQUIRE(kvm_check_cap(KVM_CAP_VM_TYPES) & BIT(KVM_X86_SW_PROTECTED_VM)); + + gmem_ctl = open("/dev/gmem_provider", O_RDWR); + __TEST_REQUIRE(gmem_ctl >= 0, + "gmem_provider module not loaded (/dev/gmem_provider absent)"); + + vm = vm_create_shape_with_one_vcpu(shape, &vcpu, guest_code); + + setup.kvm_fd = vm->fd; + setup.size = DATA_SIZE; + gmem_fd = ioctl(gmem_ctl, GMEM_PROVIDER_SETUP, &setup); + TEST_ASSERT(gmem_fd >= 0, "GMEM_PROVIDER_SETUP failed, errno %d", errno); + + hva = mmap(NULL, DATA_SIZE, PROT_READ | PROT_WRITE, MAP_SHARED, gmem_fd, 0); + TEST_ASSERT(hva != MAP_FAILED, + "mmap(provider) failed, errno %d", errno); + + r = __vm_set_user_memory_region2(vm, DATA_SLOT, KVM_MEM_GUEST_MEMFD, + DATA_GPA, DATA_SIZE, hva, gmem_fd, 0); + TEST_ASSERT(!r, "KVM_SET_USER_MEMORY_REGION2 failed: %d errno %d", r, errno); + virt_map(vm, DATA_GPA, DATA_GPA, 1); + + /* 1) Present: guest writes MAGIC and reads it back. */ + vcpu_run(vcpu); + TEST_ASSERT(get_ucall(vcpu, &uc) == UCALL_SYNC, "expected UCALL_SYNC"); + TEST_ASSERT(uc.args[1] == MAGIC, "guest read 0x%lx, want MAGIC", + (unsigned long)uc.args[1]); + pr_info("present: guest read 0x%llx\n", MAGIC); + + /* 2) Revoke: mark absent and zap the guest NPT (provider->KVM). */ + req = (struct gmem_provider_present){ .offset = 0, .len = 4096, .present = 0 }; + r = ioctl(gmem_fd, GMEM_PROVIDER_SET_PRESENT, &req); + TEST_ASSERT(!r, "revoke ioctl failed, errno %d", errno); + + /* 3) Guest re-reads -> re-fault into absent get_pfn -> must NOT see MAGIC. */ + r = _vcpu_run(vcpu); + /* + * KVM_EXIT_MEMORY_FAULT returns -1 from KVM_RUN with errno == EFAULT. + * Assert unconditionally that we observed the expected fault: + * -1 + EFAULT + exit_reason == KVM_EXIT_MEMORY_FAULT. + */ + TEST_ASSERT(r == -1 && errno == EFAULT && + vcpu->run->exit_reason == KVM_EXIT_MEMORY_FAULT, + "revoked: expected KVM_EXIT_MEMORY_FAULT (r=%d errno=%d exit_reason=%u %s)", + r, errno, vcpu->run->exit_reason, + exit_reason_str(vcpu->run->exit_reason)); + pr_info("revoked: KVM_EXIT_MEMORY_FAULT as expected\n"); + + /* 4) Restore: mark present again. */ + req.present = 1; + r = ioctl(gmem_fd, GMEM_PROVIDER_SET_PRESENT, &req); + TEST_ASSERT(!r, "restore ioctl failed, errno %d", errno); + + /* 5) Re-enter: the fault re-maps via get_pfn, guest reads MAGIC again. */ + vcpu_run(vcpu); + TEST_ASSERT(get_ucall(vcpu, &uc) == UCALL_SYNC, "expected UCALL_SYNC after restore"); + TEST_ASSERT(uc.args[1] == MAGIC, "after restore guest read 0x%lx, want MAGIC", + (unsigned long)uc.args[1]); + pr_info("restored: guest read 0x%llx again -- revoke/restore path works\n", MAGIC); + + kvm_vm_free(vm); + munmap(hva, DATA_SIZE); + close(gmem_fd); + close(gmem_ctl); + return 0; +} diff --git a/tools/testing/selftests/kvm/x86/gmem_provider_test.c b/tools/testing/selftests/kvm/x86/gmem_provider_test.c new file mode 100644 index 000000000000..d7caa11616df --- /dev/null +++ b/tools/testing/selftests/kvm/x86/gmem_provider_test.c @@ -0,0 +1,195 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * gmem_provider_test - exercise the samples/kvm gmem_provider module through a + * full SEV-SNP guest launch, then re-bind the same provider fd to a second VM + * (the live-update path). + * + * The provider module must be loaded first, in either mode: + * insmod gmem_provider.ko addr=0x5D40000000 len=0x1000000 # page-less + * insmod gmem_provider.ko # CMA fallback + * + * The test skips (KSFT_SKIP) if /dev/gmem_provider or SNP support is absent. + * + * Note: with the current kvm_gmem_populate() ABI the launch source is an + * ordinary anonymous buffer (pinned via get_user_pages_fast); the provider's + * populate() copies it into the backing. No /dev/mem mapping is needed. + */ +#include <stdio.h> +#include <stdlib.h> +#include <string.h> +#include <unistd.h> +#include <fcntl.h> +#include <errno.h> +#include <sys/ioctl.h> +#include <sys/mman.h> +#include <stdint.h> +#include <linux/kvm.h> +#include <asm/kvm.h> + +/* Mirrors samples/kvm/gmem_provider.h */ +struct gmem_provider_setup { + int32_t kvm_fd; + uint32_t pad; + uint64_t size; +}; +#define GMEM_PROVIDER_SETUP _IOW('G', 1, struct gmem_provider_setup) + +#define KSFT_PASS 0 +#define KSFT_FAIL 1 +#define KSFT_SKIP 4 + +#define GUEST_MEM_SIZE (16UL * 1024 * 1024) +#define PAGE_SIZE_4K 4096UL + +static int sev_ioctl(int vm_fd, int sev_fd, int cmd, void *data) +{ + struct kvm_sev_cmd sev_cmd = { + .id = cmd, + .data = (uint64_t)(unsigned long)data, + .sev_fd = sev_fd, + }; + return ioctl(vm_fd, KVM_MEMORY_ENCRYPT_OP, &sev_cmd); +} + +static int launch_snp_vm(int kvm_fd, int gmem_fd, void *src, int vm_num) +{ + int vm_fd, vcpu_fd, sev_fd, ret; + + printf("[VM%d] KVM_CREATE_VM (SNP)\n", vm_num); + vm_fd = ioctl(kvm_fd, KVM_CREATE_VM, KVM_X86_SNP_VM); + if (vm_fd < 0) { perror("KVM_CREATE_VM"); return -1; } + + sev_fd = open("/dev/sev", O_RDWR); + if (sev_fd < 0) { perror("open /dev/sev"); close(vm_fd); return -1; } + + struct kvm_sev_init init = { 0 }; + ret = sev_ioctl(vm_fd, sev_fd, KVM_SEV_INIT2, &init); + if (ret) { perror("KVM_SEV_INIT2"); goto out; } + + struct kvm_userspace_memory_region2 region = { + .slot = 0, + .flags = KVM_MEM_GUEST_MEMFD, + .guest_phys_addr = 0, + .memory_size = GUEST_MEM_SIZE, + .userspace_addr = (uint64_t)(unsigned long)src, + .guest_memfd = gmem_fd, + .guest_memfd_offset = 0, + }; + ret = ioctl(vm_fd, KVM_SET_USER_MEMORY_REGION2, ®ion); + if (ret) { perror("KVM_SET_USER_MEMORY_REGION2"); goto out; } + + struct kvm_memory_attributes attrs = { + .address = 0, + .size = GUEST_MEM_SIZE, + .attributes = KVM_MEMORY_ATTRIBUTE_PRIVATE, + }; + ret = ioctl(vm_fd, KVM_SET_MEMORY_ATTRIBUTES, &attrs); + if (ret) { perror("KVM_SET_MEMORY_ATTRIBUTES"); goto out; } + + struct kvm_sev_snp_launch_start start = { .policy = 0x30000 }; + ret = sev_ioctl(vm_fd, sev_fd, KVM_SEV_SNP_LAUNCH_START, &start); + if (ret) { perror("SNP_LAUNCH_START"); goto out; } + + printf("[VM%d] SNP_LAUNCH_UPDATE (code + zero, %luMB)\n", + vm_num, GUEST_MEM_SIZE >> 20); + struct kvm_sev_snp_launch_update update = { + .gfn_start = 0, + .uaddr = (uint64_t)(unsigned long)src, + .len = PAGE_SIZE_4K, + .type = KVM_SEV_SNP_PAGE_TYPE_NORMAL, + }; + ret = sev_ioctl(vm_fd, sev_fd, KVM_SEV_SNP_LAUNCH_UPDATE, &update); + if (ret) { perror("SNP_LAUNCH_UPDATE code"); goto out; } + + struct kvm_sev_snp_launch_update update_zero = { + .gfn_start = 1, + .uaddr = (uint64_t)(unsigned long)(src + PAGE_SIZE_4K), + .len = GUEST_MEM_SIZE - PAGE_SIZE_4K, + .type = KVM_SEV_SNP_PAGE_TYPE_ZERO, + }; + ret = sev_ioctl(vm_fd, sev_fd, KVM_SEV_SNP_LAUNCH_UPDATE, &update_zero); + if (ret) { perror("SNP_LAUNCH_UPDATE zero"); goto out; } + + struct kvm_sev_snp_launch_finish finish = { 0 }; + ret = sev_ioctl(vm_fd, sev_fd, KVM_SEV_SNP_LAUNCH_FINISH, &finish); + if (ret) { perror("SNP_LAUNCH_FINISH"); goto out; } + + vcpu_fd = ioctl(vm_fd, KVM_CREATE_VCPU, 0); + if (vcpu_fd < 0) { perror("KVM_CREATE_VCPU"); ret = -1; goto out; } + + printf("[VM%d] *** launch succeeded ***\n", vm_num); + close(vcpu_fd); + ret = 0; +out: + close(sev_fd); + close(vm_fd); + return ret; +} + +int main(void) +{ + int kvm_fd, gmem_ctl_fd, gmem_fd, tmp_vm; + void *src; + + setbuf(stdout, NULL); + + kvm_fd = open("/dev/kvm", O_RDWR); + if (kvm_fd < 0) { + printf("SKIP: cannot open /dev/kvm (%s)\n", strerror(errno)); + return KSFT_SKIP; + } + + gmem_ctl_fd = open("/dev/gmem_provider", O_RDWR); + if (gmem_ctl_fd < 0) { + printf("SKIP: /dev/gmem_provider not present -- load gmem_provider.ko (%s)\n", + strerror(errno)); + return KSFT_SKIP; + } + + /* Provider setup needs a VM fd; ownership transfers to each VM on bind. */ + tmp_vm = ioctl(kvm_fd, KVM_CREATE_VM, KVM_X86_SNP_VM); + if (tmp_vm < 0) { + printf("SKIP: cannot create SNP VM -- host not SNP-capable? (%s)\n", + strerror(errno)); + return KSFT_SKIP; + } + + struct gmem_provider_setup setup = { + .kvm_fd = tmp_vm, + .size = GUEST_MEM_SIZE, + }; + gmem_fd = ioctl(gmem_ctl_fd, GMEM_PROVIDER_SETUP, &setup); + if (gmem_fd < 0) { + perror("GMEM_PROVIDER_SETUP"); + return KSFT_FAIL; + } + close(tmp_vm); + printf("provider gmem_fd = %d (persists across VMs)\n", gmem_fd); + + /* Launch source: ordinary anonymous memory holding a HLT at gfn 0. */ + src = mmap(NULL, GUEST_MEM_SIZE, PROT_READ | PROT_WRITE, + MAP_PRIVATE | MAP_ANONYMOUS, -1, 0); + if (src == MAP_FAILED) { perror("mmap src"); return KSFT_FAIL; } + memset(src, 0, GUEST_MEM_SIZE); + ((uint8_t *)src)[0] = 0xf4; /* HLT */ + + if (launch_snp_vm(kvm_fd, gmem_fd, src, 1) != 0) { + printf("FAIL: VM1 launch failed\n"); + return KSFT_FAIL; + } + + usleep(100000); + + /* Re-bind the SAME provider fd to a fresh VM (live-update path). */ + if (launch_snp_vm(kvm_fd, gmem_fd, src, 2) != 0) { + printf("FAIL: VM2 re-bind launch failed\n"); + return KSFT_FAIL; + } + + printf("PASS: page-less/provider-backed SNP launch + re-bind succeeded\n"); + munmap(src, GUEST_MEM_SIZE); + close(gmem_fd); + close(gmem_ctl_fd); + close(kvm_fd); + return KSFT_PASS; +} -- 2.55.0 ^ permalink raw reply [flat|nested] 28+ messages in thread
* [RFC PATCH v2 10/11] selftests/kvm: gmem_provider iommufd tests 2026-07-20 11:03 ` [RFC PATCH v2 01/11] KVM: selftests: sev_smoke_test: Only run VM types the host offers David Woodhouse ` (7 preceding siblings ...) 2026-07-20 11:03 ` [RFC PATCH v2 09/11] selftests/kvm: gmem_provider KVM-only tests David Woodhouse @ 2026-07-20 11:03 ` David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 11/11] samples/kvm, selftests/kvm: Allow the gmem_provider NVMe DMA test on arm64 David Woodhouse 9 siblings, 0 replies; 28+ messages in thread From: David Woodhouse @ 2026-07-20 11:03 UTC (permalink / raw) To: Sean Christopherson, Paolo Bonzini, kvm Cc: linux-kernel, iommu, linux-kselftest, linux-media, dri-devel, linaro-mm-sig, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, x86, Jason Gunthorpe, Kevin Tian, Joerg Roedel, Will Deacon, Robin Murphy, Shuah Khan, Sumit Semwal, Christian König, Tom Lendacky, Tycho Andersen, Kees Cook, Ashish Kalra, Connor Williamson, Andrew Morton, Zhou Yuhang, Zi Li, Dan Williams, Eric Biggers, Colin Ian King, connordw, graf, fgriffo, David Woodhouse From: David Woodhouse <dwmw@amazon.co.uk> Three selftests that exercise the guest_memfd provider ABI from the iommufd side, going progressively closer to a real assigned-device setup: gmem_provider_iommufd_test Dual-map. A KVM memslot and an iommufd IOAS are both backed by the same provider dma-buf. The guest writes a magic pattern into the slot; the host mmap of the same provider fd sees the guest's write; the IOAS also holds a mapping of the same backing. Proves KVM and iommufd converge on the same provider phys via two different fds (the KVM guest_memfd fd and the dma-buf fd). gmem_provider_vfio_test A real vfio-pci device is bound to iommufd and attached to an IOAS that has the provider region mapped. Confirms the attach path accepts a dma-buf-backed IOAS without complaint; does not issue device DMA. Requires an unused PCI device already bound to vfio-pci and exposed as /dev/vfio/devices/vfio0 (overridable via GMEM_VFIO_CDEV). gmem_provider_nvme_dma_test Full round trip. Binds an NVMe controller to vfio-pci, attaches it to an IOAS with the provider region mapped at IOVA_DST, initialises NVMe admin queues (also in the IOAS via anonymous pages), submits Identify Controller with PRP1=IOVA_DST, and verifies the 4 KiB response via the host mmap of the same provider fd (VID/SSVID/MN check). Then revokes the target range with GMEM_PROVIDER_SET_PRESENT and re-submits Identify to the same IOVA: iommufd's iopt_revoke_notify tears down the mapping and the device DMA either faults at the IOMMU (visible as IO_PAGE_FAULT in dmesg on AMD-Vi / equivalent on VT-d, SMMUv3) or completes with an NVMe error; either way the provider region is not overwritten. Restores and confirms fresh mappings succeed. Requires an EBS/NVMe scratch device on the test host bound to vfio-pci; the tests SKIP cleanly if no such device is available. Signed-off-by: David Woodhouse (Kiro) <dwmw@amazon.co.uk> --- tools/testing/selftests/kvm/Makefile.kvm | 9 +- .../kvm/x86/gmem_provider_iommufd_test.c | 158 ++++++++ .../kvm/x86/gmem_provider_nvme_dma_test.c | 376 ++++++++++++++++++ .../kvm/x86/gmem_provider_vfio_test.c | 134 +++++++ 4 files changed, 674 insertions(+), 3 deletions(-) create mode 100644 tools/testing/selftests/kvm/x86/gmem_provider_iommufd_test.c create mode 100644 tools/testing/selftests/kvm/x86/gmem_provider_nvme_dma_test.c create mode 100644 tools/testing/selftests/kvm/x86/gmem_provider_vfio_test.c diff --git a/tools/testing/selftests/kvm/Makefile.kvm b/tools/testing/selftests/kvm/Makefile.kvm index e6e9ac9da45a..7f46843e7428 100644 --- a/tools/testing/selftests/kvm/Makefile.kvm +++ b/tools/testing/selftests/kvm/Makefile.kvm @@ -77,6 +77,12 @@ TEST_GEN_PROGS_x86 += x86/evmcs_smm_controls_test TEST_GEN_PROGS_x86 += x86/exit_on_emulation_failure_test TEST_GEN_PROGS_x86 += x86/fastops_test TEST_GEN_PROGS_x86 += x86/fix_hypercall_test +TEST_GEN_PROGS_x86 += x86/gmem_provider_test +TEST_GEN_PROGS_x86 += x86/gmem_provider_hugepage_test +TEST_GEN_PROGS_x86 += x86/gmem_provider_revoke_test +TEST_GEN_PROGS_x86 += x86/gmem_provider_iommufd_test +TEST_GEN_PROGS_x86 += x86/gmem_provider_vfio_test +TEST_GEN_PROGS_x86 += x86/gmem_provider_nvme_dma_test TEST_GEN_PROGS_x86 += x86/hwcr_msr_test TEST_GEN_PROGS_x86 += x86/hyperv_clock TEST_GEN_PROGS_x86 += x86/hyperv_cpuid @@ -161,9 +167,6 @@ TEST_GEN_PROGS_x86 += rseq_test TEST_GEN_PROGS_x86 += steal_time TEST_GEN_PROGS_x86 += system_counter_offset_test TEST_GEN_PROGS_x86 += pre_fault_memory_test -TEST_GEN_PROGS_x86 += x86/gmem_provider_test -TEST_GEN_PROGS_x86 += x86/gmem_provider_hugepage_test -TEST_GEN_PROGS_x86 += x86/gmem_provider_revoke_test # Compiled outputs used by test targets TEST_GEN_PROGS_EXTENDED_x86 += x86/nx_huge_pages_test diff --git a/tools/testing/selftests/kvm/x86/gmem_provider_iommufd_test.c b/tools/testing/selftests/kvm/x86/gmem_provider_iommufd_test.c new file mode 100644 index 000000000000..885dffa0b659 --- /dev/null +++ b/tools/testing/selftests/kvm/x86/gmem_provider_iommufd_test.c @@ -0,0 +1,158 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * gmem_provider_iommufd_test - map the SAME provider-backed region into KVM + * (as a guest_memfd memslot) AND into an iommufd IOAS (via a dma-buf exported + * from the provider fd). Prove both APIs accept the mapping and that a guest + * write is visible via the host mmap. + * + * The iommufd side exercises exactly the provider->iommufd path we just wired: + * GET_DMABUF on the provider fd -> IOMMU_IOAS_MAP_FILE, which walks + * iopt_map_dmabuf -> sym_..._iommufd_map -> gmem_provider_dma_buf_iommufd_map. + * + * The test opens the provider with GMEM_PROVIDER_FLAG_MMAP_CAPABLE at SETUP time; requires iommufd + * available at /dev/iommu. Actual IOMMU page-table programming happens once + * a device is attached; here we validate the plumbing (map succeeds; unmap + * returns the same byte count) without needing a passthrough device. + */ +#include <fcntl.h> +#include <errno.h> +#include <stdint.h> +#include <stdio.h> +#include <string.h> +#include <unistd.h> +#include <sys/ioctl.h> +#include <sys/mman.h> + +#include <linux/iommufd.h> + +#include "test_util.h" +#include "kvm_util.h" +#include "processor.h" + +/* Mirrors samples/kvm/gmem_provider.h */ +struct gmem_provider_setup { + __s32 kvm_fd; + __u32 flags; + __u64 size; +}; +#define GMEM_PROVIDER_SETUP _IOW('G', 1, struct gmem_provider_setup) +#define GMEM_PROVIDER_FLAG_MMAP_CAPABLE (1u << 0) +#define GMEM_PROVIDER_GET_DMABUF _IO('G', 3) + +#define DATA_SLOT 10 +#define DATA_GPA (1ULL << 32) +#define DATA_SIZE 0x200000ULL /* 2 MiB */ +#define IOVA_BASE (1ULL << 34) /* arbitrary */ +#define MAGIC 0xa5a5a5a5a5a5a5a5ULL + +static void guest_code(void) +{ + *(volatile uint64_t *)DATA_GPA = MAGIC; + GUEST_DONE(); +} + +int main(void) +{ + struct vm_shape shape = { + .mode = VM_MODE_DEFAULT, + .type = KVM_X86_SW_PROTECTED_VM, + }; + struct gmem_provider_setup setup = { .flags = GMEM_PROVIDER_FLAG_MMAP_CAPABLE }; + struct iommu_ioas_alloc alloc = {}; + struct iommu_ioas_map_file map = {}; + struct kvm_vcpu *vcpu; + struct kvm_vm *vm; + struct ucall uc; + int gmem_ctl, gmem_fd, dmabuf_fd, iommufd, r; + void *hva; + + TEST_REQUIRE(kvm_check_cap(KVM_CAP_VM_TYPES) & BIT(KVM_X86_SW_PROTECTED_VM)); + + gmem_ctl = open("/dev/gmem_provider", O_RDWR); + __TEST_REQUIRE(gmem_ctl >= 0, + "gmem_provider not loaded (/dev/gmem_provider absent)"); + iommufd = open("/dev/iommu", O_RDWR); + __TEST_REQUIRE(iommufd >= 0, "iommufd unavailable (/dev/iommu absent)"); + + /* + * 1) Provider fd bound to KVM VM. + */ + vm = vm_create_shape_with_one_vcpu(shape, &vcpu, guest_code); + setup.kvm_fd = vm->fd; + setup.size = DATA_SIZE; + gmem_fd = ioctl(gmem_ctl, GMEM_PROVIDER_SETUP, &setup); + TEST_ASSERT(gmem_fd >= 0, "GMEM_PROVIDER_SETUP failed errno=%d", errno); + + /* + * 2) KVM side: host mmap + guest_memfd memslot (gmem-only when mmap-capable). + */ + hva = mmap(NULL, DATA_SIZE, PROT_READ | PROT_WRITE, MAP_SHARED, gmem_fd, 0); + TEST_ASSERT(hva != MAP_FAILED, + "provider mmap failed errno=%d", errno); + r = __vm_set_user_memory_region2(vm, DATA_SLOT, KVM_MEM_GUEST_MEMFD, + DATA_GPA, DATA_SIZE, hva, gmem_fd, 0); + TEST_ASSERT(!r, "KVM_SET_USER_MEMORY_REGION2 failed r=%d errno=%d", r, errno); + virt_map(vm, DATA_GPA, DATA_GPA, 1); + + /* + * Mark the range PRIVATE so guest accesses route via the provider's + * get_pfn (i.e. exercise the guest_memfd path). Without this the + * guest write flows through the shared HVA mmap and the test proves + * nothing about the provider. + */ + vm_mem_set_private(vm, DATA_GPA, DATA_SIZE); + + /* + * 3) iommufd side: allocate IOAS, get a dma-buf from the SAME provider fd, + * and map it into the IOAS via IOMMU_IOAS_MAP_FILE. This drives + * iopt_map_dmabuf -> gmem_provider_dma_buf_iommufd_map. + */ + alloc.size = sizeof(alloc); + r = ioctl(iommufd, IOMMU_IOAS_ALLOC, &alloc); + TEST_ASSERT(!r, "IOMMU_IOAS_ALLOC failed errno=%d", errno); + pr_info("iommufd: allocated ioas id=%u\n", alloc.out_ioas_id); + + dmabuf_fd = ioctl(gmem_fd, GMEM_PROVIDER_GET_DMABUF); + TEST_ASSERT(dmabuf_fd >= 0, "GMEM_PROVIDER_GET_DMABUF failed errno=%d", errno); + pr_info("provider: exported dma-buf fd=%d\n", dmabuf_fd); + + map.size = sizeof(map); + map.flags = IOMMU_IOAS_MAP_FIXED_IOVA | + IOMMU_IOAS_MAP_READABLE | IOMMU_IOAS_MAP_WRITEABLE; + map.ioas_id = alloc.out_ioas_id; + map.fd = dmabuf_fd; + map.start = 0; + map.length = DATA_SIZE; + map.iova = IOVA_BASE; + r = ioctl(iommufd, IOMMU_IOAS_MAP_FILE, &map); + TEST_ASSERT(!r, + "IOMMU_IOAS_MAP_FILE(dma-buf) failed r=%d errno=%d\n" + " (gmem_provider_dma_buf_iommufd_map path)", + r, errno); + pr_info("iommufd: mapped provider dma-buf @ IOVA 0x%llx (0x%llx bytes)\n", + (unsigned long long)map.iova, (unsigned long long)map.length); + + /* + * 4) Run the guest. It writes MAGIC via the KVM/NPT path; we then read + * via the host mmap to confirm the dual-mapping resolves to the same + * physical backing. + */ + vcpu_run(vcpu); + TEST_ASSERT(get_ucall(vcpu, &uc) == UCALL_DONE, + "guest exit: %s", exit_reason_str(vcpu->run->exit_reason)); + + TEST_ASSERT(*(volatile uint64_t *)hva == MAGIC, + "host mmap read 0x%llx, want MAGIC (guest+host+iommufd disagree)", + (unsigned long long)*(volatile uint64_t *)hva); + pr_info("dual-map ok: guest wrote MAGIC, host mmap sees 0x%llx, " + "iommufd IOAS holds the same backing\n", + (unsigned long long)*(volatile uint64_t *)hva); + + kvm_vm_free(vm); + close(dmabuf_fd); + close(iommufd); + munmap(hva, DATA_SIZE); + close(gmem_fd); + close(gmem_ctl); + return 0; +} diff --git a/tools/testing/selftests/kvm/x86/gmem_provider_nvme_dma_test.c b/tools/testing/selftests/kvm/x86/gmem_provider_nvme_dma_test.c new file mode 100644 index 000000000000..665b19028e13 --- /dev/null +++ b/tools/testing/selftests/kvm/x86/gmem_provider_nvme_dma_test.c @@ -0,0 +1,376 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * gmem_provider_nvme_dma_test - real DMA into provider memory via a + * passthrough NVMe controller. + * + * Sends an Identify Controller (opc=0x06, CNS=1) admin command to a + * vfio-pci-bound NVMe controller, with PRP1 pointing at an IOVA that maps + * to the gmem provider's backing. Success = the provider region contains + * the 4 KiB Identify Controller data structure (verified via VID field + * matching the PCI vendor read from config space). + * + * Requires the gmem_provider module loaded, iommufd available, and an NVMe + * controller bound to vfio-pci at /dev/vfio/devices/vfio0 (overridable via + * GMEM_VFIO_CDEV). DESTROYS ANY IN-FLIGHT STATE on that controller. + */ +#include <fcntl.h> +#include <errno.h> +#include <stdint.h> +#include <stdio.h> +#include <stdlib.h> +#include <string.h> +#include <unistd.h> +#include <sys/ioctl.h> +#include <sys/mman.h> + +#include <linux/iommufd.h> +#include <linux/vfio.h> +#include <linux/kvm.h> + +#include "test_util.h" + +struct gmem_provider_setup { + __s32 kvm_fd; + __u32 flags; + __u64 size; +}; +#define GMEM_PROVIDER_SETUP _IOW('G', 1, struct gmem_provider_setup) +#define GMEM_PROVIDER_FLAG_MMAP_CAPABLE (1u << 0) +#define GMEM_PROVIDER_GET_DMABUF _IO('G', 3) + +struct gmem_provider_present { + __u64 offset; + __u64 len; + __u32 present; + __u32 pad; +}; +#define GMEM_PROVIDER_SET_PRESENT _IOW('G', 2, struct gmem_provider_present) + +#define PROVIDER_SIZE 0x200000ULL /* 2 MiB provider region */ +#define IOVA_DST 0x40000000ULL /* target for Identify Controller data */ +#define IOVA_SQ (IOVA_DST + PROVIDER_SIZE) +#define IOVA_CQ (IOVA_SQ + 0x1000) + +/* NVMe controller regs (BAR0). */ +#define NVME_REG_CAP 0x00 +#define NVME_REG_VS 0x08 +#define NVME_REG_CC 0x14 +#define NVME_REG_CSTS 0x1C +#define NVME_REG_AQA 0x24 +#define NVME_REG_ASQ 0x28 +#define NVME_REG_ACQ 0x30 +#define NVME_SQ0TDBL 0x1000 + +#define NVME_CC_EN (1U << 0) +#define NVME_CSTS_RDY (1U << 0) + +static uint32_t rd32(void *b, uint32_t o) { return *(volatile uint32_t *)((char *)b + o); } +static uint64_t rd64(void *b, uint32_t o) { return *(volatile uint64_t *)((char *)b + o); } +static void wr32(void *b, uint32_t o, uint32_t v) { *(volatile uint32_t *)((char *)b + o) = v; } +static void wr64(void *b, uint32_t o, uint64_t v) { *(volatile uint64_t *)((char *)b + o) = v; } + +static int ioas_map_user(int iommufd_fd, uint32_t ioas_id, + void *va, uint64_t iova, uint64_t len) +{ + struct iommu_ioas_map m = { + .size = sizeof(m), + .flags = IOMMU_IOAS_MAP_FIXED_IOVA | + IOMMU_IOAS_MAP_READABLE | + IOMMU_IOAS_MAP_WRITEABLE, + .ioas_id = ioas_id, + .user_va = (uintptr_t)va, + .length = len, + .iova = iova, + }; + return ioctl(iommufd_fd, IOMMU_IOAS_MAP, &m); +} + +int main(void) +{ + struct gmem_provider_setup setup = { .flags = GMEM_PROVIDER_FLAG_MMAP_CAPABLE }; + struct iommu_ioas_alloc alloc = {}; + struct iommu_ioas_map_file mapf = {}; + struct vfio_device_bind_iommufd bind = {}; + struct vfio_device_attach_iommufd_pt att = {}; + struct vfio_region_info reg = {}; + const char *cdev_path; + int gmem_ctl, gmem_fd, dmabuf_fd, iommufd_fd, vfio_fd; + void *provider_hva, *sq_buf, *cq_buf, *bar; + uint16_t pci_cmd; + uint16_t expected_vid; + uint64_t cap, cfg_off; + uint32_t doorbell_stride, aqa, cc, csts; + int i; + + gmem_ctl = open("/dev/gmem_provider", O_RDWR); + __TEST_REQUIRE(gmem_ctl >= 0, "gmem_provider module not loaded"); + iommufd_fd = open("/dev/iommu", O_RDWR); + __TEST_REQUIRE(iommufd_fd >= 0, "iommufd unavailable"); + cdev_path = getenv("GMEM_VFIO_CDEV"); + if (!cdev_path) + cdev_path = "/dev/vfio/devices/vfio0"; + vfio_fd = open(cdev_path, O_RDWR); + __TEST_REQUIRE(vfio_fd >= 0, "vfio cdev absent: %s", cdev_path); + + /* Provider fd (bare KVM VM to satisfy SETUP). */ + { + int kvm = open("/dev/kvm", O_RDWR); + int vm; + TEST_ASSERT(kvm >= 0, "open /dev/kvm errno=%d", errno); + vm = ioctl(kvm, KVM_CREATE_VM, 0); + TEST_ASSERT(vm >= 0, "KVM_CREATE_VM errno=%d", errno); + setup.kvm_fd = vm; + close(kvm); + } + setup.size = PROVIDER_SIZE; + gmem_fd = ioctl(gmem_ctl, GMEM_PROVIDER_SETUP, &setup); + TEST_ASSERT(gmem_fd >= 0, "SETUP errno=%d", errno); + + provider_hva = mmap(NULL, PROVIDER_SIZE, PROT_READ | PROT_WRITE, MAP_SHARED, + gmem_fd, 0); + TEST_ASSERT(provider_hva != MAP_FAILED, "provider mmap errno=%d", errno); + memset(provider_hva, 0xcc, 0x1000); /* poison first page */ + + /* IOAS + provider dma-buf. */ + alloc.size = sizeof(alloc); + TEST_ASSERT(!ioctl(iommufd_fd, IOMMU_IOAS_ALLOC, &alloc), "IOAS_ALLOC"); + dmabuf_fd = ioctl(gmem_fd, GMEM_PROVIDER_GET_DMABUF); + TEST_ASSERT(dmabuf_fd >= 0, "GET_DMABUF"); + mapf.size = sizeof(mapf); + mapf.flags = IOMMU_IOAS_MAP_FIXED_IOVA | IOMMU_IOAS_MAP_READABLE | + IOMMU_IOAS_MAP_WRITEABLE; + mapf.ioas_id = alloc.out_ioas_id; + mapf.fd = dmabuf_fd; + mapf.length = PROVIDER_SIZE; + mapf.iova = IOVA_DST; + TEST_ASSERT(!ioctl(iommufd_fd, IOMMU_IOAS_MAP_FILE, &mapf), "IOAS_MAP_FILE"); + + /* Admin SQ/CQ backing pages (userspace anon, IOMMU-mapped). */ + sq_buf = mmap(NULL, 0x1000, PROT_READ | PROT_WRITE, + MAP_PRIVATE | MAP_ANONYMOUS, -1, 0); + TEST_ASSERT(sq_buf != MAP_FAILED, "sq mmap"); + memset(sq_buf, 0, 0x1000); + cq_buf = mmap(NULL, 0x1000, PROT_READ | PROT_WRITE, + MAP_PRIVATE | MAP_ANONYMOUS, -1, 0); + TEST_ASSERT(cq_buf != MAP_FAILED, "cq mmap"); + memset(cq_buf, 0, 0x1000); + TEST_ASSERT(!ioas_map_user(iommufd_fd, alloc.out_ioas_id, sq_buf, IOVA_SQ, 0x1000), + "IOAS_MAP(sq) errno=%d", errno); + TEST_ASSERT(!ioas_map_user(iommufd_fd, alloc.out_ioas_id, cq_buf, IOVA_CQ, 0x1000), + "IOAS_MAP(cq) errno=%d", errno); + + /* Bind + attach. */ + bind.argsz = sizeof(bind); + bind.iommufd = iommufd_fd; + TEST_ASSERT(!ioctl(vfio_fd, VFIO_DEVICE_BIND_IOMMUFD, &bind), "BIND"); + att.argsz = sizeof(att); + att.pt_id = alloc.out_ioas_id; + TEST_ASSERT(!ioctl(vfio_fd, VFIO_DEVICE_ATTACH_IOMMUFD_PT, &att), "ATTACH"); + pr_info("vfio device attached, pt_id=%u\n", att.pt_id); + + /* mmap BAR0. */ + reg.argsz = sizeof(reg); + reg.index = VFIO_PCI_BAR0_REGION_INDEX; + TEST_ASSERT(!ioctl(vfio_fd, VFIO_DEVICE_GET_REGION_INFO, ®), "GET_REGION_INFO"); + __TEST_REQUIRE(reg.flags & VFIO_REGION_INFO_FLAG_MMAP, "BAR0 not mmap-capable"); + bar = mmap(NULL, reg.size, PROT_READ | PROT_WRITE, MAP_SHARED, vfio_fd, reg.offset); + TEST_ASSERT(bar != MAP_FAILED, "BAR0 mmap"); + pr_info("BAR0 mapped size=0x%llx\n", (unsigned long long)reg.size); + + /* Enable Bus Master. */ + reg.argsz = sizeof(reg); + reg.index = VFIO_PCI_CONFIG_REGION_INDEX; + TEST_ASSERT(!ioctl(vfio_fd, VFIO_DEVICE_GET_REGION_INFO, ®), "CFG region info"); + cfg_off = reg.offset; + TEST_ASSERT(pread(vfio_fd, &pci_cmd, 2, cfg_off + 0x04) == 2, "cfg CMD read"); + pci_cmd |= 0x0006; + TEST_ASSERT(pwrite(vfio_fd, &pci_cmd, 2, cfg_off + 0x04) == 2, "cfg CMD write"); + + /* Snapshot expected PCI vendor id (config offset 0) for the Identify VID cross-check. */ + { + uint16_t v; + + TEST_ASSERT(pread(vfio_fd, &v, 2, cfg_off + 0x00) == 2, "cfg VID read"); + expected_vid = v; + } + + cap = rd64(bar, NVME_REG_CAP); + doorbell_stride = 4U << ((cap >> 32) & 0xf); + pr_info("NVMe CAP=0x%016llx VS=0x%08x MQES=%u DSTRD=%u\n", + (unsigned long long)cap, rd32(bar, NVME_REG_VS), + (uint32_t)((cap & 0xffff) + 1), doorbell_stride); + + /* Disable + wait !RDY. */ + wr32(bar, NVME_REG_CC, 0); + for (i = 0; i < 5000 && (rd32(bar, NVME_REG_CSTS) & NVME_CSTS_RDY); i++) + usleep(1000); + + /* Admin queues: 64 entries each. */ + wr64(bar, NVME_REG_ASQ, IOVA_SQ); + wr64(bar, NVME_REG_ACQ, IOVA_CQ); + aqa = (63U << 16) | 63U; /* ACQS-1 | ASQS-1 */ + wr32(bar, NVME_REG_AQA, aqa); + cc = (6U << 16) | (4U << 20) | NVME_CC_EN; /* IOSQES=6 IOCQES=4 EN */ + wr32(bar, NVME_REG_CC, cc); + + for (i = 0; i < 10000; i++) { + csts = rd32(bar, NVME_REG_CSTS); + if (csts & NVME_CSTS_RDY) + break; + usleep(1000); + if (csts & 0x02) { /* CFS */ + TEST_FAIL("NVMe controller fatal status CSTS=0x%x", csts); + } + } + TEST_ASSERT(csts & NVME_CSTS_RDY, "controller never reached RDY, CSTS=0x%x", csts); + pr_info("NVMe ready. CSTS=0x%x\n", csts); + + /* Identify Controller SQE at slot 0. */ + { + uint32_t *s = (uint32_t *)sq_buf; + memset(s, 0, 64); + s[0] = 0x00010006; /* CID=1 << 16 | OPC=Identify */ + s[1] = 0; /* NSID */ + /* PRP1 (offset 24 in SQE = dword 6) */ + s[6] = (uint32_t)IOVA_DST; + s[7] = (uint32_t)(IOVA_DST >> 32); + s[8] = 0; + s[9] = 0; /* PRP2 */ + s[10] = 0x01; /* CDW10: CNS = Identify Controller */ + } + + /* Ring SQ0 tail doorbell. */ + asm volatile ("" ::: "memory"); + wr32(bar, NVME_SQ0TDBL, 1); + pr_info("Identify Controller submitted (SQ0 tail=1)\n"); + + /* Poll CQ0 slot 0 for a valid completion (phase tag = 1 on first pass). */ + { + volatile uint32_t *c = (volatile uint32_t *)cq_buf; + uint32_t status; + for (i = 0; i < 5000; i++) { + if (c[3] & 0x10000) /* P bit */ + break; + usleep(1000); + } + TEST_ASSERT(c[3] & 0x10000, "completion phase bit never set"); + status = (c[3] >> 17) & 0x7fff; + pr_info("CQE: DW0=0x%08x DW1=0x%08x DW2=0x%08x DW3=0x%08x status=0x%04x\n", + c[0], c[1], c[2], c[3], status); + TEST_ASSERT(status == 0, "Identify Controller status non-zero: 0x%x", status); + fprintf(stderr, "MARKER: about to write CQ doorbell\n"); fflush(stderr); + /* CQ0 head doorbell = 1 */ + wr32(bar, NVME_SQ0TDBL + doorbell_stride, 1); + fprintf(stderr, "MARKER: about to read provider\n"); fflush(stderr); + } + + /* Provider region now holds Identify Controller data. VID/SSVID at bytes 0..3 + * must equal the device's PCI vendor + subsystem vendor (both are Amazon here). */ + { + uint16_t vid = ((uint16_t *)provider_hva)[0]; + uint16_t ssvid = ((uint16_t *)provider_hva)[1]; + uint8_t first = ((uint8_t *)provider_hva)[0]; + uint8_t poison_still = ((uint8_t *)provider_hva)[8]; + char mn[41] = {}; + memcpy(mn, (char *)provider_hva + 24, 40); + fprintf(stderr, "MARKER: read done; byte0=0x%02x byte8=0x%02x\n", + first, poison_still); + fflush(stderr); + pr_info("Identify Controller in provider: VID=0x%04x SSVID=0x%04x MN='%s'\n", + vid, ssvid, mn); + fflush(stdout); + TEST_ASSERT(vid == expected_vid, + "unexpected VID 0x%04x (expected 0x%04x) in provider region -- Identify DMA landed elsewhere?\n" + " poison=0x%02x byte0=0x%02x (poison intact=DMA didn't land here)", + vid, expected_vid, poison_still, first); + pr_info("NVMe -> provider DMA VERIFIED\n"); + } + + /* + * Phase 2: revoke the target page in the provider and verify that the + * next DMA to the same IOVA fails, without touching provider memory. + * + * SET_PRESENT(present=0) marks the range absent in the provider bitmap + * and fans out dma_buf_invalidate_mappings() to every exported dma-buf. + * iommufd's iopt_revoke_notify (registered as .invalidate_mappings on + * the dynamic attach) tears down the IOMMU mapping for that range and + * marks the pages as revoked, so a subsequent device DMA to the same + * IOVA lands with no IOMMU translation and either faults at the IOMMU + * or completes with an NVMe error. Either way the provider region + * must not be overwritten. + */ + { + struct gmem_provider_present rev = { + .offset = 0, .len = 0x1000, .present = 0, + }; + volatile uint32_t *c = (volatile uint32_t *)cq_buf; + uint32_t *s = (uint32_t *)sq_buf; + uint32_t status; + int ret; + + /* Re-seed the target page so we can detect *any* write to it. */ + memset(provider_hva, 0xa5, 0x1000); + + ret = ioctl(gmem_fd, GMEM_PROVIDER_SET_PRESENT, &rev); + TEST_ASSERT(ret == 0, "SET_PRESENT(revoke) errno=%d", errno); + pr_info("Revoked provider offset=0 len=0x1000; dma-buf mappings invalidated\n"); + + /* Issue a second Identify Controller aimed at the revoked IOVA. + * Use a new CID and place the SQE in slot 1; the SQ ring will + * wrap for larger runs but we only ever use slots 0 and 1. + */ + memset(&s[16], 0, 64); + s[16] = 0x00020006; /* CID=2 << 16 | OPC=Identify */ + s[16 + 6] = (uint32_t)IOVA_DST; + s[16 + 7] = (uint32_t)(IOVA_DST >> 32); + s[16 + 10] = 0x01; /* CNS = Identify Controller */ + + asm volatile ("" ::: "memory"); + wr32(bar, NVME_SQ0TDBL, 2); + pr_info("Identify Controller (revoked target) submitted (SQ0 tail=2)\n"); + + /* Poll for CQE at slot 1 with phase bit inverted from first pass. */ + for (i = 0; i < 5000; i++) { + if ((c[4 + 3] & 0x10000)) /* CQE 1 P bit */ + break; + usleep(1000); + } + TEST_ASSERT(c[4 + 3] & 0x10000, "revoke-phase CQE never landed"); + status = (c[4 + 3] >> 17) & 0x7fff; + pr_info("post-revoke CQE: DW0=0x%08x DW3=0x%08x status=0x%04x\n", + c[4 + 0], c[4 + 3], status); + wr32(bar, NVME_SQ0TDBL + doorbell_stride, 2); + + /* Two acceptable outcomes here: (a) NVMe reports a non-zero + * status (typical when the IOMMU fault takes the transaction + * down), or (b) NVMe reports success but the target page was + * not overwritten because the IOMMU rerouted / dropped the + * write. Either proves the revocation took effect on the DMA + * side. A silent success that overwrites provider memory + * would be a bug in our revoke plumbing. + */ + { + uint8_t b0 = ((uint8_t *)provider_hva)[0]; + pr_info("post-revoke provider byte0=0x%02x (seeded 0xa5)\n", b0); + TEST_ASSERT(b0 == 0xa5 || status != 0, + "revoke leaked: status=0 and provider byte0=0x%02x != 0xa5", + b0); + pr_info("Revoked-page DMA is not observable in provider memory\n"); + } + + /* Restore and confirm a fresh IOMMU_IOAS_MAP_FILE succeeds. */ + rev.present = 1; + ret = ioctl(gmem_fd, GMEM_PROVIDER_SET_PRESENT, &rev); + TEST_ASSERT(ret == 0, "SET_PRESENT(restore) errno=%d", errno); + pr_info("Restore + revoke round trip complete\n"); + } + + munmap(bar, reg.size); + close(vfio_fd); + close(dmabuf_fd); + close(iommufd_fd); + munmap(provider_hva, PROVIDER_SIZE); + close(gmem_fd); + close(gmem_ctl); + return 0; +} diff --git a/tools/testing/selftests/kvm/x86/gmem_provider_vfio_test.c b/tools/testing/selftests/kvm/x86/gmem_provider_vfio_test.c new file mode 100644 index 000000000000..97a66a295b14 --- /dev/null +++ b/tools/testing/selftests/kvm/x86/gmem_provider_vfio_test.c @@ -0,0 +1,134 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * gmem_provider_vfio_test - real passthrough device attached to an iommufd + * IOAS holding the provider dma-buf. + * + * Steps: + * 1. SETUP provider (KVM VM shape not required for this test). + * 2. iommufd IOAS + GET_DMABUF + IOMMU_IOAS_MAP_FILE - IOAS holds the + * provider region. + * 3. Open a vfio-pci cdev (default /dev/vfio/devices/vfio0, overridable + * via GMEM_VFIO_CDEV env), VFIO_DEVICE_BIND_IOMMUFD, then + * VFIO_DEVICE_ATTACH_IOMMUFD_PT to the IOAS. + * + * On success the kernel has allocated (or reused) an iommu_domain for the + * device covering the provider region: subsequent DMA by the device to the + * mapped IOVAs would land in provider memory. Actual DMA descriptor + * programming is device-specific and left as a follow-on. + * + * Skips if either the provider module or a bound vfio-pci cdev is absent. + */ +#include <fcntl.h> +#include <errno.h> +#include <stdint.h> +#include <stdio.h> +#include <stdlib.h> +#include <string.h> +#include <unistd.h> +#include <sys/ioctl.h> + +#include <linux/iommufd.h> +#include <linux/vfio.h> +#include <linux/kvm.h> + +#include "test_util.h" + +struct gmem_provider_setup { + __s32 kvm_fd; + __u32 flags; + __u64 size; +}; +#define GMEM_PROVIDER_SETUP _IOW('G', 1, struct gmem_provider_setup) +#define GMEM_PROVIDER_FLAG_MMAP_CAPABLE (1u << 0) +#define GMEM_PROVIDER_GET_DMABUF _IO('G', 3) + +#define DATA_SIZE 0x200000ULL +#define IOVA_BASE (1ULL << 34) + +int main(void) +{ + struct gmem_provider_setup setup = { .flags = GMEM_PROVIDER_FLAG_MMAP_CAPABLE }; + struct iommu_ioas_alloc alloc = {}; + struct iommu_ioas_map_file map = {}; + struct vfio_device_bind_iommufd bind = {}; + struct vfio_device_attach_iommufd_pt att = {}; + const char *cdev_path; + int gmem_ctl, gmem_fd, dmabuf_fd, iommufd_fd, vfio_fd, r; + + gmem_ctl = open("/dev/gmem_provider", O_RDWR); + __TEST_REQUIRE(gmem_ctl >= 0, "gmem_provider module not loaded"); + iommufd_fd = open("/dev/iommu", O_RDWR); + __TEST_REQUIRE(iommufd_fd >= 0, "iommufd unavailable (/dev/iommu absent)"); + + cdev_path = getenv("GMEM_VFIO_CDEV"); + if (!cdev_path) + cdev_path = "/dev/vfio/devices/vfio0"; + vfio_fd = open(cdev_path, O_RDWR); + __TEST_REQUIRE(vfio_fd >= 0, + "no vfio-pci cdev at %s (bind a device first)", cdev_path); + pr_info("using vfio cdev: %s (fd=%d)\n", cdev_path, vfio_fd); + + /* Provider fd (KVM binding not needed here, but SETUP still requires a + * valid kvm_fd; pass /dev/kvm's own fd -- the SETUP path just fgets it + * to grab a KVM ref, and /dev/kvm's file has ->private_data==NULL, so + * the "kvm not attached" check must be relaxed OR we create a VM. + * For robustness, create a bare KVM VM. */ + { + int kvm = open("/dev/kvm", O_RDWR); + int vm; + TEST_ASSERT(kvm >= 0, "open /dev/kvm errno=%d", errno); + vm = ioctl(kvm, KVM_CREATE_VM, 0); + TEST_ASSERT(vm >= 0, "KVM_CREATE_VM errno=%d", errno); + setup.kvm_fd = vm; + close(kvm); + } + setup.size = DATA_SIZE; + gmem_fd = ioctl(gmem_ctl, GMEM_PROVIDER_SETUP, &setup); + TEST_ASSERT(gmem_fd >= 0, "GMEM_PROVIDER_SETUP errno=%d", errno); + + /* IOAS + provider dma-buf */ + alloc.size = sizeof(alloc); + r = ioctl(iommufd_fd, IOMMU_IOAS_ALLOC, &alloc); + TEST_ASSERT(!r, "IOMMU_IOAS_ALLOC errno=%d", errno); + + dmabuf_fd = ioctl(gmem_fd, GMEM_PROVIDER_GET_DMABUF); + TEST_ASSERT(dmabuf_fd >= 0, "GET_DMABUF errno=%d", errno); + + map.size = sizeof(map); + map.flags = IOMMU_IOAS_MAP_FIXED_IOVA | + IOMMU_IOAS_MAP_READABLE | IOMMU_IOAS_MAP_WRITEABLE; + map.ioas_id = alloc.out_ioas_id; + map.fd = dmabuf_fd; + map.length = DATA_SIZE; + map.iova = IOVA_BASE; + r = ioctl(iommufd_fd, IOMMU_IOAS_MAP_FILE, &map); + TEST_ASSERT(!r, "IOMMU_IOAS_MAP_FILE errno=%d", errno); + pr_info("IOAS %u holds provider dma-buf @ IOVA 0x%llx (0x%llx bytes)\n", + alloc.out_ioas_id, + (unsigned long long)map.iova, (unsigned long long)map.length); + + /* Bind VFIO cdev to iommufd */ + bind.argsz = sizeof(bind); + bind.iommufd = iommufd_fd; + r = ioctl(vfio_fd, VFIO_DEVICE_BIND_IOMMUFD, &bind); + TEST_ASSERT(!r, "VFIO_DEVICE_BIND_IOMMUFD errno=%d", errno); + pr_info("VFIO device bound: devid=%u\n", bind.out_devid); + + /* Attach the device to our IOAS */ + att.argsz = sizeof(att); + att.pt_id = alloc.out_ioas_id; + r = ioctl(vfio_fd, VFIO_DEVICE_ATTACH_IOMMUFD_PT, &att); + TEST_ASSERT(!r, + "VFIO_DEVICE_ATTACH_IOMMUFD_PT errno=%d\n" + " (real device -> IOAS attach failed)", errno); + pr_info("VFIO device attached to IOAS -> pt_id=%u (kernel-selected)\n", + att.pt_id); + pr_info("real passthrough device now has IOMMU domain covering provider region\n"); + + close(vfio_fd); + close(dmabuf_fd); + close(iommufd_fd); + close(gmem_fd); + close(gmem_ctl); + return 0; +} -- 2.55.0 ^ permalink raw reply [flat|nested] 28+ messages in thread
* [RFC PATCH v2 11/11] samples/kvm, selftests/kvm: Allow the gmem_provider NVMe DMA test on arm64 2026-07-20 11:03 ` [RFC PATCH v2 01/11] KVM: selftests: sev_smoke_test: Only run VM types the host offers David Woodhouse ` (8 preceding siblings ...) 2026-07-20 11:03 ` [RFC PATCH v2 10/11] selftests/kvm: gmem_provider iommufd tests David Woodhouse @ 2026-07-20 11:03 ` David Woodhouse 9 siblings, 0 replies; 28+ messages in thread From: David Woodhouse @ 2026-07-20 11:03 UTC (permalink / raw) To: Sean Christopherson, Paolo Bonzini, kvm Cc: linux-kernel, iommu, linux-kselftest, linux-media, dri-devel, linaro-mm-sig, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, x86, Jason Gunthorpe, Kevin Tian, Joerg Roedel, Will Deacon, Robin Murphy, Shuah Khan, Sumit Semwal, Christian König, Tom Lendacky, Tycho Andersen, Kees Cook, Ashish Kalra, Connor Williamson, Andrew Morton, Zhou Yuhang, Zi Li, Dan Williams, Eric Biggers, Colin Ian King, connordw, graf, fgriffo, David Woodhouse From: David Woodhouse <dwmw@amazon.co.uk> The sample provider and the NVMe DMA selftest have no x86 specifics. Drop the X86_64 gate on the sample's Kconfig (arm64 also has KVM_GUEST_MEMFD; the SEV RMP reset is already guarded by CONFIG_AMD_MEM_ENCRYPT) and move the NVMe DMA selftest from TEST_GEN_PROGS_x86 to TEST_GEN_PROGS_COMMON so it also builds on architectures with an IOMMU (arm64 SMMU-v3, ...). The test itself does not need to change: it uses KVM_CREATE_VM(0) as a client for the sample provider's SETUP ioctl and then exercises the iommufd + dma-buf + provider revoke path with direct NVMe MMIO, none of which is x86-specific. On arm64 it validates the same flow through the ARM SMMU. Signed-off-by: David Woodhouse (Kiro) <dwmw@amazon.co.uk> --- samples/Kconfig | 2 +- tools/testing/selftests/kvm/Makefile.kvm | 2 +- .../selftests/kvm/{x86 => }/gmem_provider_nvme_dma_test.c | 0 3 files changed, 2 insertions(+), 2 deletions(-) rename tools/testing/selftests/kvm/{x86 => }/gmem_provider_nvme_dma_test.c (100%) diff --git a/samples/Kconfig b/samples/Kconfig index 3482659638b5..d26a03dea072 100644 --- a/samples/Kconfig +++ b/samples/Kconfig @@ -328,7 +328,7 @@ source "samples/damon/Kconfig" config SAMPLE_KVM_GMEM_PROVIDER tristate "Build sample guest_memfd provider -- loadable module only" - depends on KVM_GUEST_MEMFD && X86_64 && m + depends on KVM_GUEST_MEMFD && (X86_64 || ARM64) && m help This builds a sample external guest_memfd provider module with two backing modes. Loaded with addr=/len= it backs guest memory with a diff --git a/tools/testing/selftests/kvm/Makefile.kvm b/tools/testing/selftests/kvm/Makefile.kvm index 7f46843e7428..12004a487c32 100644 --- a/tools/testing/selftests/kvm/Makefile.kvm +++ b/tools/testing/selftests/kvm/Makefile.kvm @@ -58,6 +58,7 @@ TEST_PROGS_x86 += x86/nx_huge_pages_test.sh # Compiled test targets valid on all architectures with libkvm support TEST_GEN_PROGS_COMMON = demand_paging_test TEST_GEN_PROGS_COMMON += dirty_log_test +TEST_GEN_PROGS_COMMON += gmem_provider_nvme_dma_test TEST_GEN_PROGS_COMMON += guest_print_test TEST_GEN_PROGS_COMMON += irqfd_test TEST_GEN_PROGS_COMMON += kvm_binary_stats_test @@ -82,7 +83,6 @@ TEST_GEN_PROGS_x86 += x86/gmem_provider_hugepage_test TEST_GEN_PROGS_x86 += x86/gmem_provider_revoke_test TEST_GEN_PROGS_x86 += x86/gmem_provider_iommufd_test TEST_GEN_PROGS_x86 += x86/gmem_provider_vfio_test -TEST_GEN_PROGS_x86 += x86/gmem_provider_nvme_dma_test TEST_GEN_PROGS_x86 += x86/hwcr_msr_test TEST_GEN_PROGS_x86 += x86/hyperv_clock TEST_GEN_PROGS_x86 += x86/hyperv_cpuid diff --git a/tools/testing/selftests/kvm/x86/gmem_provider_nvme_dma_test.c b/tools/testing/selftests/kvm/gmem_provider_nvme_dma_test.c similarity index 100% rename from tools/testing/selftests/kvm/x86/gmem_provider_nvme_dma_test.c rename to tools/testing/selftests/kvm/gmem_provider_nvme_dma_test.c -- 2.55.0 ^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [RFC PATCH v2 00/11] KVM: Allow alternative providers of guest_memfd backed by PFNMAP memory 2026-07-20 11:03 [RFC PATCH v2 00/11] KVM: Allow alternative providers of guest_memfd backed by PFNMAP memory David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 01/11] KVM: selftests: sev_smoke_test: Only run VM types the host offers David Woodhouse @ 2026-07-20 15:11 ` Paolo Bonzini 2026-07-20 16:39 ` David Woodhouse 2026-07-23 0:24 ` Ackerley Tng 2026-10-05 9:55 ` [RFC PATCH 0/6] KVM: guest_memfd: back guest_memfd with an imported dma-buf Fred Griffoul 3 siblings, 1 reply; 28+ messages in thread From: Paolo Bonzini @ 2026-07-20 15:11 UTC (permalink / raw) To: David Woodhouse, Sean Christopherson, kvm Cc: linux-kernel, iommu, linux-kselftest, linux-media, dri-devel, linaro-mm-sig, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, x86, Jason Gunthorpe, Kevin Tian, Joerg Roedel, Will Deacon, Robin Murphy, Shuah Khan, Sumit Semwal, Christian König, Tom Lendacky, Tycho Andersen, Kees Cook, Ashish Kalra, Connor Williamson, Andrew Morton, Zhou Yuhang, Zi Li, Dan Williams, Eric Biggers, Colin Ian King, connordw, graf, fgriffo On 7/20/26 13:03, David Woodhouse wrote: > I'll continue to bikeshed it myself, but this is at least functional > enough to show the direction and solicit further opinions... > > https://git.infradead.org/?p=users/dwmw2/linux.git;a=shortlog;h=gmem-provider-v2 In the meanwhile, I have very ignorantly applied patches 1 and 2 as testsuite fixes. :) Paolo ^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [RFC PATCH v2 00/11] KVM: Allow alternative providers of guest_memfd backed by PFNMAP memory 2026-07-20 15:11 ` [RFC PATCH v2 00/11] KVM: Allow alternative providers of guest_memfd backed by PFNMAP memory Paolo Bonzini @ 2026-07-20 16:39 ` David Woodhouse 0 siblings, 0 replies; 28+ messages in thread From: David Woodhouse @ 2026-07-20 16:39 UTC (permalink / raw) To: Paolo Bonzini, Sean Christopherson, kvm Cc: linux-kernel, iommu, linux-kselftest, linux-media, dri-devel, linaro-mm-sig, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, x86, Jason Gunthorpe, Kevin Tian, Joerg Roedel, Will Deacon, Robin Murphy, Shuah Khan, Sumit Semwal, Christian König, Tom Lendacky, Tycho Andersen, Kees Cook, Ashish Kalra, Connor Williamson, Andrew Morton, Zhou Yuhang, Zi Li, Dan Williams, Eric Biggers, Colin Ian King, connordw, graf, fgriffo [-- Attachment #1: Type: text/plain, Size: 530 bytes --] On Mon, 2026-07-20 at 17:11 +0200, Paolo Bonzini wrote: > On 7/20/26 13:03, David Woodhouse wrote: > > I'll continue to bikeshed it myself, but this is at least functional > > enough to show the direction and solicit further opinions... > > > > https://git.infradead.org/?p=users/dwmw2/linux.git;a=shortlog;h=gmem-provider-v2 > > In the meanwhile, I have very ignorantly applied patches 1 and 2 as > testsuite fixes. :) > Thanks. I think #3 can go in on its own too; will try to drum up some review for that... [-- Attachment #2: smime.p7s --] [-- Type: application/pkcs7-signature, Size: 6179 bytes --] ^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [RFC PATCH v2 00/11] KVM: Allow alternative providers of guest_memfd backed by PFNMAP memory 2026-07-20 11:03 [RFC PATCH v2 00/11] KVM: Allow alternative providers of guest_memfd backed by PFNMAP memory David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 01/11] KVM: selftests: sev_smoke_test: Only run VM types the host offers David Woodhouse 2026-07-20 15:11 ` [RFC PATCH v2 00/11] KVM: Allow alternative providers of guest_memfd backed by PFNMAP memory Paolo Bonzini @ 2026-07-23 0:24 ` Ackerley Tng 2026-07-23 9:40 ` David Woodhouse 2026-10-05 9:55 ` [RFC PATCH 0/6] KVM: guest_memfd: back guest_memfd with an imported dma-buf Fred Griffoul 3 siblings, 1 reply; 28+ messages in thread From: Ackerley Tng @ 2026-07-23 0:24 UTC (permalink / raw) To: David Woodhouse, Sean Christopherson, Paolo Bonzini, kvm Cc: linux-kernel, iommu, linux-kselftest, linux-media, dri-devel, linaro-mm-sig, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, x86, Jason Gunthorpe, Kevin Tian, Joerg Roedel, Will Deacon, Robin Murphy, Shuah Khan, Sumit Semwal, Christian König, Tom Lendacky, Tycho Andersen, Kees Cook, Ashish Kalra, Connor Williamson, Andrew Morton, Zhou Yuhang, Zi Li, Dan Williams, Eric Biggers, Colin Ian King, connordw, graf, fgriffo David Woodhouse <dwmw2@infradead.org> writes: > For dedicated hosting environments, there are a few reasons to prefer > using external PFNMAP memory for guests, rather than kernel-managed > memory: > > • No struct page overheads > > • Easily managed in 1GiB pages (or larger chunks); no fragmentation/THP > > • Faster kexec/KHO live update (every millisecond spent faffing around > with memory management is an extra millisecond of steal time taken > from guests) > Also, no need to split/merge pages when converting from private to shared if the host mappings can't depend on the refcounts in struct pages :). > Today, there is only one implementation of guest_memfd: the internal > shmem-based page cache in virt/kvm/guest_memfd.c. This series allows > for other code to provide "a guest_memfd". This is an early path- > finding proof of concept tested with AMD SEV-SNP as well as non-CoCo > guests on x86 and arm64. I'm planning to build a memory-based file > system which gives files for both guest_memfd as well as memory-based > storage to allow things like userspace VM and orchestrator state > to be passed over KHO/kexec. > Thanks for raising this, good to know that there is more interest! Vishal and I have been working on guest_memfd with HugeTLB for a while now. I'll be restarting upstream work on HugeTLB after the conversions series merges. Vishal, Frank and I brought up having guest_memfd with HugeTLB with VM_PFNMAP mappings at a guest_memfd biweekly meeting a while ago. The consensus was that guest_memfd HugeTLB must support GUP and so VM_PFNMAP can't be the only way that guest_memfd supports HugeTLB. Your proposal is true non-page-struct memory, which has been on the discussion guest_memfd biweekly discussion backlog for a while now. Google is definitely also interested in this. Frank also described some use cases [1] other than HugeTLB and the regular PAGE_SIZE pages from the buddy that is currently supported. I submitted a proposal for LPC 2026 to discuss the interface. What I had in mind was to have the KVM_GUEST_MEMFD_CREATE ioctl take a "template" fd. guest_memfd will then read off what it needs from the "template" fd (ops, like you have here), and the "template" fd could be VFIO [2] or HugeTLB. You have a provider_fd = open("/dev/gmem_provider"), how about passing that provider_fd to guest_memfd, where /dev/gmem_provider could represent the page-less memory, like: ioctl(kvm_fd, KVM_GUEST_MEMFD_CREATE, provider_fd); It would be cool if the fallback could be set up flexibly, like: ioctl(kvm_fd, KVM_GUEST_MEMFD_CREATE, [provider_fd, fallback_provider_fd, ...]); I think guest_memfd should be the wrapper around other memory providers because then the CoCo shared/private logic and special page table handling for both host and stage 2 page tables is centralized in guest_memfd, and also, KVM doesn't have to understand other, different kinds of fds other than guest_memfd. [1] https://lore.kernel.org/all/CAPTztWajm_JLpp9BjRcX=h72r25ELrXeGkOXVachybBxLJGS=g@mail.gmail.com/ [2] https://lore.kernel.org/all/CAEvNRgFpJWQ5M5sQhGpQUV3GbBq9N+MQhhaxdxa=D8ky94SCsw@mail.gmail.com/ > This series lifts guest_memfd's operations into a small ops struct and > converts KVM's own gmem to use it, so it becomes one implementation > among peers rather than the only game in town. It also adapts the > SEV-SNP fast paths to page-less operation, plumbs iommufd to map a > guest_memfd-backing dma-buf as CPU RAM (not MMIO), and adds the > provider->KVM revocation primitive that makes overcommit-style reclaim > possible. > > It's largely AI-built at the moment. My main design concerns are around > how we tell that a given file is "a guest_memfd", and how we plug into > IOMMUFD. Presenting a dma-buf was the less intrusive choice because that > path already has support for taking pages away, but I'm less convinced > that it's the cleaner choice in the long term. > IIUC from this series, you're creating a dmabuf fd out of a guest_memfd, what does "taking pages away" mean, in this case, from where/who? > > [...snip...] > ^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [RFC PATCH v2 00/11] KVM: Allow alternative providers of guest_memfd backed by PFNMAP memory 2026-07-23 0:24 ` Ackerley Tng @ 2026-07-23 9:40 ` David Woodhouse 2026-07-23 16:01 ` Ackerley Tng 0 siblings, 1 reply; 28+ messages in thread From: David Woodhouse @ 2026-07-23 9:40 UTC (permalink / raw) To: Ackerley Tng, Sean Christopherson, Paolo Bonzini, kvm, Xu Yilun Cc: linux-kernel, iommu, linux-kselftest, linux-media, dri-devel, linaro-mm-sig, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, x86, Jason Gunthorpe, Kevin Tian, Joerg Roedel, Will Deacon, Robin Murphy, Shuah Khan, Sumit Semwal, Christian König, Tom Lendacky, Tycho Andersen, Kees Cook, Ashish Kalra, Connor Williamson, Andrew Morton, Zhou Yuhang, Zi Li, Dan Williams, Eric Biggers, Colin Ian King, connordw, graf, fgriffo [-- Attachment #1: Type: text/plain, Size: 7055 bytes --] On Wed, 2026-07-22 at 17:24 -0700, Ackerley Tng wrote: > David Woodhouse <dwmw2@infradead.org> writes: > > > For dedicated hosting environments, there are a few reasons to prefer > > using external PFNMAP memory for guests, rather than kernel-managed > > memory: > > > > • No struct page overheads > > > > • Easily managed in 1GiB pages (or larger chunks); no fragmentation/THP > > > > • Faster kexec/KHO live update (every millisecond spent faffing around > > with memory management is an extra millisecond of steal time taken > > from guests) > > > > Also, no need to split/merge pages when converting from private to > shared if the host mappings can't depend on the refcounts in struct > pages :). > > > Today, there is only one implementation of guest_memfd: the internal > > shmem-based page cache in virt/kvm/guest_memfd.c. This series allows > > for other code to provide "a guest_memfd". This is an early path- > > finding proof of concept tested with AMD SEV-SNP as well as non-CoCo > > guests on x86 and arm64. I'm planning to build a memory-based file > > system which gives files for both guest_memfd as well as memory-based > > storage to allow things like userspace VM and orchestrator state > > to be passed over KHO/kexec. > > > > Thanks for raising this, good to know that there is more interest! > > Vishal and I have been working on guest_memfd with HugeTLB for a while > now. I'll be restarting upstream work on HugeTLB after the conversions > series merges. > > Vishal, Frank and I brought up having guest_memfd with HugeTLB with > VM_PFNMAP mappings at a guest_memfd biweekly meeting a while ago. The > consensus was that guest_memfd HugeTLB must support GUP and so VM_PFNMAP > can't be the only way that guest_memfd supports HugeTLB. > > Your proposal is true non-page-struct memory, which has been on the > discussion guest_memfd biweekly discussion backlog for a while > now. Google is definitely also interested in this. > > Frank also described some use cases [1] other than HugeTLB and the > regular PAGE_SIZE pages from the buddy that is currently supported. > > I submitted a proposal for LPC 2026 to discuss the interface. What I had > in mind was to have the KVM_GUEST_MEMFD_CREATE ioctl take a "template" > fd. guest_memfd will then read off what it needs from the "template" fd > (ops, like you have here), and the "template" fd could be VFIO [2] or > HugeTLB. > > You have a provider_fd = open("/dev/gmem_provider"), how about passing > that provider_fd to guest_memfd, where /dev/gmem_provider could > represent the page-less memory, like: > > ioctl(kvm_fd, KVM_GUEST_MEMFD_CREATE, provider_fd); > > It would be cool if the fallback could be set up flexibly, like: > > ioctl(kvm_fd, KVM_GUEST_MEMFD_CREATE, > [provider_fd, fallback_provider_fd, ...]); > > I think guest_memfd should be the wrapper around other memory providers > because then the CoCo shared/private logic and special page table > handling for both host and stage 2 page tables is centralized in > guest_memfd, and also, KVM doesn't have to understand other, different > kinds of fds other than guest_memfd. Yeah. My approach was to make those "other kinds of fds" essentially the same — they are "a guest_memfd", and KVM didn't have to understand anything else (after a little bit of rework to have the generic ops so that other code can create its own guest_memfds). As opposed to the approach in Xu Yilun's TDISP series, which did teach KVM about a new kind of memslot. Likewise, I used dma-buf so that IOMMUFD didn't have to understand anything new either. I like your approach, inverting the wrapping of guest_memfd/dma-buf. The provider then gives a plain dma-buf (not associated with any KVM as it is in my current code). Userspace can use that *directly* with the IOMMU, and passes it to your KVM_GUEST_MEMFD_CREATE in order to get a guest_memfd for use with a given KVM. I think a lot of the warts in my series fall away if we do it that way round. I'll have a go at reworking it that way round... and see what different warts it ends up with. The interesting part will be getting actual PFNs out of the dma-buf (which is already under active discussion). I'm not sure I understand the 'fallback provider' part. If one provider doesn't have a page for a given address, KVM would iterate over them all until it finds one that does? Is that about falling back from one THP pool to another? Is this something that the generic KVM code needs to know about, or could it be handled *within* a provider, which can chain to the next pool if it can't find a page? Wouldn't we want that done transparently for both KVM and IOMMU consumers? > > [1] https://lore.kernel.org/all/CAPTztWajm_JLpp9BjRcX=h72r25ELrXeGkOXVachybBxLJGS=g@mail.gmail.com/ > [2] https://lore.kernel.org/all/CAEvNRgFpJWQ5M5sQhGpQUV3GbBq9N+MQhhaxdxa=D8ky94SCsw@mail.gmail.com/ > > > This series lifts guest_memfd's operations into a small ops struct and > > converts KVM's own gmem to use it, so it becomes one implementation > > among peers rather than the only game in town. It also adapts the > > SEV-SNP fast paths to page-less operation, plumbs iommufd to map a > > guest_memfd-backing dma-buf as CPU RAM (not MMIO), and adds the > > provider->KVM revocation primitive that makes overcommit-style reclaim > > possible. > > > > It's largely AI-built at the moment. My main design concerns are around > > how we tell that a given file is "a guest_memfd", and how we plug into > > IOMMUFD. Presenting a dma-buf was the less intrusive choice because that > > path already has support for taking pages away, but I'm less convinced > > that it's the cleaner choice in the long term. > > > > IIUC from this series, you're creating a dmabuf fd out of a guest_memfd, > what does "taking pages away" mean, in this case, from where/who? Taken away from the guest. Ballooning pages out, either for traditional memory overcommit, or for the models where a guest donates some of its pages to the system to run a "child" which is properly isolated from the guest. The physical page is reclaimed and needs to be taken out of the KVM and IOMMU stage2 mappings. If KVM wants to touch it again it must call the provider to get a new PFN for that range. This is roughly equivalent to the MMU notifier path on userspace VMAs in the normal memslot case. IOMMUFD has no equivalent of that for userspace addresses — IOMMU_IOAS_MAP pins the pages so they *can't* go away, which is why I used dma-buf for the IOMMUFD interface. I've actually been fleshing out the overcommit path a bit more, experimenting with adding async-pf support to guest_memfd in https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=894dd0fcb34 (the sample provider just pretends its pages are slow to populate, so the guest gets to run something else while a "reclaimed" page comes back). [-- Attachment #2: smime.p7s --] [-- Type: application/pkcs7-signature, Size: 6179 bytes --] ^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [RFC PATCH v2 00/11] KVM: Allow alternative providers of guest_memfd backed by PFNMAP memory 2026-07-23 9:40 ` David Woodhouse @ 2026-07-23 16:01 ` Ackerley Tng 0 siblings, 0 replies; 28+ messages in thread From: Ackerley Tng @ 2026-07-23 16:01 UTC (permalink / raw) To: David Woodhouse, Sean Christopherson, Paolo Bonzini, kvm, Xu Yilun Cc: linux-kernel, iommu, linux-kselftest, linux-media, dri-devel, linaro-mm-sig, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, x86, Jason Gunthorpe, Kevin Tian, Joerg Roedel, Will Deacon, Robin Murphy, Shuah Khan, Sumit Semwal, Christian König, Tom Lendacky, Tycho Andersen, Kees Cook, Ashish Kalra, Connor Williamson, Andrew Morton, Zhou Yuhang, Zi Li, Dan Williams, Eric Biggers, Colin Ian King, connordw, graf, fgriffo David Woodhouse <dwmw2@infradead.org> writes: > On Wed, 2026-07-22 at 17:24 -0700, Ackerley Tng wrote: >> David Woodhouse <dwmw2@infradead.org> writes: >> >> > For dedicated hosting environments, there are a few reasons to prefer >> > using external PFNMAP memory for guests, rather than kernel-managed >> > memory: >> > >> > • No struct page overheads >> > >> > • Easily managed in 1GiB pages (or larger chunks); no fragmentation/THP >> > >> > • Faster kexec/KHO live update (every millisecond spent faffing around >> > with memory management is an extra millisecond of steal time taken >> > from guests) >> > >> >> Also, no need to split/merge pages when converting from private to >> shared if the host mappings can't depend on the refcounts in struct >> pages :). >> >> > Today, there is only one implementation of guest_memfd: the internal >> > shmem-based page cache in virt/kvm/guest_memfd.c. This series allows >> > for other code to provide "a guest_memfd". This is an early path- >> > finding proof of concept tested with AMD SEV-SNP as well as non-CoCo >> > guests on x86 and arm64. I'm planning to build a memory-based file >> > system which gives files for both guest_memfd as well as memory-based >> > storage to allow things like userspace VM and orchestrator state >> > to be passed over KHO/kexec. >> > >> >> Thanks for raising this, good to know that there is more interest! >> >> Vishal and I have been working on guest_memfd with HugeTLB for a while >> now. I'll be restarting upstream work on HugeTLB after the conversions >> series merges. >> >> Vishal, Frank and I brought up having guest_memfd with HugeTLB with >> VM_PFNMAP mappings at a guest_memfd biweekly meeting a while ago. The >> consensus was that guest_memfd HugeTLB must support GUP and so VM_PFNMAP >> can't be the only way that guest_memfd supports HugeTLB. >> >> Your proposal is true non-page-struct memory, which has been on the >> discussion guest_memfd biweekly discussion backlog for a while >> now. Google is definitely also interested in this. >> >> Frank also described some use cases [1] other than HugeTLB and the >> regular PAGE_SIZE pages from the buddy that is currently supported. >> >> I submitted a proposal for LPC 2026 to discuss the interface. What I had >> in mind was to have the KVM_GUEST_MEMFD_CREATE ioctl take a "template" >> fd. guest_memfd will then read off what it needs from the "template" fd >> (ops, like you have here), and the "template" fd could be VFIO [2] or >> HugeTLB. >> >> You have a provider_fd = open("/dev/gmem_provider"), how about passing >> that provider_fd to guest_memfd, where /dev/gmem_provider could >> represent the page-less memory, like: >> >> ioctl(kvm_fd, KVM_GUEST_MEMFD_CREATE, provider_fd); >> >> It would be cool if the fallback could be set up flexibly, like: >> >> ioctl(kvm_fd, KVM_GUEST_MEMFD_CREATE, >> [provider_fd, fallback_provider_fd, ...]); >> >> I think guest_memfd should be the wrapper around other memory providers >> because then the CoCo shared/private logic and special page table >> handling for both host and stage 2 page tables is centralized in >> guest_memfd, and also, KVM doesn't have to understand other, different >> kinds of fds other than guest_memfd. > > Yeah. My approach was to make those "other kinds of fds" essentially > the same — they are "a guest_memfd", and KVM didn't have to understand > anything else (after a little bit of rework to have the generic ops so > that other code can create its own guest_memfds). As opposed to the > approach in Xu Yilun's TDISP series, which did teach KVM about a new > kind of memslot. > > Likewise, I used dma-buf so that IOMMUFD didn't have to understand > anything new either. > > I like your approach, inverting the wrapping of guest_memfd/dma-buf. > The provider then gives a plain dma-buf (not associated with any KVM as > it is in my current code). Userspace can use that *directly* with the > IOMMU, A wrinkle we will have to iron out is kind of the opposite problem, to make sure that while guest_memfd thinks a range of offsets (PAGE_SIZE or larger) is CoCo-private, the PFNs in that range must not be mapped anywhere else in the host, or IOMMU. One of the options I was thinking of was to seal off the template fd, perhaps with an inode flag like S_IMMUTABLE, but not even allow reads, and unlike S_IMMUTABLE, the (new) flag must not be changeable by userspace. Another option would be to just close the template fd so it'll not be used, and the PFNs behind that template fd would just be completely in guest_memfd ownership. All non-guest_memfd paths to access the PFNs must be closed off, or perhaps more precisely, any path that is not checked for CoCo-shared/private status before building page table entries must be closed off. For CoCo-shared memory as tracked by guest_memfd (or non-CoCo, which would use SHARED), the PFNs can still be mmap()-ed and actually set up in userspace page tables. Then, to use CoCo-shared memory for IO, IOMMUFD would also have to take the guest_memfd (which IOMMUFD probably has to do, for CoCo-private memory for Confidential IO, since CoCo-private memory can't be mapped to userspace). Perhaps through IOMMU_IOAS_MAP_FILE? In the process of mapping, IOMMUFD would tell guest_memfd "I have some of your PFNs mapped.", and when guest_memfd gets a private->shared conversion request, guest_memfd would get IOMMUFD to drop the mappings. (perhaps some kind of MMU notifier setup) Do you plan to use this for non-CoCo? > and passes it to your KVM_GUEST_MEMFD_CREATE in order to get a > guest_memfd for use with a given KVM. I think a lot of the warts in my > series fall away if we do it that way round. > > I'll have a go at reworking it that way round... and see what different > warts it ends up with. The interesting part will be getting actual PFNs > out of the dma-buf (which is already under active discussion). > If we go with the above, IOMMUFD would get PFNs from guest_memfd instead of dma-buf. > I'm not sure I understand the 'fallback provider' part. If one provider > doesn't have a page for a given address, KVM would iterate over them > all until it finds one that does? If each provider provides a .get_pfn() op, guest_memfd's kvm_gmem_get_pfn() would try the first provider's .get_pfn() and if that errors out, try the next one. guest_memfd today uses a standard struct address_space to track the pages that it owns. guest_memfd will have to use something else that can track both pages and PFNs, and probably which provider it came from. The more you probe me, the more complicated it's sounding. :) Still worth pursuing if it's good to have though. > Is that about falling back from one > THP pool to another? Is this something that the generic KVM code needs > to know about, or could it be handled *within* a provider, which can > chain to the next pool if it can't find a page? Handling that within a provider could be nice too, so instead of guest_memfd trying .get_pfn()s in series, the provider will try a hard-coded pair/list of PFN sources and then deal with the complexity on behalf guest_memfd? The flipside is for every combination of providers, the owner of the virtualization service would have to maintain a custom kernel module for the managing pages? Maybe the flipside isn't too bad, actually, to provide more flexibility. > Wouldn't we want that > done transparently for both KVM and IOMMU consumers? > >> >> [1] https://lore.kernel.org/all/CAPTztWajm_JLpp9BjRcX=h72r25ELrXeGkOXVachybBxLJGS=g@mail.gmail.com/ >> [2] https://lore.kernel.org/all/CAEvNRgFpJWQ5M5sQhGpQUV3GbBq9N+MQhhaxdxa=D8ky94SCsw@mail.gmail.com/ >> >> > This series lifts guest_memfd's operations into a small ops struct and >> > converts KVM's own gmem to use it, so it becomes one implementation >> > among peers rather than the only game in town. It also adapts the >> > SEV-SNP fast paths to page-less operation, plumbs iommufd to map a >> > guest_memfd-backing dma-buf as CPU RAM (not MMIO), and adds the >> > provider->KVM revocation primitive that makes overcommit-style reclaim >> > possible. >> > >> > It's largely AI-built at the moment. My main design concerns are around >> > how we tell that a given file is "a guest_memfd", and how we plug into >> > IOMMUFD. Presenting a dma-buf was the less intrusive choice because that >> > path already has support for taking pages away, but I'm less convinced >> > that it's the cleaner choice in the long term. >> > >> >> IIUC from this series, you're creating a dmabuf fd out of a guest_memfd, >> what does "taking pages away" mean, in this case, from where/who? > > Taken away from the guest. Ballooning pages out, either for traditional > memory overcommit, or for the models where a guest donates some of its > pages to the system to run a "child" which is properly isolated from > the guest. > > The physical page is reclaimed and needs to be taken out of the KVM and > IOMMU stage2 mappings. Will the VMM be doing a PUNCH_HOLE on the dma-buf fd or guest_memfd to actually reclaim the page, or do you have something else in mind? > If KVM wants to touch it again it must call the > provider to get a new PFN for that range. This is roughly equivalent to > the MMU notifier path on userspace VMAs in the normal memslot case. > IOMMUFD has no equivalent of that for userspace addresses — > IOMMU_IOAS_MAP pins the pages so they *can't* go away, which is why I > used dma-buf for the IOMMUFD interface. > > I've actually been fleshing out the overcommit path a bit more, > experimenting with adding async-pf support to guest_memfd in > https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=894dd0fcb34 > (the sample provider just pretends its pages are slow to populate, so > the guest gets to run something else while a "reclaimed" page comes > back). ^ permalink raw reply [flat|nested] 28+ messages in thread
* [RFC PATCH 0/6] KVM: guest_memfd: back guest_memfd with an imported dma-buf 2026-07-20 11:03 [RFC PATCH v2 00/11] KVM: Allow alternative providers of guest_memfd backed by PFNMAP memory David Woodhouse ` (2 preceding siblings ...) 2026-07-23 0:24 ` Ackerley Tng @ 2026-10-05 9:55 ` Fred Griffoul 2026-10-05 9:55 ` [RFC PATCH 1/6] KVM: guest_memfd: Add a writable result to get_pfn() Fred Griffoul ` (5 more replies) 3 siblings, 6 replies; 28+ messages in thread From: Fred Griffoul @ 2026-10-05 9:55 UTC (permalink / raw) To: Paolo Bonzini, Sean Christopherson, Marc Zyngier, Oliver Upton, Sumit Semwal, Christian König, Jason Gunthorpe, Kevin Tian Cc: David Woodhouse, Ackerley Tng, Joey Gouly, Suzuki K Poulose, Zenghui Yu, Steffen Eiden, Catalin Marinas, Will Deacon, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, Joerg Roedel, Robin Murphy, Alex Williamson, Shuah Khan, Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers, linux-kernel, kvm, kvmarm, linux-arm-kernel, iommu, linux-media, dri-devel, linaro-mm-sig, linux-kselftest, linux-trace-kernel, x86 From: Fred Griffoul <fgriffo@amazon.co.uk> This extends David Woodhouse's "KVM: Allow alternative providers of guest_memfd backed by PFNMAP memory" RFC (v2) to let guest_memfd import a dma-buf as its page source: https://lore.kernel.org/kvm/20260720111259.122911-1-dwmw2@infradead.org/ The memory owner is a plain dma-buf exporter, linking against neither KVM nor iommufd. guest_memfd imports it for the guest (stage-2) plane and iommufd imports the same buffer for the device (DMA) plane, so one object and one invalidation protocol serve both. This is a concrete answer to the "how do we plug into iommufd without masquerading as a dma-buf" question from David's cover: the provider does not pretend to be anything; the dma-buf is the shared object. Applies on top of David's v2 series. Patches: 1 KVM: guest_memfd: Add a writable result to get_pfn() Give get_pfn() a writable out-param so a backing can map a page present but read-only (KVM_MEM_READONLY is rejected on guest_memfd slots). A guest write then exits as a memory fault. Stage-2 only; KVM's own host-mapped accesses are unaffected. 2 dma-buf: Add get_phys() to describe a physical run New exporter op: for an offset+len, return one physically contiguous backed run and an attribute word (RAM/MMIO, plus READONLY). -ENOENT for an unbacked byte, -ENODEV when revoked. Lets a page-table-building importer ask where the memory is instead of taking a scatterlist. Converts vfio-pci, the iommufd selftest exporter and the sample; iommufd behaviour is unchanged. 3 dma-buf: Add ranged mapping invalidation Invalidate a byte range rather than the whole buffer. The callback means "re-query this range" and covers loss, replacement and read-only flips. Importers without the callback still get a whole-buffer invalidation. 4 dma-buf: Allow dynamic attach without a device Allow a dynamic importer with no struct device, for a consumer (KVM) that only needs the physical layout and builds its own tables. Such an importer must not map the attachment for DMA. 5 KVM: guest_memfd: Add dma-buf backing Add GUEST_MEMFD_FLAG_USE_DMABUF and a dmabuf_fd to KVM_CREATE_GUEST_MEMFD. guest_memfd attaches as a revocable importer without a device, serves faults from get_phys() (largest aligned RAM run, read-only honoured), and zaps the range on the invalidation callback. 6 samples/kvm, selftests/kvm: Exercise dma-buf backing Turn the sample into a dma-buf exporter (a toy memory owner: one root region, a child per VM, move/donate/reclaim, absent and read-only pages, a scratch page, per-child ioctl allowlists) and move the selftests to feed one dma-buf fd to both guest_memfd and iommufd. Scope: non-confidential VMs, RAM backings only; iommufd keeps its single-range import for now. As with David's series, this is largely AI-assisted and is posted to get the interface shape reviewed. Fred Griffoul (6): KVM: guest_memfd: Add a writable result to get_pfn() dma-buf: Add get_phys() to describe a physical run dma-buf: Add ranged mapping invalidation dma-buf: Allow dynamic attach without a device KVM: guest_memfd: Add dma-buf backing samples/kvm, selftests/kvm: Exercise dma-buf backing arch/arm64/kvm/mmu.c | 13 +- arch/arm64/kvm/nested.c | 10 +- arch/x86/kvm/mmu/mmu.c | 18 +- arch/x86/kvm/svm/sev.c | 4 +- drivers/dma-buf/dma-buf.c | 92 +- drivers/iommu/iommufd/iommufd_private.h | 8 - drivers/iommu/iommufd/iommufd_test.h | 16 + drivers/iommu/iommufd/pages.c | 74 +- drivers/iommu/iommufd/selftest.c | 106 +- drivers/vfio/pci/vfio_pci_dmabuf.c | 58 +- include/linux/dma-buf.h | 70 + include/linux/kvm_host.h | 12 +- include/linux/vfio_pci_core.h | 3 - include/trace/events/dma_buf.h | 2 +- include/uapi/linux/kvm.h | 6 +- samples/kvm/gmem_provider.c | 1311 ++++++++++------- samples/kvm/gmem_provider.h | 144 +- tools/include/uapi/linux/kvm.h | 6 +- tools/testing/selftests/kvm/Makefile.kvm | 3 +- .../kvm/gmem_provider_nvme_dma_test.c | 28 +- .../testing/selftests/kvm/include/kvm_util.h | 25 +- .../testing/selftests/kvm/x86/gmem_poc_test.c | 824 +++++++++++ .../kvm/x86/gmem_provider_hugepage_test.c | 30 +- .../kvm/x86/gmem_provider_iommufd_test.c | 42 +- .../kvm/x86/gmem_provider_readonly_test.c | 172 +++ .../kvm/x86/gmem_provider_revoke_test.c | 34 +- .../selftests/kvm/x86/gmem_provider_test.c | 190 +-- .../kvm/x86/gmem_provider_vfio_test.c | 45 +- virt/kvm/Kconfig | 1 + virt/kvm/guest_memfd.c | 252 +++- 30 files changed, 2703 insertions(+), 896 deletions(-) create mode 100644 tools/testing/selftests/kvm/x86/gmem_poc_test.c create mode 100644 tools/testing/selftests/kvm/x86/gmem_provider_readonly_test.c -- 2.47.3 ^ permalink raw reply [flat|nested] 28+ messages in thread
* [RFC PATCH 1/6] KVM: guest_memfd: Add a writable result to get_pfn() 2026-10-05 9:55 ` [RFC PATCH 0/6] KVM: guest_memfd: back guest_memfd with an imported dma-buf Fred Griffoul @ 2026-10-05 9:55 ` Fred Griffoul 2026-10-05 9:55 ` [RFC PATCH 2/6] dma-buf: Add get_phys() to describe a physical run Fred Griffoul ` (4 subsequent siblings) 5 siblings, 0 replies; 28+ messages in thread From: Fred Griffoul @ 2026-10-05 9:55 UTC (permalink / raw) To: Paolo Bonzini, Sean Christopherson, Marc Zyngier, Oliver Upton, Sumit Semwal, Christian König, Jason Gunthorpe, Kevin Tian Cc: David Woodhouse, Ackerley Tng, Joey Gouly, Suzuki K Poulose, Zenghui Yu, Steffen Eiden, Catalin Marinas, Will Deacon, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, Joerg Roedel, Robin Murphy, Alex Williamson, Shuah Khan, Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers, linux-kernel, kvm, kvmarm, linux-arm-kernel, iommu, linux-media, dri-devel, linaro-mm-sig, linux-kselftest, linux-trace-kernel, x86 From: Fred Griffoul <fgriffo@amazon.co.uk> A guest_memfd backing has no way to tell KVM that a page must stay mapped but must not be written by the guest. get_pfn() returns only a frame, and KVM takes write access from the memslot flags. KVM_MEM_READONLY is not an option: KVM refuses it on guest_memfd slots, because it emulates a write to a read-only slot as MMIO, which private memory cannot support. Add a writable output to get_pfn(). KVM sets it to true before the call, and the backing may clear it. When it is false, the x86 and arm64 stage-2 fault paths map the page without write permission. A guest write to that page then exits to userspace as a memory fault; KVM first releases any page that the backing returned. Callers that do not map the page pass NULL. The result applies to stage-2 mappings only. When KVM writes guest memory through the memslot's host address, the permissions of the host mapping apply. Signed-off-by: Fred Griffoul <fgriffo@amazon.co.uk> --- arch/arm64/kvm/mmu.c | 13 ++++++++----- arch/arm64/kvm/nested.c | 10 +++++++--- arch/x86/kvm/mmu/mmu.c | 18 ++++++++++++++++-- arch/x86/kvm/svm/sev.c | 4 ++-- include/linux/kvm_host.h | 11 ++++++++--- samples/kvm/gmem_provider.c | 3 ++- virt/kvm/guest_memfd.c | 13 ++++++++----- 7 files changed, 51 insertions(+), 21 deletions(-) diff --git a/arch/arm64/kvm/mmu.c b/arch/arm64/kvm/mmu.c index 6c941aaa10c6..32e591edc69d 100644 --- a/arch/arm64/kvm/mmu.c +++ b/arch/arm64/kvm/mmu.c @@ -1607,7 +1607,7 @@ struct kvm_s2_fault_desc { static int gmem_abort(const struct kvm_s2_fault_desc *s2fd) { - bool write_fault, exec_fault; + bool write_fault, exec_fault, writable; bool perm_fault = kvm_vcpu_trap_is_permission_fault(s2fd->vcpu); enum kvm_pgtable_walk_flags flags = KVM_PGTABLE_WALK_SHARED; enum kvm_pgtable_prot prot = KVM_PGTABLE_PROT_R; @@ -1641,14 +1641,17 @@ static int gmem_abort(const struct kvm_s2_fault_desc *s2fd) /* Pairs with the smp_wmb() in kvm_mmu_invalidate_end(). */ smp_rmb(); - ret = kvm_gmem_get_pfn(kvm, s2fd->memslot, gfn, &pfn, &page, NULL); - if (ret) { + ret = kvm_gmem_get_pfn(kvm, s2fd->memslot, gfn, &pfn, &page, NULL, + &writable); + if (ret || (write_fault && !writable)) { kvm_prepare_memory_fault_exit(s2fd->vcpu, s2fd->fault_ipa, PAGE_SIZE, write_fault, exec_fault, false); - return ret; + if (!ret) + kvm_release_faultin_page(kvm, page, true, false); + return ret ?: -EFAULT; } - if (!(s2fd->memslot->flags & KVM_MEM_READONLY)) + if (!(s2fd->memslot->flags & KVM_MEM_READONLY) && writable) prot |= KVM_PGTABLE_PROT_W; if (s2fd->nested) diff --git a/arch/arm64/kvm/nested.c b/arch/arm64/kvm/nested.c index fb54f6dad995..9ad1fa029835 100644 --- a/arch/arm64/kvm/nested.c +++ b/arch/arm64/kvm/nested.c @@ -1411,12 +1411,16 @@ static int kvm_translate_vncr(struct kvm_vcpu *vcpu, bool *is_gmem) if (is_error_noslot_pfn(pfn) || (write_fault && !writable)) return -EFAULT; } else { - ret = kvm_gmem_get_pfn(vcpu->kvm, memslot, gfn, &pfn, &page, NULL); - if (ret) { + ret = kvm_gmem_get_pfn(vcpu->kvm, memslot, gfn, &pfn, &page, NULL, + &writable); + if (ret || (write_fault && !writable)) { kvm_prepare_memory_fault_exit(vcpu, vt->wr.pa, PAGE_SIZE, write_fault, false, false); - return ret; + if (!ret) + kvm_release_faultin_page(vcpu->kvm, page, true, false); + return ret ?: -EFAULT; } + vt->wr.pw &= writable; } scoped_guard(write_lock, &vcpu->kvm->mmu_lock) { diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c index 234d0a95abf5..17ac2bfdc164 100644 --- a/arch/x86/kvm/mmu/mmu.c +++ b/arch/x86/kvm/mmu/mmu.c @@ -4612,6 +4612,7 @@ static void kvm_mmu_finish_page_fault(struct kvm_vcpu *vcpu, static int kvm_mmu_faultin_pfn_gmem(struct kvm_vcpu *vcpu, struct kvm_page_fault *fault) { + bool writable; int max_order, r; if (!kvm_slot_has_gmem(fault->slot)) { @@ -4620,13 +4621,26 @@ static int kvm_mmu_faultin_pfn_gmem(struct kvm_vcpu *vcpu, } r = kvm_gmem_get_pfn(vcpu->kvm, fault->slot, fault->gfn, &fault->pfn, - &fault->refcounted_page, &max_order); + &fault->refcounted_page, &max_order, &writable); if (r) { kvm_mmu_prepare_memory_fault_exit(vcpu, fault); return r; } - fault->map_writable = !(fault->slot->flags & KVM_MEM_READONLY); + /* + * The memory's owner has the final say on writability, on top of the + * memslot flag: a page it reports read-only is mapped read-only, and a + * guest write to it exits to userspace rather than being installed. + */ + fault->map_writable = !(fault->slot->flags & KVM_MEM_READONLY) && + writable; + if (fault->write && !fault->map_writable) { + kvm_mmu_prepare_memory_fault_exit(vcpu, fault); + kvm_release_faultin_page(vcpu->kvm, fault->refcounted_page, + true, false); + fault->refcounted_page = NULL; + return -EFAULT; + } fault->max_level = kvm_max_level_for_order(max_order); return RET_PF_CONTINUE; diff --git a/arch/x86/kvm/svm/sev.c b/arch/x86/kvm/svm/sev.c index 125779c82bc4..983e19f7dbd2 100644 --- a/arch/x86/kvm/svm/sev.c +++ b/arch/x86/kvm/svm/sev.c @@ -4063,7 +4063,7 @@ static void sev_snp_init_protected_guest_state(struct kvm_vcpu *vcpu) * The new VMSA will be private memory guest memory, so retrieve the * PFN from the gmem backend. */ - if (kvm_gmem_get_pfn(vcpu->kvm, slot, gfn, &pfn, &page, NULL)) + if (kvm_gmem_get_pfn(vcpu->kvm, slot, gfn, &pfn, &page, NULL, NULL)) return; /* @@ -4996,7 +4996,7 @@ void sev_handle_rmp_fault(struct kvm_vcpu *vcpu, gpa_t gpa, u64 error_code) return; } - ret = kvm_gmem_get_pfn(kvm, slot, gfn, &pfn, &page, &order); + ret = kvm_gmem_get_pfn(kvm, slot, gfn, &pfn, &page, &order, NULL); if (ret) { pr_warn_ratelimited("SEV: Unexpected RMP fault, no backing page for private GPA 0x%llx\n", gpa); diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h index 04fa0cb126f6..f84f3ab44acf 100644 --- a/include/linux/kvm_host.h +++ b/include/linux/kvm_host.h @@ -660,9 +660,14 @@ struct kvm_gmem_ops { struct kvm_memory_slot *slot, loff_t offset); void (*unbind)(struct file *file, struct kvm *kvm, struct kvm_memory_slot *slot); + /* + * @writable: [out] clear to have KVM map the page read-only; a guest + * write then exits as a memory fault. NULL if the caller does not care. + */ int (*get_pfn)(struct file *file, struct kvm *kvm, struct kvm_memory_slot *slot, gfn_t gfn, - kvm_pfn_t *pfn, struct page **page, int *max_order); + kvm_pfn_t *pfn, struct page **page, int *max_order, + bool *writable); int (*populate)(struct file *file, struct kvm *kvm, struct kvm_memory_slot *slot, gfn_t gfn, kvm_pfn_t *pfn, struct page *src_page, int order); @@ -2651,12 +2656,12 @@ static inline bool kvm_mem_is_private(struct kvm *kvm, gfn_t gfn) #ifdef CONFIG_KVM_GUEST_MEMFD int kvm_gmem_get_pfn(struct kvm *kvm, struct kvm_memory_slot *slot, gfn_t gfn, kvm_pfn_t *pfn, struct page **page, - int *max_order); + int *max_order, bool *writable); #else static inline int kvm_gmem_get_pfn(struct kvm *kvm, struct kvm_memory_slot *slot, gfn_t gfn, kvm_pfn_t *pfn, struct page **page, - int *max_order) + int *max_order, bool *writable) { KVM_BUG_ON(1, kvm); return -EIO; diff --git a/samples/kvm/gmem_provider.c b/samples/kvm/gmem_provider.c index 9728f5a8029b..75197c088762 100644 --- a/samples/kvm/gmem_provider.c +++ b/samples/kvm/gmem_provider.c @@ -157,7 +157,8 @@ static int gmem_max_order(struct gmem_info *info, gfn_t gfn, unsigned long index static int gmem_get_pfn(struct file *file, struct kvm *kvm, struct kvm_memory_slot *slot, gfn_t gfn, - kvm_pfn_t *pfn, struct page **page, int *max_order) + kvm_pfn_t *pfn, struct page **page, int *max_order, + bool *writable) { struct gmem_info *info = to_gmem_info(file); pgoff_t index = gfn - slot->base_gfn + slot->gmem.pgoff; diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c index d284bb70fe05..0f4bf2cc5e8e 100644 --- a/virt/kvm/guest_memfd.c +++ b/virt/kvm/guest_memfd.c @@ -624,7 +624,7 @@ static void kvm_gmem_native_unbind(struct file *slot_file, struct kvm *kvm, static int kvm_gmem_native_get_pfn(struct file *file, struct kvm *kvm, struct kvm_memory_slot *slot, gfn_t gfn, kvm_pfn_t *pfn, struct page **page, - int *max_order); + int *max_order, bool *writable); static void kvm_gmem_native_release(struct file *file); static int kvm_gmem_native_mmap(struct file *file, struct vm_area_struct *vma); @@ -988,7 +988,7 @@ static struct folio *__kvm_gmem_get_pfn(struct file *file, static int kvm_gmem_native_get_pfn(struct file *file, struct kvm *kvm, struct kvm_memory_slot *slot, gfn_t gfn, kvm_pfn_t *pfn, struct page **page, - int *max_order) + int *max_order, bool *writable) { pgoff_t index = kvm_gmem_get_index(slot, gfn); struct folio *folio; @@ -1016,7 +1016,7 @@ static int kvm_gmem_native_get_pfn(struct file *file, struct kvm *kvm, int kvm_gmem_get_pfn(struct kvm *kvm, struct kvm_memory_slot *slot, gfn_t gfn, kvm_pfn_t *pfn, struct page **page, - int *max_order) + int *max_order, bool *writable) { const struct kvm_gmem_ops *ops; @@ -1029,7 +1029,10 @@ int kvm_gmem_get_pfn(struct kvm *kvm, struct kvm_memory_slot *slot, return -EFAULT; *page = NULL; - return ops->get_pfn(file, kvm, slot, gfn, pfn, page, max_order); + if (writable) + *writable = true; + return ops->get_pfn(file, kvm, slot, gfn, pfn, page, max_order, + writable); } EXPORT_SYMBOL_FOR_KVM_INTERNAL(kvm_gmem_get_pfn); @@ -1093,7 +1096,7 @@ static int kvm_gmem_populate_one(const struct kvm_gmem_ops *ops, src_page, 0); else ret = ops->get_pfn(file, kvm, slot, gfn, &pfn, - &ignored_page, NULL); + &ignored_page, NULL, NULL); if (ret) return ret; -- 2.47.3 ^ permalink raw reply [flat|nested] 28+ messages in thread
* [RFC PATCH 2/6] dma-buf: Add get_phys() to describe a physical run 2026-10-05 9:55 ` [RFC PATCH 0/6] KVM: guest_memfd: back guest_memfd with an imported dma-buf Fred Griffoul 2026-10-05 9:55 ` [RFC PATCH 1/6] KVM: guest_memfd: Add a writable result to get_pfn() Fred Griffoul @ 2026-10-05 9:55 ` Fred Griffoul 2026-10-05 10:07 ` Christian König 2026-10-05 9:55 ` [RFC PATCH 3/6] dma-buf: Add ranged mapping invalidation Fred Griffoul ` (3 subsequent siblings) 5 siblings, 1 reply; 28+ messages in thread From: Fred Griffoul @ 2026-10-05 9:55 UTC (permalink / raw) To: Paolo Bonzini, Sean Christopherson, Marc Zyngier, Oliver Upton, Sumit Semwal, Christian König, Jason Gunthorpe, Kevin Tian Cc: David Woodhouse, Ackerley Tng, Joey Gouly, Suzuki K Poulose, Zenghui Yu, Steffen Eiden, Catalin Marinas, Will Deacon, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, Joerg Roedel, Robin Murphy, Alex Williamson, Shuah Khan, Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers, linux-kernel, kvm, kvmarm, linux-arm-kernel, iommu, linux-media, dri-devel, linaro-mm-sig, linux-kselftest, linux-trace-kernel, x86 From: Fred Griffoul <fgriffo@amazon.co.uk> iommufd and KVM write physical addresses into their own page tables. To do so, they must ask the exporter which frames back an offset of a dma-buf, whether that memory is RAM or MMIO, and whether it may be written. Add a get_phys() operation. The importer passes an offset and a maximum length. The exporter reports one run: the frames that start at the offset, are backed and physically contiguous, and share one attribute word. The run never exceeds the length. The exporter may end it early, so importers must not assume that it is the longest possible run. get_phys() returns -ENOENT when the byte at the offset is not backed, and -ENODEV when the buffer is revoked. The caller holds the reservation, and either pins the attachment or handles revocation. A reported frame stays valid until an invalidation that covers it returns. The attribute word holds the memory type and a READONLY flag. Zero means writable RAM. Importers refuse unknown types, reserved bits and unknown flags, so attributes added later fail safely. Two flag bits are reserved: one for holes and one for confidential memory. Convert vfio-pci, the iommufd selftest exporter and the KVM sample. iommufd behaves as before: it maps a buffer only when one writable run covers all of it. Signed-off-by: Fred Griffoul <fgriffo@amazon.co.uk> --- drivers/dma-buf/dma-buf.c | 47 +++++++++++ drivers/iommu/iommufd/iommufd_private.h | 8 -- drivers/iommu/iommufd/iommufd_test.h | 16 ++++ drivers/iommu/iommufd/pages.c | 74 ++++------------- drivers/iommu/iommufd/selftest.c | 106 ++++++++++++++++++++---- drivers/vfio/pci/vfio_pci_dmabuf.c | 58 ++++++------- include/linux/dma-buf.h | 54 ++++++++++++ include/linux/vfio_pci_core.h | 3 - samples/kvm/gmem_provider.c | 39 ++++----- 9 files changed, 267 insertions(+), 138 deletions(-) diff --git a/drivers/dma-buf/dma-buf.c b/drivers/dma-buf/dma-buf.c index d504c636dc29..66b85d53ed22 100644 --- a/drivers/dma-buf/dma-buf.c +++ b/drivers/dma-buf/dma-buf.c @@ -1389,6 +1389,53 @@ void dma_buf_invalidate_mappings(struct dma_buf *dmabuf) } EXPORT_SYMBOL_NS_GPL(dma_buf_invalidate_mappings, "DMA_BUF"); +/** + * dma_buf_get_phys - describe the run that starts at an offset + * @attach: attachment to query + * @offset: first buffer byte to describe + * @len: maximum number of bytes to describe + * @phys: physical address and length of the run + * @attr: DMA_BUF_PHYS_ATTR_* word of the run + * + * A run is the longest stretch of backed bytes starting at @offset whose + * frames are physically contiguous and share one attribute word. On success + * *@phys starts at the byte at @offset and covers at most @len bytes; it may + * be shorter than the run. + * + * The dma-buf reservation must be held. The attachment must be pinned or have + * revocable importer operations. A frame remains valid until a covering + * invalidation callback returns; the exporter must invalidate a changed range + * before reusing its old frames. + * + * Returns: + * + * 0 on success, -ENOENT if the byte at @offset is not backed, -ENODEV if the + * buffer is revoked, -EOPNOTSUPP if the exporter cannot describe itself this + * way, or another negative error code. + */ +int dma_buf_get_phys(struct dma_buf_attachment *attach, u64 offset, u64 len, + struct phys_vec *phys, u32 *attr) +{ + u64 end; + int ret; + + if (WARN_ON_ONCE(!attach || !attach->dmabuf || !phys || !attr)) + return -EINVAL; + if (!len || check_add_overflow(offset, len, &end) || + end > attach->dmabuf->size) + return -EINVAL; + + dma_resv_assert_held(attach->dmabuf->resv); + if (!attach->dmabuf->ops->get_phys) + return -EOPNOTSUPP; + + ret = attach->dmabuf->ops->get_phys(attach, offset, len, phys, attr); + if (!ret && WARN_ON_ONCE(!phys->len || phys->len > len)) + return -EIO; + return ret; +} +EXPORT_SYMBOL_NS_GPL(dma_buf_get_phys, "DMA_BUF"); + /** * DOC: cpu access * diff --git a/drivers/iommu/iommufd/iommufd_private.h b/drivers/iommu/iommufd/iommufd_private.h index 43fbc5bed8de..5cded585c227 100644 --- a/drivers/iommu/iommufd/iommufd_private.h +++ b/drivers/iommu/iommufd/iommufd_private.h @@ -716,8 +716,6 @@ bool iommufd_should_fail(void); int __init iommufd_test_init(void); void iommufd_test_exit(void); bool iommufd_selftest_is_mock_dev(struct device *dev); -int iommufd_test_dma_buf_iommufd_map(struct dma_buf_attachment *attachment, - struct phys_vec *phys); #else static inline void iommufd_test_syz_conv_iova_id(struct iommufd_ucmd *ucmd, unsigned int ioas_id, @@ -739,11 +737,5 @@ static inline bool iommufd_selftest_is_mock_dev(struct device *dev) { return false; } -static inline int -iommufd_test_dma_buf_iommufd_map(struct dma_buf_attachment *attachment, - struct phys_vec *phys) -{ - return -EOPNOTSUPP; -} #endif #endif diff --git a/drivers/iommu/iommufd/iommufd_test.h b/drivers/iommu/iommufd/iommufd_test.h index 52b78cbcc920..28fd9c43edc4 100644 --- a/drivers/iommu/iommufd/iommufd_test.h +++ b/drivers/iommu/iommufd/iommufd_test.h @@ -31,6 +31,8 @@ enum { IOMMU_TEST_OP_PASID_CHECK_HWPT, IOMMU_TEST_OP_DMABUF_GET, IOMMU_TEST_OP_DMABUF_REVOKE, + IOMMU_TEST_OP_MD_CHECK_MAPPED, + IOMMU_TEST_OP_MD_IOVA_TO_PHYS, }; enum { @@ -193,6 +195,20 @@ struct iommu_test_cmd { __s32 dmabuf_fd; __u32 revoked; } dmabuf_revoke; + struct { + /* + * 1: every page in [iova, iova+length) must be mapped; + * 0: none of them may be. Mixed is an error. + */ + __u32 mapped; + __u32 __reserved; + __aligned_u64 iova; + __aligned_u64 length; + } check_mapped; + struct { + __aligned_u64 iova; + __aligned_u64 out_phys; /* 0 if unmapped */ + } iova_to_phys; }; __u32 last; }; diff --git a/drivers/iommu/iommufd/pages.c b/drivers/iommu/iommufd/pages.c index f9b2ae6d7e96..196d1bb330c2 100644 --- a/drivers/iommu/iommufd/pages.c +++ b/drivers/iommu/iommufd/pages.c @@ -1463,68 +1463,12 @@ static const struct dma_buf_attach_ops iopt_dmabuf_attach_revoke_ops = { .invalidate_mappings = iopt_revoke_notify, }; -/* - * iommufd and vfio have a circular dependency. Future work for a phys - * based private interconnect will remove this. - */ -/* - * Look up the exporter's phys accessor for iommufd's private-interconnect - * path. Also fills *is_cpu_ram: true if the exporter's memory is normal - * cache-coherent RAM (needs BATCH_CPU_MEMORY / IOMMU_CACHE), false for MMIO - * (needs BATCH_MMIO / IOMMU_MMIO). This will be replaced by a formal - * exporter op that returns phys + memory type together. - */ -static int -sym_vfio_pci_dma_buf_iommufd_map(struct dma_buf_attachment *attachment, - struct phys_vec *phys, bool *is_cpu_ram) -{ - typeof(&vfio_pci_dma_buf_iommufd_map) fn; - int rc; - - rc = iommufd_test_dma_buf_iommufd_map(attachment, phys); - if (rc != -EOPNOTSUPP) { - *is_cpu_ram = false; /* test hook mimics VFIO MMIO */ - return rc; - } - - /* - * Prototype: try the sample gmem provider's dma-buf exporter. This - * mirrors the vfio-pci private-interconnect hook, and (like it) is - * meant to be replaced by a formal negotiated exporter op returning - * phys + memory type. The provider serves RAM, so mark it CPU_RAM. - */ - { - extern int gmem_provider_dma_buf_iommufd_map( - struct dma_buf_attachment *, struct phys_vec *); - typeof(&gmem_provider_dma_buf_iommufd_map) gfn; - - gfn = symbol_get(gmem_provider_dma_buf_iommufd_map); - if (gfn) { - rc = gfn(attachment, phys); - symbol_put(gmem_provider_dma_buf_iommufd_map); - if (rc != -EOPNOTSUPP) { - *is_cpu_ram = true; - return rc; - } - } - } - - if (!IS_ENABLED(CONFIG_VFIO_PCI_DMABUF)) - return -EOPNOTSUPP; - - fn = symbol_get(vfio_pci_dma_buf_iommufd_map); - if (!fn) - return -EOPNOTSUPP; - rc = fn(attachment, phys); - symbol_put(vfio_pci_dma_buf_iommufd_map); - *is_cpu_ram = false; /* VFIO PCI dma-buf carries BAR (MMIO) memory */ - return rc; -} - static int iopt_map_dmabuf(struct iommufd_ctx *ictx, struct iopt_pages *pages, struct dma_buf *dmabuf) { struct dma_buf_attachment *attach; + struct phys_vec pv; + u32 attr; int rc; attach = dma_buf_dynamic_attach(dmabuf, iommufd_global_device(), @@ -1546,10 +1490,20 @@ static int iopt_map_dmabuf(struct iommufd_ctx *ictx, struct iopt_pages *pages, if (rc) goto err_detach; - rc = sym_vfio_pci_dma_buf_iommufd_map(attach, &pages->dmabuf.phys, - &pages->dmabuf.is_cpu_ram); + /* One backed, writable run covering the buffer: refuse the rest. */ + rc = dma_buf_get_phys(attach, 0, dmabuf->size, &pv, &attr); + if (rc == -ENOENT) + rc = -EOPNOTSUPP; if (rc) goto err_unpin; + if (pv.len != dmabuf->size || !dma_buf_phys_attr_known(attr) || + (attr & DMA_BUF_PHYS_ATTR_FLAGS_MASK)) { + rc = -EOPNOTSUPP; + goto err_unpin; + } + pages->dmabuf.phys = pv; + pages->dmabuf.is_cpu_ram = + dma_buf_phys_attr_type(attr) == DMA_BUF_PHYS_ATTR_RAM; dma_resv_unlock(dmabuf->resv); diff --git a/drivers/iommu/iommufd/selftest.c b/drivers/iommu/iommufd/selftest.c index af07c642a526..0899272d1e66 100644 --- a/drivers/iommu/iommufd/selftest.c +++ b/drivers/iommu/iommufd/selftest.c @@ -1962,32 +1962,31 @@ static void iommufd_test_dma_buf_release(struct dma_buf *dmabuf) kfree(priv); } -static const struct dma_buf_ops iommufd_test_dmabuf_ops = { - .attach = iommufd_test_dma_buf_attach, - .detach = iommufd_test_dma_buf_detach, - .map_dma_buf = iommufd_test_dma_buf_map, - .release = iommufd_test_dma_buf_release, - .unmap_dma_buf = iommufd_test_dma_buf_unmap, -}; - -int iommufd_test_dma_buf_iommufd_map(struct dma_buf_attachment *attachment, - struct phys_vec *phys) +static int iommufd_test_dma_buf_get_phys(struct dma_buf_attachment *attachment, + u64 offset, u64 len, + struct phys_vec *phys, u32 *attr) { struct iommufd_test_dma_buf *priv = attachment->dmabuf->priv; dma_resv_assert_held(attachment->dmabuf->resv); - - if (attachment->dmabuf->ops != &iommufd_test_dmabuf_ops) - return -EOPNOTSUPP; - if (priv->revoked) return -ENODEV; - phys->paddr = virt_to_phys(priv->memory); - phys->len = priv->length; + phys->paddr = virt_to_phys(priv->memory) + offset; + phys->len = len; + *attr = DMA_BUF_PHYS_ATTR_MMIO; return 0; } +static const struct dma_buf_ops iommufd_test_dmabuf_ops = { + .attach = iommufd_test_dma_buf_attach, + .detach = iommufd_test_dma_buf_detach, + .map_dma_buf = iommufd_test_dma_buf_map, + .release = iommufd_test_dma_buf_release, + .unmap_dma_buf = iommufd_test_dma_buf_unmap, + .get_phys = iommufd_test_dma_buf_get_phys, +}; + static int iommufd_test_dmabuf_get(struct iommufd_ucmd *ucmd, unsigned int open_flags, size_t len) @@ -2031,6 +2030,73 @@ static int iommufd_test_dmabuf_get(struct iommufd_ucmd *ucmd, return rc; } +/* + * Report the physical address the mock domain resolves @iova to, or 0 if + * it is unmapped. Lets a test check that two IOVAs share one frame (a + * scratch substitution) without knowing the frame in advance. + */ +static int iommufd_test_md_iova_to_phys(struct iommufd_ucmd *ucmd, + unsigned int mockpt_id, + unsigned long iova) +{ + struct iommu_test_cmd *cmd = ucmd->cmd; + struct iommufd_hw_pagetable *hwpt; + struct mock_iommu_domain *mock; + unsigned int page_size; + int rc; + + hwpt = get_md_pagetable(ucmd, mockpt_id, &mock); + if (IS_ERR(hwpt)) + return PTR_ERR(hwpt); + + page_size = 1 << __ffs(mock->domain.pgsize_bitmap); + if (iova % page_size) { + rc = -EINVAL; + goto out_put; + } + cmd->iova_to_phys.out_phys = + mock->domain.ops->iova_to_phys(&mock->domain, iova); + rc = iommufd_ucmd_respond(ucmd, sizeof(*cmd)); +out_put: + iommufd_put_object(ucmd->ictx, &hwpt->obj); + return rc; +} + +static int iommufd_test_md_check_mapped(struct iommufd_ucmd *ucmd, + unsigned int mockpt_id, + unsigned long iova, size_t length, + bool mapped) +{ + struct iommufd_hw_pagetable *hwpt; + struct mock_iommu_domain *mock; + unsigned int page_size; + int rc = 0; + + hwpt = get_md_pagetable(ucmd, mockpt_id, &mock); + if (IS_ERR(hwpt)) + return PTR_ERR(hwpt); + + page_size = 1 << __ffs(mock->domain.pgsize_bitmap); + if (iova % page_size || length % page_size || !length) { + rc = -EINVAL; + goto out_put; + } + + for (; length; length -= page_size, iova += page_size) { + bool is_mapped = + mock->domain.ops->iova_to_phys(&mock->domain, iova) != 0; + + if (is_mapped != mapped) { + rc = -ENOENT; + goto out_put; + } + } + +out_put: + iommufd_put_object(ucmd->ictx, &hwpt->obj); + return rc; +} + static int iommufd_test_dmabuf_revoke(struct iommufd_ucmd *ucmd, int fd, bool revoked) { @@ -2143,6 +2209,14 @@ int iommufd_test(struct iommufd_ucmd *ucmd) return iommufd_test_dmabuf_revoke(ucmd, cmd->dmabuf_revoke.dmabuf_fd, cmd->dmabuf_revoke.revoked); + case IOMMU_TEST_OP_MD_CHECK_MAPPED: + return iommufd_test_md_check_mapped(ucmd, cmd->id, + cmd->check_mapped.iova, + cmd->check_mapped.length, + cmd->check_mapped.mapped); + case IOMMU_TEST_OP_MD_IOVA_TO_PHYS: + return iommufd_test_md_iova_to_phys(ucmd, cmd->id, + cmd->iova_to_phys.iova); default: return -EOPNOTSUPP; } diff --git a/drivers/vfio/pci/vfio_pci_dmabuf.c b/drivers/vfio/pci/vfio_pci_dmabuf.c index c16f460c01d6..381c3d338e8c 100644 --- a/drivers/vfio/pci/vfio_pci_dmabuf.c +++ b/drivers/vfio/pci/vfio_pci_dmabuf.c @@ -99,46 +99,46 @@ static void vfio_pci_dma_buf_release(struct dma_buf *dmabuf) kfree(priv); } -static const struct dma_buf_ops vfio_pci_dmabuf_ops = { - .attach = vfio_pci_dma_buf_attach, - .map_dma_buf = vfio_pci_dma_buf_map, - .unmap_dma_buf = vfio_pci_dma_buf_unmap, - .release = vfio_pci_dma_buf_release, -}; - /* - * This is a temporary "private interconnect" between VFIO DMABUF and iommufd. - * It allows the two co-operating drivers to exchange the physical address of - * the BAR. This is to be replaced with a formal DMABUF system for negotiated - * interconnect types. + * Report the BAR's physical range for importers which program it into their own + * translation tables, such as iommufd. A BAR is MMIO, never cache-coherent RAM. * - * If this function succeeds the following are true: - * - There is one physical range and it is pointing to MMIO - * - When move_notify is called it means revoke, not move, vfio_dma_buf_map - * will fail if it is currently revoked + * When move_notify is called it means revoke, not move, so this fails while + * revoked and vfio_dma_buf_map() does the same. */ -int vfio_pci_dma_buf_iommufd_map(struct dma_buf_attachment *attachment, - struct phys_vec *phys) +static int vfio_pci_dma_buf_get_phys(struct dma_buf_attachment *attachment, + u64 offset, u64 len, + struct phys_vec *phys, u32 *attr) { - struct vfio_pci_dma_buf *priv; + struct vfio_pci_dma_buf *priv = attachment->dmabuf->priv; + u32 i; dma_resv_assert_held(attachment->dmabuf->resv); - - if (attachment->dmabuf->ops != &vfio_pci_dmabuf_ops) - return -EOPNOTSUPP; - - priv = attachment->dmabuf->priv; if (priv->revoked) return -ENODEV; - /* More than one range to iommufd will require proper DMABUF support */ - if (priv->nr_ranges != 1) - return -EOPNOTSUPP; - - *phys = priv->phys_vec[0]; + /* Report from @offset to the end of the BAR range containing it. */ + for (i = 0; i < priv->nr_ranges; i++) { + if (offset < priv->phys_vec[i].len) + break; + offset -= priv->phys_vec[i].len; + } + if (i == priv->nr_ranges) + return -EINVAL; + phys->paddr = priv->phys_vec[i].paddr + offset; + phys->len = min_t(u64, priv->phys_vec[i].len - offset, len); + *attr = DMA_BUF_PHYS_ATTR_MMIO; return 0; } -EXPORT_SYMBOL_FOR_MODULES(vfio_pci_dma_buf_iommufd_map, "iommufd"); + +static const struct dma_buf_ops vfio_pci_dmabuf_ops = { + .attach = vfio_pci_dma_buf_attach, + .map_dma_buf = vfio_pci_dma_buf_map, + .unmap_dma_buf = vfio_pci_dma_buf_unmap, + .release = vfio_pci_dma_buf_release, + .get_phys = vfio_pci_dma_buf_get_phys, +}; + int vfio_pci_core_fill_phys_vec(struct phys_vec *phys_vec, struct vfio_region_dma_range *dma_ranges, diff --git a/include/linux/dma-buf.h b/include/linux/dma-buf.h index d1203da56fc5..b223962e20c2 100644 --- a/include/linux/dma-buf.h +++ b/include/linux/dma-buf.h @@ -13,6 +13,7 @@ #ifndef __DMA_BUF_H__ #define __DMA_BUF_H__ +#include <linux/bitfield.h> #include <linux/iosys-map.h> #include <linux/file.h> #include <linux/err.h> @@ -23,6 +24,7 @@ #include <linux/dma-fence.h> #include <linux/wait.h> #include <linux/pci-p2pdma.h> +#include <linux/types.h> struct device; struct dma_buf; @@ -182,6 +184,29 @@ struct dma_buf_ops { struct sg_table *, enum dma_data_direction); + /** + * @get_phys: + * + * Describe the run that starts at @offset, for an importer that + * programs its own translation tables. A run is the longest stretch + * of backed bytes whose frames are physically contiguous and share + * one attribute word. Report it in *@phys, starting at the byte at + * @offset and clipped at @offset + @len, with its DMA_BUF_PHYS_ATTR_* + * word in *@attr. The exporter may stop before the end of the run; + * importers must not assume the reported run is maximal. + * + * Return 0 on success, -ENOENT if the byte at @offset is not backed, + * -ENODEV if the buffer is revoked, or another negative error. Do not + * wait for memory to become available. + * + * The dma-buf reservation is held. The attachment must be pinned or + * have revocable importer operations. A reported frame remains valid + * until a covering invalidation callback returns; an exporter must + * invalidate every change before reusing an old frame. + */ + int (*get_phys)(struct dma_buf_attachment *attach, u64 offset, u64 len, + struct phys_vec *phys, u32 *attr); + /* TODO: Add try_map_dma_buf version, to return immed with -EBUSY * if the call would block. */ @@ -576,6 +601,35 @@ void dma_buf_unmap_attachment(struct dma_buf_attachment *, struct sg_table *, enum dma_data_direction); void dma_buf_invalidate_mappings(struct dma_buf *dma_buf); bool dma_buf_attach_revocable(struct dma_buf_attachment *attach); +/* bits 0-7: memory type (a value, not flags) */ +#define DMA_BUF_PHYS_ATTR_TYPE_MASK GENMASK(7, 0) +#define DMA_BUF_PHYS_ATTR_RAM 0x00 /* cache-coherent system RAM */ +#define DMA_BUF_PHYS_ATTR_MMIO 0x01 /* device MMIO, uncached */ +/* bits 8-15: reserved for a second value field; must be zero */ +#define DMA_BUF_PHYS_ATTR_RSVD_MASK GENMASK(15, 8) +/* bits 16-31: flags; undefined bits must be zero */ +#define DMA_BUF_PHYS_ATTR_READONLY BIT(16) +/* BIT(17): reserved (hole, for importers that walk across gaps) */ +/* BIT(18): reserved (private, for confidential computing) */ +#define DMA_BUF_PHYS_ATTR_FLAGS_MASK (DMA_BUF_PHYS_ATTR_READONLY) + +static inline u32 dma_buf_phys_attr_type(u32 attrs) +{ + return FIELD_GET(DMA_BUF_PHYS_ATTR_TYPE_MASK, attrs); +} + +static inline bool dma_buf_phys_attr_known(u32 attrs) +{ + return dma_buf_phys_attr_type(attrs) <= DMA_BUF_PHYS_ATTR_MMIO && + !(attrs & DMA_BUF_PHYS_ATTR_RSVD_MASK) && + !(attrs & ~(DMA_BUF_PHYS_ATTR_TYPE_MASK | + DMA_BUF_PHYS_ATTR_RSVD_MASK | + DMA_BUF_PHYS_ATTR_FLAGS_MASK)); +} + +int dma_buf_get_phys(struct dma_buf_attachment *attach, u64 offset, u64 len, + struct phys_vec *phys, u32 *attr); + int dma_buf_begin_cpu_access(struct dma_buf *dma_buf, enum dma_data_direction dir); int dma_buf_end_cpu_access(struct dma_buf *dma_buf, diff --git a/include/linux/vfio_pci_core.h b/include/linux/vfio_pci_core.h index 9a1674c152aa..2a1d13abdb0c 100644 --- a/include/linux/vfio_pci_core.h +++ b/include/linux/vfio_pci_core.h @@ -257,7 +257,4 @@ vfio_pci_core_get_iomap(struct vfio_pci_core_device *vdev, unsigned int bar) return vdev->barmap[bar]; } -int vfio_pci_dma_buf_iommufd_map(struct dma_buf_attachment *attachment, - struct phys_vec *phys); - #endif /* VFIO_PCI_CORE_H */ diff --git a/samples/kvm/gmem_provider.c b/samples/kvm/gmem_provider.c index 75197c088762..b6824fe5d228 100644 --- a/samples/kvm/gmem_provider.c +++ b/samples/kvm/gmem_provider.c @@ -487,35 +487,30 @@ static void gmem_dma_buf_release(struct dma_buf *dmabuf) kfree(priv); } -static const struct dma_buf_ops gmem_dma_buf_ops = { - .attach = gmem_dma_buf_attach, - .map_dma_buf = gmem_dma_buf_map, - .unmap_dma_buf = gmem_dma_buf_unmap, - .release = gmem_dma_buf_release, -}; - -/* - * Private interconnect for iommufd (mirrors vfio_pci_dma_buf_iommufd_map). - * Returns the single contiguous phys range for the exported region so iommufd - * can program the IOMMU directly, bypassing the DMA API. - */ -int gmem_provider_dma_buf_iommufd_map(struct dma_buf_attachment *attach, - struct phys_vec *phys); -int gmem_provider_dma_buf_iommufd_map(struct dma_buf_attachment *attach, - struct phys_vec *phys) +/* Report this flat sample region through the generic dma-buf operation. */ +static int gmem_dma_buf_get_phys(struct dma_buf_attachment *attach, + u64 offset, u64 len, + struct phys_vec *phys, u32 *attr) { - struct gmem_dmabuf *priv; + struct gmem_dmabuf *priv = attach->dmabuf->priv; dma_resv_assert_held(attach->dmabuf->resv); - if (attach->dmabuf->ops != &gmem_dma_buf_ops) - return -EOPNOTSUPP; - priv = attach->dmabuf->priv; if (priv->revoked) return -ENODEV; - *phys = priv->phys; + + phys->paddr = priv->phys.paddr + offset; + phys->len = len; + *attr = DMA_BUF_PHYS_ATTR_RAM; return 0; } -EXPORT_SYMBOL_FOR_MODULES(gmem_provider_dma_buf_iommufd_map, "iommufd"); + +static const struct dma_buf_ops gmem_dma_buf_ops = { + .attach = gmem_dma_buf_attach, + .map_dma_buf = gmem_dma_buf_map, + .unmap_dma_buf = gmem_dma_buf_unmap, + .release = gmem_dma_buf_release, + .get_phys = gmem_dma_buf_get_phys, +}; /* Called with info->dmabufs_lock held on the revoke path. */ static void gmem_dma_buf_revoke_all(struct gmem_info *info) -- 2.47.3 ^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [RFC PATCH 2/6] dma-buf: Add get_phys() to describe a physical run 2026-10-05 9:55 ` [RFC PATCH 2/6] dma-buf: Add get_phys() to describe a physical run Fred Griffoul @ 2026-10-05 10:07 ` Christian König 2026-10-05 13:20 ` Fred Griffoul 0 siblings, 1 reply; 28+ messages in thread From: Christian König @ 2026-10-05 10:07 UTC (permalink / raw) To: Fred Griffoul, Paolo Bonzini, Sean Christopherson, Marc Zyngier, Oliver Upton, Sumit Semwal, Jason Gunthorpe, Kevin Tian Cc: David Woodhouse, Ackerley Tng, Joey Gouly, Suzuki K Poulose, Zenghui Yu, Steffen Eiden, Catalin Marinas, Will Deacon, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, Joerg Roedel, Robin Murphy, Alex Williamson, Shuah Khan, Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers, linux-kernel, kvm, kvmarm, linux-arm-kernel, iommu, linux-media, dri-devel, linaro-mm-sig, linux-kselftest, linux-trace-kernel, x86 On 10/5/26 11:55, Fred Griffoul wrote: > From: Fred Griffoul <fgriffo@amazon.co.uk> > > iommufd and KVM write physical addresses into their own page tables. > To do so, they must ask the exporter which frames back an offset of a > dma-buf, whether that memory is RAM or MMIO, and whether it may be > written. Well filling page tables by the importer is an absolutely clear NO-GO for the DMA-buf design, we have gotten down that path already and it took us years to remove this functionality again. The problem you are facing here is that dma_buf_mmap() doesn't work because you don't have a VMA. I can understand the reasoning that you don't want to have a VMA, but I don't think that this is a valid justification to add complexity to DMA-buf and especially bring an approach back which we have already deprecated. A possible solution might be Jasons patch set to directly negotiate exposing PCI BAR regions as DMA-buf, but that certainly needs more discussion. But the approach outlined here is an absolutely clear NAK from my side. Regards, Christian. > > Add a get_phys() operation. The importer passes an offset and a maximum > length. The exporter reports one run: the frames that start at the > offset, are backed and physically contiguous, and share one attribute > word. The run never exceeds the length. The exporter may end it early, > so importers must not assume that it is the longest possible run. > > get_phys() returns -ENOENT when the byte at the offset is not backed, > and -ENODEV when the buffer is revoked. The caller holds the > reservation, and either pins the attachment or handles revocation. A > reported frame stays valid until an invalidation that covers it > returns. > > The attribute word holds the memory type and a READONLY flag. Zero > means writable RAM. Importers refuse unknown types, reserved bits and > unknown flags, so attributes added later fail safely. Two flag bits are > reserved: one for holes and one for confidential memory. > > Convert vfio-pci, the iommufd selftest exporter and the KVM sample. > iommufd behaves as before: it maps a buffer only when one writable run > covers all of it. > > Signed-off-by: Fred Griffoul <fgriffo@amazon.co.uk> > --- > drivers/dma-buf/dma-buf.c | 47 +++++++++++ > drivers/iommu/iommufd/iommufd_private.h | 8 -- > drivers/iommu/iommufd/iommufd_test.h | 16 ++++ > drivers/iommu/iommufd/pages.c | 74 ++++------------- > drivers/iommu/iommufd/selftest.c | 106 ++++++++++++++++++++---- > drivers/vfio/pci/vfio_pci_dmabuf.c | 58 ++++++------- > include/linux/dma-buf.h | 54 ++++++++++++ > include/linux/vfio_pci_core.h | 3 - > samples/kvm/gmem_provider.c | 39 ++++----- > 9 files changed, 267 insertions(+), 138 deletions(-) > > diff --git a/drivers/dma-buf/dma-buf.c b/drivers/dma-buf/dma-buf.c > index d504c636dc29..66b85d53ed22 100644 > --- a/drivers/dma-buf/dma-buf.c > +++ b/drivers/dma-buf/dma-buf.c > @@ -1389,6 +1389,53 @@ void dma_buf_invalidate_mappings(struct dma_buf *dmabuf) > } > EXPORT_SYMBOL_NS_GPL(dma_buf_invalidate_mappings, "DMA_BUF"); > > +/** > + * dma_buf_get_phys - describe the run that starts at an offset > + * @attach: attachment to query > + * @offset: first buffer byte to describe > + * @len: maximum number of bytes to describe > + * @phys: physical address and length of the run > + * @attr: DMA_BUF_PHYS_ATTR_* word of the run > + * > + * A run is the longest stretch of backed bytes starting at @offset whose > + * frames are physically contiguous and share one attribute word. On success > + * *@phys starts at the byte at @offset and covers at most @len bytes; it may > + * be shorter than the run. > + * > + * The dma-buf reservation must be held. The attachment must be pinned or have > + * revocable importer operations. A frame remains valid until a covering > + * invalidation callback returns; the exporter must invalidate a changed range > + * before reusing its old frames. > + * > + * Returns: > + * > + * 0 on success, -ENOENT if the byte at @offset is not backed, -ENODEV if the > + * buffer is revoked, -EOPNOTSUPP if the exporter cannot describe itself this > + * way, or another negative error code. > + */ > +int dma_buf_get_phys(struct dma_buf_attachment *attach, u64 offset, u64 len, > + struct phys_vec *phys, u32 *attr) > +{ > + u64 end; > + int ret; > + > + if (WARN_ON_ONCE(!attach || !attach->dmabuf || !phys || !attr)) > + return -EINVAL; > + if (!len || check_add_overflow(offset, len, &end) || > + end > attach->dmabuf->size) > + return -EINVAL; > + > + dma_resv_assert_held(attach->dmabuf->resv); > + if (!attach->dmabuf->ops->get_phys) > + return -EOPNOTSUPP; > + > + ret = attach->dmabuf->ops->get_phys(attach, offset, len, phys, attr); > + if (!ret && WARN_ON_ONCE(!phys->len || phys->len > len)) > + return -EIO; > + return ret; > +} > +EXPORT_SYMBOL_NS_GPL(dma_buf_get_phys, "DMA_BUF"); > + > /** > * DOC: cpu access > * > diff --git a/drivers/iommu/iommufd/iommufd_private.h b/drivers/iommu/iommufd/iommufd_private.h > index 43fbc5bed8de..5cded585c227 100644 > --- a/drivers/iommu/iommufd/iommufd_private.h > +++ b/drivers/iommu/iommufd/iommufd_private.h > @@ -716,8 +716,6 @@ bool iommufd_should_fail(void); > int __init iommufd_test_init(void); > void iommufd_test_exit(void); > bool iommufd_selftest_is_mock_dev(struct device *dev); > -int iommufd_test_dma_buf_iommufd_map(struct dma_buf_attachment *attachment, > - struct phys_vec *phys); > #else > static inline void iommufd_test_syz_conv_iova_id(struct iommufd_ucmd *ucmd, > unsigned int ioas_id, > @@ -739,11 +737,5 @@ static inline bool iommufd_selftest_is_mock_dev(struct device *dev) > { > return false; > } > -static inline int > -iommufd_test_dma_buf_iommufd_map(struct dma_buf_attachment *attachment, > - struct phys_vec *phys) > -{ > - return -EOPNOTSUPP; > -} > #endif > #endif > diff --git a/drivers/iommu/iommufd/iommufd_test.h b/drivers/iommu/iommufd/iommufd_test.h > index 52b78cbcc920..28fd9c43edc4 100644 > --- a/drivers/iommu/iommufd/iommufd_test.h > +++ b/drivers/iommu/iommufd/iommufd_test.h > @@ -31,6 +31,8 @@ enum { > IOMMU_TEST_OP_PASID_CHECK_HWPT, > IOMMU_TEST_OP_DMABUF_GET, > IOMMU_TEST_OP_DMABUF_REVOKE, > + IOMMU_TEST_OP_MD_CHECK_MAPPED, > + IOMMU_TEST_OP_MD_IOVA_TO_PHYS, > }; > > enum { > @@ -193,6 +195,20 @@ struct iommu_test_cmd { > __s32 dmabuf_fd; > __u32 revoked; > } dmabuf_revoke; > + struct { > + /* > + * 1: every page in [iova, iova+length) must be mapped; > + * 0: none of them may be. Mixed is an error. > + */ > + __u32 mapped; > + __u32 __reserved; > + __aligned_u64 iova; > + __aligned_u64 length; > + } check_mapped; > + struct { > + __aligned_u64 iova; > + __aligned_u64 out_phys; /* 0 if unmapped */ > + } iova_to_phys; > }; > __u32 last; > }; > diff --git a/drivers/iommu/iommufd/pages.c b/drivers/iommu/iommufd/pages.c > index f9b2ae6d7e96..196d1bb330c2 100644 > --- a/drivers/iommu/iommufd/pages.c > +++ b/drivers/iommu/iommufd/pages.c > @@ -1463,68 +1463,12 @@ static const struct dma_buf_attach_ops iopt_dmabuf_attach_revoke_ops = { > .invalidate_mappings = iopt_revoke_notify, > }; > > -/* > - * iommufd and vfio have a circular dependency. Future work for a phys > - * based private interconnect will remove this. > - */ > -/* > - * Look up the exporter's phys accessor for iommufd's private-interconnect > - * path. Also fills *is_cpu_ram: true if the exporter's memory is normal > - * cache-coherent RAM (needs BATCH_CPU_MEMORY / IOMMU_CACHE), false for MMIO > - * (needs BATCH_MMIO / IOMMU_MMIO). This will be replaced by a formal > - * exporter op that returns phys + memory type together. > - */ > -static int > -sym_vfio_pci_dma_buf_iommufd_map(struct dma_buf_attachment *attachment, > - struct phys_vec *phys, bool *is_cpu_ram) > -{ > - typeof(&vfio_pci_dma_buf_iommufd_map) fn; > - int rc; > - > - rc = iommufd_test_dma_buf_iommufd_map(attachment, phys); > - if (rc != -EOPNOTSUPP) { > - *is_cpu_ram = false; /* test hook mimics VFIO MMIO */ > - return rc; > - } > - > - /* > - * Prototype: try the sample gmem provider's dma-buf exporter. This > - * mirrors the vfio-pci private-interconnect hook, and (like it) is > - * meant to be replaced by a formal negotiated exporter op returning > - * phys + memory type. The provider serves RAM, so mark it CPU_RAM. > - */ > - { > - extern int gmem_provider_dma_buf_iommufd_map( > - struct dma_buf_attachment *, struct phys_vec *); > - typeof(&gmem_provider_dma_buf_iommufd_map) gfn; > - > - gfn = symbol_get(gmem_provider_dma_buf_iommufd_map); > - if (gfn) { > - rc = gfn(attachment, phys); > - symbol_put(gmem_provider_dma_buf_iommufd_map); > - if (rc != -EOPNOTSUPP) { > - *is_cpu_ram = true; > - return rc; > - } > - } > - } > - > - if (!IS_ENABLED(CONFIG_VFIO_PCI_DMABUF)) > - return -EOPNOTSUPP; > - > - fn = symbol_get(vfio_pci_dma_buf_iommufd_map); > - if (!fn) > - return -EOPNOTSUPP; > - rc = fn(attachment, phys); > - symbol_put(vfio_pci_dma_buf_iommufd_map); > - *is_cpu_ram = false; /* VFIO PCI dma-buf carries BAR (MMIO) memory */ > - return rc; > -} > - > static int iopt_map_dmabuf(struct iommufd_ctx *ictx, struct iopt_pages *pages, > struct dma_buf *dmabuf) > { > struct dma_buf_attachment *attach; > + struct phys_vec pv; > + u32 attr; > int rc; > > attach = dma_buf_dynamic_attach(dmabuf, iommufd_global_device(), > @@ -1546,10 +1490,20 @@ static int iopt_map_dmabuf(struct iommufd_ctx *ictx, struct iopt_pages *pages, > if (rc) > goto err_detach; > > - rc = sym_vfio_pci_dma_buf_iommufd_map(attach, &pages->dmabuf.phys, > - &pages->dmabuf.is_cpu_ram); > + /* One backed, writable run covering the buffer: refuse the rest. */ > + rc = dma_buf_get_phys(attach, 0, dmabuf->size, &pv, &attr); > + if (rc == -ENOENT) > + rc = -EOPNOTSUPP; > if (rc) > goto err_unpin; > + if (pv.len != dmabuf->size || !dma_buf_phys_attr_known(attr) || > + (attr & DMA_BUF_PHYS_ATTR_FLAGS_MASK)) { > + rc = -EOPNOTSUPP; > + goto err_unpin; > + } > + pages->dmabuf.phys = pv; > + pages->dmabuf.is_cpu_ram = > + dma_buf_phys_attr_type(attr) == DMA_BUF_PHYS_ATTR_RAM; > > dma_resv_unlock(dmabuf->resv); > > diff --git a/drivers/iommu/iommufd/selftest.c b/drivers/iommu/iommufd/selftest.c > index af07c642a526..0899272d1e66 100644 > --- a/drivers/iommu/iommufd/selftest.c > +++ b/drivers/iommu/iommufd/selftest.c > @@ -1962,32 +1962,31 @@ static void iommufd_test_dma_buf_release(struct dma_buf *dmabuf) > kfree(priv); > } > > -static const struct dma_buf_ops iommufd_test_dmabuf_ops = { > - .attach = iommufd_test_dma_buf_attach, > - .detach = iommufd_test_dma_buf_detach, > - .map_dma_buf = iommufd_test_dma_buf_map, > - .release = iommufd_test_dma_buf_release, > - .unmap_dma_buf = iommufd_test_dma_buf_unmap, > -}; > - > -int iommufd_test_dma_buf_iommufd_map(struct dma_buf_attachment *attachment, > - struct phys_vec *phys) > +static int iommufd_test_dma_buf_get_phys(struct dma_buf_attachment *attachment, > + u64 offset, u64 len, > + struct phys_vec *phys, u32 *attr) > { > struct iommufd_test_dma_buf *priv = attachment->dmabuf->priv; > > dma_resv_assert_held(attachment->dmabuf->resv); > - > - if (attachment->dmabuf->ops != &iommufd_test_dmabuf_ops) > - return -EOPNOTSUPP; > - > if (priv->revoked) > return -ENODEV; > > - phys->paddr = virt_to_phys(priv->memory); > - phys->len = priv->length; > + phys->paddr = virt_to_phys(priv->memory) + offset; > + phys->len = len; > + *attr = DMA_BUF_PHYS_ATTR_MMIO; > return 0; > } > > +static const struct dma_buf_ops iommufd_test_dmabuf_ops = { > + .attach = iommufd_test_dma_buf_attach, > + .detach = iommufd_test_dma_buf_detach, > + .map_dma_buf = iommufd_test_dma_buf_map, > + .release = iommufd_test_dma_buf_release, > + .unmap_dma_buf = iommufd_test_dma_buf_unmap, > + .get_phys = iommufd_test_dma_buf_get_phys, > +}; > + > static int iommufd_test_dmabuf_get(struct iommufd_ucmd *ucmd, > unsigned int open_flags, > size_t len) > @@ -2031,6 +2030,73 @@ static int iommufd_test_dmabuf_get(struct iommufd_ucmd *ucmd, > return rc; > } > > +/* > + * Report the physical address the mock domain resolves @iova to, or 0 if > + * it is unmapped. Lets a test check that two IOVAs share one frame (a > + * scratch substitution) without knowing the frame in advance. > + */ > +static int iommufd_test_md_iova_to_phys(struct iommufd_ucmd *ucmd, > + unsigned int mockpt_id, > + unsigned long iova) > +{ > + struct iommu_test_cmd *cmd = ucmd->cmd; > + struct iommufd_hw_pagetable *hwpt; > + struct mock_iommu_domain *mock; > + unsigned int page_size; > + int rc; > + > + hwpt = get_md_pagetable(ucmd, mockpt_id, &mock); > + if (IS_ERR(hwpt)) > + return PTR_ERR(hwpt); > + > + page_size = 1 << __ffs(mock->domain.pgsize_bitmap); > + if (iova % page_size) { > + rc = -EINVAL; > + goto out_put; > + } > + cmd->iova_to_phys.out_phys = > + mock->domain.ops->iova_to_phys(&mock->domain, iova); > + rc = iommufd_ucmd_respond(ucmd, sizeof(*cmd)); > +out_put: > + iommufd_put_object(ucmd->ictx, &hwpt->obj); > + return rc; > +} > + > +static int iommufd_test_md_check_mapped(struct iommufd_ucmd *ucmd, > + unsigned int mockpt_id, > + unsigned long iova, size_t length, > + bool mapped) > +{ > + struct iommufd_hw_pagetable *hwpt; > + struct mock_iommu_domain *mock; > + unsigned int page_size; > + int rc = 0; > + > + hwpt = get_md_pagetable(ucmd, mockpt_id, &mock); > + if (IS_ERR(hwpt)) > + return PTR_ERR(hwpt); > + > + page_size = 1 << __ffs(mock->domain.pgsize_bitmap); > + if (iova % page_size || length % page_size || !length) { > + rc = -EINVAL; > + goto out_put; > + } > + > + for (; length; length -= page_size, iova += page_size) { > + bool is_mapped = > + mock->domain.ops->iova_to_phys(&mock->domain, iova) != 0; > + > + if (is_mapped != mapped) { > + rc = -ENOENT; > + goto out_put; > + } > + } > + > +out_put: > + iommufd_put_object(ucmd->ictx, &hwpt->obj); > + return rc; > +} > + > static int iommufd_test_dmabuf_revoke(struct iommufd_ucmd *ucmd, int fd, > bool revoked) > { > @@ -2143,6 +2209,14 @@ int iommufd_test(struct iommufd_ucmd *ucmd) > return iommufd_test_dmabuf_revoke(ucmd, > cmd->dmabuf_revoke.dmabuf_fd, > cmd->dmabuf_revoke.revoked); > + case IOMMU_TEST_OP_MD_CHECK_MAPPED: > + return iommufd_test_md_check_mapped(ucmd, cmd->id, > + cmd->check_mapped.iova, > + cmd->check_mapped.length, > + cmd->check_mapped.mapped); > + case IOMMU_TEST_OP_MD_IOVA_TO_PHYS: > + return iommufd_test_md_iova_to_phys(ucmd, cmd->id, > + cmd->iova_to_phys.iova); > default: > return -EOPNOTSUPP; > } > diff --git a/drivers/vfio/pci/vfio_pci_dmabuf.c b/drivers/vfio/pci/vfio_pci_dmabuf.c > index c16f460c01d6..381c3d338e8c 100644 > --- a/drivers/vfio/pci/vfio_pci_dmabuf.c > +++ b/drivers/vfio/pci/vfio_pci_dmabuf.c > @@ -99,46 +99,46 @@ static void vfio_pci_dma_buf_release(struct dma_buf *dmabuf) > kfree(priv); > } > > -static const struct dma_buf_ops vfio_pci_dmabuf_ops = { > - .attach = vfio_pci_dma_buf_attach, > - .map_dma_buf = vfio_pci_dma_buf_map, > - .unmap_dma_buf = vfio_pci_dma_buf_unmap, > - .release = vfio_pci_dma_buf_release, > -}; > - > /* > - * This is a temporary "private interconnect" between VFIO DMABUF and iommufd. > - * It allows the two co-operating drivers to exchange the physical address of > - * the BAR. This is to be replaced with a formal DMABUF system for negotiated > - * interconnect types. > + * Report the BAR's physical range for importers which program it into their own > + * translation tables, such as iommufd. A BAR is MMIO, never cache-coherent RAM. > * > - * If this function succeeds the following are true: > - * - There is one physical range and it is pointing to MMIO > - * - When move_notify is called it means revoke, not move, vfio_dma_buf_map > - * will fail if it is currently revoked > + * When move_notify is called it means revoke, not move, so this fails while > + * revoked and vfio_dma_buf_map() does the same. > */ > -int vfio_pci_dma_buf_iommufd_map(struct dma_buf_attachment *attachment, > - struct phys_vec *phys) > +static int vfio_pci_dma_buf_get_phys(struct dma_buf_attachment *attachment, > + u64 offset, u64 len, > + struct phys_vec *phys, u32 *attr) > { > - struct vfio_pci_dma_buf *priv; > + struct vfio_pci_dma_buf *priv = attachment->dmabuf->priv; > + u32 i; > > dma_resv_assert_held(attachment->dmabuf->resv); > - > - if (attachment->dmabuf->ops != &vfio_pci_dmabuf_ops) > - return -EOPNOTSUPP; > - > - priv = attachment->dmabuf->priv; > if (priv->revoked) > return -ENODEV; > > - /* More than one range to iommufd will require proper DMABUF support */ > - if (priv->nr_ranges != 1) > - return -EOPNOTSUPP; > - > - *phys = priv->phys_vec[0]; > + /* Report from @offset to the end of the BAR range containing it. */ > + for (i = 0; i < priv->nr_ranges; i++) { > + if (offset < priv->phys_vec[i].len) > + break; > + offset -= priv->phys_vec[i].len; > + } > + if (i == priv->nr_ranges) > + return -EINVAL; > + phys->paddr = priv->phys_vec[i].paddr + offset; > + phys->len = min_t(u64, priv->phys_vec[i].len - offset, len); > + *attr = DMA_BUF_PHYS_ATTR_MMIO; > return 0; > } > -EXPORT_SYMBOL_FOR_MODULES(vfio_pci_dma_buf_iommufd_map, "iommufd"); > + > +static const struct dma_buf_ops vfio_pci_dmabuf_ops = { > + .attach = vfio_pci_dma_buf_attach, > + .map_dma_buf = vfio_pci_dma_buf_map, > + .unmap_dma_buf = vfio_pci_dma_buf_unmap, > + .release = vfio_pci_dma_buf_release, > + .get_phys = vfio_pci_dma_buf_get_phys, > +}; > + > > int vfio_pci_core_fill_phys_vec(struct phys_vec *phys_vec, > struct vfio_region_dma_range *dma_ranges, > diff --git a/include/linux/dma-buf.h b/include/linux/dma-buf.h > index d1203da56fc5..b223962e20c2 100644 > --- a/include/linux/dma-buf.h > +++ b/include/linux/dma-buf.h > @@ -13,6 +13,7 @@ > #ifndef __DMA_BUF_H__ > #define __DMA_BUF_H__ > > +#include <linux/bitfield.h> > #include <linux/iosys-map.h> > #include <linux/file.h> > #include <linux/err.h> > @@ -23,6 +24,7 @@ > #include <linux/dma-fence.h> > #include <linux/wait.h> > #include <linux/pci-p2pdma.h> > +#include <linux/types.h> > > struct device; > struct dma_buf; > @@ -182,6 +184,29 @@ struct dma_buf_ops { > struct sg_table *, > enum dma_data_direction); > > + /** > + * @get_phys: > + * > + * Describe the run that starts at @offset, for an importer that > + * programs its own translation tables. A run is the longest stretch > + * of backed bytes whose frames are physically contiguous and share > + * one attribute word. Report it in *@phys, starting at the byte at > + * @offset and clipped at @offset + @len, with its DMA_BUF_PHYS_ATTR_* > + * word in *@attr. The exporter may stop before the end of the run; > + * importers must not assume the reported run is maximal. > + * > + * Return 0 on success, -ENOENT if the byte at @offset is not backed, > + * -ENODEV if the buffer is revoked, or another negative error. Do not > + * wait for memory to become available. > + * > + * The dma-buf reservation is held. The attachment must be pinned or > + * have revocable importer operations. A reported frame remains valid > + * until a covering invalidation callback returns; an exporter must > + * invalidate every change before reusing an old frame. > + */ > + int (*get_phys)(struct dma_buf_attachment *attach, u64 offset, u64 len, > + struct phys_vec *phys, u32 *attr); > + > /* TODO: Add try_map_dma_buf version, to return immed with -EBUSY > * if the call would block. > */ > @@ -576,6 +601,35 @@ void dma_buf_unmap_attachment(struct dma_buf_attachment *, struct sg_table *, > enum dma_data_direction); > void dma_buf_invalidate_mappings(struct dma_buf *dma_buf); > bool dma_buf_attach_revocable(struct dma_buf_attachment *attach); > +/* bits 0-7: memory type (a value, not flags) */ > +#define DMA_BUF_PHYS_ATTR_TYPE_MASK GENMASK(7, 0) > +#define DMA_BUF_PHYS_ATTR_RAM 0x00 /* cache-coherent system RAM */ > +#define DMA_BUF_PHYS_ATTR_MMIO 0x01 /* device MMIO, uncached */ > +/* bits 8-15: reserved for a second value field; must be zero */ > +#define DMA_BUF_PHYS_ATTR_RSVD_MASK GENMASK(15, 8) > +/* bits 16-31: flags; undefined bits must be zero */ > +#define DMA_BUF_PHYS_ATTR_READONLY BIT(16) > +/* BIT(17): reserved (hole, for importers that walk across gaps) */ > +/* BIT(18): reserved (private, for confidential computing) */ > +#define DMA_BUF_PHYS_ATTR_FLAGS_MASK (DMA_BUF_PHYS_ATTR_READONLY) > + > +static inline u32 dma_buf_phys_attr_type(u32 attrs) > +{ > + return FIELD_GET(DMA_BUF_PHYS_ATTR_TYPE_MASK, attrs); > +} > + > +static inline bool dma_buf_phys_attr_known(u32 attrs) > +{ > + return dma_buf_phys_attr_type(attrs) <= DMA_BUF_PHYS_ATTR_MMIO && > + !(attrs & DMA_BUF_PHYS_ATTR_RSVD_MASK) && > + !(attrs & ~(DMA_BUF_PHYS_ATTR_TYPE_MASK | > + DMA_BUF_PHYS_ATTR_RSVD_MASK | > + DMA_BUF_PHYS_ATTR_FLAGS_MASK)); > +} > + > +int dma_buf_get_phys(struct dma_buf_attachment *attach, u64 offset, u64 len, > + struct phys_vec *phys, u32 *attr); > + > int dma_buf_begin_cpu_access(struct dma_buf *dma_buf, > enum dma_data_direction dir); > int dma_buf_end_cpu_access(struct dma_buf *dma_buf, > diff --git a/include/linux/vfio_pci_core.h b/include/linux/vfio_pci_core.h > index 9a1674c152aa..2a1d13abdb0c 100644 > --- a/include/linux/vfio_pci_core.h > +++ b/include/linux/vfio_pci_core.h > @@ -257,7 +257,4 @@ vfio_pci_core_get_iomap(struct vfio_pci_core_device *vdev, unsigned int bar) > return vdev->barmap[bar]; > } > > -int vfio_pci_dma_buf_iommufd_map(struct dma_buf_attachment *attachment, > - struct phys_vec *phys); > - > #endif /* VFIO_PCI_CORE_H */ > diff --git a/samples/kvm/gmem_provider.c b/samples/kvm/gmem_provider.c > index 75197c088762..b6824fe5d228 100644 > --- a/samples/kvm/gmem_provider.c > +++ b/samples/kvm/gmem_provider.c > @@ -487,35 +487,30 @@ static void gmem_dma_buf_release(struct dma_buf *dmabuf) > kfree(priv); > } > > -static const struct dma_buf_ops gmem_dma_buf_ops = { > - .attach = gmem_dma_buf_attach, > - .map_dma_buf = gmem_dma_buf_map, > - .unmap_dma_buf = gmem_dma_buf_unmap, > - .release = gmem_dma_buf_release, > -}; > - > -/* > - * Private interconnect for iommufd (mirrors vfio_pci_dma_buf_iommufd_map). > - * Returns the single contiguous phys range for the exported region so iommufd > - * can program the IOMMU directly, bypassing the DMA API. > - */ > -int gmem_provider_dma_buf_iommufd_map(struct dma_buf_attachment *attach, > - struct phys_vec *phys); > -int gmem_provider_dma_buf_iommufd_map(struct dma_buf_attachment *attach, > - struct phys_vec *phys) > +/* Report this flat sample region through the generic dma-buf operation. */ > +static int gmem_dma_buf_get_phys(struct dma_buf_attachment *attach, > + u64 offset, u64 len, > + struct phys_vec *phys, u32 *attr) > { > - struct gmem_dmabuf *priv; > + struct gmem_dmabuf *priv = attach->dmabuf->priv; > > dma_resv_assert_held(attach->dmabuf->resv); > - if (attach->dmabuf->ops != &gmem_dma_buf_ops) > - return -EOPNOTSUPP; > - priv = attach->dmabuf->priv; > if (priv->revoked) > return -ENODEV; > - *phys = priv->phys; > + > + phys->paddr = priv->phys.paddr + offset; > + phys->len = len; > + *attr = DMA_BUF_PHYS_ATTR_RAM; > return 0; > } > -EXPORT_SYMBOL_FOR_MODULES(gmem_provider_dma_buf_iommufd_map, "iommufd"); > + > +static const struct dma_buf_ops gmem_dma_buf_ops = { > + .attach = gmem_dma_buf_attach, > + .map_dma_buf = gmem_dma_buf_map, > + .unmap_dma_buf = gmem_dma_buf_unmap, > + .release = gmem_dma_buf_release, > + .get_phys = gmem_dma_buf_get_phys, > +}; > > /* Called with info->dmabufs_lock held on the revoke path. */ > static void gmem_dma_buf_revoke_all(struct gmem_info *info) > -- > 2.47.3 > ^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [RFC PATCH 2/6] dma-buf: Add get_phys() to describe a physical run 2026-10-05 10:07 ` Christian König @ 2026-10-05 13:20 ` Fred Griffoul 2026-10-05 14:53 ` Christian König 0 siblings, 1 reply; 28+ messages in thread From: Fred Griffoul @ 2026-10-05 13:20 UTC (permalink / raw) To: Paolo Bonzini, Sean Christopherson, Marc Zyngier, Oliver Upton, Sumit Semwal, Christian König, Jason Gunthorpe, Kevin Tian Cc: David Woodhouse, Ackerley Tng, Joey Gouly, Suzuki K Poulose, Zenghui Yu, Steffen Eiden, Catalin Marinas, Will Deacon, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, Joerg Roedel, Robin Murphy, Alex Williamson, Shuah Khan, Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers, linux-kernel, kvm, kvmarm, linux-arm-kernel, iommu, linux-media, dri-devel, linaro-mm-sig, linux-kselftest, linux-trace-kernel, x86 On 10/5/26 Christian König wrote: > Well filling page tables by the importer is an absolutely clear NO-GO > for the DMA-buf design, we have gotten down that path already and it > took us years to remove this functionality again. Understood, thanks for the quick and clear answers on this patch and on patch 3. > The problem you are facing here is that dma_buf_mmap() doesn't work > because you don't have a VMA. Right: the memory has no struct page and is deliberately not mapped in the host, so there is no VMA to hand to KVM. I'll post a new version of this series. It drops get_phys(), the ranged invalidation and the guest_memfd dma-buf import (patches 2-5), and it does not touch drivers/dma-buf, as you suggested in the PAL discussion: https://lore.kernel.org/all/c413710b-4c28-4ed8-88ec-aeb8c4482011@amd.com/ KVM gets its frames through the KVM-private guest_memfd provider interface from David's series. iommufd gets them through a private interface with the exporter, as it already does for vfio-pci. The new version adds a small registry in iommufd for that, which replaces the symbol_get() of the sample in David's series. The dma-buf is then only the handle and the existing revoke. Thanks, Fred ^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [RFC PATCH 2/6] dma-buf: Add get_phys() to describe a physical run 2026-10-05 13:20 ` Fred Griffoul @ 2026-10-05 14:53 ` Christian König 0 siblings, 0 replies; 28+ messages in thread From: Christian König @ 2026-10-05 14:53 UTC (permalink / raw) To: Fred Griffoul, Paolo Bonzini, Sean Christopherson, Marc Zyngier, Oliver Upton, Sumit Semwal, Jason Gunthorpe, Kevin Tian Cc: David Woodhouse, Ackerley Tng, Joey Gouly, Suzuki K Poulose, Zenghui Yu, Steffen Eiden, Catalin Marinas, Will Deacon, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, Joerg Roedel, Robin Murphy, Alex Williamson, Shuah Khan, Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers, linux-kernel, kvm, kvmarm, linux-arm-kernel, iommu, linux-media, dri-devel, linaro-mm-sig, linux-kselftest, linux-trace-kernel, x86 On 10/5/26 15:20, Fred Griffoul wrote: > On 10/5/26 Christian König wrote: >> Well filling page tables by the importer is an absolutely clear NO-GO >> for the DMA-buf design, we have gotten down that path already and it >> took us years to remove this functionality again. > > Understood, thanks for the quick and clear answers on this patch and on > patch 3. > >> The problem you are facing here is that dma_buf_mmap() doesn't work >> because you don't have a VMA. > > Right: the memory has no struct page and is deliberately not mapped in > the host, so there is no VMA to hand to KVM. > > I'll post a new version of this series. It drops get_phys(), the ranged > invalidation and the guest_memfd dma-buf import (patches 2-5), and it > does not touch drivers/dma-buf, as you suggested in the PAL discussion: > > https://lore.kernel.org/all/c413710b-4c28-4ed8-88ec-aeb8c4482011@amd.com/ > > KVM gets its frames through the KVM-private guest_memfd provider > interface from David's series. iommufd gets them through a private > interface with the exporter, as it already does for vfio-pci. The new > version adds a small registry in iommufd for that, which replaces the > symbol_get() of the sample in David's series. The dma-buf is then only > the handle and the existing revoke. Yeah that approach sounds totally sane to me. We have drivers which stuff all kind of resources, not just memory, into a DMA-buf. That ranges from MMIO BARs to trigger FW operations (doorbells) over full HW blocks (ordered appends, global wave sync etc...) where two or more applications need to share which one is used. That you then have a exporter private interface is for accessing this is perfectly ok. The problem you are facing here is more that this is not limited to one driver/module but multiple ones and you don't want module inter dependencies because of the symbols. So symbol_get() is probably the right thing to do as long as we don't have something like weak symbols or similar. Regards, Christian. > > Thanks, > Fred ^ permalink raw reply [flat|nested] 28+ messages in thread
* [RFC PATCH 3/6] dma-buf: Add ranged mapping invalidation 2026-10-05 9:55 ` [RFC PATCH 0/6] KVM: guest_memfd: back guest_memfd with an imported dma-buf Fred Griffoul 2026-10-05 9:55 ` [RFC PATCH 1/6] KVM: guest_memfd: Add a writable result to get_pfn() Fred Griffoul 2026-10-05 9:55 ` [RFC PATCH 2/6] dma-buf: Add get_phys() to describe a physical run Fred Griffoul @ 2026-10-05 9:55 ` Fred Griffoul 2026-10-05 10:08 ` Christian König 2026-10-05 9:55 ` [RFC PATCH 4/6] dma-buf: Allow dynamic attach without a device Fred Griffoul ` (2 subsequent siblings) 5 siblings, 1 reply; 28+ messages in thread From: Fred Griffoul @ 2026-10-05 9:55 UTC (permalink / raw) To: Paolo Bonzini, Sean Christopherson, Marc Zyngier, Oliver Upton, Sumit Semwal, Christian König, Jason Gunthorpe, Kevin Tian Cc: David Woodhouse, Ackerley Tng, Joey Gouly, Suzuki K Poulose, Zenghui Yu, Steffen Eiden, Catalin Marinas, Will Deacon, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, Joerg Roedel, Robin Murphy, Alex Williamson, Shuah Khan, Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers, linux-kernel, kvm, kvmarm, linux-arm-kernel, iommu, linux-media, dri-devel, linaro-mm-sig, linux-kselftest, linux-trace-kernel, x86 From: Fred Griffoul <fgriffo@amazon.co.uk> dma_buf_invalidate_mappings() tells every importer that the whole buffer changed. An exporter that changes one part of its memory cannot say which bytes changed, so importers throw away mappings that are still valid. Add an exporter helper that invalidates a byte range, and an importer callback that receives it. The callback means that the address, the attributes or the backing of the range changed. If part of the range is no longer backed, get_phys() returns -ENOENT for it. Importers must stop using their old answer before the callback returns. Importers that do not implement the callback still receive a whole-buffer invalidation. Signed-off-by: Fred Griffoul <fgriffo@amazon.co.uk> --- drivers/dma-buf/dma-buf.c | 30 ++++++++++++++++++++++++++++++ include/linux/dma-buf.h | 16 ++++++++++++++++ 2 files changed, 46 insertions(+) diff --git a/drivers/dma-buf/dma-buf.c b/drivers/dma-buf/dma-buf.c index 66b85d53ed22..e7010163eb2f 100644 --- a/drivers/dma-buf/dma-buf.c +++ b/drivers/dma-buf/dma-buf.c @@ -1389,6 +1389,36 @@ void dma_buf_invalidate_mappings(struct dma_buf *dmabuf) } EXPORT_SYMBOL_NS_GPL(dma_buf_invalidate_mappings, "DMA_BUF"); +/** + * dma_buf_invalidate_mappings_range - notify attachments that a range changed + * @dmabuf: buffer whose layout changed + * @offset: first changed byte + * @length: number of changed bytes + * + * Importers with a ranged callback stop using their old mappings of the range + * before returning. Other importers receive the existing whole-buffer + * callback, which is correct but coarser. The reservation lock must be held. + */ +void dma_buf_invalidate_mappings_range(struct dma_buf *dmabuf, + unsigned long offset, + unsigned long length) +{ + struct dma_buf_attachment *attach; + + dma_resv_assert_held(dmabuf->resv); + list_for_each_entry(attach, &dmabuf->attachments, node) { + const struct dma_buf_attach_ops *ops = attach->importer_ops; + + if (!ops) + continue; + if (ops->invalidate_mappings_range) + ops->invalidate_mappings_range(attach, offset, length); + else if (ops->invalidate_mappings) + ops->invalidate_mappings(attach); + } +} +EXPORT_SYMBOL_NS_GPL(dma_buf_invalidate_mappings_range, "DMA_BUF"); + /** * dma_buf_get_phys - describe the run that starts at an offset * @attach: attachment to query diff --git a/include/linux/dma-buf.h b/include/linux/dma-buf.h index b223962e20c2..55c3fe60a0ba 100644 --- a/include/linux/dma-buf.h +++ b/include/linux/dma-buf.h @@ -485,6 +485,19 @@ struct dma_buf_attach_ops { * required behavior. */ void (*invalidate_mappings)(struct dma_buf_attachment *attach); + + /** + * @invalidate_mappings_range: [optional] a byte range changed + * + * The exporter changed the address, attributes or backing of + * [@offset, @offset + @length). The importer must stop using its old + * answer for that range before returning. + * Importers without this callback receive @invalidate_mappings for + * the whole buffer instead. + */ + void (*invalidate_mappings_range)(struct dma_buf_attachment *attach, + unsigned long offset, + unsigned long length); }; /** @@ -600,6 +613,9 @@ struct sg_table *dma_buf_map_attachment(struct dma_buf_attachment *, void dma_buf_unmap_attachment(struct dma_buf_attachment *, struct sg_table *, enum dma_data_direction); void dma_buf_invalidate_mappings(struct dma_buf *dma_buf); +void dma_buf_invalidate_mappings_range(struct dma_buf *dma_buf, + unsigned long offset, + unsigned long length); bool dma_buf_attach_revocable(struct dma_buf_attachment *attach); /* bits 0-7: memory type (a value, not flags) */ #define DMA_BUF_PHYS_ATTR_TYPE_MASK GENMASK(7, 0) -- 2.47.3 ^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [RFC PATCH 3/6] dma-buf: Add ranged mapping invalidation 2026-10-05 9:55 ` [RFC PATCH 3/6] dma-buf: Add ranged mapping invalidation Fred Griffoul @ 2026-10-05 10:08 ` Christian König 0 siblings, 0 replies; 28+ messages in thread From: Christian König @ 2026-10-05 10:08 UTC (permalink / raw) To: Fred Griffoul, Paolo Bonzini, Sean Christopherson, Marc Zyngier, Oliver Upton, Sumit Semwal, Jason Gunthorpe, Kevin Tian Cc: David Woodhouse, Ackerley Tng, Joey Gouly, Suzuki K Poulose, Zenghui Yu, Steffen Eiden, Catalin Marinas, Will Deacon, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, Joerg Roedel, Robin Murphy, Alex Williamson, Shuah Khan, Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers, linux-kernel, kvm, kvmarm, linux-arm-kernel, iommu, linux-media, dri-devel, linaro-mm-sig, linux-kselftest, linux-trace-kernel, x86 On 10/5/26 11:55, Fred Griffoul wrote: > From: Fred Griffoul <fgriffo@amazon.co.uk> > > dma_buf_invalidate_mappings() tells every importer that the whole > buffer changed. An exporter that changes one part of its memory cannot > say which bytes changed, so importers throw away mappings that are > still valid. > > Add an exporter helper that invalidates a byte range, and an importer > callback that receives it. The callback means that the address, the > attributes or the backing of the range changed. If part of the range is > no longer backed, get_phys() returns -ENOENT for it. Importers must stop > using their old answer before the callback returns. Importers that do > not implement the callback still receive a whole-buffer invalidation. Yeah that is exactly one of the reasons why we don't allow that. Clear NAK to the whole approach. See the reply to patch #2 for a detailed description. Regards, Christian. > > Signed-off-by: Fred Griffoul <fgriffo@amazon.co.uk> > --- > drivers/dma-buf/dma-buf.c | 30 ++++++++++++++++++++++++++++++ > include/linux/dma-buf.h | 16 ++++++++++++++++ > 2 files changed, 46 insertions(+) > > diff --git a/drivers/dma-buf/dma-buf.c b/drivers/dma-buf/dma-buf.c > index 66b85d53ed22..e7010163eb2f 100644 > --- a/drivers/dma-buf/dma-buf.c > +++ b/drivers/dma-buf/dma-buf.c > @@ -1389,6 +1389,36 @@ void dma_buf_invalidate_mappings(struct dma_buf *dmabuf) > } > EXPORT_SYMBOL_NS_GPL(dma_buf_invalidate_mappings, "DMA_BUF"); > > +/** > + * dma_buf_invalidate_mappings_range - notify attachments that a range changed > + * @dmabuf: buffer whose layout changed > + * @offset: first changed byte > + * @length: number of changed bytes > + * > + * Importers with a ranged callback stop using their old mappings of the range > + * before returning. Other importers receive the existing whole-buffer > + * callback, which is correct but coarser. The reservation lock must be held. > + */ > +void dma_buf_invalidate_mappings_range(struct dma_buf *dmabuf, > + unsigned long offset, > + unsigned long length) > +{ > + struct dma_buf_attachment *attach; > + > + dma_resv_assert_held(dmabuf->resv); > + list_for_each_entry(attach, &dmabuf->attachments, node) { > + const struct dma_buf_attach_ops *ops = attach->importer_ops; > + > + if (!ops) > + continue; > + if (ops->invalidate_mappings_range) > + ops->invalidate_mappings_range(attach, offset, length); > + else if (ops->invalidate_mappings) > + ops->invalidate_mappings(attach); > + } > +} > +EXPORT_SYMBOL_NS_GPL(dma_buf_invalidate_mappings_range, "DMA_BUF"); > + > /** > * dma_buf_get_phys - describe the run that starts at an offset > * @attach: attachment to query > diff --git a/include/linux/dma-buf.h b/include/linux/dma-buf.h > index b223962e20c2..55c3fe60a0ba 100644 > --- a/include/linux/dma-buf.h > +++ b/include/linux/dma-buf.h > @@ -485,6 +485,19 @@ struct dma_buf_attach_ops { > * required behavior. > */ > void (*invalidate_mappings)(struct dma_buf_attachment *attach); > + > + /** > + * @invalidate_mappings_range: [optional] a byte range changed > + * > + * The exporter changed the address, attributes or backing of > + * [@offset, @offset + @length). The importer must stop using its old > + * answer for that range before returning. > + * Importers without this callback receive @invalidate_mappings for > + * the whole buffer instead. > + */ > + void (*invalidate_mappings_range)(struct dma_buf_attachment *attach, > + unsigned long offset, > + unsigned long length); > }; > > /** > @@ -600,6 +613,9 @@ struct sg_table *dma_buf_map_attachment(struct dma_buf_attachment *, > void dma_buf_unmap_attachment(struct dma_buf_attachment *, struct sg_table *, > enum dma_data_direction); > void dma_buf_invalidate_mappings(struct dma_buf *dma_buf); > +void dma_buf_invalidate_mappings_range(struct dma_buf *dma_buf, > + unsigned long offset, > + unsigned long length); > bool dma_buf_attach_revocable(struct dma_buf_attachment *attach); > /* bits 0-7: memory type (a value, not flags) */ > #define DMA_BUF_PHYS_ATTR_TYPE_MASK GENMASK(7, 0) > -- > 2.47.3 > ^ permalink raw reply [flat|nested] 28+ messages in thread
* [RFC PATCH 4/6] dma-buf: Allow dynamic attach without a device 2026-10-05 9:55 ` [RFC PATCH 0/6] KVM: guest_memfd: back guest_memfd with an imported dma-buf Fred Griffoul ` (2 preceding siblings ...) 2026-10-05 9:55 ` [RFC PATCH 3/6] dma-buf: Add ranged mapping invalidation Fred Griffoul @ 2026-10-05 9:55 ` Fred Griffoul 2026-10-05 9:55 ` [RFC PATCH 5/6] KVM: guest_memfd: Add dma-buf backing Fred Griffoul 2026-10-05 9:55 ` [RFC PATCH 6/6] samples/kvm, selftests/kvm: Exercise " Fred Griffoul 5 siblings, 0 replies; 28+ messages in thread From: Fred Griffoul @ 2026-10-05 9:55 UTC (permalink / raw) To: Paolo Bonzini, Sean Christopherson, Marc Zyngier, Oliver Upton, Sumit Semwal, Christian König, Jason Gunthorpe, Kevin Tian Cc: David Woodhouse, Ackerley Tng, Joey Gouly, Suzuki K Poulose, Zenghui Yu, Steffen Eiden, Catalin Marinas, Will Deacon, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, Joerg Roedel, Robin Murphy, Alex Williamson, Shuah Khan, Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers, linux-kernel, kvm, kvmarm, linux-arm-kernel, iommu, linux-media, dri-devel, linaro-mm-sig, linux-kselftest, linux-trace-kernel, x86 From: Fred Griffoul <fgriffo@amazon.co.uk> dma_buf_dynamic_attach() requires a struct device: the device for which the buffer is mapped for DMA. Exporters may also check it before they accept the attachment. An importer that does not perform DMA has no such device. iommufd already attaches with a global placeholder device for this reason. KVM would need one too, although it only needs to know where the memory is so that it can map it into a guest. Allow @dev to be NULL for a dynamic importer. Such an importer learns where the memory is from get_phys() and builds its own mappings. It must not map the attachment for DMA, and dma_buf_map_attachment() refuses an attachment without a device. An exporter that needs a device to answer can refuse the attachment in its attach op, as it can for any other reason. Signed-off-by: Fred Griffoul <fgriffo@amazon.co.uk> --- drivers/dma-buf/dma-buf.c | 15 +++++++++++++-- include/trace/events/dma_buf.h | 2 +- 2 files changed, 14 insertions(+), 3 deletions(-) diff --git a/drivers/dma-buf/dma-buf.c b/drivers/dma-buf/dma-buf.c index e7010163eb2f..3d43b529a523 100644 --- a/drivers/dma-buf/dma-buf.c +++ b/drivers/dma-buf/dma-buf.c @@ -1007,6 +1007,13 @@ dma_buf_pin_on_map(struct dma_buf_attachment *attach) * Note that this can fail if the backing storage of @dmabuf is in a place not * accessible to @dev, and cannot be moved to a more suitable place. This is * indicated with the error code -EBUSY. + * + * @dev may be NULL for a dynamic importer that is not a DMA master: one that + * only asks the exporter where the memory is, through &dma_buf_ops.get_phys, + * and builds its own mappings from the answer. Such an importer must never + * call dma_buf_map_attachment(), and an exporter that needs a device to + * answer (a peer-to-peer check, say) refuses the attachment in its + * &dma_buf_ops.attach. */ struct dma_buf_attachment * dma_buf_dynamic_attach(struct dma_buf *dmabuf, struct device *dev, @@ -1016,7 +1023,7 @@ dma_buf_dynamic_attach(struct dma_buf *dmabuf, struct device *dev, struct dma_buf_attachment *attach; int ret; - if (WARN_ON(!dmabuf || !dev)) + if (WARN_ON(!dmabuf || (!dev && !importer_ops))) return ERR_PTR(-EINVAL); attach = kzalloc_obj(*attach); @@ -1175,6 +1182,9 @@ struct sg_table *dma_buf_map_attachment(struct dma_buf_attachment *attach, if (WARN_ON(!attach || !attach->dmabuf)) return ERR_PTR(-EINVAL); + /* An importer without a device cannot be a DMA master. */ + if (WARN_ON(!attach->dev)) + return ERR_PTR(-EINVAL); dma_resv_assert_held(attach->dmabuf->resv); @@ -1847,7 +1857,8 @@ static int dma_buf_debug_show(struct seq_file *s, void *unused) attach_count = 0; list_for_each_entry(attach_obj, &buf_obj->attachments, node) { - seq_printf(s, "\t%s\n", dev_name(attach_obj->dev)); + seq_printf(s, "\t%s\n", attach_obj->dev ? + dev_name(attach_obj->dev) : "none"); attach_count++; } dma_resv_unlock(buf_obj->resv); diff --git a/include/trace/events/dma_buf.h b/include/trace/events/dma_buf.h index 3bb88d05bcc8..7cabca53b798 100644 --- a/include/trace/events/dma_buf.h +++ b/include/trace/events/dma_buf.h @@ -40,7 +40,7 @@ DECLARE_EVENT_CLASS(dma_buf_attach_dev, TP_ARGS(dmabuf, attach, is_dynamic, dev), TP_STRUCT__entry( - __string( dev_name, dev_name(dev)) + __string(dev_name, dev ? dev_name(dev) : "none") __string( exp_name, dmabuf->exp_name) __field( size_t, size) __field( ino_t, ino) -- 2.47.3 ^ permalink raw reply [flat|nested] 28+ messages in thread
* [RFC PATCH 5/6] KVM: guest_memfd: Add dma-buf backing 2026-10-05 9:55 ` [RFC PATCH 0/6] KVM: guest_memfd: back guest_memfd with an imported dma-buf Fred Griffoul ` (3 preceding siblings ...) 2026-10-05 9:55 ` [RFC PATCH 4/6] dma-buf: Allow dynamic attach without a device Fred Griffoul @ 2026-10-05 9:55 ` Fred Griffoul 2026-10-05 9:55 ` [RFC PATCH 6/6] samples/kvm, selftests/kvm: Exercise " Fred Griffoul 5 siblings, 0 replies; 28+ messages in thread From: Fred Griffoul @ 2026-10-05 9:55 UTC (permalink / raw) To: Paolo Bonzini, Sean Christopherson, Marc Zyngier, Oliver Upton, Sumit Semwal, Christian König, Jason Gunthorpe, Kevin Tian Cc: David Woodhouse, Ackerley Tng, Joey Gouly, Suzuki K Poulose, Zenghui Yu, Steffen Eiden, Catalin Marinas, Will Deacon, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, Joerg Roedel, Robin Murphy, Alex Williamson, Shuah Khan, Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers, linux-kernel, kvm, kvmarm, linux-arm-kernel, iommu, linux-media, dri-devel, linaro-mm-sig, linux-kselftest, linux-trace-kernel, x86 From: Fred Griffoul <fgriffo@amazon.co.uk> A memory owner that exports a dma-buf should be able to use that one object for both guest stage-2 mappings and device IOMMU mappings. A second, KVM-specific backing interface would duplicate the description of the memory layout and the rules for invalidating it. Add GUEST_MEMFD_FLAG_USE_DMABUF and a dmabuf_fd field to KVM_CREATE_GUEST_MEMFD. This backing has its own kvm_gmem_ops. guest_memfd attaches to the buffer as a revocable importer without a device. At creation, it asks for the run at offset 0 and refuses MMIO or unknown attributes, so a buffer that KVM can never map fails early. Offset 0 may be unbacked. On a fault, guest_memfd takes the reservation once. It asks get_phys() about the PUD-sized, PMD-sized and page-sized blocks around the faulting page, in that order, and maps the first block that one aligned RAM run covers. If the page is not backed or the run is not supported, the fault returns -EFAULT. A READONLY run is mapped without stage-2 write permission. The invalidation callback only runs KVM's invalidation over the affected range. It does not allocate memory or call the exporter. A fault that read the layout during the change sees mmu_invalidate_seq move and retries. Callbacks can start as soon as guest_memfd attaches. A dma-buf-backed file is therefore added to the inode's list of guest_memfd files under filemap_invalidate_lock. mmap() uses the exporter's dma-buf mmap. fallocate() and populate are not supported. Import the DMA_BUF symbol namespace so that a CONFIG_KVM=m build passes modpost. Signed-off-by: Fred Griffoul <fgriffo@amazon.co.uk> --- include/linux/kvm_host.h | 1 + include/uapi/linux/kvm.h | 6 +- tools/include/uapi/linux/kvm.h | 6 +- virt/kvm/Kconfig | 1 + virt/kvm/guest_memfd.c | 239 +++++++++++++++++++++++++++++++-- 5 files changed, 243 insertions(+), 10 deletions(-) diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h index f84f3ab44acf..0158dc46580b 100644 --- a/include/linux/kvm_host.h +++ b/include/linux/kvm_host.h @@ -834,6 +834,7 @@ static inline u64 kvm_gmem_get_supported_flags(struct kvm *kvm) if (!kvm || kvm_arch_supports_gmem_init_shared(kvm)) flags |= GUEST_MEMFD_FLAG_INIT_SHARED; + flags |= GUEST_MEMFD_FLAG_USE_DMABUF; return flags; } diff --git a/include/uapi/linux/kvm.h b/include/uapi/linux/kvm.h index 419011097fa8..f8c83013773d 100644 --- a/include/uapi/linux/kvm.h +++ b/include/uapi/linux/kvm.h @@ -1654,11 +1654,15 @@ struct kvm_memory_attributes { #define KVM_CREATE_GUEST_MEMFD _IOWR(KVMIO, 0xd4, struct kvm_create_guest_memfd) #define GUEST_MEMFD_FLAG_MMAP (1ULL << 0) #define GUEST_MEMFD_FLAG_INIT_SHARED (1ULL << 1) +#define GUEST_MEMFD_FLAG_USE_DMABUF (1ULL << 2) struct kvm_create_guest_memfd { __u64 size; __u64 flags; - __u64 reserved[6]; + /* With USE_DMABUF, memory comes from this dma-buf's leading bytes. */ + __s32 dmabuf_fd; + __u32 pad; + __u64 reserved[5]; }; #define KVM_PRE_FAULT_MEMORY _IOWR(KVMIO, 0xd5, struct kvm_pre_fault_memory) diff --git a/tools/include/uapi/linux/kvm.h b/tools/include/uapi/linux/kvm.h index d0c0c8605976..ebeb2bc353ed 100644 --- a/tools/include/uapi/linux/kvm.h +++ b/tools/include/uapi/linux/kvm.h @@ -1644,11 +1644,15 @@ struct kvm_memory_attributes { #define KVM_CREATE_GUEST_MEMFD _IOWR(KVMIO, 0xd4, struct kvm_create_guest_memfd) #define GUEST_MEMFD_FLAG_MMAP (1ULL << 0) #define GUEST_MEMFD_FLAG_INIT_SHARED (1ULL << 1) +#define GUEST_MEMFD_FLAG_USE_DMABUF (1ULL << 2) struct kvm_create_guest_memfd { __u64 size; __u64 flags; - __u64 reserved[6]; + /* With USE_DMABUF, memory comes from this dma-buf's leading bytes. */ + __s32 dmabuf_fd; + __u32 pad; + __u64 reserved[5]; }; #define KVM_PRE_FAULT_MEMORY _IOWR(KVMIO, 0xd5, struct kvm_pre_fault_memory) diff --git a/virt/kvm/Kconfig b/virt/kvm/Kconfig index 794976b88c6f..7ac4c8a371eb 100644 --- a/virt/kvm/Kconfig +++ b/virt/kvm/Kconfig @@ -105,6 +105,7 @@ config KVM_GENERIC_MEMORY_ATTRIBUTES config KVM_GUEST_MEMFD select XARRAY_MULTI + select DMA_SHARED_BUFFER bool config HAVE_KVM_ARCH_GMEM_PREPARE diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c index 0f4bf2cc5e8e..f401cc560eb4 100644 --- a/virt/kvm/guest_memfd.c +++ b/virt/kvm/guest_memfd.c @@ -1,4 +1,7 @@ // SPDX-License-Identifier: GPL-2.0 +#include <linux/dma-buf.h> +#include <linux/dma-resv.h> +#include <linux/module.h> #include <linux/anon_inodes.h> #include <linux/backing-dev.h> #include <linux/falloc.h> @@ -44,6 +47,8 @@ struct gmem_inode { struct list_head gmem_file_list; u64 flags; + struct dma_buf *dmabuf; + struct dma_buf_attachment *attach; }; static __always_inline struct gmem_inode *GMEM_I(struct inode *inode) @@ -477,21 +482,37 @@ static const struct vm_operations_struct kvm_gmem_vm_ops = { #endif }; -static int kvm_gmem_native_mmap(struct file *file, struct vm_area_struct *vma) +static int kvm_gmem_validate_mmap(struct file *file, + struct vm_area_struct *vma) { if (!kvm_gmem_supports_mmap(file_inode(file))) return -ENODEV; - if ((vma->vm_flags & (VM_SHARED | VM_MAYSHARE)) != - (VM_SHARED | VM_MAYSHARE)) { + (VM_SHARED | VM_MAYSHARE)) return -EINVAL; - } + return 0; +} - vma->vm_ops = &kvm_gmem_vm_ops; +static int kvm_gmem_native_mmap(struct file *file, struct vm_area_struct *vma) +{ + int ret = kvm_gmem_validate_mmap(file, vma); + if (ret) + return ret; + vma->vm_ops = &kvm_gmem_vm_ops; return 0; } +static int kvm_gmem_dmabuf_mmap(struct file *file, struct vm_area_struct *vma) +{ + struct dma_buf *dmabuf = GMEM_I(file_inode(file))->dmabuf; + int ret = kvm_gmem_validate_mmap(file, vma); + + if (ret) + return ret; + return dma_buf_mmap(dmabuf, vma, vma->vm_pgoff); +} + /* File-op dispatchers — thin: they all go through the backing's ops table. */ static int kvm_gmem_fops_release(struct inode *inode, struct file *file) @@ -617,6 +638,172 @@ bool __weak kvm_arch_supports_gmem_init_shared(struct kvm *kvm) return true; } +/* ---- a dma-buf behind the file ----------------------------------------- */ + +/* + * Refuse a buffer whose first run guest_memfd can never map, such as MMIO, + * at creation rather than on every fault. A gap at offset 0 is legal. This + * is advisory: each fault checks the run it uses. + */ +static int gmem_dmabuf_check(struct dma_buf_attachment *attach, u64 size) +{ + struct phys_vec pv; + u32 attr; + int ret; + + dma_resv_lock(attach->dmabuf->resv, NULL); + ret = dma_buf_get_phys(attach, 0, size, &pv, &attr); + dma_resv_unlock(attach->dmabuf->resv); + if (ret == -ENOENT) + return 0; + if (ret) + return ret; + if (!dma_buf_phys_attr_known(attr) || + dma_buf_phys_attr_type(attr) != DMA_BUF_PHYS_ATTR_RAM) + return -EOPNOTSUPP; + return 0; +} + +static void gmem_dmabuf_invalidate_range(struct dma_buf_attachment *attach, + unsigned long offset, + unsigned long len) +{ + struct inode *inode = attach->importer_priv; + pgoff_t npages = i_size_read(inode) >> PAGE_SHIFT; + pgoff_t start = offset >> PAGE_SHIFT; + pgoff_t end = min_t(pgoff_t, DIV_ROUND_UP(offset + len, PAGE_SIZE), + npages); + + if (start >= end) + return; + /* Lock order: dma-buf reservation, file invalidation, then mmu_lock. */ + filemap_invalidate_lock(inode->i_mapping); + kvm_gmem_invalidate_start(inode, start, end); + kvm_gmem_invalidate_end(inode, start, end); + filemap_invalidate_unlock(inode->i_mapping); +} + +static void gmem_dmabuf_invalidate(struct dma_buf_attachment *attach) +{ + struct inode *inode = attach->importer_priv; + + gmem_dmabuf_invalidate_range(attach, 0, i_size_read(inode)); +} + +static const struct dma_buf_attach_ops gmem_dmabuf_attach_ops = { + .allow_peer2peer = true, + .invalidate_mappings = gmem_dmabuf_invalidate, + .invalidate_mappings_range = gmem_dmabuf_invalidate_range, +}; + +static int kvm_gmem_dmabuf_attach(struct inode *inode, int dmabuf_fd) +{ + struct gmem_inode *gi = GMEM_I(inode); + struct dma_buf_attachment *attach; + struct dma_buf *dmabuf; + int ret; + + dmabuf = dma_buf_get(dmabuf_fd); + if (IS_ERR(dmabuf)) + return PTR_ERR(dmabuf); + if (dmabuf->size < i_size_read(inode)) { + ret = -EINVAL; + goto put; + } + + gi->dmabuf = dmabuf; + attach = dma_buf_dynamic_attach(dmabuf, NULL, &gmem_dmabuf_attach_ops, + inode); + if (IS_ERR(attach)) { + ret = PTR_ERR(attach); + goto clear_put; + } + gi->attach = attach; + ret = gmem_dmabuf_check(attach, i_size_read(inode)); + if (ret) + goto detach; + return 0; + +detach: + dma_buf_detach(dmabuf, attach); + gi->attach = NULL; +clear_put: + gi->dmabuf = NULL; +put: + dma_buf_put(dmabuf); + return ret; +} + +static void kvm_gmem_dmabuf_release(struct inode *inode) +{ + struct gmem_inode *gi = GMEM_I(inode); + + if (!gi->dmabuf) + return; + dma_buf_detach(gi->dmabuf, gi->attach); + dma_buf_put(gi->dmabuf); + gi->attach = NULL; + gi->dmabuf = NULL; +} + +static int gmem_dmabuf_get_pfn(struct file *file, struct kvm *kvm, + struct kvm_memory_slot *slot, gfn_t gfn, + kvm_pfn_t *pfn, struct page **page, + int *max_order, bool *writable) +{ + static const int orders[] = { PUD_ORDER, PMD_ORDER, 0 }; + struct gmem_inode *gi = GMEM_I(file_inode(file)); + struct gmem_file *f = gmem_file_of(file); + pgoff_t index = kvm_gmem_get_index(slot, gfn); + pgoff_t npages = i_size_read(file_inode(file)) >> PAGE_SHIFT; + unsigned int oi; + int ret = -EFAULT; + + if (file != READ_ONCE(slot->gmem.file)) + return -EFAULT; + if (xa_load(&f->bindings, index) != slot) + return -EIO; + + /* A block maps at an order only if one aligned RAM run covers it. */ + dma_resv_lock(gi->dmabuf->resv, NULL); + for (oi = max_order ? 0 : ARRAY_SIZE(orders) - 1; + oi < ARRAY_SIZE(orders); oi++) { + int order = orders[oi]; + pgoff_t bs = ALIGN_DOWN(index, 1UL << order); + u64 blen = (u64)PAGE_SIZE << order; + struct phys_vec pv; + u32 attr; + + if (bs + (1UL << order) > npages) + continue; + if (dma_buf_get_phys(gi->attach, (u64)bs << PAGE_SHIFT, blen, + &pv, &attr) || + pv.len != blen || !IS_ALIGNED(pv.paddr, blen) || + !dma_buf_phys_attr_known(attr) || + dma_buf_phys_attr_type(attr) != DMA_BUF_PHYS_ATTR_RAM) + continue; + + *pfn = PHYS_PFN(pv.paddr) + (index - bs); + *page = NULL; + if (max_order) + *max_order = order; + if (writable && (attr & DMA_BUF_PHYS_ATTR_READONLY)) + *writable = false; + ret = 0; + break; + } + dma_resv_unlock(gi->dmabuf->resv); + return ret; +} + +static int kvm_gmem_dmabuf_populate(struct file *file, struct kvm *kvm, + struct kvm_memory_slot *slot, gfn_t gfn, + kvm_pfn_t *pfn, struct page *src_page, + int order) +{ + return -EOPNOTSUPP; +} + static int kvm_gmem_native_bind(struct file *file, struct kvm *kvm, struct kvm_memory_slot *slot, loff_t offset); static void kvm_gmem_native_unbind(struct file *slot_file, struct kvm *kvm, @@ -654,7 +841,17 @@ static const struct kvm_gmem_ops kvm_gmem_native_ops = { .fallocate = kvm_gmem_native_fallocate, }; -static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags) +static const struct kvm_gmem_ops kvm_gmem_dmabuf_ops = { + .bind = kvm_gmem_native_bind, + .unbind = kvm_gmem_native_unbind, + .get_pfn = gmem_dmabuf_get_pfn, + .populate = kvm_gmem_dmabuf_populate, + .release = kvm_gmem_native_release, + .mmap = kvm_gmem_dmabuf_mmap, +}; + +static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags, + int dmabuf_fd) { static const char *name = "[kvm-gmem]"; struct gmem_file *f; @@ -695,6 +892,12 @@ static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags) GMEM_I(inode)->flags = flags; + if (flags & GUEST_MEMFD_FLAG_USE_DMABUF) { + err = kvm_gmem_dmabuf_attach(inode, dmabuf_fd); + if (err) + goto err_inode; + } + file = alloc_file_pseudo(inode, kvm_gmem_mnt, name, O_RDWR, &kvm_gmem_fops); if (IS_ERR(file)) { err = PTR_ERR(file); @@ -703,12 +906,17 @@ static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags) file->f_flags |= O_LARGEFILE; file->private_data = &f->backing; - f->backing.ops = &kvm_gmem_native_ops; + f->backing.ops = flags & GUEST_MEMFD_FLAG_USE_DMABUF ? + &kvm_gmem_dmabuf_ops : &kvm_gmem_native_ops; kvm_get_kvm(kvm); f->kvm = kvm; xa_init(&f->bindings); + if (flags & GUEST_MEMFD_FLAG_USE_DMABUF) + filemap_invalidate_lock(inode->i_mapping); list_add(&f->entry, &GMEM_I(inode)->gmem_file_list); + if (flags & GUEST_MEMFD_FLAG_USE_DMABUF) + filemap_invalidate_unlock(inode->i_mapping); fd_install(fd, file); return fd; @@ -734,8 +942,11 @@ int kvm_gmem_create(struct kvm *kvm, struct kvm_create_guest_memfd *args) if (size <= 0 || !PAGE_ALIGNED(size)) return -EINVAL; + if (args->pad || + (!(flags & GUEST_MEMFD_FLAG_USE_DMABUF) && args->dmabuf_fd)) + return -EINVAL; - return __kvm_gmem_create(kvm, size, flags); + return __kvm_gmem_create(kvm, size, flags, args->dmabuf_fd); } /* @@ -1211,6 +1422,8 @@ static struct inode *kvm_gmem_alloc_inode(struct super_block *sb) mpol_shared_policy_init(&gi->policy, NULL); gi->flags = 0; + gi->dmabuf = NULL; + gi->attach = NULL; INIT_LIST_HEAD(&gi->gmem_file_list); return &gi->vfs_inode; } @@ -1225,11 +1438,19 @@ static void kvm_gmem_free_inode(struct inode *inode) kmem_cache_free(kvm_gmem_inode_cachep, GMEM_I(inode)); } +static void kvm_gmem_evict_inode(struct inode *inode) +{ + kvm_gmem_dmabuf_release(inode); + truncate_inode_pages_final(&inode->i_data); + clear_inode(inode); +} + static const struct super_operations kvm_gmem_super_operations = { .statfs = simple_statfs, .alloc_inode = kvm_gmem_alloc_inode, .destroy_inode = kvm_gmem_destroy_inode, .free_inode = kvm_gmem_free_inode, + .evict_inode = kvm_gmem_evict_inode, }; static int kvm_gmem_init_fs_context(struct fs_context *fc) @@ -1292,3 +1513,5 @@ void kvm_gmem_exit(void) rcu_barrier(); kmem_cache_destroy(kvm_gmem_inode_cachep); } + +MODULE_IMPORT_NS("DMA_BUF"); -- 2.47.3 ^ permalink raw reply [flat|nested] 28+ messages in thread
* [RFC PATCH 6/6] samples/kvm, selftests/kvm: Exercise dma-buf backing 2026-10-05 9:55 ` [RFC PATCH 0/6] KVM: guest_memfd: back guest_memfd with an imported dma-buf Fred Griffoul ` (4 preceding siblings ...) 2026-10-05 9:55 ` [RFC PATCH 5/6] KVM: guest_memfd: Add dma-buf backing Fred Griffoul @ 2026-10-05 9:55 ` Fred Griffoul 5 siblings, 0 replies; 28+ messages in thread From: Fred Griffoul @ 2026-10-05 9:55 UTC (permalink / raw) To: Paolo Bonzini, Sean Christopherson, Marc Zyngier, Oliver Upton, Sumit Semwal, Christian König, Jason Gunthorpe, Kevin Tian Cc: David Woodhouse, Ackerley Tng, Joey Gouly, Suzuki K Poulose, Zenghui Yu, Steffen Eiden, Catalin Marinas, Will Deacon, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Dave Hansen, H . Peter Anvin, Joerg Roedel, Robin Murphy, Alex Williamson, Shuah Khan, Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers, linux-kernel, kvm, kvmarm, linux-arm-kernel, iommu, linux-media, dri-devel, linaro-mm-sig, linux-kselftest, linux-trace-kernel, x86 From: Fred Griffoul <fgriffo@amazon.co.uk> Turn the sample into a dma-buf exporter, and move every supported test to dma-buf-backed guest_memfd. Doing both in one commit avoids a state where the sample and the tests use different interfaces. The sample owns one root region and creates a child descriptor for each VM. It supports moving, donating and reclaiming pages, absent pages, read-only ranges, a scratch page and per-child ioctl allowlists. The sample's get_phys() reports the run at the requested offset. It uses bitmap searches to find where presence or read-only state changes. An absent page returns -ENOENT. In scratch mode, an absent page of a root child is reported as the read-only scratch page instead. Scratch mode applies only to root children, because a mode change invalidates only them. Every change of ownership sends a ranged invalidation. Each test passes one dma-buf fd to both guest_memfd and iommufd. New tests cover: - two VMMs, checking what the guest, the device and the host see; - SET_PRESENT and SET_READONLY racing with one KVM_CREATE_GUEST_MEMFD, so that lockdep checks attachment setup; - a 2 MiB-aligned region in which every other page is replaced by the scratch page. The guest reads zeros and its own data, and KVM maps the region with 4 KiB pages. An untouched aligned region still maps at 2 MiB. The sample sizes the CMA root before children overlap, and publishes each child only after its ownership and allowlist are set. Confidential VMs are out of scope; their test is an explicit skip. Signed-off-by: Fred Griffoul <fgriffo@amazon.co.uk> --- samples/kvm/gmem_provider.c | 1287 ++++++++++------- samples/kvm/gmem_provider.h | 144 +- tools/testing/selftests/kvm/Makefile.kvm | 3 +- .../kvm/gmem_provider_nvme_dma_test.c | 28 +- .../testing/selftests/kvm/include/kvm_util.h | 25 +- .../testing/selftests/kvm/x86/gmem_poc_test.c | 824 +++++++++++ .../kvm/x86/gmem_provider_hugepage_test.c | 30 +- .../kvm/x86/gmem_provider_iommufd_test.c | 42 +- .../kvm/x86/gmem_provider_readonly_test.c | 172 +++ .../kvm/x86/gmem_provider_revoke_test.c | 34 +- .../selftests/kvm/x86/gmem_provider_test.c | 190 +-- .../kvm/x86/gmem_provider_vfio_test.c | 45 +- 12 files changed, 2091 insertions(+), 733 deletions(-) create mode 100644 tools/testing/selftests/kvm/x86/gmem_poc_test.c create mode 100644 tools/testing/selftests/kvm/x86/gmem_provider_readonly_test.c diff --git a/samples/kvm/gmem_provider.c b/samples/kvm/gmem_provider.c index b6824fe5d228..2cde5bcb71ca 100644 --- a/samples/kvm/gmem_provider.c +++ b/samples/kvm/gmem_provider.c @@ -1,43 +1,46 @@ // SPDX-License-Identifier: GPL-2.0 /* - * gmem_provider - sample external guest_memfd provider. + * gmem_provider - sample owner of page-less memory, shared as a dma-buf. * - * Demonstrates the KVM guest_memfd provider ABI with two backing modes: + * A toy owner of physical memory: one root region carved into children, one + * child per VM, ranges that MOVE between children or are DONATEd to the root + * and RECLAIMed, a shared scratch frame for revoked pages, and a per-fd + * ioctl allowlist fixed by the creator. The provider owns the frames and is + * the single authority for who may map them. * - * - External (page-less): loaded with addr=/len=, backs guest memory with a - * fixed physical range that has no struct page -- e.g. memory carved out of - * the kernel with mem= on the command line. This is the case the provider - * ABI exists for; get_pfn() returns bare PFNs KVM treats as non-refcounted. + * It is a dma-buf exporter and nothing else. A child fd hands out a dma-buf + * for the child's memory; a VMM gives that dma-buf to KVM_CREATE_GUEST_MEMFD + * for the guest and to iommufd for its devices. An ownership change here is + * one ranged invalidation of the child's dma-bufs; every importer drops the + * changed view and re-reads the layout. The owner knows nothing of KVM, or which + * importer is a guest and which is a device. * - * - CMA fallback (page-backed): when addr=/len= are not given, allocates a - * physically contiguous region via alloc_contig_pages() sized by the setup - * ioctl. Easier to run (no mem= boot param), and still exercises 2M/1G - * mappings. Not preserved across kexec/live update. - * - * On SEV-SNP hosts, bind() resets the range's RMP entries to 4K shared so a - * (new) SNP VM can re-encrypt it, which is what allows re-binding the range to - * a fresh VM across a live update. + * Backing: + * - External (page-less): loaded with addr=/len=, a fixed physical range with + * no struct page, e.g. carved out with memmap= on the command line. + * - CMA fallback: without addr=/len=, alloc_contig_pages() sized by SETUP. * * Usage: - * # page-less external range: - * insmod gmem_provider.ko addr=0x5D40000000 len=0x1000000 - * # or CMA fallback (no params); size comes from the ioctl - * insmod gmem_provider.ko - * fd = open("/dev/gmem_provider"); ioctl(fd, GMEM_PROVIDER_SETUP, {kvm_fd, size}); - * pass the returned fd + KVM_MEM_GUEST_MEMFD to KVM_SET_USER_MEMORY_REGION2. + * insmod gmem_provider.ko addr=0x180000000 len=0x42000000 + * ctl = open("/dev/gmem_provider"); + * child = ioctl(ctl, GMEM_PROVIDER_NEW_CHILD, {offset, len, allow}); # control + * dmabuf = ioctl(child, GMEM_PROVIDER_GET_DMABUF); # VMM + * gmem = ioctl(vm, KVM_CREATE_GUEST_MEMFD, {size, USE_DMABUF, dmabuf}); + * KVM_SET_USER_MEMORY_REGION2(..., KVM_MEM_GUEST_MEMFD, gmem); + * IOMMU_IOAS_MAP_FILE(ioas, dmabuf, ...); + * + * The one rule: change ownership state, drop the provider lock, then revoke. */ #include <linux/anon_inodes.h> -#include <linux/bitmap.h> #include <linux/dma-buf.h> -#include <linux/dma-buf-mapping.h> #include <linux/dma-resv.h> +#include <linux/bitmap.h> #include <linux/file.h> #include <linux/fs.h> #include <linux/gfp.h> #include <linux/highmem.h> #include <linux/io.h> -#include <linux/kvm_host.h> #include <linux/miscdevice.h> #include <linux/mm.h> #include <linux/module.h> @@ -50,7 +53,6 @@ #include "gmem_provider.h" -MODULE_IMPORT_NS("DMA_BUF"); static unsigned long long addr; module_param(addr, ullong, 0444); @@ -63,444 +65,294 @@ MODULE_PARM_DESC(len, "size in bytes of the external backing region (optional)") struct gmem_info { /* - * MUST be first: file->private_data points here. is_kvm_gmem_file() - * on the KVM side proves the reinterpretation is safe. - */ - struct kvm_gmem_backing backing; - - /* - * Protects everything below (except the immutable base_pfn/npages/ - * cma_pages fields set at setup). Ordering: info->lock is a leaf; - * do not acquire other locks under it. KVM's slots_lock is already - * held on the .bind/.unbind paths so info->lock is only needed to - * serialise those against ioctl(SET_PRESENT). + * Protects everything below except the immutable fields set at setup + * and the dma-buf list. Ordering: gmem_root.lock, then + * info->dmabufs_lock, then a dma-buf reservation. info->lock is a + * leaf: it is dropped before any invalidation. */ struct mutex lock; - /* Kept locally: kvm_gmem_ops does not carry a kvm pointer. */ - struct kvm *kvm; - - /* - * Single active memslot binding. Used at release time to zap any - * still-active guest mappings before the memory disappears. NULL - * when no memslot is bound (or the last one has been unbound). - */ - struct kvm_memory_slot *bound_slot; - bool mmap_capable; /* -> KVM_MEMSLOT_GMEM_ONLY at bind */ + bool mmap_capable; /* child and dma-buf may be mmap()ed */ unsigned long base_pfn; unsigned long npages; - struct page *cma_pages; /* non-NULL if CMA-allocated */ - gfn_t base_gfn; /* recorded at bind, for revoke */ - pgoff_t pgoff; /* provider offset (pages) of the slot */ + struct page *cma_pages; /* non-NULL if CMA-allocated (SETUP path only) */ + + /* + * Toy descriptor tree. A child created by NEW_CHILD is a sub-range + * of the root region: @root_index is its first page within the root, + * @owned marks which of its pages are currently granted to it (a page + * MOVEd out or DONATEd is not owned; a page revoked by SET_PRESENT is + * owned but absent). @allow is the ioctl allowlist. SETUP-created + * providers have no root and own everything. + */ + unsigned long root_index; + unsigned long *owned; /* NULL for SETUP-created providers */ + u32 allow; + struct list_head root_link; /* gmem_root.children */ unsigned long *absent; /* bitmap of currently-revoked pages */ - struct list_head dmabufs; /* struct gmem_dmabuf entries */ - struct mutex dmabufs_lock; + unsigned long *readonly; /* bitmap of pages the guest may not write */ + struct address_space *mapping; /* our file's: host windows live here */ + struct list_head dmabufs; /* exported dma-bufs (struct gmem_dmabuf) */ + struct mutex dmabufs_lock; /* protects @dmabufs */ }; static struct gmem_info *to_gmem_info(struct file *file) { - return container_of(file->private_data, struct gmem_info, backing); + return file->private_data; } -/* Map a backing PFN for CPU access: page-backed via kmap, page-less via memremap. */ -static void *gmem_map_pfn(kvm_pfn_t pfn) -{ - if (pfn_valid(pfn)) - return kmap_local_pfn(pfn); - return memremap(PFN_PHYS(pfn), PAGE_SIZE, MEMREMAP_WB); -} +/* + * The root of the toy descriptor tree: one backing region owned by the + * control device. Children carve sub-ranges out of it. All ownership + * transitions (NEW_CHILD, MOVE, DONATE, RECLAIM) run under root.lock, which + * is taken before any child's info->lock. + */ +static struct gmem_root { + struct mutex lock; /* protects the fields below */ + unsigned long base_pfn; + unsigned long npages; + struct page *cma_pages; + unsigned long *owned; /* pages some child currently owns */ + unsigned long *donated; /* pages parked at the root */ + struct list_head children; + bool scratch_enabled; +} gmem_root; -static void gmem_unmap_pfn(kvm_pfn_t pfn, void *vaddr) -{ - if (!vaddr) - return; - if (pfn_valid(pfn)) - kunmap_local(vaddr); - else - memunmap(vaddr); -} +/* + * One zeroed scratch frame for the whole module. When scratch mode is on, + * a revoked page of a root child reports this frame read-only to every + * importer, so a device that cannot tolerate an IOMMU fault lands here. + */ +static struct page *gmem_scratch_page; -/* Largest order KVM may map at @gfn, snapped to 4K/2M/1G. */ -static int gmem_max_order(struct gmem_info *info, gfn_t gfn, unsigned long index) +static inline unsigned long gmem_scratch_pfn(void) { - unsigned long pfn = info->base_pfn + index; - unsigned long remaining = info->npages - index; - unsigned int pud_order = PUD_SHIFT - PAGE_SHIFT; - unsigned int pmd_order = PMD_SHIFT - PAGE_SHIFT; - unsigned long absent_next; - - /* - * A hugepage may not span any revoked page. Clamp by the distance to - * the next absent bit; scanning is cheap because @absent is a plain - * bitmap and the caller already checked test_bit(index). - */ - if (info->absent) { - absent_next = find_next_bit(info->absent, info->npages, - index + 1); - remaining = min(remaining, absent_next - index); - } - - if (IS_ALIGNED(pfn, 1UL << pud_order) && - IS_ALIGNED(gfn, 1UL << pud_order) && - remaining >= (1UL << pud_order)) - return pud_order; - - if (IS_ALIGNED(pfn, 1UL << pmd_order) && - IS_ALIGNED(gfn, 1UL << pmd_order) && - remaining >= (1UL << pmd_order)) - return pmd_order; - - return 0; + return page_to_pfn(gmem_scratch_page); } -static int gmem_get_pfn(struct file *file, struct kvm *kvm, - struct kvm_memory_slot *slot, gfn_t gfn, - kvm_pfn_t *pfn, struct page **page, int *max_order, - bool *writable) +/* + * Is page @index of @info currently reachable by the guest and devices? + * Callable with or without info->lock: every change to the bitmaps is + * followed by an invalidation of the range, so a lock-free answer is either + * current or about to be superseded. + */ +static inline bool gmem_page_present(struct gmem_info *info, unsigned long index) { - struct gmem_info *info = to_gmem_info(file); - pgoff_t index = gfn - slot->base_gfn + slot->gmem.pgoff; - - if (index >= info->npages) - return -EINVAL; - - /* Revoked (absent) page: behave like not-present so the fault fails. */ + if (info->owned && !test_bit(index, info->owned)) + return false; if (info->absent && test_bit(index, info->absent)) - return -EFAULT; - - *pfn = info->base_pfn + index; - if (max_order) - *max_order = gmem_max_order(info, gfn, index); - return 0; + return false; + return true; } -static int gmem_populate(struct file *file, struct kvm *kvm, - struct kvm_memory_slot *slot, gfn_t gfn, - kvm_pfn_t *pfn, struct page *src_page, int order) -{ - struct gmem_info *info = to_gmem_info(file); - pgoff_t index = gfn - slot->base_gfn + slot->gmem.pgoff; - - if (index >= info->npages) - return -EINVAL; - - *pfn = info->base_pfn + index; +static void gmem_release(struct file *file); +static int gmem_mmap(struct file *file, struct vm_area_struct *vma); +static long gmem_fd_ioctl(struct file *file, unsigned int cmd, unsigned long arg); - if (src_page) { - void *dst, *src; - /* Map dst first: memremap() may sleep, kmap_local_page() must not. */ - dst = gmem_map_pfn(*pfn); - if (!dst) - return -ENOMEM; - src = kmap_local_page(src_page); - memcpy(dst, src, PAGE_SIZE); - kunmap_local(src); - gmem_unmap_pfn(*pfn, dst); - } - return 0; -} - -static int gmem_bind(struct file *file, struct kvm *kvm, - struct kvm_memory_slot *slot, loff_t offset) +static void gmem_release(struct file *file) { struct gmem_info *info = to_gmem_info(file); - struct kvm *old_kvm = NULL; - unsigned long start = offset >> PAGE_SHIFT; - - if (offset < 0 || !PAGE_ALIGNED(offset) || - start + slot->npages > info->npages) - return -EINVAL; /* - * An mmap-capable backing hands the VMM a host mapping onto pages - * that a hardware-encrypted VM (SEV, SEV-ES, SEV-SNP, TDX) will mark - * private in the RMP/EPT; the mmap can only ever fault on those - * pages. Refuse rather than hand the VMM a useless (and misleading) - * shared view. SW_PROTECTED_VM has no hardware encryption and is - * fine. + * Every exported dma-buf holds a reference on this file, so none can + * be alive here: no device and no guest_memfd still maps our frames. */ -#ifdef CONFIG_X86 - if (info->mmap_capable && - (kvm->arch.vm_type == KVM_X86_SEV_VM || - kvm->arch.vm_type == KVM_X86_SEV_ES_VM || - kvm->arch.vm_type == KVM_X86_SNP_VM || - kvm->arch.vm_type == KVM_X86_TDX_VM)) - return -EACCES; -#endif + WARN_ON(!list_empty(&info->dmabufs)); - /* Record the binding so the revoke ioctl can translate offset -> gfn. */ - mutex_lock(&info->lock); - info->base_gfn = slot->base_gfn; - info->pgoff = start; - info->bound_slot = slot; - mutex_unlock(&info->lock); - -#if IS_ENABLED(CONFIG_AMD_MEM_ENCRYPT) /* - * Reset the RMP for the range to 4K shared so a (new) SEV-SNP VM can - * transition it to private and re-encrypt it. PSMASH any 2M entries - * first. Harmless on non-SNP hosts, where these return -ENODEV. + * A child returns its carved range to the root. Pages it still owned + * become free again; pages it had DONATEd stay parked at the root + * (still in gmem_root.donated) until RECLAIM or module exit; pages + * MOVEd out belong to another child and are not ours to free. */ - { + if (info->owned) { unsigned long i; - unsigned long first_pmd_pfn = ALIGN(info->base_pfn + start, - PTRS_PER_PMD); - - for (i = first_pmd_pfn - (info->base_pfn + start); - i + PTRS_PER_PMD <= slot->npages; - i += PTRS_PER_PMD) - psmash(info->base_pfn + start + i); - for (i = 0; i < slot->npages; i++) { - unsigned long pfn = info->base_pfn + start + i; - int ret = rmp_make_shared(pfn, PG_LEVEL_4K); - - if (ret && ret != -ENODEV) - pr_info_once("gmem_provider: rmp_make_shared(0x%lx) = %d\n", - pfn, ret); - } + mutex_lock(&gmem_root.lock); + list_del(&info->root_link); + for_each_set_bit(i, info->owned, info->npages) + __clear_bit(info->root_index + i, gmem_root.owned); + mutex_unlock(&gmem_root.lock); + kvfree(info->owned); } -#endif - - /* - * Claim (or, on re-bind, transfer) VM ownership; pins the VM. - * kvm_put_kvm(old_kvm) MUST run outside info->lock: if the put - * drops the last ref, kvm_destroy_vm() runs inline and calls back - * into our gmem_unbind() (via kvm_gmem_unbind() on each memslot), - * which needs info->lock -- taking it here would self-deadlock. - */ - mutex_lock(&info->lock); - if (info->kvm != kvm) { - old_kvm = info->kvm; - kvm_get_kvm(kvm); - info->kvm = kvm; - } - mutex_unlock(&info->lock); - if (old_kvm) - kvm_put_kvm(old_kvm); - - /* - * Record the file on the slot (KVM's outer bind no longer does this - * for us) and mark the slot gmem-only if this backing serves host - * accesses through its own mmap. - */ - WRITE_ONCE(slot->gmem.file, file); - slot->gmem.pgoff = start; - if (info->mmap_capable) - slot->flags |= KVM_MEMSLOT_GMEM_ONLY; - - return 0; + if (info->cma_pages) + free_contig_range(info->base_pfn, info->npages); + kvfree(info->absent); + kvfree(info->readonly); + kfree(info); + module_put(THIS_MODULE); } -static void gmem_unbind(struct file *file, struct kvm *kvm, - struct kvm_memory_slot *slot) +static vm_fault_t gmem_vm_fault(struct vm_fault *vmf) { - struct gmem_info *info = to_gmem_info(file); - - mutex_lock(&info->lock); - if (info->bound_slot == slot) - info->bound_slot = NULL; - mutex_unlock(&info->lock); - -#if IS_ENABLED(CONFIG_AMD_MEM_ENCRYPT) - { - unsigned long start = slot->gmem.pgoff; - unsigned long first_pmd_pfn = ALIGN(info->base_pfn + start, - PTRS_PER_PMD); - unsigned long i; - - /* - * Symmetric with gmem_bind(): PSMASH any 2M RMP entries first - * so rmp_make_shared(PG_LEVEL_4K) can succeed on SNP hosts. - * Harmless on non-SNP: psmash() returns -ENODEV. - */ - for (i = first_pmd_pfn - (info->base_pfn + start); - i + PTRS_PER_PMD <= slot->npages; - i += PTRS_PER_PMD) - psmash(info->base_pfn + start + i); + /* The VMA's file is ours or a dma-buf's; the child rides in vm_private_data. */ + struct gmem_info *info = vmf->vma->vm_private_data; + unsigned long index = vmf->pgoff; - for (i = 0; i < slot->npages; i++) { - unsigned long pfn = info->base_pfn + start + i; - - rmp_make_shared(pfn, PG_LEVEL_4K); - } - } -#endif + if (index >= info->npages || !gmem_page_present(info, index)) + return VM_FAULT_SIGBUS; + return vmf_insert_pfn(vmf->vma, vmf->address, info->base_pfn + index); } -static void gmem_release(struct file *file); -static int gmem_mmap(struct file *file, struct vm_area_struct *vma); -static long gmem_fd_ioctl(struct file *file, unsigned int cmd, unsigned long arg); - -static const struct kvm_gmem_ops gmem_ops = { - .bind = gmem_bind, - .unbind = gmem_unbind, - .get_pfn = gmem_get_pfn, - .populate = gmem_populate, - .release = gmem_release, - .mmap = gmem_mmap, - .ioctl = gmem_fd_ioctl, +static const struct vm_operations_struct gmem_vm_ops = { + .fault = gmem_vm_fault, }; -static void gmem_release(struct file *file) -{ - struct gmem_info *info = to_gmem_info(file); - struct kvm_memory_slot *slot; - - /* - * If a memslot is still bound at close time, KVM has not yet had a - * chance to call ops->unbind. Zap the guest mappings for the range - * and clear slot->gmem.file so the eventual unbind is a no-op. This - * matches native gmem's kvm_gmem_release() and prevents the guest - * from continuing to hit backing memory after we free it below. - */ - mutex_lock(&info->lock); - slot = info->bound_slot; - if (slot && info->kvm) { - kvm_gmem_invalidate_range(info->kvm, slot->base_gfn, - slot->base_gfn + slot->npages); - WRITE_ONCE(slot->gmem.file, NULL); - info->bound_slot = NULL; - } - mutex_unlock(&info->lock); - - if (info->kvm) - kvm_put_kvm(info->kvm); - if (info->cma_pages) - free_contig_range(info->base_pfn, info->npages); - kvfree(info->absent); - kfree(info); - module_put(THIS_MODULE); -} - static int gmem_mmap(struct file *file, struct vm_area_struct *vma) { struct gmem_info *info = to_gmem_info(file); unsigned long npages = vma_pages(vma); - /* - * gmem_bind() refuses to bind an mmap-capable fd to a coco VM. A fd - * that was created without GMEM_PROVIDER_FLAG_MMAP_CAPABLE has no - * such gate at bind time, so its mmap must not succeed at any point. - */ + /* A fd created without GMEM_PROVIDER_FLAG_MMAP_CAPABLE is never host-mappable. */ if (!info->mmap_capable) return -EPERM; if (vma->vm_pgoff + npages > info->npages) return -EINVAL; - /* Page-less backing: map raw PFNs, not folios. */ + /* + * Page-less backing: raw PFNs, inserted on fault, never at mmap() + * time, so that a window torn down on a revoke comes back by itself + * once the page is the child's again. + */ vm_flags_set(vma, VM_PFNMAP | VM_IO | VM_DONTEXPAND | VM_DONTDUMP); - return remap_pfn_range(vma, vma->vm_start, info->base_pfn + vma->vm_pgoff, - npages << PAGE_SHIFT, vma->vm_page_prot); + vma->vm_private_data = info; + vma->vm_ops = &gmem_vm_ops; + return 0; } /* - * Dynamic dma-buf exporter over the provider's backing. - * - * Follows the same shape as drivers/vfio/pci/vfio_pci_dmabuf.c: a per-dmabuf - * priv holding a phys_vec, a revocable dynamic attach, and a "private - * interconnect" symbol iommufd looks up to fetch phys directly (instead of - * mapping through the DMA API). + * The dma-buf a child exports. * - * A revoke on the provider (SET_PRESENT present=0) fans out to every exported - * dma-buf via dma_buf_invalidate_mappings(), so iommufd (which registered a - * revocable importer) tears down the IOMMU mapping alongside KVM's NPT zap. + * A child may hand out several dma-bufs over its life (one per holder of the + * fd who asks); each covers the whole child and pins this file until it is + * released. Importers must accept ranged invalidation: every change to the + * child's pages is sent to them, and they re-read get_phys(). */ struct gmem_dmabuf { struct dma_buf *dmabuf; struct gmem_info *info; struct file *provider_file; /* holds info alive */ struct list_head list; /* info->dmabufs */ - struct phys_vec phys; /* single contiguous range */ - struct kref kref; - struct completion comp; - bool revoked; }; static int gmem_dma_buf_attach(struct dma_buf *dmabuf, struct dma_buf_attachment *attach) { - struct gmem_dmabuf *priv = dmabuf->priv; - - if (!attach->peer2peer) - return -EOPNOTSUPP; - if (priv->revoked) - return -ENODEV; - if (!dma_buf_attach_revocable(attach)) + /* Only importers that can be told to re-read may attach. */ + if (!attach->peer2peer || !dma_buf_attach_revocable(attach)) return -EOPNOTSUPP; return 0; } -static void gmem_dma_buf_done(struct kref *kref) -{ - struct gmem_dmabuf *priv = container_of(kref, struct gmem_dmabuf, kref); - - complete(&priv->comp); -} - static struct sg_table *gmem_dma_buf_map(struct dma_buf_attachment *attach, enum dma_data_direction dir) { - struct gmem_dmabuf *priv = attach->dmabuf->priv; - struct sg_table *sgt; - - dma_resv_assert_held(priv->dmabuf->resv); - if (priv->revoked) - return ERR_PTR(-ENODEV); - - /* RAM, not P2P MMIO: no p2pdma_provider. */ - sgt = dma_buf_phys_vec_to_sgt(attach, NULL, &priv->phys, 1, - priv->phys.len, dir); - if (IS_ERR(sgt)) - return sgt; - - kref_get(&priv->kref); - return sgt; + /* DMA-API importers are not served; use get_phys() (iommufd does). */ + return ERR_PTR(-EOPNOTSUPP); } static void gmem_dma_buf_unmap(struct dma_buf_attachment *attach, - struct sg_table *sgt, - enum dma_data_direction dir) + struct sg_table *sgt, enum dma_data_direction dir) { - struct gmem_dmabuf *priv = attach->dmabuf->priv; +} + +/* A host window through the dma-buf: the same rules as a window on the child. */ +static int gmem_dma_buf_mmap(struct dma_buf *dmabuf, struct vm_area_struct *vma) +{ + struct gmem_dmabuf *priv = dmabuf->priv; - dma_resv_assert_held(priv->dmabuf->resv); - dma_buf_free_sgt(attach, sgt, dir); - kref_put(&priv->kref, gmem_dma_buf_done); + return gmem_mmap(priv->provider_file, vma); } static void gmem_dma_buf_release(struct dma_buf *dmabuf) { struct gmem_dmabuf *priv = dmabuf->priv; + struct gmem_info *info = priv->info; - if (priv->info) { - mutex_lock(&priv->info->dmabufs_lock); - list_del_init(&priv->list); - mutex_unlock(&priv->info->dmabufs_lock); - } - if (priv->provider_file) - fput(priv->provider_file); + mutex_lock(&info->dmabufs_lock); + list_del(&priv->list); + mutex_unlock(&info->dmabufs_lock); + fput(priv->provider_file); kfree(priv); } -/* Report this flat sample region through the generic dma-buf operation. */ +/* Return the first bit whose value differs from @index. */ +static unsigned long gmem_bitmap_next_change(const unsigned long *bitmap, + unsigned long nbits, + unsigned long index) +{ + if (!bitmap) + return nbits; + if (test_bit(index, bitmap)) + return find_next_zero_bit(bitmap, nbits, index + 1); + return find_next_bit(bitmap, nbits, index + 1); +} + +/* Find a run with one present/read-only disposition, without a per-bit scan. */ +static unsigned long gmem_disposition_end(struct gmem_info *info, + unsigned long index) +{ + bool present = gmem_page_present(info, index); + bool readonly = info->readonly && test_bit(index, info->readonly); + unsigned long pos = index; + + for (;;) { + unsigned long next = info->npages; + + next = min(next, gmem_bitmap_next_change(info->owned, + info->npages, pos)); + next = min(next, gmem_bitmap_next_change(info->absent, + info->npages, pos)); + next = min(next, gmem_bitmap_next_change(info->readonly, + info->npages, pos)); + if (next >= info->npages || + gmem_page_present(info, next) != present || + (!!(info->readonly && test_bit(next, info->readonly))) != readonly) + return next; + pos = next; + } +} + +/* + * Describe the run at @offset from the child's ownership bitmaps. A present + * page extends to the next page whose presence or read-only state differs, + * found with bitmap searches. An absent page is not backed, unless scratch + * mode substitutes the shared read-only scratch frame for that one page. + * Scratch mode covers only children of the root: SET_SCRATCH invalidates + * those, so a SETUP provider must never report the scratch frame. + */ static int gmem_dma_buf_get_phys(struct dma_buf_attachment *attach, u64 offset, u64 len, struct phys_vec *phys, u32 *attr) { struct gmem_dmabuf *priv = attach->dmabuf->priv; + struct gmem_info *info = priv->info; + unsigned long index = offset >> PAGE_SHIFT; + unsigned long end, run_end; - dma_resv_assert_held(attach->dmabuf->resv); - if (priv->revoked) - return -ENODEV; + if (!PAGE_ALIGNED(offset) || !PAGE_ALIGNED(len)) + return -EINVAL; + end = (offset + len) >> PAGE_SHIFT; + + if (!gmem_page_present(info, index)) { + if (!info->owned || !READ_ONCE(gmem_root.scratch_enabled)) + return -ENOENT; + phys->paddr = PFN_PHYS(gmem_scratch_pfn()); + phys->len = PAGE_SIZE; + *attr = DMA_BUF_PHYS_ATTR_RAM | DMA_BUF_PHYS_ATTR_READONLY; + return 0; + } - phys->paddr = priv->phys.paddr + offset; - phys->len = len; + run_end = min(gmem_disposition_end(info, index), end); + phys->paddr = PFN_PHYS(info->base_pfn + index); + phys->len = (u64)(run_end - index) << PAGE_SHIFT; *attr = DMA_BUF_PHYS_ATTR_RAM; + if (info->readonly && test_bit(index, info->readonly)) + *attr |= DMA_BUF_PHYS_ATTR_READONLY; return 0; } @@ -510,23 +362,9 @@ static const struct dma_buf_ops gmem_dma_buf_ops = { .unmap_dma_buf = gmem_dma_buf_unmap, .release = gmem_dma_buf_release, .get_phys = gmem_dma_buf_get_phys, + .mmap = gmem_dma_buf_mmap, }; -/* Called with info->dmabufs_lock held on the revoke path. */ -static void gmem_dma_buf_revoke_all(struct gmem_info *info) -{ - struct gmem_dmabuf *priv; - - list_for_each_entry(priv, &info->dmabufs, list) { - dma_resv_lock(priv->dmabuf->resv, NULL); - if (!priv->revoked) { - priv->revoked = true; - dma_buf_invalidate_mappings(priv->dmabuf); - } - dma_resv_unlock(priv->dmabuf->resv); - } -} - static int gmem_provider_get_dmabuf(struct file *file) { struct gmem_info *info = to_gmem_info(file); @@ -537,27 +375,19 @@ static int gmem_provider_get_dmabuf(struct file *file) priv = kzalloc(sizeof(*priv), GFP_KERNEL); if (!priv) return -ENOMEM; - priv->info = info; priv->provider_file = get_file(file); - priv->phys.paddr = (u64)info->base_pfn << PAGE_SHIFT; - priv->phys.len = (u64)info->npages << PAGE_SHIFT; - kref_init(&priv->kref); - init_completion(&priv->comp); - INIT_LIST_HEAD(&priv->list); - exp_info.ops = &gmem_dma_buf_ops; - exp_info.size = priv->phys.len; + exp_info.size = (u64)info->npages << PAGE_SHIFT; exp_info.flags = O_RDWR; exp_info.priv = priv; - priv->dmabuf = dma_buf_export(&exp_info); if (IS_ERR(priv->dmabuf)) { fd = PTR_ERR(priv->dmabuf); + fput(priv->provider_file); kfree(priv); return fd; } - mutex_lock(&info->dmabufs_lock); list_add(&priv->list, &info->dmabufs); mutex_unlock(&info->dmabufs_lock); @@ -568,14 +398,132 @@ static int gmem_provider_get_dmabuf(struct file *file) return fd; } +/* + * Tell every importer of @info that [start, end) changed. The caller has + * already updated the bitmaps and dropped info->lock: an importer re-reads + * get_phys() from inside this. + */ +static void gmem_revoke_range(struct gmem_info *info, + unsigned long start, unsigned long end) +{ + struct gmem_dmabuf *priv; + + lockdep_assert_not_held(&info->lock); + mutex_lock(&info->dmabufs_lock); + list_for_each_entry(priv, &info->dmabufs, list) { + dma_resv_lock(priv->dmabuf->resv, NULL); + dma_buf_invalidate_mappings_range(priv->dmabuf, + (u64)start << PAGE_SHIFT, + (u64)(end - start) << PAGE_SHIFT); + dma_resv_unlock(priv->dmabuf->resv); + } + mutex_unlock(&info->dmabufs_lock); +} + +/* + * The child lost the frames behind [start, end): host windows over them go + * too. Not for a read-only flip or a scratch toggle, where the VMM may be + * the page's writer; the VMM re-maps after a grant. + */ +static void gmem_unmap_host(struct gmem_info *info, unsigned long start, + unsigned long end) +{ + loff_t off = (loff_t)start << PAGE_SHIFT, len = (loff_t)(end - start) << PAGE_SHIFT; + struct gmem_dmabuf *priv; + + if (end <= start) + return; + unmap_mapping_range(info->mapping, off, len, 1); + /* Windows through a dma-buf live on the dma-buf's own file. */ + mutex_lock(&info->dmabufs_lock); + list_for_each_entry(priv, &info->dmabufs, list) + unmap_mapping_range(priv->dmabuf->file->f_mapping, off, len, 1); + mutex_unlock(&info->dmabufs_lock); +} + +/* Which allowlist bit gates each child-fd ioctl. 0 = not gated. */ +static u32 gmem_ioctl_allow_bit(unsigned int cmd) +{ + switch (cmd) { + case GMEM_PROVIDER_SET_PRESENT: return GMEM_ALLOW_SET_PRESENT; + case GMEM_PROVIDER_SET_READONLY: return GMEM_ALLOW_SET_READONLY; + case GMEM_PROVIDER_GET_DMABUF: return GMEM_ALLOW_GET_DMABUF; + case GMEM_PROVIDER_GET_STATS: return GMEM_ALLOW_GET_STATS; + default: return 0; + } +} + +static long gmem_get_stats(struct gmem_info *info, void __user *uarg) +{ + struct gmem_provider_stats st = {}; + + mutex_lock(&info->lock); + st.region_offset = (u64)info->root_index << PAGE_SHIFT; + st.region_len = (u64)info->npages << PAGE_SHIFT; + st.owned_pages = info->owned ? + bitmap_weight(info->owned, info->npages) : info->npages; + st.absent_pages = info->absent ? + bitmap_weight(info->absent, info->npages) : 0; + st.readonly_pages = info->readonly ? + bitmap_weight(info->readonly, info->npages) : 0; + st.allow = info->allow; + mutex_unlock(&info->lock); + + return copy_to_user(uarg, &st, sizeof(st)) ? -EFAULT : 0; +} + static long gmem_fd_ioctl(struct file *file, unsigned int cmd, unsigned long arg) { struct gmem_info *info = to_gmem_info(file); struct gmem_provider_present p; unsigned long start_index, end_index; + u32 need = gmem_ioctl_allow_bit(cmd); + + /* + * The allowlist is fixed by the creator at NEW_CHILD and checked here + * before dispatch, so it also governs ioctls added later. A SETUP + * provider has no root and allows everything, as before. + */ + if (need && !(info->allow & need)) + return -EPERM; if (cmd == GMEM_PROVIDER_GET_DMABUF) return gmem_provider_get_dmabuf(file); + if (cmd == GMEM_PROVIDER_GET_STATS) + return gmem_get_stats(info, (void __user *)arg); + + if (cmd == GMEM_PROVIDER_SET_READONLY) { + struct gmem_provider_readonly r; + + if (copy_from_user(&r, (void __user *)arg, sizeof(r))) + return -EFAULT; + if (!r.len || !PAGE_ALIGNED(r.offset) || !PAGE_ALIGNED(r.len) || + r.pad) + return -EINVAL; + start_index = r.offset >> PAGE_SHIFT; + end_index = start_index + (r.len >> PAGE_SHIFT); + if (end_index > info->npages || end_index < start_index) + return -EINVAL; + + /* + * Flip the bits, then invalidate the range so importers drop + * their mappings and the next access re-reads get_phys() with + * the new permission. Making a range read-only must + * tear down writable mappings; making it writable again is + * also invalidated so a stale read-only mapping does not keep + * exiting. + */ + mutex_lock(&info->lock); + if (r.readonly) + bitmap_set(info->readonly, start_index, + end_index - start_index); + else + bitmap_clear(info->readonly, start_index, + end_index - start_index); + mutex_unlock(&info->lock); + gmem_revoke_range(info, start_index, end_index); + return 0; + } if (cmd != GMEM_PROVIDER_SET_PRESENT) return -ENOTTY; @@ -589,152 +537,495 @@ static long gmem_fd_ioctl(struct file *file, unsigned int cmd, unsigned long arg if (end_index > info->npages || end_index < start_index) return -EINVAL; - if (p.present) { - /* Restore: next guest fault calls get_pfn() and re-maps. */ - mutex_lock(&info->lock); + mutex_lock(&info->lock); + if (p.present) bitmap_clear(info->absent, start_index, end_index - start_index); - mutex_unlock(&info->lock); - } else { - unsigned long clamped_start; - - /* Revoke: mark absent, then zap the guest NPT/EPT for the range. */ - mutex_lock(&info->lock); + else bitmap_set(info->absent, start_index, end_index - start_index); + /* + * Both directions invalidate: on revoke so the guest and devices stop + * using the pages, on restore so importers that were shown a hole or + * the scratch frame re-read and map the real frames again. + */ + mutex_unlock(&info->lock); + gmem_revoke_range(info, start_index, end_index); + if (!p.present) + gmem_unmap_host(info, start_index, end_index); + return 0; +} - /* - * Translate provider offset -> guest gfn. Skip any part of - * the range that falls outside the currently bound slot; a - * naive subtraction would underflow. - */ - clamped_start = max_t(unsigned long, start_index, info->pgoff); - if (info->kvm && end_index > clamped_start) - kvm_gmem_invalidate_range(info->kvm, - info->base_gfn + clamped_start - info->pgoff, - info->base_gfn + end_index - info->pgoff); - mutex_unlock(&info->lock); - - /* - * Fan out to iommufd (and any other dma-buf importer): mark the - * exported dma-buf(s) revoked and invalidate any active mappings. - * The provider stays ignorant of scratch-page policy; that lives - * in the importer. - */ - mutex_lock(&info->dmabufs_lock); - gmem_dma_buf_revoke_all(info); - mutex_unlock(&info->dmabufs_lock); - } +static int gmem_fops_release(struct inode *inode, struct file *file) +{ + gmem_release(file); return 0; } -static long gmem_ctl_ioctl(struct file *file, unsigned int cmd, unsigned long arg) +/* A child file owns its controls, host windows and dma-buf exports. */ +static const struct file_operations gmem_provider_fops = { + .owner = THIS_MODULE, + .release = gmem_fops_release, + .mmap = gmem_mmap, + .unlocked_ioctl = gmem_fd_ioctl, + .compat_ioctl = gmem_fd_ioctl, +}; + +/* + * Allocate a child over [base_pfn, +npages) and return its fd. Shared by + * SETUP (standalone region) and NEW_CHILD (a root sub-range). The fd owns the + * module pin on success. + */ +static int gmem_new_provider_fd(unsigned long base_pfn, + unsigned long npages, struct page *cma_pages, + u32 flags, unsigned long root_index, + unsigned long *owned, u32 allow, + struct gmem_info **out) { - struct gmem_provider_setup setup; + struct file *file; struct gmem_info *info; - struct file *kvm_file; - struct page *pages = NULL; - struct kvm *kvm; - unsigned long npages; int fd, ret; - if (cmd != GMEM_PROVIDER_SETUP) - return -ENOTTY; + info = kzalloc_obj(*info); + if (!info) + return -ENOMEM; + info->base_pfn = base_pfn; + info->npages = npages; + info->cma_pages = cma_pages; + info->root_index = root_index; + info->owned = owned; + info->allow = allow; + info->absent = kvzalloc_objs(unsigned long, BITS_TO_LONGS(npages)); + info->readonly = kvzalloc_objs(unsigned long, BITS_TO_LONGS(npages)); + if (!info->absent || !info->readonly) { + ret = -ENOMEM; + goto err_free; + } + INIT_LIST_HEAD(&info->dmabufs); + INIT_LIST_HEAD(&info->root_link); + mutex_init(&info->dmabufs_lock); + mutex_init(&info->lock); - if (copy_from_user(&setup, (void __user *)arg, sizeof(setup))) - return -EFAULT; - if (setup.flags & ~GMEM_PROVIDER_FLAG_MMAP_CAPABLE) - return -EINVAL; + /* Pin THIS_MODULE while the provider fd is alive (release drops it). */ + if (!try_module_get(THIS_MODULE)) { + ret = -ENODEV; + goto err_free; + } + info->mmap_capable = !!(flags & GMEM_PROVIDER_FLAG_MMAP_CAPABLE); - kvm_file = fget(setup.kvm_fd); - if (!kvm_file) - return -EBADF; - if (!file_is_kvm(kvm_file)) { - fput(kvm_file); - return -EINVAL; + /* + * A private inode, not the shared anon one: host windows are torn + * down by file range on revoke, which must not touch other files. + */ + fd = get_unused_fd_flags(O_CLOEXEC); + if (fd < 0) { + ret = fd; + module_put(THIS_MODULE); + goto err_free; } - kvm = kvm_file->private_data; - if (!kvm) { - fput(kvm_file); - return -EINVAL; + file = anon_inode_create_getfile("[gmem-provider]", &gmem_provider_fops, + info, O_RDWR, NULL); + if (IS_ERR(file)) { + put_unused_fd(fd); + ret = PTR_ERR(file); + module_put(THIS_MODULE); + goto err_free; } - kvm_get_kvm(kvm); - fput(kvm_file); + info->mapping = file->f_mapping; + if (owned) + list_add(&info->root_link, &gmem_root.children); + fd_install(fd, file); + if (out) + *out = info; + return fd; - info = kzalloc(sizeof(*info), GFP_KERNEL); - if (!info) { - ret = -ENOMEM; - goto err_put_kvm; - } +err_free: + kvfree(info->absent); + kvfree(info->readonly); + kfree(info); + return ret; +} + +static long gmem_ctl_setup(void __user *uarg) +{ + struct gmem_provider_setup setup; + struct page *pages = NULL; + unsigned long base_pfn, npages; + int fd; + + if (copy_from_user(&setup, uarg, sizeof(setup))) + return -EFAULT; + if (setup.flags & ~GMEM_PROVIDER_FLAG_MMAP_CAPABLE) + return -EINVAL; + /* kvm_fd is legacy and ignored: binding happens in KVM_CREATE_GUEST_MEMFD. */ if (addr && len) { /* External page-less range from module params. */ - info->base_pfn = addr >> PAGE_SHIFT; - info->npages = len >> PAGE_SHIFT; + base_pfn = addr >> PAGE_SHIFT; + npages = len >> PAGE_SHIFT; } else { /* CMA fallback: allocate a contiguous, page-backed region. */ - if (!setup.size || !PAGE_ALIGNED(setup.size)) { - ret = -EINVAL; - goto err_free_info; - } + if (!setup.size || !PAGE_ALIGNED(setup.size)) + return -EINVAL; npages = setup.size >> PAGE_SHIFT; pages = alloc_contig_pages(npages, GFP_KERNEL, numa_node_id(), NULL); - if (!pages) { - ret = -ENOMEM; - goto err_free_info; - } + if (!pages) + return -ENOMEM; /* Provider path skips KVM's folio-clear; zero to avoid data leak. */ memset(page_to_virt(pages), 0, (size_t)npages << PAGE_SHIFT); - info->base_pfn = page_to_pfn(pages); - info->npages = npages; - info->cma_pages = pages; + base_pfn = page_to_pfn(pages); } - info->absent = kvzalloc(BITS_TO_LONGS(info->npages) * sizeof(unsigned long), - GFP_KERNEL); - if (!info->absent) { - ret = -ENOMEM; - goto err_free_pages; + fd = gmem_new_provider_fd(base_pfn, npages, pages, setup.flags, 0, + NULL, GMEM_ALLOW_ALL, NULL); + if (fd < 0 && pages) + free_contig_range(page_to_pfn(pages), npages); + return fd; +} + +/* + * Lazily create the root region on first NEW_CHILD. Uses the module params + * if given (page-less), else a CMA region of @size bytes. Idempotent once + * created; a later different @size is ignored. + */ +static int gmem_root_ensure(u64 size) +{ + unsigned long npages; + struct page *pages = NULL; + + lockdep_assert_held(&gmem_root.lock); + if (gmem_root.npages) + return 0; + + if (addr && len) { + gmem_root.base_pfn = addr >> PAGE_SHIFT; + npages = len >> PAGE_SHIFT; + } else { + if (!size || !PAGE_ALIGNED(size)) + return -EINVAL; + npages = size >> PAGE_SHIFT; + pages = alloc_contig_pages(npages, GFP_KERNEL, numa_node_id(), NULL); + if (!pages) + return -ENOMEM; + memset(page_to_virt(pages), 0, (size_t)npages << PAGE_SHIFT); + gmem_root.base_pfn = page_to_pfn(pages); } + gmem_root.owned = kvzalloc_objs(unsigned long, BITS_TO_LONGS(npages)); + gmem_root.donated = kvzalloc_objs(unsigned long, BITS_TO_LONGS(npages)); + if (!gmem_root.owned || !gmem_root.donated) { + kvfree(gmem_root.owned); + kvfree(gmem_root.donated); + gmem_root.owned = NULL; + gmem_root.donated = NULL; + if (pages) + free_contig_range(page_to_pfn(pages), npages); + return -ENOMEM; + } + gmem_root.cma_pages = pages; + gmem_root.npages = npages; + return 0; +} - INIT_LIST_HEAD(&info->dmabufs); - mutex_init(&info->dmabufs_lock); - mutex_init(&info->lock); +static void gmem_root_teardown(void) +{ + if (!gmem_root.npages) + return; + WARN_ON(!list_empty(&gmem_root.children)); + if (gmem_root.cma_pages) + free_contig_range(gmem_root.base_pfn, gmem_root.npages); + kvfree(gmem_root.owned); + kvfree(gmem_root.donated); + memset(&gmem_root, 0, sizeof(gmem_root)); + mutex_init(&gmem_root.lock); + INIT_LIST_HEAD(&gmem_root.children); +} + +/* Validate a page-aligned [offset, +len) against the root; return page bounds. */ +static int gmem_root_range(u64 offset, u64 length, + unsigned long *first, unsigned long *last) +{ + if (!length || !PAGE_ALIGNED(offset) || !PAGE_ALIGNED(length)) + return -EINVAL; + *first = offset >> PAGE_SHIFT; + *last = *first + (length >> PAGE_SHIFT); /* exclusive */ + if (*last <= *first || *last > gmem_root.npages) + return -EINVAL; + return 0; +} + +/* Resolve a child fd created by NEW_CHILD; returns a referenced file. */ +static struct file *gmem_get_child(int fd, struct gmem_info **infop) +{ + struct file *f = fget(fd); + struct gmem_info *info; + + if (!f) + return ERR_PTR(-EBADF); + if (f->f_op != &gmem_provider_fops) { + fput(f); + return ERR_PTR(-EINVAL); + } + info = to_gmem_info(f); + if (!info->owned) { + fput(f); + return ERR_PTR(-EINVAL); + } + *infop = info; + return f; +} +static long gmem_ctl_new_child(void __user *uarg) +{ + struct gmem_provider_new_child nc; + struct gmem_info *info; + unsigned long first, last; + unsigned long *owned; + int fd, ret; + + if (copy_from_user(&nc, uarg, sizeof(nc))) + return -EFAULT; + if ((nc.flags & ~GMEM_PROVIDER_FLAG_MMAP_CAPABLE) || nc.pad || + (nc.allow & ~GMEM_ALLOW_ALL)) + return -EINVAL; + + /* nc.kvm_fd is legacy and ignored: binding happens in KVM_CREATE_GUEST_MEMFD. */ + + mutex_lock(&gmem_root.lock); + ret = gmem_root_ensure(nc.offset + nc.len); + if (ret) + goto out_unlock; + ret = gmem_root_range(nc.offset, nc.len, &first, &last); + if (ret) + goto out_unlock; /* - * Pin THIS_MODULE while the provider fd is alive. The fd is created - * with kvm_gmem_fops (owned by kvm.ko), which does not pin us, so - * rmmod of gmem_provider is otherwise free to run behind our ops. + * A carve is the range this child may ever hold; it is the VM's whole + * view of memory and becomes its memslot. Carves may overlap: that is + * how a range can later MOVE from one VM to another. Ownership is + * per page and exclusive. The new child is granted every page of its + * carve that no other child owns and the root has not parked; the + * rest it can only receive by MOVE or RECLAIM. A carve with nothing + * to grant is refused as a likely mistake. */ - if (!try_module_get(THIS_MODULE)) { - ret = -ENODEV; - goto err_free_pages; + if (find_next_zero_bit(gmem_root.owned, last, first) >= last) { + ret = -EBUSY; + goto out_unlock; } - info->backing.ops = &gmem_ops; - info->kvm = kvm; - info->mmap_capable = !!(setup.flags & GMEM_PROVIDER_FLAG_MMAP_CAPABLE); + /* + * Allocate the ownership bitmap before the fd exists, so a failure + * here has nothing to unwind. It is handed to the new info below. + */ + owned = kvzalloc_objs(unsigned long, BITS_TO_LONGS(last - first)); + if (!owned) { + ret = -ENOMEM; + goto out_unlock; + } + /* Establish the grant before publishing the fd. */ + { + unsigned long i; - fd = anon_inode_getfd("[gmem-provider]", &kvm_gmem_fops, - &info->backing, O_RDWR | O_CLOEXEC); + for (i = first; i < last; i++) { + if (test_bit(i, gmem_root.owned) || + test_bit(i, gmem_root.donated)) + continue; + __set_bit(i - first, owned); + __set_bit(i, gmem_root.owned); + } + } + fd = gmem_new_provider_fd(gmem_root.base_pfn + first, + last - first, NULL, nc.flags, first, + owned, nc.allow, &info); if (fd < 0) { + unsigned long i; + + for_each_set_bit(i, owned, last - first) + __clear_bit(first + i, gmem_root.owned); + kvfree(owned); ret = fd; - goto err_module_put; + goto out_unlock; } + mutex_unlock(&gmem_root.lock); return fd; -err_module_put: - module_put(THIS_MODULE); +out_unlock: + mutex_unlock(&gmem_root.lock); + return ret; +} -err_free_pages: - if (pages) - free_contig_range(page_to_pfn(pages), npages); -err_free_info: - kvfree(info->absent); - kfree(info); -err_put_kvm: - kvm_put_kvm(kvm); +/* + * Take [first, last) away from @info: clear ownership, then revoke from both + * of its importers. Ownership is cleared before the revoke so a racing fault + * that slips in re-reads "not owned" and gets the scratch frame or fails. + * Caller holds gmem_root.lock; we take info->lock inside it. + */ +static void gmem_child_lose(struct gmem_info *info, + unsigned long first, unsigned long last) +{ + unsigned long s = first - info->root_index, e = last - info->root_index; + + mutex_lock(&info->lock); + bitmap_clear(info->owned, s, e - s); + mutex_unlock(&info->lock); + gmem_revoke_range(info, s, e); + gmem_unmap_host(info, s, e); + bitmap_clear(gmem_root.owned, first, last - first); +} + +/* Give [first, last) to @info and invalidate so its importers pick it up. */ +static void gmem_child_gain(struct gmem_info *info, + unsigned long first, unsigned long last) +{ + unsigned long s = first - info->root_index, e = last - info->root_index; + + mutex_lock(&info->lock); + bitmap_set(info->owned, s, e - s); + mutex_unlock(&info->lock); + gmem_revoke_range(info, s, e); + bitmap_set(gmem_root.owned, first, last - first); +} + +/* Does child @info's carved range contain [first, last)? */ +static bool gmem_child_covers(struct gmem_info *info, + unsigned long first, unsigned long last) +{ + return first >= info->root_index && + last <= info->root_index + info->npages; +} + +/* Does child @info currently own every page of [first, last)? */ +static bool gmem_child_owns(struct gmem_info *info, + unsigned long first, unsigned long last) +{ + unsigned long s = first - info->root_index, e = last - info->root_index; + + return gmem_child_covers(info, first, last) && + find_next_zero_bit(info->owned, e, s) >= e; +} + +static long gmem_ctl_move(void __user *uarg) +{ + struct gmem_provider_move mv; + struct gmem_info *src, *dst; + struct file *sf, *df; + unsigned long first, last; + long ret; + + if (copy_from_user(&mv, uarg, sizeof(mv))) + return -EFAULT; + sf = gmem_get_child(mv.src_fd, &src); + if (IS_ERR(sf)) + return PTR_ERR(sf); + df = gmem_get_child(mv.dst_fd, &dst); + if (IS_ERR(df)) { + fput(sf); + return PTR_ERR(df); + } + + mutex_lock(&gmem_root.lock); + ret = gmem_root_range(mv.offset, mv.len, &first, &last); + if (ret) + goto out; + if (src == dst || !gmem_child_owns(src, first, last) || + !gmem_child_covers(dst, first, last)) { + ret = -EINVAL; + goto out; + } + /* + * Revoke on the source first, grant on the destination second. No two + * importers hold the range at once. Both children's info->lock are + * taken in turn under gmem_root.lock, never nested with each other. + */ + gmem_child_lose(src, first, last); + gmem_child_gain(dst, first, last); + ret = 0; +out: + mutex_unlock(&gmem_root.lock); + fput(df); + fput(sf); + return ret; +} + +static long gmem_ctl_donate(void __user *uarg, bool reclaim) +{ + struct gmem_provider_donate d; + struct gmem_info *info; + struct file *f; + unsigned long first, last; + long ret; + + if (copy_from_user(&d, uarg, sizeof(d))) + return -EFAULT; + if (d.pad) + return -EINVAL; + f = gmem_get_child(d.fd, &info); + if (IS_ERR(f)) + return PTR_ERR(f); + + mutex_lock(&gmem_root.lock); + ret = gmem_root_range(d.offset, d.len, &first, &last); + if (ret) + goto out; + if (!reclaim) { + /* DONATE: the child must own it; park it at the root. */ + if (!gmem_child_owns(info, first, last)) { + ret = -EINVAL; + goto out; + } + gmem_child_lose(info, first, last); + bitmap_set(gmem_root.donated, first, last - first); + } else { + /* RECLAIM: must be donated and inside this child's range. */ + if (!gmem_child_covers(info, first, last) || + find_next_zero_bit(gmem_root.donated, last, first) < last) { + ret = -EINVAL; + goto out; + } + bitmap_clear(gmem_root.donated, first, last - first); + gmem_child_gain(info, first, last); + } + ret = 0; +out: + mutex_unlock(&gmem_root.lock); + fput(f); return ret; } +static long gmem_ctl_set_scratch(void __user *uarg) +{ + struct gmem_provider_scratch sc; + struct gmem_info *info; + + if (copy_from_user(&sc, uarg, sizeof(sc))) + return -EFAULT; + if (sc.pad || sc.enable > 1) + return -EINVAL; + + /* + * Flipping the mode changes what every revoked page reports, so every + * child is re-invalidated in full: cheap for a PoC, and it guarantees + * no importer keeps a stale hole or a stale scratch mapping. + */ + mutex_lock(&gmem_root.lock); + WRITE_ONCE(gmem_root.scratch_enabled, !!sc.enable); + list_for_each_entry(info, &gmem_root.children, root_link) + gmem_revoke_range(info, 0, info->npages); + mutex_unlock(&gmem_root.lock); + return 0; +} + +static long gmem_ctl_ioctl(struct file *file, unsigned int cmd, unsigned long arg) +{ + void __user *uarg = (void __user *)arg; + + switch (cmd) { + case GMEM_PROVIDER_SETUP: return gmem_ctl_setup(uarg); + case GMEM_PROVIDER_NEW_CHILD: return gmem_ctl_new_child(uarg); + case GMEM_PROVIDER_MOVE: return gmem_ctl_move(uarg); + case GMEM_PROVIDER_DONATE: return gmem_ctl_donate(uarg, false); + case GMEM_PROVIDER_RECLAIM: return gmem_ctl_donate(uarg, true); + case GMEM_PROVIDER_SET_SCRATCH: return gmem_ctl_set_scratch(uarg); + default: return -ENOTTY; + } +} + static const struct file_operations gmem_ctl_fops = { .owner = THIS_MODULE, .unlocked_ioctl = gmem_ctl_ioctl, @@ -749,20 +1040,34 @@ static struct miscdevice gmem_dev = { static int __init gmem_provider_init(void) { + int ret; + if ((addr || len) && (!addr || !len || !PAGE_ALIGNED(addr) || !PAGE_ALIGNED(len))) { pr_err("gmem_provider: addr= and len= must both be set and page aligned\n"); return -EINVAL; } - return misc_register(&gmem_dev); + gmem_scratch_page = alloc_page(GFP_KERNEL | __GFP_ZERO); + if (!gmem_scratch_page) + return -ENOMEM; + mutex_init(&gmem_root.lock); + INIT_LIST_HEAD(&gmem_root.children); + + ret = misc_register(&gmem_dev); + if (ret) + __free_page(gmem_scratch_page); + return ret; } module_init(gmem_provider_init); static void __exit gmem_provider_exit(void) { misc_deregister(&gmem_dev); + gmem_root_teardown(); + __free_page(gmem_scratch_page); } module_exit(gmem_provider_exit); MODULE_LICENSE("GPL"); +MODULE_IMPORT_NS("DMA_BUF"); MODULE_DESCRIPTION("Sample guest_memfd provider (external page-less range or CMA fallback)"); diff --git a/samples/kvm/gmem_provider.h b/samples/kvm/gmem_provider.h index 45f1b8257f60..1c8f5cacd2be 100644 --- a/samples/kvm/gmem_provider.h +++ b/samples/kvm/gmem_provider.h @@ -6,19 +6,18 @@ #include <linux/types.h> /* - * ioctl on /dev/gmem_provider: create a guest_memfd provider fd and return it - * for use with KVM_SET_USER_MEMORY_REGION2. + * ioctl on /dev/gmem_provider: create a memory-owner fd and return it. The + * holder may mmap it and request a dma-buf for KVM or iommufd. * * If the module was loaded with addr=/len=, the fd is backed by that fixed * page-less physical range and @size is ignored. Otherwise the fd is backed * by a @size-byte physically contiguous region from alloc_contig_pages(). */ /* Flags for struct gmem_provider_setup.flags */ -#define GMEM_PROVIDER_FLAG_MMAP_CAPABLE (1u << 0) /* fd is mmap()-able; slot becomes gmem-only. - Refused for coco VMs (SEV-SNP/TDX). */ +#define GMEM_PROVIDER_FLAG_MMAP_CAPABLE (1u << 0) /* fd and its dma-bufs are mmap()-able */ struct gmem_provider_setup { - __s32 kvm_fd; /* an open KVM VM fd */ + __s32 kvm_fd; /* ignored */ __u32 flags; /* GMEM_PROVIDER_FLAG_* */ __u64 size; /* CMA fallback size in bytes, page aligned */ }; @@ -27,10 +26,9 @@ struct gmem_provider_setup { #define GMEM_PROVIDER_SETUP _IOW(GMEM_PROVIDER_IOCTL_BASE, 1, struct gmem_provider_setup) /* - * ioctl on a provider fd (returned by SETUP): flip a byte range of the backing - * between present and absent. Revoking (present=0) marks the range absent and - * zaps the guest's NPT/EPT so the next access re-faults; get_pfn() then refuses - * the range until it is restored (present=1). Models overcommit page reclaim. + * ioctl on an owner fd: flip a byte range between present and absent. + * Revoking a range invalidates every exported dma-buf, so importers discard + * their old mappings and observe a hole until the range is restored. */ struct gmem_provider_present { __u64 offset; /* byte offset into the provider region, page aligned */ @@ -42,12 +40,134 @@ struct gmem_provider_present { #define GMEM_PROVIDER_SET_PRESENT _IOW(GMEM_PROVIDER_IOCTL_BASE, 2, struct gmem_provider_present) /* - * ioctl on a provider fd (returned by SETUP): export the backing region as a - * dynamic dma-buf and return an fd for it, suitable for + * ioctl on a provider fd: make a byte range read-only for the guest, or + * writable again. KVM maps a read-only page without write permission and a + * guest write to it exits to userspace with KVM_EXIT_MEMORY_FAULT. Existing + * mappings of the range are dropped so the change takes effect on the next + * access. + */ +struct gmem_provider_readonly { + __u64 offset; /* byte offset into the provider region, page aligned */ + __u64 len; /* byte length, page aligned */ + __u32 readonly; /* 1 = guest may not write, 0 = guest may write */ + __u32 pad; +}; + +#define GMEM_PROVIDER_SET_READONLY \ + _IOW(GMEM_PROVIDER_IOCTL_BASE, 4, struct gmem_provider_readonly) + +/* + * ioctl on an owner fd: export its memory as a dynamic dma-buf and return an + * fd suitable for * IOMMU_IOAS_MAP_FILE. Revoking the region (SET_PRESENT present=0) fans out - * to the exported dma-buf via dma_buf_invalidate_mappings(), causing iommufd + * to the exported dma-buf via dma_buf_invalidate_mappings_range(), causing importers * to tear down the IOMMU mapping so DMA to the reclaimed range faults. */ #define GMEM_PROVIDER_GET_DMABUF _IO(GMEM_PROVIDER_IOCTL_BASE, 3) +/* + * ---- Toy descriptor tree (a memory owner as a provider) ------------------- + * + * The control device owns one backing region (the "root"). NEW_CHILD carves a + * sub-range of it into a new provider fd bound to one VM. A child's pages can + * be MOVEd to another child, DONATEd to the root (absent from every consumer) + * and RECLAIMed. Every change revokes the range from KVM and from the + * child's exported dma-bufs. + * + * Every child fd carries an ioctl allowlist, fixed at NEW_CHILD by the + * creator. Ioctls not in the list fail with -EPERM. + */ + +/* Allowlist bits for struct gmem_provider_new_child.allow */ +#define GMEM_ALLOW_SET_PRESENT (1u << 0) +#define GMEM_ALLOW_SET_READONLY (1u << 1) +#define GMEM_ALLOW_GET_DMABUF (1u << 2) +#define GMEM_ALLOW_GET_STATS (1u << 3) +#define GMEM_ALLOW_ALL (GMEM_ALLOW_SET_PRESENT | \ + GMEM_ALLOW_SET_READONLY | \ + GMEM_ALLOW_GET_DMABUF | \ + GMEM_ALLOW_GET_STATS) + +/* + * ioctl on /dev/gmem_provider: carve @len bytes at @offset of the root into a + * new child fd. The range must be available. Without addr=/len=, the first + * child fixes the CMA root size; create and close a sizing child first if + * later children extend beyond it. @allow becomes the child's ioctl allowlist. + */ +struct gmem_provider_new_child { + __s32 kvm_fd; /* ignored */ + __u32 flags; /* GMEM_PROVIDER_FLAG_* */ + __u64 offset; /* byte offset into the root region, page aligned */ + __u64 len; /* byte length, page aligned */ + __u32 allow; /* GMEM_ALLOW_* */ + __u32 pad; +}; + +#define GMEM_PROVIDER_NEW_CHILD _IOW(GMEM_PROVIDER_IOCTL_BASE, 5, struct gmem_provider_new_child) + +/* + * ioctl on /dev/gmem_provider: move [@offset, @offset+@len) of the root region + * from child @src_fd to child @dst_fd. The range must currently be owned by + * @src_fd and must fall inside @dst_fd's carved range. The source is + * invalidated first, then the destination is granted, so no two children own + * the range at once. + */ +struct gmem_provider_move { + __s32 src_fd; + __s32 dst_fd; + __u64 offset; /* byte offset into the root region, page aligned */ + __u64 len; /* byte length, page aligned */ +}; + +#define GMEM_PROVIDER_MOVE _IOW(GMEM_PROVIDER_IOCTL_BASE, 6, struct gmem_provider_move) + +/* + * ioctl on /dev/gmem_provider: DONATE revokes [@offset, +@len) from the child + * that owns it and parks it at the root; RECLAIM returns a donated range to + * @fd, which must be the child whose carved range contains it. The toy does + * no scrub and no hotplug; it models only the ownership state. + */ +struct gmem_provider_donate { + __s32 fd; /* DONATE: owning child (checked); RECLAIM: recipient */ + __u32 pad; + __u64 offset; /* byte offset into the root region, page aligned */ + __u64 len; /* byte length, page aligned */ +}; + +#define GMEM_PROVIDER_DONATE _IOW(GMEM_PROVIDER_IOCTL_BASE, 7, struct gmem_provider_donate) +#define GMEM_PROVIDER_RECLAIM _IOW(GMEM_PROVIDER_IOCTL_BASE, 8, struct gmem_provider_donate) + +/* + * ioctl on /dev/gmem_provider: select what a revoked page reports to its + * consumers. With @enable=0 (default) a revoked page is absent: get_phys() + * reports it as not backed, so a guest access exits and a device DMA faults. With + * @enable=1 a revoked page reports the module's scratch frame, read-only, on + * both paths, so a device that cannot tolerate a fault lands on a harmless + * page. Scratch mode applies to children created by NEW_CHILD only; a + * SETUP provider always reports a revoked page as absent. Changing the mode + * re-invalidates every child so consumers re-read. + */ +struct gmem_provider_scratch { + __u32 enable; + __u32 pad; +}; + +#define GMEM_PROVIDER_SET_SCRATCH \ + _IOW(GMEM_PROVIDER_IOCTL_BASE, 9, struct gmem_provider_scratch) + +/* + * ioctl on a child fd: read back the child's ownership state, for tests. + */ +struct gmem_provider_stats { + __u64 region_offset; /* child's carved range within the root */ + __u64 region_len; + __u64 owned_pages; /* pages currently granted to this child */ + __u64 absent_pages; /* pages revoked (moved out, donated, or SET_PRESENT 0) */ + __u64 readonly_pages; + __u32 allow; + __u32 pad; +}; + +#define GMEM_PROVIDER_GET_STATS _IOR(GMEM_PROVIDER_IOCTL_BASE, 10, struct gmem_provider_stats) + #endif /* _SAMPLES_KVM_GMEM_PROVIDER_H */ diff --git a/tools/testing/selftests/kvm/Makefile.kvm b/tools/testing/selftests/kvm/Makefile.kvm index 12004a487c32..6accbc56ae69 100644 --- a/tools/testing/selftests/kvm/Makefile.kvm +++ b/tools/testing/selftests/kvm/Makefile.kvm @@ -78,11 +78,12 @@ TEST_GEN_PROGS_x86 += x86/evmcs_smm_controls_test TEST_GEN_PROGS_x86 += x86/exit_on_emulation_failure_test TEST_GEN_PROGS_x86 += x86/fastops_test TEST_GEN_PROGS_x86 += x86/fix_hypercall_test -TEST_GEN_PROGS_x86 += x86/gmem_provider_test TEST_GEN_PROGS_x86 += x86/gmem_provider_hugepage_test TEST_GEN_PROGS_x86 += x86/gmem_provider_revoke_test +TEST_GEN_PROGS_x86 += x86/gmem_provider_readonly_test TEST_GEN_PROGS_x86 += x86/gmem_provider_iommufd_test TEST_GEN_PROGS_x86 += x86/gmem_provider_vfio_test +TEST_GEN_PROGS_x86 += x86/gmem_poc_test TEST_GEN_PROGS_x86 += x86/hwcr_msr_test TEST_GEN_PROGS_x86 += x86/hyperv_clock TEST_GEN_PROGS_x86 += x86/hyperv_cpuid diff --git a/tools/testing/selftests/kvm/gmem_provider_nvme_dma_test.c b/tools/testing/selftests/kvm/gmem_provider_nvme_dma_test.c index 665b19028e13..b5ef6b9b1385 100644 --- a/tools/testing/selftests/kvm/gmem_provider_nvme_dma_test.c +++ b/tools/testing/selftests/kvm/gmem_provider_nvme_dma_test.c @@ -35,8 +35,8 @@ struct gmem_provider_setup { __u64 size; }; #define GMEM_PROVIDER_SETUP _IOW('G', 1, struct gmem_provider_setup) -#define GMEM_PROVIDER_FLAG_MMAP_CAPABLE (1u << 0) #define GMEM_PROVIDER_GET_DMABUF _IO('G', 3) +#define GMEM_PROVIDER_FLAG_MMAP_CAPABLE (1u << 0) struct gmem_provider_present { __u64 offset; @@ -94,7 +94,7 @@ int main(void) struct vfio_device_attach_iommufd_pt att = {}; struct vfio_region_info reg = {}; const char *cdev_path; - int gmem_ctl, gmem_fd, dmabuf_fd, iommufd_fd, vfio_fd; + int gmem_ctl, gmem_fd, dmabuf_fd, prov_fd, vm_fd = -1, iommufd_fd, vfio_fd; void *provider_hva, *sq_buf, *cq_buf, *bar; uint16_t pci_cmd; uint16_t expected_vid; @@ -120,22 +120,34 @@ int main(void) vm = ioctl(kvm, KVM_CREATE_VM, 0); TEST_ASSERT(vm >= 0, "KVM_CREATE_VM errno=%d", errno); setup.kvm_fd = vm; + vm_fd = vm; close(kvm); } setup.size = PROVIDER_SIZE; - gmem_fd = ioctl(gmem_ctl, GMEM_PROVIDER_SETUP, &setup); - TEST_ASSERT(gmem_fd >= 0, "SETUP errno=%d", errno); + prov_fd = ioctl(gmem_ctl, GMEM_PROVIDER_SETUP, &setup); + TEST_ASSERT(prov_fd >= 0, "SETUP errno=%d", errno); + + /* Export once; KVM and iommufd import the same dma-buf. */ + dmabuf_fd = ioctl(prov_fd, GMEM_PROVIDER_GET_DMABUF); + TEST_ASSERT(dmabuf_fd >= 0, "GET_DMABUF errno=%d", errno); + { + struct kvm_create_guest_memfd cgm = { + .size = PROVIDER_SIZE, + .flags = GUEST_MEMFD_FLAG_MMAP | GUEST_MEMFD_FLAG_USE_DMABUF, + .dmabuf_fd = dmabuf_fd, + }; + gmem_fd = ioctl(vm_fd, KVM_CREATE_GUEST_MEMFD, &cgm); + TEST_ASSERT(gmem_fd >= 0, "KVM_CREATE_GUEST_MEMFD(dmabuf) errno=%d", errno); + } provider_hva = mmap(NULL, PROVIDER_SIZE, PROT_READ | PROT_WRITE, MAP_SHARED, gmem_fd, 0); TEST_ASSERT(provider_hva != MAP_FAILED, "provider mmap errno=%d", errno); memset(provider_hva, 0xcc, 0x1000); /* poison first page */ - /* IOAS + provider dma-buf. */ + /* IOAS + the same dma-buf KVM imported. */ alloc.size = sizeof(alloc); TEST_ASSERT(!ioctl(iommufd_fd, IOMMU_IOAS_ALLOC, &alloc), "IOAS_ALLOC"); - dmabuf_fd = ioctl(gmem_fd, GMEM_PROVIDER_GET_DMABUF); - TEST_ASSERT(dmabuf_fd >= 0, "GET_DMABUF"); mapf.size = sizeof(mapf); mapf.flags = IOMMU_IOAS_MAP_FIXED_IOVA | IOMMU_IOAS_MAP_READABLE | IOMMU_IOAS_MAP_WRITEABLE; @@ -367,10 +379,10 @@ int main(void) munmap(bar, reg.size); close(vfio_fd); - close(dmabuf_fd); close(iommufd_fd); munmap(provider_hva, PROVIDER_SIZE); close(gmem_fd); + close(dmabuf_fd); close(gmem_ctl); return 0; } diff --git a/tools/testing/selftests/kvm/include/kvm_util.h b/tools/testing/selftests/kvm/include/kvm_util.h index 04a910164a29..c168c0880ebf 100644 --- a/tools/testing/selftests/kvm/include/kvm_util.h +++ b/tools/testing/selftests/kvm/include/kvm_util.h @@ -664,17 +664,38 @@ static inline bool is_smt_on(void) void vm_create_irqchip(struct kvm_vm *vm); -static inline int __vm_create_guest_memfd(struct kvm_vm *vm, u64 size, - u64 flags) +static inline int __vm_create_guest_memfd_dmabuf(struct kvm_vm *vm, + u64 size, u64 flags, + int dmabuf_fd) { struct kvm_create_guest_memfd guest_memfd = { .size = size, .flags = flags, + .dmabuf_fd = dmabuf_fd, }; return __vm_ioctl(vm, KVM_CREATE_GUEST_MEMFD, &guest_memfd); } +static inline int __vm_create_guest_memfd(struct kvm_vm *vm, u64 size, + u64 flags) +{ + return __vm_create_guest_memfd_dmabuf(vm, size, flags, 0); +} + +/* A guest_memfd that imports the dma-buf @dmabuf_fd for its memory. */ +static inline int vm_create_guest_memfd_dmabuf(struct kvm_vm *vm, + u64 size, u64 flags, + int dmabuf_fd) +{ + int fd = __vm_create_guest_memfd_dmabuf(vm, size, + flags | GUEST_MEMFD_FLAG_USE_DMABUF, + dmabuf_fd); + + TEST_ASSERT(fd >= 0, KVM_IOCTL_ERROR(KVM_CREATE_GUEST_MEMFD, fd)); + return fd; +} + static inline int vm_create_guest_memfd(struct kvm_vm *vm, u64 size, u64 flags) { diff --git a/tools/testing/selftests/kvm/x86/gmem_poc_test.c b/tools/testing/selftests/kvm/x86/gmem_poc_test.c new file mode 100644 index 000000000000..bdb64fe949ad --- /dev/null +++ b/tools/testing/selftests/kvm/x86/gmem_poc_test.c @@ -0,0 +1,824 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * gmem_poc_test - one provider, a control process and two VMMs. + * + * Two roles in one process, kept apart on purpose: + * + * ctl: owns /dev/gmem_provider (the root). Creates a child per VM with a + * narrowed allowlist, MOVEs a range between children, DONATEs and + * RECLAIMs, flips scratch mode. Never touches a VM. + * vmm: owns one child fd, one KVM VM and one iommufd IOAS. Creates the + * VM's guest_memfd with the child as provider fd, binds it to the + * memslot, and maps the guest_memfd's dma-buf into the IOAS on a + * mock domain. The provider revokes into guest_memfd; KVM and the + * device follow. The VMM never touches the control fd, and its + * allowlist would stop it. + * + * guest_memfd sees one flat fd per VM; every relationship between fds lives + * in the provider. The scenarios: + * + * 1. Launch: each guest runs on its child and its write lands. + * 2. Move: a range leaves A for B. A faults on it and exits to its VMM + * with KVM_EXIT_MEMORY_FAULT; A's device mapping of it is a + * hole and the rest is intact; B reads what A wrote there. + * 3. Donate: a range leaves A for the root and comes back on RECLAIM. + * 4. Read-only: a guest write to a read-only page exits; after clearing + * the bit the same write lands. + * 5. Scratch: with scratch mode on, A's donated range resolves in the + * IOAS to one frame that is none of A's, and reads as zero. + * 6. Fragment: alternate scratch and owned pages in one 2 MiB block read + * correctly and map at 4K; an untouched block maps at 2M. + * 7. Allowlist: B's fd cannot reach an ioctl the control process did not grant. + * 8. Window: a VMM window mmap()ed through the guest_memfd fd over a + * range the provider donates faults afterwards: the + * revoke tore the host PTEs down too. + * + * A's guest_memfd is created while another thread flips SET_PRESENT and + * SET_READONLY on its child, so attachment races exporter invalidation. + * + * Runs on the iommufd mock domain: no device needed. Requires + * /dev/gmem_provider (samples/kvm/gmem_provider.ko) and /dev/iommu. + */ +#include <fcntl.h> +#include <errno.h> +#include <setjmp.h> +#include <pthread.h> +#include <signal.h> +#include <stdatomic.h> +#include <stdint.h> +#include <stdio.h> +#include <string.h> +#include <unistd.h> +#include <sys/ioctl.h> +#include <sys/mman.h> + +#include <linux/iommufd.h> + +#include "test_util.h" +#include "kvm_util.h" +#include "processor.h" + +/* + * The iommufd mock domain's test ABI. Only the three probes we need; the + * iommufd selftest helper header brings the kselftest harness and its own + * bitops with it, which do not coexist with the KVM selftest library. + */ +#include "../../../../../drivers/iommu/iommufd/iommufd_test.h" + +/* Mirrors samples/kvm/gmem_provider.h */ +#define GMEM_PROVIDER_FLAG_MMAP_CAPABLE (1u << 0) +#define GMEM_ALLOW_SET_PRESENT (1u << 0) +#define GMEM_ALLOW_SET_READONLY (1u << 1) +#define GMEM_ALLOW_GET_DMABUF (1u << 2) +#define GMEM_ALLOW_GET_STATS (1u << 3) + +struct gmem_provider_present { __u64 offset, len; __u32 present, pad; }; +struct gmem_provider_readonly { __u64 offset, len; __u32 readonly, pad; }; +struct gmem_provider_new_child { + __s32 kvm_fd; __u32 flags; __u64 offset, len; __u32 allow, pad; +}; + +struct gmem_provider_move { __s32 src_fd, dst_fd; __u64 offset, len; }; + +struct gmem_provider_donate { __s32 fd; __u32 pad; __u64 offset, len; }; + +struct gmem_provider_scratch { __u32 enable, pad; }; + +struct gmem_provider_stats { + __u64 region_offset, region_len, owned_pages, absent_pages, readonly_pages; + __u32 allow, pad; +}; + +#define GMEM_PROVIDER_SET_PRESENT _IOW('G', 2, struct gmem_provider_present) +#define GMEM_PROVIDER_SET_READONLY _IOW('G', 4, struct gmem_provider_readonly) +#define GMEM_PROVIDER_NEW_CHILD _IOW('G', 5, struct gmem_provider_new_child) +#define GMEM_PROVIDER_MOVE _IOW('G', 6, struct gmem_provider_move) +#define GMEM_PROVIDER_DONATE _IOW('G', 7, struct gmem_provider_donate) +#define GMEM_PROVIDER_RECLAIM _IOW('G', 8, struct gmem_provider_donate) +#define GMEM_PROVIDER_SET_SCRATCH _IOW('G', 9, struct gmem_provider_scratch) +#define GMEM_PROVIDER_GET_DMABUF _IO('G', 3) +#define GMEM_PROVIDER_GET_STATS _IOR('G', 10, struct gmem_provider_stats) + +#define PAGE 0x1000ULL +#define CHILD_SIZE 0x800000ULL /* 8 MiB per child */ +#define GPA (1ULL << 32) /* each VM maps its child here */ +#define IOVA (1ULL << 28) /* inside the mock domain's aperture */ +#define MAGIC_A 0xa11ce000a11ce000ULL +#define MAGIC_B 0xb0bb0bb0b0bb0bb0ULL + +/* + * Layout of the root region. A and B are carved so that they OVERLAP on + * [SHARED_OFF, +SHARED_LEN): a carve is the range a child may ever hold, + * ownership within it is per page. the control process creates A owning its whole carve and + * B owning its carve minus the shared window, so the window starts with A + * and can MOVE to B. This is the carve-out shape: two VMs, two fds, + * one range that changes hands, no shared object in guest_memfd. + */ +#define A_OFF 0ULL +#define SHARED_LEN (4 * PAGE) +#define B_OFF (CHILD_SIZE - SHARED_LEN) /* B's carve begins at the window */ +#define SHARED_OFF B_OFF /* root offset of the window */ +#define ROOT_SIZE (B_OFF + CHILD_SIZE) + +/* The shared window as each VM sees it. */ +#define A_SHARED_GPA (GPA + (SHARED_OFF - A_OFF)) +#define B_SHARED_GPA (GPA + (SHARED_OFF - B_OFF)) +#define A_SHARED_IOVA (IOVA + (SHARED_OFF - A_OFF)) + +/* A page of A's that is never moved, for read-only and control checks. */ +#define RO_OFF (64 * PAGE) +#define RO_GPA (GPA + RO_OFF) + +/* A page B owns from the start (B's offset 0 is the window it lacks). */ +#define B_HOME_OFF (CHILD_SIZE / 2) +#define B_HOME_GPA (GPA + B_HOME_OFF) + +#define FRAG_OFF 0ULL +#define FRAG_LEN 0x200000ULL /* one 2 MiB-aligned block */ +#define HUGE_OFF 0x200000ULL /* untouched 2 MiB block */ +#define FRAG_MAGIC 0x5a5a5a5a5a5a5a5aULL + +/* ------------------------------------------------------------------------ */ +/* Guest: read args from a fixed GVA, act, report. */ + +struct guest_args { + uint64_t write_gpa; /* 0 = skip */ + uint64_t write_val; + uint64_t read_gpa; /* 0 = skip; value returned via GUEST_SYNC */ + uint64_t scan_gpa; + uint64_t scan_pages; + uint64_t scan_value; +}; + +static void guest_code(struct guest_args *a) +{ + if (a->write_gpa) + *(volatile uint64_t *)a->write_gpa = a->write_val; + if (a->read_gpa) + GUEST_SYNC(*(volatile uint64_t *)a->read_gpa); + if (a->scan_pages) { + uint64_t i; + + for (i = 0; i < a->scan_pages; i++) { + uint64_t value = *(volatile uint64_t *)(a->scan_gpa + i * PAGE); + + GUEST_ASSERT_EQ(value, i & 1 ? a->scan_value : 0); + } + } + GUEST_DONE(); +} + +/* ------------------------------------------------------------------------ */ +/* the control process role: the only holder of the control fd. */ + +struct vmm_control { int ctl; }; + +static void ctl_open(struct vmm_control *c) +{ + c->ctl = open("/dev/gmem_provider", O_RDWR); + __TEST_REQUIRE(c->ctl >= 0, "gmem_provider not loaded"); +} + +static int ctl_new_child(struct vmm_control *c, int kvm_fd, uint64_t off, uint64_t len, + uint32_t allow) +{ + struct gmem_provider_new_child nc = { + .kvm_fd = kvm_fd, .flags = GMEM_PROVIDER_FLAG_MMAP_CAPABLE, + .offset = off, .len = len, .allow = allow, + }; + int fd = ioctl(c->ctl, GMEM_PROVIDER_NEW_CHILD, &nc); + + TEST_ASSERT(fd >= 0, "NEW_CHILD(%#llx,%#llx) errno=%d", + (unsigned long long)off, (unsigned long long)len, errno); + return fd; +} + +static int ctl_move(struct vmm_control *c, int src, int dst, uint64_t off, uint64_t len) +{ + struct gmem_provider_move mv = { .src_fd = src, .dst_fd = dst, + .offset = off, .len = len }; + + return ioctl(c->ctl, GMEM_PROVIDER_MOVE, &mv) ? -errno : 0; +} + +static void ctl_donate(struct vmm_control *c, int fd, uint64_t off, uint64_t len, bool reclaim) +{ + struct gmem_provider_donate d = { .fd = fd, .offset = off, .len = len }; + + TEST_ASSERT(!ioctl(c->ctl, reclaim ? GMEM_PROVIDER_RECLAIM + : GMEM_PROVIDER_DONATE, &d), + "%s errno=%d", reclaim ? "RECLAIM" : "DONATE", errno); +} + +static void ctl_scratch(struct vmm_control *c, bool on) +{ + struct gmem_provider_scratch sc = { .enable = on }; + + TEST_ASSERT(!ioctl(c->ctl, GMEM_PROVIDER_SET_SCRATCH, &sc), + "SET_SCRATCH errno=%d", errno); +} + +/* ------------------------------------------------------------------------ */ +/* VMM role: one child fd, one VM, one IOAS. */ + +struct vmm { + const char *name; + int child; /* control fd: windows, stats, allowed ioctls */ + int gmem; /* the VM's memory object: memslot fd, window mmaps */ + int dmabuf; /* the child's dma-buf: what KVM and iommufd both import */ + int iommufd; + uint32_t ioas, stdev, hwpt; + struct kvm_vm *vm; + struct kvm_vcpu *vcpu; + void *hva; /* host window over the whole child */ + gva_t args_gva; +}; + +static void vmm_create(struct vmm *v) +{ + struct vm_shape shape = { .mode = VM_MODE_DEFAULT, + .type = KVM_X86_SW_PROTECTED_VM }; + + v->vm = vm_create_shape_with_one_vcpu(shape, &v->vcpu, guest_code); + v->args_gva = vm_alloc_page(v->vm); +} + +static int mock_domain_create(struct vmm *v) +{ + struct iommu_test_cmd cmd = { + .size = sizeof(cmd), .op = IOMMU_TEST_OP_MOCK_DOMAIN, .id = v->ioas, + }; + + if (ioctl(v->iommufd, IOMMU_TEST_CMD, &cmd)) + return -errno; + v->stdev = cmd.mock_domain.out_stdev_id; + v->hwpt = cmd.mock_domain.out_hwpt_id; + return 0; +} + +static bool vmm_iova_mapped(struct vmm *v, uint64_t iova) +{ + struct iommu_test_cmd cmd = { + .size = sizeof(cmd), .op = IOMMU_TEST_OP_MD_CHECK_MAPPED, .id = v->hwpt, + .check_mapped = { .mapped = true, .iova = iova, .length = PAGE }, + }; + + return !ioctl(v->iommufd, IOMMU_TEST_CMD, &cmd); +} + +static uint64_t vmm_iova_phys(struct vmm *v, uint64_t iova) +{ + struct iommu_test_cmd cmd = { + .size = sizeof(cmd), .op = IOMMU_TEST_OP_MD_IOVA_TO_PHYS, .id = v->hwpt, + .iova_to_phys = { .iova = iova }, + }; + + if (ioctl(v->iommufd, IOMMU_TEST_CMD, &cmd)) + return 0; + return cmd.iova_to_phys.out_phys; +} + +/* + * Whether the importer understands the full get_phys() contract: several + * ranges, holes, and a ranged invalidation it answers by re-reading the + * layout. Without that, iommufd maps only a buffer whose layout is one + * range (-EOPNOTSUPP otherwise), and any change to the buffer costs the + * device its whole mapping for good. Detected at attach time from B, + * whose child has a hole at launch; the device-plane checks below say + * what to expect either way. + */ +static bool ranged_import = true; + +/* Map [off, off+len) of the child into the IOAS at the matching IOVA. */ +static int vmm_map_range(struct vmm *v, uint64_t off, uint64_t len) +{ + struct iommu_ioas_map_file map = { + .size = sizeof(map), + .flags = IOMMU_IOAS_MAP_FIXED_IOVA | IOMMU_IOAS_MAP_READABLE | + IOMMU_IOAS_MAP_WRITEABLE, + .ioas_id = v->ioas, .fd = v->dmabuf, + .start = off, .length = len, .iova = IOVA + off, + }; + + return ioctl(v->iommufd, IOMMU_IOAS_MAP_FILE, &map) ? -errno : 0; +} + +/* + * the control process hands the VMM its child fd and the range it owns today. The VMM + * creates the VM's guest_memfd from it (provider fd), binds the whole view + * to the memslot, and maps the owned part into the IOAS with + * IOMMU_IOAS_MAP_FILE on the guest_memfd's dma-buf. iommufd refuses to map a range + * guest_memfd reports as a hole, so a VMM maps what it has and maps more + * when the control process tells it a range arrived (see scenario_move). + */ +struct create_race { + int child; + atomic_bool stop; + atomic_int error; +}; + +static void *create_race_thread(void *arg) +{ + struct create_race *race = arg; + struct gmem_provider_present present = { .offset = RO_OFF, .len = PAGE }; + struct gmem_provider_readonly readonly = { .offset = RO_OFF, .len = PAGE }; + + while (!atomic_load(&race->stop)) { + present.present = 0; + if (ioctl(race->child, GMEM_PROVIDER_SET_PRESENT, &present)) + break; + present.present = 1; + if (ioctl(race->child, GMEM_PROVIDER_SET_PRESENT, &present)) + break; + readonly.readonly = 1; + if (ioctl(race->child, GMEM_PROVIDER_SET_READONLY, &readonly)) + break; + readonly.readonly = 0; + if (ioctl(race->child, GMEM_PROVIDER_SET_READONLY, &readonly)) + break; + } + if (!atomic_load(&race->stop)) + atomic_store(&race->error, errno ?: EIO); + return NULL; +} + +static void finish_create_race(struct create_race *race, pthread_t thread) +{ + struct gmem_provider_present present = { + .offset = RO_OFF, .len = PAGE, .present = 1, + }; + struct gmem_provider_readonly readonly = { + .offset = RO_OFF, .len = PAGE, .readonly = 0, + }; + + atomic_store(&race->stop, true); + TEST_ASSERT(!pthread_join(thread, NULL), "pthread_join"); + TEST_ASSERT(!atomic_load(&race->error), "create race errno=%d", + atomic_load(&race->error)); + TEST_ASSERT(!ioctl(race->child, GMEM_PROVIDER_SET_PRESENT, &present), + "restore present errno=%d", errno); + TEST_ASSERT(!ioctl(race->child, GMEM_PROVIDER_SET_READONLY, &readonly), + "restore writable errno=%d", errno); +} + +static void vmm_attach(struct vmm *v, int child_fd, uint64_t own_off, + uint64_t own_len, uint64_t gmem_flags, bool race_create) +{ + struct iommu_ioas_alloc alloc = { .size = sizeof(alloc) }; + struct create_race race = { .child = child_fd }; + void *reservation; + pthread_t thread; + int r; + + v->child = child_fd; + + /* + * KVM leg: a guest_memfd whose memory the child provides. The child + * is the provider fd; guest_memfd is the VM's one memory object. The + * host window is the provider fd's mmap. + */ + v->dmabuf = ioctl(v->child, GMEM_PROVIDER_GET_DMABUF); + TEST_ASSERT(v->dmabuf >= 0, "GMEM_PROVIDER_GET_DMABUF errno=%d", errno); + if (race_create) + TEST_ASSERT(!pthread_create(&thread, NULL, create_race_thread, &race), + "pthread_create"); + v->gmem = vm_create_guest_memfd_dmabuf(v->vm, CHILD_SIZE, gmem_flags, + v->dmabuf); + if (race_create) + finish_create_race(&race, thread); + reservation = mmap(NULL, CHILD_SIZE + FRAG_LEN, PROT_NONE, + MAP_PRIVATE | MAP_ANONYMOUS | MAP_NORESERVE, -1, 0); + TEST_ASSERT(reservation != MAP_FAILED, "%s: reserve HVA errno=%d", + v->name, errno); + v->hva = (void *)(((uintptr_t)reservation + FRAG_LEN - 1) & + ~((uintptr_t)FRAG_LEN - 1)); + v->hva = mmap(v->hva, CHILD_SIZE, PROT_READ | PROT_WRITE, + MAP_SHARED | MAP_FIXED, v->child, 0); + TEST_ASSERT(v->hva != MAP_FAILED, "%s: mmap child errno=%d", v->name, + errno); + r = __vm_set_user_memory_region2(v->vm, 10, KVM_MEM_GUEST_MEMFD, GPA, + CHILD_SIZE, v->hva, v->gmem, 0); + TEST_ASSERT(!r, "%s: SET_USER_MEMORY_REGION2 r=%d errno=%d", v->name, r, errno); + virt_map(v->vm, GPA, GPA, CHILD_SIZE / PAGE); + /* Private, so every guest access goes through get_pfn(), not the HVA. */ + vm_mem_set_private(v->vm, GPA, CHILD_SIZE); + + /* DMA leg: the same guest_memfd, into an IOAS on a mock domain. */ + v->iommufd = open("/dev/iommu", O_RDWR); + __TEST_REQUIRE(v->iommufd >= 0, "iommufd unavailable"); + TEST_ASSERT(!ioctl(v->iommufd, IOMMU_IOAS_ALLOC, &alloc), "IOAS_ALLOC errno=%d", errno); + v->ioas = alloc.out_ioas_id; + + /* A mock device on a mock domain, so the IOAS is really programmed. */ + r = mock_domain_create(v); + TEST_ASSERT(!r, "%s: mock domain r=%d (CONFIG_IOMMUFD_TEST?)", v->name, r); + + r = vmm_map_range(v, own_off, own_len); + if (r == -EOPNOTSUPP && own_off) { + /* B's layout has a hole at launch; only a ranged importer takes it. */ + ranged_import = false; + pr_info(" iommufd maps one range; device checks expect full revoke\n"); + return; + } + TEST_ASSERT(!r, "%s: IOAS_MAP_FILE(dma-buf) r=%d", v->name, r); +} + +/* + * Continue the guest where it stopped. KVM_RUN returns -EFAULT for + * KVM_EXIT_MEMORY_FAULT, which several scenarios expect, so do not assert + * on it here; the expect_*() helpers check the exit. + */ +static void vmm_resume(struct vmm *v) +{ + int r = _vcpu_run(v->vcpu); + + TEST_ASSERT(!r || (errno == EFAULT && + v->vcpu->run->exit_reason == KVM_EXIT_MEMORY_FAULT), + "%s: KVM_RUN r=%d errno=%d exit=%s", v->name, r, errno, + exit_reason_str(v->vcpu->run->exit_reason)); +} + +/* Run the guest from the top with fresh args. */ +static void vmm_run(struct vmm *v, uint64_t write_gpa, uint64_t write_val, + uint64_t read_gpa) +{ + struct guest_args *a = addr_gva2hva(v->vm, v->args_gva); + + a->write_gpa = write_gpa; + a->write_val = write_val; + a->read_gpa = read_gpa; + a->scan_pages = 0; + vcpu_arch_set_entry_point(v->vcpu, guest_code); + vcpu_args_set(v->vcpu, 1, v->args_gva); + vmm_resume(v); +} + +static void vmm_scan(struct vmm *v, uint64_t gpa, uint64_t pages, + uint64_t value) +{ + struct guest_args *a = addr_gva2hva(v->vm, v->args_gva); + + memset(a, 0, sizeof(*a)); + a->scan_gpa = gpa; + a->scan_pages = pages; + a->scan_value = value; + vcpu_arch_set_entry_point(v->vcpu, guest_code); + vcpu_args_set(v->vcpu, 1, v->args_gva); + vmm_resume(v); +} + +static void expect_done(struct vmm *v) +{ + struct kvm_run *run = v->vcpu->run; + struct ucall uc; + + TEST_ASSERT(run->exit_reason != KVM_EXIT_MEMORY_FAULT, + "%s: unexpected memory fault gpa=%#llx size=%#llx flags=%#llx", + v->name, (unsigned long long)run->memory_fault.gpa, + (unsigned long long)run->memory_fault.size, + (unsigned long long)run->memory_fault.flags); + TEST_ASSERT(get_ucall(v->vcpu, &uc) == UCALL_DONE, "%s: guest exit %s", + v->name, exit_reason_str(run->exit_reason)); +} + +static uint64_t expect_sync_then_done(struct vmm *v) +{ + struct ucall uc; + uint64_t val; + + TEST_ASSERT(get_ucall(v->vcpu, &uc) == UCALL_SYNC, "%s: guest exit %s", + v->name, exit_reason_str(v->vcpu->run->exit_reason)); + val = uc.args[1]; + vmm_resume(v); + expect_done(v); + return val; +} + +static void expect_memory_fault(struct vmm *v, uint64_t gpa) +{ + struct kvm_run *run = v->vcpu->run; + + TEST_ASSERT(run->exit_reason == KVM_EXIT_MEMORY_FAULT, + "%s: want KVM_EXIT_MEMORY_FAULT, got %s", v->name, + exit_reason_str(run->exit_reason)); + TEST_ASSERT(run->memory_fault.gpa == gpa, "%s: fault gpa %#llx, want %#llx", + v->name, (unsigned long long)run->memory_fault.gpa, + (unsigned long long)gpa); +} + +static void vmm_stats(struct vmm *v, struct gmem_provider_stats *st) +{ + TEST_ASSERT(!ioctl(v->child, GMEM_PROVIDER_GET_STATS, st), + "%s: GET_STATS errno=%d", v->name, errno); +} + +static void vmm_destroy(struct vmm *v) +{ + kvm_vm_free(v->vm); + close(v->gmem); + close(v->iommufd); + munmap(v->hva, CHILD_SIZE); + close(v->child); +} + +/* ------------------------------------------------------------------------ */ +/* Scenarios. */ + +static void scenario_launch(struct vmm *a, struct vmm *b) +{ + pr_info("1. launch\n"); + vmm_run(a, GPA, MAGIC_A, 0); + expect_done(a); + vmm_run(b, B_HOME_GPA, MAGIC_B, 0); + expect_done(b); + TEST_ASSERT(*(volatile uint64_t *)a->hva == MAGIC_A, "A's write missing"); + TEST_ASSERT(*(volatile uint64_t *)(b->hva + B_HOME_OFF) == MAGIC_B, "B's write missing"); + TEST_ASSERT(vmm_iova_mapped(a, IOVA) && vmm_iova_mapped(a, A_SHARED_IOVA), + "A: IOAS not fully mapped after launch"); +} + +static void scenario_move(struct vmm_control *c, struct vmm *a, struct vmm *b) +{ + struct gmem_provider_stats st; + int r; + + pr_info("2. move A -> B\n"); + + /* A writes into the window while it still owns it. */ + vmm_run(a, A_SHARED_GPA, MAGIC_A, 0); + expect_done(a); + + /* B does not own the window yet: a read faults out to B's VMM ... */ + vmm_run(b, 0, 0, B_SHARED_GPA); + expect_memory_fault(b, B_SHARED_GPA); + /* ... and iommufd refuses to map a hole (or a layout with one in it). */ + r = vmm_map_range(b, SHARED_OFF - B_OFF, SHARED_LEN); + TEST_ASSERT(r == (ranged_import ? -EFAULT : -EOPNOTSUPP), + "B: mapping an unowned window returned %d", r); + + /* A move outside the destination's carve is refused. */ + r = ctl_move(c, a->child, b->child, RO_OFF, PAGE); + TEST_ASSERT(r == -EINVAL, "MOVE outside B's carve returned %d", r); + + /* The real move. */ + r = ctl_move(c, a->child, b->child, SHARED_OFF, SHARED_LEN); + TEST_ASSERT(!r, "MOVE returned %d", r); + + vmm_stats(a, &st); + TEST_ASSERT(st.owned_pages == CHILD_SIZE / PAGE - SHARED_LEN / PAGE, + "A owns %llu pages after move", (unsigned long long)st.owned_pages); + vmm_stats(b, &st); + TEST_ASSERT(st.owned_pages == CHILD_SIZE / PAGE, + "B owns %llu pages after move", (unsigned long long)st.owned_pages); + + /* A's device mapping: exactly the window is a hole, neighbours intact. */ + TEST_ASSERT(!vmm_iova_mapped(a, A_SHARED_IOVA), "A: moved IOVA still mapped"); + TEST_ASSERT(!vmm_iova_mapped(a, A_SHARED_IOVA + SHARED_LEN - PAGE), + "A: last moved IOVA still mapped"); + if (ranged_import) { + TEST_ASSERT(vmm_iova_mapped(a, A_SHARED_IOVA - PAGE), "A: page before window lost"); + TEST_ASSERT(vmm_iova_mapped(a, IOVA), "A: base lost"); + } else { + /* Whole-buffer revocation: the untouched pages went with the window. */ + TEST_ASSERT(!vmm_iova_mapped(a, IOVA), "A: base survived a whole-buffer revoke"); + } + + /* A's guest: the window is gone. */ + vmm_run(a, 0, 0, A_SHARED_GPA); + expect_memory_fault(a, A_SHARED_GPA); + + /* B's guest: the window is here, and it holds what A wrote. */ + vmm_run(b, 0, 0, B_SHARED_GPA); + TEST_ASSERT(expect_sync_then_done(b) == MAGIC_A, "B does not see A's write"); + + /* + * B's device: the control process told B's VMM the window arrived; now + * it maps. B's layout is one range from here on, so this works with + * either importer. + */ + r = vmm_map_range(b, SHARED_OFF - B_OFF, SHARED_LEN); + TEST_ASSERT(!r, "B: mapping the received window r=%d", r); + TEST_ASSERT(vmm_iova_mapped(b, IOVA + (SHARED_OFF - B_OFF)), "B: window not mapped"); +} + +static void scenario_donate(struct vmm_control *c, struct vmm *a) +{ + struct gmem_provider_stats st; + uint64_t off = A_OFF + 8 * PAGE, gpa = GPA + 8 * PAGE, iova = IOVA + 8 * PAGE; + + pr_info("3. donate + reclaim\n"); + ctl_donate(c, a->child, off, 2 * PAGE, false); + vmm_stats(a, &st); + TEST_ASSERT(st.owned_pages == CHILD_SIZE / PAGE - SHARED_LEN / PAGE - 2, + "A owns %llu after donate", (unsigned long long)st.owned_pages); + if (ranged_import) { + TEST_ASSERT(!vmm_iova_mapped(a, iova) && !vmm_iova_mapped(a, iova + PAGE), + "A: donated IOVAs still mapped"); + TEST_ASSERT(vmm_iova_mapped(a, iova - PAGE) && vmm_iova_mapped(a, iova + 2 * PAGE), + "A: neighbours of donation lost"); + } + vmm_run(a, 0, 0, gpa); + expect_memory_fault(a, gpa); + + ctl_donate(c, a->child, off, 2 * PAGE, true); + vmm_stats(a, &st); + TEST_ASSERT(st.owned_pages == CHILD_SIZE / PAGE - SHARED_LEN / PAGE, + "A owns %llu after reclaim", (unsigned long long)st.owned_pages); + if (ranged_import) + TEST_ASSERT(vmm_iova_mapped(a, iova), "A: reclaimed IOVA not remapped"); + vmm_run(a, 0, 0, gpa); + expect_sync_then_done(a); +} + +static void scenario_readonly(struct vmm *a) +{ + struct gmem_provider_readonly ro = { .offset = RO_OFF, .len = PAGE, .readonly = 1 }; + + pr_info("4. read-only\n"); + TEST_ASSERT(!ioctl(a->child, GMEM_PROVIDER_SET_READONLY, &ro), + "SET_READONLY errno=%d", errno); + + vmm_run(a, RO_GPA, MAGIC_A, 0); + expect_memory_fault(a, RO_GPA); + TEST_ASSERT(*(volatile uint64_t *)(a->hva + RO_OFF) != MAGIC_A, "RO write landed"); + + /* Clear the bit and let the guest retry the same instruction. */ + ro.readonly = 0; + TEST_ASSERT(!ioctl(a->child, GMEM_PROVIDER_SET_READONLY, &ro), "clear RO errno=%d", errno); + vmm_resume(a); + expect_done(a); + TEST_ASSERT(*(volatile uint64_t *)(a->hva + RO_OFF) == MAGIC_A, + "write after clearing RO did not land"); +} + +static void scenario_scratch(struct vmm_control *c, struct vmm *a) +{ + uint64_t off = A_OFF + 8 * PAGE, gpa = GPA + 8 * PAGE, iova = IOVA + 8 * PAGE; + uint64_t own0, own_before, p0, p1; + + pr_info("5. scratch\n"); + if (!ranged_import) { + pr_info(" skipped: needs the device mapping A lost in scenario 2\n"); + return; + } + own0 = vmm_iova_phys(a, IOVA); + own_before = vmm_iova_phys(a, iova - PAGE); + TEST_ASSERT(own0 && own_before, "A: expected pages unmapped before scratch test"); + + ctl_scratch(c, true); + ctl_donate(c, a->child, off, 2 * PAGE, false); + + p0 = vmm_iova_phys(a, iova); + p1 = vmm_iova_phys(a, iova + PAGE); + TEST_ASSERT(p0 && p0 == p1, "scratch: donated IOVAs -> %#llx, %#llx; want one frame", + (unsigned long long)p0, (unsigned long long)p1); + TEST_ASSERT(p0 != own0 && p0 != own_before, "scratch frame is one of A's own"); + TEST_ASSERT(vmm_iova_phys(a, IOVA) == own0, "scratch: unrelated IOVA changed"); + + /* The guest reads the scratch page: no fault, and it is zero. */ + vmm_run(a, 0, 0, gpa); + TEST_ASSERT(expect_sync_then_done(a) == 0, "scratch page must read as zero"); + + ctl_donate(c, a->child, off, 2 * PAGE, true); + ctl_scratch(c, false); + TEST_ASSERT(vmm_iova_phys(a, iova) != p0, "after reclaim IOVA still on scratch"); +} + +static void scenario_fragmented(struct vmm_control *c, struct vmm *a) +{ + uint64_t p4k_before, p2m_before, p4k_after, p2m_after; + unsigned long i; + + pr_info("6. fragmented layout\n"); + for (i = 0; i < (FRAG_LEN / PAGE); i++) + *(uint64_t *)(a->hva + FRAG_OFF + i * PAGE) = FRAG_MAGIC; + *(uint64_t *)(a->hva + HUGE_OFF) = FRAG_MAGIC; + p4k_before = vm_get_stat(a->vm, pages_4k); + p2m_before = vm_get_stat(a->vm, pages_2m); + ctl_scratch(c, true); + for (i = 0; i < (FRAG_LEN / PAGE); i += 2) + ctl_donate(c, a->child, A_OFF + FRAG_OFF + i * PAGE, PAGE, false); + + vmm_scan(a, GPA + FRAG_OFF, FRAG_LEN / PAGE, FRAG_MAGIC); + expect_done(a); + p4k_after = vm_get_stat(a->vm, pages_4k); + TEST_ASSERT(p4k_after > p4k_before, + "fragmented block did not add 4K mappings"); + /* The fragmented block's own 2M mapping was zapped; measure from here. */ + p2m_before = vm_get_stat(a->vm, pages_2m); + + vmm_run(a, 0, 0, GPA + HUGE_OFF); + TEST_ASSERT(expect_sync_then_done(a) == FRAG_MAGIC, + "untouched block data changed"); + p2m_after = vm_get_stat(a->vm, pages_2m); + pr_info(" fragmented levels: 4K %llu->%llu, 2M %llu->%llu\n", + (unsigned long long)p4k_before, (unsigned long long)p4k_after, + (unsigned long long)p2m_before, (unsigned long long)p2m_after); + TEST_ASSERT(p2m_after > p2m_before, + "untouched block did not add a 2M mapping"); + + for (i = 0; i < (FRAG_LEN / PAGE); i += 2) + ctl_donate(c, a->child, A_OFF + FRAG_OFF + i * PAGE, PAGE, true); + ctl_scratch(c, false); +} + +static void scenario_allowlist(struct vmm *b) +{ + struct gmem_provider_present p = { .offset = 0, .len = PAGE, .present = 0 }; + struct gmem_provider_stats st; + + pr_info("7. allowlist\n"); + TEST_ASSERT(ioctl(b->child, GMEM_PROVIDER_SET_PRESENT, &p) && errno == EPERM, + "B: SET_PRESENT must be denied"); + vmm_stats(b, &st); + TEST_ASSERT(!(st.allow & GMEM_ALLOW_SET_PRESENT), "B allow=%#x", st.allow); + TEST_ASSERT(st.owned_pages == CHILD_SIZE / PAGE, "B lost pages to a denied ioctl"); +} + +/* 8. A window through the gmem fd dies with the range it covers. */ +static sigjmp_buf window_jmp; +static void window_sig(int sig) +{ + siglongjmp(window_jmp, sig); +} + +static void scenario_window(struct vmm_control *c, struct vmm *a) +{ + uint64_t off = A_OFF + 8 * PAGE; + struct sigaction sa = { .sa_handler = window_sig }, old_bus, old_segv; + volatile uint64_t *win; + int sig; + + pr_info("8. host window teardown\n"); + win = mmap(NULL, PAGE, PROT_READ | PROT_WRITE, MAP_SHARED, a->gmem, off); + TEST_ASSERT(win != MAP_FAILED, "mmap(gmem, window) errno=%d", errno); + *win = MAGIC_A; + TEST_ASSERT(*(volatile uint64_t *)(a->hva + off) == MAGIC_A, + "window write not visible through the child mapping"); + + ctl_donate(c, a->child, off, PAGE, false); + + sigaction(SIGBUS, &sa, &old_bus); + sigaction(SIGSEGV, &sa, &old_segv); + sig = sigsetjmp(window_jmp, 1); + if (!sig) { + (void)*win; /* must fault: the PTE was torn down */ + sigaction(SIGBUS, &old_bus, NULL); + sigaction(SIGSEGV, &old_segv, NULL); + TEST_FAIL("window still readable after the range was donated"); + } + sigaction(SIGBUS, &old_bus, NULL); + sigaction(SIGSEGV, &old_segv, NULL); + pr_info(" window access after donate: signal %d, as expected\n", sig); + + ctl_donate(c, a->child, off, PAGE, true); + munmap((void *)win, PAGE); +} + +int main(void) +{ + struct vmm_control ctl; + struct vmm a = { .name = "A" }, b = { .name = "B" }; + int a_fd, b_fd, sizing_fd; + + TEST_REQUIRE(kvm_check_cap(KVM_CAP_VM_TYPES) & BIT(KVM_X86_SW_PROTECTED_VM)); + ctl_open(&ctl); + vmm_create(&a); + vmm_create(&b); + + /* Size a CMA-backed root for both overlapping child windows. */ + sizing_fd = ctl_new_child(&ctl, a.vm->fd, 0, ROOT_SIZE, 0); + close(sizing_fd); + + /* + * The control process carves the root. A carve is a VM's whole view; + * carves may overlap, but page ownership is exclusive. A is created + * first and gets its full carve, including the shared window. B is + * created second and gets only unowned pages, so its view has a hole + * until the control process MOVEs the window over. + * Neither VMM gets a management ioctl; those live on the control fd. + */ + a_fd = ctl_new_child(&ctl, a.vm->fd, A_OFF, CHILD_SIZE, + GMEM_ALLOW_SET_PRESENT | GMEM_ALLOW_SET_READONLY | + GMEM_ALLOW_GET_DMABUF | + GMEM_ALLOW_GET_STATS); + b_fd = ctl_new_child(&ctl, b.vm->fd, B_OFF, CHILD_SIZE, + GMEM_ALLOW_GET_DMABUF | GMEM_ALLOW_GET_STATS); + vmm_attach(&a, a_fd, 0, CHILD_SIZE, GUEST_MEMFD_FLAG_MMAP, true); + vmm_attach(&b, b_fd, SHARED_LEN, CHILD_SIZE - SHARED_LEN, 0, false); + + scenario_launch(&a, &b); + scenario_move(&ctl, &a, &b); + scenario_donate(&ctl, &a); + scenario_readonly(&a); + scenario_scratch(&ctl, &a); + scenario_fragmented(&ctl, &a); + scenario_allowlist(&b); + scenario_window(&ctl, &a); + + vmm_destroy(&a); + vmm_destroy(&b); + close(ctl.ctl); + pr_info("gmem_poc: all scenarios passed\n"); + return 0; +} diff --git a/tools/testing/selftests/kvm/x86/gmem_provider_hugepage_test.c b/tools/testing/selftests/kvm/x86/gmem_provider_hugepage_test.c index fbbef9761e64..9c3bcfbe2cae 100644 --- a/tools/testing/selftests/kvm/x86/gmem_provider_hugepage_test.c +++ b/tools/testing/selftests/kvm/x86/gmem_provider_hugepage_test.c @@ -36,8 +36,18 @@ struct gmem_provider_setup { __u64 size; }; #define GMEM_PROVIDER_SETUP _IOW('G', 1, struct gmem_provider_setup) +#define GMEM_PROVIDER_GET_DMABUF _IO('G', 3) #define GMEM_PROVIDER_FLAG_MMAP_CAPABLE (1u << 0) +/* The dma-buf for a sample-provider child: what KVM and iommufd both import. */ +static int child_dmabuf(int child_fd) +{ + int fd = ioctl(child_fd, GMEM_PROVIDER_GET_DMABUF); + + TEST_ASSERT(fd >= 0, "GMEM_PROVIDER_GET_DMABUF errno=%d", errno); + return fd; +} + #define DATA_SLOT 10 #define DATA_GPA (1ULL << 32) /* 4G: 1G-aligned */ #define DATA_SIZE ((uint64_t)SZ_1G + SZ_2M) @@ -59,7 +69,8 @@ int main(void) struct kvm_vcpu *vcpu; struct kvm_vm *vm; struct ucall uc; - int gmem_ctl, gmem_fd, r; + int gmem_ctl, gmem_fd, dmabuf_fd, r; + int prov_fd; void *resv, *hva; uint64_t p4k, p2m, p1g; @@ -73,8 +84,17 @@ int main(void) setup.kvm_fd = vm->fd; setup.size = DATA_SIZE; - gmem_fd = ioctl(gmem_ctl, GMEM_PROVIDER_SETUP, &setup); - TEST_ASSERT(gmem_fd >= 0, "GMEM_PROVIDER_SETUP failed, errno %d", errno); + prov_fd = ioctl(gmem_ctl, GMEM_PROVIDER_SETUP, &setup); + TEST_ASSERT(prov_fd >= 0, "GMEM_PROVIDER_SETUP failed, errno %d", errno); + + /* + * The provider fd is not a guest_memfd; it is the provider fd. The + * guest_memfd KVM returns is what the memslot binds, what the host + * mmaps (through the provider's mmap), and what iommufd would map. + */ + dmabuf_fd = child_dmabuf(prov_fd); + gmem_fd = vm_create_guest_memfd_dmabuf(vm, DATA_SIZE, + GUEST_MEMFD_FLAG_MMAP, dmabuf_fd); /* * Map the provider at a 1G-aligned host VA so the slot's userspace_addr @@ -86,7 +106,7 @@ int main(void) TEST_ASSERT(resv != MAP_FAILED, "reserve VA failed, errno %d", errno); hva = (void *)(((uintptr_t)resv + SZ_1G - 1) & ~((uintptr_t)SZ_1G - 1)); hva = mmap(hva, DATA_SIZE, PROT_READ | PROT_WRITE, - MAP_SHARED | MAP_FIXED, gmem_fd, 0); + MAP_SHARED | MAP_FIXED, prov_fd, 0); TEST_ASSERT(hva != MAP_FAILED, "mmap(provider) failed, errno %d", errno); @@ -125,6 +145,8 @@ int main(void) kvm_vm_free(vm); munmap(hva, DATA_SIZE); close(gmem_fd); + close(dmabuf_fd); + close(prov_fd); close(gmem_ctl); return 0; } diff --git a/tools/testing/selftests/kvm/x86/gmem_provider_iommufd_test.c b/tools/testing/selftests/kvm/x86/gmem_provider_iommufd_test.c index 885dffa0b659..b5f4afae5383 100644 --- a/tools/testing/selftests/kvm/x86/gmem_provider_iommufd_test.c +++ b/tools/testing/selftests/kvm/x86/gmem_provider_iommufd_test.c @@ -6,8 +6,8 @@ * write is visible via the host mmap. * * The iommufd side exercises exactly the provider->iommufd path we just wired: - * GET_DMABUF on the provider fd -> IOMMU_IOAS_MAP_FILE, which walks - * iopt_map_dmabuf -> sym_..._iommufd_map -> gmem_provider_dma_buf_iommufd_map. + * GMEM_PROVIDER_GET_DMABUF on the child -> IOMMU_IOAS_MAP_FILE, which walks + * iopt_map_dmabuf -> dma_buf_get_phys -> the provider's get_phys op. * * The test opens the provider with GMEM_PROVIDER_FLAG_MMAP_CAPABLE at SETUP time; requires iommufd * available at /dev/iommu. Actual IOMMU page-table programming happens once @@ -36,8 +36,17 @@ struct gmem_provider_setup { __u64 size; }; #define GMEM_PROVIDER_SETUP _IOW('G', 1, struct gmem_provider_setup) -#define GMEM_PROVIDER_FLAG_MMAP_CAPABLE (1u << 0) #define GMEM_PROVIDER_GET_DMABUF _IO('G', 3) +#define GMEM_PROVIDER_FLAG_MMAP_CAPABLE (1u << 0) + +/* The dma-buf for a sample-provider child: what KVM and iommufd both import. */ +static int child_dmabuf(int child_fd) +{ + int fd = ioctl(child_fd, GMEM_PROVIDER_GET_DMABUF); + + TEST_ASSERT(fd >= 0, "GMEM_PROVIDER_GET_DMABUF errno=%d", errno); + return fd; +} #define DATA_SLOT 10 #define DATA_GPA (1ULL << 32) @@ -64,6 +73,7 @@ int main(void) struct kvm_vm *vm; struct ucall uc; int gmem_ctl, gmem_fd, dmabuf_fd, iommufd, r; + int prov_fd; void *hva; TEST_REQUIRE(kvm_check_cap(KVM_CAP_VM_TYPES) & BIT(KVM_X86_SW_PROTECTED_VM)); @@ -80,13 +90,23 @@ int main(void) vm = vm_create_shape_with_one_vcpu(shape, &vcpu, guest_code); setup.kvm_fd = vm->fd; setup.size = DATA_SIZE; - gmem_fd = ioctl(gmem_ctl, GMEM_PROVIDER_SETUP, &setup); - TEST_ASSERT(gmem_fd >= 0, "GMEM_PROVIDER_SETUP failed errno=%d", errno); + prov_fd = ioctl(gmem_ctl, GMEM_PROVIDER_SETUP, &setup); + TEST_ASSERT(prov_fd >= 0, "GMEM_PROVIDER_SETUP failed errno=%d", errno); + + /* + * The provider fd is not a guest_memfd. Export its dma-buf and hand + * that to KVM as the provider fd; the guest_memfd KVM returns is what + * the memslot binds and what the host mmaps (mmap goes through the + * exporter). iommufd would import the very same dma-buf. + */ + dmabuf_fd = child_dmabuf(prov_fd); + gmem_fd = vm_create_guest_memfd_dmabuf(vm, DATA_SIZE, + GUEST_MEMFD_FLAG_MMAP, dmabuf_fd); /* * 2) KVM side: host mmap + guest_memfd memslot (gmem-only when mmap-capable). */ - hva = mmap(NULL, DATA_SIZE, PROT_READ | PROT_WRITE, MAP_SHARED, gmem_fd, 0); + hva = mmap(NULL, DATA_SIZE, PROT_READ | PROT_WRITE, MAP_SHARED, prov_fd, 0); TEST_ASSERT(hva != MAP_FAILED, "provider mmap failed errno=%d", errno); r = __vm_set_user_memory_region2(vm, DATA_SLOT, KVM_MEM_GUEST_MEMFD, @@ -105,16 +125,15 @@ int main(void) /* * 3) iommufd side: allocate IOAS, get a dma-buf from the SAME provider fd, * and map it into the IOAS via IOMMU_IOAS_MAP_FILE. This drives - * iopt_map_dmabuf -> gmem_provider_dma_buf_iommufd_map. + * iopt_map_dmabuf -> dma_buf_get_phys -> the get_phys op. */ alloc.size = sizeof(alloc); r = ioctl(iommufd, IOMMU_IOAS_ALLOC, &alloc); TEST_ASSERT(!r, "IOMMU_IOAS_ALLOC failed errno=%d", errno); pr_info("iommufd: allocated ioas id=%u\n", alloc.out_ioas_id); - dmabuf_fd = ioctl(gmem_fd, GMEM_PROVIDER_GET_DMABUF); - TEST_ASSERT(dmabuf_fd >= 0, "GMEM_PROVIDER_GET_DMABUF failed errno=%d", errno); - pr_info("provider: exported dma-buf fd=%d\n", dmabuf_fd); + /* IOMMU_IOAS_MAP_FILE takes the same dma-buf KVM imported. */ + pr_info("guest_memfd fd=%d imports dma-buf fd=%d\n", gmem_fd, dmabuf_fd); map.size = sizeof(map); map.flags = IOMMU_IOAS_MAP_FIXED_IOVA | @@ -127,7 +146,7 @@ int main(void) r = ioctl(iommufd, IOMMU_IOAS_MAP_FILE, &map); TEST_ASSERT(!r, "IOMMU_IOAS_MAP_FILE(dma-buf) failed r=%d errno=%d\n" - " (gmem_provider_dma_buf_iommufd_map path)", + " (dma_buf_get_phys path)", r, errno); pr_info("iommufd: mapped provider dma-buf @ IOVA 0x%llx (0x%llx bytes)\n", (unsigned long long)map.iova, (unsigned long long)map.length); @@ -153,6 +172,7 @@ int main(void) close(iommufd); munmap(hva, DATA_SIZE); close(gmem_fd); + close(prov_fd); close(gmem_ctl); return 0; } diff --git a/tools/testing/selftests/kvm/x86/gmem_provider_readonly_test.c b/tools/testing/selftests/kvm/x86/gmem_provider_readonly_test.c new file mode 100644 index 000000000000..a1b8b7ba5640 --- /dev/null +++ b/tools/testing/selftests/kvm/x86/gmem_provider_readonly_test.c @@ -0,0 +1,172 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * gmem_provider_readonly_test - exercise per-range read-only from a provider. + * + * Marks a provider-backed page read-only via an ioctl on the provider fd and + * checks that KVM honours the provider's answer: the guest can still read the + * page, a guest write exits to userspace with KVM_EXIT_MEMORY_FAULT rather than + * landing, and clearing the bit lets the write through. This is the mechanism + * a hypervisor uses to protect a page it shares with the guest, such as a + * information page a helper VM reads, without giving up the mapping. + * + * The test opens the provider with GMEM_PROVIDER_FLAG_MMAP_CAPABLE at SETUP time + * (gmem-only). Load the module with a backing region of at least DATA_SIZE. + */ +#include <fcntl.h> +#include <errno.h> +#include <stdint.h> +#include <stdio.h> +#include <string.h> +#include <unistd.h> +#include <sys/ioctl.h> +#include <sys/mman.h> + +#include "test_util.h" +#include "kvm_util.h" +#include "processor.h" + +/* Mirrors samples/kvm/gmem_provider.h */ +struct gmem_provider_setup { + __s32 kvm_fd; + __u32 flags; + __u64 size; +}; + +struct gmem_provider_readonly { + __u64 offset; + __u64 len; + __u32 readonly; + __u32 pad; +}; + +#define GMEM_PROVIDER_SETUP _IOW('G', 1, struct gmem_provider_setup) +#define GMEM_PROVIDER_GET_DMABUF _IO('G', 3) +#define GMEM_PROVIDER_FLAG_MMAP_CAPABLE (1u << 0) +#define GMEM_PROVIDER_SET_READONLY _IOW('G', 4, struct gmem_provider_readonly) + +/* The dma-buf for a sample-provider child: what KVM and iommufd both import. */ +static int child_dmabuf(int child_fd) +{ + int fd = ioctl(child_fd, GMEM_PROVIDER_GET_DMABUF); + + TEST_ASSERT(fd >= 0, "GMEM_PROVIDER_GET_DMABUF errno=%d", errno); + return fd; +} + +#define DATA_SLOT 10 +#define DATA_GPA (1ULL << 32) +#define DATA_SIZE 0x200000ULL /* 2 MiB region */ +#define MAGIC 0x1234abcdULL +#define MAGIC2 0xfeedf00dULL + +/* + * Phase 1: read the page and report it. + * Phase 2: write to it. With the page read-only this never returns to the + * guest until userspace clears the bit; then it completes and the + * guest reports what it wrote. + */ +static void guest_code(void) +{ + GUEST_SYNC(*(volatile uint64_t *)DATA_GPA); + *(volatile uint64_t *)DATA_GPA = MAGIC2; + GUEST_SYNC(*(volatile uint64_t *)DATA_GPA); + GUEST_DONE(); +} + +int main(void) +{ + struct vm_shape shape = { + .mode = VM_MODE_DEFAULT, + .type = KVM_X86_SW_PROTECTED_VM, + }; + struct gmem_provider_setup setup = { .flags = GMEM_PROVIDER_FLAG_MMAP_CAPABLE }; + struct gmem_provider_readonly req; + struct kvm_vcpu *vcpu; + struct kvm_vm *vm; + struct ucall uc; + int gmem_ctl, gmem_fd, dmabuf_fd, r; + int prov_fd; + void *hva; + + TEST_REQUIRE(kvm_check_cap(KVM_CAP_VM_TYPES) & BIT(KVM_X86_SW_PROTECTED_VM)); + + gmem_ctl = open("/dev/gmem_provider", O_RDWR); + __TEST_REQUIRE(gmem_ctl >= 0, + "gmem_provider module not loaded (/dev/gmem_provider absent)"); + + vm = vm_create_shape_with_one_vcpu(shape, &vcpu, guest_code); + + setup.kvm_fd = vm->fd; + setup.size = DATA_SIZE; + prov_fd = ioctl(gmem_ctl, GMEM_PROVIDER_SETUP, &setup); + TEST_ASSERT(prov_fd >= 0, "GMEM_PROVIDER_SETUP failed, errno %d", errno); + + /* + * The provider fd is not a guest_memfd; it is the provider fd. The + * guest_memfd KVM returns is what the memslot binds, what the host + * mmaps (through the provider's mmap), and what iommufd would map. + */ + dmabuf_fd = child_dmabuf(prov_fd); + gmem_fd = vm_create_guest_memfd_dmabuf(vm, DATA_SIZE, + GUEST_MEMFD_FLAG_MMAP, dmabuf_fd); + + hva = mmap(NULL, DATA_SIZE, PROT_READ | PROT_WRITE, MAP_SHARED, prov_fd, 0); + TEST_ASSERT(hva != MAP_FAILED, "mmap(provider) failed, errno %d", errno); + + r = __vm_set_user_memory_region2(vm, DATA_SLOT, KVM_MEM_GUEST_MEMFD, + DATA_GPA, DATA_SIZE, hva, gmem_fd, 0); + TEST_ASSERT(!r, "KVM_SET_USER_MEMORY_REGION2 failed: %d errno %d", r, errno); + virt_map(vm, DATA_GPA, DATA_GPA, 1); + + /* Seed the page from the host before the guest ever touches it. */ + *(volatile uint64_t *)hva = MAGIC; + + /* 1) Make the page read-only for the guest. */ + req = (struct gmem_provider_readonly){ .offset = 0, .len = 4096, .readonly = 1 }; + r = ioctl(prov_fd, GMEM_PROVIDER_SET_READONLY, &req); + TEST_ASSERT(!r, "set readonly ioctl failed, errno %d", errno); + + /* 2) Guest read must still work and see the host's value. */ + vcpu_run(vcpu); + TEST_ASSERT(get_ucall(vcpu, &uc) == UCALL_SYNC, "expected UCALL_SYNC"); + TEST_ASSERT(uc.args[1] == MAGIC, "guest read 0x%lx, want MAGIC", + (unsigned long)uc.args[1]); + pr_info("read-only: guest read 0x%llx\n", MAGIC); + + /* 3) Guest write must exit to userspace, not land. */ + r = _vcpu_run(vcpu); + TEST_ASSERT(r == -1 && errno == EFAULT && + vcpu->run->exit_reason == KVM_EXIT_MEMORY_FAULT, + "read-only write: expected KVM_EXIT_MEMORY_FAULT (r=%d errno=%d exit_reason=%u %s)", + r, errno, vcpu->run->exit_reason, + exit_reason_str(vcpu->run->exit_reason)); + TEST_ASSERT(*(volatile uint64_t *)hva == MAGIC, + "guest write landed on a read-only page: host sees 0x%lx", + (unsigned long)*(volatile uint64_t *)hva); + pr_info("read-only: guest write exited with KVM_EXIT_MEMORY_FAULT, page unchanged\n"); + + /* 4) Make it writable again; the retried write must complete. */ + req.readonly = 0; + r = ioctl(prov_fd, GMEM_PROVIDER_SET_READONLY, &req); + TEST_ASSERT(!r, "clear readonly ioctl failed, errno %d", errno); + + vcpu_run(vcpu); + TEST_ASSERT(get_ucall(vcpu, &uc) == UCALL_SYNC, "expected UCALL_SYNC after clear"); + TEST_ASSERT(uc.args[1] == MAGIC2, "after clear guest read 0x%lx, want MAGIC2", + (unsigned long)uc.args[1]); + TEST_ASSERT(*(volatile uint64_t *)hva == MAGIC2, + "host sees 0x%lx after guest write, want MAGIC2", + (unsigned long)*(volatile uint64_t *)hva); + pr_info("writable: guest write 0x%llx landed -- read-only path works\n", MAGIC2); + + vcpu_run(vcpu); + TEST_ASSERT(get_ucall(vcpu, &uc) == UCALL_DONE, "expected UCALL_DONE"); + + kvm_vm_free(vm); + munmap(hva, DATA_SIZE); + close(gmem_fd); + close(dmabuf_fd); + close(prov_fd); + close(gmem_ctl); + return 0; +} diff --git a/tools/testing/selftests/kvm/x86/gmem_provider_revoke_test.c b/tools/testing/selftests/kvm/x86/gmem_provider_revoke_test.c index 415972aa8a7e..df2e9fbd91d8 100644 --- a/tools/testing/selftests/kvm/x86/gmem_provider_revoke_test.c +++ b/tools/testing/selftests/kvm/x86/gmem_provider_revoke_test.c @@ -38,9 +38,19 @@ struct gmem_provider_present { __u32 pad; }; #define GMEM_PROVIDER_SETUP _IOW('G', 1, struct gmem_provider_setup) +#define GMEM_PROVIDER_GET_DMABUF _IO('G', 3) #define GMEM_PROVIDER_FLAG_MMAP_CAPABLE (1u << 0) #define GMEM_PROVIDER_SET_PRESENT _IOW('G', 2, struct gmem_provider_present) +/* The dma-buf for a sample-provider child: what KVM and iommufd both import. */ +static int child_dmabuf(int child_fd) +{ + int fd = ioctl(child_fd, GMEM_PROVIDER_GET_DMABUF); + + TEST_ASSERT(fd >= 0, "GMEM_PROVIDER_GET_DMABUF errno=%d", errno); + return fd; +} + #define DATA_SLOT 10 #define DATA_GPA (1ULL << 32) #define DATA_SIZE 0x200000ULL /* 2 MiB region */ @@ -64,7 +74,8 @@ int main(void) struct kvm_vcpu *vcpu; struct kvm_vm *vm; struct ucall uc; - int gmem_ctl, gmem_fd, r; + int gmem_ctl, gmem_fd, dmabuf_fd, r; + int prov_fd; void *hva; TEST_REQUIRE(kvm_check_cap(KVM_CAP_VM_TYPES) & BIT(KVM_X86_SW_PROTECTED_VM)); @@ -77,10 +88,19 @@ int main(void) setup.kvm_fd = vm->fd; setup.size = DATA_SIZE; - gmem_fd = ioctl(gmem_ctl, GMEM_PROVIDER_SETUP, &setup); - TEST_ASSERT(gmem_fd >= 0, "GMEM_PROVIDER_SETUP failed, errno %d", errno); + prov_fd = ioctl(gmem_ctl, GMEM_PROVIDER_SETUP, &setup); + TEST_ASSERT(prov_fd >= 0, "GMEM_PROVIDER_SETUP failed, errno %d", errno); + + /* + * The provider fd is not a guest_memfd; it is the provider fd. The + * guest_memfd KVM returns is what the memslot binds, what the host + * mmaps (through the provider's mmap), and what iommufd would map. + */ + dmabuf_fd = child_dmabuf(prov_fd); + gmem_fd = vm_create_guest_memfd_dmabuf(vm, DATA_SIZE, + GUEST_MEMFD_FLAG_MMAP, dmabuf_fd); - hva = mmap(NULL, DATA_SIZE, PROT_READ | PROT_WRITE, MAP_SHARED, gmem_fd, 0); + hva = mmap(NULL, DATA_SIZE, PROT_READ | PROT_WRITE, MAP_SHARED, prov_fd, 0); TEST_ASSERT(hva != MAP_FAILED, "mmap(provider) failed, errno %d", errno); @@ -98,7 +118,7 @@ int main(void) /* 2) Revoke: mark absent and zap the guest NPT (provider->KVM). */ req = (struct gmem_provider_present){ .offset = 0, .len = 4096, .present = 0 }; - r = ioctl(gmem_fd, GMEM_PROVIDER_SET_PRESENT, &req); + r = ioctl(prov_fd, GMEM_PROVIDER_SET_PRESENT, &req); TEST_ASSERT(!r, "revoke ioctl failed, errno %d", errno); /* 3) Guest re-reads -> re-fault into absent get_pfn -> must NOT see MAGIC. */ @@ -117,7 +137,7 @@ int main(void) /* 4) Restore: mark present again. */ req.present = 1; - r = ioctl(gmem_fd, GMEM_PROVIDER_SET_PRESENT, &req); + r = ioctl(prov_fd, GMEM_PROVIDER_SET_PRESENT, &req); TEST_ASSERT(!r, "restore ioctl failed, errno %d", errno); /* 5) Re-enter: the fault re-maps via get_pfn, guest reads MAGIC again. */ @@ -130,6 +150,8 @@ int main(void) kvm_vm_free(vm); munmap(hva, DATA_SIZE); close(gmem_fd); + close(dmabuf_fd); + close(prov_fd); close(gmem_ctl); return 0; } diff --git a/tools/testing/selftests/kvm/x86/gmem_provider_test.c b/tools/testing/selftests/kvm/x86/gmem_provider_test.c index d7caa11616df..d926b2033843 100644 --- a/tools/testing/selftests/kvm/x86/gmem_provider_test.c +++ b/tools/testing/selftests/kvm/x86/gmem_provider_test.c @@ -1,195 +1,11 @@ // SPDX-License-Identifier: GPL-2.0 -/* - * gmem_provider_test - exercise the samples/kvm gmem_provider module through a - * full SEV-SNP guest launch, then re-bind the same provider fd to a second VM - * (the live-update path). - * - * The provider module must be loaded first, in either mode: - * insmod gmem_provider.ko addr=0x5D40000000 len=0x1000000 # page-less - * insmod gmem_provider.ko # CMA fallback - * - * The test skips (KSFT_SKIP) if /dev/gmem_provider or SNP support is absent. - * - * Note: with the current kvm_gmem_populate() ABI the launch source is an - * ordinary anonymous buffer (pinned via get_user_pages_fast); the provider's - * populate() copies it into the backing. No /dev/mem mapping is needed. - */ +/* CoCo support for dma-buf-backed guest_memfd is outside this series. */ #include <stdio.h> -#include <stdlib.h> -#include <string.h> -#include <unistd.h> -#include <fcntl.h> -#include <errno.h> -#include <sys/ioctl.h> -#include <sys/mman.h> -#include <stdint.h> -#include <linux/kvm.h> -#include <asm/kvm.h> -/* Mirrors samples/kvm/gmem_provider.h */ -struct gmem_provider_setup { - int32_t kvm_fd; - uint32_t pad; - uint64_t size; -}; -#define GMEM_PROVIDER_SETUP _IOW('G', 1, struct gmem_provider_setup) - -#define KSFT_PASS 0 -#define KSFT_FAIL 1 #define KSFT_SKIP 4 -#define GUEST_MEM_SIZE (16UL * 1024 * 1024) -#define PAGE_SIZE_4K 4096UL - -static int sev_ioctl(int vm_fd, int sev_fd, int cmd, void *data) -{ - struct kvm_sev_cmd sev_cmd = { - .id = cmd, - .data = (uint64_t)(unsigned long)data, - .sev_fd = sev_fd, - }; - return ioctl(vm_fd, KVM_MEMORY_ENCRYPT_OP, &sev_cmd); -} - -static int launch_snp_vm(int kvm_fd, int gmem_fd, void *src, int vm_num) -{ - int vm_fd, vcpu_fd, sev_fd, ret; - - printf("[VM%d] KVM_CREATE_VM (SNP)\n", vm_num); - vm_fd = ioctl(kvm_fd, KVM_CREATE_VM, KVM_X86_SNP_VM); - if (vm_fd < 0) { perror("KVM_CREATE_VM"); return -1; } - - sev_fd = open("/dev/sev", O_RDWR); - if (sev_fd < 0) { perror("open /dev/sev"); close(vm_fd); return -1; } - - struct kvm_sev_init init = { 0 }; - ret = sev_ioctl(vm_fd, sev_fd, KVM_SEV_INIT2, &init); - if (ret) { perror("KVM_SEV_INIT2"); goto out; } - - struct kvm_userspace_memory_region2 region = { - .slot = 0, - .flags = KVM_MEM_GUEST_MEMFD, - .guest_phys_addr = 0, - .memory_size = GUEST_MEM_SIZE, - .userspace_addr = (uint64_t)(unsigned long)src, - .guest_memfd = gmem_fd, - .guest_memfd_offset = 0, - }; - ret = ioctl(vm_fd, KVM_SET_USER_MEMORY_REGION2, ®ion); - if (ret) { perror("KVM_SET_USER_MEMORY_REGION2"); goto out; } - - struct kvm_memory_attributes attrs = { - .address = 0, - .size = GUEST_MEM_SIZE, - .attributes = KVM_MEMORY_ATTRIBUTE_PRIVATE, - }; - ret = ioctl(vm_fd, KVM_SET_MEMORY_ATTRIBUTES, &attrs); - if (ret) { perror("KVM_SET_MEMORY_ATTRIBUTES"); goto out; } - - struct kvm_sev_snp_launch_start start = { .policy = 0x30000 }; - ret = sev_ioctl(vm_fd, sev_fd, KVM_SEV_SNP_LAUNCH_START, &start); - if (ret) { perror("SNP_LAUNCH_START"); goto out; } - - printf("[VM%d] SNP_LAUNCH_UPDATE (code + zero, %luMB)\n", - vm_num, GUEST_MEM_SIZE >> 20); - struct kvm_sev_snp_launch_update update = { - .gfn_start = 0, - .uaddr = (uint64_t)(unsigned long)src, - .len = PAGE_SIZE_4K, - .type = KVM_SEV_SNP_PAGE_TYPE_NORMAL, - }; - ret = sev_ioctl(vm_fd, sev_fd, KVM_SEV_SNP_LAUNCH_UPDATE, &update); - if (ret) { perror("SNP_LAUNCH_UPDATE code"); goto out; } - - struct kvm_sev_snp_launch_update update_zero = { - .gfn_start = 1, - .uaddr = (uint64_t)(unsigned long)(src + PAGE_SIZE_4K), - .len = GUEST_MEM_SIZE - PAGE_SIZE_4K, - .type = KVM_SEV_SNP_PAGE_TYPE_ZERO, - }; - ret = sev_ioctl(vm_fd, sev_fd, KVM_SEV_SNP_LAUNCH_UPDATE, &update_zero); - if (ret) { perror("SNP_LAUNCH_UPDATE zero"); goto out; } - - struct kvm_sev_snp_launch_finish finish = { 0 }; - ret = sev_ioctl(vm_fd, sev_fd, KVM_SEV_SNP_LAUNCH_FINISH, &finish); - if (ret) { perror("SNP_LAUNCH_FINISH"); goto out; } - - vcpu_fd = ioctl(vm_fd, KVM_CREATE_VCPU, 0); - if (vcpu_fd < 0) { perror("KVM_CREATE_VCPU"); ret = -1; goto out; } - - printf("[VM%d] *** launch succeeded ***\n", vm_num); - close(vcpu_fd); - ret = 0; -out: - close(sev_fd); - close(vm_fd); - return ret; -} - int main(void) { - int kvm_fd, gmem_ctl_fd, gmem_fd, tmp_vm; - void *src; - - setbuf(stdout, NULL); - - kvm_fd = open("/dev/kvm", O_RDWR); - if (kvm_fd < 0) { - printf("SKIP: cannot open /dev/kvm (%s)\n", strerror(errno)); - return KSFT_SKIP; - } - - gmem_ctl_fd = open("/dev/gmem_provider", O_RDWR); - if (gmem_ctl_fd < 0) { - printf("SKIP: /dev/gmem_provider not present -- load gmem_provider.ko (%s)\n", - strerror(errno)); - return KSFT_SKIP; - } - - /* Provider setup needs a VM fd; ownership transfers to each VM on bind. */ - tmp_vm = ioctl(kvm_fd, KVM_CREATE_VM, KVM_X86_SNP_VM); - if (tmp_vm < 0) { - printf("SKIP: cannot create SNP VM -- host not SNP-capable? (%s)\n", - strerror(errno)); - return KSFT_SKIP; - } - - struct gmem_provider_setup setup = { - .kvm_fd = tmp_vm, - .size = GUEST_MEM_SIZE, - }; - gmem_fd = ioctl(gmem_ctl_fd, GMEM_PROVIDER_SETUP, &setup); - if (gmem_fd < 0) { - perror("GMEM_PROVIDER_SETUP"); - return KSFT_FAIL; - } - close(tmp_vm); - printf("provider gmem_fd = %d (persists across VMs)\n", gmem_fd); - - /* Launch source: ordinary anonymous memory holding a HLT at gfn 0. */ - src = mmap(NULL, GUEST_MEM_SIZE, PROT_READ | PROT_WRITE, - MAP_PRIVATE | MAP_ANONYMOUS, -1, 0); - if (src == MAP_FAILED) { perror("mmap src"); return KSFT_FAIL; } - memset(src, 0, GUEST_MEM_SIZE); - ((uint8_t *)src)[0] = 0xf4; /* HLT */ - - if (launch_snp_vm(kvm_fd, gmem_fd, src, 1) != 0) { - printf("FAIL: VM1 launch failed\n"); - return KSFT_FAIL; - } - - usleep(100000); - - /* Re-bind the SAME provider fd to a fresh VM (live-update path). */ - if (launch_snp_vm(kvm_fd, gmem_fd, src, 2) != 0) { - printf("FAIL: VM2 re-bind launch failed\n"); - return KSFT_FAIL; - } - - printf("PASS: page-less/provider-backed SNP launch + re-bind succeeded\n"); - munmap(src, GUEST_MEM_SIZE); - close(gmem_fd); - close(gmem_ctl_fd); - close(kvm_fd); - return KSFT_PASS; + puts("SKIP: dma-buf-backed CoCo is outside this series"); + return KSFT_SKIP; } diff --git a/tools/testing/selftests/kvm/x86/gmem_provider_vfio_test.c b/tools/testing/selftests/kvm/x86/gmem_provider_vfio_test.c index 97a66a295b14..7f3c640aea78 100644 --- a/tools/testing/selftests/kvm/x86/gmem_provider_vfio_test.c +++ b/tools/testing/selftests/kvm/x86/gmem_provider_vfio_test.c @@ -5,7 +5,7 @@ * * Steps: * 1. SETUP provider (KVM VM shape not required for this test). - * 2. iommufd IOAS + GET_DMABUF + IOMMU_IOAS_MAP_FILE - IOAS holds the + * 2. guest_memfd from the provider fd; IOMMU_IOAS_MAP_FILE on its dma-buf - IOAS holds the * provider region. * 3. Open a vfio-pci cdev (default /dev/vfio/devices/vfio0, overridable * via GMEM_VFIO_CDEV env), VFIO_DEVICE_BIND_IOMMUFD, then @@ -39,8 +39,17 @@ struct gmem_provider_setup { __u64 size; }; #define GMEM_PROVIDER_SETUP _IOW('G', 1, struct gmem_provider_setup) -#define GMEM_PROVIDER_FLAG_MMAP_CAPABLE (1u << 0) #define GMEM_PROVIDER_GET_DMABUF _IO('G', 3) +#define GMEM_PROVIDER_FLAG_MMAP_CAPABLE (1u << 0) + +/* The dma-buf for a sample-provider child: what KVM and iommufd both import. */ +static int child_dmabuf(int child_fd) +{ + int fd = ioctl(child_fd, GMEM_PROVIDER_GET_DMABUF); + + TEST_ASSERT(fd >= 0, "GMEM_PROVIDER_GET_DMABUF errno=%d", errno); + return fd; +} #define DATA_SIZE 0x200000ULL #define IOVA_BASE (1ULL << 34) @@ -53,7 +62,7 @@ int main(void) struct vfio_device_bind_iommufd bind = {}; struct vfio_device_attach_iommufd_pt att = {}; const char *cdev_path; - int gmem_ctl, gmem_fd, dmabuf_fd, iommufd_fd, vfio_fd, r; + int gmem_ctl, gmem_fd, dmabuf_fd, prov_fd, vm_fd = -1, iommufd_fd, vfio_fd, r; gmem_ctl = open("/dev/gmem_provider", O_RDWR); __TEST_REQUIRE(gmem_ctl >= 0, "gmem_provider module not loaded"); @@ -80,20 +89,34 @@ int main(void) vm = ioctl(kvm, KVM_CREATE_VM, 0); TEST_ASSERT(vm >= 0, "KVM_CREATE_VM errno=%d", errno); setup.kvm_fd = vm; + vm_fd = vm; close(kvm); } setup.size = DATA_SIZE; - gmem_fd = ioctl(gmem_ctl, GMEM_PROVIDER_SETUP, &setup); - TEST_ASSERT(gmem_fd >= 0, "GMEM_PROVIDER_SETUP errno=%d", errno); + prov_fd = ioctl(gmem_ctl, GMEM_PROVIDER_SETUP, &setup); + TEST_ASSERT(prov_fd >= 0, "GMEM_PROVIDER_SETUP errno=%d", errno); + + /* + * The provider fd is the provider fd of a guest_memfd, and iommufd maps + * the dma-buf; a device and a guest import the same one. + */ + { + struct kvm_create_guest_memfd cgm; + + dmabuf_fd = child_dmabuf(prov_fd); + cgm = (struct kvm_create_guest_memfd) { + .size = DATA_SIZE, + .flags = GUEST_MEMFD_FLAG_MMAP | GUEST_MEMFD_FLAG_USE_DMABUF, + .dmabuf_fd = dmabuf_fd, + }; + gmem_fd = ioctl(vm_fd, KVM_CREATE_GUEST_MEMFD, &cgm); + TEST_ASSERT(gmem_fd >= 0, "KVM_CREATE_GUEST_MEMFD(dmabuf) errno=%d", errno); + } - /* IOAS + provider dma-buf */ + /* IOAS + the same dma-buf KVM imported. */ alloc.size = sizeof(alloc); r = ioctl(iommufd_fd, IOMMU_IOAS_ALLOC, &alloc); TEST_ASSERT(!r, "IOMMU_IOAS_ALLOC errno=%d", errno); - - dmabuf_fd = ioctl(gmem_fd, GMEM_PROVIDER_GET_DMABUF); - TEST_ASSERT(dmabuf_fd >= 0, "GET_DMABUF errno=%d", errno); - map.size = sizeof(map); map.flags = IOMMU_IOAS_MAP_FIXED_IOVA | IOMMU_IOAS_MAP_READABLE | IOMMU_IOAS_MAP_WRITEABLE; @@ -126,8 +149,8 @@ int main(void) pr_info("real passthrough device now has IOMMU domain covering provider region\n"); close(vfio_fd); - close(dmabuf_fd); close(iommufd_fd); + close(dmabuf_fd); close(gmem_fd); close(gmem_ctl); return 0; -- 2.47.3 ^ permalink raw reply [flat|nested] 28+ messages in thread
end of thread, other threads:[~2026-10-05 14:53 UTC | newest] Thread overview: 28+ messages (download: mbox.gz / follow: Atom feed) -- links below jump to the message on this page -- 2026-07-20 11:03 [RFC PATCH v2 00/11] KVM: Allow alternative providers of guest_memfd backed by PFNMAP memory David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 01/11] KVM: selftests: sev_smoke_test: Only run VM types the host offers David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 02/11] KVM: selftests: sev_init2_tests: Derive SEV availability from KVM David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 03/11] KVM: SEV: Remove struct page dependency from SNP gmem paths David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 04/11] KVM: guest_memfd: Introduce guest memory ops and route native gmem through them David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 05/11] iommufd: Look up private-interconnect phys via exporter symbols David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 06/11] iommufd: Plumb dma-buf memory-type (RAM vs MMIO) through the phys map David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 07/11] KVM: guest_memfd: Add ops-driven page revocation David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 08/11] samples/kvm: Add guest_memfd backing sample David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 09/11] selftests/kvm: gmem_provider KVM-only tests David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 10/11] selftests/kvm: gmem_provider iommufd tests David Woodhouse 2026-07-20 11:03 ` [RFC PATCH v2 11/11] samples/kvm, selftests/kvm: Allow the gmem_provider NVMe DMA test on arm64 David Woodhouse 2026-07-20 15:11 ` [RFC PATCH v2 00/11] KVM: Allow alternative providers of guest_memfd backed by PFNMAP memory Paolo Bonzini 2026-07-20 16:39 ` David Woodhouse 2026-07-23 0:24 ` Ackerley Tng 2026-07-23 9:40 ` David Woodhouse 2026-07-23 16:01 ` Ackerley Tng 2026-10-05 9:55 ` [RFC PATCH 0/6] KVM: guest_memfd: back guest_memfd with an imported dma-buf Fred Griffoul 2026-10-05 9:55 ` [RFC PATCH 1/6] KVM: guest_memfd: Add a writable result to get_pfn() Fred Griffoul 2026-10-05 9:55 ` [RFC PATCH 2/6] dma-buf: Add get_phys() to describe a physical run Fred Griffoul 2026-10-05 10:07 ` Christian König 2026-10-05 13:20 ` Fred Griffoul 2026-10-05 14:53 ` Christian König 2026-10-05 9:55 ` [RFC PATCH 3/6] dma-buf: Add ranged mapping invalidation Fred Griffoul 2026-10-05 10:08 ` Christian König 2026-10-05 9:55 ` [RFC PATCH 4/6] dma-buf: Allow dynamic attach without a device Fred Griffoul 2026-10-05 9:55 ` [RFC PATCH 5/6] KVM: guest_memfd: Add dma-buf backing Fred Griffoul 2026-10-05 9:55 ` [RFC PATCH 6/6] samples/kvm, selftests/kvm: Exercise " Fred Griffoul
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®