* [RFC PATCH 1/7] iommu: Add infrastructure for per-domain IOPT accounting
[not found] <20261001224531.765278-1-praan@google.com>
@ 2026-10-01 22:45 ` Pranjal Shrivastava
2026-10-01 22:45 ` [RFC PATCH 2/7] iommu: Implement domain-attributed page allocation Pranjal Shrivastava
` (5 subsequent siblings)
6 siblings, 0 replies; 7+ messages in thread
From: Pranjal Shrivastava @ 2026-10-01 22:45 UTC (permalink / raw)
To: Joerg Roedel, Will Deacon, Robin Murphy, Jason Gunthorpe,
Kevin Tian, Alex Williamson, David Matlack, Jonathan Corbet,
Shuah Khan, Randy Dunlap
Cc: Mostafa Saleh, Daniel Mentz, Samiullah Khawaja, iommu, kvm,
linux-kselftest, linux-doc, Logan Odell, Pranjal Shrivastava,
linux-fsdevel, linux-kernel
Currently, the IOMMU subsystem tracks the total number of pages
allocated for IO page tables globally via nr_iommu_pages, but lacks the
ability to attribute these allocations to specific domains / users.
Add an atomic `nr_pages` counter to struct iommu_domain and introduce an
anon union to store the iommu_domain (aliased with private) in ioptdesc
for the accounting infra. The union ensures that the ioptdesc layout
remains identical to struct page, maintaining the integrity of existing
static_assert checks. Utilizing the `private` field is safe here as IOMMU
page table memory is exclusively owned by the IOMMU subsystem and is not
managed by other core MM subsystems that typically claim this slot.
Co-developed-by: Logan Odell <loganodell@google.com>
Signed-off-by: Logan Odell <loganodell@google.com>
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
drivers/iommu/iommu-pages.h | 8 +++++++-
include/linux/iommu.h | 2 ++
2 files changed, 9 insertions(+), 1 deletion(-)
diff --git a/drivers/iommu/iommu-pages.h b/drivers/iommu/iommu-pages.h
index e9e605b5fa3a..edf75c81054f 100644
--- a/drivers/iommu/iommu-pages.h
+++ b/drivers/iommu/iommu-pages.h
@@ -25,7 +25,13 @@ struct ioptdesc {
u8 incoherent;
pgoff_t __index;
};
- void *_private;
+ /*
+ * Alias iommu_domain with page->private for IOPT memory accounting.
+ */
+ union {
+ void *_private;
+ struct iommu_domain *domain;
+ };
unsigned int __page_type;
atomic_t __page_refcount;
diff --git a/include/linux/iommu.h b/include/linux/iommu.h
index ac43b8b93f14..478b64f4084f 100644
--- a/include/linux/iommu.h
+++ b/include/linux/iommu.h
@@ -231,6 +231,8 @@ struct iommu_domain {
struct iommu_domain_geometry geometry;
int (*iopf_handler)(struct iopf_group *group);
+ atomic_long_t nr_pages;
+
union { /* cookie */
struct iommu_dma_cookie *iova_cookie;
struct iommu_dma_msi_cookie *msi_cookie;
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 7+ messages in thread* [RFC PATCH 2/7] iommu: Implement domain-attributed page allocation
[not found] <20261001224531.765278-1-praan@google.com>
2026-10-01 22:45 ` [RFC PATCH 1/7] iommu: Add infrastructure for per-domain IOPT accounting Pranjal Shrivastava
@ 2026-10-01 22:45 ` Pranjal Shrivastava
2026-10-01 22:45 ` [RFC PATCH 3/7] iommupt: Enable per-domain IOPT attribution Pranjal Shrivastava
` (4 subsequent siblings)
6 siblings, 0 replies; 7+ messages in thread
From: Pranjal Shrivastava @ 2026-10-01 22:45 UTC (permalink / raw)
To: Joerg Roedel, Will Deacon, Robin Murphy, Jason Gunthorpe,
Kevin Tian, Alex Williamson, David Matlack, Jonathan Corbet,
Shuah Khan, Randy Dunlap
Cc: Mostafa Saleh, Daniel Mentz, Samiullah Khawaja, iommu, kvm,
linux-kselftest, linux-doc, Logan Odell, Pranjal Shrivastava,
linux-fsdevel, linux-kernel
Implement the core logic for per-domain IOPT page accounting.
Introduce iommu_alloc_pages_node_sz_attributed() which increments
the domain's nr_pages counter & stores the domain ptr in the ioptdesc.
Update the free path to recover domain ptr and decrement the count.
Modify iommu_alloc_pages_node_sz() to be a legacy wrapper that passes a
NULL domain, ensuring backwards compatibility for non-domain use-cases.
Co-developed-by: Logan Odell <loganodell@google.com>
Signed-off-by: Logan Odell <loganodell@google.com>
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
drivers/iommu/iommu-pages.c | 35 ++++++++++++++++++++++++++++++++---
drivers/iommu/iommu-pages.h | 17 +++++++++++++++++
2 files changed, 49 insertions(+), 3 deletions(-)
diff --git a/drivers/iommu/iommu-pages.c b/drivers/iommu/iommu-pages.c
index 3bab175d8557..81fa01b763b7 100644
--- a/drivers/iommu/iommu-pages.c
+++ b/drivers/iommu/iommu-pages.c
@@ -29,8 +29,10 @@ static inline size_t ioptdesc_mem_size(struct ioptdesc *desc)
}
/**
- * iommu_alloc_pages_node_sz - Allocate a zeroed page of a given size from
- * specific NUMA node
+ * iommu_alloc_pages_node_sz_attributed - Allocate a zeroed page of a given size
+ * from specific NUMA node for a specific
+ * iommu domain
+ * @domain: IOMMU domain to which the page will belong
* @nid: memory NUMA node id
* @gfp: buddy allocator flags
* @size: Memory size to allocate, rounded up to a power of 2
@@ -40,7 +42,8 @@ static inline size_t ioptdesc_mem_size(struct ioptdesc *desc)
* returned allocation is round_up_pow_two(size) big, and is physically aligned
* to its size.
*/
-void *iommu_alloc_pages_node_sz(int nid, gfp_t gfp, size_t size)
+void *iommu_alloc_pages_node_sz_attributed(struct iommu_domain *domain, int nid,
+ gfp_t gfp, size_t size)
{
struct ioptdesc *iopt;
unsigned long pgcnt;
@@ -83,18 +86,44 @@ void *iommu_alloc_pages_node_sz(int nid, gfp_t gfp, size_t size)
mod_node_page_state(folio_pgdat(folio), NR_IOMMU_PAGES, pgcnt);
lruvec_stat_mod_folio(folio, NR_SECONDARY_PAGETABLE, pgcnt);
+ iopt->domain = domain;
+ if (domain)
+ atomic_long_add(pgcnt, &domain->nr_pages);
+
return folio_address(folio);
}
+EXPORT_SYMBOL_GPL(iommu_alloc_pages_node_sz_attributed);
+
+/**
+ * iommu_alloc_pages_node_sz - Allocate a zeroed page of a given size from
+ * specific NUMA node
+ * @nid: memory NUMA node id
+ * @gfp: buddy allocator flags
+ * @size: Memory size to allocate, rounded up to a power of 2
+ *
+ * Returns the virtual address of the allocated page. The page must be freed
+ * either by calling iommu_free_pages() or via iommu_put_pages_list(). The
+ * returned allocation is round_up_pow_two(size) big, and is physically aligned
+ * to its size.
+ */
+void *iommu_alloc_pages_node_sz(int nid, gfp_t gfp, size_t size)
+{
+ return iommu_alloc_pages_node_sz_attributed(NULL, nid, gfp, size);
+}
EXPORT_SYMBOL_GPL(iommu_alloc_pages_node_sz);
static void __iommu_free_desc(struct ioptdesc *iopt)
{
struct folio *folio = ioptdesc_folio(iopt);
const unsigned long pgcnt = folio_nr_pages(folio);
+ struct iommu_domain *domain = iopt->domain;
if (IOMMU_PAGES_USE_DMA_API)
WARN_ON_ONCE(iopt->incoherent);
+ if (domain)
+ atomic_long_sub(pgcnt, &domain->nr_pages);
+
mod_node_page_state(folio_pgdat(folio), NR_IOMMU_PAGES, -pgcnt);
lruvec_stat_mod_folio(folio, NR_SECONDARY_PAGETABLE, -pgcnt);
folio_put(folio);
diff --git a/drivers/iommu/iommu-pages.h b/drivers/iommu/iommu-pages.h
index edf75c81054f..a4a8e9ba57c5 100644
--- a/drivers/iommu/iommu-pages.h
+++ b/drivers/iommu/iommu-pages.h
@@ -55,6 +55,8 @@ static inline struct ioptdesc *virt_to_ioptdesc(void *virt)
return folio_ioptdesc(virt_to_folio(virt));
}
+void *iommu_alloc_pages_node_sz_attributed(struct iommu_domain *domain, int nid,
+ gfp_t gfp, size_t size);
void *iommu_alloc_pages_node_sz(int nid, gfp_t gfp, size_t size);
void iommu_free_pages(void *virt);
void iommu_put_pages_list(struct iommu_pages_list *list);
@@ -107,6 +109,21 @@ static inline void *iommu_alloc_pages_sz(gfp_t gfp, size_t size)
return iommu_alloc_pages_node_sz(NUMA_NO_NODE, gfp, size);
}
+/**
+ * iommu_alloc_pages_sz_attributed - Allocate a zeroed page of a given size from
+ * specific NUMA node for a specific domain
+ * @domain: iommu domain
+ * @gfp: buddy allocator flags
+ * @size: Memory size to allocate, this is rounded up to a power of 2
+ *
+ * Returns the virtual address of the allocated page.
+ */
+static inline void *iommu_alloc_pages_sz_attributed(struct iommu_domain *domain,
+ gfp_t gfp, size_t size)
+{
+ return iommu_alloc_pages_node_sz_attributed(domain, NUMA_NO_NODE, gfp, size);
+}
+
int iommu_pages_start_incoherent(void *virt, struct device *dma_dev);
int iommu_pages_start_incoherent_list(struct iommu_pages_list *list,
struct device *dma_dev);
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 7+ messages in thread* [RFC PATCH 3/7] iommupt: Enable per-domain IOPT attribution
[not found] <20261001224531.765278-1-praan@google.com>
2026-10-01 22:45 ` [RFC PATCH 1/7] iommu: Add infrastructure for per-domain IOPT accounting Pranjal Shrivastava
2026-10-01 22:45 ` [RFC PATCH 2/7] iommu: Implement domain-attributed page allocation Pranjal Shrivastava
@ 2026-10-01 22:45 ` Pranjal Shrivastava
2026-10-01 22:45 ` [RFC PATCH 4/7] vfio/type1: Expose IO page table usage via fdinfo Pranjal Shrivastava
` (3 subsequent siblings)
6 siblings, 0 replies; 7+ messages in thread
From: Pranjal Shrivastava @ 2026-10-01 22:45 UTC (permalink / raw)
To: Joerg Roedel, Will Deacon, Robin Murphy, Jason Gunthorpe,
Kevin Tian, Alex Williamson, David Matlack, Jonathan Corbet,
Shuah Khan, Randy Dunlap
Cc: Mostafa Saleh, Daniel Mentz, Samiullah Khawaja, iommu, kvm,
linux-kselftest, linux-doc, Logan Odell, Pranjal Shrivastava,
linux-fsdevel, linux-kernel
Update the _table_alloc() helper in the Generic Page Table library to
use the attributed allocator.
By leveraging the pt_iommu structure, which already contains a pointer
to the iommu_domain, the library can now automatically attribute page
table allocations for all drivers that utilize the generic_pt framework.
Thus, any IOMMU driver using the generic PT automatically gains IOPT
observability without any additional changes.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
drivers/iommu/generic_pt/iommu_pt.h | 7 +++++--
1 file changed, 5 insertions(+), 2 deletions(-)
diff --git a/drivers/iommu/generic_pt/iommu_pt.h b/drivers/iommu/generic_pt/iommu_pt.h
index 07ec2b3ab986..cf07946801d4 100644
--- a/drivers/iommu/generic_pt/iommu_pt.h
+++ b/drivers/iommu/generic_pt/iommu_pt.h
@@ -416,8 +416,9 @@ static inline struct pt_table_p *_table_alloc(struct pt_common *common,
struct pt_iommu *iommu_table = iommu_from_common(common);
struct pt_table_p *table_mem;
- table_mem = iommu_alloc_pages_node_sz(iommu_table->nid, gfp,
- log2_to_int(lg2sz));
+ table_mem = iommu_alloc_pages_node_sz_attributed(&iommu_table->domain,
+ iommu_table->nid, gfp,
+ log2_to_int(lg2sz));
if (!table_mem)
return ERR_PTR(-ENOMEM);
@@ -1177,6 +1178,8 @@ static void NS(deinit)(struct pt_iommu *iommu_table)
iommu_pages_stop_incoherent_list(&collect.pending.free_list,
iommu_table->iommu_device);
iommu_put_pages_list(&collect.pending.free_list);
+
+ WARN_ON_ONCE(atomic_long_read(&iommu_table->domain.nr_pages) != 0);
}
static const struct pt_iommu_ops NS(ops) = {
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 7+ messages in thread* [RFC PATCH 4/7] vfio/type1: Expose IO page table usage via fdinfo
[not found] <20261001224531.765278-1-praan@google.com>
` (2 preceding siblings ...)
2026-10-01 22:45 ` [RFC PATCH 3/7] iommupt: Enable per-domain IOPT attribution Pranjal Shrivastava
@ 2026-10-01 22:45 ` Pranjal Shrivastava
2026-10-01 22:45 ` [RFC PATCH 5/7] iommufd: " Pranjal Shrivastava
` (2 subsequent siblings)
6 siblings, 0 replies; 7+ messages in thread
From: Pranjal Shrivastava @ 2026-10-01 22:45 UTC (permalink / raw)
To: Joerg Roedel, Will Deacon, Robin Murphy, Jason Gunthorpe,
Kevin Tian, Alex Williamson, David Matlack, Jonathan Corbet,
Shuah Khan, Randy Dunlap
Cc: Mostafa Saleh, Daniel Mentz, Samiullah Khawaja, iommu, kvm,
linux-kselftest, linux-doc, Logan Odell, linux-fsdevel,
linux-kernel, Pranjal Shrivastava
From: Logan Odell <loganodell@google.com>
Collect the IO page table usage statistics from each IOMMU domain
attached to a legacy VFIO container and expose that data to userspace via
the new iommu-nr-pages field in fdinfo.
Introduce and implement the get_nr_pages callback for vfio-iommu-type1,
that iterates over all attached domains and aggregates their nr_pages
counters.
Only domains whose page table is implemented by the generic_pt library
account their page table memory. If any domain attached to the container
does not, get_nr_pages() returns -EOPNOTSUPP and the field is omitted
instead of reporting a partial count.
Document the new field in Documentation/filesystems/proc.rst.
Signed-off-by: Logan Odell <loganodell@google.com>
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
Documentation/filesystems/proc.rst | 17 +++++++++++++++++
drivers/vfio/container.c | 21 +++++++++++++++++++++
drivers/vfio/vfio.h | 2 ++
drivers/vfio/vfio_iommu_type1.c | 21 +++++++++++++++++++++
4 files changed, 61 insertions(+)
diff --git a/Documentation/filesystems/proc.rst b/Documentation/filesystems/proc.rst
index c102b62023cd..2d06c9930fcf 100644
--- a/Documentation/filesystems/proc.rst
+++ b/Documentation/filesystems/proc.rst
@@ -2200,6 +2200,23 @@ VFIO Device files
where 'vfio-device-syspath' is the sysfs path corresponding to the VFIO device
file.
+VFIO Container files
+~~~~~~~~~~~~~~~~~~~~
+
+::
+
+ pos: 0
+ flags: 02000002
+ mnt_id: 17
+ ino: 1103
+ iommu-nr-pages: 11
+
+where 'iommu-nr-pages' is the number of pages of memory used by the IO page
+tables of all IOMMU domains attached to the container. It is omitted when no
+IOMMU backend is set, or when any attached domain does not account its IO page
+table memory. Currently only domains using the generic_pt page table library
+account their IO page table memory.
+
3.9 /proc/<pid>/map_files - Information about memory mapped files
---------------------------------------------------------------------
This directory contains symbolic links which represent memory mapped files
diff --git a/drivers/vfio/container.c b/drivers/vfio/container.c
index 003281dbf8bc..2cb6c5373626 100644
--- a/drivers/vfio/container.c
+++ b/drivers/vfio/container.c
@@ -7,6 +7,7 @@
#include <linux/file.h>
#include <linux/slab.h>
#include <linux/fs.h>
+#include <linux/seq_file.h>
#include <linux/capability.h>
#include <linux/iommu.h>
#include <linux/miscdevice.h>
@@ -384,12 +385,32 @@ static int vfio_fops_release(struct inode *inode, struct file *filep)
return 0;
}
+#ifdef CONFIG_PROC_FS
+static void vfio_fops_show_fdinfo(struct seq_file *m, struct file *filep)
+{
+ struct vfio_container *container = filep->private_data;
+ long nr_pages = -EOPNOTSUPP;
+
+ down_read(&container->group_lock);
+ if (container->iommu_driver && container->iommu_driver->ops->get_nr_pages)
+ nr_pages = container->iommu_driver->ops->get_nr_pages(container->iommu_data);
+ up_read(&container->group_lock);
+
+ /* Only report a count that covers every domain in the container */
+ if (nr_pages >= 0)
+ seq_printf(m, "iommu-nr-pages:\t%ld\n", nr_pages);
+}
+#endif
+
static const struct file_operations vfio_fops = {
.owner = THIS_MODULE,
.open = vfio_fops_open,
.release = vfio_fops_release,
.unlocked_ioctl = vfio_fops_unl_ioctl,
.compat_ioctl = compat_ptr_ioctl,
+#ifdef CONFIG_PROC_FS
+ .show_fdinfo = vfio_fops_show_fdinfo,
+#endif
};
struct vfio_container *vfio_container_from_file(struct file *file)
diff --git a/drivers/vfio/vfio.h b/drivers/vfio/vfio.h
index 7728bc99b63d..37e8841ad0e9 100644
--- a/drivers/vfio/vfio.h
+++ b/drivers/vfio/vfio.h
@@ -226,6 +226,8 @@ struct vfio_iommu_driver_ops {
void *data, size_t count, bool write);
struct iommu_domain *(*group_iommu_domain)(void *iommu_data,
struct iommu_group *group);
+ /* Returns IO page table pages used, or -EOPNOTSUPP if not accounted */
+ long (*get_nr_pages)(void *iommu_data);
};
struct vfio_iommu_driver {
diff --git a/drivers/vfio/vfio_iommu_type1.c b/drivers/vfio/vfio_iommu_type1.c
index c8151ba54de3..fc7c1553d574 100644
--- a/drivers/vfio/vfio_iommu_type1.c
+++ b/drivers/vfio/vfio_iommu_type1.c
@@ -3249,6 +3249,26 @@ vfio_iommu_type1_group_iommu_domain(void *iommu_data,
return domain;
}
+static long vfio_iommu_type1_get_nr_pages(void *iommu_data)
+{
+ struct vfio_iommu *iommu = iommu_data;
+ struct vfio_domain *d;
+ long total = 0;
+
+ mutex_lock(&iommu->lock);
+ list_for_each_entry(d, &iommu->domain_list, next) {
+ /* Only generic_pt page tables account their pages */
+ if (!d->domain->is_iommupt) {
+ total = -EOPNOTSUPP;
+ break;
+ }
+ total += atomic_long_read(&d->domain->nr_pages);
+ }
+ mutex_unlock(&iommu->lock);
+
+ return total;
+}
+
static const struct vfio_iommu_driver_ops vfio_iommu_driver_ops_type1 = {
.name = "vfio-iommu-type1",
.owner = THIS_MODULE,
@@ -3263,6 +3283,7 @@ static const struct vfio_iommu_driver_ops vfio_iommu_driver_ops_type1 = {
.unregister_device = vfio_iommu_type1_unregister_device,
.dma_rw = vfio_iommu_type1_dma_rw,
.group_iommu_domain = vfio_iommu_type1_group_iommu_domain,
+ .get_nr_pages = vfio_iommu_type1_get_nr_pages,
};
static int __init vfio_iommu_type1_init(void)
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 7+ messages in thread* [RFC PATCH 5/7] iommufd: Expose IO page table usage via fdinfo
[not found] <20261001224531.765278-1-praan@google.com>
` (3 preceding siblings ...)
2026-10-01 22:45 ` [RFC PATCH 4/7] vfio/type1: Expose IO page table usage via fdinfo Pranjal Shrivastava
@ 2026-10-01 22:45 ` Pranjal Shrivastava
2026-10-01 22:45 ` [RFC PATCH 6/7] iommufd/selftest: Add observability test for iommu-nr-pages Pranjal Shrivastava
2026-10-01 22:45 ` [RFC PATCH 7/7] vfio/selftests: Add observability test for IO page table usage Pranjal Shrivastava
6 siblings, 0 replies; 7+ messages in thread
From: Pranjal Shrivastava @ 2026-10-01 22:45 UTC (permalink / raw)
To: Joerg Roedel, Will Deacon, Robin Murphy, Jason Gunthorpe,
Kevin Tian, Alex Williamson, David Matlack, Jonathan Corbet,
Shuah Khan, Randy Dunlap
Cc: Mostafa Saleh, Daniel Mentz, Samiullah Khawaja, iommu, kvm,
linux-kselftest, linux-doc, Logan Odell, Pranjal Shrivastava,
linux-fsdevel, linux-kernel
Collect the IO page table usage statistics from each paging HWPT object
within an iommufd context and expose that data to userspace via the new
iommu-nr-pages field in fdinfo.
Implement a show_fdinfo hook for iommufd_fops that iterates over the
context's object xarray, identifies paging HWPTs, and aggregates their
nr_pages counters. Nested HWPTs are skipped as their stage-1 page tables
are owned by userspace.
Only domains whose page table is implemented by the generic_pt library
account their page table memory. If the domain of any paging HWPT does
not, the field is omitted instead of reporting a partial count.
Document the new field in Documentation/filesystems/proc.rst.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
Documentation/filesystems/proc.rst | 17 +++++++++++++
drivers/iommu/iommufd/main.c | 38 ++++++++++++++++++++++++++++++
2 files changed, 55 insertions(+)
diff --git a/Documentation/filesystems/proc.rst b/Documentation/filesystems/proc.rst
index 2d06c9930fcf..90ba32e9e3b1 100644
--- a/Documentation/filesystems/proc.rst
+++ b/Documentation/filesystems/proc.rst
@@ -2217,6 +2217,23 @@ IOMMU backend is set, or when any attached domain does not account its IO page
table memory. Currently only domains using the generic_pt page table library
account their IO page table memory.
+IOMMUFD files
+~~~~~~~~~~~~~
+
+::
+
+ pos: 0
+ flags: 02000002
+ mnt_id: 17
+ ino: 1104
+ iommu-nr-pages: 11
+
+where 'iommu-nr-pages' is the number of pages of memory used by the IO page
+tables of all paging HW pagetables owned by the iommufd context. Nested HW
+pagetables are not included since their page tables are owned by userspace.
+The field is omitted when any paging HW pagetable does not account its IO page
+table memory.
+
3.9 /proc/<pid>/map_files - Information about memory mapped files
---------------------------------------------------------------------
This directory contains symbolic links which represent memory mapped files
diff --git a/drivers/iommu/iommufd/main.c b/drivers/iommu/iommufd/main.c
index 9a921b153162..51d0d100dc91 100644
--- a/drivers/iommu/iommufd/main.c
+++ b/drivers/iommu/iommufd/main.c
@@ -11,6 +11,7 @@
#include <linux/bug.h>
#include <linux/file.h>
#include <linux/fs.h>
+#include <linux/seq_file.h>
#include <linux/iommufd.h>
#include <linux/miscdevice.h>
#include <linux/module.h>
@@ -633,12 +634,49 @@ static int iommufd_fops_mmap(struct file *filp, struct vm_area_struct *vma)
return rc;
}
+#ifdef CONFIG_PROC_FS
+static void iommufd_fops_show_fdinfo(struct seq_file *m, struct file *filep)
+{
+ struct iommufd_ctx *ictx = filep->private_data;
+ struct iommufd_object *obj;
+ unsigned long nr_pages = 0;
+ bool supported = true;
+ unsigned long index;
+
+ xa_lock(&ictx->objects);
+ xa_for_each(&ictx->objects, index, obj) {
+ struct iommufd_hw_pagetable *hwpt;
+
+ /* Nested stage-1 page tables are owned by the guest */
+ if (obj->type != IOMMUFD_OBJ_HWPT_PAGING)
+ continue;
+
+ hwpt = container_of(obj, struct iommufd_hw_pagetable, obj);
+ if (!hwpt->domain)
+ continue;
+ /* Only generic_pt page tables account their pages */
+ if (!hwpt->domain->is_iommupt) {
+ supported = false;
+ break;
+ }
+ nr_pages += atomic_long_read(&hwpt->domain->nr_pages);
+ }
+ xa_unlock(&ictx->objects);
+
+ if (supported)
+ seq_printf(m, "iommu-nr-pages:\t%lu\n", nr_pages);
+}
+#endif
+
static const struct file_operations iommufd_fops = {
.owner = THIS_MODULE,
.open = iommufd_fops_open,
.release = iommufd_fops_release,
.unlocked_ioctl = iommufd_fops_ioctl,
.mmap = iommufd_fops_mmap,
+#ifdef CONFIG_PROC_FS
+ .show_fdinfo = iommufd_fops_show_fdinfo,
+#endif
};
/**
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 7+ messages in thread* [RFC PATCH 6/7] iommufd/selftest: Add observability test for iommu-nr-pages
[not found] <20261001224531.765278-1-praan@google.com>
` (4 preceding siblings ...)
2026-10-01 22:45 ` [RFC PATCH 5/7] iommufd: " Pranjal Shrivastava
@ 2026-10-01 22:45 ` Pranjal Shrivastava
2026-10-01 22:45 ` [RFC PATCH 7/7] vfio/selftests: Add observability test for IO page table usage Pranjal Shrivastava
6 siblings, 0 replies; 7+ messages in thread
From: Pranjal Shrivastava @ 2026-10-01 22:45 UTC (permalink / raw)
To: Joerg Roedel, Will Deacon, Robin Murphy, Jason Gunthorpe,
Kevin Tian, Alex Williamson, David Matlack, Jonathan Corbet,
Shuah Khan, Randy Dunlap
Cc: Mostafa Saleh, Daniel Mentz, Samiullah Khawaja, iommu, kvm,
linux-kselftest, linux-doc, Logan Odell, Pranjal Shrivastava,
linux-fsdevel, linux-kernel
Add a test case to validate the iommu-nr-pages fdinfo metric. The test
maps sparse IOVAs to force page table directory allocations, verifies
that the metric increments accordingly, and asserts that the count
remains stranded after unmapping, verifying the current behaviour
required for future automated reclamation logic.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
tools/testing/selftests/iommu/iommufd.c | 48 +++++++++++++++++++++++++
1 file changed, 48 insertions(+)
diff --git a/tools/testing/selftests/iommu/iommufd.c b/tools/testing/selftests/iommu/iommufd.c
index d44b34b05757..f5f4ef01ea4f 100644
--- a/tools/testing/selftests/iommu/iommufd.c
+++ b/tools/testing/selftests/iommu/iommufd.c
@@ -3613,4 +3613,52 @@ TEST_F(iommufd_device_pasid, pasid_attach)
test_cmd_mock_domain_replace(self->stdev_id, self->ioas_id);
}
+TEST_F(iommufd_ioas, iommufd_nr_pages)
+{
+ char path[256];
+ char line[128];
+ FILE *f;
+ long nr_pages_initial = -1;
+ long nr_pages_mapped = -1;
+ long nr_pages_final = -1;
+
+ if (!self->stdev_id)
+ SKIP(return, "Needs a device attached for HWPT");
+
+ snprintf(path, sizeof(path), "/proc/self/fdinfo/%d", self->fd);
+ f = fopen(path, "r");
+ ASSERT_NE(NULL, f);
+ while (fgets(line, sizeof(line), f)) {
+ if (sscanf(line, "iommu-nr-pages:\t%ld", &nr_pages_initial) == 1)
+ break;
+ }
+ fclose(f);
+ ASSERT_GE(nr_pages_initial, 0);
+
+ /* Map a page at start and one far away */
+ test_ioctl_ioas_map_fixed(buffer, PAGE_SIZE, self->base_iova);
+ test_ioctl_ioas_map_fixed(buffer, PAGE_SIZE, self->base_iova + (1UL << 30));
+
+ f = fopen(path, "r");
+ ASSERT_NE(NULL, f);
+ while (fgets(line, sizeof(line), f)) {
+ if (sscanf(line, "iommu-nr-pages:\t%ld", &nr_pages_mapped) == 1)
+ break;
+ }
+ fclose(f);
+ ASSERT_GT(nr_pages_mapped, nr_pages_initial);
+
+ test_ioctl_ioas_unmap(self->base_iova, PAGE_SIZE);
+ test_ioctl_ioas_unmap(self->base_iova + (1UL << 30), PAGE_SIZE);
+
+ f = fopen(path, "r");
+ ASSERT_NE(NULL, f);
+ while (fgets(line, sizeof(line), f)) {
+ if (sscanf(line, "iommu-nr-pages:\t%ld", &nr_pages_final) == 1)
+ break;
+ }
+ fclose(f);
+ ASSERT_EQ(nr_pages_final, nr_pages_mapped);
+}
+
TEST_HARNESS_MAIN
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 7+ messages in thread* [RFC PATCH 7/7] vfio/selftests: Add observability test for IO page table usage
[not found] <20261001224531.765278-1-praan@google.com>
` (5 preceding siblings ...)
2026-10-01 22:45 ` [RFC PATCH 6/7] iommufd/selftest: Add observability test for iommu-nr-pages Pranjal Shrivastava
@ 2026-10-01 22:45 ` Pranjal Shrivastava
6 siblings, 0 replies; 7+ messages in thread
From: Pranjal Shrivastava @ 2026-10-01 22:45 UTC (permalink / raw)
To: Joerg Roedel, Will Deacon, Robin Murphy, Jason Gunthorpe,
Kevin Tian, Alex Williamson, David Matlack, Jonathan Corbet,
Shuah Khan, Randy Dunlap
Cc: Mostafa Saleh, Daniel Mentz, Samiullah Khawaja, iommu, kvm,
linux-kselftest, linux-doc, Logan Odell, Pranjal Shrivastava,
linux-fsdevel, linux-kernel
Add a new selftest, vfio_nr_pages_test, to validate the iommu-nr-pages
fdinfo metric across all IOMMU backends (Type 1 and IOMMUFD native/compat).
The test performs a sparse mapping to force the underlying IOMMU driver
to allocate a deep page table hierarchy. It then asserts that the
nr_pages metric increases and proves the 'sticky' reclamation behavior
by verifying the page count does not drop immediately upon unmapping.
The test is skipped when the IOMMU driver does not account its page
table memory, in which case the metric is not reported.
Co-developed-by: Logan Odell <loganodell@google.com>
Signed-off-by: Logan Odell <loganodell@google.com>
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
tools/testing/selftests/vfio/Makefile | 1 +
.../selftests/vfio/vfio_nr_pages_test.c | 119 ++++++++++++++++++
2 files changed, 120 insertions(+)
create mode 100644 tools/testing/selftests/vfio/vfio_nr_pages_test.c
diff --git a/tools/testing/selftests/vfio/Makefile b/tools/testing/selftests/vfio/Makefile
index 2c32c48db509..140ef273e36b 100644
--- a/tools/testing/selftests/vfio/Makefile
+++ b/tools/testing/selftests/vfio/Makefile
@@ -13,6 +13,7 @@ TEST_GEN_PROGS += vfio_pci_device_test
TEST_GEN_PROGS += vfio_pci_device_init_perf_test
TEST_GEN_PROGS += vfio_pci_driver_test
TEST_GEN_PROGS += vfio_pci_sriov_uapi_test
+TEST_GEN_PROGS += vfio_nr_pages_test
TEST_FILES += scripts/cleanup.sh
TEST_FILES += scripts/lib.sh
diff --git a/tools/testing/selftests/vfio/vfio_nr_pages_test.c b/tools/testing/selftests/vfio/vfio_nr_pages_test.c
new file mode 100644
index 000000000000..aef6b1dd8e19
--- /dev/null
+++ b/tools/testing/selftests/vfio/vfio_nr_pages_test.c
@@ -0,0 +1,119 @@
+// SPDX-License-Identifier: GPL-2.0-only
+#include <stdio.h>
+#include <sys/mman.h>
+#include <unistd.h>
+#include <linux/sizes.h>
+
+#include <libvfio.h>
+
+#include "../kselftest_harness.h"
+
+static const char *device_bdf;
+
+static long read_nr_pages(struct iommu *iommu)
+{
+ char path[256];
+ char line[128];
+ FILE *f;
+ long val = -1;
+ int fd = iommu->container_fd != -1 ? iommu->container_fd : iommu->iommufd;
+
+ snprintf(path, sizeof(path), "/proc/self/fdinfo/%d", fd);
+ f = fopen(path, "r");
+ if (!f)
+ return -1;
+
+ while (fgets(line, sizeof(line), f)) {
+ if (sscanf(line, "iommu-nr-pages: %ld", &val) == 1)
+ break;
+ }
+
+ fclose(f);
+ return val;
+}
+
+FIXTURE(vfio_nr_pages_test) {
+ struct iommu *iommu;
+ struct vfio_pci_device *device;
+};
+
+FIXTURE_VARIANT(vfio_nr_pages_test) {
+ const char *iommu_mode;
+};
+
+#define FIXTURE_VARIANT_ADD_IOMMU_MODE(_name) \
+FIXTURE_VARIANT_ADD(vfio_nr_pages_test, _name) { \
+ .iommu_mode = #_name, \
+}
+
+FIXTURE_VARIANT_ADD_ALL_IOMMU_MODES();
+
+FIXTURE_SETUP(vfio_nr_pages_test)
+{
+ self->iommu = iommu_init(variant->iommu_mode);
+ if (!self->iommu)
+ SKIP(return, "IOMMU mode %s not supported", variant->iommu_mode);
+
+ self->device = vfio_pci_device_init(device_bdf, self->iommu);
+ if (!self->device) {
+ iommu_cleanup(self->iommu);
+ SKIP(return, "Failed to initialize VFIO device");
+ }
+}
+
+FIXTURE_TEARDOWN(vfio_nr_pages_test)
+{
+ if (self->device)
+ vfio_pci_device_cleanup(self->device);
+ if (self->iommu)
+ iommu_cleanup(self->iommu);
+}
+
+TEST_F(vfio_nr_pages_test, sparse_mapping_sticky_reclamation)
+{
+ long nr_pages_initial, nr_pages_mapped, nr_pages_final;
+ struct dma_region region1 = {0};
+ struct dma_region region2 = {0};
+
+ if (!self->iommu || !self->device)
+ SKIP(return, "Fixture setup failed");
+
+ nr_pages_initial = read_nr_pages(self->iommu);
+ if (nr_pages_initial < 0)
+ SKIP(return, "iommu-nr-pages not reported by this IOMMU");
+
+ /* Map a page at start */
+ region1.size = SZ_4K;
+ region1.vaddr = mmap_reserve(region1.size, region1.size, 0);
+ ASSERT_NE(MAP_FAILED, region1.vaddr);
+ region1.iova = 0x100000;
+ iommu_map(self->iommu, ®ion1);
+
+ /* Map another page far away (1GB stride) to force deep page tables */
+ region2.size = SZ_4K;
+ region2.vaddr = mmap_reserve(region2.size, region2.size, 0);
+ ASSERT_NE(MAP_FAILED, region2.vaddr);
+ region2.iova = 0x100000 + SZ_1G;
+ iommu_map(self->iommu, ®ion2);
+
+ nr_pages_mapped = read_nr_pages(self->iommu);
+ ASSERT_GT(nr_pages_mapped, nr_pages_initial);
+
+ /* Unmap both regions */
+ iommu_unmap(self->iommu, ®ion1);
+ iommu_unmap(self->iommu, ®ion2);
+
+ nr_pages_final = read_nr_pages(self->iommu);
+
+ /* Page tables should remain allocated */
+ ASSERT_EQ(nr_pages_final, nr_pages_mapped);
+
+ munmap(region1.vaddr, region1.size);
+ munmap(region2.vaddr, region2.size);
+}
+
+int main(int argc, char *argv[])
+{
+ device_bdf = vfio_selftests_get_bdf(&argc, argv);
+ return test_harness_run(argc, argv);
+}
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 7+ messages in thread