* [RFC PATCH 0/5] iommupt: Introduce IO page table shrinker
@ 2026-10-01 23:02 Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 1/5] iommu: Add reclaim_list xarray to struct iommu_domain Pranjal Shrivastava
` (4 more replies)
0 siblings, 5 replies; 6+ messages in thread
From: Pranjal Shrivastava @ 2026-10-01 23:02 UTC (permalink / raw)
To: Joerg Roedel, Will Deacon, Robin Murphy, Jason Gunthorpe, Kevin Tian
Cc: Mostafa Saleh, Daniel Mentz, Samiullah Khawaja, Logan Odell,
iommu, linux-mm, linux-kernel, Pranjal Shrivastava
Introduce a lockless, deferred reclamation framework for IOMMU page tables
built on the generic_pt library. As VMMs and userspace drivers map and
unmap large, sparse IOVA regions through VFIO and iommufd, page table
directories are often left allocated but completely empty. generic_pt
frees a table when a single unmap covers it entirely, but tables that
empty through a series of partial unmaps stay allocated until the domain
is destroyed. Under memory pressure, this *stranded* memory cannot be
reclaimed and has resulted in OOMs.
This series refcounts leaf directories natively in struct ioptdesc and
registers a domain-aware MM shrinker that prunes empty directories under
system memory pressure.
This is an RFC to align on the design. It has known issues, listed below,
and is not intended for merging as-is.
Design
======
- Refcounting: the __page_refcount of a leaf directory's struct ioptdesc
counts its valid leaf entries, plus one for the table itself. Map takes
a reference with atomic_inc_not_zero() before installing a leaf, and
unmap drops it after clearing one. Leaf entries are counted only in the
table that holds them, so map and unmap never update refcounts on the
shared tables above it.
- Queueing: when the refcount drops to 1, the directory is empty and is
logged with its IOVA into a per-domain xarray (reclaim_list). Only
non-DMA domains are queued.
- Claim: under memory pressure, the shrinker walks the queued directories
and claims each with atomic_cmpxchg(refcount, 1, 0). A map that finds
the refcount at 0 fails atomic_inc_not_zero() and retries, so it never
reuses a claimed directory.
- Sever: a format-agnostic top-down walk (sever_branch) clears the parent
slot with cmpxchg, severing the dead branch from the live tree.
- Free: the domain's IOTLB is flushed so the IOMMU drops any cached
pointers to the severed tables. A single synchronize_srcu() then waits
for in-flight map walkers before the pages are freed.
The cost is one atomic operation per leaf entry on map and unmap, plus an
xarray insert when a directory becomes empty.
Known issues
============
These are understood and will be addressed in the next version:
- Map vs. shrinker retry: on a claimed directory, map returns -EAGAIN,
which __map_range() already uses to re-descend into the cached table
pointer, i.e. the dead directory. It needs a distinct error and a retry
from the top.
- Unmap vs. shrinker hand-off: an unmap that fully covers a queued
directory, or a large-page map over an empty range, frees it directly
while the xarray still references it.
- xarray keys: the key is the IOVA where the emptying unmap started, so a
directory can be queued under several keys and nr_reclaimable
over-counts. Keying by the table's PFN makes entries unique.
- A failed sever leaves the refcount at 0 instead of restoring it to 1.
The top-level table, which has no parent, should never be queued.
- iommufd allocates paging domains without iommu_domain_init(), so their
reclaim_list is not initialised.
- Combined with the observability series, pages must be uncharged from
their domain at sever time, as they are freed after the domain may be
gone.
Open questions
==============
- Leaf directories with child tables: a leaf directory above the last
level (e.g. one holding a 2M leaf) can gain a child table, which is not
counted, and could then be reclaimed while the child is live. Should
child tables be counted at install, or counting be restricted to
last-level directories?
- SRCU in reclaim context: the shrinker can run from direct reclaim
inside a map's SRCU read section, where synchronize_srcu() would wait on
itself. Should the free be deferred (work item / call_srcu), or is a
different scheme preferred?
- IOTLB flush granularity: flush_iotlb_all() is used after severing.
Would a range-based paging-structure flush be preferred?
- DMA-API domains are excluded since their unmap can run in IRQ context.
Should the per-leaf atomics also be skipped for them?
Upcoming Work / Roadmap
=======================
An IO page table observability series, which attributes page table memory
to its domain and exposes it via fdinfo, is posted separately as an RFC.
The two series are independent; this one applies on v7.3-rc5 by itself.
There's an alignment session at Linux Plumbers Conference 2026 for these [1].
[1] https://lpc.events/event/20/contributions/2624/
Pranjal Shrivastava (5):
iommu: Add reclaim_list xarray to struct iommu_domain
iommupt: Implement refcounting logic for Leaf entries
iommupt: Add lockless sever_branch helper
iommupt: Introduce lockless page table shrinker
iommupt: Return real page count to the shrinker core
drivers/iommu/Makefile | 2 +-
drivers/iommu/generic_pt/iommu_pt.h | 81 ++++++++++++++++-
drivers/iommu/generic_pt/shrinker.c | 135 ++++++++++++++++++++++++++++
drivers/iommu/iommu.c | 2 +
include/linux/generic_pt/iommu.h | 29 ++++++
include/linux/iommu.h | 3 +
6 files changed, 249 insertions(+), 3 deletions(-)
create mode 100644 drivers/iommu/generic_pt/shrinker.c
base-commit: 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 6+ messages in thread
* [RFC PATCH 1/5] iommu: Add reclaim_list xarray to struct iommu_domain
2026-10-01 23:02 [RFC PATCH 0/5] iommupt: Introduce IO page table shrinker Pranjal Shrivastava
@ 2026-10-01 23:02 ` Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 2/5] iommupt: Implement refcounting logic for Leaf entries Pranjal Shrivastava
` (3 subsequent siblings)
4 siblings, 0 replies; 6+ messages in thread
From: Pranjal Shrivastava @ 2026-10-01 23:02 UTC (permalink / raw)
To: Joerg Roedel, Will Deacon, Robin Murphy, Jason Gunthorpe, Kevin Tian
Cc: Mostafa Saleh, Daniel Mentz, Samiullah Khawaja, Logan Odell,
iommu, linux-mm, linux-kernel, Pranjal Shrivastava
Introduce struct xarray reclaim_list to struct iommu_domain. This will
serve as the O(1) storage where the fast path stores empty directories by
their associated IOVA. IOVA is required by the shrinker to perform a
top-down tree walk to safely locate and prune the parent pointer. While
xa_store requires a localized spinlock, this path is only triggered when
a directory reaches a completely empty state, keeping the vast majority
of mapping operations lock-free.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
drivers/iommu/iommu.c | 2 ++
include/linux/iommu.h | 3 +++
2 files changed, 5 insertions(+)
diff --git a/drivers/iommu/iommu.c b/drivers/iommu/iommu.c
index cd1bca7ede9a..5d814854d825 100644
--- a/drivers/iommu/iommu.c
+++ b/drivers/iommu/iommu.c
@@ -2074,6 +2074,7 @@ static void iommu_domain_init(struct iommu_domain *domain, unsigned int type,
domain->owner = ops;
if (!domain->ops)
domain->ops = ops->default_domain_ops;
+ xa_init(&domain->reclaim_list);
}
static struct iommu_domain *
@@ -2126,6 +2127,7 @@ EXPORT_SYMBOL_GPL(iommu_paging_domain_alloc_flags);
void iommu_domain_free(struct iommu_domain *domain)
{
+ xa_destroy(&domain->reclaim_list);
switch (domain->cookie_type) {
case IOMMU_COOKIE_DMA_IOVA:
iommu_put_dma_cookie(domain);
diff --git a/include/linux/iommu.h b/include/linux/iommu.h
index ac43b8b93f14..b9611631ea80 100644
--- a/include/linux/iommu.h
+++ b/include/linux/iommu.h
@@ -14,6 +14,7 @@
#include <linux/err.h>
#include <linux/of.h>
#include <linux/iova_bitmap.h>
+#include <linux/xarray.h>
#include <uapi/linux/iommufd.h>
#define IOMMU_READ (1 << 0)
@@ -249,6 +250,8 @@ struct iommu_domain {
struct list_head next;
};
};
+
+ struct xarray reclaim_list;
};
static inline bool iommu_is_dma_domain(struct iommu_domain *domain)
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 6+ messages in thread
* [RFC PATCH 2/5] iommupt: Implement refcounting logic for Leaf entries
2026-10-01 23:02 [RFC PATCH 0/5] iommupt: Introduce IO page table shrinker Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 1/5] iommu: Add reclaim_list xarray to struct iommu_domain Pranjal Shrivastava
@ 2026-10-01 23:02 ` Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 3/5] iommupt: Add lockless sever_branch helper Pranjal Shrivastava
` (2 subsequent siblings)
4 siblings, 0 replies; 6+ messages in thread
From: Pranjal Shrivastava @ 2026-10-01 23:02 UTC (permalink / raw)
To: Joerg Roedel, Will Deacon, Robin Murphy, Jason Gunthorpe, Kevin Tian
Cc: Mostafa Saleh, Daniel Mentz, Samiullah Khawaja, Logan Odell,
iommu, linux-mm, linux-kernel, Pranjal Shrivastava
Implement a lockless state machine using the ioptdesc's __page_refcount
field. During an unmap, the driver executes atomic_sub_return() when
clearing a leaf PTE. If the refcount drops to the baseline state (1),
the directory is completely empty and its pointer is stored to the
domain's reclaim_list Xarray.
To protect against a concurrent background Shrinker pruning a directory
while a map thread is actively populating it, the map thread must acquire
a reference using atomic_inc_not_zero() before installing a new leaf.
If the acquisition fails, the page has been marked dead (0) by the
Shrinker. The map thread cleanly returns -EAGAIN, forcing the walker to
loop and re-allocate a fresh directory.
Only non-DMA domains are queued: DMA API unmaps can run in IRQ context
and the reclaim_list lock is not IRQ safe.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
drivers/iommu/generic_pt/iommu_pt.h | 26 ++++++++++++++++++++++++++
1 file changed, 26 insertions(+)
diff --git a/drivers/iommu/generic_pt/iommu_pt.h b/drivers/iommu/generic_pt/iommu_pt.h
index 07ec2b3ab986..c3d09b746eec 100644
--- a/drivers/iommu/generic_pt/iommu_pt.h
+++ b/drivers/iommu/generic_pt/iommu_pt.h
@@ -621,6 +621,13 @@ static int __map_range_leaf(struct pt_range *range, void *arg,
PT_WARN_ON(compute_best_pgsize(&pts, oa) !=
leaf_pgsize_lg2);
}
+
+ /* If refcount = 0, shrinker killed the page, retry map */
+ if (!atomic_inc_not_zero(&virt_to_ioptdesc(pts.table)->__page_refcount)) {
+ ret = -EAGAIN;
+ break;
+ }
+
pt_install_leaf_entry(&pts, oa, leaf_pgsize_lg2, &map->attrs);
oa += log2_to_int(leaf_pgsize_lg2);
@@ -754,6 +761,11 @@ static __always_inline int __do_map_single_page(struct pt_range *range,
if (pts.level == 0) {
if (pts.type != PT_ENTRY_EMPTY)
return -EADDRINUSE;
+
+ /* If refcount = 0, shrinker killed the page, retry map */
+ if (!atomic_inc_not_zero(&virt_to_ioptdesc(pts.table)->__page_refcount))
+ return -EAGAIN;
+
pt_install_leaf_entry(&pts, map->oa, PAGE_SHIFT,
&map->attrs);
/* No flush, not used when incoherent */
@@ -1087,6 +1099,20 @@ static __maybe_unused int __unmap_range(struct pt_range *range, void *arg,
*/
num_contig_lg2 = pt_entry_num_contig_lg2(&pts);
pt_clear_entries(&pts, num_contig_lg2);
+
+ /* Drop refcount and add to reclaim_list for the last ref */
+ /*
+ * DMA API unmaps can run in IRQ context and the reclaim_list
+ * lock is not IRQ safe, so only queue for non-DMA domains.
+ */
+ if (atomic_dec_return(&virt_to_ioptdesc(pts.table)->__page_refcount) == 1 &&
+ !iommu_is_dma_domain(&iommu_from_common(range->common)->domain)) {
+ struct pt_iommu *iommu = iommu_from_common(range->common);
+
+ xa_store(&iommu->domain.reclaim_list, range->va,
+ virt_to_ioptdesc(pts.table), GFP_ATOMIC);
+ }
+
gather_add_leaf(&unmap->pending, &pts);
num_oas += log2_to_int(num_contig_lg2);
if (pts.index < flush_start_index)
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 6+ messages in thread
* [RFC PATCH 3/5] iommupt: Add lockless sever_branch helper
2026-10-01 23:02 [RFC PATCH 0/5] iommupt: Introduce IO page table shrinker Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 1/5] iommu: Add reclaim_list xarray to struct iommu_domain Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 2/5] iommupt: Implement refcounting logic for Leaf entries Pranjal Shrivastava
@ 2026-10-01 23:02 ` Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 4/5] iommupt: Introduce lockless page table shrinker Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 5/5] iommupt: Return real page count to the shrinker core Pranjal Shrivastava
4 siblings, 0 replies; 6+ messages in thread
From: Pranjal Shrivastava @ 2026-10-01 23:02 UTC (permalink / raw)
To: Joerg Roedel, Will Deacon, Robin Murphy, Jason Gunthorpe, Kevin Tian
Cc: Mostafa Saleh, Daniel Mentz, Samiullah Khawaja, Logan Odell,
iommu, linux-mm, linux-kernel, Pranjal Shrivastava
Introduce a new sever_branch pt_iommu_op. This helper performs a
top-down walk to locate the target parent slot and safely executes a
lockless cmpxchg to 0x0. Because generic_pt abstracts all hardware PTEs
as u64, this safely utilizes the existing pt_table_install64 to cleanly
sever the branch regardless of the underlying IOMMU format.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
drivers/iommu/generic_pt/iommu_pt.h | 40 +++++++++++++++++++++++++++++
include/linux/generic_pt/iommu.h | 14 ++++++++++
2 files changed, 54 insertions(+)
diff --git a/drivers/iommu/generic_pt/iommu_pt.h b/drivers/iommu/generic_pt/iommu_pt.h
index c3d09b746eec..868242eb2d55 100644
--- a/drivers/iommu/generic_pt/iommu_pt.h
+++ b/drivers/iommu/generic_pt/iommu_pt.h
@@ -1155,6 +1155,45 @@ static size_t NS(unmap_range)(struct pt_iommu *iommu_table, dma_addr_t iova,
return unmap.unmapped;
}
+struct pt_sever_args {
+ phys_addr_t expected_phys;
+ bool success;
+};
+
+static int __sever_branch(struct pt_range *range, void *arg,
+ unsigned int level, struct pt_table_p *table)
+{
+ struct pt_state pts = pt_init(range, level, table);
+ struct pt_sever_args *sever = arg;
+
+ switch (pt_load_single_entry(&pts)) {
+ case PT_ENTRY_TABLE:
+ if (virt_to_phys(pt_table_ptr(&pts)) == sever->expected_phys) {
+ sever->success = pt_table_install64(&pts, 0x0);
+ return 1; /* Stop walking */
+ }
+ return pt_descend(&pts, arg, __sever_branch);
+ default:
+ break;
+ }
+ return 0;
+}
+
+static bool NS(sever_branch)(struct pt_iommu *iommu_table, dma_addr_t iova,
+ phys_addr_t expected_phys)
+{
+ struct pt_range range;
+ struct pt_sever_args sever = { .expected_phys = expected_phys, .success = false };
+ int ret;
+
+ ret = make_range(common_from_iommu(iommu_table), &range, iova, 1);
+ if (ret)
+ return false;
+
+ pt_walk_range(&range, __sever_branch, &sever);
+ return sever.success;
+}
+
static void NS(get_info)(struct pt_iommu *iommu_table,
struct pt_iommu_info *info)
{
@@ -1208,6 +1247,7 @@ static void NS(deinit)(struct pt_iommu *iommu_table)
static const struct pt_iommu_ops NS(ops) = {
.map_range = NS(map_range),
.unmap_range = NS(unmap_range),
+ .sever_branch = NS(sever_branch),
#if IS_ENABLED(CONFIG_IOMMUFD_DRIVER) && defined(pt_entry_is_write_dirty) && \
IS_ENABLED(CONFIG_IOMMUFD_TEST) && defined(pt_entry_make_write_dirty)
.set_dirty = NS(set_dirty),
diff --git a/include/linux/generic_pt/iommu.h b/include/linux/generic_pt/iommu.h
index dd0edd02a48a..d1ade1767a9a 100644
--- a/include/linux/generic_pt/iommu.h
+++ b/include/linux/generic_pt/iommu.h
@@ -137,6 +137,20 @@ struct pt_iommu_ops {
dma_addr_t len,
struct iommu_iotlb_gather *iotlb_gather);
+ /**
+ * @sever_branch: Locklessly sever an empty page table branch
+ * @iommu_table: Table to manipulate
+ * @iova: IO virtual address associated with the branch
+ * @expected_phys: Expected physical address of the child directory
+ *
+ * Context: Executed locklessly by a background Shrinker.
+ * Uses cmpxchg to write 0x0 to the parent slot.
+ *
+ * Returns: true if successfully severed, false if aborted.
+ */
+ bool (*sever_branch)(struct pt_iommu *iommu_table, dma_addr_t iova,
+ phys_addr_t expected_phys);
+
/**
* @set_dirty: Make the iova write dirty
* @iommu_table: Table to manipulate
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 6+ messages in thread
* [RFC PATCH 4/5] iommupt: Introduce lockless page table shrinker
2026-10-01 23:02 [RFC PATCH 0/5] iommupt: Introduce IO page table shrinker Pranjal Shrivastava
` (2 preceding siblings ...)
2026-10-01 23:02 ` [RFC PATCH 3/5] iommupt: Add lockless sever_branch helper Pranjal Shrivastava
@ 2026-10-01 23:02 ` Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 5/5] iommupt: Return real page count to the shrinker core Pranjal Shrivastava
4 siblings, 0 replies; 6+ messages in thread
From: Pranjal Shrivastava @ 2026-10-01 23:02 UTC (permalink / raw)
To: Joerg Roedel, Will Deacon, Robin Murphy, Jason Gunthorpe, Kevin Tian
Cc: Mostafa Saleh, Daniel Mentz, Samiullah Khawaja, Logan Odell,
iommu, linux-mm, linux-kernel, Pranjal Shrivastava
Under system memory pressure, intermediate IOMMU page table directories
that were left stranded by sparse unmaps consume valuable RAM. Introduce
an IO Page Table Shrinker to reclaim such memory.
The shrinker locklessly harvests empty directories from the per-domain
Xarrays, severs them from the page table tree using the format-agnostic
sever_branch helper. It employs a single, global synchronize_srcu grace
period before freeing the memory to protect against concurrent maps.
SRCU is an appropriate choice because the allocation during map may sleep
After severing, flush the domain's IOTLB so the IOMMU drops any cached
pointers to the severed tables before they are freed.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
drivers/iommu/Makefile | 2 +-
drivers/iommu/generic_pt/iommu_pt.h | 16 +++-
drivers/iommu/generic_pt/shrinker.c | 136 ++++++++++++++++++++++++++++
include/linux/generic_pt/iommu.h | 10 ++
4 files changed, 160 insertions(+), 4 deletions(-)
create mode 100644 drivers/iommu/generic_pt/shrinker.c
diff --git a/drivers/iommu/Makefile b/drivers/iommu/Makefile
index 2f05725eaab1..ad30c4b7566e 100644
--- a/drivers/iommu/Makefile
+++ b/drivers/iommu/Makefile
@@ -3,7 +3,7 @@ obj-y += arm/ iommufd/
obj-$(CONFIG_AMD_IOMMU) += amd/
obj-$(CONFIG_INTEL_IOMMU) += intel/
obj-$(CONFIG_RISCV_IOMMU) += riscv/
-obj-$(CONFIG_GENERIC_PT) += generic_pt/fmt/
+obj-$(CONFIG_GENERIC_PT) += generic_pt/fmt/ generic_pt/shrinker.o
obj-$(CONFIG_HYPERV) += hyperv/
obj-$(CONFIG_IOMMU_API) += iommu.o
obj-$(CONFIG_IOMMU_SUPPORT) += iommu-pages.o
diff --git a/drivers/iommu/generic_pt/iommu_pt.h b/drivers/iommu/generic_pt/iommu_pt.h
index 868242eb2d55..e5ee419f2f0b 100644
--- a/drivers/iommu/generic_pt/iommu_pt.h
+++ b/drivers/iommu/generic_pt/iommu_pt.h
@@ -916,7 +916,7 @@ static int check_map_range(struct pt_iommu *iommu_table, struct pt_range *range,
static int do_map(struct pt_range *range, struct pt_common *common,
bool single_page, struct pt_iommu_map_args *map)
{
- int ret;
+ int ret, idx;
/*
* The __map_single_page() fast path does not support DMA_INCOHERENT
@@ -924,17 +924,21 @@ static int do_map(struct pt_range *range, struct pt_common *common,
*/
if (single_page && !pt_feature(common, PT_FEAT_DMA_INCOHERENT)) {
+ idx = srcu_read_lock(&generic_pt_srcu);
ret = pt_walk_range(range, __map_single_page, map);
+ srcu_read_unlock(&generic_pt_srcu, idx);
if (ret != -EAGAIN)
return ret;
/* EAGAIN falls through to the full path */
}
do {
+ idx = srcu_read_lock(&generic_pt_srcu);
if (map->leaf_level == range->top_level)
ret = pt_walk_range(range, __map_range_leaf, map);
else
ret = pt_walk_range(range, __map_range, map);
+ srcu_read_unlock(&generic_pt_srcu, idx);
} while (ret == -EAGAIN);
return ret;
}
@@ -1142,13 +1146,15 @@ static size_t NS(unmap_range)(struct pt_iommu *iommu_table, dma_addr_t iova,
unmap.pending.free_list),
};
struct pt_range range;
- int ret;
+ int ret, idx;
ret = make_range(common_from_iommu(iommu_table), &range, iova, len);
if (ret)
return 0;
+ idx = srcu_read_lock(&generic_pt_srcu);
pt_walk_range(&range, __unmap_range, &unmap);
+ srcu_read_unlock(&generic_pt_srcu, idx);
gather_range_pending(&unmap.pending, iommu_table, iova, unmap.unmapped);
@@ -1170,7 +1176,8 @@ static int __sever_branch(struct pt_range *range, void *arg,
case PT_ENTRY_TABLE:
if (virt_to_phys(pt_table_ptr(&pts)) == sever->expected_phys) {
sever->success = pt_table_install64(&pts, 0x0);
- return 1; /* Stop walking */
+ /* Stop walking */
+ return 1;
}
return pt_descend(&pts, arg, __sever_branch);
default:
@@ -1231,6 +1238,8 @@ static void NS(deinit)(struct pt_iommu *iommu_table)
collect.pending.free_list),
};
+ generic_pt_shrinker_remove(iommu_table);
+
iommu_pages_list_add(&collect.pending.free_list, range.top_table);
pt_walk_range(&range, __collect_tables, &collect);
@@ -1407,6 +1416,7 @@ int pt_iommu_init(struct pt_iommu_table *fmt_table,
/* Must be last, see pt_iommu_deinit() */
iommu_table->ops = &NS(ops);
+ generic_pt_shrinker_add(iommu_table);
return 0;
}
EXPORT_SYMBOL_NS_GPL(pt_iommu_init, "GENERIC_PT_IOMMU");
diff --git a/drivers/iommu/generic_pt/shrinker.c b/drivers/iommu/generic_pt/shrinker.c
new file mode 100644
index 000000000000..66169fb5db71
--- /dev/null
+++ b/drivers/iommu/generic_pt/shrinker.c
@@ -0,0 +1,136 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * Copyright (c) 2026, Google LLC.
+ * Author: Pranjal Shrivastava <praan@google.com>
+ * IO Page table reclamation (Shrinker) for generic_pt.
+ */
+
+#include <linux/iommu.h>
+#include <linux/shrinker.h>
+#include <linux/srcu.h>
+#include <linux/list.h>
+#include <linux/mutex.h>
+#include "../iommu-pages.h"
+#include <linux/generic_pt/iommu.h>
+
+DEFINE_SRCU(generic_pt_srcu);
+EXPORT_SYMBOL_GPL(generic_pt_srcu);
+
+static LIST_HEAD(generic_pt_domains_list);
+static DEFINE_MUTEX(generic_pt_domains_list_lock);
+
+void generic_pt_shrinker_add(struct pt_iommu *iommu)
+{
+ mutex_lock(&generic_pt_domains_list_lock);
+ list_add_tail(&iommu->shrinker_list, &generic_pt_domains_list);
+ mutex_unlock(&generic_pt_domains_list_lock);
+}
+EXPORT_SYMBOL_GPL(generic_pt_shrinker_add);
+
+void generic_pt_shrinker_remove(struct pt_iommu *iommu)
+{
+ mutex_lock(&generic_pt_domains_list_lock);
+ list_del(&iommu->shrinker_list);
+ mutex_unlock(&generic_pt_domains_list_lock);
+}
+EXPORT_SYMBOL_GPL(generic_pt_shrinker_remove);
+
+static unsigned long generic_pt_shrinker_count(struct shrinker *shrink,
+ struct shrink_control *sc)
+{
+ struct pt_iommu *iommu;
+ unsigned long count = 0;
+
+ mutex_lock(&generic_pt_domains_list_lock);
+ list_for_each_entry(iommu, &generic_pt_domains_list, shrinker_list) {
+ struct iommu_domain *domain = &iommu->domain;
+
+ /* TODO: Return real nr_pages here! */
+ if (!xa_empty(&domain->reclaim_list))
+ count += 1;
+ }
+ mutex_unlock(&generic_pt_domains_list_lock);
+
+ return count;
+}
+
+static unsigned long generic_pt_shrinker_scan(struct shrinker *shrink,
+ struct shrink_control *sc)
+{
+ struct iommu_pages_list free_list = IOMMU_PAGES_LIST_INIT(free_list);
+ struct pt_iommu *iommu;
+ struct ioptdesc *ioptdesc;
+ unsigned long iova;
+ unsigned long freed_count = 0;
+
+ mutex_lock(&generic_pt_domains_list_lock);
+ list_for_each_entry(iommu, &generic_pt_domains_list, shrinker_list) {
+ struct iommu_domain *domain = &iommu->domain;
+ unsigned long prev_freed = freed_count;
+
+ xa_for_each(&domain->reclaim_list, iova, ioptdesc) {
+ if (freed_count >= sc->nr_to_scan)
+ break;
+
+ if (atomic_cmpxchg(&ioptdesc->__page_refcount, 1, 0) == 1) {
+ void *virt = folio_address(ioptdesc_folio(ioptdesc));
+
+ if (iommu->ops->sever_branch(iommu, iova, virt_to_phys(virt))) {
+ /* Successfully severed, queue for freeing */
+ xa_erase(&domain->reclaim_list, iova);
+ iommu_pages_list_add(&free_list, virt);
+ freed_count++;
+ } else {
+ /* We raced with map, abort pruning */
+ xa_erase(&domain->reclaim_list, iova);
+ }
+ } else {
+ /* Not empty anymore, remove from list */
+ xa_erase(&domain->reclaim_list, iova);
+ }
+ }
+
+ /*
+ * The IOMMU may still cache pointers to the severed tables in its
+ * paging-structure caches, flush them before the tables are freed.
+ */
+ if (freed_count != prev_freed)
+ iommu_flush_iotlb_all(domain);
+ }
+ mutex_unlock(&generic_pt_domains_list_lock);
+
+ /* If we didn't sever anything, just return */
+ if (list_empty(&free_list.pages))
+ return freed_count;
+
+ /* Wait ONE time for the entire system */
+ synchronize_srcu(&generic_pt_srcu);
+
+ /* Free pages from all domains */
+ struct page *page, *next;
+
+ list_for_each_entry_safe(page, next, &free_list.pages, lru) {
+ /* We must restore the refcount to 1 before freeing */
+ set_page_count(page, 1);
+ }
+
+ iommu_put_pages_list(&free_list);
+
+ return freed_count;
+}
+
+static int __init generic_pt_shrinker_init(void)
+{
+ struct shrinker *shrinker;
+
+ shrinker = shrinker_alloc(0, "iommu-generic-pt");
+ if (!shrinker)
+ return -ENOMEM;
+
+ shrinker->count_objects = generic_pt_shrinker_count;
+ shrinker->scan_objects = generic_pt_shrinker_scan;
+
+ shrinker_register(shrinker);
+ return 0;
+}
+subsys_initcall(generic_pt_shrinker_init);
diff --git a/include/linux/generic_pt/iommu.h b/include/linux/generic_pt/iommu.h
index d1ade1767a9a..95bf9d4f0b6e 100644
--- a/include/linux/generic_pt/iommu.h
+++ b/include/linux/generic_pt/iommu.h
@@ -64,8 +64,18 @@ struct pt_iommu {
* page table which must have dma ops that perform cache flushing.
*/
struct device *iommu_device;
+
+ /**
+ * @shrinker_list: Node for the generic_pt global shrinker list
+ */
+ struct list_head shrinker_list;
};
+extern struct srcu_struct generic_pt_srcu;
+
+void generic_pt_shrinker_add(struct pt_iommu *iommu);
+void generic_pt_shrinker_remove(struct pt_iommu *iommu);
+
static inline struct pt_iommu *iommupt_from_domain(struct iommu_domain *domain)
{
if (!IS_ENABLED(CONFIG_IOMMU_PT) || !domain->is_iommupt)
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 6+ messages in thread
* [RFC PATCH 5/5] iommupt: Return real page count to the shrinker core
2026-10-01 23:02 [RFC PATCH 0/5] iommupt: Introduce IO page table shrinker Pranjal Shrivastava
` (3 preceding siblings ...)
2026-10-01 23:02 ` [RFC PATCH 4/5] iommupt: Introduce lockless page table shrinker Pranjal Shrivastava
@ 2026-10-01 23:02 ` Pranjal Shrivastava
4 siblings, 0 replies; 6+ messages in thread
From: Pranjal Shrivastava @ 2026-10-01 23:02 UTC (permalink / raw)
To: Joerg Roedel, Will Deacon, Robin Murphy, Jason Gunthorpe, Kevin Tian
Cc: Mostafa Saleh, Daniel Mentz, Samiullah Khawaja, Logan Odell,
iommu, linux-mm, linux-kernel, Pranjal Shrivastava
Update the count op to return real nr_pages that can be reclaimed.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
drivers/iommu/generic_pt/iommu_pt.h | 1 +
drivers/iommu/generic_pt/shrinker.c | 9 ++++-----
include/linux/generic_pt/iommu.h | 5 +++++
3 files changed, 10 insertions(+), 5 deletions(-)
diff --git a/drivers/iommu/generic_pt/iommu_pt.h b/drivers/iommu/generic_pt/iommu_pt.h
index e5ee419f2f0b..4386caf3f713 100644
--- a/drivers/iommu/generic_pt/iommu_pt.h
+++ b/drivers/iommu/generic_pt/iommu_pt.h
@@ -1115,6 +1115,7 @@ static __maybe_unused int __unmap_range(struct pt_range *range, void *arg,
xa_store(&iommu->domain.reclaim_list, range->va,
virt_to_ioptdesc(pts.table), GFP_ATOMIC);
+ atomic_long_inc(&iommu->nr_reclaimable);
}
gather_add_leaf(&unmap->pending, &pts);
diff --git a/drivers/iommu/generic_pt/shrinker.c b/drivers/iommu/generic_pt/shrinker.c
index 66169fb5db71..a27e171b5e80 100644
--- a/drivers/iommu/generic_pt/shrinker.c
+++ b/drivers/iommu/generic_pt/shrinker.c
@@ -43,11 +43,7 @@ static unsigned long generic_pt_shrinker_count(struct shrinker *shrink,
mutex_lock(&generic_pt_domains_list_lock);
list_for_each_entry(iommu, &generic_pt_domains_list, shrinker_list) {
- struct iommu_domain *domain = &iommu->domain;
-
- /* TODO: Return real nr_pages here! */
- if (!xa_empty(&domain->reclaim_list))
- count += 1;
+ count += atomic_long_read(&iommu->nr_reclaimable);
}
mutex_unlock(&generic_pt_domains_list_lock);
@@ -78,15 +74,18 @@ static unsigned long generic_pt_shrinker_scan(struct shrinker *shrink,
if (iommu->ops->sever_branch(iommu, iova, virt_to_phys(virt))) {
/* Successfully severed, queue for freeing */
xa_erase(&domain->reclaim_list, iova);
+ atomic_long_dec(&iommu->nr_reclaimable);
iommu_pages_list_add(&free_list, virt);
freed_count++;
} else {
/* We raced with map, abort pruning */
xa_erase(&domain->reclaim_list, iova);
+ atomic_long_dec(&iommu->nr_reclaimable);
}
} else {
/* Not empty anymore, remove from list */
xa_erase(&domain->reclaim_list, iova);
+ atomic_long_dec(&iommu->nr_reclaimable);
}
}
diff --git a/include/linux/generic_pt/iommu.h b/include/linux/generic_pt/iommu.h
index 95bf9d4f0b6e..3c94745bded0 100644
--- a/include/linux/generic_pt/iommu.h
+++ b/include/linux/generic_pt/iommu.h
@@ -69,6 +69,11 @@ struct pt_iommu {
* @shrinker_list: Node for the generic_pt global shrinker list
*/
struct list_head shrinker_list;
+
+ /**
+ * @nr_reclaimable: Count of empty directories waiting in reclaim_list
+ */
+ atomic_long_t nr_reclaimable;
};
extern struct srcu_struct generic_pt_srcu;
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 6+ messages in thread
end of thread, other threads:[~2026-10-01 23:02 UTC | newest]
Thread overview: 6+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-01 23:02 [RFC PATCH 0/5] iommupt: Introduce IO page table shrinker Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 1/5] iommu: Add reclaim_list xarray to struct iommu_domain Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 2/5] iommupt: Implement refcounting logic for Leaf entries Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 3/5] iommupt: Add lockless sever_branch helper Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 4/5] iommupt: Introduce lockless page table shrinker Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 5/5] iommupt: Return real page count to the shrinker core Pranjal Shrivastava
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®