From: Pranjal Shrivastava <praan@google.com>
To: Joerg Roedel <joro@8bytes.org>, Will Deacon <will@kernel.org>,
Robin Murphy <robin.murphy@arm.com>,
Jason Gunthorpe <jgg@ziepe.ca>, Kevin Tian <kevin.tian@intel.com>
Cc: Mostafa Saleh <smostafa@google.com>,
Daniel Mentz <danielmentz@google.com>,
Samiullah Khawaja <skhawaja@google.com>,
Logan Odell <loganodell@google.com>,
iommu@lists.linux.dev, linux-mm@kvack.org,
linux-kernel@vger.kernel.org,
Pranjal Shrivastava <praan@google.com>
Subject: [RFC PATCH 0/5] iommupt: Introduce IO page table shrinker
Date: Thu, 1 Oct 2026 23:02:14 +0000 [thread overview]
Message-ID: <20261001230219.818128-1-praan@google.com> (raw)
Introduce a lockless, deferred reclamation framework for IOMMU page tables
built on the generic_pt library. As VMMs and userspace drivers map and
unmap large, sparse IOVA regions through VFIO and iommufd, page table
directories are often left allocated but completely empty. generic_pt
frees a table when a single unmap covers it entirely, but tables that
empty through a series of partial unmaps stay allocated until the domain
is destroyed. Under memory pressure, this *stranded* memory cannot be
reclaimed and has resulted in OOMs.
This series refcounts leaf directories natively in struct ioptdesc and
registers a domain-aware MM shrinker that prunes empty directories under
system memory pressure.
This is an RFC to align on the design. It has known issues, listed below,
and is not intended for merging as-is.
Design
======
- Refcounting: the __page_refcount of a leaf directory's struct ioptdesc
counts its valid leaf entries, plus one for the table itself. Map takes
a reference with atomic_inc_not_zero() before installing a leaf, and
unmap drops it after clearing one. Leaf entries are counted only in the
table that holds them, so map and unmap never update refcounts on the
shared tables above it.
- Queueing: when the refcount drops to 1, the directory is empty and is
logged with its IOVA into a per-domain xarray (reclaim_list). Only
non-DMA domains are queued.
- Claim: under memory pressure, the shrinker walks the queued directories
and claims each with atomic_cmpxchg(refcount, 1, 0). A map that finds
the refcount at 0 fails atomic_inc_not_zero() and retries, so it never
reuses a claimed directory.
- Sever: a format-agnostic top-down walk (sever_branch) clears the parent
slot with cmpxchg, severing the dead branch from the live tree.
- Free: the domain's IOTLB is flushed so the IOMMU drops any cached
pointers to the severed tables. A single synchronize_srcu() then waits
for in-flight map walkers before the pages are freed.
The cost is one atomic operation per leaf entry on map and unmap, plus an
xarray insert when a directory becomes empty.
Known issues
============
These are understood and will be addressed in the next version:
- Map vs. shrinker retry: on a claimed directory, map returns -EAGAIN,
which __map_range() already uses to re-descend into the cached table
pointer, i.e. the dead directory. It needs a distinct error and a retry
from the top.
- Unmap vs. shrinker hand-off: an unmap that fully covers a queued
directory, or a large-page map over an empty range, frees it directly
while the xarray still references it.
- xarray keys: the key is the IOVA where the emptying unmap started, so a
directory can be queued under several keys and nr_reclaimable
over-counts. Keying by the table's PFN makes entries unique.
- A failed sever leaves the refcount at 0 instead of restoring it to 1.
The top-level table, which has no parent, should never be queued.
- iommufd allocates paging domains without iommu_domain_init(), so their
reclaim_list is not initialised.
- Combined with the observability series, pages must be uncharged from
their domain at sever time, as they are freed after the domain may be
gone.
Open questions
==============
- Leaf directories with child tables: a leaf directory above the last
level (e.g. one holding a 2M leaf) can gain a child table, which is not
counted, and could then be reclaimed while the child is live. Should
child tables be counted at install, or counting be restricted to
last-level directories?
- SRCU in reclaim context: the shrinker can run from direct reclaim
inside a map's SRCU read section, where synchronize_srcu() would wait on
itself. Should the free be deferred (work item / call_srcu), or is a
different scheme preferred?
- IOTLB flush granularity: flush_iotlb_all() is used after severing.
Would a range-based paging-structure flush be preferred?
- DMA-API domains are excluded since their unmap can run in IRQ context.
Should the per-leaf atomics also be skipped for them?
Upcoming Work / Roadmap
=======================
An IO page table observability series, which attributes page table memory
to its domain and exposes it via fdinfo, is posted separately as an RFC.
The two series are independent; this one applies on v7.3-rc5 by itself.
There's an alignment session at Linux Plumbers Conference 2026 for these [1].
[1] https://lpc.events/event/20/contributions/2624/
Pranjal Shrivastava (5):
iommu: Add reclaim_list xarray to struct iommu_domain
iommupt: Implement refcounting logic for Leaf entries
iommupt: Add lockless sever_branch helper
iommupt: Introduce lockless page table shrinker
iommupt: Return real page count to the shrinker core
drivers/iommu/Makefile | 2 +-
drivers/iommu/generic_pt/iommu_pt.h | 81 ++++++++++++++++-
drivers/iommu/generic_pt/shrinker.c | 135 ++++++++++++++++++++++++++++
drivers/iommu/iommu.c | 2 +
include/linux/generic_pt/iommu.h | 29 ++++++
include/linux/iommu.h | 3 +
6 files changed, 249 insertions(+), 3 deletions(-)
create mode 100644 drivers/iommu/generic_pt/shrinker.c
base-commit: 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e
--
2.56.0.rc1.315.gc6ed9934b7-goog
next reply other threads:[~2026-10-01 23:02 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-01 23:02 Pranjal Shrivastava [this message]
2026-10-01 23:02 ` [RFC PATCH 1/5] iommu: Add reclaim_list xarray to struct iommu_domain Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 2/5] iommupt: Implement refcounting logic for Leaf entries Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 3/5] iommupt: Add lockless sever_branch helper Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 4/5] iommupt: Introduce lockless page table shrinker Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 5/5] iommupt: Return real page count to the shrinker core Pranjal Shrivastava
2026-10-02 15:08 ` [RFC PATCH 0/5] iommupt: Introduce IO page table shrinker Jason Gunthorpe
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261001230219.818128-1-praan@google.com \
--to=praan@google.com \
--cc=danielmentz@google.com \
--cc=iommu@lists.linux.dev \
--cc=jgg@ziepe.ca \
--cc=joro@8bytes.org \
--cc=kevin.tian@intel.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=loganodell@google.com \
--cc=robin.murphy@arm.com \
--cc=skhawaja@google.com \
--cc=smostafa@google.com \
--cc=will@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®