From: Pranjal Shrivastava <praan@google.com>
To: Jason Gunthorpe <jgg@ziepe.ca>
Cc: Joerg Roedel <joro@8bytes.org>, Will Deacon <will@kernel.org>,
Robin Murphy <robin.murphy@arm.com>,
Kevin Tian <kevin.tian@intel.com>,
Mostafa Saleh <smostafa@google.com>,
Daniel Mentz <danielmentz@google.com>,
Samiullah Khawaja <skhawaja@google.com>,
Logan Odell <loganodell@google.com>,
iommu@lists.linux.dev, linux-mm@kvack.org,
linux-kernel@vger.kernel.org
Subject: Re: [RFC PATCH 0/5] iommupt: Introduce IO page table shrinker
Date: Mon, 5 Oct 2026 04:58:50 +0000 [thread overview]
Message-ID: <asMuigioXmHZmjO2@google.com> (raw)
In-Reply-To: <20261002150844.GC3481470@ziepe.ca>
On Fri, Oct 02, 2026 at 12:08:44PM -0300, Jason Gunthorpe wrote:
> On Thu, Oct 01, 2026 at 11:02:14PM +0000, Pranjal Shrivastava wrote:
> > Introduce a lockless, deferred reclamation framework for IOMMU page tables
> > built on the generic_pt library. As VMMs and userspace drivers map and
> > unmap large, sparse IOVA regions through VFIO and iommufd, page table
> > directories are often left allocated but completely empty. generic_pt
> > frees a table when a single unmap covers it entirely, but tables that
> > empty through a series of partial unmaps stay allocated until the domain
> > is destroyed. Under memory pressure, this *stranded* memory cannot be
> > reclaimed and has resulted in OOMs.
> >
> > This series refcounts leaf directories natively in struct ioptdesc and
> > registers a domain-aware MM shrinker that prunes empty directories under
> > system memory pressure.
>
> We had talked about doing it this way
>
> But I had proposed a different, and possibly simpler, solution that
> addresses *just* the iommufd use case.
>
> After unmapping something have iommufd compute the gap in IOVA that
> contains what was unmapped and then issue a 'clean(gap)' operation to
> generic_pt.
>
> This is the same operation as unmap, except we know now that the gap
> has no PTEs so all it does is clean up the table pointers.
>
> This requires no special refcounting or anything difficult beyond
> some locking in iommufd to hold the gap stable while we clean it.
>
> Would it work for you? It seems substantially simpler, but I never
> tried to implement it.
>
I was tempted to use the interval trees too, but I started thinking
about:
a) Locking: For the gap to stay stable while we clean it, we'd need to
add some kind of serialization either through iova_rwsem or a dedicated
gap_lock to prevent a concurrent map to allocate IOVA from that gap.
Thus, every unmap pays for clean under the lock. I haven't perf-ed it
yet, but I'm not sure whether users/Guests using virtio-iommu, where any
guest DMA unmap becomes an unmap on the host, would regress.
b) Other users of IOMMU API (unmanaged domain) like the type 1 (which
was the one hurting our systems) and other in-tree drivers that use an
unmanaged domain and call iommu_map()/iommu_unmap() with their own IOVA
allocator. One of the goals was to avoid enabling the user-space or
in-kernel IOVA allocation (ab)users to cause OOMs via IOPT allocation.
> An alternative version is closer to what you have here, somehow
> connect iommufd to the shrinker and have it lock and walk the gaps
> cleaning them on shrink requests?
>
I like this one better than cleaning on every unmap, since it keeps
the cost off the map/unmap path and only does work under memory
pressure. Although, it doesn't help type1 or the in-tree iommu_map()
users.
Would you consider those users worth covering, or would you rather
they track their own holes and call clean() themselves? Happy to dig
into this at the LPC session too.
Thanks,
Praan
next prev parent reply other threads:[~2026-10-05 4:58 UTC|newest]
Thread overview: 9+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-01 23:02 Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 1/5] iommu: Add reclaim_list xarray to struct iommu_domain Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 2/5] iommupt: Implement refcounting logic for Leaf entries Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 3/5] iommupt: Add lockless sever_branch helper Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 4/5] iommupt: Introduce lockless page table shrinker Pranjal Shrivastava
2026-10-01 23:02 ` [RFC PATCH 5/5] iommupt: Return real page count to the shrinker core Pranjal Shrivastava
2026-10-02 15:08 ` [RFC PATCH 0/5] iommupt: Introduce IO page table shrinker Jason Gunthorpe
2026-10-05 4:58 ` Pranjal Shrivastava [this message]
2026-10-03 8:10 ` [syzbot ci] " syzbot ci
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=asMuigioXmHZmjO2@google.com \
--to=praan@google.com \
--cc=danielmentz@google.com \
--cc=iommu@lists.linux.dev \
--cc=jgg@ziepe.ca \
--cc=joro@8bytes.org \
--cc=kevin.tian@intel.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=loganodell@google.com \
--cc=robin.murphy@arm.com \
--cc=skhawaja@google.com \
--cc=smostafa@google.com \
--cc=will@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®