From: Kiryl Shutsemau <kas@kernel.org>
To: Xu Yilun <yilun.xu@linux.intel.com>
Cc: david@kernel.org, linux-mm@kvack.org, x86@kernel.org,
linux-coco@lists.linux.dev, linux-kernel@vger.kernel.org,
rick.p.edgecombe@intel.com, yilun.xu@intel.com,
xiaoyao.li@intel.com, sohil.mehta@intel.com,
adrian.hunter@intel.com, kishen.maloor@intel.com,
tony.lindgren@linux.intel.com, peter.fang@intel.com,
baolu.lu@linux.intel.com, zhenzhong.duan@intel.com,
chao.gao@intel.com, artem.bityutskiy@linux.intel.com,
kvm@vger.kernel.org
Subject: Re: [PATCH 4/6] x86/virt/tdx: Add extra memory to TDX module for the extensions
Date: Thu, 3 Sep 2026 12:00:57 +0100 [thread overview]
Message-ID: <aplSf00ESix5s6Ak@thinkstation> (raw)
In-Reply-To: <aplOhHWvaxqUJDqu@yilunxu-OptiPlex-7050>
On Thu, Sep 03, 2026 at 06:40:04PM +0800, Xu Yilun wrote:
> On Wed, Sep 02, 2026 at 11:54:25AM +0100, Kiryl Shutsemau wrote:
> > On Fri, Aug 28, 2026 at 03:46:16PM +0800, Xu Yilun wrote:
> > > > > > + page = alloc_contig_pages(required_pages, GFP_KERNEL, numa_mem_id(),
> > > > > > + &node_online_map);
> > > > >
> > > > > Why contiguous? TDH.EXT.MEM.ADD takes a list of page addresses and the loop
> > > > > below writes every one of them out separately.
> > > > >
> > > > > alloc_pages_bulk() fits the chunking that is already here, and a short
> > > > > return can be handled per chunk. alloc_contig_pages() isolates and migrates
> > > > > to get its range and fails TDX init outright when it cannot find one. PAMT
> > > >
> > > > Yeah, this is not the ABI requirement, but the kernel's consideration. A
> > > > brief reasoning in the commit log: avoiding permanent memory fragmentation
> > > > and buddy allocator efficiency loss.
> > > >
> > > > Also there is some discussion:
> > > >
> > > > https://lore.kernel.org/all/167d9540-2d9a-4367-bc68-b96494bc4044@intel.com/
> > > >
> > > > TL;DR
> > > > - The memory will never return to the kernel.
> > > > - There is chance that this tens of megabytes will fragment tens of
> > > > gigabytes of memory forever.
> > > > - The chance of fragmentation is actually low since at boot up, but
> > > > let the buddy allocator take care of these never-returned memory
> > > > is not necessary and lowers its efficiency.
> > >
> > > Hi Kiryl & David:
> > >
> > > I see there is another suggestion that the whole memory adding process
> > > could be a little simpler if we allocate & add pages 4k by 4k [1],
> > > rather than one-time pre-allocation. The concern of this alternative is,
> > > as said above, memory fragmentation.
> > >
> > > [1] https://lore.kernel.org/lkml/f48b83feb2ee1d3c88b5a1627cf35b4b282d3f90.camel@intel.com/
> > >
> > > And I've realized the memory fragmentation discussion is not actually
> > > closed in previous thread [2]. We need more input.
> > >
> > > [2] https://lore.kernel.org/all/167d9540-2d9a-4367-bc68-b96494bc4044@intel.com/
> > >
> > > Let me give a brief overview of the problem:
> > >
> > > Intel TDX (Trust Domain Extensions) is a feature for confidential
> > > computing. A secure firmware called "TDX module" runs in an isolated
> > > environment to provide services about security.
> > >
> > > In Linux, the host initializes TDX module at boot up time
> > > (subsys_initcall()). During the initialization, the host must donate
> > > tens of mega bytes physical memory (35M ~ 110M in the forseeable
> > > future) to the TDX module. These memory will *never be revoked* cause
> > > the TDX Module initialization is a one way path.
> > >
> > > The TDX Module doesn't require this memory be physically contiguous. But
> > > the kernel side concern is if we do PAGE_SIZE allocation, it may
> > > permanently fragment memory regions, stop them from allocating 2M huge
> > > pages. In worst case, ~50G (110M * 512) memory regions affacted.
> > >
> > > So is the physically contiguous allocation really a better choice here?
> > > We appreciate inputs from mm folks. Thanks!
> >
> > What matters for fragmentation is not contiguity, it is how many
> > pageblocks are left partially occupied by unmovable pages that are never
> > freed.
> >
> > So you can ask one pageblock at a time with
> >
> > page = alloc_pages(GFP_KERNEL | __GFP_NOWARN, order);
> >
> > with fallback to lower order if you must.
> >
> > It also fits the ABI: pageblock_order is 9 on x86, i.e. 512 pages, which
> > is exactly TDX_HPA_LIST_MAX_NR_PAGES. One allocation is one full HPA list
> > is one TDH.EXT.MEM.ADD, so the allocation loop and the chunking loop
> > become the same loop.
> >
> > But alloc_contig_pages() might be a good enough approximation for
> > per-pageblock allocation if we do it during the boot when fragmentation
> > is low.
>
> IIUC, you mean alloc_contig_pages() also gives good de-fragmentation
> that we need. But it would be slightly easier to fail cause it requires
> extra contiguity that we don't need.
alloc_contig_pages() can be more expensive than needed (or fail) since
you ask for the full allocation size to be contiguous, where you should
be okay with a set of pageblocks regardless where they are relative to
each other.
> Multiple alloc_pages(order-9) meets our requirement exactly but the
> falling back to lower order may create more fragments. And we can do
> this because of the ABI definition - an HPA_LIST could happen to hold
> an entire pageblock.
>
> If I have to choose, I prefer alloc_contig_pages(). It doesn't have to
> depend on HPA_LIST ABI details.
As I said before, as long as you do it once during the boot, it should
be good enough.
--
Kiryl Shutsemau / Kirill A. Shutemov
next prev parent reply other threads:[~2026-09-03 11:01 UTC|newest]
Thread overview: 37+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-21 3:29 [PATCH 0/6] Enable TDX module extensions Xu Yilun
2026-08-21 3:29 ` [PATCH 1/6] x86/virt/tdx: Wrap TDH.SYS.CONFIG/UPDATE operations in helpers Xu Yilun
2026-08-21 20:53 ` Edgecombe, Rick P
2026-08-24 4:52 ` Xu Yilun
2026-08-26 13:57 ` Nikolay Borisov
2026-08-21 3:29 ` [PATCH 2/6] x86/virt/tdx: Configure add-on features on TDX module init and update Xu Yilun
2026-08-21 14:38 ` Dave Hansen
2026-08-21 21:18 ` Edgecombe, Rick P
2026-08-24 6:37 ` Xu Yilun
2026-08-24 15:15 ` Dave Hansen
2026-08-24 18:34 ` Xu Yilun
2026-08-21 22:01 ` Edgecombe, Rick P
2026-08-26 5:15 ` Xu Yilun
2026-08-21 3:29 ` [PATCH 3/6] x86/virt/tdx: Detect if the extensions initialization is required Xu Yilun
2026-08-21 15:22 ` Kiryl Shutsemau
2026-08-21 22:22 ` Edgecombe, Rick P
2026-08-24 12:16 ` Kiryl Shutsemau
2026-08-21 3:29 ` [PATCH 4/6] x86/virt/tdx: Add extra memory to TDX module for the extensions Xu Yilun
2026-08-21 15:44 ` Kiryl Shutsemau
2026-08-24 9:18 ` Xu Yilun
2026-08-24 12:22 ` Kiryl Shutsemau
2026-08-28 7:46 ` Xu Yilun
2026-09-02 10:54 ` Kiryl Shutsemau
2026-09-03 10:40 ` Xu Yilun
2026-09-03 11:00 ` Kiryl Shutsemau [this message]
2026-09-03 15:05 ` Xu Yilun
2026-08-21 3:29 ` [PATCH 5/6] x86/virt/tdx: Make TDX module initialize " Xu Yilun
2026-08-21 23:55 ` Edgecombe, Rick P
2026-08-24 17:14 ` Xu Yilun
2026-08-24 17:43 ` Edgecombe, Rick P
2026-08-25 9:53 ` Xu Yilun
2026-08-24 17:58 ` Edgecombe, Rick P
2026-08-25 10:02 ` Xu Yilun
2026-08-21 3:29 ` [PATCH 6/6] x86/virt/tdx: Re-initialize the extensions on runtime TDX module update Xu Yilun
2026-08-22 0:01 ` Edgecombe, Rick P
2026-08-25 15:39 ` Xu Yilun
2026-08-25 8:35 ` Tony Lindgren
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aplSf00ESix5s6Ak@thinkstation \
--to=kas@kernel.org \
--cc=adrian.hunter@intel.com \
--cc=artem.bityutskiy@linux.intel.com \
--cc=baolu.lu@linux.intel.com \
--cc=chao.gao@intel.com \
--cc=david@kernel.org \
--cc=kishen.maloor@intel.com \
--cc=kvm@vger.kernel.org \
--cc=linux-coco@lists.linux.dev \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=peter.fang@intel.com \
--cc=rick.p.edgecombe@intel.com \
--cc=sohil.mehta@intel.com \
--cc=tony.lindgren@linux.intel.com \
--cc=x86@kernel.org \
--cc=xiaoyao.li@intel.com \
--cc=yilun.xu@intel.com \
--cc=yilun.xu@linux.intel.com \
--cc=zhenzhong.duan@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®