mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "David Hildenbrand (Arm)" <david@kernel.org>
To: Mikhail Gavrilov <mikhail.v.gavrilov@gmail.com>,
	Andrew Morton <akpm@linux-foundation.org>,
	Dave Hansen <dave.hansen@linux.intel.com>
Cc: Lorenzo Stoakes <ljs@kernel.org>,
	"Liam R . Howlett" <liam@infradead.org>,
	Vlastimil Babka <vbabka@kernel.org>,
	Mike Rapoport <rppt@kernel.org>,
	Suren Baghdasaryan <surenb@google.com>,
	Michal Hocko <mhocko@suse.com>,
	Vishal Moola <vishal.moola@gmail.com>,
	Ingo Molnar <mingo@kernel.org>,
	Lu Baolu <baolu.lu@linux.intel.com>,
	Jason Gunthorpe <jgg@nvidia.com>,
	Steven Rostedt <rostedt@goodmis.org>,
	x86@kernel.org, linux-mm@kvack.org, regressions@lists.linux.dev,
	linux-kernel@vger.kernel.org
Subject: Re: [PATCH] mm: don't defer freeing kernel page tables while booting
Date: Thu, 24 Sep 2026 09:07:45 +0200	[thread overview]
Message-ID: <41a59ba9-039e-454f-8d31-f647db802ada@kernel.org> (raw)
In-Reply-To: <20260924064321.23787-1-mikhail.v.gavrilov@gmail.com>

On 9/24/26 08:43, Mikhail Gavrilov wrote:
> Booting with a boot-time function tracer and a filter, for example
> 
>   ftrace=function ftrace_filter=pud_free_pmd_page
> 
> panics on 7.3-rc4 as soon as the tracer starts:
> 
>   [   23.531178] Starting tracer 'function'
>   [   23.675800] Oops: general protection fault, probably for non-canonical address 0xdffffc0000000038: 0000 [#1] SMP KASAN NOPTI
>   [   23.819917] KASAN: null-ptr-deref in range [0x00000000000001c0-0x00000000000001c7]
>   [   23.964025] CPU: 0 UID: 0 PID: 0 Comm: swapper Not tainted 7.3.0-rc4-fe2ec83746e5-with-fixes-v2+ #195 PREEMPT(undef)
>   [   24.252248] RIP: 0010:__queue_work+0xab/0xf00
>   [   25.981629] Call Trace:
>   [   26.125727]  <TASK>
>   [   26.413912]  ? pagetable_free_kernel+0x20/0x120
>   [   26.990283]  queue_work_on+0x97/0xf0
>   [   27.134382]  __cpa_collapse_large_pages+0x501/0x6f0
>   [   27.566662]  cpa_flush+0x394/0x620
>   [   27.998953]  change_page_attr_set_clr+0x321/0x4a0
>   [   29.151729]  set_memory_rox+0xa2/0xf0
>   [   29.584018]  create_trampoline+0x431/0x6f0
>   ...
>   [   44.343347] Kernel panic - not syncing: Attempted to kill the idle task!
> 
> The boot-time tracer is started from early_trace_init(), which runs
> before workqueue_init_early().  Making its trampoline read-only splits a
> large page, and CPA collapses it again right away.  The split table has
> been a kernel page table since commit 9e4a3ec3411b
> ("x86/mm/pat: Allocate split page tables as kernel page tables"), so the
> collapse frees it through pagetable_free_kernel(), which queues work on
> system_percpu_wq - still NULL at that point.
> 
> The deferral exists so that IOMMUs using SVA can have their paging
> structure caches flushed before a kernel page table is freed.  While
> system_state is still SYSTEM_BOOTING only the boot CPU runs and no IOMMU
> has been initialised yet (on x86 that happens from pci_iommu_init(), a
> rootfs_initcall), so there is nothing to flush.  Free the table directly
> in that case.
> 
> Fixes: 9e4a3ec3411b ("x86/mm/pat: Allocate split page tables as kernel page tables")
> Cc: stable@vger.kernel.org
> Signed-off-by: Mikhail Gavrilov <mikhail.v.gavrilov@gmail.com>
> ---
> #regzbot introduced: 9e4a3ec3411b
> 
> Tested on a Ryzen 9 7950X with a Radeon RX 7900 XTX, lockdep and KASAN
> enabled, booting with the command line above plus earlycon=efifb so
> that the oops is visible:
> 
>   7.3-rc4 (fe2ec83746e5)             panics as quoted
>   + revert of 9e4a3ec3411b           boots
>   + this patch, 9e4a3ec3411b kept    boots
> 
> With this patch the boot-time tracer works as intended and records
> pud_free_pmd_page() being called from vmap_p4d_range() during boot.
> All three kernels also carry my pud_free_pmd_page() v2 patch and
> unrelated local changes elsewhere (HID, sound/usb, Bluetooth, debugfs);
> none of them is on this path.
> 
> 9e4a3ec3411b is also in 7.2.7 and 6.18.53, together with the deferred
> pagetable_free_kernel(); I have not tried those.
> 
>  mm/pgtable-generic.c | 11 +++++++++++
>  1 file changed, 11 insertions(+)
> 
> diff --git a/mm/pgtable-generic.c b/mm/pgtable-generic.c
> index b91b1a98029c..2d9b4ba33b46 100644
> --- a/mm/pgtable-generic.c
> +++ b/mm/pgtable-generic.c
> @@ -440,6 +440,17 @@ static void kernel_pgtable_work_func(struct work_struct *work)
>  
>  void pagetable_free_kernel(struct ptdesc *pt)
>  {
> +	/*
> +	 * While the system is still booting only the boot CPU runs and no
> +	 * IOMMU has been set up, so nothing can be caching this table and
> +	 * there is nothing to flush.  The workqueue this defers to may not
> +	 * exist yet either.
> +	 */
> +	if (system_state == SYSTEM_BOOTING) {
> +		__pagetable_free(pt);
> +		return;
> +	}
> +
>  	spin_lock(&kernel_pgtable_work.lock);
>  	list_add(&pt->pt_list, &kernel_pgtable_work.list);
>  	spin_unlock(&kernel_pgtable_work.lock);

Should we instead simply skip the

	schedule_work(&kernel_pgtable_work.work);

and rely on anybody freeing stuff later to just free that one alongside?

That avoids throwing in more freeing handling.

-- 
Cheers,

David

  reply	other threads:[~2026-09-24  7:07 UTC|newest]

Thread overview: 8+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-24  6:43 Mikhail Gavrilov
2026-09-24  7:07 ` David Hildenbrand (Arm) [this message]
2026-09-24  7:28   ` Mikhail Gavrilov
2026-09-24  7:31     ` David Hildenbrand (Arm)
2026-09-24  8:57       ` Lorenzo Stoakes (ARM)
2026-09-24  8:59         ` Lorenzo Stoakes (ARM)
2026-09-24  9:29           ` Lorenzo Stoakes (ARM)
2026-09-24  7:38     ` Lorenzo Stoakes (ARM)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=41a59ba9-039e-454f-8d31-f647db802ada@kernel.org \
    --to=david@kernel.org \
    --cc=akpm@linux-foundation.org \
    --cc=baolu.lu@linux.intel.com \
    --cc=dave.hansen@linux.intel.com \
    --cc=jgg@nvidia.com \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@suse.com \
    --cc=mikhail.v.gavrilov@gmail.com \
    --cc=mingo@kernel.org \
    --cc=regressions@lists.linux.dev \
    --cc=rostedt@goodmis.org \
    --cc=rppt@kernel.org \
    --cc=surenb@google.com \
    --cc=vbabka@kernel.org \
    --cc=vishal.moola@gmail.com \
    --cc=x86@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®