From: Mike Rapoport <rppt@kernel.org>
To: Mikhail Gavrilov <mikhail.v.gavrilov@gmail.com>
Cc: Andrew Morton <akpm@linux-foundation.org>,
David Hildenbrand <david@kernel.org>,
Dave Hansen <dave.hansen@linux.intel.com>,
Lorenzo Stoakes <ljs@kernel.org>,
"Liam R . Howlett" <liam@infradead.org>,
Vlastimil Babka <vbabka@kernel.org>,
Suren Baghdasaryan <surenb@google.com>,
Michal Hocko <mhocko@suse.com>,
Vishal Moola <vishal.moola@gmail.com>,
Ingo Molnar <mingo@kernel.org>,
Lu Baolu <baolu.lu@linux.intel.com>,
Jason Gunthorpe <jgg@nvidia.com>,
Steven Rostedt <rostedt@goodmis.org>,
x86@kernel.org, linux-mm@kvack.org, regressions@lists.linux.dev,
linux-kernel@vger.kernel.org
Subject: Re: [PATCH v2] mm: don't schedule deferred kernel page table freeing while booting
Date: Thu, 24 Sep 2026 19:19:17 +0300 [thread overview]
Message-ID: <arVNhTkI3kesfRs0@kernel.org> (raw)
In-Reply-To: <20260924092307.22813-1-mikhail.v.gavrilov@gmail.com>
On Thu, Sep 24, 2026 at 02:23:07PM +0500, Mikhail Gavrilov wrote:
> Booting with a boot-time function tracer and a filter, for example
>
> ftrace=function ftrace_filter=pud_free_pmd_page
>
> panics on 7.3-rc4 as soon as the tracer starts:
>
> [ 23.531178] Starting tracer 'function'
> [ 23.675800] Oops: general protection fault, probably for non-canonical address 0xdffffc0000000038: 0000 [#1] SMP KASAN NOPTI
> [ 23.819917] KASAN: null-ptr-deref in range [0x00000000000001c0-0x00000000000001c7]
> [ 23.964025] CPU: 0 UID: 0 PID: 0 Comm: swapper Not tainted 7.3.0-rc4-fe2ec83746e5-with-fixes-v2+ #195 PREEMPT(undef)
> [ 24.252248] RIP: 0010:__queue_work+0xab/0xf00
> [ 25.981629] Call Trace:
> [ 26.125727] <TASK>
> [ 26.413912] ? pagetable_free_kernel+0x20/0x120
> [ 26.990283] queue_work_on+0x97/0xf0
> [ 27.134382] __cpa_collapse_large_pages+0x501/0x6f0
> [ 27.566662] cpa_flush+0x394/0x620
> [ 27.998953] change_page_attr_set_clr+0x321/0x4a0
> [ 29.151729] set_memory_rox+0xa2/0xf0
> [ 29.584018] create_trampoline+0x431/0x6f0
> ...
> [ 44.343347] Kernel panic - not syncing: Attempted to kill the idle task!
>
> The boot-time tracer is started from early_trace_init(), which runs
> before workqueue_init_early(). Making its trampoline read-only splits a
> large page, and CPA collapses it again right away. The split table has
> been a kernel page table since commit 9e4a3ec3411b
> ("x86/mm/pat: Allocate split page tables as kernel page tables"), so the
> collapse frees it through pagetable_free_kernel(), which queues work on
> system_percpu_wq - still NULL at that point. That commit is correct in
> itself; it only lets CPA reach pagetable_free_kernel() before the
> workqueue that function relies on exists.
>
> Keep putting the table on the list, but don't schedule the work while
> the system is still booting. The next kernel page table freed after
> boot schedules it, and the work then frees the early table too, after
> the same IOMMU flush as any other. If no kernel page table is freed
> after boot, the ones freed during boot stay on the list.
>
> Fixes: 9e4a3ec3411b ("x86/mm/pat: Allocate split page tables as kernel page tables")
> Suggested-by: David Hildenbrand (Arm) <david@kernel.org>
> Cc: stable@vger.kernel.org
> Signed-off-by: Mikhail Gavrilov <mikhail.v.gavrilov@gmail.com>
> Link: https://lore.kernel.org/20260924064321.23787-1-mikhail.v.gavrilov@gmail.com
> ---
> v2:
> - Keep the table on the list and only skip scheduling the work while
> booting, instead of freeing it directly (David Hildenbrand)
> - Say that 9e4a3ec3411b is correct in itself and only exposes the
> problem (Lorenzo Stoakes)
> v1: https://lore.kernel.org/20260924064321.23787-1-mikhail.v.gavrilov@gmail.com
>
> Tested on a Ryzen 9 7950X with a Radeon RX 7900 XTX, lockdep and KASAN
> enabled, on 7.3-rc4 (fe2ec83746e5) with the same unrelated local
> changes as noted for v1, booting with
>
> ftrace=function ftrace_filter=pud_free_pmd_page,pagetable_free_kernel,kernel_pgtable_work_func
>
> The boot that panicked without the fix completes. The table freed
> while the tracer installs itself does not show up in the trace, since
> the tracer is not live yet at that point, but the first kernel page
> table freed after boot does: systemd-modules-load freeing one from
> __cpa_collapse_large_pages() schedules the work, and
> kernel_pgtable_work_func() runs 0.8 ms later and drains the list. From
> then on every pagetable_free_kernel() in the trace (660 entries, none
> lost) is followed by a work run within a few milliseconds. So on this
> box the early tables wait until the first module is loaded, and no
> separate drain is needed.
>
> mm/pgtable-generic.c | 8 +++++++-
> 1 file changed, 7 insertions(+), 1 deletion(-)
>
> diff --git a/mm/pgtable-generic.c b/mm/pgtable-generic.c
> index b91b1a98029c..f7f504f57914 100644
> --- a/mm/pgtable-generic.c
> +++ b/mm/pgtable-generic.c
> @@ -444,6 +444,12 @@ void pagetable_free_kernel(struct ptdesc *pt)
> list_add(&pt->pt_list, &kernel_pgtable_work.list);
> spin_unlock(&kernel_pgtable_work.lock);
>
> - schedule_work(&kernel_pgtable_work.work);
> + /*
> + * The workqueue may not exist yet while the system is booting.
> + * The next kernel page table freed after boot schedules the work,
> + * which then frees this one as well.
> + */
> + if (system_state != SYSTEM_BOOTING)
> + schedule_work(&kernel_pgtable_work.work);
Maybe we'll just skip collapse on boot?
> }
> #endif
> --
> 2.55.0
>
--
Sincerely yours,
Mike.
next prev parent reply other threads:[~2026-09-24 16:19 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-24 9:23 Mikhail Gavrilov
2026-09-24 9:31 ` Lorenzo Stoakes (ARM)
2026-09-24 9:39 ` Mikhail Gavrilov
2026-09-24 16:19 ` Mike Rapoport [this message]
2026-09-24 17:57 ` Dave Hansen
2026-09-24 18:47 ` David Hildenbrand (Arm)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=arVNhTkI3kesfRs0@kernel.org \
--to=rppt@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=baolu.lu@linux.intel.com \
--cc=dave.hansen@linux.intel.com \
--cc=david@kernel.org \
--cc=jgg@nvidia.com \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=mikhail.v.gavrilov@gmail.com \
--cc=mingo@kernel.org \
--cc=regressions@lists.linux.dev \
--cc=rostedt@goodmis.org \
--cc=surenb@google.com \
--cc=vbabka@kernel.org \
--cc=vishal.moola@gmail.com \
--cc=x86@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®