From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pf1-f169.google.com (mail-pf1-f169.google.com [209.85.210.169]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B5A2B28469F for ; Fri, 20 Mar 2026 04:06:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.169 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1773979618; cv=none; b=pF0it0YYNuFJzNfpzhCMPsWRBp0JNzS8E2UiK1TYRup5Rw90//lVIi06rLMRqV0LAC1DBkhTPpTZr2Lz2X1MfkKxiNG03uDgMjJGYVCXAvJMzKdkUcRIi4x6/HLGwULWnkzTMTL5683CkC0mCv53iiYWMXOGvnqTauU4/dZ1id0= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1773979618; c=relaxed/simple; bh=RFioRWZ3eq8wc0ErrSkCtZOyQdTf4WqNRA36GtmIKuc=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=E0+puK/wAZ2/n9nlISwbz8cQ26orb3VqTy1ljaqyKxGbz8BJykpsJEFR9BXZOc1BJkJUN9wAsrfsj+nPg8KtDY+3XyQQHn4AwnrHkHloJl3PWM5V0JEoEJzUN1VvG8XDWKPh4dAlXzNn27JaJvPvcBRtPX/j8yDEWi2KerkwVe4= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=roeck-us.net; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=KLp6M1hg; arc=none smtp.client-ip=209.85.210.169 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=roeck-us.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="KLp6M1hg" Received: by mail-pf1-f169.google.com with SMTP id d2e1a72fcca58-82735a41920so701091b3a.2 for ; Thu, 19 Mar 2026 21:06:55 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1773979615; x=1774584415; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:sender:from:to:cc:subject:date:message-id :reply-to; bh=grz9eWJ1ps3EBDuJN2bzyVHKIzdPvtiDT74qagFiumU=; b=KLp6M1hgpC/l5WjXPjezZ4qApxvUirp0cLPpwNJ/cQpYkukeVX5qCeC1eg4tRNozWk JdfhJPEYERx043wwjblQhBfHJfpShpIlTIzhHUIva+ystiVt2yM9eGap4Cyjbu4C47Iy knNgDiUDxuVY10vpGeV+FT6GTsu/JSqNCDS27QNHReuahyla7IvY5e8UPYYaDR65XYEK GG1fnH0lyG8jaEffpKgqUClzDAcK0wO8DThEMnjQJ4LsTtvf2UAysRcnRc2E2O2/iyra 1QsMK+jG6guYv+T/iyr1mCndrBJGTHZOMkbKXf0ml23lnPqBF8VB5I53IZqab9gFrfss bNLw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1773979615; x=1774584415; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:sender:x-gm-gg:x-gm-message-state:from:to :cc:subject:date:message-id:reply-to; bh=grz9eWJ1ps3EBDuJN2bzyVHKIzdPvtiDT74qagFiumU=; b=NOXCt57u8V+BrSOPwCYJ7LEI2OEpXD6jwbTo1h6ZlDONdk1bmMBrzN3Us8xoSHsbJP WvTIZGH+Uph6J5mvHrdEhhIsRU4Ps58aFZEmH8kLc07owOQMApnYergbP2NnukJFPvZo hheCA7215XxqO+ZljacXzdNmnWddmaLqiX1fwhD2qBxEzYq+ZVBFxEz8PJJ61fr2DhKR VQu58QC05e/pQPXU8UepUcaZk+iZOClm3SxT7U4RCfnpSeDGXyeE35OQMo3pZCIeC5AP v1t7EWsI32MqqJShMCjjcrf+wln4E3kh6+R2v/kSg6jUIHqhk3BIHDUFLPXUNeMWGu3c n3Jw== X-Forwarded-Encrypted: i=1; AJvYcCVRh1P+52ijabmW8K1zbNKR8/QHrcvC0VRQqX3FkysKiLoXyNJJby1iYvRHkNfRcGfOuJae4gDwBePlovo=@vger.kernel.org X-Gm-Message-State: AOJu0YybmRwKN0Gkk43XJq9cF64c3IG51LwfVBu8Cj+Jeov00J5BDi+G zAPsla3SaJZemv/ZR6bCp+6hjeiJxLu485l+14AKK6jBVIGG8C2Ngz79 X-Gm-Gg: ATEYQzwmZ047hNBOpLyS+FhM7KyHa3wlnjHzYJpHV1KeY76wrEvpLJPrHnG3tsf3TS8 jouk4YulGoN0fdnyVTKKmiaecwXbJJbEshLdGWSLujtdN2byVFcaxDL7jUw2kf7T1ah85gEil86 nYfZtLB2G/+8gWouqLLbL8uyRYPKVQ2HlAmTRAaWlO6cluWICPXQoLTfV7/9NAG9LTwuiScMUhq wEQw1BPGsAhQqR/pU55fUVf3cZ4KTDJYQz/nhFd7eetM3MsJ9BNz2NjLb0gjRB4MYxchC+Zf4IZ uhMDH1xrBRiu0yATVqaYZ18W0eBqWkWJ15XpabyqCMD9PxNkJH5F7bdYwTccFjgO0pyNLFlrm8g +9UXhp06lQLV4lnZ4rAfZbEwJ7MvEwF726IB4HpMDLNTmM36SH4CjPyGBqeTzvqL42xsezr+RpZ QbzYrIHcqxv/xPE2zL4uFukm9RiUiKnnJI52mH X-Received: by 2002:a05:6a20:7f8d:b0:398:71e4:6282 with SMTP id adf61e73a8af0-39bce9b7e04mr1530953637.4.1773979614683; Thu, 19 Mar 2026 21:06:54 -0700 (PDT) Received: from server.roeck-us.net ([2600:1700:e321:62f0:da43:aeff:fecc:bfd5]) by smtp.gmail.com with ESMTPSA id 41be03b00d2f7-c74443ccb56sm650198a12.25.2026.03.19.21.06.53 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 19 Mar 2026 21:06:54 -0700 (PDT) Sender: Guenter Roeck Date: Thu, 19 Mar 2026 21:06:52 -0700 From: Guenter Roeck To: Mike Rapoport Cc: x86@kernel.org, linux-kernel@vger.kernel.org, Ard Biesheuvel , Benjamin Herrenschmidt , Borislav Petkov , Dave Hansen , Ilias Apalodimas , Ingo Molnar , "H. Peter Anvin" , Thomas Gleixner , linux-efi@vger.kernel.org, linux-mm@kvack.org, stable@vger.kernel.org Subject: Re: [PATCH v2] x86/efi: defer freeing of boot services memory Message-ID: <100b9ae1-74cc-48b3-ba63-1a72cfa2ebbd@roeck-us.net> References: <20260225065555.2471844-1-rppt@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260225065555.2471844-1-rppt@kernel.org> Hi, On Wed, Feb 25, 2026 at 08:55:55AM +0200, Mike Rapoport wrote: > From: "Mike Rapoport (Microsoft)" > > efi_free_boot_services() frees memory occupied by EFI_BOOT_SERVICES_CODE > and EFI_BOOT_SERVICES_DATA using memblock_free_late(). > > There are two issue with that: memblock_free_late() should be used for > memory allocated with memblock_alloc() while the memory reserved with > memblock_reserve() should be freed with free_reserved_area(). > > More acutely, with CONFIG_DEFERRED_STRUCT_PAGE_INIT=y > efi_free_boot_services() is called before deferred initialization of the > memory map is complete. > > Benjamin Herrenschmidt reports that this causes a leak of ~140MB of > RAM on EC2 t3a.nano instances which only have 512MB or RAM. > > If the freed memory resides in the areas that memory map for them is > still uninitialized, they won't be actually freed because > memblock_free_late() calls memblock_free_pages() and the latter skips > uninitialized pages. > > Using free_reserved_area() at this point is also problematic because > __free_page() accesses the buddy of the freed page and that again might > end up in uninitialized part of the memory map. > > Delaying the entire efi_free_boot_services() could be problematic > because in addition to freeing boot services memory it updates > efi.memmap without any synchronization and that's undesirable late in > boot when there is concurrency. > > More robust approach is to only defer freeing of the EFI boot services > memory. > > Split efi_free_boot_services() in two. First efi_unmap_boot_services() > collects ranges that should be freed into an array then > efi_free_boot_services() later frees them after deferred init is complete. > > Link: https://lore.kernel.org/all/ec2aaef14783869b3be6e3c253b2dcbf67dbc12a.camel@kernel.crashing.org > Fixes: 916f676f8dc0 ("x86, efi: Retain boot service code until after switching to virtual mode") > Cc: > Signed-off-by: Mike Rapoport (Microsoft) > Reviewed-by: Benjamin Herrenschmidt > --- > > v1: https://lore.kernel.org/all/20260223075219.2348035-1-rppt@kernel.org > * update the commit message with correct function names (Ben) > > arch/x86/include/asm/efi.h | 2 +- > arch/x86/platform/efi/efi.c | 2 +- > arch/x86/platform/efi/quirks.c | 55 +++++++++++++++++++++++++++-- > drivers/firmware/efi/mokvar-table.c | 2 +- > 4 files changed, 55 insertions(+), 6 deletions(-) > > diff --git a/arch/x86/include/asm/efi.h b/arch/x86/include/asm/efi.h > index f227a70ac91f..51b4cdbea061 100644 > --- a/arch/x86/include/asm/efi.h > +++ b/arch/x86/include/asm/efi.h > @@ -138,7 +138,7 @@ extern void __init efi_apply_memmap_quirks(void); > extern int __init efi_reuse_config(u64 tables, int nr_tables); > extern void efi_delete_dummy_variable(void); > extern void efi_crash_gracefully_on_page_fault(unsigned long phys_addr); > -extern void efi_free_boot_services(void); > +extern void efi_unmap_boot_services(void); > > void arch_efi_call_virt_setup(void); > void arch_efi_call_virt_teardown(void); > diff --git a/arch/x86/platform/efi/efi.c b/arch/x86/platform/efi/efi.c > index d00c6de7f3b7..d84c6020dda1 100644 > --- a/arch/x86/platform/efi/efi.c > +++ b/arch/x86/platform/efi/efi.c > @@ -836,7 +836,7 @@ static void __init __efi_enter_virtual_mode(void) > } > > efi_check_for_embedded_firmwares(); > - efi_free_boot_services(); > + efi_unmap_boot_services(); > > if (!efi_is_mixed()) > efi_native_runtime_setup(); > diff --git a/arch/x86/platform/efi/quirks.c b/arch/x86/platform/efi/quirks.c > index 553f330198f2..35caa5746115 100644 > --- a/arch/x86/platform/efi/quirks.c > +++ b/arch/x86/platform/efi/quirks.c > @@ -341,7 +341,7 @@ void __init efi_reserve_boot_services(void) > > /* > * Because the following memblock_reserve() is paired > - * with memblock_free_late() for this region in > + * with free_reserved_area() for this region in > * efi_free_boot_services(), we must be extremely > * careful not to reserve, and subsequently free, > * critical regions of memory (like the kernel image) or > @@ -404,17 +404,33 @@ static void __init efi_unmap_pages(efi_memory_desc_t *md) > pr_err("Failed to unmap VA mapping for 0x%llx\n", va); > } > > -void __init efi_free_boot_services(void) > +struct efi_freeable_range { > + u64 start; > + u64 end; > +}; > + > +static struct efi_freeable_range *ranges_to_free; > + > +void __init efi_unmap_boot_services(void) > { > struct efi_memory_map_data data = { 0 }; > efi_memory_desc_t *md; > int num_entries = 0; > + int idx = 0; > + size_t sz; > void *new, *new_md; > > /* Keep all regions for /sys/kernel/debug/efi */ > if (efi_enabled(EFI_DBG)) > return; > > + sz = sizeof(*ranges_to_free) * efi.memmap.nr_map + 1; Was this possibly supposed to be sz = sizeof(*ranges_to_free) * (efi.memmap.nr_map + 1); ^ ^ ? Thanks, Guenter > + ranges_to_free = kzalloc(sz, GFP_KERNEL); > + if (!ranges_to_free) { > + pr_err("Failed to allocate storage for freeable EFI regions\n"); > + return; > + } > + > for_each_efi_memory_desc(md) { > unsigned long long start = md->phys_addr; > unsigned long long size = md->num_pages << EFI_PAGE_SHIFT; > @@ -471,7 +487,15 @@ void __init efi_free_boot_services(void) > start = SZ_1M; > } > > - memblock_free_late(start, size); > + /* > + * With CONFIG_DEFERRED_STRUCT_PAGE_INIT parts of the memory > + * map are still not initialized and we can't reliably free > + * memory here. > + * Queue the ranges to free at a later point. > + */ > + ranges_to_free[idx].start = start; > + ranges_to_free[idx].end = start + size; > + idx++; > } > > if (!num_entries) > @@ -512,6 +536,31 @@ void __init efi_free_boot_services(void) > } > } > > +static int __init efi_free_boot_services(void) > +{ > + struct efi_freeable_range *range = ranges_to_free; > + unsigned long freed = 0; > + > + if (!ranges_to_free) > + return 0; > + > + while (range->start) { > + void *start = phys_to_virt(range->start); > + void *end = phys_to_virt(range->end); > + > + free_reserved_area(start, end, -1, NULL); > + freed += (end - start); > + range++; > + } > + kfree(ranges_to_free); > + > + if (freed) > + pr_info("Freeing EFI boot services memory: %ldK\n", freed / SZ_1K); > + > + return 0; > +} > +arch_initcall(efi_free_boot_services); > + > /* > * A number of config table entries get remapped to virtual addresses > * after entering EFI virtual mode. However, the kexec kernel requires > diff --git a/drivers/firmware/efi/mokvar-table.c b/drivers/firmware/efi/mokvar-table.c > index 4ff0c2926097..6842aa96d704 100644 > --- a/drivers/firmware/efi/mokvar-table.c > +++ b/drivers/firmware/efi/mokvar-table.c > @@ -85,7 +85,7 @@ static struct kobject *mokvar_kobj; > * as an alternative to ordinary EFI variables, due to platform-dependent > * limitations. The memory occupied by this table is marked as reserved. > * > - * This routine must be called before efi_free_boot_services() in order > + * This routine must be called before efi_unmap_boot_services() in order > * to guarantee that it can mark the table as reserved. > * > * Implicit inputs: > > base-commit: 6de23f81a5e08be8fbf5e8d7e9febc72a5b5f27f > -- > 2.51.