mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Breno Leitao <leitao@debian.org>
To: Kiryl Shutsemau <kas@kernel.org>
Cc: Ard Biesheuvel <ardb@kernel.org>,
	 Ilias Apalodimas <ilias.apalodimas@linaro.org>,
	Miaohe Lin <linmiaohe@huawei.com>,
	 Naoya Horiguchi <nao.horiguchi@gmail.com>,
	Andrew Morton <akpm@linux-foundation.org>,
	 linux-efi@vger.kernel.org, linux-kernel@vger.kernel.org,
	kexec@lists.infradead.org,  rneu@meta.com, riel@surriel.com,
	caggio@meta.com, anilagrawal@meta.com,  rmikey@meta.com,
	linux-mm@kvack.org, kernel-team@meta.com
Subject: Re: [PATCH RFC 0/3] efi: mm/memory-failure: keep hardware-poisoned pages out of the next kexec
Date: Mon, 27 Jul 2026 06:22:14 -0700	[thread overview]
Message-ID: <amdVuOO_6aCxMkA2@gmail.com> (raw)
In-Reply-To: <amDHmNWfQ8eid9jH@thinkstation>

On Wed, Jul 22, 2026 at 02:50:18PM +0100, Kiryl Shutsemau wrote:
> On Fri, Jul 17, 2026 at 07:03:02AM -0700, Breno Leitao wrote:
> > Problem:
> > ========
> > 
> > When a page is hard-offlined due to an uncorrectable memory error (multi
> > bit ECC), memory_failure() sets PG_hwpoison, unmaps it, removes it from
> > the buddy allocator. This information is not carried to the next kernel
> > that is kexeced. The new kernel kexecs and trip over that bad memory
> > bank _again_.
> > 
> > Why now:
> > ========
> > 
> > Several industry trends make this increasingly important:
> > 
> >     1) DRAM is getting more expensive
> >     2) soldered / on-package memory (LPDDR, HBM) is becoming more common, so a
> >        failing part can no longer simply be swapped;
> >     3) memory is kept in service far longer (at Meta, DRAM lifetime is being
> >        drastically extended).
> >     4) It is more and more common to kexec instead of full reboot
> >     5) Increase of memory per system with CXL
> > 
> > Proposed Solution:
> > ==================
> > 
> > Carry the poisoned frames to the next kernel in a new EFI configuration
> > table, LINUX_EFI_POISONED_MEMORY, modeled on the existing
> > LINUX_EFI_MEMRESERVE table.
> > 
> > EFI configuration tables already survive kexec: firmware hands the EFI
> > system table to every kernel in the chain, so a table installed once is
> > seen by all successors without a new handover channel.
> > 
> > The mechanism is architecture independent, so x86 and arm64 use the same
> > code.
> > 
> > The EFI stub installs an empty table while EFI boot service is up. A
> > configuration table can only be created there; the running kernel can only
> > append to it.
> > 
> > Each hard-offlined frame is appended; an unpoison "removes" its entry
> > so a frame that is good again is not carried forward.
> > 
> > The next kernel walks the inherited table early in
> > efi_config_parse_tables(), before memblock and the buddy allocator are
> > up, and memblock_reserve()s every recorded frame. The bad RAM is never
> > handed out.
> 
> The allocator is only half of the problem. kexec segment placement
> doesn't know about any of this: kexec_file picks destinations from
> System RAM resources (memblock on arm64), and poisoned frames are only
> ever removed from buddy. So the next kernel image, initrd or purgatory
> can be placed on top of a poisoned frame -- the relocation memcpy then
> consumes the poison and the machine checks during the very kexec this
> series is supposed to protect. Same for the table's own list pages:
> overwrite one at placement time and the next kernel parses garbage.


Agreed, and thanks. This is real and, I think, largely separable from
the cross-kexec table -- the running kernel already knows its own
poisoned frames via PG_hwpoison, so excluding them at load time protects
the immediate kexec without depending on anything carried across the
boot. I'd like to tackle it as an independent change, let me do a PoC
and see how much change it is. I suspect it is less than this one (and
probably nice to have even if this series doesn't end up anywhere)

> Hooking num_poisoned_pages_inc() also means soft-offlined pages are
> recorded. Those are functional pages, offlined predictively. Turning
> them into permanent losses for every kernel down the kexec chain does
> not seem right. I would record hard failures only.

Yea, it seems that action_result() is what we want instead of
num_poisoned_pages_inc()

> On the data structure: I am not an expert on DRAM failure modes, so I
> went reading. The field studies [1] say roughly a quarter of DRAM
> faults are multi-bit structures (row/column/bank), and the address
> interleaving means one such fault shows up to the OS as many separate
> 4K pages. A row fault lands in a window of tens of KB to about a MB; a
> column fault is one bad line per row, strided across the bank's entire
> footprint -- potentially thousands of pages scattered over gigabytes,
> reported one MCE at a time as they get touched. DDR5 on-die ECC hides
> most single-bit faults from the host, so the visible mix shifts toward
> these multi-bit modes over time.
> 
> If that is accurate, per-4K entries have no natural bound, and every
> entry becomes a separate scattered memblock_reserve() in every future
> boot.
> 
> I would consider a bitmap with one bit per 2M instead, modeled on
> struct efi_unaccepted_memory: the stub sizes it from the EFI memory map
> at cold boot, the running kernel sets a bit on hard offline, clears it
> on unpoison, and the next kernel reserves set units. It is one
> fixed-size allocation, so the whole grow-and-link machinery and the
> chain-parsing trust problem go away, and a row fault collapses into one
> or two bits.

Thanks for digging into the failure modes -- that matches my
understanding, and the bounded, chain-free bitmap is appealing for
exactly the reasons you give: a row fault collapsing to a bit or two,
one fixed allocation, and no cross-kernel chain to trust.

Let me hack a PoC with a bitmap and check how it looks like, let's see
if that looks like smoother.

--breno

      reply	other threads:[~2026-07-27 13:22 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-17 14:03 Breno Leitao
2026-07-17 14:03 ` [PATCH RFC 1/3] efi: add the LINUX_EFI_POISONED_MEMORY configuration table Breno Leitao
2026-07-17 14:03 ` [PATCH RFC 2/3] efi: record hardware-poisoned frames into the poisoned-memory table Breno Leitao
2026-07-17 14:03 ` [PATCH RFC 3/3] efi: reserve inherited poisoned frames before the allocator comes up Breno Leitao
2026-07-22 13:50 ` [PATCH RFC 0/3] efi: mm/memory-failure: keep hardware-poisoned pages out of the next kexec Kiryl Shutsemau
2026-07-27 13:22   ` Breno Leitao [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=amdVuOO_6aCxMkA2@gmail.com \
    --to=leitao@debian.org \
    --cc=akpm@linux-foundation.org \
    --cc=anilagrawal@meta.com \
    --cc=ardb@kernel.org \
    --cc=caggio@meta.com \
    --cc=ilias.apalodimas@linaro.org \
    --cc=kas@kernel.org \
    --cc=kernel-team@meta.com \
    --cc=kexec@lists.infradead.org \
    --cc=linmiaohe@huawei.com \
    --cc=linux-efi@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=nao.horiguchi@gmail.com \
    --cc=riel@surriel.com \
    --cc=rmikey@meta.com \
    --cc=rneu@meta.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®