From: Muchun Song <muchun.song@linux.dev>
To: "David Hildenbrand (Arm)" <david@kernel.org>
Cc: Lance Yang <lance.yang@linux.dev>,
osalvador@suse.de, akpm@linux-foundation.org, linux-mm@kvack.org,
linux-cxl@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH 1/1] mm/memory_hotplug: fix missing rollback in __add_pages()
Date: Thu, 1 Oct 2026 10:33:16 +0800 [thread overview]
Message-ID: <203892F4-B04A-4F69-A1B3-DC1619176C67@linux.dev> (raw)
In-Reply-To: <d4fac8af-fd71-47a2-bfe9-3c6559b92209@kernel.org>
> On Oct 1, 2026, at 00:46, David Hildenbrand (Arm) <david@kernel.org> wrote:
>
> On 9/30/26 18:04, Lance Yang wrote:
>> __add_pages() returns on a sparse_add_section() failure without removing
>> the sections already added in the same request.
>>
>> For memremap_pages(), the failed range is not counted in pgmap->nr_range,
>> so memunmap_pages() skips it. The sections already added in that range
>> retain their vmemmap mappings and subsection bits. Retrying a
>> section-aligned range can then fail with -EEXIST.
>>
>> Save the initial PFN and remove [start_pfn, pfn) on failure. For a vmemmap
>> population failure, section_activate() already cleans up the current
>> section, so the rollback excludes it. If the first section fails,
>> __remove_pages() receives an empty range and does nothing.
>>
>> Link: https://lore.kernel.org/all/BAD58999-1EDD-4A37-ABA0-DB1BD8AB3453@linux.dev/
>> Suggested-by: Muchun Song <muchun.song@linux.dev>
>> Signed-off-by: Lance Yang <lance.yang@linux.dev>
>> ---
>> No Fixes tag, as I couldn't identify the commit that introduced this issue.
>>
>> mm/memory_hotplug.c | 6 +++++-
>> 1 file changed, 5 insertions(+), 1 deletion(-)
>>
>> diff --git a/mm/memory_hotplug.c b/mm/memory_hotplug.c
>> index 796af1028ee2..ca4656698148 100644
>> --- a/mm/memory_hotplug.c
>> +++ b/mm/memory_hotplug.c
>> @@ -380,6 +380,7 @@ EXPORT_SYMBOL_GPL(pfn_to_online_page);
>> int __add_pages(int nid, unsigned long pfn, unsigned long nr_pages,
>> struct mhp_params *params)
>> {
>> + const unsigned long start_pfn = pfn;
>> const unsigned long end_pfn = pfn + nr_pages;
>> unsigned long cur_nr_pages;
>> int err;
>> @@ -413,8 +414,11 @@ int __add_pages(int nid, unsigned long pfn, unsigned long nr_pages,
>> SECTION_ALIGN_UP(pfn + 1) - pfn);
>> err = sparse_add_section(nid, pfn, cur_nr_pages, altmap,
>> params->pgmap);
>> - if (err)
>> + if (err) {
>> + __remove_pages(start_pfn, pfn - start_pfn, altmap,
>> + params->pgmap);
>> break;
>> + }
>> cond_resched();
>> }
>> vmemmap_populate_print_last();
>
> Makes sense and LGTM.
>
> Do we have a Fixes: tag? It probably dates back quite a while ... not sure about
> stable, we never saw this in practice. But if it's easy, we should just do it?
> (not sure if we ever had __remove_pages be limited to hotunplug support)
>
> I'm planning on picking this up and sending it for the next merge window (so not
> as a hotfix).
>
> Looking at this ...
>
> x86 does not really expect add_pages to fail:
>
> ret = __add_pages(nid, start_pfn, nr_pages, params);
> WARN_ON_ONCE(ret);
>
> That's probably something to clean up as well?
The warning is a historical leftover. The original code
printed an error when __add_pages() failed. Commit
10f22dde556d accidentally turned that conditional printk
into an unconditional WARN_ON(1), and commit fe8b868eccb9
subsequently changed it to WARN_ON_ONCE(ret) to avoid
warning on successful memory hot-add.
There is no no-failure contract here: __add_pages() can
legitimately return errors such as -ENOMEM, and those
errors are already propagated to the caller. Since
WARN_ON_ONCE(ret) has no effect on control flow or error
handling, it can be safely removed without changing the
failure semantics.
This also reveals a potential bug introduced by commit
ea0854170c952: when update_end_of_memory_vars() was
added, it was called without checking that ret == 0, so the
end-of-memory variables may be updated even when
__add_pages() fails. Returning immediately on error fixes
that as well.
Thanks,
Muchun
>
> --
> Cheers,
>
> David
next prev parent reply other threads:[~2026-10-01 2:33 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-30 16:04 Lance Yang
2026-09-30 16:46 ` David Hildenbrand (Arm)
2026-10-01 2:33 ` Muchun Song [this message]
2026-10-01 7:00 ` David Hildenbrand (Arm)
2026-10-01 5:57 ` Lance Yang
2026-10-01 2:05 ` Muchun Song
2026-10-01 6:26 ` Lance Yang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=203892F4-B04A-4F69-A1B3-DC1619176C67@linux.dev \
--to=muchun.song@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=david@kernel.org \
--cc=lance.yang@linux.dev \
--cc=linux-cxl@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=osalvador@suse.de \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®