mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Pratyush Yadav <pratyush@kernel.org>
To: Sourabh Jain <sourabhjain@linux.ibm.com>
Cc: George Guo <dongtai.guo@linux.dev>,
	 rppt@kernel.org, pasha.tatashin@soleen.com,
	 pratyush@kernel.org,  graf@amazon.com, changyuanl@google.com,
	 akpm@linux-foundation.org, chenhuacai@kernel.org,
	 liukexin@kylinos.cn,  guodongtai@kylinos.cn,
	kexec@lists.infradead.org,  linux-mm@kvack.org,
	loongarch@lists.linux.dev,  linux-kernel@vger.kernel.org
Subject: Re: [PATCH 1/1] liveupdate: kho: calculate per-node scratch sizes before allocation
Date: Fri, 18 Sep 2026 00:53:27 +0200	[thread overview]
Message-ID: <2vxz1par7amw.fsf@kernel.org> (raw)
In-Reply-To: <0969ede4-f617-4e24-b32e-c5e30a96880e@linux.ibm.com> (Sourabh Jain's message of "Thu, 17 Sep 2026 16:00:00 +0530")

On Thu, Sep 17 2026, Sourabh Jain wrote:

> Hello George,
>
> On 12/09/26 12:03, Sourabh Jain wrote:
>> Hello George,
>>
>> On 04/09/26 08:21, George Guo wrote:
>>> From: George Guo <guodongtai@kylinos.cn>
>>>
>>> The default percentage-based policy sizes scratch areas from the current
>>> kernel's MEMBLOCK_RSRV_KERN footprint. This is a reasonable heuristic for
>>> predicting the early memory demand of the next kernel.
>>>
>>> scratch_size_update() calculates the lowmem and global sizes before either
>>> area is allocated. However, kho_reserve_scratch() calculates each per-node
>>> size only after allocating the lowmem and global areas. Since memblock
>>> allocations are marked MEMBLOCK_RSRV_KERN, the per-node calculation
>>> includes those newly allocated scratch areas and scales them again.
>>
>> I may be missing something here, but doesn't memblock_reserved_kern_size()
>> check the node ID of the region before checking the region flag
>> (MEMBLOCK_RSRV_KERN)?
>>
>> Code snippet from memblock_reserved_kern_size()
>> ```
>> if (nid == memblock_get_region_node(r) || !numa_valid_node(nid))
>>     if (r->flags & MEMBLOCK_RSRV_KERN)
>>         total += size;
>> ```
>>
>> For a valid nid, my understanding is that the global and lowmem scratch
>> areas should not be counted because they are allocated with NUMA_NO_NODE
>> (-1). So, ideally, these regions should be excluded when calculating the
>> reserved memory for a specific node ID.
>>
>> Based on this, I am not sure that marking the lowmem and global areas as
>> MEMBLOCK_RSRV_KERN is what causes the per-node size calculation to be
>> inflated. I am looking into the code further to better understand the
>> actual cause of the issue that this patch is trying to address.
>
> I added some prints in kho_reserve_scratch() and found that the per-node size
> calculation is not impacted by the lowmem and global scratch memory allocations.
>
> KHO: Before low and global scratch allocations
> KHO: low size = 899 KB
> KHO: global size = 137 MB
> KHO: Per node 2 = 80 MB
>
> KHO: After low and global scratch allocations
> KHO: low size = 312195 KB
> KHO: global size = 441 MB
> KHO: Per node 2 = 80 MB
>
> KHO: After per node allocation
> KHO: low size = 394115 KB
> KHO: global size = 521 MB
> KHO: Per node 2 = 240 MB
>
> I only had one NUMA node (nid=2), and the per-NUMA
> allocation before and after the lowmem and global scratch
> memory allocations remained the same at 80 MB.
>
> The experiment was done on the PowerPC architecture.
>
> I am wondering how the per-NUMA allocation in your setup is
> getting inflated due to the lowmem and global scratch memory
> reservations.

I had the same question, so I asked AI. Here's what it says:

--- 8< ---

The bug was reported and tested using the KHO self-test runner
(tools/testing/selftests/kho/vmtest.sh), which builds a test kernel
using make olddefconfig with only a minimal set of CONFIG_* options.
Crucially, CONFIG_NUMA is not enabled.

When CONFIG_NUMA is disabled:

    #ifndef CONFIG_NUMA
    static inline void memblock_set_region_node(struct memblock_region *r, int nid)
    {
    }
    
    static inline int memblock_get_region_node(const struct memblock_region *r)
    {
        return 0;
    }
    #endif


struct memblock_region does not even contain an nid member.

memblock_set_region_node() is a no-op (the NUMA_NO_NODE argument is
simply discarded), and memblock_get_region_node() is hardcoded to always
return 0.

There is only one node (nid = 0), so for_each_node_state(nid, N_MEMORY)
loops once for nid = 0.

Both memblock_phys_alloc_range() and memblock_phys_alloc() mark their
allocations with MEMBLOCK_RSRV_KERN. Therefore, when
scratch_size_node(0) runs after allocating the lowmem scratch buffer,
memblock_get_region_node(r) returns 0 for that lowmem scratch buffer,
and r->flags & MEMBLOCK_RSRV_KERN is true. As a result,
scratch_size_node(0) counts the lowmem scratch area as part of Node 0's
kernel footprint and scales it by scratch_scale (200%) again.

--- >8 ---

I didn't look closer, but it does seem to make sense.

But in practice, this problem is only on CONFIG_NUMA=n and I don't think
in practice KHO or LUO is being used in non-NUMA systems. So while I
think it is worth fixing, I think we should also have a test where we
enable CONFIG_NUMA.

Your system probably has CONFIG_NUMA=y and that's why you aren't able to
reproduce this bug.

I think on NUMA systems the problem is the other way round. The
calculation for the global scratch also counts per-node allocations.

So I think the proper fix for scratch sizing is what this patch does and
then a fixup for the global scratch calculation as well.

[...]

-- 
Regards,
Pratyush Yadav

  reply	other threads:[~2026-09-17 22:53 UTC|newest]

Thread overview: 8+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-04  2:51 George Guo
2026-09-06 20:07 ` Mike Rapoport
2026-09-07 10:24   ` George Guo
2026-09-08  3:05     ` Sourabh Jain
2026-09-12  6:33 ` Sourabh Jain
2026-09-17 10:30   ` Sourabh Jain
2026-09-17 22:53     ` Pratyush Yadav [this message]
2026-09-18  9:33       ` George Guo

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=2vxz1par7amw.fsf@kernel.org \
    --to=pratyush@kernel.org \
    --cc=akpm@linux-foundation.org \
    --cc=changyuanl@google.com \
    --cc=chenhuacai@kernel.org \
    --cc=dongtai.guo@linux.dev \
    --cc=graf@amazon.com \
    --cc=guodongtai@kylinos.cn \
    --cc=kexec@lists.infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=liukexin@kylinos.cn \
    --cc=loongarch@lists.linux.dev \
    --cc=pasha.tatashin@soleen.com \
    --cc=rppt@kernel.org \
    --cc=sourabhjain@linux.ibm.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®