mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Sourabh Jain <sourabhjain@linux.ibm.com>
To: Mike Rapoport <rppt@kernel.org>
Cc: kexec@lists.infradead.org, Alexander Graf <graf@amazon.com>,
	Andrew Morton <akpm@linux-foundation.org>,
	George Guo <guodongtai@kylinos.cn>,
	Pasha Tatashin <pasha.tatashin@soleen.com>,
	Pratyush Yadav <pratyush@kernel.org>,
	"Ritesh Harjani (IBM)" <ritesh.list@gmail.com>,
	linux-kernel@vger.kernel.org, linux-mm@kvack.org
Subject: Re: [PATCH] kho: fix global scratch size calculation
Date: Sat, 26 Sep 2026 20:16:33 +0530	[thread overview]
Message-ID: <e265fdc8-e224-4815-b94e-9c974210f099@linux.ibm.com> (raw)
In-Reply-To: <areLMLTC-bO1zcXP@kernel.org>

Hello Mike,

On 26/09/26 14:36, Mike Rapoport wrote:
> Hi Sourabh,
>
> On Tue, Sep 22, 2026 at 06:42:16PM +0530, Sourabh Jain wrote:
>> KHO calculates the global scratch size based on memblock-reserved kernel
>> memory. It passes NUMA_NO_NODE to memblock_reserved_kern_size() for
>> this calculation.
>>
>> When memblock_reserved_kern_size() is called with NUMA_NO_NODE, it
>> counts both:
>>
>> - memory reserved for a specific NUMA node
>> - memory reserved with NUMA_NO_NODE
>>
>> KHO needs to distinguish between these two types of reservations.
>> When calculating the size of global scratch memory, KHO only needs to
>> account for reservations made with NUMA_NO_NODE. Reservations made for
>> a specific NUMA node must not be included in the global scratch size.
>>
>> Add memblock_reserved_size_nid() to calculate reserved memory for a
>> given reservation type and NUMA node. When NUMA_NO_NODE is passed, it
>> counts only memory reserved with NUMA_NO_NODE.
>>
>> Use the new API for lowmem, global, and per-node KHO scratch size
>> calculations. For lowmem and global scratch, count only memory
>> reservations that were made with NUMA_NO_NODE. For per-node scratch,
>> count only memory reservations that were made with the corresponding
>> NUMA node ID.
>>
>> Remove memblock_reserved_hugetlb_size() since it has the same
>> implementation as the new API and differs only in the memblock
>> reservation flag being checked. The new API handles both kernel and
>> HugeTLB reservations through its reservation type argument.
>>
>> Define the new helper as a static function in the KHO implementation,
>> since it is only used by KHO and has no users outside
>> kernel/liveupdate/kexec_handover.c.
>>
>> On powerpc, the difference can be seen in the scratch_len values
>> reported by:
>>
>> cat /sys/kernel/debug/kho/out/scratch_len
>>
>> Before this change, the global scratch allocation was 0x12000000
>> (288 MB):
>>
>> 0x1000000   (16 MB)
>> 0x12000000  (288 MB)  <- global allocation
>> 0x5000000   (80 MB)
>>
>> After this change, the global scratch allocation is 0xd000000
>> (208 MB):
>>
>> 0x1000000   (16 MB)
>> 0xd000000   (208 MB)  <- global allocation
>> 0x5000000   (80 MB)
>>
>> The 80 MB difference is the per-node reservation that was previously
>> being included in the global allocation.
>>
>> The same issue also affects lowmem scratch memory, but its impact is
>> limited because the lowmem scratch memory calculation is restricted to
>> the first 4G of memory. The changes also cover the lowmem scratch
>> memory case.
>>
>> Cc: Alexander Graf <graf@amazon.com>
>> Cc: Andrew Morton <akpm@linux-foundation.org>
>> Cc: George Guo <guodongtai@kylinos.cn>
>> Cc: Mike Rapoport <rppt@kernel.org>
>> Cc: Pasha Tatashin <pasha.tatashin@soleen.com>
>> Cc: Pratyush Yadav <pratyush@kernel.org>
>> Cc: Ritesh Harjani (IBM) <ritesh.list@gmail.com>
>> Cc: linux-kernel@vger.kernel.org
>> Cc: linux-mm@kvack.org
>> Signed-off-by: Sourabh Jain <sourabhjain@linux.ibm.com>
>> ---
>>   include/linux/memblock.h           |  1 -
>>   kernel/liveupdate/kexec_handover.c | 49 ++++++++++++++++++++++--------
>>   mm/memblock.c                      | 22 --------------
>>   3 files changed, 37 insertions(+), 35 deletions(-)
>>
>> diff --git a/include/linux/memblock.h b/include/linux/memblock.h
>> index d62db9e776cf..678fe466529a 100644
>> --- a/include/linux/memblock.h
>> +++ b/include/linux/memblock.h
>> @@ -487,7 +487,6 @@ static inline __init_memblock bool memblock_bottom_up(void)
>>   phys_addr_t memblock_phys_mem_size(void);
>>   phys_addr_t memblock_reserved_size(void);
>>   phys_addr_t memblock_reserved_kern_size(phys_addr_t limit, int nid);
>> -phys_addr_t memblock_reserved_hugetlb_size(phys_addr_t limit, int nid);
>>   unsigned long memblock_estimated_nr_free_pages(void);
>>   phys_addr_t memblock_start_of_DRAM(void);
>>   phys_addr_t memblock_end_of_DRAM(void);
>> diff --git a/kernel/liveupdate/kexec_handover.c b/kernel/liveupdate/kexec_handover.c
>> index 7c4d86daf86d..dc809e1e768c 100644
>> --- a/kernel/liveupdate/kexec_handover.c
>> +++ b/kernel/liveupdate/kexec_handover.c
>> @@ -752,6 +752,31 @@ static int __init kho_parse_scratch_size(char *p)
>>   }
>>   early_param("kho_scratch", kho_parse_scratch_size);
>>   
>> +static phys_addr_t __init_memblock memblock_reserved_size_nid(phys_addr_t limit, int nid,
>> +							      enum memblock_flags region_type)
>> +{
> Why is this in kexec_handover.c and not in memblock.c?

I kept it here because KHO is the only user of this. However, I'm happy 
to move it to
memblock.c if you think that would be more appropriate.

>> +	struct memblock_region *r;
>> +	phys_addr_t total = 0;
>> +
>> +	for_each_reserved_mem_region(r) {
>> +		phys_addr_t size = r->size;
>> +
>> +		if (r->base > limit)
>> +			break;
>> +
>> +		if (r->base + r->size > limit)
>> +			size = limit - r->base;
>> +
>> +#ifdef CONFIG_NUMA
>> +		if (nid == memblock_get_region_node(r))
>> +#endif
> Do we really need the #ifdef here?

Yes, otherwise, if CONFIG_NUMA is not set, the lowmem and global sizes 
would be 0.


- Sourabh Jain

>
>> +			if (r->flags & region_type)
>> +				total += size;
>> +	}
>> +
>> +	return total;
>> +}
>> +
>>   static void __init scratch_size_update(void)
>>   {
>>   	/*


      reply	other threads:[~2026-09-26 14:47 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-22 13:12 Sourabh Jain
2026-09-22 13:27 ` sashiko-bot
2026-09-23  4:34   ` Sourabh Jain
2026-09-26  9:06 ` Mike Rapoport
2026-09-26 14:46   ` Sourabh Jain [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=e265fdc8-e224-4815-b94e-9c974210f099@linux.ibm.com \
    --to=sourabhjain@linux.ibm.com \
    --cc=akpm@linux-foundation.org \
    --cc=graf@amazon.com \
    --cc=guodongtai@kylinos.cn \
    --cc=kexec@lists.infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=pasha.tatashin@soleen.com \
    --cc=pratyush@kernel.org \
    --cc=ritesh.list@gmail.com \
    --cc=rppt@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®