mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Gregory Price <gourry@gourry.net>
To: Arun George/Arun George <arun.george@samsung.com>
Cc: linux-mm@kvack.org, Zhigang.Luo@amd.com, balbirs@nvidia.com,
	 brendan.jackman@linux.dev, yuzenghui@huawei.com,
	apopple@nvidia.com, alucerop@amd.com,  matthew.brost@intel.com,
	akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org,
	 liam@infradead.org, vbabka@kernel.org, rppt@kernel.org,
	surenb@google.com,  mhocko@suse.com, corbet@lwn.net,
	skhan@linuxfoundation.org,  gregkh@linuxfoundation.org,
	rafael@kernel.org, dakr@kernel.org, djbw@kernel.org,
	 vishal.l.verma@intel.com, dave.jiang@intel.com,
	alison.schofield@intel.com,  osandov@osandov.com,
	jannh@google.com, pfalcato@suse.de, jackmanb@google.com,
	 hannes@cmpxchg.org, ziy@nvidia.com, pbonzini@redhat.com,
	osalvador@suse.de,  joshua.hahnjy@gmail.com, rakie.kim@sk.com,
	byungchul@sk.com, ying.huang@linux.alibaba.com,
	 kasong@tencent.com, qi.zheng@linux.dev, shakeel.butt@linux.dev,
	baohua@kernel.org,  axelrasmussen@google.com, yuanchu@google.com,
	weixugc@google.com, yury.norov@gmail.com,
	 linux@rasmusvillemoes.dk, longman@redhat.com,
	ridong.chen@linux.dev, tj@kernel.org,  mkoutny@suse.com,
	sj@kernel.org, jgg@ziepe.ca, jhubbard@nvidia.com,
	 peterx@redhat.com, baolin.wang@linux.alibaba.com,
	npache@redhat.com,  ryan.roberts@arm.com, dev.jain@arm.com,
	lance.yang@linux.dev, usama.arif@linux.dev,  xu.xin16@zte.com.cn,
	chengming.zhou@linux.dev, roman.gushchin@linux.dev,
	 muchun.song@linux.dev, linux-kernel@vger.kernel.org,
	linux-doc@vger.kernel.org,  driver-core@lists.linux.dev,
	nvdimm@lists.linux.dev, linux-cxl@vger.kernel.org,
	 linux-debuggers@vger.kernel.org, linux-fsdevel@vger.kernel.org,
	kvm@vger.kernel.org,  cgroups@vger.kernel.org,
	damon@lists.linux.dev, linux-kselftest@vger.kernel.org,
	 kernel-team@meta.com, hokyoon.lee@samsung.com,
	wj28.lee@samsung.com,  gost.dev@samsung.com,
	vikash.k5@samsung.com, d.bueso@samsung.com
Subject: Re: [PATCH v5 00/36] Private Memory NUMA Nodes
Date: Tue, 22 Sep 2026 18:04:56 -0400	[thread overview]
Message-ID: <arL2eTRaDeR3DCQx@gourry-fedora-PF4VCD3F> (raw)
In-Reply-To: <360067785.01790071202608.JavaMail.epsvc@epcpadp1new>

On Tue, Sep 22, 2026 at 03:27:44PM +0530, Arun George/Arun George wrote:
> Hi Gregory,
> 
> Thanks for sharing the working branch.
> 
> Some of us at Samsung are testing the series (both v4 & v5) on a 
> compression capable CXL expander and are observing encouraging results.
>

This is really cool, thank you so much for spending the cycles to test.

I'm working on getting a v6 out soon.  Glad to see the progress!

> Good things first! We feel that the isolation using private nodes and 
> the capability selection using NODE_PRIVATE_CAP_* work functionally well 
> for the compressed memory use cases.
> 
> On improvement points, we observed much higher 'page_faults', 
> 'allocation_stalls' etc. which contributed to the higher tail latencies 
> in some tests. I hope these are already part of the optimization plans. 
> The higher latencies might have resulted from the write_protection 
> applied to the private node, I believe.
> 

Yes, this is expected, and I think there are questions about how it
should be optimized both in a provisioning sense and a software sense.

For example, exposing a full memory expander as 100% compressible may
not make sense - given that this could many a lot of memory.  I have
some data that shows it may make more sense to split a memory expander
into two regions to better fit the compressible region to meet the
actual capacity the cold-tail is capable of consuming.

On the software side there may be some ability to adjust this at
runtime with hotplug instead of requiring a reboot, but I think that
needs to be experimented with.

> On the tools side, we used MLC and Taobench for the tests.
> 

Great! I also have some TaoBench and fio results i'm looking to share,
glad to have some additional peer review using Tao.

> Note that our intention for this phase of experiments was to test only 
> the 'private node' layer based isolation, and not the cram and below 
> layers for the compressed memory. Therefore we did not enable/utilize 
> the compression capability in the hardware for these set of experiments 
> (enabling compression would require the cram level ballooning/memory 
> shrinking and cxl level interfacing driver to the compression device). 
> That would be different set of experiments where we would be testing the 
> cram balloon shrinkers and our alternate algorithms (upstream targeted). 
> So cram and below layers were used only for enumeration of private node 
> regions for these tests and not for the run-time memory shrinkers.
> 

Agreed, i'm looking forward to getting past the isolation bits and get
more focused on the compression implementation details.

> ======= Test Methodology ==============
> 
> We explored these 3 cases:
> 
>     1) (Baseline – existing CXL infra):  DRAM 16GB + CXL 32GB. Here we 
> used the existing CXL driver infra to enable the device. No private node 
> is involved. And compression is disabled on device.
> 
>     2) (private node in default write protected path): DRAM 16GB + CXL 
> Private Node 32GB. Here private node infra is used to enable the device. 
> And compression is disabled on device.
> 
>     3) (private node with no write protection): DRAM 16GB + CXL Private 
> Node 32GB. Here private node infra is used to enable the device without 
> the write protection/fencing enabled. Compression is disabled on device. 
> We could not complete these runs as they resulted in kernel panics. 
> Guess the code path is not stable yet for this (We had hoped that case 1 
> and case 3 results would be similar). We also observed some unmovable 
> page warnings logs in 'dmesg' before crash (might be related to the 
> panic). Adding the log snippets at the end.
> 

Very similar to my test setup, seems like good signal we are of similar
mind on the use case.  (For readers: there was no coordination here).

I personally did not test private-node without write protection,
although it makes sense to test the throughput of a demotion-only
node as a baseline.  Smart.

My guess is there was probably a reclaim throughput issue, or there may
have been a bug in the v4/v5 code.  I will look at adding some pressure
tests to my suite that test demotion-only explicitly.

> ========= Results summary ===================
> 
> Private (CRAM) node performance is lower compared to a normal CXL memory 
> allocation path. VM stats shows more page_faults and allocation_stalls 
> on the private node case. Could it be the allocator waits for migration 
> path to demote pages to private node? Or the actual hot pages in dram 
> got demoted to private node to make space during overwrites? We will try 
> further analysis on this.
> 

- allocator waits for migration?
  
  Yes.  And worse, once the node is full, migration will fail and you'll
  swap directly from the top tier and reduce your reclaim behavior on
  the lower node.  It becomes very important to push proactive reclaim
  on a demotion-only tier to ensure there is sufficient headroom to
  receive demotions in the future, since direct-reclaim basically never
  targets a lower-tier on the first pass.

  There's a big discussion about whether tiered systems need to invert
  their reclaim behavior (reclaim from the lowest node first to find
  progress there, then do demotion - rather than the other way around).

- Hot pages in dram 

  Also yes.  The comparison to make isn't only against a raw CXL node,
  but also against Zswap or Zram or Swap.  How does the workload fair
  on a system with 16GB RAM and 32GB Zswap/Swap?

  This is important context.

> ========== MLC Experiments ====================
> 
> CMD:
>    $./mlc --loaded_latency -j0 -c0 -b1g -k1-15 -W5 -r
> 
> Experiment Results:
> 
> 1. (Baseline) DRAM 16GB + CXL 32GB
> 
> Inject  Latency Bandwidth
> Delay   (ns)    MB/sec
> ================
>   00000  686.18   40137.5
>   00002  686.69   40124.0
>   00008  709.74   40362.7
...
> 2. (private node in default write protected path)
>      DRAM 16GB + CXL Private Node(Uncompressed) 32GB
> 
> Inject  Latency Bandwidth
> Delay   (ns)    MB/sec
> ================
>   00000  320.09   12186.2
>   00002  543.42   22726.1
>   00008  482.11   28385.2
...
> Takeaway: Our MLC tests show a trade-off between the two configurations. 
> During low inject delay, the private Node is faster. But when inject 
> delay goes high, the baseline shows better latency. Could it be the TLB 
> cache effects?
> 

These latency results suggest to me that the memory being tested was
always DRAM, though the initial memory fault would have been CXL.

i.e. even if the memory was faulted directly onto the node via
mempolicy, the very first write to it would have caused promotion.

If you dropped the write protection you'd probably see it equals the
baseline.

The bandwidth numbers support this.  Higher stall, but faster latency
means the migration was more beneficial for latency than letting the
memory sit on CXL.

Neat.  I'm not sure these results are particularly useful in describing
how a workload would react.

> --------------------------------------------------------------------
> Case 3 dmesg log snippet:
> 
> [   82.525741] page: refcount:1 mapcount:0 mapping:0000000000000000 
> index:0x0 pfn:0x5f5800
> [   82.525753] flags: 
> 0x17ffffc0002000(reserved|node=0|zone=2|lastcpupid=0x1fffff)
> [   82.525762] raw: 0017ffffc0002000 ffefc85717d60008 ffefc85717d60008 
> 0000000000000000
> [   82.525764] raw: 0000000000000000 0000000000000000 00000001ffffffff 
> 0000000000000000
> [   82.525766] page dumped because: unmovable page
> ..............
> 

Was this memory hotplugged via cram.c or did you replicate the code and
hotplug it another way?  If so, did you hotplug it as ZONE_NORMAL or
ZONE_MOVABLE?

For a compressed tier, it must always be hotplugged as ZONE_MOVABLE,
without exception, because otherwise it can allow a GUP pin or kernel
allocation that would eventually become permanently stuck there - and by
definition of a compression tier must be 100% movable memory.

~Gregory

  reply	other threads:[~2026-09-22 22:05 UTC|newest]

Thread overview: 74+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
     [not found] <CGME20260720193443epcas5p167fa7d8490edfb4d8c0aa6a259db058f@epcas5p1.samsung.com>
2026-07-20 19:33 ` Gregory Price
2026-07-20 19:33   ` [PATCH v5 01/36] mm: refactor find_next_best_node to find_next_best_node_in Gregory Price
2026-07-21  5:46     ` Balbir Singh
2026-07-20 19:33   ` [PATCH v5 02/36] mm/page_alloc: refactor build_node_zonelist() out of build_zonelists() Gregory Price
2026-07-20 19:33   ` [PATCH v5 03/36] mm/page_alloc: let the bulk and folio allocators carry alloc_flags Gregory Price
2026-07-20 19:33   ` [PATCH v5 04/36] numa: introduce N_MEMORY_PRIVATE Gregory Price
2026-07-20 19:33   ` [PATCH v5 05/36] mm: add ZONELIST_PRIVATE(_NOFALLBACK) for N_MEMORY_PRIVATE nodes Gregory Price
2026-07-22 13:00     ` Richard Cheng
2026-07-22 13:19       ` Gregory Price
2026-07-20 19:34   ` [PATCH v5 06/36] cpuset: exclude private nodes from cpuset.mems (default-open) Gregory Price
2026-07-20 19:34   ` [PATCH v5 07/36] mm/memory_hotplug: disallow migration-driven private node hotunplug Gregory Price
2026-07-20 19:34   ` [PATCH v5 08/36] mm/mempolicy: skip private node folios when queueing for migration Gregory Price
2026-07-20 19:34   ` [PATCH v5 09/36] mm/migrate: disallow userland driven migration for private nodes Gregory Price
2026-07-20 19:34   ` [PATCH v5 10/36] mm/madvise: disallow madvise operations on private node folios Gregory Price
2026-07-20 19:34   ` [PATCH v5 11/36] mm/compaction: disallow compaction on private nodes Gregory Price
2026-07-20 19:34   ` [PATCH v5 12/36] mm/page_alloc: clear private node watermarks and system reserves Gregory Price
2026-07-20 19:34   ` [PATCH v5 13/36] mm/mempolicy: disallow NUMA Balancing prot_none on private nodes Gregory Price
2026-07-20 19:34   ` [PATCH v5 14/36] mm/damon: skip private node memory in DAMON migration and pageout Gregory Price
2026-07-21 23:46     ` SJ Park
2026-07-22 12:16       ` Gregory Price
2026-07-23  0:19         ` SJ Park
2026-07-23  3:25           ` Gregory Price
2026-07-23 13:36             ` SJ Park
2026-09-10  8:58             ` David Hildenbrand (Arm)
2026-09-10 14:06               ` SJ Park
2026-07-20 19:34   ` [PATCH v5 15/36] mm/ksm: skip KSM for managed-memory folios Gregory Price
2026-07-20 19:34   ` [PATCH v5 16/36] mm/khugepaged: skip private node folios when trying to collapse Gregory Price
2026-07-20 19:34   ` [PATCH v5 17/36] mm/vmscan: disallow reclaim of private node memory Gregory Price
2026-07-20 19:34   ` [PATCH v5 18/36] mm/gup: disallow longterm pin of private node folios Gregory Price
2026-07-20 19:34   ` [PATCH v5 19/36] proc: include N_MEMORY_PRIVATE nodes in numa_maps output Gregory Price
2026-07-20 19:34   ` [PATCH v5 20/36] mm/memcontrol: account private-node memory in per-node stats Gregory Price
2026-07-20 19:34   ` [PATCH v5 21/36] proc/kcore: include private-node RAM in the kcore RAM map Gregory Price
2026-07-20 20:05     ` Omar Sandoval
2026-07-20 19:34   ` [PATCH v5 22/36] mm/mempolicy: add MPOL_F_PRIVATE and zonelist selection Gregory Price
2026-07-20 19:34   ` [PATCH v5 23/36] mm/mempolicy: apply policy at the kernel zone for private-node binds Gregory Price
2026-07-20 19:34   ` [PATCH v5 24/36] mm/mempolicy: add in-kernel MPOL_BIND interfaces for drivers/services Gregory Price
2026-07-20 19:34   ` [PATCH v5 25/36] mm/memory_hotplug: support N_MEMORY_PRIVATE node hotplug Gregory Price
2026-08-06  3:30     ` Qiqi Li
2026-08-17  0:49       ` Gregory Price
2026-09-04  1:05         ` Qiqi Li
2026-09-04  3:18           ` Gregory Price
2026-07-20 19:34   ` [PATCH v5 26/36] mm: add NODE_PRIVATE_CAP_RECLAIM for opted-in private node reclaim Gregory Price
2026-07-22 13:36     ` Richard Cheng
2026-07-22 13:48       ` Gregory Price
2026-07-20 19:34   ` [PATCH v5 27/36] mm: add NODE_PRIVATE_CAP_USER_NUMA for userland numa controls Gregory Price
2026-07-20 19:34   ` [PATCH v5 28/36] mm: add NODE_PRIVATE_CAP_HOTUNPLUG for opted-in private nodes Gregory Price
2026-07-22 14:04     ` Gregory Price
2026-07-20 19:34   ` [PATCH v5 29/36] mm: add NODE_PRIVATE_CAP_DEMOTION for private-node tiering demotion Gregory Price
2026-07-20 19:34   ` [PATCH v5 30/36] mm: add NODE_PRIVATE_CAP_NUMA_BALANCING for private-node NUMA balancing Gregory Price
2026-07-20 19:34   ` [PATCH v5 31/36] mm: add NODE_PRIVATE_CAP_LTPIN for private node folio pinning Gregory Price
2026-07-20 19:34   ` [PATCH v5 32/36] mm/khugepaged: base private node collapse eligiblity on actor/cap bits Gregory Price
2026-07-22 13:24     ` Richard Cheng
2026-07-22 13:43       ` Gregory Price
2026-07-20 19:34   ` [PATCH v5 33/36] Documentation/mm: describe private (N_MEMORY_PRIVATE) memory nodes Gregory Price
2026-07-20 19:34   ` [PATCH v5 34/36] mm/mempolicy: add mpol_set_shared_policy_range() Gregory Price
2026-07-20 19:34   ` [PATCH v5 35/36] KVM: guest_memfd: bind backing memory to a NUMA node at creation Gregory Price
2026-07-21  3:46   ` [PATCH v5 00/36] Private Memory NUMA Nodes Balbir Singh
2026-07-21 18:16     ` Gregory Price
2026-07-22  8:29       ` Balbir Singh
2026-07-22 12:28         ` Gregory Price
2026-07-21 13:26   ` Zenghui Yu
2026-07-21 17:18     ` Gregory Price
2026-07-22 14:20   ` Richard Cheng
2026-07-22 15:40     ` Gregory Price
2026-07-22 20:56       ` Gregory Price
2026-07-24  6:29       ` Richard Cheng
2026-07-23  8:38   ` Arun George/Arun George
2026-07-23 16:53     ` Gregory Price
2026-09-22  9:57       ` Arun George/Arun George
2026-09-22 22:04         ` Gregory Price [this message]
2026-08-12 12:01   ` Pankaj Gupta
2026-08-13  1:21     ` Gregory Price
2026-08-13  8:55       ` Pankaj Gupta
2026-08-14 13:27         ` Gregory Price

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=arL2eTRaDeR3DCQx@gourry-fedora-PF4VCD3F \
    --to=gourry@gourry.net \
    --cc=Zhigang.Luo@amd.com \
    --cc=akpm@linux-foundation.org \
    --cc=alison.schofield@intel.com \
    --cc=alucerop@amd.com \
    --cc=apopple@nvidia.com \
    --cc=arun.george@samsung.com \
    --cc=axelrasmussen@google.com \
    --cc=balbirs@nvidia.com \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=brendan.jackman@linux.dev \
    --cc=byungchul@sk.com \
    --cc=cgroups@vger.kernel.org \
    --cc=chengming.zhou@linux.dev \
    --cc=corbet@lwn.net \
    --cc=d.bueso@samsung.com \
    --cc=dakr@kernel.org \
    --cc=damon@lists.linux.dev \
    --cc=dave.jiang@intel.com \
    --cc=david@kernel.org \
    --cc=dev.jain@arm.com \
    --cc=djbw@kernel.org \
    --cc=driver-core@lists.linux.dev \
    --cc=gost.dev@samsung.com \
    --cc=gregkh@linuxfoundation.org \
    --cc=hannes@cmpxchg.org \
    --cc=hokyoon.lee@samsung.com \
    --cc=jackmanb@google.com \
    --cc=jannh@google.com \
    --cc=jgg@ziepe.ca \
    --cc=jhubbard@nvidia.com \
    --cc=joshua.hahnjy@gmail.com \
    --cc=kasong@tencent.com \
    --cc=kernel-team@meta.com \
    --cc=kvm@vger.kernel.org \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-cxl@vger.kernel.org \
    --cc=linux-debuggers@vger.kernel.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=linux@rasmusvillemoes.dk \
    --cc=ljs@kernel.org \
    --cc=longman@redhat.com \
    --cc=matthew.brost@intel.com \
    --cc=mhocko@suse.com \
    --cc=mkoutny@suse.com \
    --cc=muchun.song@linux.dev \
    --cc=npache@redhat.com \
    --cc=nvdimm@lists.linux.dev \
    --cc=osalvador@suse.de \
    --cc=osandov@osandov.com \
    --cc=pbonzini@redhat.com \
    --cc=peterx@redhat.com \
    --cc=pfalcato@suse.de \
    --cc=qi.zheng@linux.dev \
    --cc=rafael@kernel.org \
    --cc=rakie.kim@sk.com \
    --cc=ridong.chen@linux.dev \
    --cc=roman.gushchin@linux.dev \
    --cc=rppt@kernel.org \
    --cc=ryan.roberts@arm.com \
    --cc=shakeel.butt@linux.dev \
    --cc=sj@kernel.org \
    --cc=skhan@linuxfoundation.org \
    --cc=surenb@google.com \
    --cc=tj@kernel.org \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=vikash.k5@samsung.com \
    --cc=vishal.l.verma@intel.com \
    --cc=weixugc@google.com \
    --cc=wj28.lee@samsung.com \
    --cc=xu.xin16@zte.com.cn \
    --cc=ying.huang@linux.alibaba.com \
    --cc=yuanchu@google.com \
    --cc=yury.norov@gmail.com \
    --cc=yuzenghui@huawei.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®