mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Arun George/Arun George <arun.george@samsung.com>
To: Gregory Price <gourry@gourry.net>
Cc: linux-mm@kvack.org, Zhigang.Luo@amd.com, balbirs@nvidia.com,
	brendan.jackman@linux.dev, yuzenghui@huawei.com,
	apopple@nvidia.com, alucerop@amd.com, matthew.brost@intel.com,
	akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org,
	liam@infradead.org, vbabka@kernel.org, rppt@kernel.org,
	surenb@google.com, mhocko@suse.com, corbet@lwn.net,
	skhan@linuxfoundation.org, gregkh@linuxfoundation.org,
	rafael@kernel.org, dakr@kernel.org, djbw@kernel.org,
	vishal.l.verma@intel.com, dave.jiang@intel.com,
	alison.schofield@intel.com, osandov@osandov.com,
	jannh@google.com, pfalcato@suse.de, jackmanb@google.com,
	hannes@cmpxchg.org, ziy@nvidia.com, pbonzini@redhat.com,
	osalvador@suse.de, joshua.hahnjy@gmail.com, rakie.kim@sk.com,
	byungchul@sk.com, ying.huang@linux.alibaba.com,
	kasong@tencent.com, qi.zheng@linux.dev, shakeel.butt@linux.dev,
	baohua@kernel.org, axelrasmussen@google.com, yuanchu@google.com,
	weixugc@google.com, yury.norov@gmail.com,
	linux@rasmusvillemoes.dk, longman@redhat.com,
	ridong.chen@linux.dev, tj@kernel.org, mkoutny@suse.com,
	sj@kernel.org, jgg@ziepe.ca, jhubbard@nvidia.com,
	peterx@redhat.com, baolin.wang@linux.alibaba.com,
	npache@redhat.com, ryan.roberts@arm.com, dev.jain@arm.com,
	lance.yang@linux.dev, usama.arif@linux.dev, xu.xin16@zte.com.cn,
	chengming.zhou@linux.dev, roman.gushchin@linux.dev,
	muchun.song@linux.dev, linux-kernel@vger.kernel.org,
	linux-doc@vger.kernel.org, driver-core@lists.linux.dev,
	nvdimm@lists.linux.dev, linux-cxl@vger.kernel.org,
	linux-debuggers@vger.kernel.org, linux-fsdevel@vger.kernel.org,
	kvm@vger.kernel.org, cgroups@vger.kernel.org,
	damon@lists.linux.dev, linux-kselftest@vger.kernel.org,
	kernel-team@meta.com, hokyoon.lee@samsung.com,
	wj28.lee@samsung.com, gost.dev@samsung.com,
	vikash.k5@samsung.com, d.bueso@samsung.com
Subject: Re: [PATCH v5 00/36] Private Memory NUMA Nodes
Date: Tue, 22 Sep 2026 15:27:44 +0530	[thread overview]
Message-ID: <360067785.01790071202608.JavaMail.epsvc@epcpadp1new> (raw)
In-Reply-To: <amJGUuBZ2857QjDB@gourry-fedora-PF4VCD3F>

Hi Gregory,

Thanks for sharing the working branch.

Some of us at Samsung are testing the series (both v4 & v5) on a 
compression capable CXL expander and are observing encouraging results.

Good things first! We feel that the isolation using private nodes and 
the capability selection using NODE_PRIVATE_CAP_* work functionally well 
for the compressed memory use cases.

On improvement points, we observed much higher 'page_faults', 
'allocation_stalls' etc. which contributed to the higher tail latencies 
in some tests. I hope these are already part of the optimization plans. 
The higher latencies might have resulted from the write_protection 
applied to the private node, I believe.

On the tools side, we used MLC and Taobench for the tests.

Note that our intention for this phase of experiments was to test only 
the 'private node' layer based isolation, and not the cram and below 
layers for the compressed memory. Therefore we did not enable/utilize 
the compression capability in the hardware for these set of experiments 
(enabling compression would require the cram level ballooning/memory 
shrinking and cxl level interfacing driver to the compression device). 
That would be different set of experiments where we would be testing the 
cram balloon shrinkers and our alternate algorithms (upstream targeted). 
So cram and below layers were used only for enumeration of private node 
regions for these tests and not for the run-time memory shrinkers.

======= Test Methodology ==============

We explored these 3 cases:

    1) (Baseline – existing CXL infra):  DRAM 16GB + CXL 32GB. Here we 
used the existing CXL driver infra to enable the device. No private node 
is involved. And compression is disabled on device.

    2) (private node in default write protected path): DRAM 16GB + CXL 
Private Node 32GB. Here private node infra is used to enable the device. 
And compression is disabled on device.

    3) (private node with no write protection): DRAM 16GB + CXL Private 
Node 32GB. Here private node infra is used to enable the device without 
the write protection/fencing enabled. Compression is disabled on device. 
We could not complete these runs as they resulted in kernel panics. 
Guess the code path is not stable yet for this (We had hoped that case 1 
and case 3 results would be similar). We also observed some unmovable 
page warnings logs in 'dmesg' before crash (might be related to the 
panic). Adding the log snippets at the end.

========= Results summary ===================

- TaoBench: ‘private node’ showed much higher latencies. We believe this 
could be due to the migration back into the DRAM for overwrites (due to 
write-protection) based on the kernel VM stats.

- MLC:  For lower inject delay, the private node shows better latency; 
but when inject delay goes high, the baseline is better.

Overall kernel VM stats suggested much higher ‘page faults, ‘allocation 
stalls’ etc. in ‘private node’ case which is corroborated with the results.

========== Setup ======================

CPU: 144 Cores (single socket), DRAM: 16GB, CXL: 32GB,
Kernel: 7.2.0-rc2+ (Private Node), 6.18.32 (Baseline)

Private node opt CAPs:
#define CRAM_NP_CAPS    (NODE_PRIVATE_CAP_RECLAIM |
                         NODE_PRIVATE_CAP_HOTUNPLUG | \
                         NODE_PRIVATE_CAP_DEMOTION |
                         NODE_PRIVATE_CAP_USER_NUMA | \
                         NODE_PRIVATE_CAP_NUMA_BALANCING |
                         NODE_PRIVATE_POLICY_WRITE_FENCE)

Enabled 'numa balancing' and 'demotion'.
   $ echo 1 > /proc/sys/kernel/numa_balancing
   $ echo true > /sys/kernel/mm/numa/demotion_enabled

========== TaoBench Experiments ====================

CMD:
$ python3 benchpress_cli.py run tao_bench_standalone -i '{"bind_mem": 0, 
"bind_cpu": 0, "memsize": 32, "test_time": 300, "warmup_time": 0, 
"clients_per_thread": 100, "set_get_ratio": "1:9"}'

1. (baseline) DRAM 16GB + CXL 32GB   ==>

$ grep -A8 "^ALL STATS" benchmark_metrics_27cfff0e/client_0.log

===================================================
Type  Avg. Latency    p50        p95        p99
----------------------------------------------------
Sets   2.75215     2.73500      4.57500     5.79100
Gets   2.84236     2.81500      4.76700     5.95100

  2. (private node in default write protected path)
      DRAM 16GB + CXL Private Node(uncompressed) 32GB   ===>

$ grep -A8 "^ALL STATS" benchmark_metrics_5181fda5/client_0.log

===================================================
Type  Avg. Latency    p50        p95        p99
----------------------------------------------------
Sets   5.58395     4.73500      12.35100     21.75900
Gets   4.83344     4.19100      10.43100     17.91900

Vmstat metrics comparison of both cases for Taobench
($cat /proc/vmstat):

-------------------------------------------------------------------
Metric           |  Baseline    |   CXL Private Node |   Ratio
-------------------------------------------------------------------
pgfault            37,121,808       171,512,616          4.6x
pgactivate          5,360,928       161,947,594         30.2x
pgreuse             4,346,396           407,397         10.7x less
pgrefill           88,775,854       436,486,518          4.9x
pgdemote_kswapd     5,933,439       129,586,750         21.8x
pgdemote_direct        24,837        35,505,114       1429x
allocstall_normal         469           153,600        327x
kswapd_low_wmark_hit_quickly 671         27,821         41.5x
pageoutrun              1,333            29,217         21.9x
numa_pte_updates   24,508,077                 0          na
numa_hint_faults   21,152,264                 0          na
numa_pages_migrated 3,330,897                 0          na
numa_hit           46,667,202       353,812,703          7.6x

---------- Takeaways -----

Private (CRAM) node performance is lower compared to a normal CXL memory 
allocation path. VM stats shows more page_faults and allocation_stalls 
on the private node case. Could it be the allocator waits for migration 
path to demote pages to private node? Or the actual hot pages in dram 
got demoted to private node to make space during overwrites? We will try 
further analysis on this.

========== MLC Experiments ====================

CMD:
   $./mlc --loaded_latency -j0 -c0 -b1g -k1-15 -W5 -r

Experiment Results:

1. (Baseline) DRAM 16GB + CXL 32GB

Inject  Latency Bandwidth
Delay   (ns)    MB/sec
================
  00000  686.18   40137.5
  00002  686.69   40124.0
  00008  709.74   40362.7
  00015  711.75   40350.0
  00050  710.72   40357.4
  00100  712.22   40316.0
  00200  695.29   40026.5
  00300  602.91   38845.3
  00400  482.28   36240.4
  00500  410.03   33606.6
  00700  275.13   26352.5
  01000  241.82   18851.7
  01300  230.91   14695.8
  01700  215.27   11403.2
  02500  217.76    7900.9
  03500  194.42    5787.0
  05000  192.17    4166.4
  09000  190.36    2475.0
  20000  188.09    1312.1

2. (private node in default write protected path)
     DRAM 16GB + CXL Private Node(Uncompressed) 32GB

Inject  Latency Bandwidth
Delay   (ns)    MB/sec
================
  00000  320.09   12186.2
  00002  543.42   22726.1
  00008  482.11   28385.2
  00015  488.26   26304.6
  00050  485.19   26898.4
  00100  477.65   24668.9
  00200  483.68   21273.8
  00300  486.44   17655.9
  00400  480.41   16470.0
  00500  481.30   13589.9
  00700  493.84    9960.8
  01000  488.66    8267.7
  01300  491.31    6433.6
  01700  476.97    6211.2
  02500  472.97    5037.0
  03500  466.05    3533.8
  05000  446.44    2783.1
  09000  417.03    1877.5
  20000  387.22    1035.3

Vmstat metrics comparison of both cases for MLC
($cat /proc/vmstat):

-------------------------------------------------------------------
Metric           |  Baseline    |   CXL Private Node |   Ratio
-------------------------------------------------------------------
pgfault           10,287,329         46540,782           4.6x
pgactivate           1141468        65,098,272          57x
pgrefill             260,699        47,592,751         182.5x
pgdemote_kswapd    1,085,783        26,063,754          24x
pgdemote_direct    4,562,539        15,809,485           3.5x
pgmigrate_fail        41,929         2,763,725          65.9x
allocstall_normal        161               425           2.6x
kswapd_low_wmark_hit_quickly  47           763          16.2x
pageoutrun                55               936          17x

Takeaway: Our MLC tests show a trade-off between the two configurations. 
During low inject delay, the private Node is faster. But when inject 
delay goes high, the baseline shows better latency. Could it be the TLB 
cache effects?

--------------------------------------------------------------------
Case 3 dmesg log snippet:

[   82.525741] page: refcount:1 mapcount:0 mapping:0000000000000000 
index:0x0 pfn:0x5f5800
[   82.525753] flags: 
0x17ffffc0002000(reserved|node=0|zone=2|lastcpupid=0x1fffff)
[   82.525762] raw: 0017ffffc0002000 ffefc85717d60008 ffefc85717d60008 
0000000000000000
[   82.525764] raw: 0000000000000000 0000000000000000 00000001ffffffff 
0000000000000000
[   82.525766] page dumped because: unmovable page
..............

We will update on the compression enabled experiments further.

~ Arun

On 23-07-2026 10:23 pm, Gregory Price wrote:
> On Thu, Jul 23, 2026 at 02:08:31PM +0530, Arun George/Arun George wrote:
>> On 21-07-2026 01:03 am, Gregory Price wrote:
>> We intend to test this series on a compression capable CXL expander
>> hardware. Since compressed ram example (cram) is not part of this
>> series, how do you suggest to do that? Do you have a version of cram
>> module compatible with this series?
>>
>> ~Arun
> 
> Hi Arun,
> > I plan to RFC the new setup for compressed ram this a bit later, but I
> will share my working branch with you for testing.
> 
> ~Gregory



  reply	other threads:[~2026-09-22 10:00 UTC|newest]

Thread overview: 74+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
     [not found] <CGME20260720193443epcas5p167fa7d8490edfb4d8c0aa6a259db058f@epcas5p1.samsung.com>
2026-07-20 19:33 ` Gregory Price
2026-07-20 19:33   ` [PATCH v5 01/36] mm: refactor find_next_best_node to find_next_best_node_in Gregory Price
2026-07-21  5:46     ` Balbir Singh
2026-07-20 19:33   ` [PATCH v5 02/36] mm/page_alloc: refactor build_node_zonelist() out of build_zonelists() Gregory Price
2026-07-20 19:33   ` [PATCH v5 03/36] mm/page_alloc: let the bulk and folio allocators carry alloc_flags Gregory Price
2026-07-20 19:33   ` [PATCH v5 04/36] numa: introduce N_MEMORY_PRIVATE Gregory Price
2026-07-20 19:33   ` [PATCH v5 05/36] mm: add ZONELIST_PRIVATE(_NOFALLBACK) for N_MEMORY_PRIVATE nodes Gregory Price
2026-07-22 13:00     ` Richard Cheng
2026-07-22 13:19       ` Gregory Price
2026-07-20 19:34   ` [PATCH v5 06/36] cpuset: exclude private nodes from cpuset.mems (default-open) Gregory Price
2026-07-20 19:34   ` [PATCH v5 07/36] mm/memory_hotplug: disallow migration-driven private node hotunplug Gregory Price
2026-07-20 19:34   ` [PATCH v5 08/36] mm/mempolicy: skip private node folios when queueing for migration Gregory Price
2026-07-20 19:34   ` [PATCH v5 09/36] mm/migrate: disallow userland driven migration for private nodes Gregory Price
2026-07-20 19:34   ` [PATCH v5 10/36] mm/madvise: disallow madvise operations on private node folios Gregory Price
2026-07-20 19:34   ` [PATCH v5 11/36] mm/compaction: disallow compaction on private nodes Gregory Price
2026-07-20 19:34   ` [PATCH v5 12/36] mm/page_alloc: clear private node watermarks and system reserves Gregory Price
2026-07-20 19:34   ` [PATCH v5 13/36] mm/mempolicy: disallow NUMA Balancing prot_none on private nodes Gregory Price
2026-07-20 19:34   ` [PATCH v5 14/36] mm/damon: skip private node memory in DAMON migration and pageout Gregory Price
2026-07-21 23:46     ` SJ Park
2026-07-22 12:16       ` Gregory Price
2026-07-23  0:19         ` SJ Park
2026-07-23  3:25           ` Gregory Price
2026-07-23 13:36             ` SJ Park
2026-09-10  8:58             ` David Hildenbrand (Arm)
2026-09-10 14:06               ` SJ Park
2026-07-20 19:34   ` [PATCH v5 15/36] mm/ksm: skip KSM for managed-memory folios Gregory Price
2026-07-20 19:34   ` [PATCH v5 16/36] mm/khugepaged: skip private node folios when trying to collapse Gregory Price
2026-07-20 19:34   ` [PATCH v5 17/36] mm/vmscan: disallow reclaim of private node memory Gregory Price
2026-07-20 19:34   ` [PATCH v5 18/36] mm/gup: disallow longterm pin of private node folios Gregory Price
2026-07-20 19:34   ` [PATCH v5 19/36] proc: include N_MEMORY_PRIVATE nodes in numa_maps output Gregory Price
2026-07-20 19:34   ` [PATCH v5 20/36] mm/memcontrol: account private-node memory in per-node stats Gregory Price
2026-07-20 19:34   ` [PATCH v5 21/36] proc/kcore: include private-node RAM in the kcore RAM map Gregory Price
2026-07-20 20:05     ` Omar Sandoval
2026-07-20 19:34   ` [PATCH v5 22/36] mm/mempolicy: add MPOL_F_PRIVATE and zonelist selection Gregory Price
2026-07-20 19:34   ` [PATCH v5 23/36] mm/mempolicy: apply policy at the kernel zone for private-node binds Gregory Price
2026-07-20 19:34   ` [PATCH v5 24/36] mm/mempolicy: add in-kernel MPOL_BIND interfaces for drivers/services Gregory Price
2026-07-20 19:34   ` [PATCH v5 25/36] mm/memory_hotplug: support N_MEMORY_PRIVATE node hotplug Gregory Price
2026-08-06  3:30     ` Qiqi Li
2026-08-17  0:49       ` Gregory Price
2026-09-04  1:05         ` Qiqi Li
2026-09-04  3:18           ` Gregory Price
2026-07-20 19:34   ` [PATCH v5 26/36] mm: add NODE_PRIVATE_CAP_RECLAIM for opted-in private node reclaim Gregory Price
2026-07-22 13:36     ` Richard Cheng
2026-07-22 13:48       ` Gregory Price
2026-07-20 19:34   ` [PATCH v5 27/36] mm: add NODE_PRIVATE_CAP_USER_NUMA for userland numa controls Gregory Price
2026-07-20 19:34   ` [PATCH v5 28/36] mm: add NODE_PRIVATE_CAP_HOTUNPLUG for opted-in private nodes Gregory Price
2026-07-22 14:04     ` Gregory Price
2026-07-20 19:34   ` [PATCH v5 29/36] mm: add NODE_PRIVATE_CAP_DEMOTION for private-node tiering demotion Gregory Price
2026-07-20 19:34   ` [PATCH v5 30/36] mm: add NODE_PRIVATE_CAP_NUMA_BALANCING for private-node NUMA balancing Gregory Price
2026-07-20 19:34   ` [PATCH v5 31/36] mm: add NODE_PRIVATE_CAP_LTPIN for private node folio pinning Gregory Price
2026-07-20 19:34   ` [PATCH v5 32/36] mm/khugepaged: base private node collapse eligiblity on actor/cap bits Gregory Price
2026-07-22 13:24     ` Richard Cheng
2026-07-22 13:43       ` Gregory Price
2026-07-20 19:34   ` [PATCH v5 33/36] Documentation/mm: describe private (N_MEMORY_PRIVATE) memory nodes Gregory Price
2026-07-20 19:34   ` [PATCH v5 34/36] mm/mempolicy: add mpol_set_shared_policy_range() Gregory Price
2026-07-20 19:34   ` [PATCH v5 35/36] KVM: guest_memfd: bind backing memory to a NUMA node at creation Gregory Price
2026-07-21  3:46   ` [PATCH v5 00/36] Private Memory NUMA Nodes Balbir Singh
2026-07-21 18:16     ` Gregory Price
2026-07-22  8:29       ` Balbir Singh
2026-07-22 12:28         ` Gregory Price
2026-07-21 13:26   ` Zenghui Yu
2026-07-21 17:18     ` Gregory Price
2026-07-22 14:20   ` Richard Cheng
2026-07-22 15:40     ` Gregory Price
2026-07-22 20:56       ` Gregory Price
2026-07-24  6:29       ` Richard Cheng
2026-07-23  8:38   ` Arun George/Arun George
2026-07-23 16:53     ` Gregory Price
2026-09-22  9:57       ` Arun George/Arun George [this message]
2026-09-22 22:04         ` Gregory Price
2026-08-12 12:01   ` Pankaj Gupta
2026-08-13  1:21     ` Gregory Price
2026-08-13  8:55       ` Pankaj Gupta
2026-08-14 13:27         ` Gregory Price

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=360067785.01790071202608.JavaMail.epsvc@epcpadp1new \
    --to=arun.george@samsung.com \
    --cc=Zhigang.Luo@amd.com \
    --cc=akpm@linux-foundation.org \
    --cc=alison.schofield@intel.com \
    --cc=alucerop@amd.com \
    --cc=apopple@nvidia.com \
    --cc=axelrasmussen@google.com \
    --cc=balbirs@nvidia.com \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=brendan.jackman@linux.dev \
    --cc=byungchul@sk.com \
    --cc=cgroups@vger.kernel.org \
    --cc=chengming.zhou@linux.dev \
    --cc=corbet@lwn.net \
    --cc=d.bueso@samsung.com \
    --cc=dakr@kernel.org \
    --cc=damon@lists.linux.dev \
    --cc=dave.jiang@intel.com \
    --cc=david@kernel.org \
    --cc=dev.jain@arm.com \
    --cc=djbw@kernel.org \
    --cc=driver-core@lists.linux.dev \
    --cc=gost.dev@samsung.com \
    --cc=gourry@gourry.net \
    --cc=gregkh@linuxfoundation.org \
    --cc=hannes@cmpxchg.org \
    --cc=hokyoon.lee@samsung.com \
    --cc=jackmanb@google.com \
    --cc=jannh@google.com \
    --cc=jgg@ziepe.ca \
    --cc=jhubbard@nvidia.com \
    --cc=joshua.hahnjy@gmail.com \
    --cc=kasong@tencent.com \
    --cc=kernel-team@meta.com \
    --cc=kvm@vger.kernel.org \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-cxl@vger.kernel.org \
    --cc=linux-debuggers@vger.kernel.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=linux@rasmusvillemoes.dk \
    --cc=ljs@kernel.org \
    --cc=longman@redhat.com \
    --cc=matthew.brost@intel.com \
    --cc=mhocko@suse.com \
    --cc=mkoutny@suse.com \
    --cc=muchun.song@linux.dev \
    --cc=npache@redhat.com \
    --cc=nvdimm@lists.linux.dev \
    --cc=osalvador@suse.de \
    --cc=osandov@osandov.com \
    --cc=pbonzini@redhat.com \
    --cc=peterx@redhat.com \
    --cc=pfalcato@suse.de \
    --cc=qi.zheng@linux.dev \
    --cc=rafael@kernel.org \
    --cc=rakie.kim@sk.com \
    --cc=ridong.chen@linux.dev \
    --cc=roman.gushchin@linux.dev \
    --cc=rppt@kernel.org \
    --cc=ryan.roberts@arm.com \
    --cc=shakeel.butt@linux.dev \
    --cc=sj@kernel.org \
    --cc=skhan@linuxfoundation.org \
    --cc=surenb@google.com \
    --cc=tj@kernel.org \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=vikash.k5@samsung.com \
    --cc=vishal.l.verma@intel.com \
    --cc=weixugc@google.com \
    --cc=wj28.lee@samsung.com \
    --cc=xu.xin16@zte.com.cn \
    --cc=ying.huang@linux.alibaba.com \
    --cc=yuanchu@google.com \
    --cc=yury.norov@gmail.com \
    --cc=yuzenghui@huawei.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®