mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Gregory Price <gourry@gourry.net>
To: linux-mm@kvack.org
Cc: Zhigang.Luo@amd.com, arun.george@samsung.com, balbirs@nvidia.com,
	 brendan.jackman@linux.dev, yuzenghui@huawei.com,
	apopple@nvidia.com, alucerop@amd.com,  matthew.brost@intel.com,
	akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org,
	 liam@infradead.org, vbabka@kernel.org, rppt@kernel.org,
	surenb@google.com,  mhocko@suse.com, corbet@lwn.net,
	skhan@linuxfoundation.org,  gregkh@linuxfoundation.org,
	rafael@kernel.org, dakr@kernel.org, djbw@kernel.org,
	 vishal.l.verma@intel.com, dave.jiang@intel.com,
	alison.schofield@intel.com,  osandov@osandov.com,
	jannh@google.com, pfalcato@suse.de, jackmanb@google.com,
	 hannes@cmpxchg.org, ziy@nvidia.com, pbonzini@redhat.com,
	osalvador@suse.de,  joshua.hahnjy@gmail.com, rakie.kim@sk.com,
	byungchul@sk.com, ying.huang@linux.alibaba.com,
	 kasong@tencent.com, qi.zheng@linux.dev, shakeel.butt@linux.dev,
	baohua@kernel.org,  axelrasmussen@google.com, yuanchu@google.com,
	weixugc@google.com, yury.norov@gmail.com,
	 linux@rasmusvillemoes.dk, longman@redhat.com,
	ridong.chen@linux.dev, tj@kernel.org,  mkoutny@suse.com,
	sj@kernel.org, jgg@ziepe.ca, jhubbard@nvidia.com,
	 peterx@redhat.com, baolin.wang@linux.alibaba.com,
	npache@redhat.com,  ryan.roberts@arm.com, dev.jain@arm.com,
	lance.yang@linux.dev, usama.arif@linux.dev,  xu.xin16@zte.com.cn,
	chengming.zhou@linux.dev, roman.gushchin@linux.dev,
	 muchun.song@linux.dev, linux-kernel@vger.kernel.org,
	linux-doc@vger.kernel.org,  driver-core@lists.linux.dev,
	nvdimm@lists.linux.dev, linux-cxl@vger.kernel.org,
	 linux-debuggers@vger.kernel.org, linux-fsdevel@vger.kernel.org,
	kvm@vger.kernel.org,  cgroups@vger.kernel.org,
	damon@lists.linux.dev, linux-kselftest@vger.kernel.org,
	 kernel-team@meta.com
Subject: Re: [PATCH v5 00/36] Private Memory NUMA Nodes
Date: Fri, 25 Sep 2026 19:02:59 -0400	[thread overview]
Message-ID: <arb0mV23QmRVDtB3@gourry-fedora-PF4VCD3F> (raw)
In-Reply-To: <20260720193431.3841992-1-gourry@gourry.net>

On Mon, Jul 20, 2026 at 03:33:54PM -0400, Gregory Price wrote:
> This series introduces the concept of "Private Memory Nodes", which
> are opted into two basic functionalities by default:
>   - page allocation (mm/page_alloc.c)
>   - OOM killing
> 

All, I wanted to preface v6 with an update, I probably won't be sending
it out on-list before LPC (i've been a bit noisey as it is, I'll give
it a rest).


But I would like to give you the highlights for those interested.

The latest working branch can be found here:
https://github.com/gourryinverse/linux/tree/scratch/gourry/managed_nodes/rfc6-nonuma

You will find the following:

- There is no depedency on exported alloc_flags (willy, i found a way),
  but there is a need for `folio_alloc_private()` to explicitly ask
  for the ALLOC_ZONELIST_PRIVATE.

- feature bits have been converted to node_state[] bits

- N_MEMORY_COMMON was added to mean "common memory" - i.e. the existing
  N_MEMORY set.  A private node is now defined as !N_MEMORY_COMMON.

- opt-out sites no longer check for !N_MEMORY_PRIVATE, they instead
  iterate over node_state[N_MEMORY_COMMON].  This is *much* cleaner.

- opt-in sites get to define bits in terms of their own service, e.g.

  /* for each reclaimable node */
  for_each_node_mask_state(..., N_MEMORY_RECLAIM) {
     ...
  }

  Or maybe even:

  for_each_reclaimable_node()    :)


- There are only 2 node features in the base set:

  N_MEMORY_RECLAIM    -  generic reclaim runs on the node
  N_MEMORY_COMPACTION -  compaction will run on the node

  You'll notice there's no userland numa control support, more on this
  in a moment.

  This is all you need to have functional device-private coherent
  memory that uses the page allocator.

  Disabling N_MEMORY_RECLAIM does not necessarily mean you're prevented
  from using vmscan.c and the LRU - it just means we'd need to expose an
  interface for a !N_MEMORY_RECLAIM node to explicitly ask for reclaim
  to operate on it.

  Tiering (Demotion) *to* a node is not supported.  We can probably
  debate whether this should mean tiering *from* a node should be
  supported or not as well.  This is easy to change.


- The working branch also includes a minimal compressed-ram service
  that only supports anonymous memory.

  Supporting file-backed memory (and shmem) is a whole different can
  of worms that deserves its own discussion.

  I will post this as a separate RFC from the base set, but it is
  available to play with on my working branch.


- I ship a basic dax driver for the base series, `anondax` which allows
  existing devices to online a private node with both compaction and
  reclaim support for testing.  This enables this use case:

  fd = open(/dev/my_device, ..., O_DIRECT);
  buf = mmap(fd, ..., MAP_SHARED)
  my_file = open(myfile, ...);
  read(buf, myfile); /* fault directly onto device memory */

  With swap (and no tiering), this gives you a private over-commitable
  node for which aspiring drivers can use to implement their own
  mmu_notifier callback stream to manage device page tables and such.


- The reason mempolicy (and user numa in general) is not supported is a
  critical relationship between the OOM killer and page allocator.

  In trivial scenarios, if a consumer of a private node overshoots and
  causes an OOM - it will be the largest consumer and be chosen as the
  victim.

  If there are many consumers - what actually happens in practice is the
  oom killer attempts to select a victim based on *task* policy... which
  makes every task a candidate.  This leads to, among other things,
  either an OOM storm and/or a full blown panic because a victim cannot
  be located (depending on the constraints).

  I've concluded that as-is, this is not tractable. But also, it is not
  actually needed if the devices provide their own chardev to provide
  a basic wrapper around folio_alloc_private().

  This unfortunately means potential users like guest_memfd are unlikely
  to be able to use the nodes without hard-coding the allocator
  interface, as opposed to a mempolicy.

  I think it's possible instead to *maybe* add N_MEMORY_USER_MIGRATION
  and allow for explicit movement between nodes, but not full blown
  mempolicies.

  The relationship between page_alloc, cpuset, mempolicy, and the oom
  killer is just too tight to allow allocations without the use of
  __GFP_THISNODE - which the private allocation interface will enforce.


As it stands, the series has been heavily minimized in both line count
and patch number.

The base series is ~22 commits with the scary diffstat below, but many
of the core mm/ patches are similar to the patch below - with nearly
all of the real complexity happening in reclaim and memory hotplug.

despite the branch name, I intend to drop RFC from here-on, everything
else required for the base series is in mm-new.


See you at LPC
~Gregory

-------

diff --git a/mm/migrate.c b/mm/migrate.c
index 7bdcdb57652f8..cd61566b44fa1 100644
--- a/mm/migrate.c
+++ b/mm/migrate.c
@@ -2266,7 +2266,7 @@ static int __add_folio_for_migration(struct folio *folio, int node,
        if (is_zero_folio(folio) || is_huge_zero_folio(folio))
                return -EFAULT;

-       if (folio_is_zone_device(folio))
+       if (!folio_is_common_memory(folio))
                return -ENOENT;

        if (folio_nid(folio) == node)
@@ -2390,7 +2390,7 @@ static int do_pages_move(struct mm_struct *mm, nodemask_t task_nodes,
                err = -ENODEV;
                if (node < 0 || node >= MAX_NUMNODES)
                        goto out_flush;
-               if (!node_state(node, N_MEMORY))
+               if (!node_state(node, N_MEMORY_COMMON))
                        goto out_flush;

                err = -EACCES;
@@ -2475,7 +2475,7 @@ static void do_pages_stat_array(struct mm_struct *mm, unsigned long nr_pages,
                if (folio) {
                        if (is_zero_folio(folio) || is_huge_zero_folio(folio))
                                err = -EFAULT;
-                       else if (folio_is_zone_device(folio))
+                       else if (!folio_is_common_memory(folio))
                                err = -ENOENT;
                        else
                                err = folio_nid(folio);


 Documentation/ABI/stable/sysfs-devices-node |  21 +++++++++++++++++
 Documentation/admin-guide/cgroup-v2.rst     |   7 ++++++
 Documentation/mm/physical_memory.rst        |  15 ++++++++++++-
 drivers/base/node.c                         |  67 +++++++++++++++++++++++++++++++++++++++++++++++++++++--
 drivers/dax/kmem.c                          |   2 +-
 include/linux/gfp.h                         |   6 +++++
 include/linux/memory_hotplug.h              |   2 +-
 include/linux/mmzone.h                      |  45 +++++++++++++++++++++++++++++++++++++
 include/linux/node.h                        |  12 ++++++++++
 include/linux/nodemask.h                    |  53 ++++++++++++++++++++++++++++++++++++++++++-
 kernel/cgroup/cpuset.c                      |  39 ++++++++++++++++++++++++++------
 kernel/sched/fair.c                         |   6 ++---
 mm/compaction.c                             |  21 +++++++++++------
 mm/damon/ops-common.c                       |  15 ++++++++-----
 mm/damon/vaddr.c                            |  20 +++++++++++++----
 mm/huge_memory.c                            |   2 +-
 mm/internal.h                               |  22 ++++++++++++++++++
 mm/khugepaged.c                             |   9 ++++++--
 mm/ksm.c                                    |   8 ++++---
 mm/madvise.c                                |   6 ++---
 mm/memory-tiers.c                           |  25 +++++++++++----------
 mm/memory_hotplug.c                         | 109 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-------------------------------
 mm/mempolicy.c                              |  51 +++++++++++++++++++++++++-----------------
 mm/migrate.c                                |   6 ++---
 mm/mm_init.c                                |  22 +++++++++---------
 mm/page_alloc.c                             | 115 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++---------------
 mm/page_alloc.h                             |   5 +++++
 mm/vmscan.c                                 |  31 +++++++++++++++-----------
 28 files changed, 585 insertions(+), 157 deletions(-)

      parent reply	other threads:[~2026-09-25 23:03 UTC|newest]

Thread overview: 75+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
     [not found] <CGME20260720193443epcas5p167fa7d8490edfb4d8c0aa6a259db058f@epcas5p1.samsung.com>
2026-07-20 19:33 ` Gregory Price
2026-07-20 19:33   ` [PATCH v5 01/36] mm: refactor find_next_best_node to find_next_best_node_in Gregory Price
2026-07-21  5:46     ` Balbir Singh
2026-07-20 19:33   ` [PATCH v5 02/36] mm/page_alloc: refactor build_node_zonelist() out of build_zonelists() Gregory Price
2026-07-20 19:33   ` [PATCH v5 03/36] mm/page_alloc: let the bulk and folio allocators carry alloc_flags Gregory Price
2026-07-20 19:33   ` [PATCH v5 04/36] numa: introduce N_MEMORY_PRIVATE Gregory Price
2026-07-20 19:33   ` [PATCH v5 05/36] mm: add ZONELIST_PRIVATE(_NOFALLBACK) for N_MEMORY_PRIVATE nodes Gregory Price
2026-07-22 13:00     ` Richard Cheng
2026-07-22 13:19       ` Gregory Price
2026-07-20 19:34   ` [PATCH v5 06/36] cpuset: exclude private nodes from cpuset.mems (default-open) Gregory Price
2026-07-20 19:34   ` [PATCH v5 07/36] mm/memory_hotplug: disallow migration-driven private node hotunplug Gregory Price
2026-07-20 19:34   ` [PATCH v5 08/36] mm/mempolicy: skip private node folios when queueing for migration Gregory Price
2026-07-20 19:34   ` [PATCH v5 09/36] mm/migrate: disallow userland driven migration for private nodes Gregory Price
2026-07-20 19:34   ` [PATCH v5 10/36] mm/madvise: disallow madvise operations on private node folios Gregory Price
2026-07-20 19:34   ` [PATCH v5 11/36] mm/compaction: disallow compaction on private nodes Gregory Price
2026-07-20 19:34   ` [PATCH v5 12/36] mm/page_alloc: clear private node watermarks and system reserves Gregory Price
2026-07-20 19:34   ` [PATCH v5 13/36] mm/mempolicy: disallow NUMA Balancing prot_none on private nodes Gregory Price
2026-07-20 19:34   ` [PATCH v5 14/36] mm/damon: skip private node memory in DAMON migration and pageout Gregory Price
2026-07-21 23:46     ` SJ Park
2026-07-22 12:16       ` Gregory Price
2026-07-23  0:19         ` SJ Park
2026-07-23  3:25           ` Gregory Price
2026-07-23 13:36             ` SJ Park
2026-09-10  8:58             ` David Hildenbrand (Arm)
2026-09-10 14:06               ` SJ Park
2026-07-20 19:34   ` [PATCH v5 15/36] mm/ksm: skip KSM for managed-memory folios Gregory Price
2026-07-20 19:34   ` [PATCH v5 16/36] mm/khugepaged: skip private node folios when trying to collapse Gregory Price
2026-07-20 19:34   ` [PATCH v5 17/36] mm/vmscan: disallow reclaim of private node memory Gregory Price
2026-07-20 19:34   ` [PATCH v5 18/36] mm/gup: disallow longterm pin of private node folios Gregory Price
2026-07-20 19:34   ` [PATCH v5 19/36] proc: include N_MEMORY_PRIVATE nodes in numa_maps output Gregory Price
2026-07-20 19:34   ` [PATCH v5 20/36] mm/memcontrol: account private-node memory in per-node stats Gregory Price
2026-07-20 19:34   ` [PATCH v5 21/36] proc/kcore: include private-node RAM in the kcore RAM map Gregory Price
2026-07-20 20:05     ` Omar Sandoval
2026-07-20 19:34   ` [PATCH v5 22/36] mm/mempolicy: add MPOL_F_PRIVATE and zonelist selection Gregory Price
2026-07-20 19:34   ` [PATCH v5 23/36] mm/mempolicy: apply policy at the kernel zone for private-node binds Gregory Price
2026-07-20 19:34   ` [PATCH v5 24/36] mm/mempolicy: add in-kernel MPOL_BIND interfaces for drivers/services Gregory Price
2026-07-20 19:34   ` [PATCH v5 25/36] mm/memory_hotplug: support N_MEMORY_PRIVATE node hotplug Gregory Price
2026-08-06  3:30     ` Qiqi Li
2026-08-17  0:49       ` Gregory Price
2026-09-04  1:05         ` Qiqi Li
2026-09-04  3:18           ` Gregory Price
2026-07-20 19:34   ` [PATCH v5 26/36] mm: add NODE_PRIVATE_CAP_RECLAIM for opted-in private node reclaim Gregory Price
2026-07-22 13:36     ` Richard Cheng
2026-07-22 13:48       ` Gregory Price
2026-07-20 19:34   ` [PATCH v5 27/36] mm: add NODE_PRIVATE_CAP_USER_NUMA for userland numa controls Gregory Price
2026-07-20 19:34   ` [PATCH v5 28/36] mm: add NODE_PRIVATE_CAP_HOTUNPLUG for opted-in private nodes Gregory Price
2026-07-22 14:04     ` Gregory Price
2026-07-20 19:34   ` [PATCH v5 29/36] mm: add NODE_PRIVATE_CAP_DEMOTION for private-node tiering demotion Gregory Price
2026-07-20 19:34   ` [PATCH v5 30/36] mm: add NODE_PRIVATE_CAP_NUMA_BALANCING for private-node NUMA balancing Gregory Price
2026-07-20 19:34   ` [PATCH v5 31/36] mm: add NODE_PRIVATE_CAP_LTPIN for private node folio pinning Gregory Price
2026-07-20 19:34   ` [PATCH v5 32/36] mm/khugepaged: base private node collapse eligiblity on actor/cap bits Gregory Price
2026-07-22 13:24     ` Richard Cheng
2026-07-22 13:43       ` Gregory Price
2026-07-20 19:34   ` [PATCH v5 33/36] Documentation/mm: describe private (N_MEMORY_PRIVATE) memory nodes Gregory Price
2026-07-20 19:34   ` [PATCH v5 34/36] mm/mempolicy: add mpol_set_shared_policy_range() Gregory Price
2026-07-20 19:34   ` [PATCH v5 35/36] KVM: guest_memfd: bind backing memory to a NUMA node at creation Gregory Price
2026-07-21  3:46   ` [PATCH v5 00/36] Private Memory NUMA Nodes Balbir Singh
2026-07-21 18:16     ` Gregory Price
2026-07-22  8:29       ` Balbir Singh
2026-07-22 12:28         ` Gregory Price
2026-07-21 13:26   ` Zenghui Yu
2026-07-21 17:18     ` Gregory Price
2026-07-22 14:20   ` Richard Cheng
2026-07-22 15:40     ` Gregory Price
2026-07-22 20:56       ` Gregory Price
2026-07-24  6:29       ` Richard Cheng
2026-07-23  8:38   ` Arun George/Arun George
2026-07-23 16:53     ` Gregory Price
2026-09-22  9:57       ` Arun George/Arun George
2026-09-22 22:04         ` Gregory Price
2026-08-12 12:01   ` Pankaj Gupta
2026-08-13  1:21     ` Gregory Price
2026-08-13  8:55       ` Pankaj Gupta
2026-08-14 13:27         ` Gregory Price
2026-09-25 23:02   ` Gregory Price [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=arb0mV23QmRVDtB3@gourry-fedora-PF4VCD3F \
    --to=gourry@gourry.net \
    --cc=Zhigang.Luo@amd.com \
    --cc=akpm@linux-foundation.org \
    --cc=alison.schofield@intel.com \
    --cc=alucerop@amd.com \
    --cc=apopple@nvidia.com \
    --cc=arun.george@samsung.com \
    --cc=axelrasmussen@google.com \
    --cc=balbirs@nvidia.com \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=brendan.jackman@linux.dev \
    --cc=byungchul@sk.com \
    --cc=cgroups@vger.kernel.org \
    --cc=chengming.zhou@linux.dev \
    --cc=corbet@lwn.net \
    --cc=dakr@kernel.org \
    --cc=damon@lists.linux.dev \
    --cc=dave.jiang@intel.com \
    --cc=david@kernel.org \
    --cc=dev.jain@arm.com \
    --cc=djbw@kernel.org \
    --cc=driver-core@lists.linux.dev \
    --cc=gregkh@linuxfoundation.org \
    --cc=hannes@cmpxchg.org \
    --cc=jackmanb@google.com \
    --cc=jannh@google.com \
    --cc=jgg@ziepe.ca \
    --cc=jhubbard@nvidia.com \
    --cc=joshua.hahnjy@gmail.com \
    --cc=kasong@tencent.com \
    --cc=kernel-team@meta.com \
    --cc=kvm@vger.kernel.org \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-cxl@vger.kernel.org \
    --cc=linux-debuggers@vger.kernel.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=linux@rasmusvillemoes.dk \
    --cc=ljs@kernel.org \
    --cc=longman@redhat.com \
    --cc=matthew.brost@intel.com \
    --cc=mhocko@suse.com \
    --cc=mkoutny@suse.com \
    --cc=muchun.song@linux.dev \
    --cc=npache@redhat.com \
    --cc=nvdimm@lists.linux.dev \
    --cc=osalvador@suse.de \
    --cc=osandov@osandov.com \
    --cc=pbonzini@redhat.com \
    --cc=peterx@redhat.com \
    --cc=pfalcato@suse.de \
    --cc=qi.zheng@linux.dev \
    --cc=rafael@kernel.org \
    --cc=rakie.kim@sk.com \
    --cc=ridong.chen@linux.dev \
    --cc=roman.gushchin@linux.dev \
    --cc=rppt@kernel.org \
    --cc=ryan.roberts@arm.com \
    --cc=shakeel.butt@linux.dev \
    --cc=sj@kernel.org \
    --cc=skhan@linuxfoundation.org \
    --cc=surenb@google.com \
    --cc=tj@kernel.org \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=vishal.l.verma@intel.com \
    --cc=weixugc@google.com \
    --cc=xu.xin16@zte.com.cn \
    --cc=ying.huang@linux.alibaba.com \
    --cc=yuanchu@google.com \
    --cc=yury.norov@gmail.com \
    --cc=yuzenghui@huawei.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®