On Mon, Jul 20, 2026 at 03:33:54PM -0400, Gregory Price wrote:
> This series introduces the concept of "Private Memory Nodes", which
> are opted into two basic functionalities by default:
>   - page allocation (mm/page_alloc.c)
>   - OOM killing
> 

All, I wanted to preface v6 with an update, I probably won't be sending
it out on-list before LPC (i've been a bit noisey as it is, I'll give
it a rest).


But I would like to give you the highlights for those interested.

The latest working branch can be found here:
https://github.com/gourryinverse/linux/tree/scratch/gourry/managed_nodes/rfc6-nonuma

You will find the following:

- There is no depedency on exported alloc_flags (willy, i found a way),
  but there is a need for `folio_alloc_private()` to explicitly ask
  for the ALLOC_ZONELIST_PRIVATE.

- feature bits have been converted to node_state[] bits

- N_MEMORY_COMMON was added to mean "common memory" - i.e. the existing
  N_MEMORY set.  A private node is now defined as !N_MEMORY_COMMON.

- opt-out sites no longer check for !N_MEMORY_PRIVATE, they instead
  iterate over node_state[N_MEMORY_COMMON].  This is *much* cleaner.

- opt-in sites get to define bits in terms of their own service, e.g.

  /* for each reclaimable node */
  for_each_node_mask_state(..., N_MEMORY_RECLAIM) {
     ...
  }

  Or maybe even:

  for_each_reclaimable_node()    :)


- There are only 2 node features in the base set:

  N_MEMORY_RECLAIM    -  generic reclaim runs on the node
  N_MEMORY_COMPACTION -  compaction will run on the node

  You'll notice there's no userland numa control support, more on this
  in a moment.

  This is all you need to have functional device-private coherent
  memory that uses the page allocator.

  Disabling N_MEMORY_RECLAIM does not necessarily mean you're prevented
  from using vmscan.c and the LRU - it just means we'd need to expose an
  interface for a !N_MEMORY_RECLAIM node to explicitly ask for reclaim
  to operate on it.

  Tiering (Demotion) *to* a node is not supported.  We can probably
  debate whether this should mean tiering *from* a node should be
  supported or not as well.  This is easy to change.


- The working branch also includes a minimal compressed-ram service
  that only supports anonymous memory.

  Supporting file-backed memory (and shmem) is a whole different can
  of worms that deserves its own discussion.

  I will post this as a separate RFC from the base set, but it is
  available to play with on my working branch.


- I ship a basic dax driver for the base series, `anondax` which allows
  existing devices to online a private node with both compaction and
  reclaim support for testing.  This enables this use case:

  fd = open(/dev/my_device, ..., O_DIRECT);
  buf = mmap(fd, ..., MAP_SHARED)
  my_file = open(myfile, ...);
  read(buf, myfile); /* fault directly onto device memory */

  With swap (and no tiering), this gives you a private over-commitable
  node for which aspiring drivers can use to implement their own
  mmu_notifier callback stream to manage device page tables and such.


- The reason mempolicy (and user numa in general) is not supported is a
  critical relationship between the OOM killer and page allocator.

  In trivial scenarios, if a consumer of a private node overshoots and
  causes an OOM - it will be the largest consumer and be chosen as the
  victim.

  If there are many consumers - what actually happens in practice is the
  oom killer attempts to select a victim based on *task* policy... which
  makes every task a candidate.  This leads to, among other things,
  either an OOM storm and/or a full blown panic because a victim cannot
  be located (depending on the constraints).

  I've concluded that as-is, this is not tractable. But also, it is not
  actually needed if the devices provide their own chardev to provide
  a basic wrapper around folio_alloc_private().

  This unfortunately means potential users like guest_memfd are unlikely
  to be able to use the nodes without hard-coding the allocator
  interface, as opposed to a mempolicy.

  I think it's possible instead to *maybe* add N_MEMORY_USER_MIGRATION
  and allow for explicit movement between nodes, but not full blown
  mempolicies.

  The relationship between page_alloc, cpuset, mempolicy, and the oom
  killer is just too tight to allow allocations without the use of
  __GFP_THISNODE - which the private allocation interface will enforce.


As it stands, the series has been heavily minimized in both line count
and patch number.

The base series is ~22 commits with the scary diffstat below, but many
of the core mm/ patches are similar to the patch below - with nearly
all of the real complexity happening in reclaim and memory hotplug.

despite the branch name, I intend to drop RFC from here-on, everything
else required for the base series is in mm-new.


See you at LPC
~Gregory

-------

diff --git a/mm/migrate.c b/mm/migrate.c
index 7bdcdb57652f8..cd61566b44fa1 100644
--- a/mm/migrate.c
+++ b/mm/migrate.c
@@ -2266,7 +2266,7 @@ static int __add_folio_for_migration(struct folio *folio, 
int node,
        if (is_zero_folio(folio) || is_huge_zero_folio(folio))
                return -EFAULT;

-       if (folio_is_zone_device(folio))
+       if (!folio_is_common_memory(folio))
                return -ENOENT;

        if (folio_nid(folio) == node)
@@ -2390,7 +2390,7 @@ static int do_pages_move(struct mm_struct *mm, nodemask_t 
task_nodes,
                err = -ENODEV;
                if (node < 0 || node >= MAX_NUMNODES)
                        goto out_flush;
-               if (!node_state(node, N_MEMORY))
+               if (!node_state(node, N_MEMORY_COMMON))
                        goto out_flush;

                err = -EACCES;
@@ -2475,7 +2475,7 @@ static void do_pages_stat_array(struct mm_struct *mm, 
unsigned long nr_pages,
                if (folio) {
                        if (is_zero_folio(folio) || is_huge_zero_folio(folio))
                                err = -EFAULT;
-                       else if (folio_is_zone_device(folio))
+                       else if (!folio_is_common_memory(folio))
                                err = -ENOENT;
                        else
                                err = folio_nid(folio);


 Documentation/ABI/stable/sysfs-devices-node |  21 +++++++++++++++++
 Documentation/admin-guide/cgroup-v2.rst     |   7 ++++++
 Documentation/mm/physical_memory.rst        |  15 ++++++++++++-
 drivers/base/node.c                         |  67 
+++++++++++++++++++++++++++++++++++++++++++++++++++++--
 drivers/dax/kmem.c                          |   2 +-
 include/linux/gfp.h                         |   6 +++++
 include/linux/memory_hotplug.h              |   2 +-
 include/linux/mmzone.h                      |  45 
+++++++++++++++++++++++++++++++++++++
 include/linux/node.h                        |  12 ++++++++++
 include/linux/nodemask.h                    |  53 
++++++++++++++++++++++++++++++++++++++++++-
 kernel/cgroup/cpuset.c                      |  39 
++++++++++++++++++++++++++------
 kernel/sched/fair.c                         |   6 ++---
 mm/compaction.c                             |  21 +++++++++++------
 mm/damon/ops-common.c                       |  15 ++++++++-----
 mm/damon/vaddr.c                            |  20 +++++++++++++----
 mm/huge_memory.c                            |   2 +-
 mm/internal.h                               |  22 ++++++++++++++++++
 mm/khugepaged.c                             |   9 ++++++--
 mm/ksm.c                                    |   8 ++++---
 mm/madvise.c                                |   6 ++---
 mm/memory-tiers.c                           |  25 +++++++++++----------
 mm/memory_hotplug.c                         | 109 
++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-------------------------------
 mm/mempolicy.c                              |  51 
+++++++++++++++++++++++++-----------------
 mm/migrate.c                                |   6 ++---
 mm/mm_init.c                                |  22 +++++++++---------
 mm/page_alloc.c                             | 115 
+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++---------------
 mm/page_alloc.h                             |   5 +++++
 mm/vmscan.c                                 |  31 +++++++++++++++-----------
 28 files changed, 585 insertions(+), 157 deletions(-)

Reply via email to