On 2026/9/22 16:02, Qi Zheng wrote:


On 9/13/26 4:37 PM, Muchun Song wrote:
The common vmemmap population path cannot yet handle optimized Device DAX
mappings on its own. It uses pfn_to_zone() to find the shared tail page,
but Device DAX populates its vmemmap at runtime before the ZONE_DEVICE span
is initialized.

Teach the common path to use device_zone() for runtime optimized vmemmap
population while retaining pfn_to_zone() for early boot. This allows the
same path to support both early boot mappings and Device DAX.

The backing PFN supplied by the Device DAX-specific population path is no
longer used, allowing the redundant lookup and population code to be
removed later.

Signed-off-by: Muchun Song <[email protected]>
---
  mm/sparse-vmemmap.c | 44 +++++++++++++++++++-------------------------
  1 file changed, 19 insertions(+), 25 deletions(-)

diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c
index 878d29a4e862..e83821768c12 100644
--- a/mm/sparse-vmemmap.c
+++ b/mm/sparse-vmemmap.c
@@ -208,18 +208,27 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
      struct page *page;
      const unsigned int order = pfn_to_section_compound_order(pfn);
  -    /*
-     * Device DAX still relies on vmemmap_populate_compound_pages() for
-     * head/first-tail allocation and tail-page reuse.
-     */
      if (!vmemmap_optimizable_pfn(pfn))
          return vmemmap_alloc_block_buf(PAGE_SIZE, node, altmap);
  -    zone = pfn_to_zone(pfn, node);
+    /*
+     * At runtime (slab available), only ZONE_DEVICE pages trigger vmemmap
+     * optimization, so device_zone() suffices. Note that pfn_to_zone()
+     * cannot be used at runtime because the zone span is not set up now.
+     */
+    zone = slab_is_available() ? device_zone(node) : pfn_to_zone(pfn, node);
      page = vmemmap_shared_tail_page(order, zone);
      if (!page)
          return NULL;
  +    /*
+     * When a PTE entry is freed, a free_pages() call occurs. This get_page() +     * pairs with put_page_testzero() on the freeing path. This can only occur
+     * when slab is available.
+     */
+    if (slab_is_available())

Would it make sense to introduce a helper function that wraps
slab_is_available() for better readability? Also, it might be worth
adding a comment to the helper function as well.

Thanks for the suggestion. I considered introducing a helper, but the
two uses of slab_is_available() depend on different properties of the
same initialization boundary. One determines how the zone is obtained,
while the other determines whether the shared vmemmap backing page can
participate in page refcounting. Any helper name would therefore be
either too generic or specific to only one of these properties.

I would prefer to keep slab_is_available() explicit and expand the
comments at both call sites. The first comment explains why early
system RAM can use pfn_to_zone(), whereas ZONE_DEVICE population after
slab becomes available must use device_zone(). The second explains why
early shared backing pages, which are allocated from memblock, cannot
yet be refcounted. Once slab becomes available, shared backing pages
are allocated from the buddy allocator and can hold one reference for
each shared PTE mapping.

I will adopt the following revisions. Does this make sense to you?

diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c
index 5e0e30c77431..406d6f7918f7 100644
--- a/mm/sparse-vmemmap.c
+++ b/mm/sparse-vmemmap.c
@@ -212,9 +212,13 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
                return vmemmap_alloc_block_buf(PAGE_SIZE, node, altmap);

        /*
-        * At runtime (slab available), only ZONE_DEVICE pages trigger vmemmap
-        * optimization, so device_zone() suffices. Note that pfn_to_zone()
-        * cannot be used at runtime because the zone span is not set up now.
+        * Before slab is available, vmemmap optimization is used for early
+        * system RAM, whose zone can be determined from the PFN.
+        *
+        * Once slab is available, only ZONE_DEVICE memory reaches this
+        * optimized population path. Its zone span has not been initialized
+        * while its vmemmap is being populated, so pfn_to_zone() cannot be
+        * used. Obtain ZONE_DEVICE directly from the node instead.
         */
        zone = slab_is_available() ? device_zone(node) : pfn_to_zone(pfn, node);
        page = vmemmap_shared_tail_page(order, zone);
@@ -222,9 +226,17 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
                return NULL;

        /*
-        * When a PTE entry is freed, a free_pages() call occurs. This get_page() -        * pairs with put_page_testzero() on the freeing path. This can only occur
-        * when slab is available.
+        * During early vmemmap population, the shared tail vmemmap backing
+        * page is allocated from memblock before its struct page can safely
+        * participate in page refcounting. Therefore, no reference can be
+        * held for each shared PTE mapping, and the mappings must be unshared
+        * before the vmemmap is depopulated.
+        *
+        * Once slab is available, the shared backing page is allocated from
+        * the buddy allocator and can be refcounted. Hold one reference for
+        * each shared PTE mapping. The architecture vmemmap teardown drops
+        * the reference through __free_pages() when removing the mapping,
+        * preventing the backing page from being freed while it is shared.
         */
        if (slab_is_available())
                get_page(page);


At least for me, encountering this always causes a moment of confusion.

+        get_page(page);
+
      return page_address(page);
  }
  @@ -231,27 +240,12 @@ static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, in
        if (pte_none(ptep_get(pte))) {
          pte_t entry;
+        void *p = vmemmap_alloc_pte(pfn, node, altmap);
  -        if (ptpfn == (unsigned long)-1) {
-            void *p = vmemmap_alloc_pte(pfn, node, altmap);
-
-            if (!p)
-                return NULL;
-            ptpfn = PHYS_PFN(__pa(p));
-        } else {
-            /*
-             * When a PTE/PMD entry is freed from the init_mm
-             * there's a free_pages() call to this page allocated
-             * above. Thus this get_page() is paired with the
-             * put_page_testzero() on the freeing path.
-             * This can only called by certain ZONE_DEVICE path,
-             * and through vmemmap_populate_compound_pages() when
-             * slab is available.
-             */
-            if (slab_is_available())
-                get_page(pfn_to_page(ptpfn));
-        }
-        entry = pfn_pte(ptpfn, PAGE_KERNEL);
+        if (!p)
+            return NULL;
+
+        entry = pfn_pte(PHYS_PFN(__pa(p)), PAGE_KERNEL);
          set_pte_at(&init_mm, addr, pte, entry);
      } else if (WARN_ON_ONCE(vmemmap_optimizable_pfn(pfn)))
          return NULL;



Reply via email to