memmap_init_zone_device() can take a noticeable amount of time when large pmem namespaces are bound or rebound, because it initializes nearly identical struct page descriptors one PFN at a time. This series reduces that ZONE_DEVICE memmap initialization overhead by reusing prepared struct page templates and, on x86, using memcpy_nontemporal() for the template copy path.
The main target is large fsdax/devdax pmem configurations, where the cost of initializing the memmap shows up directly in nd_pmem/dax_pmem bind and rebind latency. Patches 1-3 are preparatory cleanups and helper extraction. Patches 4-5 add the template-copy path for head pages and compound tails. Patches 6-8 introduce memcpy_nontemporal(), extend the x86 fixed-size memcpy_flushcache() inline cases used by that helper, and switch the template-copy path over to memcpy_nontemporal(). Patch 9 removes the remaining local opt-out predicate and always uses the template path after the first real page seeds the reusable template. Architectures without a specialized memcpy_nontemporal() backend fall back to memcpy(), so the template-copy optimization remains available without arch-specific support. On x86, memcpy_nontemporal() maps to the existing memcpy_flushcache() backend and can use the fixed-size MOVNTI paths added by this series for struct page sized copies. This version also incorporates a round of review comments from Sashiko on patches 6-8. For patch 6, the generic memcpy_nontemporal() fallback is a macro alias to memcpy(), rather than an inline wrapper, so architectures without a specialized backend keep the usual memcpy() FORTIFY coverage when object sizes remain visible at the original call site. For patch 7, I re-checked whether the new 32/48/64/80/96-byte MOVNTI cases should add a separate alignment check before staying on the inline fast path. I re-checked the Intel SDM together with the existing memcpy_flushcache() behavior, and I did not find an architectural statement that would make those new cases require a different correctness rule from the long-standing 4/8/16-byte fast paths. This series therefore keeps the new fixed-size cases inline without adding a separate alignment check. For patch 8, I re-checked whether the template-copy path needs a helper-level drain after the MOVNTI-based copies. This series keeps memcpy_nontemporal() as a copy primitive only. Callers that use it for a producer-consumer or device-visible handoff must provide the required ordering. The ZONE_DEVICE template-copy path uses it only while initializing struct page metadata, so the copy primitive itself does not grow a separate drain contract. The numbers below measure the time spent in memmap_init_zone_device() during driver bind/rebind. They are not measurements of the full nd_pmem or dax_pmem bind/rebind operation. Tested in a VM with a 100 GB fsdax namespace device configured with map=dev and a 100 GB devdax namespace (align=2097152) on Intel Ice Lake server. Test procedure: Rebind the nd_pmem and dax_pmem driver 30 times and collect the memmap initialization time from the pr_debug() output of memmap_init_zone_device(). Base(v7.2-rc1): First binding for nd_pmem driver: 1456 ms Average of subsequent rebinds: 244.28 ms First binding for dax_pmem driver: 1462 ms Average of subsequent rebinds: 273.31 ms With this series applied: First binding for nd_pmem driver: 1272 ms Average of subsequent rebinds: 96.79 ms First binding for dax_pmem driver: 1354 ms Average of subsequent rebinds: 119.04 ms This reduces the average memmap initialization time measured during rebind by about 60.4% for nd_pmem and 56.4% for dax_pmem. As an additional data point, I also ran a smaller set of measurements on the same physical x86_64 host with a 100 GB PMEM region created via the memmap= kernel command line, configured as fsdax and devdax namespaces with map=dev and 2 MiB alignment. For brevity, the individual patches keep only the VM results rather than including a second set of physical-host measurements throughout the series. The physical-host numbers below are included only as supplemental evidence that the same optimization also provides a similar benefit on a non-virtualized system. Test procedure: Reconfigure the namespace mode, rebind the nd_pmem or dax_pmem driver once, and collect the memmap initialization time from the pr_debug() output of memmap_init_zone_device(). Base (v7.2-rc1): nd_pmem / fsdax: 179 ms dax_pmem / devdax: 264 ms With this series applied: nd_pmem / fsdax: 82 ms dax_pmem / devdax: 113 ms This reduces the measured memmap initialization time during rebind by about 54.2% for nd_pmem and 57.2% for dax_pmem on that setup, which is broadly consistent with the VM results above. As another supplemental data point, I also measured the test_hmm.ko module on the same physical x86_64 host, using the test_hmm.ko setup from the previous discussion that times ten 64 GB memremap_pages()/memunmap_pages() iterations during module insertion[1]. By default, module insertion initializes two DEVICE_PRIVATE dmirror devices, so two avg memremap values are reported; each value is the average for one 64 GB chunk. This is not the primary target workload of the series, but it exercises the same large ZONE_DEVICE memmap initialization path and shows the same direction of improvement. Base (v7.2-rc1): avg memremap reported during module insertion: 116689362 ns, 116539263 ns With this series applied: avg memremap reported during module insertion: 54607108 ns, 54458236 ns This corresponds to about a 53.2% reduction based on the mean of the reported values, which is again consistent with the pmem bind/rebind results above. [1] https://lore.kernel.org/all/[email protected]/ Li Zhe (9): mm: fix stale ZONE_DEVICE refcount comment mm: factor zone-device page init helpers out of __init_zone_device_page mm: add a set_page_section_from_pfn() helper mm: add a template-based fast path for zone-device page init mm: extend the template fast path to zone-device compound tails string: introduce memcpy_nontemporal() x86/string: extend memcpy_flushcache() fixed-size fastpaths mm: use memcpy_nontemporal() in zone-device template copies mm: always use the zone-device template init path arch/x86/include/asm/string_64.h | 68 +++++++++++++++- include/linux/mm.h | 15 +++- include/linux/string.h | 12 +++ mm/mm_init.c | 131 +++++++++++++++++++++++++------ 4 files changed, 198 insertions(+), 28 deletions(-) --- v6: https://lore.kernel.org/all/[email protected]/ v5: https://lore.kernel.org/all/[email protected]/ v4: https://lore.kernel.org/all/[email protected]/ v3: https://lore.kernel.org/all/[email protected]/ v2: https://lore.kernel.org/all/[email protected]/ v1: https://lore.kernel.org/all/[email protected]/ Changelogs: v6->v7: - Rename memcpy_nt() to memcpy_nontemporal() so the generic helper name is explicit rather than abbreviated. Suggested by Muchun Song. - Drop the separate memcpy_nt_drain() helper and remove the drain calls from the ZONE_DEVICE template-copy path. Suggested by Muchun Song. - Rework patch 6 so the generic memcpy_nontemporal() fallback uses '#define memcpy_nontemporal memcpy', preserving the usual memcpy() FORTIFY coverage when object sizes remain visible at the original call site. - Summarize the off-list Sashiko review comments relayed by Andrew on the memcpy helper and x86 memcpy_flushcache() pieces. - Add patch 9 to remove the local template-enable predicate and the remaining non-template fallback path, as suggested by Muchun. This also means KASAN/KMSAN builds no longer force the per-page initialization path, and page_ref_set no longer observes every initialization-time refcount assignment. - Re-ran the VM performance test for v7. The results were close to the previously reported v5 numbers, so keep the existing performance data unchanged. - Rename pagemap_resets_refcount() to pagemap_requires_refcount_reset() and invert the predicate accordingly. Suggested by Balbir Singh. For changelogs of earlier revisions, please refer to the v6 cover letter. -- 2.20.1

