On Mon, Jul 27, 2026 at 02:23:33PM -0400, Zi Yan wrote: > On 27 Jul 2026, at 13:33, David Hildenbrand (Arm) wrote: > > > On 7/22/26 17:28, Zi Yan wrote: > >> On Wed Jul 22, 2026 at 10:26 AM EDT, Zi Yan wrote: > >>> On Wed Jul 22, 2026 at 12:42 AM EDT, Matthew Brost wrote: > >>>> Add a lightweight structural split primitive for large (compound) folios > >>>> that a driver allocated with __GFP_COMP and manages entirely by itself, > >>>> outside of the core mm's view. > >>>> > >>>> The existing split paths - split_folio() and folio_split_unmapped() - > >>>> are built for folios that the mm owns: they perform a refcount freeze, > >>>> walk and remap the rmap, and take the anon_vma / i_mmap locks, and > >>>> folio_split_unmapped() further assumes an anon, pagecache-style refcount > >>>> model (nr_pages + 1). None of that applies to a folio that is: > >>>> > >>>> - singly referenced (the caller holds the only reference), > >>>> - not mapped through the rmap (folio_mapped() == 0), > >>>> - not in the page cache or swap cache (folio->mapping == NULL), > >>>> - not on any LRU or the deferred-split list. > >>>> > >>>> For such a folio the split is purely structural: because nothing else in > >>>> the kernel can reach it, there is no need to freeze the refcount or touch > >>>> any mapping. folio_split_driver_managed() therefore performs only the > >>> > >>> I do not think so. PFN scanners like memory compaction should be able to > >>> see them, unless you mean something else about "a driver allocated with > >>> __GFP_COMP". You will need to freeze it to prevent the to-be-split > >>> folios being touched by others. > >>> > >>> > >>>> compound and split-accounting teardown via __split_unmapped_folio() and > >>>> then hands each resulting order-@new_order folio its own reference, > >>>> mirroring split_page() for compound folios. The caller keeps the original > >>>> reference on the first resulting folio and is responsible for freeing all > >>>> of them individually. > >>>> > >>>> The immediate user is TTM's GPU page pool, which allocates higher-order > >>>> compound pages, maps them into userspace via VM_PFNMAP (never through the > >>>> rmap), and needs to split them into order-0 folios under memory pressure > >>>> so pages can be backed up to shmem and freed one at a time. > >>>> > >>>> A CONFIG_TRANSPARENT_HUGEPAGE=n stub is provided so callers can build > >>>> without the split machinery; it warns and returns -EINVAL. > >>>> > >>>> Cc: Maarten Lankhorst <[email protected]> > >>>> Cc: Maxime Ripard <[email protected]> > >>>> Cc: Thomas Zimmermann <[email protected]> > >>>> Cc: David Airlie <[email protected]> > >>>> Cc: Simona Vetter <[email protected]> > >>>> Cc: Christian Koenig <[email protected]> > >>>> Cc: Huang Rui <[email protected]> > >>>> Cc: Matthew Auld <[email protected]> > >>>> Cc: Matthew Brost <[email protected]> > >>>> Cc: Andrew Morton <[email protected]> > >>>> Cc: David Hildenbrand <[email protected]> > >>>> Cc: Lorenzo Stoakes <[email protected]> > >>>> Cc: Zi Yan <[email protected]> > >>>> Cc: Baolin Wang <[email protected]> > >>>> Cc: "Liam R. Howlett" <[email protected]> > >>>> Cc: Nico Pache <[email protected]> > >>>> Cc: Ryan Roberts <[email protected]> > >>>> Cc: Dev Jain <[email protected]> > >>>> Cc: Barry Song <[email protected]> > >>>> Cc: Lance Yang <[email protected]> > >>>> Cc: Tvrtko Ursulin <[email protected]> > >>>> Cc: Dave Airlie <[email protected]> > >>>> Cc: [email protected] > >>>> Cc: [email protected] > >>>> Cc: [email protected] > >>>> Suggested-by: Matthew Wilcox <[email protected]> > >>>> Signed-off-by: Matthew Brost <[email protected]> > >>>> Assisted-by: GitHub-Copilot:claude-opus-4.8 > >>>> > >>>> --- > >>>> > >>>> The patch is based on drm-tip rather than the core MM branches to > >>>> facilitate Intel CI testing and initial review. It can be rebased onto > >>>> the core MM branches in a subsequent revision. > >>>> --- > >>>> include/linux/huge_mm.h | 8 ++++++ > >>>> mm/huge_memory.c | 63 +++++++++++++++++++++++++++++++++++++++++ > >>>> 2 files changed, 71 insertions(+) > >>>> > >>>> diff --git a/include/linux/huge_mm.h b/include/linux/huge_mm.h > >>>> index ad20f7f8c179..35661d82d54a 100644 > >>>> --- a/include/linux/huge_mm.h > >>>> +++ b/include/linux/huge_mm.h > >>>> @@ -402,6 +402,7 @@ enum split_type { > >>>> int __split_huge_page_to_list_to_order(struct page *page, struct > >>>> list_head *list, > >>>> unsigned int new_order); > >>>> int folio_split_unmapped(struct folio *folio, unsigned int new_order); > >>>> +int folio_split_driver_managed(struct folio *folio, unsigned int > >>>> new_order); > >>>> unsigned int min_order_for_split(struct folio *folio); > >>>> int split_folio_to_list(struct folio *folio, struct list_head *list); > >>>> int folio_check_splittable(struct folio *folio, unsigned int new_order, > >>>> @@ -656,6 +657,13 @@ static inline int split_folio_to_list(struct folio > >>>> *folio, struct list_head *lis > >>>> return -EINVAL; > >>>> } > >>>> > >>>> +static inline int folio_split_driver_managed(struct folio *folio, > >>>> + unsigned int new_order) > >>>> +{ > >>>> + VM_WARN_ON_ONCE_FOLIO(1, folio); > >>>> + return -EINVAL; > >>>> +} > >>>> + > >>>> static inline int folio_split(struct folio *folio, unsigned int > >>>> new_order, > >>>> struct page *page, struct list_head *list) > >>>> { > >>>> diff --git a/mm/huge_memory.c b/mm/huge_memory.c > >>>> index 2bccb0a53a0a..06f9a5f35df8 100644 > >>>> --- a/mm/huge_memory.c > >>>> +++ b/mm/huge_memory.c > >>>> @@ -4185,6 +4185,69 @@ int folio_split_unmapped(struct folio *folio, > >>>> unsigned int new_order) > >>>> return ret; > >>>> } > >>>> > >>>> +/** > >>>> + * folio_split_driver_managed() - split an exclusively-owned, off-LRU > >>>> folio > >>>> + * @folio: folio to split. Must be a large (compound) folio that is > >>>> owned > >>>> + * exclusively by the caller and is invisible to the core mm. > >>>> + * @new_order: the order of the folios after the split. > >>>> + * > >>>> + * This is a lightweight structural split for folios that a driver > >>>> allocated > >>>> + * and manages itself (for example TTM's GPU page pool, which allocates > >>>> + * higher-order compound pages with __GFP_COMP and maps them into > >>>> userspace > >>>> + * via VM_PFNMAP rather than through the rmap). Such folios are: > >> > >> Strickly speaking, these vm_insert*() compound pages are not folios, > >> since folios are supposed to be rmappable and they are either anonymous > >> memory or file-backed memory. I am working on separating them from > >> rmappable folios by replacing PG_private with PG_folio and marking all > >> pages in a folio with PG_folio in page_rmappable_folio(). > >> > >> Hopefully, we can find a better name, like refcounted_folio, later for > >> these non-rmappable compound pages. > > > > They wouldn't really be folios, I guess. They would likely be a simple > > "refcounted" memtype that allows for compound pages. > > Yes, they are not folios. But “folio” was started to replace “compound page” > and slowly becomes rmappable anon and file-backed. People outside MM still > thinks “folio” == “compound page”. But once I manage to remove PG_private > and get us PG_folio, page_folio() will return NULL for non-folio compound > pages and we will need a new type for them, “refcounted_XXX”. We can decide > XXX when I get there. :) > > > > > But what is the conclusion here? It sounds like "folio_split_" is the > > entirely > > wrong interface for these compound pages. > > We probably would allow folio_split() to be used on compound pages now > until we can make a clean distinction, e.g., using PG_folio, between them. > > Yes, folio_split() and its helper functions are meant for rmappable anon and > file-backed folios. But currently “folio” is de facto “compound page”, since > for example prep_compound_head() initializes folio fields even if it is meant > only for compound pages. > > After folio and compound page are separate concepts in the code base, probably > we can think about how to have two split functions for them and still keep > maximum code reuse. Namely, we could have split_compound() does the compound > page split (e.g., copying page flags, adjust compound head/tail/order, etc.) > and split_folio() calls split_compound() and perform extra folio operations. > The details are TBD.
I like the idea of split_compound, ideally in mm/page_alloc.c, similar to split_page(), which performs the mechanical splitting of compound pages managed on the driver side (i.e., not mappable, not on an LRU, protected by dma-resv in the case of DRM, etc.). I got yelled at for not allocating these types of pages with GFP_COMP, and I basically found that this missing MM primitive is what has prevented that approach from being practical. Using GFP_COMP makes our life quite a bit easier for managing these types of pages too, so it would be great if we could get a solution for this. So maybe I should drop this for now, keep an eye on the MM work, and circle back once the missing primitive is implemented? Thanks the help / insight, Matt > > Best Regards, > Yan, Zi
