On Tue, Sep 15, 2026 at 9:26 AM T.J. Mercier <[email protected]> wrote:
>
> On Tue, Sep 15, 2026 at 8:44 AM Christian König
> <[email protected]> wrote:
> >
> > On 9/15/26 17:33, T.J. Mercier wrote:
> > > On Tue, Sep 15, 2026 at 1:41 AM Christian König
> > > <[email protected]> wrote:
> > >>
> > >> On 9/14/26 23:22, Davidlohr Bueso wrote:
> > >>> With 4KB base pages, the system heap allocates buffers in 1MB, 64KB
> > >>> and 4KB chunks. The conventional Intel VT-d second-stage and AMD-Vi v2
> > >>> page-table formats, and Arm SMMU's 64-bit long-descriptor format with
> > >>> a 4KB translation granule, define 4KB pages and 2MB/1GB large-page
> > >>> mappings, but no 1MB leaf. A 1MB chunk therefore requires 256 4KB leaf
> > >>> entries unless it can be combined with adjacent chunks into a
> > >>> suitably aligned larger mapping.
> > >>
> > >> Yes, it was an intentional design choice to *NOT* optimize for AMD nor 
> > >> Intels IOMMU here.
> > >>
> > >>> Adding a 2MB allocation order provides naturally aligned, physically
> > >>> contiguous chunks matching the 2MB leaf size, without relying on this
> > >>> being accidental for separate smaller allocations. The new cost
> > >>> is one failed order-9 attempt (keeping current large allocation 
> > >>> semantics)
> > >>> per buffer when 2MB pages are exhausted.
> > >>
> > >> We have plenty of experience with this with TTM and the overhead this 
> > >> results in is usually not acceptable.
> > >
> > > Overhead from compaction and reclaim to get 2M pages? That is disabled
> > > for non-0 page orders here in HIGH_ORDER_GFP.
> >
> > Ah! Thanks for pointing that out, I misread the code that __GFP_RECLAIM is 
> > ORed into the mask.
> >
> > But the question is still why? Without reclaim that is pretty much an 
> > useless feature on most x86 boxes.
>
> Even on arm64 phones free 2M pages aren't likely to be available for
> very long after boot, but lately we have been doing more proactive
> reclaim triggered by userspace before launching workflows that desire
> large dma-buf allocations (and also at other times during application
> lifecycle in general). That somewhat increases the likelihood that
> high order pages will be available. I agree there's no guarantee we'll
> get any, but without __GFP_RECLAIM the attempt is pretty cheap and the
> payoff can be pretty benficial which Davidlohr's allocation and
> mapping measurements demonstrate.

Oh I forgot to add that this behavior is similar to how slab
allocations are done for slabs requiring more than an order 0 page per
slab. A large optimal order is attempted first with ~__GFP_RECLAIM,
but if that fails a fallback to a smaller min order is done. When the
optimal order allocation attempt fails, the cost is just taking the
zone lock for a quick peek at the free lists.

https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/mm/slub.c?h=v7.2#n3373

> > Regards,
> > Christian.
> >
> > >
> > >>> Two benchmarks, measured on an AMD EPYC 7313P.
> > >>
> > >> Why in the world are you testing the system heap on an AMD EPYC system?
> > >>
> > >> AMD clearly doesn't recommend using the system heap on those boxes for 
> > >> ROCm, so I'm really wondering what combination of HW you have here?
> > >>
> > >> Regards,
> > >> Christian.
> > >>
> > >>>
> > >>> (i) Mapping, into an idle NVMe function's translated DMA-FQ domain,
> > >>> measuring map/unmap: with the v1 page table restricted to 4K/2M/1G
> > >>> (amd_iommu=v2_pgsizes_only), a 1GB buffer goes from 1024 1MB chunks,
> > >>> only one of whose 2MB windows was superpage-mappable in that run, to
> > >>> 512 2MB chunks with all 512 mappable. dma_buf_map_attachment() costs
> > >>> decrease by ~16x on average with ~32x for the worst cases. Unmapping
> > >>> that buffer drops from ~500 us to 1.25 us.
> > >>>
> > >>> (ii) A microbench that measures the cost of DMA_HEAP_IOCTL_ALLOC+close
> > >>> across various thread counts decreases by factors of ~3-7x alleviating
> > >>> allocator's zone->lock contention by being PMD order and therefore is
> > >>> pcpu list eligible. Once the buffer size is large enough then the cost
> > >>> of the zeroing takes over.
> > >>>
> > >>> Signed-off-by: Davidlohr Bueso <[email protected]>
> > >>> ---
> > >>>  drivers/dma-buf/heaps/system_heap.c | 13 +++++++------
> > >>>  1 file changed, 7 insertions(+), 6 deletions(-)
> > >>>
> > >>> diff --git a/drivers/dma-buf/heaps/system_heap.c 
> > >>> b/drivers/dma-buf/heaps/system_heap.c
> > >>> index c8959eadc71d..621cf238ef97 100644
> > >>> --- a/drivers/dma-buf/heaps/system_heap.c
> > >>> +++ b/drivers/dma-buf/heaps/system_heap.c
> > >>> @@ -55,14 +55,15 @@ struct dma_heap_attachment {
> > >>>  #define HIGH_ORDER_GFP  (((GFP_HIGHUSER | __GFP_ZERO | __GFP_NOWARN \
> > >>>                                 | __GFP_NORETRY) & ~__GFP_RECLAIM) \
> > >>>                                 | __GFP_COMP)
> > >>> -static gfp_t order_flags[] = {HIGH_ORDER_GFP, HIGH_ORDER_GFP, 
> > >>> LOW_ORDER_GFP};
> > >>> +static gfp_t order_flags[] = {HIGH_ORDER_GFP, HIGH_ORDER_GFP,
> > >>> +                             HIGH_ORDER_GFP, LOW_ORDER_GFP};
> > >>>  /*
> > >>> - * The selection of the orders used for allocation (1MB, 64K, 4K) is 
> > >>> designed
> > >>> - * to match with the sizes often found in IOMMUs. Using order 4 pages 
> > >>> instead
> > >>> - * of order 0 pages can significantly improve the performance of many 
> > >>> IOMMUs
> > >>> - * by reducing TLB pressure and time spent updating page tables.
> > >>> + * The selection of the orders used for allocation (2MB, 1MB, 64K, 4K) 
> > >>> is
> > >>> + * designed to match with the sizes often found in IOMMUs. Using 
> > >>> larger order
> > >>> + * pages instead of order 0 pages can significantly improve the 
> > >>> performance of
> > >>> + * many IOMMUs by reducing TLB pressure and time spent updating page 
> > >>> tables.
> > >>>   */
> > >>> -static const unsigned int orders[] = {8, 4, 0};
> > >>> +static const unsigned int orders[] = {9, 8, 4, 0};
> > >>>  #define NUM_ORDERS ARRAY_SIZE(orders)
> > >>>
> > >>>  static int system_heap_set_page_decrypted(struct page *page)
> > >>> --
> > >>> 2.39.5
> > >>>
> > >>
> >

Reply via email to