On Tue, Sep 15, 2026 at 8:44 AM Christian König
<[email protected]> wrote:
>
> On 9/15/26 17:33, T.J. Mercier wrote:
> > On Tue, Sep 15, 2026 at 1:41 AM Christian König
> > <[email protected]> wrote:
> >>
> >> On 9/14/26 23:22, Davidlohr Bueso wrote:
> >>> With 4KB base pages, the system heap allocates buffers in 1MB, 64KB
> >>> and 4KB chunks. The conventional Intel VT-d second-stage and AMD-Vi v2
> >>> page-table formats, and Arm SMMU's 64-bit long-descriptor format with
> >>> a 4KB translation granule, define 4KB pages and 2MB/1GB large-page
> >>> mappings, but no 1MB leaf. A 1MB chunk therefore requires 256 4KB leaf
> >>> entries unless it can be combined with adjacent chunks into a
> >>> suitably aligned larger mapping.
> >>
> >> Yes, it was an intentional design choice to *NOT* optimize for AMD nor 
> >> Intels IOMMU here.
> >>
> >>> Adding a 2MB allocation order provides naturally aligned, physically
> >>> contiguous chunks matching the 2MB leaf size, without relying on this
> >>> being accidental for separate smaller allocations. The new cost
> >>> is one failed order-9 attempt (keeping current large allocation semantics)
> >>> per buffer when 2MB pages are exhausted.
> >>
> >> We have plenty of experience with this with TTM and the overhead this 
> >> results in is usually not acceptable.
> >
> > Overhead from compaction and reclaim to get 2M pages? That is disabled
> > for non-0 page orders here in HIGH_ORDER_GFP.
>
> Ah! Thanks for pointing that out, I misread the code that __GFP_RECLAIM is 
> ORed into the mask.
>
> But the question is still why? Without reclaim that is pretty much an useless 
> feature on most x86 boxes.

Even on arm64 phones free 2M pages aren't likely to be available for
very long after boot, but lately we have been doing more proactive
reclaim triggered by userspace before launching workflows that desire
large dma-buf allocations (and also at other times during application
lifecycle in general). That somewhat increases the likelihood that
high order pages will be available. I agree there's no guarantee we'll
get any, but without __GFP_RECLAIM the attempt is pretty cheap and the
payoff can be pretty benficial which Davidlohr's allocation and
mapping measurements demonstrate.


> Regards,
> Christian.
>
> >
> >>> Two benchmarks, measured on an AMD EPYC 7313P.
> >>
> >> Why in the world are you testing the system heap on an AMD EPYC system?
> >>
> >> AMD clearly doesn't recommend using the system heap on those boxes for 
> >> ROCm, so I'm really wondering what combination of HW you have here?
> >>
> >> Regards,
> >> Christian.
> >>
> >>>
> >>> (i) Mapping, into an idle NVMe function's translated DMA-FQ domain,
> >>> measuring map/unmap: with the v1 page table restricted to 4K/2M/1G
> >>> (amd_iommu=v2_pgsizes_only), a 1GB buffer goes from 1024 1MB chunks,
> >>> only one of whose 2MB windows was superpage-mappable in that run, to
> >>> 512 2MB chunks with all 512 mappable. dma_buf_map_attachment() costs
> >>> decrease by ~16x on average with ~32x for the worst cases. Unmapping
> >>> that buffer drops from ~500 us to 1.25 us.
> >>>
> >>> (ii) A microbench that measures the cost of DMA_HEAP_IOCTL_ALLOC+close
> >>> across various thread counts decreases by factors of ~3-7x alleviating
> >>> allocator's zone->lock contention by being PMD order and therefore is
> >>> pcpu list eligible. Once the buffer size is large enough then the cost
> >>> of the zeroing takes over.
> >>>
> >>> Signed-off-by: Davidlohr Bueso <[email protected]>
> >>> ---
> >>>  drivers/dma-buf/heaps/system_heap.c | 13 +++++++------
> >>>  1 file changed, 7 insertions(+), 6 deletions(-)
> >>>
> >>> diff --git a/drivers/dma-buf/heaps/system_heap.c 
> >>> b/drivers/dma-buf/heaps/system_heap.c
> >>> index c8959eadc71d..621cf238ef97 100644
> >>> --- a/drivers/dma-buf/heaps/system_heap.c
> >>> +++ b/drivers/dma-buf/heaps/system_heap.c
> >>> @@ -55,14 +55,15 @@ struct dma_heap_attachment {
> >>>  #define HIGH_ORDER_GFP  (((GFP_HIGHUSER | __GFP_ZERO | __GFP_NOWARN \
> >>>                                 | __GFP_NORETRY) & ~__GFP_RECLAIM) \
> >>>                                 | __GFP_COMP)
> >>> -static gfp_t order_flags[] = {HIGH_ORDER_GFP, HIGH_ORDER_GFP, 
> >>> LOW_ORDER_GFP};
> >>> +static gfp_t order_flags[] = {HIGH_ORDER_GFP, HIGH_ORDER_GFP,
> >>> +                             HIGH_ORDER_GFP, LOW_ORDER_GFP};
> >>>  /*
> >>> - * The selection of the orders used for allocation (1MB, 64K, 4K) is 
> >>> designed
> >>> - * to match with the sizes often found in IOMMUs. Using order 4 pages 
> >>> instead
> >>> - * of order 0 pages can significantly improve the performance of many 
> >>> IOMMUs
> >>> - * by reducing TLB pressure and time spent updating page tables.
> >>> + * The selection of the orders used for allocation (2MB, 1MB, 64K, 4K) is
> >>> + * designed to match with the sizes often found in IOMMUs. Using larger 
> >>> order
> >>> + * pages instead of order 0 pages can significantly improve the 
> >>> performance of
> >>> + * many IOMMUs by reducing TLB pressure and time spent updating page 
> >>> tables.
> >>>   */
> >>> -static const unsigned int orders[] = {8, 4, 0};
> >>> +static const unsigned int orders[] = {9, 8, 4, 0};
> >>>  #define NUM_ORDERS ARRAY_SIZE(orders)
> >>>
> >>>  static int system_heap_set_page_decrypted(struct page *page)
> >>> --
> >>> 2.39.5
> >>>
> >>
>

Reply via email to