With 4KB base pages, the system heap allocates buffers in 1MB, 64KB and 4KB chunks. The conventional Intel VT-d second-stage and AMD-Vi v2 page-table formats, and Arm SMMU's 64-bit long-descriptor format with a 4KB translation granule, define 4KB pages and 2MB/1GB large-page mappings, but no 1MB leaf. A 1MB chunk requires 256 4KB leaf entries unless it can be combined with adjacent chunks into a suitably aligned larger mapping.
Adding a 2MB allocation order provides naturally aligned, physically contiguous chunks matching the 2MB leaf size, without relying on this being accidental for separate smaller allocations. The new cost is one failed order-9 attempt (keeping current large allocation semantics) per buffer when 2MB pages are exhausted. Only x86 gets the new order. The existing orders are optimized for 4K-page Arm platforms, so they are left as they are. Document the current differences and the effect of fragmentation. Two benchmarks, measured on an AMD EPYC 7313P. (i) Mapping, into an idle NVMe function's translated DMA-FQ domain, measuring map/unmap, with the v1 page table restricted to 4K/2M/1G (amd_iommu=v2_pgsizes_only), a 1GB buffer goes from 1024 1MB chunks, only one of whose 2MB windows was superpage-mappable in that run, to 512 2MB chunks with all 512 mappable. dma_buf_map_attachment() costs decrease by ~16x on average with ~32x for the worst cases. Unmapping that buffer drops from ~500 us to 1.25 us. (ii) A microbench that measures the cost of DMA_HEAP_IOCTL_ALLOC+close across various thread counts decreases by factors of ~3-7x alleviating allocator's zone->lock contention by being PMD order and so pcpu list eligible. Once the buffer size is large enough then the cost of the zeroing takes over. Reviewed-by: T.J. Mercier <[email protected]> Signed-off-by: Davidlohr Bueso <[email protected]> --- Documentation/userspace-api/dma-buf-heaps.rst | 4 ++++ drivers/dma-buf/heaps/system_heap.c | 24 ++++++++++++++----- 2 files changed, 22 insertions(+), 6 deletions(-) diff --git a/Documentation/userspace-api/dma-buf-heaps.rst b/Documentation/userspace-api/dma-buf-heaps.rst index f56b743cdb36..03452b653b7c 100644 --- a/Documentation/userspace-api/dma-buf-heaps.rst +++ b/Documentation/userspace-api/dma-buf-heaps.rst @@ -15,6 +15,10 @@ A heap represents a specific allocator. The Linux kernel currently supports the following heaps: - The ``system`` heap allocates virtually contiguous, cacheable, buffers. + They are assembled from the largest free page blocks the heap can get + without reclaiming or compacting memory, so the physical layout of a + buffer, and with it how efficiently an IOMMU can map it, depends on how + fragmented free memory is at allocation time. - The ``system_cc_shared`` heap allocates virtually contiguous, cacheable, buffers using shared (decrypted) memory. It is only present on diff --git a/drivers/dma-buf/heaps/system_heap.c b/drivers/dma-buf/heaps/system_heap.c index c8959eadc71d..07721c136c71 100644 --- a/drivers/dma-buf/heaps/system_heap.c +++ b/drivers/dma-buf/heaps/system_heap.c @@ -55,14 +55,26 @@ struct dma_heap_attachment { #define HIGH_ORDER_GFP (((GFP_HIGHUSER | __GFP_ZERO | __GFP_NOWARN \ | __GFP_NORETRY) & ~__GFP_RECLAIM) \ | __GFP_COMP) -static gfp_t order_flags[] = {HIGH_ORDER_GFP, HIGH_ORDER_GFP, LOW_ORDER_GFP}; /* - * The selection of the orders used for allocation (1MB, 64K, 4K) is designed - * to match with the sizes often found in IOMMUs. Using order 4 pages instead - * of order 0 pages can significantly improve the performance of many IOMMUs - * by reducing TLB pressure and time spent updating page tables. + * The selection of the orders used for allocation is designed to match with + * the sizes often found in IOMMUs. Using large order pages instead of order 0 + * pages can significantly improve performance by reducing TLB pressure and + * time spent updating page tables. + * + * x86 uses 2MB, 1MB, 64K and 4K. VT-d and AMD-Vi v2 have no 1MB leaf and + * AMD-Vi v1 encodes one as 256 replicated 4K entries, so 2MB is the smallest + * chunk that installs a single entry. 1MB stays as a fallback for when 2MB + * blocks are gone but 1MB ones are not. + * + * Everywhere else uses 1MB, 64K and 4K, which is optimized for 4K-page Arm + * platforms. Here stable performance is preferred to keep the system heap + * driver independent, and be less susceptible to fragmentation. */ +#ifdef CONFIG_X86 +static const unsigned int orders[] = {9, 8, 4, 0}; +#else static const unsigned int orders[] = {8, 4, 0}; +#endif #define NUM_ORDERS ARRAY_SIZE(orders) static int system_heap_set_page_decrypted(struct page *page) @@ -387,7 +399,7 @@ static struct page *alloc_largest_available(unsigned long size, continue; if (max_order < orders[i]) continue; - flags = order_flags[i]; + flags = orders[i] ? HIGH_ORDER_GFP : LOW_ORDER_GFP; if (mem_accounting) flags |= __GFP_ACCOUNT; page = alloc_pages(flags, orders[i]); -- 2.39.5
