This series tightens the alignment requirements for buffers that are shared between confidential-computing guests and the host, and adds a common allocator for host-shared memory.
When a guest runs with private memory, buffers shared with the hypervisor are not only accessed by the guest. They are also accessed by the host kernel, and the host may manage the corresponding shared/private state at a granularity larger than the guest page size. This matters for CCA systems where the Realm stage-2 mappings managed by the RMM can still operate at 4K granularity, while the non-secure host may manage the IPA state change at a larger page size, for example 64K. In that case, allowing a guest to convert and share only a 4K subrange of a host-managed granule is unsafe. Architectures such as Arm can detect incorrect accesses to Realm physical address space PFNs through GPC faults. However, relying on that as the only line of defence is fragile and can still lead to kernel crashes. The risk is especially visible for shared buffers that are later mmapped into userspace, such as guest_memfd or dma-buf backed allocations. Once userspace can access the mapping, the kernel cannot guarantee that applications will only touch the intended 4K region rather than the whole host page mapped into their address space. Those userspace addresses may also be passed back into the kernel and accessed through the linear map, resulting in a GPC fault. To avoid this, host-shared buffers must satisfy two constraints: - the address must be aligned to the CoCo shared-granule size - the size must be a multiple of that granule size The series adds a common CoCo shared-memory layer for enforcing these constraints. It provides shared-granule geometry and range-validation helpers, byte-oriented private/shared transition helpers, and alloc_cc_shared_pages() with a node-aware variant. The allocator rounds a request to the architecture shared granule, allocates suitably aligned contiguous pages, transitions the complete allocation to shared state, and returns the transitioned size alongside the page. The corresponding free helper restores the complete allocation to private state before returning it to the buddy allocator. If private state cannot be restored safely, the allocation is deliberately leaked rather than returning potentially shared memory for unrelated use. Since a private-to-shared transition may modify memory contents, __GFP_ZERO is applied after the transition. The generic shared-granule size defaults to PAGE_SIZE. For arm64 CCA, the series queries the host IPA state change alignment through the Realm Host Interface, caches it during Realm initialization, and exposes it through the arm64 memory-encryption operations. The common allocator is used for host-shared allocations whose backing is owned by an individual caller: - GIC ITS command queues and tables - dma-direct allocations backed by CMA or the page allocator - backing allocations for the CoCo atomic DMA pools - dma-buf system_cc_shared heap allocations Hyper-V users of set_memory_encrypted() and set_memory_decrypted() are not changed by this series. Those paths are not currently used by the arm64 CCA code path, and therefore are not part of the arm64 CCA IPA state change alignment problem addressed here. NOTE: I have not added explicit MAINTAINERS entries for mm/cc_shared.c and include/linux/cc_shared.h, as I am unsure whether we need a separate section for common CoCo-related files. I will add the entries based on feedback. Changes from v7: https://lore.kernel.org/all/[email protected] * Add the following new patches: * "irqchip/gic-v3-its: Preallocate VPE L1 tables" * "mm: Assert CoCo shared allocations may sleep" * "mm: Zero memory during shared memory transitions" * Drop the arm64 RHI and shared granule size patches so that the series can be rebased on top of upstream to enable Shashiko review. Changes from v6: https://lore.kernel.org/all/[email protected] * Add a common allocator and geometry/transition helpers for CoCo host-shared memory. * Convert GIC ITS, dma-direct, atomic DMA pools, and the dma-buf system_cc_shared heap to the common allocator. * Limit dma-buf scatterlist entries to the requested buffer size so rounded backing is not exposed to importers. Changes from v5: https://lore.kernel.org/all/[email protected] * Rebased to latest kernel * Drop patch arm64: realm: Move Realm memory encryption ops to RSI code Changes from v4: https://lore.kernel.org/all/[email protected] * Rename the helpers to use CoCo terminology (mem_cc_shared_granule_size() / mem_cc_align_to_shared_granule() instead of mem_decrypt_granule_size() / mem_decrypt_align()). * Use __DMA_ATTR_ALLOC_CC_SHARED to pass CoCo shared allocation requirements down to CMA-based allocation helpers. * Add validation for restricted DMA pools to reject pools that are not aligned to the shared granule size. * Add dma-buf system heap handling for cc-shared buffers. * Split the previous combined DMA/SWIOTLB/ITS change into smaller subsystem patches covering ITS, DMA direct, SWIOTLB, restricted DMA pools, dma-buf system heap, and arm64 Realm support. * Rework arm64 Realm support by moving Realm memory encryption ops into RSI code and exposing the CCA shared granule size through arm64_mem_crypt_ops. Changes from v3: https://lore.kernel.org/all/[email protected] * Fix build error reported by kernel test robot <[email protected]> Changes from v2: https://lore.kernel.org/all/[email protected] * Rebase to latest kernel * Consider swiotlb always decrypted and don't align when allocating from swiotlb. Changes from v1: * Rename the helper to mem_encrypt_align * Improve the commit message * Handle DMA allocations from contiguous memory * Handle DMA allocations from the pool * swiotlb is still considered unencrypted. Support for an encrypted swiotlb pool is left as TODO and is independent of this series. Cc: Andrew Morton <[email protected]> Cc: Baoquan He <[email protected]> Cc: Mike Rapoport <[email protected]> Cc: Pasha Tatashin <[email protected]> Cc: Pratyush Yadav <[email protected]> Cc: Catalin Marinas <[email protected]> Cc: "Christian König" <[email protected]> Cc: Jason Gunthorpe <[email protected]> Cc: Joerg Roedel (AMD) <[email protected]> Cc: Marc Zyngier <[email protected]> Cc: Marek Szyprowski <[email protected]> Cc: Robin Murphy <[email protected]> Cc: Steven Price <[email protected]> Cc: Sumit Semwal <[email protected]> Cc: Suzuki K Poulose <[email protected]> Cc: Thomas Gleixner <[email protected]> Cc: Will Deacon <[email protected]> Cc: Russell King <[email protected]> Cc: Benjamin Gaignard <[email protected]> Cc: Brian Starkey <[email protected]> Cc: John Stultz <[email protected]> Cc: Mark Rutland <[email protected]> Cc: Radu Rendec <[email protected]> Cc: "T.J. Mercier" <[email protected]> Cc: Madhavan Srinivasan <[email protected]> Cc: Michael Ellerman <[email protected]> Cc: Nicholas Piggin <[email protected]> Cc: Christophe Leroy (CS GROUP) <[email protected]> Cc: Ritesh Harjani (IBM) <[email protected]> Cc: Shrikanth Hegde <[email protected]> Cc: Alexander Gordeev <[email protected]> Cc: Gerald Schaefer <[email protected]> Cc: Heiko Carstens <[email protected]> Cc: Vasily Gorbik <[email protected]> Cc: Christian Borntraeger <[email protected]> Cc: Sven Schnelle <[email protected]> Cc: Ingo Molnar <[email protected]> Cc: Borislav Petkov <[email protected]> Cc: Dave Hansen <[email protected]> Cc: [email protected] Cc: H. Peter Anvin <[email protected]> Cc: Kiryl Shutsemau <[email protected]> Cc: Rick Edgecombe <[email protected]> Cc: K. Y. Srinivasan <[email protected]> Cc: Haiyang Zhang <[email protected]> Cc: Wei Liu <[email protected]> Cc: Dexuan Cui <[email protected]> Cc: Long Li <[email protected]> Cc: Paolo Bonzini <[email protected]> Cc: Vitaly Kuznetsov <[email protected]> Cc: Andy Lutomirski <[email protected]> Cc: Peter Zijlstra <[email protected]> Cc: [email protected] Cc: [email protected] Cc: [email protected] Cc: [email protected] Cc: [email protected] Cc: [email protected] Cc: [email protected] Aneesh Kumar K.V (Arm) (14): mm: Add an allocator for CoCo shared memory mm: Zero memory during shared memory transitions irqchip/gic-v3-its: Resolve the default NUMA node explicitly irqchip/gic-v3-its: Allocate shared tables using CoCo shared memory allocator dma-contiguous: Derive shared alignment from DMA attributes dma-pool: Allocate CoCo atomic pools using CoCo shared memory allocator dma-direct: Align CoCo shared DMA allocations to the shared granule size swiotlb: Align shared IO TLB pools to the shared granule size swiotlb: Reject misaligned restricted DMA pools for CoCo guests dma-buf: system_heap: Limit scatterlist entries to the buffer size dma-buf: system_heap: Allocate shared buffers using CoCo shared memory allocator swiotlb: Make rounded shared pool capacity allocatable mm: Assert CoCo shared allocations may sleep irqchip/gic-v3-its: Preallocate VPE L1 tables arch/arm/mm/dma-mapping.c | 5 +- arch/arm64/mm/pageattr.c | 3 + arch/powerpc/platforms/pseries/svm.c | 2 + arch/s390/mm/init.c | 3 + arch/x86/coco/tdx/tdx.c | 3 + arch/x86/hyperv/hv_init.c | 6 +- arch/x86/hyperv/ivm.c | 4 + arch/x86/kernel/kvmclock.c | 6 +- arch/x86/mm/mem_encrypt_amd.c | 4 + drivers/dma-buf/heaps/system_heap.c | 128 +++++----- drivers/hv/connection.c | 41 ++-- drivers/hv/hv.c | 11 +- drivers/hv/hv_common.c | 2 - drivers/iommu/dma-iommu.c | 2 +- drivers/irqchip/irq-gic-v3-its.c | 153 +++++++++--- drivers/irqchip/irq-gic-v3.c | 4 +- drivers/virt/coco/pkvm-guest/arm-pkvm-guest.c | 3 + include/linux/cc_shared.h | 39 +++ include/linux/dma-map-ops.h | 9 +- include/linux/irqchip/arm-gic-v3.h | 3 +- kernel/dma/contiguous.c | 41 +++- kernel/dma/direct.c | 73 ++++-- kernel/dma/ops_helpers.c | 2 +- kernel/dma/pool.c | 25 +- kernel/dma/swiotlb.c | 80 ++++-- kernel/kexec_file.c | 3 +- mm/Makefile | 1 + mm/cc_shared.c | 232 ++++++++++++++++++ 28 files changed, 682 insertions(+), 206 deletions(-) create mode 100644 include/linux/cc_shared.h create mode 100644 mm/cc_shared.c base-commit: 704340f1cd0dcef829eb62f5b48ae95a2ce17bdf -- 2.43.0
