From: Yu Zhang <[email protected]> Sent: Friday, August 21, 2026
6:27 AM
>
> Add a para-virtualized IOMMU driver for Linux guests running on Hyper-V.
> This driver implements stage-1 IO translation within the guest OS.
> It integrates with the Linux IOMMU core, utilizing Hyper-V hypercalls
> for:
> - Capability discovery
> - Domain allocation, configuration, and deallocation
> - Device attachment and detachment
> - IOTLB invalidation
>
> The driver constructs x86-compatible stage-1 IO page tables in the
> guest memory using consolidated IO page table helpers. This allows
> the guest to manage stage-1 translations independently of vendor-
> specific drivers (like Intel VT-d or AMD IOMMU).
>
> Hyper-V consumes this stage-1 IO page table when a device domain is
> created and configured, and nests it with the host's stage-2 IO page
> tables, therefore eliminating the VM exits for guest IOMMU mapping
> operations. For unmapping operations, VM exits to perform the IOTLB
> flush are still unavoidable.
>
> Guest hibernation and resume are not supported by this initial
> implementation. The guest-owned stage-1 page tables reside in guest
> memory and are preserved in the hibernation image. However, the image
> does not by itself preserve the Hyper-V device-domain objects or their
> device attachments, and the current pvIOMMU interface does not define
> whether or how Hyper-V restores that state.
>
> To identify a device in its hypercall interface, the driver looks up the
> logical device ID prefix registered for the device's PCI domain (see the
> logical device ID registry in hv_common.c) and combines it with the PCI
> function number of the endpoint device.
>
> Co-developed-by: Wei Liu <[email protected]>
> Signed-off-by: Wei Liu <[email protected]>
> Co-developed-by: Easwar Hariharan <[email protected]>
> Signed-off-by: Easwar Hariharan <[email protected]>
> Signed-off-by: Yu Zhang <[email protected]>
> Reviewed-by: Jacob Pan <[email protected]>
> ---
> arch/x86/hyperv/hv_init.c | 4 +
> arch/x86/include/asm/mshyperv.h | 4 +
> drivers/iommu/Kconfig | 1 +
> drivers/iommu/hyperv/Kconfig | 16 +
> drivers/iommu/hyperv/Makefile | 1 +
> drivers/iommu/hyperv/hv-iommu-guest.c | 620 ++++++++++++++++++++++++++
> drivers/iommu/hyperv/iommu.h | 49 ++
> 7 files changed, 695 insertions(+)
> create mode 100644 drivers/iommu/hyperv/Kconfig
> create mode 100644 drivers/iommu/hyperv/hv-iommu-guest.c
> create mode 100644 drivers/iommu/hyperv/iommu.h
>
[snip]
> +
> +static int hv_iommu_attach_dev(struct iommu_domain *domain, struct device
> *dev,
> + struct iommu_domain *old)
> +{
> + u64 status;
> + u32 prefix;
> + unsigned long flags;
> + struct pci_dev *pdev;
> + struct hv_input_attach_device_domain *input;
> + struct hv_iommu_domain *hv_domain = to_hv_iommu_domain(domain);
> + int ret;
> +
> + pdev = to_pci_dev(dev);
> + dev_dbg(dev, "attaching to domain %d\n",
> + hv_domain->device_domain.domain_id.id);
> +
> + ret = hv_pci_lookup_dev_id(pci_domain_nr(pdev->bus), &prefix);
> + if (ret) {
> + dev_err(&pdev->dev, "no IOMMU registration for vPCI bus\n");
Nit: Given that the device id registration is independent of IOMMUs,
this message seems a bit off. Maybe just drop "IOMMU"?
> + return ret;
> + }
> +
> + local_irq_save(flags);
> +
> + input = *this_cpu_ptr(hyperv_pcpu_input_arg);
> + memset(input, 0, sizeof(*input));
> + input->device_domain = hv_domain->device_domain;
> + input->device_id.as_uint64 = (u64)prefix | PCI_FUNC(pdev->devfn);
> + status = hv_do_hypercall(HVCALL_ATTACH_DEVICE_DOMAIN, input, NULL);
> +
> + local_irq_restore(flags);
> +
> + if (!hv_result_success(status) &&
> + hv_result(status) != HV_STATUS_DEVICE_ALREADY_IN_DOMAIN) {
> + hv_status_err(status, "HVCALL_ATTACH_DEVICE_DOMAIN failed\n");
> + return hv_result_to_errno(status);
> + }
> +
> + if (domain != &hv_blocking_domain.domain &&
> + !pdev->ats_enabled &&
> + hv_iommu_ats_supported(hv_iommu_device->cap) &&
Nit: In other places where the hv_domain is available,
hv_domain->hv_iommu is used instead of directly accessing
hv_iommu_device. That's a good practice that limits references
to the static variable hv_iommu_device in case future enhancements
allow multiple IOMMUs in a VM. I'd suggest doing the same here.
> + pci_ats_supported(pdev))
> + pci_enable_ats(pdev, PAGE_SHIFT);
> +
> + return 0;
> +}
> +
[snip]
> +
> +static struct iommu_device *hv_iommu_probe_device(struct device *dev)
> +{
> + struct hv_iommu_endpoint *vdev;
> + struct hv_output_get_logical_device_property device_iommu_property =
> {0};
> +
> + if (!dev_is_pci(dev))
> + return ERR_PTR(-ENODEV);
> +
> + if (hv_iommu_get_logical_device_property(dev,
> +
> HV_LOGICAL_DEVICE_PROPERTY_PVIOMMU,
> + &device_iommu_property) ||
> + !(device_iommu_property.device_iommu & HV_DEVICE_IOMMU_ENABLED))
> + return ERR_PTR(-ENODEV);
> +
> + vdev = kzalloc_obj(*vdev, GFP_KERNEL);
> + if (!vdev)
> + return ERR_PTR(-ENOMEM);
> +
> + vdev->dev = dev;
> + vdev->hv_iommu = hv_iommu_device;
> + dev_iommu_priv_set(dev, vdev);
This struct hv_iommu_endpoint is set up, but doesn't seem to be
used for anything except to free it. Am I missing some reason that
it exists?
> +
> + return &vdev->hv_iommu->iommu;
> +}
> +
> +static void hv_iommu_release_device(struct device *dev)
> +{
> + struct hv_iommu_endpoint *vdev = dev_iommu_priv_get(dev);
> +
> + kfree(vdev);
> +}
> +
Modulo the few nits above, this all looks good to me. But I don't have
a lot of familiarity with the details of IOMMU operation and the
generic IOMMU page tables, so my review is somewhat limited.
Reviewed-by: Michael Kelley <[email protected]>