Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider: - [High] mm: unbalanced mmu_notifier in hmm_range_fault during migration --
commit 3e05399d761177583ed386d43f42e228c070f447 Author: Mika Penttilä <[email protected]> mm: enable device page migration from HMM pagewalk This patch integrates HMM pagewalk machinery into the first phase of device page migration by enabling migration in hmm_range_fault(). It updates migrate_vma_setup() to use the HMM pagewalk path and removes the older migrate_vma_collect() paths. Two new flags, MIGRATE_VMA_FAULT and MIGRATE_VMA_WRITE, are introduced for faulting missing pages and requesting write access. > diff --git a/mm/migrate_device.c b/mm/migrate_device.c > index 43e571e82ccad..1e1a592e12672 100644 > --- a/mm/migrate_device.c > +++ b/mm/migrate_device.c [ ... ] > @@ -739,10 +241,20 @@ static void migrate_vma_unmap(struct migrate_vma > *migrate) > */ > int migrate_vma_setup(struct migrate_vma *args) > { > + int ret; > long nr_pages = (args->end - args->start) >> PAGE_SHIFT; > + struct hmm_range range = { > + .notifier = NULL, > + .hmm_pfns = args->src, > + .dev_private_owner = args->pgmap_owner, [Severity: High] Does this code cause a permanent MMU notifier imbalance when args->pgmap_owner is NULL? If a caller invokes migrate_vma_setup() with args->pgmap_owner set to NULL (which is standard for migrations not involving device-private memory, such as PowerPC KVM secure guest page-in), range.dev_private_owner is also NULL. In mm/hmm.c:hmm_vma_capture_migrate_range(), the code checks: if (!hmm_vma_walk->mmu_range.owner) { mmu_notifier_range_init_owner(&hmm_vma_walk->mmu_range, MMU_NOTIFY_MIGRATE, 0, walk->vma->vm_mm, start, end, range->dev_private_owner); mmu_notifier_invalidate_range_start(&hmm_vma_walk->mmu_range); } Because mmu_range is zero-initialized, this condition is true. It calls mmu_notifier_range_init_owner(), which assigns the NULL dev_private_owner to hmm_vma_walk->mmu_range.owner. mmu_notifier_invalidate_range_start() is then called. Additionally, if the range spans multiple VMAs, invalidate_range_start is called multiple times because the sentinel remains NULL. Later, in mm/hmm.c:hmm_vma_post_range_fault(): if (hmm_select_migrate(range) && hmm_vma_walk->mmu_range.owner) { mmu_notifier_invalidate_range_end(&hmm_vma_walk->mmu_range); } Because mmu_range.owner remains NULL, this condition evaluates to false, completely skipping the required mmu_notifier_invalidate_range_end() call. A skipped invalidate_range_end() leaves the invalidation count permanently elevated, causing any subsequent operation on that memory range requiring mmu_interval_read_retry() or VMA teardown to block indefinitely. > + .migrate = args, > + .default_flags = HMM_PFN_REQ_MIGRATE > + }; > > args->start &= PAGE_MASK; > args->end &= PAGE_MASK; [ ... ] -- Sashiko AI review · https://sashiko.dev/#/patchset/[email protected]?part=10
