On Thu, 2026-09-10 at 11:58 +0200, David Hildenbrand (Arm) wrote: > On 8/5/26 08:40, Shivank Garg wrote: > > guest_memfd folios are currently always marked unmovable, so the kernel > > cannot > > perform memory compaction, offlining, etc. This is unavoidable for > > confidential VMs (SEV-SNP, TDX), since memory is encrypted and copying it > > needs firmware assistance. However, for non-confidential VMs (like > > Firecracker), we can migrate the folios. > > Yes. > > > > > This series enables folio migration for non-confidential guest_memfd and > > also lays the groundwork for migrating confidential guest_memfd later. > > Once firmware-assisted copying support is available, those VMs can be > > made movable, the confidential folio content can be copied separately, > > and the destination folio marked with FOLIO_CONTENT_COPIED[4] so > > __migrate_folio() skips the host-side folio_mc_copy(). > > > > Testing > > ------- > > Host: 7.2-rc6+(c21bb419386) + this, AMD EPYC ZEN 3, 2 NUMA nodes > > > > - KVM selftest: allocate folios on node 0, migrate them to node 1 and > > back and verify resulting NUMA node and the folio contents at each > > step. > > > > - Firecracker [1]: booted a microVM backed by guest_memfd. While the > > guest was running, forced host-side migration of its folios via > > migratepages(8) and explicit move_pages(2) of guest_memfd > > pages. Verify with /proc/firecracker_pid/numa_maps. > > > > Notes > > ----- > > - Sashiko pointed out a pre-existing ABBA deadlock between > > kvm_gmem_error_folio() and truncation. It's being addressed separately > > by Hao Zhang. [2][3] > > > > [1] > > https://github.com/firecracker-microvm/firecracker/tree/feature/secret-hiding > > In builder.rs, add GUEST_MEMFD_FLAG_MIGRATABLE to bit-2 and pass it > > instead > > of GUEST_MEMFD_FLAG_NO_DIRECT_MAP to vm.create_guest_memfd(). > > [2] https://lore.kernel.org/all/[email protected]/ > > [3] > > https://sashiko.dev/#/patchset/20260611-shivank-gmem-migrate-v1-0-2d266bfc6f95%40amd.com > > [4] > > https://lore.kernel.org/all/[email protected] > > > > Signed-off-by: Shivank Garg <[email protected]> > > --- > > Changes in v3: > > - Fix unbalanced mmu_invalidate_in_progress count unbinding dying > > guest_memfd. (Sashiko) > > - Fix maxnode handling in xapic_ipi_test selftest. > > - Add GUEST_MEMFD_FLAG_MIGRATABLE documentation > > - Replace open-coded sizeof() * 8 calculation with BITS_PER_TYPE() > > - Add get_numa_mem_nodes() and use MPOL_F_MEMS_ALLOWED for allowed NUMA > > ndoes > > instead of hardcoded NUMA node IDs. (Sashiko) > > - Extend migration selftest to verify rejection without MIGRATABLE flag and > > move repeated checks into common helpers. > > - Drop RFC tag. > > - Link to v2: > > https://lore.kernel.org/r/[email protected] > > > > Changes in v2: > > - Make folio migration opt-in through GUEST_MEMFD_FLAG_MIGRATABLE, > > preserving unmovable behavior if userspace don't explictly ask. > > (Alexandru, David, Sean) > > - Add kvm_arch_supports_gmem_migration() so arch can control whether > > GUEST_MEMFD_FLAG_MIGRATABLE is advertised. > > - Allocate movable folios with GFP_HIGHUSER_MOVABLE. (David) > > - Keep guest_memfd unevictable. (David, Sashiko, Sean) > > - Split migrate_folio() implementation and enablement as separate patches. > > - Update selftest with new flag. > > - Link to v1: > > https://lore.kernel.org/r/[email protected] > > > > --- > > Shivank Garg (9): > > KVM: guest_memfd: take the invalidate lock when unbinding a dying file > > mm: split AS_UNMOVABLE back out of AS_INACCESSIBLE > > KVM: guest_memfd: implement folio migration for non-confidential VMs > > KVM: guest_memfd: add GUEST_MEMFD_FLAG_MIGRATABLE > > > As discussed, we should for now just always enable it and not expose such a > flag. The hope is that use cases that need migration disabled can just find a > way for kvm to tell guest_memfd about that internally ... or we'll add a > GUEST_MEMFD_FLAG_UNMIGRATABLE or whatever later. >
Sure, makes sense. I'll drop GUEST_MEMFD_FLAG_UNMIGRATABLE. > With migration in place, as also discussed, it would be interesting to > investigate what it would take for these shared-only (no conversion) and > migratable guest_memfd instances to support THPs. > > I'd assume we might have to teach > > 1) guest_memfd / KVM parts about this, although I recall that most of it > should > be there > > 2) Unlock guest_memfd in khugeapged > > We disallowed khugepaged entirely in commit > > commit dd085fe9a8ebfc5d10314c60452db38d2b75e609 > Author: Deepanshu Kartikey <[email protected]> > Date: Sat Feb 14 05:45:35 2026 +0530 > > mm: thp: deny THP for files on anonymous inodes > > file_thp_enabled() incorrectly allows THP for files on anonymous inodes > (e.g. guest_memfd and secretmem). These files are created via > alloc_file_pseudo(), which does not call get_write_access() and leaves > inode->i_writecount at 0. Combined with S_ISREG(inode->i_mode) being > true, they appear as read-only regular files when > CONFIG_READ_ONLY_THP_FOR_FS is enabled, making them eligible for THP > collapse. > > It will also be interesting to figure out how well khugepaged would collapse > guest_memfd when most folios are not actually faulted into the user page > tables. > > collapse_scan_file() does not seem to worry about whether folios are actually > mapped, just if they are present in the page cache. > > Which might mean that as long as guestmemfd is mmap'ed, it might just work. Thanks for the pointers. I'll think about khugepaged implementation for this. Best regards, Shivank

