On 8/8/2026 5:52 AM, Ackerley Tng via B4 Relay wrote:
> From: Ackerley Tng <[email protected]>
>
> When memory in guest_memfd is converted from private to shared, the
> platform-specific state associated with the guest-private pages must be
> invalidated or cleaned up.
>
> Iterate over the folios in the affected range and call the
> kvm_arch_gmem_make_shared() hook for each PFN range. This allows
> architectures to update hardware metadata or encryption states to
> transition pages to the shared state.
>
> Invoke this helper after indicating to KVM's mmu code that an invalidation
> is in progress to stop in-flight page faults from succeeding.
>
> Omit support for calling the arch hook to make private, since SNP, the only
> implementer of the arch make-private hook today, would actually prefer
> making private only just before faulting memory into the NPTs.
>
> Calling the make-private arch hook would require iterating both bindings
> and the filemap to find the intersection of bindings and allocated
> folios. On top of that, SNP would need to figure out whether to actually
> make private based on whether the memory is about the be faulted, or
^
the -> to?
> whether it is a conversion.
>
[...]
>
> +#ifdef CONFIG_HAVE_KVM_ARCH_GMEM_CONVERT
> +static void kvm_gmem_make_shared(struct inode *inode, pgoff_t start, pgoff_t
> end)
> +{
> + struct folio_batch fbatch;
> + pgoff_t next = start;
> + int i;
> +
> + folio_batch_init(&fbatch);
> + while (filemap_get_folios(inode->i_mapping, &next, end - 1, &fbatch)) {
> + for (i = 0; i < folio_batch_count(&fbatch); ++i) {
> + struct folio *folio = fbatch.folios[i];
> + pgoff_t start_index, end_index;
> + kvm_pfn_t start_pfn;
> + kvm_pfn_t nr_pages;
> +
> + start_index = max(start, folio->index);
> + end_index = min(end, folio_next_index(folio));
> + /*
> + * end_index is either in folio or points to
> + * the first page of the next folio. Hence,
> + * all pages in range [start_index, end_index)
> + * are contiguous.
> + */
> + start_pfn = folio_file_pfn(folio, start_index);
> + nr_pages = end_index - start_index;
> +
> + kvm_arch_gmem_make_shared(start_pfn, nr_pages);
> + }
> +
> + folio_batch_release(&fbatch);
> + cond_resched();
> + }
> +}
> +#else
> +static void kvm_gmem_make_shared(struct inode *inode, pgoff_t start, pgoff_t
> end) {}
> +#endif
> +
> static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start,
> size_t nr_pages, uint64_t attrs,
> pgoff_t *err_index)
> @@ -599,7 +636,12 @@ static int __kvm_gmem_set_attributes(struct inode
> *inode, pgoff_t start,
>
> filter = to_private ? KVM_FILTER_SHARED : KVM_FILTER_PRIVATE;
> kvm_gmem_invalidate_start(inode, start, end, filter);
> +
> + if (!to_private)
> + kvm_gmem_make_shared(inode, start, end);
If both KVM_AMD_SEV and KVM_INTEL_TDX are enabled, HAVE_KVM_ARCH_GMEM_CONVERT
will be enabled and the logic in kvm_gmem_make_shared() introduces unnecessary
overhead for TDX.
Not sure about CSPs, but in a standard distribution kernel, it's very likely
that both are enabled, right? Should kvm_gmem_make_shared() do some
optimization or the overhead is relative small in the conversion to shared
path so that the optimization is not worth it?
> +
> mas_store_prealloc(&mas, xa_mk_value(attrs));
> +
> kvm_gmem_invalidate_end(inode, start, end);
> out:
> filemap_invalidate_unlock(mapping);
>