On 8/8/2026 5:52 AM, Ackerley Tng via B4 Relay wrote:
> From: Ackerley Tng <[email protected]>
> 
> When memory in guest_memfd is converted from private to shared, the
> platform-specific state associated with the guest-private pages must be
> invalidated or cleaned up.
> 
> Iterate over the folios in the affected range and call the
> kvm_arch_gmem_make_shared() hook for each PFN range. This allows
> architectures to update hardware metadata or encryption states to
> transition pages to the shared state.
> 
> Invoke this helper after indicating to KVM's mmu code that an invalidation
> is in progress to stop in-flight page faults from succeeding.
> 
> Omit support for calling the arch hook to make private, since SNP, the only
> implementer of the arch make-private hook today, would actually prefer
> making private only just before faulting memory into the NPTs.
> 
> Calling the make-private arch hook would require iterating both bindings
> and the filemap to find the intersection of bindings and allocated
> folios. On top of that, SNP would need to figure out whether to actually
> make private based on whether the memory is about the be faulted, or
                                                     ^
the -> to?

> whether it is a conversion.
> 

[...]

>  
> +#ifdef CONFIG_HAVE_KVM_ARCH_GMEM_CONVERT
> +static void kvm_gmem_make_shared(struct inode *inode, pgoff_t start, pgoff_t 
> end)
> +{
> +     struct folio_batch fbatch;
> +     pgoff_t next = start;
> +     int i;
> +
> +     folio_batch_init(&fbatch);
> +     while (filemap_get_folios(inode->i_mapping, &next, end - 1, &fbatch)) {
> +             for (i = 0; i < folio_batch_count(&fbatch); ++i) {
> +                     struct folio *folio = fbatch.folios[i];
> +                     pgoff_t start_index, end_index;
> +                     kvm_pfn_t start_pfn;
> +                     kvm_pfn_t nr_pages;
> +
> +                     start_index = max(start, folio->index);
> +                     end_index = min(end, folio_next_index(folio));
> +                     /*
> +                      * end_index is either in folio or points to
> +                      * the first page of the next folio. Hence,
> +                      * all pages in range [start_index, end_index)
> +                      * are contiguous.
> +                      */
> +                     start_pfn = folio_file_pfn(folio, start_index);
> +                     nr_pages = end_index - start_index;
> +
> +                     kvm_arch_gmem_make_shared(start_pfn, nr_pages);
> +             }
> +
> +             folio_batch_release(&fbatch);
> +             cond_resched();
> +     }
> +}
> +#else
> +static void kvm_gmem_make_shared(struct inode *inode, pgoff_t start, pgoff_t 
> end) {}
> +#endif
> +
>  static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start,
>                                    size_t nr_pages, uint64_t attrs,
>                                    pgoff_t *err_index)
> @@ -599,7 +636,12 @@ static int __kvm_gmem_set_attributes(struct inode 
> *inode, pgoff_t start,
>  
>       filter = to_private ? KVM_FILTER_SHARED : KVM_FILTER_PRIVATE;
>       kvm_gmem_invalidate_start(inode, start, end, filter);
> +
> +     if (!to_private)
> +             kvm_gmem_make_shared(inode, start, end);

If both KVM_AMD_SEV and KVM_INTEL_TDX are enabled, HAVE_KVM_ARCH_GMEM_CONVERT
will be enabled and the logic in kvm_gmem_make_shared() introduces unnecessary
overhead for TDX.
Not sure about CSPs, but in a standard distribution kernel, it's very likely
that both are enabled, right? Should kvm_gmem_make_shared() do some
optimization or the overhead is relative small in the conversion to shared
path so that the optimization is not worth it? 


> +
>       mas_store_prealloc(&mas, xa_mk_value(attrs));
> +
>       kvm_gmem_invalidate_end(inode, start, end);
>  out:
>       filemap_invalidate_unlock(mapping);
> 


Reply via email to