On Wed, Aug 26, 2026 at 12:44:40PM -0700, Sean Christopherson wrote:
> On Wed, Aug 26, 2026, Ackerley Tng wrote:
> > +
> >  static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start,
> >                                  size_t nr_pages, uint64_t attrs,
> >                                  pgoff_t *err_index)
> > @@ -624,7 +661,12 @@ static int __kvm_gmem_set_attributes(struct inode 
> > *inode, pgoff_t start,
> >  
> >     filter = to_private ? KVM_FILTER_SHARED : KVM_FILTER_PRIVATE;
> >     kvm_gmem_invalidate_start(inode, start, end, filter);
> > +
> > +   if (!to_private && kvm_arch_has_gmem_convert())
> > +           kvm_gmem_make_shared(inode, start, end);
> > +
> >     mas_store_prealloc(&mas, xa_mk_value(attrs));
> > +
> >     kvm_gmem_invalidate_end(inode, start, end);
> 
> The real reason I responded...
> 
> Thinking about the Secure AVIC mess made me realize zapping NPTs for SNP VMs 
> isn't
> strictly necessary in this path.  The PFN isn't changing, just the 
> attributes, and
> that's (obviously) tracked in the RMP.  KVM doesn't need to zap SPTEs to 
> induce a
> fault, because the mismatched C-bit vs. RMP status will cause an #NPF(RMP), 
> and
> AFAICT kvm_mmu_page_fault() will do the right thing.  A misbehaving guest 
> could
> continue to access the shared data (assuming we stick with lazy conversions), 
> but
> that should be fine?  E.g. it's not really any different than implicit 
> conversions.

Brilliant idea!

> 
> In other words, couldn't we do this (as an on-top optimization)?  The only 
> wrinkle
> I can think of is that it could delay reconstituion of a hugepage, especially 
> if
> we opted for eager conversion (because the guest wouldn't hit #NPFs to 
> trigger the
> hugepage promotion).
> 
> diff --git arch/x86/kvm/mmu/mmu.c arch/x86/kvm/mmu/mmu.c
> index 62f751952ad8..61f3e270ab61 100644
> --- arch/x86/kvm/mmu/mmu.c
> +++ arch/x86/kvm/mmu/mmu.c
> @@ -1670,6 +1670,7 @@ static bool __kvm_rmap_zap_gfn_range(struct kvm *kvm,
>  
>  bool kvm_unmap_gfn_range(struct kvm *kvm, struct kvm_gfn_range *range)
>  {
> +       unsigned long shared_private =  KVM_FILTER_SHARED | 
> KVM_FILTER_PRIVATE;
>         bool flush = false;
>  
>         /*
> @@ -1683,6 +1684,10 @@ bool kvm_unmap_gfn_range(struct kvm *kvm, struct 
> kvm_gfn_range *range)
>         lockdep_assert_once(kvm->mmu_invalidate_in_progress ||
>                             lockdep_is_held(&kvm->slots_lock));
>  
> +       if (gmem_in_place_conversion && !kvm_has_mirrored_tdp(kvm) &&
> +           ((range->attr_filter & shared_private) != shared_private))
> +               return false;
> +
>         if (kvm_memslots_have_rmaps(kvm))
>                 flush = __kvm_rmap_zap_gfn_range(kvm, range->slot,
>                                                  range->start, range->end,

Can this be a problem with gmem hugepages? Not sure how that is going to 
look like, so this may be covered in other ways.

But, as it exists today, if the guest converts part of a 2M page to 
shared, then this skips zapping the SPTEs from the call in 
kvm_gmem_invalidate_start(). sev_gmem_make_shared() then issues PSMASH 
to convert RMP entry to 4k entries and we end up with 2M NPT+4K RMP. If 
the guest then writes to any private page in that range, 
page_fault_can_be_fast() returns true, fast_page_fault() only checks 
permissions with spte_permission_fault() and does not do anything.  
sev_handle_rmp_fault() also does not issue a zap since it finds that the 
RMP entry is already 4k, and we end up in a loop.


- Naveen


Reply via email to