On Tue, Sep 29, 2026 at 03:10:41PM +0530, Naveen N Rao wrote:
> On Wed, Aug 26, 2026 at 12:44:40PM -0700, Sean Christopherson wrote:
> > On Wed, Aug 26, 2026, Ackerley Tng wrote:
> > > +
> > > static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start,
> > > size_t nr_pages, uint64_t attrs,
> > > pgoff_t *err_index)
> > > @@ -624,7 +661,12 @@ static int __kvm_gmem_set_attributes(struct inode
> > > *inode, pgoff_t start,
> > >
> > > filter = to_private ? KVM_FILTER_SHARED : KVM_FILTER_PRIVATE;
> > > kvm_gmem_invalidate_start(inode, start, end, filter);
> > > +
> > > + if (!to_private && kvm_arch_has_gmem_convert())
> > > + kvm_gmem_make_shared(inode, start, end);
> > > +
> > > mas_store_prealloc(&mas, xa_mk_value(attrs));
> > > +
> > > kvm_gmem_invalidate_end(inode, start, end);
> >
> > The real reason I responded...
> >
> > Thinking about the Secure AVIC mess made me realize zapping NPTs for SNP
> > VMs isn't
> > strictly necessary in this path. The PFN isn't changing, just the
> > attributes, and
> > that's (obviously) tracked in the RMP. KVM doesn't need to zap SPTEs to
> > induce a
> > fault, because the mismatched C-bit vs. RMP status will cause an #NPF(RMP),
> > and
> > AFAICT kvm_mmu_page_fault() will do the right thing. A misbehaving guest
> > could
> > continue to access the shared data (assuming we stick with lazy
> > conversions), but
> > that should be fine? E.g. it's not really any different than implicit
> > conversions.
>
> Brilliant idea!
>
> >
> > In other words, couldn't we do this (as an on-top optimization)? The only
> > wrinkle
> > I can think of is that it could delay reconstituion of a hugepage,
> > especially if
> > we opted for eager conversion (because the guest wouldn't hit #NPFs to
> > trigger the
> > hugepage promotion).
I suppose gmem could selectively zap if a range switches from mixed to
all-private. For non-confidential guests, it might even make sense for
mixed to all-shared if there's some way to avoid the forced
hugetlb-splitting in that scenario.
> >
> > diff --git arch/x86/kvm/mmu/mmu.c arch/x86/kvm/mmu/mmu.c
> > index 62f751952ad8..61f3e270ab61 100644
> > --- arch/x86/kvm/mmu/mmu.c
> > +++ arch/x86/kvm/mmu/mmu.c
> > @@ -1670,6 +1670,7 @@ static bool __kvm_rmap_zap_gfn_range(struct kvm *kvm,
> >
> > bool kvm_unmap_gfn_range(struct kvm *kvm, struct kvm_gfn_range *range)
> > {
> > + unsigned long shared_private = KVM_FILTER_SHARED |
> > KVM_FILTER_PRIVATE;
> > bool flush = false;
> >
> > /*
> > @@ -1683,6 +1684,10 @@ bool kvm_unmap_gfn_range(struct kvm *kvm, struct
> > kvm_gfn_range *range)
> > lockdep_assert_once(kvm->mmu_invalidate_in_progress ||
> > lockdep_is_held(&kvm->slots_lock));
> >
> > + if (gmem_in_place_conversion && !kvm_has_mirrored_tdp(kvm) &&
> > + ((range->attr_filter & shared_private) != shared_private))
> > + return false;
> > +
> > if (kvm_memslots_have_rmaps(kvm))
> > flush = __kvm_rmap_zap_gfn_range(kvm, range->slot,
> > range->start, range->end,
>
> Can this be a problem with gmem hugepages? Not sure how that is going to
> look like, so this may be covered in other ways.
>
> But, as it exists today, if the guest converts part of a 2M page to
> shared, then this skips zapping the SPTEs from the call in
> kvm_gmem_invalidate_start(). sev_gmem_make_shared() then issues PSMASH
> to convert RMP entry to 4k entries and we end up with 2M NPT+4K RMP. If
> the guest then writes to any private page in that range,
> page_fault_can_be_fast() returns true, fast_page_fault() only checks
> permissions with spte_permission_fault() and does not do anything.
> sev_handle_rmp_fault() also does not issue a zap since it finds that the
> RMP entry is already 4k, and we end up in a loop.
In Ackerley's hugetlb series, gmem will split the underlying folio to 4K in
that case. I assume, if we end up adopting Sean's proposed optimization, we'd
update the attr_filter to force a zap if there's a demotion, similarly to how
kvm_gmem_punch_hole() handles it in this series.
I think as a general rule of thumb though we'd decided that guessing
too hard at what hugepages will need is best left to hugepage series
itself (TDX prep work aside), because it always ends up different than
expected :)
Thanks,
Mike
>
>
> - Naveen
>