On Mon, Aug 17, 2026, Sean Christopherson wrote:
> On Sun, Aug 16, 2026, Ackerley Tng wrote:
> > Sean Christopherson <[email protected]> writes:
> > 
> > > On Thu, Aug 13, 2026, Ackerley Tng wrote:
> > >> Sean Christopherson <[email protected]> writes:
> > >> > That's why I think it's worth analyzing the cost: if it's in the
> > >> > noise, leave it alone.  If it's meaningful, figure out a not-too-gross 
> > >> > way to skip
> > >> > the entire thing if kvm_arch_gmem_make_shared() is a glorified nop in 
> > >> > the end.
> > >>
> > >> Is noise defined relative to the entire conversion process? Would this
> > >> benchmark look like
> > >>
> > >> 1. Convert 4G to shared on TDX with CONFIG_AMD_SEV defined
> > >> 2. Convert 4G to shared on TDX without CONFIG_AMD_SEV defined
> > >>
> > >> and then compare the difference in time taken?
> > >
> > > That'd work, though I was envisioning something even simpler: use rdtsc() 
> > > to
> > > count the cycles it takes to iterate over various ranges of memory.  Do 
> > > whatever
> > > is easiest for you though.
> > 
> > I made some changes to add rdtsc() for the conversion process as Sean
> > suggested [1], and exercised conversion like this [2]:
> > 
> > 1. Initialize some memory as private
> > 2. Get the guest to fault them into Secure EPTs
> > 3. Converts the memory to shared <<== this is being benchmarked
> > 4. Converts memory back to private
> > 
> > I made it build the VM once and convert 5 times:
> > 
> > ./gmem_benchmark_tdx_convert --iterations=5 --size=1g
> 
> ...
> 
> > And here's the above, tabulated:
> > 
> >   nr_pages      make_shared            total    percentage
> > ----------  ---------------  ---------------  ------------
> >          1              930            39278       2.3677%
> >          1              252            28060       0.8981%
> >          1              176            26952       0.6530%
> >          1              176            27038       0.6509%
> >          1              176            26980       0.6523%
> >          1             1072            37236       2.8789%
> >          1              316            28338       1.1151%
> >          1              176            27182       0.6475%
> >          1              176            26972       0.6525%
> >          1              176            26886       0.6546%
> >     262144         15041018       6616067680       0.2273%
> >     262144         14937462       6608542680       0.2260%
> >     262144         15138858       6599494898       0.2294%
> >     262144         15721972       6610219850       0.2378%
> >     262144         15000406       6615114540       0.2268%
> >    1048576         61902982      26400884028       0.2345%
> >    1048576         61746114      26401170984       0.2339%
> >    1048576         61096794      26404409058       0.2314%
> >    1048576         61446290      26447461896       0.2323%
> >    1048576         61774646      26444608360       0.2336%
> > 
> > Looks to me it is within noise.
> > 
> > I also actually tried measuring the conversion time from userspace with
> > CONFIG_AMD_SEV enabled and disabled. Converting a 1G-sized TD was faster
> > by 0.2%, which is in line with the above table. Interestingly, when
> > converting a 4G-sized TD, skipping kvm_gmem_make_shared() was _slower_
> > over 2 runs. I don't have an explanation for that.
> 
> Might be some cache/memory locality benefits?  Though with a conversion that 
> big,
> it could also be nothing more than bad luck.
> 
> > I think the code was correct. (If it makes a difference, I skipped
> > kvm_gmem_make_shared() using a custom guest_memfd creation time flag and
> > skipped make_shared if the flag was set on the inode.)
> > 
> > I thought adding a kvm_arch_has_gmem_make_shared(), defaulting it to
> 
> I would do kvm_arch_has_gmem_convert() for consistency with the Kconfigs, and
> because the cost of the reclaim invocation is a non-issue.

Forgot to say "why" I think it's worth adding kvm_arch_has_gmem_convert(): it's
a trivial change in the grand scheme, and I can see the ~1000 cycles of overhead
being problematic in use cases that do many one-off conversions, e.g. for guests
whose shared memory is scattered all over.

Reply via email to