On Tue, 2026-09-29 at 16:41 -0700, Ackerley Tng wrote:
> David Woodhouse <[email protected]> writes:
> 
> > On Mon, 2026-09-28 at 15:59 -0700, Ackerley Tng wrote:
> > > David Woodhouse <[email protected]> writes:
> > > 
> > > > On Fri, 2026-09-25 at 17:50 -0700, Ackerley Tng via B4 Relay wrote:
> > > > > 
> > > > > === What if my provider doesn't deal with pages?
> > > > > 
> > > > > Frank and David Woodhouse [2] have use cases for providers that use
> > > > > PFNs, and this is definitely something a generic interface must
> > > > > support.
> > > > > 
> > > > > Here are some options I can think of:
> > > > > 
> > > > > 1. Don't define .alloc_folio(), instead define .alloc_pfn()
> > > > > 2. Refactor .alloc_folio() to .alloc_pfn()
> > > > > 
> > > > > Either way, I think guest_memfd should still be the primary manager of
> > > > > the memory.
> > > > 
> > > > Thanks for working on this.
> > > > 
> > > 
> > > Thanks for the quick reply!
> > > 
> > > > I'm not sure what you mean by 'primary manager' here but my use case is
> > > 
> > > Naming is hard! See comment on revocation below.
> > > 
> > > > for the provider to be the ultimate arbiter of who owns which PFN, and
> > > > revocation of the same. Think of it like a filesystem backed by
> > > > external memory. Each guest's memory is a file, and a guest can
> > > > *donate* specific pages of its memory to other guests (which in the
> > > 
> > > Putting donation aside first - I don't have a good picture of how a
> > > guest would tell the host that it's willing to share a page - can't
> > > comment. Would love to find out more.
> > > 
> > > > context of the *interface* only means that we need revocation).
> > 
> > That's implementation-specific. Think of things like Nitro Enclaves
> > where a guest 'donates' its memory back to the hypervisor to be used by
> > another microvm. But that interface isn't the guest_memfd concern; as I
> > said, in the context of the guest_memfd interface, there is only
> > revocation: that page went away, whatever the reason.
> > 
> 
> Okay we're on the same page then. :) guest_memfd only needs to know that
> the page went away, guest_memfd cannot say no.
> 
> > > Going with this revocation part, beginning with a more basic use case:
> > > I'm thinking that the provider can notify guest_memfd that the page or
> > > PFN is going away.
> > 
> > Right. That part I have working in what I was posting.
> > [5] 
> > https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=ee95ed5aaf06
> > 
> 
> That is similar to what I had in mind except the parameters, I was
> thinking inode, offset, size.
> 
> kvm_gmem_invalidate_range() in your patch [5] does have a kvm_gmem
> prefix, but it's really a direct call to tell KVM MMU to zap the gfn
> range, guest_memfd doesn't get to update its own state.

Late here now, but it occurs to me that I need to check if that makes
sense. Does the provider *know* the GFN, or only the offset within its
own object — which might *not* be in a memslot starting at GFN0?

Tangentially: I'd actually quite like the provider to *know* about any
such offset, because if possible I'd like it to be able to be
opinionated about where in the guest pages are mapped.

> > > 
> > > Yup, guest_memfd needs some way to track PFNs in addition to folios. I
> > > think for tracking we can reference DAX, the part I haven't looked at is
> > > how to let guest_memfd handle events. Like if there's memory failure on
> > > a PFN, how do we let guest_memfd handle it?
> > 
> > Telling the guest about correctable and uncorrectable errors, you mean?
> > I would have thought that's up to the VMM, until the point where (see
> > revocation)?
> > 
> 
> Informing the guest is definitely the userspace VMM's action to
> take.
> 
> guest_memfd's role here is updating it's own tracking, and perhaps
> splitting folios. When guest_memfd supports huge pages, one thing high
> on the wishlist is to split the huge folio and zap only the page where
> there was an error and not the huge page.
> 
> Advantage of splitting: If the failure happened in the last 4K, the
> first 1G - 4K can continue to be used by the guest. If the guest never
> touches the last 4K, it doesn't have to know about the failure at all.

Right, but this is still just revocation of a single 4KiB page and
expecting both KVM and IOMMU to shatter the large page accordingly,
isn't it? That much I already had working.

(The IOMMU shatter does need to be atomic and miss-less, which I don't
think is the case on Arm but could be.)

> > 
> > I think where the memory comes *from* is an implementation detail for
> > the provider. It provides a PFN (or folio, if you must). All else is
> > not the business of KVM or IOMMUFD.
> > 
> 
> So we already have
> 
>   dmabuf --wraps--> CMA (and others)
> 
> We could have
> 
>   [a] guest_memfd --wraps--> dmabuf --wraps--> CMA
> 
> or
> 
>   [b] guest_memfd --wraps--> CMA
> 
> I agree guest_memfd shouldn't care where the PFN comes from, but this
> does impact where the provider_ops code is implemented though.
>
> 
> If we go with [b], CMA needs to find a non-dmabuf way for userspace to
> get some resource fd. If we go with [a], dmabuf folks need to support
> this :)

I'm not sure I see how CMA is relevant to the *interface*.

A given *implementation* (provider) of a guest_memfd resource might use
CMA for its backing store. Or might use hugetlbfs. Or something DAX-
like. Once the implementation has provided its guest_memfd operations
with its 'get_pfn' method, nobody else cares *where* it finds the PFNs
that it returns when we invoke that method.

> I missed considering breaking down huge pages (the folios). guest_memfd
> will definitely need to cooperate with the provider to break down the
> pages.

KVM manages this part on its own. The guest_memfd implementation only
has to tell it to revoke. It's been a while, but I don't even remember
having to do anything special.
https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=ee95ed5aaf0


> I'll need to prototype support with HugeTLBfs, probably after LPC. There
> I'll need to split pages on conversion to shared and merge on conversion
> to private.
> 
> I think you meant breaking down the mappings for the large pages in
> stage 2? guest_memfd tells KVM to do that in one of the earlier
> prototypes for guest_memfd HugeTLB pages.
> 
> > > >  • AsyncPF support for pages requested by KVM.
> > > 
> > > Is this kind of orthogonal? IIUC today KVM MMU handles the async-ness.
> > > KVM MMU checks that the page is not there, then since async was
> > > configured, KVM tells the guest to schedule in something else.
> > > 
> > > guest_memfd doesn't have support for async #PFs now. I imagine it would
> > > be something like KVM MMU firing off a kvm_gmem_get_pfn() request
> > > asynchronously, and then guest_memfd gets the page or PFN and inserts it
> > > in guest_memfd filemap, and then notifies KVM MMU?
> > > 
> > > Within kvm_gmem_get_pfn(), guest_memfd should get the page or PFN from
> > > the provider.
> > > 
> > > I think it's orthogonal because guest_memfd could
> > > have gotten a PAGE_SIZE page (existing functionality) or gotten a page
> > > from the provider.
> > 
> > Kind of, but I'm focusing on the guest_memfd provider interface, and
> > that needs to support the asychronous mode: asked for a PFN for a given
> > guest address, it returns -EAGAIN and then provides it later, and the
> > guest gets the right asyncpf behaviour:
> > https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=894dd0fcb34f
> 
> I thought the same .alloc_folio would be used, except for an async #PF
> the entire kvm_gmem_get_folio(), which calls .alloc_folio() would be
> done in some background thread, so the provider doesn't have to have a
> different .alloc_folio_async().
> 

> If the provider does need to know, then the corresponding thing in this
> series could be:
> 
>   struct guest_memfd_provider_operations {
>         void *(*attach)(struct file *resource_file);
>         void (*release)(void *provider);
>         struct folio *(*alloc_folio)(void *provider, pgoff_t index,
> -                                    struct mempolicy *mpol);
> +                                    struct mempolicy *mpol, bool nowait);
>         void (*invalidate_folio)(void *provider, struct folio *folio);
>   };
> 
> something like that?

Yeah, the 'nowait' argument on get_pfn() is how I have it working at
the moment. And then, as you say, KVM calls it again *without* nowait,
in a context which can sleep.

> Either way I think it would have to be a later series after the first
> one (for upstreaming).

I don't have a particular use case for asyncpf right now personally,
but I think we *should* get the design for the interface right. We
should do it on the IOMMU side too, because ATS+PRI exists.

Attachment: smime.p7s
Description: S/MIME cryptographic signature

Reply via email to