Gregory Price <[email protected]> writes:

> On Wed, Sep 09, 2026 at 03:41:35PM -0700, Ackerley Tng wrote:
>> Gregory Price <[email protected]> writes:
>>
>> > guest_memfd allocates its folios through a per-inode shared mempolicy.
>> >
>> > Today that policy can only be set after the fact, with mbind() on a host
>> > mmap of the fd. That requires the fd to be mappable, and it cannot reach
>> > folios that are only ever guest-faulted. Neither holds for a
>> > non-mappable (confidential) guest_memfd.
>> >
>> > Add GUEST_MEMFD_FLAG_BIND_NODE and a node field to struct
>> > kvm_create_guest_memfd. When set, KVM builds an MPOL_BIND policy for the
>> > requested node and installs it over the whole inode, so every folio is
>> > allocated there with no userspace mbind().
>> >
>>
>> Instead of a custom API to ensure all guest_memfd allocations come from
>> a single node, how about these options?
>>
>> 1. Using cgroups/cpuset to constrain allocations (could be troublesome
>>    if the guest memory is not preallocated, unless the vCPU threads are
>>    running with the cpuset config)
>>
>> 2. Process-level NUMA policy
>>
>
> for 1 and 2:
>
> the intent is to put the guest memory on the target node, not all system
> memory for a given process.  so the scope here is not the same.
>
> in fact at that granularity, the desired node may not even have eligible
> memory to host the task's memory.
>

We have a similar problem where we wanted only guest memory to be
charged a certain memcg. Currently guest_memfd memory is charge to
mm->owner, which is the thread's parent, so doing the fallocate() in a
thread wasn't good enough. The workaround was to have the fallocate()
done in a separate process (fork) and have the child process's memcg be
set to the desired memcg.

This workaround is kind of awkward, but could it technically work for
NUMA allocations?

>> 3. Why not request the guest_memfd to be mmap-able just to be able to
>>    set a memory policy?
>>
>
> the eventual intent is to enable this for fully confidential,
> host-unmapped guest, isolated to a particular memory device.
>
> Requiring a mapping to get node-placement is quite defeating the point.
>
>> 4. How about something like fbind() that takes an fd and offset range
>>    instead of mbind(), which has a prerequisite on mmap()?
>>
>
> This was a consideration - although it has other limitations and larger
> complexities associated with it.
>
> Where does the policy live for random fd's? (inode? address_space?)
>

Off the top of my head the policy would be saved at gi->policy, where
it's stored for mbind() now. fd+offset is just another way of
referencing the offsets to set the policy.

> How is a reclaimed inode's policy handled? (lost forever?)
>
> figured i'd start by reducing the scope to the narrowest and clearest
> use-case, but I did expect to have the fbind() conversation.
>
> I'm open to it, but it seems like over-engineering.
>
> If you look at tmpfs / shmem, you'll see there is the option for a
> default filesystem-wide mempolicy that can be plumbed, but i'm not sure
> there's a real usecase for per-file mempolicies that isn't literally
> guest_memfd.
>

I see, can't think of other use-cases for per-file mempolicies either.

>> This doesn't exist yet, but I'm hoping to discuss this at LPC 2026:
>>
>> 5. What if you could pass an fd representing a mount to guest_memfd at
>>    creation time, so to make all the allocations come from a single node
>>
>>     Step 1: Create a tmpfs mount, specify mpol for mount to MPOL_BIND
>>     Step 2: Get some fd representing the tmpfs mount, hand that to
>>             guest_memfd at creation time
>>     Step 3: guest_memfd allocations will always come from that tmpfs
>>             mount, and abide by that tmpfs mount's memory policy.
>>
>
> This i think this is more feasible than something like fbind, but devil
> is in the details.
>
> I definitely think it's interesting and would love to chat at LPC!
>

Great! :) Cya!

Along the above lines, for HugeTLBfs we'd do something similar, and the
mount's fd would also indicate the HugeTLB page size.

Relating to memcg above, this doesn't really help though, since tmpfs
doesn't offer a way to configure memcgs.

>> May I know more about the use case behind this new feature?
>>
>
> see above - confidential vm whose memory is quarantined to a particular
> device, without placing that same burden on the host memory.
>

Is lazy allocation a requirement for you? (as opposed to fallocate()-ing
before the guest touches pages)

> ~Gregory

Reply via email to