>> Of course I have to bitch about the naming :)
>
> Yup :)
>
>>
>> Intuitively: kernel owned vs ... user owned?
>>
>> No, it's kernel owned vs core-mm owned.
>
> I would say somebody who does:
>
> ptr = malloc(4096);
>
> Would think of that memory as 'owned' by them in the sense that they
> control the lifetime, they established its attributes, etc.
>
> So indeed, vs. user-owned.
Right, I read "contents of the VMA are owned by the kernel rather than the core
mm?" and that confused me, because the opposite of the kernel is to me not
core-mm.
Maybe it would be clearer to focus on the opposite direction, then we wouldn't
have to find a word to describe "not core-mm".
>
>>
>> Which implies core-mm is not part of the kernel?
>>
>> Yes, this is confusing. ;)
>
> I think what you're missing here is what _creates_ or _establishes_ the
> mapping.
>
> Intuitively, if I do:
>
> ptr = kmalloc(GFP_KERNEL);
>
> I, whether I am in the core kernel, or a driver, or whatever own it in any
> meaningful sense of the word.
>
> I think the issue here is you're confusing this with other things like the
> rmap and refcounting, etc.
I think it all weirdly interacts.
In !vma_is_kernel_owned(), would we only expect ordinary folios
(anon/pagecache/hugetlb/dax, maybe shared zero folio)?
That's my best guess looking at
return vma_flags_test_any(flags, VMA_PFNMAP_BIT, VMA_MIXEDMAP_BIT,
VMA_IO_BIT);
So one option would be to focus instead on that aspect (folios that are managed
by core-mm vs. random other stuff not managed by core-mm). But thinking about
it, I'd prefer if we can leave the "folio" bits out, because COW mappings also
map (some) folios. See below.
>
>
>>
>> I assume you're coming from "map_kernel_pages*", but that's rather "kernel
>> memory" and not "kernel owned".
>>
>> Usually we say "driver owned" when not talking about pagecache/anon. Or user
>> vs.
>> kernel memory.
>
> I think it would only add confusion:
>
> VDSO/VVAR, perf ring buffers, shmem mapped via PFN map, uprobes, etc. are
> all in this category and I doubt people would consider those driver-owned.
Agreed. They are all special things and won't be folios in the future. (shmem
mapped through PFN tells us to ignore its folio background and treat it just as
some PFN range).
>
>>
>> So is it really all about "is this (excluding CoW) no ordinary user memory
>> that
>> we would track through the rmap" ?
>
> VMA_MIXEDMAP_BIT mappings can be refcounted and rmapped so that's not a
> correct description.
>
> The distinction is - who put them there and who's allowed to change them
> and who owns the lifecycle.
Lol, I asked AI for better names and it told me "vma_is_special_mapping()".
Thanks, I guess.
I assume for the reverse, we really just want to say "just an ordinary core-mm
vma that you would get from a simple mmap() as long as no weird non-mm drivers
or subsystems are involved. Core MM fully manages this thing.".
* vma_is_mm_managed()
* vma_is_mm_controlled()
>>> Signed-off-by: Lorenzo Stoakes (ARM) <[email protected]>
>>> ---
>>> include/linux/mm.h | 56
>>> ++++++++++++++++++++++++++++++++++++++++-
>>> tools/testing/vma/include/dup.h | 29 ++++++++++++++++++++-
>>> 2 files changed, 83 insertions(+), 2 deletions(-)
>>>
>>> diff --git a/include/linux/mm.h b/include/linux/mm.h
>>> index 2a92193ac6a5..cab29d6e15c1 100644
>>> --- a/include/linux/mm.h
>>> +++ b/include/linux/mm.h
>>> @@ -1612,6 +1612,44 @@ static inline bool vma_is_shared_maywrite(const
>>> struct vm_area_struct *vma)
>>> return is_shared_maywrite(&vma->flags);
>>> }
>>>
>>> +/**
>>> + * vma_flags_is_kernel_owned() - Do the specified VMA flags indicate that
>>> the
>>> + * contents of the VMA are owned by the kernel rather than the core mm?
>>> + * @flags: The VMA flags to test.
>>> + *
>>> + * A kernel-owned mapping is one whose contents are established and
>>> controlled
>>> + * by the kernel, typically a driver, rather than by the core mm's fault
>>> and
>>> + * rmap machinery.
>>> + *
>>> + * The mapping may be memory-mapped I/O, kernel-allocated pages or ordinary
>>> + * pages the owner has chosen to map itself (shmem via a PFN map, for
>>> instance).
>>> + *
>>> + * In all cases the core mm must not populate, reclaim, migrate,
>>> copy-on-write
>>> + * or merge it of its own accord.
>>> + *
>>> + * Pages mapped this way are not necessarily reference counted or map
>>> counted.
>>> + *
>>> + * Returns: true if the flags indicate a kernel-owned mapping.
>>> + */
>>> +static inline bool vma_flags_is_kernel_owned(const vma_flags_t *flags)
>>> +{
>>> + return vma_flags_test_any(flags, VMA_PFNMAP_BIT, VMA_MIXEDMAP_BIT,
>>> + VMA_IO_BIT);
>>
>> I thought we have cases where we drivers insert pages and neither set
>> VMA_PFNMAP_BIT nor VMA_MIXEDMAP_BIT.
>
> There were 4 - defio, cmt_speech, uprobes and the bpf arena, and I fixed
> all of them :)
Great, that helps to identify these things. I didn't look at all patches yet,
but we should definitely document that.
>
> Other than the DAX-only case below of course.
Right, as DAX uses real folios.
>>
>> I assume vmf_insert_page_mkwrite() is fine because it is DAX doing it
>> (should we
>> limit this interface to DAX?).
>
> That is a DAX-only thing and DAX is precisely a case that should not be
> kernel-owned (and isn't!)
>
> This series actually fixes the FUSE case too, restricting this interface to
> DAX only seems like a sensible follow up as well.
>
> I could also add a patch to this series to do that too if you wanted?
We can do a follow up. Not giving others the chance to abuse these interfaces
would be great.
--
Cheers,
David