On 8/26/26 18:30, Lorenzo Stoakes (ARM) wrote:
> secretmem accounts folios by treating memory as if it were mlock()'d and
> thus limited by the RLIMIT_MEMLOCK limit.
> 
> However the folios are unevictable and remain so until the inode is
> evicted, eliminating usual mlock() semantics - mapping folios then
> unmapping them does not clear their unevictable state, since it depends on
> AS_UNEVICTABLE, not PG_mlocked.
> 
> A user can therefore easily work around the RLIMIT_MEMLOCK limit - simply
> map then unmap and VmLck no longer counts the secretmem range. Worse,
> folios are not accounted in the process's RSS, meaning the OOM killer won't
> know to kill the process.
> 
> Repeatedly mapping/unmapping (or forking) can then result in the
> consumption of all available system memory with unevictable folios and
> cause system instability.
> 
> A secretmem fd can be passed between processes and over fork so a
> per-process limit simply does not make sense, so follow the precedent set
> by io_uring, perf, skbuff, iommufd and xdp by tracking the number of locked
> pages in user_struct->locked_vm.
> 
> Since the scope tracked is actually inode lifetime, the RLIMIT_MEMLOCK
> applies per-user not per-process, so it doesn't make sense to bypass for
> users with CAP_IPC_LOCK, therefore remove this bypass.
> 
> There is simply no reason to carry on marking the mapping as mlock()'d
> since it's misleading and the lifecycle is now correctly handled, so remove
> this too.
> 
> Note that secretmem does not support any form of truncation (including hole
> punching) and the folios are unreclaimable, so the folios need only be
> accounted on fault and unaccounted on inode destruction.
> 
> __secretmem_account_pages() is more or less a duplicate of the code that
> io_uring etc. use, but since this is a bug fix that needs backporting,
> defer any de-duplication efforts to a follow-up.
> 
> test_mlock_limit() asserts mlock_future_ok() on mmap(), however this has
> been removed, so remove the test altogether for the fix. A new test will be
> sent separately for upstream.
> 
> Reported-by: Daehyeon Ko <[email protected]>
> Closes: 
> https://lore.kernel.org/linux-mm/[email protected]/
> Fixes: 1507f51255c9 ("mm: introduce memfd_secret system call to create 
> "secret" memory areas")
> Cc: [email protected]
> Reviewed-by: Mike Rapoport (Microsoft) <[email protected]>
> Acked-by: David Hildenbrand (Arm) <[email protected]>
> Signed-off-by: Lorenzo Stoakes (ARM) <[email protected]>
> ---

LGTM thanks

-- 
Cheers,

David

Reply via email to