On 8/26/26 18:30, Lorenzo Stoakes (ARM) wrote: > secretmem accounts folios by treating memory as if it were mlock()'d and > thus limited by the RLIMIT_MEMLOCK limit. > > However the folios are unevictable and remain so until the inode is > evicted, eliminating usual mlock() semantics - mapping folios then > unmapping them does not clear their unevictable state, since it depends on > AS_UNEVICTABLE, not PG_mlocked. > > A user can therefore easily work around the RLIMIT_MEMLOCK limit - simply > map then unmap and VmLck no longer counts the secretmem range. Worse, > folios are not accounted in the process's RSS, meaning the OOM killer won't > know to kill the process. > > Repeatedly mapping/unmapping (or forking) can then result in the > consumption of all available system memory with unevictable folios and > cause system instability. > > A secretmem fd can be passed between processes and over fork so a > per-process limit simply does not make sense, so follow the precedent set > by io_uring, perf, skbuff, iommufd and xdp by tracking the number of locked > pages in user_struct->locked_vm. > > Since the scope tracked is actually inode lifetime, the RLIMIT_MEMLOCK > applies per-user not per-process, so it doesn't make sense to bypass for > users with CAP_IPC_LOCK, therefore remove this bypass. > > There is simply no reason to carry on marking the mapping as mlock()'d > since it's misleading and the lifecycle is now correctly handled, so remove > this too. > > Note that secretmem does not support any form of truncation (including hole > punching) and the folios are unreclaimable, so the folios need only be > accounted on fault and unaccounted on inode destruction. > > __secretmem_account_pages() is more or less a duplicate of the code that > io_uring etc. use, but since this is a bug fix that needs backporting, > defer any de-duplication efforts to a follow-up. > > test_mlock_limit() asserts mlock_future_ok() on mmap(), however this has > been removed, so remove the test altogether for the fix. A new test will be > sent separately for upstream. > > Reported-by: Daehyeon Ko <[email protected]> > Closes: > https://lore.kernel.org/linux-mm/[email protected]/ > Fixes: 1507f51255c9 ("mm: introduce memfd_secret system call to create > "secret" memory areas") > Cc: [email protected] > Reviewed-by: Mike Rapoport (Microsoft) <[email protected]> > Acked-by: David Hildenbrand (Arm) <[email protected]> > Signed-off-by: Lorenzo Stoakes (ARM) <[email protected]> > ---
LGTM thanks -- Cheers, David

