uschindler commented on PR #16684: URL: https://github.com/apache/lucene/pull/16684#issuecomment-5831759083
**That is what Claude Max (Fable 5.1 High) says:** Short version: neither POSIX nor Linux has a notion of "advice on a file region". Advice is attached to an address range in the calling process, and on Linux that means to the VMA (`vm_area_struct`). What's shared between your two mappings is the page cache of the inode (and, if both come from the same `open()`, the `struct file` with its readahead state). So the question for each advice type is: does it change a per-VMA flag, act on page-table entries of that VMA, or act on the shared pages/inode? **POSIX** `posix_madvise()` is defined on a range of the process's address space, is purely advisory, and must not change program semantics. The spec says nothing about two mappings of the same file with conflicting advice; the only guarantee you get is MAP_SHARED coherence (both see the same data). Everything else is implementation-defined. Side note: glibc's `posix_madvise(POSIX_MADV_DONTNEED)` is a no-op on Linux, precisely because Linux's `MADV_DONTNEED` is destructive on anonymous memory and POSIX's version must not be. **Linux, by advice type** Per-VMA flag, independent for each mapping: - `MADV_NORMAL` / `MADV_RANDOM` / `MADV_SEQUENTIAL` set `VM_RAND_READ` / `VM_SEQ_READ` on the VMA. They're consulted in `filemap_fault` → `do_sync_mmap_readahead`: `VM_SEQ_READ` forces a full-window readahead, `VM_RAND_READ` suppresses it, no hint uses the heuristic (`mmap_miss` counting). Since readahead fills the shared page cache, a `MADV_RANDOM` mapping still profits from pages that a `MADV_SEQUENTIAL` mapping of the same range pulled in — the flag only governs what faults *through this VMA* trigger. - Subtlety: `file_ra_state` (window size and the `mmap_miss` counter that disables readahead after ~100 misses) lives in `struct file`, not the VMA. Two mappings created from the same fd share it; two `open()`s of the same file don't. A random-access mapping without a hint can thus throttle readahead for an un-hinted mapping of the same fd — while a `MADV_SEQUENTIAL` mapping bypasses the miss check entirely. - `MADV_HUGEPAGE` / `MADV_NOHUGEPAGE`, `MADV_DONTFORK`, `MADV_WIPEONFORK`, `MADV_DODUMP`/`DONTDUMP`, `MADV_MERGEABLE`: all VMA flags, no cross-mapping effect. Acts on this VMA's page tables only: - `MADV_DONTNEED` on a file-backed shared mapping just zaps the PTEs of that range (`zap_page_range_single`). The page cache is untouched; the other mapping keeps its PTEs and the next access through the first one simply refaults. It is *not* a "drop this from cache" operation — that's `posix_fadvise(POSIX_FADV_DONTNEED)`, which works on the inode. - `MADV_POPULATE_READ`/`WRITE` (5.14+) prefault this VMA; the page-cache side effect benefits everyone. - `MADV_FREE` returns `EINVAL` on file-backed mappings. Acts on shared state, visible through every mapping: - `MADV_WILLNEED` on file mappings is basically forced readahead into the page cache (no PTEs installed). All mappings benefit. - `MADV_COLD` / `MADV_PAGEOUT` (5.4+) walk the PTEs of this VMA and deactivate/reclaim the underlying folios — but they skip folios that are mapped elsewhere (the `mapcount != 1` / `folio_likely_mapped_shared` check), so a second mapping of the same range effectively pins pages against `MADV_PAGEOUT` from the first. Also, since ~5.11 `MADV_PAGEOUT` on file pages requires you to be the file's owner or have write permission/`CAP_SYS_ADMIN`. - `MADV_REMOVE` punches a hole in the file itself. `MADV_COLLAPSE` (6.1+) collapses the page-cache folios, so it affects all mappings. **Practical consequences** Conflicting advice on overlapping mappings doesn't "fight"; each fault path just looks at the VMA it came through. The one real cross-talk is the shared `file_ra_state`, and the one surprise is that a duplicate mapping defeats `MADV_PAGEOUT`. Two other things worth knowing when you create many mappings: adjacent mappings of the same file with identical flags can be merged into one VMA, and `madvise` on a sub-range then splits it again — which is what eats into `vm.max_map_count`. And the flag-changing advices (`NORMAL`/`RANDOM`/`SEQUENTIAL`/`HUGEPAGE`) take the mmap lock in write mode, so hammering them concurrently from many threads serializes against page faults in the whole process. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
