MM VMA Operations
SLAB_TYPESAFE_BY_RCU and VMA Recycling
Dereferencing a parent/owner pointer from a SLAB_TYPESAFE_BY_RCU object
after dropping the object's refcount causes use-after-free when the object has
been recycled to a different owner. The owner can exit and free its backing
structure in the window between the refcount drop and the dereference.
The VMA cache is created with SLAB_TYPESAFE_BY_RCU (see vma_state_init()
in mm/vma_init.c), which means a VMA's slab memory remains valid through an
RCU read-side critical section even after vm_area_free(), but the VMA can be
reallocated to a completely different mm_struct during that window.
Per-VMA lock lookup protocol (see lock_vma_under_rcu() and
vma_start_read() in mm/mmap_lock.c):
mas_walk()underrcu_read_lock()finds a VMA in the maple treevma_start_read()incrementsvma->vm_refcnt- If
vma->vm_mm != mm(VMA was recycled), the refcount must be dropped -- butvma_refcount_put()dereferencesvma->vm_mmforrcuwait_wake_up() - The foreign
mmmust be stabilized withmmgrab()before callingvma_refcount_put(), then released withmmdrop()afterward
REPORT as bugs: Code paths in lock_vma_under_rcu(), lock_next_vma(),
or vma_start_read() that call vma_refcount_put() on a VMA whose vm_mm
does not match the caller's mm without first stabilizing the foreign mm
via mmgrab().
VMA Anonymous vs File-backed Classification
Using vma->vm_file to determine whether a VMA is file-backed causes
incorrect dispatch for VMAs that have a vm_file but are treated as
anonymous (e.g., private mappings of /dev/zero). This leads to BUG_ON
crashes, unaligned page offsets, or wrong code paths being taken.
How VMA classification works (see include/linux/mm.h):
vma_is_anonymous(vma)returns!vma->vm_ops-- this is the canonical test for anonymous VMAsvma_set_anonymous(vma)setsvma->vm_ops = NULLbut does NOT clearvma->vm_file- A VMA can have
vma->vm_file != NULLAND be anonymous (vm_ops == NULL)
VMAs where vm_file is set but the VMA is anonymous:
- Private mappings of
/dev/zero:mmap_zero_private_success()indrivers/char/mem.ccallsvma_set_anonymous(vma)for private mappings, leavingvm_filepointing to the/dev/zerofile. Shared mappings take a different path viashmem_zero_setup()which setsvm_ops = &shmem_anon_vm_ops - Any driver
mmaphandler that callsvma_set_anonymous()after the VMA is created with a file reference
Correct usage:
- To test "is this VMA file-backed?": use
!vma_is_anonymous(vma), NOTvma->vm_file != NULL - To test "is this VMA anonymous?": use
vma_is_anonymous(vma), NOTvma->vm_file == NULL - To access the backing file of a file-backed VMA: check
!vma_is_anonymous(vma)first, then usevma->vm_file
REPORT as bugs: Code that uses vma->vm_file (or !vma->vm_file) as
a proxy for file-backed (or anonymous) VMA classification in dispatch logic,
conditionals, or assertions. The correct test is vma_is_anonymous().
VMA Split/Merge Critical Section
Page table structural changes performed outside the vma_prepare()/
vma_complete() critical section race with concurrent page faults (via VMA
lock) and rmap walks (via file/anon rmap locks). The result is use-after-free,
page table corruption, or re-establishment of state that was just torn down.
The VMA-modifying paths -- __split_vma(), commit_merge(), and
vma_shrink() in mm/vma.c -- share a critical section:
vma_start_write()(acquire per-VMA lock, before or at entry)vma_prepare()(acquire file rmapi_mmap_lock_writeand anon_vma lock)- Page table structural changes:
vma_adjust_trans_huge(),hugetlb_split() - VMA range update (
vm_start/vm_end/vm_pgoff) vma_complete()(release locks acquired in step 2)
__split_vma() additionally calls vm_ops->may_split() before this sequence.
Rules:
vm_ops->may_split()must only validate whether the split is permitted (e.g., alignment checks). It must not modify page tables or other shared state, because it runs before the VMA and rmap locks are acquired- Any page table unsharing, splitting, or teardown required by a VMA split
must happen between
vma_prepare()andvma_complete(), where the VMA write lock and file/anon rmap write locks prevent concurrent page table walks (except hardware walks andgup_fast()) - When calling helpers that normally acquire their own locks (e.g.,
hugetlb_unshare_pmds()), use atake_locks=falsepath and assert that the needed locks are already held (seehugetlb_split()inmm/hugetlb.c)
Per-VMA Lock Exclusion via vma_start_write()
mmap_write_lock() alone does NOT exclude per-VMA lock holders — per-VMA
read locks acquired before mmap_write_lock() remain held, because the
seqcount in vma_start_read() only prevents new acquisitions, not
revocation of existing ones. Only vma_start_write(vma) drains existing
per-VMA read lock holders. VMA-locked operations (MADV_DONTNEED,
page faults) modify page tables at all levels including PMDs via
PT_RECLAIM (pmd_clear() in try_to_free_pte()), so any
mmap_write_lock holder accessing page tables before vma_start_write()
races with per-VMA locked paths.
check_pmd_still_valid() / find_pmd_or_thp_or_none(): These
functions walk page tables (mm_find_pmd → PGD→P4D→PUD→PMD) and then
read the PMD value via pmdp_get_lockless(pmd) in check_pmd_state().
A concurrent per-VMA locked MADV_DONTNEED can call try_to_free_pte()
→ pmd_clear() + free_pte() between the PMD read and subsequent use
of the result — the check succeeds, the caller proceeds assuming a valid
PMD, but the PMD has been cleared and the PTE page freed underneath it.
Code that calls these functions before vma_start_write() and then acts
on the result (e.g., proceeding to pmd_lock + pmdp_collapse_flush on
the assumption the PMD is still populated) is a bug — even though the
PMD pointer remains valid (it's in the PUD page which isn't freed),
the value and the PTE page it pointed to are gone.
REPORT as bugs: Functions holding mmap_write_lock that access page
tables or PTE pages before calling vma_start_write(vma). When a patch
adds new per-VMA lock users (e.g., converting a path from mmap_read_lock
to per-VMA lock), search with grep_functions/find_callers for
mmap_write_lock holders that access page tables for the same VMA and
verify each calls vma_start_write() before the access.
VMA Flags Modification API
Key distinction: vm_flags_set() ORs (adds bits, never clears),
vm_flags_reset() replaces (sets to exact value), vm_flags_init() replaces
without locking (VMA not yet in tree). vm_flags_clear() removes specific
bits. vm_flags_mod() adds and removes in one operation. See
include/linux/mm.h.
Common mistake: vm_flags_set(vma, new_flags) to replace flags -- because
it ORs, stale flags silently survive. Use vm_flags_reset() for exact
replacement. Stale VM_WRITE/VM_MAYWRITE creates security holes.
File Reference Ownership During mmap Callbacks
mmap uses split ownership: ksys_mmap_pgoff() holds one file reference
(fput at end), VMA gets its own via get_file() in __mmap_new_file_vma().
When a callback replaces the file (f_op->mmap_prepare() replacing
desc->vm_file, or legacy f_op->mmap() replacing vma->vm_file), the
replacement already carries its own reference.
REPORT as bugs: unconditional get_file() on a file that may have been
swapped by a callback -- the replacement gets a leaked extra reference. See
map->file_doesnt_need_get in call_mmap_prepare() in mm/vma.c and
shmem_zero_setup() in mm/shmem.c.
Quick Checks
- mmap_lock ordering: Taking the wrong lock type deadlocks or corrupts the
VMA tree. Write lock (
mmap_write_lock()) for VMA structural changes (insert/delete/split/merge, modifying vm_flags/vm_page_prot). Read lock (mmap_read_lock()) for VMA lookup, page fault handling, read-only traversal. See the "Lock ordering in mm" comment block at the top ofmm/rmap.c - Failable mmap lock reacquisition:
mmap_write_lock_killable()/mmap_read_lock_killable()return-EINTRon kill. Ignoring the return means continuing without the lock. Check in retry loops and lock upgrade sequences. See__get_user_pages_locked()inmm/gup.c - VMA merge anon_vma propagation: merging an unfaulted VMA with a
faulted one requires
dup_anon_vma()(seevma_expand()inmm/vma.c). Merge-timeanon_vmaproperty checks (e.g.,list_is_singular()inis_mergeable_anon_vma()) must apply to the VMA that has theanon_vma, not unconditionally to the destination -- the three cases (dst unfaulted/src faulted, dst faulted/src unfaulted, both faulted) are asymmetric. Seevma_is_fork_child()inmm/vma.c - VMA interval tree uses pgoff, not PFN:
mapping->i_mmapis keyed byvm_pgoff;vma_address()expectspgoff_t. Passing a raw PFN searches the wrong coordinate space. REPORT as bugs: raw PFN tovma_interval_tree_foreach()orvma_address() - VMA merge/modify error handling:
vma_modify()/vma_merge_new_range()may return error or a different VMA. Original VMA may be freed on success. On failure,vmg->start/end/pgoffmay be mutated and not restored — save originals or checkvmg_nomem(). Seemadvise_walk_vmas()inmm/madvise.c - VMA flag ordering vs merging: flags not in
VM_IGNORE_MERGEmust be set in proposedvm_flagsbeforevma_merge_new_range(). Setting flags post-merge viavm_flags_set()silently breaks future merges (is_mergeable_vma()XORs flags). Seeksm_vma_flags()inmm/ksm.c - VMA merge side effects vs page table operations:
vma_complete()triggersuprobe_mmap()which installs PTEs. Callers that subsequently move/overwrite page tables must setskip_vma_uprobeinstruct vma_merge_struct(seemm/vma.h), or orphaned PTEs leak memory - Fork-time VMA flag divergence:
dup_mmap()clears__VM_UFFD_FLAGSandVM_LOCKED_MASKon the child VMA. Fork-time flag checks (e.g.,vma_needs_copy()checkingVM_UFFD_WP) must use the destination VMA, not the source. Combined mask checks must verify all flags have the same source-vs-destination semantics - VM_ACCOUNT preservation during VMA manipulation: clearing
VM_ACCOUNTon a surviving VMA (e.g.,MREMAP_DONTUNMAP, partial unmap) leaks committed memory permanently —do_vmi_munmap()only uncharges VMAs withVM_ACCOUNT. Reviewvm_flags_clear()calls includingVM_ACCOUNT - VMA iteration on external mm_struct: call
check_stable_address_space(mm)after mmap lock, before traversal. Ondup_mmap()failure, maple tree slots containXA_ZERO_ENTRYmarkers and the mm is flaggedMMF_UNSTABLE. OOM reaper also setsMMF_UNSTABLE. Seeunuse_mm()inmm/swapfile.c - VMA operation results assigned to struct members:
vma_merge_extend(),vma_merge_new_range(),copy_vma()return NULL on failure. Assigning directly to a struct member (e.g.,vrm->vma = vma_merge_extend(...)) clobbers the original VMA pointer before the NULL check. Assign to a local first, NULL-check, then update the struct member on success - VMA merge functions invalidate input on success:
vma_merge_new_range(),vma_merge_existing_range(),vma_modify()may free the original VMA on success. Callers must use the returned VMA, not the original. Discarding the return value and using the original is use-after-free vma_modify*()error returns in VMA iteration loops:vma_modify_flags()etc. returnERR_PTR(-ENOMEM)on merge/split failure. Assigning back to a VMA loop variable withoutIS_ERR()check dereferences the error pointer. Even when the merge is best-effort (VMA unchanged on failure), the error return corrupts iteration. CheckIS_ERR()or usegive_up_on_oom- VMA lock refcount balance on error paths:
__vma_enter_locked()addsVMA_LOCK_OFFSETtovm_refcntthen waits for readers. When usingTASK_KILLABLE/TASK_INTERRUPTIBLE, the-EINTRpath must subtract the offset back. Leaked offset permanently blocks VMA detach/free - VMA addresses used as boolean flags:
vm_startcan legitimately be zero, soif (addr_var)to mean "was this set" silently fails for zero-address VMAs. Use an explicitboolflag or direct comparisons. Same for anyunsigned longaddress/offset that can be zero - Maple state RCU lifetime:
ma_statecaches RCU-protected node pointers. Afterrcu_read_unlock(), invalidate withmas_set()ormas_reset()before reuse. Easy to miss whenvma_start_read()drops RCU internally on failure. Seelock_vma_under_rcu()inmm/mmap_lock.c mm_structflexible array sizing: trailing flexible array packs cpumask and mm_cid regions. Static definitions (init_mm,efi_mm) must useMM_STRUCT_FLEXIBLE_ARRAY_INIT. Adding a new region requires updatingmm_cache_init()(dynamic),MM_STRUCT_FLEXIBLE_ARRAY_INIT(static), and all staticmm_structdefinitions- Memfd file creation API layering: calling
shmem_file_setup()orhugetlb_file_setup()directly for memfd produces files missingO_LARGEFILE, fmode flags, and security init. Usememfd_alloc_file(). REPORT as bugs: memfd creation via directshmem_file_setup()/hugetlb_file_setup()(non-memfd callers like DRM/SGX/SysV are fine) - VMA lock vs mmap_lock assertions:
mmap_assert_locked(mm)fires when only a VMA lock is held. Paths reachable under per-VMA locks must usevma_assert_locked(vma)(accepts either VMA lock or mmap_lock). Legacymmap_assert_locked()in page table walk/zap paths is likely incorrect - VM Committed Memory Accounting:
security_vm_enough_memory_mm()is not just a check -- on success it incrementsvm_committed_asviavm_acct_memory()inmm/util.c. Every error path after a successful call must invokevm_unacct_memory(). A leaked charge permanently inflatesvm_committed_as, causing-ENOMEMunder strict overcommit (vm.overcommit_memory=2)