MM Page Table Operations
PTE State Consistency
Incorrect PTE flag combinations cause data corruption (dirty data silently dropped), security holes (writable pages that should be read-only), and kernel crashes on architectures that trap invalid combinations. Review any code that constructs or modifies PTEs for these invariants.
Invariants (software-enforced, not hardware):
- Writable PTEs must be dirty: a clean+writable PTE is invalid
- For shared mappings,
can_change_shared_pte_writable()inmm/mprotect.cenforces this by only returning true whenpte_dirty(pte)(clean shared PTEs need a write-fault for filesystem writenotify) - For private/anonymous mappings, code paths use
pte_mkwrite(pte_mkdirty(entry))to set both together (seedo_anonymous_page()inmm/memory.c,migrate_vma_insert_page()inmm/migrate_device.c) - Exception -- MADV_FREE:
madvise_free_pte_range()inmm/madvise.cclears the dirty bit viaclear_young_dirty_ptes()but preserves write permission, intentionally creating a clean+writable PTE. This allows the page to be reclaimed without writeback (it's clean and lazyfree), but if the process writes new data before reclaim, the page becomes dirty again without a full write-protect fault. On x86,pte_mkclean()only clears_PAGE_DIRTY_BITSand does not touch_PAGE_RW, so hardware sets dirty directly with no fault at all. On arm64,pte_mkclean()setsPTE_RDONLYbut preservesPTE_WRITE; with FEAT_HAFDBS hardware clearsPTE_RDONLYon write (no fault), without it a minor fault resolves quickly sincepte_write()is still true
- For shared mappings,
- Dirty implies accessed as a software convention:
pte_mkdirty()does NOT set the accessed bit (x86, arm64), so code paths must set both explicitly - Non-accessed+writable is invalid on architectures without hardware A/D bit management (on x86, hardware sets accessed automatically on first access)
Migration entries (include/linux/swapops.h):
- Encode A/D bits via
SWP_MIG_YOUNG_BITandSWP_MIG_DIRTY_BIT - Only available when
migration_entry_supports_ad()returns true (depends on whether the architecture's swap offset has enough free bits; controlled byswap_migration_ad_supportedinmm/swapfile.c) make_migration_entry_young()/make_migration_entry_dirty()preserve original PTE state into the migration entryremove_migration_pte()inmm/migrate.crestores A/D bits: dirty is set only if both the migration entry AND the folio are dirty (avoids re-dirtying a folio that was cleaned during migration)
NUMA balancing (see change_pte_range() in mm/mprotect.c):
- Skips PTEs already
pte_protnone()to avoid double-faulting - Checks
folio_can_map_prot_numa()before applying NUMA hint faults
Swap entries (see try_to_unmap_one() in mm/rmap.c):
- Only exclusive, soft-dirty, and uffd-wp flags survive in swap PTEs; all other PTE state is lost on swap-out
pte_swp_clear_flags()ininclude/linux/swapops.hstrips these flags to extract the bare swap entry for comparison (seepte_same_as_swp()inmm/swapfile.c)
Non-present PTE type dispatch (see check_pte() in
mm/page_vma_mapped.c, softleaf_type() in include/linux/leafops.h):
Non-present PTEs encode several distinct swap entry types via the
softleaf_type / swp_type field. Each type has different semantics and must
be handled in the correct branch of any dispatch logic. Accepting an entry type
in the wrong branch causes semantic confusion (e.g., treating a
device-exclusive entry as a migration entry), which may silently produce wrong
behavior even if the types share the PFN-encoding property.
The distinct non-present PTE categories are:
- Migration (
SOFTLEAF_MIGRATION_READ,_READ_EXCLUSIVE,_WRITE): page temporarily unmapped during folio migration; checked viasoftleaf_is_migration() - Device-private (
SOFTLEAF_DEVICE_PRIVATE_READ,_WRITE): page migrated to un-addressable device memory (HMM); checked viasoftleaf_is_device_private() - Device-exclusive (
SOFTLEAF_DEVICE_EXCLUSIVE): CPU access temporarily revoked for device atomic operations, page remains in host memory; checked viasoftleaf_is_device_exclusive() - HW poison (
SOFTLEAF_HWPOISON): page has uncorrectable memory error; checked viasoftleaf_is_hwpoison()/is_hwpoison_entry() - Marker (
SOFTLEAF_MARKER): metadata-only entry (e.g., uffd-wp marker, poison marker); checked viasoftleaf_is_marker()
When reviewing code that adds a new swap entry type or modifies dispatch logic
over non-present PTEs, verify that each branch accepts only the entry types
whose semantics match that branch's purpose. A common mistake is grouping
device-exclusive with migration (both involve temporarily unmapped pages with
PFNs) even though their refcount behavior, resolution paths, and semantics
are entirely different. softleaf_has_pfn() in include/linux/leafops.h
shows which types encode a PFN -- sharing this property does not make types
interchangeable in dispatch logic.
Flag transfer on non-present-to-present PTE reconstruction:
Every code path that converts a non-present PTE (swap, migration, or device-exclusive entry) to a present PTE must carry over the soft-dirty and uffd-wp bits. These bits have different encodings in swap PTEs vs present PTEs, so they require explicit read-then-write transfer:
if (pte_swp_soft_dirty(old_pte))
newpte = pte_mksoft_dirty(newpte);
if (pte_swp_uffd_wp(old_pte))
newpte = pte_mkuffd_wp(newpte);
This pattern is required in do_swap_page(), restore_exclusive_pte() in
mm/memory.c, remove_migration_pte(), try_to_map_unused_to_zeropage()
in mm/migrate.c, and unuse_pte() in mm/swapfile.c. When converting
between non-present entries (swap-to-swap), use the swap-side writers
instead: pte_swp_mksoft_dirty() and pte_swp_mkuffd_wp() (see
copy_nonpresent_pte() in mm/memory.c)
Soft dirty vs hardware dirty in PTE move/remap:
Soft dirty (pte_mksoft_dirty() / pte_swp_mksoft_dirty()) is a
userspace-visible tracking bit for /proc/pid/pagemap and CRIU, distinct
from hardware dirty (pte_mkdirty()). PTE move operations (mremap,
userfaultfd UFFDIO_MOVE) must set soft dirty on the destination to signal
the mapping changed, while preserving the source PTE's hardware dirty state.
Common mistakes:
- Using
pte_mkdirty()when the intent is to mark the PTE as "touched" for userspace tracking -- this should bepte_mksoft_dirty() - Handling present PTEs but forgetting
pte_swp_mksoft_dirty()for swap PTEs - Using
#ifdef CONFIG_MEM_SOFT_DIRTYinstead ofpgtable_supports_soft_dirty(), which also handles runtime detection (e.g., RISC-V ISA extensions)
See move_soft_dirty_pte() in mm/mremap.c for the reference implementation
handling both present and swap cases.
Special vs Normal Page Table Mappings
Marking a normal refcounted folio's page table entry as "special" causes
vm_normal_page() (and vm_normal_page_pmd() / vm_normal_page_pud())
to return NULL, hiding the folio from page table walkers, GUP, and refcount
management. GUP-fast checks pte_special() / pmd_special() /
pud_special() early and bails out, falling back to slow GUP.
Invariant (see __vm_normal_page() in mm/memory.c):
- Normal refcounted folios must NOT have their page table entry marked
special (
pte_mkspecial()/pmd_mkspecial()/pud_mkspecial()) - Only raw PFN mappings (VM_PFNMAP, VM_MIXEDMAP without struct page),
devmap entries (
pte_mkdevmap()/pmd_mkdevmap()/pud_mkdevmap()), and the shared zero folios may be marked special - Use
folio_mk_pmd()/folio_mk_pud()when constructing entries for normal refcounted folios; these helpers produce a plain huge entry without setting the special bit (seeinclude/linux/mm.h) - Use
pfn_pmd()+pmd_mkspecial()orpfn_pud()+pud_mkspecial()only for raw PFN mappings
Common mistake: When a vmf_insert_folio_*() function reuses a
PFN-oriented helper (e.g., vmf_insert_pfn_pud()), the helper's
unconditional pXd_mkspecial() call applies to the folio mapping too.
The fix is to split the entry-construction logic so the folio path uses
folio_mk_pXd() without the special bit, while the PFN path retains
pXd_mkspecial() (see insert_pmd() and insert_pud() in
mm/huge_memory.c for the correct pattern using struct folio_or_pfn).
Page Table Entry to Folio/Page Conversion Preconditions
Applying a present-entry conversion function to a non-present page table
entry (migration entry, swap entry, or poisoned entry) interprets swap
metadata bits as a physical page frame number, producing a bogus
struct page * that causes an invalid address dereference.
Functions that require a present entry:
pmd_folio(pmd)/pmd_page(pmd)-- expand topfn_to_page(pmd_pfn(pmd)), which is only valid whenpmd_present(pmd)orpmd_trans_huge(pmd)orpmd_devmap(pmd)(seeinclude/linux/pgtable.h)pte_page(pte)/vm_normal_page()-- only valid whenpte_present(pte)
Correct conversion for non-present entries:
- Migration entries:
softleaf_from_pmd()orsoftleaf_from_pte()to extract thesoftleaf_t, thensoftleaf_to_folio()to get the folio (seeinclude/linux/leafops.h) - Alternatively, if the folio comparison is not needed for a non-present entry (e.g., because a migration entry is locked and cannot refer to the target folio), skip the conversion entirely
Review pattern: When code handles multiple PMD/PTE states in a combined
conditional (e.g., if (pmd_trans_huge(*pmd) || pmd_is_migration_entry(*pmd))),
verify that subsequent operations like pmd_folio() or pmd_page() are
guarded to execute only on the present-entry cases. A common mistake is
adding a non-present entry type to an existing condition without adjusting
the folio extraction that follows.
PTE Batching
Batching consecutive PTEs that map the same large folio into a single
set_ptes() call propagates the first PTE's permission bits to all entries
in the batch, because set_ptes() only advances the PFN and preserves all
other bits. If the batch includes PTEs with different permissions (e.g.,
writable vs read-only), the result silently overwrites the intended
permissions, causing security bypasses.
folio_pte_batch() in mm/util.c is a simplified wrapper that calls
folio_pte_batch_flags() in mm/internal.h with flags=0. With no flags,
differences in writable, dirty, and soft-dirty bits are ignored and PTEs
with different permissions are batched together.
FPB flags (defined as fpb_t in mm/internal.h) control which PTE bits
are compared vs ignored during batching:
| Flag | Effect |
|---|---|
FPB_RESPECT_WRITE |
Include the writable bit in comparison; PTEs with different write permissions will not batch |
FPB_RESPECT_DIRTY |
Include the dirty bit in comparison |
FPB_RESPECT_SOFT_DIRTY |
Include the soft-dirty bit in comparison |
FPB_MERGE_WRITE |
After batching, if any PTE was writable, set the writable bit on the output PTE |
FPB_MERGE_YOUNG_DIRTY |
After batching, merge young and dirty bits from all PTEs into the output |
folio_pte_batch()(no flags): safe only when the caller does not stamp the first PTE's permission bits onto other entries (e.g.,zap_present_ptes()which clears all PTEs, orfolio_unmap_pte_batch()which unmaps)folio_pte_batch_flags()withFPB_RESPECT_WRITE: required when the caller usesset_ptes()to write the batched PTE value back (seemove_ptes()inmm/mremap.c,change_pte_range()inmm/mprotect.c)
REPORT as bugs: Code that uses folio_pte_batch() (without flags) to
determine a batch count and then passes that count to set_ptes(), because
the first PTE's writable/dirty/soft-dirty bits will be stamped onto all
entries in the batch.
Batched PTE operation boundaries:
Passing an uncapped max_nr to folio_pte_batch() causes out-of-bounds reads
past the end of a page table. The max_nr parameter must be capped so that
scanning stays within a single page table and a single VMA. The standard
expression is (pmd_addr_end(addr, vma->vm_end) - addr) >> PAGE_SHIFT. Code
that reaches PTE-level iteration through the standard walker hierarchy
(zap_pmd_range() -> zap_pte_range()) receives a pre-capped end. Code
that operates directly at PTE level via page_vma_mapped_walk() must perform
its own PMD boundary capping.
REPORT as bugs: Any caller of folio_pte_batch() that derives max_nr
from folio_nr_pages() without capping at the PMD boundary.
page_vma_mapped_walk() Non-Present Entries
Calling PTE accessor functions (pte_young(), pte_dirty(), pte_write(),
ptep_clear_flush_young(), etc.) on a non-present entry returned by
page_vma_mapped_walk() produces undefined results because swap entries
encode bits differently than present PTEs. This class of bug went undetected
because the non-present entries only appear with device-exclusive or
device-private ZONE_DEVICE pages.
page_vma_mapped_walk() in mm/page_vma_mapped.c can return true with
pvmw.pte pointing to non-present entries. The check_pte() helper accepts
three PTE types when PVMW_MIGRATION is not set:
| Entry type | pte_present() |
How to identify |
|---|---|---|
| Normal present PTE | true | pte_present(ptep_get(pvmw.pte)) |
| Device-exclusive swap entry | false | softleaf_is_device_exclusive(...) |
| Device-private swap entry | false | softleaf_is_device_private(...) |
Rules for rmap walk callbacks (functions passed to rmap_walk() or using
page_vma_mapped_walk()):
- When
pvmw.pteis set, always checkpte_present(ptep_get(pvmw.pte))before calling present-PTE accessors (pte_pfn(),pte_dirty(),pte_write(),pte_young(),pte_soft_dirty(),pte_uffd_wp()) - Non-present PFN swap PTEs (device-exclusive and device-private entries)
require converting to
softleaf_tfirst viasoftleaf_from_pte(), then usingsoftleaf_to_pfn()for PFN andsoftleaf_is_device_private_write()for writability. Swap-PTE flag accessors (pte_swp_soft_dirty(),pte_swp_uffd_wp()) still acceptpte_tbut read different bit positions than their present-PTE counterparts (pte_soft_dirty(),pte_uffd_wp()), so using the wrong family silently reads wrong bits. Seetry_to_migrate_one()andtry_to_unmap_one()inmm/rmap.cfor the correct dispatching pattern - Non-present PFN swap PTEs represent pages that are "old" and "clean" from the CPU's perspective; MMU notifiers handle device-side access tracking
Large Folio PTE Installation
When pte_range_none() returns false during large folio installation (some
PTEs already populated), the handler must ensure forward progress:
- Page cache folios (
finish_fault()): fall back to single-PTE install - Freshly allocated anon folios (
do_anonymous_page()): release folio, retry at smaller size viaalloc_anon_folio() - PMD-level (
do_set_pmd()): returnVM_FAULT_FALLBACK
REPORT as bugs: returning VM_FAULT_NOPAGE (retry) when
pte_range_none() fails for a page cache folio without falling back to
single-PTE -- this creates a livelock (hung process, no warning).
Page Table Walker Callbacks
pmd_entry in struct mm_walk_ops (see include/linux/pagewalk.h) receives
every non-empty PMD including pmd_trans_huge(). Failing to handle THP PMDs
causes silent data skipping or crashes from treating a huge-page PFN as a
page table pointer. When pmd_entry is defined without pte_entry, the
walker does NOT descend to PTEs -- the callback must walk PTEs internally.
Return values: 0 = continue, > 0 = stop (returned to caller), < 0 =
error. walk_lock specifies locking: PGWALK_RDLOCK (mmap_lock read),
PGWALK_WRLOCK (walker write-locks VMAs), PGWALK_WRLOCK_VERIFY /
PGWALK_VMA_RDLOCK_VERIFY (assert already locked).
In pmd_entry callbacks: read PMD locklessly with pmdp_get_lockless(),
reread under pmd_lock() for THP; check pte_offset_map_lock() return for
NULL; call folio_get() before releasing PTL if returning a folio reference.
Quick Checks
- TLB flushes after PTE modifications: Missing a TLB flush after making a
PTE less permissive lets userspace keep stale write access, causing data
corruption or security bypass. Required for writable-to-readonly and
present-to-not-present transitions. Not needed for not-present-to-present or
more-permissive transitions (callers pair
ptep_set_access_flags()withupdate_mmu_cache()). Seechange_pte_range()inmm/mprotect.candzap_pte_range()inmm/memory.c - VM_WRITE gate for writable PTEs: writable PTEs require
VM_WRITEinvma->vm_flags. Usemaybe_mkwrite()(include/linux/mm.h). Verify in fork/COW, userfaultfd install, and any PTE construction path — VMA permissions can change viamprotect()between mapping and installation - VMA flag and PTE/PMD flag consistency: clearing a
vm_flagsbit (e.g.,VM_UFFD_WP,VM_SOFT_DIRTY) requires clearing the corresponding PTE/PMD bits across all forms (present, swap, PTE markers). Error-prone when VMA flag clearing and page table walk are in different code paths. Seeclear_uffd_wp_pmd()inmm/huge_memory.c flush_tlb_batched_pending()after PTL re-acquisition: after dropping and re-acquiring PTL, callflush_tlb_batched_pending(mm)— reclaim on another CPU may have batched TLB flushes while the lock was released. Seeflush_tlb_batched_pending()inmm/rmap.c- Page table removal vs GUP-fast: clearing a PUD/PMD to free a page
table page requires
tlb_remove_table_sync_one()ortlb_remove_table()before reuse. GUP-fast walks locklessly underlocal_irq_save()and can follow stale entries into freed page tables without synchronization. Seemm/mmu_gather.c vma_start_write()before page table access undermmap_write_lock:mmap_write_lockdoes NOT exclude per-VMA lock readers (e.g., madvise underlock_vma_under_rcu()). Callvma_start_write(vma)before checking/modifying page tables to drain per-VMA lock holders. Seecollapse_huge_page()inmm/khugepaged.c. Critical: PTE-level zap operations cross granularity boundaries. VMA-lockedMADV_DONTNEEDcallszap_page_range_single_batched()→unmap_page_range()→zap_pmd_range()→zap_pte_range()→try_to_free_pte()(inmm/pt_reclaim.c) →pmd_clear(). When all PTEs in a page table are zapped, PT_RECLAIM frees the PTE page and clears the PMD entry. Code that read the PMD beforevma_start_write()now holds a stale pointer to freed memory — this is use-after-free (kernel panic), not just stale data. Do NOT dismiss PMD-level accesses beforevma_start_write()as "different granularity" from PTE-level zap operations — the zap path modifies PMDs too- Page fault path lock constraints:
->fault/->page_mkwriterun undermmap_lock, nested belowi_rwsemandsb_start_write. Fault handlers must not wait on freeze protection (ABBA deadlock). Copy user data withcopy_folio_from_iter_atomic()and retry outside the lock. Seegeneric_perform_write()inmm/filemap.c pte_unmap_unlockpointer must be within the kmap'd PTE page: After a PTE iteration loop, the iterated pointer may point one-past-the-end of the PTE page. OnCONFIG_HIGHPTEsystems,pte_unmap()callskunmap_local(), which derives the page address viaPAGE_MASK. If the pointer crosses a page boundary, it unmaps the wrong page. Save the start pointer frompte_offset_map_lock()or passptep - 1after the loop. Only triggers on 32-bit HIGHMEM architecturespte_unmap()LIFO ordering: multiple PTE mappings must be unmapped in reverse order. Invisible on 64-bit; triggers WARNING on 32-bit HIGHPTE wherepte_unmap()callskunmap_local()pmd_present()afterpmd_trans_huge_lock(): succeeds for both present THP PMDs and non-present PMD leaf entries (migration, device-private). Must checkpmd_present()beforepmd_folio()/pmd_page()or any function assuming a present PMD- Page table state after lock drop and retry: after dropping and
reacquiring PTL, concurrent threads may have repopulated empty entries.
Decisions to free page table structures must be re-validated. See
zap_pte_range()direct_reclaimflag inmm/memory.candtry_to_free_pte()inmm/pt_reclaim.c - Kernel page table population synchronization:
pgd_populate()/p4d_populate()do NOT sync to other processes' kernel page tables. Usepgd_populate_kernel()/p4d_populate_kernel()which callarch_sync_kernel_mappings(). Affects vmemmap, percpu, KASAN shadow - Lazy MMU mode pairing and hazards: (1) PTE reads after writes inside
lazy mode may return stale data — bracket with leave/enter. (2) Error
paths must not skip
arch_leave_lazy_mmu_mode()— usebreaknotreturn. No-op on most configs; bugs only manifest on Xen PV, sparc, powerpc book3s64, arm64 - Lazy MMU mode implies possible atomic context: disables preemption
on some architectures (sparc, powerpc).
pte_fn_tcallbacks and PTE loops inside lazy MMU mode must not sleep. Allocations needGFP_ATOMIC/GFP_NOWAITor pre-allocation. Invisible on x86/arm64 - Non-present PTE swap entry type dispatch: see the full section in PTE State Consistency above. Verify each dispatch branch accepts only semantically matching entry types — do not group device-exclusive with migration despite both having PFNs
arch_sync_kernel_mappings()on error paths: loops that accumulatepgtbl_mod_maskand callarch_sync_kernel_mappings()after must usebreak(notreturn) on errors. Earlyreturnskips the sync, leaving other processes' kernel page tables stale. See__apply_to_page_range()andvmap_range_noflush()- Page table walker iterator advancement: in
do { } whilepage table loops, advance the pointer unconditionally in thewhileclause (e.g.,} while (pte++, addr += PAGE_SIZE, addr != end)), not inside a conditional body. Usecontinueto skip entries sowhilestill advances. Placingptr++inside anifstalls the walker when false - Mapcount symmetry for non-present PTE entries: non-present swap PTEs
holding a folio reference (device-private, device-exclusive, migration) must
keep mapcount symmetric: if mapcount is maintained during creation, teardown
(
zap_nonpresent_ptes()) must remove it. Device-private/exclusive maintain mapcount; migration entries are managed bytry_to_migrate_one()itself - PTE batch loop bounds: loops batching consecutive PTEs must not rely
solely on
pte_same()against a synthetic expected PTE. On XEN PV,pte_advance_pfn()can producepte_none()for PFNs without valid machine frames, causing false matches and overrunning the folio. Bound iteration independently using folio metadata (folio_nr_pages()etc.) pte_offset_map/pte_unmappairing:pte_unmap()must receive the exact pointer frompte_offset_map(), never a pointer to a localpte_tcopy. Both arepte_t *so the compiler won't warn. OnCONFIG_HIGHPTE, passing a stack address topte_unmap()unmaps the wrong mapping. Common mistake:pte_t orig = ptep_get(pte)thenpte_unmap(&orig)- Memory hotplug lock for kernel page table walks: walking
init_mmpage tables needsget_online_mems()/put_online_mems(), not justmmap_lock. Hot-remove frees intermediate PUDs/PMDs for direct-map and vmemmap ranges, causing use-after-free in concurrent walkers. Acquire hotplug lock beforemmap_lockfor ordering - Bit-based locking barrier pairing: when a bit flag is used for mutual
exclusion (trylock pattern), the unlock must use
clear_bit_unlock()(release semantics), notclear_bit()(relaxed, no barrier). The lock side must usetest_and_set_bit_lock()(acquire semantics). Plainclear_bit()allows stores to be reordered past the unlock on weakly-ordered architectures (arm64). MM uses this forPGDAT_RECLAIM_LOCKEDinmm/vmscan.c walk_page_range()default skipsVM_PFNMAPVMAs: without.test_walk, defaultwalk_page_test()skipsVM_PFNMAPsilently. Callers needing all VMAs must provide.test_walkreturning 0. A 0 return fromwalk_page_range()may mean "skipped", not "handled"ACTION_AGAINin page walk callbacks:ACTION_AGAINretries with no limit.pte_offset_map_lock()returns NULL non-transiently for migration entries — settingACTION_AGAINon this failure creates an infinite loop. Return 0 to skip gracefully.walk_pte_range()already handles retry internally; callbacks should not duplicate it- Page comparison for zeropage remapping must use
pages_identical(): rawmemchr_inv()/memcmp()miss architecture metadata. On arm64 MTE, byte-identical pages with different tags cause mismatch faults after remapping toZERO_PAGE(0).pages_identical()has an arm64 override rejecting MTE-tagged pages. REPORT as bugs:memchr_inv()/memcmp()for zeropage/merge decisions