From f166586f74dd5d9cbadaabf86ef81c8ddf6cafa7 Mon Sep 17 00:00:00 2001 From: Nathan Gao Date: Thu, 3 Sep 2026 17:28:27 -0700 Subject: [PATCH] mm/damon/ops-common: use a page-aligned address in damon_ptep_mkold() __damon_va_prepare_access_check() picks a random byte address within the region and stores it in r->sampling_addr. damon_va_mkold() passes it into a page table walk, which hands it to damon_ptep_mkold() as the address of the page to sample: damon_va_mkold(mm, r->sampling_addr) damon_va_walk_page_range(mm, addr, addr + 1) damon_mkold_pmd_entry() damon_ptep_mkold(pte, vma, addr) ptep_test_and_clear_young(vma, addr, pte) mmu_notifier_clear_young(mm, addr, addr + PAGE_SIZE) For arm64, before commit 6f0e1142173a ("arm64: mm: support batch clearing of the young flag for large folios"), the contpte helper walked exactly CONT_PTES entries from the aligned-down page table pointer and used @addr only to pass down to each entry, so an unaligned value was harmless: ptep = contpte_align_down(ptep); addr = ALIGN_DOWN(addr, CONT_PTE_SIZE); for (i = 0; i < CONT_PTES; i++, ptep++, addr += PAGE_SIZE) Now the range to walk is derived from @addr instead: end = addr + nr * PAGE_SIZE, rounded up to CONT_PTE_SIZE. For a sample in the last page of a contpte block, the sub-page offset puts end just past the block boundary, so the round-up lands a whole block further and the walk clears PTE_AF in CONT_PTES entries beyond the sampled block. For the last block in a page table page, those entries are past the end of that page, so the walk writes into the page that follows. Triggered by the full 7.1/7.2 kernel selftest suite on arm64 (EC2 c/m6g.4xlarge). The kernel sometimes crashes at or shortly after the DAMON test. What the overrun does depends on the page that happens to follow the page table, so there is no single signature. If that page is read-only, the write faults in the sampling path itself: Unable to handle kernel write to read-only memory at virtual address ffff0003c5d2d000 FSC = 0x0f: level 3 permission fault CM = 0, WnR = 1, TnD = 0, TagAccess = 0 CPU: 10 UID: 0 PID: 3487 Comm: kdamond.2 pc : contpte_test_and_clear_young_ptes+0x70/0xc0 lr : damon_ptep_mkold+0x1e8/0x1f8 Call trace: contpte_test_and_clear_young_ptes+0x70/0xc0 (P) damon_mkold_pmd_entry+0x150/0x170 walk_pmd_range+0x110/0x2b0 walk_pud_range+0x10c/0x208 walk_pgd_range+0x134/0x258 __walk_page_range+0x98/0x1b0 walk_page_range_vma_unsafe+0x90/0x148 walk_page_range_vma+0x28/0x40 damon_va_walk_page_range+0x114/0x2b8 damon_va_prepare_access_checks+0xec/0x1a8 kdamond_fn+0x534/0x770 kthread+0x128/0x138 ret_from_fork+0x10/0x20 Otherwise the page is writable, the PTE_AF clearing succeeds silently and the damage only surfaces later, in whatever happened to own the page, so the backtrace is unrelated to DAMON and differs between runs. Pass a page-aligned address to the ptep_test_and_clear_young() call in damon_ptep_mkold(), which is the only place DAMON can reach contpte_test_and_clear_young_ptes() from. Nothing else sees the aligned address, and r->sampling_addr itself is left as is, so the sampling and region bookkeeping semantics are unchanged. Link: https://lore.kernel.org/20260904002829.116381-1-sj@kernel.org Fixes: 6f0e1142173a ("arm64: mm: support batch clearing of the young flag for large folios") Signed-off-by: Nathan Gao Signed-off-by: SJ Park Signed-off-by: Andrew Morton Reviewed-by: SJ Park Reviewed-by: Baolin Wang Cc: Baolin Wang Cc: David Hildenbrand (Arm) Cc: Ryan Roberts Cc: --- mm/damon/ops-common.c | 7 ++++++- 1 file changed, 6 insertions(+), 1 deletion(-) diff --git a/mm/damon/ops-common.c b/mm/damon/ops-common.c index fbda70d8ea4d..8fc61d06d358 100644 --- a/mm/damon/ops-common.c +++ b/mm/damon/ops-common.c @@ -61,7 +61,12 @@ void damon_ptep_mkold(pte_t *pte, struct vm_area_struct *vma, unsigned long addr * device aspects. */ if (likely(pte_present(pteval))) - young |= ptep_test_and_clear_young(vma, addr, pte); + /* + * Arch implementation of ptep_test_and_clear_young() may + * require aligned @addr + */ + young |= ptep_test_and_clear_young(vma, PAGE_ALIGN_DOWN(addr), + pte); young |= mmu_notifier_clear_young(vma->vm_mm, addr, addr + PAGE_SIZE); if (young) folio_set_young(folio);