mm/damon/vaddr: avoid hw-driven pte updates during damon_hugetlb_mkold()

damon_hugetlb_mkold() reads the page table entry into a local variable,
unsets the accessed bit in the variable, and updates the page table entry
with the updated variable value.  If hardware updates the same page table
entry in parallel, the hw updates could be lost.  For example,
hardware-updated dirty bits might be lost.

Avoid the parallel updates by clearing the page table entry when reading
it together, using huge_ptep_get_and_clear().  If a parallel write to the
memory is made after the clearing, the hw will see the page table entry is
cleared, trigger page fault and wait until it is handled.  The page fault
handling will wait for damon_hugetlb_mkold() due to the page table lock.

Because hugetlbfs is an in-memory file system and hugetlb pages cannot be
reclaimed, no critical issue is expected to my best knowledge.  But
definitely this is a nasty bug that should be fixed sooner rather than
later.

The issue was discovered [1] by Sashiko.

Link: https://lore.kernel.org/20260907170358.100168-1-sj@kernel.org
Link: https://lore.kernel.org/20260830160545.98969-1-sj@kernel.org [1]
Fixes: 49f4203aae ("mm/damon: add access checking for hugetlb pages")
Signed-off-by: SJ Park <sj@kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Cc: Baolin Wang <baolin.wang@linux.alibaba.com>
Cc: <stable@vger.kernel.org> # 5.17.x
This commit is contained in:
SJ Park 2026-09-07 10:03:56 -07:00 committed by Andrew Morton
parent 90179da203
commit 39c0ceedd5

View File

@ -293,22 +293,29 @@ static int damon_mkold_pmd_entry(pmd_t *pmd, unsigned long addr,
}
#ifdef CONFIG_HUGETLB_PAGE
static bool damon_hugetlb_ptep_mkold(pte_t *pte, struct mm_struct *mm,
struct vm_area_struct *vma, unsigned long addr, pte_t *entry)
{
unsigned long psize = huge_page_size(hstate_vma(vma));
if (!pte_young(*entry))
return false;
*entry = huge_ptep_get_and_clear(mm, addr, pte, psize);
*entry = pte_mkold(*entry);
set_huge_pte_at(mm, addr, pte, *entry, psize);
return true;
}
static void damon_hugetlb_mkold(pte_t *pte, struct mm_struct *mm,
struct vm_area_struct *vma, unsigned long addr)
{
bool referenced = false;
pte_t entry = huge_ptep_get(mm, addr, pte);
struct folio *folio = pfn_folio(pte_pfn(entry));
unsigned long psize = huge_page_size(hstate_vma(vma));
folio_get(folio);
if (pte_young(entry)) {
referenced = true;
entry = pte_mkold(entry);
set_huge_pte_at(mm, addr, pte, entry, psize);
}
referenced = damon_hugetlb_ptep_mkold(pte, mm, vma, addr, &entry);
if (mmu_notifier_clear_young(mm, addr,
addr + huge_page_size(hstate_vma(vma))))
referenced = true;