Commit Graph

1481928 Commits

Author SHA1 Message Date
SJ Park
eb64948249 mm/damon/core: reset invalid quota->charge_target_from
DAMOS can suddenly stop working if a target process that the quota is just
fully charged on is terminated.  Fix by catching and processing the corner
case.

When DAMOS quota is fully charged, the target and the region to continue
applying the action in the next round is saved in
damos_quota->charge_{target,addr}_from.  In the next round, DAMOS iterates
targets and regions from the beginning.  It skips applying the action to
the regions until it visits and skips the saved target/region.

Virtual address space targets become invalid if the process is terminated.
Trying to apply the scheme to invalid target is just a waste of time. 
Hence commit 6e4930e333 ("mm/damon/core: fix wasteful CPU calls by
skipping non-existent targets") made the logic to skip invalid targets. 
However, it does skip before the charged target/region skipping/updating.

Let's suppose the user runs DAMOS for multiple virtual address spaces with
a quota.  The quota exceeded in the middle of a virtual address space. 
And the process of the address space is terminated.  Then the
charge_target_from points to the invalid target.  The pointer update logic
is skipped for the invalid target, so the charge_target_from is never
updated.  DAMOS action to every target/region is skipped.  From the user's
perspective, it would look like suddenly DAMOS has stopped working.

No critical leak or crash can happen.  The user could reinstall the
scheme.  But this makes use of DAMOS under certain setups quite
unreliable.

When the invalid target is found, further check the corner case and reset
the pointer.

This issue was discovered [1] by Sashiko.

Link: https://lore.kernel.org/20260910142846.172957-1-sj@kernel.org
Link: https://lore.kernel.org/20260830064708.40CA61F000E9@smtp.kernel.org [1]
Fixes: 6e4930e333 ("mm/damon/core: fix wasteful CPU calls by skipping non-existent targets")
Signed-off-by: SJ Park <sj@kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Cc: Enze Li <lienze@kylinos.cn>
Cc: <stable@vger.kernel.org> # 7.0.x
2026-09-18 08:20:14 -07:00
Baolin Wang
44fcc0bfb0 MAINTAINERS: add Baoquan and Baolin as MGLRU reviewers
Baoquan and I have been contributing MGLRU patches and helping review
MGLRU related patches for some time.  We will continue to follow MGLRU
changes, so we'd like to be CC'd on MGLRU related patches.

Link: https://lore.kernel.org/06e20ef4f603a4ffeafcdbf623ce2936281457bf.1788831480.git.baolin.wang@linux.alibaba.com
Signed-off-by: Baolin Wang <baolin.wang@linux.alibaba.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Acked-by: Barry Song <baohua@kernel.org>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
Acked-by: Baoquan He <baoquan.he@linux.dev>
Acked-by: Kairui Song <kasong@tencent.com>
Cc: Axel Rasmussen <axelrasmussen@google.com>
Cc: Shakeel Butt <shakeel.butt@linux.dev>
Cc: Wei Xu <weixugc@google.com>
Cc: Yuanchu Xie <yuanchu@google.com>
2026-09-18 08:20:14 -07:00
Jinjiang Tu
b6ac0b3f60 mm/rmap: fix missing barrier between anon_vma init and vma->anon_vma publish
On arm64 server, we find that a task trying to grab the anon_vma lock
triggers hungtask.

INFO: task main:2354726 blocked for more than 120 seconds.
      Tainted: G            E     5.10.0-0021.aarch64 #1
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
task:main            state:D stack:    0 pid:2354726 ppid:2350673 flags:0x00000a01
Call trace:
 __switch_to+0x7c/0xbc
 __schedule+0x3b4/0x8a0
 schedule+0x50/0xe0
 rwsem_down_write_slowpath+0x3cc/0x6cc
 down_write+0x60/0x260
 __anon_vma_prepare+0x6c/0x210
 do_anonymous_page+0x258/0x660
 handle_pte_fault+0x188/0x214
 __handle_mm_fault+0x1b0/0x380
 handle_mm_fault+0xf4/0x284
 do_page_fault+0x19c/0x494
 do_translation_fault+0xcc/0xf8
 do_mem_abort+0x48/0xac
 el0_da+0x44/0x80
 el0_sync_handler+0x88/0xb4
 el0_sync+0x160/0x180

After analyzing the vmcore, we found the anon_vma->root->rwsem.count is
-1.  There is another anon_vma whose anon_vma->root->rwsem.count is 1, the
anon_vma->root->rwsem.owner shows the lock is held, but the stack of the
task shows the task doesn't hold the anon_vma lock.

After adding more debugging info, we found __anon_vma_prepare() reuses
anon_vma and triggers the UAF of anon_vma->root due to missing memory
barrier, leading to locking and unlocking two different anon_vma->root,
thus leading to an anon_vma will never be unlocked, and another anon_vma
couldn't be locked anymore.

This race requires two adjacent VMAs that are not merged but are
anon_vma-compatible (e.g., they differ in VMA_ACCESS_FLAGS that can be
changed by mprotect()).  Two threads fault on each VMA concurrently, both
calling __anon_vma_prepare() with only mmap_lock held for reading.

    THREAD A                             THREAD B
__anon_vma_prepare                __anon_vma_prepare
 find_mergeable_anon_vma() -> NULL
 anon_vma = anon_vma_alloc();
   anon_vma->root = anon_vma;
 // the two stores may be reordered
 vma->anon_vma = anon_vma;
                                   // finds A's anon_vma
                                   anon_vma = find_mergeable_anon_vma(vma);
                                   anon_vma_lock_write(anon_vma);
                                     // may still see the old root
                                     down_write(&anon_vma->root->rwsem);
                                   anon_vma_unlock_write(anon_vma);
                                     // see the new root, never unlock old
                                     up_write(&anon_vma->root->rwsem);

thread A triggers page fault and calls __anon_vma_prepare() to prepare
anon_vma for the faulting vma.  __anon_vma_prepare() allocates and
initializes a new anon_vma, and then publishes it to the vma with a plain
store.  anon_vma_prepare() only requires the mmap_lock to be held for
reading, so two threads can fault on adjacent VMAs at the same time. 
While thread A publishes a new anon_vma, thread B could find the anon_vma
via find_mergeable_anon_vma() and then locks anon_vma->root->rwsem.

The store to anon_vma->root in anon_vma_alloc() and the store to
vma->anon_vma can be reordered.  The anon_vma_lock_write() and spin_lock()
only provide acquire semantics, which do not prevent prior stores from
being reordered after them.  The release semantics of the corresponding
spin_unlock() and anon_vma_unlock_write() come too late, the store to
vma->anon_vma is already published before they take effect.  As a result,
thread B can observe the following order:

    vma->anon_vma = anon_vma;
    anon_vma->root = anon_vma;

The anon_vma slab is SLAB_TYPESAFE_BY_RCU, so a newly allocated anon_vma
may reuse memory from a previously freed one.  The constructor
(anon_vma_ctor) does not reset anon_vma->root, and __put_anon_vma()
doesn't clear it either, so the old root value persists until
anon_vma_alloc() overwrites it.  If that store isn't visible, thread B
reads a root that points to the old anon_vma and locks it.

As a result, thread B can call anon_vma_lock_write() with the old root,
and call anon_vma_unlock_write() with the new root, leading to an anon_vma
will never be unlocked, and another anon_vma couldn't be locked anymore
(its count is dropped from 0 to -1 due to wrong unlock).

To fix it, change the plain store `vma->anon_vma = anon_vma` to store
release, so that the fields of anon_vma are visible before anon_vma is
published to vma->anon_vma.

At read side, the load of anon_vma and anon_vma->root have address
dependency.  According to Documentation/memory-barriers.txt and some
investigations, only Alpha needs address-dependency barriers and it has
been handled by READ_ONCE() in reusable_anon_vma().

We reproduced this issue in v5.10 with KSM enabled.  The kernel doesn't
merge commit cf7e7a3503 ("mm: prevent KSM from breaking VMA merging for
new VMAs"), so there are many adjacent VMAs that aren't merged but are
compatible for anon_vma.

Without this fix, our production environment could reproduce this issue
about 2-5 times each month.  After adding a smp_mb() before
anon_vma_lock_write(anon_vma) in __anon_vma_prepare(), which is different
to this patch, this issue hasn't been reproduced for one month.

Link: https://lore.kernel.org/20260908122924.554373-1-tujinjiang@huawei.com
Fixes: 5c341ee1df ("mm: track the root (oldest) anon_vma")
Signed-off-by: Jinjiang Tu <tujinjiang@huawei.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Reviewed-by: Lance Yang <lance.yang@linux.dev>
Reviewed-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Cc: Minchan Kim <minchan@kernel.org>
Cc: Harry Yoo <harry@kernel.org>
Cc: Hiroyouki Kamezawa <kamezawa.hiroyu@jp.fujitsu.com>
Cc: Jann Horn <jannh@google.com>
Cc: Jinjiang Tu <tujinjiang@huawei.com>
Cc: Kefeng Wang <wangkefeng.wang@huawei.com>
Cc: Larry Woodman <lwoodman@redhat.com>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Nanyong Sun <sunnanyong@huawei.com>
Cc: Rik van Riel <riel@surriel.com>
Cc: <stable@vger.kernel.org>
2026-09-18 08:20:14 -07:00
Jaewook You
9bdad082d4 mm/hugetlb: preserve mremap address delta when skipping page tables
move_hugetlb_page_tables() optimizes mremap() by advancing to the last
entry in the page table when the source page table does not exist, either
initially or after unsharing a PMD table.  The common loop increment then
steps to the first entry in the next page table.

However, the code advances both the source and destination addresses to
the last entries in their respective page tables, which is wrong.  The
destination address must be advanced only by the same amount as the source
address.

If the source and destination offsets within their page tables differ, the
destination address can be advanced too far, causing follow-up issues. 
Fix this by advancing the destination address by the source advance
distance.

With a reproducer, we were able to trigger a kernel panic on x86-64.  With
this fix in place, we can no longer reproduce the issue.

Link: https://lore.kernel.org/20260914132352.472-1-jaewook376@gmail.com
Fixes: e95a985178 ("hugetlb: skip to end of PT page mapping when pte not present")
Fixes: 4ddb4d91b8 ("hugetlb: do not update address in huge_pmd_unshare")
Signed-off-by: Jaewook You <jaewook376@gmail.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Cc: Johan Hovold <johan@kernel.org>
Cc: Muchun Song <muchun.song@linux.dev>
Cc: Oscar Salvador <osalvador@suse.de>
Cc: <stable@vger.kernel.org>
Assisted-by: LLM
2026-09-18 08:20:14 -07:00
Liew Rui Yan
b3723b596b mm/damon/core: fix unconditionally skip last region
Once quota set, the charge_{target,addr}_from unconditionally skips and
resets at the last region of the tracked target, so the last region can be
skipped even when it has not been processed.

Example:

    1. Target has 2 regions: R1 (0-100 bytes) and R2 (100-200 bytes).
    2. Quota is configured to process only 100 bytes per window.
    3. Window 1: Processes R1 (0-100).  Quota is full.  charge_{target,
       addr}_from is saved at (Target, 100).
    4. Window 2: The loop reaches R2.  Because R2 is
       damon_last_region(t), the old code unconditionally returns true,
       skipping R2 entirely and resetting the charge_{target,addr}_from.

    Result: R2 is permanently skipped even though it has never been
    processed.

However, it is important to note that this is a very minor issue.  This is
because it is triggered only when the previous window saved/kept
charge_{target,addr}_from, and in the next window, all regions except the
last region were skipped by damos_skip_charged_region().

Fix this by only resetting the charge_{target,addr}_from when last region
is reached, only skipping when it is applied or cannot split.

Link: https://lore.kernel.org/20260908134739.96919-1-sj@kernel.org
Fixes: 50585192bc ("mm/damon/schemes: skip already charged targets and regions")
Signed-off-by: Liew Rui Yan <aethernet65535@gmail.com>
Reviewed-by: SJ Park <sj@kernel.org>
Signed-off-by: SJ Park <sj@kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Cc: <stable@vger.kernel.org> # v5.16.x
2026-09-18 08:20:14 -07:00
SJ Park
39c0ceedd5 mm/damon/vaddr: avoid hw-driven pte updates during damon_hugetlb_mkold()
damon_hugetlb_mkold() reads the page table entry into a local variable,
unsets the accessed bit in the variable, and updates the page table entry
with the updated variable value.  If hardware updates the same page table
entry in parallel, the hw updates could be lost.  For example,
hardware-updated dirty bits might be lost.

Avoid the parallel updates by clearing the page table entry when reading
it together, using huge_ptep_get_and_clear().  If a parallel write to the
memory is made after the clearing, the hw will see the page table entry is
cleared, trigger page fault and wait until it is handled.  The page fault
handling will wait for damon_hugetlb_mkold() due to the page table lock.

Because hugetlbfs is an in-memory file system and hugetlb pages cannot be
reclaimed, no critical issue is expected to my best knowledge.  But
definitely this is a nasty bug that should be fixed sooner rather than
later.

The issue was discovered [1] by Sashiko.

Link: https://lore.kernel.org/20260907170358.100168-1-sj@kernel.org
Link: https://lore.kernel.org/20260830160545.98969-1-sj@kernel.org [1]
Fixes: 49f4203aae ("mm/damon: add access checking for hugetlb pages")
Signed-off-by: SJ Park <sj@kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Cc: Baolin Wang <baolin.wang@linux.alibaba.com>
Cc: <stable@vger.kernel.org> # 5.17.x
2026-09-18 08:20:14 -07:00
Liew Rui Yan
90179da203 mm/damon/core: allow esz to be set to zero
When the temporal quota goal tuner determines that the goal has been
achieved (score >= 10000), it sets esz_bp to zero so that the esz becomes
zero.  However, damos_set_effective_quota() clamps the esz to
min_region_sz when quota->ms is set.

This is a minor issue, the main problem is that it doesn't match the
description in the documentation, which state that if the goal has already
been [over-]achieved, the quota will be set to zero.

Fix this by set quota (esz) as minimum as possible.

Link: https://lore.kernel.org/20260908135413.97570-1-sj@kernel.org
Fixes: 8bbde987c2 ("mm/damon/core: disallow time-quota setting zero esz")
Signed-off-by: SJ Park <sj@kernel.org>
Signed-off-by: Liew Rui Yan <aethernet65535@gmail.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Reviewed-by: SJ Park <sj@kernel.org>
Cc: <stable@vger.kernel.org> # v7.1.x
2026-09-18 08:20:13 -07:00
Nathan Gao
f166586f74 mm/damon/ops-common: use a page-aligned address in damon_ptep_mkold()
__damon_va_prepare_access_check() picks a random byte address within the
region and stores it in r->sampling_addr.  damon_va_mkold() passes it into
a page table walk, which hands it to damon_ptep_mkold() as the address of
the page to sample:

  damon_va_mkold(mm, r->sampling_addr)
    damon_va_walk_page_range(mm, addr, addr + 1)
      damon_mkold_pmd_entry()
        damon_ptep_mkold(pte, vma, addr)
          ptep_test_and_clear_young(vma, addr, pte)
          mmu_notifier_clear_young(mm, addr, addr + PAGE_SIZE)

For arm64, before commit 6f0e114217 ("arm64: mm: support batch clearing
of the young flag for large folios"), the contpte helper walked exactly
CONT_PTES entries from the aligned-down page table pointer and used @addr
only to pass down to each entry, so an unaligned value was harmless:

        ptep = contpte_align_down(ptep);
        addr = ALIGN_DOWN(addr, CONT_PTE_SIZE);
        for (i = 0; i < CONT_PTES; i++, ptep++, addr += PAGE_SIZE)

Now the range to walk is derived from @addr instead: end = addr + nr *
PAGE_SIZE, rounded up to CONT_PTE_SIZE.  For a sample in the last page of
a contpte block, the sub-page offset puts end just past the block
boundary, so the round-up lands a whole block further and the walk clears
PTE_AF in CONT_PTES entries beyond the sampled block.  For the last block
in a page table page, those entries are past the end of that page, so the
walk writes into the page that follows.

Triggered by the full 7.1/7.2 kernel selftest suite on arm64 (EC2
c/m6g.4xlarge).  The kernel sometimes crashes at or shortly after the
DAMON test.

What the overrun does depends on the page that happens to follow the page
table, so there is no single signature.  If that page is read-only, the
write faults in the sampling path itself:

  Unable to handle kernel write to read-only memory at virtual address ffff0003c5d2d000
    FSC = 0x0f: level 3 permission fault
    CM = 0, WnR = 1, TnD = 0, TagAccess = 0
  CPU: 10 UID: 0 PID: 3487 Comm: kdamond.2
  pc : contpte_test_and_clear_young_ptes+0x70/0xc0
  lr : damon_ptep_mkold+0x1e8/0x1f8
  Call trace:
   contpte_test_and_clear_young_ptes+0x70/0xc0 (P)
   damon_mkold_pmd_entry+0x150/0x170
   walk_pmd_range+0x110/0x2b0
   walk_pud_range+0x10c/0x208
   walk_pgd_range+0x134/0x258
   __walk_page_range+0x98/0x1b0
   walk_page_range_vma_unsafe+0x90/0x148
   walk_page_range_vma+0x28/0x40
   damon_va_walk_page_range+0x114/0x2b8
   damon_va_prepare_access_checks+0xec/0x1a8
   kdamond_fn+0x534/0x770
   kthread+0x128/0x138
   ret_from_fork+0x10/0x20

Otherwise the page is writable, the PTE_AF clearing succeeds silently and
the damage only surfaces later, in whatever happened to own the page, so
the backtrace is unrelated to DAMON and differs between runs.

Pass a page-aligned address to the ptep_test_and_clear_young() call in
damon_ptep_mkold(), which is the only place DAMON can reach
contpte_test_and_clear_young_ptes() from.  Nothing else sees the aligned
address, and r->sampling_addr itself is left as is, so the sampling and
region bookkeeping semantics are unchanged.

Link: https://lore.kernel.org/20260904002829.116381-1-sj@kernel.org
Fixes: 6f0e114217 ("arm64: mm: support batch clearing of the young flag for large folios")
Signed-off-by: Nathan Gao <zcgao@amazon.com>
Signed-off-by: SJ Park <sj@kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Reviewed-by: SJ Park <sj@kernel.org>
Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
Cc: Baolin Wang <baolin.wang@linux.alibaba.com>
Cc: David Hildenbrand (Arm) <david@kernel.org>
Cc: Ryan Roberts <ryan.roberts@arm.com>
Cc: <stable@vger.kernel.org>
2026-09-18 08:20:13 -07:00
Joseph Qi
525c0edc03 ocfs2: make ocfs2_calc_xattr_init() return void
ocfs2_calc_xattr_init() used to read the default ACL off the parent inode
itself, so it could return an error from ocfs2_xattr_get_nolock().  Commit
bd7c05fb4a ("ocfs2: fix circular locking dependency in
ocfs2_init_acl()") moved that lookup before the transaction starts and
deleted the error path, but left the now vestigial 'int ret = 0'
declaration and both 'return ret' statements behind, along with an
unreachable error branch in ocfs2_mknod().

Drop the leftover variable and convert the return type to void, so the
callee states that it always succeeds and the caller no longer carries a
check that can never trigger.

No functional change.

Link: https://lore.kernel.org/20260904023751.3703334-1-joseph.qi@linux.alibaba.com
Fixes: bd7c05fb4a ("ocfs2: fix circular locking dependency in ocfs2_init_acl()")
Signed-off-by: Joseph Qi <joseph.qi@linux.alibaba.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Reported-by: kernel test robot <lkp@intel.com>
Closes: https://lore.kernel.org/oe-kbuild-all/202609040247.8B3lmoqX-lkp@intel.com/
Cc: Mark Fasheh <mark@fasheh.com>
Cc: Joel Becker <jlbec@evilplan.org>
Cc: Junxiao Bi <junxiao.bi@oracle.com>
Cc: Changwei Ge <gechangwei@live.cn>
Cc: Jun Piao <piaojun@huawei.com>
Cc: Heming Zhao <heming.zhao@suse.com>
2026-09-18 08:20:13 -07:00
Haowen Bai
8d50c2f37b mailmap: update Haowen Bai's email address
Map Haowen Bai's former Meizu and current Ugreen addresses to
baihaowen88@gmail.com as the canonical public address.

Link: https://lore.kernel.org/20260903132854.1930923-1-calvin.bai@ugreen.com
Signed-off-by: Haowen Bai <calvin.bai@ugreen.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-09-18 08:20:13 -07:00
Joshua Hahn
8c7fdc0b4c selftests/cgroup: account for zswap shrinker writeback
The test_no_invasive_cgroup_shrink selftest checks that when a cgroup has
zswapped out more memory than memory.zswap.max, it does not trigger
writeback for other cgroups.  To do this, it compares the writeback count
in a control cgroup and makes sure that it is 0, and then checks the
writeback count in an aggressor cgroup who does expect to see writeback.

However, when the zswap shrinker is enabled, the victim cgroup can see
legitimate writebacks not triggered by the aggressor.  In some Meta CI
tests, we have seen this failure mode happen.

Instead of checking that the victim cgroup has 0 writeback, compare the
writeback values before and after the aggressor runs and check that the
victim cgroup did not perform any additional writeback.  Note that this
can still lead to probabilistic failures if writebacks take longer than
5 seconds, but this should fix the systematic failure case and make
"not ok test_no_invasive_cgroup_shrink" less likely.

Link: https://lore.kernel.org/20260902194521.3652178-1-joshua.hahnjy@gmail.com
Fixes: b5ba474f3f ("zswap: shrink zswap pool based on memory pressure")
Signed-off-by: Joshua Hahn <joshua.hahnjy@gmail.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Reported-by: Krush Chavan <krushchavan@outlook.com>
Suggested-by: Nhat Pham <nphamcs@gmail.com>
Cc: <stable@vger.kernel.org>
2026-09-18 08:20:13 -07:00
Longlong Xia
a363c62a65 mm/hugetlb: do not dissolve gigantic pages without runtime support
dissolve_free_hugetlb_folio() doesn't check
hstate_is_gigantic_no_runtime(h) though remove_hugetlb_folio()/
update_and_free_hugetlb_folio() silently bail for such folios, so it frees
a still-listed folio and, on vmemmap restore failure, the
add_hugetlb_folio() rollback corrupts the free list.

Link: https://lore.kernel.org/20260823044118.1097121-2-xialonglong2025@163.com
Fixes: 6eb4e88a6d ("hugetlb: create remove_hugetlb_page() to separate functionality")
Signed-off-by: Longlong Xia <xialonglong@kylinos.cn>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Assisted-by: Codex:gpt-5.6-sol
Acked-by: Muchun Song <muchun.song@linux.dev>
Cc: David Hildenbrand <david@kernel.org>
Cc: Miaohe Lin <linmiaohe@huawei.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Oscar Salvador <osalvador@suse.de>
Cc: <stable@vger.kernel.org>
2026-09-17 22:42:40 -07:00
Ackerley Tng
7891fbb951 mm/folio: EXPORT_SYMBOL_FOR_KVM(lru_cache_drain_for_folio)
To simplify independent development in the KVM and MM subsystems, now
export to KVM the lru_cache_drain_for_folio() which MM added in 7.3-rc1.

Link: https://lore.kernel.org/lkml/bd6c9c74-e374-a9d3-ba1f-8b6f430894fc@google.com/T/#u
Link: https://lore.kernel.org/02876cea-5727-2ca4-bead-73659ea6fec4@google.com
Signed-off-by: Ackerley Tng <ackerleytng@google.com>
Signed-off-by: Hugh Dickins <hughd@google.com>
Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Suggested-by: David Hildenbrand <david@kernel.org>
Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
Cc: Sean Christopherson <seanjc@google.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-09-08 23:28:06 -07:00
Jiayuan Chen
932cfb25e7 mm/shrinker: fix bogus set_shrinker_bit() with cgroup.memory=nokmem
With cgroup.memory=nokmem, shrinker_memcg_alloc() bails out early and
never allocates an id, so shrinker->id keeps the 0 it got from the
kzalloc() in shrinker_alloc().  __list_lru_init() then copies that 0 into
lru->shrinker_id, where it looks like a valid bit index.

Nothing calls expand_shrinker_info() on nokmem either, so shrinker_nr_max
stays 0 and every memcg ends up with an empty map (map_nr_max == 0).

deferred_split_folio() hands a real memcg to __list_lru_add() regardless
of whether the lru is memcg aware, so the first THP queued in a cgroup
does set_shrinker_bit(memcg, nid, 0) and trips the bounds check:

WARNING: mm/shrinker.c:212 at set_shrinker_bit+0x7d/0x90, CPU#126
Call Trace:
 <TASK>
 deferred_split_folio+0x18c/0x220
 map_anon_folio_pmd_nopf+0xdd/0x130
 map_anon_folio_pmd_pf+0x14/0xb0
 do_huge_pmd_anonymous_page+0x1a1/0x620
 __handle_mm_fault+0xea9/0x10d0
 handle_mm_fault+0xe5/0x320
 do_user_addr_fault+0x1cc/0x870
 exc_page_fault+0x81/0x1b0
 asm_exc_page_fault+0x27/0x30
 </TASK>

Harmless, the WARN_ON_ONCE() is what keeps the out of bounds unit[] read
from happening, but the id should not look valid in the first place. 
Clear it before returning.

Two other spots could paper over this: drop the id in __list_lru_init()
when nokmem turns memcg_aware off, or make deferred_split_folio() pass
NULL like list_lru_add_obj() does.  Both leave shrinker->id lying around
for the next caller, so fix it where the id is handed out.

Link: https://lore.kernel.org/20260902073800.305481-1-jiayuan.chen@linux.dev
Fixes: fafaeceb89 ("mm: switch deferred split shrinker to list_lru")
Signed-off-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Acked-by: Shakeel Butt <shakeel.butt@linux.dev>
Cc: Usama Arif <usama.arif@linux.dev>
Cc: Dave Chinner <david@fromorbit.com>
Cc: Johannes Weiner <hannes@cmpxchg.org>
Cc: Kairui Song <kasong@tencent.com>
Cc: Muchun Song <muchun.song@linux.dev>
Cc: Roman Gushchin <roman.gushchin@linux.dev>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-09-08 23:28:06 -07:00
Lorenzo Stoakes (ARM)
6cc27d8219 mm/vma: correctly unaccount on mmap_prepare() failure
__mmap_setup() accounts memory for relevant mappings via:

security_vm_enough_memory_mm()
  -> __vm_enough_memory()
    -> vm_acct_memory()

If __mmap_setup() fails, this indicates that this accounting did not take
place, and thus it's appropriate for __mmap_region() to jump to
abort_munmap.

However if call_mmap_prepare() fails, it also jumps there and any accounted
memory is not correctly unaccounted.

Fix this by handling each error separately.

Link: https://lore.kernel.org/20260902-fix-unaccount-mmap_prepare-v1-1-ea070189fdfb@kernel.org
Fixes: c84bf6dd2b ("mm: introduce new .mmap_prepare() file callback")
Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Cc: Jann Horn <jannh@google.com>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Pedro Falcato <pfalcato@suse.de>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-09-08 23:28:06 -07:00
Shakeel Butt
e14a345480 mm/mlock: use the IRQ-safe accessor for NR_MLOCK in __munlock_folio()
NR_MLOCK is updated from interrupt context.  __free_pages_prepare() clears
a stray PG_mlocked and adjusts NR_MLOCK, and a folio can reach it with the
flag still set from a bio completion handler:

  __free_pages_ok+0x6af/0x7a0
  <IRQ>
  __bio_release_pages+0xde/0x260
  __iomap_dio_bio_end_io+0x16e/0x1a0
  blk_update_request+0x14b/0x3d0
  blk_mq_end_request+0x18/0x30
  blk_done_softirq+0x49/0x60

The folio gets there like this.  A MAP_SHARED file mapping is mlocked, so
its page cache folios carry PG_mlocked, and an O_DIRECT write sourced from
that mapping GUP-pins those same folios.  munlock() then runs
mlock_vma_pages_range(), which clears VM_LOCKED before walking the page
tables to munlock each folio.  A concurrent hole punch reaches the folio
through the rmap (i_mmap_rwsem, not mmap_lock) and can land inside that
window: __folio_remove_rmap() -> munlock_vma_folio() sees VM_LOCKED
already clear, so it neither queues the folio on the mlock batch nor takes
a reference, and the pte it clears makes the pending mlock_pte_range()
walk skip the folio at its !pte_present() check.  filemap_remove_folio()
then drops the page cache reference, leaving the bio's pin as the last
one, released from the completion handler above.

So __zone_stat_mod_folio() here needs interrupts disabled, not merely
preemption, and __munlock_folio() has a path where they are not: when the
folio has already been taken off the LRU by somebody else the function
jumps straight to the counter update without taking the lruvec lock.  The
read-modify-write of the per-CPU NR_MLOCK diff can then be interrupted by
the softirq above, and one of the two decrements is lost, leaving Mlocked
in /proc/meminfo permanently overstated.

Use zone_stat_mod_folio().  mod_zone_state()'s this_cpu_try_cmpxchg() is
atomic against a same-CPU interrupt and retries, and on the path where the
lruvec lock is held its cost is negligible next to the lock itself.

The UNEVICTABLE_PG* events are deliberately left on the __ accessors: they
occupy different vm_event_states slots from the UNEVICTABLE_PGCLEARED that
__free_pages_prepare() bumps, and nothing updates those two from interrupt
context.

Link: https://lore.kernel.org/20260901180109.3797944-1-shakeel.butt@linux.dev
Fixes: 2fbb0c10d1 ("mm/munlock: mlock_page() munlock_page() batch by pagevec")
Signed-off-by: Shakeel Butt <shakeel.butt@linux.dev>
Reported-by: syzbot+cd2073ee6d958a8d0fcd@syzkaller.appspotmail.com
Closes: https://lore.kernel.org/linux-mm/6a931c5a.08e933ee.dbf97.0093.GAE@google.com/
Acked-by: Hugh Dickins <hughd@google.com>
Cc: Jann Horn <jannh@google.com>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Lorenzo Stoakes <ljs@kernel.org>
Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
Cc: Pedro Falcato <pfalcato@suse.de>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-09-05 11:47:49 -07:00
Andrew Morton
0791a234b3 remove old lib/alloc_tag.c
This was moved into mm/, but the original lib/ file somehow remained. 
Remove it.

Reported-by: Suren Baghdasaryan <surenb@google.com>
Cc: Lorenzo Stoakes <ljs@kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-09-05 11:47:49 -07:00
Seunguk Shin
8e2b861403 fs/dax: check zero or empty entry before converting xarray entry
Calling dax_to_folio() with empty entry causes kernel panic below when
booting a VM with DAX enabled storage.

This patch checks empty entry before calling dax_to_folio() on
dax_associate_entry(), dax_disassociate_entry(), and dax_busy_page().

Commit 98c183a4fc ("fs/dax: don't disassociate zero page entries") added
guards in the associate and disassociate paths, but the guards still come
after dax_to_folio(), and dax_busy_page() still has the same problem.

[    0.737679] EXT4-fs (pmem0p1): mounted filesystem 79676804-7c8b-491a-b2a6-9bae3c72af70 ro with ordered data mode. Quota mode: disabled.
[    0.737891] VFS: Mounted root (ext4 filesystem) readonly on device 259:1.
[    0.739119] devtmpfs: mounted
[    0.739476] Freeing unused kernel memory: 1920K
[    0.740156] Run /sbin/init as init process
[    0.740229]   with arguments:
[    0.740286]     /sbin/init
[    0.740321]   with environment:
[    0.740369]     HOME=/
[    0.740400]     TERM=linux
[    0.743162] Unable to handle kernel paging request at virtual address fffffdffbf000008
[    0.743285] Mem abort info:
[    0.743316]   ESR = 0x0000000096000006
[    0.743371]   EC = 0x25: DABT (current EL), IL = 32 bits
[    0.743444]   SET = 0, FnV = 0
[    0.743489]   EA = 0, S1PTW = 0
[    0.743545]   FSC = 0x06: level 2 translation fault
[    0.743610] Data abort info:
[    0.743656]   ISV = 0, ISS = 0x00000006, ISS2 = 0x00000000
[    0.743720]   CM = 0, WnR = 0, TnD = 0, TagAccess = 0
[    0.743785]   GCS = 0, Overlay = 0, DirtyBit = 0, Xs = 0
[    0.743848] swapper pgtable: 4k pages, 48-bit VAs, pgdp=00000000b9d17000
[    0.743931] [fffffdffbf000008] pgd=10000000bfa3d403, p4d=10000000bfa3d403, pud=1000000040bfe403, pmd=0000000000000000
[    0.744070] Internal error: Oops: 0000000096000006 [#1]  SMP
[    0.748888] CPU: 0 UID: 0 PID: 1 Comm: init Not tainted 6.18.4 #1 NONE
[    0.749421] pstate: 004000c5 (nzcv daIF +PAN -UAO -TCO -DIT -SSBS BTYPE=--)
[    0.749969] pc : dax_disassociate_entry.constprop.0+0x20/0x50
[    0.750444] lr : dax_insert_entry+0xcc/0x408
[    0.750802] sp : ffff80008000b9e0
[    0.751083] x29: ffff80008000b9e0 x28: 0000000000000000 x27: 0000000000000000
[    0.751682] x26: 0000000001963d01 x25: ffff0000004f7d90 x24: 0000000000000000
[    0.752264] x23: 0000000000000000 x22: ffff80008000bcc8 x21: 0000000000000011
[    0.752836] x20: ffff80008000ba90 x19: 0000000001963d01 x18: 0000000000000000
[    0.753407] x17: 0000000000000000 x16: 0000000000000000 x15: 0000000000000000
[    0.753970] x14: ffffbf3154b9ae70 x13: 0000000000000000 x12: ffffbf3154b9ae70
[    0.754548] x11: ffffffffffffffff x10: 0000000000000000 x9 : 0000000000000000
[    0.755122] x8 : 000000000000000d x7 : 000000000000001f x6 : 0000000000000000
[    0.755707] x5 : 0000000000000000 x4 : 0000000000000000 x3 : fffffdffc0000000
[    0.756287] x2 : 0000000000000008 x1 : 0000000040000000 x0 : fffffdffbf000000
[    0.756871] Call trace:
[    0.757107]  dax_disassociate_entry.constprop.0+0x20/0x50 (P)
[    0.757592]  dax_iomap_pte_fault+0x4fc/0x808
[    0.757951]  dax_iomap_fault+0x28/0x30
[    0.758258]  ext4_dax_huge_fault+0x80/0x2dc
[    0.758594]  ext4_dax_fault+0x10/0x3c
[    0.758892]  __do_fault+0x38/0x12c
[    0.759175]  __handle_mm_fault+0x530/0xcf0
[    0.759518]  handle_mm_fault+0xe4/0x230
[    0.759833]  do_page_fault+0x17c/0x4dc
[    0.760144]  do_translation_fault+0x30/0x38
[    0.760483]  do_mem_abort+0x40/0x8c
[    0.760771]  el0_ia+0x4c/0x170
[    0.761032]  el0t_64_sync_handler+0xd8/0xdc
[    0.761371]  el0t_64_sync+0x168/0x16c
[    0.761677] Code: f9453021 f2dfbfe3 cb813080 8b001860 (f9400401)
[    0.762168] ---[ end trace 0000000000000000 ]---
[    0.762550] note: init[1] exited with irqs disabled
[    0.762631] Kernel panic - not syncing: Attempted to kill init! exitcode=0x0000000b

Link: https://lore.kernel.org/m2y0enxtzk.fsf@arm.com
Fixes: 38607c62b3 ("fs/dax: properly refcount fs dax pages")
Signed-off-by: Seunguk Shin <seunguk.shin@arm.com>
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Alistair Popple <apopple@nvidia.com>
Reported-by: Kiara Grouwstra <cinereal@riseup.net>
Cc: Al Viro <viro@zeniv.linux.org.uk>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-09-05 11:47:49 -07:00
Qi Zheng
641aade99f fs: fix missed removal of super_fs_objects_eligible()
Commit 0ef8faff49 ("fs: push nr_cached_objects memcg gating into
individual filesystems") was meant to drop the blanket memcg gate in
fs/super.c and let each ->nr_cached_objects() implementation decide for
itself whether it is meaningful in per-memcg reclaim.  However, when
that patch was applied the removal of super_fs_objects_eligible() and
its two call sites in super_cache_scan() / super_cache_count() was
lost, so the helper is still gating every ->nr_cached_objects() hook
and 0ef8faff49 is effectively a no-op.

Consequences of the leftover gate:

   - XFS's inode-reclaim hook, which is intentionally driven from
     per-memcg contexts to free memcg-charged slab, is still
     short-circuited in fs/super.c  exactly the regression from
     commit 0baad6f9b9 ("fs/super: skip non-memcg-aware
     nr_cached_objects in memcg slab shrink") that 0ef8faff49 was
     written to undo. Memcg-charged XFS inode slab therefore keeps
     piling up under per-memcg pressure until global reclaim kicks in.

   - Any future ->nr_cached_objects()/->free_cached_objects() that
     grows memcg awareness is likewise blocked before it can run, so
     filesystems cannot opt in to per-memcg reclaim on their own 
     defeating the whole point of pushing the gating decision down
     into the callbacks.

Drop the leftover helper and its call sites so the intent of
0ef8faff49 actually takes effect.

Link: https://lore.kernel.org/cover.1786955972.git.zhengqi.arch@bytedance.com
Link: https://lore.kernel.org/3b038d373c70ebac7cdabfb0035bb91d1d6e6cfe.1786955972.git.zhengqi.arch@bytedance.com
Link: https://lore.kernel.org/all/20260715103516.2410175-1-usama.arif@linux.dev/ [0]
Fixes: 0ef8faff49 ("fs: push nr_cached_objects memcg gating into individual filesystems")
Signed-off-by: Qi Zheng <zhengqi.arch@bytedance.com>
Acked-by: Usama Arif <usama.arif@linux.dev>
Cc: Baolin Wang <baolin.wang@linux.alibaba.com>
Cc: Christian Brauner <brauner@kernel.org>
Cc: David Hildenbrand <david@kernel.org>
Cc: Hugh Dickins <hughd@google.com>
Cc: Johannes Weiner <hannes@cmpxchg.org>
Cc: Michal Hocko <mhocko@kernel.org>
Cc: Muchun Song <muchun.song@linux.dev>
Cc: Roman Gushchin <roman.gushchin@linux.dev>
Cc: Shakeel Butt <shakeel.butt@linux.dev>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-09-05 11:47:49 -07:00
Wenjie Qi
848d2ce2fc mm: filemap: retain mapped dropbehind folios
Fault-around can map ready dropbehind folios without going through the
normal page-cache lookup that clears dropbehind.  A mapping represents a
competing cached user, so retain the folio instead of forcibly unmapping
it when writeback completes.

For a mapped folio, folio_unmap_invalidate() can call
unmap_mapping_folio(), which takes i_mmap_rwsem and may sleep.  Retaining
mapped folios avoids this path when folio_end_dropbehind() runs in
non-preemptible task context.

Tal was able to trigger a sleeping-in-atomic warning due to this [1].

Unmapped dropbehind folios continue through the existing invalidation path.

Link: https://lore.kernel.org/4aba05e1a2c3b61cb337d373eb9b7a8db4ddd822.1788024049.git.qiwenjie@xiaomi.com
Link: https://lore.kernel.org/076bb01b-6fcf-4691-be8c-0e8507c9fe64@columbia.edu [1]
Fixes: fb7d3bc414 ("mm/filemap: drop streaming/uncached pages when writeback completes")
Signed-off-by: Wenjie Qi <qiwenjie@xiaomi.com>
Reviewed-by: Matthew Wilcox (Oracle) <willy@infradead.org>
Reviewed-by: Tal Zussman <tz2294@columbia.edu>
Tested-by: Tal Zussman <tz2294@columbia.edu>
Cc: Barry Song <baohua@kernel.org>
Cc: Jan Kara <jack@suse.cz>
Cc: Jens Axboe <axboe@kernel.dk>
Cc: Trond Myklebust <trond.myklebust@hammerspace.com>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-09-05 11:47:48 -07:00
Christopher Obbard
e1d469a8d6 mailmap: update entry for Christopher Obbard
I have changed employer; update my mailmap entry to point at my new email
address.

Link: https://lore.kernel.org/20260829-update-mail-oss-qualcomm-v2-1-1670c515f225@oss.qualcomm.com
Signed-off-by: Christopher Obbard <chris.obbard@oss.qualcomm.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-09-05 11:47:48 -07:00
Shakeel Butt
6e673d0879 memcg: avoid charging the root memcg from obj_cgroup_charge_pages()
obj_cgroup_charge_pages() resolves the objcg to its memcg and calls
try_charge_memcg(), which does not short circuit the root memcg.  That
memcg can be the root memcg: obj_cgroup_is_root() reflects the memcg the
objcg was created for and is never updated, while memcg_reparent_objcgs()
does redirect objcg->memcg to the parent on rmdir.  An objcg of a dying
child of root therefore passes every obj_cgroup_is_root() filter but
resolves to the root memcg.

Folios keep the objcg they were charged with, so this is easy to reach
through zswap: allocate anon memory in a cgroup, move the task out, remove
the cgroup, then write to the root cgroup's memory.reclaim.  The reclaimed
folios are charged through the reparented objcg and end up in
refill_stock() with the root memcg:

  WARNING: mm/memcontrol.c:2198 at refill_stock+0x644/0x940
   refill_stock+0x644/0x940
   try_charge_memcg+0x12d6/0x1570
   __obj_cgroup_charge+0x35/0xf0
   obj_cgroup_charge+0x1de/0x210
   obj_cgroup_charge_zswap+0x83/0x270
   zswap_store+0x1620/0x2000
   swap_writeout+0x94c/0x14c0
   shrink_folio_list+0x3388/0x52b0
   [...]
   try_to_free_mem_cgroup_pages+0x30d/0x830
   user_proactive_reclaim+0x504/0x840
   memory_reclaim+0x1f/0x30

Beyond the warning, the charge is asymmetric: obj_cgroup_uncharge_pages()
skips refill_stock() for the root memcg, so the root's page counter grows
and is never uncharged.  It is not user visible, since memory.current is
not exposed on the root, but it is a leak.

Use try_charge(), which returns early for the root memcg, restoring the
symmetry with obj_cgroup_uncharge_pages().

The above sequence was scripted into a standalone reproducer (zswap on,
swap on a virtio disk, 512MB of anon memory faulted in inside a child of
the root cgroup, the task then migrated to the root cgroup, the child
removed, followed by "echo 600M swappiness=max > memory.reclaim" on the
root) and run in a CONFIG_DEBUG_VM=y VM.  It reproduces the splat on the
first zswap store of a reparented folio, with the same call chain as the
report.  With this patch applied the splat is gone while the zswap store
count over the run is unchanged, so the same path is still exercised. 
cgroup selftests test_zswap, test_kmem and test_memcontrol show no new
failures.

Link: https://lore.kernel.org/20260829023251.474083-1-shakeel.butt@linux.dev
Fixes: 20d6c17252 ("memcg: avoid refill_stock for root memcg")
Signed-off-by: Shakeel Butt <shakeel.butt@linux.dev>
Reported-by: Farhad Alemi <farhad.alemi@berkeley.edu>
Closes: https://lore.kernel.org/all/CA+0ovCgWzUMK+nNbbtH7eV65Ca=fDN4Ozu7iASgryjvv8Tk8zQ@mail.gmail.com/
Reviewed-by: Muchun Song <muchun.song@linux.dev>
Reviewed-by: Johannes Weiner <hannes@cmpxchg.org>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Roman Gushchin <roman.gushchin@linux.dev>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-09-05 11:47:48 -07:00
Nhat Pham
12e9ac7bc5 mm, swap: fix SWAP_USAGE_OFFLIST_BIT collision with real usage count
SWAP_USAGE_OFFLIST_BIT is embedded in the si->inuse_pages usage counter,
and is meant to sit above any value that counter can reach.  However, it
is defined from BITS_PER_TYPE(atomic_t), so it is bit 30.  On a system
with 4 KiB pages the flag collides with the usage count once that count
reaches 4 TiB.

swap_usage_in_pages() masks bit 30 out, so whenever the real count has
that bit set, every caller of it reads 4 TiB low:

* /proc/swaps understates Used by 4 TiB.

* A raw count of exactly 2^30 masks to zero, so try_to_unuse() takes its
  "if (!swap_usage_in_pages(si)) goto success;" early exit and swapoff
  tears the device down while pages are still swapped out. Nothing in
  the rest of swapoff aborts the teardown, so those pages are lost.

Independently of swapoff, the collision also corrupts the counter and the
plist.  On a device in normal use, a free that leaves bit 30 set in the
count makes swap_usage_sub() see the flag where there is only count, and
call add_to_avail_list().  It clears the bit with
fetch_and(~SWAP_USAGE_OFFLIST_BIT), leaving the stored count 4 TiB below
the real one, and calls plist_add() on a device that is already listed,
tripping the WARN_ON(!plist_node_empty(node)) in plist_add() and linking
the node a second time.

Change the definition of SWAP_USAGE_OFFLIST_BIT to be based on
atomic_long_t instead.  Note that the usage counter field itself is of
this same type, so it is still a valid bit.

Link: https://lore.kernel.org/20260828191433.3304458-1-nphamcs@gmail.com
Fixes: b228386cf2 ("mm, swap: clean up plist removal and adding")
Signed-off-by: Nhat Pham <nphamcs@gmail.com>
Reported-by: Sashiko <sashiko-bot@kernel.org>
Closes: https://sashiko.dev/#/patchset/20260825153238.2695446-1-nphamcs%40gmail.com
Suggested-by: Andrew Morton <akpm@linux-foundation.org>
Reviewed-by: Andrew Morton <akpm@linux-foundation.org>
Acked-by: Kairui Song <kasong@tencent.com>
Cc: Baoquan He <baoquan.he@linux.dev>
Cc: Barry Song <baohua@kernel.org>
Cc: Chris Li <chrisl@kernel.org>
Cc: Gregory Price <gourry@gourry.net>
Cc: Johannes Weiner <hannes@cmpxchg.org>
Cc: Joshua Hahn <joshua.hahnjy@gmail.com>
Cc: Kemeng Shi <shikemeng@huaweicloud.com>
Cc: Shakeel Butt <shakeel.butt@linux.dev>
Cc: Youngjun Park <youngjun.park@lge.com>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-09-05 11:47:48 -07:00
Coiby Xu
e1d56f0465 mailmap: map Coiby Xu's address
Point to my gmail address as I've left Red Hat.

Link: https://lore.kernel.org/20260828084106.1494733-1-coiby.xu@gmail.com
Signed-off-by: Coiby Xu <coiby.xu@gmail.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-09-05 11:47:47 -07:00
Lorenzo Stoakes (ARM)
397432cab1 mm/mremap: account mm->locked_vm correctly for MREMAP_DONTUNMAP
When a VMA is mremap()'d with MREMAP_DONTUNMAP set, that results in the
VMA being copied, but the source VMA not being unmapped.

If the VMA is mlock()'d this is a legal operation, though the source VMA
has its VMA_LOCKED_BIT cleared.

However this is done in dontunmap_complete(), after mm->locked_vm was
incremented via vrm_stat_account(), resulting in double-counting.

Worse, this is not even corrected when source VMA is unmapped, due to the
VMA_LOCKED_BIT flag having been cleared.

This all works fine in the usual mremap() case (without MREMAP_DONTUNMAP),
as the source VMA is unmapped with VMA_LOCKED_BIT intact, at which time
mm->locked_vm is decremented accordingly.

Resolve the issue by invoking vrm_stat_account() only after
dontunmap_complete() has run.

Note that MREMAP_DONTUNMAP requires old_len == new_len, so no need to
account for a delta in size in this case.

The bug was introduced by commit b714ccb02a ("mm/mremap: complete
refactor of move_vma()") which incorrectly reordered the accounting and
the clearing of the VMA_LOCKED_BIT flag.

Link: https://lore.kernel.org/20260828-mremap-fix-locked-vm-v1-1-c80be7505d1e@kernel.org
Fixes: b714ccb02a ("mm/mremap: complete refactor of move_vma()")
Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Reported-by: sashiko-bot <sashiko-bot@kernel.org>
Closes: https://sashiko.dev/#/patchset/20260825-fix-mremap-dontunmap-pgoff-v1-1-39a40b2c98b3@kernel.org
Reported-by: Kunwu Chan <kunwu.chan@gmail.com>
Closes: https://lore.kernel.org/all/20260828094823.594279-1-kunwu.chan@linux.dev/
Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Tested-by: Kunwu Chan <kunwu.chan@gmail.com>
Reviewed-by: Kunwu Chan <kunwu.chan@gmail.com>
Cc: Jann Horn <jannh@google.com>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Pedro Falcato <pfalcato@suse.de>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-09-05 11:47:47 -07:00
Lorenzo Stoakes (ARM)
e384abeb55 mm/huge_memory: bypass THP tuneables for huge pfnmap mappings
The sysfs THP tuneables at /sys/kernel/mm/transparent_huge_pages/ rather
confusingly only control the behaviour of THP in some instances.

They are not applicable to MADV_COLLAPSE operations, nor to DAX mappings.

Long-term, THP is predicated upon compaction being able to obtain large
folios to populate THP ranges.

However, vm_normal_folio() returns NULL for PFN map mappings, thus their
reference count is maintained by the driver, not core mm.

As a consequence, the folios are not subject to reclaim nor compaction, so
are not truly part of the THP mechanism at all.

However, since commit 5dd40721f1 ("mm: allow THP orders for PFNMAPs")
introduced the ability to establish huge PFN maps, they have been subject
to THP tuneables.

This is incorrect - if a huge PFN map is available (defined by
vma->vm_ops->huge_fault being non-NULL for a VMA_PFNMAP_BIT VMA), then it
should be mapped huge upon fault-in.

Correct this by explicitly checking for this while ensuring that smaps
continues to accurately report THPeligible statistics.

While here, abstract the entire file-backed THP check in
vma_can_map_huge_file(), with sensible separation of logic into helper
functions.

Note that drm_gem_shmem_mmap() and panthor_gem_mmap() establish huge PFN
maps of shmem folios, however they are marked unevictable in
drm_gem_get_pages(), and in any case would fail the reference check in
__remove_mapping() even if they weren't.

Failing to map huge PFN maps has resulted in significant real-world
performance degradation, see links for details.

[ziy@nvidia.com: rename some functions]
  Link: https://lore.kernel.org/DL1HIHWYJ7TB.1CY76SJS0V03L@nvidia.com
Link: https://lore.kernel.org/20260827-hugepfn-allowable-orders-v1-1-94819c8807c8@kernel.org
Fixes: 5dd40721f1 ("mm: allow THP orders for PFNMAPs")
Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Signed-off-by: Zi Yan <ziy@nvidia.com>
Reported-by: Cedric Le Goater <clg@redhat.com>
Closes: https://lore.kernel.org/linux-mm/20260805055544.1568534-1-clg@redhat.com/
Reported-by: Saravanan D <saravanand@crusoe.ai>
Closes: https://lore.kernel.org/linux-mm/20260821070520.25759-1-saravanand@crusoe.ai/
Reviewed-by: Zi Yan <ziy@nvidia.com>
Tested-by: Saravanan D <saravanand@crusoe.ai>
Tested-by: Lance Yang <lance.yang@linux.dev>
Reviewed-by: SJ Park <sj@kernel.org>
Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
Cc: Barry Song <baohua@kernel.org>
Cc: David Hildenbrand <david@kernel.org>
Cc: Dev Jain <dev.jain@arm.com>
Cc: Jason Gunthorpe <jgg@ziepe.ca>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Peter Xu <peterx@redhat.com>
Cc: Ryan Roberts <ryan.roberts@arm.com>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-09-05 11:47:47 -07:00
Linus Torvalds
0d9ff90a54 SCSI fixes on 20260905
2 enhancements to add support and MCQ for additional Intel 4.0
 controller types. The rest are all driver fixes, the largest of which is
 the mpi3mr target use after free fix, follwed by a similar TOCTOU fix
 for io_uring passthrough in bsg.
 
 Signed-off-by: James E.J. Bottomley <James.Bottomley@HansenPartnership.com>
 -----BEGIN PGP SIGNATURE-----
 
 iLgEABMIAGAWIQTnYEDbdso9F2cI+arnQslM7pishQUCapw/mRsUgAAAAAAEAA5t
 YW51MiwyLjUrMS4xMiwyLDImHGphbWVzLmJvdHRvbWxleUBoYW5zZW5wYXJ0bmVy
 c2hpcC5jb20ACgkQ50LJTO6YrIXurAD7BFNaHlTRLIlurYSeMYOV0ZVQQXR9GsW8
 1KQhJW1W6GABAIr7L81Gm1jTUa+CuXixW2N9tv8PBsXx627eWyRZ7tyb
 =9pG9
 -----END PGP SIGNATURE-----

Merge tag 'scsi-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/jejb/scsi

Pull SCSI fixes from James Bottomley:
 "Two enhancements to add support and MCQ for additional Intel 4.0
  controller types.

  The rest are all driver fixes, the largest of which is the mpi3mr
  target use after free fix, follwed by a similar TOCTOU fix for
  io_uring passthrough in bsg"

* tag 'scsi-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/jejb/scsi:
  scsi: megaraid_sas: Limit NVMe request size to the PRP chain frame
  scsi: bsg: Fix TOCTOU in io_uring passthrough command setup
  scsi: bsg: Cap io_uring sense copy to max_response_len
  scsi: mpt3sas: Avoid out-of-bounds cpumask_of_node() call in _base_assign_reply_queues()
  scsi: mpi3mr: Fix use-after-free on tgt_dev->starget during target device refresh/update
  scsi: target: iscsi: Reserve a terminator byte for the login payload
  scsi: target: iscsi: Fix hang for aborted WRITE_PENDING commands
  scsi: ufs: ufs-pci: Add MCQ support for Intel UFS 4.0 controllers
  scsi: ufs: ufs-pci: Add support for Intel UFS 4.0 HS-Gear5
  scsi: sg: Report request-table problems when any status is set
  scsi: mpi3mr: Fix target device refcount leak in mpi3mr_sas_port_add()
  scsi: mpi3mr: Fix NULL pointer dereference in mpi3mr_sas_port_add()
  scsi: ufs: ufs-qcom: Fix sequential read variance
  scsi: ufs: ufs-qcom: Restore HS/LS link startup mode for Qualcomm UFS controller v6.2+
  scsi: ibmvfc: Document protocol parameter of ibmvfc_alloc_target()
  scsi: ibmvfc: Fix kernel-doc name for ibmvfc_scsi_relogin()
  scsi: pm8001: Use rollback index when freeing MSI-X vectors
  scsi: fnic: Initialize the NVMe local port info before registering
2026-09-05 09:25:50 -07:00
Linus Torvalds
d0fc310b4d block-7.3-20260905
-----BEGIN PGP SIGNATURE-----
 
 iQJEBAABCAAuFiEEwPw5LcreJtl1+l5K99NY+ylx4KYFAmqb+cYQHGF4Ym9lQGtl
 cm5lbC5kawAKCRD301j7KXHgpjBkEACTFzHAVopJbtKT6+Rg9esQNPUDfeJkJy/L
 3v8Vi4R7tozAB0IKc58RxV2YFMvga5teWJkAnd33983/MbwCzj9B0oRjSmnHf8K5
 pq4gu1f5pdyfRXGGAnI6ZMom1MNsfuWWiZmD8vQuQ+q4qNVSnQBg0UjrDggDlcW6
 o6EtyjgqAwaaGs+sWgxgy0sYWV7TMiCx4+AZR0TDm8cN3LXyOkOp2abgR35/tnDB
 fs3kUPTkBC4rCZK2uVUhgF6Wctcd2qIF6AEP+bBbWifSCI/jmqAYHk/0IM1xpn8e
 XXPO43X5Iad5iiMWMHlku9G7ZjTo/K2bLc1n9F4IlNZgVPAcF9qtYok5Uc3ghldg
 /qOsclI8D2feQ5j6u060FdnN99+TSHcS3h4roa8jIPNojmh5orw2xRiDCvSVG3au
 +UfEUWnf0JjsmJfX9HCQyV6oTB/7IeiSI+4akHXobCnsG4n9BL8kZ1HR8qwPWsYj
 HPeLHraPljX2slDj+X9EFA8AyxgU33HC9JbjvLlN4L2amX+7Jtwb3A7MlKrGEuUP
 gzybtHk5g17/mcYVj9N2k4tHECR1aZChTQcfNdCctMVB3vXk3KKkFsZstJapCVmT
 fSKx7pdrrNwo3WoKzn8xp/+2gcrRRo4AhSCg9oVfloc5nSFkyfPpq0vPspsUj3ib
 4ze4Sy2TiQ==
 =4ZE+
 -----END PGP SIGNATURE-----

Merge tag 'block-7.3-20260905' of git://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux

Pull block fixes from Jens Axboe:

 - NVMe fixes via Keith:
     - nvme-tcp fixes for an out-of-bounds write on an over-long PDU
     - nvmet-tcp, nvmet-rdma and nvme-rdma leak and cleanup-ordering
       fixes
     - FDP placement id array racy access fix
     - nvme-fc double free of fabrics options on nvme_add_ctrl()
       failure, and a secret leak failure
     - Fault injection opcode filtering
     - stale namespace removal during scan
     - Various other smaller fixes and cleanups

 - Flag zoned disks with GENHD_FL_NO_PART

 - Save the page offset gaps in a cloned bio

 - Fix dma_alignment for large or unreported limits in loop and zloop

 - Clear VM_MAYWRITE on a read-only ublk char device mmap

* tag 'block-7.3-20260905' of git://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux: (25 commits)
  nvme-tcp.h: drop kernel-doc comments, fix a few descriptions
  nvme-fc: fix double free of fabrics options when nvme_add_ctrl() fails
  nvmet: reject namespace enable without device path
  nvmet-auth: Synchronize timeout work during SQ teardown
  MAINTAINERS: update nvme entry
  nvmet-tcp: reject unsolicited H2CData PDUs
  nvme-tcp: defer TLS inline send to io_work
  nvmet-tcp: fix out-of-bounds write when receiving an over-long PDU
  nvme-tcp: return -EPROTO for a C2HData on a write
  nvmet: print namespace IDs as unsigned 32bit value
  nvme: print namespace IDs as unsigned 32bit value
  nvme: remove stale namespaces by NSID range during scan
  nvme: add missing SRCU grace period in error path
  nvme-fabrics: fix DHCHAP secret leak on parse failure
  ublk: clear VM_MAYWRITE on read-only ublk char device mmap
  loop, zloop: fix dma_alignment for large or unreported limits
  block: save page offset gaps in cloned bio
  block: flag zoned disks with GENHD_FL_NO_PART
  nvmet-rdma: fix queue leak when connect backlog is exceeded
  nvme: add opcode filtering for fault injection
  ...
2026-09-05 08:58:55 -07:00
Linus Torvalds
4d7d9486c0 integrity-v7.3-rc2
-----BEGIN PGP SIGNATURE-----
 
 iIoEABYKADIWIQQdXVVFGN5XqKr1Hj7LwZzRsCrn5QUCapsP/BQcem9oYXJAbGlu
 dXguaWJtLmNvbQAKCRDLwZzRsCrn5b2gAQC3ms2HRoZolscMWqnUNoi5SmPpwcV2
 v/ojwDc1TnS9HAEA/604QYihEvRQzKQwEyF6W6b83w22tyWKhDW1a0d7KQ4=
 =/Hd6
 -----END PGP SIGNATURE-----

Merge tag 'integrity-v7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/zohar/linux-integrity

Pull IMA fixes from Mimi Zohar:

 - Instantiating the ima_file_truncate and ima_path_truncate LSM hooks
   resulted in configfs locking issues.

   configfs files should not be measured, appraised, or audited in the
   first place, so the builtin policies are updated to exclude them.

 - IMA audit messages include the filename, which could result in a page
   fault when the filename doesn't exist

 - Un-hide the IMA_MEASURE_PCR_IDX Kconfig prompt

* tag 'integrity-v7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/zohar/linux-integrity:
  ima: allow users to specify the pcr index with IMA_MEASURE_PCR_IDX
  ima: Check for ERR_PTR from dentry_path() in validate_hash_algo()
  ima: don't measure/appraise files on configfs
  configfs: move CONFIGFS_MAGIC definition to magic.h
2026-09-04 19:36:11 -07:00
Linus Torvalds
654ae5d73c drm fixes for 7.3-rc2
core:
 - Fix drm_crtc_commit leak when PAGE_FLIP_EVENT is used,
 
 dma-buf:
 - Publish the dma-buf only after copy_to_user succeeds
 - fix some kernel-doc warnings
 
 atomic-state-helpers:
 - set pixel_blend_mode to prop default on reset
 
 sysfb:
 - Fix integer overflow
 - fix constant comparison bug
 
 pagemap:
 - Prevent double migration of device pages
 - Reset migration page count on eviction retry
 - dma-unmap pages before handling migration errors
 - use after free fixes
 
 prime:
 - fix prime exports tracing
 
 amdgpu:
  - Fix for drm_amdgpu_info_device with mixed 64 bit kernel and 32 bit userspace
 - plane blend mode fixes
 - SR-IOV fix
 - GFX8 fix
 - MES queue reset fix
 - GPUVM fixes
 - DCN 6 warning fix
 - DCN 3.5/3.6 fix
 - DML fix
 - Backlight fix
 - Colorop fix
 - DC get_estimated_bw() fix
 - devcoredump fix
 - Userq fixes
 - APU PSP fix
 - Cursor fix
 
 amdkfd:
 - MES queue eviction fix
 - MQD debugfs fix
 
 xe:
 - oa uapi error handling fix
 - drm info message to report FLAT_CSS base misalignment.
 
 i915:
 - Drop an accidentally duplicated panel fitter call in DP MST
 - Fix DDI clock programming for Cx0 and LT PHY
 - Fix PTL CDCLK handling at probe, causing a glitch
 - Fix dg2_power_well_count() return type
 - Fix a NULL pointer deref at forced probe
 - Fix selective fetch disable
 
 amdxdna:
 - out-of-bounds access fix
 - reject commands chains with no commands
 - handle chained mapping BO failures
 - refuse to flush an imported BO
 
 ethosu:
 - handle mmio mapping failures
 - handle storage modes only on hardware that supports it
 - fix job completion fence cleanup
 
 fastrpc:
 - Publish the dma-buf only after copy_to_user succeeds
 
 gud:
 - Improve TV modes and rotation handling
 
 nouveau:
 - use-after-free fixes
 - add missing scanline position support
 - HDMI and DP fixes
 - null pointer dereference fix
 - dmem accounting fixes for large folios
 - use write-combined maps for coherent
 
 qaic:
 - out-of-bounds access fix
 
 tegra:
 - Add blend mode properties
 
 virtio:
 - exit path and error handling fixes
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEEKbZHaGwW9KfbeusDHTzWXnEhr4FAmqbH88ACgkQDHTzWXnE
 hr7stA//SAJOADL8CuoBzSyAX7zoqhVErYk798+4r5tPFJ5CzjkZVxGRGur7XWbm
 atezEEKaMTEz2BDVb1JDNRI1X1Yq9GtfWM860hUXmYHCegeO49B9lDnS1v0HZweD
 PQvPhCqvJpOOF6D8sjYFjVNvi0OrY0JVRlzsnMNmTIlw0xbg01lqn/oWgRA7qdvq
 Zo1k1yEPlK3jV2YuX674n7NoioFqeiSWBo9PzIX+yghagg2LrrS1Plx4RFA/B6bm
 Czt/x7WPE7lvoZqyDGBwlAY/dta3bagCFkGoDwb2Q1B3MjYXKcKBS0aGL9PyOgu/
 /9dNqvR4aDu9CXvNwb3kNbqjJL7DdFBCzwm78PNc43TizkR4WXCBXCNxJOf93e6h
 Bwx0GamXQJeGI6xNvQpEssUxezuS3wdoNZ0Rbk3nxMXlvf7OB/sgwkNYVCNKkk/V
 dSMOr1XB9pBGmtuWFPOf1kq/is4P4Ns/m8Rutfp5SBJNU9Air5ECblRNNkqFuOtS
 582QAM+7xp6zIbepWALu8TTNQMsNKlDwiNc3JOH3Ks3wZ0wExXMTIaSJ7NHOxYj/
 B9gWEN+g1LreFHDaCzR1xRetO1bIHNGMNhuqChXj8K8vuDoVPP8FAvKzKEy/iyIl
 CRQXLPSV8LjqKL5d5YW3QFvwUSwJwGr3fXCWd1Voo26dL4DOEx0=
 =vHdE
 -----END PGP SIGNATURE-----

Merge tag 'drm-fixes-2026-09-05' of https://gitlab.freedesktop.org/drm/kernel

Pull drm fixes from Dave Airlie:
 "Lots of scattered fixes: nouveau has a bunch of display fixes for
  blackwell GPUs that should mean we light up monitors properly and fix
  some desktop rendering problems, amdgpu and intel display changes as
  usual.

  There also changes to the core pagemap, then the usual amouny of AI
  inspired validation fixes.

  core:
   - Fix drm_crtc_commit leak when PAGE_FLIP_EVENT is used

  dma-buf:
   - Publish the dma-buf only after copy_to_user succeeds
   - fix some kernel-doc warnings

  atomic-state-helpers:
   - set pixel_blend_mode to prop default on reset

  sysfb:
   - Fix integer overflow
   - fix constant comparison bug

  pagemap:
   - Prevent double migration of device pages
   - Reset migration page count on eviction retry
   - dma-unmap pages before handling migration errors
   - use after free fixes

  prime:
   - fix prime exports tracing

  amdgpu:
   - Fix for drm_amdgpu_info_device with mixed 64 bit kernel and 32 bit
     userspace
   - plane blend mode fixes
   - SR-IOV fix
   - GFX8 fix
   - MES queue reset fix
   - GPUVM fixes
   - DCN 6 warning fix
   - DCN 3.5/3.6 fix
   - DML fix
   - Backlight fix
   - Colorop fix
   - DC get_estimated_bw() fix
   - devcoredump fix
   - Userq fixes
   - APU PSP fix
   - Cursor fix

  amdkfd:
   - MES queue eviction fix
   - MQD debugfs fix

  xe:
   - oa uapi error handling fix
   - drm info message to report FLAT_CSS base misalignment

  i915:
   - Drop an accidentally duplicated panel fitter call in DP MST
   - Fix DDI clock programming for Cx0 and LT PHY
   - Fix PTL CDCLK handling at probe, causing a glitch
   - Fix dg2_power_well_count() return type
   - Fix a NULL pointer deref at forced probe
   - Fix selective fetch disable

  amdxdna:
   - out-of-bounds access fix
   - reject commands chains with no commands
   - handle chained mapping BO failures
   - refuse to flush an imported BO

  ethosu:
   - handle mmio mapping failures
   - handle storage modes only on hardware that supports it
   - fix job completion fence cleanup

  fastrpc:
   - Publish the dma-buf only after copy_to_user succeeds

  gud:
   - Improve TV modes and rotation handling

  nouveau:
   - use-after-free fixes
   - add missing scanline position support
   - HDMI and DP fixes
   - null pointer dereference fix
   - dmem accounting fixes for large folios
   - use write-combined maps for coherent

  qaic:
   - out-of-bounds access fix

  tegra:
   - Add blend mode properties

  virtio:
   - exit path and error handling fixes

* tag 'drm-fixes-2026-09-05' of https://gitlab.freedesktop.org/drm/kernel: (83 commits)
  drm/xe/vram: report FLAT_CCS base misalignment
  MAINTAINERS, mailmap: use Aditya Garg's linux.dev account
  drm/amd/display: use plane color_mgmt_changed to track colorop changes
  drm/amdgpu/userq: fix struct drm_amdgpu_info_device padding for 32bit compile
  drm/amd/display: Fix cursor disable with horizontally split planes
  drm/amdgpu/userq: dont overwrite the error of subsequent map call
  drm/amdgpu: Skip accessing psp rum time db for APUs
  drm/amdgpu: update the fw version for gfx12 userqueues
  drm/amdgpu: update the fw version for gfx11 userqueues
  drm/amdgpu: fix byte/dword unit mismatch in coredump IB dump
  drm/amdkfd: fix scope of mqd_mgr dereference in pqm_debugfs_mqds
  drm/amd/display: fix division by zero in get_estimated_bw()
  drm/amd/display: use halving distribution for all encode-to-linear curves
  drm/amd/display: Fix backlight control for luminance-capable OLED
  drm/amd/display: Remove const Qualifier From Non-Pointer Fields
  drm/amd/display: Set gpuvm min page size to 4K on dcn35/36
  drm/amd/display: Fix DCN5/6 DML2 compilation warnings
  drm/amdgpu: fix Idle BOs list in VM debugfs status info
  drm/amdgpu: use AMDGPU_GPU_PAGE_SHIFT instead of PAGE_SHIFT
  drm/amdgpu: Update queue reset support version
  ...
2026-09-04 13:42:16 -07:00
Linus Torvalds
3f17a52d47 arm64 fixes for -rc2
- Disable interrupts during page-table walk in show_pte()
 
 - Fix kexec_file_load() with 52-bit capable kernels on machines without
   52-bit addressing
 
 - Fix MIDR matching in CPU errata handling for KVM guests
 
 - Avoid reading MTE-specific ID registers when MTE support is disabled
 -----BEGIN PGP SIGNATURE-----
 
 iQFEBAABCgAuFiEEPxTL6PPUbjXGY88ct6xw3ITBYzQFAmqayVUQHHdpbGxAa2Vy
 bmVsLm9yZwAKCRC3rHDchMFjNIG6B/47THEr7Wqq00c1s7loGtwGJiN8dMcfSMpp
 r86zeLL37erJQ9K/OeaUFR0bQEgfh7gXSqXXJ8N1wqAObrAfzek7X0lxxXMkVz4p
 tpNPFgEgP9jwvtYbKH2W9apmP8xxT7MJHF/FnLQVkdVdJBBU+nmrpYcEz37e7O6a
 PmSdl4grWL6AG/CifSCnGvyVFsWVzLeaDgJSAXQWgalefZJzar8dki6W7HfZULo0
 I9uizj52+lG2tVpIz6MUcy9k1cwOyTDl061qrD4/oD6aC67U6aS3tRPmwX3qKF7a
 rvuwEBwRckmKoK1bxzd3wjzkEH/Urhom1Op8SwJ72s1ex6m5EUJ4
 =QWAE
 -----END PGP SIGNATURE-----

Merge tag 'arm64-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/arm64/linux

Pull arm64 fixes from Will Deacon:
 "Nothing Earth-shattering, but worthwhile fixes nonetheless:

   - Disable interrupts during page-table walk in show_pte()

   - Fix kexec_file_load() with 52-bit capable kernels on machines
     without 52-bit addressing

   - Fix MIDR matching in CPU errata handling for KVM guests

   - Avoid reading MTE-specific ID registers when MTE support is
     disabled"

* tag 'arm64-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/arm64/linux:
  arm64: Don't read GMID_EL1 when MTE is disabled
  arm64: errata: pass REVIDR when matching target implementation CPUs
  arm64: trans_pgd: clone only the linear map that exists at runtime
  arm64: mm: Fix the lockless page-table walk in show_pte()
2026-09-04 13:32:46 -07:00
Linus Torvalds
408802f1e6 A small fixup for the new nearfull_sync mount option, a potential
use-after-free fix (marked for stable) and a patch that eliminates
 the last use of PageWriteback macro in the tree.
 -----BEGIN PGP SIGNATURE-----
 
 iQFHBAABCgAxFiEEydHwtzie9C7TfviiSn/eOAIR84sFAmqa/nETHGlkcnlvbW92
 QGdtYWlsLmNvbQAKCRBKf944AhHzi7qLB/4gfbxdP5lystLHoDbwSo+ceM15Apl+
 THXVTBGspHNe5w07/fk1NfSk8KccqN66cCh9W23JMZt2PQr+n5/0Azp+ZkcL+koO
 CNNhvachvs+E3J5cRNHvvP3PCQurOO0tCO4vGGHRt6j3VTrWdwKkgVHGHu2hZ49q
 GLkm82eUKYkZdV80FV31q1ZdXHQBCAuxkBgRQNbqlc9yj3OA6UoxLsAaB3uxvmKq
 sifdvIyWFt/+SdntzoM6Dt4vo6P0/RiQJdIXKLj+fxHFiJEeX0IJaTpZvA7FOHOt
 DyRMhb3lTzQBDjZydNQO15XcjfifskuqeQxxSQmISv7lASiUkq5Tn+uj
 =p61q
 -----END PGP SIGNATURE-----

Merge tag 'ceph-for-7.3-rc2' of https://github.com/ceph/ceph-client

Pull ceph fixes from Ilya Dryomov:
 "A small fixup for the new nearfull_sync mount option, a potential
  use-after-free fix (marked for stable) and a patch that eliminates
  the last use of PageWriteback macro in the tree"

* tag 'ceph-for-7.3-rc2' of https://github.com/ceph/ceph-client:
  ceph: apply nearfull_sync option on remount
  libceph: remove pinning assertion in ceph_msg_data_iter_next()
  ceph: lock mutex in ceph_mds_check_access()
2026-09-04 13:27:58 -07:00
Julian Braha
6903878d46 ima: allow users to specify the pcr index with IMA_MEASURE_PCR_IDX
The IMA_MEASURE_PCR_IDX option is currently not visible in the kconfig
frontend, so it always uses its default, 10. This means that the
'range 8 14' is dead code, and users are unable to specify the pcr index
value.

In a previous discussion, Mimi explained that users should be able to use
this config option to specify the pcr index. [1]

Let's add a prompt for users to specify the pcr index, when EXPERT is
enabled.

This dead range was found by kconfirm, a static analysis tool for Kconfig.

Link: https://lore.kernel.org/all/1feff118-4afa-4b9c-86f1-271a7a88208f@gmail.com/T/#mc4efa2491b4937eb7c9e532c29ffba516a70e662 [1]
Signed-off-by: Julian Braha <julianbraha@gmail.com>
Signed-off-by: Mimi Zohar <zohar@linux.ibm.com>
2026-09-04 14:01:00 -04:00
Linus Torvalds
986c24e0fe hid-for-linus-2026090401
-----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEoEVH9lhNrxiMPSyI7MXwXhnZSjYFAmqa4YwACgkQ7MXwXhnZ
 SjaNsQ/9Ff0KeKgaQUZHLE47SpOlXKWQaDJmrodDkngQh+9KaZj3NgmZD5BR2Z2p
 5v/6dhs4gFoFzQtXjuR0GvDTWzu0bx1IV3IsXhqTdhQ2fLAeu8IxUL0DEOtg3lUO
 Vlf67UagvmC+K01UWkbloS3f8dEt8tg3CXyg0Uy7f23QgBFa/TbtQpXTNGlMHOqv
 G6qE1PPBqGmPI74E/5uusI8L3tw4t8A4ylHi3UcQhTxGaUGK+Ew8GCeDsIUrwrzM
 A8Um5GHdBCWZAqluT8HPnBI2wgnUR+pvda4UdqMSYkBJW2Rz1FFaOhgLkHt+azRX
 F7RhjuxcBlaZsXIaCmIZEW6rEr0QIeUPeFK6ML3uswLtFdh/yWASUMo84Ev08Z9N
 iB7qm0+S9AZSDknINAtRRcOXsOgjvug00xMf6zcUvcP66mP1Rj/PnOGb5Lqm4icp
 SiXBF+CpN0qn3h8TWG5+GvEX0AnGcmkpL0Vx7noVHJeK8Z+Yroozv+vGi1/29pxo
 ML4QEUIYV3Uj0rU1Azgd/rKiaxnizpczeJ5ViW4+ozpT4nPHjTxcz5kVsqJ6gGV3
 XTsV9YW13xgjuB0objDDeGjYRku7MtTWUfdQiKCE71a+L9nYWr5/q2mp+8QTKWqy
 /XnK5I8o2dGWhMWCPmwoqCNO/f4FGT+J2Ok6yisG8jJ/9j6Dovk=
 =a+wk
 -----END PGP SIGNATURE-----

Merge tag 'hid-for-linus-2026090401' of git://git.kernel.org/pub/scm/linux/kernel/git/hid/hid

Pull HID fixes from Benjamin Tissoires:

 - hid-hyperv build fixes on certain configs (Jiri Kosina)

 - HID-BPF fix and selftests now that the bpf verifier is more
   restrictive (Benjamin Tissoires)

 - Some AI detected fixes for OOB, errors and validation (Ibrahim
   Hashimov, Shen Yongchao, Wei Jie Law)

 - various device fixes (Dave Carey and Vadim Klishko)

* tag 'hid-for-linus-2026090401' of git://git.kernel.org/pub/scm/linux/kernel/git/hid/hid:
  HID: bpf: serialize device reference release in struct_ops destroy path
  HID: rmi: fix OOB access with undersized RMI reports
  selftests/hid: prepare test_rdesc_fixup_get_data_overflow for the new verifier
  selftests/hid: Add a test to ensure we can write fields in hid_device
  HID: bpf: mark struct hid_device as safe BPF pointer
  HID: wacom: validate report length in wacom_intuos_pro2_bt_irq
  HID: multitouch: Fix stale MT slots when contact count drops to zero
  HID: i2c-hid: Add a quirk for a Cirque I2C device.
  HID: hyperv: make pointer arithmetics understandable for FORTIFY_SOURCE
  HID: hyperv: fix build breakage with certain configs
2026-09-04 09:25:38 -07:00
Linus Torvalds
36ec09e263 sound fixes for 7.3-rc2
A collection of small fixes since 7.3-rc1.  Quite a few fixes are
 for ALSA core for issues that have been detected by the things you
 know well.  Additionally a series of hardening for runtime PM, and
 usual quirk updates, and some other misc driver fixes are included.
 
 * Core:
 - Fixes for PCM races
 - UMP parser NULL dereference fix
 - Fix error handling in rawmidi ioctl
 
 * USB- and HD-audio:
 - Implement missing runtime PM guards across multiple interfaces
 - Fix for OOB access in US-122L MIDI driver
 - Double-free fix for CAIAQ driver
 - Quirks for HD-audio Realtek & Cirrus codecs, Conexant S3-resume,
   USB Audient devices
 
 * Others:
 - Fix of logical mistakes in dummy driver mixer and selftest code
 - Lock init fix in the legacy harmony driver
 -----BEGIN PGP SIGNATURE-----
 
 iQJCBAABCAAsFiEEIXTw5fNLNI7mMiVaLtJE4w1nLE8FAmqaiQ0OHHRpd2FpQHN1
 c2UuZGUACgkQLtJE4w1nLE+d2Q/9EnlQ0Sr+MYS81pYzxjWNSKzvwtRw5h3B8Hjo
 bBTflzPH+iD0AI5Y0HJ31wMrkhJPSDkznQ76foZiSgTTJh85LhoG9HlZlpIxdPW4
 fNS9/N28JKRZM+qTd5P7UvGbKv9hBpMAYQPkgmAiCZ5+47oQLhgBU5THpn0Mwhxo
 JpmULLjKuSGKCf+b/SY3MY+UF7CotiQL5L5uTF83JSm8T7DdjFwzUgk6YCTzc7/b
 WbuWa5TjIg6smSzCQaip8WoE/KimLJ++zKwk8tFH2mpWcNthmTdzQGosZbp7k3l5
 G01va200DdRE7ROXdkyao7jj8FSkex23NZyQmTDYMvQz9YmGMDhsDps+Aiv11hfY
 vlwGHHHuUO/gZcjrMB+JO5MXVxsFvyNOI7L0bFMAcX8BQwcVmXLujOI+B+iY/3Ut
 TMncDaaeKDbP/dGNROVFds2nO0pUsR3Fip16Xczy4lds6QQ8XaHefUIN4jzbltHc
 hkrCKGdyXuzpBKIlSag7xE7wSAWSLGgNnRkN7614cZSK8DRJpkrtef60Uc7W8hEU
 17c9cmhOlUM6C5BvIEmdyzN6Y0tKXdJaG/gz1Wlx2iMnsNyZ9zO7ZJq5c5HU+Txl
 3559GWY4sMDtm/A/aG9daMo1Rym3dqRwqdW/DvfhLXythLCA2YRTQIG5bHunGFYU
 jok+fBc=
 =uXXN
 -----END PGP SIGNATURE-----

Merge tag 'sound-7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/tiwai/sound

Pull sound fixes from Takashi Iwai:
 "A collection of small fixes since 7.3-rc1.

  Quite a few fixes are for ALSA core for issues that have been detected
  by the things you know well. Additionally a series of hardening for
  runtime PM, and usual quirk updates, and some other misc driver fixes
  are included.

  Core:
   - Fixes for PCM races
   - UMP parser NULL dereference fix
   - Fix error handling in rawmidi ioctl

  USB- and HD-audio:
   - Implement missing runtime PM guards across multiple interfaces
   - Fix for OOB access in US-122L MIDI driver
   - Double-free fix for CAIAQ driver
   - Quirks for HD-audio Realtek & Cirrus codecs, Conexant S3-resume,
     USB Audient devices

  Others:
   - Fix of logical mistakes in dummy driver mixer and selftest code
   - Lock init fix in the legacy harmony driver"

* tag 'sound-7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/tiwai/sound: (23 commits)
  ALSA: caiaq: Fix potential double-free at error path
  selftests/alsa: Fix the step check for INTEGER controls
  ALSA: hda/realtek: Fix cold-boot headset misdetection on Acer Aspire A515-57G
  ALSA: rawmidi: Return the error from snd_rawmidi_input_params()
  ALSA: ump: do not touch legacy_rmidi before it exists
  ALSA: hda/cs420x: Add CS4208 fixup for MacBookAir 7,2
  ALSA: dummy: Report a change when one capture switch channel moves
  ALSA: usb-audio: Add mixer map quirk for Audient iD24
  ALSA: hda: restore MFG widget enumeration after core split
  ALSA: usb-audio: fix OOB write in snd_usbmidi_us122l_output()
  ALSA: pcm: Serialize PCM mmap with buffer reallocation to fix page UAF
  ALSA: harmony: initialize locks before requesting IRQ
  ALSA: hda/realtek: Add quirk for VAIO VJS131
  ALSA: pcm: Fix race between non-atomic ops and trigger-start
  ALSA: hda/realtek: Add quirk for Acer Predator PHN16-72
  ALSA: hda/realtek: Add quirk for Lenovo Yoga Slim 9 14ILL10
  ALSA: hda/conexant:Fix abnormal Mic/Speaker functionality on SN6140 after S3 wake-up
  ALSA: usb-audio: Guard FCP protocol transfers
  ALSA: usb-audio: Add PM guards to RME Digiface controls
  ALSA: usb-audio: Guard Scarlett2 protocol transfers
  ...
2026-09-04 09:17:05 -07:00
Linus Torvalds
3e66602704 ata fixes for 7.3-rc2
- Work around lost interrupts on Marvell 88SE61xx.
    The Marvell AHCI controller requires you to clear interrupts in the
    opposite order from what is specified in the AHCI specification in
    order to not lose interrupts (Hajo)
 
  - Do not raise UNIT ATTENTION for depopulation commands.
    The libata completion function unconditionally sets sense data with
    sense key UNIT ATTENTION (UA) for depopulation commands. The SCSI
    layer will fail a command when seeing this sense data. UA is only
    supposed to be raised if the capacity actually changed. Since these
    commands are currently only supported as passthrough commands, the
    user is expected to revalidate the device, which will detect a
    capacity change anyway. Thus drop the unconditional UA until a
    better solution has been implemented (Damien)
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQRN+ES/c4tHlMch3DzJZDGjmcZNcgUCaprYjgAKCRDJZDGjmcZN
 cj0LAQCuA58xAXmZBFIsfHysezSKBNn3ZjwuiKKwX5DYzK1fiQEAxZWAnV0fE0wQ
 8bebZXF12uoG2+PD22ZIcUxneKeyXQg=
 =qHGE
 -----END PGP SIGNATURE-----

Merge tag 'ata-7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux

Pull ata fixes from Niklas Cassel:

 - Work around lost interrupts on Marvell 88SE61xx

   The Marvell AHCI controller requires you to clear interrupts in the
   opposite order from what is specified in the AHCI specification in
   order to not lose interrupts (Hajo)

 - Do not raise UNIT ATTENTION for depopulation commands

   The libata completion function unconditionally sets sense data with
   sense key UNIT ATTENTION (UA) for depopulation commands. The SCSI
   layer will fail a command when seeing this sense data. UA is only
   supposed to be raised if the capacity actually changed.

   Since these commands are currently only supported as passthrough
   commands, the user is expected to revalidate the device, which will
   detect a capacity change anyway. Thus drop the unconditional UA until
   a better solution has been implemented (Damien)

* tag 'ata-7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux:
  ata: libata-scsi: do not raise UA for storage element depopulation and restoration
  ata: ahci: work around lost interrupts on Marvell 88SE61xx
2026-09-04 09:00:03 -07:00
Linus Torvalds
58f93a4b73 - Fix a tree connection use-after-free in smb2_tree_connect() by
balancing references across concurrent connect, disconnect, and
    session logoff paths.
 
  - Validate source and target ranges in COPYCHUNK requests before range
    locking and copy operations.
 
  - Fix an oplock break notification UAF by acquiring a connection
    reference under ksmbd_inode lock and releasing it after the
    notification work completes.
 
  - Fix the sparc build by using an unsigned int for the atomic work
    state, ensuring xchg() uses a supported four-byte operation.
 -----BEGIN PGP SIGNATURE-----
 
 iQJKBAABCgA0FiEE6NzKS6Uv/XAAGHgyZwv7A1FEIQgFAmqan48WHGxpbmtpbmpl
 b25Aa2VybmVsLm9yZwAKCRBnC/sDUUQhCP41D/4gNSDDjbGg/Du5nNlNfd7x0/Ql
 ARgdajcnTT/2rUrBkXb3hfNi7BvSHz8jShb7acnwcs9VbxF7cWMk0r+tkjpsONI9
 hwUAXiqOQNkiUJZez+29WgiVuIqNjWSB9WKDGcA7J364Vnwm4M5a8y9wZHfV9ReQ
 aXmVlNlWP6zosrXv4Ex2Eb1bUaYnB822ZrsMQBdiZireUlVUyi/MeWOIrjxt7xPJ
 /nMTJcyDNIFjJJIQMZ/LjzIzvD82QO4LP3F8rlvHD2UIMQik6m0UXF3wPUZGPMbp
 X3o/sPoZHF1KrVpOG4SR5Lvy6KHLtGDSP7bVVhw4ahdtUnyeAeHx/xo9+OKbsiBO
 E4Sji35E8ZyIZ/xHMtOfSfA74W9ia0A0olWyG/mviptJ0RD8unddJP+D/L+EAD0I
 2pQyt9YfwsXm/7FoQXzfbtyi8Z2gl5Jp+xNr6DOyzkIulsBxFOBjgqrUEaTV2/Dc
 abqnVPLo4X9zqiE0HdSNS/go2STx9iox02blZ2wBmYNxEy4X22IheXbyN+wdtn8Q
 Ts8PrE1W0DO8PjWL9TG8okRYY3tRd0AcAVffi6P+QqkRsQNRCduv0mnAsSY1EuU0
 r/M2i5RpQ+JIBLQc4C53zOaAbBCyY1MSXUCgAs1BVyyUdqtjdPJbzyzNn6FDYVPA
 HrPxSl+fmtBsH0C25A==
 =2ZL4
 -----END PGP SIGNATURE-----

Merge tag 'ksmbd-for-7.3-rc2-part2' of git://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/smb

Pull smb server fixes from Namjae Jeon:

 - Fix a tree connection use-after-free in smb2_tree_connect() by
   balancing references across concurrent connect, disconnect, and
   session logoff paths.

 - Validate source and target ranges in COPYCHUNK requests before range
   locking and copy operations.

 - Fix an oplock break notification UAF by acquiring a connection
   reference under ksmbd_inode lock and releasing it after the
   notification work completes.

 - Fix the sparc build by using an unsigned int for the atomic work
   state, ensuring xchg() uses a supported four-byte operation.

* tag 'ksmbd-for-7.3-rc2-part2' of git://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/smb:
  ksmbd: fix tree connection use-after-free in smb2_tree_connect()
  ksmbd: validate COPYCHUNK source and target ranges
  ksmbd: fix use-after-free in oplock break notification
  ksmbd: fix sparc build with atomic work state
2026-09-04 08:42:14 -07:00
Linus Torvalds
421066905c Probes fixes for v7.3-rc1:
- kprobes: Protect kprobe_blacklist with RCU
   . RCU-protect kprobe_blacklist and use kfree_rcu() to prevent UAF
     races during module unloading and enable safe atomic lookups.
 - tracing/probes: Fix multi-probe field use-after-free and BTF parsing
   . Multi-probe UAF fix: Duplicate field and type strings on
     trace_probe_event to prevent UAF when freeing primary probe.
   . BTF member lookup fixes:
     - Check the containing inner struct/union kflag when resolving
       anonymous members to ensure correct bitfield offset calculation.
     - Prevent unnamed bitfields from being pushed to anon_stack in
       btf_find_struct_member(), avoiding false lookup errors.
     - Fix code block indentation in get_bitoffset_of_field().
 - uprobes: Error pointer safety
   . Guard free_trace_uprobe() with IS_ERR_OR_NULL() to avoid crashing
     during automatic cleanup when an error pointer is returned.
 -----BEGIN PGP SIGNATURE-----
 
 iQFPBAABCgA5FiEEh7BulGwFlgAOi5DV2/sHvwUrPxsFAmqahXQbHG1hc2FtaS5o
 aXJhbWF0c3VAZ21haWwuY29tAAoJENv7B78FKz8bOJIH/1RuAq2y8fvfqWKwDBNG
 9CrSIMmZ0915s4LVSGQrrjNYfpj2rFYkEMsJcFo2pavKwWNyaxXFjXu8Vy9JGckx
 VFAHA52x2QaEYdwBeoo/Jd3+7Ks/3zH1XwfSILFa0PMn86/JCKHx/+5ah6Sk4vcu
 /he61Auyp6lJtvv88n95j1evCJNouU6lJ3fnvm8mNYTeLOIvPZ3qku6SsiOqNdeQ
 Ln8bcNP2Iis33PqfeydiRWv9nPog/ifH4a9WJ+fdqKA+06AHKVsfHB+fP0hxyUDp
 TS4oxKrIk6HIbEQRjgo8YcOPWurHAm0GQ2ZslgLFuWDE9rDQCtPdOYv/YzJe0ByG
 nHc=
 =hX9F
 -----END PGP SIGNATURE-----

Merge tag 'probes-fixes-v7.3-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace

Pull probes fixes from Masami Hiramatsu:

 - Protect kprobe_blacklist with RCU

   RCU-protect kprobe_blacklist and use kfree_rcu() to prevent UAF races
   during module unloading and enable safe atomic lookups.

 - Fix multi-probe field use-after-free

   Duplicate field and type strings on trace_probe_event to prevent UAF
   when freeing primary probe

 - Fix probe BTF member lookup:

   Check the containing inner struct/union kflag when resolving
   anonymous members to ensure correct bitfield offset calculation

   Prevent unnamed bitfields from being pushed to anon_stack in
   btf_find_struct_member(), avoiding false lookup errors

   Fix code block indentation in get_bitoffset_of_field()

 - uprobes error pointer safety

   Guard free_trace_uprobe() with IS_ERR_OR_NULL() to avoid crashing
   during automatic cleanup when an error pointer is returned

* tag 'probes-fixes-v7.3-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace:
  kprobes: Protect kprobe_blacklist with RCU
  tracing/probes: Fix use-after-free on field name/type of events with multiple probes
  tracing/probes: Fix code indent in get_bitoffset_of_field()
  tracing/probes: Fix BTF kflag check for anonymous struct member access
  tracing/probes: Fix anon_stack check for unnamed bitfields in btf_find_struct_member
  uprobes: guard trace cleanup against error pointers
2026-09-04 08:24:09 -07:00
Linus Torvalds
65119e86fe pmdomain providers:
- mediatek: Fix Kconfig for Airoha power domains
  - qcom: Revert adding the missing power domains for Eliza
 
 cpuidle:
  - psci: Fix support for probe deferral by dropping the faux device
  - dt_idle_genpd: Free the original name allocation
 -----BEGIN PGP SIGNATURE-----
 
 iQJEBAABCgAuFiEEugLDXPmKSktSkQsV/iaEJXNYjCkFAmqafHMQHHVsZmhAa2Vy
 bmVsLm9yZwAKCRD+JoQlc1iMKRjID/948vUtJNrC8lwohdIxMbycSq6tyDTxGVip
 HiiF8MZu6L6agpW0sNd+9LXdmcTkzIujfTsQedmcl9Xq4ZmWTND4fyOSn29QnomD
 wyCI3W19C4SKUU9sFLJGDwXam97CJD6IG0N+0S6djKEIjuAAm2F6H/BfBjVJKqIA
 Wtp2vOdy5jMJd0+U4+odPG2EqW314G/sNA0EHDgbLEizE9riyumYh/OjjSRDOfD1
 vWobP8riqxxaDj+ij8tr0r4lWpDm/cKHCRSdONdLCh8x2ozIuHtq7T+KfCeWH1q5
 mxcgVJy6hSLMP/kFoNYTL7rxq110IaNIIAqPLLz/ywpQ13TbZHKvJd7jTM0ty8xo
 GhbiiDEDGB0s7Cdo2sZV8cZ2o+DniiDa+tLg1lNUfAH0vZp+uaiqX2tVohHLvZuW
 ELyg0cljC7bNZVvp0TCB4Or0wnE2p5Dlfr2wu/ygHOerA2m7vJOIr9UxNGDtbhHJ
 OBz8rRGKJpVAN1XlOwJNcumDhtm0ipEJ7hYpOiq2LTv7vefSg1/GxuZy50rPfDKK
 4LwAiF8AomlAhcIwkkxv6odflCSzemTNGouZ3XXSh3ETExQOg+bJhqw37cBbP8bD
 BF0uPOfEM8SqNcjejuC1M1XQW3U1sHmIzRc4B8HNUJUzjBBgrxeEV2KmhrHfQtFU
 +2kq/wPO9Q==
 =QG7D
 -----END PGP SIGNATURE-----

Merge tag 'pmdomain-v7.3-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/ulfh/linux-pm

Pull pmdomain and cpuidle fixes from Ulf Hansson:
 "pmdomain providers:
   - mediatek: Fix Kconfig for Airoha power domains
   - qcom: Revert adding the missing power domains for Eliza

  cpuidle:
   - psci: Fix support for probe deferral by dropping the faux device
   - dt_idle_genpd: Free the original name allocation"

* tag 'pmdomain-v7.3-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/ulfh/linux-pm:
  cpuidle: dt_idle_genpd: kfree() the original name allocation
  pmdomain: airoha: fix unselectable AIROHA_CPU_PM_DOMAIN kconfig
  cpuidle: psci: Fix support for probe deferral by dropping the faux device
  Revert "pmdomain: qcom: rpmhpd: Add missing MXC and MMCX power domains for Eliza"
2026-09-04 08:17:41 -07:00
Dave Airlie
c96294afbc A small fix on the error handling of an OA uapi and the
addition of a drm_info message to report FLAT_CSS base misalignment.
 -----BEGIN PGP SIGNATURE-----
 
 iQEzBAABCgAdFiEEbSBwaO7dZQkcLOKj+mJfZA7rE8oFAmqZ1RgACgkQ+mJfZA7r
 E8qyAwf9FrxuliHzaaSz0vxrIlL4LzCYarKbLc9quiSAXEu0TM5QhIYlJJu8PSrp
 ChZGAutqwG+0x8o/+ztbt+5ij21+FVWOK/PnGEmctBevd+bPPRKAWYghVgSOkFww
 OuUwUotEDicqIM+Ml8qjDUXgWNRgkeoLknKZH2XTWpRZxPAxZYcC5P6k+DVAN4pC
 zcurjW9gbcTci2OP9No8EtxuY8+3YCz/Jtwd/Sx1nw0gqoD5l2yPziwdKMOjd3XU
 imml0dnBEWQYPUliVojj2onZKM9ujR4JHSpFCKAHs5c8UN5srlIEFzh14QgR+9Ux
 4ZDXzmbzc3Y3z4oM+dsJOiY3LDUJTA==
 =XZIm
 -----END PGP SIGNATURE-----

Merge tag 'drm-xe-fixes-2026-09-03' of https://gitlab.freedesktop.org/drm/xe/kernel into drm-fixes

A small fix on the error handling of an OA uapi and the
addition of a drm_info message to report FLAT_CSS base misalignment.

Signed-off-by: Dave Airlie <airlied@redhat.com>

From: Rodrigo Vivi <rodrigo.vivi@intel.com>
Link: https://patch.msgid.link/apnVOtDv4WAIoj_X@intel.com
2026-09-04 20:36:14 +10:00
Dave Airlie
7f78fe856e amd-drm-fixes-7.3-2026-09-03:
amdgpu:
 - SR-IOV fix
 - GFX8 fix
 - MES queue reset fix
 - GPUVM fixes
 - DCN 6 warning fix
 - DCN 3.5/3.6 fix
 - DML fix
 - Backlight fix
 - Colorop fix
 - DC get_estimated_bw() fix
 - devcoredump fix
 - Userq fixes
 - APU PSP fix
 - Cursor fix
 
 amdkfd:
 - MES queue eviction fix
 - MQD debugfs fix
 
 UAPI:
 - Fix for drm_amdgpu_info_device with mixed 64 bit kernel and 32 bit userspace
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQQgO5Idg2tXNTSZAr293/aFa7yZ2AUCapmnvwAKCRC93/aFa7yZ
 2DbfAQCapkI0p5iRMd/2fk2JcdhhaHfTtwdNEKyiHx7Z8Fyo7wD/egYUpCbhpy4W
 6bavqT8G5Gkn4+myqJmD9bIVoWmdlAA=
 =tJ2Q
 -----END PGP SIGNATURE-----

Merge tag 'amd-drm-fixes-7.3-2026-09-03' of https://gitlab.freedesktop.org/drm/amdgpu/kernel into drm-fixes

amd-drm-fixes-7.3-2026-09-03:

amdgpu:
- SR-IOV fix
- GFX8 fix
- MES queue reset fix
- GPUVM fixes
- DCN 6 warning fix
- DCN 3.5/3.6 fix
- DML fix
- Backlight fix
- Colorop fix
- DC get_estimated_bw() fix
- devcoredump fix
- Userq fixes
- APU PSP fix
- Cursor fix

amdkfd:
- MES queue eviction fix
- MQD debugfs fix

UAPI:
- Fix for drm_amdgpu_info_device with mixed 64 bit kernel and 32 bit userspace

Signed-off-by: Dave Airlie <airlied@redhat.com>

From: Alex Deucher <alexander.deucher@amd.com>
Link: https://patch.msgid.link/20260903174712.584320-1-alexander.deucher@amd.com
2026-09-04 20:34:31 +10:00
Dave Airlie
5ff6e2f8a7 Merge tag 'drm-intel-fixes-2026-09-03' of https://gitlab.freedesktop.org/drm/i915/kernel into drm-fixes
drm/i915 fixes for v7.3-rc2:
- Drop an accidentally duplicated panel fitter call in DP MST
- Fix DDI clock programming for Cx0 and LT PHY
- Fix PTL CDCLK handling at probe, causing a glitch
- Fix dg2_power_well_count() return type
- Fix a NULL pointer deref at forced probe
- Fix selective fetch disable

Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Jani Nikula <jani.nikula@intel.com>
Link: https://patch.msgid.link/affe11af9d5eb9dc6f906441495cb843f9d4817c@intel.com
2026-09-04 15:58:07 +10:00
Dave Airlie
42bc1b92c9 A whole bunch of fixes for various drivers
- Fix drm_crtc_commit leak when PAGE_FLIP_EVENT is used,
 - amd: plane blend mode fixes
 - amdxdna: out-of-bounds access fix, reject commands chains with no
   commands, handle chained mapping BO failures, refuse to flush an
   imported BO
 - atomic-state-helpers: set pixel_blend_mode to prop default on reset
 - dma-buf: Publish the dma-buf only after copy_to_user succeeds, fix
   some kernel-doc warnings
 - ethosu: handle mmio mapping failures, handle storage modes only on
   hardware that supports it, fix job completion fence cleanup
 - fastrpc: Publish the dma-buf only after copy_to_user succeeds
 - gud: Improve TV modes and rotation handling
 - nouveau: use-after-free fixes, add scanline position support, HDMI
   and DP fixes, null pointer dereference fix, dmem accounting fixes for
   large folios, use write-combined maps for coherent
 - pagemap: Prevent double migration of device pages, Reset migration
   page count on eviction retry, dma-unmap pages before handling
   migration errors, use after free fixes
 - prime: fix prime exports tracing
 - qaic: out-of-bounds access fix
 - sysfb: Fix integer overflow, fix constant comparison bug
 - tegra: Add blend mode properties
 - virtio: exit path and error handling fixes
 -----BEGIN PGP SIGNATURE-----
 
 iJUEABMJAB0WIQTkHFbLp4ejekA/qfgnX84Zoj2+dgUCapk9RQAKCRAnX84Zoj2+
 du3SAX9yHaEcnGDqW4cDdwSG04Q/8om+V24gOepm5HEC1Tfwwn/SthpMUArHyT4e
 PyIlngsBf0PbIyv63k4tURGMDck7RDUDZCf5YtiUR1HdhXUffTCYBClz+TdZCLpq
 qLyXJFtOqg==
 =jD2Q
 -----END PGP SIGNATURE-----

Merge tag 'drm-misc-fixes-2026-09-03' of https://gitlab.freedesktop.org/drm/misc/kernel into drm-fixes

A whole bunch of fixes for various drivers

- Fix drm_crtc_commit leak when PAGE_FLIP_EVENT is used,
- amd: plane blend mode fixes
- amdxdna: out-of-bounds access fix, reject commands chains with no
  commands, handle chained mapping BO failures, refuse to flush an
  imported BO
- atomic-state-helpers: set pixel_blend_mode to prop default on reset
- dma-buf: Publish the dma-buf only after copy_to_user succeeds, fix
  some kernel-doc warnings
- ethosu: handle mmio mapping failures, handle storage modes only on
  hardware that supports it, fix job completion fence cleanup
- fastrpc: Publish the dma-buf only after copy_to_user succeeds
- gud: Improve TV modes and rotation handling
- nouveau: use-after-free fixes, add scanline position support, HDMI
  and DP fixes, null pointer dereference fix, dmem accounting fixes for
  large folios, use write-combined maps for coherent
- pagemap: Prevent double migration of device pages, Reset migration
  page count on eviction retry, dma-unmap pages before handling
  migration errors, use after free fixes
- prime: fix prime exports tracing
- qaic: out-of-bounds access fix
- sysfb: Fix integer overflow, fix constant comparison bug
- tegra: Add blend mode properties
- virtio: exit path and error handling fixes

Signed-off-by: Dave Airlie <airlied@redhat.com>

From: Maxime Ripard <mripard@redhat.com>
Link: https://patch.msgid.link/apk9X5SkRLS9g4RF@houat
2026-09-04 13:32:49 +10:00
Jens Axboe
00ef2248c5 nvme fixes for Linux 7.3
- Harden the tcp host and target against malformed PDUs: reject C2HData
    for a non-read command, bound an over-long PDU before copying it, and
    reject unsolicited H2CData (Yehyeong, Shivam)
  - Fix circular locking on TLS queues (Xixin)
  - Fix a soft lockup when scanning sparse namespace ID space (Mohamed)
  - Fix racy access to the FDP placement id array (Kanchan)
  - RDMA host and target fixes for a double cleanup on the queue_rq
    error path and a queue leak when the connect backlog is exceeded
    (Xixin)
  - Authentication fixes: drain the target's expiry work before the SQ
    is freed, and release the DH-CHAP secret when parsing fails (Kazuki,
    Xu Rao)
  - Fix nvme-fc options double free when nvme_add_ctrl() fails (Niklas)
  - Add missing SRCU grace period to nvme_alloc_ns() error path (Tristan)
  - Skip zoned limits update when the zone info query failed (Chao)
  - Reject enabling a target namespace with no device path (Seokgyu)
  - Add opcode filtering for fault injection (Mohamed)
  - Drop the kernel-doc comments from nvme-tcp.h (Randy)
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCAAdFiEE3Fbyvv+648XNRdHTPe3zGtjzRgkFAmqZ46MACgkQPe3zGtjz
 RglHuRAAhB+37HZ4wsYuwIud6IwF4rrXMv/EU1C87Nt6ORXiTG892ECT+4G4ruZo
 GTxFl6UB+i8GRw+RivMpZAWoLtEclU1GZ/ijPpgQ4+QKva60q29/2oQQiqT+x5us
 Nhez9uuC1hywxY+HfVDfvx44ISXwPG/8ZTrGMfuEO8kbmiczY8X5LxnSicgWLTc8
 PTdXvSDw5mtinzCPozRDVuFRTcDwUhM3NzVIv7MVXgEyY1sCwcPbmj8GdCkFco1G
 WnIpWP3WplnO7yNcwMN+sSFvq8PjBFfW+LJ/25WRLQFvTpF5zwejPXNWPt90USU4
 hA9kXua1CRHuApKUeNSLMQ3PG2pbB+NjIN31ZpBnveXplmIudoQuT+wLwltBg3te
 9doRCqCGLMCNu+qPFquUOOr9+6+36pBlFTynNy5WGqen0YwY4/UEG4PEU649J3Ve
 N10KY+ttHgi3bY6JdcCVlDhdzGW/rSobNz1GL4IpqPcYdZsvMqSyLrQpZfdKfp5m
 4KWYyiVmydq6ixHpZF1yEM9y2+RZGi9AtOi6CCY5pNDt7lS257htsNSTX9T3pWcH
 IeknLIuseNEW1IgUa7RPUrEkQJsiQx7eb6wI13otJULA60T40rISshuC8yGZblni
 4zu5PQiaecjqXSWQfenEi6tvPuu9jV7YsuvK5x5qs21DnFy7pUo=
 =YJ4v
 -----END PGP SIGNATURE-----

Merge tag 'nvme-7.3-2026-09-03' of git://git.infradead.org/nvme into block-7.3

Pull NVMe fixes from Keith:

"- Harden the tcp host and target against malformed PDUs: reject C2HData
   for a non-read command, bound an over-long PDU before copying it, and
   reject unsolicited H2CData (Yehyeong, Shivam)
 - Fix circular locking on TLS queues (Xixin)
 - Fix a soft lockup when scanning sparse namespace ID space (Mohamed)
 - Fix racy access to the FDP placement id array (Kanchan)
 - RDMA host and target fixes for a double cleanup on the queue_rq
   error path and a queue leak when the connect backlog is exceeded
   (Xixin)
 - Authentication fixes: drain the target's expiry work before the SQ
   is freed, and release the DH-CHAP secret when parsing fails (Kazuki,
   Xu Rao)
 - Fix nvme-fc options double free when nvme_add_ctrl() fails (Niklas)
 - Add missing SRCU grace period to nvme_alloc_ns() error path (Tristan)
 - Skip zoned limits update when the zone info query failed (Chao)
 - Reject enabling a target namespace with no device path (Seokgyu)
 - Add opcode filtering for fault injection (Mohamed)
 - Drop the kernel-doc comments from nvme-tcp.h (Randy)"

* tag 'nvme-7.3-2026-09-03' of git://git.infradead.org/nvme: (21 commits)
  nvme-tcp.h: drop kernel-doc comments, fix a few descriptions
  nvme-fc: fix double free of fabrics options when nvme_add_ctrl() fails
  nvmet: reject namespace enable without device path
  nvmet-auth: Synchronize timeout work during SQ teardown
  MAINTAINERS: update nvme entry
  nvmet-tcp: reject unsolicited H2CData PDUs
  nvme-tcp: defer TLS inline send to io_work
  nvmet-tcp: fix out-of-bounds write when receiving an over-long PDU
  nvme-tcp: return -EPROTO for a C2HData on a write
  nvmet: print namespace IDs as unsigned 32bit value
  nvme: print namespace IDs as unsigned 32bit value
  nvme: remove stale namespaces by NSID range during scan
  nvme: add missing SRCU grace period in error path
  nvme-fabrics: fix DHCHAP secret leak on parse failure
  nvmet-rdma: fix queue leak when connect backlog is exceeded
  nvme: add opcode filtering for fault injection
  nvme: fix racy access to FDP placement id array
  nvme: set ns->head in nvme_alloc_ns_head
  nvme-rdma: fix -EIO cleanup order in queue_rq
  nvme: skip the zoned limits update if the zone info query failed
  ...
2026-09-03 19:48:07 -06:00
Linus Torvalds
bc35965f69 18 hotfixes. 13 are cc:stable. 15 are for MM.
All are singletons - please see the changelogs for details.
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQTTMBEPP41GrTpTJgfdBJ7gKXxAjgUCapoUuQAKCRDdBJ7gKXxA
 jgscAP9iRyonROgpsNKC9H8EsAL7QhZNxjwc5PWs0bN6J50LOwD/Um6G7b1P8cxs
 j7kGpxbQYI0RWxxLUBLTQiPbDrvn6wY=
 =TJzY
 -----END PGP SIGNATURE-----

Merge tag 'mm-hotfixes-stable-2026-09-03-17-45' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm

Pull misc fixes from Andrew Morton:
 "18 hotfixes.  13 are cc:stable.  15 are for MM.

  All are singletons - please see the changelogs for details.

  There are no fixes (yet) for all the stuff we added in the most recent
  merge window. Hopefully a good sign"

* tag 'mm-hotfixes-stable-2026-09-03-17-45' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm:
  mm/secretmem: properly account locked pages
  mm/mremap: reset unfaulted VMA page offset for MREMAP_DONTUNMAP
  MAINTAINERS: add Kiryl as a THP reviewer
  MAINTAINERS: cover all of RAID
  MAINTAINERS: mailmap: update entries for Thorsten Blum
  MAINTAINERS: remove Lorenzo as THP co-maintainer
  Revert "once: don't use a work queue to reset sleepable static key"
  mm/hugetlb: fix missing migratable flag on same-node hugetlb migration
  mm/mempolicy: fix sleeping allocation in alloc_pages_bulk_weighted_interleave()
  mm/huge_memory: transfer the pmd dirty bit to the folio on zap
  MAINTAINERS: add Lance Yang as a hung task detector co-maintainer
  userfaultfd: reset err to be 0 when move_pages_ptes succeeded
  mm: fix incorrect vm_flags usage when checking allowable orders for tmpfs
  mm/hugetlb: keep max_huge_pages when dissolving surplus folios
  mm/migrate_device: avoid out-of-bounds writes for compound folios
  mm/hugetlb_cgroup: call page_counter_set_max() outside VM_BUG_ON()
  memcg: make the v1 soft limit knob inert
  mm/hugetlb_cma: fix null nodemask dereference in hugetlb_cma_alloc_frozen_folio
2026-09-03 17:59:19 -07:00
Randy Dunlap
fd9beb8870 nvme-tcp.h: drop kernel-doc comments, fix a few descriptions
Expand @fei into @feil and @feih because the field was split due to it
not being 32-bit aligned.

Struct member @hdr was described twice in struct nvme_tcp_rsp_pdu, so
drop one of them.

These structs are defined in a spec outside of the kernel, so kernel-doc
comments for them aren't needed here as well.

This avoids kernel-doc warnings:

Warning: include/linux/nvme-tcp.h:95 struct member 'rsvd2' not described in 'nvme_tcp_icreq_pdu'
Warning: include/linux/nvme-tcp.h:113 struct member 'rsvd' not described in 'nvme_tcp_icresp_pdu'
Warning: include/linux/nvme-tcp.h:128 struct member 'feil' not described in 'nvme_tcp_term_pdu'
Warning: include/linux/nvme-tcp.h:128 struct member 'feiu' not described in 'nvme_tcp_term_pdu'
Warning: include/linux/nvme-tcp.h:128 struct member 'rsvd' not described in 'nvme_tcp_term_pdu'
Warning: include/linux/nvme-tcp.h:169 struct member 'rsvd' not described in 'nvme_tcp_r2t_pdu'
Warning: include/linux/nvme-tcp.h:187 struct member 'rsvd' not described in 'nvme_tcp_data_pdu'

Signed-off-by: Randy Dunlap <rdunlap@infradead.org>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Keith Busch <kbusch@kernel.org>
2026-09-03 14:15:11 -07:00
Niklas Cassel
56e6279266 nvme-fc: fix double free of fabrics options when nvme_add_ctrl() fails
nvmf_create_ctrl() owns the fabrics options and frees them whenever
->create_ctrl() returns an error, so a transport must not free them on
its own error paths.  nvme-fc tracks this by testing ctrl->ctrl.opts in
nvme_fc_ctrl_free(), which requires nvme_fc_init_ctrl() to clear that
pointer on every error exit.

The coupling is implicit, and commit 1a9e218195 ("nvme: split device
add from initialization") broke it by adding a second error exit.  When
nvme_add_ctrl() fails, nvme_fc_init_ctrl() jumps to out_put_ctrl:, past
the "ctrl->ctrl.opts = NULL" that only sits on the fail_ctrl: path, so
nvme_fc_ctrl_free() frees the options and nvmf_create_ctrl() frees them
a second time:

  BUG: KASAN: slab-use-after-free in nvmf_free_options+0x30/0x190
   nvmf_free_options+0x30/0x190 drivers/nvme/host/fabrics.c:1284
   nvmf_create_ctrl drivers/nvme/host/fabrics.c:1374 [inline]
  Freed by task 5534:
   nvme_fc_ctrl_free drivers/nvme/host/fc.c:2374 [inline]
   nvme_fc_init_ctrl+0xe17/0x1450 drivers/nvme/host/fc.c:3605

nvme_add_ctrl() fails when dev_set_name() cannot allocate, so this is
reachable under memory pressure or fault injection.  Without KASAN the
options are freed twice.

Rather than clear the pointer on the second exit as well, derive
ownership the way nvme-tcp, nvme-rdma and nvme-loop do, from list
membership: their free_ctrl leaves the options alone unless the
controller made it onto the transport list.

The list cannot simply be populated on the success path as it is there.
nvme-fc runs the initial connect synchronously via flush_delayed_work(),
and the controller has to be reachable on rport->ctrl_list for the whole
of it: nvme_fc_unregister_remoteport() needs to find it to signal
connectivity loss, nvme_fc_match_disconn_ls() matches an incoming
Disconnect Association LS against ctrl->association_id, which is only
assigned during that window, nvme_fc_resume_controller() needs it on
remoteport re-registration, and nvme_fc_existing_controller() uses it to
reject a duplicate connect racing the one in flight.

Keep the insertion where it is and add a fail_unlist: label, falling
into fail_ctrl:, for the error paths that run after it.  The earlier
error paths never reach the insertion and keep using fail_ctrl:
directly, so the list is only touched where the controller is actually
on it.

nvme_fc_ctrl_free() cannot use the plain "goto free_ctrl" the other
transports use, because it still has to put_device(), release the rport
reference and free the ida entry for resources taken before the
insertion.  Sample list_empty() under rport->lock instead.

ctrl->ctrl.opts also stays valid for the whole teardown now.  That is
not the bug being fixed, but it removes some fragility around the old
idiom: nvme_free_ctrl() calls nvme_auth_free() before ->free_ctrl(), and
ctrl_max_dhchaps() dereferences ctrl->opts without a NULL check when
ctrl->dhchap_ctxs is set, which nvme-fc permits since NVMF_ALLOWED_OPTS
allows the dhchap options.  The nvme sysfs attributes that dereference
ctrl->opts, such as hostnqn and address, evaluate their is_visible()
test once at device_add() time and stay readable until
cdev_device_del().

Fixes: 1a9e218195 ("nvme: split device add from initialization")
Cc: stable@vger.kernel.org
Reported-by: syzbot+f58e57380a6083c4041d@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=f58e57380a6083c4041d
Signed-off-by: Niklas Cassel <cassel@kernel.org>
Tested-by: Rihyeon Kim <rihyeon8648@gmail.com>
Reviewed-by: Hannes Reinecke <hare@kernel.org>
Signed-off-by: Keith Busch <kbusch@kernel.org>
2026-09-03 14:15:11 -07:00
Seokgyu Choi
09d0c07bd9 nvmet: reject namespace enable without device path
A newly allocated namespace has a NULL device_path until userspace
configures the device_path attribute.

If buffered_io is enabled before device_path is configured,
nvmet_bdev_ns_enable() returns -ENOTBLK and nvmet_ns_enable() falls
back to nvmet_file_ns_enable(). The latter passes the NULL
device_path to filp_open(), causing a NULL pointer dereference in
getname_kernel().

Reject namespace enable when device_path has not been configured.

Reported-by: syzbot+f613f9f010ec98eb9d86@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=f613f9f010ec98eb9d86
Signed-off-by: Seokgyu Choi <tjrrb0313@gmail.com>
Reviewed-by: Sagi Grimberg <sagi@grimberg.me>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Keith Busch <kbusch@kernel.org>
2026-09-03 14:15:11 -07:00
Kazuki Hanai
eaa948c0e1 nvmet-auth: Synchronize timeout work during SQ teardown
nvmet_auth_sq_free() cancels auth_expired_work with
cancel_delayed_work(). If the work has already started, cancellation does
not wait for the callback. Transport teardown can consequently free or
reuse the queue containing struct nvmet_sq while
nvmet_auth_expired_work() still accesses that SQ.

Add a teardown-specific helper that synchronously drains the delayed work
before freeing authentication state, and use it from nvmet_sq_destroy().
Keep the non-synchronous helper for in-band authentication state cleanup,
where the SQ owner remains alive.

Fixes: 1a70200f40 ("nvmet-auth: expire authentication sessions")
Cc: stable@vger.kernel.org
Signed-off-by: Kazuki Hanai <hnkz.64@gmail.com>
Reviewed-by: Sagi Grimberg <sagi@grimberg.me>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Keith Busch <kbusch@kernel.org>
2026-09-03 14:15:11 -07:00
Keith Busch
5cdd07a688 MAINTAINERS: update nvme entry
Update Jens' entry to match the mail address of his other entries.

Acked-by: Jens Axboe <axboe@kernel.dk>
Signed-off-by: Keith Busch <kbusch@kernel.org>
2026-09-03 14:15:10 -07:00