Commit Graph

1484066 Commits

Author SHA1 Message Date
Linus Torvalds
5ccda18d1b Perf events fixes:
- Fixes for KVM guest PEBS virtualization (Sean Christopherson)
 
  - Fixes for various Intel PMUs related to PEBS data-source
    (Dapeng Mi)
 
  - Fix Intel Panther Cove event scheduling constraints (Dapeng Mi)
 
  - Fix Intel DMR/NVL OMR extra registers event scheduling (Dapeng Mi)
 
  - Rename two confusingly named PMU attributes (Dapeng Mi)
 
  - Fix a refcount leak in attach_perf_ctx_data() (Namhyung Kim)
 
  - Fix NULL pointer dereference crash in __perf_pmu_sched_task()
    (Puranjay Mohan)
 
  - Fix CPU-wide event scheduling (Puranjay Mohan)
 
  - Fix x86 LBR branch entry generation (Puranjay Mohan)
 
 Signed-off-by: Ingo Molnar <mingo@kernel.org>
 -----BEGIN PGP SIGNATURE-----
 
 iQJFBAABCgAvFiEEBpT5eoXrXCwVQwEKEnMQ0APhK1gFAmq4zGoRHG1pbmdvQGtl
 cm5lbC5vcmcACgkQEnMQ0APhK1guAw//U+WiZoWkWv6oFgAg+KrQFPr/OAzyIYOe
 Q355hrxdb2yh5+YGkLl/OCte/wIg1JiUq54gUBHPjEBUK8lxenNK2kOo6/Zd8tDx
 oGPyX6t0tpiP1O/PMII3yyd6Q7DYVC08TqOYW68r7Jv/fO2qCtBWZY3hxe0otU2/
 z+jqcQno3o8DbzfJnlDJb5honJos8CaT5FA+FkZkvjBF4dMWVErQowefdz2Zd78H
 Bu/wxmPG+Hej7ownQalps+ZA52KuQboRwJ+NCmas+fbFCcaWRbdDKTPbvEvzBsHT
 nybzlmBdneikrM0aNAXtBxJCLNzVrB6ffmKNFnd3CyfXmNtd/zleoqcKZpII2xO3
 UqlnwWR+u2JMGscvWiAAtlkdI0K5P4U1sHObhvzyXwgi3LgRZWJoLoIyiYWGcfnW
 Mxe3MCTy0wS+AEKRDvzg+rRr4jqTfl0BsHyzaTyTmWwx0LEg+7zwdryeixpXBbwH
 qC/SWkr4sfQlO/kVZHGa1sdHTZocRLcLl3YHsLiwQygu3rjNes+olKyPTtbvt9/a
 ZePBdcP+84K146VUczNJRAI4qt5b/X2aMvIYYxVvfDF6kiL0GbXgAfhdNVm5BIZF
 u5e2Sb6AX3w78V+YaMX1DuF+ZaFrpQd2UKOkBJPLvFnHTQr2eP2Fa6XfVMXnMcDy
 U7APgunF6D8=
 =Ggeg
 -----END PGP SIGNATURE-----

Merge tag 'perf-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip

Pull perf events fixes from Ingo Molnar:

 - Fixes for KVM guest PEBS virtualization (Sean Christopherson)

 - Fixes for various Intel PMUs related to PEBS data-source (Dapeng Mi)

 - Fix Intel Panther Cove event scheduling constraints (Dapeng Mi)

 - Fix Intel DMR/NVL OMR extra registers event scheduling (Dapeng Mi)

 - Rename two confusingly named PMU attributes (Dapeng Mi)

 - Fix a refcount leak in attach_perf_ctx_data() (Namhyung Kim)

 - Fix NULL pointer dereference crash in __perf_pmu_sched_task()
   (Puranjay Mohan)

 - Fix CPU-wide event scheduling (Puranjay Mohan)

 - Fix x86 LBR branch entry generation (Puranjay Mohan)

* tag 'perf-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
  perf/core: Fill branch entries with a single assignment
  perf/core: Run sched_task() for PMUs with only CPU-wide events
  perf/core: Fix NULL pmu_ctx passed to pmu->sched_task()
  perf/core: Fix a refcount leak in attach_perf_ctx_data()
  perf/x86/intel: Rename NVL offcore_rsp attribute to offmodule_rsp
  perf/x86/intel: Rename DMR offcore_rsp attribute to offmodule_rsp
  perf/x86/intel: Fix precise OMR event scheduling for DMR/NVL
  perf/x86/intel: Constrain Panther Cove UOPS_DISPATCHED events to PMCs 0-3
  perf/x86/intel: Delete dead NVL PEBS data-source initcall
  perf/x86/intel: Fix Panther Cove PEBS data-source snoop states
  perf/x86/intel: Remove incorrect Panther Cove PEBS data-source constraints
  perf/x86/intel: Remove incorrect LionCove PEBS data-source constraints
  perf/x86/intel: Update arw_latency_data() mem-op direction handling
  perf/x86/intel: Fix DKT PEBS load/store direction for latency events, to fix sample classification
  perf/x86/intel: Fix CMT PEBS load/store direction for latency events, to fix sample classification
  perf/x86/intel: Fix GRT PEBS load/store direction for latency events, to fix sample classification
  perf/x86/intel: Make @data a mandatory param for intel_guest_get_msrs()
  perf/x86/intel: Don't pointlessly context switch DS_AREA (and PEBS config) if PEBS is unused
  perf/x86/intel: Don't write PEBS_ENABLED on host<=>guest xfers if CPU has PEBS isolation, to fix stuck PEBS_ENABLED
  perf/x86/intel: Ensure KVM guest PEBS path doesn't set unwanted PERF_GLOBAL_CTRL bits
2026-09-27 08:15:58 -07:00
Linus Torvalds
fd179f8a05 ata fixes for 7.3-rc5
- Extend the quirk "no LPM on ATI" quirk, that is currently only
    applied for Samsung drives, to include AMD controllers as well.
    The AMD AHCI controllers are newer versions of the ATI AHCI
    controllers, and these controllers still have LPM issues with
    Samsung drives - LPM works with drives from other vendors (me)
 
  - Fix errors in the libata.force parameter documentation (me)
 
  - Verify the sense data descriptor lengths for ATA PASS-THROUGH
    command, so that a malicious device cannot write past the buffer
    length (Matthias)
 
  - Mention the libata for-next branch in MAINTAINERS such that the
    git ls-remote command done by get_maintainer.pl --self-test=scm
    can verify it (Matthias)
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQRN+ES/c4tHlMch3DzJZDGjmcZNcgUCargFGAAKCRDJZDGjmcZN
 cm3QAP97hnUqqmFfM20u4KSVo/1KV93S/uJ4kVGdsDf5HNvsZgEApw/OlB7lZ8dE
 hZ3otv4Yv1bxC4rdOs6RIJrz8aDCBQU=
 =q1Z4
 -----END PGP SIGNATURE-----

Merge tag 'ata-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux

Pull ata fixes from Niklas Cassel:

 - Extend the quirk "no LPM on ATI" quirk, that is currently only
   applied for Samsung drives, to include AMD controllers as well.

   The AMD AHCI controllers are newer versions of the ATI AHCI
   controllers, and these controllers still have LPM issues with
   Samsung drives - LPM works with drives from other vendors (me)

 - Fix errors in the libata.force parameter documentation (me)

 - Verify the sense data descriptor lengths for ATA PASS-THROUGH
   command, so that a malicious device cannot write past the buffer
   length (Matthias)

 - Mention the libata for-next branch in MAINTAINERS such that the
   git ls-remote command done by get_maintainer.pl --self-test=scm
   can verify it (Matthias)

* tag 'ata-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux:
  MAINTAINERS: name the libata/linux for-next branch
  ata: libata-scsi: bound the ATA passthru sense descriptor writes
  ata: libata: Correct libata.force parameter documentation
  ata: libata-core: Extend Samsung LPM quirk to AMD controllers
2026-09-26 11:14:35 -07:00
Linus Torvalds
fddfc3ec31 pci-v7.3-fixes-2
-----BEGIN PGP SIGNATURE-----
 
 iQJIBAABCgAyFiEEgMe7l+5h9hnxdsnuWYigwDrT+vwFAmq38q4UHGJoZWxnYWFz
 QGdvb2dsZS5jb20ACgkQWYigwDrT+vwVFhAAhO4zr6WHzB/Tx87FGSxYnjrwzOmH
 nS/w0MD0FxPkAzhbkPaXkdgAks10dBNF/umQUXvWBwgZYbR9v0CoZBSm8viuq/8C
 tYOdOm4u4whsHVPU1QJPV5nUSGb2Eiq7vza+2kS1Ll0xy/3/24yVn+IU8sy4EcgX
 IUTwApHfBGq6se88LxtdT5ctAUCRSP8XAeFlKXqfPig9yrrQOZ0cJrA5nhbe4jCL
 jJMFjgH12OvR55ZeKyZianz/lCZ8NcyATblM0EUCVKpx1gv4BcPfbprcycgnyLbw
 R+SoPl0wJDWKGqwuCFYkLMaT4XsLYnJv+6C/aQQ8CJjovAVUzaeDXwlJzPP3hH+N
 iIZp3b5tULHK/c/AykjoRQPq/rVy9ywMvlqmafIeR4+BEJBVYlUGLhum14pnjAQo
 B/wLCSmxrP7K7E+71ODunUdpts6HsH+GbKqSjBS46BAY/Vw+icq0v5RxaIaa6WCv
 BEkVEIfcJ/pnlmxMCoQEVu0bjCrIIaTe7EflO4BDhEIrWTxSq19vCVlXB+XgsHUI
 DWPyCFnAWFG5bZi50PKkrwTOhmcRMRgAjc69DOZ400chbuMAj9PxpomYBEeWvnss
 N+DLTLL2Yp4Y8/tMJIxjNB2eAbiSKql2EoGIpyjQruDRmg1Bcug5gBY90RjRr204
 afoNJYCbtPnd0mY=
 =hhCg
 -----END PGP SIGNATURE-----

Merge tag 'pci-v7.3-fixes-2' of git://git.kernel.org/pub/scm/linux/kernel/git/pci/pci

Pull PCI fixes from Bjorn Helgaas:

 - Make BAR resize work even for devices where no upstream bridge is
   visible to the OS, which fixes an amdgpu regression on SolidRun
   HoneyComb, which doesn't expose Root Ports to the OS (Liz Fong-Jones)

 - Omit bus properties in dynamic OF nodes when a bridge has no
   subordinate bus, which fixes early boot hangs caused by NULL pointer
   dereferences with CONFIG_PCI_DYNAMIC_OF_NODES enabled (Angel J)

 - Disable enhanced atomics on AMD NBIO 7.7 and 7.11 to avoid silent
   data corruption on 64-bit DMAs (Mario Limonciello)

* tag 'pci-v7.3-fixes-2' of git://git.kernel.org/pub/scm/linux/kernel/git/pci/pci:
  x86/PCI: Disable enhanced atomics on AMD NBIO 7.7 and 7.11
  PCI: of_property: Omit bus properties without a subordinate bus
  PCI: Fix BAR resize for devices on a root bus
2026-09-26 09:59:59 -07:00
Linus Torvalds
efb27d4767 Probes fixes for v7.3-rc4:
- kprobes: Fix permanent hang when flushing the kprobe optimizer
   Fix a deadlock when disabling kprobe optimization via sysctl or debugfs
   where flushers hung waiting for optimizer_completion. Replaced the
   completion with an optimizer_passes counter and wait_var_event_mutex()
   under kprobe_mutex so concurrent flushers can wait and wake up safely.
 
 - fprobe: Terminate the fgraph_data list when the reservation is not filled
   Fix an issue where unused shadow stack data left uninitialized by
   fprobe_fgraph_entry() was misparsed as stale fprobe headers on return.
   Explicitly write a zero word to terminate the list and update
   read_fprobe_header() to handle the zeroed slot properly.
 
 - ftracetest: Fix unique symbol check in kprobe_non_uniq_symbol.tc
   Fix false test failures in kprobe_non_uniq_symbol.tc on architectures
   like s390 where a symbol exists once in core kernel but also in modules.
   Anchor the /proc/kallsyms search regex to the end of the line so that
   module symbols are not incorrectly counted.
 -----BEGIN PGP SIGNATURE-----
 
 iQFPBAABCgA5FiEEh7BulGwFlgAOi5DV2/sHvwUrPxsFAmq3llUbHG1hc2FtaS5o
 aXJhbWF0c3VAZ21haWwuY29tAAoJENv7B78FKz8bAhEH/0EAamjv7/EDUoUq+BOO
 a2gnlYqvr+zcrDVQLNgiYbvTRDfIFPOdB2LpY7Rguee3747qeL7kkNATD10WFr1F
 5lXe5LaLncNIrvHDdtcT5eER5ePAuSDMSL5CwnJrRvXJw42iFsqegZ07nrvc9HFS
 5zw7Ej9VnFJFxeXIY3J4U92wkntLJ3JhsNheomOtQmEZU1g5ZPAbdq0icNrC3CAb
 FTezYk60VG0CT/gNTSd8JFnI4P5vKlZpkFFCLLMmptW4yQU9+ZdT55JRgY7wJ8ye
 w5rLBTRJrFhDgNAjoehMkoAtxik3dKgede7qlwxUk0sv+gwXssw28b5zabRBJgee
 XQ8=
 =cHwO
 -----END PGP SIGNATURE-----

Merge tag 'probes-fixes-v7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace

Pull probe fixes from Masami Hiramatsu:

 - kprobes: Fix permanent hang when flushing the kprobe optimizer

   Fix a deadlock when disabling kprobe optimization via sysctl or
   debugfs where flushers hung waiting for optimizer_completion.
   Replaced the completion with an optimizer_passes counter and
   wait_var_event_mutex() under kprobe_mutex so concurrent flushers can
   wait and wake up safely.

 - fprobe: Terminate the fgraph_data list when the reservation is not
   filled

   Fix an issue where unused shadow stack data left uninitialized by
   fprobe_fgraph_entry() was misparsed as stale fprobe headers on
   return. Explicitly write a zero word to terminate the list and update
   read_fprobe_header() to handle the zeroed slot properly.

 - ftracetest: Fix unique symbol check in kprobe_non_uniq_symbol.tc

   Fix false test failures in kprobe_non_uniq_symbol.tc on architectures
   like s390 where a symbol exists once in core kernel but also in
   modules. Anchor the /proc/kallsyms search regex to the end of the
   line so that module symbols are not incorrectly counted.

* tag 'probes-fixes-v7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace:
  kprobes: Fix permanent hang when flushing the kprobe optimizer
  fprobe: Terminate the fgraph_data list when the reservation is not filled
  selftests/ftrace: Fix unique symbol check in kprobe_non_uniq_symbol.tc
2026-09-26 08:36:29 -07:00
Linus Torvalds
eff8d2791c Arm:
- Invalidate the ITS translation cache when the guest changes the
   base address of the ITS tables (Fuad Tabba)
 
 - Skip saving ITS devices with device IDs that are out-of-bounds
   rather than failing the entire ITS save ioctl (Fuad Tabba)
 
 - Close race between VM teardown and invalidations of nested MMUs
   when handling MMU operations that are allowed to block
   (Lorenzo Stoakes)
 
 - Various fixes for the handling of the host's untrusted SVE
   configuration in pKVM (Fuad Tabba)
 
 - Make sure that empty SMCCC ranges based at 0 are rejected by the
   kvm_smccc_set_filter() (Karl Mehltretter)
 
 - Revoke the host mapping for pKVM's private stack pages, along
   with a new sanity check that all mappings in the hyp's private
   VA range have been correctly marked as hyp-owned (Fuad Tabba)
 
 - Lifetime fixes for the array of shadow stage-2 MMUs, ensuring that
   concurrent vCPU initialization cannot relocate in-use MMUs. Defer
   the freeing of shadow stage-2 MMUs to the point that no other users
   (e.g. MMU notifier) could reference them (Marc Zyngier)
 
 - Drop useless WARN when rejecting an unsupported ioctl for pKVM
   (Fuad Tabba)
 
 - Fix the steal_time selftest to install correctly-sized mappings for
   non-4K hosts (Sebastian Ott)
 
 - Correct mapping of fine-grained trap for GCSPOPX instruction
   (Mark Brown)
 
 - Fix KVM_BUG_ON() due to missing handling of DBGBXVR<n> from 32-bit
   guests (Karl Mehltretter)
 
 RISC-V:
 
 - Synchronize hrtimer during VCPU teardown.
 
 - Fix the conversion between vsip and hvip values.
 
 - Serialize IMSIC attributes with vCPU migration.
 
 - Release unused page after MMU invalidation.
 
 - Propagate interrupted G-stage faults to KVM user-space as EINTR.
 
 - Fix nested acceleration hfence entry update order.
 
 - Fix sdata leak and stale snapshot_addr in snapshot_set_shmem.
 
 - Preserve firmware counter value across PMU counter stop/start.
 
 - Report PMU snapshot write failure to the guest.
 
 - Fix perf-backed counter accounting across PMU stop and read.
 
 - Correctly propagate error of a hart status SBI call.
 
 s390:
 
 - Ensure that accesses through kvm_arch_set_irq_inatomic mark as dirty
   the pages that contain indicator and summary bits.
 
 - Fix compile warning for kvm_s390_update_cmma_dirty().
 
 - Fix incorrect propagation of ENOENT from _gaccess_shadow_fault() to
   userspace.
 
 - Move s390_kvm_mmu_commit_memory_region() into
   s390_kvm_mmu_prepare_memory_region() so that it can fail instead of WARN.
 
 - Add missing srcu in kvm_s390_set_irq_state().
 
 - Fix potential races in storage functions.
 
 - Fix race in _destroy_pages_crste().
 
 - Fix issues in the handling of KVM interrupt and page resources, when a
   queue that is assigned to a mediated device (mdev) is removed from the
   host's AP configuration.
 
 - Fix loop condition in uv_find_secrets.
 
 - Prevent potential out-of-bounds read.
 
 x86:
 
 - Fix a brown paper bag bug where KVM would incorrectly treat Intel PMU MSRs
   as valid on AMD.
 
 - Fix a regression in the hardware disable selftest where it checked the wrong
   macro when detecting glibc support (breaks at least musl).
 
 - Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is especially
   important for KVM_BUG_ON() flows, which often guard more dangerous bugs.
 
 - Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a bug
   where KVM would let userspace run a broken setup with stale vmcs12 pages.
 
 - Fix a class of bugs where KVM would fail to fill kvm_run exit fields if
   getting nested pages failed.
 
 - Treat reserved entries in the memory attributes xarray as "no attributes",
   to fix false positives when checking for mixed attributes.
 
 - Fix memcg accounting for the memory attributes xarray (the xarray library
   subtly requires the xarray to be configured for accounting upfront; the gfp
   flags taken at runtime are used only rarely).
 
 - Don't pre-reserve xarray entries when storing empty attributes, as storing
   NULL must not require memory allocation (KVM and other subsystems heavily
   rely on this behavior).
 
 - Fix a memory leak and a cache maintenance issue related to doing intra-host
   migration on an SEV guest.
 -----BEGIN PGP SIGNATURE-----
 
 iQFIBAABCAAyFiEE8TM4V0tmI4mGbHaCv/vSX3jHroMFAmq3Ta0UHHBib256aW5p
 QHJlZGhhdC5jb20ACgkQv/vSX3jHroMcygf/Z8GovgUACbWs/PSgVxZr5++MsbnM
 /N+vUP0VhsBxWCt191fBmKSBMNyTL734CKe0fkmisIAIPxf63uA3lttfMshNXKct
 VJ3TQkbA7qiSEGf2Um1dp4k6egqDDbTS1a4NF7CM7AEL4Wm3YDfqv2h4rZicEBsd
 Spy4yV8sKoiEHMnduhF2m5r0cWkrj194e2dolJy0p1WiNeEykf/nMTmNq5JWusBh
 TDDzZf1p2BianOmQ6+fkzAfWDSIdv0OPmoH7iqvQ/xhvm4dLWq2oeOz5Q2/uWAp0
 BqU4Z5zBrGRUWyXM0cDZEqBfmNH4v7oR9PfLgvD4qpMoqONuMj1+mIO68A==
 =XtLD
 -----END PGP SIGNATURE-----

Merge tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm

Pull kvm fixes from Paolo Bonzini:
 "Arm:

   - Invalidate the ITS translation cache when the guest changes the
     base address of the ITS tables (Fuad Tabba)

   - Skip saving ITS devices with device IDs that are out-of-bounds
     rather than failing the entire ITS save ioctl (Fuad Tabba)

   - Close race between VM teardown and invalidations of nested MMUs
     when handling MMU operations that are allowed to block (Lorenzo
     Stoakes)

   - Various fixes for the handling of the host's untrusted SVE
     configuration in pKVM (Fuad Tabba)

   - Make sure that empty SMCCC ranges based at 0 are rejected by the
     kvm_smccc_set_filter() (Karl Mehltretter)

   - Revoke the host mapping for pKVM's private stack pages, along with
     a new sanity check that all mappings in the hyp's private VA range
     have been correctly marked as hyp-owned (Fuad Tabba)

   - Lifetime fixes for the array of shadow stage-2 MMUs, ensuring that
     concurrent vCPU initialization cannot relocate in-use MMUs. Defer
     the freeing of shadow stage-2 MMUs to the point that no other users
     (e.g. MMU notifier) could reference them (Marc Zyngier)

   - Drop useless WARN when rejecting an unsupported ioctl for pKVM
     (Fuad Tabba)

   - Fix the steal_time selftest to install correctly-sized mappings for
     non-4K hosts (Sebastian Ott)

   - Correct mapping of fine-grained trap for GCSPOPX instruction (Mark
     Brown)

   - Fix KVM_BUG_ON() due to missing handling of DBGBXVR<n> from 32-bit
     guests (Karl Mehltretter)

  RISC-V:

   - Synchronize hrtimer during VCPU teardown

   - Fix the conversion between vsip and hvip values

   - Serialize IMSIC attributes with vCPU migration

   - Release unused page after MMU invalidation

   - Propagate interrupted G-stage faults to KVM user-space as EINTR

   - Fix nested acceleration hfence entry update order

   - Fix sdata leak and stale snapshot_addr in snapshot_set_shmem

   - Preserve firmware counter value across PMU counter stop/start

   - Report PMU snapshot write failure to the guest

   - Fix perf-backed counter accounting across PMU stop and read

   - Correctly propagate error of a hart status SBI call

  s390:

   - Ensure that accesses through kvm_arch_set_irq_inatomic mark as
     dirty the pages that contain indicator and summary bits

   - Fix compile warning for kvm_s390_update_cmma_dirty()

   - Fix incorrect propagation of ENOENT from _gaccess_shadow_fault() to
     userspace

   - Move s390_kvm_mmu_commit_memory_region() into
     s390_kvm_mmu_prepare_memory_region() so that it can fail instead of
     WARN

   - Add missing srcu in kvm_s390_set_irq_state()

   - Fix potential races in storage functions

   - Fix race in _destroy_pages_crste()

   - Fix issues in the handling of KVM interrupt and page resources,
     when a queue that is assigned to a mediated device (mdev) is
     removed from the host's AP configuration

   - Fix loop condition in uv_find_secrets

   - Prevent potential out-of-bounds read

  x86:

   - Fix a brown paper bag bug where KVM would incorrectly treat Intel
     PMU MSRs as valid on AMD

   - Fix a regression in the hardware disable selftest where it checked
     the wrong macro when detecting glibc support (breaks at least musl)

   - Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is
     especially important for KVM_BUG_ON() flows, which often guard more
     dangerous bugs

   - Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a
     bug where KVM would let userspace run a broken setup with stale
     vmcs12 pages

   - Fix a class of bugs where KVM would fail to fill kvm_run exit
     fields if getting nested pages failed

   - Treat reserved entries in the memory attributes xarray as "no
     attributes", to fix false positives when checking for mixed
     attributes

   - Fix memcg accounting for the memory attributes xarray (the xarray
     library subtly requires the xarray to be configured for accounting
     upfront; the gfp flags taken at runtime are used only rarely)

   - Don't pre-reserve xarray entries when storing empty attributes, as
     storing NULL must not require memory allocation (KVM and other
     subsystems heavily rely on this behavior)

   - Fix a memory leak and a cache maintenance issue related to doing
     intra-host migration on an SEV guest"

* tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (54 commits)
  KVM: SEV: Do cache maintenance on the source VM during intra-host migration
  KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV
  KVM: Don't pre-reserve xarray entries when storing empty/NULL attributes
  KVM: Ensure memory attributes xarray nodes are accounted to the caller's memcg
  KVM: Don't treat reserved xarray entries as having memory attributes
  KVM: x86: Fill kvm_run exit fields in common get_nested_state_pages() error paths
  KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages fails
  KVM: arm64: Fix AArch32 DBGBXVR<n> handling
  KVM: arm64: Fix FGT mapping for HFGITR_EL2.nGCSEPP
  KVM: selftests: fix steal_time for arm64 with host page size > 4K
  KVM: arm64: Don't WARN on an unknown VM ioctl in protected mode
  KVM: arm64: nv: Delay freeing of shadow S2 structures until VM destruction
  KVM: arm64: nv: Fix life cycle of the nested_mmus array
  KVM: arm64: Check every private mapping is hyp-owned at pKVM init
  KVM: arm64: Move the private VA allocation cursor to __io_map_next
  KVM: arm64: Match hyp text by physical address in fix_host_ownership()
  KVM: arm64: Transfer the hyp stack pages out of the host stage-2
  KVM: arm64: selftests: Test empty SMCCC filter range at base 0
  KVM: arm64: Return -EINVAL for an empty SMCCC filter range at base 0
  KVM: arm64: Derive GUEST_HAS_SVE from the SVE feature bit at EL2
  ...
2026-09-26 08:26:12 -07:00
Paolo Bonzini
c2f24f140c KVM fixes for 7.3-rcN
- Fix a brown paper bag bug where KVM would incorrectly treat Intel PMU MSRs
    as valid on AMD.
 
  - Fix a regression in the hardware disable selftest where it checked the wrong
    macro when detecting glibc support (breaks at least musl).
 
  - Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is especially
    important for KVM_BUG_ON() flows, which often guard more dangerous bugs.
 
  - Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a bug
    where KVM would let userspace run a broken setup with stale vmcs12 pages.
 
  - Fix a class of bugs where KVM would fail to fill kvm_run exit fields if
    getting nested pages failed.
 
  - Treat reserved entries in the memory attributes xarray as "no attributes",
    to fix false positives when checking for mixed attributes.
 
  - Fix memcg accounting for the memory attributes xarray (the xarray library
    subtly requires the xarray to be configured for accounting upfront; the gfp
    flags taken at runtime are used only rarely).
 
  - Don't pre-reserve xarray entries when storing empty attributes, as storing
    NULL must not require memory allocation (KVM and other subsystems heavily
    rely on this behavior).
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEKTobbabEP7vbhhN9OlYIJqCjN/0FAmq3A78ACgkQOlYIJqCj
 N/1uVA//TVDznH49htPNV4dh7vfiE8oF9ka4e9KTPA2eMdDeNVInNXBz0ZbJCdNN
 Hm7iZZfk8yx0JCCGXdp8EoYg66iwulBr0KZXRP2U2LoyiRELA1VtGGOLUVP3PXjF
 hElXD56+RrkJJSucINlFU7kSfl2hY/Jbaq87HVBLpXp8iY3stBbLdgMT1ReLdorO
 s9i2uJRqHKjtSG330qDYEG1TtRAdUehgRbxbbJVuOEYtU37tCvX6ai8dEjZAvxYo
 1VZOD05ITo12HUXPiZjIMoNTujQnZKmGYlBYczWlfOmkY/tN6SV8M41srV3ZcK0k
 XGOQM6tOzbVTWaSH/pW1krklnOFAbnL6QzHzb3enHZfUN8AikHPDyfsJWzt/s0f/
 ATTQk03OYL4fJMQSQ778EI6S4PcOzY9HiBxDXM5ySpUv1uyXw8ZFsQSdH/TYTMhi
 Dce42r5rlGa1vAKHwI1Xya2su3REhWzxYxGpUOg/PtRRqWvL2RMrHArzEd6zu+Ti
 jPinuQuXYcT1FTCa5S39NWhU4gWpjw8Pvs9VuUYSCDYxZCaqQ2YheYgsAcgPqgfE
 hS43XVmmOB/tXGsvu+/YuEfbzNV/mHOuJBcjBnnsBP3d6EFB02FPKg4/LYiXujYo
 OocUbSB7bKd87W3B+Ys3QmgOtZh4qX/aLYusa2vQ4Dp9ujXqcWs=
 =LG0B
 -----END PGP SIGNATURE-----

Merge tag 'kvm-x86-fixes-7.3-rc5' of https://github.com/kvm-x86/linux into HEAD

KVM fixes for 7.3-rcN

 - Fix a brown paper bag bug where KVM would incorrectly treat Intel PMU MSRs
   as valid on AMD.

 - Fix a regression in the hardware disable selftest where it checked the wrong
   macro when detecting glibc support (breaks at least musl).

 - Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is especially
   important for KVM_BUG_ON() flows, which often guard more dangerous bugs.

 - Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a bug
   where KVM would let userspace run a broken setup with stale vmcs12 pages.

 - Fix a class of bugs where KVM would fail to fill kvm_run exit fields if
   getting nested pages failed.

 - Treat reserved entries in the memory attributes xarray as "no attributes",
   to fix false positives when checking for mixed attributes.

 - Fix memcg accounting for the memory attributes xarray (the xarray library
   subtly requires the xarray to be configured for accounting upfront; the gfp
   flags taken at runtime are used only rarely).

 - Don't pre-reserve xarray entries when storing empty attributes, as storing
   NULL must not require memory allocation (KVM and other subsystems heavily
   rely on this behavior).
2026-09-26 00:40:07 -04:00
Sean Christopherson
93de2a6a4b KVM: SEV: Do cache maintenance on the source VM during intra-host migration
Manually perform cache maintenance on the source VM during intra-host
migration to ensure no stale data is left in CPU caches after the VM is
destroyed.  Because the source VM is "converted" to a non-SEV VM, KVM's
memory reclaim flows won't trigger cache maintenance, e.g. when all guest
memory is reclaimed in response to detaching from the mmu_notifier.

Note, relying on the destination VM to do cache maintenance isn't an option
as KVM doesn't require identical guest memory configurations, i.e. the
source VM may have access to memory that the destination VM does not.
Enforcing equivalent memory configurations is infeasible, as it would
require a *deep* comparison of memslots, e.g. to verify that not only are
the memslot identical, but what the memslots point at is also identical.

Fixes: b56639318b ("KVM: SEV: Add support for SEV intra host migration")
Cc: stable@vger.kernel.org
Reported-by: Stefan Teodorescu <fane@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260923163721.1584779-3-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-09-26 00:39:55 -04:00
Sean Christopherson
12c1f6e03f KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV
Unconditionally free SEV's "have run CPUs" cpumask in the VM destroy path,
i.e. even for what appear to be non-SEV VMs, as an SEV VM becomes a non-SEV
VM if its state is intra-host migrated.  Alternatively, the mask could be
freed in sev_migrate_from() when "converting" the source VM, but that gets
annoying because ideally KVM would nullify the mask to guard against UAF,
and nullifying the mask would need be conditioned on CPUMASK_OFFSTACK=y.

Freeing the mask during sev_migrate_from() is also not robust against other
KVM bugs, though that's kind of a moot point since any such bugs would show
up even if sev->active is never set.  I.e. KVM must get that side of things
correct.  But, that's not a great reason to add more code just to make
things marginally less robust.

Fixes: 6f38f8c574 ("KVM: SVM: Flush cache only on CPUs running SEV guest")
Cc: stable@vger.kernel.org
Reported-by: Stefan Teodorescu <fane@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260923163721.1584779-2-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-09-26 00:39:55 -04:00
Linus Torvalds
6812ce4e43 drm fixes for 7.3-rc5
client:
 - fix restore of partially initialized client
 
 i915:
 - Fix incorrect RCU teardown order leading to endless loop
 - Fix DP MST TU and FEC handling for disconnected streams
 - Fix selective fetch disable, again
 - Fix export namespace for kunit helpers
 - Workaround eDP flicker on a specific laptop model
 
 xe:
 - CRI throttle reasons report
 - TLB invalidation at wedge
 - SVM eviction and VM close
 - Display corruption on LNL on Xen PV
 - W/a fix and addition
 
 amdgpu:
 - Display ref count fix
 - Userq fixes
 - VCN 4, 5 reset fixes
 - Fixes for various error paths
 - Stack frame size fixes for various combinations of compilers and configs
 
 amdkfd:
 - Possible UAF fix
 
 nouveau:
 - runtime suspend/resume fixes for newer firmware
 - rcu free the scheduler
 - fix VRAM pinning
 - fix double free
 - fix reference leaks
 - fix runtime PM leak
 - fix cursor list usage problems
 - fix HDMI config rejection without SCDC
 
 virtio:
 - fix a bunch of object/memory leaks in failure paths
 - add pixel blend mode property to cursor plane
 - revert prime buffers import
 - sync shmem backing on guest transfers
 
 imagination:
 - propogate map failures properly
 - fix page count in map interface
 - clamp freelist reconstruction requests
 
 ivpu:
 - use separate flag for job timeout
 
 bridge:
 - samsung-dsim: fix TE GPIO lifetime for host attach
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEEKbZHaGwW9KfbeusDHTzWXnEhr4FAmq29FYACgkQDHTzWXnE
 hr4pnw/+JhcAdts2saCNt/j7EtD6DqXJvyfo4BVyjYCSVP3YEZuZCrpBJh02UUUo
 sU8MgjmzwXHRxflTk8DLQMMKhN7cvFmzdREro5x7gJkCi2m/x3eanafQJS0iyvqN
 rl+K+rVeEmrOELN7auNFSl3kBtvR9yvbPZtBKNNe1IhCSq/iWKm5I4HLwP85diJZ
 OHOPuXXh+QFhtXsbwH3p1NsbWb/Kiovbevqitd/X/9CFZRMiiwuy5QFX8xOAW4CR
 0id0njQPWVbLupuqZRGSDMAsqiVq/ClKIb8X0uMzU92y6l8oMsamPXxbs7Z6pfsT
 d8k9+zMP0LXXkwdS/sTI/9axJ5LZmd4mqEL+iAkoKkO2EGnDerzyklOCWtk38qvc
 qpjhQDW77AGLoQbISKZzI2Ce9B3RotdDghyd0Bt++GJegKQAAvevEQKNzWVvlMd3
 aQcpI/EKf1PXJnX44LtDqNmoEmmEy4fKOICu52GzpPYOGaSQL25tuA7/KGNX35D8
 hY7KiJRYaPcU0hAwAPLvQDHUvFBhDJD4BHGO01YOZ0fZzBARwE57Hwp3hxkHez64
 CbcVZ5bzTREmZ3KHKI79OMhKUfCbq+E6YoJ1jxSDs902LWTrULx/iMCA+7NTZL1U
 f/9nIe2fkA0/D+++CraKd+HYkfdxvZZwuRuZtp5jWqRLxB4bdes=
 =oRa2
 -----END PGP SIGNATURE-----

Merge tag 'drm-fixes-2026-09-26' of https://gitlab.freedesktop.org/drm/kernel

Pull drm fixes from Dave Airlie:
 "While most of this is AI inspired fixes for error handling paths,
  leaks and use after frees, there are some normal things.

  nouveau has probably the biggest changes with some fixes to stabilise
  runtime suspend/resume on 570 firmware which regressed after we moved
  from 535, there are some fixes to stackframe issues seen with amdgpu,
  and otherwise the usual bunch of i915/xe/amdgpu fixes, and some
  virtio-gpu fixes.

  Hopefully it will start to quiten down a bit from here.

  client:
   - fix restore of partially initialized client

  i915:
   - Fix incorrect RCU teardown order leading to endless loop
   - Fix DP MST TU and FEC handling for disconnected streams
   - Fix selective fetch disable, again
   - Fix export namespace for kunit helpers
   - Workaround eDP flicker on a specific laptop model

  xe:
   - CRI throttle reasons report
   - TLB invalidation at wedge
   - SVM eviction and VM close
   - Display corruption on LNL on Xen PV
   - W/a fix and addition

  amdgpu:
   - Display ref count fix
   - Userq fixes
   - VCN 4, 5 reset fixes
   - Fixes for various error paths
   - Stack frame size fixes for various combinations of compilers and
     configs

  amdkfd:
   - Possible UAF fix

  nouveau:
   - runtime suspend/resume fixes for newer firmware
   - rcu free the scheduler
   - fix VRAM pinning
   - fix double free
   - fix reference leaks
   - fix runtime PM leak
   - fix cursor list usage problems
   - fix HDMI config rejection without SCDC

  virtio:
   - fix a bunch of object/memory leaks in failure paths
   - add pixel blend mode property to cursor plane
   - revert prime buffers import
   - sync shmem backing on guest transfers

  imagination:
   - propogate map failures properly
   - fix page count in map interface
   - clamp freelist reconstruction requests

  ivpu:
   - use separate flag for job timeout

  bridge:
   - samsung-dsim: fix TE GPIO lifetime for host attach"

* tag 'drm-fixes-2026-09-26' of https://gitlab.freedesktop.org/drm/kernel: (60 commits)
  drm/amd/display: Bump frame warning limit for all builds of dml
  drm/imagination: clamp freelist reconstruction requests
  drm/imagination: Fix page count for page table for map() interface
  drm/imagination: Propagate map failures correctly from pvr_mmu_map_sgl()
  drm/amd/display: Bump frame warning limit for clang builds of dml
  drm/amd/display: Relax DML frame limit with UBSAN
  drm/amdgpu: Fix runtime PM leak in amdgpu_debugfs_test_ib_show()
  drm/amdgpu: Fix last_update fence leak in amdgpu_vm_init()
  drm/amdgpu: Fix acpi device leak in amdgpu_acpi_enumerate_xcc()
  drm/amdgpu: Fix vmid_wait fence leak in amdgpu_ring_init()
  drm/amdkfd: fix use-after-free and multi-container gap in kfd_dev_mapping
  drm/amdgpu/vcn4.0.3: fix video_timeout unit mismatch in jpeg reset wait
  drm/amdgpu/vcn5.0.1: fix video_timeout unit mismatch in jpeg reset wait
  drm/amdgpu/userq: fix double jiffies conversion in hang detect timeout
  drm/amdgpu: move userq fence wait out of signalling section
  drm/amd/display: Fix dc stream excess put in dm_update_crtc_state()
  drm/xe: Add wa_14025941587 to xe2, xe3 and xe3p platforms
  drm/xe: harden adjust_idledly() against divide-by-zero and overflow
  drm/xe: Limit sg segment size to PAGE_SIZE on Xen PV
  drm/i915: fix incorrect RCU teardown order
  ...
2026-09-25 16:00:18 -07:00
Linus Torvalds
75467f60a3 ipe/stable-7.3 PR 20260925
-----BEGIN PGP SIGNATURE-----
 
 iHUEABYIAB0WIQQzmBmZPBN6m/hUJmnyomI6a/yO7QUCarbccwAKCRDyomI6a/yO
 7YKCAQDsDhTU3Rq7HCUgCkGYpZnsTAqXo3JeDP66Ku6Y0bOJjAEAhyRelXr+tKFW
 aYpuOn3qF+tihTJ52RWzAhQzEQ1BRQY=
 =QEqL
 -----END PGP SIGNATURE-----

Merge tag 'ipe-pr-20260925' of git://git.kernel.org/pub/scm/linux/kernel/git/wufan/ipe

Pull IPE fixes from Fan Wu:
 "Two fixes for use-after-free issues found by recent LLM-assisted code
  analysis.

   - move successful policy load auditing under the new policy
     directory's inode lock, preventing a concurrent policy deletion
     from freeing the policy while it is still being audited

   - protect the dm-verity root hash with RCU, preventing policy
     evaluation from racing with root hash replacement during preresume"

* tag 'ipe-pr-20260925' of git://git.kernel.org/pub/scm/linux/kernel/git/wufan/ipe:
  ipe: protect the dm-verity root hash with RCU
  ipe: fix use-after-free when auditing a newly loaded policy
2026-09-25 15:55:03 -07:00
Dave Airlie
a9ed3aa9b8 A number of fixes:
- bridge:
     - samsung-dsim: fix GPIO lifetime
   - client: Null pointer dereference fix
   - imagination: error handling fix, page handling fix
   - nouveau: fix reference leaks, double-frees, out-of-bounds accesses,
     use-after-frees, don't reject config without SCDC,  a number of
     workarounds
   - virtio: fix memory leak, reference leaks, null pointer dereference,
     add pixel blend mode, cache coherency fix
 -----BEGIN PGP SIGNATURE-----
 
 iJUEABMJAB0WIQTkHFbLp4ejekA/qfgnX84Zoj2+dgUCarU2wwAKCRAnX84Zoj2+
 djqrAYCyMUsPcuCxM0A/RWQj9DpZ2W2P6UNoglWsC8QdUtE1MRmx2BbdXFN0OMJn
 UiJ33s4BfRLkJWTCBoHbTHl1jjZTQeRY0jIvFYUEIUXeRuXYQg5P8SAktSz6+ZRc
 xBaV09iwgA==
 =/tKR
 -----END PGP SIGNATURE-----

Merge tag 'drm-misc-fixes-2026-09-24' of https://gitlab.freedesktop.org/drm/misc/kernel into drm-fixes

A number of fixes:
  - bridge:
    - samsung-dsim: fix GPIO lifetime
  - client: Null pointer dereference fix
  - imagination: error handling fix, page handling fix
  - nouveau: fix reference leaks, double-frees, out-of-bounds accesses,
    use-after-frees, don't reject config without SCDC,  a number of
    workarounds
  - virtio: fix memory leak, reference leaks, null pointer dereference,
    add pixel blend mode, cache coherency fix

Signed-off-by: Dave Airlie <airlied@redhat.com>

From: Maxime Ripard <self@mripard.dev>
Link: https://patch.msgid.link/arU22zzqUGDEco1y@houat
2026-09-26 07:58:24 +10:00
Linus Torvalds
049380360c SCSI fixes on 20260925
Mostly small driver fixes.  The biggest fix is the one to the block zone
 handling which might trip for real or virtual hardware if the number of
 zones is > 2^32.
 
 Signed-off-by: James E.J. Bottomley <James.Bottomley@HansenPartnership.com>
 -----BEGIN PGP SIGNATURE-----
 
 iLgEABMIAGAWIQTnYEDbdso9F2cI+arnQslM7pishQUCarblBBsUgAAAAAAEAA5t
 YW51MiwyLjUrMS4xMiwyLDImHGphbWVzLmJvdHRvbWxleUBoYW5zZW5wYXJ0bmVy
 c2hpcC5jb20ACgkQ50LJTO6YrIXG4wEA0mYHxKLKHGzKMTZhLIQc9LqJ+oAFQfwk
 sUfhX5kqt+QA+QGx/jrg6DKgC6LLeti8xRD+hkeSCwSPCt98Tg2/ba/O
 =pQUi
 -----END PGP SIGNATURE-----

Merge tag 'scsi-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/jejb/scsi

Pull SCSI fixes from James Bottomley:
 "Mostly small driver fixes. The biggest fix is the one to the block
  zone handling which might trip for real or virtual hardware if the
  number of zones is > 2^32"

* tag 'scsi-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/jejb/scsi:
  scsi: megaraid_sas: Protect megasas_get_ctrl_info() in megasas_resume()
  scsi: sd_zbc: Reject disks with too many zones
  scsi: block: Fix zones_cond out-of-bounds write on zone report
  scsi: leapraid: Avoid -Wformat-security warning
  scsi: devinfo: Add BLIST_SKIP_IO_HINTS for EMC Symmetrix
  scsi: libiscsi_tcp: Check the data direction of a Data-In PDU
  scsi: ufs: pltfrm: Add quirk for R-Car S4 lacking lanes-per-direction
  scsi: ufs: core: Keep internal commands dispatchable during error handling
2026-09-25 14:41:52 -07:00
Fan Wu
2776e9c285 ipe: protect the dm-verity root hash with RCU
ipe_bdev_setintegrity() frees the old root hash when dm-verity publishes
a new one on ->preresume, while policy evaluation can still be
dereferencing it.

Protect the root hash with RCU. The evaluation path already runs under
rcu_read_lock().

Fixes: e155858dd9 ("ipe: add support for dm-verity as a trust provider")
Cc: stable@vger.kernel.org
Assisted-by: LLM
[FW: remove model name according to latest guideline]
Signed-off-by: Fan Wu <wufan@kernel.org>
2026-09-25 13:30:27 -07:00
Fan Wu
9814077275 ipe: fix use-after-free when auditing a newly loaded policy
new_policy() audits the policy after ipe_new_policyfs_node() publishes it
and drops the new directory's inode lock. A concurrent delete can free
the policy while ipe_audit_policy_load() is still using it.

Audit the successful load under that lock.

Fixes: f44554b506 ("audit,ipe: add IPE auditing support")
Cc: stable@vger.kernel.org
Assisted-by: LLM
[FW: remove model name according to latest guideline]
Signed-off-by: Fan Wu <wufan@kernel.org>
2026-09-25 13:30:27 -07:00
Linus Torvalds
f14572c203 smb client fixes for v7.3-rc5
- Fix leaked server handles and dropped errors in the SMB2 compound
    create path: a parsing error reported as success, an earlier CREATE
    left open when a later command fails, the cached directory open
    losing the FID needed for cleanup, and SMB2_open() not closing the
    handle after a create-context parse failure
 
  - Fix out-of-bounds reads when parsing create contexts from a
    malicious server: bound each context by its Next field, parse the
    lease and QFid contexts from their declared offsets and validate
    the POSIX create context length
 
  - Fix a double credit decrement, and its warning, when a compound
    send fails and triggers a reconnect; found by syzbot
 
  - Fix a dentry and server handle leak in cifs_atomic_open() when an
    O_CREAT open resolves to a symlink or other non-regular inode
 
  - Use GFP_KERNEL in the DFS get_targets() path
 
  - Minor update to the POSIX extension specification references
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQTcqRusfSdYROJQwGkpVtNKoQNdYwUCarbTMwAKCRApVtNKoQNd
 Y4KIAQC58KxOQ51ppRaqMGyL/Rk6duW254zritcmacI49ekHSgD9FZfkYe/MpLFH
 eKnZ4mxjXkH7EcHmkDC2rfh5cj+edgg=
 =HCa6
 -----END PGP SIGNATURE-----

Merge tag 'cifs-fixes-7.3-rc5' of https://git.manguebit.org/linux

Pull smb client fixes from Paulo Alcantara:

 - Fix leaked server handles and dropped errors in the SMB2 compound
   create path: a parsing error reported as success, an earlier CREATE
   left open when a later command fails, the cached directory open
   losing the FID needed for cleanup, and SMB2_open() not closing the
   handle after a create-context parse failure

 - Fix out-of-bounds reads when parsing create contexts from a
   malicious server: bound each context by its Next field, parse the
   lease and QFid contexts from their declared offsets and validate
   the POSIX create context length

 - Fix a double credit decrement, and its warning, when a compound
   send fails and triggers a reconnect; found by syzbot

 - Fix a dentry and server handle leak in cifs_atomic_open() when an
   O_CREAT open resolves to a symlink or other non-regular inode

 - Use GFP_KERNEL in the DFS get_targets() path

 - Minor update to the POSIX extension specification references

* tag 'cifs-fixes-7.3-rc5' of https://git.manguebit.org/linux:
  smb: client: use finish_no_open() for non-regular inodes
  smb: client: use GFP_KERNEL in get_targets()
  smb: client: update POSIX extension specification references
  smb: client: preserve create-context parsing errors
  smb: client: close completed creates on compound wait errors
  smb: client: clean up failed cached directory opens
  smb: client: close handle after create-context parsing failure
  smb: client: validate POSIX create context length
  smb: client: fix create context out-of-bounds reads
  smb: client: delete compound mids on send failure before unlock
2026-09-25 13:30:04 -07:00
Linus Torvalds
aa98230e41 vfs-7.3-rc5.fixes
Please consider pulling these changes from the signed vfs-7.3-rc5.fixes tag.
 
 Thanks!
 Christian
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQRAhzRXHqcMeLMyaSiRxhvAZXjcogUCarab9AAKCRCRxhvAZXjc
 osTRAP90MigSgX2U/USBn8zgSTs9Key89pWuwpctsvELYKUSzwEA12UDBswbtgmq
 EGT6S2HSa6bifJMXG6AB+nMCcSluuAs=
 =KEgE
 -----END PGP SIGNATURE-----

Merge tag 'vfs-7.3-rc5.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs

Pull vfs fixes from Christian Brauner:

 - Revert "put_mnt_ns(): leave mounts connected". This allows the
   creation of reference count cycles in a very trivial way. We can't
   bring this in until we have fixed the underlying cause

 - vfs: Don't create the private nullfs instance for kthreads under
   namespace_sem to avoid false lockdeps complaints

 - binfmt_misc:
     - Copy the name into a stack buffer and look up the copy in
       bpf_binprm_select_interp()
     - bpf_binprm_set_interp() and bpf_binprm_set_interp_arg(): Check
       the private copy instead so the string that gets staged is the
       kstring that was checked

 - netfs:
     - Make netfs_read_gaps() use separate sink folios rather than one
       reused sink folio to discard unwanted data so that cifs checksum
       checking sees all the data that was fetched
     - Trim reads down to i_size so afs symlinks read correctly from the
       cache
     - Wrap the direct mempool ->alloc() calls the GFP_KERNEL paths make
       in alloc_hooks() via a new mempool_alloc_noreserve() helper

 - iov_iter: Use iov_iter_alignment() for the start and length check
   added to iov_iter_extract_bvecs() this cycle. It used iter_iov_addr()
   and iter_iov_len() which are only valid for ITER_UBUF and ITER_IOVEC
   iterators

 - super: Make iterate_supers_type() deletion-safe

 - inode: Stop evict_inodes() from rescanning the same inodes

 - writeback: Bound the cleanup_offline_cgwb() rescans

 - ntfs3: Use d_instantiate_new() in ntfs_create_inode()

 - ovl: Fix a use-after-free in the ovl_do_mkdir() debug print

 - dcache: Unpoison the inline name buffer in __d_alloc() for KMSAN

 - autofs: Fix a pipe file reference leak in autofs_kill_sb()

 - bpf: Drop the path_unlink and path_rmdir hooks from the list of hooks
   for which the verifier rewrites bpf_{set,remove}_dentry_xattr() to
   the _locked variants

 - squashfs: Range check the xz dictionary size before shifting by it

 - selftests: Add the missing eventfd, open_tree_ns, openat2 and xattr
   filesystems selftests to TARGETS and drop the stale openat2 entry
   left behind when those tests moved

* tag 'vfs-7.3-rc5.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs:
  netfs: Fix missing alloc tagging of direct mempool allocations
  bpf: fs/xattr: don't assume the inode is locked in path_unlink/path_rmdir
  autofs: fix sbi->pipe file reference leak in autofs_kill_sb()
  dcache: unpoison the inline name buffer in __d_alloc()
  ovl: fix UAF in ovl_do_mkdir() debug print
  super: make iterate_supers_type() deletion-safe
  Revert "put_mnt_ns(): leave mounts connected"
  Revert "selftests/filesystems: add mntns cleanup test"
  binfmt_misc: fix racy checks in bpf set_interp kfuncs
  binfmt_misc: fix OOB read in bpf_binprm_select_interp()
  fs: don't create the private nullfs mount under namespace_sem
  writeback: bound cleanup_offline_cgwb() rescans by rotating scanned inodes
  fs: avoid repeated scans in evict_inodes()
  netfs, afs: Fix symlink reading
  netfs: Fix netfs_read_gaps() to use separate sink folios
  squashfs: Add dictionary size range check to prevent shift-out-of-bounds
  fs/ntfs3: use d_instantiate_new() in ntfs_create_inode() and murder syzbot's "WARNING in do_new_mount" saga
  selftests/filesystems: fix missing and stale TARGETS entries
  block: Fix start and length check added to iov_iter_extract_bvecs()
2026-09-25 11:12:06 -07:00
Linus Torvalds
a2ff1b626e \n
-----BEGIN PGP SIGNATURE-----
 
 iQEzBAABCAAdFiEEq1nRK9aeMoq1VSgcnJ2qBz9kQNkFAmq2Qs8ACgkQnJ2qBz9k
 QNnBEAf5AXGzuQPeOHZQ0+kXhe0y+2Epj1oqSH0YIDI2263nav9go3wjz4uyL/J2
 0bbALzEX0Em0n0OfYOSKPdOx22AFZdZVKrFiACXeaUQmK1R0zwLOJPxDvpyWhHQS
 8Rn+HxMVxdqbobIClHQvGsU3EkBZR+d2ZSDfG4A3Gqc4O2vkgB3CvDvw5d7OdDNB
 5Et2tydekYTSUH2JZvJxlzwvbfsvvgSEEVupYILYZvMOv01EMxOTLCCv/8KWOFNE
 Xjen/maE+1ks+nGmTgDQC2bOALB3nCQhoMlL31nsYzP89Dg/C0IGCbVX25yUvlzX
 0FXMPKVVEEvpZVhTc5Hy7HYhDatC5Q==
 =IeuV
 -----END PGP SIGNATURE-----

Merge tag 'fs_for_v7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/jack/linux-fs

Pull isofs fix from Jan Kara:
 "A fix for reading tightly packed isofs directories"

* tag 'fs_for_v7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/jack/linux-fs:
  isofs: Fix handling of directories with tight blocks
2026-09-25 11:07:44 -07:00
Linus Torvalds
b9dbb658e2 Thermal control fix for 7.3-rc5
Fix a step-wise thermal governor issue that causes thermal mitigation to
 contiune forever after the temperature has dropped below the trip point
 threshold in some cases (Manaf Meethalavalappu Pallikunhi).
 -----BEGIN PGP SIGNATURE-----
 
 iQFGBAABCAAwFiEEcM8Aw/RY0dgsiRUR7l+9nS/U47UFAmq2l84SHHJqd0Byand5
 c29ja2kubmV0AAoJEO5fvZ0v1OO1NAAH/2pYr1erbmTfHAkmka5b+4/nJsAONclz
 izc6GPbLdM7PgR7LnAxAlJzQCz9wIljxtOHYKviZmlrnXbVwQxLZLI/GLpF5LN5Z
 au79wSH/q1UHsBbLUIUKp887WVEvxeS/jC6//YFUkjY+gVCiRpvKUHkLr8WzQYue
 DD1cQNR+nLquYao04+Src02VfbjKhRouMoHNhSkJXOm4IbkPMA75ENBw6LGfWshf
 tm1uSXghzog0kNNU3rqjLOdMdIpp1Ea/l1wCfg8mkwd+AVZS2G3eItQbrLQmXPWg
 yjKWpmTPggletPB4bz7LjMwkHwI+xzLIWSbr0j5zRqEopedSHhc0MXk=
 =VygM
 -----END PGP SIGNATURE-----

Merge tag 'thermal-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm

Pull thermal control fix from Rafael Wysocki:
 "Fix a step-wise thermal governor issue that causes thermal mitigation
  to contiune forever after the temperature has dropped below the trip
  point threshold in some cases (Manaf Meethalavalappu Pallikunhi)"

* tag 'thermal-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm:
  thermal: gov_step_wise: Fix stale mitigation vote with non-zero lower bounds
2026-09-25 10:11:15 -07:00
Linus Torvalds
4ba51ef66a Power management fix for 7.3-rc5
Address a hibernation regression introduced during the 7.2 development
 cycle that causes the image memory preallocation to deadlock if it
 depends on frozen kernel threads (Florian Schmaus).
 -----BEGIN PGP SIGNATURE-----
 
 iQFGBAABCAAwFiEEcM8Aw/RY0dgsiRUR7l+9nS/U47UFAmq2mDESHHJqd0Byand5
 c29ja2kubmV0AAoJEO5fvZ0v1OO1wFAH/2e9iz++Qjb9SODpGX/2Fz07qb4SBRCP
 Zf0r74G7qfoPczcLuiKu8irb1FSvlwr1nlcygWcF0gYLg9TJCaaQ7JIxHL/9ePV0
 lbINk+4ozu5S6AbMh7O5wpv3n+nwBtg1wZZP3kSY4hxQ5zymACuEzwBTGT9vfQVA
 GpkrasgQVTOyt1gWAO8Ak3WX3z1EaBqzl8DsCm/75PVq2Wy1I801JagtYT6KArf6
 kqH19SdiihrTl+2k/k2Vvc14H9XYfXab77ShWv1JgKF7X1QqaO8QWVFQulj0bpgc
 LLSlpiQXqBYfz469bPGyS2U6tyFioynqoiWQfdxC6zBGcHv1Jl3LdDg=
 =mfRN
 -----END PGP SIGNATURE-----

Merge tag 'pm-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm

Pull power management fix from Rafael Wysocki:
 "Address a hibernation regression introduced during the 7.2 development
  cycle that causes the image memory preallocation to deadlock if it
  depends on frozen kernel threads (Florian Schmaus)"

* tag 'pm-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm:
  PM: hibernate: Freeze kernel threads after image preallocation
2026-09-25 09:47:05 -07:00
Linus Torvalds
547463efb9 s390 fixes for 7.3-rc5
- Fix several bugs in PCI error recovery SCLP reporting: don't report
   success on skipped recovery, report errors when no pdev is associated,
   add missing device lock, and fix struct pci_dev reference leak in
   zpci_report_status()
 
 - Fix several bugs in CIO code: fix use of invalid SCHIB data, guard PMCW
   field accesses, check device number valid bit in PMWC before accessing
   other fields, and fix NULL pointer dereference in
   ccw_device_get_util_str()
 
 - Fix virtual vs physical address confusion in channel measurement
   facility code on kernels with CONFIG_RANDOMIZE_IDENTITY_BASE=y
 
 - Fix couple of bugs in s390dbf: fix copy of failed static debug areas,
   skip view registration on failure, and reject NULL pointer in
   debug_dump()
 
 - Fix sriov_numvfs attribute name in zPCI documentation
 
 - Fix typos in comments
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEECMNfWEw3SLnmiLkZIg7DeRspbsIFAmq2V1AACgkQIg7DeRsp
 bsJ7zg/7BrnPs0WyKv9NZvVUudYbz3nUIAleJo6HcDK2SUDr+EX3W4U7OjsXL6Fv
 H76hx5AqF6ddtUOJ2uevCAb3PKm0eA7vmK9YiERAclrBA+tWGW0xBwCqqyc84k38
 CPxOBr87Ae1jIIcfDxfBWujtKTFi/4BpSvKvKgg3mnd2egny4UnHnABx5ZcClEee
 p6aVjCKycQFHwNLMoJA1BxvGSpcXdiKXbm6VRJ4gCaBnVuzkLUv1ZMp/Pi1fWVbV
 E1AEztZ/03Tfry51q7wQLv6wD5lR3Z/Py2vZ4YIbelQwVMAMI+rFtgd5oFfo9qkF
 gcFJ+faE9m0E/nsoaaVCaa2Ja4+hV/vZ5ik0bq1ljEqEvInaa2XTuxNAljId351Y
 JrLvCdaip0FkMNMqEym0iydMAlpc491GAhzNrT3yvq6t03VVpxprPPA6VIwHg+zk
 UBMP+QdxtnaV8AAUzDD9E255d88N4DECBICUgf1XjFm/I7xLJFWyUvbzcqNn9VnY
 EGVBOeQmoByhCL9lN0+nzEM1sgT+hQB/C3WmLqiuLxrxjnc8dVr8rVHcPXFgARPw
 Ht3W9jxdTdl4VBwckOWmbvG4L6spXFLdEzI51cP46L1RKWhFjLl/6RoY8fUHf4Vz
 YnHutQevSoPI2D6G4v8K1FBz9wclxCojBPbXCmevfbL9B4ChPoQ=
 =r3uP
 -----END PGP SIGNATURE-----

Merge tag 's390-7.3-4' of git://git.kernel.org/pub/scm/linux/kernel/git/s390/linux

Pull s390 fixes from Heiko Carstens:

 - Fix several bugs in PCI error recovery SCLP reporting: don't report
   success on skipped recovery, report errors when no pdev is
   associated, add missing device lock, and fix struct pci_dev reference
   leak in zpci_report_status()

 - Fix several bugs in CIO code: fix use of invalid SCHIB data, guard
   PMCW field accesses, check device number valid bit in PMWC before
   accessing other fields, and fix NULL pointer dereference in
   ccw_device_get_util_str()

 - Fix virtual vs physical address confusion in channel measurement
   facility code on kernels with CONFIG_RANDOMIZE_IDENTITY_BASE=y

 - Fix couple of bugs in s390dbf: fix copy of failed static debug areas,
   skip view registration on failure, and reject NULL pointer in
   debug_dump()

 - Fix sriov_numvfs attribute name in zPCI documentation

 - Fix typos in comments

* tag 's390-7.3-4' of git://git.kernel.org/pub/scm/linux/kernel/git/s390/linux:
  s390/cio: Fix NULL pointer dereference in ccw_device_get_util_str()
  s390/debug: Fix NULL pointer dereference in debug_info_copy()
  s390/debug: Do not register views for failed static debug areas
  s390/debug: Reject NULL debug info in debug_dump()
  s390/cmf: Fix virtual vs physical address confusion
  s390/pci: Don't report recovery success on skipped recovery
  s390/pci: Report SCLP status on error events when no pdev is associated
  s390/pci: Fix missing device lock in zpci_report_status()
  s390/pci: Fix leak of struct pci_dev reference in zpci_report_status()
  s390/cio: Guard PMCW field accesses with dnv check
  s390/cio: Check pmcw.dnv before pmcw.ena in I/O entry points
  s390/cio: Fix cio_update_schib() to not cache invalid schib
  s390/pci/docs: Fix sriov_numvfs attribute name
  s390: Fix typos in comments
2026-09-25 09:39:33 -07:00
Linus Torvalds
80e466f0c8 gpio fixes for v7.4-rc5
- fix a regression introduced by moving GPIO hog handling into GPIOLIB
   core where of_node_name was used if line name property was missing on
   DT systems
 - fix kernel stack leak to user-space in error path in GPIO character
   device code
 - fix runtime PM leaks in gpio-xilinx and gpio-arizona
 - fix several register programming bugs in gpio-tps65219
 - fix interrupt storm on resume in gpio-mvebu
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEkeUTLeW1Rh17omX8BZ0uy/82hMMFAmq2KYwACgkQBZ0uy/82
 hMOkFhAAjIaH6nUhhUBykRyIbBVAbxYFWrps/qLw9Wlg1wDteUwwMWpdHWivsi2M
 Ch0zCR1/op812cu47dTsnNteHuK9oGoL0IHXGweylNJVU7IXTiUx6sUu+kLlT2wK
 /yIH6Zk78rULodV/80S9aNbFNNHOu7Pu/AEAPPJjGIKK6UVxT7o2o04jG1xORVbr
 X11hLjzKe1RpVKo3fhBFglzW1T4YVpulorsgkMNabTG4hbf5pArtkaQGPHVvH3Hk
 LNpTFJ3Jmoczo4UgsKAaUso2LoC19hoIUMRBVoOwscTLlPoIE6MTH0bLj+U9npId
 TUhGpVYIvfy4B6V2+FU+9YnWTkDE7efmnmm0l+GnNp7FmDTxTO0aVJoouPNEeUgu
 Zu1CJxaKYTWuTm5mdk2Tgw5G+3YtzcGS6iW3dCFvshyYSdHHkvPQ6ys2blJJyeEd
 2MRoZQZdSzMDmCfw5Y+RxF1jp1BRjIID2NH1p/cxUokFvMDgFJU+MFxdOo6ySnC/
 +1EEeA7Gak4StJUDa2Y4+uI0PMokWAEXs3ixlQUmgIpgGVtBIzI17psj4RDumdGD
 6HIG/31Pedg3LtUkkzC8P6vfdREzzquWUPBr9F7dApur4/MMJZyeu3pUwFtx0AKs
 +T5qgUwwsoKxLmg2EAdg5pThzLVgqgtrSYuw8esqfG8tNRzoRgc=
 =Mp0f
 -----END PGP SIGNATURE-----

Merge tag 'gpio-fixes-for-v7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/brgl/linux

Pull gpio fixes from Bartosz Golaszewski:

 - fix a regression introduced by moving GPIO hog handling into GPIOLIB
   core where of_node_name was used if line name property was missing on
   DT systems

 - fix kernel stack leak to user-space in error path in GPIO character
   device code

 - fix runtime PM leaks in gpio-xilinx and gpio-arizona

 - fix several register programming bugs in gpio-tps65219

 - fix interrupt storm on resume in gpio-mvebu

* tag 'gpio-fixes-for-v7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/brgl/linux:
  gpio: tps65219: Fix TPS65214 GPIO direction programming
  gpio: tps65219: Use the variant-specific direction callback
  gpio: tps65219: Fix GPIO input value reads
  gpio: zynq: fix runtime PM leak on request error path
  gpio: cdev: fix kernel stack leak to user-space in error path
  gpiolib: use of_node_name if line-name is missing
  gpio: mvebu: keep resume masks within the irqchip cache
  gpio: arizona: Fix runtime PM leak in arizona_gpio_direction_out()
2026-09-25 09:35:07 -07:00
Hao Ge
b78b728e21
netfs: Fix missing alloc tagging of direct mempool allocations
Commit 1d78d56c43 ("netfs: Fix folio_queue ENOMEM in writeback by
adding a mempool") added a mempool for the folio_queues and made the
request, subrequest and folio_queue allocations distinguish between
writeback and everything else.  Writeback is part of memory reclaim
and must not fail due to ENOMEM, so it allocates under GFP_NOFS
through mempool_alloc(), which may dip into the pool's reserve and,
if that runs empty, wait for elements to be returned.  The
GFP_KERNEL paths, which can return -ENOMEM to their callers, invoke
the pool's ->alloc() callback directly instead.

The direct call, however, skips the alloc_hooks() wrapper that the
mempool_alloc() macro provides.  The pool callbacks, mempool_alloc_slab()
and mempool_kmalloc(), call kmem_cache_alloc_noprof() and kmalloc_noprof()
and rely on current->alloc_tag having been set by the caller.  With
CONFIG_MEM_ALLOC_PROFILING_DEBUG=y this leads to

    current->alloc_tag not set
    WARNING: ./include/linux/alloc_tag.h:161 at __alloc_tagging_slab_alloc_hook
    alloc_tag was not set
    WARNING: ./include/linux/alloc_tag.h:166 at __alloc_tagging_slab_free_hook

at allocation and free time respectively, as reported when reading
files on a CIFS mount.  The allocations are also missing from
/proc/allocinfo.

Wrap the direct ->alloc() invocations in alloc_hooks() with a new
mempool_alloc_noreserve() helper in include/linux/mempool.h, next to
the other alloc_hooks()-wrapped macros such as mempool_alloc().  The
GFP_KERNEL paths keep their failable allocation semantics, they just
get tagged now.

Fixes: 1d78d56c43 ("netfs: Fix folio_queue ENOMEM in writeback by adding a mempool")
Reported-by: Erhard Furtner <erhard_f@mailbox.org>
Closes: https://lore.kernel.org/all/0b004319-9ef7-437c-a4dd-174d6a9a83db@mailbox.org/
Tested-by: Erhard Furtner <erhard_f@mailbox.org>
Suggested-by: Suren Baghdasaryan <surenb@google.com>
Cc: stable@vger.kernel.org
Signed-off-by: Hao Ge <hao.ge@linux.dev>
Link: https://patch.msgid.link/20260923063759.34667-1-hao.ge@linux.dev
Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-09-25 17:30:39 +02:00
Andrea Parri
35d442ed1f
bpf: fs/xattr: don't assume the inode is locked in path_unlink/path_rmdir
bpf_lsm_has_d_inode_locked() makes the verifier rewrite
bpf_[set|remove]_dentry_xattr() to the _locked variants, which assume
that the caller already holds the inode's i_rwsem.  The path_unlink and
path_rmdir hooks are listed, but security_path_unlink() and
security_path_rmdir() run before vfs_unlink()/vfs_rmdir() take the
victim inode's i_rwsem, so a sleepable BPF LSM program attached to
either hook mutates the victim's xattrs without the lock held.

Drop the two path hooks from d_inode_locked_hooks so that the verifier
keeps the locking bpf_[set|remove]_dentry_xattr() variants, which take
the lock themselves.

Fixes: 5646729279 ("bpf: fs/xattr: Add BPF kfuncs to set and remove xattrs")
Cc: stable@vger.kernel.org
Signed-off-by: Andrea Parri <parri.andrea@gmail.com>
Link: https://patch.msgid.link/20260922145530.369775-1-parri.andrea@gmail.com
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-09-25 17:22:09 +02:00
Hui Peng
aa5e44b29f
autofs: fix sbi->pipe file reference leak in autofs_kill_sb()
When autofs_fill_super() fails before clearing AUTOFS_SBI_CATATONIC (for
example, when find_get_pid() fails on an invalid pgrp mount option, or
when an fs_context is closed before mounting), deactivate_locked_super()
invokes autofs_kill_sb() -> autofs_catatonic_mode(sbi).

Because AUTOFS_SBI_CATATONIC is still set in sbi->flags,
autofs_catatonic_mode() returns early without calling fput(sbi->pipe),
permanently leaking the pipe struct file reference.

Explicitly release sbi->pipe in autofs_kill_sb() if it is still non-NULL
after autofs_catatonic_mode().

Fixes: ebc921ca9b ("autofs: copy autofs4 to autofs")
Signed-off-by: Hui Peng <benquike@gmail.com>
Link: https://patch.msgid.link/20260919204808.2812930-1-benquike@gmail.com
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-09-25 16:51:46 +02:00
Andrea Parri
5bfa9f1a9d kprobes: Fix permanent hang when flushing the kprobe optimizer
Writing 0 to /proc/sys/debug/kprobes-optimization while a kprobe is
jump-optimized never returns. The writer sleeps in D state forever with
kprobe_sysctl_mutex held, so any later read or write of that sysctl
hangs as well. For example, with vfs_read+9 as an optimizable address
in this build:

  # cd /sys/kernel/tracing
  # echo 'p:myprobe vfs_read+9' >> kprobe_events
  # echo 1 > events/kprobes/myprobe/enable
  # # wait until /sys/kernel/debug/kprobes/list shows [OPTIMIZED]
  # echo 0 > /proc/sys/debug/kprobes-optimization

  INFO: task sh:246 blocked for more than 10 seconds.
  Call Trace:
   <TASK>
   __schedule+0x1176/0x4f70
   schedule+0xdc/0x2c0
   schedule_timeout+0x17b/0x260
   wait_for_completion+0x173/0x3c0
   wait_for_kprobe_optimizer_locked+0xbc/0x130
   proc_kprobes_optimization_handler+0x156/0x1b0
   proc_sys_call_handler+0x324/0x490
   vfs_write+0x52d/0xfe0
   ksys_write+0xff/0x200
   do_syscall_64+0x106/0x630
   entry_SYSCALL_64_after_hwframe+0x77/0x7f
   </TASK>
  ...
  INFO: task cat:265 is blocked on a mutex likely owned by task sh:246.

wait_for_kprobe_optimizer_locked() reinitializes optimizer_completion,
asks the optimizer thread to flush and sleeps in wait_for_completion().
The thread drains the (un)optimizing lists, but calls complete() only
if completion_done() is true, i.e. if the completion is already done,
which never happens while someone waits. disarm_all_kprobes() and
kprobe_trace_self_tests_init() wait the same way.

Calling complete() unconditionally would not be enough: the waiter
drops kprobe_mutex while it sleeps, and nothing else serializes the
sysctl handler against the debugfs "enabled" file. A second flusher
that still finds the lists non-empty, e.g. because a disabled probe is
queued for unoptimizing, reinitializes the completion under the first:

  sysctl write                      debugfs "enabled" write
  unoptimize_all_kprobes()
    wait_for_kprobe_optimizer_locked()
      init_completion(c)
      mutex_unlock(&kprobe_mutex)
      wait_for_completion(c)
                                    disarm_all_kprobes()
                                      wait_for_kprobe_optimizer_locked()
                                        init_completion(c)
                                          // c->wait is reset, the first
                                          // waiter is off the queue
                                        mutex_unlock(&kprobe_mutex)
                                        wait_for_completion(c)
  kprobe_optimizer()
    complete(c)
      // wakes the debugfs writer only

where c is &optimizer_completion. Lining up the two writes during an
optimizer pass loses the sysctl writer this way.

Replace the completion with a counter of optimizer passes, bumped at the
end of each pass and signalled with wake_up_var_locked(), both under
kprobe_mutex. A flusher samples the count and waits with
wait_var_event_mutex(), which drops kprobe_mutex only while sleeping, so
a new count means a whole pass ran in the meantime. Nothing is
reinitialized, so several flushers can sleep in the wait at once.

Link: https://lore.kernel.org/all/20260924092142.199198-1-parri.andrea@gmail.com/

Fixes: 73c12f2094 ("kprobes: Use dedicated kthread for kprobe optimizer")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Andrea Parri <parri.andrea@gmail.com>
Signed-off-by: Masami Hiramatsu (Google) <mhiramat@kernel.org>
2026-09-25 23:35:23 +09:00
Drif Abdelmalek Mohamed Said
76d8e69724
dcache: unpoison the inline name buffer in __d_alloc()
syzbot reported:

    BUG: KMSAN: uninit-value in dentry_string_cmp fs/dcache.c:291 [inline]
    BUG: KMSAN: uninit-value in dentry_cmp fs/dcache.c:322 [inline]
    BUG: KMSAN: uninit-value in __d_lookup_rcu+0x37d/0x5e0 fs/dcache.c:2522

     dentry_string_cmp fs/dcache.c:291 [inline]
     dentry_cmp fs/dcache.c:322 [inline]
     __d_lookup_rcu+0x37d/0x5e0 fs/dcache.c:2522
     lookup_fast+0x194/0xa40 fs/namei.c:1854
     lookup_fast_for_open fs/namei.c:4545 [inline]
     open_last_lookups fs/namei.c:4579 [inline]
     path_openat+0x9ef/0x6540 fs/namei.c:4856
     do_file_open+0x2aa/0x680 fs/namei.c:4888
     do_sys_openat2+0x17c/0x390 fs/open.c:1395
     do_sys_open fs/open.c:1401 [inline]
     __do_sys_openat fs/open.c:1417 [inline]
     __se_sys_openat fs/open.c:1412 [inline]
     __x64_sys_openat+0x240/0x300 fs/open.c:1412
     x64_sys_call+0x2445/0x3ea0 arch/x86/include/generated/asm/syscalls_64.h:258
     do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
     do_syscall_64+0x15d/0x3c0 arch/x86/entry/syscall_64.c:94
     entry_SYSCALL_64_after_hwframe+0x77/0x7f

Uninit was stored to memory at:
     copy_name fs/dcache.c:3031 [inline]
     __d_move+0xd29/0x21f0 fs/dcache.c:3099
     d_move+0x71/0xf0 fs/dcache.c:3147
     vfs_rename+0x2619/0x2770 fs/namei.c:6085
     filename_renameat2+0xa59/0x1230 fs/namei.c:6188
     __do_sys_rename fs/namei.c:6232 [inline]
     __se_sys_rename+0xc5/0x5c0 fs/namei.c:6228
     __x64_sys_rename+0x78/0xb0 fs/namei.c:6228
     x64_sys_call+0x329/0x3ea0 arch/x86/include/generated/asm/syscalls_64.h:83
     do_syscall_x64 arch/x86/entry/syscall_64.c:63
     do_syscall_64+0x15d/0x3c0 arch/x86/entry/syscall_64.c:94
     entry_SYSCALL_64_after_hwframe+0x77/0x7f

Uninit was created at:
     slab_post_alloc_hook mm/slub.c:4617 [inline]
     slab_alloc_node mm/slub.c:4939 [inline]
     kmem_cache_alloc_lru_noprof+0x376/0x1230 mm/s
     __d_alloc+0x52/0x9f0 fs/dcache.c:1902
     d_alloc+0x57/0x300 fs/dcache.c:1981
     lookup_one_qstr_excl+0x19d/0x7a0 fs/namei.c:1806
     __start_renaming+0x341/0x850 fs/namei.c:3888
     filename_renameat2+0x625/0x1230 fs/namei.c:6163
     __do_sys_rename fs/namei.c:6232 [inline]
     __se_sys_rename+0xc5/0x5c0 fs/namei.c:6228
     __x64_sys_rename+0x78/0xb0 fs/namei.c:6228
     x64_sys_call+0x329/0x3ea0 arch/x86/include/generated/asm/syscalls_64.h:83
     do_syscall_x64 arch/x86/entry/syscall_64.c:63
     do_syscall_64+0x15d/0x3c0 arch/x86/entry/syscall_64.c:94
     entry_SYSCALL_64_after_hwframe+0x77/0x7f

The race is between a concurrent open() and rename() of the same path.

__d_alloc() only stores the name itself and its terminating NUL, so the
rest of the inline buffer (d_shortname, DNAME_INLINE_LEN bytes) is left
uninitialized.  copy_name(), called from rename(), copies that buffer as a
whole, so the uninitialized tail is propagated into the dentry that is
being moved.  Meanwhile __d_lookup_rcu(), called from open(), is an
optimistic lockless lookup: it checks d_name.hash_len first and leaves the
seqcount retry to its caller, so it can end up comparing against a dentry
whose name a rename is rewriting in place, using a stale (longer) length.
The comparison then runs past the terminating NUL and reads bytes of the
uninitialized tail, which KMSAN reports.

The read is harmless by design: it stays inside the buffer, the name is
still NUL-terminated, and the result is thrown away by the seqcount retry.
It is not specific to KMSAN either - with CONFIG_DCACHE_WORD_ACCESS
enabled the very same bytes are read by read_word_at_a_time(), which is
__no_sanitize_or_inline and therefore invisible to KMSAN.  KMSAN builds
only see the instrumented byte-at-a-time dentry_string_cmp() because
CONFIG_DCACHE_WORD_ACCESS is disabled when KMSAN is enabled on x86:

commit 7cf8f44a5a ("x86: fs: kmsan: disable CONFIG_DCACHE_WORD_ACCESS")

Zeroing the inline buffer would hide the report, but it would add a
memset() to a hot allocation path just to initialize bytes that are never
used as part of a name.  Instead, tell KMSAN the inline buffer is
initialized: kmsan_unpoison_memory() compiles to nothing unless
CONFIG_KMSAN is set, and doing it at allocation time is enough for every
dentry, because copy_name() and swap_names() copy the whole buffer and
thus propagate its shadow.

Reported-by: syzbot+7ff3adde89dd795ad4c4@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=7ff3adde89dd795ad4c4
Signed-off-by: Drif Abdelmalek Mohamed Said <drifabdelmalekmohamedsaid@gmail.com>

Changes in v3:
- Annotate for KMSAN instead of zeroing, as suggested in review: the read
  is harmless, so unpoison the inline buffer in __d_alloc() with
  kmsan_unpoison_memory() (a no-op unless CONFIG_KMSAN) rather than adding
  a memset() to the dentry allocation path.
- Document why only KMSAN builds report this at all: with
  CONFIG_DCACHE_WORD_ACCESS the same read goes through
  read_word_at_a_time(), which KMSAN does not instrument.
- Rewrite the commit message; the previous one had several truncated
  lines.

Link: https://patch.msgid.link/20260918224204.3056-1-drifabdelmalekmohamedsaid@gmail.com
Reviewed-by: Jan Kara <jack@suse.cz>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-09-25 15:54:23 +02:00
Amir Goldstein
ae146bc1ab
ovl: fix UAF in ovl_do_mkdir() debug print
ovl_do_mkdir() prints the input dentry with %pd after vfs_mkdir().
Since commit fe497f0759 ("VFS: change vfs_mkdir() to unlock on
failure."), vfs_mkdir() calls end_creating() on the input dentry on
failure and may replace it on success, so the post-call %pd can
use-after-free the dentry when CONFIG_OVERLAY_FS_DEBUG is enabled.

Print the dentry before the call and only the result afterward.

Reported-by: syzbot+ced26b784bf977d223dd@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=ced26b784bf977d223dd
Fixes: fe497f0759 ("VFS: change vfs_mkdir() to unlock on failure.")
Signed-off-by: Amir Goldstein <amir73il@gmail.com>
Link: https://patch.msgid.link/20260921104013.40475-1-amir73il@gmail.com
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-09-25 15:52:26 +02:00
Matthias Goergens
113dcdfadf MAINTAINERS: name the libata/linux for-next branch
The T: entry for LIBATA SUBSYSTEM (Serial and Parallel ATA drivers)
names libata/linux without a branch.  The repository's HEAD pointer
points to branch master, which has no active development.  Active
development is on the for-next branch.  Name the branch so the entry
identifies where development happens.

Documentation/process/submitting-patches.rst sends contributors to the
T: entry to find the tree to prepare patches against, so a branch-less
entry whose HEAD is already in mainline points them to the wrong branch.

Reviewed-by: Damien Le Moal <dlemoal@kernel.org>
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
Link: https://lore.kernel.org/r/20260925052329.2683619-1-matthias.goergens@gmail.com
Signed-off-by: Niklas Cassel <cassel@kernel.org>
2026-09-25 11:08:37 +02:00
Dave Airlie
0f50dab8b4 amd-drm-fixes-7.3-2026-09-24:
amdgpu:
 - Display ref count fix
 - Userq fixes
 - VCN 4, 5 reset fixes
 - Fixes for various error paths
 - Stack frame size fixes for various combinations of compilers and configs
 
 amdkfd:
 - Possible UAF fix
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQQgO5Idg2tXNTSZAr293/aFa7yZ2AUCarVdBwAKCRC93/aFa7yZ
 2IiRAP9QnDZ0lLR36p0yPHlMVl2opdMfkqeLkdHd6BgDn/3/mwEA780cxI7gxbFH
 LS7QLSW9L68xjdAixoUWi/UYCIG4sAQ=
 =VbtV
 -----END PGP SIGNATURE-----

Merge tag 'amd-drm-fixes-7.3-2026-09-24' of https://gitlab.freedesktop.org/drm/amdgpu/kernel into drm-fixes

amd-drm-fixes-7.3-2026-09-24:

amdgpu:
- Display ref count fix
- Userq fixes
- VCN 4, 5 reset fixes
- Fixes for various error paths
- Stack frame size fixes for various combinations of compilers and configs

amdkfd:
- Possible UAF fix

Signed-off-by: Dave Airlie <airlied@redhat.com>

From: Alex Deucher <alexander.deucher@amd.com>
Link: https://patch.msgid.link/20260924172938.634777-1-alexander.deucher@amd.com
2026-09-25 18:04:49 +10:00
Dave Airlie
c5d1690ff1 Fixes in:
- CRI throttle reasons report (Sk)
  - TLB invalidation at wedge (Shuicheng)
  - SVM eviction and VM close (Brost)
  - Display corruption on LNL on Xen PV (Szymon)
  - W/a fix and addition (Tilak)
 -----BEGIN PGP SIGNATURE-----
 
 iQEzBAABCgAdFiEEbSBwaO7dZQkcLOKj+mJfZA7rE8oFAmq1KCQACgkQ+mJfZA7r
 E8pjZAgAoGb4wVoZ9Qm1qCWRXiUJWoD63X9/s0cYe9qDoStsfJ3oppKuuy6A2lZQ
 SvaTLExbRLtYJiPU0RWQcphiay5lnjQzkjIrGMt8cMA/9fVPgdXfL13PXeJuJ4yM
 uQzv9+eWuptjyHxNbCuW0pEnxWZ8yyR/lMbyX68+hdJl7TgSDWHEllTkv8+xuNPn
 APyIizk69bI3VPzBcl0GcoW3VOtabgUzZiWkuuehuVf+Bl2vwZElRScjLq5gNZ3x
 IzNOQPE1z6d0v2drK/bk7SzfEt41o/J9RqKnoBguhW66XFbRU3/sVn924yCpThHQ
 74e9Wa2DP2pdJZ6htJcnrNRl9ZNi9A==
 =cXRT
 -----END PGP SIGNATURE-----

Merge tag 'drm-xe-fixes-2026-09-24' of https://gitlab.freedesktop.org/drm/xe/kernel into drm-fixes

Fixes in:
 - CRI throttle reasons report (Sk)
 - TLB invalidation at wedge (Shuicheng)
 - SVM eviction and VM close (Brost)
 - Display corruption on LNL on Xen PV (Szymon)
 - W/a fix and addition (Tilak)

Signed-off-by: Dave Airlie <airlied@redhat.com>

From: Rodrigo Vivi <rodrigo.vivi@intel.com>
Link: https://patch.msgid.link/arUoUf9LpsJpJouN@intel.com
2026-09-25 16:50:21 +10:00
Linus Torvalds
165768bb70 firewire fixes for 7.3-rc5
Fix a race in the cdev layer that can cause a fw_iso_resource_auto object
 to transition back to a previous state. This can happen when a file
 descriptor is closed while the work item for the object is running. The
 race can leak several memory objects, including client object itself.
 This fix should be applied to 7.2 kernel or later.
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQQE66IEYNDXNBPeGKSsLtaWM8LwEwUCarWseAAKCRCsLtaWM8Lw
 E2x/AQCr6GAMO1G8/7mUgEj3X4CHCzkS/N3bVwb6TJ18/e5gxgEAuM9Ry+FU6IP4
 0NCz6dbcn9dta3YGnlfgAQ9HtTMwnAg=
 =BLzx
 -----END PGP SIGNATURE-----

Merge tag 'firewire-fixes-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/ieee1394/linux1394

Pull firewire fix from Takashi Sakamoto:
 "Fix a race in the cdev layer that can cause a fw_iso_resource_auto
  object to transition back to a previous state. This can happen when a
  file descriptor is closed while the work item for the object is
  running. The race can leak several memory objects, including client
  object itself"

* tag 'firewire-fixes-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/ieee1394/linux1394:
  firewire: cdev: fix back-transition for iso_resource_auto client resource
2026-09-24 17:09:56 -07:00
Dave Airlie
fd0ba310f9 Merge tag 'drm-intel-fixes-2026-09-24' of https://gitlab.freedesktop.org/drm/i915/kernel into drm-fixes
drm/i915 fixes for v7.3-rc5:
- Fix incorrect RCU teardown order leading to endless loop
- Fix DP MST TU and FEC handling for disconnected streams
- Fix selective fetch disable, again
- Fix export namespace for kunit helpers
- Workaround eDP flicker on a specific laptop model

Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Jani Nikula <jani.nikula@intel.com>
Link: https://patch.msgid.link/9b8ff50ad5c3ff9ca120207b1e7aa95d5a16423f@intel.com
2026-09-25 09:47:46 +10:00
Linus Torvalds
ee9c669f9b sched_ext: Fixes for v7.3-rc4
- A task reenqueued while its dispatch was still completing had its queued
   state clobbered by the dispatcher, dropping every later dispatch of the
   task. Wait for the in-flight dispatch to settle first.
 
 - A wakeup activation on another CPU marked the destination runqueue as
   mid-wakeup, stranding a pending local reenqueue. If the scheduler was
   unloaded first, the stale request pointed into freed memory that the next
   scheduler dereferenced.
 
 - ops.dequeue() ran with the source dispatch queue's lock held, so a
   scheduler iterating that queue from the callback deadlocked the CPU.
 
 - Schedulers with their own CPU ID mapping had no way to learn a task's
   initial CPU mask and rebuilt it themselves, which went wrong across
   sub-scheduler enable and re-home. Pass it to ops.enable().
 
 - A bypass dispatch event counter missed the dispatches made by the
   end-of-dispatch fallback and under-reported.
 
 - Selftests for the dequeue locking and initial mask changes.
 -----BEGIN PGP SIGNATURE-----
 
 iIQEABYKACwWIQTfIjM1kS57o3GsC/uxYfJx3gVYGQUCarWSLQ4cdGpAa2VybmVs
 Lm9yZwAKCRCxYfJx3gVYGY+sAP9y6nJh6vIvFh/X9FlJtWlNo0mncOKhy93E8jii
 8CKnPQEAhvX3+Gcdl+imTh4Z915kdsEByBjTTPPOXnQxI8BKUAY=
 =VhBF
 -----END PGP SIGNATURE-----

Merge tag 'sched_ext-for-7.3-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext

Pull sched_ext fixes from Tejun Heo:

 - A task reenqueued while its dispatch was still completing had its
   queued state clobbered by the dispatcher, dropping every later
   dispatch of the task. Wait for the in-flight dispatch to settle
   first

 - A wakeup activation on another CPU marked the destination runqueue as
   mid-wakeup, stranding a pending local reenqueue. If the scheduler was
   unloaded first, the stale request pointed into freed memory that the
   next scheduler dereferenced

 - ops.dequeue() ran with the source dispatch queue's lock held, so a
   scheduler iterating that queue from the callback deadlocked the CPU

 - Schedulers with their own CPU ID mapping had no way to learn a task's
   initial CPU mask and rebuilt it themselves, which went wrong across
   sub-scheduler enable and re-home. Pass it to ops.enable()

 - A bypass dispatch event counter missed the dispatches made by the
   end-of-dispatch fallback and under-reported

 - Selftests for the dequeue locking and initial mask changes

* tag 'sched_ext-for-7.3-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext:
  sched_ext: Count SCX_EV_SUB_BYPASS_DISPATCH in the dispatch fallback
  selftests/sched_ext: Check the cmask cid-form ops.enable() receives
  sched_ext: Pass the initial cmask to cid-form ops.enable()
  selftests/sched_ext: Test that ops.dequeue() can iterate the consumed DSQ
  sched_ext: Don't run ops.dequeue() with a DSQ lock held
  sched_ext: Derive SCX_RQ_IN_WAKEUP from the core enqueue flags
  sched_ext: Wait for SCX_OPSS_DISPATCHING before reenqueueing a task
2026-09-24 15:04:11 -07:00
Linus Torvalds
e8dfd03a1c cgroup: Fixes for v7.3-rc4
- With local event accounting, a fork rejected by the pids controller
   updated pids.events without notifying its pollers.
 
 - A cgroup selftest failed to compile with fortification enabled because
   an O_TMPFILE open lacked its mode argument.
 -----BEGIN PGP SIGNATURE-----
 
 iIQEABYKACwWIQTfIjM1kS57o3GsC/uxYfJx3gVYGQUCarWSKg4cdGpAa2VybmVs
 Lm9yZwAKCRCxYfJx3gVYGYoiAQD6JUCqDjv2Hr1YeMFlsoYVQZ7tNmljVnPD2tZ6
 9WmD/AEAotdlmzM8egOuDqAi2s+UMJPm8vCZuxVzewiGqhYELgk=
 =0Ljp
 -----END PGP SIGNATURE-----

Merge tag 'cgroup-for-7.3-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup

Pull cgroup fixes from Tejun Heo:

 - With local event accounting, a fork rejected by the pids controller
   updated pids.events without notifying its pollers

 - A cgroup selftest failed to compile with fortification enabled
   because an O_TMPFILE open lacked its mode argument

* tag 'cgroup-for-7.3-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup:
  cgroup/pids: Restore pids.events notifications in local mode
  selftests: cgroup: give the O_TMPFILE open in get_temp_fd() a mode
2026-09-24 14:26:12 -07:00
Linus Torvalds
f2c53ea949 Including fixes from Bluetooth, NFC and Netfilter.
Every week in this release is record-setting for number of posted
 patches. It doesn't seem like we're creating any regressions with
 all these fixes, 3 Fixes tags here point to 7.2 commits but none
 are true regression fixes. We're trying to keep the count down,
 nonetheless.
 
 Previous releases - regressions:
 
  - net: don't require the hwtstamp NDOs when a PHY provides timestamping
 
  - ipv6: fix dst leak for uncached routes
 
  - vrf: stop corrupting skb->csum when capturing CHECKSUM_COMPLETE
    packets
 
 Previous releases - always broken:
 
  - packet: use ubuf_info completion for TX_RING packets
 
  - arp: terminate device name before lookup
 
  - ipv6: do not let ipv6_find_hdr() return an offset past the packet end
 
  - udp: remove a disconnected socket from the 4-tuple hash table
 
  - sctp: discard the rest of the packet on a stale-cookie error
 
  - eth: mlx5: Bridge, fix remaining switchdev ownership gaps on merged
    eswitch
 
 Signed-off-by: Jakub Kicinski <kuba@kernel.org>
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEE6jPA+I1ugmIBA4hXMUZtbf5SIrsFAmq1bCQACgkQMUZtbf5S
 IrsgCBAAiHOxV3pR0pNw+K/dO5maYhuUV4w+Xcd6CbjRnxzYtImVC1q2ieez5dEe
 JrQrVK76t2wuEgGSyiWcckIufPL8xYg5nYaOUbkYMC2QRtDaqhCo+u2jXu26/Lr8
 O3sI2cj86XDZN+xKTvmLld1zIRyD6MhYxWz+gIicArdBqTsr5JsWkwhJE7SH013a
 iVGu0sIxTgCNu/AusD6SdtwFte26dqvadelI8yuDhBR6rM9pi4EWXChleg7ohCsu
 DoZTlaEjEiFavat+6ni99oz8MAT0WVBX0IzTqZa6TCnbCw4qUEbqN8G+CNgC/VPe
 Sq5FRoo6vJjlFviwgWPbpje2FoOYtk6CX1eTavx8gZ4pPMTzGkT62crS+ETwLO5H
 mRbE+ZjG8UDTUKpIY5P/IQYdznHSc9Ny/pHNlZhFmDdntJOp1lnegx5CmZkyuZxC
 Qc0hoFxKdjUrpQ1n0ygT3/P6MT9vXMwbGP52VLIxaB1oaYCKFV71NOp0ET5fCnYb
 CBK8cdN0kcaSLrpi3MnPamyQoNZfH78BiuKsLDu/TWg3i2HGmFTlp0/LJ0nkQcCU
 4KlCmiJNOkHjbQKHNtYMUOZQcZu5witEv33kb7gr581b0RR+n0lZTaXcnRIyXYdb
 ORqIa1Khg5mW3DLaIc9IrzEGGAW5zXUHc5OX7t8J2rVRQWhBEh4=
 =YiiK
 -----END PGP SIGNATURE-----

Merge tag 'net-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net

Pull networking fixes from Jakub Kicinski:
 "Including fixes from Bluetooth, NFC and Netfilter.

  Every week in this release is record-setting for number of posted
  patches. It doesn't seem like we're creating any regressions with all
  these fixes, three 'Fixes' tags here point to 7.2 commits but none are
  true regression fixes. We're trying to keep the count down,
  nonetheless.

  Previous releases - regressions:

   - net: don't require the hwtstamp NDOs when a PHY provides
     timestamping

   - ipv6: fix dst leak for uncached routes

   - vrf: stop corrupting skb->csum when capturing CHECKSUM_COMPLETE
     packets

  Previous releases - always broken:

   - packet: use ubuf_info completion for TX_RING packets

   - arp: terminate device name before lookup

   - ipv6: do not let ipv6_find_hdr() return an offset past the packet
     end

   - udp: remove a disconnected socket from the 4-tuple hash table

   - sctp: discard the rest of the packet on a stale-cookie error

   - eth: mlx5: Bridge, fix remaining switchdev ownership gaps on merged
     eswitch"

[ And lots of other random network driver fixes ]

* tag 'net-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (189 commits)
  tcp: prevent collapsing skbs across boundary in rtx queue
  vlan: ensure sufficient headroom in vlan_dev_hard_header()
  net/sched: sch_teql: fix shadowed err in __teql_resolve()
  bridge: check llc_mac_hdr_init() return value in br_send_bpdu()
  llc: fix skb UAF and leaks on llc_mac_hdr_init() failure
  llc: reserve device headroom for allocated frames
  gve: DQO: reject TSO packets with an out of range MSS
  gve: fix TX drop when GSO MSS is too small for hw
  gve: DQO: fix header length used by gve_can_send_tso() for UDP GSO
  net: flush skb_defer_nodes in dev_cpu_dead()
  net: ethernet: stmmac: dwmac-rk: fix bulk clock leak when the PHY clock fails
  af_packet: fix integer overflow in prb_calc_retire_blk_tmo()
  tipc: Fix a data race on mon->peer_cnt in mon_timeout()
  net: phy: intel-xway: workaround 100BASE-TX Link-Up issue
  net/smc: fix UAF on lgr list traversal in smcr_port_err()
  net/rds: size a connection's path set by the transport it ends up with
  nfp: hold IPsec RX state under the XArray lock
  net: ena: fix MMIO read buffer leak on probe failure
  net: ena: fix PHC cleanup on probe failure
  net/sched: act_ct: fix helper UAF due to extensions realloc
  ...
2026-09-24 11:45:33 -07:00
Willem de Bruijn
fc6d80eb50 tcp: prevent collapsing skbs across boundary in rtx queue
tcp_write_collapse_fence() sets TCP_SKB_CB(skb)->eor = 1 on
tcp_write_queue_tail(sk) to prevent skbs queued after a switch to
device encryption from being collapsed into earlier skbs.

The fence is a no-op if all earlier data has already been transmitted
when the switch happens: sk->sk_write_queue is empty. The not yet
acknowledged earlier skbs wait in sk->tcp_rtx_queue with eor 0.

On a subsequent retransmit or SACK shift, tcp_retrans_try_collapse() or
tcp_shift_skb_data() can then merge an skb queued after the switch into
one queued before it.

Both users of the fence are affected:

- psp: devices only encrypt skbs with skb->decrypted set. The merged skb
  keeps decrypted = 0 from the earlier skb, so merged data sent after
  psp_sock_assoc_set_tx() is retransmitted in cleartext.

- tls device offload: the merged skb straddles the start marker set in
  tls_set_device_offload(). The software fallback (fill_sg_in() returns
  -EINVAL) and the mlx5, nfp and funeth drivers cannot handle such an
  skb and drop it. Every retransmit rebuilds the same skb, so the
  connection stalls.

Fix this in two places, for defense in depth:

1. Fall back to tcp_rtx_queue_tail(sk) in tcp_write_collapse_fence()
   when tcp_write_queue_tail(sk) is NULL.

2. Check !skb_cmp_decrypted(to, from) in tcp_skb_can_collapse(), as
   tcp_skb_can_collapse_rx() does on receive. skb_shift(), which both
   collapse paths call, already has a DEBUG_NET_WARN_ON_ONCE() for this
   condition.

Fixes: e8f6979981 ("net/tls: Add generic NIC offload infrastructure")
Cc: stable@vger.kernel.org
Signed-off-by: Willem de Bruijn <willemb@google.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Daniel Zahka <daniel.zahka@gmail.com>
Link: https://patch.msgid.link/20260924154427.953800-1-willemdebruijn.kernel@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 11:03:42 -07:00
Jakub Kicinski
1078a38344 Merge branch 'vlan-ensure-sufficient-headroom-in-vlan_dev_hard_header'
Eric Dumazet says:

====================
vlan: ensure sufficient headroom in vlan_dev_hard_header()

Callers that only reserve ETH_HLEN or less, or skbs allocated before
dynamic device/headroom changes (such as toggling VLAN_FLAG_REORDER_HDR
or bonding/team switching slaves), can reach vlan_dev_hard_header() with
insufficient headroom and trigger skb_under_panic().

When vlan_dev_hard_header() returns -ENOMEM upon skb_cow_head() failure,
a few callers of dev_hard_header() / llc_mac_hdr_init() had pre-existing
error-handling bugs:
====================

Link: https://patch.msgid.link/20260924082951.1599377-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 11:01:04 -07:00
Eric Dumazet
cd5dd68267 vlan: ensure sufficient headroom in vlan_dev_hard_header()
Callers that only reserve ETH_HLEN or less (such as llc_alloc_frame()),
or skbs allocated before dynamic device/headroom changes (e.g. toggling
VLAN_FLAG_REORDER_HDR or bonding/team switching slaves), can reach
vlan_dev_hard_header() with insufficient headroom and trigger
skb_under_panic().

Use skb_cow_head() in vlan_dev_hard_header() when VLAN_FLAG_REORDER_HDR
is not set to ensure sufficient headroom for the VLAN header(s) and the
underlying device hard header.

Use READ_ONCE() to read dev->hard_header_len and dev->needed_headroom as
they can be updated concurrently under RTNL (e.g. in
vlan_transfer_features()) while vlan_dev_hard_header() runs locklessly on
the transmit path. Also avoid LL_RESERVED_SPACE(dev) here so that the
extra HH_DATA_MOD alignment padding does not trigger unnecessary
pskb_expand_head() reallocations on inner stacked VLAN devices after the
outer VLAN header has been pushed.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Reported-by: Zixuan Chai <petalzu987@gmail.com>
Closes: https://lore.kernel.org/netdev/cover.1789987105.git.petalzu987@gmail.com/
Link: https://lore.kernel.org/netdev/179022851638.2160803.1808206741379444999@kernel.org/
Cc: Hangbin Liu <liuhangbin@gmail.com>
Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260924082951.1599377-5-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 11:00:59 -07:00
Eric Dumazet
907b978e82 net/sched: sch_teql: fix shadowed err in __teql_resolve()
__teql_resolve() declares an inner 'int err;' inside the
'if (neigh_event_send(n, skb_res) == 0)' block, shadowing the outer
'int err = 0;'. As a result, a negative return from dev_hard_header()
is written to the inner variable and __teql_resolve() still returns 0.

Remove the shadowed variable and set the outer err to -EINVAL when
dev_hard_header() returns a negative error.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Closes: https://lore.kernel.org/netdev/179022851638.2160803.1808206741379444999@kernel.org/
Cc: Jamal Hadi Salim <jhs@mojatatu.com>
Cc: Jiri Pirko <jiri@resnulli.us>
Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260924082951.1599377-4-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 11:00:59 -07:00
Eric Dumazet
ac704ff08e bridge: check llc_mac_hdr_init() return value in br_send_bpdu()
If llc_mac_hdr_init() fails (for instance if the port device type does
not support LLC or dev_hard_header() fails), br_send_bpdu() should drop
the skb instead of resetting the mac header to the LLC payload and
transmitting a malformed frame.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Closes: https://lore.kernel.org/netdev/179022851638.2160803.1808206741379444999@kernel.org/
Cc: Nikolay Aleksandrov <razor@blackwall.org>
Cc: Ido Schimmel <idosch@nvidia.com>
Cc: bridge@lists.linux.dev
Signed-off-by: Eric Dumazet <edumazet@google.com>
Acked-by: Nikolay Aleksandrov <razor@blackwall.org>
Link: https://patch.msgid.link/20260924082951.1599377-3-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 11:00:59 -07:00
Eric Dumazet
72f9dd522f llc: fix skb UAF and leaks on llc_mac_hdr_init() failure
In llc_conn_ac_resend_i_xxx_x_set_0_or_send_rr(), if llc_mac_hdr_init()
fails, kfree_skb(skb) is called instead of kfree_skb(nskb). This leaks
the newly allocated nskb, reads from the freed skb via LLC_I_GET_NR(pdu),
and double-frees skb when llc_conn_state_process() drops its reference.

In llc_sap_action_send_xid_r() and llc_sap_action_send_test_r(), nskb is
leaked if llc_mac_hdr_init() returns an error.

Free nskb in all three error paths.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Closes: https://lore.kernel.org/netdev/179022851638.2160803.1808206741379444999@kernel.org/
Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260924082951.1599377-2-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 11:00:59 -07:00
Zixuan Chai
72b5b9a28b llc: reserve device headroom for allocated frames
llc_alloc_frame() reserves link-layer headroom using the device type.
This is insufficient for stacked Ethernet devices such as VLAN devices,
where vlan_dev_hard_header() pushes a VLAN header before the lower
device's Ethernet header. An LLC response on such a device can
therefore underflow skb headroom in eth_header().

Use LL_RESERVED_SPACE() to account for the device's actual required
headroom while preserving the existing LLC device-type check.

Fixes: bf9ae5386b ("llc: use dev_hard_header")
Cc: stable@vger.kernel.org
Reported-by: VEGA <vega@nebusec.ai>
Signed-off-by: Zixuan Chai <petalzu987@gmail.com>
Signed-off-by: Ren Wei <weir@nebusec.ai>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260924012613.2533934-1-weir@nebusec.ai
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 10:58:16 -07:00
Jakub Kicinski
488089055b Merge branch 'gve-dqo-fix-handling-of-out-of-range-tso-mss'
Eric Dumazet says:

====================
gve: DQO: fix handling of out of range TSO MSS

The DQO TX path assumes that the MSS of a TSO packet is within the
range supported by the device, [88, 9728].

This holds for locally generated traffic, but not for packets coming
from a tap or from a packet socket: virtio_net_hdr_to_skb() takes
gso_size from user space and only enforces a minimum, layer 2
forwarding does not check the MTU of GSO packets, and
gso_features_check() bounds skb->len and gso_segs but never gso_size.

Patch 1, from Eddie Phillips, deals with the lower bound. It moves the
existing test out of gve_prep_tso() into gve_features_check_dqo(), so
that these packets are segmented in software instead of being dropped.

Patch 2 deals with the upper bound, which is currently not checked at
all. gve_tx_fill_tso_ctx_desc() stores gso_size into a 14 bits wide
field, so that an MSS of 16384 silently becomes zero. Falling back to
software segmentation is not an option here, because skb_segment()
splits at gso_size regardless of the MTU, and would only replace an
invalid TSO packet by non TSO packets larger than the 9728 bytes the
device supports. These packets are dropped instead.

As noted in patch 2, oversized non TSO packets can still reach the
device whenever the stack segments in software. This is not specific
to gve and is better fixed in the core, so a patch for
__is_skb_forwardable() will be sent separately for net-next.
====================

Link: https://patch.msgid.link/20260924004252.1196328-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 10:57:01 -07:00
Eric Dumazet
296c83b5cc gve: DQO: reject TSO packets with an out of range MSS
gve_prep_tso() notes that the device requires the MSS to be <= 9728,
but does not enforce it, assuming the 9K MTU enforced by the hypervisor
and the 64KB limit on TSO sizes are enough.

This does not hold for packets that were not generated locally.
A guest behind a tap, or any packet socket user, can provide an
arbitrary gso_size in virtio_net_hdr. Layer 2 forwarding does not check
the MTU for GSO packets (is_skb_forwardable()), and gso_features_check()
only bounds skb->len and gso_segs, never gso_size.

Such a packet reaches gve_tx_fill_tso_ctx_desc(), which puts gso_size
into the mss field of the TSO context descriptor. This field is 14 bits
wide, so a gso_size of 16384 is silently turned into an MSS of zero.

Drop these packets from gve_prep_tso(), and make sure that
gve_features_check_dqo() leaves their GSO bits alone: skb_segment()
splits at gso_size regardless of the MTU, so falling back to software
segmentation would give the device non TSO packets bigger than the
9728 bytes it supports.

Note that the device can still be given oversized non TSO packets when
the stack segments in software for other reasons, for instance after
TSO has been disabled with ethtool. This is a generic issue, because
the MTU check is skipped for GSO packets in the forwarding path, and
is addressed separately.

Fixes: a57e5de476 ("gve: DQO: Add TX path")
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Harshitha Ramamurthy <hramamurthy@google.com>
Link: https://patch.msgid.link/20260924004252.1196328-3-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 10:56:59 -07:00
Eddie Phillips
3b430ea623 gve: fix TX drop when GSO MSS is too small for hw
The device has a strict requirement that the minimum MSS
(gso_size) for TSO/GSO packets must be at least 88 bytes. If a packet
below this threshold is pushed to the hardware, it can cause
hardware to silently drop the packet, leading to increased latency
and retransmissions.

Currently, this is validated too late in the transmit pipeline
(gve_prep_tso), leading to silent drops.

Fix this by moving the validation into the .ndo_features_check
callback (gve_features_check_dqo). If we detect a GSO packet with
a gso_size smaller than GVE_TX_MIN_TSO_MSS_DQO, we clear the GSO
feature flags for this packet.

Fixes: a57e5de476 ("gve: DQO: Add TX path")
Signed-off-by: Eddie Phillips <eddiephillips@google.com>
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Harshitha Ramamurthy <hramamurthy@google.com>
Link: https://patch.msgid.link/20260924004252.1196328-2-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 10:56:59 -07:00
Linus Torvalds
415f204422 Landlock fix for v7.3-rc5
-----BEGIN PGP SIGNATURE-----
 
 iIYEABYKAC4WIQSVyBthFV4iTW/VU1/l49DojIL20gUCarVI0xAcbWljQGRpZ2lr
 b2QubmV0AAoJEOXj0OiMgvbSEVgA+gNbC9CVCbCo0oufZbVpQwlwtuuEKEVpvxZx
 q7oFYKCsAP9svCujGCXRHOmWhAAwe+wpXNb43l8coFn+pCfA8x3fCw==
 =NOli
 -----END PGP SIGNATURE-----

Merge tag 'landlock-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux

Pull Landlock fixes from Mickaël Salaün:
 "This mainly fixes the Landlock tracepoint support merged this cycle so
  that denial and rule events report the intended policy context,
  whether through tracefs or BTF-visible callbacks.

  The size of this all is mainly from propagating the corrected contract
  through event definitions and producers, adding new tests for the
  reported context, and updating the documentation.

  Also improve annotation and fix a GCC 16 build warning"

* tag 'landlock-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux:
  landlock: Widen ruleset versions to 64 bits
  landlock: Add counted_by in landlock_domain
  landlock: Fix tracepoint contract documentation
  selftests/landlock: Test network denial context
  selftests/landlock: Test filesystem denial blockers
  landlock: Report the effective signal number
  landlock: Report the actual ptrace tracer
  landlock: Fix network denial trace context
  landlock: Fix rule tracepoint context
  landlock: Fix filesystem denial blocker reporting
  landlock: Fix tracepoint fixed-width type names
  landlock: Work around gcc-16 -Wuninitialized warning
2026-09-24 10:53:59 -07:00
Eric Dumazet
83769c23fb gve: DQO: fix header length used by gve_can_send_tso() for UDP GSO
gve_can_send_tso() computes how many buffers each segment of a GSO
packet would span, and for this it needs the length of the headers
that the device replicates in front of every segment.

It unconditionally uses skb_tcp_all_headers(), which reads the doff
field of the TCP header. SKB_GSO_UDP_L4 packets have no TCP header:
tcp_hdrlen() then reads one byte of the UDP payload, and header_len
can be anything in [0, 60] instead of the transport offset plus the
eight bytes of the UDP header that gve_prep_tso() programs into the
TSO context descriptor.

A wrong header length shifts all the segment boundaries computed in
the loop, so the number of buffers per segment can be over or under
estimated. In the first case, GSO is needlessly disabled for this
packet by gve_features_check_dqo() and the stack has to segment it.
In the second case, the driver hands the device a packet whose
segments span more than GVE_TX_MAX_DATA_DESCS buffers.

Use the UDP header length for SKB_GSO_UDP_L4 packets, matching what
gve_prep_tso() does.

Fixes: 014c607f86 ("gve: add support for UDP GSO for DQO format")
Closes: https://lore.kernel.org/netdev/CANn89i+MS4L60sFQ49=-f-mibeveUfcrpVkD5X+Qy6SOnEpd6w@mail.gmail.com/
Signed-off-by: Eric Dumazet <edumazet@google.com>
Cc: Ankit Garg <nktgrg@google.com>
Cc: Harshitha Ramamurthy <hramamurthy@google.com>
Cc: Joshua Washington <joshwash@google.com>
Cc: Willem de Bruijn <willemb@google.com>
Reviewed-by: Ankit Garg <nktgrg@google.com>
Reviewed-by: Harshitha Ramamurthy <hramamurthy@google.com>
Link: https://patch.msgid.link/20260923145942.731365-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 10:51:29 -07:00
Eric Dumazet
06e3f54e8b net: flush skb_defer_nodes in dev_cpu_dead()
When a CPU goes offline, dev_cpu_dead() drains its softnet queues
(completion_queue, output_queue, poll_list, process_queue, and
input_pkt_queue), but leaves net_hotdata.skb_defer_nodes untouched.

If oldcpu goes offline while holding pending skbs in its
skb_defer_nodes lists (e.g. below the sysctl_skb_defer_max >> 1 IPI
threshold, or if the IPI races with CPU teardown), those skbs remain
stranded until oldcpu is brought back online. If any of these skbs
hold page_pool fragments, page_pool_destroy() will stall indefinitely
waiting for inflight pages to be returned when a netdev or driver is
torn down while oldcpu is offline.

Additionally, if smp_call_function_single_async() fails in
kick_defer_list_purge() because the target CPU went offline, reset
defer_ipi_scheduled to 0 so future IPI kicks are not blocked when the
CPU comes back online.

Also, if oldcpu was the last online CPU on its NUMA node, drain that
node's slot across all CPUs so no skbs deferred from that node remain
stranded on idle remote CPUs (or if the node itself is subsequently
offlined).

Finally, in skb_attempt_defer_free(), re-check cpu_online(cpu) and
whether the caller migrated CPUs after llist_add(), flushing the node
list if so, to close the preemption TOCTOU race against CPU/node
teardown.

Fixes: 68822bdf76 ("net: generalize skb freeing deferral to per-cpu lists")
Fixes: 5628f3fe3b ("net: add NUMA awareness to skb_attempt_defer_free()")
Closes: https://lore.kernel.org/netdev/20260916003430.3612956-1-kris.pan@intel.com/
Signed-off-by: Eric Dumazet <edumazet@google.com>
Cc: Kris Pan <kris.pan@intel.com>
Link: https://patch.msgid.link/20260923130318.607255-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 10:48:10 -07:00
Coia Prant
8db67bb6a1 net: ethernet: stmmac: dwmac-rk: fix bulk clock leak when the PHY clock fails
gmac_clk_enable() enables the bulk clocks first and then the optional
PHY clock. If clk_prepare_enable() on the PHY clock fails, the function
returns without rolling back the bulk clocks, and bsp_priv->clk_enabled
stays false, so the later gmac_clk_enable(bsp_priv, false) becomes a
no-op and the bulk clock references are leaked.

Add the missing clk_bulk_disable_unprepare() on that failure path.

Fixes: ea449f7fa0 ("net: ethernet: stmmac: dwmac-rk: rework optional clock handling")
Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Reviewed-by: Heiko Stuebner <heiko@sntech.de>
Acked-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Signed-off-by: Coia Prant <coiaprant@gmail.com>
Link: https://patch.msgid.link/20260923123713.3137146-1-coiaprant@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 10:47:24 -07:00
Dairui Zhang
56d82862a0 af_packet: fix integer overflow in prb_calc_retire_blk_tmo()
prb_calc_retire_blk_tmo() computes in 32-bit int arithmetic:

        mbits = (blk_size_in_bytes * 8) / (1024 * 1024);

If I'm reading the validation right, tp_block_size is user
controlled and packet_set_ring() only rejects values that are <= 0
as int or not page aligned, so a 256MiB block goes right through
(and alloc_one_pg_vec_page() even has a vzalloc fallback for it).
0x10000000 * 8 wraps to INT_MIN, and on a NIC reporting 1 Gbps
(div == 1) the function ends up returning -2047.

The condition is actually (8 * size) mod 2^32 >= 2^31 && div == 1,
so the trigger set is [256,512), [768,1024), [1280,1536) and
[1792,2048) MiB. Other sizes wrap to non-negative values and faster
links divide the unsigned value back below 2^31, which is why this
doesn't blow up for everyone.

What makes it fatal is what happens next in init_prb_bdqc():

        p1->interval_ktime = ms_to_ktime(prb_calc_retire_blk_tmo(...));
        hrtimer_start(&p1->retire_blk_timer, p1->interval_ktime,
                      HRTIMER_MODE_REL_SOFT);

A negative relative timeout expires immediately. The callback
unconditionally returns HRTIMER_RESTART, and hrtimer_forward() turns
the negative interval into hrtimer_resolution:

        if (interval < hrtimer_resolution)
                interval = hrtimer_resolution;

So the SOFT timer re-fires at the maximum rate forever, holding
sk_receive_queue.lock each pass. One CPU spins in softirq until the
socket is closed. Repeat with more rings and the machine is gone.

The overflow itself is ancient - it was introduced together with
TPACKET_V3 in f6fb8f100b ("af-packet: TPACKET_V3 flexible buffer
implementation."). Its effect prior to f7460d2989 ("net:
af_packet: Use hrtimer to do the retire operation", v6.18) was not
as clear-cut, though: the return value was stored into an unsigned
short retire_blk_tov, so a negative result was truncated, and a
0-jiffy delay loop could be programmed as well. Neither is nearly
as detrimental as the immediate maximum-rate spin the hrtimer
conversion turned it into.

(Unrelated to CVE-2019-20812 - that one was the ethtool failure path
returning 0, which now returns DEFAULT_PRB_RETIRE_TOV.)

Reproducer, needs CAP_NET_RAW (a --network host container has it by
default) and a 1 Gbps NIC (QEMU e1000 works):

        int fd = socket(AF_PACKET, SOCK_RAW, htons(ETH_P_ALL));
        bind(fd, ...);
        int v = TPACKET_V3;
        setsockopt(fd, SOL_PACKET, PACKET_VERSION, &v, sizeof(v));
        struct tpacket_req3 req = {
                .tp_block_size = 0x10000000,
                .tp_block_nr = 1,
                .tp_frame_size = 2048,
                .tp_frame_nr = 0x10000000 / 2048,
                .tp_retire_blk_tov = 0,
        };
        setsockopt(fd, SOL_PACKET, PACKET_RX_RING, &req, sizeof(req));

Compute in 64 bits instead. The operands are already bounded by the
existing validation, so nothing else changes. If you'd prefer a
different fix, just say so and I'll respin.

Fixes: f6fb8f100b ("af-packet: TPACKET_V3 flexible buffer implementation.")
Cc: stable@vger.kernel.org
Signed-off-by: Dairui Zhang <zhangdairui@gmail.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260923050101.1510064-1-zhangdairui@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 10:40:50 -07:00