Commit Graph

24044 Commits

Author SHA1 Message Date
Linus Torvalds
b1fa457bdd cgroup: Fixes for v7.3-rc4
- A cpuset partition could claim CPUs an ancestor partition already held
   exclusively. Restore the rejection an earlier change had turned into a
   warning.
 -----BEGIN PGP SIGNATURE-----
 
 iIQEABYKACwWIQTfIjM1kS57o3GsC/uxYfJx3gVYGQUCarjXAw4cdGpAa2VybmVs
 Lm9yZwAKCRCxYfJx3gVYGctVAQD3UPz2537MRkVZtcy4QfgyObMRXEaTGEBBJV+g
 ougCLgD/ZOVFjyK/bD3wXgtpiRZq9NbL+VvPQfrOjqz3Q/wVww8=
 =Vyrg
 -----END PGP SIGNATURE-----

Merge tag 'cgroup-for-7.3-rc4-fixes-2' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup

Pull cgroup fix from Tejun Heo:

 - A cpuset partition could claim CPUs an ancestor partition already
   held exclusively. Restore the rejection an earlier change had turned
   into a warning.

* tag 'cgroup-for-7.3-rc4-fixes-2' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup:
  cgroup/cpuset: Return PERR_NOCPUS in remote_partition_enable() on subpartitions_cpus conflict
2026-09-27 08:51:45 -07:00
Hui Peng
31c88350b7 cgroup/cpuset: Return PERR_NOCPUS in remote_partition_enable() on subpartitions_cpus conflict
When a remote partition is created underneath an existing local partition
via a non-partition (PRS_MEMBER) intermediate cgroup, update_prstate() sees
parent->partition_root_state == PRS_MEMBER and calls
remote_partition_enable().

Commit 86888c7bd1 ("cgroup/cpuset: Add warnings to catch inconsistency
in exclusive CPUs") replaced the cpumask_intersects(tmp->new_cpus,
subpartitions_cpus) error check in remote_partition_enable() with
WARN_ON_ONCE(). As a result, remote_partition_enable() emits a warning
and proceeds to enable the remote partition on CPUs that are already
owned by the ancestor local partition in subpartitions_cpus.

This can be reproduced on Linux 7.3.0-rc3 with:

  mkdir -p /tmp/cg1
  mount -t cgroup2 none /tmp/cg1
  echo "+cpuset" > /tmp/cg1/cgroup.subtree_control

  mkdir /tmp/cg1/A
  echo 1 > /tmp/cg1/A/cpuset.cpus
  echo 1 > /tmp/cg1/A/cpuset.cpus.exclusive
  echo root > /tmp/cg1/A/cpuset.cpus.partition
  echo "+cpuset" > /tmp/cg1/A/cgroup.subtree_control

  mkdir /tmp/cg1/A/B
  echo 1 > /tmp/cg1/A/B/cpuset.cpus
  echo 1 > /tmp/cg1/A/B/cpuset.cpus.exclusive
  echo "+cpuset" > /tmp/cg1/A/B/cgroup.subtree_control

  mkdir /tmp/cg1/A/B/D
  echo 1 > /tmp/cg1/A/B/D/cpuset.cpus
  echo 1 > /tmp/cg1/A/B/D/cpuset.cpus.exclusive
  echo root > /tmp/cg1/A/B/D/cpuset.cpus.partition

which triggers:

  WARNING: kernel/cgroup/cpuset.c:1594 at remote_partition_enable+0x1c1/0x300

and leaves both /tmp/cg1/A and /tmp/cg1/A/B/D as active root partitions
claiming exclusive CPU 1.

Fix this by returning PERR_NOCPUS when tmp->new_cpus intersects
subpartitions_cpus in remote_partition_enable(), matching the error code
used by remote_cpus_update() for the same subpartitions_cpus conflict, and
add a regression test case to
tools/testing/selftests/cgroup/test_cpuset_prs.sh.

Tested in QEMU on Linux 7.3.0-rc3 using the reproducer above and
tools/testing/selftests/cgroup/test_cpuset_prs.sh.

Fixes: 86888c7bd1 ("cgroup/cpuset: Add warnings to catch inconsistency in exclusive CPUs")
Suggested-by: Guopeng Zhang <guopeng.zhang@linux.dev>
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Hui Peng <benquike@gmail.com>
Reviewed-by: Waiman Long <longman@redhat.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-09-26 08:07:26 -10:00
Linus Torvalds
efb27d4767 Probes fixes for v7.3-rc4:
- kprobes: Fix permanent hang when flushing the kprobe optimizer
   Fix a deadlock when disabling kprobe optimization via sysctl or debugfs
   where flushers hung waiting for optimizer_completion. Replaced the
   completion with an optimizer_passes counter and wait_var_event_mutex()
   under kprobe_mutex so concurrent flushers can wait and wake up safely.
 
 - fprobe: Terminate the fgraph_data list when the reservation is not filled
   Fix an issue where unused shadow stack data left uninitialized by
   fprobe_fgraph_entry() was misparsed as stale fprobe headers on return.
   Explicitly write a zero word to terminate the list and update
   read_fprobe_header() to handle the zeroed slot properly.
 
 - ftracetest: Fix unique symbol check in kprobe_non_uniq_symbol.tc
   Fix false test failures in kprobe_non_uniq_symbol.tc on architectures
   like s390 where a symbol exists once in core kernel but also in modules.
   Anchor the /proc/kallsyms search regex to the end of the line so that
   module symbols are not incorrectly counted.
 -----BEGIN PGP SIGNATURE-----
 
 iQFPBAABCgA5FiEEh7BulGwFlgAOi5DV2/sHvwUrPxsFAmq3llUbHG1hc2FtaS5o
 aXJhbWF0c3VAZ21haWwuY29tAAoJENv7B78FKz8bAhEH/0EAamjv7/EDUoUq+BOO
 a2gnlYqvr+zcrDVQLNgiYbvTRDfIFPOdB2LpY7Rguee3747qeL7kkNATD10WFr1F
 5lXe5LaLncNIrvHDdtcT5eER5ePAuSDMSL5CwnJrRvXJw42iFsqegZ07nrvc9HFS
 5zw7Ej9VnFJFxeXIY3J4U92wkntLJ3JhsNheomOtQmEZU1g5ZPAbdq0icNrC3CAb
 FTezYk60VG0CT/gNTSd8JFnI4P5vKlZpkFFCLLMmptW4yQU9+ZdT55JRgY7wJ8ye
 w5rLBTRJrFhDgNAjoehMkoAtxik3dKgede7qlwxUk0sv+gwXssw28b5zabRBJgee
 XQ8=
 =cHwO
 -----END PGP SIGNATURE-----

Merge tag 'probes-fixes-v7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace

Pull probe fixes from Masami Hiramatsu:

 - kprobes: Fix permanent hang when flushing the kprobe optimizer

   Fix a deadlock when disabling kprobe optimization via sysctl or
   debugfs where flushers hung waiting for optimizer_completion.
   Replaced the completion with an optimizer_passes counter and
   wait_var_event_mutex() under kprobe_mutex so concurrent flushers can
   wait and wake up safely.

 - fprobe: Terminate the fgraph_data list when the reservation is not
   filled

   Fix an issue where unused shadow stack data left uninitialized by
   fprobe_fgraph_entry() was misparsed as stale fprobe headers on
   return. Explicitly write a zero word to terminate the list and update
   read_fprobe_header() to handle the zeroed slot properly.

 - ftracetest: Fix unique symbol check in kprobe_non_uniq_symbol.tc

   Fix false test failures in kprobe_non_uniq_symbol.tc on architectures
   like s390 where a symbol exists once in core kernel but also in
   modules. Anchor the /proc/kallsyms search regex to the end of the
   line so that module symbols are not incorrectly counted.

* tag 'probes-fixes-v7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace:
  kprobes: Fix permanent hang when flushing the kprobe optimizer
  fprobe: Terminate the fgraph_data list when the reservation is not filled
  selftests/ftrace: Fix unique symbol check in kprobe_non_uniq_symbol.tc
2026-09-26 08:36:29 -07:00
Linus Torvalds
eff8d2791c Arm:
- Invalidate the ITS translation cache when the guest changes the
   base address of the ITS tables (Fuad Tabba)
 
 - Skip saving ITS devices with device IDs that are out-of-bounds
   rather than failing the entire ITS save ioctl (Fuad Tabba)
 
 - Close race between VM teardown and invalidations of nested MMUs
   when handling MMU operations that are allowed to block
   (Lorenzo Stoakes)
 
 - Various fixes for the handling of the host's untrusted SVE
   configuration in pKVM (Fuad Tabba)
 
 - Make sure that empty SMCCC ranges based at 0 are rejected by the
   kvm_smccc_set_filter() (Karl Mehltretter)
 
 - Revoke the host mapping for pKVM's private stack pages, along
   with a new sanity check that all mappings in the hyp's private
   VA range have been correctly marked as hyp-owned (Fuad Tabba)
 
 - Lifetime fixes for the array of shadow stage-2 MMUs, ensuring that
   concurrent vCPU initialization cannot relocate in-use MMUs. Defer
   the freeing of shadow stage-2 MMUs to the point that no other users
   (e.g. MMU notifier) could reference them (Marc Zyngier)
 
 - Drop useless WARN when rejecting an unsupported ioctl for pKVM
   (Fuad Tabba)
 
 - Fix the steal_time selftest to install correctly-sized mappings for
   non-4K hosts (Sebastian Ott)
 
 - Correct mapping of fine-grained trap for GCSPOPX instruction
   (Mark Brown)
 
 - Fix KVM_BUG_ON() due to missing handling of DBGBXVR<n> from 32-bit
   guests (Karl Mehltretter)
 
 RISC-V:
 
 - Synchronize hrtimer during VCPU teardown.
 
 - Fix the conversion between vsip and hvip values.
 
 - Serialize IMSIC attributes with vCPU migration.
 
 - Release unused page after MMU invalidation.
 
 - Propagate interrupted G-stage faults to KVM user-space as EINTR.
 
 - Fix nested acceleration hfence entry update order.
 
 - Fix sdata leak and stale snapshot_addr in snapshot_set_shmem.
 
 - Preserve firmware counter value across PMU counter stop/start.
 
 - Report PMU snapshot write failure to the guest.
 
 - Fix perf-backed counter accounting across PMU stop and read.
 
 - Correctly propagate error of a hart status SBI call.
 
 s390:
 
 - Ensure that accesses through kvm_arch_set_irq_inatomic mark as dirty
   the pages that contain indicator and summary bits.
 
 - Fix compile warning for kvm_s390_update_cmma_dirty().
 
 - Fix incorrect propagation of ENOENT from _gaccess_shadow_fault() to
   userspace.
 
 - Move s390_kvm_mmu_commit_memory_region() into
   s390_kvm_mmu_prepare_memory_region() so that it can fail instead of WARN.
 
 - Add missing srcu in kvm_s390_set_irq_state().
 
 - Fix potential races in storage functions.
 
 - Fix race in _destroy_pages_crste().
 
 - Fix issues in the handling of KVM interrupt and page resources, when a
   queue that is assigned to a mediated device (mdev) is removed from the
   host's AP configuration.
 
 - Fix loop condition in uv_find_secrets.
 
 - Prevent potential out-of-bounds read.
 
 x86:
 
 - Fix a brown paper bag bug where KVM would incorrectly treat Intel PMU MSRs
   as valid on AMD.
 
 - Fix a regression in the hardware disable selftest where it checked the wrong
   macro when detecting glibc support (breaks at least musl).
 
 - Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is especially
   important for KVM_BUG_ON() flows, which often guard more dangerous bugs.
 
 - Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a bug
   where KVM would let userspace run a broken setup with stale vmcs12 pages.
 
 - Fix a class of bugs where KVM would fail to fill kvm_run exit fields if
   getting nested pages failed.
 
 - Treat reserved entries in the memory attributes xarray as "no attributes",
   to fix false positives when checking for mixed attributes.
 
 - Fix memcg accounting for the memory attributes xarray (the xarray library
   subtly requires the xarray to be configured for accounting upfront; the gfp
   flags taken at runtime are used only rarely).
 
 - Don't pre-reserve xarray entries when storing empty attributes, as storing
   NULL must not require memory allocation (KVM and other subsystems heavily
   rely on this behavior).
 
 - Fix a memory leak and a cache maintenance issue related to doing intra-host
   migration on an SEV guest.
 -----BEGIN PGP SIGNATURE-----
 
 iQFIBAABCAAyFiEE8TM4V0tmI4mGbHaCv/vSX3jHroMFAmq3Ta0UHHBib256aW5p
 QHJlZGhhdC5jb20ACgkQv/vSX3jHroMcygf/Z8GovgUACbWs/PSgVxZr5++MsbnM
 /N+vUP0VhsBxWCt191fBmKSBMNyTL734CKe0fkmisIAIPxf63uA3lttfMshNXKct
 VJ3TQkbA7qiSEGf2Um1dp4k6egqDDbTS1a4NF7CM7AEL4Wm3YDfqv2h4rZicEBsd
 Spy4yV8sKoiEHMnduhF2m5r0cWkrj194e2dolJy0p1WiNeEykf/nMTmNq5JWusBh
 TDDzZf1p2BianOmQ6+fkzAfWDSIdv0OPmoH7iqvQ/xhvm4dLWq2oeOz5Q2/uWAp0
 BqU4Z5zBrGRUWyXM0cDZEqBfmNH4v7oR9PfLgvD4qpMoqONuMj1+mIO68A==
 =XtLD
 -----END PGP SIGNATURE-----

Merge tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm

Pull kvm fixes from Paolo Bonzini:
 "Arm:

   - Invalidate the ITS translation cache when the guest changes the
     base address of the ITS tables (Fuad Tabba)

   - Skip saving ITS devices with device IDs that are out-of-bounds
     rather than failing the entire ITS save ioctl (Fuad Tabba)

   - Close race between VM teardown and invalidations of nested MMUs
     when handling MMU operations that are allowed to block (Lorenzo
     Stoakes)

   - Various fixes for the handling of the host's untrusted SVE
     configuration in pKVM (Fuad Tabba)

   - Make sure that empty SMCCC ranges based at 0 are rejected by the
     kvm_smccc_set_filter() (Karl Mehltretter)

   - Revoke the host mapping for pKVM's private stack pages, along with
     a new sanity check that all mappings in the hyp's private VA range
     have been correctly marked as hyp-owned (Fuad Tabba)

   - Lifetime fixes for the array of shadow stage-2 MMUs, ensuring that
     concurrent vCPU initialization cannot relocate in-use MMUs. Defer
     the freeing of shadow stage-2 MMUs to the point that no other users
     (e.g. MMU notifier) could reference them (Marc Zyngier)

   - Drop useless WARN when rejecting an unsupported ioctl for pKVM
     (Fuad Tabba)

   - Fix the steal_time selftest to install correctly-sized mappings for
     non-4K hosts (Sebastian Ott)

   - Correct mapping of fine-grained trap for GCSPOPX instruction (Mark
     Brown)

   - Fix KVM_BUG_ON() due to missing handling of DBGBXVR<n> from 32-bit
     guests (Karl Mehltretter)

  RISC-V:

   - Synchronize hrtimer during VCPU teardown

   - Fix the conversion between vsip and hvip values

   - Serialize IMSIC attributes with vCPU migration

   - Release unused page after MMU invalidation

   - Propagate interrupted G-stage faults to KVM user-space as EINTR

   - Fix nested acceleration hfence entry update order

   - Fix sdata leak and stale snapshot_addr in snapshot_set_shmem

   - Preserve firmware counter value across PMU counter stop/start

   - Report PMU snapshot write failure to the guest

   - Fix perf-backed counter accounting across PMU stop and read

   - Correctly propagate error of a hart status SBI call

  s390:

   - Ensure that accesses through kvm_arch_set_irq_inatomic mark as
     dirty the pages that contain indicator and summary bits

   - Fix compile warning for kvm_s390_update_cmma_dirty()

   - Fix incorrect propagation of ENOENT from _gaccess_shadow_fault() to
     userspace

   - Move s390_kvm_mmu_commit_memory_region() into
     s390_kvm_mmu_prepare_memory_region() so that it can fail instead of
     WARN

   - Add missing srcu in kvm_s390_set_irq_state()

   - Fix potential races in storage functions

   - Fix race in _destroy_pages_crste()

   - Fix issues in the handling of KVM interrupt and page resources,
     when a queue that is assigned to a mediated device (mdev) is
     removed from the host's AP configuration

   - Fix loop condition in uv_find_secrets

   - Prevent potential out-of-bounds read

  x86:

   - Fix a brown paper bag bug where KVM would incorrectly treat Intel
     PMU MSRs as valid on AMD

   - Fix a regression in the hardware disable selftest where it checked
     the wrong macro when detecting glibc support (breaks at least musl)

   - Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is
     especially important for KVM_BUG_ON() flows, which often guard more
     dangerous bugs

   - Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a
     bug where KVM would let userspace run a broken setup with stale
     vmcs12 pages

   - Fix a class of bugs where KVM would fail to fill kvm_run exit
     fields if getting nested pages failed

   - Treat reserved entries in the memory attributes xarray as "no
     attributes", to fix false positives when checking for mixed
     attributes

   - Fix memcg accounting for the memory attributes xarray (the xarray
     library subtly requires the xarray to be configured for accounting
     upfront; the gfp flags taken at runtime are used only rarely)

   - Don't pre-reserve xarray entries when storing empty attributes, as
     storing NULL must not require memory allocation (KVM and other
     subsystems heavily rely on this behavior)

   - Fix a memory leak and a cache maintenance issue related to doing
     intra-host migration on an SEV guest"

* tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (54 commits)
  KVM: SEV: Do cache maintenance on the source VM during intra-host migration
  KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV
  KVM: Don't pre-reserve xarray entries when storing empty/NULL attributes
  KVM: Ensure memory attributes xarray nodes are accounted to the caller's memcg
  KVM: Don't treat reserved xarray entries as having memory attributes
  KVM: x86: Fill kvm_run exit fields in common get_nested_state_pages() error paths
  KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages fails
  KVM: arm64: Fix AArch32 DBGBXVR<n> handling
  KVM: arm64: Fix FGT mapping for HFGITR_EL2.nGCSEPP
  KVM: selftests: fix steal_time for arm64 with host page size > 4K
  KVM: arm64: Don't WARN on an unknown VM ioctl in protected mode
  KVM: arm64: nv: Delay freeing of shadow S2 structures until VM destruction
  KVM: arm64: nv: Fix life cycle of the nested_mmus array
  KVM: arm64: Check every private mapping is hyp-owned at pKVM init
  KVM: arm64: Move the private VA allocation cursor to __io_map_next
  KVM: arm64: Match hyp text by physical address in fix_host_ownership()
  KVM: arm64: Transfer the hyp stack pages out of the host stage-2
  KVM: arm64: selftests: Test empty SMCCC filter range at base 0
  KVM: arm64: Return -EINVAL for an empty SMCCC filter range at base 0
  KVM: arm64: Derive GUEST_HAS_SVE from the SVE feature bit at EL2
  ...
2026-09-26 08:26:12 -07:00
Paolo Bonzini
c2f24f140c KVM fixes for 7.3-rcN
- Fix a brown paper bag bug where KVM would incorrectly treat Intel PMU MSRs
    as valid on AMD.
 
  - Fix a regression in the hardware disable selftest where it checked the wrong
    macro when detecting glibc support (breaks at least musl).
 
  - Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is especially
    important for KVM_BUG_ON() flows, which often guard more dangerous bugs.
 
  - Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a bug
    where KVM would let userspace run a broken setup with stale vmcs12 pages.
 
  - Fix a class of bugs where KVM would fail to fill kvm_run exit fields if
    getting nested pages failed.
 
  - Treat reserved entries in the memory attributes xarray as "no attributes",
    to fix false positives when checking for mixed attributes.
 
  - Fix memcg accounting for the memory attributes xarray (the xarray library
    subtly requires the xarray to be configured for accounting upfront; the gfp
    flags taken at runtime are used only rarely).
 
  - Don't pre-reserve xarray entries when storing empty attributes, as storing
    NULL must not require memory allocation (KVM and other subsystems heavily
    rely on this behavior).
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEKTobbabEP7vbhhN9OlYIJqCjN/0FAmq3A78ACgkQOlYIJqCj
 N/1uVA//TVDznH49htPNV4dh7vfiE8oF9ka4e9KTPA2eMdDeNVInNXBz0ZbJCdNN
 Hm7iZZfk8yx0JCCGXdp8EoYg66iwulBr0KZXRP2U2LoyiRELA1VtGGOLUVP3PXjF
 hElXD56+RrkJJSucINlFU7kSfl2hY/Jbaq87HVBLpXp8iY3stBbLdgMT1ReLdorO
 s9i2uJRqHKjtSG330qDYEG1TtRAdUehgRbxbbJVuOEYtU37tCvX6ai8dEjZAvxYo
 1VZOD05ITo12HUXPiZjIMoNTujQnZKmGYlBYczWlfOmkY/tN6SV8M41srV3ZcK0k
 XGOQM6tOzbVTWaSH/pW1krklnOFAbnL6QzHzb3enHZfUN8AikHPDyfsJWzt/s0f/
 ATTQk03OYL4fJMQSQ778EI6S4PcOzY9HiBxDXM5ySpUv1uyXw8ZFsQSdH/TYTMhi
 Dce42r5rlGa1vAKHwI1Xya2su3REhWzxYxGpUOg/PtRRqWvL2RMrHArzEd6zu+Ti
 jPinuQuXYcT1FTCa5S39NWhU4gWpjw8Pvs9VuUYSCDYxZCaqQ2YheYgsAcgPqgfE
 hS43XVmmOB/tXGsvu+/YuEfbzNV/mHOuJBcjBnnsBP3d6EFB02FPKg4/LYiXujYo
 OocUbSB7bKd87W3B+Ys3QmgOtZh4qX/aLYusa2vQ4Dp9ujXqcWs=
 =LG0B
 -----END PGP SIGNATURE-----

Merge tag 'kvm-x86-fixes-7.3-rc5' of https://github.com/kvm-x86/linux into HEAD

KVM fixes for 7.3-rcN

 - Fix a brown paper bag bug where KVM would incorrectly treat Intel PMU MSRs
   as valid on AMD.

 - Fix a regression in the hardware disable selftest where it checked the wrong
   macro when detecting glibc support (breaks at least musl).

 - Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is especially
   important for KVM_BUG_ON() flows, which often guard more dangerous bugs.

 - Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a bug
   where KVM would let userspace run a broken setup with stale vmcs12 pages.

 - Fix a class of bugs where KVM would fail to fill kvm_run exit fields if
   getting nested pages failed.

 - Treat reserved entries in the memory attributes xarray as "no attributes",
   to fix false positives when checking for mixed attributes.

 - Fix memcg accounting for the memory attributes xarray (the xarray library
   subtly requires the xarray to be configured for accounting upfront; the gfp
   flags taken at runtime are used only rarely).

 - Don't pre-reserve xarray entries when storing empty attributes, as storing
   NULL must not require memory allocation (KVM and other subsystems heavily
   rely on this behavior).
2026-09-26 00:40:07 -04:00
Linus Torvalds
aa98230e41 vfs-7.3-rc5.fixes
Please consider pulling these changes from the signed vfs-7.3-rc5.fixes tag.
 
 Thanks!
 Christian
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQRAhzRXHqcMeLMyaSiRxhvAZXjcogUCarab9AAKCRCRxhvAZXjc
 osTRAP90MigSgX2U/USBn8zgSTs9Key89pWuwpctsvELYKUSzwEA12UDBswbtgmq
 EGT6S2HSa6bifJMXG6AB+nMCcSluuAs=
 =KEgE
 -----END PGP SIGNATURE-----

Merge tag 'vfs-7.3-rc5.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs

Pull vfs fixes from Christian Brauner:

 - Revert "put_mnt_ns(): leave mounts connected". This allows the
   creation of reference count cycles in a very trivial way. We can't
   bring this in until we have fixed the underlying cause

 - vfs: Don't create the private nullfs instance for kthreads under
   namespace_sem to avoid false lockdeps complaints

 - binfmt_misc:
     - Copy the name into a stack buffer and look up the copy in
       bpf_binprm_select_interp()
     - bpf_binprm_set_interp() and bpf_binprm_set_interp_arg(): Check
       the private copy instead so the string that gets staged is the
       kstring that was checked

 - netfs:
     - Make netfs_read_gaps() use separate sink folios rather than one
       reused sink folio to discard unwanted data so that cifs checksum
       checking sees all the data that was fetched
     - Trim reads down to i_size so afs symlinks read correctly from the
       cache
     - Wrap the direct mempool ->alloc() calls the GFP_KERNEL paths make
       in alloc_hooks() via a new mempool_alloc_noreserve() helper

 - iov_iter: Use iov_iter_alignment() for the start and length check
   added to iov_iter_extract_bvecs() this cycle. It used iter_iov_addr()
   and iter_iov_len() which are only valid for ITER_UBUF and ITER_IOVEC
   iterators

 - super: Make iterate_supers_type() deletion-safe

 - inode: Stop evict_inodes() from rescanning the same inodes

 - writeback: Bound the cleanup_offline_cgwb() rescans

 - ntfs3: Use d_instantiate_new() in ntfs_create_inode()

 - ovl: Fix a use-after-free in the ovl_do_mkdir() debug print

 - dcache: Unpoison the inline name buffer in __d_alloc() for KMSAN

 - autofs: Fix a pipe file reference leak in autofs_kill_sb()

 - bpf: Drop the path_unlink and path_rmdir hooks from the list of hooks
   for which the verifier rewrites bpf_{set,remove}_dentry_xattr() to
   the _locked variants

 - squashfs: Range check the xz dictionary size before shifting by it

 - selftests: Add the missing eventfd, open_tree_ns, openat2 and xattr
   filesystems selftests to TARGETS and drop the stale openat2 entry
   left behind when those tests moved

* tag 'vfs-7.3-rc5.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs:
  netfs: Fix missing alloc tagging of direct mempool allocations
  bpf: fs/xattr: don't assume the inode is locked in path_unlink/path_rmdir
  autofs: fix sbi->pipe file reference leak in autofs_kill_sb()
  dcache: unpoison the inline name buffer in __d_alloc()
  ovl: fix UAF in ovl_do_mkdir() debug print
  super: make iterate_supers_type() deletion-safe
  Revert "put_mnt_ns(): leave mounts connected"
  Revert "selftests/filesystems: add mntns cleanup test"
  binfmt_misc: fix racy checks in bpf set_interp kfuncs
  binfmt_misc: fix OOB read in bpf_binprm_select_interp()
  fs: don't create the private nullfs mount under namespace_sem
  writeback: bound cleanup_offline_cgwb() rescans by rotating scanned inodes
  fs: avoid repeated scans in evict_inodes()
  netfs, afs: Fix symlink reading
  netfs: Fix netfs_read_gaps() to use separate sink folios
  squashfs: Add dictionary size range check to prevent shift-out-of-bounds
  fs/ntfs3: use d_instantiate_new() in ntfs_create_inode() and murder syzbot's "WARNING in do_new_mount" saga
  selftests/filesystems: fix missing and stale TARGETS entries
  block: Fix start and length check added to iov_iter_extract_bvecs()
2026-09-25 11:12:06 -07:00
Linus Torvalds
ee9c669f9b sched_ext: Fixes for v7.3-rc4
- A task reenqueued while its dispatch was still completing had its queued
   state clobbered by the dispatcher, dropping every later dispatch of the
   task. Wait for the in-flight dispatch to settle first.
 
 - A wakeup activation on another CPU marked the destination runqueue as
   mid-wakeup, stranding a pending local reenqueue. If the scheduler was
   unloaded first, the stale request pointed into freed memory that the next
   scheduler dereferenced.
 
 - ops.dequeue() ran with the source dispatch queue's lock held, so a
   scheduler iterating that queue from the callback deadlocked the CPU.
 
 - Schedulers with their own CPU ID mapping had no way to learn a task's
   initial CPU mask and rebuilt it themselves, which went wrong across
   sub-scheduler enable and re-home. Pass it to ops.enable().
 
 - A bypass dispatch event counter missed the dispatches made by the
   end-of-dispatch fallback and under-reported.
 
 - Selftests for the dequeue locking and initial mask changes.
 -----BEGIN PGP SIGNATURE-----
 
 iIQEABYKACwWIQTfIjM1kS57o3GsC/uxYfJx3gVYGQUCarWSLQ4cdGpAa2VybmVs
 Lm9yZwAKCRCxYfJx3gVYGY+sAP9y6nJh6vIvFh/X9FlJtWlNo0mncOKhy93E8jii
 8CKnPQEAhvX3+Gcdl+imTh4Z915kdsEByBjTTPPOXnQxI8BKUAY=
 =VhBF
 -----END PGP SIGNATURE-----

Merge tag 'sched_ext-for-7.3-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext

Pull sched_ext fixes from Tejun Heo:

 - A task reenqueued while its dispatch was still completing had its
   queued state clobbered by the dispatcher, dropping every later
   dispatch of the task. Wait for the in-flight dispatch to settle
   first

 - A wakeup activation on another CPU marked the destination runqueue as
   mid-wakeup, stranding a pending local reenqueue. If the scheduler was
   unloaded first, the stale request pointed into freed memory that the
   next scheduler dereferenced

 - ops.dequeue() ran with the source dispatch queue's lock held, so a
   scheduler iterating that queue from the callback deadlocked the CPU

 - Schedulers with their own CPU ID mapping had no way to learn a task's
   initial CPU mask and rebuilt it themselves, which went wrong across
   sub-scheduler enable and re-home. Pass it to ops.enable()

 - A bypass dispatch event counter missed the dispatches made by the
   end-of-dispatch fallback and under-reported

 - Selftests for the dequeue locking and initial mask changes

* tag 'sched_ext-for-7.3-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext:
  sched_ext: Count SCX_EV_SUB_BYPASS_DISPATCH in the dispatch fallback
  selftests/sched_ext: Check the cmask cid-form ops.enable() receives
  sched_ext: Pass the initial cmask to cid-form ops.enable()
  selftests/sched_ext: Test that ops.dequeue() can iterate the consumed DSQ
  sched_ext: Don't run ops.dequeue() with a DSQ lock held
  sched_ext: Derive SCX_RQ_IN_WAKEUP from the core enqueue flags
  sched_ext: Wait for SCX_OPSS_DISPATCHING before reenqueueing a task
2026-09-24 15:04:11 -07:00
Linus Torvalds
e8dfd03a1c cgroup: Fixes for v7.3-rc4
- With local event accounting, a fork rejected by the pids controller
   updated pids.events without notifying its pollers.
 
 - A cgroup selftest failed to compile with fortification enabled because
   an O_TMPFILE open lacked its mode argument.
 -----BEGIN PGP SIGNATURE-----
 
 iIQEABYKACwWIQTfIjM1kS57o3GsC/uxYfJx3gVYGQUCarWSKg4cdGpAa2VybmVs
 Lm9yZwAKCRCxYfJx3gVYGYoiAQD6JUCqDjv2Hr1YeMFlsoYVQZ7tNmljVnPD2tZ6
 9WmD/AEAotdlmzM8egOuDqAi2s+UMJPm8vCZuxVzewiGqhYELgk=
 =0Ljp
 -----END PGP SIGNATURE-----

Merge tag 'cgroup-for-7.3-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup

Pull cgroup fixes from Tejun Heo:

 - With local event accounting, a fork rejected by the pids controller
   updated pids.events without notifying its pollers

 - A cgroup selftest failed to compile with fortification enabled
   because an O_TMPFILE open lacked its mode argument

* tag 'cgroup-for-7.3-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup:
  cgroup/pids: Restore pids.events notifications in local mode
  selftests: cgroup: give the O_TMPFILE open in get_temp_fd() a mode
2026-09-24 14:26:12 -07:00
Linus Torvalds
f2c53ea949 Including fixes from Bluetooth, NFC and Netfilter.
Every week in this release is record-setting for number of posted
 patches. It doesn't seem like we're creating any regressions with
 all these fixes, 3 Fixes tags here point to 7.2 commits but none
 are true regression fixes. We're trying to keep the count down,
 nonetheless.
 
 Previous releases - regressions:
 
  - net: don't require the hwtstamp NDOs when a PHY provides timestamping
 
  - ipv6: fix dst leak for uncached routes
 
  - vrf: stop corrupting skb->csum when capturing CHECKSUM_COMPLETE
    packets
 
 Previous releases - always broken:
 
  - packet: use ubuf_info completion for TX_RING packets
 
  - arp: terminate device name before lookup
 
  - ipv6: do not let ipv6_find_hdr() return an offset past the packet end
 
  - udp: remove a disconnected socket from the 4-tuple hash table
 
  - sctp: discard the rest of the packet on a stale-cookie error
 
  - eth: mlx5: Bridge, fix remaining switchdev ownership gaps on merged
    eswitch
 
 Signed-off-by: Jakub Kicinski <kuba@kernel.org>
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEE6jPA+I1ugmIBA4hXMUZtbf5SIrsFAmq1bCQACgkQMUZtbf5S
 IrsgCBAAiHOxV3pR0pNw+K/dO5maYhuUV4w+Xcd6CbjRnxzYtImVC1q2ieez5dEe
 JrQrVK76t2wuEgGSyiWcckIufPL8xYg5nYaOUbkYMC2QRtDaqhCo+u2jXu26/Lr8
 O3sI2cj86XDZN+xKTvmLld1zIRyD6MhYxWz+gIicArdBqTsr5JsWkwhJE7SH013a
 iVGu0sIxTgCNu/AusD6SdtwFte26dqvadelI8yuDhBR6rM9pi4EWXChleg7ohCsu
 DoZTlaEjEiFavat+6ni99oz8MAT0WVBX0IzTqZa6TCnbCw4qUEbqN8G+CNgC/VPe
 Sq5FRoo6vJjlFviwgWPbpje2FoOYtk6CX1eTavx8gZ4pPMTzGkT62crS+ETwLO5H
 mRbE+ZjG8UDTUKpIY5P/IQYdznHSc9Ny/pHNlZhFmDdntJOp1lnegx5CmZkyuZxC
 Qc0hoFxKdjUrpQ1n0ygT3/P6MT9vXMwbGP52VLIxaB1oaYCKFV71NOp0ET5fCnYb
 CBK8cdN0kcaSLrpi3MnPamyQoNZfH78BiuKsLDu/TWg3i2HGmFTlp0/LJ0nkQcCU
 4KlCmiJNOkHjbQKHNtYMUOZQcZu5witEv33kb7gr581b0RR+n0lZTaXcnRIyXYdb
 ORqIa1Khg5mW3DLaIc9IrzEGGAW5zXUHc5OX7t8J2rVRQWhBEh4=
 =YiiK
 -----END PGP SIGNATURE-----

Merge tag 'net-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net

Pull networking fixes from Jakub Kicinski:
 "Including fixes from Bluetooth, NFC and Netfilter.

  Every week in this release is record-setting for number of posted
  patches. It doesn't seem like we're creating any regressions with all
  these fixes, three 'Fixes' tags here point to 7.2 commits but none are
  true regression fixes. We're trying to keep the count down,
  nonetheless.

  Previous releases - regressions:

   - net: don't require the hwtstamp NDOs when a PHY provides
     timestamping

   - ipv6: fix dst leak for uncached routes

   - vrf: stop corrupting skb->csum when capturing CHECKSUM_COMPLETE
     packets

  Previous releases - always broken:

   - packet: use ubuf_info completion for TX_RING packets

   - arp: terminate device name before lookup

   - ipv6: do not let ipv6_find_hdr() return an offset past the packet
     end

   - udp: remove a disconnected socket from the 4-tuple hash table

   - sctp: discard the rest of the packet on a stale-cookie error

   - eth: mlx5: Bridge, fix remaining switchdev ownership gaps on merged
     eswitch"

[ And lots of other random network driver fixes ]

* tag 'net-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (189 commits)
  tcp: prevent collapsing skbs across boundary in rtx queue
  vlan: ensure sufficient headroom in vlan_dev_hard_header()
  net/sched: sch_teql: fix shadowed err in __teql_resolve()
  bridge: check llc_mac_hdr_init() return value in br_send_bpdu()
  llc: fix skb UAF and leaks on llc_mac_hdr_init() failure
  llc: reserve device headroom for allocated frames
  gve: DQO: reject TSO packets with an out of range MSS
  gve: fix TX drop when GSO MSS is too small for hw
  gve: DQO: fix header length used by gve_can_send_tso() for UDP GSO
  net: flush skb_defer_nodes in dev_cpu_dead()
  net: ethernet: stmmac: dwmac-rk: fix bulk clock leak when the PHY clock fails
  af_packet: fix integer overflow in prb_calc_retire_blk_tmo()
  tipc: Fix a data race on mon->peer_cnt in mon_timeout()
  net: phy: intel-xway: workaround 100BASE-TX Link-Up issue
  net/smc: fix UAF on lgr list traversal in smcr_port_err()
  net/rds: size a connection's path set by the transport it ends up with
  nfp: hold IPsec RX state under the XArray lock
  net: ena: fix MMIO read buffer leak on probe failure
  net: ena: fix PHC cleanup on probe failure
  net/sched: act_ct: fix helper UAF due to extensions realloc
  ...
2026-09-24 11:45:33 -07:00
Linus Torvalds
415f204422 Landlock fix for v7.3-rc5
-----BEGIN PGP SIGNATURE-----
 
 iIYEABYKAC4WIQSVyBthFV4iTW/VU1/l49DojIL20gUCarVI0xAcbWljQGRpZ2lr
 b2QubmV0AAoJEOXj0OiMgvbSEVgA+gNbC9CVCbCo0oufZbVpQwlwtuuEKEVpvxZx
 q7oFYKCsAP9svCujGCXRHOmWhAAwe+wpXNb43l8coFn+pCfA8x3fCw==
 =NOli
 -----END PGP SIGNATURE-----

Merge tag 'landlock-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux

Pull Landlock fixes from Mickaël Salaün:
 "This mainly fixes the Landlock tracepoint support merged this cycle so
  that denial and rule events report the intended policy context,
  whether through tracefs or BTF-visible callbacks.

  The size of this all is mainly from propagating the corrected contract
  through event definitions and producers, adding new tests for the
  reported context, and updating the documentation.

  Also improve annotation and fix a GCC 16 build warning"

* tag 'landlock-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux:
  landlock: Widen ruleset versions to 64 bits
  landlock: Add counted_by in landlock_domain
  landlock: Fix tracepoint contract documentation
  selftests/landlock: Test network denial context
  selftests/landlock: Test filesystem denial blockers
  landlock: Report the effective signal number
  landlock: Report the actual ptrace tracer
  landlock: Fix network denial trace context
  landlock: Fix rule tracepoint context
  landlock: Fix filesystem denial blocker reporting
  landlock: Fix tracepoint fixed-width type names
  landlock: Work around gcc-16 -Wuninitialized warning
2026-09-24 10:53:59 -07:00
Jakub Kicinski
33a61f0923 Included fixes:
* add selftest coverage for peer VPN address validation
 * reject multicast, broadcast and loopback peer VPN addresses, which
   can never identify a peer
 * reject MP peers left with no usable VPN address, as they can never
   be selected for TX
 * reject duplicate peer VPN addresses, which made peer lookup return
   an arbitrary peer
 * fix stale entry left in the VPN address hashtable when an address
   is cleared
 * fix torn IPv6 address read on lockless TX when the unusable local
   source is cleared in place
 * fix torn IPv6 address read on lockless TX when a new local endpoint
   is learned in place
 * fix dst cache being populated with a route resolved from an already
   replaced bind
 * fix stale route being reused after the socket mark or UDP source
   port changed
 * fix bogus validation of an unspecified local source address, which
   must instead be left to route source autoselection
 * fix IPv6 link-local peer endpoints losing their scope id when
   configured via netlink, breaking route lookup
 -----BEGIN PGP SIGNATURE-----
 
 iJEEABYIADkWIQQr0db7q+Rc7Zog28Fc8QQzwdnOtwUCarECxxsUgAAAAAAEAA5t
 YW51MiwyLjUrMS4xMiwyLDIACgkQXPEEM8HZzrcvTAEA/r3oFnZ4EOtiOyA2OpZ/
 Iya+WAJ7LShj7tLymzujMe8BAIH6J2RyrWXnfDFvtSG8HGDEe4EBzVpP56tX0ZVn
 xc0N
 =fg0G
 -----END PGP SIGNATURE-----

Merge tag 'ovpn-net-20260921' of https://github.com/OpenVPN/ovpn-net-next

Antonio Quartulli says:

====================
Included fixes:
* add selftest coverage for peer VPN address validation
* reject multicast, broadcast and loopback peer VPN addresses, which
  can never identify a peer
* reject MP peers left with no usable VPN address, as they can never
  be selected for TX
* reject duplicate peer VPN addresses, which made peer lookup return
  an arbitrary peer
* fix stale entry left in the VPN address hashtable when an address
  is cleared
* fix torn IPv6 address read on lockless TX when the unusable local
  source is cleared in place
* fix torn IPv6 address read on lockless TX when a new local endpoint
  is learned in place
* fix dst cache being populated with a route resolved from an already
  replaced bind
* fix stale route being reused after the socket mark or UDP source
  port changed
* fix bogus validation of an unspecified local source address, which
  must instead be left to route source autoselection
* fix IPv6 link-local peer endpoints losing their scope id when
  configured via netlink, breaking route lookup

* tag 'ovpn-net-20260921' of https://github.com/OpenVPN/ovpn-net-next:
  selftests: ovpn: validate peer VPN addresses
  ovpn: reject invalid peer VPN addresses
  ovpn: reject multipeer peers without VPN addresses
  ovpn: reject duplicate peer VPN addresses
  ovpn: always unhash old VPN addresses before rehashing
  ovpn: replace bind when clearing stale local source
  ovpn: replace bind when learning local endpoint
  ovpn: validate peer state before caching UDP dst
  ovpn: track UDP socket route key for peer dst cache
  ovpn: skip UDP source validation for unspecified addresses
  ovpn: preserve IPv6 scope id for netlink peer endpoints
====================

Link: https://patch.msgid.link/20260921102215.3599702-1-antonio@openvpn.net
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 09:50:49 -07:00
Linus Torvalds
5fc5768c7c bpf-fixes
-----BEGIN PGP SIGNATURE-----
 
 iQJRBAABCgA7FiEE+soXsSLHKoYyzcli6rmadz2vbToFAmq1NvAdHGFsZXhlaS5z
 dGFyb3ZvaXRvdkBnbWFpbC5jb20ACgkQ6rmadz2vbTpmpA//fqEoC6Sq1zxo3ADH
 hV0Z9ewkNTjjH85QnispjcRkAhSHG3JscNXKRXm1NmNkwJsHJ4TDIKcjABYDjquD
 wqfvL9hLXPsvud0M/PR6/CZeBAWXpukkaxeYedY+83ttTHjzR0tDq0Ne9yvIV+Nc
 hS0qFIIHs8C6l2nyuNSxvgrv216orG8qd0Bi3tpDfsLqCsLLEmDyQ1H+ZzpJF2xW
 RV3oMcggCeo305m8+uiofQGf8hmmRrmA7SEfF+Qe08ab2GOn9glINfVTbTE3AdQW
 fkvAk8Zuio3hwwMBHDWWYXrKO0N3ykzcDk4V6JPUWiH1dOANf1tS3G1uJyxb4tsv
 dHVdA0xL3sg7YnuSywfb82vTXvQz5QEFeDYxxLB2fMJe8LcCfgF1lkIr1DbdrMjI
 hlozGCs3p/GVIVNhjGVPezWsUvzK4PIuKVM6U1qWojtAhXU/bJCvP0iBU83RlBL7
 S0GGSC/qvkuPed9hAp2pnrrRo/9GEp1PZq8AXsp3nti0OaRxgxpjjQQFzZEWlOrR
 bErtbKziHzl2xARvCRycNZ+QWT4ZdjmK++8pba7hOuFkRclOV4KRWk4OthBvcADu
 NfNsVXdXpOnesg7TNIUJ8heVAYPjKaBHJBuG2Prghrnc95uqEQwOfbYT15zAbe6I
 l3HqGerSa7WtbrnCwsf1DvyPPII=
 =ZxKg
 -----END PGP SIGNATURE-----

Merge tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf

Pull bpf fixes from Alexei Starovoitov:

 - Fix bpf_skb_change_tail() to drop the checksum offload instead of
   rejecting the trim of CHECKSUM_PARTIAL skbs (Daniel Borkmann)

 - Add KF_PERFMON kfunc flag and require CAP_PERFMON for kfuncs that
   read arbitrary memory and for untrusted read-only memory reads
   (Daniel Borkmann)

 - Clear scalar delta on narrowing stack spill (Daniel Borkmann)

 - Set up the frame pointer for the exception callback in arm64 JIT, and
   zero-fill other CPUs when BPF_F_CPU update creates a per-cpu hash
   element (Donggeun Yoo)

 - Various fixes (Emil Tsalapatis):
     - Fix bounds check underflow for skb-backed dynptrs
     - Fix rx_queue_mapping context access code generation in bpf_sock
     - Reject packet pointer arguments to subprogs that may mutate the
       packet
     - Reject ALU instructions that see arena and non-arena operands on
       different code paths

 - Fix copied_seq double-counting on sockmap self-redirect
   (Geliang Tang)

 - Fix divide-by-zero in btf_struct_walk() on a flexible array of
   zero-sized elements, fix out-of-bounds read of rtt_min in sock_ops
   (Jiayuan Chen)

 - Fix bpf_sock_destroy() out-of-bounds read of sk_protocol on TIME_WAIT
   and request socks, and sleeping under RCU when destroying a listener
   with pending children (Jiayuan Chen)

 - Fix JEQ/JNE with immediate operand in MIPS32 JIT and missing zero
   extension of BSWAP 16/32 in MIPS64 JIT (Johan Almbladh)

 - Avoid soft lockup in htab lookup[_and_delete] batch operations on
   large maps (Jose Fernandez)

 - Various fixes (Kumar Kartikeya Dwivedi):
     - Verify global subprogs in each sleepability context they are
       called from
     - Make post-verification instruction rewrites killable
     - Preserve packet pointer displacement in regsafe()
     - Apply CO-RE relocations before subprogram validation, restrict
       CO-RE poisoning to relocatable instructions, and reject truncated
       ldimm64 CO-RE relocations in libbpf
     - Assign lock identity to callback map values
     - Compare stack frames in regs_exact()
     - Bound ownership depth through local kptrs and graph roots

 - Fix u32 overflow in map batch operations when the map size exceeds
   4GB (Masoud Aghasi)

 - Fix UAF in bpf memalloc due to concurrent consumption of ttrace lists
   in alloc_bulk() (Pu Lehui)

 - Allow gotox as the terminal instruction of a program or a subprogram
   (Siddharth Chintamaneni)

 - Disallow bpf_skb_pull_data() for LWT_SEG6LOCAL, skip unsettled links
   in link iterator, and reject dev-bound-only programs on other devices
   (Weiming Shi)

 - Reject non-negative stack offsets in stack_slot_obj_get_spi()
   (Xu Yunxiang)

 - Check params size before reading reserved fields in
   bpf_crypto_ctx_create() (Yuqi Xu)

 - Reject max_entries > INT_MAX in sock_map_alloc() (Zhao Gongyi)

 - Use a 32-bit compare in xsk_map_gen_lookup() (Zhiling Zou)

 - Use kvfree() in xdp_test_run_teardown() (Zhixing Chen)

* tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf: (58 commits)
  selftests/bpf: Test per-cpu initialization of a BPF_F_CPU created element
  bpf: Zero-fill other CPUs when BPF_F_CPU creates a per-cpu hash element
  bpf: Fix BSWAP 32 and 16 on MIPS64
  bpf: Fix immediate JMP JEQ/JNE on MIPS32
  bpf: Reject dev-bound-only programs on other devices
  bpf, sockmap: Reject max_entries > INT_MAX in sock_map_alloc
  selftests/bpf: Test for mixed arena/nonarena code paths
  bpf: Prevent variable arena/non-arena register contents
  selftests/bpf: Test rejection of pkt args to mutating subprogs
  bpf: Reject pkt arguments in mutating subprogs
  selftests/bpf: Add selftests for rx_queue_mapping context access
  bpf: Fix bpf_sock context code generation
  selftests/bpf: Test dynptr slices past end of skb
  bpf: Fix bounds check for skb-backed dynptrs
  selftests/bpf: Reject iterator destruction through fp+0
  bpf: Reject non-negative offsets in stack_slot_obj_get_spi()
  bpf: Check params size before reading reserved fields
  selftests/bpf: Check local object ownership depth
  bpf: Bound ownership depth through local kptrs and graph roots
  selftests/bpf: Cover frame changes in bounded loops
  ...
2026-09-24 08:25:26 -07:00
Donggeun Yoo
3422808f4e
selftests/bpf: Test per-cpu initialization of a BPF_F_CPU created element
The existing cpu_flag subtests always prime a key with BPF_F_ALL_CPUS
before any BPF_F_CPU write, so the create path is never covered.

Add a subtest that creates the element with BPF_F_CPU on a map with
max_entries 1, so the key can only reuse the element the previous key
released, and check that the CPUs the update did not name read back
zero.  Run it for PERCPU_HASH preallocated and BPF_F_NO_PREALLOC,
whose per-cpu areas come from different allocators, and for
LRU_PERCPU_HASH.

Under BPF_F_NO_PREALLOC the reuse is only guaranteed on the cpu that
ran the delete, so pin the thread across the pair, and name a cpu other
than that one in map_flags.

Signed-off-by: Donggeun Yoo <donggeunyoo.kernel@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260924102321.2120434-3-donggeunyoo.kernel@gmail.com
2026-09-24 14:24:52 +00:00
Jakub Kicinski
90c2e97ff8 NFC fixes for net 7.3-rc5
Mostly security fixes.
 
 fix use-after-free in nfc_get_local_general_bytes
 llcp: Fix list corruption / refcount desync in nfc_llcp_recv_dm()
 llcp: Fix race condition in accept_queue lifecycle
 llcp: drop truncated I/RR/RNR PDUs in nfc_llcp_recv_hdlc()
 llcp: fix -ENOMEM on connect with zero-length service name
 llcp: fix WKS SAP hijacking via prefix match in nfc_llcp_wks_sap()
 llcp: fix sdreq TLV list leak on parse/alloc/send failure
 llcp: fix slab-out-of-bounds reads when logging service names
 selftests/nci: Fix out-of-bounds store on thread join
 selftests: nci: Correct pthread_create return value check
 selftests: nci: Fix uninitialized family ID on missing attribute
 virtual_ncidev: Add missing ioctl compat handler
 nfcmrvl: validate helper command length before pull
 pn533: fix OOB read in pn533_acr122_is_rx_frame_valid()
 port100: reject frames whose declared length exceeds the received data
 st21nfca: validate ISO15693 inventory length
 st21nfca: validate received frame size
 trf7970a: power down on startup RX gain failure
 
 Signed-off-by: David Heidelberg <david@ixit.cz>
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCAAdFiEE13oJz+7cK71TpwR0YAI/xNNJIHIFAmq0M08ACgkQYAI/xNNJ
 IHKXbw//cxnjUWiGiCWQKNGIrPQhO8xnxp29/adTfr5cpaCf9XS1Ah742JnsvV0J
 jiKwrV2VdhIedF0gCU51pfMdgV4VWWup2KQJbl820m8EaRmaKY6Spao/HxznnXtF
 H1JqYir2ISVADjX4hgQpqQIH2Vrz2o2QAkcsY+jjjydxxajmrrQh1SrJeEGom366
 eDHAPmO9MTheJZ8gFXtuz4JdoipVWwY0Ta8D3mmVfgX/Jv5ESsPa4Ie4gQvQ1Vua
 9bvYSgqrCGi4jtzEOxXNL1KfufjsbW9I64anwhxwjKJ+sL+/GJ0VfjN4Dgq9dnub
 H4dUJs9vUCMLY+Ns6Wo+2pJQ49Ults2SdVjH7kvdJ2FiUn+FvoPDLrQY00P/lhaB
 MPLrOCFaNVIzDKX98lxAzT/6c6xrNb5lTrkf17XHzcWa2H4eZj30pBaFfYrfiEH5
 dlASR5aSy+GuXa8J/i7+jpQ+xl0GEg9tb0GJsvsWTDAftAl8/9CvcZ/TCbjKeb9L
 oGeStPcl/i+MB9wGZLOga8lox0DX5PMNkIcp+RXIjVvPDYiCBh1a3lWHtz65Qx4m
 9VlNcIpUzN14VUTkWm1i29DVHpW8V5c/eLmoDHzBUFAVSVr2M0CkFwhPYOByqIzt
 lIcpLuTLXrAplzIFWMbi6gU8FSRAFX2I+vVr6rkgWIz2+QCPmYM=
 =lO2R
 -----END PGP SIGNATURE-----

Merge tag 'nfc-7.3-rc5' of https://codeberg.org/linux-nfc/linux

David Heidelberg says:

====================
NFC fixes for net 7.3-rc5

* tag 'nfc-7.3-rc5' of https://codeberg.org/linux-nfc/linux:
  nfc: pn533: fix OOB read in pn533_acr122_is_rx_frame_valid()
  nfc: llcp: fix slab-out-of-bounds reads when logging service names
  nfc: llcp: fix WKS SAP hijacking via prefix match in nfc_llcp_wks_sap()
  nfc: llcp: fix -ENOMEM on connect with zero-length service name
  nfc: st21nfca: validate ISO15693 inventory length
  nfc: fix use-after-free in nfc_get_local_general_bytes
  nfc: trf7970a: power down on startup RX gain failure
  nfc: llcp: fix sdreq TLV list leak on parse/alloc/send failure
  nfc: llcp: drop truncated I/RR/RNR PDUs in nfc_llcp_recv_hdlc()
  nfc: virtual_ncidev: Add missing ioctl compat handler
  selftests/nci: Fix out-of-bounds store on thread join
  selftests: nci: Fix uninitialized family ID on missing attribute
  nfc: llcp: Fix race condition in accept_queue lifecycle
  selftests: nci: Correct pthread_create return value check
  nfc: port100: reject frames whose declared length exceeds the received data
  nfc: llcp: Fix list corruption / refcount desync in nfc_llcp_recv_dm()
  nfc: st21nfca: validate received frame size
  nfc: nfcmrvl: validate helper command length before pull
====================

Link: https://patch.msgid.link/adeaccc1-cc04-4bb9-a28a-61a61d75ba14@ixit.cz
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23 19:45:33 -07:00
Chris Gellermann
6be581aeff
selftests/nci: Fix out-of-bounds store on thread join
The NCI test collects the exit status of its helper threads by passing
the address of an int to pthread_join():

	int status;
	...
	pthread_join(thread_t, (void **) &status);

pthread_join() stores a void pointer to the memory location. On 64-bit
systems, a void pointer is wider than an int, so the store overruns the
4 bytes of space allocated on the stack for the integer and corrupts the
adjacent stack. On our CHERI system, this caused a fault due to a
capability bounds violation.

Fix this by introducing a helper that joins a thread through a void
pointer and converts the result back to an integer, which is what the
helper threads return.

While here, also fix the logic in disconnect_tag() if the helper thread
creation failed. Previously, it would have joined a thread that was
never created when pthread_create() failed.

Fixes: f595cf1242 ("selftests: Add nci suite")
Signed-off-by: Chris Gellermann <christian.gellermann@codasip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260904095915.3372241-1-christian.gellermann@codasip.com
Signed-off-by: David Heidelberg <david@ixit.cz>
2026-09-23 22:13:16 +02:00
Chaithanya Lagisetty
eda518d2cd
selftests: nci: Fix uninitialized family ID on missing attribute
get_family_id() walks the generic netlink CTRL_CMD_GETFAMILY reply
looking for the CTRL_ATTR_FAMILY_ID attribute and returns the parsed
value in the local variable "id". If the reply does not carry that
attribute, the parsing loop never assigns "id" and the function returns
an indeterminate stack value, which the caller stores in self->fid and
uses for subsequent netlink requests.

Initialize "id" to 0 so a missing attribute yields a deterministic
(invalid) family ID instead of a garbage value.

Fixes: f595cf1242 ("selftests: Add nci suite")
Signed-off-by: Chaithanya Lagisetty <nagachaithanya9911@gmail.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Link: https://patch.msgid.link/20260901070618.3299012-1-nagachaithanya9911@gmail.com
Signed-off-by: David Heidelberg <david@ixit.cz>
2026-09-23 22:13:16 +02:00
Lei Zhu
3d8afc5243
selftests: nci: Correct pthread_create return value check
The pthread_create() functions returns 0 on success and a positive value on
failure. Modify the return value check to correctly detect failure cases.

Fixes: 72696bd8a0 ("selftests: nci: Extract the start/stop discovery function")
Signed-off-by: Lei Zhu <zhulei@kylinos.cn>
Link: https://patch.msgid.link/20260729072426.303484-1-zhulei_szu@163.com
Signed-off-by: David Heidelberg <david@ixit.cz>
2026-09-23 22:13:16 +02:00
Christian Brauner
76ebb69da6 Revert "selftests/filesystems: add mntns cleanup test"
This reverts commit 3452eecbcc.

The test checks that the mounts of a destroyed mount namespace stay
connected. The commit that made them stay connected is reverted next
because it leaks superblocks and loop devices, so drop the test with it.

Link: https://patch.msgid.link/20260917-work-put_mnt_ns-revert-v1-1-34d9b8679e68@kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-09-23 09:03:18 +02:00
Emil Tsalapatis
a9e86dd9de
selftests/bpf: Test for mixed arena/nonarena code paths
Add a selftest to confirm the verifier rejects ALU operations
that return arena or non-arena results depending on code path.

Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260922172028.6269-9-emil@etsalapatis.com
2026-09-22 19:34:05 +00:00
Emil Tsalapatis
1ed69a54d3
selftests/bpf: Test rejection of pkt args to mutating subprogs
Add a selftests that ensures that PTR_TO_PACKET arguments can
only be passed to subprogs that will never adjust the underlying
packet memory, and are rejected otherwise.

Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260922172028.6269-7-emil@etsalapatis.com
2026-09-22 19:34:04 +00:00
Emil Tsalapatis
dec0c209a6
selftests/bpf: Add selftests for rx_queue_mapping context access
Add tests to ensure the verifier properly tracks the 0 bit state
and width of the rx_queue_mapping field read from struct sock.

Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260922172028.6269-5-emil@etsalapatis.com
2026-09-22 19:34:04 +00:00
Emil Tsalapatis
4fd72eb9f1
selftests/bpf: Test dynptr slices past end of skb
Add a selftest to ensure dynptr slices cannot include
past the end of the linear area of an skb.

Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260922172028.6269-3-emil@etsalapatis.com
2026-09-22 19:34:04 +00:00
Linus Torvalds
fe2ec83746 hid-for-linus-2026092202
-----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCAAdFiEEL65usyKPHcrRDEicpmLzj2vtYEkFAmqyY7EACgkQpmLzj2vt
 YEnprxAAofmpQUorlic97WALU2CamXN/hdzLGlra2RPW6eJyxQUlwqR8xcUptx3e
 14OM5X4WDHfFD1lRFx7/flzn2J5h5CLXUWnNUp5DR7W0GLQSnSwOiuawfb5/Enkd
 aqhgsqC2nfEjxQ3iqdsY1MTASN/MEjolNU66tP31DQH5D+CGY6eZtivNwFHaHjJI
 Ps4/vr2nuddSCsXaWRKwdBQ9wMr8g+F7dIeI8vqAJlWg5IdN0tu3ClAuHUqMr6X3
 zMHJ7xkkrRg7S40xeQsajDiPGRRWPkC6G3lrJ1OSFOP1IEMKT5dLLg5HB2wLwfzI
 Xzw4CJZiSIaNKyMQLQJcik8wZZbc60A2ulDscbpk9PLlr9nweBc+7wzVZ0mOQ6Lo
 rD2iYG7ROG684Jt5HVd9SPrIgjCDnlOXDBT9JLyPCw38eEysEHOdSlwA0CSJyR21
 cD0UmfIUgFDXxR/omcPbZDVng8683eRR3W/Tj2XPBOyKZYF/JtnIsrbowIN2dmGe
 PKnsrtp2LW5/xiO/lTfEsiiHL8WoxtXj/Mwoq/NuZm4TvZ/C3nd3ZiEaL5UQeYYc
 AyTFOqBe4DoLauyS2M50gitvDpSBFS9i5xrjCukalpJGfTUlX8lQjU7kdH7O5rbh
 VlMjlYqPNSyd4Yo/nveqIS3ELrKGs5z3oDSxwSB05h1nCgn3dgo=
 =k0Hr
 -----END PGP SIGNATURE-----

Merge tag 'hid-for-linus-2026092202' of git://git.kernel.org/pub/scm/linux/kernel/git/hid/hid

Pull HID fixes from Jiri Kosina:

 - new device IDs/quirks (Logitech G502X, Elecom M-XT4DRBK, Steelseries
   Arctis 7, Asus Rog Z13 Folio, Lenovo Yoga Slim Gen 11)

 - fixes for various code issues found by LLMs

* tag 'hid-for-linus-2026092202' of git://git.kernel.org/pub/scm/linux/kernel/git/hid/hid:
  selftests/hid: add unnumbered variant to the hid_bpf tests
  HID: bpf: fix __hid_bpf_hw_check_params report length
  selftests/hid: add define for commonly used buf size
  HID: amd_sfh: Validate PCI BAR size before mapping
  HID: elecom: fix bus type for M-XGL20DLBK
  HID: elecom: Add support for ELECOM M-XT4DRBK (018E)
  HID: corsair-void: Fix firmware event packet description
  HID: i2c-hid: Add i2c-hid-quirk-bad-input-size quirk for 0911:5288 device
  HID: hid-oxp: use cancel_delayed_work_sync() in remove
  HID: i2c-hid: add reset quirk for Lenovo Yoga Slim 7x Gen 11 keyboard
  HID: roccat: fix locking in roccat_connect() and roccat_disconnect()
  HID: steelseries: Add support for Arctis 7 (2018)
  HID: logitech-hidpp: Add support for G502 X Lightspeed USB mouse
  HID: fix semantic patch and improve its performance
  HID: alps: fix use-after-free on input2 registration failure
  HID: alps: unregister DualPoint Stick input device on remove
  HID: winwing: fix use-after-free in force feedback teardown
  HID: multitouch: Add report ID mismatch quirk for ASUS ROG Z13 Folio
  HID: wacom: fix OOB read in wacom_wac_pen_serial_enforce()
  HID: quirks: add ALWAYS_POLL quirk for SDINNOVATION gaming keyboard
2026-09-22 10:30:17 -07:00
Linus Torvalds
a52a93358a 14 hotfixes. 10 are cc:stable. 11 are for MM.
Five are DAMON fixes.  One fixes an arm64 contpte bug where DAMON can
 write past the end of a page-table page, resulting in memory corruption
 and possible crashes.
 
 Two are hugetlb fixes.  One fixes an mremap() address calculation bug
 which can panic x86-64.
 
 There's also a missing anon_vma publication barrier which can result in
 hung tasks, and a writeback fix to keep long cgroup writeback drains from
 delaying Tasks-RCU grace periods.
 
 The remainder are smaller fixes and maintenance changes.
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQTTMBEPP41GrTpTJgfdBJ7gKXxAjgUCarHHNAAKCRDdBJ7gKXxA
 jhb7AP4yE/k7RrZC6zWg4M9ejI3fVlHI9+EG1mCLGiV57jZ4mAEA9LLSZryOD6Nc
 zsaeCtZhHcEdxW6EhO18hMfH4b7oJA0=
 =alhC
 -----END PGP SIGNATURE-----

Merge tag 'mm-hotfixes-stable-2026-09-21-17-08' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm

Pull MM fixes from Andrew Morton:
 "14 hotfixes.  10 are cc:stable.  11 are for MM.

  Five DAMON fixes: one fixes an arm64 contpte bug where DAMON can write
  past the end of a page-table page, resulting in memory corruption and
  possible crashes.

  Two hugetlb fixes: one fixes an mremap() address calculation bug which
  can panic x86-64.

  There's also a missing anon_vma publication barrier which can result
  in hung tasks, and a writeback fix to keep long cgroup writeback
  drains from delaying Tasks-RCU grace periods.

  The remainder are smaller fixes and maintenance changes"

* tag 'mm-hotfixes-stable-2026-09-21-17-08' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm:
  MAINTAINERS: update Xu Xin's email
  writeback: report a Tasks-RCU quiescent state per cgwb drain pass
  mm/damon/core: reset invalid quota->charge_target_from
  MAINTAINERS: add Baoquan and Baolin as MGLRU reviewers
  mm/rmap: fix missing barrier between anon_vma init and vma->anon_vma publish
  mm/hugetlb: preserve mremap address delta when skipping page tables
  mm/damon/core: fix unconditionally skip last region
  mm/damon/vaddr: avoid hw-driven pte updates during damon_hugetlb_mkold()
  mm/damon/core: allow esz to be set to zero
  mm/damon/ops-common: use a page-aligned address in damon_ptep_mkold()
  ocfs2: make ocfs2_calc_xattr_init() return void
  mailmap: update Haowen Bai's email address
  selftests/cgroup: account for zswap shrinker writeback
  mm/hugetlb: do not dissolve gigantic pages without runtime support
2026-09-22 10:03:33 -07:00
Jakub Kicinski
a87529034b selftests: net: nl_nlctrl: check the op ids in the policy map
Validate that the op map in the policy dump is correct.

Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Link: https://patch.msgid.link/20260918222949.4190284-2-kuba@kernel.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-22 15:21:32 +02:00
Yuya Kusakabe
be581d6635 selftests: net: fix CONFIG_SYSCTL sort order in configs
Commit 8d75c338f0 ("sysctl: remove CONFIG_PROC_SYSCTL, it just mirrors
CONFIG_SYSCTL") renamed CONFIG_PROC_SYSCTL to CONFIG_SYSCTL in place,
which left the entry out of alphabetical order in the net and
packetdrill configs. The netdev CI check for sorted selftest configs
now fails for every patch that touches either file.

Fixes: 8d75c338f0 ("sysctl: remove CONFIG_PROC_SYSCTL, it just mirrors CONFIG_SYSCTL")
Signed-off-by: Yuya Kusakabe <yuya.kusakabe@gmail.com>
Reviewed-by: Joel Granados <joel.granados@kernel.org>
Link: https://patch.msgid.link/20260918-selftests-net-config-sort-v1-1-968ea6e8c1b7@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-21 17:35:09 -07:00
Xu Yunxiang
8244668cbb selftests/bpf: Reject iterator destruction through fp+0
Add a verifier regression test that initializes a numeric iterator at fp-8
and attempts to destroy it through fp+0. The verifier must reject the
non-negative offset instead of treating it as the initialized stack slot.

Check the offset diagnostic to ensure rejection happens at the stack
object address check. The numeric iterator destroy operation is a no-op;
this test checks verifier rejection and does not run the program.

Signed-off-by: Xu Yunxiang <xyx2021@mail.ustc.edu.cn>
Signed-off-by: Andrii Nakryiko <andrii@kernel.org>
Reviewed-by: Sun Jian <sun.jian.kdev@gmail.com>
Link: https://lore.kernel.org/bpf/20260920210423.345636-3-xyx2021@mail.ustc.edu.cn
2026-09-21 15:00:59 -07:00
Ralf Lici
0062080268 selftests: ovpn: validate peer VPN addresses
Exercise peer VPN address validation through both peer creation and
update. Check missing, unspecified, duplicate, multicast, broadcast,
loopback, IPv4-compatible and IPv4-mapped addresses.

Temporarily configure a peer with both address families to verify that
either family can be cleared while the other remains configured, then
restore the original addresses before running the existing traffic
tests.

Extend ovpn-cli peer updates with an optional VPN address and preserve
peer creation errors so the negative tests can observe rejected netlink
requests.

Signed-off-by: Ralf Lici <ralf@mandelbit.com>
Signed-off-by: Antonio Quartulli <antonio@openvpn.net>
2026-09-21 11:39:13 +02:00
Linus Torvalds
156fa7417f x86 fixes:
- Reject the loading of a potentially problematic microcode version
    on Intel Granite Rapids systems (Chang S. Bae)
 
  - On FRED, reconstruct the proper #GP context for rejected INT
    instructions, to fix a signal ABI regression (Matthew Schwartz)
 
  - Add a test for this signal ABI regression the x86
    self-test suite (Matthew Schwartz)
 
  - Don't emit the new and not yet properly supported EGPR instructions
    (%r16-%r31) on CONFIG_X86_NATIVE_CPU=y builds (Chang S. Bae)
 
 Signed-off-by: Ingo Molnar <mingo@kernel.org>
 -----BEGIN PGP SIGNATURE-----
 
 iQJFBAABCgAvFiEEBpT5eoXrXCwVQwEKEnMQ0APhK1gFAmqvrhgRHG1pbmdvQGtl
 cm5lbC5vcmcACgkQEnMQ0APhK1inbQ/9G4Q0h9esQ945n0uYRYvVLuj2P8DvE3XN
 9d5IJ5pIYdhJCvlv8yXyv1u3sq59QYNmUgpBMghHJonhqobtnchvBSiznoAMgi8L
 yjRlcaFBPX9qjFQV1W9WJQuUHhW/12QFy9Bo9n1uXEv6pYspYVDwodCaS+rEvkTF
 0wK3TjETv90TravGjFscYTt2VLAiy+cd/FxkUA0sBaMLjhFpUyAPHGJe/vcf8vT3
 rEXkY8EeSjv5F6eKMTDVacrSZu2c9wco1bJIsSFEcxZGFzhDbuCOGiDCkwmB3uuD
 ReKbEUC0KmF8qd4Ubf0dMGpbHT9LDu9Ggex619cHFOHMupPubP+yym7aCpm3v7Sb
 gZe6eV6VrSREJ+oBVKsqMrWXsg1YDj6uJhC/5K3S78xBeRq64kKsmBbC93Zp20TC
 QpSAxKIFqp8C+tfKzpBmcSbcipRUfATJxLjp94QaQDp467jjJCvpDJdc12mV7WR9
 0gMtFjxUaIe8BGX7s2PuWPgZ8+CIH0hZQuttxUU4QWvM2RGU+QrDaMwTeAY7wme4
 xomc0wpXQb5enrjmFu8betlD0xfjgt7k6eW0njezRpWSk3wyHDd5w/6Vfn8pA6Fx
 AQV4UuszXzQmG+W2pdIjLhnjmsjEeWMpmwgzn0rkcqujWHeh8XOSvQIjULFrFtFv
 srSbum6vEfk=
 =MP1H
 -----END PGP SIGNATURE-----

Merge tag 'x86-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip

Pull x86 fixes from Ingo Molnar:

 - Reject the loading of a potentially problematic microcode version
   on Intel Granite Rapids systems (Chang S. Bae)

 - On FRED, reconstruct the proper #GP context for rejected INT
   instructions, to fix a signal ABI regression (Matthew Schwartz)

 - Add a test for this signal ABI regression the x86
   self-test suite (Matthew Schwartz)

 - Don't emit the new and not yet properly supported EGPR instructions
   (%r16-%r31) on CONFIG_X86_NATIVE_CPU=y builds (Chang S. Bae)

* tag 'x86-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
  x86/build/64: Prevent native builds from generating EGPR use
  selftests/x86: Check signal state for rejected software interrupts
  x86/fred: Reconstruct the #GP context for rejected INT instructions
  x86/microcode/intel: Reject problematic loading on Granite Rapids systems
2026-09-20 09:49:10 -07:00
Mickaël Salaün
c8dcb17205
selftests/landlock: Test network denial context
Network denial events now report the port from their one authoritative
validated address through one signed field.  Verify this value directly
without inferring a source or destination role.

Exercise a nonzero IPv4 TCP bind, an explicit-zero IPv4 UDP bind, a
synthetic-zero IPv6 UDP autobind, and a family-only AF_UNSPEC UDP send.
Require exactly one event with the exact policy blocker for each shape.
The value -1 distinguishes a family-only address with no validated port
from the two valid port-zero cases.

Test coverage for security/landlock is 91.6% of 2625 lines according to
LLVM 22.

Cc: Günther Noack <gnoack@google.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Link: https://patch.msgid.link/20260918185036.608651-9-mic@digikod.net
[mic: Add test coverage]
Signed-off-by: Mickaël Salaün <mic@digikod.net>
2026-09-20 14:13:31 +02:00
Mickaël Salaün
c6dea91d84
selftests/landlock: Test filesystem denial blockers
Filesystem denial traces identify the policy change needed to allow a
request, so require exact blocker values rather than merely nonempty
output. Pin a READ_DIR denial to exactly one event with
blockers=read_dir. Pin a REFER-only mount denial to EPERM and exactly
one event with blockers=change_topology.

The mount child retains CAP_SYS_ADMIN so Landlock is the only expected
source of EPERM.  This prevents a later capability failure from masking
a Landlock regression; the trace-collecting parent remains unsandboxed.

Cc: Günther Noack <gnoack@google.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Link: https://patch.msgid.link/20260918185036.608651-8-mic@digikod.net
Signed-off-by: Mickaël Salaün <mic@digikod.net>
2026-09-20 11:08:05 +02:00
Mickaël Salaün
98b04ab00f
landlock: Fix network denial trace context
Network denial events report source and destination ports reconstructed
from audit data. Their zero values are ambiguous, and neither identifies
the complete endpoint that Landlock checked.

Carry the checked sockaddr and its signed length in a private trace-only
context. For an enabled event, validate the length and copy only the
initialized prefix into zeroed local storage. This prevents a typed BPF
program from reading uninitialized bytes while exposing the socket
family, socket, address, and length.

Replace the source and destination trace-record fields with one signed
port derived from the checked address. A value of -1 means that no port
was checked, zero is a valid port, and positive values use host
endianness. Bind blockers select the bind address; connect and send
blockers select the destination.

Cc: Günther Noack <gnoack@google.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Fixes: 01ce260f5c ("landlock: Add landlock_deny_access_fs and landlock_deny_access_net")
Link: https://patch.msgid.link/20260918185036.608651-5-mic@digikod.net
Signed-off-by: Mickaël Salaün <mic@digikod.net>
2026-09-20 11:08:02 +02:00
Mickaël Salaün
1a985d3890
landlock: Fix rule tracepoint context
Name each event after the identity it reports. Add-rule events describe
UAPI rule insertion, so rename them after LANDLOCK_RULE_PATH_BENEATH and
LANDLOCK_RULE_NET_PORT. Check-rule events describe matches in internal
rule trees, so rename them after LANDLOCK_KEY_INODE and
LANDLOCK_KEY_NET_PORT. This remains accurate if multiple UAPI rule types
share one lookup and stored rule. Keep denial event names based on
filesystem and network families because they describe final access
decisions.

Use u64 for growable access masks passed by value to add-rule and
check-rule typed BTF callbacks. CO-RE can relocate pointer-reached
fields, but it cannot widen a scalar callback slot declared by a BPF
program. Keep native access_mask_t for internal state and trace records.

For add-rule callbacks, report the normalized per-call contribution
passed to landlock_insert_rule() and expose the complete validated flags
value. Put the ruleset and flags first as a common invocation prefix.
This distinguishes duplicate and effective-zero additions without
recovering arguments from saved syscall registers.

Cc: Günther Noack <gnoack@google.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Fixes: 63747c9477 ("landlock: Add landlock_add_rule_fs and landlock_add_rule_net tracepoints")
Fixes: 3f1f106e4c ("landlock: Add tracepoints for rule checking")
Link: https://patch.msgid.link/20260918185036.608651-4-mic@digikod.net
Signed-off-by: Mickaël Salaün <mic@digikod.net>
2026-09-20 11:08:02 +02:00
Jamal Hadi Salim
1e24c4f2ee selftests: tc-testing: add a lateral-drift hfsc classify-walk test
The classify-loop fix bounds a walk's non-descending hops, so the guard
must not misfire on a legal walk that reaches its leaf through a
level-drift lateral chain. Add a case that builds exactly that chain and
asserts traffic still reaches the chain's own leaf.

A lateral hop can only exist because a bind was legal when it was made
and a later class add raised the target's level, so the setup binds each
hop while the target is still a leaf and only then deepens it: bind
1:1 -> 1:2 while 1:2 is a leaf, add 1:20 under 1:2, add 1:3 and bind
1:2 -> 1:3 while 1:3 is a leaf, then add 1:30 and 1:31 under 1:3 and
bind 1:3 -> 1:31. The walk root -> 1:1 -> 1:2 -> 1:3 -> 1:31 then takes
two lateral hops and must reach leaf 1:31.

The default class is 1:30, distinct from the asserted leaf, and the
verify pattern is anchored to the 1:31 stats line, so neither a
fall-through to the default nor a nonzero count on another class can
satisfy the check. On the patched kernel the test passes; with the bound
forced to zero the walk falls to the default and 1:31 stays idle, so the
test fails.

Reviewed-by: Victor Nogueira <victor@mojatatu.com>
Tested-by: hybris <hybris@mojatatu.ai>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-CTUU.v3.20260916184908@mojatatu.com.2
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-19 16:42:11 -07:00
Tejun Heo
d781d1b78a selftests/sched_ext: Check the cmask cid-form ops.enable() receives
cid-form ops.enable() now hands the task's cmask to the scheduler and
set_cmask() repeats it right after. Add a cid-form selftest that checks both
against p->cpus_ptr, that they match each other, that the initial
set_cmask() lands before set_weight() and before the task first becomes
runnable, and that set_cmask() never precedes enable(), across class-switch
enables, fork-path enables and live affinity changes.

v2: Mismatch details returned through a caller-local struct instead of
globals, alloc_words validated in the header check, loop bounded by nr_cids
directly (Andrea Righi).

Signed-off-by: Tejun Heo <tj@kernel.org>
2026-09-19 04:09:37 -10:00
Kumar Kartikeya Dwivedi
0288ed6748
selftests/bpf: Check local object ownership depth
Build raw program BTF records to exercise local object ownership without
constructing a runtime chain deep enough to threaten the kernel stack.

Cover referenced-kptr and percpu-kptr self-cycles, a two-type cycle, and a
cycle mixing a graph root with a referenced kptr. The existing graph-only
check accepts these local-kptr cycles and over-limit kptr chains; the fix
rejects them with -ELOOP.

Pin the eight-record depth boundary with a terminal plain object. Exercise
both parent-first and child-first BTF orders, and a shared suffix reached
first through a shorter path. These cases require cached suffix depths to
be checked against the remaining depth budget on each path.

Check list and rbtree chains of three, four, eight, and nine types. This
covers the old graph-only depth boundary and the new explicit bound. The
existing linked-list BTF tests still reject pure graph cycles and now accept
the longer acyclic layouts previously rejected by the conservative rule.

Although bpf_percpu_obj_new() currently rejects types with special fields,
require the percpu cycle to fail at BTF load so future support cannot bypass
the ownership bound. Keep a positive control for a non-owning kptr, which
remains outside the ownership graph.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260914132444.2564218-3-memxor@gmail.com
2026-09-19 05:26:43 +00:00
Kumar Kartikeya Dwivedi
bfc888f045
bpf: Bound ownership depth through local kptrs and graph roots
Program-allocated objects can own other local objects through referenced
kptrs. bpf_obj_free_fields() follows those pointers through
__bpf_obj_drop_impl() synchronously, before the object storage is freed
through RCU. A self-referential local kptr type therefore permits arbitrarily
deep object chains, and dropping the head can exhaust the kernel stack.
Long acyclic type chains have the same problem.

btf_check_and_fixup_fields() still assumes referenced kptrs only point to
kernel types and checks ownership through list and rbtree roots only. Its
existing rule is sufficient for graph-only cycles: the target of each graph
edge must contain a node, so every type in a cycle has both a root and a
node. The rule rejects such a type owning another root, breaking every
cycle. It also limits graph-only chains to three types, or two if the first
type contains a node, and conservatively rejects longer acyclic chains.
The missing local-kptr edges, rather than a missed graph-only cycle, are the
bug introduced by support for bpf_kptr_xchg() into local kptrs.

Replace that restriction with one bounded ownership walk covering graph
roots and local referenced kptrs. Run it after all BTF records have been
fixed up, reject cycles and paths deeper than eight record-bearing types,
and cache each type's suffix depth while checking it against the remaining
budget. This also permits the longer acyclic graph-only layouts rejected
by the old rule; update their existing BTF tests accordingly.

Keep the bound independent of MAX_CALL_FRAMES because recursive destruction
can run below a BPF call chain. A plain local pointee without special-field
metadata adds only a final non-recursing drop. Non-owning kptrs and
kernel-BTF kptrs do not recurse through local records and remain outside the
walk. Include local percpu-kptr edges too, although allocation of percpu
objects with special fields is currently forbidden, so that relaxing that
restriction cannot bypass the ownership bound.

btf_check_and_fixup_fields() continues to initialize graph_root.value_rec,
including for separately allocated map records. The ownership relationships
belong to immutable program BTF and only need validation at BTF load time.

Fixes: b0966c7245 ("bpf: Support bpf_kptr_xchg into local kptr")
Reported-by: Nicholas Carlini <npc@anthropic.com>
Suggested-by: Nicholas Carlini <npc@anthropic.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260914132444.2564218-2-memxor@gmail.com
2026-09-19 05:26:43 +00:00
Kumar Kartikeya Dwivedi
0705878489
selftests/bpf: Cover frame changes in bounded loops
Add a finite loop whose progress is represented only by changing the
frame number of a stack pointer. The loop first reads zero from the
caller's stack, switches to the same offset in the callee's stack, and
exits after reading one on its next iteration.

Force frequent checkpoints so the test exercises infinite-loop detection,
and check that the program returns one when run.

Without the frameno comparison in regs_exact(), the program is rejected
with an "infinite loop detected" diagnostic instead of loading
successfully.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Tested-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260919014213.1840880-3-memxor@gmail.com
2026-09-19 05:25:14 +00:00
Linus Torvalds
a077be4fde arm64 fixes for -rc4
- Fix hypercall arguments when resetting EL2 vectors during hibernation
 
 - Fix hibernation with 52-bit capable kernels on machines without
   52-bit addressing, similarly to the recent kexec fix
 
 - Fix a bunch of clumsy codegen issues with our per-cpu accessors
 
 - Fix MTE ptrace documentation to reflect the de-facto ABI behaviour
 
 - Fix pthread_join() usage in MTE selftest
 
 - Fix port selection in the Arm CMN PMU driver
 -----BEGIN PGP SIGNATURE-----
 
 iQFEBAABCgAuFiEEPxTL6PPUbjXGY88ct6xw3ITBYzQFAmqtMZ4QHHdpbGxAa2Vy
 bmVsLm9yZwAKCRC3rHDchMFjNNujB/4pa4pdivXRMOH39HSDu+8i3KlREGpTA9dq
 LJvwUjlgoIw0d8/b5sMSAxJi8RA29IENJPqrU97eBBLKgC2gq1oMwjOaHNTV8XUj
 gal9hwuZcfO5NIJi61TkyLxn++ysVYXD3J0tbhSXbqz2oeg/jxwj9SsFfQux34GY
 GeGb2cWvEXznLgN0h+vZiNh6+FsQRN+dizpiHKLh+wpbZ/vIzuVTQEo6w/+BEUWg
 R4YQlvpj/2yAC0BBdPb9gALUWFbjea6zpX5gM5/PgLoqW0CIJvGXahZ7T5kDoXBT
 Xp5zsKvcC8dxM5nCJIJ7xRPg75v97toYrrl2AoNQN9bgOynQAI5k
 =j0NQ
 -----END PGP SIGNATURE-----

Merge tag 'arm64-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/arm64/linux

Pull arm64 fixes from Will Deacon:
 "In this batch we've got a couple of hibernation fixes, a couple of
  minor MTE fixes, some per-cpu codegen fixes (which were found as part
  of Mark's series adding preemptible this_cpu_*() operations) and a fix
  for the Arm CMN PMU driver.

  Summary:

   - Fix hypercall arguments when resetting EL2 vectors during
     hibernation

   - Fix hibernation with 52-bit capable kernels on machines without
     52-bit addressing, similarly to the recent kexec fix

   - Fix a bunch of clumsy codegen issues with our per-cpu accessors

   - Fix MTE ptrace documentation to reflect the de-facto ABI behaviour

   - Fix pthread_join() usage in MTE selftest

   - Fix port selection in the Arm CMN PMU driver"

* tag 'arm64-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/arm64/linux:
  arm64: mte: Fix PTRACE_{PEEK,POKE}MTETAGS error documentation
  kselftest/arm64: Fix size of thread_data values for pthread_join()
  arm64: percpu: Fix LSE operations on {8,16}-bit types
  arm64: percpu: Fix this_cpu_and() mask generation
  arm64: percpu: Fix this_cpu_write() casting
  arm64: hibernate: clone only the linear map that exists at runtime
  perf/arm-cmn: Fix wp_dev_sel2 setting for multi-DTM configurations
  arm64: hibernate: pass HVC_SET_VECTORS args to the resume hvc
2026-09-18 09:28:30 -07:00
Joshua Hahn
8c7fdc0b4c selftests/cgroup: account for zswap shrinker writeback
The test_no_invasive_cgroup_shrink selftest checks that when a cgroup has
zswapped out more memory than memory.zswap.max, it does not trigger
writeback for other cgroups.  To do this, it compares the writeback count
in a control cgroup and makes sure that it is 0, and then checks the
writeback count in an aggressor cgroup who does expect to see writeback.

However, when the zswap shrinker is enabled, the victim cgroup can see
legitimate writebacks not triggered by the aggressor.  In some Meta CI
tests, we have seen this failure mode happen.

Instead of checking that the victim cgroup has 0 writeback, compare the
writeback values before and after the aggressor runs and check that the
victim cgroup did not perform any additional writeback.  Note that this
can still lead to probabilistic failures if writebacks take longer than
5 seconds, but this should fix the systematic failure case and make
"not ok test_no_invasive_cgroup_shrink" less likely.

Link: https://lore.kernel.org/20260902194521.3652178-1-joshua.hahnjy@gmail.com
Fixes: b5ba474f3f ("zswap: shrink zswap pool based on memory pressure")
Signed-off-by: Joshua Hahn <joshua.hahnjy@gmail.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Reported-by: Krush Chavan <krushchavan@outlook.com>
Suggested-by: Nhat Pham <nphamcs@gmail.com>
Cc: <stable@vger.kernel.org>
2026-09-18 08:20:13 -07:00
Matthew Schwartz
96443a53bc selftests/x86: Check signal state for rejected software interrupts
Add a test of the signal ABI for INT instructions in both 32-bit and
64-bit processes. Check the signal number, trap number, error code,
si_code, si_addr, instruction pointer and RF/TF state against legacy
IDT behavior. Include a 15-byte prefixed INT to check that IP uses the
hardware instruction length. Exercise both INT3 encodings, INT4, UD2
and HLT to cover the unchanged trap and fault paths.

Run each instruction with TF clear and set. Resume at a known NOP after
handling the signal and check that single-stepping traps after the NOP.

Also drive INT 0x2d under ptrace, which resumes through the fault frame
rather than sigreturn and so exposes a stale FRED software event flag.
Start from an INT3 stop, whose FRED frame has no software event flag,
instead of the syscall frame of raise(SIGSTOP). Single-step into the INT
and check that the fault reports its address. Then suppress SIGSEGV and
resume at the NOP, once with PTRACE_SINGLESTEP and once with PTRACE_CONT
and TF set. Section 6.2.3 of the Intel FRED specification [1] specifies
the immediate single-step trap caused by returning with both that flag
and TF set. Check that each trap occurs after the NOP, rather than at
its address.

Report whether the CPU supports FRED, since a pass looks the same on
either entry path. INT 0x80 with IA32 emulation disabled and a 64-bit
tracer of a 32-bit tracee are not covered.

Both variants pass all 29 checks on a non-FRED AMD host and on Panther
Lake with FRED enabled and the fix applied. With the same binaries on
unpatched Panther Lake, 16 signal-context checks fail and the first
ptrace check reports the IP after the INT. The two dependent ptrace
resume checks are not reached.

[1] Intel Flexible Return and Event Delivery (FRED) Specification,
revision 9.0 (346446-009US), section 6.2.3.

Signed-off-by: Matthew Schwartz <matthew.schwartz@linux.dev>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Link: https://cdrdv2.intel.com/v1/dl/getContent/678938 # [1]
Link: https://patch.msgid.link/20260917230907.2080792-3-matthew.schwartz@linux.dev
2026-09-18 12:33:18 +02:00
Kumar Kartikeya Dwivedi
04ae4ffc57 selftests/bpf: Check callback map value lock identity
Add a verifier test which retains a map value from an outer callback and
then acquires a lock through an inner callback value before attempting to
release the outer callback value. Both values can denote different elements,
so the verifier must reject the mismatched unlock.

Also exercise callbacks reached through two inner-map lookups. The lookup
results share inner_map_meta but may refer to different one-element arrays,
so their callback values must retain distinct lock identities.

Extend the existing spin_lock failure table and reuse its array and
inner-map fixtures to keep these cases alongside the other lock identity
tests. Update the nested callback reference-leak expectation for the extra
callback value ID.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Link: https://patch.msgid.link/20260917233222.2542500-10-memxor@gmail.com
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
2026-09-17 18:01:43 -07:00
Kumar Kartikeya Dwivedi
3440505aca selftests/bpf: Test CO-RE instruction poisoning restrictions
Add raw CO-RE relocations that fail to resolve their target enum value.
Place each supported and unsupported instruction form in dead code.
Unsupported targets must fail relocation with a diagnostic even when they
are unreachable. Supported ALU immediates, memory accesses, and ldimm64
instructions must still be poisoned and removed as dead code, allowing the
program to load. Check that both halves of ldimm64 are poisoned.

Load every instruction stream without relocations first to ensure that
rejection is caused by the relocation rather than the original program.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260917233222.2542500-8-memxor@gmail.com
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
2026-09-17 18:01:43 -07:00
Kumar Kartikeya Dwivedi
968ee7c06b selftests/bpf: Test early in-kernel CO-RE relocation
Add a raw program load with CO-RE relocation metadata but no func_info or
line_info. Place the relocation in dead code and require the poisoning log,
proving that the kernel processes standalone CO-RE metadata instead of
silently skipping it.

Also give a subprogram a relocatable immediate as its terminal instruction.
Require the relocation's poisoning log before check_subprogs() rejects the
resulting fall-through. With the old ordering, check_subprogs() rejects the
original terminal instruction before CO-RE can emit the substitution log, so
the test continues to distinguish the ordering after relocation target
validation is tightened.

Submit a trailing ldimm64 first slot with CO-RE metadata and require the early
structural diagnostic. This exercises the check that protects relocation
processing instead of the later regular instruction validation.

Load the standalone instruction stream without relocation metadata first to
ensure that CO-RE processing causes its poisoning diagnostic. Encode the fixed
BTF metadata directly with the selftest BTF helpers.

Suggested-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260917233222.2542500-6-memxor@gmail.com
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
2026-09-17 18:01:42 -07:00
Kumar Kartikeya Dwivedi
2059d9af54 selftests/bpf: Test packet pointer class displacement pruning
Add two paths whose packet pointer ranges are individually compatible at
a join but whose members have different relative displacements. The first
path proves an eight-byte access through one member. On the second path,
the same guard only proves that the access starts before data_end.

An affected verifier prunes the second path and accepts the program. With
packet pointer class displacement preserved, it explores that path and
rejects the out-of-bounds access.

Read the unknown offset and branch selector directly from XDP context
fields, and force state checkpoints so the pruning attempt does not
depend on the verifier checkpoint heuristics.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Link: https://patch.msgid.link/20260917233222.2542500-4-memxor@gmail.com
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
2026-09-17 18:01:42 -07:00
Jamal Hadi Salim
960ab631f3 selftests/tc-testing: add u32 manual table handle IDR tests
35fc: create a manual table with handle 801:, then add an auto-allocated
table. Before the fix, the auto allocation reuses id 1 and hands out the
same handle 0x80100000, aliasing the manual table; the test requires the
manual 801: handle to keep exactly one entry in the dump.

a6e8: with a live u32 table keeping the tc_u_common alive, add and delete
a manual table with handle 901:, then re-add it. Unpatched, the delete
leaks the raw-keyed IDR entry and the re-add fails with -ENOSPC; the
test requires the re-add to succeed.

Reviewed-by: Victor Nogueira <victor@mojatatu.com>
Tested-by: hybris <hybris@mojatatu.ai>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/QDISC-LQFE.v1.20260911041746.2@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17 17:00:28 -07:00
Kumar Kartikeya Dwivedi
a452e729b7 selftests/bpf: Test global subprog callback contexts
Exercise global subprogram verification from workqueue callbacks, which
can run in a sleepable context even when the containing program is not
sleepable. An unprotected callback must not let the global subprogram use
implicit RCU protection inherited from the program.

Add a negative case which loads an RCU-protected task kptr in a global
subprogram reached from a workqueue callback. It fails on an unfixed
kernel because the program is incorrectly accepted. Also cover a
workqueue callback protected by an explicit RCU read-side critical
section, where the same global subprogram remains valid.

Call the same harmless global subprogram directly from the main program
and from an unprotected callback. Mark it __weak __noinline so both calls
survive optimization, and check that its instruction statistics account
for both verification contexts. This also verifies that global calls
from callbacks are not rejected wholesale. Workqueue callbacks return
zero explicitly after the global call, as required by their contract.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260914131923.2544250-3-memxor@gmail.com
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
2026-09-17 10:55:55 -07:00
Linus Torvalds
b5a051f6b8 Including fixes from Netfilter, Bluetooth, IPSec and WiFi.
Previous releases - regressions:
 
   - netfilter: hold reference on ct until flow is released
 
   - bridge:
     - move switchdev call outside rcu
     - vlan: fix bugs caused by switchdev deletion errors
 
   - wifi:
     - mac80211: reset state when starting AP fails
     - cfg80211: don't free driver-owned scan requests
 
   - tcp: don't call skb_clone_and_charge_r() for close()d listener in tcp_v6_do_rcv().
 
   - mptcp: return sk_wait_data() errors from recvmsg()
 
   - xfrm: serialize state GC with device state flush
 
   - drop_monitor: synchronize tracepoint unregistration on error path
 
   - bluetooth:
     - eir: validate service data length before reading UUID
     - hci_sync: serialize local codec list cleanup
     - RFCOMM: avoid socket lock inversion in listener cleanup
 
   - eth: lan743x: fix RX checksum use-after-free
 
   - eth: mvpp2: prevent buffer overflow in page_pool allocation
 
 Previous releases - always broken:
 
   - core: lock the socket in sock_gettstamp()
 
   - neighbour: enforce min/max to NDTPA_INTERVAL_PROBE_TIME_MS.
 
   - sched: codel: bound the dropping loop per dequeue call
 
   - wifi: mac80211: include TIM bitmap control for buffered S1G mcast traffic
 
   - psp: avoid conflicts with skb->decrypted and sk_validate_xmit_skb()
 
   - xfrm: fix stack OOB read in iptfs_skb_reset_frag_walk()
 
   - bluetooth: hci_qca: do not write to the serial port after it is closed
 
   - dsa: mxl862xx: disable the stats poll on teardown
 
   - eth: stmmac: fix TSO header length truncation
 
   - eth: ip_tunnel: initialize `options_len` before referencing options
 
 Signed-off-by: Paolo Abeni <pabeni@redhat.com>
 -----BEGIN PGP SIGNATURE-----
 
 iQJKBAABCgA0FiEEg1AjqC77wbdLX2LbKSR5jcyPE6QFAmqsEOsWHHBhb2xvLmFi
 ZW5pQGdtYWlsLmNvbQAKCRApJHmNzI8TpKMnD/4vxx/YloxnyDssUB95WcTfiHau
 XD+YDfTTf4Qy3DwKnMiesZ4w2547FXG+LwJZGpAGWpRSE99OawFHvx5gyn3Vf4IU
 HEsOuZEYMwhSeeMJxGdCg8K6McNEQx+aAD7D8gLoisnJr/DCKBJtxFgKVEjaKUfP
 aCG+FmbBqdS7hwBVbYREwmqwQSMnNzRWLd7/10/oYMcLfvgEsIkZisRErAW2BgZj
 mzGw/IcG+QRldDPGDwLwLzfEG9o1a2JctSRrQ3uLHZ3VOdmpnSkmf25s2IaEFAn6
 xfIhGPMgYIBDrakn/Ci4fAF0L98FtBq4Sa21HlvPMBst7rcce5x49ddSlIUWxRVZ
 Fbvs/0IMN0cEpYGbVJxE8iF3yo+t8XvsMdGS1JXd/ycaL9lpF+9gCzNvChSxba3N
 ik7BGAmlg76rqMuzeeMbWqMCmOcCBhQsb7iZXjNStJoiVY+UT2QfgvJ1TqF8Qrar
 Eu/xFkaqk8i/7jrx4ujceg9XpRt1Y3Y3Pq0sgdzlzTm162AV/kg1fvwHc6ORIML7
 71XpBSGvjyTHoh95ob/w3/5fec/ekzvWlCBL/cYIwBJ4R6n7Vk5e0Y1yZDffPSo3
 xQR0r4YdLewkjQdtnGEih6FSyWsu6nYpVSuwHbof1bQSzvlIM6TAabJmtJckFTld
 1bR2C9UVw5mQYMy8PA==
 =Gw4S
 -----END PGP SIGNATURE-----

Merge tag 'net-7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net

Pull networking fixes from Paolo Abeni:
 "Including fixes from Netfilter, Bluetooth, IPSec and WiFi.

  Previous releases - regressions:

   - netfilter: hold reference on ct until flow is released

   - bridge:
      - move switchdev call outside rcu
      - vlan: fix bugs caused by switchdev deletion errors

   - wifi:
      - mac80211: reset state when starting AP fails
      - cfg80211: don't free driver-owned scan requests

   - tcp: don't call skb_clone_and_charge_r() for close()d listener in
     tcp_v6_do_rcv()

   - mptcp: return sk_wait_data() errors from recvmsg()

   - xfrm: serialize state GC with device state flush

   - drop_monitor: synchronize tracepoint unregistration on error path

   - bluetooth:
      - eir: validate service data length before reading UUID
      - hci_sync: serialize local codec list cleanup
      - RFCOMM: avoid socket lock inversion in listener cleanup

   - eth:
      - lan743x: fix RX checksum use-after-free
      - mvpp2: prevent buffer overflow in page_pool allocation

  Previous releases - always broken:

   - core: lock the socket in sock_gettstamp()

   - neighbour: enforce min/max to NDTPA_INTERVAL_PROBE_TIME_MS.

   - sched: codel: bound the dropping loop per dequeue call

   - wifi: mac80211: include TIM bitmap control for buffered S1G mcast
     traffic

   - psp: avoid conflicts with skb->decrypted and sk_validate_xmit_skb()

   - xfrm: fix stack OOB read in iptfs_skb_reset_frag_walk()

   - bluetooth: hci_qca: do not write to the serial port after it is
     closed

   - dsa: mxl862xx: disable the stats poll on teardown

   - eth:
      - stmmac: fix TSO header length truncation
      - ip_tunnel: initialize `options_len` before referencing options"

* tag 'net-7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (159 commits)
  mptcp: fix bad accounting in __mptcp_subflow_push_pending()
  mptcp: close race between scheduler and state change
  mptcp: avoid unneeded actions on subflow reset
  net: skbuff: do not leave stale header offsets after pskb_carve()
  selftests: net: packetdrill: test exclusion of old ACK from TCP fast path
  tcp: exclude old ACKs from tcp fast path
  dpll: reject a reference sync pin which is not on the pin's dpll
  net: mvpp2: prevent buffer overflow in page_pool allocation
  net: macb: fix ordering around PTP timestamp read
  selftests: drv-net: psp: test PSP and TCP ULP mutual exclusion
  net: psp: avoid conflicts with skb->decrypted and sk_validate_xmit_skb()
  net: stmmac: preserve real_num_tx_queues on mqprio setup failure
  net: stmmac: propagate FPE preemption-class mapping errors
  net: wwan: t7xx: validate the netif index in t7xx_ccmni_recv_skb()
  net: wwan: mhi_wwan_mbim: check skb_copy_bits() return value
  net: wwan: mhi_wwan_mbim: guard against a cyclic NDP chain
  net: ethernet: cortina: Ack RX overrun interrupt correctly
  net: lock the socket in sock_gettstamp()
  eth: fbnic: ring the doorbell if a burst ends in a drop
  net: netsec: fix device_node reference leak on phy_np
  ...
2026-09-17 10:40:48 -07:00
fangqiurong
9ec7ba20c9 selftests/sched_ext: Test that ops.dequeue() can iterate the consumed DSQ
Add a scheduler whose ops.dequeue() iterates the source user DSQ with
bpf_iter_scx_dsq. The iteration takes the DSQ's raw spinlock; on a
kernel that runs ops.dequeue() while the consume path still holds that
lock, the first task consumed self-deadlocks the CPU with IRQs
disabled. The watchdog cannot recover from that state, so on an
unfixed kernel this test wedges the system instead of failing cleanly.
On a fixed kernel the scheduler runs clean and the test passes.

Signed-off-by: fangqiurong <fangqiurong@kylinos.cn>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-09-17 07:20:17 -10:00
Inbal Schussheim
d841cd7513 selftests: net: packetdrill: test exclusion of old ACK from TCP fast path
Add a packetdrill test for an in-sequence data segment carrying an
excessively old ACK.

Verify that the segment falls through from the TCP fast path to the slow
path, where the existing ACK validation rejects it and sends a challenge
ACK. The payload is not accepted and RCV.NXT remains unchanged.

Based on the reproducer from Commit 3d501dd326
("tcp: do not accept ACK of bytes we never sent").

Signed-off-by: Inbal Schussheim <inbal.lipshtat@mail.huji.ac.il>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260914090408.1435080-3-inbal.lipshtat@mail.huji.ac.il
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-17 15:17:58 +02:00