Commit Graph

4314 Commits

Author SHA1 Message Date
Mikhail Zaslonko
0945285e6c s390/debug: Fix race between debug area resize and event logging
Trace functions check for non-NULL id->areas without lock to minimize
overhead. This opens a race window where a NULL pointer dereference
occurs if id->areas is set to NULL (e.g. via echo 0 > ../pages) after
the check and before id->lock is taken.

Fix this by rechecking id->areas under lock.

Signed-off-by: Mikhail Zaslonko <zaslonko@linux.ibm.com>
Reviewed-by: Peter Oberparleiter <oberpar@linux.ibm.com>
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-09-08 15:53:40 +02:00
Mikhail Zaslonko
22d4210bf9 s390/debug: Do not repeat parameter override notice on debug_set_level()
Commit a2cec68637 ("s390/debug: Add s390dbf kernel parameter") calls
debug_get_param() from both debug_info_create() and debug_set_level().
Since debug_get_param() emits the override notice unconditionally, and
drivers typically call debug_set_level() right after debug_register(),
the same line is printed twice per debug area:

  s390dbf: 0.0.1234: override level to 6
  s390dbf: 0.0.1234: override level to 6

For areas registered per device this is multiplied by the device count.
With 's390dbf=0.0.*:6' a system with many DASDs emits a large number of
redundant lines during boot.

Add a quiet parameter to debug_get_param() and pass quiet=true from
debug_set_level(), where the override has already been announced during
registration. The remaining callers keep printing the notice.

Signed-off-by: Mikhail Zaslonko <zaslonko@linux.ibm.com>
Reviewed-by: Peter Oberparleiter <oberpar@linux.ibm.com>
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-09-08 15:53:40 +02:00
Mikhail Zaslonko
b1eb31d533 s390/debug: Fix NULL pointer dereference in debug_set_level()
Commit a2cec68637 ("s390/debug: Add s390dbf kernel parameter")
incorrectly removed a null-id check from debug_set_level(), introducing
a possible NULL pointer dereference for debug-API users that put
debug_register() results unchecked into debug_set_level().

Fix this by moving the check from the internal _debug_set_level()
variant back to the external debug_set_level() wrapper.

Fixes: a2cec68637 ("s390/debug: Add s390dbf kernel parameter")

Signed-off-by: Mikhail Zaslonko <zaslonko@linux.ibm.com>
Reviewed-by: Peter Oberparleiter <oberpar@linux.ibm.com>
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-09-08 15:53:40 +02:00
Thomas Richter
9ecc4d0338 s390/pai: Support CPU hotplug for PMU PAI
The command 'perf stat -e pai_crypto/CRYPTO_ALL/ -- <command>'
crashes the kernel when CPUs are hotplug added during that run.

Root cause is the missing allocation of per-CPU data structures
for that new CPU. The allocation is dynamic and the first
event that has task context creates such a structure for
each online CPU. This is not sufficient. CPUs may be offline
during event creation and can be set online during the
perf run time. For example commands

 # echo 0 > /sys/devices/system/cpu/cpu1/online
 # perf stat -e cycles -i -- stress-ng -t10s --matrix X
 # sleep 1
 # echo 1 > /sys/devices/system/cpu/cpu1/online

Currently without a CPU hotplug handler, that new CPU has no
per-CPU data infrastructure. The scheduler runs PMU call back
function pai_add() to install the PMU support for that CPU before
the task is being scheduled on that new CPU.
In pai_add() instructions

    mp = this_cpu_ptr(pai_root[idx].mapptr);
    cpump = mp->mapptr;

return a NULL pointer and the result is a kernel panic as variable
cpump is used inside that function.

Add CPU hotplug support for CPU add and delete and create
the necessary per-CPU data infrastructure during CPU hotplug
add processing. Same for CPU hotplug remove.
This is done when the CPU is offline to ensure the data structures
are available when CPU is made online and tasks are scheduled on it.

[hca@linux.ibm.com: fixup error path in pai_init()]

Cc: stable@vger.kernel.org # v6.19
Fixes: 582cc1b28e ("s390/pai_ext: Enable per-task and system-wide sampling event")
Fixes: 9f66572f28 ("s390/pai_crypto: Enable per-task and system-wide sampling event")
Signed-off-by: Thomas Richter <tmricht@linux.ibm.com>
Reviewed-by: Jan Polensky <japo@linux.ibm.com>
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-09-08 15:53:40 +02:00
Thomas Richter
e8df39dacb s390/pai: Move locking to event init and delete
Move mutex locking from per CPU allocation to event allocation.
No functional change.

Signed-off-by: Thomas Richter <tmricht@linux.ibm.com>
Reviewed-by: Sumanth Korikkar <sumanthk@linux.ibm.com>
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-09-08 15:53:40 +02:00
Thomas Richter
4eef4ab3aa s390/pai: Use PAI PMU index as parameter replacing event
Use PAI PMU index value as function argument instead of pointer
to struct perf_event. Only that index value is used inside
functions pai_alloc_cpu() and pai_event_destroy_cpu().
No functional change.

Signed-off-by: Thomas Richter <tmricht@linux.ibm.com>
Reviewed-by: Sumanth Korikkar <sumanthk@linux.ibm.com>
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-09-08 15:53:40 +02:00
Mete Durlu
bb06e5a2a0 s390/topology: Switch to common cpu capacity code
s390 implementation of cpu capacity management infrastructure code does
not do anything different than its common code counterpart. Switch to
common code functions and remove the smp_cpu_*_capacity() functions.
Make s390 code better align with other architectures which utilize
cpu_capacity. No functional changes.

Allow cpu_capacity attributes inside sysfs to accurately reflect cpu
capacity.
ex:

$ cat /sys/devices/system/cpu/cpu0/polarization
vertical:high
$ cat /sys/devices/system/cpu/cpu0/cpu_capacity
1024

$ cat /sys/devices/system/cpu/cpu40/polarization
vertical:low
$ cat /sys/devices/system/cpu/cpu40/cpu_capacity
128

Prior to commit 6bceea7a1e ("arch_topology: Relocate cpu_scale to
topology.[h|c]") cpu_capacity attribute was only available to the common
arch_topology driver's users. Reflect the correct values to the newly
made available attributes.

Signed-off-by: Mete Durlu <meted@linux.ibm.com>
Reviewed-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
2026-08-31 16:25:39 +02:00
Heiko Carstens
8cff0ac216 s390/pai: Reduce excessive debug feature size
The pai debug feature is registered with 256 areas, where each area
contains 32 pages. This sums up to a total of 32MiB. The code does not use
any debug exceptions, which means that 255 of those areas are never
used. In addition all existing debug feature calls have a lower level (5)
than the default level (3).

This in turn means that without user interaction the debug feature is
unused.

Reduce the number of areas to 1, and also reduce the number of pages for
the remaining area to 1. Since user interaction is required, the user can
also increase the size of the remaining area, instead of wasting memory by
default.

This reduces the total size of the debug feature to 4KiB.

Fixes: a3f8423622 ("s390/pai_crypto: Add PAI crypto characteristics table for parameters")
Reviewed-by: Thomas Richter <tmricht@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
2026-08-31 16:25:39 +02:00
Thomas Richter
f3110e969a s390/pai: Handle multiple PMU stop callback invocations
Handle the following scenario:
The kernel protects itself against a very high sampling load and
throttles the sampling using:

  perf_event_throttle() --> PMU->stop()

Shortly later the scheduler may terminate the task and removes it from the
CPU. It again calls

  PMU->stop()

which results in two invocations of PMU->stop() called back to back.
Protect against this and check the PERF_HES_STOPPED bit on function
entry.  If it is already set return.
Clear bit PERF_HES_STOPPED in PMU->start().

Prohibit ioctl(fd, PERF_EVENT_IOC_PERIOD, ...) call for this event.
It sets perf_event::event_limit to a positive value and causes
perf_event_overflow() to invoke pai_stop() call back function when
perf_event::event_limit hits zero. This is not supported because the
sample events CRYPTO_ALL and NNPA_ALL are only taken at schedule out
of a task.

Use list_for_each_entry_safe() for safe iteration over syswide_list
in pai_have_samples().

Fixes: 9f66572f28 ("s390/pai_crypto: Enable per-task and system-wide sampling event")
Fixes: 582cc1b28e ("s390/pai_ext: Enable per-task and system-wide sampling event")
Cc: stable@vger.kernel.org # v6.19+
Signed-off-by: Thomas Richter <tmricht@linux.ibm.com>
Reviewed-by: Sumanth Korikkar <sumanthk@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
2026-08-31 16:25:39 +02:00
Sumanth Korikkar
439077c39d s390/diag324: Preserve -EBUSY return code
When diag324 reports -EBUSY, the error code is
overwritten by the result of copy_to_user() and put_user(). As a result,
the ioctl may incorrectly return success instead of -EBUSY.

Preserve the original diag324 return code and only return -EFAULT when
copying data to userspace fails.

Fixes: 90e6f191e1 ("s390/diag324: Retrieve power readings via diag 0x324")
Signed-off-by: Sumanth Korikkar <sumanthk@linux.ibm.com>
Reviewed-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
2026-08-31 16:25:39 +02:00
Vasily Gorbik
37f61b71cb s390/ipl: Fix NULL deref in dump_reipl without re-IPL parm block
Unlike kdump, which passes the re-IPL parameter block through os_info,
the stand-alone dump passes it through the IPL parm block address and
checksum in lowcore.

Some IPL types, like HMC FTP boot or QEMU direct kernel boot, might not
provide an IPL parameter block. In this case reipl_type_init() selects
IPL_TYPE_UNKNOWN and reipl_block_actual remains NULL. Nevertheless,
dump_reipl_run() unconditionally dereferences it when preparing the
lowcore fields. This may happen to work by chance when address zero
contains readable lowcore data. A zero IPL parameter block address is
then stored in lowcore, causing the stand-alone dumper to enter disabled
wait after completing the dump.

Explicitly store a zero IPL parameter block address and checksum when no
re-IPL parameter block is available. This does not change the behavior:
the stand-alone dumper completes the dump and halts, while valid re-IPL
parameter blocks continue to be handled as before.

Fixes: 099b765139 ("[S390] Automatic IPL after dump")
Reviewed-by: Mikhail Zaslonko <zaslonko@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
2026-08-31 16:25:38 +02:00
Vasily Gorbik
7f91887111 s390/ipl: Fix NULL deref in kdump without re-IPL parm block
Some IPL types, like HMC FTP boot or QEMU direct kernel boot, might
not provide an IPL parameter block. In this case, reipl_type_init()
selects IPL_TYPE_UNKNOWN, and reipl_block_actual remains NULL.

kdump passes the re-IPL parameter block to the dump kernel through
os_info. Before commit 3b9678472b ("s390/ipl: correct kdump reipl
block checksum calculation"), the os_info entry was added only for
IPL types which initialized reipl_block_actual. That commit moved the
os_info update to machine_crash_shutdown(), making it unconditional. As
a result, set_os_info_reipl_block() dereferences reipl_block_actual for
IPL_TYPE_UNKNOWN. This may happen to work by chance when address zero
contains readable lowcore data and the resulting empty os_info entry is
ignored by the dump kernel.

Skip the os_info update when no re-IPL parameter block is available.
Kdump then collect the dump and reboot without setting re-IPL parameter
block.

Fixes: 3b9678472b ("s390/ipl: correct kdump reipl block checksum calculation")
Reviewed-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
2026-08-31 16:25:38 +02:00
Heiko Carstens
ca1f4a5eca s390/time: Use jiffies instead of jiffies_64
Christoph Schlameuss and Alexander Egorenkov reported a data-race
reported by KCSAN when jiffies_64 is read:

==================================================================
BUG: KCSAN: data-race in do_account_vtime / tick_do_update_jiffies64

write to 0x0000016599ea8600 of 8 bytes by interrupt on cpu 6:
 tick_do_update_jiffies64+0x140/0x250
=============================================================>
BUG: KCSAN: data-race in do_account_vtime / tick_do_update_ji>

write to 0x0000016599ea8600 of 8 bytes by interrupt on cpu 6:
 tick_do_update_jiffies64+0x140/0x250
 tick_nohz_handler+0x2e6/0x300
 __run_hrtimer+0x156/0x4d0
 __hrtimer_run_queues+0xd2/0x150
...
 system_call+0x72/0x90

read to 0x0000016599ea8600 of 8 bytes by interrupt on cpu 12:
 do_account_vtime+0x7d6/0x860
 vtime_flush+0x26/0xe0
 update_process_times+0x32/0x160
 tick_nohz_handler+0x12a/0x300
...
 system_call+0x72/0x90

value changed: 0x00000000ffffaa6c -> 0x00000000ffffaa6d
...
=============================================================>

Problem is that jiffies_64 instead of jiffies is used. Both are at the
same address, but only jiffies is of volatile type, which prevents this
warning.

Change the vtime code so jiffies instead of jiffies_64 is used
everywhere. This addresses also the inconsistency that both jiffies and
jiffies_64 were used in the original patch which introduced this.

Fixes: f341b8dff9 ("s390/vtime: limit MT scaling value updates")
Reported-by: Christoph Schlameuss <schlameuss@linux.ibm.com>
Reported-by: Alexander Egorenkov <egorenar@linux.ibm.com>
Reviewed-by: Alexander Egorenkov <egorenar@linux.ibm.com>
Tested-by: Alexander Egorenkov <egorenar@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
2026-08-31 16:25:37 +02:00
Linus Torvalds
7bb6284aa7 Arm:
* Add support for 'slot' based PMU events, paired with new UAPI that
   compels the user to select a specific PMU implementation
 
 * Lazy save/restore of vCPU state for pKVM, along with various fixes
   and cleanups to the management of vCPU state between the untrusted
   host and pKVM hypervisor
 
 * Disable traps of EL1 registers for nested hypervisors when FEAT_NV2p1
   is present, guaranteeing that EL2-specific register bits are stateful
   in the EL1 counterpart
 
 * Leverage FEAT_NV3 to avoid unnecessary ERET/TLBI traps when the scope
   of those instructions remains 'in host' (i.e. L1 kernel/userspace)
 
 * Pile of fixes for the management of the VNCR pseudo-TLB, such as
   under-invalidations and races with concurrent TLBIs on other vCPUs
 
 * Consolidate the non-protected and pKVM view of ICH_VTR_EL2 to a
   runtime-patched constant, allowing the same data to be shared with
   pKVM prior to dropping host privileges
 
 * Considerable pile of LLM-assisted fixes around the shop but mostly in
   the VGIC, our in-kernel generator of bugs (and sometimes interrupts)
 
 LoongArch:
 
 * Advertise already-supported capabilities.
 
 * Some bug fixes about timer and MMIO.
 
 * Some hardening about interrupt injection.
 
 * Replace kvm_err() with kvm_pr_unimpl().
 
 * Add FPU/LSX/LASX test cases for selftests.
 
 RISC-V:
 
 * Svadu/Zicfiss/Zicfilp FWFT support for Guest
 
 * Use try_cmpxchg for IMSIC MRIF RMW
 
 * More arch-specific tracepoints in KVM RISC-V
 
 * Eager page splitting when enabling dirty logging
 
 * Optimize hfence request handling for SMP Guests
 
 * Improve dirty log clearing by skipping zero bits in mask
 
 * Guard HFENCE range loops against overflow
 
 * CPU PM notifiers in KVM RISC-V for non-retentive idle states
 
 * Fix kernel-mode vector context save/restore for Guest
 
 s390:
 
 * Fixes for vfio-ap
 
 * Fixes for the gmap rework
 
 * Fixes for vsie
 
 * AI triggered fixes all over
 
 * diag9c tracing
 
 * code move preparation for the additional arm64 support
 
 * enable CONTEXT_ANALYSIS
 
 x86:
 
 * Perform spring cleaning on x86.{c,h} and asm/kvm_host.h, by adding regs.c
   (the kvm_cache_regs.h => regs.h is already applied) and msrs.{c,h}, and moving
   relevant code out of x86.c.
 
 * Split kvm_mmu in three parts, respectively to describe the format of page
   tables, walking the guest page tables and building the page tables.  Always
   use the same page table walker kvm->arch.gva_walk as the entry point to
   convert a guest's virtual address, where the previous code used two
   different kvm_mmu structs depending on whether the walk included nested
   EPT/NPT or not.  Make page fault vmexits reuse the permission checking
   machinery that is used for guest page faults.  This is both a cleanup
   and a baby step towards supporting XS/XU memory permissions.
 
 * Document some of the "fun" gotchas with the APIC base when creating IRQCHIPs
   on x86.
 
 * Remove a defunct masterclock update from kvm_xen_shared_info_init().  It
   could result in incorrect kvmclock due to triggering an unnecessary
   switch to/from masterclock mode.
 
 * Skip Xen runstate time updates if time has effectively gone backwards, so
   that the guest doesn't report 100% steal time for a very, very long time.
 
 * Drop KVM's runtime updates of the Xen PV timing CPUID leaf, as KVM was
   updating the wrong sub-leaf, and upstream KVM will soon provide all the
   information needed by userspace to populate the CPUID field itself.
 
 * Fix a bug where KVM would walk a newly created rmap without holding the rmap
   lock (or mmu_lock) during aging.
 
 * Fix a bug where aging TDP MMU SPTEs could clobber FROZEN SPTEs.
 
 * Fix a variety of #DB priority bugs.
 
 * Fix a class of races related to enabling Hyper-V emulation on a vCPU after
   the vCPU is visible to the rest of KVM.
 
 * Use static calls for nested virtualization ops.
 
 * Move more KVM-internal code out of x86's kvm_host.h.
 
 * Enumerate support for a variety of Zhaoxin instructions that don't require
   explicit virtualization.
 
 * Fix missing EFER validation bugs, including in the KVM_SET_SREGS* path.
 
 * Harden kvm_vcpu_map() against double-mapping and thus leaking references.
 
 * Misc fixes and cleanups, e.g. for largely benign syzkaller splats.
 
 x86 (Intel):
 
 * Zero a vCPU's entry in VMX's Posted Interrupt Descriptor table used for IPI
   virtualization when the vCPU is freed, to fix a use-after-free where hardware
   will write to a freed vCPU's PID.
 
 * Service local TLB flushes on a failed nested VM-Enter to fix a bug where KVM
   could miss a TLB on a future, successful VM-Enter with the same L2 VPID.
 
 * Cap the maximum value shoved into the VMX Preemption Timer to workaround an
   erratum that affects all existing Intel CPUs that support CPUID 0x15.
 
 * Fix VPID virtualization bugs where KVM would fail to flush hardware TLBs.
 
 * Harden the TDX "populate" ioctls against bad input, and to prepare
   for supporting in-place private<=>shared conversion.
 
 x86 (AMD):
 
 * Forcefully invalidate SNP VMSA pages if their backing guest_memfd page is
   zapped/invalidated, e.g. due to a PUNCH_HOLE in response to a Page-State
   Change request.
 
 * Remove a dying VM from the GA Log notifier list before the VM is actually
   destroyed, to fix a potential use-after-free.
 
 * While FOLL_WRITE was needed in the past to trigger CoW unsharing, nowadays
   FOLL_LONGTERM does that already even without FOLL_WRITE, and in fact,
   get_user_pages() actually disallows FOLL_WRITE together with FOLL_LONGTERM.
   So don't pass FOLL_WRITE when registering encrypted memory regions, i.e. when
   pinning SEV/SEV-ES guest memory, to fix a regression with file-backed memory
   introduced by KVM's (correct) usage of long-term pins.
 
   (This was reviewed by mm maintainers; for more information, see commit
   ee1a586dd1).
 
 * Allocate full pages for SEV/SEV-ES {DE,EN}CRYPT ops on SNP-enabled hosts to
   fix a data corruption issue due to the PSP driver assigning to-be-written
   pages to firmware (as required by the SNP specs).
 
 * Unconditionally intercept ICBEP so that KVM generates the correct guest RIP
   when handling an ICEBP-induced TASK_SWITCH #VMEXIT.
 
 * Harden the SNP "populate" ioctls against bad input, and to prepare
   for supporting in-place private<=>shared conversion.
 
 Generic:
 
 * Remove kvm_debugfs_dir if kvm_init() fails after creating KVM's debugfs.
 
 * Add a per-VM bitmap to track which vCPU IDs have been "claimed" but for
   which the vCPU isn't yet online, and use the bitmap to reject duplicate IDs
   before calling into arch code.  This allows arch code to consume vcpu_id
   without having to worry about cross-vCPU clobbering (at least s390 and x86
   have had related bugs).
 
 * Rework the so called "prepare" and "invalidate" guest_memfd hooks to prepare
   for in-place private<=>shared conversion, and clean up a few warts along the
   way.
 
 Selftests:
 
 * Automatically allocate a full page for L2 guest stacks on x86 instead of
   requiring test-specific L1 guest code to carve out a portion of the L1
   stack for L2 usage, and to ensure the L2 stack also adheres to the x86-64
   calling convention ABI.
 
 * Add a selftest to verify {Guest,Host}-Only behavior in x86's mediated PMU.
 
 * Clean up nested SVM's handling of GPRs on L2<=>L1 transitions, reuse the
   functionality for nested VMX, and drop the ucall hack that was fudging
   around the lack of GPR switching on nVMX.
 
 * Add a stress test to verify KVM doesn't clobber/drop #PF state, e.g. CR2,
   across save/restore, including when L2 is active.
 
 * Add a test to verify KVM_CREATE_VM accepts exactly what is reported by
   KVM_CAP_VM_TYPES.
 
 * Misc selftests fixes and cleanups
 
 * Fix several issues with seeding the pRNG, and rework the pRNG APIs to that
   the pRNG can be sanely used in host code, not just guest code.
 
 * Add an IRQ test to validate virtual IRQ deliverty for IRQs wired up via
   KVM_IRQFD + KVM_SET_GSI_ROUTING, with optional support for triggering IRQs
   via writes to an assigned VFIO device.
 
 * Add syscall wrappers to assert success on a variety of pthreads and CPU
   affinity APIs.
 
 * Set vCPU pthread affinity as early as possible to reduce contention issues
   that were surfaced by PREEMPT_LAZY, which result in runtimes of over a
   minute on large hosts, versus the expected ~5 seconds.
 
 * Rework the PMU counters test to run each testcase using a single VM with
   many vCPUs for each sub-testcase, instead of using a unique VM for each
   sub-testcase.  This cuts the runtime by ~20x.
 
 Miscellaneous:
 
 * MAINTAINERS updates for vfio-ap, guest_memfd, kvm-x86.  Mostly representing
   the status quo more accurately, but also... welcome David Hildenbrand
   as guest_memfd reviewer!
 -----BEGIN PGP SIGNATURE-----
 
 iQFIBAABCAAyFiEE8TM4V0tmI4mGbHaCv/vSX3jHroMFAmqMgosUHHBib256aW5p
 QHJlZGhhdC5jb20ACgkQv/vSX3jHroP8bwf+ORImBMDM3QEmybZM3I+N2+xqSuHP
 QHttbmqGbsFK/RUeH96/X/+P9waqaz3uVeUQ6Qp2r0ryqwKtLt8YvIxKDp+M0vVJ
 n+iukk1xulBEc28aGdKHn9G4wayAwDA/9f7CvJ23hojaJfLScbF3OlFkDd7y5DpO
 x15Rtg9folYUjjop3LDML4N9/9Qmk4KRvVZ4ZVv6IB4uGJJ72fLd5dbBMyDk+BLl
 Lz9N1xVTXcnJXJmrjMB4/QNt/HiQJdun8LcokJZyykta7Xx6aY7OGZv+VmCeq47K
 e9Hk8mD4AyAdVbVvIntROSeBJOrlsgWJXAPHea6wEdRQnS3XVEBuZC0yDA==
 =rfLh
 -----END PGP SIGNATURE-----

Merge tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm

Pull kvm updates from Paolo Bonzini:
 "ARM64:

   - Add support for 'slot' based PMU events, paired with new UAPI that
     compels the user to select a specific PMU implementation

   - Lazy save/restore of vCPU state for pKVM, along with various fixes
     and cleanups to the management of vCPU state between the untrusted
     host and pKVM hypervisor

   - Disable traps of EL1 registers for nested hypervisors when
     FEAT_NV2p1 is present, guaranteeing that EL2-specific register bits
     are stateful in the EL1 counterpart

   - Leverage FEAT_NV3 to avoid unnecessary ERET/TLBI traps when the
     scope of those instructions remains 'in host' (i.e. L1
     kernel/userspace)

   - Pile of fixes for the management of the VNCR pseudo-TLB, such as
     under-invalidations and races with concurrent TLBIs on other vCPUs

   - Consolidate the non-protected and pKVM view of ICH_VTR_EL2 to a
     runtime-patched constant, allowing the same data to be shared with
     pKVM prior to dropping host privileges

   - Considerable pile of LLM-assisted fixes around the shop but mostly
     in the VGIC, our in-kernel generator of bugs (and sometimes
     interrupts)

  LoongArch:

   - Advertise already-supported capabilities

   - Some bug fixes about timer and MMIO

   - Some hardening about interrupt injection

   - Replace kvm_err() with kvm_pr_unimpl()

   - Add FPU/LSX/LASX test cases for selftests

  RISC-V:

   - Svadu/Zicfiss/Zicfilp FWFT support for Guest

   - Use try_cmpxchg for IMSIC MRIF RMW

   - More arch-specific tracepoints in KVM RISC-V

   - Eager page splitting when enabling dirty logging

   - Optimize hfence request handling for SMP Guests

   - Improve dirty log clearing by skipping zero bits in mask

   - Guard HFENCE range loops against overflow

   - CPU PM notifiers in KVM RISC-V for non-retentive idle states

   - Fix kernel-mode vector context save/restore for Guest

  s390:

   - Fixes for vfio-ap

   - Fixes for the gmap rework

   - Fixes for vsie

   - AI triggered fixes all over

   - diag9c tracing

   - code move preparation for the additional arm64 support

   - enable CONTEXT_ANALYSIS

  x86:

   - Perform spring cleaning on x86.{c,h} and asm/kvm_host.h, by adding
     regs.c (the kvm_cache_regs.h => regs.h is already applied) and
     msrs.{c,h}, and moving relevant code out of x86.c

   - Split kvm_mmu in three parts, respectively to describe the format
     of page tables, walking the guest page tables and building the page
     tables. Always use the same page table walker kvm->arch.gva_walk as
     the entry point to convert a guest's virtual address, where the
     previous code used two different kvm_mmu structs depending on
     whether the walk included nested EPT/NPT or not. Make page fault
     vmexits reuse the permission checking machinery that is used for
     guest page faults. This is both a cleanup and a baby step towards
     supporting XS/XU memory permissions

   - Document some of the "fun" gotchas with the APIC base when creating
     IRQCHIPs on x86

   - Remove a defunct masterclock update from kvm_xen_shared_info_init().
     It could result in incorrect kvmclock due to triggering an
     unnecessary switch to/from masterclock mode

   - Skip Xen runstate time updates if time has effectively gone
     backwards, so that the guest doesn't report 100% steal time for
     a very, very long time

   - Drop KVM's runtime updates of the Xen PV timing CPUID leaf, as KVM
     was updating the wrong sub-leaf, and upstream KVM will soon provide
     all the information needed by userspace to populate the CPUID field
     itself

   - Fix a bug where KVM would walk a newly created rmap without holding
     the rmap lock (or mmu_lock) during aging

   - Fix a bug where aging TDP MMU SPTEs could clobber FROZEN SPTEs

   - Fix a variety of #DB priority bugs

   - Fix a class of races related to enabling Hyper-V emulation on a
     vCPU after the vCPU is visible to the rest of KVM

   - Use static calls for nested virtualization ops

   - Move more KVM-internal code out of x86's kvm_host.h

   - Enumerate support for a variety of Zhaoxin instructions that don't
     require explicit virtualization

   - Fix missing EFER validation bugs, including in the KVM_SET_SREGS*
     path

   - Harden kvm_vcpu_map() against double-mapping and thus leaking
     references

   - Misc fixes and cleanups, e.g. for largely benign syzkaller splats

  x86 (Intel):

   - Zero a vCPU's entry in VMX's Posted Interrupt Descriptor table used
     for IPI virtualization when the vCPU is freed, to fix a
     use-after-free where hardware will write to a freed vCPU's PID

   - Service local TLB flushes on a failed nested VM-Enter to fix a bug
     where KVM could miss a TLB on a future, successful VM-Enter with
     the same L2 VPID

   - Cap the maximum value shoved into the VMX Preemption Timer to
     workaround an erratum that affects all existing Intel CPUs that
     support CPUID 0x15

   - Fix VPID virtualization bugs where KVM would fail to flush hardware
     TLBs

   - Harden the TDX "populate" ioctls against bad input, and to prepare
     for supporting in-place private<=>shared conversion

  x86 (AMD):

   - Forcefully invalidate SNP VMSA pages if their backing guest_memfd
     page is zapped/invalidated, e.g. due to a PUNCH_HOLE in response to
     a Page-State Change request

   - Remove a dying VM from the GA Log notifier list before the VM is
     actually destroyed, to fix a potential use-after-free

   - While FOLL_WRITE was needed in the past to trigger CoW unsharing,
     nowadays FOLL_LONGTERM does that already even without FOLL_WRITE,
     and in fact, get_user_pages() actually disallows FOLL_WRITE
     together with FOLL_LONGTERM. So don't pass FOLL_WRITE when
     registering encrypted memory regions, i.e. when pinning SEV/SEV-ES
     guest memory, to fix a regression with file-backed memory
     introduced by KVM's (correct) usage of long-term pins

     (This was reviewed by mm maintainers; for more information, see
     commit ee1a586dd1 "KVM: SEV: Drop FOLL_WRITE for encrypted region
     registration")

   - Allocate full pages for SEV/SEV-ES {DE,EN}CRYPT ops on SNP-enabled
     hosts to fix a data corruption issue due to the PSP driver
     assigning to-be-written pages to firmware (as required by the SNP
     specs)

   - Unconditionally intercept ICBEP so that KVM generates the correct
     guest RIP when handling an ICEBP-induced TASK_SWITCH #VMEXIT

   - Harden the SNP "populate" ioctls against bad input, and to prepare
     for supporting in-place private<=>shared conversion

  Generic:

   - Remove kvm_debugfs_dir if kvm_init() fails after creating KVM's
     debugfs

   - Add a per-VM bitmap to track which vCPU IDs have been "claimed" but
     for which the vCPU isn't yet online, and use the bitmap to reject
     duplicate IDs before calling into arch code. This allows arch code
     to consume vcpu_id without having to worry about cross-vCPU
     clobbering (at least s390 and x86 have had related bugs)

   - Rework the so called "prepare" and "invalidate" guest_memfd hooks
     to prepare for in-place private<=>shared conversion, and clean up a
     few warts along the way

  Selftests:

   - Automatically allocate a full page for L2 guest stacks on x86
     instead of requiring test-specific L1 guest code to carve out a
     portion of the L1 stack for L2 usage, and to ensure the L2 stack
     also adheres to the x86-64 calling convention ABI

   - Add a selftest to verify {Guest,Host}-Only behavior in x86's
     mediated PMU

   - Clean up nested SVM's handling of GPRs on L2<=>L1 transitions,
     reuse the functionality for nested VMX, and drop the ucall hack
     that was fudging around the lack of GPR switching on nVMX

   - Add a stress test to verify KVM doesn't clobber/drop #PF state,
     e.g. CR2, across save/restore, including when L2 is active

   - Add a test to verify KVM_CREATE_VM accepts exactly what is reported
     by KVM_CAP_VM_TYPES

   - Misc selftests fixes and cleanups

   - Fix several issues with seeding the pRNG, and rework the pRNG APIs
     to that the pRNG can be sanely used in host code, not just guest
     code

   - Add an IRQ test to validate virtual IRQ deliverty for IRQs wired up
     via KVM_IRQFD + KVM_SET_GSI_ROUTING, with optional support for
     triggering IRQs via writes to an assigned VFIO device

   - Add syscall wrappers to assert success on a variety of pthreads and
     CPU affinity APIs

   - Set vCPU pthread affinity as early as possible to reduce contention
     issues that were surfaced by PREEMPT_LAZY, which result in runtimes
     of over a minute on large hosts, versus the expected ~5 seconds

   - Rework the PMU counters test to run each testcase using a single VM
     with many vCPUs for each sub-testcase, instead of using a unique VM
     for each sub-testcase. This cuts the runtime by ~20x

  Miscellaneous:

   - MAINTAINERS updates for vfio-ap, guest_memfd, kvm-x86. Mostly
     representing the status quo more accurately, but also... welcome
     David Hildenbrand as guest_memfd reviewer!"

* tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: (413 commits)
  KVM: arm64: Validate GICv5 timer PPIs before claiming ownership
  KVM: arm64: vgic: Reject out-of-range GICv5 PPI IDs
  KVM: arm64: vgic: Prevent speculative SPI array underflow
  KVM: arm64: vgic: Free gic_kvm_info on initialization failure
  KVM: arm64: Avoid mismatched accesses to 'struct kvm_nvhe_init_params'
  s390/vfio-ap: Fix NULL deref in status_show() during queue probe
  s390/vfio-ap: Fix hot-unplug skipped when last AP adapter or domain removed
  s390/vfio-ap: fix potential use of uninitialized apm_filtered bitmap
  s390/vfio-ap: Fix control domain removal in vfio_ap_mdev_cfg_remove
  s390/vfio-ap: Fix required lock not held during update of ap_matrix_mdev object
  s390/vfio-ap: Fix missing lock required to access list of ap_matrix_mdev objects
  s390/vfio-ap: Fix dereference matrix_mdev->kvm without checking for NULL
  s390/vfio-ap: Fix stale do_remove flag across iterations in vfio_ap_mdev_cfg_remove
  RISC-V: KVM: fix vcpu vector context handling for kernel-mode vector
  riscv: vector: allow non-preemptible kernel-mode vector with IRQs off
  riscv: vector: refactor riscv_v_start_kernel_context
  KVM: s390: gmap: Make prefix handling optional
  KVM: s390: gmap: Make CMMA optional
  KVM: s390: gmap: Make storage keys optional
  KVM: s390: Prepare gmap for a second KVM implementation
  ...
2026-08-25 11:48:04 -07:00
Linus Torvalds
66ec24c5d7 s390 updates for 7.3 merge window
- Add a cpuidle driver with polling and enabled wait states using the
   existing CPU idle infrastructure and idle governor to improve latency
   for frequent sleep/wakeup cycles. Remove the obsolete tick delay
   heuristic and generic arch_needs_cpu() hook. Add the corresponding
   driver entry to MAINTAINERS
 
 - Add kCFI support using the generic support provided by Clang
 
 - Enable Clang CONTEXT_ANALYSIS for various architecture code and for
   char, PCI, CIO and virtio drivers. Add required lock annotations,
   exclude unsupported mm helpers and remove conditional PCI locking
 
 - Fix secure storage access exception handling and reintroduce
   DCACHE_WORD_ACCESS previously removed as a workaround
 
 - Fix cpum_cf perf crashes when CPUs are brought online while per-task
   events are active. Allocate and remove per-CPU counter data from CPU
   hotplug callbacks
 
 - Fix a deadlock when an s390dbf debug area is unregistered while one
   of its debugfs files is being written to
 
 - Fix MVIY_PERCPU() with binutils older than 2.39, where an assembler
   macro silently omitted an instruction needed to repair interrupted
   operations after CPU migration
 
 - Remove/replace cond_resched() calls which are no-ops with the supported
   s390 preemption models
 
 - Fix AP queue depth and maximum message length decoding according to
   the architecture. Current hardware is not affected, but future hardware
   could report values which were handled incorrectly
 
 - Reflect the configured CPU state in cpu_enabled_mask so deconfigured
   CPUs are not presented as available for onlining
 
 - Restore the vDSO GNU_EH_FRAME program header which was lost when the
   build switched to direct linker invocation, and mark it read-only
 
 - Add SCLP action qualifiers used by Spyre for card initialization,
   recoverable error and telemetry reporting
 
 - Move KMSAN interrupt flag helpers out of line to fix
   -Wstatic-in-inline build warnings
 
 - Use level-specific page table entry accessors for hugetlb entries and
   ptep_get() when accessing crashed kernel memory in kdump
 
 - Make forced AP bus rescans killable so that a user process blocked
   behind an ongoing scan can still be terminated with SIGKILL
 
 - Rework pkey ioctl error paths to remove duplicated cleanup code and
   avoid freeing error pointers
 
 - Allow the protected guest SWIOTLB buffer to be allocated outside the
   first 2GB. Also enable dynamic SWIOTLB growth and the coherent atomic
   pool fallback to improve I/O behavior when the initial pool is exhausted
 
 - Add program check statistics and spinlock contention tracepoints.
   Increase the lockdep chain capacity to keep lockdep enabled for complex
   code paths such as btrfs
 
 - Simplify IPL, trap and syscall code and remove the obsolete unistd_32.h
   generation entry
 -----BEGIN PGP SIGNATURE-----
 
 iQEzBAABCgAdFiEE3QHqV+H2a8xAv27vjYWKoQLXFBgFAmqLHuoACgkQjYWKoQLX
 FBhf2Qf+JlV+jQM1Lvn/Dj16vuQ77a4aP5C/OnLGMaTrrzbX420qU04yvC96v2Xu
 ux01aDU9VakonE74IT0NmrNo1VDUk8nSvIWUTB6GH7KvK76VEZN5Kkyn8TmeRmE0
 bZ0Fg7MgnhwdYijFDiX9w4rLyirwxs7vkScdJdJd0iKEdoZHXojGSjPDvmSpXght
 FgCszt+YOqu9MMf9B5oGAl+P40mgPTlm6M+ygoe2dX7qPQBUHLbDPTgZiWnKdXi2
 LPx0QPEha921ePDWrWz2HEqNetMfwGl12iertXddf1uzuK6LLObi0M5QrGw/ZbOy
 UJFM+AjFekTQyZPSunD4NWyCjglqrA==
 =XF7Y
 -----END PGP SIGNATURE-----

Merge tag 's390-7.3-1' of git://git.kernel.org/pub/scm/linux/kernel/git/s390/linux

Pull s390 updates from Vasily Gorbik:

 - Add a cpuidle driver with polling and enabled wait states using the
   existing CPU idle infrastructure and idle governor to improve latency
   for frequent sleep/wakeup cycles. Remove the obsolete tick delay
   heuristic and generic arch_needs_cpu() hook. Add the corresponding
   driver entry to MAINTAINERS

 - Add kCFI support using the generic support provided by Clang

 - Enable Clang CONTEXT_ANALYSIS for various architecture code and for
   char, PCI, CIO and virtio drivers. Add required lock annotations,
   exclude unsupported mm helpers and remove conditional PCI locking

 - Fix secure storage access exception handling and reintroduce
   DCACHE_WORD_ACCESS previously removed as a workaround

 - Fix cpum_cf perf crashes when CPUs are brought online while per-task
   events are active. Allocate and remove per-CPU counter data from CPU
   hotplug callbacks

 - Fix a deadlock when an s390dbf debug area is unregistered while one
   of its debugfs files is being written to

 - Fix MVIY_PERCPU() with binutils older than 2.39, where an assembler
   macro silently omitted an instruction needed to repair interrupted
   operations after CPU migration

 - Remove/replace cond_resched() calls which are no-ops with the
   supported s390 preemption models

 - Fix AP queue depth and maximum message length decoding according to
   the architecture. Current hardware is not affected, but future
   hardware could report values which were handled incorrectly

 - Reflect the configured CPU state in cpu_enabled_mask so deconfigured
   CPUs are not presented as available for onlining

 - Restore the vDSO GNU_EH_FRAME program header which was lost when the
   build switched to direct linker invocation, and mark it read-only

 - Add SCLP action qualifiers used by Spyre for card initialization,
   recoverable error and telemetry reporting

 - Move KMSAN interrupt flag helpers out of line to fix
   -Wstatic-in-inline build warnings

 - Use level-specific page table entry accessors for hugetlb entries and
   ptep_get() when accessing crashed kernel memory in kdump

 - Make forced AP bus rescans killable so that a user process blocked
   behind an ongoing scan can still be terminated with SIGKILL

 - Rework pkey ioctl error paths to remove duplicated cleanup code and
   avoid freeing error pointers

 - Allow the protected guest SWIOTLB buffer to be allocated outside the
   first 2GB. Also enable dynamic SWIOTLB growth and the coherent atomic
   pool fallback to improve I/O behavior when the initial pool is
   exhausted

 - Add program check statistics and spinlock contention tracepoints.
   Increase the lockdep chain capacity to keep lockdep enabled for
   complex code paths such as btrfs

 - Simplify IPL, trap and syscall code and remove the obsolete
   unistd_32.h generation entry

* tag 's390-7.3-1' of git://git.kernel.org/pub/scm/linux/kernel/git/s390/linux: (59 commits)
  s390/percpu: Fix MVIY_PERCPU() with older binutils
  s390/debug: Fix deadlock during unregister
  s390/cpum_cf: Handle CPU hotplug via prepare/dead callbacks
  s390: Enable CONTEXT_ANALYSIS for various directories
  s390/mm: Add __context_unsafe() attribute to gmap helper functions
  s390/mm: Add __context_unsafe() attribute to do_secure_storage_access()
  s390/sysinfo: Add context analysis attributes
  s390/irqflags: Add out-of-line definitions of arch_local_irq_*() for KMSAN
  s390/virtio: Enable CONTEXT_ANALYSIS
  s390/cio: Enable CONTEXT_ANALYSIS
  s390/vfio_ccw: Add __must_hold() attribute to vfio_ccw_sch_quiesce()
  s390/pci: Enable CONTEXT_ANALYSIS
  s390/pci: Rework __zpci_event_availability() to remove conditional locking
  s390/pci: Rework __zpci_event_error() to remove conditional locking
  s390/char: Enable CONTEXT_ANALYSIS
  s390/con3215: Add __must_hold() attribute to raw3215_make_room()
  s390/ap: Fix MAPML computation
  s390/cio: Remove cond_resched() calls
  s390: Remove cond_resched() calls
  KVM: s390: Remove cond_resched() calls
  ...
2026-08-23 10:26:44 -07:00
Linus Torvalds
3424d8c18a Generic entry code updates:
- Make syscall user dispatching configurable
 
     Not all architectures can makes use of syscall user dispatching. Allow
     them to disable the feature completely.
 
   - Consolidate stack randomization for the generic entry code and the
     architectures using it.
 
     Stack randomization on syscall entry was sprinkled throughout the
     architecture specific low level entry code and in some cases at the
     wrong points, e.g. before establishing state, which violates the
     non-instrumentable constraints of that code.
 
     Clean this up by integrating stack randomization into the generic entry
     code helpers so that it is invoked at the earliest possible point right
     after establishing state and converting all generic entry code using
     architecture over.
 
   - Clean up the syscall number handling in the generic entry code. It
     works correctly for architectures which have a separate return value
     storage in pt_regs, but fails to distinguish the case where user space
     handed in -1 as syscall number from the case where the entry code
     rejects it by returning -1 to the callers. Aside of that the return
     value functionality of those interfaces is not really intuitive.
 
     Fix this by separating the decision to reject a syscall (user dispatch,
     ptrace, seccomp ...) from the potential modification of the syscall
     number through these mechanisms.
 
     This solves most of the problems for architectures which do not have a
     separate return value storage in pt_regs except for the case where a
     tracepoint has a BPF script or a probe attached which overwrite both
     the syscall number and the return value. But that's a problem which
     cannot be solved in the generic code, that only can be addressed by
     separating the storage model in the affected architectures.
 -----BEGIN PGP SIGNATURE-----
 
 iQJEBAABCgAuFiEEQp8+kY+LLUocC4bMphj1TA10mKEFAmqCs10QHHRnbHhAa2Vy
 bmVsLm9yZwAKCRCmGPVMDXSYoaf6D/0ZBG1Yb0/C/6lrI185qPu38aGOROuAcxP+
 RV1O1x6C83w2hCLBH8LeswY2x4/iGbdftne/hfmvu8eNCE5MzBfYvXhLL4If75Tc
 IJ6C8uummnDmrT1TFuWHryTAfjyF28gt0+GGq0Zy5Hyz9b4CTJqOMx5u6KV4cZuJ
 odoNQpE/GlWo40wCSTYP/Tt5xONrogk2pMQtFyV8JEoaXkdYSj/V815yojEmofYU
 fmgPPO5/vOnZzE4b29gZyndXnU1Boah7r1l5fg7c9za376yCEEzh/ApPhovHyY0A
 t8zjnrtooZ27IUKbcsyycrAM14asfcmViDNDgaCj8ttBioQaCnxO1BpKWjVxEZhE
 AbM6q3Q66ER4Df6GNhZjPqT5Lr7E7+vLLarhXLWztsGQklIx4AFbrsa73hA20UC9
 1PSeMd45JSxH3yA8vMauXAGHFK1tD1V8Lgofu69+2Z3jtKB+aU0fqWeL1jesSEM0
 oCGhUb3hIC1pz3KVA0MGmNTm0yyQJYTGZL7wADYNV5NbxJVqXgo37qa/0n94Gf/4
 TG3OwY4Sb/H/sve7v/eY4IvxVh+xs3dLZP8ZoqMlPCp9JIxc6iNoe6VHqPI7PFnM
 fXwDtsy+bRF/SKnB/32qxnR7UJqmdNH3XIjd+lXWliKt6UYoC79/MEKN5DmJcO9P
 CykZUWa72A==
 =XUd9
 -----END PGP SIGNATURE-----

Merge tag 'core-entry-2026-08-17' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip

Pull generic entry code updates from Thomas Gleixner:

 - Make syscall user dispatching configurable

   Not all architectures can makes use of syscall user dispatching.
   Allow them to disable the feature completely.

 - Consolidate stack randomization for the generic entry code and the
   architectures using it.

   Stack randomization on syscall entry was sprinkled throughout the
   architecture specific low level entry code and in some cases at the
   wrong points, e.g. before establishing state, which violates the
   non-instrumentable constraints of that code.

   Clean this up by integrating stack randomization into the generic
   entry code helpers so that it is invoked at the earliest possible
   point right after establishing state and converting all generic entry
   code using architecture over.

 - Clean up the syscall number handling in the generic entry code. It
   works correctly for architectures which have a separate return value
   storage in pt_regs, but fails to distinguish the case where user
   space handed in -1 as syscall number from the case where the entry
   code rejects it by returning -1 to the callers. Aside of that the
   return value functionality of those interfaces is not really
   intuitive.

   Fix this by separating the decision to reject a syscall (user
   dispatch, ptrace, seccomp ...) from the potential modification of the
   syscall number through these mechanisms.

   This solves most of the problems for architectures which do not have
   a separate return value storage in pt_regs except for the case where
   a tracepoint has a BPF script or a probe attached which overwrite
   both the syscall number and the return value. But that's a problem
   which cannot be solved in the generic code, that only can be
   addressed by separating the storage model in the affected
   architectures.

* tag 'core-entry-2026-08-17' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: (23 commits)
  entry, treewide: Make syscall_enter_from_user_mode[_work]() indicate syscall execution
  entry: Make return type of syscall_trace_enter() bool
  entry: Rework trace_syscall_enter()
  entry: Rework syscall_audit_enter()
  syscall_user_dispatch: Introduce ARCH_SUPPORTS_SYSCALL_USER_DISPATCH
  entry: Fix seccomp bypass after ptrace with TSYNC
  x86/entry: Simplify the syscall number logic
  x86/entry: Get rid of the sys_ni_syscall() indirection
  x86/entry: Make syscall functions static
  ptrace, treewide: Rename ptrace_report_syscall_entry() to ptrace_report_syscall_permit_entry()
  seccomp, treewide: Rename and convert __secure_computing() to return boolean
  entry: Use syscall number instead of rereading it
  entry: Remove syscall_enter_from_user_mode()
  x86/syscall: Use [syscall_]enter_from_user_mode_randomize_stack()
  s390/syscall: Use enter_from_user_mode_randomize_stack()
  riscv/syscall: Use syscall_enter_from_user_mode_randomize_stack()
  powerpc/syscall: Use syscall_enter_from_user_mode_randomize_stack()
  loongarch/syscall: Use syscall_enter_from_user_mode_randomize_stack()
  entry: Provide [syscall_]enter_from_user_mode_randomize_stack()
  randomize_kstack: Provide add_random_kstack_offset_irqsoff()
  ...
2026-08-18 15:00:56 -07:00
Paolo Bonzini
1526a27e79 Merge tag 'kvm-s390-next-7.3-1' of git://git.kernel.org/pub/scm/linux/kernel/git/kvms390/linux into HEAD
KVM: s390: Features and Fixes for 7.3

- merged kvms390/master to pick up additional fixes that came too late
  for 7.2
- Fixes for vfio-ap
- Fixes for the gmap rework
- Fixes for vsie
- AI triggered fixes all over
- diag9c tracing
- code move preparation for the additional arm64 support
- enable CONTEXT_ANALYSIS
- update to vfio maintainer file location
2026-08-18 13:10:00 +02:00
Linus Torvalds
cd051cfe1e vfs-7.3-rc1.failfs
Please consider pulling these changes from the signed vfs-7.3-rc1.failfs tag.
 
 Thanks!
 Christian
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQRAhzRXHqcMeLMyaSiRxhvAZXjcogUCan7RJAAKCRCRxhvAZXjc
 ouz+AQCXKHb1Ay9ra1RG+dGu8mCpVZLebMt/+VO0/beMCqiqWAD8ChgvsFqObmr5
 8vLKOnzsSMeglRYGPL81h3xnaILRIQk=
 =0qwg
 -----END PGP SIGNATURE-----

Merge tag 'vfs-7.3-rc1.failfs' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs

Pull failfs filesystem from Christian Brauner:
 "Add failfs and expose a FD_FAILFS_ROOT sentinel.

  This allows userspace to shed their filesystem state completely. A
  process with its root or working directory in failfs must anchor every
  path lookup at an explicit file descriptor. Absolute paths, absolute
  symlinks and AT_FDCWD-relative lookups simply fail.

  Failfs is the counterpart to nullfs. nullfs says adds a permanently
  empty, immutable directory whose lookups fail with ENOENT but which
  can be opened, read, stat'd and mounted upon. Failfs on the other hand
  fails every operation. The root cannot be opened at all. A single
  instance is mounted during early boot via kern_mount(), which makes it
  logically distinct from every mount namespace.

  This is accompanied by a new fchroot() system call which makes
  chrooting via a file descriptor a first class concept. It's possible
  to chroot into failfs as an unprivileged user provided the task has no
  new privileges set"

* tag 'vfs-7.3-rc1.failfs' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs:
  Documentation: add failfs documentation
  selftests/filesystems: add failfs selftests
  arch: hookup fchroot() system call
  fs: support FD_FAILFS_ROOT in fchroot()
  fs: add fchroot()
  fs: support FD_FAILFS_ROOT in fchdir()
  fs: add failfs
2026-08-17 09:15:52 -07:00
Peter Oberparleiter
445c31ac63 s390/debug: Fix deadlock during unregister
Unregistering an s390dbf debug area while one of the associated debugfs
files is being written to can cause a deadlock:

$ echo >.../vmur/level    $ rmmod vmur
===================================================
debugfs write
debugfs_file_get()
                          debug_unregister()
                          mutex_lock(debug_mutex)
                          debugfs_remove()
                          wait for debugfs_file_put()
debug_file_ops.write()
debug_input()
mutex_lock(debug_mutex) ==> DEADLOCK

Fix this by splitting debug_unregister() into an s390dbf and debugfs
part, and running only the s390dbf part with debug_mutex locked.

Fixes: 9372a82892 ("s390/debug: fix debug area life cycle")
Signed-off-by: Peter Oberparleiter <oberpar@linux.ibm.com>
Reviewed-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-08-13 16:56:45 +02:00
Thomas Richter
337bd95507 s390/cpum_cf: Handle CPU hotplug via prepare/dead callbacks
The command 'perf stat -e cycles -- <command>' crashes the kernel
when CPUs are hotplug added during that run.

Root cause is the allocation of struct cpu_cf_events at first
event initialization. The allocation is dynamic and the first
event that has task context creates such a structure for
each online CPU. This is not sufficient. CPUs may be offline
during event creation and can be set online during the
perf run time. For example commands

 # echo 0 > /sys/devices/system/cpu/cpu1/online
 # perf stat -e cycles -i -- stress-ng -t10s --matrix X
 # sleep 1
 # echo 1 > /sys/devices/system/cpu/cpu1/online

create an event for CPUs 0,2-X. Since the events are created with
task-context, the scheduler will eventually schedule the program
on CPU1. This CPU has not created and initialized any per
CPU event infrastructure as that CPU was not online at the time
of the perf invocation. Thus when the scheduler runs stress-ng
on CPU1, the function cpumf_pmu_add() refers to a NULL pointer:

 struct cpu_cf_events *cpuhw = this_cpu_cfhw();

This function call is invoked after the task stress-ng has been
made runnable on CPU1. And this_cpu_cfhw() returns NULL.

The result is a panic:
Unable to handle kernel pointer dereference in virtual kernel address space
Failing address: 0000000000000000 TEID: 0000000000000483
....
Krnl PSW : 0404d00180000000 000003ef8291fd0c (cpumf_pmu_add+0x3c/0x80)
....
Call Trace:
 [<000003ef8291fd0c>] cpumf_pmu_add+0x3c/0x80
 [<000003ef82bb5e3e>] event_sched_in+0xae/0x190
 [<000003ef82bb60d6>] merge_sched_in+0x1b6/0x390
 [<000003ef82bb65b8>] visit_groups_merge.constprop.0.isra.0+0x308/0x5b0
 [<000003ef82bb689a>] pmu_groups_sched_in+0x3a/0x50
 [<000003ef82bb6a30>] ctx_sched_in+0x180/0x260
 [<000003ef82bb780c>] perf_event_context_sched_in+0x11c/0x2d0
 [<000003ef82bb79ee>] __perf_event_task_sched_in+0x2e/0xc0
 [<000003ef82994834>] finish_task_switch.isra.0+0x1a4/0x250
....
Last Breaking-Event-Address:
 [<000003ef8291f1d8>] this_cpu_cfhw+0x38/0x40

The issue arises only in per-task context when the CPUMF facility is
used and the scheduler picks a random CPU for such a process to run on.
The scheduler enables the CPUMF infrastructure via PMU callback
functions pmu::add() and pmu::del().

Introduce a CPU hotplug prepare/dead callback pair which creates and
removes the per CPU counter data while the CPU is offline. Count the
users which track every CPU (cpu == -1), that is perf_event_open()
events with task context and /dev/hwctr device sessions, in the new
counter cpu_cf_root::tskcnt, protected by pmc_reserve_mutex.
This ensures the infrastructure is available when
new CPU is selected to run the per-task context process.

In cpum_cf_free_root() and cpum_cf_free_cpu() ensure the reference
pointer to data structures is set to NULL before the data is freed
to prevent interrupt handlers to access stale data.

[gor@linux.ibm.com: change commit message]
Fixes: 9b9cf3c77e ("s390/cpum_cf: rework PER_CPU_DEFINE of struct cpu_cf_events")
Cc: stable@vger.kernel.org # v6.5+
Suggested-by: Heiko Carstens <hca@linux.ibm.com>
Suggested-by: Christian Borntraeger <borntraeger@linux.ibm.com>
Assisted-by: Claude:claude-sonnet-5
Signed-off-by: Thomas Richter <tmricht@linux.ibm.com>
Acked-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-08-13 16:56:45 +02:00
Heiko Carstens
29a63aa52d s390: Enable CONTEXT_ANALYSIS for various directories
Enable CONTEXT_ANALYSIS for various directories which do not generate
any warnings (anymore).

Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-08-11 16:42:18 +02:00
Heiko Carstens
1f45a8663a s390/sysinfo: Add context analysis attributes
Add context analysis attributes to service_level_start() and
service_level_stop() to specify that those functions only
acquire or release a lock.

Addresses the following warnings:

arch/s390/kernel/sysinfo.c:331:1: warning: rw_semaphore 'service_level_sem' is still held at the end of function
arch/s390/kernel/sysinfo.c:329:2: note: rw_semaphore acquired here
  329 |         down_read(&service_level_sem);
arch/s390/kernel/sysinfo.c:340:2: warning: releasing rw_semaphore 'service_level_sem' that was not held
  340 |         up_read(&service_level_sem);

Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-08-11 16:42:18 +02:00
Ilya Leoshkevich
3f7c9f9c36 s390/irqflags: Add out-of-line definitions of arch_local_irq_*() for KMSAN
Inline KMSAN arch_local_irq_*() definitions run afoul of
-Wstatic-in-inline. Move them out-of-line. Make sure decompressor and
non-GPL modules see the out-of-line definitions.

Cc: Boqun Feng <boqun@kernel.org>
Reported-by: kernel test robot <lkp@intel.com>
Closes: https://lore.kernel.org/oe-kbuild-all/202607131219.euJHPSJ5-lkp@intel.com/
Suggested-by: Heiko Carstens <hca@linux.ibm.com>
Fixes: 1b301f5f28 ("s390/irqflags: do not instrument arch_local_irq_*() with KMSAN")
Signed-off-by: Ilya Leoshkevich <iii@linux.ibm.com>
Reviewed-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-08-11 16:42:18 +02:00
Christian Borntraeger
546dde823a KVM: s390: Fix memory corruption by not reinjecting CK machine checks
Channel-subsystem damage machine checks are for the host channel
subsystem. The guest channel subsystem is emulated in the userspace VMM.
There is no point in forwarding such machine checks into the guest.

This also simplifies the machine check reinjection and avoids kfree of a
stack variable as reported by sashiko.  There might be still machine
checks that have the ck bit set with another bit (like instruction
damage), mask out the CK bit in s390_backup_mcck_info(), like the CP and
ED bits already are.

Fixes: 4d62fcc0b6 ("KVM: s390: Inject machine check into the guest")
Cc: stable@vger.kernel.org
Signed-off-by: Christian Borntraeger <borntraeger@linux.ibm.com>
Acked-by: Heiko Carstens <hca@linux.ibm.com>
Acked-by: Claudio Imbrenda <imbrenda@linux.ibm.com>
Signed-off-by: Claudio Imbrenda <imbrenda@linux.ibm.com>
Message-ID: <20260806145835.31818-1-borntraeger@linux.ibm.com>
2026-08-11 10:45:56 +02:00
Heiko Carstens
de8ca0119a s390: Remove cond_resched() calls
Since [1] cond_resched() is a no-op on s390. Remove all calls.

[1] commit 7dadeaa6e8 ("sched: Further restrict the preemption modes")

Reviewed-by: Vasily Gorbik <gor@linux.ibm.com>
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-08-05 15:16:54 +02:00
Heiko Carstens
97d86fc847 KVM: s390: Remove cond_resched() calls
Since [1] cond_resched() is a no-op on s390. Remove all calls.

This also entirely removes uv_call_sched() and replaces all call sites
with uv_call(), since both functions are identical after the removal
of cond_resched().

[1] commit 7dadeaa6e8 ("sched: Further restrict the preemption modes")

Reviewed-by: Claudio Imbrenda <imbrenda@linux.ibm.com>
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-08-05 15:16:54 +02:00
Heiko Carstens
4133c9d3f3 s390/diag: Generate CFI type information for assembly functions
Use SYM_TYPED_FUNC_START to generate __kcfi_typeid_ symbols for assembler
functions which are called indirectly. All assembler functions contained in
text_amode31.S are called indirectly and require such annotations.

Reviewed-by: Jens Remus <jremus@linux.ibm.com>
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
Tested-by: Nathan Chancellor <nathan@kernel.org>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-08-05 15:16:54 +02:00
Heiko Carstens
9ac687a280 s390: Add ftrace_stub_graph
This is the s390 variant of commit f3a0c23f25 ("riscv: Add
ftrace_stub_graph"):

"Commit 883bbbffa5 ("ftrace,kcfi: Separate ftrace_stub() and
ftrace_stub_graph()") added a separate ftrace_stub_graph function for
CFI_CLANG. Add the stub to fix FUNCTION_GRAPH_TRACER compatibility
with CFI."

Reviewed-by: Jens Remus <jremus@linux.ibm.com>
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
Tested-by: Nathan Chancellor <nathan@kernel.org>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-08-05 15:16:54 +02:00
Mete Durlu
faa39c2eca s390/smp: Reflect (de)configured CPUs to cpu_enabled_mask
On s390, CPUs can be in a state where it is not possible to hotplug
them online before certain prequisite steps. For example the CPUs which
get introduced during runtime of a system can posses a "deconfigured"
state which prevents them from being hotplugged online before they get
configured. Another case is when users set the configured state of CPUs
themselves via "chcpu" or sysfs attributes.

On s390 available CPUs are being registered as new devices via
smp_add_core() either during boot or after a CPU rescan (for newly added
CPUs during runtime). Registered CPUs are marked as enabled without
considering the configure states. Add necessary checks to smp_add_core()
and userspace configure attribute handler. Reflect the configured CPUs
to cpu_enabled_mask to correctly represent which CPUs can be hotplugged
online.

Signed-off-by: Mete Durlu <meted@linux.ibm.com>
Acked-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-07-31 13:31:30 +02:00
Heiko Carstens
aab0734094 KVM: s390: pv: Use VM_SPARSE area for guest variable storage area
The guest variable storage area is allocated with vmalloc and then
donated to the ultravisor. Any kernel access to that area will result
in a secure storage access exception (aka fault).

This is a problem if such a memory area is read via /proc/kcore. This
causes an exception via vread_iter() and results in an unexpected short
read. Avoid this by allocating a custom VM_SPARSE area. If such an area
is read, vread_iter() returns zeroes for the entire area.

Note that the function which frees the area does not update ptes. This
is intentional to allow for deferred / lazy pte updates and TLB flushing
like the generic vfree() code is doing that. See vunmap_pte_range().

This assumes that s390 will gain full support for lazy_mmu_mode_enable()
and lazy_mmu_mode_disable() in the future, since as of now the used
ptep_get_and_clear() in vunmap_pte_range() does indeed invalidate and
flush every single pte entry, but only for s390.

Tested-by: Christian Borntraeger <borntraeger@linux.ibm.com>
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
Reviewed-by: Christian Borntraeger <borntraeger@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-07-30 00:29:10 +02:00
Mete Durlu
bc5d0909ee s390/ipl: Improve readability
Use explicit decleration on all shutdown_action/shutdown_trigger
declerations and reformat shutdown_actions_list decleration to
improve readability. No functional changes.

Signed-off-by: Mete Durlu <meted@linux.ibm.com>
Acked-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-07-30 00:29:09 +02:00
Mete Durlu
654e97c9d7 s390/ipl: Use ARRAY_SIZE macro
Use ARRAY_SIZE macro instead of reimplementing it.

Signed-off-by: Mete Durlu <meted@linux.ibm.com>
Acked-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-07-30 00:29:09 +02:00
Jens Remus
6e17a45b3c s390/vdso: Use symbolic constants for the PHDR permission flags
While at it explicitly specify GNU_EH_FRAME PHDR to be read-only.

Inspired by x86 commit 8717b02b8c ("x86/entry/vdso: Include
GNU_PROPERTY and GNU_STACK PHDRs").

Reviewed-by: Ilya Leoshkevich <iii@linux.ibm.com>
Acked-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Jens Remus <jremus@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-07-30 00:29:08 +02:00
Jens Remus
dc161efb6d s390/vdso: Pass --eh-frame-hdr to the linker
Commit 2b2a25845d ("s390/vdso: Use $(LD) instead of $(CC) to link
vDSO") accidentally broke the GNU_EH_FRAME program table entry in
the vDSO, causing it to be empty:

  $ readelf --program-headers arch/s390/kernel/vdso/vdso.so
  ...
  Program Headers:
    Type           Offset             VirtAddr           PhysAddr
                   FileSiz            MemSiz              Flags  Align
  ...
    GNU_EH_FRAME   0x0000000000000000 0x0000000000000000 0x0000000000000000
                   0x0000000000000000 0x0000000000000000         0x8
  ...

Originally, the compiler would implicitly add --eh-frame-hdr when
invoking the linker, but when this Makefile was converted from invoking
the linker via the compiler, to invoking it directly, the option was
missed.

This is the s390 variant of x86 commit cd01544a26 ("x86/vdso: Pass
--eh-frame-hdr to the linker").

Fixes: 2b2a25845d ("s390/vdso: Use $(LD) instead of $(CC) to link vDSO")
Reviewed-by: Ilya Leoshkevich <iii@linux.ibm.com>
Acked-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Jens Remus <jremus@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-07-30 00:29:08 +02:00
Sven Schnelle
5c92744c1f s390/syscalls: Use define instead of '1' to indicate PER trap
Make the code a bit easier to read by defining SYSCALL_PER_TRAP
instead of passing '1' to __do_syscall().

Suggested-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Sven Schnelle <svens@linux.ibm.com>
Reviewed-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-07-30 00:29:08 +02:00
Sven Schnelle
f588c5cb0b s390/traps: Remove PIF_GUEST_FAULT
PIF_GUEST_FAULT is only used to pass information whether a fault was
caused when executing SIE or when executing host code. Instead of
using ptregs for this, just pass the flag directly as argument to
__do_pgm_check(). This also saves the time required to read the flag
from ptregs, although this likely isn't much as it is already in the
data cache.

Signed-off-by: Sven Schnelle <svens@linux.ibm.com>
Reviewed-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-07-30 00:29:08 +02:00
Christian Brauner
79d27fd718
arch: hookup fchroot() system call
Wire up the fchroot() system call as number 472 on (nearly) all
architectures and sync the mirrored copies of the syscall tables and
the asm-generic unistd.h under tools/.

Link: https://patch.msgid.link/20260724-work-failfs-v2-5-485dabbae185@kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-07-27 17:18:00 +02:00
Linus Torvalds
d326f83e81 Lots of fixes, double the count even for the "new normal".
Largely due to my time off followed by a networking conference
 which distracted most maintainers (less so the AI generators).
 
 Including fixes from Bluetooth and WiFi.
 
 Current release - regressions:
 
  - wifi: mt76: fix MAC address for non OF pcie cards
 
 Current release - new code bugs:
 
  - mptcp: fix BUILD_BUG_ON on legacy ARM config
 
  - wifi: cfg80211: guard optional PMSR nominal time
 
 Previous releases - regressions:
 
  - qrtr: ns: raise node count limit to 512, we arbitrarily picked
    256 as a limit, turns out it was too low for real world deployments
 
  - vhost-net: fix TX stall when vhost owns virtio-net header
 
  - eth: amd-xgbe: fix MAC_AUTO_SW handling in CL37 AN
 
  - wifi: ath12k: fix low MLO RX throughput on WCN7850
 
 Previous releases - always broken:
 
  - number of random AI fixes for SCTP, RDS and TIPC protocols
 
  - more AI-looking fixes for WiFi drivers
 
  - number of fixes for missing pointer reloading after skb pull
 
  - reject BPF redirect use from qdisc qevent block
 
  - tcp: initialize standalone TCP-AO response padding
 
  - vsock/virtio: collapse receive queue under memory pressure to avoid
    client OOMing the host with tiny messages
 
  - ipv4: icmp: fill flow parameters in icmp_route_lookup decoy lookup,
    make sure the ICMP response routing follows the routing policy
 
  - gro: fix double aggregation of flush-marked skbs
 
  - ovpn: fix various refcount bugs
 
  - tls: device: push pending open record on splice EOF
 
  - eth: mlx5:
   - use sender devcom for MPV master-up
   - fix MCIA register buffer overflow on 32 dword reads
 
 Signed-off-by: Jakub Kicinski <kuba@kernel.org>
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEE6jPA+I1ugmIBA4hXMUZtbf5SIrsFAmpiX68ACgkQMUZtbf5S
 IrvYsRAAkuMhUpz0Ss9aF7rBY8iTp4SofSvFeVe06ywraUfqPuflGlak07t1Lz/i
 G4MuKXN0q8m+B0EZddfMeYw6rCGd0SCtFAkxUI3dd+pu4hssgioaCPL193drSsfC
 /lYeacjVL45jNrQvAWwKsRaAs3xdwzxWf0ddIXWvVWbdDsVfIf/mYahSS3TvniWw
 MQtEbWPnFwPvOrHzb+1ChLELCtig/yvK+3xS9JrwOkjUF4BczOUgqrYlG5MWerXP
 f/JDLsegPcoZaTycW5F5fshY05umeRQza/zCFqMKQNcQux49fjREnYxBuyTacVCo
 0cxhsNbKOhvBpBFNsHA6TjUbDxuiyL8L/g3e7VOlQFxI4hX3IMsnsP+UrSdE2zyG
 lgFAQ6HIcelgFnzFcwp9YEGsiZ5nDoJKe5aBcgftzTFPx3Plh1UeCrNjYtJawcjk
 1POovopI+G6eszwluVOoucUdDD3wf0jPgDqvdOcI9P9FVTsFmvRESsfen7NbdjG0
 v5mk9+sasWL1dns6mre6nt5is4QWSg7PDjufQUhuPKSSEnld+csEgmyxmUm0/FgL
 krUZLHdx0Yj9yIOAIYAvz8QoW9jHIyK05Mr7CoL4a/9RJ4rtxjb+3CT9qebeyd49
 jK5uzYX6tPHvILFK4CgZwcE/z9S+DoxCAuDEp6LhfstKsJW4KIM=
 =DIF8
 -----END PGP SIGNATURE-----

Merge tag 'net-7.2-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net

Pull networking fixes from Jakub Kicinski:
 "Lots of fixes, double the count even for the 'new normal'. Largely due
  to my time off followed by a networking conference which distracted
  most maintainers (less so the AI generators).

  Including fixes from Bluetooth and WiFi.

  Current release - regressions:

   - wifi: mt76: fix MAC address for non OF pcie cards

  Current release - new code bugs:

   - mptcp: fix BUILD_BUG_ON on legacy ARM config

   - wifi: cfg80211: guard optional PMSR nominal time

  Previous releases - regressions:

   - qrtr: ns: raise node count limit to 512, we arbitrarily picked
     256 as a limit, turns out it was too low for real world deployments

   - vhost-net: fix TX stall when vhost owns virtio-net header

   - eth: amd-xgbe: fix MAC_AUTO_SW handling in CL37 AN

   - wifi: ath12k: fix low MLO RX throughput on WCN7850

  Previous releases - always broken:

   - number of random AI fixes for SCTP, RDS and TIPC protocols

   - more AI-looking fixes for WiFi drivers

   - number of fixes for missing pointer reloading after skb pull

   - reject BPF redirect use from qdisc qevent block

   - tcp: initialize standalone TCP-AO response padding

   - vsock/virtio: collapse receive queue under memory pressure to avoid
     client OOMing the host with tiny messages

   - ipv4: icmp: fill flow parameters in icmp_route_lookup decoy lookup,
     make sure the ICMP response routing follows the routing policy

   - gro: fix double aggregation of flush-marked skbs

   - ovpn: fix various refcount bugs

   - tls: device: push pending open record on splice EOF

   - eth: mlx5:
      - use sender devcom for MPV master-up
      - fix MCIA register buffer overflow on 32 dword reads"

* tag 'net-7.2-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (234 commits)
  drop_monitor: perform u64_stats updates under IRQ-disabled section
  drop_monitor: fix size calculations for 64-bit attributes
  net: drop_monitor: fix info leak in NET_DM_ATTR_PAYLOAD
  mptcp: fix BUILD_BUG_ON on legacy ARM config
  selftests: mptcp: userspace_pm: fix undefined variable port
  mptcp: fix stale skb->sk reference on subflow close
  mptcp: pm: userspace: fix use-after-free in get_local_id
  mptcp: decrement subflows counter on failed passive join
  mac802154: hold an interface reference across the scan worker
  sctp: don't free the ASCONF's own transport in DEL-IP processing
  phonet: check register_netdevice_notifier() error in phonet_device_init()
  phonet: pep: fix use-after-free in pep_get_sb()
  bnge/bng_re: fix ring ID widths
  tipc: fix integer overflow in tipc_recvmsg() and tipc_recvstream()
  net: airoha: fix ETS channel derivation in airoha_tc_setup_qdisc_ets()
  mctp: check register_netdevice_notifier() error in mctp_device_init()
  ptp: netc: explicitly clear TMR_OFF during initialization
  rds: tcp: unregister sysctl before tearing down listen socket
  ipv6: Change allocation flags to match rcu_read_lock section requirements
  net: slip: serialize receive against buffer reallocation
  ...
2026-07-23 12:58:08 -07:00
Sven Schnelle
9de445d829 s390/ptff: Export ptff_function_mask[]
Export the ptff_function_mask to make ptff_query() usable in modules.

Signed-off-by: Sven Schnelle <svens@linux.ibm.com>
Acked-by: Heiko Carstens <hca@linux.ibm.com>
Link: https://patch.msgid.link/20260714130342.1971700-2-svens@linux.ibm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-07-23 07:02:34 -07:00
Thomas Gleixner
05c033db7e entry, treewide: Make syscall_enter_from_user_mode[_work]() indicate syscall execution
The return values of syscall_enter_from_user_mode[_work]() are
non-intuitive. Both functions return the syscall number which should be
invoked by the architecture specific syscall entry code. The returned
number can be:

  - the unmodified syscall number which was handed in by the caller

  - a modified syscall number (ptrace, seccomp, trace/probe/bpf)

That has an additional twist. If the return value is -1L then the caller is
not allowed to modify the return value as that indicates that the modifying
entity requests to abort the syscall and set the return value already. That
can obviously not be differentiated from a syscall which handed in -1 as
syscall number.

The most trivial way to deal with that is:

    set_return_value(regs, -ENOSYS);
    nr = syscall_enter_from_user_mode(regs, nr);
    if (valid(nr))
    	handle_syscall(regs, nr);

That's what LOONGARCH, RISCV, and X86 do. But PowerPC and S390 do not
preset the return value, so when user space hands in -1 and there is
nothing setting the return value in the entry work code, then the syscall
is skipped but the return value is whatever random data has been in the
return value register.

Change the return values of syscall_enter_from_user_mode[_work]() to
boolean and return false, when either ptrace or seccomp request to skip the
syscall. If they return true, update the syscall number as it might have
been changed.

That results in slightly different behaviour of the architectures versus
tracing.

If the syscall tracepoint has probe/BPF attached, those might set the
syscall number to -1 and also set the return value. PowerPC and S390 will
then overwrite that value with -ENOSYS. The other architectures will just
ignore it like any other invalid syscall and use the modified one.

Originally-by: Michal Suchánek <msuchanek@suse.de>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Tested-by: Michal Suchánek <msuchanek@suse.de>
Link: https://patch.msgid.link/20260712141346.772209074@kernel.org
2026-07-20 20:38:40 +02:00
Sumanth Korikkar
49145bce53 s390/perf_cpum_cf: Add missing array_index_nospec() to __hw_perf_event_init()
ev variable is userspace controlled via event->attr.config and used
as an array index after bounds checking, but without speculation
barriers.

Add the missing array_index_nospec() call to prevent speculative
execution.

Cc: stable@vger.kernel.org
Fixes: 212188a596 ("[S390] perf: add support for s390x CPU counters")
Signed-off-by: Sumanth Korikkar <sumanthk@linux.ibm.com>
Reviewed-by: Ilya Leoshkevich <iii@linux.ibm.com>
Acked-by: Thomas Richter <tmricht@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-07-15 17:35:42 +02:00
Thomas Gleixner
05a194d8bd s390/syscall: Use enter_from_user_mode_randomize_stack()
enter_from_user_mode_randomize_stack() replaces enter_from_user_mode() and
the subsequent invocation of add_random_kstack_offset_irqsoff().

As a bonus this avoids the overhead of get/put_cpu_var() in
add_random_kstack_offset().

No functional change.

Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Tested-by: Mukesh Kumar Chaurasiya (IBM) <mkchauras@gmail.com>
Reviewed-by: Sven Schnelle <svens@linux.ibm.com>
Reviewed-by: Radu Rendec <radu@rendec.net>
Reviewed-by: Jinjie Ruan <ruanjinjie@huawei.com>
Reviewed-by: Mukesh Kumar Chaurasiya (IBM) <mkchauras@gmail.com>
Reviewed-by: Philippe Mathieu-Daudé <philmd@oss.qualcomm.com>
Link: https://patch.msgid.link/20260707190254.030598804@kernel.org
2026-07-12 12:38:01 +02:00
Sven Schnelle
a7325d0d77 s390/traps: Add exception statistics
Add a new debugfs file which displays the number of exceptions (program
checks) per CPU. This is helpful for debugging purposes.

The statistics are typically available at
/sys/kernel/debug/s390/exceptions.

[ hca@linux.ibm.com: Forward ported code, changed file location ]

Suggested-by: Christian Borntraeger <borntraeger@linux.ibm.com>
Signed-off-by: Sven Schnelle <svens@linux.ibm.com>
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
Tested-by: Christian Borntraeger <borntraeger@linux.ibm.com>
Reviewed-by: Christian Borntraeger <borntraeger@linux.ibm.com>>
Signed-off-by: Alexander Gordeev <agordeev@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-07-08 17:03:28 +02:00
Mete Durlu
71eabd104e s390/tick: Remove CIF_NOHZ_DELAY flag
Remove obsolete tick delay heuristic [1]. The upcoming cpuidle driver
handles frequent sleep/wakeup cycles more effectively.

[1] https://lore.kernel.org/all/20090929122533.402715150@de.ibm.com/

Suggested-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Mete Durlu <meted@linux.ibm.com>
Reviewed-by: Heiko Carstens <hca@linux.ibm.com>
Acked-by: Rafael J. Wysocki (Intel) <rafael@kernel.org>
Signed-off-by: Alexander Gordeev <agordeev@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-07-08 17:03:28 +02:00
Bastian Blank
7d5c2f6791 s390: Add build salt to the vDSO
The vDSO needs to have a unique build id in a similar manner
to the kernel and modules. Use the build salt macro.

Signed-off-by: Bastian Blank <waldi@debian.org>
Reviewed-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Alexander Gordeev <agordeev@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-07-08 17:02:48 +02:00
Heiko Carstens
b7577fe4c4 s390/diag: Add missing array_index_nospec() call to memtop_get_page_count()
'level' is user space controlled and used to read from an array. Add the
missing array_index_nospec() call to prevent speculative execution.

Cc: stable@vger.kernel.org
Fixes: 0d30871739 ("s390/diag: Add memory topology information via diag310")
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
Reviewed-by: Mete Durlu <meted@linux.ibm.com>
Signed-off-by: Vasily Gorbik <gor@linux.ibm.com>
2026-07-08 17:02:47 +02:00
Heiko Carstens
6938846872 s390/idle: Add missing EXPORT_SYMBOL_GPL()
Uwe Kleine-König reported this build breakage caused by a recent commit
which provides arch specific kcpustat_field_idle()/kcpustat_field_iowait()
functions:

ERROR: modpost: "arch_kcpustat_field_idle" [drivers/leds/trigger/ledtrig-activity.ko] undefined!
ERROR: modpost: "arch_kcpustat_field_iowait" [drivers/leds/trigger/ledtrig-activity.ko] undefined!

Fix this by adding the missing EXPORT_SYMBOL_GPL().

Fixes: 670e057744 ("s390/idle: Provide arch specific kcpustat_field_idle()/kcpustat_field_iowait()")
Reported-by: Uwe Kleine-König <u.kleine-koenig@baylibre.com>
Closes: https://lore.kernel.org/r/ajKsG0JP6qTssQBX@monoceros
Acked-by: Alexander Gordeev <agordeev@linux.ibm.com>
Tested-by: Uwe Kleine-König <u.kleine-koenig@baylibre.com>
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
Signed-off-by: Alexander Gordeev <agordeev@linux.ibm.com>
2026-06-18 14:44:44 +02:00
Alexander Gordeev
9fb519794e Merge branch 'idle-time-acc' into features
Heiko Carstens says:

===================
This is supposed to improve s390 idle time accounting, and brings it
back to the state it was before arch_cpu_idle_time() was removed from
s390 [3].

In result all cpu time accounting is done by the s390 architecture backend
again, instead of having a mix of architecure specific and common code
accounting (common code: idle, s390 architecture: everything else).
===================

Signed-off-by: Alexander Gordeev <agordeev@linux.ibm.com>
2026-06-16 16:21:46 +02:00
Linus Torvalds
25a01b5155 s390 updates for 7.2 merge window
- Use CIO device online variable instead of the internal FSM state to
   determine device availability during purge operations
 
 - Remove extra check of task_stack_page() because try_get_task_stack()
   already takes care of that when reading /proc/<pid>/wchan
 
 - Allow user-space to use the new SCLP action qualifier 4 for to
   provide NVMe SMART log data to the platform.
 
 - Send AP CHANGE uevents on successful bind and successful association
   to notify user-space about SE operations on AP queue devices
 
 - Add an s390dbf kernel parameter to configure debug log levels and
   area sizes during early boot
 
 - On arm64 the empty zero page is going to be mapped read-only.
   Do the same for s390 with an explicit set_memory_ro() call
 
 - Improve s390-specific bcr_serialize() and cpu_relax() implementations
 
 - Remove all unused variables to avoid allmodconfig W=1 build fails
   with latest clang-23
 
 - Cleanup default Kconfig values for s390 selftests
 
 - Add a s390-tod trace clock to allow comparing trace timestamps
   between different systems or virtual machines on s390
 
 - Remove the s390 implementation of strlcat() in favor of the
   generic variant
 
 - Make consistent the calling order between page_table_check_pte_clear()
   and secure page conversion across all code paths
 
 - Rearrange some fields within AP and zcrypt structs to reduce
   memory consumption and unused holes
 
 - Shorten GR_NUM and VX_NUM macros and move them to a separate header
 
 - Replace __get_free_page() with kmalloc() in few sources
 
 - Introduce an infrastructure for more efficient this_cpu operations.
   Eliminate conditional branches when PREEMPT_NONE is removed
 
 - Enable Rust support
 
 - Use z10 as minimum architecture level, similar to the boot code,
   to enforce a defined architecture level set
 
 - Improve and convert various mem*() helper functions to C. For that
   add .noinstr.text section to avoid orphaned warnings from the linker
 
 - Fix the function pointer type in __ret_from_fork() to correct
   the indirect call to match kernel thread return type of int
 
 - Revert support for DCACHE_WORD_ACCESS to avoid an endless exception
   loop on read from donated Ultravisor pages at unaligned addresses
 -----BEGIN PGP SIGNATURE-----
 
 iI0EABYKADUWIQQrtrZiYVkVzKQcYivNdxKlNrRb8AUCai/rTRccYWdvcmRlZXZA
 bGludXguaWJtLmNvbQAKCRDNdxKlNrRb8KNOAPwMpGVtXcKF4HftCv49X0WpqKbU
 tdYO9hbq9wanIGpgIgEAk5vggxe74pj+palTbtCDteVjDpnSp811x8gfmLlPrgU=
 =qzZ2
 -----END PGP SIGNATURE-----

Merge tag 's390-7.2-1' of gitolite.kernel.org:pub/scm/linux/kernel/git/s390/linux

Pull s390 updates from Alexander Gordeev:

 - Use CIO device online variable instead of the internal FSM state to
   determine device availability during purge operations

 - Remove extra check of task_stack_page() because try_get_task_stack()
   already takes care of that when reading /proc/<pid>/wchan

 - Allow user-space to use the new SCLP action qualifier 4 for to
   provide NVMe SMART log data to the platform.

 - Send AP CHANGE uevents on successful bind and successful association
   to notify user-space about SE operations on AP queue devices

 - Add an s390dbf kernel parameter to configure debug log levels and
   area sizes during early boot

 - On arm64 the empty zero page is going to be mapped read-only. Do the
   same for s390 with an explicit set_memory_ro() call

 - Improve s390-specific bcr_serialize() and cpu_relax() implementations

 - Remove all unused variables to avoid allmodconfig W=1 build fails
   with latest clang-23

 - Cleanup default Kconfig values for s390 selftests

 - Add a s390-tod trace clock to allow comparing trace timestamps
   between different systems or virtual machines on s390

 - Remove the s390 implementation of strlcat() in favor of the generic
   variant

 - Make consistent the calling order between
   page_table_check_pte_clear() and secure page conversion across all
   code paths

 - Rearrange some fields within AP and zcrypt structs to reduce memory
   consumption and unused holes

 - Shorten GR_NUM and VX_NUM macros and move them to a separate header

 - Replace __get_free_page() with kmalloc() in few sources

 - Introduce an infrastructure for more efficient this_cpu operations.
   Eliminate conditional branches when PREEMPT_NONE is removed

 - Enable Rust support

 - Use z10 as minimum architecture level, similar to the boot code, to
   enforce a defined architecture level set

 - Improve and convert various mem*() helper functions to C. For that
   add .noinstr.text section to avoid orphaned warnings from the linker

 - Fix the function pointer type in __ret_from_fork() to correct the
   indirect call to match kernel thread return type of int

 - Revert support for DCACHE_WORD_ACCESS to avoid an endless exception
   loop on read from donated Ultravisor pages at unaligned addresses

* tag 's390-7.2-1' of gitolite.kernel.org:pub/scm/linux/kernel/git/s390/linux: (52 commits)
  s390: Revert support for DCACHE_WORD_ACCESS
  s390/process: Fix kernel thread function pointer type
  s390/tishift: Convert __ashlti3(), __ashrti3(), __lshrti3() to C
  s390/memmove: Optimize backward copy case
  s390/string: Convert memset(16|32|64)() to C
  s390/string: Convert memcpy() to C
  s390/string: Convert memset() to C
  s390/string: Convert memmove() to C
  s390/string: Add -ffreestanding compile option to string.o
  s390: Add .noinstr.text to boot and purgatory linker scripts
  s390/purgatory: Enforce z10 minimum architecture level
  s390: Enable Rust support
  s390/cmpxchg: Fix KASAN stack-out-of-bounds in atomic helpers
  rust: helpers: Add memchr wrapper for string operations
  rust/bindgen_parameters: Mark s390 types as opaque to prevent repr conflicts
  s390/jump_label: Implement ARCH_STATIC_BRANCH_JUMP_ASM and ARCH_STATIC_BRANCH_ASM macros
  s390/bug: Provide ARCH_WARN_ASM for Rust WARN/BUG support
  s390/ap: Fix locking issue in SE bind and associate sysfs functions
  s390/percpu: Provide arch_this_cpu_write() implementation
  s390/percpu: Provide arch_this_cpu_read() implementation
  ...
2026-06-16 05:08:13 +05:30
Heiko Carstens
968ec3f960 s390/idle: Remove idle time and count sysfs files
Remove the s390 specific idle_time_us and idle_count per cpu sysfs
files. They do not provide any additional value. The risk that there
are existing applications which rely on these architecture specific
files should be very low.

However if it turns out such applications exist, this can be easily
reverted.

Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
Acked-by: Frederic Weisbecker <frederic@kernel.org>
Signed-off-by: Alexander Gordeev <agordeev@linux.ibm.com>
2026-06-15 16:33:40 +02:00