mirror of
https://github.com/torvalds/linux.git
synced 2026-10-04 02:09:03 +02:00
673dab7eac
53557 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
673dab7eac |
Scheduler fixes:
- Fix LLC mis-scheduling bugs (Tim Chen, Lu Wang)
- Fix cache-grouping related scheduling statistics UAF bugs (Tim Chen)
- Skip kernel threads for cache aware scheduling to rubustify the code
(Chen Yu)
- Refresh LLC capacity across CPU hotplug, to fix capacity
underestimation bug (Davi Chaves Azevedo)
- Account PSI IRQ time to the execution context, not the scheduling
context, to fix proxy scheduling accounting bug (Zhan Xusheng)
Signed-off-by: Ingo Molnar <mingo@kernel.org>
-----BEGIN PGP SIGNATURE-----
iQJFBAABCgAvFiEEBpT5eoXrXCwVQwEKEnMQ0APhK1gFAmq4zcIRHG1pbmdvQGtl
cm5lbC5vcmcACgkQEnMQ0APhK1iN/Q/9HCvfdQ0tTjZHLVhL5f7ZnIjwxVNzmKTK
ysqfXacW3KIwuulXjEDXgAXRi85nKiAxV4sHVp8WaDfpu4NEVzekXZV7dHPf8fPZ
9LCfFJrrYhcAhKHwEIEyGKn7uSSYzFNVjUxHVfc7W7JtqNtrXOYmYKtJuaMdE8ya
fxV/wGHw3HHCmDv/v4BurqOgNyPFEzhN6vl875J0qEdsauxakwvlcIJsAWDstbxi
4kDF+47lLyDZoolrWsVCl8BoxERznK8JelmNVx0FkEpKi2DnLarWjkpAnjfGkIFg
lz2jxdQnTW5v60BTLw4yAUIam8SAzbDNnfKexzT3oG1UgdyqEOMZ4bpynOhoCu5b
3c/2E7Ah7zfMnyQrItnZ7JaXiCrHgPM+aHxjskmMV08Zr54S5hbTxws7uwfl/MA0
5r64GzrStcjswNZKKQycDFCwtkZIcsBmSyJvJqb3RPm4NVYHIBJWDd1YLlH7pM8R
BuNiW0zM9IdXl6EgD2gX41WoeESobaPPPb67b0DQ6WRSZkz7UAF0Xzfz/xu66WlO
AwMMFHF1qhAcBkUBYU9R+hGSQoaTClf7f2stuSr4zVebAZSAzGLRQob5s/ldmfel
2c1se8w3fPUC29eVB7AQMFsh1xFMzkR53r6zZEKJsG5MHu06fdKD2FhdQcfQDAob
K8Y4Yh4LhOo=
=2067
-----END PGP SIGNATURE-----
Merge tag 'sched-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull scheduler fixes from Ingo Molnar:
- Fix LLC mis-scheduling bugs (Tim Chen, Lu Wang)
- Fix cache-grouping related scheduling statistics UAF bugs (Tim Chen)
- Skip kernel threads for cache aware scheduling to rubustify the code
(Chen Yu)
- Refresh LLC capacity across CPU hotplug, to fix capacity
underestimation bug (Davi Chaves Azevedo)
- Account PSI IRQ time to the execution context, not the scheduling
context, to fix proxy scheduling accounting bug (Zhan Xusheng)
* tag 'sched-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
sched/core: Account PSI IRQ time to the execution context, not the scheduling context
sched/cache: Refresh LLC capacity across CPU hotplug, to fix capacity underestimation bug
sched/cache: Skip kernel threads for cache aware scheduling to rubustify the code
sched/cache: Introduce task_struct->sched_cache_grp to fix UAF
sched/cache: Decouple sched_cache_group from mm to fix UAF
sched/cache: Honor migrate_llc_task semantics in active load balance, to fix LLC mis-scheduling bug
sched/cache: Keep nr_pref_llc_running in the runnable domain, to fix LLC mis-scheduling bug
|
||
|
|
5ccda18d1b |
Perf events fixes:
- Fixes for KVM guest PEBS virtualization (Sean Christopherson)
- Fixes for various Intel PMUs related to PEBS data-source
(Dapeng Mi)
- Fix Intel Panther Cove event scheduling constraints (Dapeng Mi)
- Fix Intel DMR/NVL OMR extra registers event scheduling (Dapeng Mi)
- Rename two confusingly named PMU attributes (Dapeng Mi)
- Fix a refcount leak in attach_perf_ctx_data() (Namhyung Kim)
- Fix NULL pointer dereference crash in __perf_pmu_sched_task()
(Puranjay Mohan)
- Fix CPU-wide event scheduling (Puranjay Mohan)
- Fix x86 LBR branch entry generation (Puranjay Mohan)
Signed-off-by: Ingo Molnar <mingo@kernel.org>
-----BEGIN PGP SIGNATURE-----
iQJFBAABCgAvFiEEBpT5eoXrXCwVQwEKEnMQ0APhK1gFAmq4zGoRHG1pbmdvQGtl
cm5lbC5vcmcACgkQEnMQ0APhK1guAw//U+WiZoWkWv6oFgAg+KrQFPr/OAzyIYOe
Q355hrxdb2yh5+YGkLl/OCte/wIg1JiUq54gUBHPjEBUK8lxenNK2kOo6/Zd8tDx
oGPyX6t0tpiP1O/PMII3yyd6Q7DYVC08TqOYW68r7Jv/fO2qCtBWZY3hxe0otU2/
z+jqcQno3o8DbzfJnlDJb5honJos8CaT5FA+FkZkvjBF4dMWVErQowefdz2Zd78H
Bu/wxmPG+Hej7ownQalps+ZA52KuQboRwJ+NCmas+fbFCcaWRbdDKTPbvEvzBsHT
nybzlmBdneikrM0aNAXtBxJCLNzVrB6ffmKNFnd3CyfXmNtd/zleoqcKZpII2xO3
UqlnwWR+u2JMGscvWiAAtlkdI0K5P4U1sHObhvzyXwgi3LgRZWJoLoIyiYWGcfnW
Mxe3MCTy0wS+AEKRDvzg+rRr4jqTfl0BsHyzaTyTmWwx0LEg+7zwdryeixpXBbwH
qC/SWkr4sfQlO/kVZHGa1sdHTZocRLcLl3YHsLiwQygu3rjNes+olKyPTtbvt9/a
ZePBdcP+84K146VUczNJRAI4qt5b/X2aMvIYYxVvfDF6kiL0GbXgAfhdNVm5BIZF
u5e2Sb6AX3w78V+YaMX1DuF+ZaFrpQd2UKOkBJPLvFnHTQr2eP2Fa6XfVMXnMcDy
U7APgunF6D8=
=Ggeg
-----END PGP SIGNATURE-----
Merge tag 'perf-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull perf events fixes from Ingo Molnar:
- Fixes for KVM guest PEBS virtualization (Sean Christopherson)
- Fixes for various Intel PMUs related to PEBS data-source (Dapeng Mi)
- Fix Intel Panther Cove event scheduling constraints (Dapeng Mi)
- Fix Intel DMR/NVL OMR extra registers event scheduling (Dapeng Mi)
- Rename two confusingly named PMU attributes (Dapeng Mi)
- Fix a refcount leak in attach_perf_ctx_data() (Namhyung Kim)
- Fix NULL pointer dereference crash in __perf_pmu_sched_task()
(Puranjay Mohan)
- Fix CPU-wide event scheduling (Puranjay Mohan)
- Fix x86 LBR branch entry generation (Puranjay Mohan)
* tag 'perf-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
perf/core: Fill branch entries with a single assignment
perf/core: Run sched_task() for PMUs with only CPU-wide events
perf/core: Fix NULL pmu_ctx passed to pmu->sched_task()
perf/core: Fix a refcount leak in attach_perf_ctx_data()
perf/x86/intel: Rename NVL offcore_rsp attribute to offmodule_rsp
perf/x86/intel: Rename DMR offcore_rsp attribute to offmodule_rsp
perf/x86/intel: Fix precise OMR event scheduling for DMR/NVL
perf/x86/intel: Constrain Panther Cove UOPS_DISPATCHED events to PMCs 0-3
perf/x86/intel: Delete dead NVL PEBS data-source initcall
perf/x86/intel: Fix Panther Cove PEBS data-source snoop states
perf/x86/intel: Remove incorrect Panther Cove PEBS data-source constraints
perf/x86/intel: Remove incorrect LionCove PEBS data-source constraints
perf/x86/intel: Update arw_latency_data() mem-op direction handling
perf/x86/intel: Fix DKT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Fix CMT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Fix GRT PEBS load/store direction for latency events, to fix sample classification
perf/x86/intel: Make @data a mandatory param for intel_guest_get_msrs()
perf/x86/intel: Don't pointlessly context switch DS_AREA (and PEBS config) if PEBS is unused
perf/x86/intel: Don't write PEBS_ENABLED on host<=>guest xfers if CPU has PEBS isolation, to fix stuck PEBS_ENABLED
perf/x86/intel: Ensure KVM guest PEBS path doesn't set unwanted PERF_GLOBAL_CTRL bits
|
||
|
|
efb27d4767 |
Probes fixes for v7.3-rc4:
- kprobes: Fix permanent hang when flushing the kprobe optimizer Fix a deadlock when disabling kprobe optimization via sysctl or debugfs where flushers hung waiting for optimizer_completion. Replaced the completion with an optimizer_passes counter and wait_var_event_mutex() under kprobe_mutex so concurrent flushers can wait and wake up safely. - fprobe: Terminate the fgraph_data list when the reservation is not filled Fix an issue where unused shadow stack data left uninitialized by fprobe_fgraph_entry() was misparsed as stale fprobe headers on return. Explicitly write a zero word to terminate the list and update read_fprobe_header() to handle the zeroed slot properly. - ftracetest: Fix unique symbol check in kprobe_non_uniq_symbol.tc Fix false test failures in kprobe_non_uniq_symbol.tc on architectures like s390 where a symbol exists once in core kernel but also in modules. Anchor the /proc/kallsyms search regex to the end of the line so that module symbols are not incorrectly counted. -----BEGIN PGP SIGNATURE----- iQFPBAABCgA5FiEEh7BulGwFlgAOi5DV2/sHvwUrPxsFAmq3llUbHG1hc2FtaS5o aXJhbWF0c3VAZ21haWwuY29tAAoJENv7B78FKz8bAhEH/0EAamjv7/EDUoUq+BOO a2gnlYqvr+zcrDVQLNgiYbvTRDfIFPOdB2LpY7Rguee3747qeL7kkNATD10WFr1F 5lXe5LaLncNIrvHDdtcT5eER5ePAuSDMSL5CwnJrRvXJw42iFsqegZ07nrvc9HFS 5zw7Ej9VnFJFxeXIY3J4U92wkntLJ3JhsNheomOtQmEZU1g5ZPAbdq0icNrC3CAb FTezYk60VG0CT/gNTSd8JFnI4P5vKlZpkFFCLLMmptW4yQU9+ZdT55JRgY7wJ8ye w5rLBTRJrFhDgNAjoehMkoAtxik3dKgede7qlwxUk0sv+gwXssw28b5zabRBJgee XQ8= =cHwO -----END PGP SIGNATURE----- Merge tag 'probes-fixes-v7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace Pull probe fixes from Masami Hiramatsu: - kprobes: Fix permanent hang when flushing the kprobe optimizer Fix a deadlock when disabling kprobe optimization via sysctl or debugfs where flushers hung waiting for optimizer_completion. Replaced the completion with an optimizer_passes counter and wait_var_event_mutex() under kprobe_mutex so concurrent flushers can wait and wake up safely. - fprobe: Terminate the fgraph_data list when the reservation is not filled Fix an issue where unused shadow stack data left uninitialized by fprobe_fgraph_entry() was misparsed as stale fprobe headers on return. Explicitly write a zero word to terminate the list and update read_fprobe_header() to handle the zeroed slot properly. - ftracetest: Fix unique symbol check in kprobe_non_uniq_symbol.tc Fix false test failures in kprobe_non_uniq_symbol.tc on architectures like s390 where a symbol exists once in core kernel but also in modules. Anchor the /proc/kallsyms search regex to the end of the line so that module symbols are not incorrectly counted. * tag 'probes-fixes-v7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace: kprobes: Fix permanent hang when flushing the kprobe optimizer fprobe: Terminate the fgraph_data list when the reservation is not filled selftests/ftrace: Fix unique symbol check in kprobe_non_uniq_symbol.tc |
||
|
|
4ba51ef66a |
Power management fix for 7.3-rc5
Address a hibernation regression introduced during the 7.2 development cycle that causes the image memory preallocation to deadlock if it depends on frozen kernel threads (Florian Schmaus). -----BEGIN PGP SIGNATURE----- iQFGBAABCAAwFiEEcM8Aw/RY0dgsiRUR7l+9nS/U47UFAmq2mDESHHJqd0Byand5 c29ja2kubmV0AAoJEO5fvZ0v1OO1wFAH/2e9iz++Qjb9SODpGX/2Fz07qb4SBRCP Zf0r74G7qfoPczcLuiKu8irb1FSvlwr1nlcygWcF0gYLg9TJCaaQ7JIxHL/9ePV0 lbINk+4ozu5S6AbMh7O5wpv3n+nwBtg1wZZP3kSY4hxQ5zymACuEzwBTGT9vfQVA GpkrasgQVTOyt1gWAO8Ak3WX3z1EaBqzl8DsCm/75PVq2Wy1I801JagtYT6KArf6 kqH19SdiihrTl+2k/k2Vvc14H9XYfXab77ShWv1JgKF7X1QqaO8QWVFQulj0bpgc LLSlpiQXqBYfz469bPGyS2U6tyFioynqoiWQfdxC6zBGcHv1Jl3LdDg= =mfRN -----END PGP SIGNATURE----- Merge tag 'pm-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm Pull power management fix from Rafael Wysocki: "Address a hibernation regression introduced during the 7.2 development cycle that causes the image memory preallocation to deadlock if it depends on frozen kernel threads (Florian Schmaus)" * tag 'pm-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm: PM: hibernate: Freeze kernel threads after image preallocation |
||
|
|
5bfa9f1a9d |
kprobes: Fix permanent hang when flushing the kprobe optimizer
Writing 0 to /proc/sys/debug/kprobes-optimization while a kprobe is
jump-optimized never returns. The writer sleeps in D state forever with
kprobe_sysctl_mutex held, so any later read or write of that sysctl
hangs as well. For example, with vfs_read+9 as an optimizable address
in this build:
# cd /sys/kernel/tracing
# echo 'p:myprobe vfs_read+9' >> kprobe_events
# echo 1 > events/kprobes/myprobe/enable
# # wait until /sys/kernel/debug/kprobes/list shows [OPTIMIZED]
# echo 0 > /proc/sys/debug/kprobes-optimization
INFO: task sh:246 blocked for more than 10 seconds.
Call Trace:
<TASK>
__schedule+0x1176/0x4f70
schedule+0xdc/0x2c0
schedule_timeout+0x17b/0x260
wait_for_completion+0x173/0x3c0
wait_for_kprobe_optimizer_locked+0xbc/0x130
proc_kprobes_optimization_handler+0x156/0x1b0
proc_sys_call_handler+0x324/0x490
vfs_write+0x52d/0xfe0
ksys_write+0xff/0x200
do_syscall_64+0x106/0x630
entry_SYSCALL_64_after_hwframe+0x77/0x7f
</TASK>
...
INFO: task cat:265 is blocked on a mutex likely owned by task sh:246.
wait_for_kprobe_optimizer_locked() reinitializes optimizer_completion,
asks the optimizer thread to flush and sleeps in wait_for_completion().
The thread drains the (un)optimizing lists, but calls complete() only
if completion_done() is true, i.e. if the completion is already done,
which never happens while someone waits. disarm_all_kprobes() and
kprobe_trace_self_tests_init() wait the same way.
Calling complete() unconditionally would not be enough: the waiter
drops kprobe_mutex while it sleeps, and nothing else serializes the
sysctl handler against the debugfs "enabled" file. A second flusher
that still finds the lists non-empty, e.g. because a disabled probe is
queued for unoptimizing, reinitializes the completion under the first:
sysctl write debugfs "enabled" write
unoptimize_all_kprobes()
wait_for_kprobe_optimizer_locked()
init_completion(c)
mutex_unlock(&kprobe_mutex)
wait_for_completion(c)
disarm_all_kprobes()
wait_for_kprobe_optimizer_locked()
init_completion(c)
// c->wait is reset, the first
// waiter is off the queue
mutex_unlock(&kprobe_mutex)
wait_for_completion(c)
kprobe_optimizer()
complete(c)
// wakes the debugfs writer only
where c is &optimizer_completion. Lining up the two writes during an
optimizer pass loses the sysctl writer this way.
Replace the completion with a counter of optimizer passes, bumped at the
end of each pass and signalled with wake_up_var_locked(), both under
kprobe_mutex. A flusher samples the count and waits with
wait_var_event_mutex(), which drops kprobe_mutex only while sleeping, so
a new count means a whole pass ran in the meantime. Nothing is
reinitialized, so several flushers can sleep in the wait at once.
Link: https://lore.kernel.org/all/20260924092142.199198-1-parri.andrea@gmail.com/
Fixes:
|
||
|
|
ee9c669f9b |
sched_ext: Fixes for v7.3-rc4
- A task reenqueued while its dispatch was still completing had its queued state clobbered by the dispatcher, dropping every later dispatch of the task. Wait for the in-flight dispatch to settle first. - A wakeup activation on another CPU marked the destination runqueue as mid-wakeup, stranding a pending local reenqueue. If the scheduler was unloaded first, the stale request pointed into freed memory that the next scheduler dereferenced. - ops.dequeue() ran with the source dispatch queue's lock held, so a scheduler iterating that queue from the callback deadlocked the CPU. - Schedulers with their own CPU ID mapping had no way to learn a task's initial CPU mask and rebuilt it themselves, which went wrong across sub-scheduler enable and re-home. Pass it to ops.enable(). - A bypass dispatch event counter missed the dispatches made by the end-of-dispatch fallback and under-reported. - Selftests for the dequeue locking and initial mask changes. -----BEGIN PGP SIGNATURE----- iIQEABYKACwWIQTfIjM1kS57o3GsC/uxYfJx3gVYGQUCarWSLQ4cdGpAa2VybmVs Lm9yZwAKCRCxYfJx3gVYGY+sAP9y6nJh6vIvFh/X9FlJtWlNo0mncOKhy93E8jii 8CKnPQEAhvX3+Gcdl+imTh4Z915kdsEByBjTTPPOXnQxI8BKUAY= =VhBF -----END PGP SIGNATURE----- Merge tag 'sched_ext-for-7.3-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext Pull sched_ext fixes from Tejun Heo: - A task reenqueued while its dispatch was still completing had its queued state clobbered by the dispatcher, dropping every later dispatch of the task. Wait for the in-flight dispatch to settle first - A wakeup activation on another CPU marked the destination runqueue as mid-wakeup, stranding a pending local reenqueue. If the scheduler was unloaded first, the stale request pointed into freed memory that the next scheduler dereferenced - ops.dequeue() ran with the source dispatch queue's lock held, so a scheduler iterating that queue from the callback deadlocked the CPU - Schedulers with their own CPU ID mapping had no way to learn a task's initial CPU mask and rebuilt it themselves, which went wrong across sub-scheduler enable and re-home. Pass it to ops.enable() - A bypass dispatch event counter missed the dispatches made by the end-of-dispatch fallback and under-reported - Selftests for the dequeue locking and initial mask changes * tag 'sched_ext-for-7.3-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext: sched_ext: Count SCX_EV_SUB_BYPASS_DISPATCH in the dispatch fallback selftests/sched_ext: Check the cmask cid-form ops.enable() receives sched_ext: Pass the initial cmask to cid-form ops.enable() selftests/sched_ext: Test that ops.dequeue() can iterate the consumed DSQ sched_ext: Don't run ops.dequeue() with a DSQ lock held sched_ext: Derive SCX_RQ_IN_WAKEUP from the core enqueue flags sched_ext: Wait for SCX_OPSS_DISPATCHING before reenqueueing a task |
||
|
|
e8dfd03a1c |
cgroup: Fixes for v7.3-rc4
- With local event accounting, a fork rejected by the pids controller updated pids.events without notifying its pollers. - A cgroup selftest failed to compile with fortification enabled because an O_TMPFILE open lacked its mode argument. -----BEGIN PGP SIGNATURE----- iIQEABYKACwWIQTfIjM1kS57o3GsC/uxYfJx3gVYGQUCarWSKg4cdGpAa2VybmVs Lm9yZwAKCRCxYfJx3gVYGYoiAQD6JUCqDjv2Hr1YeMFlsoYVQZ7tNmljVnPD2tZ6 9WmD/AEAotdlmzM8egOuDqAi2s+UMJPm8vCZuxVzewiGqhYELgk= =0Ljp -----END PGP SIGNATURE----- Merge tag 'cgroup-for-7.3-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup Pull cgroup fixes from Tejun Heo: - With local event accounting, a fork rejected by the pids controller updated pids.events without notifying its pollers - A cgroup selftest failed to compile with fortification enabled because an O_TMPFILE open lacked its mode argument * tag 'cgroup-for-7.3-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup: cgroup/pids: Restore pids.events notifications in local mode selftests: cgroup: give the O_TMPFILE open in get_temp_fd() a mode |
||
|
|
1765a153d9 |
cgroup/pids: Restore pids.events notifications in local mode
A fork rejected by the pids controller increments the counter reported by
pids.events. When local event accounting is selected, however, pids_event()
returns after notifying only events_local_file, leaving pids.events pollers
asleep.
On legacy hierarchies, pids.events.local does not exist. With
pids_localevents, pids.events reports the same local counter. In both
cases, pids.events changes without generating a notification.
This can be reproduced with a pids_localevents mount:
mkdir /tmp/test
mount -t cgroup2 -o pids_localevents none /tmp/test
mkdir /tmp/test/t
echo 1 > /tmp/test/t/pids.max
cat /tmp/test/t/pids.events # max 0
timeout 3 inotifywait -e modify /tmp/test/t/pids.events &
sh -c 'echo $$ > /tmp/test/t/cgroup.procs; (true &)' 2>/dev/null
wait
cat /tmp/test/t/pids.events # max 1
Without this patch, inotifywait times out without reporting an event.
Notify pids.events before returning from the local event path.
Fixes:
|
||
|
|
5fc5768c7c |
bpf-fixes
-----BEGIN PGP SIGNATURE-----
iQJRBAABCgA7FiEE+soXsSLHKoYyzcli6rmadz2vbToFAmq1NvAdHGFsZXhlaS5z
dGFyb3ZvaXRvdkBnbWFpbC5jb20ACgkQ6rmadz2vbTpmpA//fqEoC6Sq1zxo3ADH
hV0Z9ewkNTjjH85QnispjcRkAhSHG3JscNXKRXm1NmNkwJsHJ4TDIKcjABYDjquD
wqfvL9hLXPsvud0M/PR6/CZeBAWXpukkaxeYedY+83ttTHjzR0tDq0Ne9yvIV+Nc
hS0qFIIHs8C6l2nyuNSxvgrv216orG8qd0Bi3tpDfsLqCsLLEmDyQ1H+ZzpJF2xW
RV3oMcggCeo305m8+uiofQGf8hmmRrmA7SEfF+Qe08ab2GOn9glINfVTbTE3AdQW
fkvAk8Zuio3hwwMBHDWWYXrKO0N3ykzcDk4V6JPUWiH1dOANf1tS3G1uJyxb4tsv
dHVdA0xL3sg7YnuSywfb82vTXvQz5QEFeDYxxLB2fMJe8LcCfgF1lkIr1DbdrMjI
hlozGCs3p/GVIVNhjGVPezWsUvzK4PIuKVM6U1qWojtAhXU/bJCvP0iBU83RlBL7
S0GGSC/qvkuPed9hAp2pnrrRo/9GEp1PZq8AXsp3nti0OaRxgxpjjQQFzZEWlOrR
bErtbKziHzl2xARvCRycNZ+QWT4ZdjmK++8pba7hOuFkRclOV4KRWk4OthBvcADu
NfNsVXdXpOnesg7TNIUJ8heVAYPjKaBHJBuG2Prghrnc95uqEQwOfbYT15zAbe6I
l3HqGerSa7WtbrnCwsf1DvyPPII=
=ZxKg
-----END PGP SIGNATURE-----
Merge tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf
Pull bpf fixes from Alexei Starovoitov:
- Fix bpf_skb_change_tail() to drop the checksum offload instead of
rejecting the trim of CHECKSUM_PARTIAL skbs (Daniel Borkmann)
- Add KF_PERFMON kfunc flag and require CAP_PERFMON for kfuncs that
read arbitrary memory and for untrusted read-only memory reads
(Daniel Borkmann)
- Clear scalar delta on narrowing stack spill (Daniel Borkmann)
- Set up the frame pointer for the exception callback in arm64 JIT, and
zero-fill other CPUs when BPF_F_CPU update creates a per-cpu hash
element (Donggeun Yoo)
- Various fixes (Emil Tsalapatis):
- Fix bounds check underflow for skb-backed dynptrs
- Fix rx_queue_mapping context access code generation in bpf_sock
- Reject packet pointer arguments to subprogs that may mutate the
packet
- Reject ALU instructions that see arena and non-arena operands on
different code paths
- Fix copied_seq double-counting on sockmap self-redirect
(Geliang Tang)
- Fix divide-by-zero in btf_struct_walk() on a flexible array of
zero-sized elements, fix out-of-bounds read of rtt_min in sock_ops
(Jiayuan Chen)
- Fix bpf_sock_destroy() out-of-bounds read of sk_protocol on TIME_WAIT
and request socks, and sleeping under RCU when destroying a listener
with pending children (Jiayuan Chen)
- Fix JEQ/JNE with immediate operand in MIPS32 JIT and missing zero
extension of BSWAP 16/32 in MIPS64 JIT (Johan Almbladh)
- Avoid soft lockup in htab lookup[_and_delete] batch operations on
large maps (Jose Fernandez)
- Various fixes (Kumar Kartikeya Dwivedi):
- Verify global subprogs in each sleepability context they are
called from
- Make post-verification instruction rewrites killable
- Preserve packet pointer displacement in regsafe()
- Apply CO-RE relocations before subprogram validation, restrict
CO-RE poisoning to relocatable instructions, and reject truncated
ldimm64 CO-RE relocations in libbpf
- Assign lock identity to callback map values
- Compare stack frames in regs_exact()
- Bound ownership depth through local kptrs and graph roots
- Fix u32 overflow in map batch operations when the map size exceeds
4GB (Masoud Aghasi)
- Fix UAF in bpf memalloc due to concurrent consumption of ttrace lists
in alloc_bulk() (Pu Lehui)
- Allow gotox as the terminal instruction of a program or a subprogram
(Siddharth Chintamaneni)
- Disallow bpf_skb_pull_data() for LWT_SEG6LOCAL, skip unsettled links
in link iterator, and reject dev-bound-only programs on other devices
(Weiming Shi)
- Reject non-negative stack offsets in stack_slot_obj_get_spi()
(Xu Yunxiang)
- Check params size before reading reserved fields in
bpf_crypto_ctx_create() (Yuqi Xu)
- Reject max_entries > INT_MAX in sock_map_alloc() (Zhao Gongyi)
- Use a 32-bit compare in xsk_map_gen_lookup() (Zhiling Zou)
- Use kvfree() in xdp_test_run_teardown() (Zhixing Chen)
* tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf: (58 commits)
selftests/bpf: Test per-cpu initialization of a BPF_F_CPU created element
bpf: Zero-fill other CPUs when BPF_F_CPU creates a per-cpu hash element
bpf: Fix BSWAP 32 and 16 on MIPS64
bpf: Fix immediate JMP JEQ/JNE on MIPS32
bpf: Reject dev-bound-only programs on other devices
bpf, sockmap: Reject max_entries > INT_MAX in sock_map_alloc
selftests/bpf: Test for mixed arena/nonarena code paths
bpf: Prevent variable arena/non-arena register contents
selftests/bpf: Test rejection of pkt args to mutating subprogs
bpf: Reject pkt arguments in mutating subprogs
selftests/bpf: Add selftests for rx_queue_mapping context access
bpf: Fix bpf_sock context code generation
selftests/bpf: Test dynptr slices past end of skb
bpf: Fix bounds check for skb-backed dynptrs
selftests/bpf: Reject iterator destruction through fp+0
bpf: Reject non-negative offsets in stack_slot_obj_get_spi()
bpf: Check params size before reading reserved fields
selftests/bpf: Check local object ownership depth
bpf: Bound ownership depth through local kptrs and graph roots
selftests/bpf: Cover frame changes in bounded loops
...
|
||
|
|
c3a66e5f5b
|
bpf: Zero-fill other CPUs when BPF_F_CPU creates a per-cpu hash element
pcpu_init_value() initializes the per-cpu area of a newly created
[lru_]percpu_hash element. The area is recycled, so when the value
comes from a BPF program (onallcpus == false) it writes the running
CPU's slot and zeroes the rest.
bpf_percpu_hash_update() passes onallcpus == true, which delegates to
pcpu_copy_value(). pcpu_copy_value() writes only the CPU named in
map_flags when BPF_F_CPU is set, so on the create path the other slots
keep the recycled element's values:
update(k1, 0xdeadc0de, BPF_F_ALL_CPUS) every CPU holds 0xdeadc0de
delete(k1) element back on the freelist
update(k2, 0xc0ffee, BPF_F_CPU | 0) creates, writes CPU 0 only
lookup(k2) CPU 0 0xc0ffee, rest 0xdeadc0de
Zero-fill the other CPUs on that arm too.
Fixes:
|
||
|
|
a0bb6fac53 |
sched/core: Account PSI IRQ time to the execution context, not the scheduling context
psi_account_irqtime() has two callers which share rq->psi_irq_time, and they disagree about the context: __schedule() passes the outgoing rq->curr, sched_tick() passes rq->donor. Under proxy execution the donor is blocked on a mutex while rq->curr burns the CPU. The tick charges PSI_IRQ_FULL to the donor's cgroup and advances the timestamp, so the call from __schedule() then finds delta <= 0 and charges nothing. The delta is not counted twice, it lands on the wrong cgroup. Pass rq->curr, which is what the call read before commit |
||
|
|
4409a85735 |
sched_ext: Count SCX_EV_SUB_BYPASS_DISPATCH in the dispatch fallback
When a descendant scheduler enters bypass mode, its tasks are parked in
the bypass DSQs of the nearest non-bypassing ancestor, which is then
responsible for running them. On behalf of such a non-bypassing host,
scx_dispatch_sched() consumes those bypass DSQs from two places: the
attempt made every SCX_BYPASS_HOST_NTH dispatches, and the
end-of-dispatch fallback that keeps the CPU from going idle while
bypassed descendants still have tasks queued.
The former increments SCX_EV_SUB_BYPASS_DISPATCH but the latter does
not, even though both perform the same scx_consume_dispatch_q() on the
same bypass DSQ. The descendant bypass dispatches done by the fallback
are therefore missing from the counter exposed via sysfs,
scx_dump_state() and the scx_bpf_events() kfunc, which under-reports the
actual number of such dispatches.
Add the missing __scx_add_event() so the fallback counts them too. When
@sch itself is bypassing, scx_dispatch_sched() takes the earlier
self-bypass branch and returns before reaching these host paths; that
mode is accounted for by SCX_EV_BYPASS_DISPATCH at enqueue time and is
intentionally left unchanged.
Fixes:
|
||
|
|
3d8d741009 |
perf/core: Run sched_task() for PMUs with only CPU-wide events
perf_pmu_sched_task() returns early when cpuctx->task_ctx is set and
leaves the work to perf_ctx_sched_task_cb(), which only walks
ctx->pmu_ctx_list. A PMU whose events are all CPU-wide is not on that
list, so nothing calls its sched_task(). With
perf record -b -e cycles -a -- ls
armv8pmu_sched_task() is skipped on every switch to a task that has a
perf context but no event on that PMU, and BRBE records leak across the
task boundary. intel_pmu_lbr_add() calls perf_sched_cb_inc()
unconditionally too, so LBR records leak the same way on x86.
Drop the early return and skip only the CPCs that
perf_ctx_sched_task_cb() handles. That one needs a gate of its own to
make the split exact: it tests cpc->sched_cb_usage, which
perf_sched_cb_inc() sets per CPU for every branch stack user, so a task
with an event for that PMU pinned to another CPU would be handled twice.
On x86 the second __intel_pmu_lbr_restore() finds lbr_stack_state ==
LBR_NONE and calls intel_pmu_lbr_reset(), throwing away the callstack
the first one restored.
cpc->task_epc is set only while a task context is scheduled in, and
there is one epc per PMU on ctx->pmu_ctx_list, so the two gates are
inverses.
For the CPCs perf_pmu_sched_task() picks up, the callback now runs
outside the perf_ctx_disable() and perf_ctx_enable() pair in
perf_event_context_sched_in(). __perf_pmu_sched_task() disables the PMU
around the call itself.
Fixes:
|
||
|
|
36bb85cf36 |
perf/core: Fix NULL pmu_ctx passed to pmu->sched_task()
perf_pmu_sched_task() returns early when cpuctx->task_ctx is set, and
cpc->task_epc is only non-NULL while a task context is scheduled in on
this CPU. __perf_pmu_sched_task() therefore always passes NULL:
Unable to handle kernel NULL pointer dereference at virtual address 00
pc : armv8pmu_sched_task+0x14/0x50
Call trace:
armv8pmu_sched_task+0x14/0x50 (P)
perf_pmu_sched_task+0xac/0x108
__perf_event_task_sched_out+0x6c/0xe0
Pass &cpc->epc instead, the CPU-wide context for this PMU, which the
function already dereferences a few lines up to find pmu.
armv8pmu_sched_task() is the only in-tree implementation that
dereferences the argument, and it only reads ->pmu, so the oops needs
BRBE, added in v6.17.
Fixes:
|
||
|
|
cca4980630 |
perf/core: Fix a refcount leak in attach_perf_ctx_data()
The attach_perf_ctx_data() can race on global and !global cases. The
global case is protected by global_ctx_data_rwsem and shares a single
reference count using perf_ctx_data.global field.
But when it races with !global case, it may miss to set the global field
and result in a reference count leak.
CPU1 CPU2
----------------------------------------------------------------
attach_task_ctx_data(.global=1) attach_task_ctx_data(.global=0)
cd1 = alloc_perf_ctx_data(); cd2 = alloc_perf_ctx_data();
// { .global = 0, .refcount = 1 };
try_cmpxchg(); // success,
// task->perf_ctx_data = cd2
try_cmpxhg(); // fail; old = cd2
refcount_inc_not_zero(&old->refcount); // success
// old.refcount = 2
free_perf_ctx_data(cd1);
Then later detach_global_ctx_data() will see the data but it's not
marked as global, so it won't call detach_task_ctx_data().
Fixes:
|
||
|
|
6db1ce73e9
|
bpf: Reject dev-bound-only programs on other devices
__bpf_offload_dev_match() falls back to comparing offdev pointers after an
exact netdev mismatch. Bound-only programs normally have NULL offdevs, so
unrelated netdevs compare equal. A bound-only program on an
offload-registered netdev can instead inherit a real offdev and match a
sibling port. With CAP_BPF and CAP_NET_ADMIN, a caller can use
bpf(BPF_LINK_CREATE) with a different target ifindex to run metadata kfuncs
specialized for the bound driver on the target driver's xdp_buff. Running a
veth-bound program on tun reads beyond tun's bare stack xdp_buff as a
veth_xdp_buff.
Oops: general protection fault, probably for non-canonical address
KASAN: null-ptr-deref in range [0x0000000000000010-0x0000000000000017]
RIP: 0010:veth_xdp_rx_timestamp (drivers/net/veth.c:1673)
Call Trace:
...
tun_build_skb (drivers/net/tun.c:1739)
tun_get_user (drivers/net/tun.c:1856)
tun_chr_write_iter (drivers/net/tun.c:2091)
vfs_write (fs/read_write.c:595 fs/read_write.c:687)
ksys_write (fs/read_write.c:739)
do_syscall_64 (arch/x86/entry/syscall_64.c:84)
entry_SYSCALL_64_after_hwframe (arch/x86/entry/entry_64.S:121)
Kernel panic - not syncing: Fatal exception in interrupt
Restrict non-offloaded programs to exact netdev matches and retain the
shared-offdev fallback only for genuinely offloaded multi-port programs.
Fixes:
|
||
|
|
f85f5917aa
|
bpf: Prevent variable arena/non-arena register contents
The verifier marks ALU instructions that include at least
one arena operand with needs_zext: These instructions are
fixed up after verification to be ALU32 instructions to
ensure that the result is a valid offset into an arena.
However, different code paths may provide two non-arena
64-bit arguments to the same instruction. The result of
the operation in that code path is wrong, since it is
now unexpectedly truncated to 32 bits and zero-extended.
Add logic to the verifier to ensure every instruction either
always has at least one PTR_TO_ARENA argument, or never does.
Since needs_zext already tracks the first scenario, add a
prevent_zext field in bpf_insn_aux to track the latter.
Reject instructions that use arena arguments and have prevent_zext
set, or do not have arena arguments and have needs_zext set.
Fixes:
|
||
|
|
a6c1edfbe2
|
bpf: Reject pkt arguments in mutating subprogs
The verifier tracks changes in how PTR_TO_PACKET registers'
bounds are modified across subprog boundaries. PTR_TO_PACKET
registers are actually passed as PTR_TO_MEM, which is assumed
valid for the entire call. This is not the case with packet memory,
where a pskb_* call may invalidate its memory region.
Reject BPF code that passes PTR_TO_PACKET pointers to subprogs that
may mutate a packet. We cannot pass the pointer as a true PTR_TO_PACKET
because we would also need to somehow pass the PTR_TO_PACKET_META
or PTR_TO_PACKET_END to the subprog. Since we cannot avoid representing
the pointer in the subprog as PTR_TO_MEM, only permit it if the
subprog is guaranteed not to mutate the packet.
Fixes: 80f281664f5a ("bpf: Support pointers in global func args")
Reported-by: Nicholas Carlini <nicholas@carlini.com>
Suggested-by: Nicholas Carlini <nicholas@carlini.com>
Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260922172028.6269-6-emil@etsalapatis.com
|
||
|
|
41112a787f |
PM: hibernate: Freeze kernel threads after image preallocation
Commit |
||
|
|
3cb0243767 |
sched/cache: Refresh LLC capacity across CPU hotplug, to fix capacity underestimation bug
The scheduler scales LLC capacity by the fraction of cache-sharing CPUs
covered by a domain:
llc_bytes = cache_size * span_weight / shared_weight
During CPU teardown, sched_cpu_deactivate() rebuilds scheduler domains
before cacheinfo_cpu_pre_down() removes the CPU from shared_cpu_map. The
new domains therefore use the old sharing weight. The later call to
sched_update_llc_bytes() looks up the departing CPU's sd_llc, which has
already been detached, and returns without correcting the surviving CPUs.
On a Ryzen 5 7535U with twelve logical CPUs sharing a 16 MiB LLC,
offlining one SMT sibling left the remaining CPUs with:
llc_bytes = floor(16777216 * 11 / 12) = 15379114 bytes
The correct capacity is still 16777216 bytes. On systems with active
cache-aware scheduling, an underestimated capacity can cause
exceed_llc_capacity() to reject aggregation for a process whose footprint
would fit. Unchanged cpuset partitions sharing the physical cache can
also retain stale capacity when a CPU comes online in another partition.
Pass the cache-sharing mask already retained by cacheinfo to the
scheduler update. Refresh every surviving CPU using its own LLC domain
so that each partition receives the correct share. This also preserves
the correction needed as cache-sharing maps grow during boot.
Keep the existing CPU-hotplug and scheduler-domain synchronization. The
update remains on the hotplug path; no steady-state scheduling operation
or persistent allocation is added.
Fixes:
|
||
|
|
65efcccddc |
sched/cache: Skip kernel threads for cache aware scheduling to rubustify the code
Kernel thread should not be covered by cache aware scheduling as
it borrows the statistics from the user space thread. Filter the
kernel thread in account_mm_sched().
In theory a kernel thread does not have any valid
cache group, so !grp should gate the kernel thread.
Add the PF_KTHREAD check explicitly here for safety
reasons, to guard against future modifications and
to pair with task_tick_cache().
Fixes:
|
||
|
|
b636fef85b |
sched/cache: Introduce task_struct->sched_cache_grp to fix UAF
Add a sched_cache_grp pointer to task_struct so that scheduler code
can access the cache group directly via the task, without going
through mm->sched_cache_grp. This decouples the scheduler's hot-path
accesses from the mm_struct.
Each task holds its own refcount on the sched_cache_group, separate
from the reference held by its mm_struct. The reference is acquired
in copy_mm() (fork) and exec_mmap() (exec), and released in exit_mm().
This fixes the use-after-free when account_mm_sched() reaches the group
through a task whose mm is being switched, as reported by Hyunwoo:
https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
Convert all scheduler code in fair.c and exit.c to use
p->sched_cache_grp instead of p->mm->sched_cache_grp.
Keep the fork/exec/exit reference management out of the generic mm
paths: add sched_cache_fork(), sched_cache_fork_cleanup(),
sched_cache_exec_mmap() and sched_cache_exit_mm() in
kernel/sched/cache_sched.c (with empty stubs for !CONFIG_SCHED_CACHE),
so fs/exec.c, kernel/fork.c and kernel/exit.c each call one helper
instead of open-coding the refcounting under #ifdef. Also add
sched_cache_group_get() and task_cache_group_get().
Fixes:
|
||
|
|
28f9c0e0a0 |
sched/cache: Decouple sched_cache_group from mm to fix UAF
Currently the sched cache grouping is by mm and the scheduling statistics
sched_cache_stat lives in the mm structure. This ties the life cycle
of scheduling stats with mm.
In account_mm_sched(), the scheduling stats are accessed by
task->mm->sc_stat. However, a task may be switching mm on one CPU when
another CPU is running account_mm_sched(), and possibly accessing the
old mm that was freed. This problem was found when running tests with
KASAN by Hyunwoo:
https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
Instead of serializing the mm access by introducing extra acquisition of
rq lock in the mm free path, extract sched_cache_stat from mm_struct,
rename it as sched_cache_group and manage its life cycle apart from
mm_struct with its own ref counting. This allows us in the next patch
access sched_cache_group directly from task, and add a refcount
on sched_cache_group when a task links to it. This prevents the use
after free issue when accessing stale and released old mm and its
sched cache stat a task switches to a new mm while account_mm_sched()
is done elsewhere.
The other benefit of this restructure is in the future, the grouping of
tasks to a LLC would have the flexibility to be associated with a user
defined grouping, or cgroup, cookie group, numa_group or others instead
of just with a single mm address space.
Rename sched_cache_stat to sched_cache_group and turn it into a refcounted
object allocated from mm_struct. The mm_struct now holds a pointer
(sched_cache_grp) to this object instead of embedding it.
Fixes:
|
||
|
|
d6013e2465 |
sched/cache: Honor migrate_llc_task semantics in active load balance, to fix LLC mis-scheduling bug
Cache aware scheduling introduced the migrate_llc_task migration type to direct
tasks toward their preferred LLC, but its semantics can be lost when passive
load balance falls back to active load balance (ALB). This may allow ALB to
select a candidate whose preferred LLC does not match the destination, moving
it away from its preferred LLC.
Example scenario:
src_rq has two runnable tasks, p1 and p2. p1 prefers dst_rq (dst_llc), while p2
prefers src_rq (src_llc). In this case, migrate_llc_task is set because src_rq
has at least one task, p1, that wants to migrate to dst_rq. In ALB,
can_migrate_task() finds p2 and returns true for it, thus moving p2 out of its
preferred LLC.
Solution:
The CPU stopper in ALB constructs a fresh lb_env that does not inherit
migration_type from the passive load-balance pass. Two approaches are
possible:
(a) Add a new member to struct rq so ALB can inherit migrate_llc_task
from the passive LB that triggered it.
(b) Define a new flag LBF_ACTIVE_LB_LLC and select the stopper callback
at kick time to preserve the migration semantics across the
asynchronous boundary.
We choose (b) because it avoids passing migration_type through the
stopper, which would affect the meaning of migration_type for
delayed-dequeue tasks.
Fixes:
|
||
|
|
0d6526f82c |
sched/cache: Keep nr_pref_llc_running in the runnable domain, to fix LLC mis-scheduling bug
alb_break_llc() decides whether to break LLC preference during active
load balance. It does so by testing that every runnable fair task on the
source rq prefers its LLC:
env->src_rq->nr_pref_llc_running == env->src_rq->cfs.h_nr_runnable
But the two counters cover different sets. nr_pref_llc_running is updated
in account_llc_enqueue()/account_llc_dequeue(), next to cfs_rq->nr_queued,
so it follows queued tasks. h_nr_runnable is updated in set_delayed()/
clear_delayed() and drops delay-dequeued tasks.
So under DELAY_DEQUEUE, a preferring task that goes to sleep stays counted
in nr_pref_llc_running while h_nr_runnable falls. The equality then breaks,
alb_break_llc() returns false, and active balance is free to pull a task
off its preferred LLC. Active balance only moves runnable tasks, and this
is the only LLC check it consults: once the stopper runs, LBF_ACTIVE_LB
skips the per-task test in can_migrate_task(). The runnable set is the one
we want.
Fix it on the counter side. A task should be counted in
nr_pref_llc_running exactly while it is both queued on its preferred LLC
(pref_llc_queued) and runnable (!sched_delayed). Define that membership
once in task_pref_llc_runnable(), and adjust the counter only through
pref_llc_running_inc()/pref_llc_running_dec() from the four sites that
change either input: account_llc_enqueue(), account_llc_dequeue(),
set_delayed() and clear_delayed(). Gating every update on the same
predicate keeps the delay, wake and dequeue paths from double-counting
or underflowing; see the comments at those sites for the ordering.
nr_llc_running and sd->llc_counts are not touched and stay on queued
semantics.
Fixes:
|
||
|
|
79a9172f3a |
bpf: Reject non-negative offsets in stack_slot_obj_get_spi()
bpf_get_spi() computes (-off - 1) / BPF_REG_SIZE using C division,
which truncates toward zero. For off == 0, this produces spi 0, the
same index used by the valid stack slot at fp-8.
stack_slot_obj_get_spi() currently checks alignment and the resulting
spi bounds, but does not reject the non-negative offset itself. It can
therefore validate a PTR_TO_STACK register holding fp+0 against an
iterator stored at fp-8 even though the runtime receives the actual fp+0
pointer. An effectful iterator kfunc can then interpret memory outside
the BPF stack as iterator state.
Reject non-negative offsets before converting the offset to an spi. All
valid stack objects begin at a negative offset from the frame pointer.
Fixes:
|
||
|
|
0a15ba6b0c |
Timer race fixes:
- Fix timer signal <-> exec() race, to prevent UAF (Thomas Gleixner)
- Clean up POSIX CPU timers right after de_thread(), to prevent UAF
(Hyunwoo Kim)
- Fix POSIX CPU timers race between expiry and timer_settime(),
to prevent UAF (Thomas Gleixner)
Signed-off-by: Ingo Molnar <mingo@kernel.org>
-----BEGIN PGP SIGNATURE-----
iQJFBAABCgAvFiEEBpT5eoXrXCwVQwEKEnMQ0APhK1gFAmqvrMYRHG1pbmdvQGtl
cm5lbC5vcmcACgkQEnMQ0APhK1gqUQ/+NkruN984bFynF/eZ0/2DFv91AAUP8zgH
/S3PBlwuSbYFN9JVhngDMwxQkamE56weJbFc+0QuvVT5UVw/vX9BS4QOvvzN+f8D
FEN3UqD0d1B8OwlPNTw0sFPwJDdPctTinfKOhNjNQe6RLFsNARvGyaKDIDroWTfV
dxuJ/7Ecs+5m1bmGJnPEC+IH/OnV9BEEl1NdZb+INKpBlui9LCsw4rRIj/8dPK/H
UNhvXpykKrJCDftbCzAFSNryuzcJgq4kHtMbsqiUL6y50AB69eHGi/Y0xYBAEr1h
NiDPq2PAMmH1NCCMsTtqbJZMqgCr+7DSZiCFn7bZPwg0V5tV4PFZD484q0sCbiej
Fwg+arHd0icnceIcWMsBWPUVOSLxZaWdp9a2Tj3Ill06//b5bEDBJBbpecS+so3t
8W6IvdoCYm7sz50mohnjOdx7biHPu0yhwgj+EoAV3nZKoALQAAcI7+HJzSWpGnJi
HIO0zylRAZCjk9H3QNWO+LdWgifc8DysAZOWpmbuwGgp8q483IDRDtme/kMt3+D1
1qTHa1TD/tPo8UmmgyVJQ7e1hCxBkGuuBBu5Y3/qkUEOQM6B/H2Ji7stxsLpW3JL
HLzC3kL2SBBVBO2ljqiH5IhVAL10Qm5vPxaCOjExdBt1vjMxN1DowZqD4rT/tw58
ArpD4zmr8VQ=
=KhWJ
-----END PGP SIGNATURE-----
Merge tag 'timers-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull timer race fixes from Ingo Molnar:
- Fix timer signal <-> exec() race, to prevent UAF (Thomas Gleixner)
- Clean up POSIX CPU timers right after de_thread(), to prevent UAF
(Hyunwoo Kim)
- Fix POSIX CPU timers race between expiry and timer_settime(),
to prevent UAF (Thomas Gleixner)
* tag 'timers-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
posix-cpu-timers: Prevent freeing a timer which is queued on the expiry list
exec: Cleanup POSIX timers right after de_thread()
signal: Prevent exec() race
|
||
|
|
fecbe78ac0 |
Scheduler fix:
- Avoid false positive migration warning for proxy donors
(Andrea Righi)
Signed-off-by: Ingo Molnar <mingo@kernel.org>
-----BEGIN PGP SIGNATURE-----
iQJFBAABCgAvFiEEBpT5eoXrXCwVQwEKEnMQ0APhK1gFAmqvqssRHG1pbmdvQGtl
cm5lbC5vcmcACgkQEnMQ0APhK1gPzBAAhE9hcpkBV6vNW9xzBDKdGPxRjch2vcfu
lQbx7da1yC2rHU7RDlcdfEgkhAqTRwEAXKz26zth3sDHLk43tusuNtqG3swELhNW
YAGbHjdn87snQlmP8AOK4EaT3uE5NjWqRSIJWMK+AhWSwqO/zzK91T6yZyQY0fM4
3Edp8CFo3IeyOzCR96vsob2x1fFQhuR//5fWul3uuB0EZ8VA5FhH3ene6VCm7P/1
dauxX0rMTb2730qNXA8cHROZq+bwhqTZOUaoOZ33WxnRPkvY9mV/hZsN8JnJuzBE
ogzyyorcl8dFH8qOapos9Cp3tQj9GkTX7mXWDbuUflt/8uOXtQMf75kFGBVI5NUw
2xNfgubTYGo6Qc1C+wyOhGYJ6T5Al+083pV/vPc4y7Z1i7RgZ94QypMuZuaR/RqS
z+RQcdNQlXiRIAIILGJqq1xdbaCJvbVx3tiFZkhPse6ioOF6UNGbQiNExAq3v5BU
ocvhBuf9p/uvRmfs+ZtQNqAAjZUL7tQPvdFAsjxKjuI2Z5YGPROy1L9NdYKfzryM
yWOEkV2mdn97CwzDS+auC0HmPkGqf8we2VI5Ub4R35UPqDcV6vc5BAnwzAbuxAmc
K7mQOMDEaRkbVQ0MoSMdidIXLL0AcN+BROBUtHv8kqJur6nADE3K3HZNdhr2ViLN
eisNI6m8pLk=
=BEte
-----END PGP SIGNATURE-----
Merge tag 'sched-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull scheduler fix from Ingo Molnar:
- Avoid false positive migration warning for proxy donors
(Andrea Righi)
* tag 'sched-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
sched/core: Avoid false migration warning for proxy donors
|
||
|
|
abb91eed94 |
Perf events fixes:
- Fix crash when probing CS CALL instructions (Jinke Han) - Fix NULL pointer crash during module unload (Vinay Belgaumkar) Signed-off-by: Ingo Molnar <mingo@kernel.org> -----BEGIN PGP SIGNATURE----- iQJFBAABCgAvFiEEBpT5eoXrXCwVQwEKEnMQ0APhK1gFAmqvqiURHG1pbmdvQGtl cm5lbC5vcmcACgkQEnMQ0APhK1hiohAAjcvOY7M998pX1tmo1Egw9kAuI0odtNOX weQ4Wq7K3X+tg+q1wVUmE/N/y/WLYWNatjwvjc8TClYdQEiaqBMH4TSLAuY3TOxa TxwdTC20uSN4EPZ0iRhKzm3biPzzzRq4M9hhV+WfGcK7ieXRn4Q7d9S3DDe5oEfG lI4l/RefIBiINVPC7dNM7xpS/7XBEPnzNeshOMwp6ZsPziZJizgC8C7RhQFAPszo Ho36KKFQqMNouCSybQl1GxLyPw+oGtneWESHXrF6Mhp+bcx40fMtJxKNyVLsVWuM wo0Ry843pCbDofOIqg7m0AufWUhz7B4MttTXXrU2/BYMHEbxlgms1AO7lnSGzPmn vP0JHZnT3y34P5uvGaVho7t9QKKbuY47cKNmsiiLuXiBQQuKnvDmDln1Mu1KS3fg a1LI8kv043iLqAjsIWMVtKlRGUX36f4NXUWrxvO/tmup3ocJhG1oYmEe/RaFddVX 5qRbHn7Z1w8jAWldODYSrXkpMRgMtClQuqHdjxZQt5DPZSORzoFxTvSfJLdUUtVx CUiT34zcLPHfvGD+ctHe5kVesgDHCxp/z1tKrJBunB+XZVl6rKHycfiGjaUXQRcR NUp6iFlfay8b79EkMsLC8p6uAmzX2q+gEwNVy8eevBSXdsYqcY5islj3ob7sBHsH vk4gDJikpE0= =onkS -----END PGP SIGNATURE----- Merge tag 'perf-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip Pull perf events fixes from Ingo Molnar: - Fix crash when probing CS CALL instructions (Jinke Han) - Fix NULL pointer crash during module unload (Vinay Belgaumkar) * tag 'perf-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: perf: Fix null pointer access in is_include_guest_event() x86/kprobes: Fix crash when probing CS CALL instructions |
||
|
|
bdab18633a |
- Also allocate a default private futex hash on vfork()
as well, to avoid races with (private) futex waiters
(Peter Zijlstra)
Signed-off-by: Ingo Molnar <mingo@kernel.org>
-----BEGIN PGP SIGNATURE-----
iQJFBAABCgAvFiEEBpT5eoXrXCwVQwEKEnMQ0APhK1gFAmqvp20RHG1pbmdvQGtl
cm5lbC5vcmcACgkQEnMQ0APhK1giDg//VpFKPXgma7fW5E+T0kjWHWD7FS1vYyJ4
nMunKME8PAV0A4S7946iUAWHjViVfksTmfvcXepSFCKZKLf55g8s2aZ1VoigMtYX
bQxNrnOidebSg9npLs9NV4sjvvNy+k03gcNN+OK+ZRfYCVRqbLgfCm5RJbaId5y4
+EuizVmtJuNav6HwAEIbU4LXGIdwSL9Pn8Zitkz/H7g0ZEkv+VA4h6NttvKDNfV7
OltcUFejtU7Z1lItfH+PP9pMcjPA6OPI2h7LLmUcZGUhhQQtW2Gb1fffz45PiH8T
FB7PdFt22TLG5c6hLB7zbrFFillWKn3l3Ihi/IxqVmMulhk6RVTPdCVJcj1cGZxF
9NXZ+L81poKwEETaIk52v5jm9qNF+kHbXJuCjPbmEdPxtxfv+Ma4zXom/xKkuXpx
qf01GXxelUPdCAl0cT0pzRvbEfHNIOsE2Id7f+59jB+L8ZbYEch04cIVRqCQcOgi
9B+lXj36fkFBV6wmuP5SWShWsgMsMpgSzOz5mUdWjY3Ocn9QuzDxPkk50Cm7uL1i
q/HpGv3T849HFnCH7+mi8wSiX33La+N297+AkGO7U7h5QFldPokvbAijwfmF+1Nw
CebsAfsgJsN1GZacDTH0jmRxh4I0Wa32yiCAldeGS2w9EMEAFbA87iwsOgjwHiQl
JOv/s06Pg7o=
=qQHe
-----END PGP SIGNATURE-----
Merge tag 'locking-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull futex fix from Ingo Molnar:
- Also allocate a default private futex hash on vfork() as well, to
avoid races with (private) futex waiters (Peter Zijlstra)
* tag 'locking-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
futex: Also allocate private hash on vfork()
|
||
|
|
a11212910c
|
bpf: Check params size before reading reserved fields
bpf_crypto_ctx_create() is a kfunc whose second argument is declared
with the __sz annotation, so the verifier only guarantees that
params__sz bytes of params are valid. The function nevertheless reads
params->reserved[0] and params->reserved[1] (offsets 14 and 15) before
comparing params__sz against the size of struct bpf_crypto_params, so a
BPF program can pass a shorter buffer and have the kernel read past the
region that was validated for it.
Move the size check in front of the reserved field reads.
Fixes:
|
||
|
|
c21eaa72f0 |
posix-cpu-timers: Prevent freeing a timer which is queued on the expiry list
Kijo analyzed another race in the POSIX CPU timer code: Commit |
||
|
|
3bd46666cf |
sched_ext: Pass the initial cmask to cid-form ops.enable()
The cid-form API has an obvious hole. A task's cid mask is only visible through ops.set_cmask(), which fires on affinity changes and class switches but not when a task enters a scheduler through fork, sub-sched enable or re-home, and there is no p->cpus_ptr equivalent to fall back on. Schedulers work around it by seeding the mask in ops.init_task() from p->cpus_ptr cid by cid, which is subtly wrong: on sub-sched enable and re-home, an affinity change between init_task() and enable() is delivered to the sched the task is still on, and nothing corrects the new sched's copy afterwards. Fix it by adding struct scx_enable_args to cid-form ops.enable() carrying the task's cmask, built in the per-cpu scratch under the rq lock as the task enters the scheduler, and calling set_cmask() with the same mask right after enable(), ahead of set_weight(). A scheduler can then track affinity in set_cmask() alone, and scx_qmap drops its init_task() seed. set_cmask() no longer fires for a cid-form task before it is enabled, and the class-switch republish in switching_to_scx() is limited to the cpu form. This changes the cid-form ops.enable() signature, which is fine as the cid-form API is still considered unreleased. An args struct rather than a bare cmask argument leaves room for more initial state without another signature change, and the cmask travels as a plain arena address because BTF can't type arena struct members yet. v2: The cmask travels as a u64 arena address, cmask_arena_addr, instead of a kernel-typed pointer, with the typing limitation and the planned typed alias documented (Sashiko review). v3: The initial set_cmask() is delivered before set_weight() so that the mask is in place when weight-dependent state is derived (Andrea Righi). Selftest added. Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com> |
||
|
|
bfc888f045
|
bpf: Bound ownership depth through local kptrs and graph roots
Program-allocated objects can own other local objects through referenced
kptrs. bpf_obj_free_fields() follows those pointers through
__bpf_obj_drop_impl() synchronously, before the object storage is freed
through RCU. A self-referential local kptr type therefore permits arbitrarily
deep object chains, and dropping the head can exhaust the kernel stack.
Long acyclic type chains have the same problem.
btf_check_and_fixup_fields() still assumes referenced kptrs only point to
kernel types and checks ownership through list and rbtree roots only. Its
existing rule is sufficient for graph-only cycles: the target of each graph
edge must contain a node, so every type in a cycle has both a root and a
node. The rule rejects such a type owning another root, breaking every
cycle. It also limits graph-only chains to three types, or two if the first
type contains a node, and conservatively rejects longer acyclic chains.
The missing local-kptr edges, rather than a missed graph-only cycle, are the
bug introduced by support for bpf_kptr_xchg() into local kptrs.
Replace that restriction with one bounded ownership walk covering graph
roots and local referenced kptrs. Run it after all BTF records have been
fixed up, reject cycles and paths deeper than eight record-bearing types,
and cache each type's suffix depth while checking it against the remaining
budget. This also permits the longer acyclic graph-only layouts rejected
by the old rule; update their existing BTF tests accordingly.
Keep the bound independent of MAX_CALL_FRAMES because recursive destruction
can run below a BPF call chain. A plain local pointee without special-field
metadata adds only a final non-recursing drop. Non-owning kptrs and
kernel-BTF kptrs do not recurse through local records and remain outside the
walk. Include local percpu-kptr edges too, although allocation of percpu
objects with special fields is currently forbidden, so that relaxing that
restriction cannot bypass the ownership bound.
btf_check_and_fixup_fields() continues to initialize graph_root.value_rec,
including for separately allocated map records. The ownership relationships
belong to immutable program BTF and only need validation at BTF load time.
Fixes:
|
||
|
|
8901cee931
|
bpf: Compare stack frames in regs_exact()
regs_exact() compares the register state up to id, followed by the ID
mappings, but does not compare frameno. The PTR_TO_STACK case in regsafe()
checks frameno separately, which is bypassed when exact comparison is
requested. Consequently, infinite-loop detection can treat pointers to
different stack frames as the same pointer and reject a finite loop.
For example, initialize fp-8 to zero in the caller and to one in the
callee, then pass the caller's fp-8 to the callee as r1:
loop:
r0 = *(u64 *)(r1 + 0);
if r0 != 0 goto done;
r1 = r10;
r1 += -8;
goto loop;
done:
exit;
The loop terminates after reading the callee's slot on its second
iteration. At the loop header, however, the only relevant difference is
r1's frameno, so exact comparison incorrectly reports an infinite loop.
The same problem occurs when the pointer is spilled to the stack.
Move frameno into the type-specific metadata union, ahead of id, so the
existing prefix comparison in regs_exact() covers it. Ordinary stack
pointers do not use another union member. Iterator and IRQ stack-slot
states use their dedicated union views and do not need a frame lookup.
This also keeps bpf_reg_state at 80 bytes.
Since frameno now shares storage with other pointer metadata, it is only
meaningful for PTR_TO_STACK registers. Return NULL from bpf_func() for
other register types. process_iter_arg(), get_constant_map_key() and
is_dynptr_reg_valid_init() look up the frame before checking the register
type and would otherwise index frame[] with a byte of the register's map
or BTF pointer. They dereference the frame only after their type check.
Move the states_maybe_looping() boundary from frameno to precise after the
field relocation. Its prefix comparison continues to cover the complete
value state and now includes frameno.
Continue to ignore precise. Precision marks control whether pruning may
ignore scalar ranges; they do not change the represented values, and exact
comparison already compares those ranges unconditionally. Marks can also
change through backtracking while an ancestor state is still being
explored.
Fixes:
|
||
|
|
fe3c73d7bc |
sched/core: Avoid false migration warning for proxy donors
Proxy execution can move a blocked donor's scheduling context to the
lock owner's CPU even when the donor is migration-disabled. The donor
does not execute there, and its original execution CPU remains recorded
in wake_cpu.
set_task_cpu() warns unconditionally for migration-disabled tasks, so a
subsequent proxy migration or the wakeup path returning the donor home
triggers a false positive: moving a blocked scheduling context does not
violate the migration-disabled execution context.
For example, creating a mutex owner on CPU1 and a migration-disabled
waiter on CPU0 can trigger the following warning:
proxy_migrate_repro: donor blocking on CPU0 with migration disabled
proxy_migrate_repro: donor moved from CPU0 to CPU1
WARNING: kernel/sched/core.c:3389 at set_task_cpu+0x1d3/0x280
...
Call Trace:
try_to_wake_up+0x43f/0x780
__mutex_unlock_slowpath+0x330/0x540
owner_fn+0x9f/0xc0 [proxy_migrate_repro]
...
proxy_migrate_repro: donor woke on CPU0, task_cpu=0
proxy_migrate_repro: completed
Exclude blocked proxy donors from the warning. The proxy wakeup path
restores an executable placement before clearing the blocked state.
Fixes:
|
||
|
|
88aed0422f |
perf: Fix null pointer access in is_include_guest_event()
A typical module unload occurring event when there is an active perf
connection leads to freeing of the pmu pointer. The call log is something
like:
..
__pmu_detach_event
pmu_detach_event
pmu_detach_events
perf_pmu_unregister
..
__pmu_detach_event() sets event->pmu to null. When the perf connection
finally is closed, the following stack trace is observed:
Oops: general protection fault, kernel NULL pointer dereference
...
RIP: 0010:_free_event+0x3e/0x370
...
Call Trace:
...
perf_event_release_kernel+0x260/0x2d0
perf_release+0x12/0x20
A call to mediated_pmu_unaccount_event() inside _free_event() is the root
cause of this crash. Adding a check inside is_include_guest_event() ensures
we don't accidentally access a null pmu ptr. In addition to this, we will
now call mediated_pmu_unaccount_event() before clearing the pmu ptr so that
nr_include_guest_events counts are maintained correctly.
Fixes:
|
||
|
|
71919742c8 |
bpf: Assign lock identity to callback map values
A nested bpf_for_each_map_elem() callback can unlock a different element
of the same map:
static long inner(void *map, int *key, struct value *v,
struct value **outer_value)
{
bpf_spin_lock(&v->lock);
bpf_spin_unlock(&(*outer_value)->lock);
return 0;
}
static long outer(void *map, int *key, struct value *v, void *ctx)
{
bpf_for_each_map_elem(map, inner, &v, 0);
return 0;
}
Both callback values currently have ID zero and the same map_ptr.
process_spin_lock() compares those two fields, so it accepts the unlock
even though the two callbacks can receive different map elements.
Assign a fresh ID to every callback map value in the for-each,
timer/workqueue, and task-work constructors. Copies of one callback
argument retain its ID, so locking and unlocking through that argument
continues to work. Distinct callbacks also get distinct IDs for
single-element arrays, including inner arrays sharing inner_map_meta.
Preserve map_uid for every inner-map lookup and compare it through
check_ids() during state pruning. This preserves relationships between
maps, keys, and values while allowing equivalent states with different
lookup IDs to match. It avoids field-specific rules for when an inner map
needs an identity.
Move map_uid out of the metadata union and next to the other IDs, so
register comparisons can use the existing memcmp() ranges and remap the
IDs separately. Clear it when resetting a register or converting a map
lookup result to a socket pointer. Shrink frameno to u8, which is enough
for MAX_CALL_FRAMES, to make room without growing bpf_reg_state.
Fixes:
|
||
|
|
c26e97721b |
bpf: Apply CO-RE relocations before subprogram validation
check_subprogs() verifies that each subprogram ends in an exit or an
unconditional jump before in-kernel CO-RE relocations are applied. An
unresolved relocation can then replace that terminal instruction with an
invalid helper call. The resulting fall-through into another subprogram
breaks the CFG invariant used by postorder and stack liveness analysis,
which can write past their per-subprogram arrays.
Apply CO-RE relocations immediately after preparing the program BTF, before
subprogram discovery and validation. Keep func_info and line_info validation
after subprogram discovery because those records depend on the complete
subprogram layout.
Reject an ldimm64 first slot at the end of the instruction stream before
CO-RE can inspect its missing second slot. check_subprogs() previously
rejected this form before relocation processing because it is not a valid
subprogram terminator. Moving CO-RE ahead of check_subprogs() removes that
implicit protection, so perform an explicit check before applying
relocations.
Include core_relo_cnt when deciding whether to prepare program BTF. A load
that supplied only CO-RE relocation metadata previously skipped both BTF
setup and relocation processing.
Fixes:
|
||
|
|
fd16449a9b |
bpf: Preserve packet pointer class displacement in regsafe()
regsafe() maps packet pointer IDs between states and checks that each current register range is a subset of the corresponding explored register range. It does not, however, preserve the displacement between registers that share a packet pointer ID. This is unsound because packet range is shared by ID. A bounds check on one class member updates every member, and a later access can consume the range through another member. Commit |
||
|
|
261b61d373 |
bpf: Make post-verification instruction rewrites killable
After do_check() returns, the verifier runs several instruction rewrite
passes. Some of them patch or remove one instruction at a time. Each
operation moves the remaining instruction and auxiliary-data arrays and
adjusts all branch offsets, making the overall work quadratic in the
program length.
A privileged loader can submit 131072 unconditional jumps by zero followed
by a valid return. Verification finishes quickly, but bpf_opt_remove_nops()
then spends a long time removing each jump separately. Since this
post-verification work neither checks for signals nor reschedules, a pending
SIGKILL cannot terminate the task until the rewrite finishes.
Make bpf_patch_insn_data() and verifier_remove_insns() common cancellation
and rescheduling points. These helpers run from BPF_PROG_LOAD process
context, and bpf_patch_insn_data() can already sleep while reallocating
auxiliary data.
Report interrupted constant blinding as -EINTR and propagate it through
both JIT paths, including kernels that permit interpreter fallback.
Other blinding failures retain the existing fallback behavior.
This does not reduce the quadratic cost of the rewrite passes, but it makes
the work preemptible and allows a killed loader to be torn down promptly.
Fixes:
|
||
|
|
1d653a1839 |
fprobe: Terminate the fgraph_data list when the reservation is not filled
fprobe_fgraph_entry() reserves shadow stack space for every fprobe with
an exit handler, but only fills it for those whose entry handler returns
0. fgraph_reserve_data() does not clear the area, so fprobe_return()
parses the unused tail as headers left over from an earlier call, and an
exit handler can run twice or despite its entry handler asking to skip
it.
Write a zero word after the last entry to terminate the walk. A zeroed
slot does not decode to a NULL fprobe on the arches that encode the
header into one unsigned long, since arch_decode_fprobe_header_fp() ORs
in FPROBE_HEADER_MSB_PATTERN, so make read_fprobe_header() return NULL
for a zeroed slot.
Link: https://lore.kernel.org/all/20260917212407.384468-1-devnexen@gmail.com/
Fixes:
|
||
|
|
50e80e2bb5 |
bpf: Skip unsettled links in link iterator
bpf_link_prime() inserts a link into link_idr before anon_inode_getfile()
succeeds and before bpf_link_settle() publishes the ID in link->id.
bpf_link_by_id() treats such an ID-zero link as unsettled, but the link
iterator takes a reference without this check.
If anon_inode_getfile() then fails, the creator removes the ID and frees
its still-private link directly. The iterator is left with a dangling
reference and its next bpf_link_put() accesses freed memory.
Treat ID-zero entries as transient in bpf_link_get_curr_or_next(), just as
bpf_link_by_id() does.
BUG: KASAN: slab-use-after-free in bpf_link_put
Write of size 8 by task exp/384
Call Trace:
bpf_link_put kernel/bpf/syscall.c:3372
bpf_link_seq_next kernel/bpf/link_iter.c:33
bpf_seq_read kernel/bpf/bpf_iter.c:158
vfs_read fs/read_write.c:572
ksys_read fs/read_write.c:716
do_syscall_64 arch/x86/entry/syscall_64.c:84
entry_SYSCALL_64_after_hwframe arch/x86/entry/entry_64.S:121
Kernel panic - not syncing: KASAN: panic_on_warn set ...
Fixes:
|
||
|
|
40c2096961 |
bpf: Verify global subprogs in each sleepability context
Global subprograms are verified independently with a fresh verifier root. do_check_common() currently seeds that root's in_sleepable state from the program, even though a global subprogram can also run from callbacks whose execution context differs from the program's main entry point. In particular, workqueue and task-work callbacks are sleepable even when the containing program is not. A global subprogram of that program is therefore verified as non-sleepable, making in_rcu_cs() true and allowing loads of RCU-protected kptrs to produce trusted MEM_RCU pointers. The same subprogram can then be called from a sleepable callback without a classic RCU reader. It can retain such a pointer while the object is freed and use it after free. The verifier's execution-context predicates are complementary. A state is sleepable only when in_sleepable is set and no RCU, preemption, IRQ, or lock region is active. Each condition which prevents sleeping also provides RCU protection, while in_rcu_cs() treats a non-sleepable state as implicitly protected. Use this relationship to represent a global subprogram caller with only the result of in_sleepable_context(). A protected sleepable caller is normalized to in_sleepable=false at the independent verification root. This both prevents sleepable operations and makes in_rcu_cs() true without copying caller-owned lock state. Track only the contexts in which each global subprogram is actually reached. Verify it once if all reachable calls use the same context, and twice only if both sleepable and non-sleepable calls reach it. Calls found while verifying globals or asynchronous callbacks mark further contexts for checking. Repeat the existing subprogram walk until all called contexts have been verified; unreachable global calls remain unchecked. Accumulate instruction counts over those verification passes. Preserve the total recorded before each pass, since path accounting has already added this pass's synchronous instructions and its root total must also include asynchronous subprograms. This makes an unprotected callback verify the global subprogram as sleepable, turning its RCU-protected kptr load into an untrusted pointer. Protected callers and global subprograms which do not depend on implicit RCU protection remain valid. Fixes: |
||
|
|
cb86607ada |
sched_ext: Don't run ops.dequeue() with a DSQ lock held
ops.dequeue() is invoked with the source user DSQ's lock still held on
the consume and move paths (scx_consume_dispatch_q(),
move_task_between_dsqs()). A BPF scheduler which locks the source user
DSQ from ops.dequeue() - e.g. by iterating it with bpf_iter_scx_dsq -
self-deadlocks.
ops.dequeue() can only call the "any" kfuncs and none of them can lock a
builtin DSQ, so the global and bypass paths can't deadlock; however,
all DSQ locks share one lockdep class, so iterating any user DSQ from
ops.dequeue() on those paths trips the recursion check.
Move the invocation after the DSQ unlock on all three paths.
SCX_TASK_IN_CUSTODY is cleared under the lock serializing the transfer
so that the callback is invoked exactly once.
Fixes:
|
||
|
|
df5cdc2c83 |
sched_ext: Derive SCX_RQ_IN_WAKEUP from the core enqueue flags
schedule_deferred_locked() skips scheduling a deferred action while
SCX_RQ_IN_WAKEUP is set and relies on the task_woken_scx() call that follows
a wakeup enqueue to run it. enqueue_task_scx() sets the flag from the merged
enqueue flags, which include the flags stashed for a remote activation.
move_remote_task_to_local_dsq() thus sets SCX_RQ_IN_WAKEUP on the
destination rq when the moved task was woken up, although no
task_woken_scx() follows that activation.
An IMMED insert into a busy destination requests a local reenqueue during
that enqueue. The request gets linked but not scheduled and stays pending
until an unrelated wakeup or preemption on that CPU runs the deferred
actions. The IMMED task sits behind the running task in the meantime. If
nothing runs them before the scheduler is disabled, the request outlives the
scheduler and points into its freed per-cpu area, which the next scheduler
dereferences from run_deferred().
Test the core enqueue flags for the wakeup bit. Only the core's wakeup path
is followed by task_woken_scx().
Fixes:
|
||
|
|
4aec9ad1c6 |
dma-mapping fixes for Linux 7.3
A few fixes for the DMA-mapping code:
- resolved regression in accessing encrypted memory by IOMMU-backed
devices (Aneesh Kumar K.V),
- improved failure handling and removed rare bug in swiotlb/highmem
(Donggeun Yoo).
-----BEGIN PGP SIGNATURE-----
iHUEABYKAB0WIQSrngzkoBtlA8uaaJ+Jp1EFxbsSRAUCaquwMwAKCRCJp1EFxbsS
RMT6AP0elpdaZXNY0KwUBTwU95H604J+donqriepHABIBhIDEQD9GWZqNf/m1gEI
tR5lHQ3+NGs0Q7Vd2ed1vSe82HQSsgU=
=QxUC
-----END PGP SIGNATURE-----
Merge tag 'dma-mapping-7.3-2026-09-17' of git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux
Pull dma-mapping fixes from Marek Szyprowski:
"A few fixes for the DMA-mapping code:
- resolved regression in accessing encrypted memory by IOMMU-backed
devices (Aneesh Kumar K.V)
- improved failure handling and removed rare bug in swiotlb/highmem
(Donggeun Yoo)"
* tag 'dma-mapping-7.3-2026-09-17' of git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux:
x86/mm: Don't force unencrypted DMA for IOMMU-backed devices
dma-mapping: don't trace the DMA address when the allocation fails
swiotlb: use the adjusted address for the highmem page lookup
dma-coherent: report a failed reserved memory assignment
|
||
|
|
d2710c8d93 |
signal: Prevent exec() race
Hyunwoo debugged the following KASAN UAF splat:
BUG: KASAN: slab-use-after-free in __send_signal_locked+0xb27/0xba0
Write of size 8 at addr ffff888007ed80c8 by task poc/79
...
Call Trace:
__send_signal_locked+0xb27/0xba0
do_send_sig_info+0xa7/0x160
do_send_specific+0x76/0xa0
__x64_sys_tgkill+0x193/0x270
...
Allocated by task 80:
do_timer_create+0x1a4/0x1030
__x64_sys_timer_create+0x145/0x190
...
Freed by task 12:
kmem_cache_free_bulk+0x1f8/0x4a0
kvfree_rcu_bulk+0x14f/0x1c0
kfree_rcu_work+0x128/0x1a0
...
Last potentially related work creation:
kvfree_call_rcu+0x39/0x390
__flush_itimer_signals+0x211/0x320
flush_itimer_signals+0x47/0x90
begin_new_exec+0xa6b/0x28c0
It turned out that this happens with a non-leader exec() as Hyunwoo
explained:
de_thread() calls exchange_tids() before release_task(leader), so the
struct pid held by a SIGEV_THREAD_ID timer created against the leader's tid
now points to the thread which called execve(). pid_task() returns that
thread and lock_task_sighand() on it succeeds.
If the timer signal is blocked, its sigqueue stays queued on the leader's
task::pending. The next expiry of that timer can then run while
release_task() flushes the queue.
posixtimer_send_sigqueue() checks whether the sigqueue is already queued
with a plain list_empty(), which only reads list_head::next.
list_del_init() is not atomic and INIT_LIST_HEAD() stores list_head::next
before list_head::prev, so the check can pass in between. list_add_tail()
queues the entry on the task::pending of the live thread, and the
list_head::prev store from the flush then overwrites the list_head::prev
link that list_add_tail() has just set.
__flush_itimer_signals() does not undo that either. With list_head::prev
pointing at the entry itself, its list_del_init() only stores the same
values again, so the entry is not removed from the list. It is still there
after the last reference is dropped and the timer is freed by RCU, and the
list_add_tail() of a later tgkill() follows that list_head::prev into the
freed timer.
This problem surfaced with the recent commit which moved the sigqueue flush
out of the sighand lock held region.
Hyonwoo proposed to fix this by using list_del_init_careful(), but that
just papers over the problem. After some disucssions and various attempts
to solve it, Eric pointed out that there is no reason to flush
task::pending late in release_task() and it should be done in
exit_signals() already.
As nothing can collect and deliver signals which are queued in a dying
task's pending queue, there is no reason to delay it further.
But it has to be ensured that no signals can be queued into it after that
point. exit_signals() sets PF_EXITING in task::flags, which can be used as
an indicator for this.
Cure it by:
- Preventing signal queueing for task private signals (PIDTYPE_PID) when
the task has PF_EXITING set in __send_signal_locked() and in
posixtimer_send_sigqueue().
- Protecting the unlocked setting of PF_EXITING in exit_signals() for the
task group empty and the group exit case with sighand lock
- Flushing task::pending signals right there.
Optimize that by moving the whole pending list to an on-stack list head
under sighand lock and free the signals without the lock held.
There has been quite some discussion about the lockless flush and the
non-leader exec case on weakly ordered systems. The problem is that a third
party which tries to send a posix timer signal relies on the PID lookup to
find the target task and that lookup might result in the new leader when
the signal was originaly directed to the old leader. In case that the
signal was queued on the old leader then the lockless flush raised a
concern over the following situation:
old_leader new_leader third party
A: flush_list() // list_del_init() stores to sigqueue
LOCK (tasklist)
old_leader->exit_state = EXIT_ZOMBIE;
B: UNLOCK (tasklist)
C: LOCK (tasklist)
if (old_leader->exit_state)
transfer_tids()
D: store PID
posix_timer_send_sigqueue()
// Observes #D so t = new_leader
E: t = get_target()
F: LOCK (sighand)
G: if (list_empty(sigqueue))
list_add(sigqueue)
The concern was that the third party might observe #D but not observe #A
and therefore would proceed to #G while the list_del() stores (#A) in
flush_list() are not visible yet, which could result in list corruption.
That would be possible if looking at it solely from a RELEASE+ACQUIRE
ordering point of view, but B-C is a UNLOCK+LOCK hand-over, which is not
the same as RELEASE+ACQUIRE:
RELEASE+ACQUIRE: RCpc, only the CPUs involved agree on the ordering
UNLOCK+LOCK: RCtso, the hand-over is store-ordering
As B-C is UNLOCK+LOCK, which is RCtso and that does impose store order,
A stores must happen before the D store.
Combine with E-F, which has a data dependency from the LOAD to the LOCK and
thereby constraints later LOADs, those sigqueue loads in G that come after
F must in fact observe the A stores.
Fixes:
|
||
|
|
b61b6f95d6 |
futex: Also allocate private hash on vfork()
As Jann demonstrated, it is entirely feasible to access the mm through vfork().
Therefore we need to allocate a private hash on vfork() as well as any other
CLONE_VM user.
Specifically, it must be avoided to have (private) futex waiters before
allocating the private hash.
Fixes:
|
||
|
|
7de9a6fb44 |
sched_ext: Wait for SCX_OPSS_DISPATCHING before reenqueueing a task
|