Please consider pulling these changes from the signed vfs-7.3-rc5.fixes tag.
Thanks!
Christian
-----BEGIN PGP SIGNATURE-----
iHUEABYKAB0WIQRAhzRXHqcMeLMyaSiRxhvAZXjcogUCarab9AAKCRCRxhvAZXjc
osTRAP90MigSgX2U/USBn8zgSTs9Key89pWuwpctsvELYKUSzwEA12UDBswbtgmq
EGT6S2HSa6bifJMXG6AB+nMCcSluuAs=
=KEgE
-----END PGP SIGNATURE-----
Merge tag 'vfs-7.3-rc5.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs
Pull vfs fixes from Christian Brauner:
- Revert "put_mnt_ns(): leave mounts connected". This allows the
creation of reference count cycles in a very trivial way. We can't
bring this in until we have fixed the underlying cause
- vfs: Don't create the private nullfs instance for kthreads under
namespace_sem to avoid false lockdeps complaints
- binfmt_misc:
- Copy the name into a stack buffer and look up the copy in
bpf_binprm_select_interp()
- bpf_binprm_set_interp() and bpf_binprm_set_interp_arg(): Check
the private copy instead so the string that gets staged is the
kstring that was checked
- netfs:
- Make netfs_read_gaps() use separate sink folios rather than one
reused sink folio to discard unwanted data so that cifs checksum
checking sees all the data that was fetched
- Trim reads down to i_size so afs symlinks read correctly from the
cache
- Wrap the direct mempool ->alloc() calls the GFP_KERNEL paths make
in alloc_hooks() via a new mempool_alloc_noreserve() helper
- iov_iter: Use iov_iter_alignment() for the start and length check
added to iov_iter_extract_bvecs() this cycle. It used iter_iov_addr()
and iter_iov_len() which are only valid for ITER_UBUF and ITER_IOVEC
iterators
- super: Make iterate_supers_type() deletion-safe
- inode: Stop evict_inodes() from rescanning the same inodes
- writeback: Bound the cleanup_offline_cgwb() rescans
- ntfs3: Use d_instantiate_new() in ntfs_create_inode()
- ovl: Fix a use-after-free in the ovl_do_mkdir() debug print
- dcache: Unpoison the inline name buffer in __d_alloc() for KMSAN
- autofs: Fix a pipe file reference leak in autofs_kill_sb()
- bpf: Drop the path_unlink and path_rmdir hooks from the list of hooks
for which the verifier rewrites bpf_{set,remove}_dentry_xattr() to
the _locked variants
- squashfs: Range check the xz dictionary size before shifting by it
- selftests: Add the missing eventfd, open_tree_ns, openat2 and xattr
filesystems selftests to TARGETS and drop the stale openat2 entry
left behind when those tests moved
* tag 'vfs-7.3-rc5.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs:
netfs: Fix missing alloc tagging of direct mempool allocations
bpf: fs/xattr: don't assume the inode is locked in path_unlink/path_rmdir
autofs: fix sbi->pipe file reference leak in autofs_kill_sb()
dcache: unpoison the inline name buffer in __d_alloc()
ovl: fix UAF in ovl_do_mkdir() debug print
super: make iterate_supers_type() deletion-safe
Revert "put_mnt_ns(): leave mounts connected"
Revert "selftests/filesystems: add mntns cleanup test"
binfmt_misc: fix racy checks in bpf set_interp kfuncs
binfmt_misc: fix OOB read in bpf_binprm_select_interp()
fs: don't create the private nullfs mount under namespace_sem
writeback: bound cleanup_offline_cgwb() rescans by rotating scanned inodes
fs: avoid repeated scans in evict_inodes()
netfs, afs: Fix symlink reading
netfs: Fix netfs_read_gaps() to use separate sink folios
squashfs: Add dictionary size range check to prevent shift-out-of-bounds
fs/ntfs3: use d_instantiate_new() in ntfs_create_inode() and murder syzbot's "WARNING in do_new_mount" saga
selftests/filesystems: fix missing and stale TARGETS entries
block: Fix start and length check added to iov_iter_extract_bvecs()
-----BEGIN PGP SIGNATURE-----
iQEzBAABCAAdFiEEq1nRK9aeMoq1VSgcnJ2qBz9kQNkFAmq2Qs8ACgkQnJ2qBz9k
QNnBEAf5AXGzuQPeOHZQ0+kXhe0y+2Epj1oqSH0YIDI2263nav9go3wjz4uyL/J2
0bbALzEX0Em0n0OfYOSKPdOx22AFZdZVKrFiACXeaUQmK1R0zwLOJPxDvpyWhHQS
8Rn+HxMVxdqbobIClHQvGsU3EkBZR+d2ZSDfG4A3Gqc4O2vkgB3CvDvw5d7OdDNB
5Et2tydekYTSUH2JZvJxlzwvbfsvvgSEEVupYILYZvMOv01EMxOTLCCv/8KWOFNE
Xjen/maE+1ks+nGmTgDQC2bOALB3nCQhoMlL31nsYzP89Dg/C0IGCbVX25yUvlzX
0FXMPKVVEEvpZVhTc5Hy7HYhDatC5Q==
=IeuV
-----END PGP SIGNATURE-----
Merge tag 'fs_for_v7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/jack/linux-fs
Pull isofs fix from Jan Kara:
"A fix for reading tightly packed isofs directories"
* tag 'fs_for_v7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/jack/linux-fs:
isofs: Fix handling of directories with tight blocks
Fix a step-wise thermal governor issue that causes thermal mitigation to
contiune forever after the temperature has dropped below the trip point
threshold in some cases (Manaf Meethalavalappu Pallikunhi).
-----BEGIN PGP SIGNATURE-----
iQFGBAABCAAwFiEEcM8Aw/RY0dgsiRUR7l+9nS/U47UFAmq2l84SHHJqd0Byand5
c29ja2kubmV0AAoJEO5fvZ0v1OO1NAAH/2pYr1erbmTfHAkmka5b+4/nJsAONclz
izc6GPbLdM7PgR7LnAxAlJzQCz9wIljxtOHYKviZmlrnXbVwQxLZLI/GLpF5LN5Z
au79wSH/q1UHsBbLUIUKp887WVEvxeS/jC6//YFUkjY+gVCiRpvKUHkLr8WzQYue
DD1cQNR+nLquYao04+Src02VfbjKhRouMoHNhSkJXOm4IbkPMA75ENBw6LGfWshf
tm1uSXghzog0kNNU3rqjLOdMdIpp1Ea/l1wCfg8mkwd+AVZS2G3eItQbrLQmXPWg
yjKWpmTPggletPB4bz7LjMwkHwI+xzLIWSbr0j5zRqEopedSHhc0MXk=
=VygM
-----END PGP SIGNATURE-----
Merge tag 'thermal-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm
Pull thermal control fix from Rafael Wysocki:
"Fix a step-wise thermal governor issue that causes thermal mitigation
to contiune forever after the temperature has dropped below the trip
point threshold in some cases (Manaf Meethalavalappu Pallikunhi)"
* tag 'thermal-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm:
thermal: gov_step_wise: Fix stale mitigation vote with non-zero lower bounds
Address a hibernation regression introduced during the 7.2 development
cycle that causes the image memory preallocation to deadlock if it
depends on frozen kernel threads (Florian Schmaus).
-----BEGIN PGP SIGNATURE-----
iQFGBAABCAAwFiEEcM8Aw/RY0dgsiRUR7l+9nS/U47UFAmq2mDESHHJqd0Byand5
c29ja2kubmV0AAoJEO5fvZ0v1OO1wFAH/2e9iz++Qjb9SODpGX/2Fz07qb4SBRCP
Zf0r74G7qfoPczcLuiKu8irb1FSvlwr1nlcygWcF0gYLg9TJCaaQ7JIxHL/9ePV0
lbINk+4ozu5S6AbMh7O5wpv3n+nwBtg1wZZP3kSY4hxQ5zymACuEzwBTGT9vfQVA
GpkrasgQVTOyt1gWAO8Ak3WX3z1EaBqzl8DsCm/75PVq2Wy1I801JagtYT6KArf6
kqH19SdiihrTl+2k/k2Vvc14H9XYfXab77ShWv1JgKF7X1QqaO8QWVFQulj0bpgc
LLSlpiQXqBYfz469bPGyS2U6tyFioynqoiWQfdxC6zBGcHv1Jl3LdDg=
=mfRN
-----END PGP SIGNATURE-----
Merge tag 'pm-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm
Pull power management fix from Rafael Wysocki:
"Address a hibernation regression introduced during the 7.2 development
cycle that causes the image memory preallocation to deadlock if it
depends on frozen kernel threads (Florian Schmaus)"
* tag 'pm-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm:
PM: hibernate: Freeze kernel threads after image preallocation
- Fix several bugs in PCI error recovery SCLP reporting: don't report
success on skipped recovery, report errors when no pdev is associated,
add missing device lock, and fix struct pci_dev reference leak in
zpci_report_status()
- Fix several bugs in CIO code: fix use of invalid SCHIB data, guard PMCW
field accesses, check device number valid bit in PMWC before accessing
other fields, and fix NULL pointer dereference in
ccw_device_get_util_str()
- Fix virtual vs physical address confusion in channel measurement
facility code on kernels with CONFIG_RANDOMIZE_IDENTITY_BASE=y
- Fix couple of bugs in s390dbf: fix copy of failed static debug areas,
skip view registration on failure, and reject NULL pointer in
debug_dump()
- Fix sriov_numvfs attribute name in zPCI documentation
- Fix typos in comments
-----BEGIN PGP SIGNATURE-----
iQIzBAABCgAdFiEECMNfWEw3SLnmiLkZIg7DeRspbsIFAmq2V1AACgkQIg7DeRsp
bsJ7zg/7BrnPs0WyKv9NZvVUudYbz3nUIAleJo6HcDK2SUDr+EX3W4U7OjsXL6Fv
H76hx5AqF6ddtUOJ2uevCAb3PKm0eA7vmK9YiERAclrBA+tWGW0xBwCqqyc84k38
CPxOBr87Ae1jIIcfDxfBWujtKTFi/4BpSvKvKgg3mnd2egny4UnHnABx5ZcClEee
p6aVjCKycQFHwNLMoJA1BxvGSpcXdiKXbm6VRJ4gCaBnVuzkLUv1ZMp/Pi1fWVbV
E1AEztZ/03Tfry51q7wQLv6wD5lR3Z/Py2vZ4YIbelQwVMAMI+rFtgd5oFfo9qkF
gcFJ+faE9m0E/nsoaaVCaa2Ja4+hV/vZ5ik0bq1ljEqEvInaa2XTuxNAljId351Y
JrLvCdaip0FkMNMqEym0iydMAlpc491GAhzNrT3yvq6t03VVpxprPPA6VIwHg+zk
UBMP+QdxtnaV8AAUzDD9E255d88N4DECBICUgf1XjFm/I7xLJFWyUvbzcqNn9VnY
EGVBOeQmoByhCL9lN0+nzEM1sgT+hQB/C3WmLqiuLxrxjnc8dVr8rVHcPXFgARPw
Ht3W9jxdTdl4VBwckOWmbvG4L6spXFLdEzI51cP46L1RKWhFjLl/6RoY8fUHf4Vz
YnHutQevSoPI2D6G4v8K1FBz9wclxCojBPbXCmevfbL9B4ChPoQ=
=r3uP
-----END PGP SIGNATURE-----
Merge tag 's390-7.3-4' of git://git.kernel.org/pub/scm/linux/kernel/git/s390/linux
Pull s390 fixes from Heiko Carstens:
- Fix several bugs in PCI error recovery SCLP reporting: don't report
success on skipped recovery, report errors when no pdev is
associated, add missing device lock, and fix struct pci_dev reference
leak in zpci_report_status()
- Fix several bugs in CIO code: fix use of invalid SCHIB data, guard
PMCW field accesses, check device number valid bit in PMWC before
accessing other fields, and fix NULL pointer dereference in
ccw_device_get_util_str()
- Fix virtual vs physical address confusion in channel measurement
facility code on kernels with CONFIG_RANDOMIZE_IDENTITY_BASE=y
- Fix couple of bugs in s390dbf: fix copy of failed static debug areas,
skip view registration on failure, and reject NULL pointer in
debug_dump()
- Fix sriov_numvfs attribute name in zPCI documentation
- Fix typos in comments
* tag 's390-7.3-4' of git://git.kernel.org/pub/scm/linux/kernel/git/s390/linux:
s390/cio: Fix NULL pointer dereference in ccw_device_get_util_str()
s390/debug: Fix NULL pointer dereference in debug_info_copy()
s390/debug: Do not register views for failed static debug areas
s390/debug: Reject NULL debug info in debug_dump()
s390/cmf: Fix virtual vs physical address confusion
s390/pci: Don't report recovery success on skipped recovery
s390/pci: Report SCLP status on error events when no pdev is associated
s390/pci: Fix missing device lock in zpci_report_status()
s390/pci: Fix leak of struct pci_dev reference in zpci_report_status()
s390/cio: Guard PMCW field accesses with dnv check
s390/cio: Check pmcw.dnv before pmcw.ena in I/O entry points
s390/cio: Fix cio_update_schib() to not cache invalid schib
s390/pci/docs: Fix sriov_numvfs attribute name
s390: Fix typos in comments
- fix a regression introduced by moving GPIO hog handling into GPIOLIB
core where of_node_name was used if line name property was missing on
DT systems
- fix kernel stack leak to user-space in error path in GPIO character
device code
- fix runtime PM leaks in gpio-xilinx and gpio-arizona
- fix several register programming bugs in gpio-tps65219
- fix interrupt storm on resume in gpio-mvebu
-----BEGIN PGP SIGNATURE-----
iQIzBAABCgAdFiEEkeUTLeW1Rh17omX8BZ0uy/82hMMFAmq2KYwACgkQBZ0uy/82
hMOkFhAAjIaH6nUhhUBykRyIbBVAbxYFWrps/qLw9Wlg1wDteUwwMWpdHWivsi2M
Ch0zCR1/op812cu47dTsnNteHuK9oGoL0IHXGweylNJVU7IXTiUx6sUu+kLlT2wK
/yIH6Zk78rULodV/80S9aNbFNNHOu7Pu/AEAPPJjGIKK6UVxT7o2o04jG1xORVbr
X11hLjzKe1RpVKo3fhBFglzW1T4YVpulorsgkMNabTG4hbf5pArtkaQGPHVvH3Hk
LNpTFJ3Jmoczo4UgsKAaUso2LoC19hoIUMRBVoOwscTLlPoIE6MTH0bLj+U9npId
TUhGpVYIvfy4B6V2+FU+9YnWTkDE7efmnmm0l+GnNp7FmDTxTO0aVJoouPNEeUgu
Zu1CJxaKYTWuTm5mdk2Tgw5G+3YtzcGS6iW3dCFvshyYSdHHkvPQ6ys2blJJyeEd
2MRoZQZdSzMDmCfw5Y+RxF1jp1BRjIID2NH1p/cxUokFvMDgFJU+MFxdOo6ySnC/
+1EEeA7Gak4StJUDa2Y4+uI0PMokWAEXs3ixlQUmgIpgGVtBIzI17psj4RDumdGD
6HIG/31Pedg3LtUkkzC8P6vfdREzzquWUPBr9F7dApur4/MMJZyeu3pUwFtx0AKs
+T5qgUwwsoKxLmg2EAdg5pThzLVgqgtrSYuw8esqfG8tNRzoRgc=
=Mp0f
-----END PGP SIGNATURE-----
Merge tag 'gpio-fixes-for-v7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/brgl/linux
Pull gpio fixes from Bartosz Golaszewski:
- fix a regression introduced by moving GPIO hog handling into GPIOLIB
core where of_node_name was used if line name property was missing on
DT systems
- fix kernel stack leak to user-space in error path in GPIO character
device code
- fix runtime PM leaks in gpio-xilinx and gpio-arizona
- fix several register programming bugs in gpio-tps65219
- fix interrupt storm on resume in gpio-mvebu
* tag 'gpio-fixes-for-v7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/brgl/linux:
gpio: tps65219: Fix TPS65214 GPIO direction programming
gpio: tps65219: Use the variant-specific direction callback
gpio: tps65219: Fix GPIO input value reads
gpio: zynq: fix runtime PM leak on request error path
gpio: cdev: fix kernel stack leak to user-space in error path
gpiolib: use of_node_name if line-name is missing
gpio: mvebu: keep resume masks within the irqchip cache
gpio: arizona: Fix runtime PM leak in arizona_gpio_direction_out()
Commit 1d78d56c43 ("netfs: Fix folio_queue ENOMEM in writeback by
adding a mempool") added a mempool for the folio_queues and made the
request, subrequest and folio_queue allocations distinguish between
writeback and everything else. Writeback is part of memory reclaim
and must not fail due to ENOMEM, so it allocates under GFP_NOFS
through mempool_alloc(), which may dip into the pool's reserve and,
if that runs empty, wait for elements to be returned. The
GFP_KERNEL paths, which can return -ENOMEM to their callers, invoke
the pool's ->alloc() callback directly instead.
The direct call, however, skips the alloc_hooks() wrapper that the
mempool_alloc() macro provides. The pool callbacks, mempool_alloc_slab()
and mempool_kmalloc(), call kmem_cache_alloc_noprof() and kmalloc_noprof()
and rely on current->alloc_tag having been set by the caller. With
CONFIG_MEM_ALLOC_PROFILING_DEBUG=y this leads to
current->alloc_tag not set
WARNING: ./include/linux/alloc_tag.h:161 at __alloc_tagging_slab_alloc_hook
alloc_tag was not set
WARNING: ./include/linux/alloc_tag.h:166 at __alloc_tagging_slab_free_hook
at allocation and free time respectively, as reported when reading
files on a CIFS mount. The allocations are also missing from
/proc/allocinfo.
Wrap the direct ->alloc() invocations in alloc_hooks() with a new
mempool_alloc_noreserve() helper in include/linux/mempool.h, next to
the other alloc_hooks()-wrapped macros such as mempool_alloc(). The
GFP_KERNEL paths keep their failable allocation semantics, they just
get tagged now.
Fixes: 1d78d56c43 ("netfs: Fix folio_queue ENOMEM in writeback by adding a mempool")
Reported-by: Erhard Furtner <erhard_f@mailbox.org>
Closes: https://lore.kernel.org/all/0b004319-9ef7-437c-a4dd-174d6a9a83db@mailbox.org/
Tested-by: Erhard Furtner <erhard_f@mailbox.org>
Suggested-by: Suren Baghdasaryan <surenb@google.com>
Cc: stable@vger.kernel.org
Signed-off-by: Hao Ge <hao.ge@linux.dev>
Link: https://patch.msgid.link/20260923063759.34667-1-hao.ge@linux.dev
Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
bpf_lsm_has_d_inode_locked() makes the verifier rewrite
bpf_[set|remove]_dentry_xattr() to the _locked variants, which assume
that the caller already holds the inode's i_rwsem. The path_unlink and
path_rmdir hooks are listed, but security_path_unlink() and
security_path_rmdir() run before vfs_unlink()/vfs_rmdir() take the
victim inode's i_rwsem, so a sleepable BPF LSM program attached to
either hook mutates the victim's xattrs without the lock held.
Drop the two path hooks from d_inode_locked_hooks so that the verifier
keeps the locking bpf_[set|remove]_dentry_xattr() variants, which take
the lock themselves.
Fixes: 5646729279 ("bpf: fs/xattr: Add BPF kfuncs to set and remove xattrs")
Cc: stable@vger.kernel.org
Signed-off-by: Andrea Parri <parri.andrea@gmail.com>
Link: https://patch.msgid.link/20260922145530.369775-1-parri.andrea@gmail.com
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
When autofs_fill_super() fails before clearing AUTOFS_SBI_CATATONIC (for
example, when find_get_pid() fails on an invalid pgrp mount option, or
when an fs_context is closed before mounting), deactivate_locked_super()
invokes autofs_kill_sb() -> autofs_catatonic_mode(sbi).
Because AUTOFS_SBI_CATATONIC is still set in sbi->flags,
autofs_catatonic_mode() returns early without calling fput(sbi->pipe),
permanently leaking the pipe struct file reference.
Explicitly release sbi->pipe in autofs_kill_sb() if it is still non-NULL
after autofs_catatonic_mode().
Fixes: ebc921ca9b ("autofs: copy autofs4 to autofs")
Signed-off-by: Hui Peng <benquike@gmail.com>
Link: https://patch.msgid.link/20260919204808.2812930-1-benquike@gmail.com
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
syzbot reported:
BUG: KMSAN: uninit-value in dentry_string_cmp fs/dcache.c:291 [inline]
BUG: KMSAN: uninit-value in dentry_cmp fs/dcache.c:322 [inline]
BUG: KMSAN: uninit-value in __d_lookup_rcu+0x37d/0x5e0 fs/dcache.c:2522
dentry_string_cmp fs/dcache.c:291 [inline]
dentry_cmp fs/dcache.c:322 [inline]
__d_lookup_rcu+0x37d/0x5e0 fs/dcache.c:2522
lookup_fast+0x194/0xa40 fs/namei.c:1854
lookup_fast_for_open fs/namei.c:4545 [inline]
open_last_lookups fs/namei.c:4579 [inline]
path_openat+0x9ef/0x6540 fs/namei.c:4856
do_file_open+0x2aa/0x680 fs/namei.c:4888
do_sys_openat2+0x17c/0x390 fs/open.c:1395
do_sys_open fs/open.c:1401 [inline]
__do_sys_openat fs/open.c:1417 [inline]
__se_sys_openat fs/open.c:1412 [inline]
__x64_sys_openat+0x240/0x300 fs/open.c:1412
x64_sys_call+0x2445/0x3ea0 arch/x86/include/generated/asm/syscalls_64.h:258
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0x15d/0x3c0 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
Uninit was stored to memory at:
copy_name fs/dcache.c:3031 [inline]
__d_move+0xd29/0x21f0 fs/dcache.c:3099
d_move+0x71/0xf0 fs/dcache.c:3147
vfs_rename+0x2619/0x2770 fs/namei.c:6085
filename_renameat2+0xa59/0x1230 fs/namei.c:6188
__do_sys_rename fs/namei.c:6232 [inline]
__se_sys_rename+0xc5/0x5c0 fs/namei.c:6228
__x64_sys_rename+0x78/0xb0 fs/namei.c:6228
x64_sys_call+0x329/0x3ea0 arch/x86/include/generated/asm/syscalls_64.h:83
do_syscall_x64 arch/x86/entry/syscall_64.c:63
do_syscall_64+0x15d/0x3c0 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
Uninit was created at:
slab_post_alloc_hook mm/slub.c:4617 [inline]
slab_alloc_node mm/slub.c:4939 [inline]
kmem_cache_alloc_lru_noprof+0x376/0x1230 mm/s
__d_alloc+0x52/0x9f0 fs/dcache.c:1902
d_alloc+0x57/0x300 fs/dcache.c:1981
lookup_one_qstr_excl+0x19d/0x7a0 fs/namei.c:1806
__start_renaming+0x341/0x850 fs/namei.c:3888
filename_renameat2+0x625/0x1230 fs/namei.c:6163
__do_sys_rename fs/namei.c:6232 [inline]
__se_sys_rename+0xc5/0x5c0 fs/namei.c:6228
__x64_sys_rename+0x78/0xb0 fs/namei.c:6228
x64_sys_call+0x329/0x3ea0 arch/x86/include/generated/asm/syscalls_64.h:83
do_syscall_x64 arch/x86/entry/syscall_64.c:63
do_syscall_64+0x15d/0x3c0 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
The race is between a concurrent open() and rename() of the same path.
__d_alloc() only stores the name itself and its terminating NUL, so the
rest of the inline buffer (d_shortname, DNAME_INLINE_LEN bytes) is left
uninitialized. copy_name(), called from rename(), copies that buffer as a
whole, so the uninitialized tail is propagated into the dentry that is
being moved. Meanwhile __d_lookup_rcu(), called from open(), is an
optimistic lockless lookup: it checks d_name.hash_len first and leaves the
seqcount retry to its caller, so it can end up comparing against a dentry
whose name a rename is rewriting in place, using a stale (longer) length.
The comparison then runs past the terminating NUL and reads bytes of the
uninitialized tail, which KMSAN reports.
The read is harmless by design: it stays inside the buffer, the name is
still NUL-terminated, and the result is thrown away by the seqcount retry.
It is not specific to KMSAN either - with CONFIG_DCACHE_WORD_ACCESS
enabled the very same bytes are read by read_word_at_a_time(), which is
__no_sanitize_or_inline and therefore invisible to KMSAN. KMSAN builds
only see the instrumented byte-at-a-time dentry_string_cmp() because
CONFIG_DCACHE_WORD_ACCESS is disabled when KMSAN is enabled on x86:
commit 7cf8f44a5a ("x86: fs: kmsan: disable CONFIG_DCACHE_WORD_ACCESS")
Zeroing the inline buffer would hide the report, but it would add a
memset() to a hot allocation path just to initialize bytes that are never
used as part of a name. Instead, tell KMSAN the inline buffer is
initialized: kmsan_unpoison_memory() compiles to nothing unless
CONFIG_KMSAN is set, and doing it at allocation time is enough for every
dentry, because copy_name() and swap_names() copy the whole buffer and
thus propagate its shadow.
Reported-by: syzbot+7ff3adde89dd795ad4c4@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=7ff3adde89dd795ad4c4
Signed-off-by: Drif Abdelmalek Mohamed Said <drifabdelmalekmohamedsaid@gmail.com>
Changes in v3:
- Annotate for KMSAN instead of zeroing, as suggested in review: the read
is harmless, so unpoison the inline buffer in __d_alloc() with
kmsan_unpoison_memory() (a no-op unless CONFIG_KMSAN) rather than adding
a memset() to the dentry allocation path.
- Document why only KMSAN builds report this at all: with
CONFIG_DCACHE_WORD_ACCESS the same read goes through
read_word_at_a_time(), which KMSAN does not instrument.
- Rewrite the commit message; the previous one had several truncated
lines.
Link: https://patch.msgid.link/20260918224204.3056-1-drifabdelmalekmohamedsaid@gmail.com
Reviewed-by: Jan Kara <jack@suse.cz>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
Fix a race in the cdev layer that can cause a fw_iso_resource_auto object
to transition back to a previous state. This can happen when a file
descriptor is closed while the work item for the object is running. The
race can leak several memory objects, including client object itself.
This fix should be applied to 7.2 kernel or later.
-----BEGIN PGP SIGNATURE-----
iHUEABYKAB0WIQQE66IEYNDXNBPeGKSsLtaWM8LwEwUCarWseAAKCRCsLtaWM8Lw
E2x/AQCr6GAMO1G8/7mUgEj3X4CHCzkS/N3bVwb6TJ18/e5gxgEAuM9Ry+FU6IP4
0NCz6dbcn9dta3YGnlfgAQ9HtTMwnAg=
=BLzx
-----END PGP SIGNATURE-----
Merge tag 'firewire-fixes-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/ieee1394/linux1394
Pull firewire fix from Takashi Sakamoto:
"Fix a race in the cdev layer that can cause a fw_iso_resource_auto
object to transition back to a previous state. This can happen when a
file descriptor is closed while the work item for the object is
running. The race can leak several memory objects, including client
object itself"
* tag 'firewire-fixes-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/ieee1394/linux1394:
firewire: cdev: fix back-transition for iso_resource_auto client resource
- A task reenqueued while its dispatch was still completing had its queued
state clobbered by the dispatcher, dropping every later dispatch of the
task. Wait for the in-flight dispatch to settle first.
- A wakeup activation on another CPU marked the destination runqueue as
mid-wakeup, stranding a pending local reenqueue. If the scheduler was
unloaded first, the stale request pointed into freed memory that the next
scheduler dereferenced.
- ops.dequeue() ran with the source dispatch queue's lock held, so a
scheduler iterating that queue from the callback deadlocked the CPU.
- Schedulers with their own CPU ID mapping had no way to learn a task's
initial CPU mask and rebuilt it themselves, which went wrong across
sub-scheduler enable and re-home. Pass it to ops.enable().
- A bypass dispatch event counter missed the dispatches made by the
end-of-dispatch fallback and under-reported.
- Selftests for the dequeue locking and initial mask changes.
-----BEGIN PGP SIGNATURE-----
iIQEABYKACwWIQTfIjM1kS57o3GsC/uxYfJx3gVYGQUCarWSLQ4cdGpAa2VybmVs
Lm9yZwAKCRCxYfJx3gVYGY+sAP9y6nJh6vIvFh/X9FlJtWlNo0mncOKhy93E8jii
8CKnPQEAhvX3+Gcdl+imTh4Z915kdsEByBjTTPPOXnQxI8BKUAY=
=VhBF
-----END PGP SIGNATURE-----
Merge tag 'sched_ext-for-7.3-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext
Pull sched_ext fixes from Tejun Heo:
- A task reenqueued while its dispatch was still completing had its
queued state clobbered by the dispatcher, dropping every later
dispatch of the task. Wait for the in-flight dispatch to settle
first
- A wakeup activation on another CPU marked the destination runqueue as
mid-wakeup, stranding a pending local reenqueue. If the scheduler was
unloaded first, the stale request pointed into freed memory that the
next scheduler dereferenced
- ops.dequeue() ran with the source dispatch queue's lock held, so a
scheduler iterating that queue from the callback deadlocked the CPU
- Schedulers with their own CPU ID mapping had no way to learn a task's
initial CPU mask and rebuilt it themselves, which went wrong across
sub-scheduler enable and re-home. Pass it to ops.enable()
- A bypass dispatch event counter missed the dispatches made by the
end-of-dispatch fallback and under-reported
- Selftests for the dequeue locking and initial mask changes
* tag 'sched_ext-for-7.3-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext:
sched_ext: Count SCX_EV_SUB_BYPASS_DISPATCH in the dispatch fallback
selftests/sched_ext: Check the cmask cid-form ops.enable() receives
sched_ext: Pass the initial cmask to cid-form ops.enable()
selftests/sched_ext: Test that ops.dequeue() can iterate the consumed DSQ
sched_ext: Don't run ops.dequeue() with a DSQ lock held
sched_ext: Derive SCX_RQ_IN_WAKEUP from the core enqueue flags
sched_ext: Wait for SCX_OPSS_DISPATCHING before reenqueueing a task
- With local event accounting, a fork rejected by the pids controller
updated pids.events without notifying its pollers.
- A cgroup selftest failed to compile with fortification enabled because
an O_TMPFILE open lacked its mode argument.
-----BEGIN PGP SIGNATURE-----
iIQEABYKACwWIQTfIjM1kS57o3GsC/uxYfJx3gVYGQUCarWSKg4cdGpAa2VybmVs
Lm9yZwAKCRCxYfJx3gVYGYoiAQD6JUCqDjv2Hr1YeMFlsoYVQZ7tNmljVnPD2tZ6
9WmD/AEAotdlmzM8egOuDqAi2s+UMJPm8vCZuxVzewiGqhYELgk=
=0Ljp
-----END PGP SIGNATURE-----
Merge tag 'cgroup-for-7.3-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup
Pull cgroup fixes from Tejun Heo:
- With local event accounting, a fork rejected by the pids controller
updated pids.events without notifying its pollers
- A cgroup selftest failed to compile with fortification enabled
because an O_TMPFILE open lacked its mode argument
* tag 'cgroup-for-7.3-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup:
cgroup/pids: Restore pids.events notifications in local mode
selftests: cgroup: give the O_TMPFILE open in get_temp_fd() a mode
Every week in this release is record-setting for number of posted
patches. It doesn't seem like we're creating any regressions with
all these fixes, 3 Fixes tags here point to 7.2 commits but none
are true regression fixes. We're trying to keep the count down,
nonetheless.
Previous releases - regressions:
- net: don't require the hwtstamp NDOs when a PHY provides timestamping
- ipv6: fix dst leak for uncached routes
- vrf: stop corrupting skb->csum when capturing CHECKSUM_COMPLETE
packets
Previous releases - always broken:
- packet: use ubuf_info completion for TX_RING packets
- arp: terminate device name before lookup
- ipv6: do not let ipv6_find_hdr() return an offset past the packet end
- udp: remove a disconnected socket from the 4-tuple hash table
- sctp: discard the rest of the packet on a stale-cookie error
- eth: mlx5: Bridge, fix remaining switchdev ownership gaps on merged
eswitch
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
-----BEGIN PGP SIGNATURE-----
iQIzBAABCgAdFiEE6jPA+I1ugmIBA4hXMUZtbf5SIrsFAmq1bCQACgkQMUZtbf5S
IrsgCBAAiHOxV3pR0pNw+K/dO5maYhuUV4w+Xcd6CbjRnxzYtImVC1q2ieez5dEe
JrQrVK76t2wuEgGSyiWcckIufPL8xYg5nYaOUbkYMC2QRtDaqhCo+u2jXu26/Lr8
O3sI2cj86XDZN+xKTvmLld1zIRyD6MhYxWz+gIicArdBqTsr5JsWkwhJE7SH013a
iVGu0sIxTgCNu/AusD6SdtwFte26dqvadelI8yuDhBR6rM9pi4EWXChleg7ohCsu
DoZTlaEjEiFavat+6ni99oz8MAT0WVBX0IzTqZa6TCnbCw4qUEbqN8G+CNgC/VPe
Sq5FRoo6vJjlFviwgWPbpje2FoOYtk6CX1eTavx8gZ4pPMTzGkT62crS+ETwLO5H
mRbE+ZjG8UDTUKpIY5P/IQYdznHSc9Ny/pHNlZhFmDdntJOp1lnegx5CmZkyuZxC
Qc0hoFxKdjUrpQ1n0ygT3/P6MT9vXMwbGP52VLIxaB1oaYCKFV71NOp0ET5fCnYb
CBK8cdN0kcaSLrpi3MnPamyQoNZfH78BiuKsLDu/TWg3i2HGmFTlp0/LJ0nkQcCU
4KlCmiJNOkHjbQKHNtYMUOZQcZu5witEv33kb7gr581b0RR+n0lZTaXcnRIyXYdb
ORqIa1Khg5mW3DLaIc9IrzEGGAW5zXUHc5OX7t8J2rVRQWhBEh4=
=YiiK
-----END PGP SIGNATURE-----
Merge tag 'net-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net
Pull networking fixes from Jakub Kicinski:
"Including fixes from Bluetooth, NFC and Netfilter.
Every week in this release is record-setting for number of posted
patches. It doesn't seem like we're creating any regressions with all
these fixes, three 'Fixes' tags here point to 7.2 commits but none are
true regression fixes. We're trying to keep the count down,
nonetheless.
Previous releases - regressions:
- net: don't require the hwtstamp NDOs when a PHY provides
timestamping
- ipv6: fix dst leak for uncached routes
- vrf: stop corrupting skb->csum when capturing CHECKSUM_COMPLETE
packets
Previous releases - always broken:
- packet: use ubuf_info completion for TX_RING packets
- arp: terminate device name before lookup
- ipv6: do not let ipv6_find_hdr() return an offset past the packet
end
- udp: remove a disconnected socket from the 4-tuple hash table
- sctp: discard the rest of the packet on a stale-cookie error
- eth: mlx5: Bridge, fix remaining switchdev ownership gaps on merged
eswitch"
[ And lots of other random network driver fixes ]
* tag 'net-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (189 commits)
tcp: prevent collapsing skbs across boundary in rtx queue
vlan: ensure sufficient headroom in vlan_dev_hard_header()
net/sched: sch_teql: fix shadowed err in __teql_resolve()
bridge: check llc_mac_hdr_init() return value in br_send_bpdu()
llc: fix skb UAF and leaks on llc_mac_hdr_init() failure
llc: reserve device headroom for allocated frames
gve: DQO: reject TSO packets with an out of range MSS
gve: fix TX drop when GSO MSS is too small for hw
gve: DQO: fix header length used by gve_can_send_tso() for UDP GSO
net: flush skb_defer_nodes in dev_cpu_dead()
net: ethernet: stmmac: dwmac-rk: fix bulk clock leak when the PHY clock fails
af_packet: fix integer overflow in prb_calc_retire_blk_tmo()
tipc: Fix a data race on mon->peer_cnt in mon_timeout()
net: phy: intel-xway: workaround 100BASE-TX Link-Up issue
net/smc: fix UAF on lgr list traversal in smcr_port_err()
net/rds: size a connection's path set by the transport it ends up with
nfp: hold IPsec RX state under the XArray lock
net: ena: fix MMIO read buffer leak on probe failure
net: ena: fix PHC cleanup on probe failure
net/sched: act_ct: fix helper UAF due to extensions realloc
...
tcp_write_collapse_fence() sets TCP_SKB_CB(skb)->eor = 1 on
tcp_write_queue_tail(sk) to prevent skbs queued after a switch to
device encryption from being collapsed into earlier skbs.
The fence is a no-op if all earlier data has already been transmitted
when the switch happens: sk->sk_write_queue is empty. The not yet
acknowledged earlier skbs wait in sk->tcp_rtx_queue with eor 0.
On a subsequent retransmit or SACK shift, tcp_retrans_try_collapse() or
tcp_shift_skb_data() can then merge an skb queued after the switch into
one queued before it.
Both users of the fence are affected:
- psp: devices only encrypt skbs with skb->decrypted set. The merged skb
keeps decrypted = 0 from the earlier skb, so merged data sent after
psp_sock_assoc_set_tx() is retransmitted in cleartext.
- tls device offload: the merged skb straddles the start marker set in
tls_set_device_offload(). The software fallback (fill_sg_in() returns
-EINVAL) and the mlx5, nfp and funeth drivers cannot handle such an
skb and drop it. Every retransmit rebuilds the same skb, so the
connection stalls.
Fix this in two places, for defense in depth:
1. Fall back to tcp_rtx_queue_tail(sk) in tcp_write_collapse_fence()
when tcp_write_queue_tail(sk) is NULL.
2. Check !skb_cmp_decrypted(to, from) in tcp_skb_can_collapse(), as
tcp_skb_can_collapse_rx() does on receive. skb_shift(), which both
collapse paths call, already has a DEBUG_NET_WARN_ON_ONCE() for this
condition.
Fixes: e8f6979981 ("net/tls: Add generic NIC offload infrastructure")
Cc: stable@vger.kernel.org
Signed-off-by: Willem de Bruijn <willemb@google.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Daniel Zahka <daniel.zahka@gmail.com>
Link: https://patch.msgid.link/20260924154427.953800-1-willemdebruijn.kernel@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Eric Dumazet says:
====================
vlan: ensure sufficient headroom in vlan_dev_hard_header()
Callers that only reserve ETH_HLEN or less, or skbs allocated before
dynamic device/headroom changes (such as toggling VLAN_FLAG_REORDER_HDR
or bonding/team switching slaves), can reach vlan_dev_hard_header() with
insufficient headroom and trigger skb_under_panic().
When vlan_dev_hard_header() returns -ENOMEM upon skb_cow_head() failure,
a few callers of dev_hard_header() / llc_mac_hdr_init() had pre-existing
error-handling bugs:
====================
Link: https://patch.msgid.link/20260924082951.1599377-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Callers that only reserve ETH_HLEN or less (such as llc_alloc_frame()),
or skbs allocated before dynamic device/headroom changes (e.g. toggling
VLAN_FLAG_REORDER_HDR or bonding/team switching slaves), can reach
vlan_dev_hard_header() with insufficient headroom and trigger
skb_under_panic().
Use skb_cow_head() in vlan_dev_hard_header() when VLAN_FLAG_REORDER_HDR
is not set to ensure sufficient headroom for the VLAN header(s) and the
underlying device hard header.
Use READ_ONCE() to read dev->hard_header_len and dev->needed_headroom as
they can be updated concurrently under RTNL (e.g. in
vlan_transfer_features()) while vlan_dev_hard_header() runs locklessly on
the transmit path. Also avoid LL_RESERVED_SPACE(dev) here so that the
extra HH_DATA_MOD alignment padding does not trigger unnecessary
pskb_expand_head() reallocations on inner stacked VLAN devices after the
outer VLAN header has been pushed.
Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Reported-by: Zixuan Chai <petalzu987@gmail.com>
Closes: https://lore.kernel.org/netdev/cover.1789987105.git.petalzu987@gmail.com/
Link: https://lore.kernel.org/netdev/179022851638.2160803.1808206741379444999@kernel.org/
Cc: Hangbin Liu <liuhangbin@gmail.com>
Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260924082951.1599377-5-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
__teql_resolve() declares an inner 'int err;' inside the
'if (neigh_event_send(n, skb_res) == 0)' block, shadowing the outer
'int err = 0;'. As a result, a negative return from dev_hard_header()
is written to the inner variable and __teql_resolve() still returns 0.
Remove the shadowed variable and set the outer err to -EINVAL when
dev_hard_header() returns a negative error.
Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Closes: https://lore.kernel.org/netdev/179022851638.2160803.1808206741379444999@kernel.org/
Cc: Jamal Hadi Salim <jhs@mojatatu.com>
Cc: Jiri Pirko <jiri@resnulli.us>
Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260924082951.1599377-4-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
In llc_conn_ac_resend_i_xxx_x_set_0_or_send_rr(), if llc_mac_hdr_init()
fails, kfree_skb(skb) is called instead of kfree_skb(nskb). This leaks
the newly allocated nskb, reads from the freed skb via LLC_I_GET_NR(pdu),
and double-frees skb when llc_conn_state_process() drops its reference.
In llc_sap_action_send_xid_r() and llc_sap_action_send_test_r(), nskb is
leaked if llc_mac_hdr_init() returns an error.
Free nskb in all three error paths.
Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Closes: https://lore.kernel.org/netdev/179022851638.2160803.1808206741379444999@kernel.org/
Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260924082951.1599377-2-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
llc_alloc_frame() reserves link-layer headroom using the device type.
This is insufficient for stacked Ethernet devices such as VLAN devices,
where vlan_dev_hard_header() pushes a VLAN header before the lower
device's Ethernet header. An LLC response on such a device can
therefore underflow skb headroom in eth_header().
Use LL_RESERVED_SPACE() to account for the device's actual required
headroom while preserving the existing LLC device-type check.
Fixes: bf9ae5386b ("llc: use dev_hard_header")
Cc: stable@vger.kernel.org
Reported-by: VEGA <vega@nebusec.ai>
Signed-off-by: Zixuan Chai <petalzu987@gmail.com>
Signed-off-by: Ren Wei <weir@nebusec.ai>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260924012613.2533934-1-weir@nebusec.ai
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Eric Dumazet says:
====================
gve: DQO: fix handling of out of range TSO MSS
The DQO TX path assumes that the MSS of a TSO packet is within the
range supported by the device, [88, 9728].
This holds for locally generated traffic, but not for packets coming
from a tap or from a packet socket: virtio_net_hdr_to_skb() takes
gso_size from user space and only enforces a minimum, layer 2
forwarding does not check the MTU of GSO packets, and
gso_features_check() bounds skb->len and gso_segs but never gso_size.
Patch 1, from Eddie Phillips, deals with the lower bound. It moves the
existing test out of gve_prep_tso() into gve_features_check_dqo(), so
that these packets are segmented in software instead of being dropped.
Patch 2 deals with the upper bound, which is currently not checked at
all. gve_tx_fill_tso_ctx_desc() stores gso_size into a 14 bits wide
field, so that an MSS of 16384 silently becomes zero. Falling back to
software segmentation is not an option here, because skb_segment()
splits at gso_size regardless of the MTU, and would only replace an
invalid TSO packet by non TSO packets larger than the 9728 bytes the
device supports. These packets are dropped instead.
As noted in patch 2, oversized non TSO packets can still reach the
device whenever the stack segments in software. This is not specific
to gve and is better fixed in the core, so a patch for
__is_skb_forwardable() will be sent separately for net-next.
====================
Link: https://patch.msgid.link/20260924004252.1196328-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
gve_prep_tso() notes that the device requires the MSS to be <= 9728,
but does not enforce it, assuming the 9K MTU enforced by the hypervisor
and the 64KB limit on TSO sizes are enough.
This does not hold for packets that were not generated locally.
A guest behind a tap, or any packet socket user, can provide an
arbitrary gso_size in virtio_net_hdr. Layer 2 forwarding does not check
the MTU for GSO packets (is_skb_forwardable()), and gso_features_check()
only bounds skb->len and gso_segs, never gso_size.
Such a packet reaches gve_tx_fill_tso_ctx_desc(), which puts gso_size
into the mss field of the TSO context descriptor. This field is 14 bits
wide, so a gso_size of 16384 is silently turned into an MSS of zero.
Drop these packets from gve_prep_tso(), and make sure that
gve_features_check_dqo() leaves their GSO bits alone: skb_segment()
splits at gso_size regardless of the MTU, so falling back to software
segmentation would give the device non TSO packets bigger than the
9728 bytes it supports.
Note that the device can still be given oversized non TSO packets when
the stack segments in software for other reasons, for instance after
TSO has been disabled with ethtool. This is a generic issue, because
the MTU check is skipped for GSO packets in the forwarding path, and
is addressed separately.
Fixes: a57e5de476 ("gve: DQO: Add TX path")
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Harshitha Ramamurthy <hramamurthy@google.com>
Link: https://patch.msgid.link/20260924004252.1196328-3-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The device has a strict requirement that the minimum MSS
(gso_size) for TSO/GSO packets must be at least 88 bytes. If a packet
below this threshold is pushed to the hardware, it can cause
hardware to silently drop the packet, leading to increased latency
and retransmissions.
Currently, this is validated too late in the transmit pipeline
(gve_prep_tso), leading to silent drops.
Fix this by moving the validation into the .ndo_features_check
callback (gve_features_check_dqo). If we detect a GSO packet with
a gso_size smaller than GVE_TX_MIN_TSO_MSS_DQO, we clear the GSO
feature flags for this packet.
Fixes: a57e5de476 ("gve: DQO: Add TX path")
Signed-off-by: Eddie Phillips <eddiephillips@google.com>
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Harshitha Ramamurthy <hramamurthy@google.com>
Link: https://patch.msgid.link/20260924004252.1196328-2-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
-----BEGIN PGP SIGNATURE-----
iIYEABYKAC4WIQSVyBthFV4iTW/VU1/l49DojIL20gUCarVI0xAcbWljQGRpZ2lr
b2QubmV0AAoJEOXj0OiMgvbSEVgA+gNbC9CVCbCo0oufZbVpQwlwtuuEKEVpvxZx
q7oFYKCsAP9svCujGCXRHOmWhAAwe+wpXNb43l8coFn+pCfA8x3fCw==
=NOli
-----END PGP SIGNATURE-----
Merge tag 'landlock-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux
Pull Landlock fixes from Mickaël Salaün:
"This mainly fixes the Landlock tracepoint support merged this cycle so
that denial and rule events report the intended policy context,
whether through tracefs or BTF-visible callbacks.
The size of this all is mainly from propagating the corrected contract
through event definitions and producers, adding new tests for the
reported context, and updating the documentation.
Also improve annotation and fix a GCC 16 build warning"
* tag 'landlock-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux:
landlock: Widen ruleset versions to 64 bits
landlock: Add counted_by in landlock_domain
landlock: Fix tracepoint contract documentation
selftests/landlock: Test network denial context
selftests/landlock: Test filesystem denial blockers
landlock: Report the effective signal number
landlock: Report the actual ptrace tracer
landlock: Fix network denial trace context
landlock: Fix rule tracepoint context
landlock: Fix filesystem denial blocker reporting
landlock: Fix tracepoint fixed-width type names
landlock: Work around gcc-16 -Wuninitialized warning
gve_can_send_tso() computes how many buffers each segment of a GSO
packet would span, and for this it needs the length of the headers
that the device replicates in front of every segment.
It unconditionally uses skb_tcp_all_headers(), which reads the doff
field of the TCP header. SKB_GSO_UDP_L4 packets have no TCP header:
tcp_hdrlen() then reads one byte of the UDP payload, and header_len
can be anything in [0, 60] instead of the transport offset plus the
eight bytes of the UDP header that gve_prep_tso() programs into the
TSO context descriptor.
A wrong header length shifts all the segment boundaries computed in
the loop, so the number of buffers per segment can be over or under
estimated. In the first case, GSO is needlessly disabled for this
packet by gve_features_check_dqo() and the stack has to segment it.
In the second case, the driver hands the device a packet whose
segments span more than GVE_TX_MAX_DATA_DESCS buffers.
Use the UDP header length for SKB_GSO_UDP_L4 packets, matching what
gve_prep_tso() does.
Fixes: 014c607f86 ("gve: add support for UDP GSO for DQO format")
Closes: https://lore.kernel.org/netdev/CANn89i+MS4L60sFQ49=-f-mibeveUfcrpVkD5X+Qy6SOnEpd6w@mail.gmail.com/
Signed-off-by: Eric Dumazet <edumazet@google.com>
Cc: Ankit Garg <nktgrg@google.com>
Cc: Harshitha Ramamurthy <hramamurthy@google.com>
Cc: Joshua Washington <joshwash@google.com>
Cc: Willem de Bruijn <willemb@google.com>
Reviewed-by: Ankit Garg <nktgrg@google.com>
Reviewed-by: Harshitha Ramamurthy <hramamurthy@google.com>
Link: https://patch.msgid.link/20260923145942.731365-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
When a CPU goes offline, dev_cpu_dead() drains its softnet queues
(completion_queue, output_queue, poll_list, process_queue, and
input_pkt_queue), but leaves net_hotdata.skb_defer_nodes untouched.
If oldcpu goes offline while holding pending skbs in its
skb_defer_nodes lists (e.g. below the sysctl_skb_defer_max >> 1 IPI
threshold, or if the IPI races with CPU teardown), those skbs remain
stranded until oldcpu is brought back online. If any of these skbs
hold page_pool fragments, page_pool_destroy() will stall indefinitely
waiting for inflight pages to be returned when a netdev or driver is
torn down while oldcpu is offline.
Additionally, if smp_call_function_single_async() fails in
kick_defer_list_purge() because the target CPU went offline, reset
defer_ipi_scheduled to 0 so future IPI kicks are not blocked when the
CPU comes back online.
Also, if oldcpu was the last online CPU on its NUMA node, drain that
node's slot across all CPUs so no skbs deferred from that node remain
stranded on idle remote CPUs (or if the node itself is subsequently
offlined).
Finally, in skb_attempt_defer_free(), re-check cpu_online(cpu) and
whether the caller migrated CPUs after llist_add(), flushing the node
list if so, to close the preemption TOCTOU race against CPU/node
teardown.
Fixes: 68822bdf76 ("net: generalize skb freeing deferral to per-cpu lists")
Fixes: 5628f3fe3b ("net: add NUMA awareness to skb_attempt_defer_free()")
Closes: https://lore.kernel.org/netdev/20260916003430.3612956-1-kris.pan@intel.com/
Signed-off-by: Eric Dumazet <edumazet@google.com>
Cc: Kris Pan <kris.pan@intel.com>
Link: https://patch.msgid.link/20260923130318.607255-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
gmac_clk_enable() enables the bulk clocks first and then the optional
PHY clock. If clk_prepare_enable() on the PHY clock fails, the function
returns without rolling back the bulk clocks, and bsp_priv->clk_enabled
stays false, so the later gmac_clk_enable(bsp_priv, false) becomes a
no-op and the bulk clock references are leaked.
Add the missing clk_bulk_disable_unprepare() on that failure path.
Fixes: ea449f7fa0 ("net: ethernet: stmmac: dwmac-rk: rework optional clock handling")
Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Reviewed-by: Heiko Stuebner <heiko@sntech.de>
Acked-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Signed-off-by: Coia Prant <coiaprant@gmail.com>
Link: https://patch.msgid.link/20260923123713.3137146-1-coiaprant@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
prb_calc_retire_blk_tmo() computes in 32-bit int arithmetic:
mbits = (blk_size_in_bytes * 8) / (1024 * 1024);
If I'm reading the validation right, tp_block_size is user
controlled and packet_set_ring() only rejects values that are <= 0
as int or not page aligned, so a 256MiB block goes right through
(and alloc_one_pg_vec_page() even has a vzalloc fallback for it).
0x10000000 * 8 wraps to INT_MIN, and on a NIC reporting 1 Gbps
(div == 1) the function ends up returning -2047.
The condition is actually (8 * size) mod 2^32 >= 2^31 && div == 1,
so the trigger set is [256,512), [768,1024), [1280,1536) and
[1792,2048) MiB. Other sizes wrap to non-negative values and faster
links divide the unsigned value back below 2^31, which is why this
doesn't blow up for everyone.
What makes it fatal is what happens next in init_prb_bdqc():
p1->interval_ktime = ms_to_ktime(prb_calc_retire_blk_tmo(...));
hrtimer_start(&p1->retire_blk_timer, p1->interval_ktime,
HRTIMER_MODE_REL_SOFT);
A negative relative timeout expires immediately. The callback
unconditionally returns HRTIMER_RESTART, and hrtimer_forward() turns
the negative interval into hrtimer_resolution:
if (interval < hrtimer_resolution)
interval = hrtimer_resolution;
So the SOFT timer re-fires at the maximum rate forever, holding
sk_receive_queue.lock each pass. One CPU spins in softirq until the
socket is closed. Repeat with more rings and the machine is gone.
The overflow itself is ancient - it was introduced together with
TPACKET_V3 in f6fb8f100b ("af-packet: TPACKET_V3 flexible buffer
implementation."). Its effect prior to f7460d2989 ("net:
af_packet: Use hrtimer to do the retire operation", v6.18) was not
as clear-cut, though: the return value was stored into an unsigned
short retire_blk_tov, so a negative result was truncated, and a
0-jiffy delay loop could be programmed as well. Neither is nearly
as detrimental as the immediate maximum-rate spin the hrtimer
conversion turned it into.
(Unrelated to CVE-2019-20812 - that one was the ethtool failure path
returning 0, which now returns DEFAULT_PRB_RETIRE_TOV.)
Reproducer, needs CAP_NET_RAW (a --network host container has it by
default) and a 1 Gbps NIC (QEMU e1000 works):
int fd = socket(AF_PACKET, SOCK_RAW, htons(ETH_P_ALL));
bind(fd, ...);
int v = TPACKET_V3;
setsockopt(fd, SOL_PACKET, PACKET_VERSION, &v, sizeof(v));
struct tpacket_req3 req = {
.tp_block_size = 0x10000000,
.tp_block_nr = 1,
.tp_frame_size = 2048,
.tp_frame_nr = 0x10000000 / 2048,
.tp_retire_blk_tov = 0,
};
setsockopt(fd, SOL_PACKET, PACKET_RX_RING, &req, sizeof(req));
Compute in 64 bits instead. The operands are already bounded by the
existing validation, so nothing else changes. If you'd prefer a
different fix, just say so and I'll respin.
Fixes: f6fb8f100b ("af-packet: TPACKET_V3 flexible buffer implementation.")
Cc: stable@vger.kernel.org
Signed-off-by: Dairui Zhang <zhangdairui@gmail.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260923050101.1510064-1-zhangdairui@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
mon_timeout() evaluates dom_size(mon->peer_cnt) before it takes mon->lock,
while mon->peer_cnt is updated under that lock by tipc_mon_add_peer() and
tipc_mon_remove_peer(). The value can therefore be stale, and the decision
whether the local domain has to be recomputed can be based on an outdated
member count.
Read mon->peer_cnt inside the write_lock_bh(&mon->lock) protected region.
Fixes: 35c55c9877 ("tipc: add neighbor monitoring framework")
Signed-off-by: Ginger Li <ginger.jzllee@gmail.com>
Reviewed-by: Tung Nguyen <tung.quang.nguyen@est.tech>
Link: https://patch.msgid.link/20260922080909.21123-1-ginger.jzllee@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
MaxLinear GSW12x/GSW14x Ethernet Switch Errata Sheet states:
"An issue has been sporadically observed after device power-on on the first
link-up attempt in 100BASE-TX mode resulting in either the link-up taking a
long time, or failing to link-up altogether...
Workaround:
After power-on, enable Cable Diagnostic Mode for all ports and disable
it..."
Implement the proposed workaround unconditionally in the Intel XWAY driver
(MaxLinear GSW1xx switches incorporate Intel XWAY PHYs) because the
diagnostic bits have the same meaning even in older integral PHYs such as
GPY111/PEF7071/PHY11G. So it's not clear how to distinguish the affected
newer integrated PHYs, but the workaround should not hurt the older PHYs.
Cc: stable@vger.kernel.org
Fixes: 22335939ec ("net: dsa: add driver for MaxLinear GSW1xx switch family")
Signed-off-by: Alexander Sverdlin <alexander.sverdlin@siemens.com>
Reviewed-by: Andrew Lunn <andrew@lunn.ch>
Link: https://patch.msgid.link/20260922075251.23386-1-alexander.sverdlin@siemens.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
smcr_port_err() traverses smc_lgr_list.list without holding
smc_lgr_list.lock, allowing a concurrent smc_lgr_terminate_sched()
to free an lgr while it is still being dereferenced.
Hold smc_lgr_list.lock across the traversal. Update
smc_ib_gid_check() to call smcr_port_err() after releasing the lock.
Fixes: 541afa10c1 ("net/smc: add smcr_port_err() and smcr_link_down() processing")
Reviewed-by: Mahanta Jambigi <mjambigi@linux.ibm.com>
Signed-off-by: Sidraya Jayagond <sidraya@linux.ibm.com>
Reviewed-by: Dust Li <dust.li@linux.alibaba.com>
Link: https://patch.msgid.link/20260922073149.474762-1-sidraya@linux.ibm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
__rds_conn_create() computes npaths from the caller's transport before
it decides whether a connection to one of the host's own addresses is
to be handled by the loopback transport instead. That substitution is
what an RDS/TCP socket sending to a local address gets, and after it
the path init loop still runs for the TCP transport's RDS_MPATH_WORKERS
paths and allocates an ordered workqueue for each, while
rds_loop_conn_alloc() only ever provides transport data for path 0.
rds_conn_destroy() sizes its teardown from c_trans, by then the
loopback transport, so it visits path 0 only - and
rds_conn_path_destroy() would skip the other paths anyway, since it
returns before destroy_workqueue() for a path without transport data.
kfree(c_path) then drops the last pointers to seven workqueues. That
repeats for every such connection, on every netns teardown or module
unload, and every distinct local destination address is a separate
connection.
Recompute npaths once the transport is final, so that creation and
destruction agree on the set of paths. The c_path array stays sized
for the caller's transport; the unused entries are freed with it.
Fixes: 4716af3897 ("net/rds: Give each connection path its own workqueue")
Signed-off-by: Allison Henderson <achender@kernel.org>
Link: https://patch.msgid.link/20260921215027.174657-1-achender@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
nfp_net_ipsec_rx() drops the XArray lock before taking a reference to the
xfrm_state it found. The delete path can erase the entry and drop the last
state reference in that interval. RX can then try to increment a zero
refcount after the state has been queued for destruction.
The driver queues firmware invalidation asynchronously; the delete path
does not wait for the command to complete or drain pending RX processing.
The XFRM garbage collector waits for an RCU grace period before freeing
the state. That delays reclamation but does not make acquiring a reference
from zero valid.
Take the xfrm_state reference before releasing the XArray lock so
xa_erase() cannot run between lookup and reference acquisition.
Fixes: 57f273adbc ("nfp: add framework to support ipsec offloading")
Reported-by: Changyul Lee <lcy8047@gmail.com>
Signed-off-by: Sang-Hoon Choi <csh0052@gmail.com>
Link: https://patch.msgid.link/179001455912.44752.17153022439349797877.idr-bug-92@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Guangshuo Li says:
====================
net: ena: fix resource cleanup on probe failure
This series fixes two resource leaks in the ena_probe() error path after
ena_device_init() has successfully initialized device resources.
Patch 1 adds the missing PHC cleanup.
Patch 2 adds the missing MMIO read request cleanup.
====================
Link: https://patch.msgid.link/20260921154202.471662-1-lgs201920130244@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
ena_device_init() initializes the MMIO read mechanism with
ena_com_mmio_reg_read_request_init(), which allocates a coherent DMA
buffer for MMIO read responses.
The normal removal path releases this buffer through
ena_com_mmio_reg_read_request_destroy(). However, if ena_probe() fails
after ena_device_init() succeeds, the error path destroys the admin
resources and eventually frees ena_dev without destroying the MMIO read
request, leaving the coherent DMA buffer allocated.
Call ena_com_mmio_reg_read_request_destroy() in the probe error path
before releasing the remaining device resources.
This issue was found by manual code inspection.
Fixes: 1738cd3ed3 ("net: ena: Add a driver for Amazon Elastic Network Adapters (ENA)")
Cc: stable@vger.kernel.org
Signed-off-by: Guangshuo Li <lgs201920130244@gmail.com>
Link: https://patch.msgid.link/20260921154202.471662-3-lgs201920130244@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
ena_probe() initializes the PHC as part of ena_device_init(), but the
probe failure path does not destroy it before freeing the PHC private
data.
The normal removal path calls ena_phc_destroy() through
ena_destroy_device() before ena_phc_free(). However, if probe fails
after ena_device_init() succeeds, the error path reaches ena_phc_free()
without unregistering the PTP clock or destroying the device PHC
resources.
Call ena_phc_destroy() in the probe error path before freeing the PHC
private data.
This issue was found by manual code inspection.
Cc: stable@vger.kernel.org tags and describe this as a consistency cleanup
Fixes: e0ea34158e ("net: ena: Add PHC support in the ENA driver")
Cc: stable@vger.kernel.org
Signed-off-by: Guangshuo Li <lgs201920130244@gmail.com>
Cc: stable
Link: https://patch.msgid.link/20260921154202.471662-2-lgs201920130244@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Ilya Maximets says:
====================
ovs, net/sched: fixes for UAF after conntrack extension realloc
One clean up change plus two fixes for the UAF on helper extension
realloc x 2. First half for OVS and the second half for the similar
code in act_ct. This should cover all the known cases of this problem
in these two modules.
====================
Link: https://patch.msgid.link/20260921145655.3167436-1-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
While calling the helpers, a raw pointer to the extensions area is
wired into expectations list:
-> nf_ct_helper()
-> helper->help()
-> nf_ct_expect_related_report()
-> nf_ct_expect_insert()
-> hlist_add_head_rcu(&exp->lnode, &master_help->expectations)
In case the connection is not confirmed yet, more extensions can be
added afterwards with *_ext_add() calls reallocating the extension
space and leaving the now invalid pointer in the expectations list
that is later accessed while removing the expectation.
Make sure that helpers are called at the end after all the other
extensions are already added.
Note that the helper rejection now leaves the mark and labels set,
but that's not different from how the NAT was handled before or how
the mark and the labels were handled on confirmation failure. And
there are no atomicity guarantees provided by the API anyway.
Fixes: a21b06e731 ("net: sched: add helper support in act_ct")
Cc: stable@vger.kernel.org
Reported-by: Axel Mierczuk <axel.mierczuk@1password.com>
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Xin Long <lucien.xin@gmail.com>
Reviewed-by: Jamal Hadi Salim <jhs@mojatatu.com>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Link: https://patch.msgid.link/20260921145655.3167436-7-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
This variable can only become 'true' when the connection is not
confirmed, but it is only checked when it is confirmed. So, it can be
treated as being always false and just removed.
Fixes: a21b06e731 ("net: sched: add helper support in act_ct")
Cc: stable@vger.kernel.org
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Reviewed-by: Xin Long <lucien.xin@gmail.com>
Reviewed-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/20260921145655.3167436-6-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
In a case where skb with an unconfirmed ct entry gets cloned, we may
end up processing both again but with different sets of extensions.
The series of events:
1. The first clone wants to commit and runs the helpers wiring up
the extension pointer into the expectation list.
2. Then it looses the confirmation keeping the entry unconfirmed.
3. Second clone now wants to commit labels or run NAT and adds the
new extension for that breaking the pointer in the expectation
list causing UAF on the destruction path later.
While this is possible to trigger, there should be no practical
network pipeline where we need to process both clones without
modifications in the same zone. So, let's just reset the entry in
case for some reason we got an skb with a shared one. This doesn't
affect any known use cases, but avoids any potential problems with
sharing and modification of the unconfirmed ct entry.
Unlike openvswitch module, act_ct allows for NAT without commit.
Changing that would be a uAPI break. So, act_ct needs to reset on NAT
regardless of the commit flag to avoid reallocation of the extension
space. This, however, doesn't really change the picture for sensible
networking cases as there should be no need to run the same packet
twice (before and after the clone) through conntrack without packet
header or zone changes and without commit.
The fixes tag points to the introduction of helpers, since that's the
main UAF trigger for the sharing.
Fixes: a21b06e731 ("net: sched: add helper support in act_ct")
Cc: stable@vger.kernel.org
Reported-by: Axel Mierczuk <axel.mierczuk@1password.com>
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Reviewed-by: Xin Long <lucien.xin@gmail.com>
Reviewed-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/20260921145655.3167436-5-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
While calling the helpers, a raw pointer to the extensions area is
wired into expectations list:
-> nf_ct_helper()
-> helper->help()
-> nf_ct_expect_related_report()
-> nf_ct_expect_insert()
-> hlist_add_head_rcu(&exp->lnode, &master_help->expectations)
In case the connection is not confirmed yet, more extensions can be
added afterwards with *_ext_add() calls reallocating the extension
space and leaving the now invalid pointer in the expectations list
that is later accessed while removing the expectation.
Make sure that helpers are called at the end after all the other
extensions are already added.
Note that the helper rejection now leaves the mark and labels set,
but that's not different from how the NAT was handled before or how
the mark and the labels were handled on confirmation failure. And
there are no atomicity guarantees provided by the API anyway.
Fixes: cae3a26275 ("openvswitch: Allow attaching helpers to ct action")
Cc: stable@vger.kernel.org
Reported-by: Axel Mierczuk <axel.mierczuk@1password.com>
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Link: https://patch.msgid.link/20260921145655.3167436-4-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
This variable can only become 'true' when the connection is not
confirmed, but it is only checked when it is confirmed. So, it can be
treated as being always false and just removed.
Fixes: 3c1860543f ("openvswitch: add nf_ct_is_confirmed check before assigning the helper")
Cc: stable@vger.kernel.org
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Link: https://patch.msgid.link/20260921145655.3167436-3-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
In a case where skb with an unconfirmed ct entry gets cloned, we may
end up committing both but with different sets of extensions.
The series of events:
1. The first clone wants to commit and runs the helpers wiring up
the extension pointer into the expectation list.
2. Then it looses the confirmation keeping the entry unconfirmed.
3. Second clone now wants to commit labels and adds the new extension
for that breaking the pointer in the expectation list causing
UAF on the destruction path later.
While this is possible to trigger, there should be no practical
network pipeline where committing both clones without modifications
into the same zone is needed. So, let's just reset the entry in case
for some reason we got an skb with a shared one during commit. This
doesn't affect any known use cases, but avoids any potential problems
with sharing and modification of the unconfirmed ct entry.
The fixes tag points to the introduction of helpers, since that's the
main UAF trigger for the sharing.
Fixes: cae3a26275 ("openvswitch: Allow attaching helpers to ct action")
Cc: stable@vger.kernel.org
Reported-by: Axel Mierczuk <axel.mierczuk@1password.com>
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Link: https://patch.msgid.link/20260921145655.3167436-2-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
* add selftest coverage for peer VPN address validation
* reject multicast, broadcast and loopback peer VPN addresses, which
can never identify a peer
* reject MP peers left with no usable VPN address, as they can never
be selected for TX
* reject duplicate peer VPN addresses, which made peer lookup return
an arbitrary peer
* fix stale entry left in the VPN address hashtable when an address
is cleared
* fix torn IPv6 address read on lockless TX when the unusable local
source is cleared in place
* fix torn IPv6 address read on lockless TX when a new local endpoint
is learned in place
* fix dst cache being populated with a route resolved from an already
replaced bind
* fix stale route being reused after the socket mark or UDP source
port changed
* fix bogus validation of an unspecified local source address, which
must instead be left to route source autoselection
* fix IPv6 link-local peer endpoints losing their scope id when
configured via netlink, breaking route lookup
-----BEGIN PGP SIGNATURE-----
iJEEABYIADkWIQQr0db7q+Rc7Zog28Fc8QQzwdnOtwUCarECxxsUgAAAAAAEAA5t
YW51MiwyLjUrMS4xMiwyLDIACgkQXPEEM8HZzrcvTAEA/r3oFnZ4EOtiOyA2OpZ/
Iya+WAJ7LShj7tLymzujMe8BAIH6J2RyrWXnfDFvtSG8HGDEe4EBzVpP56tX0ZVn
xc0N
=fg0G
-----END PGP SIGNATURE-----
Merge tag 'ovpn-net-20260921' of https://github.com/OpenVPN/ovpn-net-next
Antonio Quartulli says:
====================
Included fixes:
* add selftest coverage for peer VPN address validation
* reject multicast, broadcast and loopback peer VPN addresses, which
can never identify a peer
* reject MP peers left with no usable VPN address, as they can never
be selected for TX
* reject duplicate peer VPN addresses, which made peer lookup return
an arbitrary peer
* fix stale entry left in the VPN address hashtable when an address
is cleared
* fix torn IPv6 address read on lockless TX when the unusable local
source is cleared in place
* fix torn IPv6 address read on lockless TX when a new local endpoint
is learned in place
* fix dst cache being populated with a route resolved from an already
replaced bind
* fix stale route being reused after the socket mark or UDP source
port changed
* fix bogus validation of an unspecified local source address, which
must instead be left to route source autoselection
* fix IPv6 link-local peer endpoints losing their scope id when
configured via netlink, breaking route lookup
* tag 'ovpn-net-20260921' of https://github.com/OpenVPN/ovpn-net-next:
selftests: ovpn: validate peer VPN addresses
ovpn: reject invalid peer VPN addresses
ovpn: reject multipeer peers without VPN addresses
ovpn: reject duplicate peer VPN addresses
ovpn: always unhash old VPN addresses before rehashing
ovpn: replace bind when clearing stale local source
ovpn: replace bind when learning local endpoint
ovpn: validate peer state before caching UDP dst
ovpn: track UDP socket route key for peer dst cache
ovpn: skip UDP source validation for unspecified addresses
ovpn: preserve IPv6 scope id for netlink peer endpoints
====================
Link: https://patch.msgid.link/20260921102215.3599702-1-antonio@openvpn.net
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
wx->ptp_tx_skb is shared between the Tx path, the PTP auxiliary
worker and the timestamp cleanup paths. The
WX_STATE_PTP_TX_IN_PROGRESS bit prevents multiple Tx paths from
submitting timestamp requests, but does not serialize the worker
against cleanup.
As a result, wx_ptp_clear_tx_timestamp() can free an skb after
wx_ptp_tx_hwtstamp_work() has obtained its pointer. The worker may
then pass the freed skb to skb_tstamp_tx() and release the same
reference again.
The cleanup path may also clear the in-progress bit while the worker
is still processing the old skb. This allows the Tx path to publish a
new skb which the worker can subsequently overwrite with NULL,
leaking its reference.
Add a dedicated spinlock to protect publication and consumption of
the Tx timestamp skb. Detach the skb and clear the in-progress bit
while holding the lock, then deliver the timestamp and release the skb
after dropping it. Use the same locked cleanup in the quiesce path,
but keep the detach there free of register accesses: quiesce runs
during PCIe error recovery, where MMIO is not reliable, and it
deliberately did not touch the device before. The lock is taken with
interrupts disabled, because netpoll can call ndo_start_xmit() with
hard interrupts already off.
When handling a Tx DMA mapping failure, keep the transmit path
reference until after comparing the skb under the lock. This prevents
skb address reuse from making the error path mistake a newer timestamp
request for the failed one.
Fixes: 06e75161b9 ("net: wangxun: Add support for PTP clock")
Reported-by: Sashiko <sashiko-bot@kernel.org>
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/6C7EC12D69217315%2B20260818074721.45536-1-jiawenwu%40trustnetic.com
Signed-off-by: Jiawen Wu <jiawenwu@trustnetic.com>
Link: https://patch.msgid.link/77431AF9A0E369F3+20260921071549.1141804-1-jiawenwu@trustnetic.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
phylink_bringup_phy() stores the PHY in pl->phydev before its last
fallible step: on a MAC whose phylink ops implement LPI,
phy_eee_rx_clock_stop() can fail with a real MDIO error. The callers
unwind with phy_detach(), which knows nothing about pl->phydev, so a
pointer to a PHY that is no longer attached outlives the failed
connect.
What that costs depends on how the caller got here.
phylink_connect_phy() goes through phylink_attach_phy(), which refuses
to attach while pl->phydev is set, turning a transient MDIO error into
a permanent -EBUSY. The SFP path is worse than that: sfp_sm_probe_phy()
answers the failure with phy_device_remove() and phy_device_free(), and
it assigns sfp->mod_phy only past that error return, so nothing clears
pl->phydev and it is left pointing at a freed phy_device that
phylink_resolve() and the ethtool helpers go on reading.
phylink_fwnode_phy_connect() has no such check, so a later connect
overwrites the stale pointer and hides the problem. A disconnect does
not: phylink_disconnect_phy() hands that pointer to phy_disconnect(),
and the second phy_detach() on the same PHY drops references the first
one already released.
Found while making a DSA port survive a PHY whose driver arrives after
the switch probes: keeping the port across a failed connect and
retrying is what makes this window reachable.
Publish the pointer after the last call that can fail instead of
unwinding it afterwards. Nothing between the two points reads
pl->phydev, and the registration that follows cannot fail:
phy_request_interrupt() falls back to polling on its own. The PHY-side
state keeps the order it had, so no MDIO operation moves relative to
another.
Fixes: 03abf2a7c6 ("net: phylink: add EEE management")
Signed-off-by: Aleksei Sviridkin <f@lex.la>
Link: https://patch.msgid.link/20260920222044.1752860-1-f@lex.la
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The MeiG Smart SRM821 5G module (0x2dee:0x4d53) crashes and drops off
the USB bus when it receives a Zero Length Packet (ZLP) after sending
or receiving an NTB of exactly 16384 bytes (tx_max).
According to the MBIM specification, devices do not require a ZLP
if the NTB size is exactly dwNtbOutMaxSize. However, the cdc_mbim
driver defaults to sending ZLPs for devices not explicitly whitelisted
to accommodate non-conformant hardware. This default behavior breaks
the strictly conformant MeiG SRM821 module.
Add this device to the ZLP conformance whitelist (cdc_mbim_info) so
the driver will pad the NTB to avoid sending ZLPs, preventing the
device firmware from crashing.
Cc: stable@vger.kernel.org
Signed-off-by: Ming Wang <wangming01@loongson.cn>
Link: https://patch.msgid.link/20260920074500.826121-1-wangming01@loongson.cn
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
A fork rejected by the pids controller increments the counter reported by
pids.events. When local event accounting is selected, however, pids_event()
returns after notifying only events_local_file, leaving pids.events pollers
asleep.
On legacy hierarchies, pids.events.local does not exist. With
pids_localevents, pids.events reports the same local counter. In both
cases, pids.events changes without generating a notification.
This can be reproduced with a pids_localevents mount:
mkdir /tmp/test
mount -t cgroup2 -o pids_localevents none /tmp/test
mkdir /tmp/test/t
echo 1 > /tmp/test/t/pids.max
cat /tmp/test/t/pids.events # max 0
timeout 3 inotifywait -e modify /tmp/test/t/pids.events &
sh -c 'echo $$ > /tmp/test/t/cgroup.procs; (true &)' 2>/dev/null
wait
cat /tmp/test/t/pids.events # max 1
Without this patch, inotifywait times out without reporting an event.
Notify pids.events before returning from the local event path.
Fixes: 3f26a885a0 ("cgroup/pids: Add pids.events.local")
Cc: stable@vger.kernel.org # v6.11+
Signed-off-by: Guopeng Zhang <zhangguopeng@kylinos.cn>
Signed-off-by: Tejun Heo <tj@kernel.org>