Aleksei Sviridkin says:
====================
net: dsa: mt7530: fix two crashes on driver unbind
Unbinding the MT7530 driver from an MT7531 dereferences NULL in
regulator_disable(). On a Netcraze NC-1012 (MT7981B + MT7531, 6.18.44):
# echo mdio-bus:1f > /sys/bus/mdio_bus/drivers/mt7530-mdio/unbind
oopses there, and the build it was found on sets CONFIG_PANIC_ON_OOPS, so
the board goes down with it. Fix that and the same command gets as far as
mt7530_remove_common(), which disposes interrupt mappings the switch's own
regmap-irq chip still owns; the regmap-irq thread then faults in
handle_nested_irq() later in the same teardown. rmmod reaches both, since
mdio_module_driver() calls .remove on module exit.
Patch 1 is the regulator one. mt7530_probe() requests the core and io
supplies only for ID_MT7530 and mt7530_setup() enables them under the same
test, but mt7530_remove() disables them unconditionally, so on an MT7621 or
an MT7531 both pointers are still NULL from devm_kzalloc(). It reaches the
MDIO front end only.
Patch 2 is the interrupt one, and it reaches further.
mt7530_remove_common() disposes the per-PHY interrupt mappings by hand from
.remove, while the regmap-irq chip that owns the domain is devm-registered
and its parent interrupt is only freed once .remove has returned.
regmap_del_irq_chip() disposes the same mappings itself, in an order that
cannot race, so the driver's call adds nothing but a window. That helper is
called from both front ends, so the defect also covers the MMIO parts -
MT7988, EN7581, AN7583 and EN7528 - which have no regulators and never meet
the first defect at all.
The order is not arbitrary. On an MT7531 the regulator fault happens in the
first thing mt7530_remove() does with the switch, so execution never
reaches the interrupt defect. The second only became visible once the first
was fixed, which is also how both came to be found on one board.
Found and verified there. Without patch 1 the unbind panics in
regulator_disable(); with patch 1 alone the panic moves on to
handle_nested_irq(); with both, three unbind/bind cycles run, two back to
back and a third after a pause. In the two whose dmesg was captured, each
unbind removes the switch from the driver directory and takes lan1 to lan4
with it, each bind brings them back, and lan1 relinks at 1Gbps/full after
both binds, lan4 after the second. uptime rose from 58 to 202 seconds
across the three without resetting and pstore gained no new record. The
third cycle stayed unbound long enough to read the descriptors: no mt7530
line in /proc/interrupts and no irq/79, irq/80 or irq/81 directory, and the
next bind reuses those three numbers - regmap-irq freeing and disposing
what the driver no longer touches. The kernel under test was identified by
the sha256 of its ELF notes section, read from /sys/kernel/notes on the
running board and computed in advance from the flashed image.
What hardware could not answer here. There is no MT7530 or MT7621 part on
this bench, so the ID_MT7530 branch that patch 1 adds was checked by
reading the generated code rather than by running it, and no MMIO part was
available to exercise patch 2 on that front end either. One unrelated WARN
remains across the unbind, from sysfs_remove_link() under
dsa_user_destroy(); it is a separate DSA teardown-ordering defect and is
not addressed here.
v1: https://lore.kernel.org/20260914202421.2737079-1-f@lex.la
====================
Link: https://patch.msgid.link/20260918015020.2518315-1-f@lex.la
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
mt7530_remove_common() disposes the per-PHY interrupt mappings from
.remove, but the regmap-irq chip that owns the domain is devm-registered,
so its parent interrupt is only freed once .remove has returned. The
switch's own regmap-irq thread can therefore still dispatch on a mapping
that is already gone: irq_find_mapping() returns 0, irq_to_desc() returns
NULL and handle_nested_irq() locks desc->lock without checking it. The
attached PHYs have not given those interrupts back yet either, which the
kernel warns about a moment before the fault.
regmap_del_irq_chip() disposes the same mappings itself, after freeing the
parent interrupt and before removing the domain, so there is nothing left
for the driver to do here. Until it runs the descriptors stay alive, and a
late dispatch on one of them is harmless: dsa_unregister_switch() has freed
the PHY handlers by then, so handle_nested_irq() finds no action and
returns.
Fixes: 254f6b272e ("dsa: mt7530: Utilize REGMAP_IRQ for interrupt handling")
Signed-off-by: Aleksei Sviridkin <f@lex.la>
Link: https://patch.msgid.link/20260918015020.2518315-3-f@lex.la
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The core and io supplies are only requested for ID_MT7530: both the
devm_regulator_get() in probe and the regulator_enable() in
mt7530_setup() are guarded by the switch id, but mt7530_remove()
disables them unconditionally. On an MT7621 or an MT7531 both pointers
are still NULL from devm_kzalloc(), so rmmod or a sysfs unbind calls
regulator_disable() on NULL.
Fixes: ddda1ac116 ("net: dsa: mt7530: support the 7530 switch on the Mediatek MT7621 SoC")
Signed-off-by: Aleksei Sviridkin <f@lex.la>
Link: https://patch.msgid.link/20260918015020.2518315-2-f@lex.la
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
fixes the following crash on driver initialization:
|BUG: Kernel NULL pointer dereference on read at 0x00000158
|Faulting instruction address: 0xc0566b40
|Oops: Kernel access of bad area, sig: 11 [#1]
|BE PAGE_SIZE=4K PowerPC 44x Platform
|Modules linked in:
|CPU: 0 UID: 0 PID: 1 Comm: swapper/0 Tainted: GW 7.3.0-rc3+ #1
|Tainted: [W]=WARN
|Hardware name: MyBook Live APM821XX 0x12c41c83 PowerPC 44x Platform
|NIP: c0566b40 LR: c05648f8 CTR: c04c1e9c
|REGS: c1053a20 TRAP: 0300 Tainted: GW (7.3.0-rc3+)
|MSR: 0002b000 <CE,EE,FP,ME> CR: 24008808 XER: 00000000
|DEAR: 00000158 ESR: 00000000
|GPR00: c05648f8 c1053b10 c1063600 c1030000 c5ab3000 00000000 [...]
|GPR08: 00000002 00000000 00000000 c1053b40 84002808 00000000 [...]
|GPR16: cfffd210 00000002 c0beafcc cfffc960 00000000 c1030644 [...]
|GPR24: c0beafbc c1037000 00000000 0000000a 00000000 c1030000 [...]
|NIP [c0566b40] phy_link_topo_add_phy+0x2c/0x1d0
|LR [c05648f8] phy_attach_direct+0x1a4/0x368
|Call Trace:
|[c1053b10] [c0811e04] klist_put+0x54/0xb4 (unreliable)
|[c1053b40] [c05648f8] phy_attach_direct+0x1a4/0x368
|[c1053b70] [c0564ae8] phy_connect_direct+0x2c/0x60
|[c1053b90] [c056c7a8] of_phy_connect+0x50/0x74
|[c1053bc0] [c0572a88] emac_probe+0xd50/0x119c
|[c1053c90] [c04cb770] platform_probe+0x74/0xa4
|[c1053cb0] [c04c8f08] really_probe+0x120/0x2b0
|[c1053cd0] [c04c9254] __driver_probe_device+0x1bc/0x1fc
|[c1053d00] [c04c9334] driver_probe_device+0x38/0xa8
|[c1053d30] [c04c9560] __driver_attach+0xf4/0x10c
|[c1053d50] [c04c6af8] bus_for_each_dev+0x68/0xd0
|[c1053d90] [c04c7c34] bus_add_driver+0xcc/0x1ec
|[c1053dc0] [c04c9f6c] driver_register+0xcc/0x110
|[c1053de0] [c0aa3f40] emac_init+0x1c4/0x200
This bug showed up starting with v7.3-rc1. At the NIP in
phy_link_topo_add_phy() is a netdev_need_ops_lock() check.
This was added by the following
commit ded86da4bb ("net: ethtool: relax ethnl_req_get_phydev() locking assertion")
The bug shows up because at the time of_phy_connect() was called, the
netdev_ops were *not yet* determined. My fix is to move the code that
sets netdev_ops+commac.ops+ethtool_ops further up as emac_init_config()
derives that by looking at the device-tree and sets the required
dev->phy_mode accordingly.
During review, the Sashiko bot's AI stated that the commac assignment
became a dead store. Great catch! To keep the original behavior as-is,
one mentioned option "set dev->commac.ops = &emac_commac_ops only in
the non-gige case?" sounded like a great plan. So the dev->commac.ops
assignment for the non-gige-case moves into the else block.
This patch was tested on a WD MyBook Live (RGMII). the device now works again.
Fixes: ded86da4bb ("net: ethtool: relax ethnl_req_get_phydev() locking assertion")
Signed-off-by: Christian Lamparter <chunkeey@gmail.com>
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/49cd7343bc0e507c022071f3e2b5662b053dca73.1790007431.git.chunkeey@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
skbmark_decode(), skbprio_decode() and skbtcindex_decode() read fixed-size
values from the TLV payload without validating its length.
A malformed IFE frame can declare a shorter payload, causing the decoders
to consume bytes beyond the declared metadata value:
[TLV type=IFE_META_SKBMARK len=4]
-> dlen == 0, but decode reads 4 bytes
The decoder may therefore set skb metadata from unintended input.
Validate the payload length before decoding and return -EINVAL for
invalid lengths. Read the values with get_unaligned_be32() and
get_unaligned_be16(), as TLV payloads are not guaranteed to be
aligned. Teach tcf_ife_decode() to log a decoder error separately
from an unknown metaid; both are counted as overlimits and decoding
continues with the remaining metadata.
The metadata length issue was found by an automated audit of the IFE
decode path at v6.18-rc7 and reproduced with a userspace sanitizer
model of the decode path. Compile-tested on x86_64 with defconfig and
NET_ACT_IFE=y: act_ife.o and the three act_meta_*.o build
warning-free.
Fixes: 084e2f6566 ("Support to encoding decoding skb mark on IFE action")
Fixes: 200e10f469 ("Support to encoding decoding skb prio on IFE action")
Fixes: 408fbc22ef ("net sched ife action: Introduce skb tcindex metadata encap decap")
Assisted-by: Hawkeye:GLM-5.3-flash
Assisted-by: Qoder:Qwen3.8-Max
Signed-off-by: Fang Xieyan <fangxy@xiaopeng.com>
Link: https://patch.msgid.link/20260921125441.81459-1-fangxy@xiaopeng.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Creating a MACsec device with MAC offload over an LRO-capable lower
device triggers a warning in rtmsg_ifinfo_build_skb() when IPv4
forwarding is enabled by default.
register_netdevice() invokes inetdev_init(), which disables LRO and emits
a NETDEV_FEAT_CHANGE notification. This reaches macsec_fill_info() before
macsec_add_dev() initializes the SecY. key_len is still zero, so
macsec_fill_info() returns -EMSGSIZE and trips the WARN_ON in
rtmsg_ifinfo_build_skb(), even though the skb has enough space.
Even without the warning, notifications during registration can report
uninitialized SecY attributes, including the SCI. This ordering has existed
since the driver was introduced.
Initialize the SecY and apply the new-link attributes before registration.
Move MAC address inheritance into macsec_newlink() so the SCI can also be
initialized before registration-time notifications report it. Move the
per-CPU statistics and metadata destination allocation into ndo_init(),
and release partial allocations on failure.
Fixes: c09440f7dc ("macsec: introduce IEEE 802.1AE driver")
Reported-by: syzbot+f2f6312ad1b5a0bfe316@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=f2f6312ad1b5a0bfe316
Suggested-by: Sabrina Dubroca <sd@queasysnail.net>
Link: https://lists.openwall.net/linux-kernel/2026/08/19/552
Signed-off-by: Haseeb Malik <haseebulhaq55@gmail.com>
Reviewed-by: Sabrina Dubroca <sd@queasysnail.net>
Link: https://patch.msgid.link/20260921-fix-macsec-net-v3-1-accf94f93f5e@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
rds_ib_map_frmr() stores the caller's scatterlist in the MR before DMA
mapping and registration can fail. On failure, __rds_rdma_map() unpins
the pages and frees the scatterlist, but rds_ib_free_frmr() can still
return the MR to the pool with the stale pointer set.
This leaves the pool with a dangling scatterlist and can lead to local
privilege escalation. KASAN detects the resulting use-after-free when the
MR is later torn down:
BUG: KASAN: slab-use-after-free in __rds_ib_teardown_mr
Read of size 8
Call Trace:
__rds_ib_teardown_mr
rds_ib_unreg_frmr
rds_ib_flush_mr_pool
rds_ib_flush_mrs
rds_free_mr
rds_setsockopt
Store the scatterlist in the MR only after DMA mapping succeeds. If DMA
mapping fails, return directly while the MR fields remain clear; the caller
keeps ownership of the scatterlist and its pinned pages. If a later
registration step fails, unmap the scatterlist and clear the MR fields
before returning.
Fixes: 1659185fb4 ("RDS: IB: Support Fastreg MR (FRMR) memory registration mode")
Cc: stable@vger.kernel.org
Signed-off-by: Dongliang Qin <cccccccccccc777777@gmail.com>
Reviewed-by: Allison Henderson <achender@kernel.org>
Link: https://patch.msgid.link/20260922031546.3874605-1-cccccccccccc777777@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
bna: prevent IOC timer rearm during teardown
bnad_pci_remove() and the probe disable_ioceth path call
timer_delete_sync() for ioc_timer, sem_timer and hb_timer, but not for
iocpf_timer. bnad_iocpf_timeout() then takes bnad->bna_lock after
free_netdev() has freed the struct bnad.
Deleting iocpf_timer last does not fix this. sem_timer and
iocpf_timer rearm each other: bnad_iocpf_sem_timeout() can arm
iocpf_timer, and bnad_iocpf_timeout() arms sem_timer from
bfa_ioc_hw_sem_get() when the semaphore is busy.
timer_delete_sync() only waits out its own callback.
bnad_ioceth_disable() can time out and leave that callback live.
Shut all four IOC timers down with timer_shutdown_sync() on both
paths, so a later mod_timer() is ignored.
This issue was identified during our ongoing static-analysis research
while reviewing kernel code.
Fixes: 1d32f76962 ("bna: IOC failure auto recovery fix")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Co-developed-by: Ijae Kim <ae878000@gmail.com>
Signed-off-by: Ijae Kim <ae878000@gmail.com>
Signed-off-by: Myeonghun Pak <mhun512@gmail.com>
Link: https://patch.msgid.link/20260922014605.588040-1-mhun512@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
says:
====================
net: atl1c/atl1e/atl1: fix soft lockup on out-of-range tx consumer index
atl1c_clean_tx() reads a hardware-maintained tx consumer index and
walks a software index towards it:
while (next_to_clean != hw_next_to_clean) {
...
if (++next_to_clean == tpd_ring->count)
next_to_clean = 0;
}
next_to_clean only ever takes values in [0, tpd_ring->count). If the
hardware read returns a value outside that range - seen as 0xffff
while the PCIe link/MAC is resetting, e.g. during a neighboring
device's reboot - the loop condition can never become false, and the
NAPI thread spins forever. This produced a real soft lockup on
current hardware:
watchdog: BUG: soft lockup - CPU#12 stuck for 354s! [napi/eth%d-0]
RIP: 0010:atl1c_clean_tx+0x142/0x2d0 [atl1c]
Patch 1 fixes this in atl1c. atl1e and atl1 (atlx) share the exact
same loop shape, reading their own hardware-maintained consumer index
with no bounds check either, and are just as reachable from the same
kind of PCIe link event. Patches 2 and 3 apply the same guard to each.
====================
Link: https://patch.msgid.link/20260921091334.3571525-1-tamas@rimpianto.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Same issue as atl1c (see the first commit in this series, "net:
atl1c: fix soft lockup on out-of-range tpd_cons read"): the hardware
can report an out-of-range cmb_tpd_next_to_clean (seen as 0xffff)
while the PCIe link/MAC is resetting. An out-of-range value can
never be reached and the loop below would spin forever. Treat it as
"nothing new to clean" instead.
Fixes: f3cc28c797 ("Add Attansic L1 ethernet driver.")
Cc: stable@vger.kernel.org
Signed-off-by: Gajdos Tamás <tamas@rimpianto.com>
Link: https://patch.msgid.link/20260921091334.3571525-4-tamas@rimpianto.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Same issue as atl1c (see the first commit in this series, "net:
atl1c: fix soft lockup on out-of-range tpd_cons read"): the hardware
can report an out-of-range hw_next_to_clean (seen as 0xffff) while
the PCIe link/MAC is resetting. An out-of-range value can never be
reached and the loop below would spin forever. Treat it as "nothing
new to clean" instead.
Fixes: a6a5325239 ("atl1e: Atheros L1E Gigabit Ethernet driver")
Cc: stable@vger.kernel.org
Signed-off-by: Gajdos Tamás <tamas@rimpianto.com>
Link: https://patch.msgid.link/20260921091334.3571525-3-tamas@rimpianto.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
The hardware can report an out-of-range tpd_cons (seen as 0xffff)
while the PCIe link/MAC is resetting. An out-of-range value can
never be reached and the loop below would spin forever. To avoid
a soft lockup treat it as "nothing new to clean" instead.
Reproduced on two machines, same NIC (Qualcomm Atheros AR8151 v2.0,
4-port), triggered by rebooting a Mikrotik CCR2004 PCIe card that
the ports are directly linked to:
- Ubuntu 26.04.1 LTS, kernel 7.0.0-31-generic. The link-flap
precursor, before the lockup was captured with a full trace
elsewhere:
atl1c 0000:05:00.0 enp5s0f0: NETDEV WATCHDOG: CPU: 4: transmit queue 2 timed out 489984 ms
atl1c 0000:05:00.0: MAC state machine can't be idle since disabled for 10ms second
atl1c 0000:05:00.0: atl1c: enp5s0f0 NIC Link is Up<65535 Mbps Full Duplex>
65535 (0xffff) here is the same value tpd_cons reads back once the
loop below gets stuck.
- Proxmox VE, kernel 7.0.14-11-pve. Same NIC/trigger, this time
caught by the soft lockup watchdog with a full stack trace:
watchdog: BUG: soft lockup - CPU#12 stuck for 354s! [napi/eth%d-0:329]
CPU: 12 UID: 0 PID: 329 Comm: napi/eth%d-0 Tainted: P O L 7.0.14-11-pve #1 PREEMPT(lazy)
RIP: 0010:atl1c_clean_tx+0x142/0x2d0 [atl1c]
Call Trace:
<TASK>
__napi_poll+0x32/0x1e0
napi_threaded_poll_loop+0x286/0x2e0
napi_threaded_poll+0xfd/0x140
kthread+0xf7/0x130
ret_from_fork+0x2da/0x3a0
ret_from_fork_asm+0x1a/0x30
</TASK>
Fixes: 43250ddd75 ("atl1c: Atheros L1C Gigabit Ethernet driver")
Cc: stable@vger.kernel.org
Signed-off-by: Gajdos Tamás <tamas@rimpianto.com>
Link: https://patch.msgid.link/20260921091334.3571525-2-tamas@rimpianto.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Create a bonding device (i.e. bond0) in active-backup mode, 2 slaves.
Active slave: offload capable interface (i.e. eth1), primary interface.
Backup slave: non-offload capable interface(i.e. eth2).
Configure strongswan service swantl.conf child SA "hw_offload = crypto"
Start strongswan service
IPSec Crytpo Offload is enabled on top of bond0. i.e.
ip xfrm state |grep offload
crypto offload parameters: dev bond0 dir out mode crypto
crypto offload parameters: dev bond0 dir in mode crypto
Active slave eth1 takes adavantage of IPSec Crypto Offload capability.
If active slave eth1 is down for any reason (i.e. eth1 link down):
ip link set down dev eth1
non-offload capable interface eth2 failover to becomes active slave.
The existing SAs can continue use software IPsec after failover.
Traffic still keeps going properly.
However if eth1 link had not recovered yet, strongswan service does
new child SA rekey, or uses swanctl command to do new child SA rekey,
it will fail because active slave eth2 doesn't support crypto offload.
In bond_ipsec_add_sa routine, it returns -EINVAL now, which is
treated as fatal error by xfrm_dev_state_add routine in kernel xfrm.
To make the non-offload active slave survive the child SA rekey, need
to make bond_ipsec_add_sa routine returns -EOPNOTSUPP instead when
active slave doesn't support IPsec Crypto offload, the xfrm will
gracefully fallback to create new SA using Software IPsec.
Network traffic can keep going.
After offload capable interface eth1 link is up, becomes active slave,
next time strongswan child SA rekey will create a new SA which enables
crypto offload again.
Fixes: 18cb261afd ("bonding: support hardware encryption offload to slaves")
Signed-off-by: David Dai <zdai@linux.ibm.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Link: https://patch.msgid.link/20260918211155.1664493-1-zdai@linux.ibm.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
ic_dhcp_init_options() appends the hostname (option 12), vendor-class
(option 60) and client-ID (option 61) options into the fixed 312-byte
bootp_pkt.exten[] buffer. Only the client-ID branch checked the
remaining space; the hostname and vendor-class writes were unbounded.
A 64-byte hostname together with the maximum 252-byte dhcpclass=
identifier needs 18 + (2 + 64) + (2 + 252) = 338 of the 312 available
bytes even before the terminating END marker, so the vendor-class memcpy
runs past the end of exten[]. With CONFIG_FORTIFY_SOURCE this is
reported as a field-spanning write and, when the kernel is booted with
panic_on_warn=1, aborts boot with a panic.
Route the optional options through a common helper that makes sure the
option, its 2-byte header and the END marker all fit and drops an option
that would not. Configurations with short options keep sending exactly
the same bytes as before.
Fixes: 130c0f47fd ("ipconfig: send host-name in DHCP requests")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Assisted-by: LLM
Signed-off-by: Yuqi Xu <xuyuqiabc@gmail.com>
Reviewed-by: Ren Wei <weir@nebusec.ai>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/7808dfbfa2162dfd0b19f59aff5742d6e0db2abb.1789798023.git.xuyuqiabc@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Mostly security fixes.
fix use-after-free in nfc_get_local_general_bytes
llcp: Fix list corruption / refcount desync in nfc_llcp_recv_dm()
llcp: Fix race condition in accept_queue lifecycle
llcp: drop truncated I/RR/RNR PDUs in nfc_llcp_recv_hdlc()
llcp: fix -ENOMEM on connect with zero-length service name
llcp: fix WKS SAP hijacking via prefix match in nfc_llcp_wks_sap()
llcp: fix sdreq TLV list leak on parse/alloc/send failure
llcp: fix slab-out-of-bounds reads when logging service names
selftests/nci: Fix out-of-bounds store on thread join
selftests: nci: Correct pthread_create return value check
selftests: nci: Fix uninitialized family ID on missing attribute
virtual_ncidev: Add missing ioctl compat handler
nfcmrvl: validate helper command length before pull
pn533: fix OOB read in pn533_acr122_is_rx_frame_valid()
port100: reject frames whose declared length exceeds the received data
st21nfca: validate ISO15693 inventory length
st21nfca: validate received frame size
trf7970a: power down on startup RX gain failure
Signed-off-by: David Heidelberg <david@ixit.cz>
-----BEGIN PGP SIGNATURE-----
iQIzBAABCAAdFiEE13oJz+7cK71TpwR0YAI/xNNJIHIFAmq0M08ACgkQYAI/xNNJ
IHKXbw//cxnjUWiGiCWQKNGIrPQhO8xnxp29/adTfr5cpaCf9XS1Ah742JnsvV0J
jiKwrV2VdhIedF0gCU51pfMdgV4VWWup2KQJbl820m8EaRmaKY6Spao/HxznnXtF
H1JqYir2ISVADjX4hgQpqQIH2Vrz2o2QAkcsY+jjjydxxajmrrQh1SrJeEGom366
eDHAPmO9MTheJZ8gFXtuz4JdoipVWwY0Ta8D3mmVfgX/Jv5ESsPa4Ie4gQvQ1Vua
9bvYSgqrCGi4jtzEOxXNL1KfufjsbW9I64anwhxwjKJ+sL+/GJ0VfjN4Dgq9dnub
H4dUJs9vUCMLY+Ns6Wo+2pJQ49Ults2SdVjH7kvdJ2FiUn+FvoPDLrQY00P/lhaB
MPLrOCFaNVIzDKX98lxAzT/6c6xrNb5lTrkf17XHzcWa2H4eZj30pBaFfYrfiEH5
dlASR5aSy+GuXa8J/i7+jpQ+xl0GEg9tb0GJsvsWTDAftAl8/9CvcZ/TCbjKeb9L
oGeStPcl/i+MB9wGZLOga8lox0DX5PMNkIcp+RXIjVvPDYiCBh1a3lWHtz65Qx4m
9VlNcIpUzN14VUTkWm1i29DVHpW8V5c/eLmoDHzBUFAVSVr2M0CkFwhPYOByqIzt
lIcpLuTLXrAplzIFWMbi6gU8FSRAFX2I+vVr6rkgWIz2+QCPmYM=
=lO2R
-----END PGP SIGNATURE-----
Merge tag 'nfc-7.3-rc5' of https://codeberg.org/linux-nfc/linux
David Heidelberg says:
====================
NFC fixes for net 7.3-rc5
* tag 'nfc-7.3-rc5' of https://codeberg.org/linux-nfc/linux:
nfc: pn533: fix OOB read in pn533_acr122_is_rx_frame_valid()
nfc: llcp: fix slab-out-of-bounds reads when logging service names
nfc: llcp: fix WKS SAP hijacking via prefix match in nfc_llcp_wks_sap()
nfc: llcp: fix -ENOMEM on connect with zero-length service name
nfc: st21nfca: validate ISO15693 inventory length
nfc: fix use-after-free in nfc_get_local_general_bytes
nfc: trf7970a: power down on startup RX gain failure
nfc: llcp: fix sdreq TLV list leak on parse/alloc/send failure
nfc: llcp: drop truncated I/RR/RNR PDUs in nfc_llcp_recv_hdlc()
nfc: virtual_ncidev: Add missing ioctl compat handler
selftests/nci: Fix out-of-bounds store on thread join
selftests: nci: Fix uninitialized family ID on missing attribute
nfc: llcp: Fix race condition in accept_queue lifecycle
selftests: nci: Correct pthread_create return value check
nfc: port100: reject frames whose declared length exceeds the received data
nfc: llcp: Fix list corruption / refcount desync in nfc_llcp_recv_dm()
nfc: st21nfca: validate received frame size
nfc: nfcmrvl: validate helper command length before pull
====================
Link: https://patch.msgid.link/adeaccc1-cc04-4bb9-a28a-61a61d75ba14@ixit.cz
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Only the entries below dev->num_tc are valid in dev->tc_to_txq[], and
dev->prio_tc_map[] may only name classes below it. netdev_set_num_tc()
lowers dev->num_tc without touching either array.
netdev_txq_to_tc() walks all TC_MAX_QUEUE slots and
netdev_get_prio_tc_map() returns the entry as it stands, so a leftover
entry is handed out as a traffic class >= dev->num_tc. Taking that
class from netdev_txq_to_tc(), __netif_set_xps_queue() rejects only a
negative one and indexes an XPS map sized for dev->num_tc classes:
tci = j * num_tc + tc;
RCU_INIT_POINTER(new_dev_maps->attr_map[tci], map);
attr_map[] holds nr_ids * num_tc entries and j runs over the ids named
in the mask, so a class that is not below num_tc pushes tci past the end
of the map for the last ids and the store overruns it.
Any caller that lowers num_tc leaves such entries behind, and
mqprio_destroy() tears down with netdev_set_num_tc(dev, 0) rather than
netdev_reset_tc(). After mqprio with 8 classes then 1, tc_to_txq[1..7]
still describe txq 1..7. The splat is from an XPS write to txq 2 on a
veth with 8 rx queues: attr_map[] has 8 * 1 entries, tci = j + 2, and
j == 6 stores one past the end of the 88-byte map:
BUG: KASAN: slab-out-of-bounds in __netif_set_xps_queue (net/core/dev.c:2954)
Write of size 8 at addr ffff88813016bc58 by task xps_oob/634
__netif_set_xps_queue (net/core/dev.c:2954)
xps_rxqs_store (net/core/net-sysfs.c:1880)
netdev_queue_attr_store (net/core/net-sysfs.c:1390)
Allocated by task 634:
__kmalloc_noprof (mm/slub.c:5439)
__netif_set_xps_queue (net/core/dev.c:2937)
The buggy address is located 0 bytes to the right of
allocated 88-byte region [ffff88813016bc00, ffff88813016bc58)
Reject a class the map has no room for.
Fixes: 184c449f91 ("net: Add support for XPS with QoS via traffic classes")
Signed-off-by: Norbert Szetei <norbert@doyensec.com>
Link: https://patch.msgid.link/162DD16F-54C6-444A-9E09-0B8CB3D591F2@doyensec.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
vxlan_na_create() samples LL_RESERVED_SPACE() to size the reply skb and then
samples it again to reserve headroom. A concurrent vxlan_changelink() can
update needed_headroom between the two reads, creating a TOCTOU race. The
second value can exceed the allocation and make the Ethernet header write out
of bounds.
The race is reproducible on the unpatched kernel. It occurred when
vxlan_na_create() generated a neighbour reply while vxlan_changelink() changed
the link headroom. KASAN caught a four-byte write two bytes beyond a 704-byte
skbuff_small_head allocation.
Snapshot the headroom once and use that value for both allocation and
reservation.
Fixes: 4b29dba9c0 ("vxlan: fix nonfunctional neigh_reduce()")
Signed-off-by: Sanghyun Park <sanghyun.park.cnu@gmail.com>
Link: https://patch.msgid.link/20260918032842.502409-2-sanghyun.park.cnu@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
ipgre_netlink_parms() can enable collect_md on an existing GRE, GRETAP
or ERSPAN device. Unlike newlink, changelink does not enforce metadata
tunnel uniqueness. Converting a non-metadata device can therefore
replace the metadata receive entry for another device of the same type
in the same netns. Deleting either device then clears the shared entry,
breaking metadata receive lookup for the surviving device.
If parameter validation fails after collect_md is set, deleting the
modified device can also clear an entry it never owned.
Reject enabling metadata mode in both changelink callbacks before any
encapsulation or tunnel parameters are modified. Allow requests that
repeat the metadata attribute on an existing metadata device.
Fixes: 2e15ea390e ("ip_gre: Add support to collect tunnel metadata.")
Signed-off-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Link: https://patch.msgid.link/20260921031859.9283-1-xuanqiang.luo@linux.dev
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
airoha_npu_remove() calls cancel_work_sync() on each core's wdt_work,
but the watchdog IRQ that queues it is requested with devm_request_irq()
and is freed only after .remove() returns. airoha_npu_wdt_handler() can
therefore schedule_work() again once the cancel has returned. struct
airoha_npu, which contains the work, is devm_kzalloc()'d and is freed in
that same unwind, so the late work dereferences freed memory.
Register the work with devm_work_autocancel() before devm_request_irq()
and drop .remove(). Devres runs in reverse order, so the IRQ is freed
before cancel_work_sync(), including when probe fails. A cancel left in
.remove() cannot get that order. Initializing the work first also stops
a pending watchdog interrupt from queuing an uninitialized work item.
Probe currently calls INIT_WORK() only after devm_request_irq().
This issue was identified during our ongoing static-analysis research
while reviewing kernel code.
Fixes: 23290c7bc1 ("net: airoha: Introduce Airoha NPU support")
Cc: stable@vger.kernel.org # 6.15+
Co-developed-by: Ijae Kim <ae878000@gmail.com>
Signed-off-by: Ijae Kim <ae878000@gmail.com>
Signed-off-by: Myeonghun Pak <mhun512@gmail.com>
Acked-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Link: https://patch.msgid.link/20260922000914.542068-1-mhun512@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Florian Fainelli says:
====================
net: bcmgenet: Collection of bug fixes
This patch series contains a collection of bug fixes found during a
LLM-assisted coding session.
====================
Link: https://patch.msgid.link/20260921220021.281418-1-florian.fainelli@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
bcmgenet_get_coalesce() reads DMA_RING0_TIMEOUT to calculate
rx_coalesce_usecs without masking out bits outside DMA_TIMEOUT_MASK
(16 bits). If upper bits are non-zero or contain status/flags, the
computed value of rx_coalesce_usecs returned to userspace via ethtool
becomes corrupted.
Mask the register read with DMA_TIMEOUT_MASK before computing the
timeout in microseconds.
Fixes: 4a29645bfe ("net: bcmgenet: Implement RX coalescing control knobs")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Florian Fainelli <florian.fainelli@broadcom.com>
Link: https://patch.msgid.link/20260921220021.281418-6-florian.fainelli@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
bcmgenet_set_mac_addr() did not check whether the provided MAC address is a
valid Ethernet address before applying it. Userspace could configure an
invalid address (such as all zeroes or a multicast address) while the
interface is down.
Add a call to is_valid_ether_addr() and return -EADDRNOTAVAIL if the MAC
address is not valid.
Fixes: 1c1008c793 ("net: bcmgenet: add main driver file")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Florian Fainelli <florian.fainelli@broadcom.com>
Link: https://patch.msgid.link/20260921220021.281418-5-florian.fainelli@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
bcmgenet_power_up() had an early check for bcmgenet_has_ext(priv) before
dispatching by power mode. GENET V1 does not have the EXT block (unlike
GENET V2+), which causes bcmgenet_power_up() to immediately return 0.
As a consequence, when waking up from GENET_POWER_WOL_MAGIC on GENET V1,
bcmgenet_wol_power_up_cfg() is never invoked to disable the WoL clock,
clear wake event masks, and restore normal PHY and MAC operations.
Move the bcmgenet_has_ext() checks to the GENET_POWER_PASSIVE and
GENET_POWER_CABLE_SENSE cases where the EXT registers are actually
accessed, allowing GENET_POWER_WOL_MAGIC cleanup to execute on all
hardware versions.
Fixes: c3ae64ae0c ("net: bcmgenet: handle GENET_POWER_WOL_MAGIC")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Florian Fainelli <florian.fainelli@broadcom.com>
Link: https://patch.msgid.link/20260921220021.281418-4-florian.fainelli@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
bcmgenet_gstrings_stats statically defines ethtool statistics for queues
0 through GENET_MAX_MQ_CNT (4). However, bcmgenet_probe() only initialized
the u64_stats_sync seq counter up to priv->hw_params->rx_queues and
priv->hw_params->tx_queues.
Since priv->hw_params->rx_queues is 0 across all hardware versions (and
priv->hw_params->tx_queues is 0 on GENET V1), rings 1..4 have uninitialized
u64_stats_sync structures. When ethtool -S is run on 32-bit kernels,
bcmgenet_get_ethtool_stats() reads stats from rx_rings[1..4], causing
lockdep warnings due to the uninitialized sequence counters.
Initialize the sequence counters for all GENET_MAX_MQ_CNT + 1 queues.
Fixes: ffc2c8c4a7 ("net: bcmgenet: Initialize u64 stats seq counter")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Florian Fainelli <florian.fainelli@broadcom.com>
Link: https://patch.msgid.link/20260921220021.281418-3-florian.fainelli@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
When bcmgenet was converted to 64-bit statistics, STAT_RTNL members were
switched to point into struct rtnl_link_stats64, whose fields are 64-bit
(__u64) regardless of architecture.
However, bcmgenet_get_ethtool_stats() retained a legacy check:
if (sizeof(unsigned long) != sizeof(u32) &&
s->stat_sizeof == sizeof(unsigned long))
On 32-bit systems, sizeof(unsigned long) == sizeof(u32), causing this
condition to evaluate to false. As a result, 64-bit RTNL stats fields were
read via *(u32 *)p. On 32-bit Big-Endian systems (such as MIPS BE), this
reads the high 32 bits and returns 0 until the counter exceeds 4GB; on
32-bit Little-Endian systems (such as 32-bit ARM), the value is truncated
to 32 bits.
Fix this by checking if s->stat_sizeof == sizeof(u64) so 64-bit fields are
always read as 64-bit values.
Fixes: 59aa6e3072 ("net: bcmgenet: switch to use 64bit statistics")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Florian Fainelli <florian.fainelli@broadcom.com>
Link: https://patch.msgid.link/20260921220021.281418-2-florian.fainelli@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The seg6, ioam6 and rpl lwtunnels size their skb_cow_head() request as
the length they are about to push plus dst_dev_overhead(), then push the
new headers and rebuild the mac header below them with
skb_mac_header_rebuild(). That rebuild needs skb->mac_len of headroom,
but dst_dev_overhead() leaves LL_RESERVED_SPACE() of the egress device,
16 bytes for plain Ethernet.
Where the mac header is longer than that, as it is on ingress through a
VLAN device with reorder_hdr off, the rebuild runs out of room:
skb_set_mac_header(skb, -skb->mac_len) computes a negative offset,
stores it unchecked in the u16 skb->mac_header, and the memmove that
follows writes skb->mac_len bytes about 64 KB past skb->head.
Forwarding plain ping6 traffic through such a device reproduces it on
all five seg6 encapsulation modes and on the rpl and ioam6 inline paths;
skb->mac_header comes back as 65534 on a 704-byte head.
Return the larger of the two. The helper already returns skb->mac_len
when it has no dst, so this only makes the other branch agree, and it
covers every caller rather than each call site in turn.
Fixes: 40475b6376 ("net: ipv6: seg6_iptunnel: mitigate 2-realloc issue")
Fixes: dce525185b ("net: ipv6: ioam6_iptunnel: mitigate 2-realloc issue")
Fixes: 985ec6f5e6 ("net: ipv6: rpl_iptunnel: mitigate 2-realloc issue")
Suggested-by: Andrea Mayer <andrea.mayer@uniroma2.it>
Signed-off-by: Yuya Kusakabe <yuya.kusakabe@gmail.com>
Reviewed-by: Justin Iurman <justin.iurman@gmail.com>
Reviewed-by: Gabriel Goller <g.goller@proxmox.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Andrea Mayer <andrea.mayer@uniroma2.it>
Link: https://patch.msgid.link/20260922-seg6-maclen-headroom-v3-1-7b2f982ef79d@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
In stmmac_rx(), when napi_build_skb() fails the descriptor page is
recycled back to the page pool with page_pool_recycle_direct(), but
buf->page is left pointing at the recycled page, unlike every other
consumption site in the function which clears the pointer after handing
the page away.
With the stale pointer stmmac_rx_refill() skips the replacement
allocation and programs the already-recycled page back into the RX
descriptor.
Clear buf->page on the napi_build_skb() failure path to keep the buffer
lifecycle consistent with the other consumption sites.
Fixes: df542f6693 ("net: stmmac: Switch to zero-copy in non-XDP RX path")
Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/20260921-stmmac-fix-napi-build-skb-error-v1-1-3d54bf6d9bb6@oss.qualcomm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Some of the tg3 NICs we use (BCM57762) reset the SRAM MAC address to the
placeholder address on link flaps, tg3_chip_reset, etc. We've typically
fixed it from userspace, but since e4c00ba727 ("tg3: replace
placeholder MAC address with device property") was merged, tg3 just
fails probe as we don't have a way to get it through the generic
device_get_mac_address infrastructure as fallback on our systems.
Make the driver assign a random address in this condition instead of
being fatal for probe. Set deferred_probe_reason through dev_warn_probe
if the address isn't yet available from the provider.
Fixes: e4c00ba727 ("tg3: replace placeholder MAC address with device property")
Suggested-by: Jakub Kicinski <kuba@kernel.org>
Link: https://lore.kernel.org/netdev/20260909191751.651aa5c4@kernel.org/
Signed-off-by: Ivan Delalande <colona@arista.com>
Link: https://patch.msgid.link/20260918224715.GA654128@visor
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The ARP ioctl copies a user-provided struct arpreq into a stack object. Its
arp_dev field may contain IFNAMSIZ bytes without a NUL terminator.
Such input is passed to dev_get_by_name_rcu() or __dev_get_by_name(), where
strcmp() can read past the end of the stack object when a matching
alternative interface name exists.
Terminate the field before the lookup to prevent the out-of-bounds read.
Fixes: 36fbf1e52b ("net: rtnetlink: add linkprop commands to add and delete alternative ifnames")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Signed-off-by: Zijie Huang <milkory@outlook.com>
Signed-off-by: Ren Wei <weir@nebusec.ai>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/fabf02a70787d17299e4b3153eadffaf20d154b3.1789910973.git.milkory@outlook.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The RTL931x indirect access engine has a separate nine-bit extended page
field. The driver leaves it at zero, and otto_emdio_run_cmd() therefore
programs extended page zero for every Clause 22 transaction. This
overrides page selection made through PHY register 30, causing accesses
to private PHY pages to hit extended page zero instead.
Set the field to its 0x1ff "do not change" value for RTL931x Clause 22
reads and writes. This preserves extended page selection made through
PHY register 30 and restores access to its private register pages.
Fixes: 5ebdcac59a ("net: mdio: realtek-rtl9300: Add support for RTL931x")
Signed-off-by: Jonas Jelonek <jonas@jonasjelonek.de>
Acked-by: Markus Stockhausen <markus.stockhausen@gmx.de>
Link: https://patch.msgid.link/20260918211955.3955777-1-jonas@jonasjelonek.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Commit 7a9bc9e3f4 ("fou: Don't allow 0 for FOU_ATTR_IPPROTO.") added
NLA_POLICY_MIN(NLA_U8, 1) to fou_nl_policy[FOU_ATTR_IPPROTO], which
rejects an explicitly supplied FOU_ATTR_IPPROTO == 0 attribute with
-ERANGE.
However, FOU_ATTR_IPPROTO is an optional netlink attribute. When a user
sends FOU_CMD_ADD with FOU_ATTR_TYPE set to FOU_ENCAP_DIRECT and omits
FOU_ATTR_IPPROTO entirely, nla_policy validation succeeds and
parse_nl_config() leaves cfg->protocol as 0 (from memset(cfg, 0,
sizeof(*cfg))). fou_create() then creates a FOU_ENCAP_DIRECT socket with
fou->protocol == 0.
In fou_udp_recv(), returning -fou->protocol to udp_queue_rcv_one_skb()
triggers IP protocol resubmission when fou->protocol > 0, whereas
returning 0 tells the UDP tunnel layer that the skb was consumed without
freeing it. When fou->protocol == 0, every packet received on the socket
returns 0 from fou_udp_recv() and leaks the sk_buff.
Reject FOU_ENCAP_DIRECT when !cfg->protocol in fou_create() so that
creating a direct encapsulation port without FOU_ATTR_IPPROTO fails with
-EINVAL while leaving FOU_CMD_DEL and FOU_CMD_GET (which share
parse_nl_config()) unaffected.
Tested in QEMU against Linux 7.3.0-rc3 by sending a FOU_CMD_ADD Generic
Netlink request with FOU_ATTR_PORT = 5555 and FOU_ATTR_TYPE =
FOU_ENCAP_DIRECT while omitting FOU_ATTR_IPPROTO. On the unfixed kernel,
FOU_CMD_ADD succeeds (err = 0), FOU_CMD_GET reports fou->type = 1 and
fou->protocol = 0, and sending 4000 UDP packets to 127.0.0.1:5555 leaks
all 4000 sk_buffs (SUnreclaim in /proc/meminfo grows from 41456 kB to
59008 kB, +17552 kB); with this patch applied, FOU_CMD_ADD is rejected
with -EINVAL (-22).
Fixes: 23461551c0 ("fou: Support for foo-over-udp RX path")
Fixes: 7a9bc9e3f4 ("fou: Don't allow 0 for FOU_ATTR_IPPROTO.")
Cc: stable@vger.kernel.org
Signed-off-by: Hui Peng <benquike@gmail.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Link: https://patch.msgid.link/20260921045920.1613098-1-benquike@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The 88E6191X and 88E6193X are 6393 family devices that share
mv88e6393x_ops with the 88E6393X and are marked as ptp_support. Marvell's
UMSD driver describes both as parts without AVB (88E6193X: "BGA package -
No AVB, No Routing, No Cut-through"), and the register access confirms it
on an 88E6193X: the whole indirect AVB register space behind Global 2
registers 0x16 and 0x17 reads zero, for every port, block and address,
with the 6390 and with the 6352 command encoding. Writes to the TAI
registers, including the clock period register and the TAI global
configuration register, read back as zero.
Since commit 7e3c18097a ("net: dsa: mv88e6xxx: read cycle counter
period from hardware") the PTP setup reads the TAI clock period, so the
switch fails to probe:
mv88e6xxx ...: unexpected cycle counter period of 0 ps
Add mv88e6191x_ops, a copy of mv88e6393x_ops without avb_ops and ptp_ops,
use it for the 88E6191X and the 88E6193X and stop setting ptp_support for
them. The 88E6393X is unchanged.
Tested on an 88E6193X (Sophos XGS 107w): the switch probes and the ports
work. I do not have an 88E6191X, it is changed because UMSD describes it
the same way.
Fixes: de776d0d31 ("net: dsa: mv88e6xxx: add support for mv88e6393x family")
Suggested-by: Andrew Lunn <andrew@lunn.ch>
Signed-off-by: Nicolo Giuliani <nicolo.giuliani6@studio.unibo.it>
Reviewed-by: Andrew Lunn <andrew@lunn.ch>
Link: https://patch.msgid.link/20260921-send-net-v2-1-031ad720f140@studio.unibo.it
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
In seg6_genl_policy, SEG6_ATTR_DST is defined with .type = NLA_BINARY and
.len = sizeof(struct in6_addr). For NLA_BINARY, .len only enforces the
maximum payload length and permits shorter payloads (e.g., 0 bytes).
When seg6_genl_set_tunsrc() copies sizeof(struct in6_addr) bytes via
kmemdup(val, sizeof(*val), GFP_KERNEL), a short SEG6_ATTR_DST attribute
triggers a 16-byte out-of-bounds read past skb->tail into uninitialized
skb->head memory, which is stored in sdata->tun_src and leaked back to
userspace via SEG6_CMD_GET_TUNSRC.
Switch SEG6_ATTR_DST in seg6_genl_policy to
NLA_POLICY_EXACT_LEN(sizeof(struct in6_addr)) so that generic netlink
validation rejects any attribute whose length is not exactly
sizeof(struct in6_addr) with -ERANGE.
Tested in QEMU against Linux 7.3.0-rc3 by sending a SEG6_CMD_SET_TUNSRC
Generic Netlink message with a 0-byte SEG6_ATTR_DST attribute followed
by SEG6_CMD_GET_TUNSRC. On the unfixed kernel, SEG6_CMD_SET_TUNSRC
succeeds (err = 0) and SEG6_CMD_GET_TUNSRC leaks 16 bytes of
uninitialized kernel heap memory (tun_src =
836a61ecc4d25a1042a8d60411cfb378); with this patch applied,
SEG6_CMD_SET_TUNSRC is rejected by netlink policy validation with
-ERANGE (-34) and tun_src remains zeroed.
Fixes: 915d7e5e59 ("ipv6: sr: add code base for control plane support of SR-IPv6")
Cc: stable@vger.kernel.org
Signed-off-by: Hui Peng <benquike@gmail.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Reviewed-by: Justin Iurman <justin.iurman@gmail.com>
Reviewed-by: Andrea Mayer <andrea.mayer@uniroma2.it>
Link: https://patch.msgid.link/20260921044025.1535982-1-benquike@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The VRF device is an Ethernet device but it can have non-Ethernet ports
such as IP tunnels. Before the cited commit, capturing packets from such
ports on the VRF device resulted in these packets being detected as
malformed since they lack an Ethernet header.
The cited commit fixed it by pushing a dummy Ethernet header to such
packets before the capture and pulling it afterwards. In the case of
CHECKSUM_COMPLETE packets it also updated skb->csum with the checksum of
the dummy Ethernet header. This is wrong as skb->csum should not include
the checksum of the Ethernet header ("checksum of the _whole_ packet as
seen by netif_rx()").
This also means that L4 protocols receive a corrupted skb->csum and
potentially drop the packet, as is the case with UDP packets whose
checksum was completed by software.
Fix by removing the unnecessary call to skb_postpush_rcsum().
Fixes: 0489390882 ("vrf: add mac header for tunneled packets when sniffer is attached")
Reported-by: Stefano Sasso <stesasso@gmail.com>
Closes: https://lore.kernel.org/netdev/CALtE316UtL3x7LL6uxfXzx8rW6AbzYPeDOb478hqJCr_-dj=Wg@mail.gmail.com/
Signed-off-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: David Ahern <dsahern@kernel.org>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Andrea Mayer <andrea.mayer@uniroma2.it>
Link: https://patch.msgid.link/20260922131239.2509494-1-idosch@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
skb_maybe_pull_tail() subtracts skb_headlen(skb) from the unsigned max
argument and passes the result to __pskb_pull_tail() as a signed int. The
function does not ensure that max is at least skb_headlen(skb).
This can happen while parsing IPv6 extension headers when an skb already
has a linear area larger than MAX_IPV6_HDR_LEN. Once the parser needs data
beyond the linear area, max - skb_headlen(skb) wraps and is converted to a
negative delta. __pskb_pull_tail() then passes that negative length to
skb_copy_bits(), where it can become a very large copy length.
Pass the requested length itself as the pull bound at the three
extension-header call sites, so the delta can no longer go negative.
Fixes: 1431fb31ec ("xen-netback: fix fragment detection in checksum setup")
Suggested-by: Eric Dumazet <edumazet@google.com>
Signed-off-by: Shihuang Liu <shlomojune6@gmail.com>
Link: https://patch.msgid.link/20260919133604.50948-1-shlomojune6@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
veth_set_channels() tears down XDP resources for removed RX queues
without clearing rq->xdp_prog. If the program is then detached or
replaced, those queues keep the old pointer after bpf_prog_put().
A later channel increase can re-enable NAPI and run the freed program.
BUG: unable to handle page fault for address: ffffc90000256048
Oops: Oops: 0000 [#1] SMP KASAN NOPTI
RIP: veth_xdp_rcv_skb (include/linux/filter.h:779
include/net/xdp.h:696 drivers/net/veth.c:820)
Call Trace:
veth_xdp_rcv (drivers/net/veth.c:941)
veth_poll (drivers/net/veth.c:986)
__napi_poll (net/core/dev.c:7787)
net_rx_action (net/core/dev.c:7850 net/core/dev.c:8007)
handle_softirqs (kernel/softirq.c:645)
Kernel panic - not syncing: Fatal exception in interrupt
Fixes: 4752eeb3d8 ("veth: implement support for set_channel ethtool op")
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Reviewed-by: Jason Xing <kerneljasonxing@gmail.com>
Link: https://patch.msgid.link/20260921231856.1798630-1-kuba@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
bcmgenet_netif_stop() and the Wake-on-LAN branch of bcmgenet_suspend()
both disable the Tx queues first and stop Tx NAPI several steps later. A
completion in flight calls netif_tx_wake_queue() in between, and nothing
stops the queue again, so a transmit can reach the rings after they have
been freed.
Close is safe because dev_deactivate_many() stops the qdisc first.
bcmgenet_suspend() does not, so stop Tx NAPI before the queues on both
paths.
KASAN on a Raspberry Pi CM4, driven from an MTU change because suspend
freezes user space before the callback runs:
BUG: KASAN: use-after-free in bcmgenet_xmit+0x17f8/0x2258
Write of size 8 at addr ffffff8055844a68 by task ksoftirqd/0/14
bcmgenet_xmit+0x17f8/0x2258
dev_hard_start_xmit+0x13c/0x588
sch_direct_xmit+0x108/0x340
__dev_queue_xmit+0x1190/0x3848
Fixes: 254f3239dd ("net: bcmgenet: revise suspend/resume")
Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
Link: https://patch.msgid.link/20260922130639.1660797-1-nb@tipi-net.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
I have been contributing to and reviewing the GENET driver for a while
now. Florian asked if I would like to formalize this commitment, so add
myself as a maintainer.
Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
Acked-by: Florian Fainelli <florian.fainelli@broadcom.com>
Acked-by: Justin Chen <justin.chen@broadcom.com>
Link: https://patch.msgid.link/20260922073140.1471858-1-nb@tipi-net.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
br_mdb_flush_pgs() keeps a pointer-to-pointer cursor while walking
mp->ports. br_multicast_del_pg() can re-enter the same MDB entry through
br_multicast_sg_del_exclude_ports() and unlink other port groups. If the
cursor points into one of those groups, the next iteration dereferences a
stale cursor and can leave mp->ports pointing at freed memory.
A following RTM_GETMDB exposes the dangling pointer:
BUG: KASAN: slab-use-after-free in br_mdb_dump
Read of size 8
br_mdb_dump
rtnl_mdb_dump
rtnl_dumpit
netlink_dump
Reset the cursor to mp->ports after every deletion. The deletion removes at
least the selected group, so the restarted walk always makes progress.
Fixes: a6acb535af ("bridge: mdb: Add MDB bulk deletion support")
Cc: stable@vger.kernel.org
Signed-off-by: Fourie Zhang <fouriezhang@tencent.com>
Acked-by: Nikolay Aleksandrov <razor@blackwall.org>
Link: https://patch.msgid.link/20260920110852.60293-1-fouriezhang@tencent.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
tcf_gate_get_fill_size returns only the TCA_GATE_PARMS size, but
tcf_gate_dump also emits three 64-bit timestamps, the clock id, flags,
priority and the variable-length TCA_GATE_ENTRY_LIST nest. The per-entry
nest is unbounded: parse_gate_list places no cap on the number of
sched-entries, so a gate with many entries can push the real dump well
past the skb that tca_get_fill allocates from this size.
RTM_NEWACTION then fails the add-notify with -EINVAL while the action is
already committed to the IDR, and a subsequent RTM_GETACTION on the
installed gate also returns -EINVAL because its dump no longer fits.
Fix this by accounting for the missing fields in tcf_gate_get_fill_size
along with all elements in the entries list.
Note that sizing the reply from the action lets an oversized gate
install cleanly for the first time: with the input unbounded by
parse_gate_list, the sized skb can now grow well above
NLMSG_GOODSIZE per netlink request (a transient GFP_KERNEL allocation
reachable only with namespace-local CAP_NET_ADMIN). Overload from a
malicious netns admin is hardening material, not net, per the
discussion at
https://lore.kernel.org/netdev/20260914191108.55a1a4f1@kernel.org/;
a follow-up patch for net-next will cap the sched-entry count.
Fixes: 4e76e75d6a ("net sched actions: calculate add/delete event message size")
Reported-by: Sashiko <sashiko-bot@kernel.org>
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260824153903.4143642-1-victor@mojatatu.com
Tested-by: hybris <hybris@mojatatu.ai>
Co-developed-by: Jamal Hadi Salim <jhs@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Signed-off-by: Victor Nogueira <victor@mojatatu.com>
Link: https://patch.msgid.link/QDISC-3BLH.v1.20260914203033@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
frame->ccid.datalen is read directly from the USB response frame
and used, unchecked, as an index into frame->data[]. A malicious or
malfunctioning device can set this field to an arbitrary value,
causing the driver to read far outside the received buffer.
Bound ccid.datalen against the maximum possible ACR122 frame size
before using it. This replaces the existing datalen == 0 check,
since datalen < 2 already covers that case and additionally
rejects datalen == 1, which would still underflow the
"datalen - 2" offset used below.
Fixes: 9815c7cf22 ("NFC: pn533: Separate physical layer from the core implementation")
Reported-by: syzbot+1853daab1a47603d4678@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=1853daab1a47603d4678
Tested-by: syzbot+1853daab1a47603d4678@syzkaller.appspotmail.com
Assisted-by: LLM
Signed-off-by: Deepanshu Kartikey <kartikey406@gmail.com>
Link: https://patch.msgid.link/20260923035627.6210-1-kartikey406@gmail.com
Signed-off-by: David Heidelberg <david@ixit.cz>
nfc_llcp_wks_sap() and nfc_llcp_build_sdreq_tlv() pass non-null-
terminated strings to pr_debug() using the %s format specifier.
The buffers are allocated via kmemdup() or come from netlink
attributes and are not guaranteed to be null-terminated, causing
__dynamic_pr_debug() to read beyond the allocated region:
KASAN: slab-out-of-bounds Read in __dynamic_pr_debug
Fix both call sites by using %.*s with the explicit length to limit
the output to the actual length of the string.
Fixes: d9b8d8e19b ("NFC: llcp: Service Name Lookup netlink interface")
Reported-by: syzbot+1e3df0852e82c21ca418@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=1e3df0852e82c21ca418
Signed-off-by: Ömer Mete Kaya <omermetekaya0@gmail.com>
Link: https://patch.msgid.link/20260908161952.731468-1-omermetekaya0@gmail.com
Signed-off-by: David Heidelberg <david@ixit.cz>
nfc_llcp_wks_sap() compares only service_name_len bytes, so a short
service_name like "u" matches longer WKS strings like "urn:nfc:sn:snep".
Fix by requiring exact length match before strncmp().
Fixes: d646960f79 ("NFC: Initial LLCP support")
Signed-off-by: Ömer Mete Kaya <omermetekaya0@gmail.com>
Link: https://patch.msgid.link/20260909121437.33744-1-omermetekaya0@gmail.com
Signed-off-by: David Heidelberg <david@ixit.cz>
When service_name_len is 0, kmemdup() returns ZERO_SIZE_PTR which
passes the NULL check, causing nfc_llcp_send_connect() to attempt
building a zero-length service name TLV and fail with -ENOMEM.
Fix by setting service_name to NULL directly when service_name_len is 0.
Fixes: d646960f79 ("NFC: Initial LLCP support")
Signed-off-by: Ömer Mete Kaya <omermetekaya0@gmail.com>
Link: https://patch.msgid.link/20260909122029.34081-1-omermetekaya0@gmail.com
Signed-off-by: David Heidelberg <david@ixit.cz>
The ISO15693 inventory helper removes a two-byte prefix without checking
that it exists, then accepts a one-byte remainder before reading data[1] as
the DSFID.
Require the prefix and at least two remaining bytes before copying the UID
data and reading the DSFID.
Fixes: 7974728094 ("NFC: st21nfca: Add ISO15693 Reader/Writer support")
Signed-off-by: Pengpeng Hou <pengpeng@iscas.ac.cn>
Link: https://patch.msgid.link/20260830132958.6397-1-pengpeng@iscas.ac.cn
Signed-off-by: David Heidelberg <david@ixit.cz>
Commit 6709d4b7bc ("net: nfc: Fix use-after-free caused by
nfc_llcp_find_local") attempted to fix a use-after-free (UAF) issue by
invoking nfc_llcp_local_put(local) after accessing local->gb. However,
if the reference count drops to zero, local is freed immediately,
leading to a use-after-free when callers access the returned pointer.
Alternative approaches using dynamic allocation (e.g. kmemdup) introduced
memory leaks because callers consistently treat the returned pointer as
borrowed memory.
Fix this properly by refactoring nfc_llcp_general_bytes() and
nfc_get_local_general_bytes() to accept a caller-provided output buffer
(out_gb) and its maximum length (gb_max_len). The general bytes are
safely copied into out_gb before calling nfc_llcp_local_put(local),
ensuring safe lifetime management without ownership transfer complications.
Update all callers across drivers (microread, pn533, pn544, st21nfca,
digital_dep, and nci) to provide their own destination buffers and pass
them to nfc_get_local_general_bytes().
Fixes: 6709d4b7bc ("net: nfc: Fix use-after-free caused by nfc_llcp_find_local")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Assisted-by: LLM
Signed-off-by: Luxiao Xu <rakukuip@gmail.com>
Signed-off-by: Ren Wei <weir@nebusec.ai>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/3cbaac3bee23f8ff3a3284ed32d347696eb1d208.1788841683.git.rakukuip@gmail.com
Signed-off-by: David Heidelberg <david@ixit.cz>
trf7970a_startup() powers up the device before applying the optional RX
gain reduction. If the register read or write fails, it returns without
undoing that power-up. Probe's unwind only drops the separate regulator
references acquired by probe, leaving the additional VIN enable from
startup unbalanced. The system resume caller also has no power-down on
this error.
Call trf7970a_power_down() before returning the RX gain error to deassert
the enable GPIOs, release the startup VIN reference and restore the
powered-off state. Runtime PM has not been enabled yet, so the full
shutdown helper is not appropriate here. Preserve the original SPI error.
This issue was identified during our ongoing static-analysis research while
reviewing kernel code.
Fixes: 5d69351820 ("NFC: trf7970a: Create device-tree parameter for RX gain reduction")
Cc: stable@vger.kernel.org
Assisted-by: OpenAI:GPT-5.6
Co-developed-by: Ijae Kim <ae878000@gmail.com>
Signed-off-by: Ijae Kim <ae878000@gmail.com>
Signed-off-by: Myeonghun Pak <mhun512@gmail.com>
Reviewed-by: Paul Geurts <paul.geurts@prodrive-technologies.com>
Link: https://patch.msgid.link/20260913042625.31296-1-mhun512@gmail.com
Signed-off-by: David Heidelberg <david@ixit.cz>
nfc_genl_llc_sdreq() builds a list of TLV nodes while walking nested
netlink attrs, but 3 error paths (nested-attr parse failure, TLV alloc
ENOMEM, nfc_llcp_send_snl_sdreq() failure) all skip freeing what was
already queued.
Route them through a new free_list label, mirroring the SDRES path in
the same file which already does this. Harmless on the success path
too -- send_snl_sdreq() drains the list as it moves nodes, so it's
already empty by the time free_list runs.
Fixes: d9b8d8e19b ("NFC: llcp: Service Name Lookup netlink interface")
Assisted-by: Claude:claude-opus-4
Signed-off-by: Cong Nguyen <congnt264@gmail.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260914121129.2098606-1-congnt264@gmail.com
Signed-off-by: David Heidelberg <david@ixit.cz>
nfc_llcp_recv_hdlc() reads the sequence byte skb->data[2], via
nfc_llcp_ns()/nfc_llcp_nr(), before any length check. The receive path
only guarantees the two-byte LLCP header -- __nfc_llcp_recv() checks it
with pskb_may_pull() and nfc_llcp_recv_agf() admits two-byte inner PDUs
-- so a two-byte I, RR or RNR PDU reads one byte of uninitialised skb
tailroom. The byte becomes N(R)/N(S); a peer can already set those with
a well-formed PDU, so this is acting on uninitialised memory, not new
peer control.
Guard the read with pskb_may_pull(), as commit 95674f506c ("nfc: llcp:
reject PDUs shorter than the LLCP header") did for the two-byte header,
so the sequence byte is present and linear before it is read. RR and RNR
PDUs are LLCP_HEADER_SIZE + LLCP_SEQUENCE_SIZE bytes and an I PDU is
longer, so no valid frame is rejected; a truncated PDU is malformed, so
return without a DM reply.
Fixes: d646960f79 ("NFC: Initial LLCP support")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Aamir Ahmed <elb12345@hotmail.co.uk>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/AS8P251MB0001789BBF04B72745C7D96BC8BA2@AS8P251MB0001.EURP251.PROD.OUTLOOK.COM
Signed-off-by: David Heidelberg <david@ixit.cz>