The VRF device is an Ethernet device but it can have non-Ethernet ports
such as IP tunnels. Before the cited commit, capturing packets from such
ports on the VRF device resulted in these packets being detected as
malformed since they lack an Ethernet header.
The cited commit fixed it by pushing a dummy Ethernet header to such
packets before the capture and pulling it afterwards. In the case of
CHECKSUM_COMPLETE packets it also updated skb->csum with the checksum of
the dummy Ethernet header. This is wrong as skb->csum should not include
the checksum of the Ethernet header ("checksum of the _whole_ packet as
seen by netif_rx()").
This also means that L4 protocols receive a corrupted skb->csum and
potentially drop the packet, as is the case with UDP packets whose
checksum was completed by software.
Fix by removing the unnecessary call to skb_postpush_rcsum().
Fixes: 0489390882 ("vrf: add mac header for tunneled packets when sniffer is attached")
Reported-by: Stefano Sasso <stesasso@gmail.com>
Closes: https://lore.kernel.org/netdev/CALtE316UtL3x7LL6uxfXzx8rW6AbzYPeDOb478hqJCr_-dj=Wg@mail.gmail.com/
Signed-off-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: David Ahern <dsahern@kernel.org>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Andrea Mayer <andrea.mayer@uniroma2.it>
Link: https://patch.msgid.link/20260922131239.2509494-1-idosch@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
skb_maybe_pull_tail() subtracts skb_headlen(skb) from the unsigned max
argument and passes the result to __pskb_pull_tail() as a signed int. The
function does not ensure that max is at least skb_headlen(skb).
This can happen while parsing IPv6 extension headers when an skb already
has a linear area larger than MAX_IPV6_HDR_LEN. Once the parser needs data
beyond the linear area, max - skb_headlen(skb) wraps and is converted to a
negative delta. __pskb_pull_tail() then passes that negative length to
skb_copy_bits(), where it can become a very large copy length.
Pass the requested length itself as the pull bound at the three
extension-header call sites, so the delta can no longer go negative.
Fixes: 1431fb31ec ("xen-netback: fix fragment detection in checksum setup")
Suggested-by: Eric Dumazet <edumazet@google.com>
Signed-off-by: Shihuang Liu <shlomojune6@gmail.com>
Link: https://patch.msgid.link/20260919133604.50948-1-shlomojune6@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
veth_set_channels() tears down XDP resources for removed RX queues
without clearing rq->xdp_prog. If the program is then detached or
replaced, those queues keep the old pointer after bpf_prog_put().
A later channel increase can re-enable NAPI and run the freed program.
BUG: unable to handle page fault for address: ffffc90000256048
Oops: Oops: 0000 [#1] SMP KASAN NOPTI
RIP: veth_xdp_rcv_skb (include/linux/filter.h:779
include/net/xdp.h:696 drivers/net/veth.c:820)
Call Trace:
veth_xdp_rcv (drivers/net/veth.c:941)
veth_poll (drivers/net/veth.c:986)
__napi_poll (net/core/dev.c:7787)
net_rx_action (net/core/dev.c:7850 net/core/dev.c:8007)
handle_softirqs (kernel/softirq.c:645)
Kernel panic - not syncing: Fatal exception in interrupt
Fixes: 4752eeb3d8 ("veth: implement support for set_channel ethtool op")
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Reviewed-by: Jason Xing <kerneljasonxing@gmail.com>
Link: https://patch.msgid.link/20260921231856.1798630-1-kuba@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
bcmgenet_netif_stop() and the Wake-on-LAN branch of bcmgenet_suspend()
both disable the Tx queues first and stop Tx NAPI several steps later. A
completion in flight calls netif_tx_wake_queue() in between, and nothing
stops the queue again, so a transmit can reach the rings after they have
been freed.
Close is safe because dev_deactivate_many() stops the qdisc first.
bcmgenet_suspend() does not, so stop Tx NAPI before the queues on both
paths.
KASAN on a Raspberry Pi CM4, driven from an MTU change because suspend
freezes user space before the callback runs:
BUG: KASAN: use-after-free in bcmgenet_xmit+0x17f8/0x2258
Write of size 8 at addr ffffff8055844a68 by task ksoftirqd/0/14
bcmgenet_xmit+0x17f8/0x2258
dev_hard_start_xmit+0x13c/0x588
sch_direct_xmit+0x108/0x340
__dev_queue_xmit+0x1190/0x3848
Fixes: 254f3239dd ("net: bcmgenet: revise suspend/resume")
Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
Link: https://patch.msgid.link/20260922130639.1660797-1-nb@tipi-net.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
I have been contributing to and reviewing the GENET driver for a while
now. Florian asked if I would like to formalize this commitment, so add
myself as a maintainer.
Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
Acked-by: Florian Fainelli <florian.fainelli@broadcom.com>
Acked-by: Justin Chen <justin.chen@broadcom.com>
Link: https://patch.msgid.link/20260922073140.1471858-1-nb@tipi-net.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
br_mdb_flush_pgs() keeps a pointer-to-pointer cursor while walking
mp->ports. br_multicast_del_pg() can re-enter the same MDB entry through
br_multicast_sg_del_exclude_ports() and unlink other port groups. If the
cursor points into one of those groups, the next iteration dereferences a
stale cursor and can leave mp->ports pointing at freed memory.
A following RTM_GETMDB exposes the dangling pointer:
BUG: KASAN: slab-use-after-free in br_mdb_dump
Read of size 8
br_mdb_dump
rtnl_mdb_dump
rtnl_dumpit
netlink_dump
Reset the cursor to mp->ports after every deletion. The deletion removes at
least the selected group, so the restarted walk always makes progress.
Fixes: a6acb535af ("bridge: mdb: Add MDB bulk deletion support")
Cc: stable@vger.kernel.org
Signed-off-by: Fourie Zhang <fouriezhang@tencent.com>
Acked-by: Nikolay Aleksandrov <razor@blackwall.org>
Link: https://patch.msgid.link/20260920110852.60293-1-fouriezhang@tencent.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
tcf_gate_get_fill_size returns only the TCA_GATE_PARMS size, but
tcf_gate_dump also emits three 64-bit timestamps, the clock id, flags,
priority and the variable-length TCA_GATE_ENTRY_LIST nest. The per-entry
nest is unbounded: parse_gate_list places no cap on the number of
sched-entries, so a gate with many entries can push the real dump well
past the skb that tca_get_fill allocates from this size.
RTM_NEWACTION then fails the add-notify with -EINVAL while the action is
already committed to the IDR, and a subsequent RTM_GETACTION on the
installed gate also returns -EINVAL because its dump no longer fits.
Fix this by accounting for the missing fields in tcf_gate_get_fill_size
along with all elements in the entries list.
Note that sizing the reply from the action lets an oversized gate
install cleanly for the first time: with the input unbounded by
parse_gate_list, the sized skb can now grow well above
NLMSG_GOODSIZE per netlink request (a transient GFP_KERNEL allocation
reachable only with namespace-local CAP_NET_ADMIN). Overload from a
malicious netns admin is hardening material, not net, per the
discussion at
https://lore.kernel.org/netdev/20260914191108.55a1a4f1@kernel.org/;
a follow-up patch for net-next will cap the sched-entry count.
Fixes: 4e76e75d6a ("net sched actions: calculate add/delete event message size")
Reported-by: Sashiko <sashiko-bot@kernel.org>
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260824153903.4143642-1-victor@mojatatu.com
Tested-by: hybris <hybris@mojatatu.ai>
Co-developed-by: Jamal Hadi Salim <jhs@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Signed-off-by: Victor Nogueira <victor@mojatatu.com>
Link: https://patch.msgid.link/QDISC-3BLH.v1.20260914203033@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
em_text_dump() allocates struct tcf_em_text on the stack without zeroing
it. strscpy() writes the algorithm name and a NUL terminator into
conf.algo[], leaving the remaining bytes uninitialised. nla_put_nohdr()
then copies the full struct to the netlink response.
KMSAN on Linux 7.2-rc6 reports two kernel-infoleak splats from this path,
one triggered via "tc filter show" and one via a raw RTM_GETTFILTER dump:
BUG: KMSAN: kernel-infoleak in _copy_to_iter+0x1c9/0x2620
nla_put_nohdr+0x83/0x130
em_text_dump+0x291/0x550
Local variable conf created at: em_text_dump+0x5d/0x550
Bytes 168-179 of 199 are uninitialized
I am not certain whether this constitutes a real security problem in
practice: the test was conducted in a controlled KMSAN environment and
the leaked stack bytes may or may not carry sensitive data on actual
production kernels. I am reporting it because KMSAN flagged it as a
kernel-infoleak and the fix is straightforward. I can provide a
userspace reproducer on request.
The original code used strncpy() which zero-pads to the destination size.
Commit b04202d606 ("net/sched: replace strncpy with strscpy") replaced
it with strscpy(), which does not pad, creating this condition.
Zero-initialising the struct closes it.
Fixes: b04202d606 ("net/sched: replace strncpy with strscpy")
Link: https://lore.kernel.org/netdev/20250327143733.187438-1-richard120310@gmail.com/
Assisted-by: Claude:claude-sonnet-4-6 [KMSAN]
Signed-off-by: Bernard Ladenthin <bernard.ladenthin@gmail.com>
Acked-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/20260918133953.12494-1-bernard.ladenthin@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
roundup_pow_of_two() is undefined for zero. ethtool permits a zero ring
size to reach the driver, where the minimum-size check should reject it.
Leave zero unchanged while rounding nonzero ring sizes. The minimum-size
check then rejects zero deterministically without changing the established
behavior for other values.
Fixes: 6cbf18a05c ("eth: fbnic: support ring size configuration")
Reported-by: Sashiko <netdev-bot+sashiko@kernel.org>
Link: https://lore.kernel.org/netdev/178971206933.22033.236948278674126701@kernel.org/
Suggested-by: Alexander Duyck <alexanderduyck@fb.com>
Signed-off-by: Björn Töpel <bjorn@kernel.org>
Reviewed-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260918114641.1281172-1-bjorn@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Maxime Chevallier says:
====================
net: stmmac: More selftest-related fixes
This is V4 of stmmac selftest fixes, addressing Sashiko's issues over
the MTU patch. This lead to the introduction of a new one. The
dev_add_pack races have been addressed, however the double-vlan issue
stayed there. Ovidiu is actively working on it, let's wait for his work
to land before fixing that.
This is another round of stmmac selftest fixes, mostly about the selftests
themselves but a few things were discovered w.r.t MTU and buffer size
handling, see patch 5 anf 6.
After this is merged, I consider the selftests to be now reliable enough
to run them nightly on every stmmac series that's sent, and I'll be requiring
clean selftests for new glue drivers.
Since V3, the testing farm grew ! I've been running this on :
- Altera CycloneV (dwmac-socfpga, dwmac1000 IP, v3.70a)
- NXP imx8mp (dwmac-imx, dwmac4, v5.10a)
- Allwinner H2S (dwmac-sun8i, dwmac1000)
- Amlogic S905X3 (dwmac-meson8b, dwmac1000, v3.70a)
- STM32mp157a (dwmac-stm32, dwmac4, v4.20a)
- SiFive JH7110 (dwmac-starfive, dwmac4, v5.20)
- Motorcomm YT8061 (PCIe, dwmac-motorcomm, dwmac4)
- Qualcomm IPQ8064 (dwmac-ipq806x, dwmac1000)
- Altera AgileX5 (dwmac-socfpga, dwxgmac2 !) (NEW)
- Generic dwmac1000 (Loongson 2K0300, dwmac 3.70a) (NEW)
- Rockchip RK3566 (dwmac-rk, dwmac4) (NEW)
It's becoming cumbersome to list the test results here, they can be
found, updated daily, here :
https://minimaxwell.github.io/stmmac-ci/
====================
Link: https://patch.msgid.link/20260917215339.2022523-1-maxime.chevallier@bootlin.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
On dwmac1000, we currently only support single-descriptor frames. The
Jumbo test started failing when NET_IP_ALIGN was added to align the IP
header, as this tests tries to send the biggest possible frame.
On dwmac1000 the DMA transfer is aligned on 4-bytes, so adding a 2-byte
shift at the start-of-buffer address means it takes a whole extra 4-byte
DMA burst to receive the Jumbo packet, causing it to spill over the next
descriptor.
This doesn't seem to happen on dwmac4 and xgmac that appear to correctly
handle unaligned xfers (only tested on dwmac4)
Let's account for that in the Jumbo test, reduce the size of our big
packet by the align size.
Fixes: 23680bf5f8 ("net: stmmac: restore NET_IP_ALIGN in the RX DMA offset")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/20260917215339.2022523-8-maxime.chevallier@bootlin.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
When picking the buffsize to use based on the MTU, we shouldn't check
only the MTU value, but also :
- ETH_HLEN for the L2 header,
- up to 2 VLAN tags,
- the FCS,
The default bufsize is 1536 bytes, which is enough to contain all the
above so this hasn't surfaced before, but the addition of NET_IP_ALIGN
to the start of buffer address tripped the Jumbo selftest, leading to
this discovery.
With that, we don't need the '>=' checks on the buffer len, we can use
more consistent comparison operators in stmmac_set_bfsize.
Fixes: 286a837217 ("stmmac: add CHAINED descriptor mode support (V4)")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/20260917215339.2022523-7-maxime.chevallier@bootlin.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
DMA bufsize selection isn't made on the MTU but the actual frame length,
so including the L2 header. On DWMAC4, if the len is exactly BUF_SIZE_8KiB,
the next larger size is incorrectly selected.
Lets fix the comparison and while at it, rename the parameter from len
to mtu.
Fixes: c3efed5ad1 ("net: stmmac: Enable dwmac4 jumbo frame more than 8KiB").
Signed-off-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Link: https://patch.msgid.link/20260917215339.2022523-6-maxime.chevallier@bootlin.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
While we use vlan_vid_add to trigger the tag filtering machinery
in the driver, there's no netdev associated to the VLAN. This causes the
skb to arrive with empty skb->vlan_tci fields, as the packet is marked
OTHERHOST in __netif_receive_skb_core(), and we fail our validation.
Let's use the proxy mechanism introduced for DSA, that registers a
ETH_P_ALL packet handler that runs earlier, before the vlan netdev
lookup, then filters for the correct ethertype before passing an skb
clone to our validation function.
As we may receive external frames with the right tag from the outside,
let's move the address check in the vlan validation function earlier.
Fixes: 091810dbde ("net: stmmac: Introduce selftests support")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/20260917215339.2022523-5-maxime.chevallier@bootlin.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The S-TAG offload insertion incorrectly checks the dvlan (double vlan)
DMA cap, which is different than S-TAG support. Use
NETIF_F_HW_VLAN_STAG_TX to check if the feature is supported instead.
Note that this flag isn't set in stmmac yet, but contrary to ARP
offload, this is a feature that has a chance to get there eventually so
let's leave the selftest here for now. It'll report -EOPNOTSUPP in the
meantime.
Fixes: 091810dbde ("net: stmmac: Introduce selftests support")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/20260917215339.2022523-4-maxime.chevallier@bootlin.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The EEE selftest is a 2-step test :
- It validates that we enter in LPI mode with the
irq_tx_path_in_lpi_mode_n counter
- It then validates that we exit LPI when sending a frame, with the
irq_tx_path_exit_lpi_mode_n counter.
The current state of the test lacks 2 main things :
- We don't know exactly when was the previous frame sent (it's from the
previous selftest)
- The timeout is hardcoded, while the LPI is entered after a
user-configurable delay. On top of that, the timeout loop uses a
pre-decrement iterator (--retries) that actually only iterate nine
times, so 900ms while the default LPI value is 1 second.
Let's therefore make it more deterministic :
- Send a frame at the beginning of the test
- Wait for more than the lpi timer value, we timeout after about twice
the value,
- Then send another frame, and verify that we do go out of LPI, also
with a timeout.
As LPI timer can get pretty high, bail out if LPI timer is over 5
seconds.
Note that the test's goal isn't to validate the LPI timer value itself,
only that we enter/leave LPI mode.
Fixes: 091810dbde ("net: stmmac: Introduce selftests support")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/20260917215339.2022523-3-maxime.chevallier@bootlin.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Most stmmac selftests rely on dev_add_pack() to add custom handlers,
that validate the packets sent to ourselves through MAC loopback.
However, when the stmmac-driven interface is a DSA CPU conduit, all
frames that are received have ETH_P_XDSA as a protocol, even though they
don't actually contain any tag as they come from the loopback and not
the switch.
This will prevent any incoming packet to match our packet handlers.
Let's register a ETH_P_ALL packet handler when we detect that we're a
DSA conduit, and use a proxy packet handler to filter the h_proto.
As this allows external frames to be received through our .func(), the
packet handler is added after the dev->addr field is populated in our
selftest attributes.
Note that we may still receive incoming packets from the switch, but
these frames shouldn't interfere with the very specific frames used for
selftests, and stmmac selftests in general aren't safe against external
traffic interferences.
This was validated on a WPQ864 devkit for IPQ8064, that has the SoC
connected to a QCA8k switch.
The ARP offload's packet handler is left alone, this feature is just not
implemented in stmmac and due for removal.
Fixes: 091810dbde ("net: stmmac: Introduce selftests support")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/20260917215339.2022523-2-maxime.chevallier@bootlin.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Unlike bind-rx, which configures shared NIC RX queues to steer incoming
traffic into the caller's dmabuf and requires CAP_NET_ADMIN
(uns-admin-perm), bind-tx only DMA-maps the caller's dmabuf so the caller
can transmit from it on their own sockets without affecting other traffic
or device configuration.
Add a comment in netdev.yaml and above netdev_nl_bind_tx_doit() to make it
explicit that NETDEV_CMD_BIND_TX is unprivileged by design.
Signed-off-by: Mina Almasry <almasrymina@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
Acked-by: Daniel Borkmann <daniel@iogearbox.net>
Link: https://patch.msgid.link/20260921195545.493253-1-almasrymina@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Take the association or transport reference before rearming a timer in the
timer handlers.
The existing code calls mod_timer() before taking the reference needed by
the rearmed timer without holding the sock lock. This creates a race with
timer cleanup: if the timer is deleted after mod_timer() returns but before
the reference is taken, the cleanup path can drop the timer's reference and
destroy the transport or association. The timer handler then takes a
reference on the already freed object and eventually drops it, causing a
refcount underflow.
Hold the object before mod_timer() and drop the reference if mod_timer()
reports that the timer was already pending in timer handlers. Apply the
same ordering to the proto-unreachable path, which can rearm a transport
timer outside the timer handlers without holding the sock lock.
Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Reported-by: Tangxin Xie <xietangxin@h-partners.com>
Signed-off-by: Xin Long <lucien.xin@gmail.com>
Link: https://patch.msgid.link/c31b5e3ee2b7274e804f5eba2f21e2412e7eef7a.1790013825.git.lucien.xin@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Willem de Bruijn says:
====================
packet: fix PACKET_TX_RING data corruption on skb_orphan
When transmitting packets via PACKET_TX_RING, tpacket_snd links user
ring buffer pages as skb frags and releases the slot on skb->destructor
(tpacket_destruct_skb).
skb_orphan() invokes the destructor while the skb is still alive.
This marks the slot as TP_STATUS_AVAILABLE prematurely, allowing
userspace to overwrite the slot and causing data corruption.
This series fixes the issue by switching PACKET_TX_RING to standard
ubuf_info zerocopy completion, ensuring ring slots are released only
after all payload references are freed or copied.
Virtio-net needs a separate solution, because deferring the release
can cause deadlock in its !use_napi mode.
- Patch 1 addresses the virtio-net special case.
- Patch 2 converts tpacket_snd to standard ubuf_info completion
Patch 1 must be applied, and backported, before patch 2. Both carry
the same Fixes tag for that reason.
v1: https://lore.kernel.org/netdev/20260914214229.1674102-1-willemdebruijn.kernel@gmail.com/
====================
Link: https://patch.msgid.link/20260919004748.1463985-1-willemdebruijn.kernel@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
tpacket_snd sends skbs with frags pointing into its ring slots. Slots
are released when skb->destructor is called.
A call to skb_orphan calls skb->destructor before the skb is freed.
This can cause the slot to be reused while still linked into the skb.
Switch to standard zerocopy completion (ubuf_info) so the slot is only
released once all references to the payload are freed or copied.
Restore skb->destructor to standard sock_wfree.
The ubuf_info completion callback can be called with a NULL skb, but
only from net_zcopy_put and related API, used by zerocopy implementations
that hold their own reference on the uarg, such as MSG_ZEROCOPY. This
uarg is only ever completed from skb_zcopy_clear, so skb is always set.
To prevent userspace from aliasing in-flight state on shared ring
slots, allocate tpacket_uarg per packet, rather than per slot. This
adds a small allocation to the transmit path. Use standard kmalloc to
allow backporting to stable kernels.
The uarg holds an sk_wmem_alloc reference, rather than an sk_refcnt
reference. packet_free_tx_ring waits on sk_wmem_alloc before freeing
the ring pages. Always allocate vec->deferred for tx_ring so page-backed
rings also wait on sk_wmem_alloc when skb_copy_ubufs drops page refs
before calling tpacket_ubuf_complete.
Drop the tx_ring.pg_vec test that tpacket_destruct_skb performed before
accessing the slot. The sk_wmem_alloc reference now guarantees that the
slot is valid. The test is also not sufficient by itself, as it reads
pg_vec without pg_vec_lock, so it can race with packet_set_ring.
As a result a slot is released when its payload is copied, which can
be before transmission (e.g., in skb_orphan_frags_rx). If copied
before skb_tx_timestamp() is called, no slot timestamp is recorded,
similar to when skb_orphan() was called early in the datapath before
this patch.
Revert the now unused previous skb_zcopy_.._nouarg infra.
Depends on commit 992cc9f94c ("net/packet: defer vmalloc TX_RING
free until skbs finish").
Reported-by: Katherine Leaver <kleaver@janestreet.com>
Reported-by: Bjoern Doebel <doebel@amazon.de>
Closes: https://lore.kernel.org/netdev/20260909085542.3370986-1-doebel@amazon.de/
Fixes: 5cd8d46ea1 ("packet: copy user buffers before orphan or clone")
Cc: stable@vger.kernel.org
Signed-off-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260919004748.1463985-3-willemdebruijn.kernel@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Virtio-net without NAPI frees completed skbs lazily on the next
start_xmit. Senders waiting for in-flight zerocopy buffers can
deadlock if they cannot transmit more packets, as then no
completed packets will be freed.
When !use_napi, virtio-net already calls skb_orphan to avoid waiting
up for transmitted skbs to be freed. For zerocopy packets that
require deep copying on orphan (i.e. those that do not set
SKBFL_DONT_ORPHAN, such as PACKET_TX_RING), call skb_orphan_frags
before orphaning to release the buffers.
This fixes the tpacket_snd slot reuse bug on skb_orphan for
virtio-net, and prevents PACKET_TX_RING from running out of slots.
This fix also touches vhost_net zerocopy packets, which also do not
set SKBFL_DONT_ORPHAN. This is fine: vhost_net packets only encounter
virtio-net in nested virtualization, and only if napi_tx is
explicitly disabled (it has been default-enabled since Linux 4.12).
In that rare case, copying the frags is desirable anyway to prevent
holding guest descriptors pinned across unbounded intervals.
This is a prerequisite for the next patch, which converts
PACKET_TX_RING to standard zerocopy completion. Without this patch
first, a bounded ring sender can stall indefinitely behind a
virtio-net virtqueue that cannot reclaim.
Fixes: 5cd8d46ea1 ("packet: copy user buffers before orphan or clone")
Cc: stable@vger.kernel.org
Cc: mst@redhat.com
Cc: jasowangio@gmail.com
Signed-off-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260919004748.1463985-2-willemdebruijn.kernel@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
When an association is in COOKIE-ECHOED state and the peer sends a
bundled [ERROR(Stale Cookie)][DATA] packet from one of its non-primary
addresses, processing the ERROR chunk takes the non-fatal stale-cookie
retry path sctp_sf_do_5_2_6_stale(), which queues
SCTP_CMD_DEL_NON_PRIMARY while keeping the association alive.
sctp_cmd_del_non_primary() removes every non-primary transport -
including the very transport this packet arrived on, which is still
referenced by the receive lookup and shared by all chunks of the
packet via chunk->transport.
sctp_assoc_rm_peer() does redirect asoc->peer.last_data_from away from
the removed transport, but right afterwards the bundled DATA chunk
makes sctp_assoc_bh_rcv() re-register
asoc->peer.last_data_from = chunk->transport unconditionally, undoing
the redirection with the just-removed transport.
Once the packet is done, the receive reference is dropped and the
transport is RCU-freed, while the surviving association keeps the
dangling last_data_from. A later FWD-TSN (or the delayed SACK timer)
makes sctp_gen_sack() dereference it (->param_flags and friends), and
sctp_make_sack()/sctp_outq_select_transport() may write to the freed
object and link it into the live transport list. This is a
use-after-free triggerable by any malicious SCTP peer (or a local
unprivileged user acting as one) with no capabilities required:
BUG: KASAN: slab-use-after-free in sctp_do_sm+0x498a/0x5660
Read of size 4 at addr ffff88800e1e356c by task poc/115
Call Trace: sctp_do_sm <- sctp_assoc_bh_rcv <- sctp_inq_push <-
sctp_rcv <- ip_protocol_deliver_rcu <- ip_rcv
Allocated: sctp_transport_new <- sctp_assoc_add_peer <-
sctp_process_init (INIT-ACK processing)
Freed: kfree <- sctp_transport_destroy_rcu <- rcu_core
(call_rcu queued by sctp_transport_put at end of sctp_rcv)
The buggy address is located 364 bytes inside of freed 1024-byte
region [ffff88800e1e3400, ffff88800e1e3800), cache kmalloc-1k
Note that commit 03a9d10ecf ("sctp: drop a chunk if its transport
was removed") only covers the window between the receive lookup and
the chunk processing (e.g. an ASCONF DEL-IP racing the socket backlog);
here the transport is removed *while* the packet is being processed,
by an earlier chunk of the same packet, so the drop in sctp_inq_push()
does not reach this path. Verified with the bundled [ERROR(Stale
Cookie)][DATA] + FWD-TSN reproducer: the KASAN report above still
fires with that commit applied, and is gone with this patch on top.
Fix it by discarding the rest of the packet on this path, as suggested
by Xin. After the stale-cookie ERROR has sent the association back to
COOKIE-WAIT and removed the non-primary transports, the remaining
chunks of the packet can only run against the restarted handshake
while referencing the removed arrival transport through
chunk->transport: besides the last_data_from registration above,
sctp_cmd_setup_t2() and the sctp_make_*() reply builders would also
copy that pointer into association-lifetime state that
sctp_assoc_rm_peer() has already sanitized. Let the peer retransmit
them, in line with what sctp_inq_push() does for chunks whose
transport was removed before processing.
Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Suggested-by: Xin Long <lucien.xin@gmail.com>
Reported-by: TencentOS Corvus AI <corvus@tencent.com>
Cc: stable@vger.kernel.org
Signed-off-by: Aohan Mei <henrymei@tencent.com>
Acked-by: Xin Long <lucien.xin@gmail.com>
Link: https://patch.msgid.link/20260921093707.1432184-1-ljp1205831794@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
When the reverse entry is found but its counter is already being
released, refcount_inc_not_zero() fails and the reference taken by
mlx5_tc_ct_entry_get() is never dropped before falling through to
create_counter. Drop it so the reverse entry is not kept alive forever
by a shared counter lookup that did not use it.
Fixes: 1edae2335a ("net/mlx5e: CT: Use the same counter for both directions")
Cc: stable@vger.kernel.org
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Reviewed-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260917113131.2149024-1-vulab@iscas.ac.cn
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
mlx5_esw_bridge_vport_unlink() returns -EINVAL when the port isn't
tracked by this instance's br_offloads. This is reachable on a sibling
instance that registered its notifier after the port was already
enslaved: it never saw the NETDEV_CHANGEUPPER link event, so
peer_link() never created a peer port for it, but it does see the
later unlink event and fails. Return 0 instead, and give
mlx5_esw_bridge_vport_peer_unlink() the same merged_eswitch capability
guard peer_link() already has, since without it peer_link() likewise
never creates a port to unlink.
This also matters beyond the -EINVAL itself:
mlx5_esw_bridge_switchdev_port_event() runs on the per-netns
netdev_chain, and notifier_from_errno(-EINVAL) sets NOTIFY_STOP_MASK,
which call_netdevice_notifiers_info() checks to stop calling further
listeners on that chain - so the old -EINVAL silently dropped the
event for any listener registered later on the same chain, even
though none of it was visible to user space since
__netdev_upper_dev_unlink() discards the return value.
Fixes: c358ea1741 ("net/mlx5: Bridge, allow merged eswitch connectivity")
Signed-off-by: Bernardo Soares <bsoares.it@gmail.com>
Reviewed-by: Mark Bloch <mbloch@nvidia.com>
Link: https://patch.msgid.link/20260918095931.29792-3-bsoares.it@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
mlx5 registers the bridge offload switchdev notifiers once per eswitch
instance, but the notifier chains are global, so every instance sees
every event and must filter out the ones that aren't its own. The
existing filter, mlx5_esw_bridge_dev_same_hw(), only checks that the
event netdevice sits on the same HCA - intentional for merged eswitch,
where one bridge can span representors of several eswitches on one
HCA - but same-HCA doesn't mean the instance actually has that port:
peer ports are only created reactively from NETDEV_CHANGEUPPER, so an
instance brought up after a sibling PF's port was already enslaved has
none. The port object and attribute handlers claim the event anyway
once same-HW passes, then fail the port lookup and return -EINVAL,
which gets reported to user space even though the owning instance
already handled it (e.g. "bridge vlan add ... RTNETLINK answers:
Invalid argument"). Fix by filtering on the tracked port instead.
The same gap exists in the generic recursive lower-device walk used by
attribute changes on a bridge with more than one representor enslaved
directly: mlx5_esw_bridge_lower_rep_vport_num_vhca_id_get() is entered
with the bridge master netdevice, falls through to its generic
netdev_for_each_lower_dev() loop, and returns as soon as the recursion
into any one lower device yields a non-NULL rep - the underlying base
case, mlx5_esw_bridge_rep_vport_num_vhca_id_get(), only checks
mlx5_esw_bridge_dev_same_hw(), not ownership by the calling instance's
br_offloads. mlx5_esw_bridge_lag_rep_get(), used for the LAG-master
case, already filters on mlx5_esw_bridge_dev_same_esw() per candidate
and so cannot select a sibling's rep; it is not the source of this bug.
On a merged-eswitch HCA with a bridge spanning representors of more
than one eswitch instance directly, the walk can return a sibling's rep
instead of continuing to the one the calling instance actually owns, so
the attribute change fails the same way as above. Fix by checking
mlx5_esw_bridge_port_exists() at the point each rep is picked, same as
the previous fix did for the notifier filter.
Fixes: c358ea1741 ("net/mlx5: Bridge, allow merged eswitch connectivity")
Signed-off-by: Bernardo Soares <bsoares.it@gmail.com>
Cc: Vlad Buslov <vladbu@nvidia.com>
Cc: Saeed Mahameed <saeedm@nvidia.com>
Reviewed-by: Mark Bloch <mbloch@nvidia.com>
Link: https://patch.msgid.link/20260918095931.29792-2-bsoares.it@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
usb_get_from_anchor() hands over a reference to the URB, which the caller
must release. lan78xx_submit_deferred_urbs() never does, so every deferred
Tx URB keeps an extra reference: the counter grows on each suspend/resume
cycle and the URBs are never freed when the buffers are released. Drop
the reference after submitting, and on the path that drops the packet
instead of submitting it.
Fixes: 5f4cc6e251 ("lan78xx: Fix race conditions in suspend/resume handling")
Cc: stable@vger.kernel.org
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Link: https://patch.msgid.link/20260917115811.2150119-1-vulab@iscas.ac.cn
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
xpcs_init_clks() takes references with clk_bulk_get_optional() and then
enables them with clk_bulk_prepare_enable(). If the enable step fails,
the function returns without dropping the references.
xpcs_create() handles the failure through out_free_data, which calls
xpcs_free_data() but never xpcs_clear_clks(), so the clk references are
leaked.
Add the missing clk_bulk_put() on the enable failure path. The
prepare/enable side is already rolled back by
clk_bulk_prepare_enable() itself.
Fixes: f6bb3e9d98 ("net: pcs: xpcs: Add Synopsys DW xPCS platform device driver")
Signed-off-by: Coia Prant <coiaprant@gmail.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260919172021.2336748-1-coiaprant@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
The op-to-policy map a CTRL_CMD_GETPOLICY dump returns is the only way
for userspace to find out which policy index belongs to which command.
ctrl_dumppolicy_put_op() tags the nest with doit->cmd, but an op which
only has a dumpit has no doit and every path which fills the split ops
in zeroes it out, so those entries all claim to be command 0. nlctrl's
own CTRL_CMD_GETPOLICY and NETDEV_CMD_QSTATS_GET are both in that group:
[{'family-id': 16, 'op-policy': {'do': 0, 'dump': 0, 'op-id': 3}},
{'family-id': 16, 'op-policy': {'dump': 1, 'op-id': 0}},
ctrl_fill_info() gets this right - it uses the iterator's cmd for
CTRL_ATTR_OP_ID - so the two introspection interfaces of the same family
contradict each other today.
Pass the command in rather than reconstructing it from
doit->cmd | dumpit->cmd inside the helper, both callers already have it.
Fixes: 26588edbef ("genetlink: support split policies in ctrl_dumppolicy_put_op()")
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Link: https://patch.msgid.link/20260918222949.4190284-1-kuba@kernel.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Fix 3 leaks in macb_alloc() error paths:
- Tx buffer allocated but crossing a 4G boundary: Tx leaked.
- Rx buffer allocation fails: Tx leaked.
- Rx buffer allocated but crossing a 4G boundary: Tx & Rx leaked.
This is because our error handling calls macb_free(bp) which in turn
frees the buffers stored in bp->queues[0], but nothing has been stored
in there. Fix by storing allocated buffers into bp->queues[0] ASAP.
Fixes: 78d901897b ("net: macb: single dma_alloc_coherent() for DMA descriptors")
Cc: stable@vger.kernel.org
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Link: https://patch.msgid.link/20260918-macb-alloc-leak-v1-1-aba9a3d4f6e3@bootlin.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
hns_dsaf_find_platform_device() returns the mdio platform device with its
reference count incremented. hns_mac_register_phy() never drops that
reference, so the mdio device can not be released.
Release the reference on both the deferred probe and the normal path.
Fixes: 1d1afa2ebf ("net: hns: register phy device in each mac initial sequence")
Cc: stable@vger.kernel.org
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260917110828.2148390-1-vulab@iscas.ac.cn
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
pm_runtime_get_sync() leaves the runtime PM usage counter incremented even
when it fails, but the error path in netcp_probe() does not call
pm_runtime_put_noidle() to balance it, leaking a reference each time
resume fails.
Use pm_runtime_resume_and_get() instead, which automatically drops the
usage counter on failure, fixing the leak.
Fixes: 84640e27f2 ("net: netcp: Add Keystone NetCP core ethernet driver")
Signed-off-by: bui duc phuc <phucduc.bui@gmail.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260918042804.13101-1-phucduc.bui@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
emac_tx_mem_map() writes TX_DESC_0_OWN into the ring descriptor for
every slot beyond old_head as soon as that slot's memset()'d local
copy is committed with "*tx_desc_addr = tx_desc", i.e. before the
buffers for that slot have necessarily all been mapped successfully.
If emac_tx_map_frag() then fails on a later fragment, the err_free_skb
path calls emac_free_tx_buf() to unmap and drop the skb, but leaves
the already-written descriptor memory untouched, and tx_ring->head is
never advanced past old_head (the "tx_ring->head = head" store is
skipped by the goto).
So a slot between old_head and the rolled-back head can be left with
TX_DESC_0_OWN set and buffer_addr_{1,2} pointing at DMA mappings that
emac_free_tx_buf() just tore down, while software considers that slot
free again. The next successful emac_tx_mem_map() call only rebuilds
old_head itself; if the DMA engine auto-advances into the following
descriptor once it finishes old_head's packet, it will fetch that
stale, already-unmapped address.
emac_tx_clean_desc() already treats emac_free_tx_buf() and clearing
the descriptor as a pair when reclaiming completed descriptors; do
the same in the mapping failure path.
Fixes: bfec6d7f20 ("net: spacemit: Add K1 Ethernet MAC")
Signed-off-by: Muhammad Bilal <meatuni001@gmail.com>
Reviewed-by: Vivian Wang <wangruikang@iscas.ac.cn>
Reviewed-by: Troy Mitchell <troy.mitchell@linux.spacemit.com>
Link: https://patch.msgid.link/20260919191937.271202-1-meatuni001@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
of_clk_get() returns a clock with its reference count incremented, but
read_dts_node() only uses it to read the rate and never calls clk_put().
The clock is not stored anywhere, so the reference cannot be released
later either.
Release the clock once its rate has been read, which also covers the
error path taken when the rate is zero.
Fixes: 414fd46e77 ("fsl/fman: Add FMan support")
Cc: stable@vger.kernel.org
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260917110135.2148068-1-vulab@iscas.ac.cn
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
tg3_get_invariants() can register an MDIO bus and connect a PHY for
USE_PHYLIB devices. If tg3_init_one() later fails, its common error path
releases the mappings and netdev without undoing those PHYLIB resources.
Disconnect the PHY and unregister the MDIO bus before the remaining
teardown. Guard PHY cleanup with USE_PHYLIB to match tg3_phy_init(), and
call tg3_mdio_fini() unconditionally to match tg3_mdio_init(). The existing
IS_CONNECTED and MDIOBUS_INITED flags make both helpers safe when
initialization only completed partially.
This issue was identified during our ongoing static-analysis research while
reviewing kernel code.
Fixes: 158d7abdae ("tg3: Add mdio bus registration")
Assisted-by: OpenAI:GPT-5.6
Co-developed-by: Ijae Kim <ae878000@gmail.com>
Signed-off-by: Ijae Kim <ae878000@gmail.com>
Signed-off-by: Myeonghun Pak <mhun512@gmail.com>
Link: https://patch.msgid.link/20260917183336.36239-1-mhun512@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Shardul Bankar says:
====================
udp: two fixes for the 4-tuple hash table
Two ways a UDP socket ends up in the wrong place in the 4-tuple hash table.
The patches are independent, with different Fixes: tags and no dependency
between them.
Patch 1: a socket that connects a second time is not relocated, so it stays
filed under its first peer's hash and packets for it fall back to scoring
the hash2 chain for its address and port.
Patch 2: a socket bound to a specific address and port is not taken out of
the table when it disconnects, because __udp_disconnect() only does that
via ->rehash() or ->unhash() and neither runs for it.
Both are in code shared by IPv4 and IPv6.
Patch 1's cost, with N sockets sharing a port and one of them misfiled,
200k packets sent to its 4-tuple:
N without with
200 1,061,652 2,093,259 pps
500 522,553 2,055,078 pps
1000 279,729 2,136,606 pps
Correctly filed sockets measure ~2.1M pps throughout, so the cost scales
with the number of sockets on the port, as the fallback scan does. For
comparison, commit 78c91ae2c6 ("ipv4/udp: Add 4-tuple hash for connected
socket") measured 290,860 pps without the table and 1,889,658 with it at
500 connected sockets.
Patch 2's cost is not in throughput. Its stale entry keeps hash4_cnt raised
for the life of the socket, so every packet for that address and port is
sent through the 4-tuple lookup first; on IPv6 the entry is also matchable,
because __udp_disconnect() does not clear sk_v6_daddr. That last one is a
separate defect, which I will send on its own.
Neither patch has a selftest. Nothing in tree reports which 4-tuple bucket
a socket is filed under, so a test can only measure the cost indirectly.
What I did instead was add pr_info() to the hash4 paths and a knob that
dumps bucket occupancy, then run the same scenarios on two kernels
differing only by these patches; that is where the numbers above come from.
The instrumentation, the reproducers and the benchmark are at [1]. If
exposing the bucket through diag would be welcome, that would make both
defects testable in tree and I am glad to do it for net-next.
Removing the connect(AF_UNSPEC) limitation described in 644f9108f3 is a
side effect of patch 1 fixing the general case. I can make it narrower if
you would rather that limitation stayed.
Tooling, per Documentation/process/generated-content.rst: this series was
developed in an assisted session with an LLM. The assistant did most of
the code reading, wrote the instrumentation and reproducers behind [1],
drafted these changelogs, and ran the A/B builds and the regression
suites below. Every claim in these messages was checked against the
source, and the IPv6 behaviour described in patch 2 was confirmed at
runtime.
Tested on x86-64, IPv4 and IPv6. No regressions across reuseport_bpf,
reuseport_bpf_cpu, reuseport_addr_any.sh, reuseport_dualstack,
udpgso_bench.sh, udpgro_bench.sh and socket.
[1] https://github.com/shardulsdk-mpiric/linux/tree/udp-hash4-fix-verification
To: Willem de Bruijn <willemdebruijn.kernel@gmail.com>
To: "David S. Miller" <davem@davemloft.net>
To: Eric Dumazet <edumazet@google.com>
To: Jakub Kicinski <kuba@kernel.org>
To: Paolo Abeni <pabeni@redhat.com>
To: Simon Horman <horms@kernel.org>
To: Philo Lu <lulie@linux.alibaba.com>
To: Fred Chen <fred.cc@alibaba-inc.com>
To: Yubing Qiu <yubing.qiuyubing@alibaba-inc.com>
Cc: Kuniyuki Iwashima <kuniyu@google.com>
Cc: Willem de Bruijn <willemb@google.com>
Cc: Cambda Zhu <cambda@linux.alibaba.com>
Cc: Janak Bhatt <janak@mpiric.us>
Cc: Kalpan Jani <kalpan.jani@mpiricsoftware.com>
Cc: Shardul Bankar <shardulsb08@gmail.com>
Cc: netdev@vger.kernel.org
Cc: linux-kernel@vger.kernel.org
Signed-off-by: Shardul Bankar <shardul.b@mpiricsoftware.com>
====================
Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-0-718891af0d7a@mpiricsoftware.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
A UDP socket bound to a specific address and port keeps its entry in the
4-tuple hash table after it is disconnected:
sk binds to 127.0.0.1:21001
sk connects to 127.0.0.2:20001 // filed in the 4-tuple table
sk disconnects, connect(AF_UNSPEC) // still filed, peer now 0.0.0.0:0
__udp_disconnect() takes a socket out of that table only as a side effect
of ->rehash() or ->unhash(), and it skips ->rehash() when
SOCK_BINDADDR_LOCK is set and ->unhash() when SOCK_BINDPORT_LOCK is set.
commit 6996a2d2d0 ("udp: Unhash auto-bound connected sk from 4-tuple hash
table when disconnected.") fixed the same end state for a wildcard-bound
socket, by a path this one does not take.
The entry is counted whether or not anything hits it. hash4_cnt on the
hash2 slot stays raised for as long as the socket lives, so udp_has_hash4()
keeps sending every packet for that address and port through the 4-tuple
lookup first.
On IPv6 it can also be hit. __udp_disconnect() does not clear sk_v6_daddr,
so udp_v6_rehash() files the entry under the peer the socket was connected
to with a zero dport, and inet6_match() compares that same
field: a datagram from the former peer with a zero source port matches,
and source port zero is accepted on receive. On IPv4 the peer is cleared,
so a match would need a zero source address as well, which the routing
layer rejects as martian. The stale sk_v6_daddr is a separate defect, not
addressed here; removing the entry closes this path either way.
The entry can also be relocated. __udp_disconnect() clears sk_bound_dev_if,
so a subsequent SO_BINDTODEVICE calls ->rehash(), and because the receive
address is still specific udp_lib_rehash() moves the entry instead of
removing it, into the bucket that (rcv_saddr, num, 0, 0) hashes to -- a
pure function of the address and port, so every socket reaching this state
on one address and port collects in one bucket. The bucket cannot be chosen
from outside, as udp_ehashfn() is seeded with a per-boot secret. This last
one became reachable only with commit 644f9108f3 ("udp: Make rehash4
independent in udp_lib_rehash()"), which moved the hash4 handling out of a
branch a disconnected socket does not take; the stale entry itself dates
from the commit in Fixes.
Take the socket out of the table before __udp_disconnect() runs, while it
still matches how it was filed. This also reaches the wildcard case ahead
of udp_lib_rehash()'s udp_unhash4() branch, leaving that branch unreachable
from udp_disconnect(); removing it belongs in net-next. udp_disconnect()
and udp_abort() are the only UDP entries into __udp_disconnect(), which is
shared with raw, ping and l2tp sockets that are not struct udp_sock:
ping_prot.obj_size is sizeof(struct inet_sock), so udp_hashed4() on one
would read past the allocation.
Fixes: 78c91ae2c6 ("ipv4/udp: Add 4-tuple hash for connected socket")
Assisted-by: LLM
Signed-off-by: Shardul Bankar <shardul.b@mpiricsoftware.com>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-2-718891af0d7a@mpiricsoftware.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
A connected UDP socket that connects again to a different peer is not
re-filed in the 4-tuple hash table:
sk binds to 127.0.0.1:21001
sk connects to 127.0.0.2:20001 // filed under hash(sk, peer1)
sk connects to 127.0.0.3:20002 // still filed under hash(sk, peer1)
packet from 127.0.0.3:20002 // hash(sk, peer2) misses, so the
// lookup falls back to scoring the
// hash2 chain for this address
// and port
udp_lib_hash4() returns early when the socket is already hashed, assuming
->rehash() relocates it. ->rehash() runs from __ip{4,6}_datagram_connect()
only while the receive address is unset, which a second connect never is:
the first connect assigns it, whether the socket was bound to a specific
address or to the wildcard. commit 644f9108f3 ("udp: Make rehash4
independent in udp_lib_rehash()") added that early return and named
connect(AF_UNSPEC) as the way around it. That workaround does not help a
socket with both SOCK_BINDADDR_LOCK and SOCK_BINDPORT_LOCK set, because
__udp_disconnect() skips ->rehash() for the first and ->unhash() for the
second.
Delivery is correct either way.
Relocate the socket when the hash it is filed under differs from the one
requested, which is what commit 78c91ae2c6 ("ipv4/udp: Add 4-tuple hash
for connected socket") did before the early return became unconditional. It
is done here under hslot->lock, which that version did not take, to match
udp_lib_rehash() and udp_lib_unhash(). hslot2 is unchanged, so hash4_cnt
needs no adjustment, as in udp_lib_rehash(). A first connect is unaffected,
and IPv6 shares the code.
With 500 sockets on the port, a re-connected socket measured 522,553 pps
without this change and 2,055,078 with it. The UDP side was noted as
remaining work in [1].
Link: https://lore.kernel.org/netdev/apnHqmYZQ4yzOP4N@v4bel/ [1]
Fixes: 644f9108f3 ("udp: Make rehash4 independent in udp_lib_rehash()")
Assisted-by: LLM
Signed-off-by: Shardul Bankar <shardul.b@mpiricsoftware.com>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-1-718891af0d7a@mpiricsoftware.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Core:
- hci_conn: fix CIS hold ownership on reuse
- hci_sock: reject out-of-range OCF values
- hci_sock: validate event length before filtering
- L2CAP: validate frame length before control and FCS access
- RFCOMM: fix NULL dereference of dlc->session in RFCOMM_CONNINFO
- RFCOMM: Reject short EA=0 frames in rfcomm_recv_frame()
- ISO: release unused CIS holds after channel attach
- ISO: balance the parent hold in hci_bind_bis()
- SMP: reject Security Request over BR/EDR
- MGMT: fix race in read_unconf_index_list()
- MGMT: Dequeue pending mesh_send_sync entries on cancel
- BNEP: fix out-of-bounds reads on short RX/TX frames and control fallthrough
Drivers:
- btintel_pcie: validate device-supplied DMA indices
- btnxpuart: Fix skb leak in nxp_process_fw_dump()
-----BEGIN PGP SIGNATURE-----
iQJNBAABCgA3FiEE7E6oRXp8w05ovYr/9JCA4xAyCykFAmqxN44ZHGx1aXoudm9u
LmRlbnR6QGludGVsLmNvbQAKCRD0kIDjEDILKS7mD/43nt9IlCMp4fRi5eT3iV6J
phP/zJTiikgOMv87kTI0Q9OXY8Xl1nhIGrTiypXQIJNJGTW/OTHtMF+N55pA7pF4
exD0US01bKUcuopztHeP1Yk08CMKAtE9VTPLi/PdAQz6KTo7wcH7yFwfsO6avIk4
T/IXi9b8iKPBMKizjgUe8Uo6wrFoByD/o2VwTSQwtZOgWppUkCvVKKx10Y8EYJRG
UbDS8rpPtWN3t0fK5F4yVfkTP+9sA7uWb9bF2UoLcmov2ie33M50zf5giVofiNrX
pT6RWTCPfjkPdDro66l6wV4M8prEUVok0QhlqeYOUXSZJSOTqqXBTOorioABSeHz
i4QWH+KsqUFbaOJQuoF3v5jAccY8ZxKczRODK8Xt1+bY9O73lqpuWjQvoi98E/7i
vNCSU/bDpRS0MjYDTgQTG9c0rv9qxKI0GBC8gDVI78/nAzUIxdJ09/wDBm17aUpt
uxaBwYE1lhJwT0Lf+dwJrwVuTm5q5XotdRIsOwxZgfF25Ewk8ygUlEsQu/ioHmS+
q+9SL/mHb/wPI7j0YeVG2l7cHiHOZDEn0FW++WlgtF1Oj/HWUsz2UN/qdmpYZo1h
rGX7D2DdZQo8IOEmZZn7wqLNOxjPjFVkxgQtgWd5OhE0O+7/e2H4mNeDJyWs6rJp
lT1BJZnPQTFtwP41x7Hxng==
=cLTc
-----END PGP SIGNATURE-----
Merge tag 'for-net-2026-09-21' of git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth
Luiz Augusto von Dentz says:
====================
bluetooth pull request for net:
Core:
- hci_conn: fix CIS hold ownership on reuse
- hci_sock: reject out-of-range OCF values
- hci_sock: validate event length before filtering
- L2CAP: validate frame length before control and FCS access
- RFCOMM: fix NULL dereference of dlc->session in RFCOMM_CONNINFO
- RFCOMM: Reject short EA=0 frames in rfcomm_recv_frame()
- ISO: release unused CIS holds after channel attach
- ISO: balance the parent hold in hci_bind_bis()
- SMP: reject Security Request over BR/EDR
- MGMT: fix race in read_unconf_index_list()
- MGMT: Dequeue pending mesh_send_sync entries on cancel
- BNEP: fix out-of-bounds reads on short RX/TX frames and control fallthrough
Drivers:
- btintel_pcie: validate device-supplied DMA indices
- btnxpuart: Fix skb leak in nxp_process_fw_dump()
* tag 'for-net-2026-09-21' of git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth:
Bluetooth: RFCOMM: Reject short EA=0 frames in rfcomm_recv_frame()
Bluetooth: RFCOMM: fix NULL dereference of dlc->session in RFCOMM_CONNINFO
Bluetooth: btintel_pcie: validate device-supplied DMA indices
Bluetooth: bnep: fix out-of-bounds reads on short RX/TX frames and control fallthrough
Bluetooth: mgmt: fix race in read_unconf_index_list()
Bluetooth: L2CAP: validate frame length before control and FCS access
Bluetooth: ISO: balance the parent hold in hci_bind_bis()
Bluetooth: hci_sock: validate event length before filtering
Bluetooth: hci_sock: reject out-of-range OCF values
Bluetooth: ISO: release unused CIS holds after channel attach
Bluetooth: hci_conn: fix CIS hold ownership on reuse
Bluetooth: mgmt: Dequeue pending mesh_send_sync entries on cancel
Bluetooth: btnxpuart: Fix skb leak in nxp_process_fw_dump()
Bluetooth: SMP: reject Security Request over BR/EDR
====================
Link: https://patch.msgid.link/20260921135807.3459373-1-luiz.dentz@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
ip6_route_output_flags(), ip6_rt_put_flags(), and ip6_dst_check()
detect an uncached route by list_empty(&rt->dst.rt_uncached),
which replaced the static DST_NOCACHE flag check in commit
a4c2fd7f78 ("net: remove DST_NOCACHE flag").
When a device is unregistered, rt6_uncached_list_flush_dev()
unlinks uncached routes tied to the device from rt6_uncached_list.
Previously, they were moved to another list with list_move()
(__list_del_entry() + list_add()), and since commit 98aa546af5
("inet: remove (struct uncached_list)->quarantine"), the routes
are just unlinked with list_del_init().
If list_del_init() runs concurrently, list_empty() evaluates to
true; ip6_route_output_flags() calls dst_hold_safe() incorrectly
and ip6_rt_put_flags() skips ip6_rt_put(), leaking dst, and thus
dev tied via rt->from as well.
The same race is partially fixed by commit 9a6f0c4d57 ("dst:
fix races in rt6_uncached_list_del() and rt_del_uncached_list()").
Let's check rt6->dst.rt_uncached_list instead.
Note that IPv4 does not have the same issue.
Fixes: 98aa546af5 ("inet: remove (struct uncached_list)->quarantine")
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Reviewed-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260920191558.2990636-1-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
IPv6 XFRM policies may use different source and destination prefix
lengths. mlx5e_ipsec_policy_mask() builds the corresponding masks
independently, but setup_fte_addr6() installs each mask in the opposite
address field.
When the prefix lengths differ, this makes the source match use the
destination prefix and the destination match use the source prefix. The
resulting hardware rule can both miss traffic covered by the policy and
match traffic outside it.
Install each mask in its corresponding match field.
Fixes: ca7992f52c ("net/mlx5e: Properly match IPsec subnet addresses")
Cc: stable@vger.kernel.org
Signed-off-by: Andrea Parri <parri.andrea@gmail.com>
Reviewed-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260917115542.177675-1-parri.andrea@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Commit 8d75c338f0 ("sysctl: remove CONFIG_PROC_SYSCTL, it just mirrors
CONFIG_SYSCTL") renamed CONFIG_PROC_SYSCTL to CONFIG_SYSCTL in place,
which left the entry out of alphabetical order in the net and
packetdrill configs. The netdev CI check for sorted selftest configs
now fails for every patch that touches either file.
Fixes: 8d75c338f0 ("sysctl: remove CONFIG_PROC_SYSCTL, it just mirrors CONFIG_SYSCTL")
Signed-off-by: Yuya Kusakabe <yuya.kusakabe@gmail.com>
Reviewed-by: Joel Granados <joel.granados@kernel.org>
Link: https://patch.msgid.link/20260918-selftests-net-config-sort-v1-1-968ea6e8c1b7@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The receive fixup subtracts the Ethernet CRC from the reported packet
length, but compares that payload length against the whole remaining
receive buffer. The following copy starts after the three-byte header,
and the cursor advance consumes both that header and the four-byte CRC.
Require the payload to fit after SR_RX_OVERHEAD before copying it or
advancing to the next packet. The loop already ensures that the
remaining buffer is larger than the overhead, so the subtraction is
safe.
The issue was found by our static-analysis tool.
Fixes: c9b37458e9 ("USB2NET : SR9700 : One chip USB 1.1 USB2NET SR9700Device Driver Support")
Reviewed-by: Ethan Nelson-Moore <enelsonmoore@gmail.com>
Tested-by: Ethan Nelson-Moore <enelsonmoore@gmail.com>
Signed-off-by: Pengpeng Hou <hppiscas@163.com>
Link: https://patch.msgid.link/20260920034745.18468-1-hppiscas@163.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
dpll_pin_ref_sync_state_set() looks up the reference sync pin in the
pin->ref_sync_pins xarray, which is keyed by the sync pin's id (see
dpll_pin_ref_sync_pair_add() using xa_insert() with ref_sync_pin->id).
The pin id to operate on is supplied by userspace via DPLL_A_PIN_ID.
The lookup however used xa_find() with a ULONG_MAX limit, which returns
the first present entry with an index greater than or equal to the
requested id, not the entry stored exactly at that id. If userspace
passes an id that is not paired as a reference sync pin, but another
pin with a higher id is present in the xarray, xa_find() silently
returns that wrong pin and the subsequent ref_sync_set() operates on
it. The request only fails when the given id is larger than every
present key.
Use xa_load() for an exact-key lookup instead, mirroring the deletion
path in dpll_pin_ref_sync_pair_del().
Fixes: 58256a26bf ("dpll: add reference sync get/set")
Signed-off-by: Ivan Vecera <ivecera@redhat.com>
Link: https://patch.msgid.link/20260917143736.526221-1-ivecera@redhat.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Removing the legacy ioctl fallback made both hwtstamp NDOs mandatory. A
device that only timestamps in its PHY implements neither, so
SIOCSHWTSTAMP fails with EOPNOTSUPP before anything looks at the PHY and
PTP stops working there.
The check only ever picked the legacy path. That path is gone, so drop it
and test where the NDOs are actually called.
SIOCGHWTSTAMP is new here, not restored. The old path went through
phy_mii_ioctl(), which only handled SIOCSHWTSTAMP.
Such a device now returns -ENODEV while absent instead of -EOPNOTSUPP,
like the ones that do implement the NDOs.
Fixes: 5062245a5a ("net: remove legacy way to get/set HW timestamp config")
Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
Reviewed-by: Kory Maincent <kory.maincent@bootlin.com>
Link: https://patch.msgid.link/20260918095540.34286-1-nb@tipi-net.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Before the cited commit, fib6_nh_flush_exceptions() always set
from->exception_bucket_flushed = 1 under rt6_exception_lock to
prevent rt6_insert_exception() from inserting a new exception
for a dying fib6_info.
The flag was replaced with the FIB6_EXCEPTION_BUCKET_FLUSHED
bit stored in nh->rt6i_exception_bucket.
The problem is that now the bit is only set when the bucket
is not NULL and fib6_nh_flush_exceptions() is called from
fib6_nh_release() after fib6_ref has already reached zero.
If rt6_insert_exception() is called while the target fib6_info
is being removed via fib6_purge_rt(), a new exception could be
created successfully because rt6_flush_exceptions() no longer
sets the bit.
This creates a reference cycle between the fib6_info and the
exception route, leaking the fib6_info, its nexthop device,
and all per-CPU routes in fib6_nh->rt6i_pcpu, which stalls netdev
unregistration.
[ 34.680602] unregister_netdevice: waiting for gre6 to become free. Usage count = 68
[ 44.920675] unregister_netdevice: waiting for gre6 to become free. Usage count = 68
[ 55.176582] unregister_netdevice: waiting for gre6 to become free. Usage count = 68
Let's call fib6_drop_pcpu_from() before rt6_flush_exceptions(),
to set fib6_destroying before rt6_exception_lock, and check
f6i->fib6_destroying in rt6_insert_exception().
Note that FIB6_EXCEPTION_BUCKET_FLUSHED logic is dead and
we can clean it up in net-next.
Fixes: cc5c073a69 ("ipv6: Move exception bucket to fib6_nh")
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260918082209.2853582-1-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The original code used DMA_ATTR_FORCE_CONTIGUOUS, which could exhaust
the CMA pool when a large number of VFs were requested.
Fix this by switching to the DMA streaming API. This is equivalent on
Octeon platforms, which provide full I/O coherency via the SMMU.
Cc: Leon Romanovsky <leon@kernel.org>
Fixes: 73d33dbc07 ("octeontx2-af: Use DMA_ATTR_FORCE_CONTIGUOUS attribute in DMA alloc")
Signed-off-by: Ratheesh Kannoth <rkannoth@marvell.com>
Reviewed-by: Leon Romanovsky <leon@kernel.org>
Link: https://patch.msgid.link/20260916022111.1083017-1-rkannoth@marvell.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
-----BEGIN PGP SIGNATURE-----
iQIzBAABCgAdFiEEjF9xRqF1emXiQiqU1w0aZmrPKyEFAmqtHzMACgkQ1w0aZmrP
KyGtkQ//bMKQGEudQKCMCtQmPqaHyW1ajmAo17aumfczaE/nSjqrsGY5Ahw7OKlQ
4otcdPvI4qpV9sLTg41KFaIHIC5sozxt4Q3m3RNB2TbyCkGn9xpSpZxM5IpHybvE
83tVjSA0wpfIqxBEqKUqk8Z9AXtBLo/JocdfYry+6JUyj4PM76X2ViKpzaPbpoMU
1mndfLAYtADIIvs3805CmfdJmOkoSV6XCEsiNutPrJhiRfN4xJZ9leP9xb1zA0IQ
cnqiaw1xkTcFyWCicu4MqOkEALRknr9SL2yX1S9wx5Q6WHwU9JXUeQTlvfv7OoVP
uxuMlNr3WcbwHC9e1GfOHapzjrYgnvEe2Z79i2GFh51Ci+5L9Yr9XCQ/fc6G5NNZ
3W52kh35s3lXq32hll9Tkr7pf4cKLBA+IAJ19VNlRfMrPB0cz4EqbIZ6xNNuLqdh
DbEb3VgTT2dHwuGxEshJVmSfzfR+VeHBG2ZRlRmZElfhViHEwgPaAxkaJhNpPyub
qmHbZCXK0BVp/UrGHDm5rmHJtdkwprXY9YceZBRfW8Fr2Ler4rWQvy+uo9sRRFgN
oF9B6qzSl78THBv3UDB3U+aWuDv0I+VlDub0DKf0k9iWg8OoMlM9yw6OE9ULHt3m
LRvSAfF0LnF/z7byzumUYMVjpVuIVueDzasGrEZjWbHHE7MCsp0=
=RfSQ
-----END PGP SIGNATURE-----
Merge tag 'nf-26-09-18' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf
Pablo Neira Ayuso says:
====================
Netfilter/IPVS fixes for net
The following patchset contains Netfilter/IPVS fixes for net, they are:
1) Set on HW_DEAD after HW_PENDING is cleared in the flowtable offload
to ensure GC does not zap it, from Jérémy Jean.
2) Hold the nfnetlink_queue mutex while removing the queue instance
from the netlink notifier that handles NETLINK_URELEASE to fix a
possible race with the UNBIND command. From Florian Westphal.
3) Reject route with NULL rt6i_idev in ip6t_rpfilter. From Weiming Shi.
4) Reject rtinfo->addrnr set to zero from ip6t_rt .checkentry path.
This also fortifies the datapath loop as per Florian's request.
From Luxiao Xu.
5) Fix checksuming in nft_synproxy for IPv6, from Karl Mehltretter.
6) Revalidate ihl before calling icmp_send() in IPVS,
from Julian Anastasov.
7) Fix suspicious RCU usage splat in ctnetlink with expectations.
8) Check for expired catchall elements in the insert and deactivate
path. From Aohan Mei.
* tag 'nf-26-09-18' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf:
netfilter: nf_tables: skip expired catchall elements on insert and delete
netfilter: ctnetlink: fix suspicious RCU usage in expect_iter_name
ipvs: revalidate ihl before icmp_send
netfilter: nft_synproxy: use the family-aware checksum helper
netfilter: ip6t_rt: fix zero-address non-strict match out-of-bounds read
netfilter: ip6t_rpfilter: reject routes without inet6_dev
netfilter: nfnetlink_queue: hold nfnl mutex in event notifier
netfilter: flowtable: publish HW_DEAD after worker is done
====================
Link: https://patch.msgid.link/20260918112844.194503-1-pablo@netfilter.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>