Commit Graph

1483268 Commits

Author SHA1 Message Date
Florian Fainelli
cbbc1aee77 net: bcmgenet: do not skip WoL power up on GENET V1
bcmgenet_power_up() had an early check for bcmgenet_has_ext(priv) before
dispatching by power mode. GENET V1 does not have the EXT block (unlike
GENET V2+), which causes bcmgenet_power_up() to immediately return 0.

As a consequence, when waking up from GENET_POWER_WOL_MAGIC on GENET V1,
bcmgenet_wol_power_up_cfg() is never invoked to disable the WoL clock,
clear wake event masks, and restore normal PHY and MAC operations.

Move the bcmgenet_has_ext() checks to the GENET_POWER_PASSIVE and
GENET_POWER_CABLE_SENSE cases where the EXT registers are actually
accessed, allowing GENET_POWER_WOL_MAGIC cleanup to execute on all
hardware versions.

Fixes: c3ae64ae0c ("net: bcmgenet: handle GENET_POWER_WOL_MAGIC")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Florian Fainelli <florian.fainelli@broadcom.com>
Link: https://patch.msgid.link/20260921220021.281418-4-florian.fainelli@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23 19:07:00 -07:00
Florian Fainelli
3aeaa609fd net: bcmgenet: initialize u64 stats seq counter for all queues
bcmgenet_gstrings_stats statically defines ethtool statistics for queues
0 through GENET_MAX_MQ_CNT (4). However, bcmgenet_probe() only initialized
the u64_stats_sync seq counter up to priv->hw_params->rx_queues and
priv->hw_params->tx_queues.

Since priv->hw_params->rx_queues is 0 across all hardware versions (and
priv->hw_params->tx_queues is 0 on GENET V1), rings 1..4 have uninitialized
u64_stats_sync structures. When ethtool -S is run on 32-bit kernels,
bcmgenet_get_ethtool_stats() reads stats from rx_rings[1..4], causing
lockdep warnings due to the uninitialized sequence counters.

Initialize the sequence counters for all GENET_MAX_MQ_CNT + 1 queues.

Fixes: ffc2c8c4a7 ("net: bcmgenet: Initialize u64 stats seq counter")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Florian Fainelli <florian.fainelli@broadcom.com>
Link: https://patch.msgid.link/20260921220021.281418-3-florian.fainelli@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23 19:06:49 -07:00
Florian Fainelli
0e2bec77ea net: bcmgenet: fix 64-bit RTNL stats reading in ethtool on 32-bit systems
When bcmgenet was converted to 64-bit statistics, STAT_RTNL members were
switched to point into struct rtnl_link_stats64, whose fields are 64-bit
(__u64) regardless of architecture.

However, bcmgenet_get_ethtool_stats() retained a legacy check:
  if (sizeof(unsigned long) != sizeof(u32) &&
      s->stat_sizeof == sizeof(unsigned long))

On 32-bit systems, sizeof(unsigned long) == sizeof(u32), causing this
condition to evaluate to false. As a result, 64-bit RTNL stats fields were
read via *(u32 *)p. On 32-bit Big-Endian systems (such as MIPS BE), this
reads the high 32 bits and returns 0 until the counter exceeds 4GB; on
32-bit Little-Endian systems (such as 32-bit ARM), the value is truncated
to 32 bits.

Fix this by checking if s->stat_sizeof == sizeof(u64) so 64-bit fields are
always read as 64-bit values.

Fixes: 59aa6e3072 ("net: bcmgenet: switch to use 64bit statistics")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Florian Fainelli <florian.fainelli@broadcom.com>
Link: https://patch.msgid.link/20260921220021.281418-2-florian.fainelli@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23 19:06:49 -07:00
Yuya Kusakabe
87cd6b717e net: ipv6: keep room for the mac header in dst_dev_overhead()
The seg6, ioam6 and rpl lwtunnels size their skb_cow_head() request as
the length they are about to push plus dst_dev_overhead(), then push the
new headers and rebuild the mac header below them with
skb_mac_header_rebuild().  That rebuild needs skb->mac_len of headroom,
but dst_dev_overhead() leaves LL_RESERVED_SPACE() of the egress device,
16 bytes for plain Ethernet.

Where the mac header is longer than that, as it is on ingress through a
VLAN device with reorder_hdr off, the rebuild runs out of room:
skb_set_mac_header(skb, -skb->mac_len) computes a negative offset,
stores it unchecked in the u16 skb->mac_header, and the memmove that
follows writes skb->mac_len bytes about 64 KB past skb->head.
Forwarding plain ping6 traffic through such a device reproduces it on
all five seg6 encapsulation modes and on the rpl and ioam6 inline paths;
skb->mac_header comes back as 65534 on a 704-byte head.

Return the larger of the two.  The helper already returns skb->mac_len
when it has no dst, so this only makes the other branch agree, and it
covers every caller rather than each call site in turn.

Fixes: 40475b6376 ("net: ipv6: seg6_iptunnel: mitigate 2-realloc issue")
Fixes: dce525185b ("net: ipv6: ioam6_iptunnel: mitigate 2-realloc issue")
Fixes: 985ec6f5e6 ("net: ipv6: rpl_iptunnel: mitigate 2-realloc issue")
Suggested-by: Andrea Mayer <andrea.mayer@uniroma2.it>
Signed-off-by: Yuya Kusakabe <yuya.kusakabe@gmail.com>
Reviewed-by: Justin Iurman <justin.iurman@gmail.com>
Reviewed-by: Gabriel Goller <g.goller@proxmox.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Andrea Mayer <andrea.mayer@uniroma2.it>
Link: https://patch.msgid.link/20260922-seg6-maclen-headroom-v3-1-7b2f982ef79d@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23 19:04:47 -07:00
Lorenzo Bianconi
0a7822e34a net: stmmac: clear stale buf->page after recycling on skb build failure
In stmmac_rx(), when napi_build_skb() fails the descriptor page is
recycled back to the page pool with page_pool_recycle_direct(), but
buf->page is left pointing at the recycled page, unlike every other
consumption site in the function which clears the pointer after handing
the page away.

With the stale pointer stmmac_rx_refill() skips the replacement
allocation and programs the already-recycled page back into the RX
descriptor.

Clear buf->page on the napi_build_skb() failure path to keep the buffer
lifecycle consistent with the other consumption sites.

Fixes: df542f6693 ("net: stmmac: Switch to zero-copy in non-XDP RX path")
Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/20260921-stmmac-fix-napi-build-skb-error-v1-1-3d54bf6d9bb6@oss.qualcomm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23 19:02:26 -07:00
Ivan Delalande
4eb3f195ef tg3: use random MAC address when tg3_get_device_address fails
Some of the tg3 NICs we use (BCM57762) reset the SRAM MAC address to the
placeholder address on link flaps, tg3_chip_reset, etc. We've typically
fixed it from userspace, but since e4c00ba727 ("tg3: replace
placeholder MAC address with device property") was merged, tg3 just
fails probe as we don't have a way to get it through the generic
device_get_mac_address infrastructure as fallback on our systems.

Make the driver assign a random address in this condition instead of
being fatal for probe. Set deferred_probe_reason through dev_warn_probe
if the address isn't yet available from the provider.

Fixes: e4c00ba727 ("tg3: replace placeholder MAC address with device property")
Suggested-by: Jakub Kicinski <kuba@kernel.org>
Link: https://lore.kernel.org/netdev/20260909191751.651aa5c4@kernel.org/
Signed-off-by: Ivan Delalande <colona@arista.com>
Link: https://patch.msgid.link/20260918224715.GA654128@visor
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23 19:01:15 -07:00
Zijie Huang
d8b6529e80 net: arp: terminate device name before lookup
The ARP ioctl copies a user-provided struct arpreq into a stack object. Its
arp_dev field may contain IFNAMSIZ bytes without a NUL terminator.

Such input is passed to dev_get_by_name_rcu() or __dev_get_by_name(), where
strcmp() can read past the end of the stack object when a matching
alternative interface name exists.

Terminate the field before the lookup to prevent the out-of-bounds read.

Fixes: 36fbf1e52b ("net: rtnetlink: add linkprop commands to add and delete alternative ifnames")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Signed-off-by: Zijie Huang <milkory@outlook.com>
Signed-off-by: Ren Wei <weir@nebusec.ai>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/fabf02a70787d17299e4b3153eadffaf20d154b3.1789910973.git.milkory@outlook.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23 19:00:15 -07:00
Jonas Jelonek
89a8a1eef2 net: mdio: realtek-rtl9300: fix RTL931x C22 extended page selection
The RTL931x indirect access engine has a separate nine-bit extended page
field. The driver leaves it at zero, and otto_emdio_run_cmd() therefore
programs extended page zero for every Clause 22 transaction. This
overrides page selection made through PHY register 30, causing accesses
to private PHY pages to hit extended page zero instead.

Set the field to its 0x1ff "do not change" value for RTL931x Clause 22
reads and writes. This preserves extended page selection made through
PHY register 30 and restores access to its private register pages.

Fixes: 5ebdcac59a ("net: mdio: realtek-rtl9300: Add support for RTL931x")
Signed-off-by: Jonas Jelonek <jonas@jonasjelonek.de>
Acked-by: Markus Stockhausen <markus.stockhausen@gmx.de>
Link: https://patch.msgid.link/20260918211955.3955777-1-jonas@jonasjelonek.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23 18:54:19 -07:00
Hui Peng
d22609f3d1 fou: reject omitted FOU_ATTR_IPPROTO on FOU_ENCAP_DIRECT
Commit 7a9bc9e3f4 ("fou: Don't allow 0 for FOU_ATTR_IPPROTO.") added
NLA_POLICY_MIN(NLA_U8, 1) to fou_nl_policy[FOU_ATTR_IPPROTO], which
rejects an explicitly supplied FOU_ATTR_IPPROTO == 0 attribute with
-ERANGE.

However, FOU_ATTR_IPPROTO is an optional netlink attribute. When a user
sends FOU_CMD_ADD with FOU_ATTR_TYPE set to FOU_ENCAP_DIRECT and omits
FOU_ATTR_IPPROTO entirely, nla_policy validation succeeds and
parse_nl_config() leaves cfg->protocol as 0 (from memset(cfg, 0,
sizeof(*cfg))). fou_create() then creates a FOU_ENCAP_DIRECT socket with
fou->protocol == 0.

In fou_udp_recv(), returning -fou->protocol to udp_queue_rcv_one_skb()
triggers IP protocol resubmission when fou->protocol > 0, whereas
returning 0 tells the UDP tunnel layer that the skb was consumed without
freeing it. When fou->protocol == 0, every packet received on the socket
returns 0 from fou_udp_recv() and leaks the sk_buff.

Reject FOU_ENCAP_DIRECT when !cfg->protocol in fou_create() so that
creating a direct encapsulation port without FOU_ATTR_IPPROTO fails with
-EINVAL while leaving FOU_CMD_DEL and FOU_CMD_GET (which share
parse_nl_config()) unaffected.

Tested in QEMU against Linux 7.3.0-rc3 by sending a FOU_CMD_ADD Generic
Netlink request with FOU_ATTR_PORT = 5555 and FOU_ATTR_TYPE =
FOU_ENCAP_DIRECT while omitting FOU_ATTR_IPPROTO. On the unfixed kernel,
FOU_CMD_ADD succeeds (err = 0), FOU_CMD_GET reports fou->type = 1 and
fou->protocol = 0, and sending 4000 UDP packets to 127.0.0.1:5555 leaks
all 4000 sk_buffs (SUnreclaim in /proc/meminfo grows from 41456 kB to
59008 kB, +17552 kB); with this patch applied, FOU_CMD_ADD is rejected
with -EINVAL (-22).

Fixes: 23461551c0 ("fou: Support for foo-over-udp RX path")
Fixes: 7a9bc9e3f4 ("fou: Don't allow 0 for FOU_ATTR_IPPROTO.")
Cc: stable@vger.kernel.org
Signed-off-by: Hui Peng <benquike@gmail.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Link: https://patch.msgid.link/20260921045920.1613098-1-benquike@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23 18:43:08 -07:00
Nicolo Giuliani
6b491af01a net: dsa: mv88e6xxx: 88E6191X and 88E6193X have no PTP
The 88E6191X and 88E6193X are 6393 family devices that share
mv88e6393x_ops with the 88E6393X and are marked as ptp_support. Marvell's
UMSD driver describes both as parts without AVB (88E6193X: "BGA package -
No AVB, No Routing, No Cut-through"), and the register access confirms it
on an 88E6193X: the whole indirect AVB register space behind Global 2
registers 0x16 and 0x17 reads zero, for every port, block and address,
with the 6390 and with the 6352 command encoding. Writes to the TAI
registers, including the clock period register and the TAI global
configuration register, read back as zero.

Since commit 7e3c18097a ("net: dsa: mv88e6xxx: read cycle counter
period from hardware") the PTP setup reads the TAI clock period, so the
switch fails to probe:

mv88e6xxx ...: unexpected cycle counter period of 0 ps

Add mv88e6191x_ops, a copy of mv88e6393x_ops without avb_ops and ptp_ops,
use it for the 88E6191X and the 88E6193X and stop setting ptp_support for
them. The 88E6393X is unchanged.

Tested on an 88E6193X (Sophos XGS 107w): the switch probes and the ports
work. I do not have an 88E6191X, it is changed because UMSD describes it
the same way.

Fixes: de776d0d31 ("net: dsa: mv88e6xxx: add support for mv88e6393x family")
Suggested-by: Andrew Lunn <andrew@lunn.ch>
Signed-off-by: Nicolo Giuliani <nicolo.giuliani6@studio.unibo.it>
Reviewed-by: Andrew Lunn <andrew@lunn.ch>
Link: https://patch.msgid.link/20260921-send-net-v2-1-031ad720f140@studio.unibo.it
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23 18:39:46 -07:00
Hui Peng
2d959c75c2 ipv6: sr: enforce exact attribute length for SEG6_ATTR_DST
In seg6_genl_policy, SEG6_ATTR_DST is defined with .type = NLA_BINARY and
.len = sizeof(struct in6_addr). For NLA_BINARY, .len only enforces the
maximum payload length and permits shorter payloads (e.g., 0 bytes).
When seg6_genl_set_tunsrc() copies sizeof(struct in6_addr) bytes via
kmemdup(val, sizeof(*val), GFP_KERNEL), a short SEG6_ATTR_DST attribute
triggers a 16-byte out-of-bounds read past skb->tail into uninitialized
skb->head memory, which is stored in sdata->tun_src and leaked back to
userspace via SEG6_CMD_GET_TUNSRC.

Switch SEG6_ATTR_DST in seg6_genl_policy to
NLA_POLICY_EXACT_LEN(sizeof(struct in6_addr)) so that generic netlink
validation rejects any attribute whose length is not exactly
sizeof(struct in6_addr) with -ERANGE.

Tested in QEMU against Linux 7.3.0-rc3 by sending a SEG6_CMD_SET_TUNSRC
Generic Netlink message with a 0-byte SEG6_ATTR_DST attribute followed
by SEG6_CMD_GET_TUNSRC. On the unfixed kernel, SEG6_CMD_SET_TUNSRC
succeeds (err = 0) and SEG6_CMD_GET_TUNSRC leaks 16 bytes of
uninitialized kernel heap memory (tun_src =
836a61ecc4d25a1042a8d60411cfb378); with this patch applied,
SEG6_CMD_SET_TUNSRC is rejected by netlink policy validation with
-ERANGE (-34) and tun_src remains zeroed.

Fixes: 915d7e5e59 ("ipv6: sr: add code base for control plane support of SR-IPv6")
Cc: stable@vger.kernel.org
Signed-off-by: Hui Peng <benquike@gmail.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Reviewed-by: Justin Iurman <justin.iurman@gmail.com>
Reviewed-by: Andrea Mayer <andrea.mayer@uniroma2.it>
Link: https://patch.msgid.link/20260921044025.1535982-1-benquike@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23 18:34:13 -07:00
Ido Schimmel
ab7aa05c06 vrf: Stop corrupting skb->csum when capturing CHECKSUM_COMPLETE packets
The VRF device is an Ethernet device but it can have non-Ethernet ports
such as IP tunnels. Before the cited commit, capturing packets from such
ports on the VRF device resulted in these packets being detected as
malformed since they lack an Ethernet header.

The cited commit fixed it by pushing a dummy Ethernet header to such
packets before the capture and pulling it afterwards. In the case of
CHECKSUM_COMPLETE packets it also updated skb->csum with the checksum of
the dummy Ethernet header. This is wrong as skb->csum should not include
the checksum of the Ethernet header ("checksum of the _whole_ packet as
seen by netif_rx()").

This also means that L4 protocols receive a corrupted skb->csum and
potentially drop the packet, as is the case with UDP packets whose
checksum was completed by software.

Fix by removing the unnecessary call to skb_postpush_rcsum().

Fixes: 0489390882 ("vrf: add mac header for tunneled packets when sniffer is attached")
Reported-by: Stefano Sasso <stesasso@gmail.com>
Closes: https://lore.kernel.org/netdev/CALtE316UtL3x7LL6uxfXzx8rW6AbzYPeDOb478hqJCr_-dj=Wg@mail.gmail.com/
Signed-off-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: David Ahern <dsahern@kernel.org>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Andrea Mayer <andrea.mayer@uniroma2.it>
Link: https://patch.msgid.link/20260922131239.2509494-1-idosch@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23 17:47:10 -07:00
Shihuang Liu
3b4e0b0c00 net: skbuff: fix pull-bound underflow in skb_checksum_setup_ipv6()
skb_maybe_pull_tail() subtracts skb_headlen(skb) from the unsigned max
argument and passes the result to __pskb_pull_tail() as a signed int.  The
function does not ensure that max is at least skb_headlen(skb).

This can happen while parsing IPv6 extension headers when an skb already
has a linear area larger than MAX_IPV6_HDR_LEN.  Once the parser needs data
beyond the linear area, max - skb_headlen(skb) wraps and is converted to a
negative delta.  __pskb_pull_tail() then passes that negative length to
skb_copy_bits(), where it can become a very large copy length.

Pass the requested length itself as the pull bound at the three
extension-header call sites, so the delta can no longer go negative.

Fixes: 1431fb31ec ("xen-netback: fix fragment detection in checksum setup")
Suggested-by: Eric Dumazet <edumazet@google.com>
Signed-off-by: Shihuang Liu <shlomojune6@gmail.com>
Link: https://patch.msgid.link/20260919133604.50948-1-shlomojune6@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23 17:41:22 -07:00
Jakub Kicinski
7104a37071 veth: manage XDP program pointers during channel resize
veth_set_channels() tears down XDP resources for removed RX queues
without clearing rq->xdp_prog.  If the program is then detached or
replaced, those queues keep the old pointer after bpf_prog_put().
A later channel increase can re-enable NAPI and run the freed program.

  BUG: unable to handle page fault for address: ffffc90000256048
  Oops: Oops: 0000 [#1] SMP KASAN NOPTI
  RIP: veth_xdp_rcv_skb (include/linux/filter.h:779
                         include/net/xdp.h:696 drivers/net/veth.c:820)
  Call Trace:
   veth_xdp_rcv (drivers/net/veth.c:941)
   veth_poll (drivers/net/veth.c:986)
   __napi_poll (net/core/dev.c:7787)
   net_rx_action (net/core/dev.c:7850 net/core/dev.c:8007)
   handle_softirqs (kernel/softirq.c:645)
  Kernel panic - not syncing: Fatal exception in interrupt

Fixes: 4752eeb3d8 ("veth: implement support for set_channel ethtool op")
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Reviewed-by: Jason Xing <kerneljasonxing@gmail.com>
Link: https://patch.msgid.link/20260921231856.1798630-1-kuba@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23 16:53:57 -07:00
Nicolai Buchwitz
7e87508b5c net: bcmgenet: stop Tx NAPI before disabling the queues
bcmgenet_netif_stop() and the Wake-on-LAN branch of bcmgenet_suspend()
both disable the Tx queues first and stop Tx NAPI several steps later. A
completion in flight calls netif_tx_wake_queue() in between, and nothing
stops the queue again, so a transmit can reach the rings after they have
been freed.

Close is safe because dev_deactivate_many() stops the qdisc first.
bcmgenet_suspend() does not, so stop Tx NAPI before the queues on both
paths.

KASAN on a Raspberry Pi CM4, driven from an MTU change because suspend
freezes user space before the callback runs:

  BUG: KASAN: use-after-free in bcmgenet_xmit+0x17f8/0x2258
  Write of size 8 at addr ffffff8055844a68 by task ksoftirqd/0/14
   bcmgenet_xmit+0x17f8/0x2258
   dev_hard_start_xmit+0x13c/0x588
   sch_direct_xmit+0x108/0x340
   __dev_queue_xmit+0x1190/0x3848

Fixes: 254f3239dd ("net: bcmgenet: revise suspend/resume")
Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
Link: https://patch.msgid.link/20260922130639.1660797-1-nb@tipi-net.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23 16:52:54 -07:00
Nicolai Buchwitz
fdfec06ac1 MAINTAINERS: add Nicolai Buchwitz as GENET maintainer
I have been contributing to and reviewing the GENET driver for a while
now. Florian asked if I would like to formalize this commitment, so add
myself as a maintainer.

Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
Acked-by: Florian Fainelli <florian.fainelli@broadcom.com>
Acked-by: Justin Chen <justin.chen@broadcom.com>
Link: https://patch.msgid.link/20260922073140.1471858-1-nb@tipi-net.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23 16:52:14 -07:00
Fourie Zhang
ab1404ac81 net: bridge: mdb: restart port group walk after deletion
br_mdb_flush_pgs() keeps a pointer-to-pointer cursor while walking
mp->ports. br_multicast_del_pg() can re-enter the same MDB entry through
br_multicast_sg_del_exclude_ports() and unlink other port groups. If the
cursor points into one of those groups, the next iteration dereferences a
stale cursor and can leave mp->ports pointing at freed memory.

A following RTM_GETMDB exposes the dangling pointer:

  BUG: KASAN: slab-use-after-free in br_mdb_dump
  Read of size 8
    br_mdb_dump
    rtnl_mdb_dump
    rtnl_dumpit
    netlink_dump

Reset the cursor to mp->ports after every deletion. The deletion removes at
least the selected group, so the restarted walk always makes progress.

Fixes: a6acb535af ("bridge: mdb: Add MDB bulk deletion support")
Cc: stable@vger.kernel.org
Signed-off-by: Fourie Zhang <fouriezhang@tencent.com>
Acked-by: Nikolay Aleksandrov <razor@blackwall.org>
Link: https://patch.msgid.link/20260920110852.60293-1-fouriezhang@tencent.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23 16:51:03 -07:00
Victor Nogueira
cfa165cbfb net/sched: act_gate: budget the per-entry list in get_fill_size
tcf_gate_get_fill_size returns only the TCA_GATE_PARMS size, but
tcf_gate_dump also emits three 64-bit timestamps, the clock id, flags,
priority and the variable-length TCA_GATE_ENTRY_LIST nest. The per-entry
nest is unbounded: parse_gate_list places no cap on the number of
sched-entries, so a gate with many entries can push the real dump well
past the skb that tca_get_fill allocates from this size.

RTM_NEWACTION then fails the add-notify with -EINVAL while the action is
already committed to the IDR, and a subsequent RTM_GETACTION on the
installed gate also returns -EINVAL because its dump no longer fits.

Fix this by accounting for the missing fields in tcf_gate_get_fill_size
along with all elements in the entries list.

Note that sizing the reply from the action lets an oversized gate
install cleanly for the first time: with the input unbounded by
parse_gate_list, the sized skb can now grow well above
NLMSG_GOODSIZE per netlink request (a transient GFP_KERNEL allocation
reachable only with namespace-local CAP_NET_ADMIN). Overload from a
malicious netns admin is hardening material, not net, per the
discussion at
https://lore.kernel.org/netdev/20260914191108.55a1a4f1@kernel.org/;
a follow-up patch for net-next will cap the sched-entry count.

Fixes: 4e76e75d6a ("net sched actions: calculate add/delete event message size")
Reported-by: Sashiko <sashiko-bot@kernel.org>
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260824153903.4143642-1-victor@mojatatu.com
Tested-by: hybris <hybris@mojatatu.ai>
Co-developed-by: Jamal Hadi Salim <jhs@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Signed-off-by: Victor Nogueira <victor@mojatatu.com>
Link: https://patch.msgid.link/QDISC-3BLH.v1.20260914203033@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23 16:42:50 -07:00
Bernard Ladenthin
9c572a8303 net/sched: fix potential stack infoleak in em_text_dump()
em_text_dump() allocates struct tcf_em_text on the stack without zeroing
it.  strscpy() writes the algorithm name and a NUL terminator into
conf.algo[], leaving the remaining bytes uninitialised.  nla_put_nohdr()
then copies the full struct to the netlink response.

KMSAN on Linux 7.2-rc6 reports two kernel-infoleak splats from this path,
one triggered via "tc filter show" and one via a raw RTM_GETTFILTER dump:

  BUG: KMSAN: kernel-infoleak in _copy_to_iter+0x1c9/0x2620
    nla_put_nohdr+0x83/0x130
    em_text_dump+0x291/0x550
  Local variable conf created at: em_text_dump+0x5d/0x550
  Bytes 168-179 of 199 are uninitialized

I am not certain whether this constitutes a real security problem in
practice: the test was conducted in a controlled KMSAN environment and
the leaked stack bytes may or may not carry sensitive data on actual
production kernels.  I am reporting it because KMSAN flagged it as a
kernel-infoleak and the fix is straightforward.  I can provide a
userspace reproducer on request.

The original code used strncpy() which zero-pads to the destination size.
Commit b04202d606 ("net/sched: replace strncpy with strscpy") replaced
it with strscpy(), which does not pad, creating this condition.
Zero-initialising the struct closes it.

Fixes: b04202d606 ("net/sched: replace strncpy with strscpy")
Link: https://lore.kernel.org/netdev/20250327143733.187438-1-richard120310@gmail.com/
Assisted-by: Claude:claude-sonnet-4-6 [KMSAN]
Signed-off-by: Bernard Ladenthin <bernard.ladenthin@gmail.com>
Acked-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/20260918133953.12494-1-bernard.ladenthin@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22 19:14:25 -07:00
Björn Töpel
0160953d8e eth: fbnic: Avoid rounding zero ring sizes
roundup_pow_of_two() is undefined for zero. ethtool permits a zero ring
size to reach the driver, where the minimum-size check should reject it.

Leave zero unchanged while rounding nonzero ring sizes. The minimum-size
check then rejects zero deterministically without changing the established
behavior for other values.

Fixes: 6cbf18a05c ("eth: fbnic: support ring size configuration")
Reported-by: Sashiko <netdev-bot+sashiko@kernel.org>
Link: https://lore.kernel.org/netdev/178971206933.22033.236948278674126701@kernel.org/
Suggested-by: Alexander Duyck <alexanderduyck@fb.com>
Signed-off-by: Björn Töpel <bjorn@kernel.org>
Reviewed-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260918114641.1281172-1-bjorn@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22 18:53:11 -07:00
Jakub Kicinski
f764634137 Merge branch 'net-stmmac-more-selftest-related-fixes'
Maxime Chevallier says:

====================
net: stmmac: More selftest-related fixes

This is V4 of stmmac selftest fixes, addressing Sashiko's issues over
the MTU patch. This lead to the introduction of a new one. The
dev_add_pack races have been addressed, however the double-vlan issue
stayed there. Ovidiu is actively working on it, let's wait for his work
to land before fixing that.

This is another round of stmmac selftest fixes, mostly about the selftests
themselves but a few things were discovered w.r.t MTU and buffer size
handling, see patch 5 anf 6.

After this is merged, I consider the selftests to be now reliable enough
to run them nightly on every stmmac series that's sent, and I'll be requiring
clean selftests for new glue drivers.

Since V3, the testing farm grew ! I've been running this on :

 - Altera CycloneV (dwmac-socfpga, dwmac1000 IP, v3.70a)
 - NXP imx8mp (dwmac-imx, dwmac4, v5.10a)
 - Allwinner H2S (dwmac-sun8i, dwmac1000)
 - Amlogic S905X3 (dwmac-meson8b, dwmac1000, v3.70a)
 - STM32mp157a (dwmac-stm32, dwmac4, v4.20a)
 - SiFive JH7110 (dwmac-starfive, dwmac4, v5.20)
 - Motorcomm YT8061 (PCIe, dwmac-motorcomm, dwmac4)
 - Qualcomm IPQ8064 (dwmac-ipq806x, dwmac1000)
 - Altera AgileX5 (dwmac-socfpga, dwxgmac2 !) (NEW)
 - Generic dwmac1000 (Loongson 2K0300, dwmac 3.70a) (NEW)
 - Rockchip RK3566 (dwmac-rk, dwmac4) (NEW)

It's becoming cumbersome to list the test results here, they can be
found, updated daily, here :

https://minimaxwell.github.io/stmmac-ci/
====================

Link: https://patch.msgid.link/20260917215339.2022523-1-maxime.chevallier@bootlin.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22 18:51:20 -07:00
Maxime Chevallier
c4ac6e94eb net: stmmac: selftests: Account for alignment shift on dwmac1000 for Jumbo test
On dwmac1000, we currently only support single-descriptor frames. The
Jumbo test started failing when NET_IP_ALIGN was added to align the IP
header, as this tests tries to send the biggest possible frame.

On dwmac1000 the DMA transfer is aligned on 4-bytes, so adding a 2-byte
shift at the start-of-buffer address means it takes a whole extra 4-byte
DMA burst to receive the Jumbo packet, causing it to spill over the next
descriptor.

This doesn't seem to happen on dwmac4 and xgmac that appear to correctly
handle unaligned xfers (only tested on dwmac4)

Let's account for that in the Jumbo test, reduce the size of our big
packet by the align size.

Fixes: 23680bf5f8 ("net: stmmac: restore NET_IP_ALIGN in the RX DMA offset")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/20260917215339.2022523-8-maxime.chevallier@bootlin.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22 18:51:17 -07:00
Maxime Chevallier
b8a26d46c0 net: stmmac: size the RX buffers from the frame length, not the MTU
When picking the buffsize to use based on the MTU, we shouldn't check
only the MTU value, but also :
 - ETH_HLEN for the L2 header,
 - up to 2 VLAN tags,
 - the FCS,

The default bufsize is 1536 bytes, which is enough to contain all the
above so this hasn't surfaced before, but the addition of NET_IP_ALIGN
to the start of buffer address tripped the Jumbo selftest, leading to
this discovery.

With that, we don't need the '>=' checks on the buffer len, we can use
more consistent comparison operators in stmmac_set_bfsize.

Fixes: 286a837217 ("stmmac: add CHAINED descriptor mode support (V4)")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/20260917215339.2022523-7-maxime.chevallier@bootlin.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22 18:51:17 -07:00
Maxime Chevallier
b42e701277 net: stmmac: dwmac4: Use the correct bufzise when the len is exactly 8K
DMA bufsize selection isn't made on the MTU but the actual frame length,
so including the L2 header. On DWMAC4, if the len is exactly BUF_SIZE_8KiB,
the next larger size is incorrectly selected.

Lets fix the comparison and while at it, rename the parameter from len
to mtu.

Fixes: c3efed5ad1 ("net: stmmac: Enable dwmac4 jumbo frame more than 8KiB").
Signed-off-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Link: https://patch.msgid.link/20260917215339.2022523-6-maxime.chevallier@bootlin.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22 18:51:17 -07:00
Maxime Chevallier
960db6f657 net: stmmac: selftests: Capture all packets for vlan checks
While we use vlan_vid_add to trigger the tag filtering machinery
in the driver, there's no netdev associated to the VLAN. This causes the
skb to arrive with empty skb->vlan_tci fields, as the packet is marked
OTHERHOST in __netif_receive_skb_core(), and we fail our validation.

Let's use the proxy mechanism introduced for DSA, that registers a
ETH_P_ALL packet handler that runs earlier, before the vlan netdev
lookup, then filters for the correct ethertype before passing an skb
clone to our validation function.

As we may receive external frames with the right tag from the outside,
let's move the address check in the vlan validation function earlier.

Fixes: 091810dbde ("net: stmmac: Introduce selftests support")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/20260917215339.2022523-5-maxime.chevallier@bootlin.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22 18:51:17 -07:00
Maxime Chevallier
ba804b23d7 net: stmmac: selftests: Check the dev->features for S-TAG offload testing
The S-TAG offload insertion incorrectly checks the dvlan (double vlan)
DMA cap, which is different than S-TAG support. Use
NETIF_F_HW_VLAN_STAG_TX to check if the feature is supported instead.

Note that this flag isn't set in stmmac yet, but contrary to ARP
offload, this is a feature that has a chance to get there eventually so
let's leave the selftest here for now. It'll report -EOPNOTSUPP in the
meantime.

Fixes: 091810dbde ("net: stmmac: Introduce selftests support")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/20260917215339.2022523-4-maxime.chevallier@bootlin.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22 18:51:17 -07:00
Maxime Chevallier
c8c1795aa8 net: stmmac: selftests: Validate EEE based on the actual LPI timer value
The EEE selftest is a 2-step test :
 - It validates that we enter in LPI mode with the
   irq_tx_path_in_lpi_mode_n counter
 - It then validates that we exit LPI when sending a frame, with the
   irq_tx_path_exit_lpi_mode_n counter.

The current state of the test lacks 2 main things :

 - We don't know exactly when was the previous frame sent (it's from the
   previous selftest)

 - The timeout is hardcoded, while the LPI is entered after a
   user-configurable delay. On top of that, the timeout loop uses a
   pre-decrement iterator (--retries) that actually only iterate nine
   times, so 900ms while the default LPI value is 1 second.

Let's therefore make it more deterministic :

 - Send a frame at the beginning of the test
 - Wait for more than the lpi timer value, we timeout after about twice
   the value,
 - Then send another frame, and verify that we do go out of LPI, also
   with a timeout.

As LPI timer can get pretty high, bail out if LPI timer is over 5
seconds.

Note that the test's goal isn't to validate the LPI timer value itself,
only that we enter/leave LPI mode.

Fixes: 091810dbde ("net: stmmac: Introduce selftests support")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/20260917215339.2022523-3-maxime.chevallier@bootlin.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22 18:51:16 -07:00
Maxime Chevallier
d68acbf935 net: stmmac: selftests: Support running selftests on DSA conduits
Most stmmac selftests rely on dev_add_pack() to add custom handlers,
that validate the packets sent to ourselves through MAC loopback.

However, when the stmmac-driven interface is a DSA CPU conduit, all
frames that are received have ETH_P_XDSA as a protocol, even though they
don't actually contain any tag as they come from the loopback and not
the switch.

This will prevent any incoming packet to match our packet handlers.

Let's register a ETH_P_ALL packet handler when we detect that we're a
DSA conduit, and use a proxy packet handler to filter the h_proto.

As this allows external frames to be received through our .func(), the
packet handler is added after the dev->addr field is populated in our
selftest attributes.

Note that we may still receive incoming packets from the switch, but
these frames shouldn't interfere with the very specific frames used for
selftests, and stmmac selftests in general aren't safe against external
traffic interferences.

This was validated on a WPQ864 devkit for IPQ8064, that has the SoC
connected to a QCA8k switch.

The ARP offload's packet handler is left alone, this feature is just not
implemented in stmmac and due for removal.

Fixes: 091810dbde ("net: stmmac: Introduce selftests support")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/20260917215339.2022523-2-maxime.chevallier@bootlin.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22 18:51:16 -07:00
Mina Almasry
73b3fc67a4 net: devmem: document that bind-tx is unprivileged by design
Unlike bind-rx, which configures shared NIC RX queues to steer incoming
traffic into the caller's dmabuf and requires CAP_NET_ADMIN
(uns-admin-perm), bind-tx only DMA-maps the caller's dmabuf so the caller
can transmit from it on their own sockets without affecting other traffic
or device configuration.

Add a comment in netdev.yaml and above netdev_nl_bind_tx_doit() to make it
explicit that NETDEV_CMD_BIND_TX is unprivileged by design.

Signed-off-by: Mina Almasry <almasrymina@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
Acked-by: Daniel Borkmann <daniel@iogearbox.net>
Link: https://patch.msgid.link/20260921195545.493253-1-almasrymina@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22 18:40:15 -07:00
Xin Long
cae23ae3f7 sctp: hold asoc or transport before mod_timer() in timer handlers
Take the association or transport reference before rearming a timer in the
timer handlers.

The existing code calls mod_timer() before taking the reference needed by
the rearmed timer without holding the sock lock. This creates a race with
timer cleanup: if the timer is deleted after mod_timer() returns but before
the reference is taken, the cleanup path can drop the timer's reference and
destroy the transport or association. The timer handler then takes a
reference on the already freed object and eventually drops it, causing a
refcount underflow.

Hold the object before mod_timer() and drop the reference if mod_timer()
reports that the timer was already pending in timer handlers. Apply the
same ordering to the proto-unreachable path, which can rearm a transport
timer outside the timer handlers without holding the sock lock.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Reported-by: Tangxin Xie <xietangxin@h-partners.com>
Signed-off-by: Xin Long <lucien.xin@gmail.com>
Link: https://patch.msgid.link/c31b5e3ee2b7274e804f5eba2f21e2412e7eef7a.1790013825.git.lucien.xin@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22 18:36:36 -07:00
Jakub Kicinski
6836807c5f Merge branch 'packet-fix-packet_tx_ring-data-corruption-on-skb_orphan'
Willem de Bruijn says:

====================
packet: fix PACKET_TX_RING data corruption on skb_orphan

When transmitting packets via PACKET_TX_RING, tpacket_snd links user
ring buffer pages as skb frags and releases the slot on skb->destructor
(tpacket_destruct_skb).

skb_orphan() invokes the destructor while the skb is still alive.
This marks the slot as TP_STATUS_AVAILABLE prematurely, allowing
userspace to overwrite the slot and causing data corruption.

This series fixes the issue by switching PACKET_TX_RING to standard
ubuf_info zerocopy completion, ensuring ring slots are released only
after all payload references are freed or copied.

Virtio-net needs a separate solution, because deferring the release
can cause deadlock in its !use_napi mode.

- Patch 1 addresses the virtio-net special case.
- Patch 2 converts tpacket_snd to standard ubuf_info completion

Patch 1 must be applied, and backported, before patch 2. Both carry
the same Fixes tag for that reason.

v1: https://lore.kernel.org/netdev/20260914214229.1674102-1-willemdebruijn.kernel@gmail.com/
====================

Link: https://patch.msgid.link/20260919004748.1463985-1-willemdebruijn.kernel@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22 18:33:03 -07:00
Willem de Bruijn
9518405613 packet: use ubuf_info completion for TX_RING packets
tpacket_snd sends skbs with frags pointing into its ring slots. Slots
are released when skb->destructor is called.

A call to skb_orphan calls skb->destructor before the skb is freed.
This can cause the slot to be reused while still linked into the skb.

Switch to standard zerocopy completion (ubuf_info) so the slot is only
released once all references to the payload are freed or copied.
Restore skb->destructor to standard sock_wfree.

The ubuf_info completion callback can be called with a NULL skb, but
only from net_zcopy_put and related API, used by zerocopy implementations
that hold their own reference on the uarg, such as MSG_ZEROCOPY. This
uarg is only ever completed from skb_zcopy_clear, so skb is always set.

To prevent userspace from aliasing in-flight state on shared ring
slots, allocate tpacket_uarg per packet, rather than per slot. This
adds a small allocation to the transmit path. Use standard kmalloc to
allow backporting to stable kernels.

The uarg holds an sk_wmem_alloc reference, rather than an sk_refcnt
reference. packet_free_tx_ring waits on sk_wmem_alloc before freeing
the ring pages. Always allocate vec->deferred for tx_ring so page-backed
rings also wait on sk_wmem_alloc when skb_copy_ubufs drops page refs
before calling tpacket_ubuf_complete.

Drop the tx_ring.pg_vec test that tpacket_destruct_skb performed before
accessing the slot. The sk_wmem_alloc reference now guarantees that the
slot is valid. The test is also not sufficient by itself, as it reads
pg_vec without pg_vec_lock, so it can race with packet_set_ring.

As a result a slot is released when its payload is copied, which can
be before transmission (e.g., in skb_orphan_frags_rx). If copied
before skb_tx_timestamp() is called, no slot timestamp is recorded,
similar to when skb_orphan() was called early in the datapath before
this patch.

Revert the now unused previous skb_zcopy_.._nouarg infra.

Depends on commit 992cc9f94c ("net/packet: defer vmalloc TX_RING
free until skbs finish").

Reported-by: Katherine Leaver <kleaver@janestreet.com>
Reported-by: Bjoern Doebel <doebel@amazon.de>
Closes: https://lore.kernel.org/netdev/20260909085542.3370986-1-doebel@amazon.de/
Fixes: 5cd8d46ea1 ("packet: copy user buffers before orphan or clone")
Cc: stable@vger.kernel.org
Signed-off-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260919004748.1463985-3-willemdebruijn.kernel@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22 18:32:57 -07:00
Willem de Bruijn
07e1a9408b virtio_net: copy zerocopy frags in start_xmit without NAPI
Virtio-net without NAPI frees completed skbs lazily on the next
start_xmit. Senders waiting for in-flight zerocopy buffers can
deadlock if they cannot transmit more packets, as then no
completed packets will be freed.

When !use_napi, virtio-net already calls skb_orphan to avoid waiting
up for transmitted skbs to be freed. For zerocopy packets that
require deep copying on orphan (i.e. those that do not set
SKBFL_DONT_ORPHAN, such as PACKET_TX_RING), call skb_orphan_frags
before orphaning to release the buffers.

This fixes the tpacket_snd slot reuse bug on skb_orphan for
virtio-net, and prevents PACKET_TX_RING from running out of slots.

This fix also touches vhost_net zerocopy packets, which also do not
set SKBFL_DONT_ORPHAN. This is fine: vhost_net packets only encounter
virtio-net in nested virtualization, and only if napi_tx is
explicitly disabled (it has been default-enabled since Linux 4.12).
In that rare case, copying the frags is desirable anyway to prevent
holding guest descriptors pinned across unbounded intervals.

This is a prerequisite for the next patch, which converts
PACKET_TX_RING to standard zerocopy completion. Without this patch
first, a bounded ring sender can stall indefinitely behind a
virtio-net virtqueue that cannot reclaim.

Fixes: 5cd8d46ea1 ("packet: copy user buffers before orphan or clone")
Cc: stable@vger.kernel.org
Cc: mst@redhat.com
Cc: jasowangio@gmail.com
Signed-off-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260919004748.1463985-2-willemdebruijn.kernel@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22 18:32:57 -07:00
Aohan Mei
4498467a8a sctp: discard the rest of the packet on a stale-cookie error
When an association is in COOKIE-ECHOED state and the peer sends a
bundled [ERROR(Stale Cookie)][DATA] packet from one of its non-primary
addresses, processing the ERROR chunk takes the non-fatal stale-cookie
retry path sctp_sf_do_5_2_6_stale(), which queues
SCTP_CMD_DEL_NON_PRIMARY while keeping the association alive.
sctp_cmd_del_non_primary() removes every non-primary transport -
including the very transport this packet arrived on, which is still
referenced by the receive lookup and shared by all chunks of the
packet via chunk->transport.

sctp_assoc_rm_peer() does redirect asoc->peer.last_data_from away from
the removed transport, but right afterwards the bundled DATA chunk
makes sctp_assoc_bh_rcv() re-register
asoc->peer.last_data_from = chunk->transport unconditionally, undoing
the redirection with the just-removed transport.

Once the packet is done, the receive reference is dropped and the
transport is RCU-freed, while the surviving association keeps the
dangling last_data_from.  A later FWD-TSN (or the delayed SACK timer)
makes sctp_gen_sack() dereference it (->param_flags and friends), and
sctp_make_sack()/sctp_outq_select_transport() may write to the freed
object and link it into the live transport list.  This is a
use-after-free triggerable by any malicious SCTP peer (or a local
unprivileged user acting as one) with no capabilities required:

  BUG: KASAN: slab-use-after-free in sctp_do_sm+0x498a/0x5660
  Read of size 4 at addr ffff88800e1e356c by task poc/115
  Call Trace: sctp_do_sm <- sctp_assoc_bh_rcv <- sctp_inq_push <-
              sctp_rcv <- ip_protocol_deliver_rcu <- ip_rcv
  Allocated: sctp_transport_new <- sctp_assoc_add_peer <-
             sctp_process_init (INIT-ACK processing)
  Freed: kfree <- sctp_transport_destroy_rcu <- rcu_core
         (call_rcu queued by sctp_transport_put at end of sctp_rcv)
  The buggy address is located 364 bytes inside of freed 1024-byte
  region [ffff88800e1e3400, ffff88800e1e3800), cache kmalloc-1k

Note that commit 03a9d10ecf ("sctp: drop a chunk if its transport
was removed") only covers the window between the receive lookup and
the chunk processing (e.g. an ASCONF DEL-IP racing the socket backlog);
here the transport is removed *while* the packet is being processed,
by an earlier chunk of the same packet, so the drop in sctp_inq_push()
does not reach this path.  Verified with the bundled [ERROR(Stale
Cookie)][DATA] + FWD-TSN reproducer: the KASAN report above still
fires with that commit applied, and is gone with this patch on top.

Fix it by discarding the rest of the packet on this path, as suggested
by Xin.  After the stale-cookie ERROR has sent the association back to
COOKIE-WAIT and removed the non-primary transports, the remaining
chunks of the packet can only run against the restarted handshake
while referencing the removed arrival transport through
chunk->transport: besides the last_data_from registration above,
sctp_cmd_setup_t2() and the sctp_make_*() reply builders would also
copy that pointer into association-lifetime state that
sctp_assoc_rm_peer() has already sanitized.  Let the peer retransmit
them, in line with what sctp_inq_push() does for chunks whose
transport was removed before processing.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Suggested-by: Xin Long <lucien.xin@gmail.com>
Reported-by: TencentOS Corvus AI <corvus@tencent.com>
Cc: stable@vger.kernel.org
Signed-off-by: Aohan Mei <henrymei@tencent.com>
Acked-by: Xin Long <lucien.xin@gmail.com>
Link: https://patch.msgid.link/20260921093707.1432184-1-ljp1205831794@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22 18:31:01 -07:00
Wentao Liang
0bf6bb567f net/mlx5: Fix rev_entry reference leak in mlx5_tc_ct_shared_counter_get()
When the reverse entry is found but its counter is already being
released, refcount_inc_not_zero() fails and the reference taken by
mlx5_tc_ct_entry_get() is never dropped before falling through to
create_counter.  Drop it so the reverse entry is not kept alive forever
by a shared counter lookup that did not use it.

Fixes: 1edae2335a ("net/mlx5e: CT: Use the same counter for both directions")
Cc: stable@vger.kernel.org
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Reviewed-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260917113131.2149024-1-vulab@iscas.ac.cn
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22 18:10:17 -07:00
Jakub Kicinski
1c34493df9 Merge branch 'net-mlx5-bridge-fix-remaining-switchdev-ownership-gaps-on-merged-eswitch'
Bernardo Soares says:

====================
net/mlx5: Bridge, fix remaining switchdev ownership gaps on merged eswitch
====================

Link: https://patch.msgid.link/20260918095931.29792-1-bsoares.it@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22 18:09:23 -07:00
Bernardo Soares
2e51097c98 net/mlx5: Bridge, don't fail unlink of untracked/unsupported peer ports
mlx5_esw_bridge_vport_unlink() returns -EINVAL when the port isn't
tracked by this instance's br_offloads. This is reachable on a sibling
instance that registered its notifier after the port was already
enslaved: it never saw the NETDEV_CHANGEUPPER link event, so
peer_link() never created a peer port for it, but it does see the
later unlink event and fails. Return 0 instead, and give
mlx5_esw_bridge_vport_peer_unlink() the same merged_eswitch capability
guard peer_link() already has, since without it peer_link() likewise
never creates a port to unlink.

This also matters beyond the -EINVAL itself:
mlx5_esw_bridge_switchdev_port_event() runs on the per-netns
netdev_chain, and notifier_from_errno(-EINVAL) sets NOTIFY_STOP_MASK,
which call_netdevice_notifiers_info() checks to stop calling further
listeners on that chain - so the old -EINVAL silently dropped the
event for any listener registered later on the same chain, even
though none of it was visible to user space since
__netdev_upper_dev_unlink() discards the return value.

Fixes: c358ea1741 ("net/mlx5: Bridge, allow merged eswitch connectivity")
Signed-off-by: Bernardo Soares <bsoares.it@gmail.com>
Reviewed-by: Mark Bloch <mbloch@nvidia.com>
Link: https://patch.msgid.link/20260918095931.29792-3-bsoares.it@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22 18:09:18 -07:00
Bernardo Soares
35e6f970f5 net/mlx5: Bridge, don't fail switchdev events of sibling eswitch ports
mlx5 registers the bridge offload switchdev notifiers once per eswitch
instance, but the notifier chains are global, so every instance sees
every event and must filter out the ones that aren't its own. The
existing filter, mlx5_esw_bridge_dev_same_hw(), only checks that the
event netdevice sits on the same HCA - intentional for merged eswitch,
where one bridge can span representors of several eswitches on one
HCA - but same-HCA doesn't mean the instance actually has that port:
peer ports are only created reactively from NETDEV_CHANGEUPPER, so an
instance brought up after a sibling PF's port was already enslaved has
none. The port object and attribute handlers claim the event anyway
once same-HW passes, then fail the port lookup and return -EINVAL,
which gets reported to user space even though the owning instance
already handled it (e.g. "bridge vlan add ... RTNETLINK answers:
Invalid argument"). Fix by filtering on the tracked port instead.

The same gap exists in the generic recursive lower-device walk used by
attribute changes on a bridge with more than one representor enslaved
directly: mlx5_esw_bridge_lower_rep_vport_num_vhca_id_get() is entered
with the bridge master netdevice, falls through to its generic
netdev_for_each_lower_dev() loop, and returns as soon as the recursion
into any one lower device yields a non-NULL rep - the underlying base
case, mlx5_esw_bridge_rep_vport_num_vhca_id_get(), only checks
mlx5_esw_bridge_dev_same_hw(), not ownership by the calling instance's
br_offloads. mlx5_esw_bridge_lag_rep_get(), used for the LAG-master
case, already filters on mlx5_esw_bridge_dev_same_esw() per candidate
and so cannot select a sibling's rep; it is not the source of this bug.
On a merged-eswitch HCA with a bridge spanning representors of more
than one eswitch instance directly, the walk can return a sibling's rep
instead of continuing to the one the calling instance actually owns, so
the attribute change fails the same way as above. Fix by checking
mlx5_esw_bridge_port_exists() at the point each rep is picked, same as
the previous fix did for the notifier filter.

Fixes: c358ea1741 ("net/mlx5: Bridge, allow merged eswitch connectivity")
Signed-off-by: Bernardo Soares <bsoares.it@gmail.com>
Cc: Vlad Buslov <vladbu@nvidia.com>
Cc: Saeed Mahameed <saeedm@nvidia.com>
Reviewed-by: Mark Bloch <mbloch@nvidia.com>
Link: https://patch.msgid.link/20260918095931.29792-2-bsoares.it@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22 18:09:17 -07:00
Wentao Liang
17741334d0 net: usb: lan78xx: Fix URB reference leak in lan78xx_submit_deferred_urbs()
usb_get_from_anchor() hands over a reference to the URB, which the caller
must release. lan78xx_submit_deferred_urbs() never does, so every deferred
Tx URB keeps an extra reference: the counter grows on each suspend/resume
cycle and the URBs are never freed when the buffers are released. Drop
the reference after submitting, and on the path that drops the packet
instead of submitting it.

Fixes: 5f4cc6e251 ("lan78xx: Fix race conditions in suspend/resume handling")
Cc: stable@vger.kernel.org
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Link: https://patch.msgid.link/20260917115811.2150119-1-vulab@iscas.ac.cn
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-22 15:54:42 +02:00
Coia Prant
9892d71cf0 net: pcs: xpcs: fix clock reference leak on xpcs_init_clks failure
xpcs_init_clks() takes references with clk_bulk_get_optional() and then
enables them with clk_bulk_prepare_enable(). If the enable step fails,
the function returns without dropping the references.

xpcs_create() handles the failure through out_free_data, which calls
xpcs_free_data() but never xpcs_clear_clks(), so the clk references are
leaked.

Add the missing clk_bulk_put() on the enable failure path. The
prepare/enable side is already rolled back by
clk_bulk_prepare_enable() itself.

Fixes: f6bb3e9d98 ("net: pcs: xpcs: Add Synopsys DW xPCS platform device driver")
Signed-off-by: Coia Prant <coiaprant@gmail.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260919172021.2336748-1-coiaprant@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-22 15:39:06 +02:00
Jakub Kicinski
a87529034b selftests: net: nl_nlctrl: check the op ids in the policy map
Validate that the op map in the policy dump is correct.

Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Link: https://patch.msgid.link/20260918222949.4190284-2-kuba@kernel.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-22 15:21:32 +02:00
Jakub Kicinski
261e8a37ec genetlink: report the real command id for dump-only ops in policy dumps
The op-to-policy map a CTRL_CMD_GETPOLICY dump returns is the only way
for userspace to find out which policy index belongs to which command.
ctrl_dumppolicy_put_op() tags the nest with doit->cmd, but an op which
only has a dumpit has no doit and every path which fills the split ops
in zeroes it out, so those entries all claim to be command 0.  nlctrl's
own CTRL_CMD_GETPOLICY and NETDEV_CMD_QSTATS_GET are both in that group:

  [{'family-id': 16, 'op-policy': {'do': 0, 'dump': 0, 'op-id': 3}},
   {'family-id': 16, 'op-policy': {'dump': 1, 'op-id': 0}},

ctrl_fill_info() gets this right - it uses the iterator's cmd for
CTRL_ATTR_OP_ID - so the two introspection interfaces of the same family
contradict each other today.

Pass the command in rather than reconstructing it from
doit->cmd | dumpit->cmd inside the helper, both callers already have it.

Fixes: 26588edbef ("genetlink: support split policies in ctrl_dumppolicy_put_op()")
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Link: https://patch.msgid.link/20260918222949.4190284-1-kuba@kernel.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-22 15:21:32 +02:00
Théo Lebrun
23d42b9a3b net: macb: fix dma_alloc_coherent() leak on macb_alloc() error paths
Fix 3 leaks in macb_alloc() error paths:
- Tx buffer allocated but crossing a 4G boundary: Tx leaked.
- Rx buffer allocation fails: Tx leaked.
- Rx buffer allocated but crossing a 4G boundary: Tx & Rx leaked.

This is because our error handling calls macb_free(bp) which in turn
frees the buffers stored in bp->queues[0], but nothing has been stored
in there. Fix by storing allocated buffers into bp->queues[0] ASAP.

Fixes: 78d901897b ("net: macb: single dma_alloc_coherent() for DMA descriptors")
Cc: stable@vger.kernel.org
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Link: https://patch.msgid.link/20260918-macb-alloc-leak-v1-1-aba9a3d4f6e3@bootlin.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-22 13:33:50 +02:00
Wentao Liang
999e8295bc net: hisilicon: hns_dsaf_mac: fix mdio device leak in hns_mac_register_phy()
hns_dsaf_find_platform_device() returns the mdio platform device with its
reference count incremented. hns_mac_register_phy() never drops that
reference, so the mdio device can not be released.

Release the reference on both the deferred probe and the normal path.

Fixes: 1d1afa2ebf ("net: hns: register phy device in each mac initial sequence")
Cc: stable@vger.kernel.org
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260917110828.2148390-1-vulab@iscas.ac.cn
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-22 13:13:10 +02:00
bui duc phuc
ac4334522e net: ethernet: ti: netcp: fix pm_runtime usage counter leak on error
pm_runtime_get_sync() leaves the runtime PM usage counter incremented even
when it fails, but the error path in netcp_probe() does not call
pm_runtime_put_noidle() to balance it, leaking a reference each time
resume fails.

Use pm_runtime_resume_and_get() instead, which automatically drops the
usage counter on failure, fixing the leak.

Fixes: 84640e27f2 ("net: netcp: Add Keystone NetCP core ethernet driver")
Signed-off-by: bui duc phuc <phucduc.bui@gmail.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260918042804.13101-1-phucduc.bui@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-22 13:10:01 +02:00
Muhammad Bilal
2d14720beb net: spacemit: clear TX descriptor on fragment mapping failure
emac_tx_mem_map() writes TX_DESC_0_OWN into the ring descriptor for
every slot beyond old_head as soon as that slot's memset()'d local
copy is committed with "*tx_desc_addr = tx_desc", i.e. before the
buffers for that slot have necessarily all been mapped successfully.
If emac_tx_map_frag() then fails on a later fragment, the err_free_skb
path calls emac_free_tx_buf() to unmap and drop the skb, but leaves
the already-written descriptor memory untouched, and tx_ring->head is
never advanced past old_head (the "tx_ring->head = head" store is
skipped by the goto).

So a slot between old_head and the rolled-back head can be left with
TX_DESC_0_OWN set and buffer_addr_{1,2} pointing at DMA mappings that
emac_free_tx_buf() just tore down, while software considers that slot
free again. The next successful emac_tx_mem_map() call only rebuilds
old_head itself; if the DMA engine auto-advances into the following
descriptor once it finishes old_head's packet, it will fetch that
stale, already-unmapped address.

emac_tx_clean_desc() already treats emac_free_tx_buf() and clearing
the descriptor as a pair when reclaiming completed descriptors; do
the same in the mapping failure path.

Fixes: bfec6d7f20 ("net: spacemit: Add K1 Ethernet MAC")
Signed-off-by: Muhammad Bilal <meatuni001@gmail.com>
Reviewed-by: Vivian Wang <wangruikang@iscas.ac.cn>
Reviewed-by: Troy Mitchell <troy.mitchell@linux.spacemit.com>
Link: https://patch.msgid.link/20260919191937.271202-1-meatuni001@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-22 13:02:35 +02:00
Wentao Liang
a644f09b20 fsl/fman: Fix clk reference leak in read_dts_node()
of_clk_get() returns a clock with its reference count incremented, but
read_dts_node() only uses it to read the rate and never calls clk_put().
The clock is not stored anywhere, so the reference cannot be released
later either.

Release the clock once its rate has been read, which also covers the
error path taken when the rate is zero.

Fixes: 414fd46e77 ("fsl/fman: Add FMan support")
Cc: stable@vger.kernel.org
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260917110135.2148068-1-vulab@iscas.ac.cn
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-22 12:31:26 +02:00
Myeonghun Pak
a92e1a412c tg3: clean up PHYLIB resources on probe failure
tg3_get_invariants() can register an MDIO bus and connect a PHY for
USE_PHYLIB devices. If tg3_init_one() later fails, its common error path
releases the mappings and netdev without undoing those PHYLIB resources.

Disconnect the PHY and unregister the MDIO bus before the remaining
teardown. Guard PHY cleanup with USE_PHYLIB to match tg3_phy_init(), and
call tg3_mdio_fini() unconditionally to match tg3_mdio_init(). The existing
IS_CONNECTED and MDIOBUS_INITED flags make both helpers safe when
initialization only completed partially.

This issue was identified during our ongoing static-analysis research while
reviewing kernel code.

Fixes: 158d7abdae ("tg3: Add mdio bus registration")
Assisted-by: OpenAI:GPT-5.6
Co-developed-by: Ijae Kim <ae878000@gmail.com>
Signed-off-by: Ijae Kim <ae878000@gmail.com>
Signed-off-by: Myeonghun Pak <mhun512@gmail.com>
Link: https://patch.msgid.link/20260917183336.36239-1-mhun512@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-22 11:15:00 +02:00
Paolo Abeni
177a59665d Merge branch 'udp-two-fixes-for-the-4-tuple-hash-table'
Shardul Bankar says:

====================
udp: two fixes for the 4-tuple hash table

Two ways a UDP socket ends up in the wrong place in the 4-tuple hash table.
The patches are independent, with different Fixes: tags and no dependency
between them.

Patch 1: a socket that connects a second time is not relocated, so it stays
filed under its first peer's hash and packets for it fall back to scoring
the hash2 chain for its address and port.

Patch 2: a socket bound to a specific address and port is not taken out of
the table when it disconnects, because __udp_disconnect() only does that
via ->rehash() or ->unhash() and neither runs for it.

Both are in code shared by IPv4 and IPv6.

Patch 1's cost, with N sockets sharing a port and one of them misfiled,
200k packets sent to its 4-tuple:

	N	without		with
	200	1,061,652	2,093,259 pps
	500	  522,553	2,055,078 pps
	1000	  279,729	2,136,606 pps

Correctly filed sockets measure ~2.1M pps throughout, so the cost scales
with the number of sockets on the port, as the fallback scan does. For
comparison, commit 78c91ae2c6 ("ipv4/udp: Add 4-tuple hash for connected
socket") measured 290,860 pps without the table and 1,889,658 with it at
500 connected sockets.

Patch 2's cost is not in throughput. Its stale entry keeps hash4_cnt raised
for the life of the socket, so every packet for that address and port is
sent through the 4-tuple lookup first; on IPv6 the entry is also matchable,
because __udp_disconnect() does not clear sk_v6_daddr. That last one is a
separate defect, which I will send on its own.

Neither patch has a selftest. Nothing in tree reports which 4-tuple bucket
a socket is filed under, so a test can only measure the cost indirectly.
What I did instead was add pr_info() to the hash4 paths and a knob that
dumps bucket occupancy, then run the same scenarios on two kernels
differing only by these patches; that is where the numbers above come from.
The instrumentation, the reproducers and the benchmark are at [1]. If
exposing the bucket through diag would be welcome, that would make both
defects testable in tree and I am glad to do it for net-next.

Removing the connect(AF_UNSPEC) limitation described in 644f9108f3 is a
side effect of patch 1 fixing the general case. I can make it narrower if
you would rather that limitation stayed.

Tooling, per Documentation/process/generated-content.rst: this series was
developed in an assisted session with an LLM. The assistant did most of
the code reading, wrote the instrumentation and reproducers behind [1],
drafted these changelogs, and ran the A/B builds and the regression
suites below. Every claim in these messages was checked against the
source, and the IPv6 behaviour described in patch 2 was confirmed at
runtime.

Tested on x86-64, IPv4 and IPv6. No regressions across reuseport_bpf,
reuseport_bpf_cpu, reuseport_addr_any.sh, reuseport_dualstack,
udpgso_bench.sh, udpgro_bench.sh and socket.

[1] https://github.com/shardulsdk-mpiric/linux/tree/udp-hash4-fix-verification

To: Willem de Bruijn <willemdebruijn.kernel@gmail.com>
To: "David S. Miller" <davem@davemloft.net>
To: Eric Dumazet <edumazet@google.com>
To: Jakub Kicinski <kuba@kernel.org>
To: Paolo Abeni <pabeni@redhat.com>
To: Simon Horman <horms@kernel.org>
To: Philo Lu <lulie@linux.alibaba.com>
To: Fred Chen <fred.cc@alibaba-inc.com>
To: Yubing Qiu <yubing.qiuyubing@alibaba-inc.com>
Cc: Kuniyuki Iwashima <kuniyu@google.com>
Cc: Willem de Bruijn <willemb@google.com>
Cc: Cambda Zhu <cambda@linux.alibaba.com>
Cc: Janak Bhatt <janak@mpiric.us>
Cc: Kalpan Jani <kalpan.jani@mpiricsoftware.com>
Cc: Shardul Bankar <shardulsb08@gmail.com>
Cc: netdev@vger.kernel.org
Cc: linux-kernel@vger.kernel.org

Signed-off-by: Shardul Bankar <shardul.b@mpiricsoftware.com>
====================

Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-0-718891af0d7a@mpiricsoftware.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-22 10:58:20 +02:00
Shardul Bankar
9e95b1a94c udp: remove a disconnected socket from the 4-tuple hash table
A UDP socket bound to a specific address and port keeps its entry in the
4-tuple hash table after it is disconnected:

    sk binds to 127.0.0.1:21001
    sk connects to 127.0.0.2:20001	// filed in the 4-tuple table
    sk disconnects, connect(AF_UNSPEC)	// still filed, peer now 0.0.0.0:0

__udp_disconnect() takes a socket out of that table only as a side effect
of ->rehash() or ->unhash(), and it skips ->rehash() when
SOCK_BINDADDR_LOCK is set and ->unhash() when SOCK_BINDPORT_LOCK is set.
commit 6996a2d2d0 ("udp: Unhash auto-bound connected sk from 4-tuple hash
table when disconnected.") fixed the same end state for a wildcard-bound
socket, by a path this one does not take.

The entry is counted whether or not anything hits it. hash4_cnt on the
hash2 slot stays raised for as long as the socket lives, so udp_has_hash4()
keeps sending every packet for that address and port through the 4-tuple
lookup first.

On IPv6 it can also be hit. __udp_disconnect() does not clear sk_v6_daddr,
so udp_v6_rehash() files the entry under the peer the socket was connected
to with a zero dport, and inet6_match() compares that same
field: a datagram from the former peer with a zero source port matches,
and source port zero is accepted on receive. On IPv4 the peer is cleared,
so a match would need a zero source address as well, which the routing
layer rejects as martian. The stale sk_v6_daddr is a separate defect, not
addressed here; removing the entry closes this path either way.

The entry can also be relocated. __udp_disconnect() clears sk_bound_dev_if,
so a subsequent SO_BINDTODEVICE calls ->rehash(), and because the receive
address is still specific udp_lib_rehash() moves the entry instead of
removing it, into the bucket that (rcv_saddr, num, 0, 0) hashes to -- a
pure function of the address and port, so every socket reaching this state
on one address and port collects in one bucket. The bucket cannot be chosen
from outside, as udp_ehashfn() is seeded with a per-boot secret. This last
one became reachable only with commit 644f9108f3 ("udp: Make rehash4
independent in udp_lib_rehash()"), which moved the hash4 handling out of a
branch a disconnected socket does not take; the stale entry itself dates
from the commit in Fixes.

Take the socket out of the table before __udp_disconnect() runs, while it
still matches how it was filed. This also reaches the wildcard case ahead
of udp_lib_rehash()'s udp_unhash4() branch, leaving that branch unreachable
from udp_disconnect(); removing it belongs in net-next. udp_disconnect()
and udp_abort() are the only UDP entries into __udp_disconnect(), which is
shared with raw, ping and l2tp sockets that are not struct udp_sock:
ping_prot.obj_size is sizeof(struct inet_sock), so udp_hashed4() on one
would read past the allocation.

Fixes: 78c91ae2c6 ("ipv4/udp: Add 4-tuple hash for connected socket")
Assisted-by: LLM
Signed-off-by: Shardul Bankar <shardul.b@mpiricsoftware.com>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-2-718891af0d7a@mpiricsoftware.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-22 10:57:53 +02:00