of_clk_get() returns a clock with its reference count incremented, but
read_dts_node() only uses it to read the rate and never calls clk_put().
The clock is not stored anywhere, so the reference cannot be released
later either.
Release the clock once its rate has been read, which also covers the
error path taken when the rate is zero.
Fixes: 414fd46e77 ("fsl/fman: Add FMan support")
Cc: stable@vger.kernel.org
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260917110135.2148068-1-vulab@iscas.ac.cn
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
tg3_get_invariants() can register an MDIO bus and connect a PHY for
USE_PHYLIB devices. If tg3_init_one() later fails, its common error path
releases the mappings and netdev without undoing those PHYLIB resources.
Disconnect the PHY and unregister the MDIO bus before the remaining
teardown. Guard PHY cleanup with USE_PHYLIB to match tg3_phy_init(), and
call tg3_mdio_fini() unconditionally to match tg3_mdio_init(). The existing
IS_CONNECTED and MDIOBUS_INITED flags make both helpers safe when
initialization only completed partially.
This issue was identified during our ongoing static-analysis research while
reviewing kernel code.
Fixes: 158d7abdae ("tg3: Add mdio bus registration")
Assisted-by: OpenAI:GPT-5.6
Co-developed-by: Ijae Kim <ae878000@gmail.com>
Signed-off-by: Ijae Kim <ae878000@gmail.com>
Signed-off-by: Myeonghun Pak <mhun512@gmail.com>
Link: https://patch.msgid.link/20260917183336.36239-1-mhun512@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Shardul Bankar says:
====================
udp: two fixes for the 4-tuple hash table
Two ways a UDP socket ends up in the wrong place in the 4-tuple hash table.
The patches are independent, with different Fixes: tags and no dependency
between them.
Patch 1: a socket that connects a second time is not relocated, so it stays
filed under its first peer's hash and packets for it fall back to scoring
the hash2 chain for its address and port.
Patch 2: a socket bound to a specific address and port is not taken out of
the table when it disconnects, because __udp_disconnect() only does that
via ->rehash() or ->unhash() and neither runs for it.
Both are in code shared by IPv4 and IPv6.
Patch 1's cost, with N sockets sharing a port and one of them misfiled,
200k packets sent to its 4-tuple:
N without with
200 1,061,652 2,093,259 pps
500 522,553 2,055,078 pps
1000 279,729 2,136,606 pps
Correctly filed sockets measure ~2.1M pps throughout, so the cost scales
with the number of sockets on the port, as the fallback scan does. For
comparison, commit 78c91ae2c6 ("ipv4/udp: Add 4-tuple hash for connected
socket") measured 290,860 pps without the table and 1,889,658 with it at
500 connected sockets.
Patch 2's cost is not in throughput. Its stale entry keeps hash4_cnt raised
for the life of the socket, so every packet for that address and port is
sent through the 4-tuple lookup first; on IPv6 the entry is also matchable,
because __udp_disconnect() does not clear sk_v6_daddr. That last one is a
separate defect, which I will send on its own.
Neither patch has a selftest. Nothing in tree reports which 4-tuple bucket
a socket is filed under, so a test can only measure the cost indirectly.
What I did instead was add pr_info() to the hash4 paths and a knob that
dumps bucket occupancy, then run the same scenarios on two kernels
differing only by these patches; that is where the numbers above come from.
The instrumentation, the reproducers and the benchmark are at [1]. If
exposing the bucket through diag would be welcome, that would make both
defects testable in tree and I am glad to do it for net-next.
Removing the connect(AF_UNSPEC) limitation described in 644f9108f3 is a
side effect of patch 1 fixing the general case. I can make it narrower if
you would rather that limitation stayed.
Tooling, per Documentation/process/generated-content.rst: this series was
developed in an assisted session with an LLM. The assistant did most of
the code reading, wrote the instrumentation and reproducers behind [1],
drafted these changelogs, and ran the A/B builds and the regression
suites below. Every claim in these messages was checked against the
source, and the IPv6 behaviour described in patch 2 was confirmed at
runtime.
Tested on x86-64, IPv4 and IPv6. No regressions across reuseport_bpf,
reuseport_bpf_cpu, reuseport_addr_any.sh, reuseport_dualstack,
udpgso_bench.sh, udpgro_bench.sh and socket.
[1] https://github.com/shardulsdk-mpiric/linux/tree/udp-hash4-fix-verification
To: Willem de Bruijn <willemdebruijn.kernel@gmail.com>
To: "David S. Miller" <davem@davemloft.net>
To: Eric Dumazet <edumazet@google.com>
To: Jakub Kicinski <kuba@kernel.org>
To: Paolo Abeni <pabeni@redhat.com>
To: Simon Horman <horms@kernel.org>
To: Philo Lu <lulie@linux.alibaba.com>
To: Fred Chen <fred.cc@alibaba-inc.com>
To: Yubing Qiu <yubing.qiuyubing@alibaba-inc.com>
Cc: Kuniyuki Iwashima <kuniyu@google.com>
Cc: Willem de Bruijn <willemb@google.com>
Cc: Cambda Zhu <cambda@linux.alibaba.com>
Cc: Janak Bhatt <janak@mpiric.us>
Cc: Kalpan Jani <kalpan.jani@mpiricsoftware.com>
Cc: Shardul Bankar <shardulsb08@gmail.com>
Cc: netdev@vger.kernel.org
Cc: linux-kernel@vger.kernel.org
Signed-off-by: Shardul Bankar <shardul.b@mpiricsoftware.com>
====================
Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-0-718891af0d7a@mpiricsoftware.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
A UDP socket bound to a specific address and port keeps its entry in the
4-tuple hash table after it is disconnected:
sk binds to 127.0.0.1:21001
sk connects to 127.0.0.2:20001 // filed in the 4-tuple table
sk disconnects, connect(AF_UNSPEC) // still filed, peer now 0.0.0.0:0
__udp_disconnect() takes a socket out of that table only as a side effect
of ->rehash() or ->unhash(), and it skips ->rehash() when
SOCK_BINDADDR_LOCK is set and ->unhash() when SOCK_BINDPORT_LOCK is set.
commit 6996a2d2d0 ("udp: Unhash auto-bound connected sk from 4-tuple hash
table when disconnected.") fixed the same end state for a wildcard-bound
socket, by a path this one does not take.
The entry is counted whether or not anything hits it. hash4_cnt on the
hash2 slot stays raised for as long as the socket lives, so udp_has_hash4()
keeps sending every packet for that address and port through the 4-tuple
lookup first.
On IPv6 it can also be hit. __udp_disconnect() does not clear sk_v6_daddr,
so udp_v6_rehash() files the entry under the peer the socket was connected
to with a zero dport, and inet6_match() compares that same
field: a datagram from the former peer with a zero source port matches,
and source port zero is accepted on receive. On IPv4 the peer is cleared,
so a match would need a zero source address as well, which the routing
layer rejects as martian. The stale sk_v6_daddr is a separate defect, not
addressed here; removing the entry closes this path either way.
The entry can also be relocated. __udp_disconnect() clears sk_bound_dev_if,
so a subsequent SO_BINDTODEVICE calls ->rehash(), and because the receive
address is still specific udp_lib_rehash() moves the entry instead of
removing it, into the bucket that (rcv_saddr, num, 0, 0) hashes to -- a
pure function of the address and port, so every socket reaching this state
on one address and port collects in one bucket. The bucket cannot be chosen
from outside, as udp_ehashfn() is seeded with a per-boot secret. This last
one became reachable only with commit 644f9108f3 ("udp: Make rehash4
independent in udp_lib_rehash()"), which moved the hash4 handling out of a
branch a disconnected socket does not take; the stale entry itself dates
from the commit in Fixes.
Take the socket out of the table before __udp_disconnect() runs, while it
still matches how it was filed. This also reaches the wildcard case ahead
of udp_lib_rehash()'s udp_unhash4() branch, leaving that branch unreachable
from udp_disconnect(); removing it belongs in net-next. udp_disconnect()
and udp_abort() are the only UDP entries into __udp_disconnect(), which is
shared with raw, ping and l2tp sockets that are not struct udp_sock:
ping_prot.obj_size is sizeof(struct inet_sock), so udp_hashed4() on one
would read past the allocation.
Fixes: 78c91ae2c6 ("ipv4/udp: Add 4-tuple hash for connected socket")
Assisted-by: LLM
Signed-off-by: Shardul Bankar <shardul.b@mpiricsoftware.com>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-2-718891af0d7a@mpiricsoftware.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
A connected UDP socket that connects again to a different peer is not
re-filed in the 4-tuple hash table:
sk binds to 127.0.0.1:21001
sk connects to 127.0.0.2:20001 // filed under hash(sk, peer1)
sk connects to 127.0.0.3:20002 // still filed under hash(sk, peer1)
packet from 127.0.0.3:20002 // hash(sk, peer2) misses, so the
// lookup falls back to scoring the
// hash2 chain for this address
// and port
udp_lib_hash4() returns early when the socket is already hashed, assuming
->rehash() relocates it. ->rehash() runs from __ip{4,6}_datagram_connect()
only while the receive address is unset, which a second connect never is:
the first connect assigns it, whether the socket was bound to a specific
address or to the wildcard. commit 644f9108f3 ("udp: Make rehash4
independent in udp_lib_rehash()") added that early return and named
connect(AF_UNSPEC) as the way around it. That workaround does not help a
socket with both SOCK_BINDADDR_LOCK and SOCK_BINDPORT_LOCK set, because
__udp_disconnect() skips ->rehash() for the first and ->unhash() for the
second.
Delivery is correct either way.
Relocate the socket when the hash it is filed under differs from the one
requested, which is what commit 78c91ae2c6 ("ipv4/udp: Add 4-tuple hash
for connected socket") did before the early return became unconditional. It
is done here under hslot->lock, which that version did not take, to match
udp_lib_rehash() and udp_lib_unhash(). hslot2 is unchanged, so hash4_cnt
needs no adjustment, as in udp_lib_rehash(). A first connect is unaffected,
and IPv6 shares the code.
With 500 sockets on the port, a re-connected socket measured 522,553 pps
without this change and 2,055,078 with it. The UDP side was noted as
remaining work in [1].
Link: https://lore.kernel.org/netdev/apnHqmYZQ4yzOP4N@v4bel/ [1]
Fixes: 644f9108f3 ("udp: Make rehash4 independent in udp_lib_rehash()")
Assisted-by: LLM
Signed-off-by: Shardul Bankar <shardul.b@mpiricsoftware.com>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-1-718891af0d7a@mpiricsoftware.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Core:
- hci_conn: fix CIS hold ownership on reuse
- hci_sock: reject out-of-range OCF values
- hci_sock: validate event length before filtering
- L2CAP: validate frame length before control and FCS access
- RFCOMM: fix NULL dereference of dlc->session in RFCOMM_CONNINFO
- RFCOMM: Reject short EA=0 frames in rfcomm_recv_frame()
- ISO: release unused CIS holds after channel attach
- ISO: balance the parent hold in hci_bind_bis()
- SMP: reject Security Request over BR/EDR
- MGMT: fix race in read_unconf_index_list()
- MGMT: Dequeue pending mesh_send_sync entries on cancel
- BNEP: fix out-of-bounds reads on short RX/TX frames and control fallthrough
Drivers:
- btintel_pcie: validate device-supplied DMA indices
- btnxpuart: Fix skb leak in nxp_process_fw_dump()
-----BEGIN PGP SIGNATURE-----
iQJNBAABCgA3FiEE7E6oRXp8w05ovYr/9JCA4xAyCykFAmqxN44ZHGx1aXoudm9u
LmRlbnR6QGludGVsLmNvbQAKCRD0kIDjEDILKS7mD/43nt9IlCMp4fRi5eT3iV6J
phP/zJTiikgOMv87kTI0Q9OXY8Xl1nhIGrTiypXQIJNJGTW/OTHtMF+N55pA7pF4
exD0US01bKUcuopztHeP1Yk08CMKAtE9VTPLi/PdAQz6KTo7wcH7yFwfsO6avIk4
T/IXi9b8iKPBMKizjgUe8Uo6wrFoByD/o2VwTSQwtZOgWppUkCvVKKx10Y8EYJRG
UbDS8rpPtWN3t0fK5F4yVfkTP+9sA7uWb9bF2UoLcmov2ie33M50zf5giVofiNrX
pT6RWTCPfjkPdDro66l6wV4M8prEUVok0QhlqeYOUXSZJSOTqqXBTOorioABSeHz
i4QWH+KsqUFbaOJQuoF3v5jAccY8ZxKczRODK8Xt1+bY9O73lqpuWjQvoi98E/7i
vNCSU/bDpRS0MjYDTgQTG9c0rv9qxKI0GBC8gDVI78/nAzUIxdJ09/wDBm17aUpt
uxaBwYE1lhJwT0Lf+dwJrwVuTm5q5XotdRIsOwxZgfF25Ewk8ygUlEsQu/ioHmS+
q+9SL/mHb/wPI7j0YeVG2l7cHiHOZDEn0FW++WlgtF1Oj/HWUsz2UN/qdmpYZo1h
rGX7D2DdZQo8IOEmZZn7wqLNOxjPjFVkxgQtgWd5OhE0O+7/e2H4mNeDJyWs6rJp
lT1BJZnPQTFtwP41x7Hxng==
=cLTc
-----END PGP SIGNATURE-----
Merge tag 'for-net-2026-09-21' of git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth
Luiz Augusto von Dentz says:
====================
bluetooth pull request for net:
Core:
- hci_conn: fix CIS hold ownership on reuse
- hci_sock: reject out-of-range OCF values
- hci_sock: validate event length before filtering
- L2CAP: validate frame length before control and FCS access
- RFCOMM: fix NULL dereference of dlc->session in RFCOMM_CONNINFO
- RFCOMM: Reject short EA=0 frames in rfcomm_recv_frame()
- ISO: release unused CIS holds after channel attach
- ISO: balance the parent hold in hci_bind_bis()
- SMP: reject Security Request over BR/EDR
- MGMT: fix race in read_unconf_index_list()
- MGMT: Dequeue pending mesh_send_sync entries on cancel
- BNEP: fix out-of-bounds reads on short RX/TX frames and control fallthrough
Drivers:
- btintel_pcie: validate device-supplied DMA indices
- btnxpuart: Fix skb leak in nxp_process_fw_dump()
* tag 'for-net-2026-09-21' of git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth:
Bluetooth: RFCOMM: Reject short EA=0 frames in rfcomm_recv_frame()
Bluetooth: RFCOMM: fix NULL dereference of dlc->session in RFCOMM_CONNINFO
Bluetooth: btintel_pcie: validate device-supplied DMA indices
Bluetooth: bnep: fix out-of-bounds reads on short RX/TX frames and control fallthrough
Bluetooth: mgmt: fix race in read_unconf_index_list()
Bluetooth: L2CAP: validate frame length before control and FCS access
Bluetooth: ISO: balance the parent hold in hci_bind_bis()
Bluetooth: hci_sock: validate event length before filtering
Bluetooth: hci_sock: reject out-of-range OCF values
Bluetooth: ISO: release unused CIS holds after channel attach
Bluetooth: hci_conn: fix CIS hold ownership on reuse
Bluetooth: mgmt: Dequeue pending mesh_send_sync entries on cancel
Bluetooth: btnxpuart: Fix skb leak in nxp_process_fw_dump()
Bluetooth: SMP: reject Security Request over BR/EDR
====================
Link: https://patch.msgid.link/20260921135807.3459373-1-luiz.dentz@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
ip6_route_output_flags(), ip6_rt_put_flags(), and ip6_dst_check()
detect an uncached route by list_empty(&rt->dst.rt_uncached),
which replaced the static DST_NOCACHE flag check in commit
a4c2fd7f78 ("net: remove DST_NOCACHE flag").
When a device is unregistered, rt6_uncached_list_flush_dev()
unlinks uncached routes tied to the device from rt6_uncached_list.
Previously, they were moved to another list with list_move()
(__list_del_entry() + list_add()), and since commit 98aa546af5
("inet: remove (struct uncached_list)->quarantine"), the routes
are just unlinked with list_del_init().
If list_del_init() runs concurrently, list_empty() evaluates to
true; ip6_route_output_flags() calls dst_hold_safe() incorrectly
and ip6_rt_put_flags() skips ip6_rt_put(), leaking dst, and thus
dev tied via rt->from as well.
The same race is partially fixed by commit 9a6f0c4d57 ("dst:
fix races in rt6_uncached_list_del() and rt_del_uncached_list()").
Let's check rt6->dst.rt_uncached_list instead.
Note that IPv4 does not have the same issue.
Fixes: 98aa546af5 ("inet: remove (struct uncached_list)->quarantine")
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Reviewed-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260920191558.2990636-1-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
IPv6 XFRM policies may use different source and destination prefix
lengths. mlx5e_ipsec_policy_mask() builds the corresponding masks
independently, but setup_fte_addr6() installs each mask in the opposite
address field.
When the prefix lengths differ, this makes the source match use the
destination prefix and the destination match use the source prefix. The
resulting hardware rule can both miss traffic covered by the policy and
match traffic outside it.
Install each mask in its corresponding match field.
Fixes: ca7992f52c ("net/mlx5e: Properly match IPsec subnet addresses")
Cc: stable@vger.kernel.org
Signed-off-by: Andrea Parri <parri.andrea@gmail.com>
Reviewed-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260917115542.177675-1-parri.andrea@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Commit 8d75c338f0 ("sysctl: remove CONFIG_PROC_SYSCTL, it just mirrors
CONFIG_SYSCTL") renamed CONFIG_PROC_SYSCTL to CONFIG_SYSCTL in place,
which left the entry out of alphabetical order in the net and
packetdrill configs. The netdev CI check for sorted selftest configs
now fails for every patch that touches either file.
Fixes: 8d75c338f0 ("sysctl: remove CONFIG_PROC_SYSCTL, it just mirrors CONFIG_SYSCTL")
Signed-off-by: Yuya Kusakabe <yuya.kusakabe@gmail.com>
Reviewed-by: Joel Granados <joel.granados@kernel.org>
Link: https://patch.msgid.link/20260918-selftests-net-config-sort-v1-1-968ea6e8c1b7@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The receive fixup subtracts the Ethernet CRC from the reported packet
length, but compares that payload length against the whole remaining
receive buffer. The following copy starts after the three-byte header,
and the cursor advance consumes both that header and the four-byte CRC.
Require the payload to fit after SR_RX_OVERHEAD before copying it or
advancing to the next packet. The loop already ensures that the
remaining buffer is larger than the overhead, so the subtraction is
safe.
The issue was found by our static-analysis tool.
Fixes: c9b37458e9 ("USB2NET : SR9700 : One chip USB 1.1 USB2NET SR9700Device Driver Support")
Reviewed-by: Ethan Nelson-Moore <enelsonmoore@gmail.com>
Tested-by: Ethan Nelson-Moore <enelsonmoore@gmail.com>
Signed-off-by: Pengpeng Hou <hppiscas@163.com>
Link: https://patch.msgid.link/20260920034745.18468-1-hppiscas@163.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
dpll_pin_ref_sync_state_set() looks up the reference sync pin in the
pin->ref_sync_pins xarray, which is keyed by the sync pin's id (see
dpll_pin_ref_sync_pair_add() using xa_insert() with ref_sync_pin->id).
The pin id to operate on is supplied by userspace via DPLL_A_PIN_ID.
The lookup however used xa_find() with a ULONG_MAX limit, which returns
the first present entry with an index greater than or equal to the
requested id, not the entry stored exactly at that id. If userspace
passes an id that is not paired as a reference sync pin, but another
pin with a higher id is present in the xarray, xa_find() silently
returns that wrong pin and the subsequent ref_sync_set() operates on
it. The request only fails when the given id is larger than every
present key.
Use xa_load() for an exact-key lookup instead, mirroring the deletion
path in dpll_pin_ref_sync_pair_del().
Fixes: 58256a26bf ("dpll: add reference sync get/set")
Signed-off-by: Ivan Vecera <ivecera@redhat.com>
Link: https://patch.msgid.link/20260917143736.526221-1-ivecera@redhat.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Removing the legacy ioctl fallback made both hwtstamp NDOs mandatory. A
device that only timestamps in its PHY implements neither, so
SIOCSHWTSTAMP fails with EOPNOTSUPP before anything looks at the PHY and
PTP stops working there.
The check only ever picked the legacy path. That path is gone, so drop it
and test where the NDOs are actually called.
SIOCGHWTSTAMP is new here, not restored. The old path went through
phy_mii_ioctl(), which only handled SIOCSHWTSTAMP.
Such a device now returns -ENODEV while absent instead of -EOPNOTSUPP,
like the ones that do implement the NDOs.
Fixes: 5062245a5a ("net: remove legacy way to get/set HW timestamp config")
Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
Reviewed-by: Kory Maincent <kory.maincent@bootlin.com>
Link: https://patch.msgid.link/20260918095540.34286-1-nb@tipi-net.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Before the cited commit, fib6_nh_flush_exceptions() always set
from->exception_bucket_flushed = 1 under rt6_exception_lock to
prevent rt6_insert_exception() from inserting a new exception
for a dying fib6_info.
The flag was replaced with the FIB6_EXCEPTION_BUCKET_FLUSHED
bit stored in nh->rt6i_exception_bucket.
The problem is that now the bit is only set when the bucket
is not NULL and fib6_nh_flush_exceptions() is called from
fib6_nh_release() after fib6_ref has already reached zero.
If rt6_insert_exception() is called while the target fib6_info
is being removed via fib6_purge_rt(), a new exception could be
created successfully because rt6_flush_exceptions() no longer
sets the bit.
This creates a reference cycle between the fib6_info and the
exception route, leaking the fib6_info, its nexthop device,
and all per-CPU routes in fib6_nh->rt6i_pcpu, which stalls netdev
unregistration.
[ 34.680602] unregister_netdevice: waiting for gre6 to become free. Usage count = 68
[ 44.920675] unregister_netdevice: waiting for gre6 to become free. Usage count = 68
[ 55.176582] unregister_netdevice: waiting for gre6 to become free. Usage count = 68
Let's call fib6_drop_pcpu_from() before rt6_flush_exceptions(),
to set fib6_destroying before rt6_exception_lock, and check
f6i->fib6_destroying in rt6_insert_exception().
Note that FIB6_EXCEPTION_BUCKET_FLUSHED logic is dead and
we can clean it up in net-next.
Fixes: cc5c073a69 ("ipv6: Move exception bucket to fib6_nh")
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260918082209.2853582-1-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The original code used DMA_ATTR_FORCE_CONTIGUOUS, which could exhaust
the CMA pool when a large number of VFs were requested.
Fix this by switching to the DMA streaming API. This is equivalent on
Octeon platforms, which provide full I/O coherency via the SMMU.
Cc: Leon Romanovsky <leon@kernel.org>
Fixes: 73d33dbc07 ("octeontx2-af: Use DMA_ATTR_FORCE_CONTIGUOUS attribute in DMA alloc")
Signed-off-by: Ratheesh Kannoth <rkannoth@marvell.com>
Reviewed-by: Leon Romanovsky <leon@kernel.org>
Link: https://patch.msgid.link/20260916022111.1083017-1-rkannoth@marvell.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
-----BEGIN PGP SIGNATURE-----
iQIzBAABCgAdFiEEjF9xRqF1emXiQiqU1w0aZmrPKyEFAmqtHzMACgkQ1w0aZmrP
KyGtkQ//bMKQGEudQKCMCtQmPqaHyW1ajmAo17aumfczaE/nSjqrsGY5Ahw7OKlQ
4otcdPvI4qpV9sLTg41KFaIHIC5sozxt4Q3m3RNB2TbyCkGn9xpSpZxM5IpHybvE
83tVjSA0wpfIqxBEqKUqk8Z9AXtBLo/JocdfYry+6JUyj4PM76X2ViKpzaPbpoMU
1mndfLAYtADIIvs3805CmfdJmOkoSV6XCEsiNutPrJhiRfN4xJZ9leP9xb1zA0IQ
cnqiaw1xkTcFyWCicu4MqOkEALRknr9SL2yX1S9wx5Q6WHwU9JXUeQTlvfv7OoVP
uxuMlNr3WcbwHC9e1GfOHapzjrYgnvEe2Z79i2GFh51Ci+5L9Yr9XCQ/fc6G5NNZ
3W52kh35s3lXq32hll9Tkr7pf4cKLBA+IAJ19VNlRfMrPB0cz4EqbIZ6xNNuLqdh
DbEb3VgTT2dHwuGxEshJVmSfzfR+VeHBG2ZRlRmZElfhViHEwgPaAxkaJhNpPyub
qmHbZCXK0BVp/UrGHDm5rmHJtdkwprXY9YceZBRfW8Fr2Ler4rWQvy+uo9sRRFgN
oF9B6qzSl78THBv3UDB3U+aWuDv0I+VlDub0DKf0k9iWg8OoMlM9yw6OE9ULHt3m
LRvSAfF0LnF/z7byzumUYMVjpVuIVueDzasGrEZjWbHHE7MCsp0=
=RfSQ
-----END PGP SIGNATURE-----
Merge tag 'nf-26-09-18' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf
Pablo Neira Ayuso says:
====================
Netfilter/IPVS fixes for net
The following patchset contains Netfilter/IPVS fixes for net, they are:
1) Set on HW_DEAD after HW_PENDING is cleared in the flowtable offload
to ensure GC does not zap it, from Jérémy Jean.
2) Hold the nfnetlink_queue mutex while removing the queue instance
from the netlink notifier that handles NETLINK_URELEASE to fix a
possible race with the UNBIND command. From Florian Westphal.
3) Reject route with NULL rt6i_idev in ip6t_rpfilter. From Weiming Shi.
4) Reject rtinfo->addrnr set to zero from ip6t_rt .checkentry path.
This also fortifies the datapath loop as per Florian's request.
From Luxiao Xu.
5) Fix checksuming in nft_synproxy for IPv6, from Karl Mehltretter.
6) Revalidate ihl before calling icmp_send() in IPVS,
from Julian Anastasov.
7) Fix suspicious RCU usage splat in ctnetlink with expectations.
8) Check for expired catchall elements in the insert and deactivate
path. From Aohan Mei.
* tag 'nf-26-09-18' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf:
netfilter: nf_tables: skip expired catchall elements on insert and delete
netfilter: ctnetlink: fix suspicious RCU usage in expect_iter_name
ipvs: revalidate ihl before icmp_send
netfilter: nft_synproxy: use the family-aware checksum helper
netfilter: ip6t_rt: fix zero-address non-strict match out-of-bounds read
netfilter: ip6t_rpfilter: reject routes without inet6_dev
netfilter: nfnetlink_queue: hold nfnl mutex in event notifier
netfilter: flowtable: publish HW_DEAD after worker is done
====================
Link: https://patch.msgid.link/20260918112844.194503-1-pablo@netfilter.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
While rfcomm_recv_frame() verifies that skb->len is at least
sizeof(*hdr) + 1 (4 bytes: 3-byte header + 1-byte FCS), an RFCOMM frame
with an extended 2-byte length field (!__test_ea(hdr->len)) has a 4-byte
header plus a 1-byte FCS (5 bytes minimum, sizeof(*hdr) + 2).
When a 4-byte RFCOMM frame with EA == 0 arrives:
1. The initial skb->len < sizeof(*hdr) + 1 check passes (4 < 4 is false).
2. Trimming the FCS byte decrements skb->len to 3.
3. If __check_fcs() succeeds, skb_pull(skb, 4) fails (4 > 3) and returns
NULL without advancing skb->data.
4. Because the return value of skb_pull() is ignored, the un-pulled
3-byte struct rfcomm_hdr remains at skb->data and is either queued as
application payload via rfcomm_recv_data() or parsed as a multiplexer
control command via rfcomm_recv_mcc() on DLCI 0.
Fix this by extending the length check in rfcomm_recv_frame() to also
require skb->len >= sizeof(*hdr) + 2 when !__test_ea(hdr->len).
Fixes: b230e5bf50 ("Bluetooth: RFCOMM: validate skb length in rfcomm_recv_frame")
Assisted-by: LLM
Signed-off-by: Hui Peng <benquike@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
The RFCOMM_CONNINFO getsockopt handler accepts a socket that is not
connected as long as deferred setup is enabled:
if (sk->sk_state != BT_CONNECTED &&
!rfcomm_pi(sk)->dlc->defer_setup) {
err = -ENOTCONN;
break;
}
l2cap_sk = rfcomm_pi(sk)->dlc->session->sock->sk;
dlc->defer_setup is set in rfcomm_sock_init() when rfcomm_connect_ind()
creates a child socket for an incoming connection on a listening socket
that has BT_DEFER_SETUP enabled. It is never cleared afterwards. The
session, however, can go away underneath it.
rfcomm_recv_disc() forces the dlc state before tearing it down:
d->state = BT_CLOSED;
__rfcomm_dlc_close(d, err);
The RFCOMM_DEFER_SETUP early return in __rfcomm_dlc_close() only covers
BT_CONNECT, BT_CONFIG, BT_OPEN and BT_CONNECT2, so with the state
already BT_CLOSED that switch does not match and the function falls
through to rfcomm_dlc_unlink(), which sets d->session = NULL, while
d->defer_setup stays 1.
A getsockopt(SOL_RFCOMM, RFCOMM_CONNINFO) on the accepted socket after
that point therefore skips the -ENOTCONN path -- sk->sk_state is
BT_CLOSED, but dlc->defer_setup is still set -- and dereferences the
NULL session. No race is needed: once the DISC has been processed, the
dereference is unconditional.
Reproduced on a KASAN kernel under QEMU with a BR/EDR peer emulated over
/dev/vhci: the peer brings up an ACL link, opens L2CAP on the RFCOMM
PSM, starts a session and sends SABM for a channel bound with
BT_DEFER_SETUP, and sends DISC for that dlci after the socket has been
accepted. getsockopt(SOL_RFCOMM, RFCOMM_CONNINFO) on the accepted
socket then hits:
Oops: general protection fault, probably for non-canonical address
0xdffffc0000000002: 0000 [#1] SMP KASAN PTI
KASAN: null-ptr-deref in range [0x0000000000000010-0x0000000000000017]
CPU: 1 UID: 0 PID: 150 Comm: init Tainted: G B 7.3.0-rc3-g5dd1818b15d9
Hardware name: QEMU Standard PC (i440FX + PIIX, 1996)
RIP: 0010:rfcomm_sock_getsockopt+0x529/0x780
Call Trace:
<TASK>
do_sock_getsockopt+0x3ad/0x7d0
__sys_getsockopt+0x10e/0x1b0
__x64_sys_getsockopt+0xc2/0x160
do_syscall_64+0xda/0x4b0
entry_SYSCALL_64_after_hwframe+0x77/0x7f
</TASK>
0x10 is the offset of sock in struct rfcomm_session;
rfcomm_sock_getsockopt_old() is inlined into rfcomm_sock_getsockopt().
Commit 43a556b2fd ("Bluetooth: RFCOMM: take rfcomm_mutex for the
deferred setup accept") fixed the same "a remote DISC clears the session
while deferred setup is still flagged" problem in rfcomm_dlc_accept();
this is the remaining instance of it, in the getsockopt path.
Deferred setup only leaves a socket usable here once it has reached
BT_CONNECT2, so restrict the exception to that state and check that a
session is actually present before following it.
Fixes: bb23c0ab82 ("Bluetooth: Add support for deferring RFCOMM connection setup")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Hui Peng <benquike@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
In btintel_pcie_msix_rx_handle(), the driver processes RX completion
descriptors (urbd1) written by the PCIe device into DMA-coherent memory.
urbd1->frbd_tag (a 16-bit field fully controlled by the device firmware
via DMA) is used directly as an array index into rxq->bufs[] without any
bounds check. rxq->bufs[] has only BTINTEL_PCIE_RX_DESCS_COUNT (64)
entries, while frbd_tag can be any value 0-65535. A malicious or
malfunctioning device can write an out-of-range frbd_tag, causing the
driver to dereference an out-of-bounds data_buf pointer.
Additionally, cr_hia is read from a DMA-shared index array also writable
by the device; if the device sets cr_hia >= rxq->count, the while-loop
never terminates because cr_tia is wrapped via modulo rxq->count and can
never equal an out-of-range cr_hia.
Add bounds validation for cr_hia and frbd_tag in the RX path, and cr_hia
in the TX path. Log invalid values with bt_dev_err before returning.
Fixes: c2b636b3f7 ("Bluetooth: btintel_pcie: Add support for PCIe transport")
Signed-off-by: Ravindra <ravindra@intel.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
Fix multiple out-of-bounds reads in Bluetooth BNEP frame processing:
1. In bnep_rx_frame() and bnep_ctrl_frame() (net/bluetooth/bnep/core.c),
use pskb_may_pull() to verify the BNEP header, control type byte,
filter count, and extension headers exist before reading them, and
return 0 after handling BNEP_CONTROL instead of falling through to
Ethernet frame submission when no extension headers follow.
2. In bnep_net_xmit() (net/bluetooth/bnep/netdev.c), verify skb->len >=
ETH_HLEN with pskb_may_pull() before reading the 14-byte Ethernet
header to prevent an out-of-bounds heap read and infoleak on short
AF_PACKET TX frames.
Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Assisted-by: LLM
Signed-off-by: Hui Peng <benquike@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
The classify-loop fix bounds a walk's non-descending hops, so the guard
must not misfire on a legal walk that reaches its leaf through a
level-drift lateral chain. Add a case that builds exactly that chain and
asserts traffic still reaches the chain's own leaf.
A lateral hop can only exist because a bind was legal when it was made
and a later class add raised the target's level, so the setup binds each
hop while the target is still a leaf and only then deepens it: bind
1:1 -> 1:2 while 1:2 is a leaf, add 1:20 under 1:2, add 1:3 and bind
1:2 -> 1:3 while 1:3 is a leaf, then add 1:30 and 1:31 under 1:3 and
bind 1:3 -> 1:31. The walk root -> 1:1 -> 1:2 -> 1:3 -> 1:31 then takes
two lateral hops and must reach leaf 1:31.
The default class is 1:30, distinct from the asserted leaf, and the
verify pattern is anchored to the 1:31 stats line, so neither a
fall-through to the default nor a nonzero count on another class can
satisfy the check. On the patched kernel the test passes; with the bound
forced to zero the walk falls to the default and 1:31 stays idle, so the
test fails.
Reviewed-by: Victor Nogueira <victor@mojatatu.com>
Tested-by: hybris <hybris@mojatatu.ai>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-CTUU.v3.20260916184908@mojatatu.com.2
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
hfsc_classify() applies the "filter may only point downwards" level check
only when the filter result carries no bound class. A filter created with
a flowid gets res.class set once at bind time, so the check never runs for
it during classification. hfsc_adjust_levels() can later raise a class's
level without revalidating existing bindings, leaving two binds that were
each legal at bind time pointing at each other; the classify walk then
bounces between two interior classes forever with the qdisc lock held and
BH disabled — a soft lockup from a single packet. The stuck walk trips
the watchdog:
watchdog: BUG: soft lockup - CPU#3 stuck for 13s! [ping:444]
RIP: 0010:u32_classify+0x542/0x17f0
...
tcf_classify+0x66/0xa0
hfsc_enqueue+0x166/0xdf0
Bound the traversal with a budget of non-descending hops, the only way a
configured walk can move without descending the class tree once levels
drift after bind time. The budget is cumulative over the whole walk and
is deliberately not reset on a descending hop: a chain that alternates a
descent with a lateral hop would return the budget every lap and never
trip. Descending hops never decrement it, so legitimately deep trees are
unaffected and a terminating lateral chain still classifies normally.
Drop the packet with a rate-limited warning once the budget is exhausted,
mirroring the merged HTB fix.
This is a follow-up to commit 729c4896ab ("net/sched: sch_htb: limit
htb_classify inner-class filter hops"), which bounded the same classify
loop on the HTB side but left the HFSC walk unbounded.
Conditions to recreate the bug:
- CONFIG_NET_SCHED, CONFIG_NET_SCH_HFSC, CONFIG_NET_CLS_U32,
CONFIG_LOCKUP_DETECTOR.
- Build a cycle with two legal-at-bind-time flowid binds and a level
drift: class X 1:1 (child of root) with leaf child 1:10; class Y 1:2
(sibling of X) with children 1:20 and 1:200; root u32 filter flowid
1:1; filter on X flowid 1:2 (legal when Y is a leaf); after Y's level
rises to 2, filter on Y flowid 1:1 (legal then). Send one packet (ping
on the device). Unfixed kernel: classify spins with the qdisc lock
held; with softlockup_panic=1 it panics.
- Reachable from unprivileged user via unshare -Urn (CAP_NET_ADMIN).
Fixes: a2f7922713 ("net_sched: sch_hfsc: fix classification loops")
Reported-by: Sashiko (gemini + nipa) <sashiko-bot@kernel.org>
Closes: https://lore.kernel.org/netdev/QDISC-CTUU.v2.20260913192614@mojatatu.com/
Link: https://sashiko.dev/#/patchset/QDISC-CTUU.v2.20260913192614@mojatatu.com
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/QDISC-CTUU.v2.20260913192614%40mojatatu.com
Reviewed-by: Victor Nogueira <victor@mojatatu.com>
Tested-by: hybris <hybris@mojatatu.ai>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-CTUU.v3.20260916184908@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
__vlan_insert_inner_tag() only guarantees head room via skb_cow_head(),
never that mac_len bytes of MAC header are present. Its ETH_HLEN
wrappers - __vlan_insert_tag() under skb_vlan_push(), and
vlan_insert_tag() under validate_xmit_vlan() on the generic transmit
path - therefore rewrite the first 16 bytes at skb->data: a 12-byte
memmove plus two 2-byte stores at +12 and +14. No caller supplies the
bound, while the pop helpers use skb_ensure_writable()/pskb_may_pull().
An IFF_TUN device has hard_header_len == 0, so packet_snd() accepts a
one-byte AF_PACKET/SOCK_RAW frame. The first vlan push only sets a
hwaccel tag; the next - clsact "action vlan push" or
bpf_skb_vlan_push() - enters the helper with skb->len still 1. The
head comes from skbuff_small_head without __GFP_ZERO, so each push
drags bytes from beyond skb->tail into the frame. After three the
one-byte send leaves as 13 bytes carrying 11 bytes of uninitialised
slab:
0000: 5a b3 62 12 80 88 ff ff 00 b3 62 12 81
`------------------------------'
only 0x5a was sent; the rest is slab, here the top 56 bits of a
linear-map address
Require the MAC header the helper rewrites to be present, so such a
frame is dropped rather than transmitted.
Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Cc: stable@vger.kernel.org
Reported-by: co+0ea1ac045375cf05@bugs.sh
Signed-off-by: Xiang Mei <xmei5@asu.edu>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260915083152.705309-1-xmei5@asu.edu
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
catc_rx_done() walks a multi-packet URB, reading a two-byte length from
each packet header. Its bound, pkt_len > urb->actual_length, ignores the
header offset and compares against the whole transfer rather than the
bytes left from pkt_start, so a crafted packet header makes
skb_copy_to_linear_data() read past the buffer.
A length below ETH_HLEN is also accepted, including zero, and
eth_type_trans() then reads a MAC header from the uninitialised tailroom
of a shorter skb. The is_f5u011 branch takes its length straight from
the transfer, so a zero-length URB reaches the same path.
Track the bytes remaining from the current packet, and reject a header
that does not fit, a length past what is left, and a length below an
Ethernet header.
A transfer shorter than an Ethernet header, including a zero-length one,
previously became a runt skb passed to netif_rx() and counted as
received; it is now counted in rx_length_errors and ends the walk.
Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Signed-off-by: Aamir Ahmed <elb12345@hotmail.co.uk>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/AS8P251MB00015FD7716F38C345619B56C8BB2@AS8P251MB0001.EURP251.PROD.OUTLOOK.COM
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
tcf_action_delete() drops the reference held by its lookup before calling
tcf_idr_delete_index() with the saved action index. An unlocked
classifier can remove that action and reserve the same IDR slot with
ERR_PTR(-EBUSY) in between.
tcf_idr_delete_index() only checks the lookup result for NULL. It
therefore treats the reservation as a tc_action and dereferences
tcfa_bindcnt. A hardware execution breakpoint was used to schedule the
interleaving without changing the kernel source. KASAN reported this
decoded trace:
BUG: KASAN: null-ptr-deref in tca_action_gd+0x5b9/0x1010
Read of size 4 at addr 0000000000000010 by task poc/150
Oops: general protection fault, probably for non-canonical address 0xdffffc0000000002
RIP: tca_action_gd+0x5c0/0x1010:
arch_atomic_read at arch/x86/include/asm/atomic.h:23
raw_atomic_read at include/linux/atomic/atomic-arch-fallback.h:457
atomic_read at include/linux/atomic/atomic-instrumented.h:33
tcf_idr_delete_index at net/sched/act_api.c:766
tcf_action_delete at net/sched/act_api.c:1859
tcf_del_notify at net/sched/act_api.c:2014
tca_action_gd at net/sched/act_api.c:2064
R13: 0000000000000010 R15: fffffffffffffff0
Kernel panic - not syncing: Fatal exception
R15 contains ERR_PTR(-EBUSY), and adding the tcfa_bindcnt offset produces
the address in R13. With the guard applied, the same reproducer returned
-ENOENT without a KASAN report or panic. Treat error pointers as absent
and return -ENOENT.
Fixes: 0190c1d452 ("net: sched: atomically check-allocate action")
Cc: stable@vger.kernel.org
Reported-by: Xiang Mei <xmei5@asu.edu>
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Link: https://patch.msgid.link/20260914065123.4109709-2-bestswngs@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Commit fa8fca8871 ("ipv4: validate IPV4_DEVCONF attributes properly")
added validation of IFLA_INET_CONF attributes, and in the process
changed the call of nla_for_each_nested() to nla_parse_nested(). A
side effect of this change is that the IFLA_INET_CONF option is now
tested for NLA_F_NESTED being set, and fails if it is not. Prior to the
commit there was no check of NLA_F_NESTED.
Change nla_parse_nested() to nla_parse(). This restores the previous
functionality of not checking NLA_F_NESTED, thereby allowing code that
(incorrectly) doesn't set NLA_F_NESTED to continue to work.
This issue was identified because keepalived started logging errors when
it was configuring macvlans that it created.
Fixes: fa8fca8871 ("ipv4: validate IPV4_DEVCONF attributes properly")
Signed-off-by: Quentin Armitage <quentin@armitage.org.uk>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260915213320.1527029-2-quentin@armitage.org.uk
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Commit 339ccec8d4 ("net/mlx5: Enable MACsec offload feature for VLAN
interface") added NETIF_F_HW_MACSEC unconditionally to vlan_features so
that VLAN devices could inherit MACsec offload support.
mlx5e_build_nic_netdev subsequently copies vlan_features into
hw_features and features. As a result, all mlx5e NIC netdevices
advertise MACsec hardware offload, even when the firmware does not
support it and the driver does not install macsec_ops.
Set the MACsec feature bits in mlx5e_macsec_build_netdev, after device
capabilities have been validated. This preserves MACsec-over-VLAN
support and the ethtool feature control on capable devices, without
advertising either on unsupported hardware.
Fixes: 339ccec8d4 ("net/mlx5: Enable MACsec offload feature for VLAN interface")
Cc: stable@vger.kernel.org
Reviewed-by: Tariq Toukan <tariqt@nvidia.com>
Signed-off-by: Ralf Lici <ralf@mandelbit.com>
Link: https://patch.msgid.link/20260917122724.654639-1-ralf@mandelbit.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
sctp_assoc_update_retran_path() can loop forever when every remaining
transport, including retran_path, is SCTP_UNCONFIRMED: the state check
runs before the wraparound test, so the loop cannot observe that it has
completed a full pass.
Fix this by considering a transport only when it is not UNCONFIRMED,
then checking whether the walk has returned to retran_path. This makes
the full-pass termination independent of the transport state while
preserving the existing fallback selection semantics.
Also restore the NULL guard around the retran_path assignment. In the
all-UNCONFIRMED case there is no eligible replacement transport, and
installing NULL would leave later retransmit-path users and the debug
print with a NULL path.
Fixes: 4c47af4d5e ("net: sctp: rework multihoming retransmission path selection to rfc4960")
Signed-off-by: Yiqi Sun <sunyiqixm@gmail.com>
Acked-by: Xin Long <lucien.xin@gmail.com>
Link: https://patch.msgid.link/20260915095017.942213-1-sunyiqixm@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Alexander Duyck says:
====================
eth: fbnic: a collection of fixes
This series collects a handful of independent fbnic fixes for issues on
released kernels, plus one core ethtool fix needed by the fbnic offline
self test.
The first patch keeps rtnl_lock held on the ethtool ioctl path for the self
test. Since the ioctl path became rtnl-optional for ops-locked drivers,
fbnic's offline self test (which brings the interface down and up via
netif_close()/netif_open()) runs holding only the instance lock, tripping a
lockdep splat / ASSERT_RTNL and reconfiguring the device without the lock
it requires. A similar issue was found with Broadcom drivers so we expanded
the scope for v2 to just have the rtnl lock held for all selftest calls.
The second addresses a comparison issue in that we were limiting the
maximum number of standalone Tx queues to one less than the maximum number
of Tx queues. To resolve this it was just a matter of replacing a "<" with
a "<=".
The third addresses an indexing issue with netdev queues on fbnic in which
the NAPI vector was assumed to be findable as the Rx index modulo the
number of NAPI vectors. However this is actually not the case for if Tx
only and Rx only queues are setup. To resolve this we make use of the
cached NAPI pointer in the netdev Rx queues themselves.
The fourth patch fixes a NULL pointer dereference on unbind after a failed
PCIe error recovery: fbnic_pm_suspend() frees the napi vectors via a direct
ndo_stop() while leaving netif_running() true, and when slot_reset ->
resume fails the data path is never re-allocated. To prevent the panic we
reset num_napi to 0 before we free the IRQs which prevents walking the
unallocated napi vectors when we unbind the interface later.
The last two patches address the FW mailbox. One sets AW_FLUSH_MODE
alongside AW_FLUSH when tearing down the Rx ring, so the write pipeline
actually drains the staged requests instead of hanging on the BME halt.
The other handles completions flagged with FW_ERR on both mailboxes, which
the driver previously ignored. This resulted in us parsing a stale Rx page,
and spinning the capabilities poll to a timeout on a healthy ring.
====================
Link: https://patch.msgid.link/178941996343.7700.9376081102002673062.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The firmware can complete a mailbox descriptor while also setting FW_ERR
to indicate it could not process the request, for example on a mailbox
DMA error. The completion carries no valid data.
The driver did not check FW_ERR. On the Rx mailbox it would sync and
parse the stale page as a normal message, and on the Tx mailbox it
silently freed the request. If the initial capabilities exchange in
fbnic_mbx_poll_tx_ready() hit FW_ERR -- on the Tx request or on the Rx
response descriptor -- no response was parsed and the poll spun until it
timed out even though the ring was healthy.
Check FW_ERR on both mailboxes. Count it per-mailbox in
fbnic_fw_mbx.resp_error, which is also shown in debugfs, warn (rate
limited, since the bit is firmware controlled), and drop the Rx page
instead of parsing it.
In fbnic_mbx_poll_tx_ready() re-issue the capabilities request when
either the Tx or the Rx resp_error counter advances, so a FW_ERR on the
request or on its response triggers a retry rather than a timeout. A
valid capabilities response is honored before the retry check, so a
response parsed in the same poll as an unrelated FW_ERR is not discarded.
The counters are mailbox-wide rather than keyed to the capabilities
request; that is sufficient here because the exchange runs during
bring-up before any other mailbox traffic, and any spurious retry is
bounded by the existing 10s timeout.
Fixes: da3cde0820 ("eth: fbnic: Add FW communication mechanism")
Signed-off-by: Alexander Duyck <alexanderduyck@fb.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942023343.7700.9423398932961964439.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
When tearing down the FW mailbox Rx ring, fbnic_mbx_reset_desc_ring()
writes AW_CFG with FLUSH set and everything else, BME included, cleared.
Clearing BME halts the device's writes to the host but leaves the staged
requests parked in the PUL write pipeline rather than draining them, so
on the write path FLUSH alone never terminates the outstanding requests
and the flush the firmware waits on never completes.
Add the FLUSH_MODE definition and set both bits so the staged writes
drain out of the pipeline on their own. BME stays cleared, so nothing
lands on the host; it is restored later in fbnic_mbx_init_desc_ring()
when the ring is rebuilt, once the outstanding writes are gone.
The read path is unaffected. AR_CFG has no equivalent mode bit and
AR_FLUSH terminates the outstanding reads by itself, so it is left as
is.
Both writes remain plain stores rather than read-modify-writes. That is
deliberate: the matching write in fbnic_mbx_init_desc_ring() restores
BME and the TLP attributes, and clears both flush bits as a side effect.
Fixes: 3b12f00ddd ("fbnic: Gate AXI read/write enabling on FW mailbox")
Signed-off-by: Alexander Duyck <alexanderduyck@fb.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942022583.7700.11050671998277309744.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
fbn->num_napi is the count of live napi vectors, each of which owns an
IRQ. The PM path had freed them without clearing the count.
fbnic_pm_suspend() tears the datapath down via ndo_stop() and frees the
IRQs, but leaves netif_running() true so resume knows to re-open. Resume
rebuilds the datapath in __fbnic_pm_resume() and fbnic_reset_queues() sets
num_napi and __fbnic_open() re-allocates the vectors.
When the datapath is torn down but never rebuilt, num_napi is left
pointing at freed vectors under 2 different scenarios:
- a PCIe error recovery that fails (fbnic_err_slot_reset() ->
__fbnic_pm_resume() returns an error -> PCI_ERS_RESULT_DISCONNECT), so
.resume never runs; or
- an __fbnic_open() that fails partway on resume and unwinds, freeing
the vectors after fbnic_reset_queues() has already set num_napi.
The netdev is then running with num_napi > 0 but napi[] freed, and the
eventual remove/unbind close re-enters fbnic_down() -> fbnic_dbg_down()
and dereferences the freed vectors:
BUG: kernel NULL pointer dereference, address: 0000000000000210
RIP: fbnic_dbg_down+0x28
Clear num_napi when the vectors are freed: in the suspend teardown (a
good resume re-establishes it before __fbnic_open()) and on the resume
open failure. A redundant ndo_stop() then walks an empty napi[]. The
normal ndo_stop() down/up cycle is untouched and keeps num_napi for the
next ndo_open().
Fixes: bc6107771b ("eth: fbnic: Allocate a netdevice and napi vectors with queues")
Signed-off-by: Alexander Duyck <alexanderduyck@fb.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942021809.7700.10804028989308077839.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The queue management ndos pick the napi vector for an Rx queue with:
nv = fbn->napi[idx % fbn->num_napi];
The issue is this is only correct in the cases where there are no
standalone Tx vectors. In those cases we were allocating the Tx vectors
first and then the Rx so the queues would be pointing to Tx NAPI vectors
instead of the Rx ones.
The mapping the ndos want is already recorded. fbnic_set_netif_napi()
publishes it with netif_queue_set_napi(), which stores the napi pointer
in netdev_rx_queue.napi, and fbnic_reset_netif_napi() clears it again.
Both run under the netdev instance lock that the queue management ndos
also hold, so the pointer can be read directly.
Use it and drop the divide. The pointer is NULL exactly while the
datapath is down, so fbnic_queue_mem_alloc() can reject that case rather
than reaching into freed state: netdev_rx_queue_restart() calls it
before it tests netif_running(), and fbnic_pm_suspend() leaves
netif_running() true across a PCIe recovery that never completes, so a
queue restart can arrive after fbnic_stop() has freed the rings and the
vectors. fbnic_stop() clears the association in
fbnic_reset_netif_queues() before fbnic_free_napi_vectors(), so the
NULL is always published first. fbnic_queue_start() and
fbnic_queue_stop() need no check of their own, as
netdev_rx_queue_reconfig() only reaches them once fbnic_queue_mem_alloc()
has succeeded under the same instance lock.
Fixes: da43127a8e ("eth: fbnic: support queue ops / zero-copy Rx")
Signed-off-by: Alexander Duyck <alexanderduyck@fb.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942021136.7700.4391219358260544104.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Standalone channels use one NAPI vector for each Tx and Rx queue.
fbnic's allocation path excludes FBNIC_MAX_TXQS from that layout. A
64-Tx/64-Rx configuration therefore records 128 vectors but allocates
only 64, leaving NULL entries that resource setup dereferences.
Include the maximum vector count in standalone allocation.
Fixes: bc6107771b ("eth: fbnic: Allocate a netdevice and napi vectors with queues")
Signed-off-by: Björn Töpel <bjorn@kernel.org>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942020457.7700.13129750616387075931.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
An offline self test that brings the interface down and back up with
netif_close() / netif_open() requires rtnl_lock for both. Since the
ethtool IOCTL path became rtnl-optional for ops-locked drivers, the
ETHTOOL_TEST ioctl runs holding only the netdev instance lock, so on an
ops-locked driver the self test now tears the device down without
rtnl_lock.
With lockdep this reproduces deterministically on every offline self
test on such a driver; note the sole lock held is the instance lock, not
rtnl:
WARNING: suspicious RCU usage
net/core/netpoll.c:207 suspicious rcu_dereference_protected() usage!
1 lock held by ethtool/107:
#0: (&dev->lock){+.+.}, at: dev_ethtool
Call Trace:
netpoll_poll_disable
__dev_close_many
netif_close_many
netif_close
fbnic_self_test
dev_ethtool_locked
dev_ethtool
dev_ioctl
sock_ioctl
__x64_sys_ioctl
Without lockdep the same condition trips ASSERT_RTNL() in
__dev_close_many() / __dev_open(); that check only samples the global
rtnl state, so it can be masked by a concurrent rtnl holder, but the
device is still being reconfigured without the lock it requires.
The ethtool self_test is a legacy ioctl-only command, so an ETHTOOL_TEST
case is only needed on the ioctl path. Add an opt-in bit for drivers whose
self test needs rtnl_lock and set it on the ops-locked drivers whose
offline self test tears the interface down and up:
- fbnic (ops-locked via queue_mgmt_ops): fbnic_self_test() offline path
uses netif_close() / netif_open().
- bnxt (ops-locked via queue_mgmt_ops): bnxt_self_test() offline path
goes through bnxt_close_nic() / bnxt_half_open_nic() /
bnxt_half_close_nic() / bnxt_open_nic(), which close and reopen the
device.
Fixes: f994752b11 ("net: ethtool: optionally skip rtnl_lock on IOCTL path")
Signed-off-by: Alexander Duyck <alexanderduyck@fb.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942019771.7700.338431553546884773.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
lanphy_write_reg_data() does not advance the data pointer while iterating
over the register table. As a result, it writes the first entry num times
and leaves the remaining errata registers unconfigured.
Single-entry tables are unaffected, but tables with multiple entries
leave every entry after the first unapplied.
Advance the data pointer after each successful write so every table entry
is applied in order.
Fixes: c8732e9339 ("net: phy: micrel: lan8842 errata")
Cc: stable@vger.kernel.org
Signed-off-by: Abhishek Ojha <abhishek.ojha@savoirfairelinux.com>
Reviewed-by: Andrew Lunn <andrew@lunn.ch>
Link: https://patch.msgid.link/20260916231928.1336305-1-abhishek.ojha@savoirfairelinux.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
ipv6_find_hdr() walks the extension header chain, skipping each header by
the length that header itself declares. ipv6_optlen() returns up to 2048,
and the skip is never checked against skb->len, so the offset stored in
*offset can point past the end of the packet.
openvswitch installs that offset as the transport header, and
update_ipv6_checksum() then reads and writes the transport checksum field
out of bounds:
BUG: KASAN: slab-use-after-free in inet_proto_csum_replace16+0x445/0x470
Read of size 2 at addr ffff88810b754b06 by task ovs_ipv6_oob/629
CPU: 4 UID: 1000 PID: 629 Comm: ovs_ipv6_oob Tainted: G N 7.3.0-rc3+ #348
Call Trace:
inet_proto_csum_replace16+0x445/0x470
set_ipv6_addr+0x3dd/0x460
do_execute_actions+0x6a3d/0x7c40
ovs_execute_actions+0xfd/0x480
ovs_packet_cmd_execute+0xc38/0xf20
genl_rcv_msg+0x59e/0x870
netlink_rcv_skb+0x18b/0x450
genl_rcv+0x2d/0x40
netlink_unicast+0x6bc/0xa20
The buggy address belongs to the object at ffff88810b754980
which belongs to the cache skbuff_small_head of size 704
The buggy address is located 390 bytes inside of
freed 704-byte region [ffff88810b754980, ffff88810b754c40)
Other callers use that offset too, so bound it here rather than in one
caller.
Reject a header whose declared length does not fit in the packet.
ipv6_find_hdr() already fails with -EBADMSG on a malformed chain, so this
adds no new failure mode.
Fixes: f8f626754e ("ipv6: Move ipv6_find_hdr() out of Netfilter code.")
Suggested-by: Ilya Maximets <i.maximets@ovn.org>
Suggested-by: Eric Dumazet <edumazet@google.com>
Cc: stable@vger.kernel.org
Signed-off-by: Norbert Szetei <norbert@doyensec.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Ilya Maximets <i.maximets@ovn.org>
Link: https://patch.msgid.link/8F80BA1A-DDFD-432D-9075-242A3435FEB5@doyensec.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
X4D is an X4 controller instance as an IP block in an SoC.
It has the same feature set as X4.
Signed-off-by: Andy Moreton <andy.moreton@amd.com>
Reviewed-by: Pieter Jansen van Vuuren <pieter.jansen-van-vuuren@amd.com>
Reviewed-by: Alejandro Lucero <alucerop@amd.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260916125641.12238-1-alucerop@amd.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
If register_netdev() fails for one of the MTK_MAX_DEVS devices in
mtk_probe(), the error path jumps to err_deinit_ppe, skipping
mtk_unreg_dev(). The previously registered net_devices are then freed by
mtk_free_dev() while still in NETREG_REGISTERED state, hitting the
BUG_ON(dev->reg_state != NETREG_UNREGISTERED).
Route the register_netdev() failure to err_unreg_netdev so the net_devices
registered so far are properly unregistered before being freed.
Fixes: 8a8a9e89f8 ("net: ethernet: mediatek: cleanup error path inside mtk_hw_init")
Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Link: https://patch.msgid.link/20260916-mtk_eth_soc-netdev-fix-v1-1-5dac50eb65b1@oss.qualcomm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The REMCSUM option carries an absolute checksum start and checksum field
offset. gue_remcsum() passes them to skb_remcsum_process(), whose
partial path stores offset - start in the u16 skb->csum_offset variable.
If offset is less than start, this underflows.
A forwarded packet can retain CHECKSUM_PARTIAL and reach a NETIF_F_HW_CSUM
driver which trusts the metadata, leading skb_copy_and_csum_dev() to write
two bytes about 64 KiB beyond the destination buffer.
Reject reversed tuples in validate_gue_flags(), after the existing length
validation, so all GUE parsers enforce the ordering in one place.
Fixes: fe881ef11c ("gue: Use checksum partial with remote checksum offload")
Signed-off-by: Jérémy Jean <Jeremy.Jean@oss.cyber.gouv.fr>
Link: https://patch.msgid.link/20260915124806.2852293-2-Jeremy.Jean@oss.cyber.gouv.fr
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
flow_offload_alloc() returns NULL when the conntrack entry is dying
(e.g. raced with a conntrack flush) or when the GFP_ATOMIC allocation
fails; both are expected under load and neither is a kernel bug. This
path runs from softirq on every committed packet, so with
panic_on_warn=1 an unprivileged user can panic the box just by racing
a conntrack flush against a `tc ... action ct commit` classifier.
Reproduced with a custom repro under QEMU: a small, fixed set of UDP
flows through `tc filter ... action ct commit` on lo, raced against
threads flooding bare ctnetlink CT_DELETE (flush) requests. Hits
WARNING: net/sched/act_ct.c:437 (tcf_ct_flow_table_add(), inlined
into tcf_ct_act() in this build) within ~15s on the unpatched kernel;
same setup is clean on the patched kernel. The fix itself is
behavior-preserving: both branches already did `goto err_alloc`
before and after, only the WARN is removed.
Fixes: 64ff70b80f ("net/sched: act_ct: Offload established connections to flow table")
Reported-by: syzbot+6cc37aba98dac721c415@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=6cc37aba98dac721c415
Signed-off-by: Nguyen Ngoc Thang <ngocthang2710.1999@gmail.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260915150816.36487-1-ngocthang2710.1999@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Commit 48a5fe3877 ("tipc: fix bc_ackers underflow on duplicate
GRP_ACK_MSG") rejected duplicate/stale ACKs in tipc_group_proto_rcv()
by returning early when less_eq(acked, m->bc_acked).
However, that check remains incomplete in two ways:
1. When grp->bc_ackers is zero (e.g. on a quiet group, when replicast
ACKs were not requested, or after all expected members have already
acknowledged), an unexpected GRP_ACK_MSG with acked > m->bc_acked
passes less_eq() and unconditionally decrements grp->bc_ackers.
Because bc_ackers is a u16, this wraps to 65535, causing
tipc_group_bc_cong() to permanently report congestion and blocking
all future group broadcasts on the socket.
2. During an active broadcast round (grp->bc_ackers > 0), the sender
transmits packet S and advances grp->bc_snd_nxt to S + 1. Receivers
increment their expected counter to S + 1 upon consuming packet S,
so the only valid ACK value for the current round is strictly
acked == grp->bc_snd_nxt.
However, tipc_group_update_bc_members() initializes each member's
m->bc_acked to prev = grp->bc_snd_nxt - 1 (S - 1 before increment).
This leaves a 2-sequence gap (S - 1 to S + 1) in sequence space.
An incoming ACK is therefore neither rejected as duplicate nor
prevented from decrementing grp->bc_ackers if an unexpected or stale
value (such as S) is received. A member sending acked = S followed
by acked = S + 1 could decrement grp->bc_ackers twice in the same
round, prematurely clearing bc_ackers or underflowing it.
Fix this by:
- Dropping GRP_ACK_MSG immediately if grp->bc_ackers is zero.
- Requiring acked == grp->bc_snd_nxt and rejecting duplicates where
m->bc_acked == acked. Because replicast broadcast rounds are strictly
sequential, only grp->bc_snd_nxt can be acknowledged, and each member
can acknowledge at most once per round.
Note that a related pre-existing issue in tipc_group_delete_member()
(where grp->bc_ackers decrementing to zero upon member departure does
not restore *grp->open or trigger a socket wakeup) will be addressed
in a separate patch.
Fixes: 48a5fe3877 ("tipc: fix bc_ackers underflow on duplicate GRP_ACK_MSG")
Fixes: 2f487712b8 ("tipc: guarantee that group broadcast doesn't bypass group unicast")
Reported-by: James Burton <jamesburton@meta.com>
Cc: stable@vger.kernel.org
Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260913044233.193927-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
nft_setelem_catchall_insert() looks up duplicates with
nft_set_elem_active() only, while nft_set_catchall_lookup() and the
dump path additionally skip expired elements.
Once a catchall element with a timeout expires, this predicate drift
makes it invisible to userspace dumps, yet it still blocks
re-insertion: with NLM_F_EXCL the request fails with -EEXIST, and
without it the request reports success but silently inserts nothing.
The stale entry only goes away when the (user-tunable) gc interval
elapses, so the catchall rule may silently stop matching for an
arbitrarily long time after its first expiration.
The delete path shows the same drift: nft_setelem_catchall_deactivate()
picks the first active-next entry in the catchall list, so with an
expired entry still pending GC it retires the stale entry instead of
the fresh one, and it deactivates an element that userspace no longer
sees instead of failing with -ENOENT.
Align both walks with the lookup and dump predicates: only an element
that is active and not expired counts as a duplicate or delete
candidate, using the per-netns timestamp taken at transaction start,
in line with the set backend .insert/.deactivate and catchall GC sync
paths.
Reported-by: TencentOS Corvus AI <corvus@tencent.com>
Cc: stable@vger.kernel.org
Fixes: aaa31047a6 ("netfilter: nftables: add catch-all set element support")
Assisted-by: CodeBuddy:Kimi-K3
Signed-off-by: Aohan Mei <henrymei@tencent.com>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
expect_iter_name() is invoked by nf_ct_expect_iterate_net() under
spin_lock_bh(&nf_conntrack_expect_lock). It does not hold
rcu_read_lock().
When accessing exp->helper with rcu_dereference() in syzbot's report,
lockdep warns:
=============================
WARNING: suspicious RCU usage
syzkaller #0 Not tainted
-----------------------------
net/netfilter/nf_conntrack_netlink.c:3393 suspicious rcu_dereference_check() usage!
locks held by syz-executor381/5628: 2, last CPU#1:
#0: ffffffff9aee42a0 (nfnl_subsys_ctnetlink_exp){+.+.}-{4:4},
at: nfnetlink_rcv_msg+0xa69/0x12b0
#1: ffffffff8ea74d58 (nf_conntrack_expect_lock){+...}-{3:3},
at: nf_ct_expect_iterate_net+0x38/0x180
Call Trace:
<TASK>
dump_stack_lvl+0xe8/0x150
lockdep_rcu_suspicious+0x140/0x1d0
expect_iter_name+0xfb/0x100
nf_ct_expect_iterate_net+0xf2/0x180
ctnetlink_del_expect+0x45d/0x640
nfnetlink_rcv_msg+0xcc2/0x12b0
netlink_rcv_skb+0x226/0x4a0
nfnetlink_rcv+0x2b9/0x28c0
netlink_unicast+0x7bd/0x940
netlink_sendmsg+0x813/0xb40
____sys_sendmsg+0x54e/0x850
___sys_sendmsg+0x2a5/0x360
__sys_sendmsg+0x2a5/0x360
do_syscall_64+0x166/0x520
entry_SYSCALL_64_after_hwframe+0x77/0x7f
Use rcu_dereference_protected() with lockdep_is_held() on
nf_conntrack_expect_lock instead, similar to expect_iter_me() in
nf_conntrack_helper.c.
Fixes: f017941060 ("netfilter: nf_conntrack_expect: use expect->helper")
Reported-by: syzbot+4bd730aede2791e40bdf@syzkaller.appspotmail.com
Closes: https://lore.kernel.org/netdev/6aa4a377.f81106d8.2ab401.0024.GAE@google.com/T/#u
Signed-off-by: Naman Gulati <namangulati@google.com>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
While the outer IP header is already pulled into the skb head, we must
be careful and revalidate the embedded headers after reading them from
the skb frags to prevent possible out-of-bounds access.
One such place reported by Sashiko is ip_vs_in_icmp() where local
process can change the ihl field and after pskb_may_pull() we can see
larger value. Even if icmp_send() has checks to prevent out-of-bounds
access, play safe and add check to drop the packet if the ihl field is
changed. As the outer headers are pulled, make sure the transport
header is updated too, it was used before commit 7fcc2fe39f ("net:
icmp: avoid invalid transport header access in icmp_send tracepoint")
Fixes: f2edb9f770 ("ipvs: implement passive PMTUD for IPIP packets")
Link: https://sashiko.dev/#/patchset/20260806105211.34622-1-ja%40ssi.bg
Signed-off-by: Julian Anastasov <ja@ssi.bg>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
nft_synproxy_do_eval() verifies the TCP checksum before it switches on
skb->protocol. It uses nf_ip_checksum(), which constructs an IPv4
pseudo header and relies on the IPv4 header checksum when folding the
whole skb. Neither operation is valid for an IPv6 packet.
A correctly checksummed IPv6 segment can therefore fail verification
when it reaches the hook as CHECKSUM_NONE or, at NF_INET_LOCAL_IN,
CHECKSUM_COMPLETE. nft_synproxy_do_eval() returns NF_DROP before
nft_synproxy_eval_v6() can send a SYN-ACK.
nft_synproxy_validate() deliberately admits NFPROTO_IPV6 and
NFPROTO_INET, and the xtables counterpart ip6t_SYNPROXY.c already calls
nf_ip6_checksum().
Use nf_checksum() with nft_pf() so the checksum helper dispatches to the
packet family's implementation.
Fixes: ad49d86e07 ("netfilter: nf_tables: Add synproxy support")
Assisted-by: LLM
Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
rt_mt6_check() permits rules to be configured with rtinfo->addrnr == 0
even when address matching (IP6T_RT_FST_MASK) is requested.
In the IP6T_RT_FST_NSTRICT path, rt_mt6() evaluates packet routing
addresses against rtinfo->addrs[i] and terminates backwards at the bottom
of the loop:
if (ipv6_addr_equal(ap, &rtinfo->addrs[i])) {
i++;
}
if (i == rtinfo->addrnr)
break;
When addrnr is 0, if the first packet address matches rtinfo->addrs[0],
i is incremented to 1. Because i is now strictly greater than addrnr (0),
the loop termination condition (i == rtinfo->addrnr) is bypassed and will
never be satisfied.
If a crafted IPv6 packet contains matching routing addresses, i will
advance past IP6T_RT_HOPS (16). The subsequent call to ipv6_addr_equal()
reads beyond struct ip6t_rt, triggering UBSAN/KASAN out-of-bounds warnings
or kernel panics.
Fix this by:
1. Rejecting rules in rt_mt6_check() where IP6T_RT_FST_MASK is set but
rtinfo->addrnr is zero.
2. In rt_mt6(), moving the termination condition (i < rtinfo->addrnr)
into the for-loop header condition and removing the backwards break
at the end of the loop body.
Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Suggested-by: Florian Westphal <fw@strlen.de>
Assisted-by: LLM
Signed-off-by: Luxiao Xu <rakukuip@gmail.com>
Signed-off-by: Ren Wei <weir@nebusec.ai>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
ip6_route_lookup() can return an error-free route whose rt6i_idev is
NULL. Lowering an external nexthop device's MTU below IPV6_MIN_MTU tears
down its inet6_dev while fib6_ifdown() leaves routes using nexthop objects
in the FIB. An unprivileged user can construct this state with rtnetlink
in a private user and network namespace, then trigger a NULL dereference
through an IPv6 rpfilter lookup:
Oops: general protection fault, probably for non-canonical address
0xdffffc0000000000
KASAN: null-ptr-deref in range [0x0000000000000000-0x0000000000000007]
RIP: rpfilter_mt (net/ipv6/netfilter/ip6t_rpfilter.c:75)
Call Trace:
ip6t_do_table (net/ipv6/netfilter/ip6_tables.c:316)
nf_hook_slow (net/netfilter/core.c:619)
ipv6_rcv (net/ipv6/ip6_input.c:351)
__netif_receive_skb_one_core (net/core/dev.c:6216)
process_backlog (net/core/dev.c:6680)
__napi_poll (net/core/dev.c:7739)
net_rx_action (net/core/dev.c:7959)
handle_softirqs (kernel/softirq.c:622)
do_softirq.part.0 (kernel/softirq.c:523)
__local_bh_enable_ip (kernel/softirq.c:450)
__dev_queue_xmit (net/core/dev.c:4913)
packet_sendmsg (net/packet/af_packet.c:3139)
__sys_sendto (net/socket.c:2252)
__x64_sys_sendto (net/socket.c:2259)
do_syscall_64 (arch/x86/entry/syscall_64.c:94)
entry_SYSCALL_64_after_hwframe (arch/x86/entry/entry_64.S:121)
Kernel panic - not syncing: Fatal exception in interrupt
Reject routes without an inet6_dev immediately after lookup. Such routes
are not eligible for reverse-path filtering, and the check protects all
later rt6i_idev dereferences.
Fixes: e26f9a480f ("netfilter: add ipv6 reverse path filter match")
Reported-by: co+459f67f4d8af8ce6@bugs.sh
Closes: https://lore.kernel.org/all/VtWUkE8QzJt5CroTj2V2v3ZQ0gwbXZ7nq7I3@bugs.sh/
Suggested-by: Florian Westphal <fw@strlen.de>
Assisted-by: Claude:gpt-5
Cc: stable@vger.kernel.org
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
We must serialize the release notifier and the config netlink function.
A concurrent thread can issue close() which can call the release function
while unrelated socket processes UNBIND request for same portid:
Oops: general protection fault, [..]
RIP: 0010:__instance_destroy+0x60/0x210 [nfnetlink_queue]
Call Trace:
nfqnl_recv_config+0x9b0/0xdc0 [nfnetlink_queue]
nfnetlink_rcv_msg+0x7c2/0xeb0
? __pfx_nfnetlink_rcv_msg+0x10/0x10
After this, parallel UNBIND and URELEASE events are impossible.
This change isn't nice, but its the shortest fix given instances
are not refcounted and the nfnetlink config callback drops the
rcu read lock early due to need for sleeping allocations.
Fixes: 7af4cc3fa1 ("[NETFILTER]: Add "nfnetlink_queue" netfilter queue handler over nfnetlink")
Signed-off-by: Florian Westphal <fw@strlen.de>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
flow_offload_work_del() sets NF_FLOW_HW_DEAD before the work handler
clears NF_FLOW_HW_PENDING. Once a flow is both HW_DYING and HW_DEAD, a
concurrent garbage collection pass can remove it and schedule it for RCU
freeing.
The offload worker holds neither an RCU read lock nor a reference to the
flow. If it is preempted after publishing HW_DEAD, the RCU callback can
free the flow before the worker resumes and clears HW_PENDING, resulting
in a use-after-free.
Move HW_DEAD publication to the common worker epilogue after the pending
bit is cleared, making it the final flow access by destroy work. Order all
preceding flow accesses before publishing the bit that allows garbage
collection to free the object.
Fixes: 2c8897953f ("netfilter: flowtable: Add pending bit for offload work")
Assisted-by: Codex:gpt-5
Signed-off-by: Jérémy Jean <Jeremy.Jean@oss.cyber.gouv.fr>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>