Most stmmac selftests rely on dev_add_pack() to add custom handlers,
that validate the packets sent to ourselves through MAC loopback.
However, when the stmmac-driven interface is a DSA CPU conduit, all
frames that are received have ETH_P_XDSA as a protocol, even though they
don't actually contain any tag as they come from the loopback and not
the switch.
This will prevent any incoming packet to match our packet handlers.
Let's register a ETH_P_ALL packet handler when we detect that we're a
DSA conduit, and use a proxy packet handler to filter the h_proto.
As this allows external frames to be received through our .func(), the
packet handler is added after the dev->addr field is populated in our
selftest attributes.
Note that we may still receive incoming packets from the switch, but
these frames shouldn't interfere with the very specific frames used for
selftests, and stmmac selftests in general aren't safe against external
traffic interferences.
This was validated on a WPQ864 devkit for IPQ8064, that has the SoC
connected to a QCA8k switch.
The ARP offload's packet handler is left alone, this feature is just not
implemented in stmmac and due for removal.
Fixes: 091810dbde ("net: stmmac: Introduce selftests support")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/20260917215339.2022523-2-maxime.chevallier@bootlin.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Unlike bind-rx, which configures shared NIC RX queues to steer incoming
traffic into the caller's dmabuf and requires CAP_NET_ADMIN
(uns-admin-perm), bind-tx only DMA-maps the caller's dmabuf so the caller
can transmit from it on their own sockets without affecting other traffic
or device configuration.
Add a comment in netdev.yaml and above netdev_nl_bind_tx_doit() to make it
explicit that NETDEV_CMD_BIND_TX is unprivileged by design.
Signed-off-by: Mina Almasry <almasrymina@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
Acked-by: Daniel Borkmann <daniel@iogearbox.net>
Link: https://patch.msgid.link/20260921195545.493253-1-almasrymina@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Take the association or transport reference before rearming a timer in the
timer handlers.
The existing code calls mod_timer() before taking the reference needed by
the rearmed timer without holding the sock lock. This creates a race with
timer cleanup: if the timer is deleted after mod_timer() returns but before
the reference is taken, the cleanup path can drop the timer's reference and
destroy the transport or association. The timer handler then takes a
reference on the already freed object and eventually drops it, causing a
refcount underflow.
Hold the object before mod_timer() and drop the reference if mod_timer()
reports that the timer was already pending in timer handlers. Apply the
same ordering to the proto-unreachable path, which can rearm a transport
timer outside the timer handlers without holding the sock lock.
Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Reported-by: Tangxin Xie <xietangxin@h-partners.com>
Signed-off-by: Xin Long <lucien.xin@gmail.com>
Link: https://patch.msgid.link/c31b5e3ee2b7274e804f5eba2f21e2412e7eef7a.1790013825.git.lucien.xin@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Willem de Bruijn says:
====================
packet: fix PACKET_TX_RING data corruption on skb_orphan
When transmitting packets via PACKET_TX_RING, tpacket_snd links user
ring buffer pages as skb frags and releases the slot on skb->destructor
(tpacket_destruct_skb).
skb_orphan() invokes the destructor while the skb is still alive.
This marks the slot as TP_STATUS_AVAILABLE prematurely, allowing
userspace to overwrite the slot and causing data corruption.
This series fixes the issue by switching PACKET_TX_RING to standard
ubuf_info zerocopy completion, ensuring ring slots are released only
after all payload references are freed or copied.
Virtio-net needs a separate solution, because deferring the release
can cause deadlock in its !use_napi mode.
- Patch 1 addresses the virtio-net special case.
- Patch 2 converts tpacket_snd to standard ubuf_info completion
Patch 1 must be applied, and backported, before patch 2. Both carry
the same Fixes tag for that reason.
v1: https://lore.kernel.org/netdev/20260914214229.1674102-1-willemdebruijn.kernel@gmail.com/
====================
Link: https://patch.msgid.link/20260919004748.1463985-1-willemdebruijn.kernel@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
tpacket_snd sends skbs with frags pointing into its ring slots. Slots
are released when skb->destructor is called.
A call to skb_orphan calls skb->destructor before the skb is freed.
This can cause the slot to be reused while still linked into the skb.
Switch to standard zerocopy completion (ubuf_info) so the slot is only
released once all references to the payload are freed or copied.
Restore skb->destructor to standard sock_wfree.
The ubuf_info completion callback can be called with a NULL skb, but
only from net_zcopy_put and related API, used by zerocopy implementations
that hold their own reference on the uarg, such as MSG_ZEROCOPY. This
uarg is only ever completed from skb_zcopy_clear, so skb is always set.
To prevent userspace from aliasing in-flight state on shared ring
slots, allocate tpacket_uarg per packet, rather than per slot. This
adds a small allocation to the transmit path. Use standard kmalloc to
allow backporting to stable kernels.
The uarg holds an sk_wmem_alloc reference, rather than an sk_refcnt
reference. packet_free_tx_ring waits on sk_wmem_alloc before freeing
the ring pages. Always allocate vec->deferred for tx_ring so page-backed
rings also wait on sk_wmem_alloc when skb_copy_ubufs drops page refs
before calling tpacket_ubuf_complete.
Drop the tx_ring.pg_vec test that tpacket_destruct_skb performed before
accessing the slot. The sk_wmem_alloc reference now guarantees that the
slot is valid. The test is also not sufficient by itself, as it reads
pg_vec without pg_vec_lock, so it can race with packet_set_ring.
As a result a slot is released when its payload is copied, which can
be before transmission (e.g., in skb_orphan_frags_rx). If copied
before skb_tx_timestamp() is called, no slot timestamp is recorded,
similar to when skb_orphan() was called early in the datapath before
this patch.
Revert the now unused previous skb_zcopy_.._nouarg infra.
Depends on commit 992cc9f94c ("net/packet: defer vmalloc TX_RING
free until skbs finish").
Reported-by: Katherine Leaver <kleaver@janestreet.com>
Reported-by: Bjoern Doebel <doebel@amazon.de>
Closes: https://lore.kernel.org/netdev/20260909085542.3370986-1-doebel@amazon.de/
Fixes: 5cd8d46ea1 ("packet: copy user buffers before orphan or clone")
Cc: stable@vger.kernel.org
Signed-off-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260919004748.1463985-3-willemdebruijn.kernel@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Virtio-net without NAPI frees completed skbs lazily on the next
start_xmit. Senders waiting for in-flight zerocopy buffers can
deadlock if they cannot transmit more packets, as then no
completed packets will be freed.
When !use_napi, virtio-net already calls skb_orphan to avoid waiting
up for transmitted skbs to be freed. For zerocopy packets that
require deep copying on orphan (i.e. those that do not set
SKBFL_DONT_ORPHAN, such as PACKET_TX_RING), call skb_orphan_frags
before orphaning to release the buffers.
This fixes the tpacket_snd slot reuse bug on skb_orphan for
virtio-net, and prevents PACKET_TX_RING from running out of slots.
This fix also touches vhost_net zerocopy packets, which also do not
set SKBFL_DONT_ORPHAN. This is fine: vhost_net packets only encounter
virtio-net in nested virtualization, and only if napi_tx is
explicitly disabled (it has been default-enabled since Linux 4.12).
In that rare case, copying the frags is desirable anyway to prevent
holding guest descriptors pinned across unbounded intervals.
This is a prerequisite for the next patch, which converts
PACKET_TX_RING to standard zerocopy completion. Without this patch
first, a bounded ring sender can stall indefinitely behind a
virtio-net virtqueue that cannot reclaim.
Fixes: 5cd8d46ea1 ("packet: copy user buffers before orphan or clone")
Cc: stable@vger.kernel.org
Cc: mst@redhat.com
Cc: jasowangio@gmail.com
Signed-off-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260919004748.1463985-2-willemdebruijn.kernel@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
When an association is in COOKIE-ECHOED state and the peer sends a
bundled [ERROR(Stale Cookie)][DATA] packet from one of its non-primary
addresses, processing the ERROR chunk takes the non-fatal stale-cookie
retry path sctp_sf_do_5_2_6_stale(), which queues
SCTP_CMD_DEL_NON_PRIMARY while keeping the association alive.
sctp_cmd_del_non_primary() removes every non-primary transport -
including the very transport this packet arrived on, which is still
referenced by the receive lookup and shared by all chunks of the
packet via chunk->transport.
sctp_assoc_rm_peer() does redirect asoc->peer.last_data_from away from
the removed transport, but right afterwards the bundled DATA chunk
makes sctp_assoc_bh_rcv() re-register
asoc->peer.last_data_from = chunk->transport unconditionally, undoing
the redirection with the just-removed transport.
Once the packet is done, the receive reference is dropped and the
transport is RCU-freed, while the surviving association keeps the
dangling last_data_from. A later FWD-TSN (or the delayed SACK timer)
makes sctp_gen_sack() dereference it (->param_flags and friends), and
sctp_make_sack()/sctp_outq_select_transport() may write to the freed
object and link it into the live transport list. This is a
use-after-free triggerable by any malicious SCTP peer (or a local
unprivileged user acting as one) with no capabilities required:
BUG: KASAN: slab-use-after-free in sctp_do_sm+0x498a/0x5660
Read of size 4 at addr ffff88800e1e356c by task poc/115
Call Trace: sctp_do_sm <- sctp_assoc_bh_rcv <- sctp_inq_push <-
sctp_rcv <- ip_protocol_deliver_rcu <- ip_rcv
Allocated: sctp_transport_new <- sctp_assoc_add_peer <-
sctp_process_init (INIT-ACK processing)
Freed: kfree <- sctp_transport_destroy_rcu <- rcu_core
(call_rcu queued by sctp_transport_put at end of sctp_rcv)
The buggy address is located 364 bytes inside of freed 1024-byte
region [ffff88800e1e3400, ffff88800e1e3800), cache kmalloc-1k
Note that commit 03a9d10ecf ("sctp: drop a chunk if its transport
was removed") only covers the window between the receive lookup and
the chunk processing (e.g. an ASCONF DEL-IP racing the socket backlog);
here the transport is removed *while* the packet is being processed,
by an earlier chunk of the same packet, so the drop in sctp_inq_push()
does not reach this path. Verified with the bundled [ERROR(Stale
Cookie)][DATA] + FWD-TSN reproducer: the KASAN report above still
fires with that commit applied, and is gone with this patch on top.
Fix it by discarding the rest of the packet on this path, as suggested
by Xin. After the stale-cookie ERROR has sent the association back to
COOKIE-WAIT and removed the non-primary transports, the remaining
chunks of the packet can only run against the restarted handshake
while referencing the removed arrival transport through
chunk->transport: besides the last_data_from registration above,
sctp_cmd_setup_t2() and the sctp_make_*() reply builders would also
copy that pointer into association-lifetime state that
sctp_assoc_rm_peer() has already sanitized. Let the peer retransmit
them, in line with what sctp_inq_push() does for chunks whose
transport was removed before processing.
Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Suggested-by: Xin Long <lucien.xin@gmail.com>
Reported-by: TencentOS Corvus AI <corvus@tencent.com>
Cc: stable@vger.kernel.org
Signed-off-by: Aohan Mei <henrymei@tencent.com>
Acked-by: Xin Long <lucien.xin@gmail.com>
Link: https://patch.msgid.link/20260921093707.1432184-1-ljp1205831794@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
When the reverse entry is found but its counter is already being
released, refcount_inc_not_zero() fails and the reference taken by
mlx5_tc_ct_entry_get() is never dropped before falling through to
create_counter. Drop it so the reverse entry is not kept alive forever
by a shared counter lookup that did not use it.
Fixes: 1edae2335a ("net/mlx5e: CT: Use the same counter for both directions")
Cc: stable@vger.kernel.org
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Reviewed-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260917113131.2149024-1-vulab@iscas.ac.cn
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
mlx5_esw_bridge_vport_unlink() returns -EINVAL when the port isn't
tracked by this instance's br_offloads. This is reachable on a sibling
instance that registered its notifier after the port was already
enslaved: it never saw the NETDEV_CHANGEUPPER link event, so
peer_link() never created a peer port for it, but it does see the
later unlink event and fails. Return 0 instead, and give
mlx5_esw_bridge_vport_peer_unlink() the same merged_eswitch capability
guard peer_link() already has, since without it peer_link() likewise
never creates a port to unlink.
This also matters beyond the -EINVAL itself:
mlx5_esw_bridge_switchdev_port_event() runs on the per-netns
netdev_chain, and notifier_from_errno(-EINVAL) sets NOTIFY_STOP_MASK,
which call_netdevice_notifiers_info() checks to stop calling further
listeners on that chain - so the old -EINVAL silently dropped the
event for any listener registered later on the same chain, even
though none of it was visible to user space since
__netdev_upper_dev_unlink() discards the return value.
Fixes: c358ea1741 ("net/mlx5: Bridge, allow merged eswitch connectivity")
Signed-off-by: Bernardo Soares <bsoares.it@gmail.com>
Reviewed-by: Mark Bloch <mbloch@nvidia.com>
Link: https://patch.msgid.link/20260918095931.29792-3-bsoares.it@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
mlx5 registers the bridge offload switchdev notifiers once per eswitch
instance, but the notifier chains are global, so every instance sees
every event and must filter out the ones that aren't its own. The
existing filter, mlx5_esw_bridge_dev_same_hw(), only checks that the
event netdevice sits on the same HCA - intentional for merged eswitch,
where one bridge can span representors of several eswitches on one
HCA - but same-HCA doesn't mean the instance actually has that port:
peer ports are only created reactively from NETDEV_CHANGEUPPER, so an
instance brought up after a sibling PF's port was already enslaved has
none. The port object and attribute handlers claim the event anyway
once same-HW passes, then fail the port lookup and return -EINVAL,
which gets reported to user space even though the owning instance
already handled it (e.g. "bridge vlan add ... RTNETLINK answers:
Invalid argument"). Fix by filtering on the tracked port instead.
The same gap exists in the generic recursive lower-device walk used by
attribute changes on a bridge with more than one representor enslaved
directly: mlx5_esw_bridge_lower_rep_vport_num_vhca_id_get() is entered
with the bridge master netdevice, falls through to its generic
netdev_for_each_lower_dev() loop, and returns as soon as the recursion
into any one lower device yields a non-NULL rep - the underlying base
case, mlx5_esw_bridge_rep_vport_num_vhca_id_get(), only checks
mlx5_esw_bridge_dev_same_hw(), not ownership by the calling instance's
br_offloads. mlx5_esw_bridge_lag_rep_get(), used for the LAG-master
case, already filters on mlx5_esw_bridge_dev_same_esw() per candidate
and so cannot select a sibling's rep; it is not the source of this bug.
On a merged-eswitch HCA with a bridge spanning representors of more
than one eswitch instance directly, the walk can return a sibling's rep
instead of continuing to the one the calling instance actually owns, so
the attribute change fails the same way as above. Fix by checking
mlx5_esw_bridge_port_exists() at the point each rep is picked, same as
the previous fix did for the notifier filter.
Fixes: c358ea1741 ("net/mlx5: Bridge, allow merged eswitch connectivity")
Signed-off-by: Bernardo Soares <bsoares.it@gmail.com>
Cc: Vlad Buslov <vladbu@nvidia.com>
Cc: Saeed Mahameed <saeedm@nvidia.com>
Reviewed-by: Mark Bloch <mbloch@nvidia.com>
Link: https://patch.msgid.link/20260918095931.29792-2-bsoares.it@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
usb_get_from_anchor() hands over a reference to the URB, which the caller
must release. lan78xx_submit_deferred_urbs() never does, so every deferred
Tx URB keeps an extra reference: the counter grows on each suspend/resume
cycle and the URBs are never freed when the buffers are released. Drop
the reference after submitting, and on the path that drops the packet
instead of submitting it.
Fixes: 5f4cc6e251 ("lan78xx: Fix race conditions in suspend/resume handling")
Cc: stable@vger.kernel.org
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Link: https://patch.msgid.link/20260917115811.2150119-1-vulab@iscas.ac.cn
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
xpcs_init_clks() takes references with clk_bulk_get_optional() and then
enables them with clk_bulk_prepare_enable(). If the enable step fails,
the function returns without dropping the references.
xpcs_create() handles the failure through out_free_data, which calls
xpcs_free_data() but never xpcs_clear_clks(), so the clk references are
leaked.
Add the missing clk_bulk_put() on the enable failure path. The
prepare/enable side is already rolled back by
clk_bulk_prepare_enable() itself.
Fixes: f6bb3e9d98 ("net: pcs: xpcs: Add Synopsys DW xPCS platform device driver")
Signed-off-by: Coia Prant <coiaprant@gmail.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260919172021.2336748-1-coiaprant@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
The op-to-policy map a CTRL_CMD_GETPOLICY dump returns is the only way
for userspace to find out which policy index belongs to which command.
ctrl_dumppolicy_put_op() tags the nest with doit->cmd, but an op which
only has a dumpit has no doit and every path which fills the split ops
in zeroes it out, so those entries all claim to be command 0. nlctrl's
own CTRL_CMD_GETPOLICY and NETDEV_CMD_QSTATS_GET are both in that group:
[{'family-id': 16, 'op-policy': {'do': 0, 'dump': 0, 'op-id': 3}},
{'family-id': 16, 'op-policy': {'dump': 1, 'op-id': 0}},
ctrl_fill_info() gets this right - it uses the iterator's cmd for
CTRL_ATTR_OP_ID - so the two introspection interfaces of the same family
contradict each other today.
Pass the command in rather than reconstructing it from
doit->cmd | dumpit->cmd inside the helper, both callers already have it.
Fixes: 26588edbef ("genetlink: support split policies in ctrl_dumppolicy_put_op()")
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Link: https://patch.msgid.link/20260918222949.4190284-1-kuba@kernel.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Fix 3 leaks in macb_alloc() error paths:
- Tx buffer allocated but crossing a 4G boundary: Tx leaked.
- Rx buffer allocation fails: Tx leaked.
- Rx buffer allocated but crossing a 4G boundary: Tx & Rx leaked.
This is because our error handling calls macb_free(bp) which in turn
frees the buffers stored in bp->queues[0], but nothing has been stored
in there. Fix by storing allocated buffers into bp->queues[0] ASAP.
Fixes: 78d901897b ("net: macb: single dma_alloc_coherent() for DMA descriptors")
Cc: stable@vger.kernel.org
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Link: https://patch.msgid.link/20260918-macb-alloc-leak-v1-1-aba9a3d4f6e3@bootlin.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
hns_dsaf_find_platform_device() returns the mdio platform device with its
reference count incremented. hns_mac_register_phy() never drops that
reference, so the mdio device can not be released.
Release the reference on both the deferred probe and the normal path.
Fixes: 1d1afa2ebf ("net: hns: register phy device in each mac initial sequence")
Cc: stable@vger.kernel.org
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260917110828.2148390-1-vulab@iscas.ac.cn
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
pm_runtime_get_sync() leaves the runtime PM usage counter incremented even
when it fails, but the error path in netcp_probe() does not call
pm_runtime_put_noidle() to balance it, leaking a reference each time
resume fails.
Use pm_runtime_resume_and_get() instead, which automatically drops the
usage counter on failure, fixing the leak.
Fixes: 84640e27f2 ("net: netcp: Add Keystone NetCP core ethernet driver")
Signed-off-by: bui duc phuc <phucduc.bui@gmail.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260918042804.13101-1-phucduc.bui@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
emac_tx_mem_map() writes TX_DESC_0_OWN into the ring descriptor for
every slot beyond old_head as soon as that slot's memset()'d local
copy is committed with "*tx_desc_addr = tx_desc", i.e. before the
buffers for that slot have necessarily all been mapped successfully.
If emac_tx_map_frag() then fails on a later fragment, the err_free_skb
path calls emac_free_tx_buf() to unmap and drop the skb, but leaves
the already-written descriptor memory untouched, and tx_ring->head is
never advanced past old_head (the "tx_ring->head = head" store is
skipped by the goto).
So a slot between old_head and the rolled-back head can be left with
TX_DESC_0_OWN set and buffer_addr_{1,2} pointing at DMA mappings that
emac_free_tx_buf() just tore down, while software considers that slot
free again. The next successful emac_tx_mem_map() call only rebuilds
old_head itself; if the DMA engine auto-advances into the following
descriptor once it finishes old_head's packet, it will fetch that
stale, already-unmapped address.
emac_tx_clean_desc() already treats emac_free_tx_buf() and clearing
the descriptor as a pair when reclaiming completed descriptors; do
the same in the mapping failure path.
Fixes: bfec6d7f20 ("net: spacemit: Add K1 Ethernet MAC")
Signed-off-by: Muhammad Bilal <meatuni001@gmail.com>
Reviewed-by: Vivian Wang <wangruikang@iscas.ac.cn>
Reviewed-by: Troy Mitchell <troy.mitchell@linux.spacemit.com>
Link: https://patch.msgid.link/20260919191937.271202-1-meatuni001@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
of_clk_get() returns a clock with its reference count incremented, but
read_dts_node() only uses it to read the rate and never calls clk_put().
The clock is not stored anywhere, so the reference cannot be released
later either.
Release the clock once its rate has been read, which also covers the
error path taken when the rate is zero.
Fixes: 414fd46e77 ("fsl/fman: Add FMan support")
Cc: stable@vger.kernel.org
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260917110135.2148068-1-vulab@iscas.ac.cn
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
tg3_get_invariants() can register an MDIO bus and connect a PHY for
USE_PHYLIB devices. If tg3_init_one() later fails, its common error path
releases the mappings and netdev without undoing those PHYLIB resources.
Disconnect the PHY and unregister the MDIO bus before the remaining
teardown. Guard PHY cleanup with USE_PHYLIB to match tg3_phy_init(), and
call tg3_mdio_fini() unconditionally to match tg3_mdio_init(). The existing
IS_CONNECTED and MDIOBUS_INITED flags make both helpers safe when
initialization only completed partially.
This issue was identified during our ongoing static-analysis research while
reviewing kernel code.
Fixes: 158d7abdae ("tg3: Add mdio bus registration")
Assisted-by: OpenAI:GPT-5.6
Co-developed-by: Ijae Kim <ae878000@gmail.com>
Signed-off-by: Ijae Kim <ae878000@gmail.com>
Signed-off-by: Myeonghun Pak <mhun512@gmail.com>
Link: https://patch.msgid.link/20260917183336.36239-1-mhun512@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Shardul Bankar says:
====================
udp: two fixes for the 4-tuple hash table
Two ways a UDP socket ends up in the wrong place in the 4-tuple hash table.
The patches are independent, with different Fixes: tags and no dependency
between them.
Patch 1: a socket that connects a second time is not relocated, so it stays
filed under its first peer's hash and packets for it fall back to scoring
the hash2 chain for its address and port.
Patch 2: a socket bound to a specific address and port is not taken out of
the table when it disconnects, because __udp_disconnect() only does that
via ->rehash() or ->unhash() and neither runs for it.
Both are in code shared by IPv4 and IPv6.
Patch 1's cost, with N sockets sharing a port and one of them misfiled,
200k packets sent to its 4-tuple:
N without with
200 1,061,652 2,093,259 pps
500 522,553 2,055,078 pps
1000 279,729 2,136,606 pps
Correctly filed sockets measure ~2.1M pps throughout, so the cost scales
with the number of sockets on the port, as the fallback scan does. For
comparison, commit 78c91ae2c6 ("ipv4/udp: Add 4-tuple hash for connected
socket") measured 290,860 pps without the table and 1,889,658 with it at
500 connected sockets.
Patch 2's cost is not in throughput. Its stale entry keeps hash4_cnt raised
for the life of the socket, so every packet for that address and port is
sent through the 4-tuple lookup first; on IPv6 the entry is also matchable,
because __udp_disconnect() does not clear sk_v6_daddr. That last one is a
separate defect, which I will send on its own.
Neither patch has a selftest. Nothing in tree reports which 4-tuple bucket
a socket is filed under, so a test can only measure the cost indirectly.
What I did instead was add pr_info() to the hash4 paths and a knob that
dumps bucket occupancy, then run the same scenarios on two kernels
differing only by these patches; that is where the numbers above come from.
The instrumentation, the reproducers and the benchmark are at [1]. If
exposing the bucket through diag would be welcome, that would make both
defects testable in tree and I am glad to do it for net-next.
Removing the connect(AF_UNSPEC) limitation described in 644f9108f3 is a
side effect of patch 1 fixing the general case. I can make it narrower if
you would rather that limitation stayed.
Tooling, per Documentation/process/generated-content.rst: this series was
developed in an assisted session with an LLM. The assistant did most of
the code reading, wrote the instrumentation and reproducers behind [1],
drafted these changelogs, and ran the A/B builds and the regression
suites below. Every claim in these messages was checked against the
source, and the IPv6 behaviour described in patch 2 was confirmed at
runtime.
Tested on x86-64, IPv4 and IPv6. No regressions across reuseport_bpf,
reuseport_bpf_cpu, reuseport_addr_any.sh, reuseport_dualstack,
udpgso_bench.sh, udpgro_bench.sh and socket.
[1] https://github.com/shardulsdk-mpiric/linux/tree/udp-hash4-fix-verification
To: Willem de Bruijn <willemdebruijn.kernel@gmail.com>
To: "David S. Miller" <davem@davemloft.net>
To: Eric Dumazet <edumazet@google.com>
To: Jakub Kicinski <kuba@kernel.org>
To: Paolo Abeni <pabeni@redhat.com>
To: Simon Horman <horms@kernel.org>
To: Philo Lu <lulie@linux.alibaba.com>
To: Fred Chen <fred.cc@alibaba-inc.com>
To: Yubing Qiu <yubing.qiuyubing@alibaba-inc.com>
Cc: Kuniyuki Iwashima <kuniyu@google.com>
Cc: Willem de Bruijn <willemb@google.com>
Cc: Cambda Zhu <cambda@linux.alibaba.com>
Cc: Janak Bhatt <janak@mpiric.us>
Cc: Kalpan Jani <kalpan.jani@mpiricsoftware.com>
Cc: Shardul Bankar <shardulsb08@gmail.com>
Cc: netdev@vger.kernel.org
Cc: linux-kernel@vger.kernel.org
Signed-off-by: Shardul Bankar <shardul.b@mpiricsoftware.com>
====================
Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-0-718891af0d7a@mpiricsoftware.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
A UDP socket bound to a specific address and port keeps its entry in the
4-tuple hash table after it is disconnected:
sk binds to 127.0.0.1:21001
sk connects to 127.0.0.2:20001 // filed in the 4-tuple table
sk disconnects, connect(AF_UNSPEC) // still filed, peer now 0.0.0.0:0
__udp_disconnect() takes a socket out of that table only as a side effect
of ->rehash() or ->unhash(), and it skips ->rehash() when
SOCK_BINDADDR_LOCK is set and ->unhash() when SOCK_BINDPORT_LOCK is set.
commit 6996a2d2d0 ("udp: Unhash auto-bound connected sk from 4-tuple hash
table when disconnected.") fixed the same end state for a wildcard-bound
socket, by a path this one does not take.
The entry is counted whether or not anything hits it. hash4_cnt on the
hash2 slot stays raised for as long as the socket lives, so udp_has_hash4()
keeps sending every packet for that address and port through the 4-tuple
lookup first.
On IPv6 it can also be hit. __udp_disconnect() does not clear sk_v6_daddr,
so udp_v6_rehash() files the entry under the peer the socket was connected
to with a zero dport, and inet6_match() compares that same
field: a datagram from the former peer with a zero source port matches,
and source port zero is accepted on receive. On IPv4 the peer is cleared,
so a match would need a zero source address as well, which the routing
layer rejects as martian. The stale sk_v6_daddr is a separate defect, not
addressed here; removing the entry closes this path either way.
The entry can also be relocated. __udp_disconnect() clears sk_bound_dev_if,
so a subsequent SO_BINDTODEVICE calls ->rehash(), and because the receive
address is still specific udp_lib_rehash() moves the entry instead of
removing it, into the bucket that (rcv_saddr, num, 0, 0) hashes to -- a
pure function of the address and port, so every socket reaching this state
on one address and port collects in one bucket. The bucket cannot be chosen
from outside, as udp_ehashfn() is seeded with a per-boot secret. This last
one became reachable only with commit 644f9108f3 ("udp: Make rehash4
independent in udp_lib_rehash()"), which moved the hash4 handling out of a
branch a disconnected socket does not take; the stale entry itself dates
from the commit in Fixes.
Take the socket out of the table before __udp_disconnect() runs, while it
still matches how it was filed. This also reaches the wildcard case ahead
of udp_lib_rehash()'s udp_unhash4() branch, leaving that branch unreachable
from udp_disconnect(); removing it belongs in net-next. udp_disconnect()
and udp_abort() are the only UDP entries into __udp_disconnect(), which is
shared with raw, ping and l2tp sockets that are not struct udp_sock:
ping_prot.obj_size is sizeof(struct inet_sock), so udp_hashed4() on one
would read past the allocation.
Fixes: 78c91ae2c6 ("ipv4/udp: Add 4-tuple hash for connected socket")
Assisted-by: LLM
Signed-off-by: Shardul Bankar <shardul.b@mpiricsoftware.com>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-2-718891af0d7a@mpiricsoftware.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
A connected UDP socket that connects again to a different peer is not
re-filed in the 4-tuple hash table:
sk binds to 127.0.0.1:21001
sk connects to 127.0.0.2:20001 // filed under hash(sk, peer1)
sk connects to 127.0.0.3:20002 // still filed under hash(sk, peer1)
packet from 127.0.0.3:20002 // hash(sk, peer2) misses, so the
// lookup falls back to scoring the
// hash2 chain for this address
// and port
udp_lib_hash4() returns early when the socket is already hashed, assuming
->rehash() relocates it. ->rehash() runs from __ip{4,6}_datagram_connect()
only while the receive address is unset, which a second connect never is:
the first connect assigns it, whether the socket was bound to a specific
address or to the wildcard. commit 644f9108f3 ("udp: Make rehash4
independent in udp_lib_rehash()") added that early return and named
connect(AF_UNSPEC) as the way around it. That workaround does not help a
socket with both SOCK_BINDADDR_LOCK and SOCK_BINDPORT_LOCK set, because
__udp_disconnect() skips ->rehash() for the first and ->unhash() for the
second.
Delivery is correct either way.
Relocate the socket when the hash it is filed under differs from the one
requested, which is what commit 78c91ae2c6 ("ipv4/udp: Add 4-tuple hash
for connected socket") did before the early return became unconditional. It
is done here under hslot->lock, which that version did not take, to match
udp_lib_rehash() and udp_lib_unhash(). hslot2 is unchanged, so hash4_cnt
needs no adjustment, as in udp_lib_rehash(). A first connect is unaffected,
and IPv6 shares the code.
With 500 sockets on the port, a re-connected socket measured 522,553 pps
without this change and 2,055,078 with it. The UDP side was noted as
remaining work in [1].
Link: https://lore.kernel.org/netdev/apnHqmYZQ4yzOP4N@v4bel/ [1]
Fixes: 644f9108f3 ("udp: Make rehash4 independent in udp_lib_rehash()")
Assisted-by: LLM
Signed-off-by: Shardul Bankar <shardul.b@mpiricsoftware.com>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-1-718891af0d7a@mpiricsoftware.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Core:
- hci_conn: fix CIS hold ownership on reuse
- hci_sock: reject out-of-range OCF values
- hci_sock: validate event length before filtering
- L2CAP: validate frame length before control and FCS access
- RFCOMM: fix NULL dereference of dlc->session in RFCOMM_CONNINFO
- RFCOMM: Reject short EA=0 frames in rfcomm_recv_frame()
- ISO: release unused CIS holds after channel attach
- ISO: balance the parent hold in hci_bind_bis()
- SMP: reject Security Request over BR/EDR
- MGMT: fix race in read_unconf_index_list()
- MGMT: Dequeue pending mesh_send_sync entries on cancel
- BNEP: fix out-of-bounds reads on short RX/TX frames and control fallthrough
Drivers:
- btintel_pcie: validate device-supplied DMA indices
- btnxpuart: Fix skb leak in nxp_process_fw_dump()
-----BEGIN PGP SIGNATURE-----
iQJNBAABCgA3FiEE7E6oRXp8w05ovYr/9JCA4xAyCykFAmqxN44ZHGx1aXoudm9u
LmRlbnR6QGludGVsLmNvbQAKCRD0kIDjEDILKS7mD/43nt9IlCMp4fRi5eT3iV6J
phP/zJTiikgOMv87kTI0Q9OXY8Xl1nhIGrTiypXQIJNJGTW/OTHtMF+N55pA7pF4
exD0US01bKUcuopztHeP1Yk08CMKAtE9VTPLi/PdAQz6KTo7wcH7yFwfsO6avIk4
T/IXi9b8iKPBMKizjgUe8Uo6wrFoByD/o2VwTSQwtZOgWppUkCvVKKx10Y8EYJRG
UbDS8rpPtWN3t0fK5F4yVfkTP+9sA7uWb9bF2UoLcmov2ie33M50zf5giVofiNrX
pT6RWTCPfjkPdDro66l6wV4M8prEUVok0QhlqeYOUXSZJSOTqqXBTOorioABSeHz
i4QWH+KsqUFbaOJQuoF3v5jAccY8ZxKczRODK8Xt1+bY9O73lqpuWjQvoi98E/7i
vNCSU/bDpRS0MjYDTgQTG9c0rv9qxKI0GBC8gDVI78/nAzUIxdJ09/wDBm17aUpt
uxaBwYE1lhJwT0Lf+dwJrwVuTm5q5XotdRIsOwxZgfF25Ewk8ygUlEsQu/ioHmS+
q+9SL/mHb/wPI7j0YeVG2l7cHiHOZDEn0FW++WlgtF1Oj/HWUsz2UN/qdmpYZo1h
rGX7D2DdZQo8IOEmZZn7wqLNOxjPjFVkxgQtgWd5OhE0O+7/e2H4mNeDJyWs6rJp
lT1BJZnPQTFtwP41x7Hxng==
=cLTc
-----END PGP SIGNATURE-----
Merge tag 'for-net-2026-09-21' of git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth
Luiz Augusto von Dentz says:
====================
bluetooth pull request for net:
Core:
- hci_conn: fix CIS hold ownership on reuse
- hci_sock: reject out-of-range OCF values
- hci_sock: validate event length before filtering
- L2CAP: validate frame length before control and FCS access
- RFCOMM: fix NULL dereference of dlc->session in RFCOMM_CONNINFO
- RFCOMM: Reject short EA=0 frames in rfcomm_recv_frame()
- ISO: release unused CIS holds after channel attach
- ISO: balance the parent hold in hci_bind_bis()
- SMP: reject Security Request over BR/EDR
- MGMT: fix race in read_unconf_index_list()
- MGMT: Dequeue pending mesh_send_sync entries on cancel
- BNEP: fix out-of-bounds reads on short RX/TX frames and control fallthrough
Drivers:
- btintel_pcie: validate device-supplied DMA indices
- btnxpuart: Fix skb leak in nxp_process_fw_dump()
* tag 'for-net-2026-09-21' of git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth:
Bluetooth: RFCOMM: Reject short EA=0 frames in rfcomm_recv_frame()
Bluetooth: RFCOMM: fix NULL dereference of dlc->session in RFCOMM_CONNINFO
Bluetooth: btintel_pcie: validate device-supplied DMA indices
Bluetooth: bnep: fix out-of-bounds reads on short RX/TX frames and control fallthrough
Bluetooth: mgmt: fix race in read_unconf_index_list()
Bluetooth: L2CAP: validate frame length before control and FCS access
Bluetooth: ISO: balance the parent hold in hci_bind_bis()
Bluetooth: hci_sock: validate event length before filtering
Bluetooth: hci_sock: reject out-of-range OCF values
Bluetooth: ISO: release unused CIS holds after channel attach
Bluetooth: hci_conn: fix CIS hold ownership on reuse
Bluetooth: mgmt: Dequeue pending mesh_send_sync entries on cancel
Bluetooth: btnxpuart: Fix skb leak in nxp_process_fw_dump()
Bluetooth: SMP: reject Security Request over BR/EDR
====================
Link: https://patch.msgid.link/20260921135807.3459373-1-luiz.dentz@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
ip6_route_output_flags(), ip6_rt_put_flags(), and ip6_dst_check()
detect an uncached route by list_empty(&rt->dst.rt_uncached),
which replaced the static DST_NOCACHE flag check in commit
a4c2fd7f78 ("net: remove DST_NOCACHE flag").
When a device is unregistered, rt6_uncached_list_flush_dev()
unlinks uncached routes tied to the device from rt6_uncached_list.
Previously, they were moved to another list with list_move()
(__list_del_entry() + list_add()), and since commit 98aa546af5
("inet: remove (struct uncached_list)->quarantine"), the routes
are just unlinked with list_del_init().
If list_del_init() runs concurrently, list_empty() evaluates to
true; ip6_route_output_flags() calls dst_hold_safe() incorrectly
and ip6_rt_put_flags() skips ip6_rt_put(), leaking dst, and thus
dev tied via rt->from as well.
The same race is partially fixed by commit 9a6f0c4d57 ("dst:
fix races in rt6_uncached_list_del() and rt_del_uncached_list()").
Let's check rt6->dst.rt_uncached_list instead.
Note that IPv4 does not have the same issue.
Fixes: 98aa546af5 ("inet: remove (struct uncached_list)->quarantine")
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Reviewed-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260920191558.2990636-1-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
IPv6 XFRM policies may use different source and destination prefix
lengths. mlx5e_ipsec_policy_mask() builds the corresponding masks
independently, but setup_fte_addr6() installs each mask in the opposite
address field.
When the prefix lengths differ, this makes the source match use the
destination prefix and the destination match use the source prefix. The
resulting hardware rule can both miss traffic covered by the policy and
match traffic outside it.
Install each mask in its corresponding match field.
Fixes: ca7992f52c ("net/mlx5e: Properly match IPsec subnet addresses")
Cc: stable@vger.kernel.org
Signed-off-by: Andrea Parri <parri.andrea@gmail.com>
Reviewed-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260917115542.177675-1-parri.andrea@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Commit 8d75c338f0 ("sysctl: remove CONFIG_PROC_SYSCTL, it just mirrors
CONFIG_SYSCTL") renamed CONFIG_PROC_SYSCTL to CONFIG_SYSCTL in place,
which left the entry out of alphabetical order in the net and
packetdrill configs. The netdev CI check for sorted selftest configs
now fails for every patch that touches either file.
Fixes: 8d75c338f0 ("sysctl: remove CONFIG_PROC_SYSCTL, it just mirrors CONFIG_SYSCTL")
Signed-off-by: Yuya Kusakabe <yuya.kusakabe@gmail.com>
Reviewed-by: Joel Granados <joel.granados@kernel.org>
Link: https://patch.msgid.link/20260918-selftests-net-config-sort-v1-1-968ea6e8c1b7@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The receive fixup subtracts the Ethernet CRC from the reported packet
length, but compares that payload length against the whole remaining
receive buffer. The following copy starts after the three-byte header,
and the cursor advance consumes both that header and the four-byte CRC.
Require the payload to fit after SR_RX_OVERHEAD before copying it or
advancing to the next packet. The loop already ensures that the
remaining buffer is larger than the overhead, so the subtraction is
safe.
The issue was found by our static-analysis tool.
Fixes: c9b37458e9 ("USB2NET : SR9700 : One chip USB 1.1 USB2NET SR9700Device Driver Support")
Reviewed-by: Ethan Nelson-Moore <enelsonmoore@gmail.com>
Tested-by: Ethan Nelson-Moore <enelsonmoore@gmail.com>
Signed-off-by: Pengpeng Hou <hppiscas@163.com>
Link: https://patch.msgid.link/20260920034745.18468-1-hppiscas@163.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
dpll_pin_ref_sync_state_set() looks up the reference sync pin in the
pin->ref_sync_pins xarray, which is keyed by the sync pin's id (see
dpll_pin_ref_sync_pair_add() using xa_insert() with ref_sync_pin->id).
The pin id to operate on is supplied by userspace via DPLL_A_PIN_ID.
The lookup however used xa_find() with a ULONG_MAX limit, which returns
the first present entry with an index greater than or equal to the
requested id, not the entry stored exactly at that id. If userspace
passes an id that is not paired as a reference sync pin, but another
pin with a higher id is present in the xarray, xa_find() silently
returns that wrong pin and the subsequent ref_sync_set() operates on
it. The request only fails when the given id is larger than every
present key.
Use xa_load() for an exact-key lookup instead, mirroring the deletion
path in dpll_pin_ref_sync_pair_del().
Fixes: 58256a26bf ("dpll: add reference sync get/set")
Signed-off-by: Ivan Vecera <ivecera@redhat.com>
Link: https://patch.msgid.link/20260917143736.526221-1-ivecera@redhat.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Removing the legacy ioctl fallback made both hwtstamp NDOs mandatory. A
device that only timestamps in its PHY implements neither, so
SIOCSHWTSTAMP fails with EOPNOTSUPP before anything looks at the PHY and
PTP stops working there.
The check only ever picked the legacy path. That path is gone, so drop it
and test where the NDOs are actually called.
SIOCGHWTSTAMP is new here, not restored. The old path went through
phy_mii_ioctl(), which only handled SIOCSHWTSTAMP.
Such a device now returns -ENODEV while absent instead of -EOPNOTSUPP,
like the ones that do implement the NDOs.
Fixes: 5062245a5a ("net: remove legacy way to get/set HW timestamp config")
Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
Reviewed-by: Kory Maincent <kory.maincent@bootlin.com>
Link: https://patch.msgid.link/20260918095540.34286-1-nb@tipi-net.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Before the cited commit, fib6_nh_flush_exceptions() always set
from->exception_bucket_flushed = 1 under rt6_exception_lock to
prevent rt6_insert_exception() from inserting a new exception
for a dying fib6_info.
The flag was replaced with the FIB6_EXCEPTION_BUCKET_FLUSHED
bit stored in nh->rt6i_exception_bucket.
The problem is that now the bit is only set when the bucket
is not NULL and fib6_nh_flush_exceptions() is called from
fib6_nh_release() after fib6_ref has already reached zero.
If rt6_insert_exception() is called while the target fib6_info
is being removed via fib6_purge_rt(), a new exception could be
created successfully because rt6_flush_exceptions() no longer
sets the bit.
This creates a reference cycle between the fib6_info and the
exception route, leaking the fib6_info, its nexthop device,
and all per-CPU routes in fib6_nh->rt6i_pcpu, which stalls netdev
unregistration.
[ 34.680602] unregister_netdevice: waiting for gre6 to become free. Usage count = 68
[ 44.920675] unregister_netdevice: waiting for gre6 to become free. Usage count = 68
[ 55.176582] unregister_netdevice: waiting for gre6 to become free. Usage count = 68
Let's call fib6_drop_pcpu_from() before rt6_flush_exceptions(),
to set fib6_destroying before rt6_exception_lock, and check
f6i->fib6_destroying in rt6_insert_exception().
Note that FIB6_EXCEPTION_BUCKET_FLUSHED logic is dead and
we can clean it up in net-next.
Fixes: cc5c073a69 ("ipv6: Move exception bucket to fib6_nh")
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260918082209.2853582-1-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The original code used DMA_ATTR_FORCE_CONTIGUOUS, which could exhaust
the CMA pool when a large number of VFs were requested.
Fix this by switching to the DMA streaming API. This is equivalent on
Octeon platforms, which provide full I/O coherency via the SMMU.
Cc: Leon Romanovsky <leon@kernel.org>
Fixes: 73d33dbc07 ("octeontx2-af: Use DMA_ATTR_FORCE_CONTIGUOUS attribute in DMA alloc")
Signed-off-by: Ratheesh Kannoth <rkannoth@marvell.com>
Reviewed-by: Leon Romanovsky <leon@kernel.org>
Link: https://patch.msgid.link/20260916022111.1083017-1-rkannoth@marvell.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
-----BEGIN PGP SIGNATURE-----
iQIzBAABCgAdFiEEjF9xRqF1emXiQiqU1w0aZmrPKyEFAmqtHzMACgkQ1w0aZmrP
KyGtkQ//bMKQGEudQKCMCtQmPqaHyW1ajmAo17aumfczaE/nSjqrsGY5Ahw7OKlQ
4otcdPvI4qpV9sLTg41KFaIHIC5sozxt4Q3m3RNB2TbyCkGn9xpSpZxM5IpHybvE
83tVjSA0wpfIqxBEqKUqk8Z9AXtBLo/JocdfYry+6JUyj4PM76X2ViKpzaPbpoMU
1mndfLAYtADIIvs3805CmfdJmOkoSV6XCEsiNutPrJhiRfN4xJZ9leP9xb1zA0IQ
cnqiaw1xkTcFyWCicu4MqOkEALRknr9SL2yX1S9wx5Q6WHwU9JXUeQTlvfv7OoVP
uxuMlNr3WcbwHC9e1GfOHapzjrYgnvEe2Z79i2GFh51Ci+5L9Yr9XCQ/fc6G5NNZ
3W52kh35s3lXq32hll9Tkr7pf4cKLBA+IAJ19VNlRfMrPB0cz4EqbIZ6xNNuLqdh
DbEb3VgTT2dHwuGxEshJVmSfzfR+VeHBG2ZRlRmZElfhViHEwgPaAxkaJhNpPyub
qmHbZCXK0BVp/UrGHDm5rmHJtdkwprXY9YceZBRfW8Fr2Ler4rWQvy+uo9sRRFgN
oF9B6qzSl78THBv3UDB3U+aWuDv0I+VlDub0DKf0k9iWg8OoMlM9yw6OE9ULHt3m
LRvSAfF0LnF/z7byzumUYMVjpVuIVueDzasGrEZjWbHHE7MCsp0=
=RfSQ
-----END PGP SIGNATURE-----
Merge tag 'nf-26-09-18' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf
Pablo Neira Ayuso says:
====================
Netfilter/IPVS fixes for net
The following patchset contains Netfilter/IPVS fixes for net, they are:
1) Set on HW_DEAD after HW_PENDING is cleared in the flowtable offload
to ensure GC does not zap it, from Jérémy Jean.
2) Hold the nfnetlink_queue mutex while removing the queue instance
from the netlink notifier that handles NETLINK_URELEASE to fix a
possible race with the UNBIND command. From Florian Westphal.
3) Reject route with NULL rt6i_idev in ip6t_rpfilter. From Weiming Shi.
4) Reject rtinfo->addrnr set to zero from ip6t_rt .checkentry path.
This also fortifies the datapath loop as per Florian's request.
From Luxiao Xu.
5) Fix checksuming in nft_synproxy for IPv6, from Karl Mehltretter.
6) Revalidate ihl before calling icmp_send() in IPVS,
from Julian Anastasov.
7) Fix suspicious RCU usage splat in ctnetlink with expectations.
8) Check for expired catchall elements in the insert and deactivate
path. From Aohan Mei.
* tag 'nf-26-09-18' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf:
netfilter: nf_tables: skip expired catchall elements on insert and delete
netfilter: ctnetlink: fix suspicious RCU usage in expect_iter_name
ipvs: revalidate ihl before icmp_send
netfilter: nft_synproxy: use the family-aware checksum helper
netfilter: ip6t_rt: fix zero-address non-strict match out-of-bounds read
netfilter: ip6t_rpfilter: reject routes without inet6_dev
netfilter: nfnetlink_queue: hold nfnl mutex in event notifier
netfilter: flowtable: publish HW_DEAD after worker is done
====================
Link: https://patch.msgid.link/20260918112844.194503-1-pablo@netfilter.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
While rfcomm_recv_frame() verifies that skb->len is at least
sizeof(*hdr) + 1 (4 bytes: 3-byte header + 1-byte FCS), an RFCOMM frame
with an extended 2-byte length field (!__test_ea(hdr->len)) has a 4-byte
header plus a 1-byte FCS (5 bytes minimum, sizeof(*hdr) + 2).
When a 4-byte RFCOMM frame with EA == 0 arrives:
1. The initial skb->len < sizeof(*hdr) + 1 check passes (4 < 4 is false).
2. Trimming the FCS byte decrements skb->len to 3.
3. If __check_fcs() succeeds, skb_pull(skb, 4) fails (4 > 3) and returns
NULL without advancing skb->data.
4. Because the return value of skb_pull() is ignored, the un-pulled
3-byte struct rfcomm_hdr remains at skb->data and is either queued as
application payload via rfcomm_recv_data() or parsed as a multiplexer
control command via rfcomm_recv_mcc() on DLCI 0.
Fix this by extending the length check in rfcomm_recv_frame() to also
require skb->len >= sizeof(*hdr) + 2 when !__test_ea(hdr->len).
Fixes: b230e5bf50 ("Bluetooth: RFCOMM: validate skb length in rfcomm_recv_frame")
Assisted-by: LLM
Signed-off-by: Hui Peng <benquike@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
The RFCOMM_CONNINFO getsockopt handler accepts a socket that is not
connected as long as deferred setup is enabled:
if (sk->sk_state != BT_CONNECTED &&
!rfcomm_pi(sk)->dlc->defer_setup) {
err = -ENOTCONN;
break;
}
l2cap_sk = rfcomm_pi(sk)->dlc->session->sock->sk;
dlc->defer_setup is set in rfcomm_sock_init() when rfcomm_connect_ind()
creates a child socket for an incoming connection on a listening socket
that has BT_DEFER_SETUP enabled. It is never cleared afterwards. The
session, however, can go away underneath it.
rfcomm_recv_disc() forces the dlc state before tearing it down:
d->state = BT_CLOSED;
__rfcomm_dlc_close(d, err);
The RFCOMM_DEFER_SETUP early return in __rfcomm_dlc_close() only covers
BT_CONNECT, BT_CONFIG, BT_OPEN and BT_CONNECT2, so with the state
already BT_CLOSED that switch does not match and the function falls
through to rfcomm_dlc_unlink(), which sets d->session = NULL, while
d->defer_setup stays 1.
A getsockopt(SOL_RFCOMM, RFCOMM_CONNINFO) on the accepted socket after
that point therefore skips the -ENOTCONN path -- sk->sk_state is
BT_CLOSED, but dlc->defer_setup is still set -- and dereferences the
NULL session. No race is needed: once the DISC has been processed, the
dereference is unconditional.
Reproduced on a KASAN kernel under QEMU with a BR/EDR peer emulated over
/dev/vhci: the peer brings up an ACL link, opens L2CAP on the RFCOMM
PSM, starts a session and sends SABM for a channel bound with
BT_DEFER_SETUP, and sends DISC for that dlci after the socket has been
accepted. getsockopt(SOL_RFCOMM, RFCOMM_CONNINFO) on the accepted
socket then hits:
Oops: general protection fault, probably for non-canonical address
0xdffffc0000000002: 0000 [#1] SMP KASAN PTI
KASAN: null-ptr-deref in range [0x0000000000000010-0x0000000000000017]
CPU: 1 UID: 0 PID: 150 Comm: init Tainted: G B 7.3.0-rc3-g5dd1818b15d9
Hardware name: QEMU Standard PC (i440FX + PIIX, 1996)
RIP: 0010:rfcomm_sock_getsockopt+0x529/0x780
Call Trace:
<TASK>
do_sock_getsockopt+0x3ad/0x7d0
__sys_getsockopt+0x10e/0x1b0
__x64_sys_getsockopt+0xc2/0x160
do_syscall_64+0xda/0x4b0
entry_SYSCALL_64_after_hwframe+0x77/0x7f
</TASK>
0x10 is the offset of sock in struct rfcomm_session;
rfcomm_sock_getsockopt_old() is inlined into rfcomm_sock_getsockopt().
Commit 43a556b2fd ("Bluetooth: RFCOMM: take rfcomm_mutex for the
deferred setup accept") fixed the same "a remote DISC clears the session
while deferred setup is still flagged" problem in rfcomm_dlc_accept();
this is the remaining instance of it, in the getsockopt path.
Deferred setup only leaves a socket usable here once it has reached
BT_CONNECT2, so restrict the exception to that state and check that a
session is actually present before following it.
Fixes: bb23c0ab82 ("Bluetooth: Add support for deferring RFCOMM connection setup")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Hui Peng <benquike@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
In btintel_pcie_msix_rx_handle(), the driver processes RX completion
descriptors (urbd1) written by the PCIe device into DMA-coherent memory.
urbd1->frbd_tag (a 16-bit field fully controlled by the device firmware
via DMA) is used directly as an array index into rxq->bufs[] without any
bounds check. rxq->bufs[] has only BTINTEL_PCIE_RX_DESCS_COUNT (64)
entries, while frbd_tag can be any value 0-65535. A malicious or
malfunctioning device can write an out-of-range frbd_tag, causing the
driver to dereference an out-of-bounds data_buf pointer.
Additionally, cr_hia is read from a DMA-shared index array also writable
by the device; if the device sets cr_hia >= rxq->count, the while-loop
never terminates because cr_tia is wrapped via modulo rxq->count and can
never equal an out-of-range cr_hia.
Add bounds validation for cr_hia and frbd_tag in the RX path, and cr_hia
in the TX path. Log invalid values with bt_dev_err before returning.
Fixes: c2b636b3f7 ("Bluetooth: btintel_pcie: Add support for PCIe transport")
Signed-off-by: Ravindra <ravindra@intel.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
Fix multiple out-of-bounds reads in Bluetooth BNEP frame processing:
1. In bnep_rx_frame() and bnep_ctrl_frame() (net/bluetooth/bnep/core.c),
use pskb_may_pull() to verify the BNEP header, control type byte,
filter count, and extension headers exist before reading them, and
return 0 after handling BNEP_CONTROL instead of falling through to
Ethernet frame submission when no extension headers follow.
2. In bnep_net_xmit() (net/bluetooth/bnep/netdev.c), verify skb->len >=
ETH_HLEN with pskb_may_pull() before reading the 14-byte Ethernet
header to prevent an out-of-bounds heap read and infoleak on short
AF_PACKET TX frames.
Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Assisted-by: LLM
Signed-off-by: Hui Peng <benquike@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
The classify-loop fix bounds a walk's non-descending hops, so the guard
must not misfire on a legal walk that reaches its leaf through a
level-drift lateral chain. Add a case that builds exactly that chain and
asserts traffic still reaches the chain's own leaf.
A lateral hop can only exist because a bind was legal when it was made
and a later class add raised the target's level, so the setup binds each
hop while the target is still a leaf and only then deepens it: bind
1:1 -> 1:2 while 1:2 is a leaf, add 1:20 under 1:2, add 1:3 and bind
1:2 -> 1:3 while 1:3 is a leaf, then add 1:30 and 1:31 under 1:3 and
bind 1:3 -> 1:31. The walk root -> 1:1 -> 1:2 -> 1:3 -> 1:31 then takes
two lateral hops and must reach leaf 1:31.
The default class is 1:30, distinct from the asserted leaf, and the
verify pattern is anchored to the 1:31 stats line, so neither a
fall-through to the default nor a nonzero count on another class can
satisfy the check. On the patched kernel the test passes; with the bound
forced to zero the walk falls to the default and 1:31 stays idle, so the
test fails.
Reviewed-by: Victor Nogueira <victor@mojatatu.com>
Tested-by: hybris <hybris@mojatatu.ai>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-CTUU.v3.20260916184908@mojatatu.com.2
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
hfsc_classify() applies the "filter may only point downwards" level check
only when the filter result carries no bound class. A filter created with
a flowid gets res.class set once at bind time, so the check never runs for
it during classification. hfsc_adjust_levels() can later raise a class's
level without revalidating existing bindings, leaving two binds that were
each legal at bind time pointing at each other; the classify walk then
bounces between two interior classes forever with the qdisc lock held and
BH disabled — a soft lockup from a single packet. The stuck walk trips
the watchdog:
watchdog: BUG: soft lockup - CPU#3 stuck for 13s! [ping:444]
RIP: 0010:u32_classify+0x542/0x17f0
...
tcf_classify+0x66/0xa0
hfsc_enqueue+0x166/0xdf0
Bound the traversal with a budget of non-descending hops, the only way a
configured walk can move without descending the class tree once levels
drift after bind time. The budget is cumulative over the whole walk and
is deliberately not reset on a descending hop: a chain that alternates a
descent with a lateral hop would return the budget every lap and never
trip. Descending hops never decrement it, so legitimately deep trees are
unaffected and a terminating lateral chain still classifies normally.
Drop the packet with a rate-limited warning once the budget is exhausted,
mirroring the merged HTB fix.
This is a follow-up to commit 729c4896ab ("net/sched: sch_htb: limit
htb_classify inner-class filter hops"), which bounded the same classify
loop on the HTB side but left the HFSC walk unbounded.
Conditions to recreate the bug:
- CONFIG_NET_SCHED, CONFIG_NET_SCH_HFSC, CONFIG_NET_CLS_U32,
CONFIG_LOCKUP_DETECTOR.
- Build a cycle with two legal-at-bind-time flowid binds and a level
drift: class X 1:1 (child of root) with leaf child 1:10; class Y 1:2
(sibling of X) with children 1:20 and 1:200; root u32 filter flowid
1:1; filter on X flowid 1:2 (legal when Y is a leaf); after Y's level
rises to 2, filter on Y flowid 1:1 (legal then). Send one packet (ping
on the device). Unfixed kernel: classify spins with the qdisc lock
held; with softlockup_panic=1 it panics.
- Reachable from unprivileged user via unshare -Urn (CAP_NET_ADMIN).
Fixes: a2f7922713 ("net_sched: sch_hfsc: fix classification loops")
Reported-by: Sashiko (gemini + nipa) <sashiko-bot@kernel.org>
Closes: https://lore.kernel.org/netdev/QDISC-CTUU.v2.20260913192614@mojatatu.com/
Link: https://sashiko.dev/#/patchset/QDISC-CTUU.v2.20260913192614@mojatatu.com
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/QDISC-CTUU.v2.20260913192614%40mojatatu.com
Reviewed-by: Victor Nogueira <victor@mojatatu.com>
Tested-by: hybris <hybris@mojatatu.ai>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-CTUU.v3.20260916184908@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
__vlan_insert_inner_tag() only guarantees head room via skb_cow_head(),
never that mac_len bytes of MAC header are present. Its ETH_HLEN
wrappers - __vlan_insert_tag() under skb_vlan_push(), and
vlan_insert_tag() under validate_xmit_vlan() on the generic transmit
path - therefore rewrite the first 16 bytes at skb->data: a 12-byte
memmove plus two 2-byte stores at +12 and +14. No caller supplies the
bound, while the pop helpers use skb_ensure_writable()/pskb_may_pull().
An IFF_TUN device has hard_header_len == 0, so packet_snd() accepts a
one-byte AF_PACKET/SOCK_RAW frame. The first vlan push only sets a
hwaccel tag; the next - clsact "action vlan push" or
bpf_skb_vlan_push() - enters the helper with skb->len still 1. The
head comes from skbuff_small_head without __GFP_ZERO, so each push
drags bytes from beyond skb->tail into the frame. After three the
one-byte send leaves as 13 bytes carrying 11 bytes of uninitialised
slab:
0000: 5a b3 62 12 80 88 ff ff 00 b3 62 12 81
`------------------------------'
only 0x5a was sent; the rest is slab, here the top 56 bits of a
linear-map address
Require the MAC header the helper rewrites to be present, so such a
frame is dropped rather than transmitted.
Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Cc: stable@vger.kernel.org
Reported-by: co+0ea1ac045375cf05@bugs.sh
Signed-off-by: Xiang Mei <xmei5@asu.edu>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260915083152.705309-1-xmei5@asu.edu
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
catc_rx_done() walks a multi-packet URB, reading a two-byte length from
each packet header. Its bound, pkt_len > urb->actual_length, ignores the
header offset and compares against the whole transfer rather than the
bytes left from pkt_start, so a crafted packet header makes
skb_copy_to_linear_data() read past the buffer.
A length below ETH_HLEN is also accepted, including zero, and
eth_type_trans() then reads a MAC header from the uninitialised tailroom
of a shorter skb. The is_f5u011 branch takes its length straight from
the transfer, so a zero-length URB reaches the same path.
Track the bytes remaining from the current packet, and reject a header
that does not fit, a length past what is left, and a length below an
Ethernet header.
A transfer shorter than an Ethernet header, including a zero-length one,
previously became a runt skb passed to netif_rx() and counted as
received; it is now counted in rx_length_errors and ends the walk.
Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Signed-off-by: Aamir Ahmed <elb12345@hotmail.co.uk>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/AS8P251MB00015FD7716F38C345619B56C8BB2@AS8P251MB0001.EURP251.PROD.OUTLOOK.COM
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
tcf_action_delete() drops the reference held by its lookup before calling
tcf_idr_delete_index() with the saved action index. An unlocked
classifier can remove that action and reserve the same IDR slot with
ERR_PTR(-EBUSY) in between.
tcf_idr_delete_index() only checks the lookup result for NULL. It
therefore treats the reservation as a tc_action and dereferences
tcfa_bindcnt. A hardware execution breakpoint was used to schedule the
interleaving without changing the kernel source. KASAN reported this
decoded trace:
BUG: KASAN: null-ptr-deref in tca_action_gd+0x5b9/0x1010
Read of size 4 at addr 0000000000000010 by task poc/150
Oops: general protection fault, probably for non-canonical address 0xdffffc0000000002
RIP: tca_action_gd+0x5c0/0x1010:
arch_atomic_read at arch/x86/include/asm/atomic.h:23
raw_atomic_read at include/linux/atomic/atomic-arch-fallback.h:457
atomic_read at include/linux/atomic/atomic-instrumented.h:33
tcf_idr_delete_index at net/sched/act_api.c:766
tcf_action_delete at net/sched/act_api.c:1859
tcf_del_notify at net/sched/act_api.c:2014
tca_action_gd at net/sched/act_api.c:2064
R13: 0000000000000010 R15: fffffffffffffff0
Kernel panic - not syncing: Fatal exception
R15 contains ERR_PTR(-EBUSY), and adding the tcfa_bindcnt offset produces
the address in R13. With the guard applied, the same reproducer returned
-ENOENT without a KASAN report or panic. Treat error pointers as absent
and return -ENOENT.
Fixes: 0190c1d452 ("net: sched: atomically check-allocate action")
Cc: stable@vger.kernel.org
Reported-by: Xiang Mei <xmei5@asu.edu>
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Link: https://patch.msgid.link/20260914065123.4109709-2-bestswngs@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Commit fa8fca8871 ("ipv4: validate IPV4_DEVCONF attributes properly")
added validation of IFLA_INET_CONF attributes, and in the process
changed the call of nla_for_each_nested() to nla_parse_nested(). A
side effect of this change is that the IFLA_INET_CONF option is now
tested for NLA_F_NESTED being set, and fails if it is not. Prior to the
commit there was no check of NLA_F_NESTED.
Change nla_parse_nested() to nla_parse(). This restores the previous
functionality of not checking NLA_F_NESTED, thereby allowing code that
(incorrectly) doesn't set NLA_F_NESTED to continue to work.
This issue was identified because keepalived started logging errors when
it was configuring macvlans that it created.
Fixes: fa8fca8871 ("ipv4: validate IPV4_DEVCONF attributes properly")
Signed-off-by: Quentin Armitage <quentin@armitage.org.uk>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260915213320.1527029-2-quentin@armitage.org.uk
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Commit 339ccec8d4 ("net/mlx5: Enable MACsec offload feature for VLAN
interface") added NETIF_F_HW_MACSEC unconditionally to vlan_features so
that VLAN devices could inherit MACsec offload support.
mlx5e_build_nic_netdev subsequently copies vlan_features into
hw_features and features. As a result, all mlx5e NIC netdevices
advertise MACsec hardware offload, even when the firmware does not
support it and the driver does not install macsec_ops.
Set the MACsec feature bits in mlx5e_macsec_build_netdev, after device
capabilities have been validated. This preserves MACsec-over-VLAN
support and the ethtool feature control on capable devices, without
advertising either on unsupported hardware.
Fixes: 339ccec8d4 ("net/mlx5: Enable MACsec offload feature for VLAN interface")
Cc: stable@vger.kernel.org
Reviewed-by: Tariq Toukan <tariqt@nvidia.com>
Signed-off-by: Ralf Lici <ralf@mandelbit.com>
Link: https://patch.msgid.link/20260917122724.654639-1-ralf@mandelbit.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
sctp_assoc_update_retran_path() can loop forever when every remaining
transport, including retran_path, is SCTP_UNCONFIRMED: the state check
runs before the wraparound test, so the loop cannot observe that it has
completed a full pass.
Fix this by considering a transport only when it is not UNCONFIRMED,
then checking whether the walk has returned to retran_path. This makes
the full-pass termination independent of the transport state while
preserving the existing fallback selection semantics.
Also restore the NULL guard around the retran_path assignment. In the
all-UNCONFIRMED case there is no eligible replacement transport, and
installing NULL would leave later retransmit-path users and the debug
print with a NULL path.
Fixes: 4c47af4d5e ("net: sctp: rework multihoming retransmission path selection to rfc4960")
Signed-off-by: Yiqi Sun <sunyiqixm@gmail.com>
Acked-by: Xin Long <lucien.xin@gmail.com>
Link: https://patch.msgid.link/20260915095017.942213-1-sunyiqixm@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Alexander Duyck says:
====================
eth: fbnic: a collection of fixes
This series collects a handful of independent fbnic fixes for issues on
released kernels, plus one core ethtool fix needed by the fbnic offline
self test.
The first patch keeps rtnl_lock held on the ethtool ioctl path for the self
test. Since the ioctl path became rtnl-optional for ops-locked drivers,
fbnic's offline self test (which brings the interface down and up via
netif_close()/netif_open()) runs holding only the instance lock, tripping a
lockdep splat / ASSERT_RTNL and reconfiguring the device without the lock
it requires. A similar issue was found with Broadcom drivers so we expanded
the scope for v2 to just have the rtnl lock held for all selftest calls.
The second addresses a comparison issue in that we were limiting the
maximum number of standalone Tx queues to one less than the maximum number
of Tx queues. To resolve this it was just a matter of replacing a "<" with
a "<=".
The third addresses an indexing issue with netdev queues on fbnic in which
the NAPI vector was assumed to be findable as the Rx index modulo the
number of NAPI vectors. However this is actually not the case for if Tx
only and Rx only queues are setup. To resolve this we make use of the
cached NAPI pointer in the netdev Rx queues themselves.
The fourth patch fixes a NULL pointer dereference on unbind after a failed
PCIe error recovery: fbnic_pm_suspend() frees the napi vectors via a direct
ndo_stop() while leaving netif_running() true, and when slot_reset ->
resume fails the data path is never re-allocated. To prevent the panic we
reset num_napi to 0 before we free the IRQs which prevents walking the
unallocated napi vectors when we unbind the interface later.
The last two patches address the FW mailbox. One sets AW_FLUSH_MODE
alongside AW_FLUSH when tearing down the Rx ring, so the write pipeline
actually drains the staged requests instead of hanging on the BME halt.
The other handles completions flagged with FW_ERR on both mailboxes, which
the driver previously ignored. This resulted in us parsing a stale Rx page,
and spinning the capabilities poll to a timeout on a healthy ring.
====================
Link: https://patch.msgid.link/178941996343.7700.9376081102002673062.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The firmware can complete a mailbox descriptor while also setting FW_ERR
to indicate it could not process the request, for example on a mailbox
DMA error. The completion carries no valid data.
The driver did not check FW_ERR. On the Rx mailbox it would sync and
parse the stale page as a normal message, and on the Tx mailbox it
silently freed the request. If the initial capabilities exchange in
fbnic_mbx_poll_tx_ready() hit FW_ERR -- on the Tx request or on the Rx
response descriptor -- no response was parsed and the poll spun until it
timed out even though the ring was healthy.
Check FW_ERR on both mailboxes. Count it per-mailbox in
fbnic_fw_mbx.resp_error, which is also shown in debugfs, warn (rate
limited, since the bit is firmware controlled), and drop the Rx page
instead of parsing it.
In fbnic_mbx_poll_tx_ready() re-issue the capabilities request when
either the Tx or the Rx resp_error counter advances, so a FW_ERR on the
request or on its response triggers a retry rather than a timeout. A
valid capabilities response is honored before the retry check, so a
response parsed in the same poll as an unrelated FW_ERR is not discarded.
The counters are mailbox-wide rather than keyed to the capabilities
request; that is sufficient here because the exchange runs during
bring-up before any other mailbox traffic, and any spurious retry is
bounded by the existing 10s timeout.
Fixes: da3cde0820 ("eth: fbnic: Add FW communication mechanism")
Signed-off-by: Alexander Duyck <alexanderduyck@fb.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942023343.7700.9423398932961964439.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
When tearing down the FW mailbox Rx ring, fbnic_mbx_reset_desc_ring()
writes AW_CFG with FLUSH set and everything else, BME included, cleared.
Clearing BME halts the device's writes to the host but leaves the staged
requests parked in the PUL write pipeline rather than draining them, so
on the write path FLUSH alone never terminates the outstanding requests
and the flush the firmware waits on never completes.
Add the FLUSH_MODE definition and set both bits so the staged writes
drain out of the pipeline on their own. BME stays cleared, so nothing
lands on the host; it is restored later in fbnic_mbx_init_desc_ring()
when the ring is rebuilt, once the outstanding writes are gone.
The read path is unaffected. AR_CFG has no equivalent mode bit and
AR_FLUSH terminates the outstanding reads by itself, so it is left as
is.
Both writes remain plain stores rather than read-modify-writes. That is
deliberate: the matching write in fbnic_mbx_init_desc_ring() restores
BME and the TLP attributes, and clears both flush bits as a side effect.
Fixes: 3b12f00ddd ("fbnic: Gate AXI read/write enabling on FW mailbox")
Signed-off-by: Alexander Duyck <alexanderduyck@fb.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942022583.7700.11050671998277309744.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
fbn->num_napi is the count of live napi vectors, each of which owns an
IRQ. The PM path had freed them without clearing the count.
fbnic_pm_suspend() tears the datapath down via ndo_stop() and frees the
IRQs, but leaves netif_running() true so resume knows to re-open. Resume
rebuilds the datapath in __fbnic_pm_resume() and fbnic_reset_queues() sets
num_napi and __fbnic_open() re-allocates the vectors.
When the datapath is torn down but never rebuilt, num_napi is left
pointing at freed vectors under 2 different scenarios:
- a PCIe error recovery that fails (fbnic_err_slot_reset() ->
__fbnic_pm_resume() returns an error -> PCI_ERS_RESULT_DISCONNECT), so
.resume never runs; or
- an __fbnic_open() that fails partway on resume and unwinds, freeing
the vectors after fbnic_reset_queues() has already set num_napi.
The netdev is then running with num_napi > 0 but napi[] freed, and the
eventual remove/unbind close re-enters fbnic_down() -> fbnic_dbg_down()
and dereferences the freed vectors:
BUG: kernel NULL pointer dereference, address: 0000000000000210
RIP: fbnic_dbg_down+0x28
Clear num_napi when the vectors are freed: in the suspend teardown (a
good resume re-establishes it before __fbnic_open()) and on the resume
open failure. A redundant ndo_stop() then walks an empty napi[]. The
normal ndo_stop() down/up cycle is untouched and keeps num_napi for the
next ndo_open().
Fixes: bc6107771b ("eth: fbnic: Allocate a netdevice and napi vectors with queues")
Signed-off-by: Alexander Duyck <alexanderduyck@fb.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942021809.7700.10804028989308077839.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>