gmac_clk_enable() enables the bulk clocks first and then the optional
PHY clock. If clk_prepare_enable() on the PHY clock fails, the function
returns without rolling back the bulk clocks, and bsp_priv->clk_enabled
stays false, so the later gmac_clk_enable(bsp_priv, false) becomes a
no-op and the bulk clock references are leaked.
Add the missing clk_bulk_disable_unprepare() on that failure path.
Fixes: ea449f7fa0 ("net: ethernet: stmmac: dwmac-rk: rework optional clock handling")
Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Reviewed-by: Heiko Stuebner <heiko@sntech.de>
Acked-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Signed-off-by: Coia Prant <coiaprant@gmail.com>
Link: https://patch.msgid.link/20260923123713.3137146-1-coiaprant@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
prb_calc_retire_blk_tmo() computes in 32-bit int arithmetic:
mbits = (blk_size_in_bytes * 8) / (1024 * 1024);
If I'm reading the validation right, tp_block_size is user
controlled and packet_set_ring() only rejects values that are <= 0
as int or not page aligned, so a 256MiB block goes right through
(and alloc_one_pg_vec_page() even has a vzalloc fallback for it).
0x10000000 * 8 wraps to INT_MIN, and on a NIC reporting 1 Gbps
(div == 1) the function ends up returning -2047.
The condition is actually (8 * size) mod 2^32 >= 2^31 && div == 1,
so the trigger set is [256,512), [768,1024), [1280,1536) and
[1792,2048) MiB. Other sizes wrap to non-negative values and faster
links divide the unsigned value back below 2^31, which is why this
doesn't blow up for everyone.
What makes it fatal is what happens next in init_prb_bdqc():
p1->interval_ktime = ms_to_ktime(prb_calc_retire_blk_tmo(...));
hrtimer_start(&p1->retire_blk_timer, p1->interval_ktime,
HRTIMER_MODE_REL_SOFT);
A negative relative timeout expires immediately. The callback
unconditionally returns HRTIMER_RESTART, and hrtimer_forward() turns
the negative interval into hrtimer_resolution:
if (interval < hrtimer_resolution)
interval = hrtimer_resolution;
So the SOFT timer re-fires at the maximum rate forever, holding
sk_receive_queue.lock each pass. One CPU spins in softirq until the
socket is closed. Repeat with more rings and the machine is gone.
The overflow itself is ancient - it was introduced together with
TPACKET_V3 in f6fb8f100b ("af-packet: TPACKET_V3 flexible buffer
implementation."). Its effect prior to f7460d2989 ("net:
af_packet: Use hrtimer to do the retire operation", v6.18) was not
as clear-cut, though: the return value was stored into an unsigned
short retire_blk_tov, so a negative result was truncated, and a
0-jiffy delay loop could be programmed as well. Neither is nearly
as detrimental as the immediate maximum-rate spin the hrtimer
conversion turned it into.
(Unrelated to CVE-2019-20812 - that one was the ethtool failure path
returning 0, which now returns DEFAULT_PRB_RETIRE_TOV.)
Reproducer, needs CAP_NET_RAW (a --network host container has it by
default) and a 1 Gbps NIC (QEMU e1000 works):
int fd = socket(AF_PACKET, SOCK_RAW, htons(ETH_P_ALL));
bind(fd, ...);
int v = TPACKET_V3;
setsockopt(fd, SOL_PACKET, PACKET_VERSION, &v, sizeof(v));
struct tpacket_req3 req = {
.tp_block_size = 0x10000000,
.tp_block_nr = 1,
.tp_frame_size = 2048,
.tp_frame_nr = 0x10000000 / 2048,
.tp_retire_blk_tov = 0,
};
setsockopt(fd, SOL_PACKET, PACKET_RX_RING, &req, sizeof(req));
Compute in 64 bits instead. The operands are already bounded by the
existing validation, so nothing else changes. If you'd prefer a
different fix, just say so and I'll respin.
Fixes: f6fb8f100b ("af-packet: TPACKET_V3 flexible buffer implementation.")
Cc: stable@vger.kernel.org
Signed-off-by: Dairui Zhang <zhangdairui@gmail.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260923050101.1510064-1-zhangdairui@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
mon_timeout() evaluates dom_size(mon->peer_cnt) before it takes mon->lock,
while mon->peer_cnt is updated under that lock by tipc_mon_add_peer() and
tipc_mon_remove_peer(). The value can therefore be stale, and the decision
whether the local domain has to be recomputed can be based on an outdated
member count.
Read mon->peer_cnt inside the write_lock_bh(&mon->lock) protected region.
Fixes: 35c55c9877 ("tipc: add neighbor monitoring framework")
Signed-off-by: Ginger Li <ginger.jzllee@gmail.com>
Reviewed-by: Tung Nguyen <tung.quang.nguyen@est.tech>
Link: https://patch.msgid.link/20260922080909.21123-1-ginger.jzllee@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
MaxLinear GSW12x/GSW14x Ethernet Switch Errata Sheet states:
"An issue has been sporadically observed after device power-on on the first
link-up attempt in 100BASE-TX mode resulting in either the link-up taking a
long time, or failing to link-up altogether...
Workaround:
After power-on, enable Cable Diagnostic Mode for all ports and disable
it..."
Implement the proposed workaround unconditionally in the Intel XWAY driver
(MaxLinear GSW1xx switches incorporate Intel XWAY PHYs) because the
diagnostic bits have the same meaning even in older integral PHYs such as
GPY111/PEF7071/PHY11G. So it's not clear how to distinguish the affected
newer integrated PHYs, but the workaround should not hurt the older PHYs.
Cc: stable@vger.kernel.org
Fixes: 22335939ec ("net: dsa: add driver for MaxLinear GSW1xx switch family")
Signed-off-by: Alexander Sverdlin <alexander.sverdlin@siemens.com>
Reviewed-by: Andrew Lunn <andrew@lunn.ch>
Link: https://patch.msgid.link/20260922075251.23386-1-alexander.sverdlin@siemens.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
smcr_port_err() traverses smc_lgr_list.list without holding
smc_lgr_list.lock, allowing a concurrent smc_lgr_terminate_sched()
to free an lgr while it is still being dereferenced.
Hold smc_lgr_list.lock across the traversal. Update
smc_ib_gid_check() to call smcr_port_err() after releasing the lock.
Fixes: 541afa10c1 ("net/smc: add smcr_port_err() and smcr_link_down() processing")
Reviewed-by: Mahanta Jambigi <mjambigi@linux.ibm.com>
Signed-off-by: Sidraya Jayagond <sidraya@linux.ibm.com>
Reviewed-by: Dust Li <dust.li@linux.alibaba.com>
Link: https://patch.msgid.link/20260922073149.474762-1-sidraya@linux.ibm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
__rds_conn_create() computes npaths from the caller's transport before
it decides whether a connection to one of the host's own addresses is
to be handled by the loopback transport instead. That substitution is
what an RDS/TCP socket sending to a local address gets, and after it
the path init loop still runs for the TCP transport's RDS_MPATH_WORKERS
paths and allocates an ordered workqueue for each, while
rds_loop_conn_alloc() only ever provides transport data for path 0.
rds_conn_destroy() sizes its teardown from c_trans, by then the
loopback transport, so it visits path 0 only - and
rds_conn_path_destroy() would skip the other paths anyway, since it
returns before destroy_workqueue() for a path without transport data.
kfree(c_path) then drops the last pointers to seven workqueues. That
repeats for every such connection, on every netns teardown or module
unload, and every distinct local destination address is a separate
connection.
Recompute npaths once the transport is final, so that creation and
destruction agree on the set of paths. The c_path array stays sized
for the caller's transport; the unused entries are freed with it.
Fixes: 4716af3897 ("net/rds: Give each connection path its own workqueue")
Signed-off-by: Allison Henderson <achender@kernel.org>
Link: https://patch.msgid.link/20260921215027.174657-1-achender@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
nfp_net_ipsec_rx() drops the XArray lock before taking a reference to the
xfrm_state it found. The delete path can erase the entry and drop the last
state reference in that interval. RX can then try to increment a zero
refcount after the state has been queued for destruction.
The driver queues firmware invalidation asynchronously; the delete path
does not wait for the command to complete or drain pending RX processing.
The XFRM garbage collector waits for an RCU grace period before freeing
the state. That delays reclamation but does not make acquiring a reference
from zero valid.
Take the xfrm_state reference before releasing the XArray lock so
xa_erase() cannot run between lookup and reference acquisition.
Fixes: 57f273adbc ("nfp: add framework to support ipsec offloading")
Reported-by: Changyul Lee <lcy8047@gmail.com>
Signed-off-by: Sang-Hoon Choi <csh0052@gmail.com>
Link: https://patch.msgid.link/179001455912.44752.17153022439349797877.idr-bug-92@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Guangshuo Li says:
====================
net: ena: fix resource cleanup on probe failure
This series fixes two resource leaks in the ena_probe() error path after
ena_device_init() has successfully initialized device resources.
Patch 1 adds the missing PHC cleanup.
Patch 2 adds the missing MMIO read request cleanup.
====================
Link: https://patch.msgid.link/20260921154202.471662-1-lgs201920130244@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
ena_device_init() initializes the MMIO read mechanism with
ena_com_mmio_reg_read_request_init(), which allocates a coherent DMA
buffer for MMIO read responses.
The normal removal path releases this buffer through
ena_com_mmio_reg_read_request_destroy(). However, if ena_probe() fails
after ena_device_init() succeeds, the error path destroys the admin
resources and eventually frees ena_dev without destroying the MMIO read
request, leaving the coherent DMA buffer allocated.
Call ena_com_mmio_reg_read_request_destroy() in the probe error path
before releasing the remaining device resources.
This issue was found by manual code inspection.
Fixes: 1738cd3ed3 ("net: ena: Add a driver for Amazon Elastic Network Adapters (ENA)")
Cc: stable@vger.kernel.org
Signed-off-by: Guangshuo Li <lgs201920130244@gmail.com>
Link: https://patch.msgid.link/20260921154202.471662-3-lgs201920130244@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
ena_probe() initializes the PHC as part of ena_device_init(), but the
probe failure path does not destroy it before freeing the PHC private
data.
The normal removal path calls ena_phc_destroy() through
ena_destroy_device() before ena_phc_free(). However, if probe fails
after ena_device_init() succeeds, the error path reaches ena_phc_free()
without unregistering the PTP clock or destroying the device PHC
resources.
Call ena_phc_destroy() in the probe error path before freeing the PHC
private data.
This issue was found by manual code inspection.
Cc: stable@vger.kernel.org tags and describe this as a consistency cleanup
Fixes: e0ea34158e ("net: ena: Add PHC support in the ENA driver")
Cc: stable@vger.kernel.org
Signed-off-by: Guangshuo Li <lgs201920130244@gmail.com>
Cc: stable
Link: https://patch.msgid.link/20260921154202.471662-2-lgs201920130244@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Ilya Maximets says:
====================
ovs, net/sched: fixes for UAF after conntrack extension realloc
One clean up change plus two fixes for the UAF on helper extension
realloc x 2. First half for OVS and the second half for the similar
code in act_ct. This should cover all the known cases of this problem
in these two modules.
====================
Link: https://patch.msgid.link/20260921145655.3167436-1-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
While calling the helpers, a raw pointer to the extensions area is
wired into expectations list:
-> nf_ct_helper()
-> helper->help()
-> nf_ct_expect_related_report()
-> nf_ct_expect_insert()
-> hlist_add_head_rcu(&exp->lnode, &master_help->expectations)
In case the connection is not confirmed yet, more extensions can be
added afterwards with *_ext_add() calls reallocating the extension
space and leaving the now invalid pointer in the expectations list
that is later accessed while removing the expectation.
Make sure that helpers are called at the end after all the other
extensions are already added.
Note that the helper rejection now leaves the mark and labels set,
but that's not different from how the NAT was handled before or how
the mark and the labels were handled on confirmation failure. And
there are no atomicity guarantees provided by the API anyway.
Fixes: a21b06e731 ("net: sched: add helper support in act_ct")
Cc: stable@vger.kernel.org
Reported-by: Axel Mierczuk <axel.mierczuk@1password.com>
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Xin Long <lucien.xin@gmail.com>
Reviewed-by: Jamal Hadi Salim <jhs@mojatatu.com>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Link: https://patch.msgid.link/20260921145655.3167436-7-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
This variable can only become 'true' when the connection is not
confirmed, but it is only checked when it is confirmed. So, it can be
treated as being always false and just removed.
Fixes: a21b06e731 ("net: sched: add helper support in act_ct")
Cc: stable@vger.kernel.org
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Reviewed-by: Xin Long <lucien.xin@gmail.com>
Reviewed-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/20260921145655.3167436-6-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
In a case where skb with an unconfirmed ct entry gets cloned, we may
end up processing both again but with different sets of extensions.
The series of events:
1. The first clone wants to commit and runs the helpers wiring up
the extension pointer into the expectation list.
2. Then it looses the confirmation keeping the entry unconfirmed.
3. Second clone now wants to commit labels or run NAT and adds the
new extension for that breaking the pointer in the expectation
list causing UAF on the destruction path later.
While this is possible to trigger, there should be no practical
network pipeline where we need to process both clones without
modifications in the same zone. So, let's just reset the entry in
case for some reason we got an skb with a shared one. This doesn't
affect any known use cases, but avoids any potential problems with
sharing and modification of the unconfirmed ct entry.
Unlike openvswitch module, act_ct allows for NAT without commit.
Changing that would be a uAPI break. So, act_ct needs to reset on NAT
regardless of the commit flag to avoid reallocation of the extension
space. This, however, doesn't really change the picture for sensible
networking cases as there should be no need to run the same packet
twice (before and after the clone) through conntrack without packet
header or zone changes and without commit.
The fixes tag points to the introduction of helpers, since that's the
main UAF trigger for the sharing.
Fixes: a21b06e731 ("net: sched: add helper support in act_ct")
Cc: stable@vger.kernel.org
Reported-by: Axel Mierczuk <axel.mierczuk@1password.com>
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Reviewed-by: Xin Long <lucien.xin@gmail.com>
Reviewed-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/20260921145655.3167436-5-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
While calling the helpers, a raw pointer to the extensions area is
wired into expectations list:
-> nf_ct_helper()
-> helper->help()
-> nf_ct_expect_related_report()
-> nf_ct_expect_insert()
-> hlist_add_head_rcu(&exp->lnode, &master_help->expectations)
In case the connection is not confirmed yet, more extensions can be
added afterwards with *_ext_add() calls reallocating the extension
space and leaving the now invalid pointer in the expectations list
that is later accessed while removing the expectation.
Make sure that helpers are called at the end after all the other
extensions are already added.
Note that the helper rejection now leaves the mark and labels set,
but that's not different from how the NAT was handled before or how
the mark and the labels were handled on confirmation failure. And
there are no atomicity guarantees provided by the API anyway.
Fixes: cae3a26275 ("openvswitch: Allow attaching helpers to ct action")
Cc: stable@vger.kernel.org
Reported-by: Axel Mierczuk <axel.mierczuk@1password.com>
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Link: https://patch.msgid.link/20260921145655.3167436-4-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
This variable can only become 'true' when the connection is not
confirmed, but it is only checked when it is confirmed. So, it can be
treated as being always false and just removed.
Fixes: 3c1860543f ("openvswitch: add nf_ct_is_confirmed check before assigning the helper")
Cc: stable@vger.kernel.org
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Link: https://patch.msgid.link/20260921145655.3167436-3-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
In a case where skb with an unconfirmed ct entry gets cloned, we may
end up committing both but with different sets of extensions.
The series of events:
1. The first clone wants to commit and runs the helpers wiring up
the extension pointer into the expectation list.
2. Then it looses the confirmation keeping the entry unconfirmed.
3. Second clone now wants to commit labels and adds the new extension
for that breaking the pointer in the expectation list causing
UAF on the destruction path later.
While this is possible to trigger, there should be no practical
network pipeline where committing both clones without modifications
into the same zone is needed. So, let's just reset the entry in case
for some reason we got an skb with a shared one during commit. This
doesn't affect any known use cases, but avoids any potential problems
with sharing and modification of the unconfirmed ct entry.
The fixes tag points to the introduction of helpers, since that's the
main UAF trigger for the sharing.
Fixes: cae3a26275 ("openvswitch: Allow attaching helpers to ct action")
Cc: stable@vger.kernel.org
Reported-by: Axel Mierczuk <axel.mierczuk@1password.com>
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Link: https://patch.msgid.link/20260921145655.3167436-2-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
* add selftest coverage for peer VPN address validation
* reject multicast, broadcast and loopback peer VPN addresses, which
can never identify a peer
* reject MP peers left with no usable VPN address, as they can never
be selected for TX
* reject duplicate peer VPN addresses, which made peer lookup return
an arbitrary peer
* fix stale entry left in the VPN address hashtable when an address
is cleared
* fix torn IPv6 address read on lockless TX when the unusable local
source is cleared in place
* fix torn IPv6 address read on lockless TX when a new local endpoint
is learned in place
* fix dst cache being populated with a route resolved from an already
replaced bind
* fix stale route being reused after the socket mark or UDP source
port changed
* fix bogus validation of an unspecified local source address, which
must instead be left to route source autoselection
* fix IPv6 link-local peer endpoints losing their scope id when
configured via netlink, breaking route lookup
-----BEGIN PGP SIGNATURE-----
iJEEABYIADkWIQQr0db7q+Rc7Zog28Fc8QQzwdnOtwUCarECxxsUgAAAAAAEAA5t
YW51MiwyLjUrMS4xMiwyLDIACgkQXPEEM8HZzrcvTAEA/r3oFnZ4EOtiOyA2OpZ/
Iya+WAJ7LShj7tLymzujMe8BAIH6J2RyrWXnfDFvtSG8HGDEe4EBzVpP56tX0ZVn
xc0N
=fg0G
-----END PGP SIGNATURE-----
Merge tag 'ovpn-net-20260921' of https://github.com/OpenVPN/ovpn-net-next
Antonio Quartulli says:
====================
Included fixes:
* add selftest coverage for peer VPN address validation
* reject multicast, broadcast and loopback peer VPN addresses, which
can never identify a peer
* reject MP peers left with no usable VPN address, as they can never
be selected for TX
* reject duplicate peer VPN addresses, which made peer lookup return
an arbitrary peer
* fix stale entry left in the VPN address hashtable when an address
is cleared
* fix torn IPv6 address read on lockless TX when the unusable local
source is cleared in place
* fix torn IPv6 address read on lockless TX when a new local endpoint
is learned in place
* fix dst cache being populated with a route resolved from an already
replaced bind
* fix stale route being reused after the socket mark or UDP source
port changed
* fix bogus validation of an unspecified local source address, which
must instead be left to route source autoselection
* fix IPv6 link-local peer endpoints losing their scope id when
configured via netlink, breaking route lookup
* tag 'ovpn-net-20260921' of https://github.com/OpenVPN/ovpn-net-next:
selftests: ovpn: validate peer VPN addresses
ovpn: reject invalid peer VPN addresses
ovpn: reject multipeer peers without VPN addresses
ovpn: reject duplicate peer VPN addresses
ovpn: always unhash old VPN addresses before rehashing
ovpn: replace bind when clearing stale local source
ovpn: replace bind when learning local endpoint
ovpn: validate peer state before caching UDP dst
ovpn: track UDP socket route key for peer dst cache
ovpn: skip UDP source validation for unspecified addresses
ovpn: preserve IPv6 scope id for netlink peer endpoints
====================
Link: https://patch.msgid.link/20260921102215.3599702-1-antonio@openvpn.net
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
wx->ptp_tx_skb is shared between the Tx path, the PTP auxiliary
worker and the timestamp cleanup paths. The
WX_STATE_PTP_TX_IN_PROGRESS bit prevents multiple Tx paths from
submitting timestamp requests, but does not serialize the worker
against cleanup.
As a result, wx_ptp_clear_tx_timestamp() can free an skb after
wx_ptp_tx_hwtstamp_work() has obtained its pointer. The worker may
then pass the freed skb to skb_tstamp_tx() and release the same
reference again.
The cleanup path may also clear the in-progress bit while the worker
is still processing the old skb. This allows the Tx path to publish a
new skb which the worker can subsequently overwrite with NULL,
leaking its reference.
Add a dedicated spinlock to protect publication and consumption of
the Tx timestamp skb. Detach the skb and clear the in-progress bit
while holding the lock, then deliver the timestamp and release the skb
after dropping it. Use the same locked cleanup in the quiesce path,
but keep the detach there free of register accesses: quiesce runs
during PCIe error recovery, where MMIO is not reliable, and it
deliberately did not touch the device before. The lock is taken with
interrupts disabled, because netpoll can call ndo_start_xmit() with
hard interrupts already off.
When handling a Tx DMA mapping failure, keep the transmit path
reference until after comparing the skb under the lock. This prevents
skb address reuse from making the error path mistake a newer timestamp
request for the failed one.
Fixes: 06e75161b9 ("net: wangxun: Add support for PTP clock")
Reported-by: Sashiko <sashiko-bot@kernel.org>
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/6C7EC12D69217315%2B20260818074721.45536-1-jiawenwu%40trustnetic.com
Signed-off-by: Jiawen Wu <jiawenwu@trustnetic.com>
Link: https://patch.msgid.link/77431AF9A0E369F3+20260921071549.1141804-1-jiawenwu@trustnetic.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
phylink_bringup_phy() stores the PHY in pl->phydev before its last
fallible step: on a MAC whose phylink ops implement LPI,
phy_eee_rx_clock_stop() can fail with a real MDIO error. The callers
unwind with phy_detach(), which knows nothing about pl->phydev, so a
pointer to a PHY that is no longer attached outlives the failed
connect.
What that costs depends on how the caller got here.
phylink_connect_phy() goes through phylink_attach_phy(), which refuses
to attach while pl->phydev is set, turning a transient MDIO error into
a permanent -EBUSY. The SFP path is worse than that: sfp_sm_probe_phy()
answers the failure with phy_device_remove() and phy_device_free(), and
it assigns sfp->mod_phy only past that error return, so nothing clears
pl->phydev and it is left pointing at a freed phy_device that
phylink_resolve() and the ethtool helpers go on reading.
phylink_fwnode_phy_connect() has no such check, so a later connect
overwrites the stale pointer and hides the problem. A disconnect does
not: phylink_disconnect_phy() hands that pointer to phy_disconnect(),
and the second phy_detach() on the same PHY drops references the first
one already released.
Found while making a DSA port survive a PHY whose driver arrives after
the switch probes: keeping the port across a failed connect and
retrying is what makes this window reachable.
Publish the pointer after the last call that can fail instead of
unwinding it afterwards. Nothing between the two points reads
pl->phydev, and the registration that follows cannot fail:
phy_request_interrupt() falls back to polling on its own. The PHY-side
state keeps the order it had, so no MDIO operation moves relative to
another.
Fixes: 03abf2a7c6 ("net: phylink: add EEE management")
Signed-off-by: Aleksei Sviridkin <f@lex.la>
Link: https://patch.msgid.link/20260920222044.1752860-1-f@lex.la
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The MeiG Smart SRM821 5G module (0x2dee:0x4d53) crashes and drops off
the USB bus when it receives a Zero Length Packet (ZLP) after sending
or receiving an NTB of exactly 16384 bytes (tx_max).
According to the MBIM specification, devices do not require a ZLP
if the NTB size is exactly dwNtbOutMaxSize. However, the cdc_mbim
driver defaults to sending ZLPs for devices not explicitly whitelisted
to accommodate non-conformant hardware. This default behavior breaks
the strictly conformant MeiG SRM821 module.
Add this device to the ZLP conformance whitelist (cdc_mbim_info) so
the driver will pad the NTB to avoid sending ZLPs, preventing the
device firmware from crashing.
Cc: stable@vger.kernel.org
Signed-off-by: Ming Wang <wangming01@loongson.cn>
Link: https://patch.msgid.link/20260920074500.826121-1-wangming01@loongson.cn
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
When tcp_send_synack() replaces the cloned SYN skb at the head of the
retransmit queue with a copy, it frees the original with
tcp_rtx_queue_unlink_and_free() and only repairs tp->highest_sack.
tp->retransmit_skb_hint keeps pointing at the freed
skbuff_fclone_cache object.
The dangling hint is read in tcp_verify_retransmit_hint() and used as
the root of the rbtree walk in tcp_xmit_retransmit_queue(). An
unprivileged TFO client (sendmsg(MSG_FASTOPEN)) can arm the hint with
an attacker-supplied ICMP fragmentation-needed message, after which a
simultaneous open frees the armed SYN skb:
BUG: KASAN: slab-use-after-free in tcp_mark_skb_lost (net/ipv4/tcp_input.c:1316)
Read of size 4 at addr ffff88800604d928 by task swapper/1/0
Call Trace:
tcp_mark_skb_lost (net/ipv4/tcp_input.c:1316)
tcp_simple_retransmit (net/ipv4/tcp_input.c:3158)
tcp_v4_err (net/ipv4/tcp_ipv4.c:587)
Sync the hint to the copy.
Fixes: c31b70c996 ("tcp: Add logic to check for SYN w/ data in tcp_simple_retransmit")
Reported-by: Kimi Security Team <bug-report@moonshot.ai>
Tested-by: Weiming Shi <shiweiming@moonshot.ai>
Signed-off-by: Yilin Zhang <yilinzhang@moonshot.ai>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/8a9dff4063a2745653b7e88ceb745d75efa16e68.1790224474.git.yilinzhang@moonshot.ai
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
When a socket transmits a packet with MCTP_TAG_PREALLOC set,
mctp_lookup_prealloc_tag() iterates over the per-netns &mns->keys list
and matches netid, req_tag, peer_addr, and manual_alloc, without
checking whether tmp->sk == &msk->sk. This allows any MCTP socket in the
same network namespace to use and consume another socket's preallocated
tag.
Iterate the socket's own tag list (&msk->keys via sklist) instead of the
namespace-wide &mns->keys list in mctp_lookup_prealloc_tag(), ensuring
that only tags allocated by msk are matched.
Tested in QEMU against Linux 7.3.0-rc3 by allocating a manual tag
(0x18) on socket A via SIOCMCTPALLOCTAG for peer EID 9 and sending a
4-byte message with MCTP_TAG_PREALLOC from socket B in the same network
namespace. On the unfixed kernel, sendto(sock_b) using socket A's
preallocated tag succeeds (ret = 4); with this patch applied,
sendto(sock_b) fails with -ENOENT (errno = 2) while sendto(sock_a)
succeeds (ret = 4).
Fixes: 63ed1aab3d ("mctp: Add SIOCMCTP{ALLOC,DROP}TAG ioctls for tag control")
Suggested-by: Jeremy Kerr <jk@codeconstruct.com.au>
Cc: stable@vger.kernel.org
Signed-off-by: Hui Peng <benquike@gmail.com>
Link: https://patch.msgid.link/20260921051002.1656692-1-benquike@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Aleksei Sviridkin says:
====================
net: dsa: mt7530: fix two crashes on driver unbind
Unbinding the MT7530 driver from an MT7531 dereferences NULL in
regulator_disable(). On a Netcraze NC-1012 (MT7981B + MT7531, 6.18.44):
# echo mdio-bus:1f > /sys/bus/mdio_bus/drivers/mt7530-mdio/unbind
oopses there, and the build it was found on sets CONFIG_PANIC_ON_OOPS, so
the board goes down with it. Fix that and the same command gets as far as
mt7530_remove_common(), which disposes interrupt mappings the switch's own
regmap-irq chip still owns; the regmap-irq thread then faults in
handle_nested_irq() later in the same teardown. rmmod reaches both, since
mdio_module_driver() calls .remove on module exit.
Patch 1 is the regulator one. mt7530_probe() requests the core and io
supplies only for ID_MT7530 and mt7530_setup() enables them under the same
test, but mt7530_remove() disables them unconditionally, so on an MT7621 or
an MT7531 both pointers are still NULL from devm_kzalloc(). It reaches the
MDIO front end only.
Patch 2 is the interrupt one, and it reaches further.
mt7530_remove_common() disposes the per-PHY interrupt mappings by hand from
.remove, while the regmap-irq chip that owns the domain is devm-registered
and its parent interrupt is only freed once .remove has returned.
regmap_del_irq_chip() disposes the same mappings itself, in an order that
cannot race, so the driver's call adds nothing but a window. That helper is
called from both front ends, so the defect also covers the MMIO parts -
MT7988, EN7581, AN7583 and EN7528 - which have no regulators and never meet
the first defect at all.
The order is not arbitrary. On an MT7531 the regulator fault happens in the
first thing mt7530_remove() does with the switch, so execution never
reaches the interrupt defect. The second only became visible once the first
was fixed, which is also how both came to be found on one board.
Found and verified there. Without patch 1 the unbind panics in
regulator_disable(); with patch 1 alone the panic moves on to
handle_nested_irq(); with both, three unbind/bind cycles run, two back to
back and a third after a pause. In the two whose dmesg was captured, each
unbind removes the switch from the driver directory and takes lan1 to lan4
with it, each bind brings them back, and lan1 relinks at 1Gbps/full after
both binds, lan4 after the second. uptime rose from 58 to 202 seconds
across the three without resetting and pstore gained no new record. The
third cycle stayed unbound long enough to read the descriptors: no mt7530
line in /proc/interrupts and no irq/79, irq/80 or irq/81 directory, and the
next bind reuses those three numbers - regmap-irq freeing and disposing
what the driver no longer touches. The kernel under test was identified by
the sha256 of its ELF notes section, read from /sys/kernel/notes on the
running board and computed in advance from the flashed image.
What hardware could not answer here. There is no MT7530 or MT7621 part on
this bench, so the ID_MT7530 branch that patch 1 adds was checked by
reading the generated code rather than by running it, and no MMIO part was
available to exercise patch 2 on that front end either. One unrelated WARN
remains across the unbind, from sysfs_remove_link() under
dsa_user_destroy(); it is a separate DSA teardown-ordering defect and is
not addressed here.
v1: https://lore.kernel.org/20260914202421.2737079-1-f@lex.la
====================
Link: https://patch.msgid.link/20260918015020.2518315-1-f@lex.la
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
mt7530_remove_common() disposes the per-PHY interrupt mappings from
.remove, but the regmap-irq chip that owns the domain is devm-registered,
so its parent interrupt is only freed once .remove has returned. The
switch's own regmap-irq thread can therefore still dispatch on a mapping
that is already gone: irq_find_mapping() returns 0, irq_to_desc() returns
NULL and handle_nested_irq() locks desc->lock without checking it. The
attached PHYs have not given those interrupts back yet either, which the
kernel warns about a moment before the fault.
regmap_del_irq_chip() disposes the same mappings itself, after freeing the
parent interrupt and before removing the domain, so there is nothing left
for the driver to do here. Until it runs the descriptors stay alive, and a
late dispatch on one of them is harmless: dsa_unregister_switch() has freed
the PHY handlers by then, so handle_nested_irq() finds no action and
returns.
Fixes: 254f6b272e ("dsa: mt7530: Utilize REGMAP_IRQ for interrupt handling")
Signed-off-by: Aleksei Sviridkin <f@lex.la>
Link: https://patch.msgid.link/20260918015020.2518315-3-f@lex.la
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The core and io supplies are only requested for ID_MT7530: both the
devm_regulator_get() in probe and the regulator_enable() in
mt7530_setup() are guarded by the switch id, but mt7530_remove()
disables them unconditionally. On an MT7621 or an MT7531 both pointers
are still NULL from devm_kzalloc(), so rmmod or a sysfs unbind calls
regulator_disable() on NULL.
Fixes: ddda1ac116 ("net: dsa: mt7530: support the 7530 switch on the Mediatek MT7621 SoC")
Signed-off-by: Aleksei Sviridkin <f@lex.la>
Link: https://patch.msgid.link/20260918015020.2518315-2-f@lex.la
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
fixes the following crash on driver initialization:
|BUG: Kernel NULL pointer dereference on read at 0x00000158
|Faulting instruction address: 0xc0566b40
|Oops: Kernel access of bad area, sig: 11 [#1]
|BE PAGE_SIZE=4K PowerPC 44x Platform
|Modules linked in:
|CPU: 0 UID: 0 PID: 1 Comm: swapper/0 Tainted: GW 7.3.0-rc3+ #1
|Tainted: [W]=WARN
|Hardware name: MyBook Live APM821XX 0x12c41c83 PowerPC 44x Platform
|NIP: c0566b40 LR: c05648f8 CTR: c04c1e9c
|REGS: c1053a20 TRAP: 0300 Tainted: GW (7.3.0-rc3+)
|MSR: 0002b000 <CE,EE,FP,ME> CR: 24008808 XER: 00000000
|DEAR: 00000158 ESR: 00000000
|GPR00: c05648f8 c1053b10 c1063600 c1030000 c5ab3000 00000000 [...]
|GPR08: 00000002 00000000 00000000 c1053b40 84002808 00000000 [...]
|GPR16: cfffd210 00000002 c0beafcc cfffc960 00000000 c1030644 [...]
|GPR24: c0beafbc c1037000 00000000 0000000a 00000000 c1030000 [...]
|NIP [c0566b40] phy_link_topo_add_phy+0x2c/0x1d0
|LR [c05648f8] phy_attach_direct+0x1a4/0x368
|Call Trace:
|[c1053b10] [c0811e04] klist_put+0x54/0xb4 (unreliable)
|[c1053b40] [c05648f8] phy_attach_direct+0x1a4/0x368
|[c1053b70] [c0564ae8] phy_connect_direct+0x2c/0x60
|[c1053b90] [c056c7a8] of_phy_connect+0x50/0x74
|[c1053bc0] [c0572a88] emac_probe+0xd50/0x119c
|[c1053c90] [c04cb770] platform_probe+0x74/0xa4
|[c1053cb0] [c04c8f08] really_probe+0x120/0x2b0
|[c1053cd0] [c04c9254] __driver_probe_device+0x1bc/0x1fc
|[c1053d00] [c04c9334] driver_probe_device+0x38/0xa8
|[c1053d30] [c04c9560] __driver_attach+0xf4/0x10c
|[c1053d50] [c04c6af8] bus_for_each_dev+0x68/0xd0
|[c1053d90] [c04c7c34] bus_add_driver+0xcc/0x1ec
|[c1053dc0] [c04c9f6c] driver_register+0xcc/0x110
|[c1053de0] [c0aa3f40] emac_init+0x1c4/0x200
This bug showed up starting with v7.3-rc1. At the NIP in
phy_link_topo_add_phy() is a netdev_need_ops_lock() check.
This was added by the following
commit ded86da4bb ("net: ethtool: relax ethnl_req_get_phydev() locking assertion")
The bug shows up because at the time of_phy_connect() was called, the
netdev_ops were *not yet* determined. My fix is to move the code that
sets netdev_ops+commac.ops+ethtool_ops further up as emac_init_config()
derives that by looking at the device-tree and sets the required
dev->phy_mode accordingly.
During review, the Sashiko bot's AI stated that the commac assignment
became a dead store. Great catch! To keep the original behavior as-is,
one mentioned option "set dev->commac.ops = &emac_commac_ops only in
the non-gige case?" sounded like a great plan. So the dev->commac.ops
assignment for the non-gige-case moves into the else block.
This patch was tested on a WD MyBook Live (RGMII). the device now works again.
Fixes: ded86da4bb ("net: ethtool: relax ethnl_req_get_phydev() locking assertion")
Signed-off-by: Christian Lamparter <chunkeey@gmail.com>
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/49cd7343bc0e507c022071f3e2b5662b053dca73.1790007431.git.chunkeey@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
skbmark_decode(), skbprio_decode() and skbtcindex_decode() read fixed-size
values from the TLV payload without validating its length.
A malformed IFE frame can declare a shorter payload, causing the decoders
to consume bytes beyond the declared metadata value:
[TLV type=IFE_META_SKBMARK len=4]
-> dlen == 0, but decode reads 4 bytes
The decoder may therefore set skb metadata from unintended input.
Validate the payload length before decoding and return -EINVAL for
invalid lengths. Read the values with get_unaligned_be32() and
get_unaligned_be16(), as TLV payloads are not guaranteed to be
aligned. Teach tcf_ife_decode() to log a decoder error separately
from an unknown metaid; both are counted as overlimits and decoding
continues with the remaining metadata.
The metadata length issue was found by an automated audit of the IFE
decode path at v6.18-rc7 and reproduced with a userspace sanitizer
model of the decode path. Compile-tested on x86_64 with defconfig and
NET_ACT_IFE=y: act_ife.o and the three act_meta_*.o build
warning-free.
Fixes: 084e2f6566 ("Support to encoding decoding skb mark on IFE action")
Fixes: 200e10f469 ("Support to encoding decoding skb prio on IFE action")
Fixes: 408fbc22ef ("net sched ife action: Introduce skb tcindex metadata encap decap")
Assisted-by: Hawkeye:GLM-5.3-flash
Assisted-by: Qoder:Qwen3.8-Max
Signed-off-by: Fang Xieyan <fangxy@xiaopeng.com>
Link: https://patch.msgid.link/20260921125441.81459-1-fangxy@xiaopeng.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Creating a MACsec device with MAC offload over an LRO-capable lower
device triggers a warning in rtmsg_ifinfo_build_skb() when IPv4
forwarding is enabled by default.
register_netdevice() invokes inetdev_init(), which disables LRO and emits
a NETDEV_FEAT_CHANGE notification. This reaches macsec_fill_info() before
macsec_add_dev() initializes the SecY. key_len is still zero, so
macsec_fill_info() returns -EMSGSIZE and trips the WARN_ON in
rtmsg_ifinfo_build_skb(), even though the skb has enough space.
Even without the warning, notifications during registration can report
uninitialized SecY attributes, including the SCI. This ordering has existed
since the driver was introduced.
Initialize the SecY and apply the new-link attributes before registration.
Move MAC address inheritance into macsec_newlink() so the SCI can also be
initialized before registration-time notifications report it. Move the
per-CPU statistics and metadata destination allocation into ndo_init(),
and release partial allocations on failure.
Fixes: c09440f7dc ("macsec: introduce IEEE 802.1AE driver")
Reported-by: syzbot+f2f6312ad1b5a0bfe316@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=f2f6312ad1b5a0bfe316
Suggested-by: Sabrina Dubroca <sd@queasysnail.net>
Link: https://lists.openwall.net/linux-kernel/2026/08/19/552
Signed-off-by: Haseeb Malik <haseebulhaq55@gmail.com>
Reviewed-by: Sabrina Dubroca <sd@queasysnail.net>
Link: https://patch.msgid.link/20260921-fix-macsec-net-v3-1-accf94f93f5e@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
rds_ib_map_frmr() stores the caller's scatterlist in the MR before DMA
mapping and registration can fail. On failure, __rds_rdma_map() unpins
the pages and frees the scatterlist, but rds_ib_free_frmr() can still
return the MR to the pool with the stale pointer set.
This leaves the pool with a dangling scatterlist and can lead to local
privilege escalation. KASAN detects the resulting use-after-free when the
MR is later torn down:
BUG: KASAN: slab-use-after-free in __rds_ib_teardown_mr
Read of size 8
Call Trace:
__rds_ib_teardown_mr
rds_ib_unreg_frmr
rds_ib_flush_mr_pool
rds_ib_flush_mrs
rds_free_mr
rds_setsockopt
Store the scatterlist in the MR only after DMA mapping succeeds. If DMA
mapping fails, return directly while the MR fields remain clear; the caller
keeps ownership of the scatterlist and its pinned pages. If a later
registration step fails, unmap the scatterlist and clear the MR fields
before returning.
Fixes: 1659185fb4 ("RDS: IB: Support Fastreg MR (FRMR) memory registration mode")
Cc: stable@vger.kernel.org
Signed-off-by: Dongliang Qin <cccccccccccc777777@gmail.com>
Reviewed-by: Allison Henderson <achender@kernel.org>
Link: https://patch.msgid.link/20260922031546.3874605-1-cccccccccccc777777@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
bna: prevent IOC timer rearm during teardown
bnad_pci_remove() and the probe disable_ioceth path call
timer_delete_sync() for ioc_timer, sem_timer and hb_timer, but not for
iocpf_timer. bnad_iocpf_timeout() then takes bnad->bna_lock after
free_netdev() has freed the struct bnad.
Deleting iocpf_timer last does not fix this. sem_timer and
iocpf_timer rearm each other: bnad_iocpf_sem_timeout() can arm
iocpf_timer, and bnad_iocpf_timeout() arms sem_timer from
bfa_ioc_hw_sem_get() when the semaphore is busy.
timer_delete_sync() only waits out its own callback.
bnad_ioceth_disable() can time out and leave that callback live.
Shut all four IOC timers down with timer_shutdown_sync() on both
paths, so a later mod_timer() is ignored.
This issue was identified during our ongoing static-analysis research
while reviewing kernel code.
Fixes: 1d32f76962 ("bna: IOC failure auto recovery fix")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Co-developed-by: Ijae Kim <ae878000@gmail.com>
Signed-off-by: Ijae Kim <ae878000@gmail.com>
Signed-off-by: Myeonghun Pak <mhun512@gmail.com>
Link: https://patch.msgid.link/20260922014605.588040-1-mhun512@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
says:
====================
net: atl1c/atl1e/atl1: fix soft lockup on out-of-range tx consumer index
atl1c_clean_tx() reads a hardware-maintained tx consumer index and
walks a software index towards it:
while (next_to_clean != hw_next_to_clean) {
...
if (++next_to_clean == tpd_ring->count)
next_to_clean = 0;
}
next_to_clean only ever takes values in [0, tpd_ring->count). If the
hardware read returns a value outside that range - seen as 0xffff
while the PCIe link/MAC is resetting, e.g. during a neighboring
device's reboot - the loop condition can never become false, and the
NAPI thread spins forever. This produced a real soft lockup on
current hardware:
watchdog: BUG: soft lockup - CPU#12 stuck for 354s! [napi/eth%d-0]
RIP: 0010:atl1c_clean_tx+0x142/0x2d0 [atl1c]
Patch 1 fixes this in atl1c. atl1e and atl1 (atlx) share the exact
same loop shape, reading their own hardware-maintained consumer index
with no bounds check either, and are just as reachable from the same
kind of PCIe link event. Patches 2 and 3 apply the same guard to each.
====================
Link: https://patch.msgid.link/20260921091334.3571525-1-tamas@rimpianto.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Same issue as atl1c (see the first commit in this series, "net:
atl1c: fix soft lockup on out-of-range tpd_cons read"): the hardware
can report an out-of-range cmb_tpd_next_to_clean (seen as 0xffff)
while the PCIe link/MAC is resetting. An out-of-range value can
never be reached and the loop below would spin forever. Treat it as
"nothing new to clean" instead.
Fixes: f3cc28c797 ("Add Attansic L1 ethernet driver.")
Cc: stable@vger.kernel.org
Signed-off-by: Gajdos Tamás <tamas@rimpianto.com>
Link: https://patch.msgid.link/20260921091334.3571525-4-tamas@rimpianto.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Same issue as atl1c (see the first commit in this series, "net:
atl1c: fix soft lockup on out-of-range tpd_cons read"): the hardware
can report an out-of-range hw_next_to_clean (seen as 0xffff) while
the PCIe link/MAC is resetting. An out-of-range value can never be
reached and the loop below would spin forever. Treat it as "nothing
new to clean" instead.
Fixes: a6a5325239 ("atl1e: Atheros L1E Gigabit Ethernet driver")
Cc: stable@vger.kernel.org
Signed-off-by: Gajdos Tamás <tamas@rimpianto.com>
Link: https://patch.msgid.link/20260921091334.3571525-3-tamas@rimpianto.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
The hardware can report an out-of-range tpd_cons (seen as 0xffff)
while the PCIe link/MAC is resetting. An out-of-range value can
never be reached and the loop below would spin forever. To avoid
a soft lockup treat it as "nothing new to clean" instead.
Reproduced on two machines, same NIC (Qualcomm Atheros AR8151 v2.0,
4-port), triggered by rebooting a Mikrotik CCR2004 PCIe card that
the ports are directly linked to:
- Ubuntu 26.04.1 LTS, kernel 7.0.0-31-generic. The link-flap
precursor, before the lockup was captured with a full trace
elsewhere:
atl1c 0000:05:00.0 enp5s0f0: NETDEV WATCHDOG: CPU: 4: transmit queue 2 timed out 489984 ms
atl1c 0000:05:00.0: MAC state machine can't be idle since disabled for 10ms second
atl1c 0000:05:00.0: atl1c: enp5s0f0 NIC Link is Up<65535 Mbps Full Duplex>
65535 (0xffff) here is the same value tpd_cons reads back once the
loop below gets stuck.
- Proxmox VE, kernel 7.0.14-11-pve. Same NIC/trigger, this time
caught by the soft lockup watchdog with a full stack trace:
watchdog: BUG: soft lockup - CPU#12 stuck for 354s! [napi/eth%d-0:329]
CPU: 12 UID: 0 PID: 329 Comm: napi/eth%d-0 Tainted: P O L 7.0.14-11-pve #1 PREEMPT(lazy)
RIP: 0010:atl1c_clean_tx+0x142/0x2d0 [atl1c]
Call Trace:
<TASK>
__napi_poll+0x32/0x1e0
napi_threaded_poll_loop+0x286/0x2e0
napi_threaded_poll+0xfd/0x140
kthread+0xf7/0x130
ret_from_fork+0x2da/0x3a0
ret_from_fork_asm+0x1a/0x30
</TASK>
Fixes: 43250ddd75 ("atl1c: Atheros L1C Gigabit Ethernet driver")
Cc: stable@vger.kernel.org
Signed-off-by: Gajdos Tamás <tamas@rimpianto.com>
Link: https://patch.msgid.link/20260921091334.3571525-2-tamas@rimpianto.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Create a bonding device (i.e. bond0) in active-backup mode, 2 slaves.
Active slave: offload capable interface (i.e. eth1), primary interface.
Backup slave: non-offload capable interface(i.e. eth2).
Configure strongswan service swantl.conf child SA "hw_offload = crypto"
Start strongswan service
IPSec Crytpo Offload is enabled on top of bond0. i.e.
ip xfrm state |grep offload
crypto offload parameters: dev bond0 dir out mode crypto
crypto offload parameters: dev bond0 dir in mode crypto
Active slave eth1 takes adavantage of IPSec Crypto Offload capability.
If active slave eth1 is down for any reason (i.e. eth1 link down):
ip link set down dev eth1
non-offload capable interface eth2 failover to becomes active slave.
The existing SAs can continue use software IPsec after failover.
Traffic still keeps going properly.
However if eth1 link had not recovered yet, strongswan service does
new child SA rekey, or uses swanctl command to do new child SA rekey,
it will fail because active slave eth2 doesn't support crypto offload.
In bond_ipsec_add_sa routine, it returns -EINVAL now, which is
treated as fatal error by xfrm_dev_state_add routine in kernel xfrm.
To make the non-offload active slave survive the child SA rekey, need
to make bond_ipsec_add_sa routine returns -EOPNOTSUPP instead when
active slave doesn't support IPsec Crypto offload, the xfrm will
gracefully fallback to create new SA using Software IPsec.
Network traffic can keep going.
After offload capable interface eth1 link is up, becomes active slave,
next time strongswan child SA rekey will create a new SA which enables
crypto offload again.
Fixes: 18cb261afd ("bonding: support hardware encryption offload to slaves")
Signed-off-by: David Dai <zdai@linux.ibm.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Link: https://patch.msgid.link/20260918211155.1664493-1-zdai@linux.ibm.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
ic_dhcp_init_options() appends the hostname (option 12), vendor-class
(option 60) and client-ID (option 61) options into the fixed 312-byte
bootp_pkt.exten[] buffer. Only the client-ID branch checked the
remaining space; the hostname and vendor-class writes were unbounded.
A 64-byte hostname together with the maximum 252-byte dhcpclass=
identifier needs 18 + (2 + 64) + (2 + 252) = 338 of the 312 available
bytes even before the terminating END marker, so the vendor-class memcpy
runs past the end of exten[]. With CONFIG_FORTIFY_SOURCE this is
reported as a field-spanning write and, when the kernel is booted with
panic_on_warn=1, aborts boot with a panic.
Route the optional options through a common helper that makes sure the
option, its 2-byte header and the END marker all fit and drops an option
that would not. Configurations with short options keep sending exactly
the same bytes as before.
Fixes: 130c0f47fd ("ipconfig: send host-name in DHCP requests")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Assisted-by: LLM
Signed-off-by: Yuqi Xu <xuyuqiabc@gmail.com>
Reviewed-by: Ren Wei <weir@nebusec.ai>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/7808dfbfa2162dfd0b19f59aff5742d6e0db2abb.1789798023.git.xuyuqiabc@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Mostly security fixes.
fix use-after-free in nfc_get_local_general_bytes
llcp: Fix list corruption / refcount desync in nfc_llcp_recv_dm()
llcp: Fix race condition in accept_queue lifecycle
llcp: drop truncated I/RR/RNR PDUs in nfc_llcp_recv_hdlc()
llcp: fix -ENOMEM on connect with zero-length service name
llcp: fix WKS SAP hijacking via prefix match in nfc_llcp_wks_sap()
llcp: fix sdreq TLV list leak on parse/alloc/send failure
llcp: fix slab-out-of-bounds reads when logging service names
selftests/nci: Fix out-of-bounds store on thread join
selftests: nci: Correct pthread_create return value check
selftests: nci: Fix uninitialized family ID on missing attribute
virtual_ncidev: Add missing ioctl compat handler
nfcmrvl: validate helper command length before pull
pn533: fix OOB read in pn533_acr122_is_rx_frame_valid()
port100: reject frames whose declared length exceeds the received data
st21nfca: validate ISO15693 inventory length
st21nfca: validate received frame size
trf7970a: power down on startup RX gain failure
Signed-off-by: David Heidelberg <david@ixit.cz>
-----BEGIN PGP SIGNATURE-----
iQIzBAABCAAdFiEE13oJz+7cK71TpwR0YAI/xNNJIHIFAmq0M08ACgkQYAI/xNNJ
IHKXbw//cxnjUWiGiCWQKNGIrPQhO8xnxp29/adTfr5cpaCf9XS1Ah742JnsvV0J
jiKwrV2VdhIedF0gCU51pfMdgV4VWWup2KQJbl820m8EaRmaKY6Spao/HxznnXtF
H1JqYir2ISVADjX4hgQpqQIH2Vrz2o2QAkcsY+jjjydxxajmrrQh1SrJeEGom366
eDHAPmO9MTheJZ8gFXtuz4JdoipVWwY0Ta8D3mmVfgX/Jv5ESsPa4Ie4gQvQ1Vua
9bvYSgqrCGi4jtzEOxXNL1KfufjsbW9I64anwhxwjKJ+sL+/GJ0VfjN4Dgq9dnub
H4dUJs9vUCMLY+Ns6Wo+2pJQ49Ults2SdVjH7kvdJ2FiUn+FvoPDLrQY00P/lhaB
MPLrOCFaNVIzDKX98lxAzT/6c6xrNb5lTrkf17XHzcWa2H4eZj30pBaFfYrfiEH5
dlASR5aSy+GuXa8J/i7+jpQ+xl0GEg9tb0GJsvsWTDAftAl8/9CvcZ/TCbjKeb9L
oGeStPcl/i+MB9wGZLOga8lox0DX5PMNkIcp+RXIjVvPDYiCBh1a3lWHtz65Qx4m
9VlNcIpUzN14VUTkWm1i29DVHpW8V5c/eLmoDHzBUFAVSVr2M0CkFwhPYOByqIzt
lIcpLuTLXrAplzIFWMbi6gU8FSRAFX2I+vVr6rkgWIz2+QCPmYM=
=lO2R
-----END PGP SIGNATURE-----
Merge tag 'nfc-7.3-rc5' of https://codeberg.org/linux-nfc/linux
David Heidelberg says:
====================
NFC fixes for net 7.3-rc5
* tag 'nfc-7.3-rc5' of https://codeberg.org/linux-nfc/linux:
nfc: pn533: fix OOB read in pn533_acr122_is_rx_frame_valid()
nfc: llcp: fix slab-out-of-bounds reads when logging service names
nfc: llcp: fix WKS SAP hijacking via prefix match in nfc_llcp_wks_sap()
nfc: llcp: fix -ENOMEM on connect with zero-length service name
nfc: st21nfca: validate ISO15693 inventory length
nfc: fix use-after-free in nfc_get_local_general_bytes
nfc: trf7970a: power down on startup RX gain failure
nfc: llcp: fix sdreq TLV list leak on parse/alloc/send failure
nfc: llcp: drop truncated I/RR/RNR PDUs in nfc_llcp_recv_hdlc()
nfc: virtual_ncidev: Add missing ioctl compat handler
selftests/nci: Fix out-of-bounds store on thread join
selftests: nci: Fix uninitialized family ID on missing attribute
nfc: llcp: Fix race condition in accept_queue lifecycle
selftests: nci: Correct pthread_create return value check
nfc: port100: reject frames whose declared length exceeds the received data
nfc: llcp: Fix list corruption / refcount desync in nfc_llcp_recv_dm()
nfc: st21nfca: validate received frame size
nfc: nfcmrvl: validate helper command length before pull
====================
Link: https://patch.msgid.link/adeaccc1-cc04-4bb9-a28a-61a61d75ba14@ixit.cz
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Only the entries below dev->num_tc are valid in dev->tc_to_txq[], and
dev->prio_tc_map[] may only name classes below it. netdev_set_num_tc()
lowers dev->num_tc without touching either array.
netdev_txq_to_tc() walks all TC_MAX_QUEUE slots and
netdev_get_prio_tc_map() returns the entry as it stands, so a leftover
entry is handed out as a traffic class >= dev->num_tc. Taking that
class from netdev_txq_to_tc(), __netif_set_xps_queue() rejects only a
negative one and indexes an XPS map sized for dev->num_tc classes:
tci = j * num_tc + tc;
RCU_INIT_POINTER(new_dev_maps->attr_map[tci], map);
attr_map[] holds nr_ids * num_tc entries and j runs over the ids named
in the mask, so a class that is not below num_tc pushes tci past the end
of the map for the last ids and the store overruns it.
Any caller that lowers num_tc leaves such entries behind, and
mqprio_destroy() tears down with netdev_set_num_tc(dev, 0) rather than
netdev_reset_tc(). After mqprio with 8 classes then 1, tc_to_txq[1..7]
still describe txq 1..7. The splat is from an XPS write to txq 2 on a
veth with 8 rx queues: attr_map[] has 8 * 1 entries, tci = j + 2, and
j == 6 stores one past the end of the 88-byte map:
BUG: KASAN: slab-out-of-bounds in __netif_set_xps_queue (net/core/dev.c:2954)
Write of size 8 at addr ffff88813016bc58 by task xps_oob/634
__netif_set_xps_queue (net/core/dev.c:2954)
xps_rxqs_store (net/core/net-sysfs.c:1880)
netdev_queue_attr_store (net/core/net-sysfs.c:1390)
Allocated by task 634:
__kmalloc_noprof (mm/slub.c:5439)
__netif_set_xps_queue (net/core/dev.c:2937)
The buggy address is located 0 bytes to the right of
allocated 88-byte region [ffff88813016bc00, ffff88813016bc58)
Reject a class the map has no room for.
Fixes: 184c449f91 ("net: Add support for XPS with QoS via traffic classes")
Signed-off-by: Norbert Szetei <norbert@doyensec.com>
Link: https://patch.msgid.link/162DD16F-54C6-444A-9E09-0B8CB3D591F2@doyensec.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
vxlan_na_create() samples LL_RESERVED_SPACE() to size the reply skb and then
samples it again to reserve headroom. A concurrent vxlan_changelink() can
update needed_headroom between the two reads, creating a TOCTOU race. The
second value can exceed the allocation and make the Ethernet header write out
of bounds.
The race is reproducible on the unpatched kernel. It occurred when
vxlan_na_create() generated a neighbour reply while vxlan_changelink() changed
the link headroom. KASAN caught a four-byte write two bytes beyond a 704-byte
skbuff_small_head allocation.
Snapshot the headroom once and use that value for both allocation and
reservation.
Fixes: 4b29dba9c0 ("vxlan: fix nonfunctional neigh_reduce()")
Signed-off-by: Sanghyun Park <sanghyun.park.cnu@gmail.com>
Link: https://patch.msgid.link/20260918032842.502409-2-sanghyun.park.cnu@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
ipgre_netlink_parms() can enable collect_md on an existing GRE, GRETAP
or ERSPAN device. Unlike newlink, changelink does not enforce metadata
tunnel uniqueness. Converting a non-metadata device can therefore
replace the metadata receive entry for another device of the same type
in the same netns. Deleting either device then clears the shared entry,
breaking metadata receive lookup for the surviving device.
If parameter validation fails after collect_md is set, deleting the
modified device can also clear an entry it never owned.
Reject enabling metadata mode in both changelink callbacks before any
encapsulation or tunnel parameters are modified. Allow requests that
repeat the metadata attribute on an existing metadata device.
Fixes: 2e15ea390e ("ip_gre: Add support to collect tunnel metadata.")
Signed-off-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Link: https://patch.msgid.link/20260921031859.9283-1-xuanqiang.luo@linux.dev
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
airoha_npu_remove() calls cancel_work_sync() on each core's wdt_work,
but the watchdog IRQ that queues it is requested with devm_request_irq()
and is freed only after .remove() returns. airoha_npu_wdt_handler() can
therefore schedule_work() again once the cancel has returned. struct
airoha_npu, which contains the work, is devm_kzalloc()'d and is freed in
that same unwind, so the late work dereferences freed memory.
Register the work with devm_work_autocancel() before devm_request_irq()
and drop .remove(). Devres runs in reverse order, so the IRQ is freed
before cancel_work_sync(), including when probe fails. A cancel left in
.remove() cannot get that order. Initializing the work first also stops
a pending watchdog interrupt from queuing an uninitialized work item.
Probe currently calls INIT_WORK() only after devm_request_irq().
This issue was identified during our ongoing static-analysis research
while reviewing kernel code.
Fixes: 23290c7bc1 ("net: airoha: Introduce Airoha NPU support")
Cc: stable@vger.kernel.org # 6.15+
Co-developed-by: Ijae Kim <ae878000@gmail.com>
Signed-off-by: Ijae Kim <ae878000@gmail.com>
Signed-off-by: Myeonghun Pak <mhun512@gmail.com>
Acked-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Link: https://patch.msgid.link/20260922000914.542068-1-mhun512@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Florian Fainelli says:
====================
net: bcmgenet: Collection of bug fixes
This patch series contains a collection of bug fixes found during a
LLM-assisted coding session.
====================
Link: https://patch.msgid.link/20260921220021.281418-1-florian.fainelli@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
bcmgenet_get_coalesce() reads DMA_RING0_TIMEOUT to calculate
rx_coalesce_usecs without masking out bits outside DMA_TIMEOUT_MASK
(16 bits). If upper bits are non-zero or contain status/flags, the
computed value of rx_coalesce_usecs returned to userspace via ethtool
becomes corrupted.
Mask the register read with DMA_TIMEOUT_MASK before computing the
timeout in microseconds.
Fixes: 4a29645bfe ("net: bcmgenet: Implement RX coalescing control knobs")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Florian Fainelli <florian.fainelli@broadcom.com>
Link: https://patch.msgid.link/20260921220021.281418-6-florian.fainelli@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
bcmgenet_set_mac_addr() did not check whether the provided MAC address is a
valid Ethernet address before applying it. Userspace could configure an
invalid address (such as all zeroes or a multicast address) while the
interface is down.
Add a call to is_valid_ether_addr() and return -EADDRNOTAVAIL if the MAC
address is not valid.
Fixes: 1c1008c793 ("net: bcmgenet: add main driver file")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Florian Fainelli <florian.fainelli@broadcom.com>
Link: https://patch.msgid.link/20260921220021.281418-5-florian.fainelli@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
bcmgenet_power_up() had an early check for bcmgenet_has_ext(priv) before
dispatching by power mode. GENET V1 does not have the EXT block (unlike
GENET V2+), which causes bcmgenet_power_up() to immediately return 0.
As a consequence, when waking up from GENET_POWER_WOL_MAGIC on GENET V1,
bcmgenet_wol_power_up_cfg() is never invoked to disable the WoL clock,
clear wake event masks, and restore normal PHY and MAC operations.
Move the bcmgenet_has_ext() checks to the GENET_POWER_PASSIVE and
GENET_POWER_CABLE_SENSE cases where the EXT registers are actually
accessed, allowing GENET_POWER_WOL_MAGIC cleanup to execute on all
hardware versions.
Fixes: c3ae64ae0c ("net: bcmgenet: handle GENET_POWER_WOL_MAGIC")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Florian Fainelli <florian.fainelli@broadcom.com>
Link: https://patch.msgid.link/20260921220021.281418-4-florian.fainelli@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
bcmgenet_gstrings_stats statically defines ethtool statistics for queues
0 through GENET_MAX_MQ_CNT (4). However, bcmgenet_probe() only initialized
the u64_stats_sync seq counter up to priv->hw_params->rx_queues and
priv->hw_params->tx_queues.
Since priv->hw_params->rx_queues is 0 across all hardware versions (and
priv->hw_params->tx_queues is 0 on GENET V1), rings 1..4 have uninitialized
u64_stats_sync structures. When ethtool -S is run on 32-bit kernels,
bcmgenet_get_ethtool_stats() reads stats from rx_rings[1..4], causing
lockdep warnings due to the uninitialized sequence counters.
Initialize the sequence counters for all GENET_MAX_MQ_CNT + 1 queues.
Fixes: ffc2c8c4a7 ("net: bcmgenet: Initialize u64 stats seq counter")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Florian Fainelli <florian.fainelli@broadcom.com>
Link: https://patch.msgid.link/20260921220021.281418-3-florian.fainelli@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
When bcmgenet was converted to 64-bit statistics, STAT_RTNL members were
switched to point into struct rtnl_link_stats64, whose fields are 64-bit
(__u64) regardless of architecture.
However, bcmgenet_get_ethtool_stats() retained a legacy check:
if (sizeof(unsigned long) != sizeof(u32) &&
s->stat_sizeof == sizeof(unsigned long))
On 32-bit systems, sizeof(unsigned long) == sizeof(u32), causing this
condition to evaluate to false. As a result, 64-bit RTNL stats fields were
read via *(u32 *)p. On 32-bit Big-Endian systems (such as MIPS BE), this
reads the high 32 bits and returns 0 until the counter exceeds 4GB; on
32-bit Little-Endian systems (such as 32-bit ARM), the value is truncated
to 32 bits.
Fix this by checking if s->stat_sizeof == sizeof(u64) so 64-bit fields are
always read as 64-bit values.
Fixes: 59aa6e3072 ("net: bcmgenet: switch to use 64bit statistics")
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Signed-off-by: Florian Fainelli <florian.fainelli@broadcom.com>
Link: https://patch.msgid.link/20260921220021.281418-2-florian.fainelli@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The seg6, ioam6 and rpl lwtunnels size their skb_cow_head() request as
the length they are about to push plus dst_dev_overhead(), then push the
new headers and rebuild the mac header below them with
skb_mac_header_rebuild(). That rebuild needs skb->mac_len of headroom,
but dst_dev_overhead() leaves LL_RESERVED_SPACE() of the egress device,
16 bytes for plain Ethernet.
Where the mac header is longer than that, as it is on ingress through a
VLAN device with reorder_hdr off, the rebuild runs out of room:
skb_set_mac_header(skb, -skb->mac_len) computes a negative offset,
stores it unchecked in the u16 skb->mac_header, and the memmove that
follows writes skb->mac_len bytes about 64 KB past skb->head.
Forwarding plain ping6 traffic through such a device reproduces it on
all five seg6 encapsulation modes and on the rpl and ioam6 inline paths;
skb->mac_header comes back as 65534 on a 704-byte head.
Return the larger of the two. The helper already returns skb->mac_len
when it has no dst, so this only makes the other branch agree, and it
covers every caller rather than each call site in turn.
Fixes: 40475b6376 ("net: ipv6: seg6_iptunnel: mitigate 2-realloc issue")
Fixes: dce525185b ("net: ipv6: ioam6_iptunnel: mitigate 2-realloc issue")
Fixes: 985ec6f5e6 ("net: ipv6: rpl_iptunnel: mitigate 2-realloc issue")
Suggested-by: Andrea Mayer <andrea.mayer@uniroma2.it>
Signed-off-by: Yuya Kusakabe <yuya.kusakabe@gmail.com>
Reviewed-by: Justin Iurman <justin.iurman@gmail.com>
Reviewed-by: Gabriel Goller <g.goller@proxmox.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Andrea Mayer <andrea.mayer@uniroma2.it>
Link: https://patch.msgid.link/20260922-seg6-maclen-headroom-v3-1-7b2f982ef79d@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>