tcf_action_delete() drops the reference held by its lookup before calling
tcf_idr_delete_index() with the saved action index. An unlocked
classifier can remove that action and reserve the same IDR slot with
ERR_PTR(-EBUSY) in between.
tcf_idr_delete_index() only checks the lookup result for NULL. It
therefore treats the reservation as a tc_action and dereferences
tcfa_bindcnt. A hardware execution breakpoint was used to schedule the
interleaving without changing the kernel source. KASAN reported this
decoded trace:
BUG: KASAN: null-ptr-deref in tca_action_gd+0x5b9/0x1010
Read of size 4 at addr 0000000000000010 by task poc/150
Oops: general protection fault, probably for non-canonical address 0xdffffc0000000002
RIP: tca_action_gd+0x5c0/0x1010:
arch_atomic_read at arch/x86/include/asm/atomic.h:23
raw_atomic_read at include/linux/atomic/atomic-arch-fallback.h:457
atomic_read at include/linux/atomic/atomic-instrumented.h:33
tcf_idr_delete_index at net/sched/act_api.c:766
tcf_action_delete at net/sched/act_api.c:1859
tcf_del_notify at net/sched/act_api.c:2014
tca_action_gd at net/sched/act_api.c:2064
R13: 0000000000000010 R15: fffffffffffffff0
Kernel panic - not syncing: Fatal exception
R15 contains ERR_PTR(-EBUSY), and adding the tcfa_bindcnt offset produces
the address in R13. With the guard applied, the same reproducer returned
-ENOENT without a KASAN report or panic. Treat error pointers as absent
and return -ENOENT.
Fixes: 0190c1d452 ("net: sched: atomically check-allocate action")
Cc: stable@vger.kernel.org
Reported-by: Xiang Mei <xmei5@asu.edu>
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Link: https://patch.msgid.link/20260914065123.4109709-2-bestswngs@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Commit fa8fca8871 ("ipv4: validate IPV4_DEVCONF attributes properly")
added validation of IFLA_INET_CONF attributes, and in the process
changed the call of nla_for_each_nested() to nla_parse_nested(). A
side effect of this change is that the IFLA_INET_CONF option is now
tested for NLA_F_NESTED being set, and fails if it is not. Prior to the
commit there was no check of NLA_F_NESTED.
Change nla_parse_nested() to nla_parse(). This restores the previous
functionality of not checking NLA_F_NESTED, thereby allowing code that
(incorrectly) doesn't set NLA_F_NESTED to continue to work.
This issue was identified because keepalived started logging errors when
it was configuring macvlans that it created.
Fixes: fa8fca8871 ("ipv4: validate IPV4_DEVCONF attributes properly")
Signed-off-by: Quentin Armitage <quentin@armitage.org.uk>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260915213320.1527029-2-quentin@armitage.org.uk
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Commit 339ccec8d4 ("net/mlx5: Enable MACsec offload feature for VLAN
interface") added NETIF_F_HW_MACSEC unconditionally to vlan_features so
that VLAN devices could inherit MACsec offload support.
mlx5e_build_nic_netdev subsequently copies vlan_features into
hw_features and features. As a result, all mlx5e NIC netdevices
advertise MACsec hardware offload, even when the firmware does not
support it and the driver does not install macsec_ops.
Set the MACsec feature bits in mlx5e_macsec_build_netdev, after device
capabilities have been validated. This preserves MACsec-over-VLAN
support and the ethtool feature control on capable devices, without
advertising either on unsupported hardware.
Fixes: 339ccec8d4 ("net/mlx5: Enable MACsec offload feature for VLAN interface")
Cc: stable@vger.kernel.org
Reviewed-by: Tariq Toukan <tariqt@nvidia.com>
Signed-off-by: Ralf Lici <ralf@mandelbit.com>
Link: https://patch.msgid.link/20260917122724.654639-1-ralf@mandelbit.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
sctp_assoc_update_retran_path() can loop forever when every remaining
transport, including retran_path, is SCTP_UNCONFIRMED: the state check
runs before the wraparound test, so the loop cannot observe that it has
completed a full pass.
Fix this by considering a transport only when it is not UNCONFIRMED,
then checking whether the walk has returned to retran_path. This makes
the full-pass termination independent of the transport state while
preserving the existing fallback selection semantics.
Also restore the NULL guard around the retran_path assignment. In the
all-UNCONFIRMED case there is no eligible replacement transport, and
installing NULL would leave later retransmit-path users and the debug
print with a NULL path.
Fixes: 4c47af4d5e ("net: sctp: rework multihoming retransmission path selection to rfc4960")
Signed-off-by: Yiqi Sun <sunyiqixm@gmail.com>
Acked-by: Xin Long <lucien.xin@gmail.com>
Link: https://patch.msgid.link/20260915095017.942213-1-sunyiqixm@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Alexander Duyck says:
====================
eth: fbnic: a collection of fixes
This series collects a handful of independent fbnic fixes for issues on
released kernels, plus one core ethtool fix needed by the fbnic offline
self test.
The first patch keeps rtnl_lock held on the ethtool ioctl path for the self
test. Since the ioctl path became rtnl-optional for ops-locked drivers,
fbnic's offline self test (which brings the interface down and up via
netif_close()/netif_open()) runs holding only the instance lock, tripping a
lockdep splat / ASSERT_RTNL and reconfiguring the device without the lock
it requires. A similar issue was found with Broadcom drivers so we expanded
the scope for v2 to just have the rtnl lock held for all selftest calls.
The second addresses a comparison issue in that we were limiting the
maximum number of standalone Tx queues to one less than the maximum number
of Tx queues. To resolve this it was just a matter of replacing a "<" with
a "<=".
The third addresses an indexing issue with netdev queues on fbnic in which
the NAPI vector was assumed to be findable as the Rx index modulo the
number of NAPI vectors. However this is actually not the case for if Tx
only and Rx only queues are setup. To resolve this we make use of the
cached NAPI pointer in the netdev Rx queues themselves.
The fourth patch fixes a NULL pointer dereference on unbind after a failed
PCIe error recovery: fbnic_pm_suspend() frees the napi vectors via a direct
ndo_stop() while leaving netif_running() true, and when slot_reset ->
resume fails the data path is never re-allocated. To prevent the panic we
reset num_napi to 0 before we free the IRQs which prevents walking the
unallocated napi vectors when we unbind the interface later.
The last two patches address the FW mailbox. One sets AW_FLUSH_MODE
alongside AW_FLUSH when tearing down the Rx ring, so the write pipeline
actually drains the staged requests instead of hanging on the BME halt.
The other handles completions flagged with FW_ERR on both mailboxes, which
the driver previously ignored. This resulted in us parsing a stale Rx page,
and spinning the capabilities poll to a timeout on a healthy ring.
====================
Link: https://patch.msgid.link/178941996343.7700.9376081102002673062.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The firmware can complete a mailbox descriptor while also setting FW_ERR
to indicate it could not process the request, for example on a mailbox
DMA error. The completion carries no valid data.
The driver did not check FW_ERR. On the Rx mailbox it would sync and
parse the stale page as a normal message, and on the Tx mailbox it
silently freed the request. If the initial capabilities exchange in
fbnic_mbx_poll_tx_ready() hit FW_ERR -- on the Tx request or on the Rx
response descriptor -- no response was parsed and the poll spun until it
timed out even though the ring was healthy.
Check FW_ERR on both mailboxes. Count it per-mailbox in
fbnic_fw_mbx.resp_error, which is also shown in debugfs, warn (rate
limited, since the bit is firmware controlled), and drop the Rx page
instead of parsing it.
In fbnic_mbx_poll_tx_ready() re-issue the capabilities request when
either the Tx or the Rx resp_error counter advances, so a FW_ERR on the
request or on its response triggers a retry rather than a timeout. A
valid capabilities response is honored before the retry check, so a
response parsed in the same poll as an unrelated FW_ERR is not discarded.
The counters are mailbox-wide rather than keyed to the capabilities
request; that is sufficient here because the exchange runs during
bring-up before any other mailbox traffic, and any spurious retry is
bounded by the existing 10s timeout.
Fixes: da3cde0820 ("eth: fbnic: Add FW communication mechanism")
Signed-off-by: Alexander Duyck <alexanderduyck@fb.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942023343.7700.9423398932961964439.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
When tearing down the FW mailbox Rx ring, fbnic_mbx_reset_desc_ring()
writes AW_CFG with FLUSH set and everything else, BME included, cleared.
Clearing BME halts the device's writes to the host but leaves the staged
requests parked in the PUL write pipeline rather than draining them, so
on the write path FLUSH alone never terminates the outstanding requests
and the flush the firmware waits on never completes.
Add the FLUSH_MODE definition and set both bits so the staged writes
drain out of the pipeline on their own. BME stays cleared, so nothing
lands on the host; it is restored later in fbnic_mbx_init_desc_ring()
when the ring is rebuilt, once the outstanding writes are gone.
The read path is unaffected. AR_CFG has no equivalent mode bit and
AR_FLUSH terminates the outstanding reads by itself, so it is left as
is.
Both writes remain plain stores rather than read-modify-writes. That is
deliberate: the matching write in fbnic_mbx_init_desc_ring() restores
BME and the TLP attributes, and clears both flush bits as a side effect.
Fixes: 3b12f00ddd ("fbnic: Gate AXI read/write enabling on FW mailbox")
Signed-off-by: Alexander Duyck <alexanderduyck@fb.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942022583.7700.11050671998277309744.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
fbn->num_napi is the count of live napi vectors, each of which owns an
IRQ. The PM path had freed them without clearing the count.
fbnic_pm_suspend() tears the datapath down via ndo_stop() and frees the
IRQs, but leaves netif_running() true so resume knows to re-open. Resume
rebuilds the datapath in __fbnic_pm_resume() and fbnic_reset_queues() sets
num_napi and __fbnic_open() re-allocates the vectors.
When the datapath is torn down but never rebuilt, num_napi is left
pointing at freed vectors under 2 different scenarios:
- a PCIe error recovery that fails (fbnic_err_slot_reset() ->
__fbnic_pm_resume() returns an error -> PCI_ERS_RESULT_DISCONNECT), so
.resume never runs; or
- an __fbnic_open() that fails partway on resume and unwinds, freeing
the vectors after fbnic_reset_queues() has already set num_napi.
The netdev is then running with num_napi > 0 but napi[] freed, and the
eventual remove/unbind close re-enters fbnic_down() -> fbnic_dbg_down()
and dereferences the freed vectors:
BUG: kernel NULL pointer dereference, address: 0000000000000210
RIP: fbnic_dbg_down+0x28
Clear num_napi when the vectors are freed: in the suspend teardown (a
good resume re-establishes it before __fbnic_open()) and on the resume
open failure. A redundant ndo_stop() then walks an empty napi[]. The
normal ndo_stop() down/up cycle is untouched and keeps num_napi for the
next ndo_open().
Fixes: bc6107771b ("eth: fbnic: Allocate a netdevice and napi vectors with queues")
Signed-off-by: Alexander Duyck <alexanderduyck@fb.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942021809.7700.10804028989308077839.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The queue management ndos pick the napi vector for an Rx queue with:
nv = fbn->napi[idx % fbn->num_napi];
The issue is this is only correct in the cases where there are no
standalone Tx vectors. In those cases we were allocating the Tx vectors
first and then the Rx so the queues would be pointing to Tx NAPI vectors
instead of the Rx ones.
The mapping the ndos want is already recorded. fbnic_set_netif_napi()
publishes it with netif_queue_set_napi(), which stores the napi pointer
in netdev_rx_queue.napi, and fbnic_reset_netif_napi() clears it again.
Both run under the netdev instance lock that the queue management ndos
also hold, so the pointer can be read directly.
Use it and drop the divide. The pointer is NULL exactly while the
datapath is down, so fbnic_queue_mem_alloc() can reject that case rather
than reaching into freed state: netdev_rx_queue_restart() calls it
before it tests netif_running(), and fbnic_pm_suspend() leaves
netif_running() true across a PCIe recovery that never completes, so a
queue restart can arrive after fbnic_stop() has freed the rings and the
vectors. fbnic_stop() clears the association in
fbnic_reset_netif_queues() before fbnic_free_napi_vectors(), so the
NULL is always published first. fbnic_queue_start() and
fbnic_queue_stop() need no check of their own, as
netdev_rx_queue_reconfig() only reaches them once fbnic_queue_mem_alloc()
has succeeded under the same instance lock.
Fixes: da43127a8e ("eth: fbnic: support queue ops / zero-copy Rx")
Signed-off-by: Alexander Duyck <alexanderduyck@fb.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942021136.7700.4391219358260544104.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Standalone channels use one NAPI vector for each Tx and Rx queue.
fbnic's allocation path excludes FBNIC_MAX_TXQS from that layout. A
64-Tx/64-Rx configuration therefore records 128 vectors but allocates
only 64, leaving NULL entries that resource setup dereferences.
Include the maximum vector count in standalone allocation.
Fixes: bc6107771b ("eth: fbnic: Allocate a netdevice and napi vectors with queues")
Signed-off-by: Björn Töpel <bjorn@kernel.org>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942020457.7700.13129750616387075931.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
An offline self test that brings the interface down and back up with
netif_close() / netif_open() requires rtnl_lock for both. Since the
ethtool IOCTL path became rtnl-optional for ops-locked drivers, the
ETHTOOL_TEST ioctl runs holding only the netdev instance lock, so on an
ops-locked driver the self test now tears the device down without
rtnl_lock.
With lockdep this reproduces deterministically on every offline self
test on such a driver; note the sole lock held is the instance lock, not
rtnl:
WARNING: suspicious RCU usage
net/core/netpoll.c:207 suspicious rcu_dereference_protected() usage!
1 lock held by ethtool/107:
#0: (&dev->lock){+.+.}, at: dev_ethtool
Call Trace:
netpoll_poll_disable
__dev_close_many
netif_close_many
netif_close
fbnic_self_test
dev_ethtool_locked
dev_ethtool
dev_ioctl
sock_ioctl
__x64_sys_ioctl
Without lockdep the same condition trips ASSERT_RTNL() in
__dev_close_many() / __dev_open(); that check only samples the global
rtnl state, so it can be masked by a concurrent rtnl holder, but the
device is still being reconfigured without the lock it requires.
The ethtool self_test is a legacy ioctl-only command, so an ETHTOOL_TEST
case is only needed on the ioctl path. Add an opt-in bit for drivers whose
self test needs rtnl_lock and set it on the ops-locked drivers whose
offline self test tears the interface down and up:
- fbnic (ops-locked via queue_mgmt_ops): fbnic_self_test() offline path
uses netif_close() / netif_open().
- bnxt (ops-locked via queue_mgmt_ops): bnxt_self_test() offline path
goes through bnxt_close_nic() / bnxt_half_open_nic() /
bnxt_half_close_nic() / bnxt_open_nic(), which close and reopen the
device.
Fixes: f994752b11 ("net: ethtool: optionally skip rtnl_lock on IOCTL path")
Signed-off-by: Alexander Duyck <alexanderduyck@fb.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942019771.7700.338431553546884773.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
lanphy_write_reg_data() does not advance the data pointer while iterating
over the register table. As a result, it writes the first entry num times
and leaves the remaining errata registers unconfigured.
Single-entry tables are unaffected, but tables with multiple entries
leave every entry after the first unapplied.
Advance the data pointer after each successful write so every table entry
is applied in order.
Fixes: c8732e9339 ("net: phy: micrel: lan8842 errata")
Cc: stable@vger.kernel.org
Signed-off-by: Abhishek Ojha <abhishek.ojha@savoirfairelinux.com>
Reviewed-by: Andrew Lunn <andrew@lunn.ch>
Link: https://patch.msgid.link/20260916231928.1336305-1-abhishek.ojha@savoirfairelinux.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
ipv6_find_hdr() walks the extension header chain, skipping each header by
the length that header itself declares. ipv6_optlen() returns up to 2048,
and the skip is never checked against skb->len, so the offset stored in
*offset can point past the end of the packet.
openvswitch installs that offset as the transport header, and
update_ipv6_checksum() then reads and writes the transport checksum field
out of bounds:
BUG: KASAN: slab-use-after-free in inet_proto_csum_replace16+0x445/0x470
Read of size 2 at addr ffff88810b754b06 by task ovs_ipv6_oob/629
CPU: 4 UID: 1000 PID: 629 Comm: ovs_ipv6_oob Tainted: G N 7.3.0-rc3+ #348
Call Trace:
inet_proto_csum_replace16+0x445/0x470
set_ipv6_addr+0x3dd/0x460
do_execute_actions+0x6a3d/0x7c40
ovs_execute_actions+0xfd/0x480
ovs_packet_cmd_execute+0xc38/0xf20
genl_rcv_msg+0x59e/0x870
netlink_rcv_skb+0x18b/0x450
genl_rcv+0x2d/0x40
netlink_unicast+0x6bc/0xa20
The buggy address belongs to the object at ffff88810b754980
which belongs to the cache skbuff_small_head of size 704
The buggy address is located 390 bytes inside of
freed 704-byte region [ffff88810b754980, ffff88810b754c40)
Other callers use that offset too, so bound it here rather than in one
caller.
Reject a header whose declared length does not fit in the packet.
ipv6_find_hdr() already fails with -EBADMSG on a malformed chain, so this
adds no new failure mode.
Fixes: f8f626754e ("ipv6: Move ipv6_find_hdr() out of Netfilter code.")
Suggested-by: Ilya Maximets <i.maximets@ovn.org>
Suggested-by: Eric Dumazet <edumazet@google.com>
Cc: stable@vger.kernel.org
Signed-off-by: Norbert Szetei <norbert@doyensec.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Ilya Maximets <i.maximets@ovn.org>
Link: https://patch.msgid.link/8F80BA1A-DDFD-432D-9075-242A3435FEB5@doyensec.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
X4D is an X4 controller instance as an IP block in an SoC.
It has the same feature set as X4.
Signed-off-by: Andy Moreton <andy.moreton@amd.com>
Reviewed-by: Pieter Jansen van Vuuren <pieter.jansen-van-vuuren@amd.com>
Reviewed-by: Alejandro Lucero <alucerop@amd.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260916125641.12238-1-alucerop@amd.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
If register_netdev() fails for one of the MTK_MAX_DEVS devices in
mtk_probe(), the error path jumps to err_deinit_ppe, skipping
mtk_unreg_dev(). The previously registered net_devices are then freed by
mtk_free_dev() while still in NETREG_REGISTERED state, hitting the
BUG_ON(dev->reg_state != NETREG_UNREGISTERED).
Route the register_netdev() failure to err_unreg_netdev so the net_devices
registered so far are properly unregistered before being freed.
Fixes: 8a8a9e89f8 ("net: ethernet: mediatek: cleanup error path inside mtk_hw_init")
Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Link: https://patch.msgid.link/20260916-mtk_eth_soc-netdev-fix-v1-1-5dac50eb65b1@oss.qualcomm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The REMCSUM option carries an absolute checksum start and checksum field
offset. gue_remcsum() passes them to skb_remcsum_process(), whose
partial path stores offset - start in the u16 skb->csum_offset variable.
If offset is less than start, this underflows.
A forwarded packet can retain CHECKSUM_PARTIAL and reach a NETIF_F_HW_CSUM
driver which trusts the metadata, leading skb_copy_and_csum_dev() to write
two bytes about 64 KiB beyond the destination buffer.
Reject reversed tuples in validate_gue_flags(), after the existing length
validation, so all GUE parsers enforce the ordering in one place.
Fixes: fe881ef11c ("gue: Use checksum partial with remote checksum offload")
Signed-off-by: Jérémy Jean <Jeremy.Jean@oss.cyber.gouv.fr>
Link: https://patch.msgid.link/20260915124806.2852293-2-Jeremy.Jean@oss.cyber.gouv.fr
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
flow_offload_alloc() returns NULL when the conntrack entry is dying
(e.g. raced with a conntrack flush) or when the GFP_ATOMIC allocation
fails; both are expected under load and neither is a kernel bug. This
path runs from softirq on every committed packet, so with
panic_on_warn=1 an unprivileged user can panic the box just by racing
a conntrack flush against a `tc ... action ct commit` classifier.
Reproduced with a custom repro under QEMU: a small, fixed set of UDP
flows through `tc filter ... action ct commit` on lo, raced against
threads flooding bare ctnetlink CT_DELETE (flush) requests. Hits
WARNING: net/sched/act_ct.c:437 (tcf_ct_flow_table_add(), inlined
into tcf_ct_act() in this build) within ~15s on the unpatched kernel;
same setup is clean on the patched kernel. The fix itself is
behavior-preserving: both branches already did `goto err_alloc`
before and after, only the WARN is removed.
Fixes: 64ff70b80f ("net/sched: act_ct: Offload established connections to flow table")
Reported-by: syzbot+6cc37aba98dac721c415@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=6cc37aba98dac721c415
Signed-off-by: Nguyen Ngoc Thang <ngocthang2710.1999@gmail.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260915150816.36487-1-ngocthang2710.1999@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Commit 48a5fe3877 ("tipc: fix bc_ackers underflow on duplicate
GRP_ACK_MSG") rejected duplicate/stale ACKs in tipc_group_proto_rcv()
by returning early when less_eq(acked, m->bc_acked).
However, that check remains incomplete in two ways:
1. When grp->bc_ackers is zero (e.g. on a quiet group, when replicast
ACKs were not requested, or after all expected members have already
acknowledged), an unexpected GRP_ACK_MSG with acked > m->bc_acked
passes less_eq() and unconditionally decrements grp->bc_ackers.
Because bc_ackers is a u16, this wraps to 65535, causing
tipc_group_bc_cong() to permanently report congestion and blocking
all future group broadcasts on the socket.
2. During an active broadcast round (grp->bc_ackers > 0), the sender
transmits packet S and advances grp->bc_snd_nxt to S + 1. Receivers
increment their expected counter to S + 1 upon consuming packet S,
so the only valid ACK value for the current round is strictly
acked == grp->bc_snd_nxt.
However, tipc_group_update_bc_members() initializes each member's
m->bc_acked to prev = grp->bc_snd_nxt - 1 (S - 1 before increment).
This leaves a 2-sequence gap (S - 1 to S + 1) in sequence space.
An incoming ACK is therefore neither rejected as duplicate nor
prevented from decrementing grp->bc_ackers if an unexpected or stale
value (such as S) is received. A member sending acked = S followed
by acked = S + 1 could decrement grp->bc_ackers twice in the same
round, prematurely clearing bc_ackers or underflowing it.
Fix this by:
- Dropping GRP_ACK_MSG immediately if grp->bc_ackers is zero.
- Requiring acked == grp->bc_snd_nxt and rejecting duplicates where
m->bc_acked == acked. Because replicast broadcast rounds are strictly
sequential, only grp->bc_snd_nxt can be acknowledged, and each member
can acknowledge at most once per round.
Note that a related pre-existing issue in tipc_group_delete_member()
(where grp->bc_ackers decrementing to zero upon member departure does
not restore *grp->open or trigger a socket wakeup) will be addressed
in a separate patch.
Fixes: 48a5fe3877 ("tipc: fix bc_ackers underflow on duplicate GRP_ACK_MSG")
Fixes: 2f487712b8 ("tipc: guarantee that group broadcast doesn't bypass group unicast")
Reported-by: James Burton <jamesburton@meta.com>
Cc: stable@vger.kernel.org
Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260913044233.193927-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
fib_select_multipath() compares nexthop_nh->nh_saddr against the flow
source address with no lock held, while fib_info_update_nhc_saddr()
stores a new value from another CPU as soon as the preferred source
address of the egress device changes.
Commit 195374d893 ("ipv4: fib: annotate races around nh->nh_saddr_genid
and nh->nh_saddr") added WRITE_ONCE() on the store side and READ_ONCE()
in fib_result_prefsrc() after syzbot reported
BUG: KCSAN: data-race in fib_select_path / fib_select_path
but it only covered that reader. fib_select_multipath(), reached from
fib_select_path(), is a second lockless reader of nh->nh_saddr and was
left bare.
Moreover, nh_saddr is only meaningful when nh_saddr_genid matches
dev_addr_genid, as established by commit 436c3b66ec ("ipv4: Invalidate
nexthop cache nh_saddr more correctly."). fib_select_multipath()
skips that validation, so it can score a nexthop using a stale source
address and skew the ECMP selection.
Annotate both reads with READ_ONCE() and refresh the cached source
address via fib_info_update_nhc_saddr() when the genid does not match,
mirroring fib_result_prefsrc().
Fixes: 32607a332c ("ipv4: prefer multipath nexthop that matches source address")
Signed-off-by: Linkui Xiao <xiaolinkui@kylinos.cn>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260916125316.988044-1-xiaolinkui@126.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
txgbe_fdir_filter_restore() reprograms every filter from
txgbe->fdir_filter_list after a reset. It extracts the ring part of
filter->action with ethtool_get_flow_spec_ring() and maps it onto a
PF rx ring, silently dropping the VF part of the cookie that
txgbe_add_ethtool_fdir_entry() stores there (input->action =
fsp->ring_cookie).
For a rule directed at a VF, restore therefore reprograms the filter
to the PF queue with the same ring index: after any down/up or
txgbe_reinit_locked(), traffic matching the rule is steered to the
PF instead of the VF.
Handle VF rules the same way txgbe_add_ethtool_fdir_entry() does:
validate vf against wx->num_vfs and ring against
wx->num_rx_queues_per_pool, and map the ring onto the absolute
queue index ((vf - 1) * wx->num_rx_queues_per_pool) + ring.
Fixes: 7a91722e0d ("net: txgbe: Support the FDIR rules assigned to VFs")
Cc: stable@vger.kernel.org
Signed-off-by: Zhang Yunfei <zhangyunfei1@kylinos.cn>
Link: https://patch.msgid.link/20260911091123.798931-1-zhangyunfei1@kylinos.cn
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
__vsock_net_mode_string() returns success without updating new_mode when
the transfer length is zero. Its caller then reads the uninitialized enum
and may permanently store a stack-derived value in the write-once child
mode.
Return before calling __vsock_net_mode_string() when *lenp is zero so
that the helper is never invoked with nothing to parse and new_mode is
never read uninitialized. This also prevents an empty write from
locking the current mode.
Fixes: eafb64f40c ("vsock: add netns to vsock core")
Cc: stable@vger.kernel.org
Reviewed-by: Luigi Leonardi <leonardi@redhat.com>
Signed-off-by: Aldo Ariel Panzardo <qwe.aldo@gmail.com>
Reviewed-by: Stefano Garzarella <sgarzare@redhat.com>
Reviewed-by: Bobby Eshleman <bobbyeshleman@meta.com>
Link: https://patch.msgid.link/20260915173050.3176344-1-qwe.aldo@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Tariq Toukan says:
====================
net/mlx5: SD LAG and devcom stability fixes
This series by Shay fixes four bugs in the Socket Direct LAG and devcom
subsystems, all related to initialization/teardown ordering and
concurrent access to the LAG device.
====================
Link: https://patch.msgid.link/20260915113459.3934760-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
In a shared-FDB LAG the master device creates the bond IB device; the
other LAG members do not create their own, they populate a port inside
the master's IB device. mlx5_lag_reload_ib_reps_unlocked() reloaded the
members' IB reps in iteration order, with no guarantee the master is
reloaded first. When a non-master member is reloaded before the master,
it tries to populate its port in an IB device that has not been
recreated yet.
Hence, reload the master's IB reps first, then every other member.
Fixes: 2b204cdb12 ("net/mlx5: LAG, use xa_alloc to manage LAG device indices")
Signed-off-by: Shay Drory <shayd@nvidia.com>
Reviewed-by: Akiva Goldberger <agoldberger@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260915113459.3934760-4-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
mlx5_lag_shared_fdb_create() sets sd_fdb_active on every group member
before reloading the representors, so mlx5_lag_is_active() is already
true and the guard in mlx5_esw_offloads_rep_load() does not skip the
VF/SF reps. If the reload then fails, the error path clears
sd_fdb_active and destroys the shared FDB, leaving the reps loaded
while SD LAG is inactive - the state cited commit was written
to prevent.
Unload the reps in the error path as well.
Fixes: 68c2dd59a6 ("net/mlx5: E-Switch, Tie rep load/unload to SD LAG state")
Signed-off-by: Shay Drory <shayd@nvidia.com>
Reviewed-by: Akiva Goldberger <agoldberger@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260915113459.3934760-3-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
mlx5_devcom_comp_get_size() returns the component's kref count. That
kref is bumped in mlx5_devcom_register_component() under comp_list_lock,
before the comp_dev is linked onto comp_dev_list_head under comp->sem.
The event broadcast (mlx5_devcom_locked_send_event()) walks that list.
Hence, a caller can read the expected size, but send_event won't be sent
to all peers. In the SD group registration path, this lets a member
broadcast its role-election event over an incomplete list, electing a
primary that never completes the group, is never marked ready, and
leaves the group with a stale primary.
Track the number of linked comp_devs in a dedicated counter, maintained
under comp->sem together with the list add/remove, and return it from
mlx5_devcom_comp_get_size().
Fixes: 9bb1ac8073 ("net/mlx5: devcom, Add component size getter")
Signed-off-by: Shay Drory <shayd@nvidia.com>
Reviewed-by: Akiva Goldberger <agoldberger@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260915113459.3934760-2-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The rsrc_alloc debugfs reader writes rows directly to userspace without
respecting the caller's read count. It also uses the current row length as
the userspace stride, which can corrupt output when rows have different
widths.
Use seq_file to handle userspace buffer sizes, offsets, and partial reads,
and write output columns directly to the seq_file buffer.
Fixes: 23205e6d06 ("octeontx2-af: Dump current resource provisioning status")
Signed-off-by: Heyang Tan <thy15333007817@163.com>
Reviewed-by: Ratheesh Kannoth <rkannoth@marvell.com>
Link: https://patch.msgid.link/20260914020521.146-1-thy15333007817@163.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
35fc: create a manual table with handle 801:, then add an auto-allocated
table. Before the fix, the auto allocation reuses id 1 and hands out the
same handle 0x80100000, aliasing the manual table; the test requires the
manual 801: handle to keep exactly one entry in the dump.
a6e8: with a live u32 table keeping the tc_u_common alive, add and delete
a manual table with handle 901:, then re-add it. Unpatched, the delete
leaks the raw-keyed IDR entry and the re-add fails with -ENOSPC; the
test requires the re-add to succeed.
Reviewed-by: Victor Nogueira <victor@mojatatu.com>
Tested-by: hybris <hybris@mojatatu.ai>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/QDISC-LQFE.v1.20260911041746.2@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
A u32 hash table created with an explicit handle ('tc filter add ...
handle 801: u32 divisor N') keys its IDR entry on the raw handle, while
the destroy paths free it under handle2id(handle). The two key domains
disagree for handles in the 0x800..0xFFF htid range:
handle2id() folds them back into the auto-allocated id space (1..0x7FF).
A manual table therefore leaves its raw-keyed IDR entry unreachable on
delete (a permanent leak), and its delete can drop the idr entry of an
unrelated live auto table. A later auto allocation can then hand out a
handle that aliases the live manual table; u32_lookup_ht() first-match
routes lookups and TCA_U32_LINK for that htid to the wrong table.
Key the divisor-path alloc on handle2id(handle) so allocation and
removal share one key domain. A manual handle that maps onto an id
already in use is rejected with -ENOSPC, and auto allocation skips ids
held by live manual tables.
Conditions to recreate:
ip link add test0 type dummy
tc qdisc add dev test0 clsact
tc filter add dev test0 ingress protocol ip pref 1 \
handle 801: u32 divisor 16
tc filter add dev test0 ingress protocol ip pref 2 u32 divisor 16
tc -d filter show dev test0 ingress | grep 'fh 801:'
# unpatched: two live tables with handle 0x80100000 (the pref 2 root
# hnode is auto-allocated id 1); patched: the auto hnode takes id 2.
Also tested with a poc with a live u32 table on the block, add/delete a manual
table 'handle 901: u32 divisor 1' twice; unpatched, the re-add fails with
-ENOSPC because the raw key leaked on the first delete.
Fixes: 73af53d820 ("net: sched: cls_u32: Fix u32's systematic failure to free IDR entries for hnodes.")
Reported-by: Sashiko (gemini + nipa) <sashiko-bot@kernel.org>
Closes: https://sashiko.dev/#/patchset/20260822222049.114526-1-jhs@mojatatu.com
Reviewed-by: Victor Nogueira <victor@mojatatu.com>
Tested-by: hybris <hybris@mojatatu.ai>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/QDISC-LQFE.v1.20260911041746.1@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The payload page pool pointer contains an error pointer when its
allocation fails. The cleanup path passes that error pointer to
page_pool_destroy() instead of destroying the header page pool. This
can dereference the error pointer and leave the header page pool
allocated.
Destroy the header page pool instead.
Fixes: 8a11010fdd ("eth: fbnic: allocate unreadable page pool for the payloads")
Reported-by: Sashiko <netdev-bot+sashiko@kernel.org>
Link: https://lore.kernel.org/netdev/178915061000.219967.7726187707862333281@kernel.org/
Signed-off-by: Björn Töpel <bjorn@kernel.org>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260915104917.3978113-1-bjorn@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
My Meta email address is no longer active. Map it to my current address
so that git and get_maintainer.pl stop pointing at a dead address for my
contributions.
Signed-off-by: Wei Wang <weiwan@google.com>
Link: https://patch.msgid.link/20260915201412.2201757-1-weiwan@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Previous releases - regressions:
- netfilter: hold reference on ct until flow is released
- bridge:
- move switchdev call outside rcu
- vlan: fix bugs caused by switchdev deletion errors
- wifi:
- mac80211: reset state when starting AP fails
- cfg80211: don't free driver-owned scan requests
- tcp: don't call skb_clone_and_charge_r() for close()d listener in tcp_v6_do_rcv().
- mptcp: return sk_wait_data() errors from recvmsg()
- xfrm: serialize state GC with device state flush
- drop_monitor: synchronize tracepoint unregistration on error path
- bluetooth:
- eir: validate service data length before reading UUID
- hci_sync: serialize local codec list cleanup
- RFCOMM: avoid socket lock inversion in listener cleanup
- eth: lan743x: fix RX checksum use-after-free
- eth: mvpp2: prevent buffer overflow in page_pool allocation
Previous releases - always broken:
- core: lock the socket in sock_gettstamp()
- neighbour: enforce min/max to NDTPA_INTERVAL_PROBE_TIME_MS.
- sched: codel: bound the dropping loop per dequeue call
- wifi: mac80211: include TIM bitmap control for buffered S1G mcast traffic
- psp: avoid conflicts with skb->decrypted and sk_validate_xmit_skb()
- xfrm: fix stack OOB read in iptfs_skb_reset_frag_walk()
- bluetooth: hci_qca: do not write to the serial port after it is closed
- dsa: mxl862xx: disable the stats poll on teardown
- eth: stmmac: fix TSO header length truncation
- eth: ip_tunnel: initialize `options_len` before referencing options
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
-----BEGIN PGP SIGNATURE-----
iQJKBAABCgA0FiEEg1AjqC77wbdLX2LbKSR5jcyPE6QFAmqsEOsWHHBhb2xvLmFi
ZW5pQGdtYWlsLmNvbQAKCRApJHmNzI8TpKMnD/4vxx/YloxnyDssUB95WcTfiHau
XD+YDfTTf4Qy3DwKnMiesZ4w2547FXG+LwJZGpAGWpRSE99OawFHvx5gyn3Vf4IU
HEsOuZEYMwhSeeMJxGdCg8K6McNEQx+aAD7D8gLoisnJr/DCKBJtxFgKVEjaKUfP
aCG+FmbBqdS7hwBVbYREwmqwQSMnNzRWLd7/10/oYMcLfvgEsIkZisRErAW2BgZj
mzGw/IcG+QRldDPGDwLwLzfEG9o1a2JctSRrQ3uLHZ3VOdmpnSkmf25s2IaEFAn6
xfIhGPMgYIBDrakn/Ci4fAF0L98FtBq4Sa21HlvPMBst7rcce5x49ddSlIUWxRVZ
Fbvs/0IMN0cEpYGbVJxE8iF3yo+t8XvsMdGS1JXd/ycaL9lpF+9gCzNvChSxba3N
ik7BGAmlg76rqMuzeeMbWqMCmOcCBhQsb7iZXjNStJoiVY+UT2QfgvJ1TqF8Qrar
Eu/xFkaqk8i/7jrx4ujceg9XpRt1Y3Y3Pq0sgdzlzTm162AV/kg1fvwHc6ORIML7
71XpBSGvjyTHoh95ob/w3/5fec/ekzvWlCBL/cYIwBJ4R6n7Vk5e0Y1yZDffPSo3
xQR0r4YdLewkjQdtnGEih6FSyWsu6nYpVSuwHbof1bQSzvlIM6TAabJmtJckFTld
1bR2C9UVw5mQYMy8PA==
=Gw4S
-----END PGP SIGNATURE-----
Merge tag 'net-7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net
Pull networking fixes from Paolo Abeni:
"Including fixes from Netfilter, Bluetooth, IPSec and WiFi.
Previous releases - regressions:
- netfilter: hold reference on ct until flow is released
- bridge:
- move switchdev call outside rcu
- vlan: fix bugs caused by switchdev deletion errors
- wifi:
- mac80211: reset state when starting AP fails
- cfg80211: don't free driver-owned scan requests
- tcp: don't call skb_clone_and_charge_r() for close()d listener in
tcp_v6_do_rcv()
- mptcp: return sk_wait_data() errors from recvmsg()
- xfrm: serialize state GC with device state flush
- drop_monitor: synchronize tracepoint unregistration on error path
- bluetooth:
- eir: validate service data length before reading UUID
- hci_sync: serialize local codec list cleanup
- RFCOMM: avoid socket lock inversion in listener cleanup
- eth:
- lan743x: fix RX checksum use-after-free
- mvpp2: prevent buffer overflow in page_pool allocation
Previous releases - always broken:
- core: lock the socket in sock_gettstamp()
- neighbour: enforce min/max to NDTPA_INTERVAL_PROBE_TIME_MS.
- sched: codel: bound the dropping loop per dequeue call
- wifi: mac80211: include TIM bitmap control for buffered S1G mcast
traffic
- psp: avoid conflicts with skb->decrypted and sk_validate_xmit_skb()
- xfrm: fix stack OOB read in iptfs_skb_reset_frag_walk()
- bluetooth: hci_qca: do not write to the serial port after it is
closed
- dsa: mxl862xx: disable the stats poll on teardown
- eth:
- stmmac: fix TSO header length truncation
- ip_tunnel: initialize `options_len` before referencing options"
* tag 'net-7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (159 commits)
mptcp: fix bad accounting in __mptcp_subflow_push_pending()
mptcp: close race between scheduler and state change
mptcp: avoid unneeded actions on subflow reset
net: skbuff: do not leave stale header offsets after pskb_carve()
selftests: net: packetdrill: test exclusion of old ACK from TCP fast path
tcp: exclude old ACKs from tcp fast path
dpll: reject a reference sync pin which is not on the pin's dpll
net: mvpp2: prevent buffer overflow in page_pool allocation
net: macb: fix ordering around PTP timestamp read
selftests: drv-net: psp: test PSP and TCP ULP mutual exclusion
net: psp: avoid conflicts with skb->decrypted and sk_validate_xmit_skb()
net: stmmac: preserve real_num_tx_queues on mqprio setup failure
net: stmmac: propagate FPE preemption-class mapping errors
net: wwan: t7xx: validate the netif index in t7xx_ccmni_recv_skb()
net: wwan: mhi_wwan_mbim: check skb_copy_bits() return value
net: wwan: mhi_wwan_mbim: guard against a cyclic NDP chain
net: ethernet: cortina: Ack RX overrun interrupt correctly
net: lock the socket in sock_gettstamp()
eth: fbnic: ring the doorbell if a burst ends in a drop
net: netsec: fix device_node reference leak on phy_np
...
A collection of small fixes. Most of them are device-specific fixes
while there are a few core fixes. The continued flux, but not too
scaring yet. Some highlights below.
ALSA Core:
- Fix potential UAF after asynchronous card release
- Fix a race condition in PCM timer initialization order
USB-Audio:
- Hardening fixes for issues reported by fuzzer for 6fire, bcd2000,
and implicit FB packets
- Fix double list addition in implicit FB handling
- Quirks for AVerMedia GC553Pro and Behringer FCA1616
HD-Audio:
- Quirks / fixes for HP OmniBook 7, OMEN 15, and Victus 15 laptops
ASoC:
- Support for DAI link codec channel mask to avoid mismatches
- Fix HDMI-codec channel status change report
- Fixes for various codecs and platforms: Realtek rt712/rt721
(calibration, reset fixes), Cirrus Logic (empty EFI variable
validation, capture channel fixup), AMD ACP SoundWire (bounds
checks, refactorings), ADAU1977 (OF match table support, SPI
cleanups), ES8336 (Huawei Matebook B3-420 quirk), UX500 (macro fix)
-----BEGIN PGP SIGNATURE-----
iQJCBAABCAAsFiEEIXTw5fNLNI7mMiVaLtJE4w1nLE8FAmqrwIYOHHRpd2FpQHN1
c2UuZGUACgkQLtJE4w1nLE8WeQ//eN0ufLXXqy5U6X7kri8eM4vjRDZa5v1z8oD/
oHBdkgPmSAOJOddCVHuyPyx5BP4reBgKbukxUTBjuJvVRD4iOYfvXSW1LzVCVcly
f60rJl1/Ck3RcXfVabNV9eeGCFtAwR00U/3oH7z0eDhjwisX2F4ucAMmgxyJUKTh
pLdMPFsI/XW5CWeVbPVObGlOB8Bf/c79wy5NAnpixAkR1WaXUSLKoF5djO9qIQex
eTIknPmMYLLSfzFfO0TWY1PPRPz5qJHDr6Acer0VMTHyZF0yg2bR+gRInbM1at2H
c/lg8m899ZbobSCiHFEJdPH/W5x3iHqFi3hKVtKtQ5niot+gWTirQpjUCZ2HxNOp
5fYEyieSJ0W/t2NfNW0SP+DUOvKaIxY9VUuGCW7z/YgjW5DvSxipd2FOC4Av0s6F
l67vYazzPzf9d34NIM333FHeSZ4WMXVKKTfB34CQD93lVB1GladNA6lFFer5zuqG
Sc8q+YGUF65OZEnbslANIDvqPG0eMNlqBn/iyqXQO+P/C6oXWIlqFm4DUtM5a8wB
3Vf/e9L4mzXwKZfoW2xsJLgfpAO0faYiTAO6fr7wjBPrRl7VoenOlsCshN5GLTmM
b4bQ29LZrpk9e3e8BdGdZzT/kfQEZ31O5sawaj6f+a2JG09xCi9qxklxV5wFvZ5E
FTO+HgE=
=PnHM
-----END PGP SIGNATURE-----
Merge tag 'sound-7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/tiwai/sound
Pull sound fixes from Takashi Iwai:
"A collection of small fixes. Most of them are device-specific fixes
while there are a few core fixes. The continued flux, but not too
scaring yet. Some highlights below.
ALSA Core:
- Fix potential UAF after asynchronous card release
- Fix a race condition in PCM timer initialization order
USB-Audio:
- Hardening fixes for issues reported by fuzzer for 6fire, bcd2000,
and implicit FB packets
- Fix double list addition in implicit FB handling
- Quirks for AVerMedia GC553Pro and Behringer FCA1616
HD-Audio:
- Quirks / fixes for HP OmniBook 7, OMEN 15, and Victus 15 laptops
ASoC:
- Support for DAI link codec channel mask to avoid mismatches
- Fix HDMI-codec channel status change report
- Fixes for various codecs and platforms: Realtek rt712/rt721
(calibration, reset fixes), Cirrus Logic (empty EFI variable
validation, capture channel fixup), AMD ACP SoundWire (bounds
checks, refactorings), ADAU1977 (OF match table support, SPI
cleanups), ES8336 (Huawei Matebook B3-420 quirk), UX500 (macro
fix)"
* tag 'sound-7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/tiwai/sound: (33 commits)
ASoC: adau1977-i2c: add OF match table for I2C
ASoC: adau1977-spi: drop __maybe_unused and of_match_ptr()
ASoC: adau1977: make the Kconfig symbols user selectable
ASoC: amd: acp: fix card name length warning in SOF SoundWire machine driver
ASoC: amd: acp: fix ffs() operator precedence for SoundWire link ID
ASoC: amd: acp: refactor codec config count in SOF SoundWire machine driver
ASoC: amd: acp: bounds-check SoundWire link ID in machine drivers
ASoC: cs-amp-lib: Prevent NULL pointer if efi variable is zero length
ASoC: codecs: rt712-sdca-dmic: fix uninitialized stream_config->type
ASoC: hdmi-codec: Report a change when the channel status moves
ASoC: ux500: Parenthesize MSP_{RX,TX}_CLKPOL_BIT() arguments
ASoC: rt721: Reset codec to fix abnormal sound
ALSA: usb-audio: fix list_add double-add in push_back_to_ready_list
ALSA: hda: trace PCM open only after assigning a stream
ALSA: usb-audio: skip the broken mute control on AVerMedia GC553Pro
ALSA: hda/realtek: Enable mute LEDs on HP OmniBook 7 17-dc0xxx
ALSA: 6fire: fix OOB write from device-reported iso length
ALSA: usb-audio: Add capture quirk for Behringer FCA1616
ALSA: hda/realtek: Add mute LED quirk for HP OMEN 15-ax
ASoC: Intel: sof_es8336: Add a quirk for Huawei Matebook B3-420
...
- fix fwnode reference leak on failure in shared GPIO handling
- fix regression in OF_POPULATED logic after the unification of GPIO hog
handling between OF, ACPI and machine variants
- don't call free_irq() if no IRQ is installed in gpio-virtuser
-----BEGIN PGP SIGNATURE-----
iQIzBAABCgAdFiEEkeUTLeW1Rh17omX8BZ0uy/82hMMFAmqrs8YACgkQBZ0uy/82
hMPTsw/+Jm+8Z0tuCuryhBsDwMiqS4Gnr8TahCO3aB3UyMqApkiTS+hlqWGMRWbF
Np07I4uu3ca6ohutKN/RN6K5hrjlJz6FhU1kliJ9RrK1a3bjLH4rbYoXBVsYeIpC
mjLx9lyt6RmS4RaHPV75xPEmsAdxWbMyar6SQfiZT2t96Czsrph/VggX9kMbnXt3
SgKTM5SyHrKmw4DgnQmZ4OcWt8p2edW+5DO+jxRmPlWUvYE/q91yemaedw5wBEoz
ftrvbuIr+JRKKOSugjbswwBbJ0pVUkMm+hwkyAfWSPp83aF+sm2rUB46Gkr6tfQ7
oJVqVewAV6vqW/XoAnB8vr2KO3As5HFEx8xLYZpEf9RlOIDsu8R3HOooFkvO0flP
EsQSFccdX4WEHZoSc81iJl/TJjoM2gJtBZqqOvkHr2RZ7zdKPcnW81VHJGkSyZiW
o70PMzcQ+FAu7I/o5Q+cqsw1eG68wHhRYZG0q38mJCRznjlEop9yBqCAx3yi+1mF
cIhB37FlfRkFYGXByrqUFhi9V8gHWdDTChIQFJBMx1c99mBpT/eHD2hbGmac42LP
QI4sbywoN6L3m+UJpuJwS7KoO4NtJaVN444nqnS5C8bxBGRdXV3JQHelzwhSSED9
cq+MyMQm5hjS+i6SrADGGg+PytUF9w5Z7yRq2LFNUFwcS0TaHpk=
=NiRo
-----END PGP SIGNATURE-----
Merge tag 'gpio-fixes-for-v7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/brgl/linux
Pull gpio fixes from Bartosz Golaszewski:
- fix fwnode reference leak on failure in shared GPIO handling
- fix regression in OF_POPULATED logic after the unification of GPIO
hog handling between OF, ACPI and machine variants
- don't call free_irq() if no IRQ is installed in gpio-virtuser
* tag 'gpio-fixes-for-v7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/brgl/linux:
gpio: virtuser: skip free_irq when no IRQ is installed
gpiolib: of: don't mark hog nodes OF_POPULATED before a chip is found
gpiolib: Put fwnode reference on failure
Matthieu Baerts says:
====================
mptcp: misc fixes for v7.3-rc4
Here are two unrelated fixes:
- Patch 1: avoid unneeded actions on subflow reset. A fix for another
fix introduced in v6.12 and targeting a commit from v5.7.
- Patch 2: close a possible race when scheduling a closing path. A fix
for another fix introduced in v6.0 and targeting v5.10.
- Patch 3: fix bad accounting when __subflow_push_pending returns an
error. A fix for v6.6.
====================
Link: https://patch.msgid.link/20260917-net-mptcp-misc-fixes-7-3-rc4-v2-0-0cf5c72667c8@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
If __subflow_push_pending() errors out we should avoid updating the
copied byte counters, to avoid mismatch push call later on.
Fixes: 0fa1b3783a ("mptcp: use get_send wrapper")
Cc: stable@vger.kernel.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Reviewed-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Link: https://patch.msgid.link/20260917-net-mptcp-misc-fixes-7-3-rc4-v2-3-0cf5c72667c8@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The mptcp scheduler may race with subflow sockets state change: data
transmission on the selected socket may fail and a later release could
try to use mss_now reset to 0 for a divide operation.
Address the issue by explicitly checking for the critical scenario.
Fixes: c886d70286 ("mptcp: do not queue data on closed subflows")
Cc: stable@vger.kernel.org
Reported-by: Shardul Bankar <shardul.b@mpiricsoftware.com>
Reported-by: Xinyang Ge <xinyang@anthropic.com>
Closes: https://lore.kernel.org/20260525194828.1137119-1-shardul.b@mpiricsoftware.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Reviewed-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Link: https://patch.msgid.link/20260917-net-mptcp-misc-fixes-7-3-rc4-v2-2-0cf5c72667c8@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Once in a blue moon, the mptcp receive path can recursively call
mptcp_data_ready() via state change under unlucky error conditions, and
then try to hold the data lock again.
Break the recursion loop explicitly checking for the exceptional
condition.
Add a new flag instead of using an existing one like 'closing', to exit
early in subflow_state_change(), and explicitly flush the RX queue at
reset time.
This avoids unneeded processing to check for available data -- calling
get_mapping_status() and more on a dying subflow -- but also in error
reporting and worker scheduling.
Note that we must consume the currently peeked skb before invoking
mptcp_dss_corruption to avoid consuming it again after the eventual
reset has freed it.
Fixes: e32d262c89 ("mptcp: handle consistently DSS corruption")
Cc: stable@vger.kernel.org
Reported-by: Xinyang Ge <xinyang@anthropic.com>
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Reviewed-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Link: https://patch.msgid.link/20260917-net-mptcp-misc-fixes-7-3-rc4-v2-1-0cf5c72667c8@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
A few fixes for the DMA-mapping code:
- resolved regression in accessing encrypted memory by IOMMU-backed
devices (Aneesh Kumar K.V),
- improved failure handling and removed rare bug in swiotlb/highmem
(Donggeun Yoo).
-----BEGIN PGP SIGNATURE-----
iHUEABYKAB0WIQSrngzkoBtlA8uaaJ+Jp1EFxbsSRAUCaquwMwAKCRCJp1EFxbsS
RMT6AP0elpdaZXNY0KwUBTwU95H604J+donqriepHABIBhIDEQD9GWZqNf/m1gEI
tR5lHQ3+NGs0Q7Vd2ed1vSe82HQSsgU=
=QxUC
-----END PGP SIGNATURE-----
Merge tag 'dma-mapping-7.3-2026-09-17' of git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux
Pull dma-mapping fixes from Marek Szyprowski:
"A few fixes for the DMA-mapping code:
- resolved regression in accessing encrypted memory by IOMMU-backed
devices (Aneesh Kumar K.V)
- improved failure handling and removed rare bug in swiotlb/highmem
(Donggeun Yoo)"
* tag 'dma-mapping-7.3-2026-09-17' of git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux:
x86/mm: Don't force unencrypted DMA for IOMMU-backed devices
dma-mapping: don't trace the DMA address when the allocation fails
swiotlb: use the adjusted address for the highmem page lookup
dma-coherent: report a failed reserved memory assignment
pskb_carve_inside_header() and pskb_carve_inside_nonlinear() remove
the first bytes of a packet and reallocate skb->head.
All the headers that were present before the operation are gone,
but both functions call skb_headers_offset_update(skb, 0), which
is a no-op : skb->mac_header, skb->network_header,
skb->transport_header and skb->csum_start keep their old values and
now describe bytes which are no longer there.
Both helpers size the new head from the old skb_end_offset(), so the
stale offsets still land inside the new allocation. They point past
skb_tail_pointer() though, to bytes that were never initialized.
pskb_carve_inside_nonlinear() is the worst case, because it leaves a
zombie skb with an empty linear part (skb->data ==
skb_tail_pointer(skb), skb_headlen(skb) == 0), while
skb_mac_header_was_set() is still true and skb->mac_header is way
ahead of skb->data.
The only user of pskb_extract() is rds_tcp_data_recv(), and the
carved skb is queued on tinc->ti_skb_list. When the RDS incoming
message is released, rds_tcp_inc_free() calls skb_queue_purge(),
which frees the skbs with SKB_DROP_REASON_QUEUE_PURGE. This is
visible from drop_monitor, which then tries to pull back to the
(bogus) mac header :
skbuff: __skb_pull(len=234)
skb len=6968 data_len=6968 headroom=0 headlen=0 tailroom=0
end-tail=384 mac=(234,14) mac_len=14 net=(248,40) trans=288
shinfo(txflags=0 nr_frags=1 gso(size=1428 type=16 segs=5))
csum(0x100120 start=288 offset=16 ip_summed=3 complete_sw=0 valid=1 level=0)
hash(0x7b446c6c sw=0 l4=1) proto=0x86dd pkttype=0 iif=60
kernel BUG at ./include/linux/skbuff.h:2847!
Add skb_carve_reset_headers() to mark the mac and transport headers
as not set, reset the network header, clear skb->mac_len, and drop
a now meaningless CHECKSUM_PARTIAL (csum_start no longer describes
anything).
Invalidate the inner offsets as well. Unlike mac_header and
transport_header they have no "unset" sentinel, so a leftover
non-zero value still looks like a real header. Zero
skb->inner_mac_header, skb->inner_network_header,
skb->inner_transport_header, skb->inner_protocol and
skb->encapsulation, so that all the header state is invalidated in
one place.
v2: fixed an inaccurate changelog. The stale offsets stay inside the
new skb->head, which is never smaller than the old one, they
simply point past skb_tail_pointer() to bytes that are gone.
Thanks to Xuanqiang Luo for insisting on this.
Also invalidate the inner header state, as suggested by the
netdev AI review :
https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260911114922.621937-1-edumazet%40google.com
Fixes: 6fa01ccd88 ("skbuff: Add pskb_extract() helper function")
Reported-by: syzbot+586af68eb819833c2d91@syzkaller.appspotmail.com
Closes: https://lore.kernel.org/netdev/6aa3e9d3.f2639fcc.29487d.0028.GAE@google.com/
Cc: Xuanqiang Luo <xuanqiang.luo@linux.dev>
Cc: Allison Henderson <achender@kernel.org>
Cc: rds-devel@oss.oracle.com
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Link: https://patch.msgid.link/20260915130423.3956471-1-edumazet@google.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Add a packetdrill test for an in-sequence data segment carrying an
excessively old ACK.
Verify that the segment falls through from the TCP fast path to the slow
path, where the existing ACK validation rejects it and sends a challenge
ACK. The payload is not accepted and RCV.NXT remains unchanged.
Based on the reproducer from Commit 3d501dd326
("tcp: do not accept ACK of bytes we never sent").
Signed-off-by: Inbal Schussheim <inbal.lipshtat@mail.huji.ac.il>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260914090408.1435080-3-inbal.lipshtat@mail.huji.ac.il
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Exclude old ACKs before SND.UNA from the tcp fast path
as well as ACKs after SND.NXT.
Such ACKs will fall through to the slow path, where tcp_ack()
performs the appropriate validation and challenge ACK handling
according to RFC5961 and Commit 3d501dd326 ("tcp: do not
accept ACK of bytes we never sent").
This prevents old ACKs from being accepted
or modifying connection state as part of the fast path before
appropriate ACK validation is applied.
In particular, this prevents payload carried by a segment with
an excessively old ACK from advancing RCV.NXT before the ACK
is rejected.
Fixes: 31770e34e4 ("tcp: Revert "tcp: remove header prediction"")
Reported-by: Amit Klein <amit.klein@mail.huji.ac.il>
Reported-by: Tamir Shahar <tamir.shahar1@mail.huji.ac.il>
Reported-by: Inbal Schussheim <inbal.lipshtat@mail.huji.ac.il>
Suggested-by: Eric Dumazet <edumazet@google.com>
Cc: stable@vger.kernel.org
Signed-off-by: Inbal Schussheim <inbal.lipshtat@mail.huji.ac.il>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260914090408.1435080-2-inbal.lipshtat@mail.huji.ac.il
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
dpll_pin_ref_sync_state_set() resolves the partner's driver private data
with dpll_pin_on_dpll_priv() and passes the result to ref_sync_get() and
ref_sync_set() without looking at it. The helper returns NULL when the
partner holds no ref on that dpll. Of the two drivers implementing the
feature only zl3073x dereferences the pointer (sync_pin->id); ice ignores
it, so ice cannot fault here.
The NULL is a teardown race, not a steady state - zl3073x registers every
input pin with every channel, so the partner is normally present on the
dpll the base pin resolves to. zl3073x_dev_stop() unregisters pins one at
a time, taking and dropping dpll_lock for each, and between the partner's
turn and the base pin's the partner is out of that dpll's pin_refs while
still registered with the channels not yet torn down, so
dpll_pin_available() keeps passing. That path is not only driver removal:
devlink reload and devlink dev flash both run zl3073x_dev_stop().
Reproduced by holding that state open with a mock dpll device, which is
where the frame name comes from:
BUG: kernel NULL pointer dereference, address: 0000000000000000
Oops: Oops: 0000 [#1] SMP NOPTI
RIP: 0010:mock_ref_sync_get+0x5/0x30
Call Trace:
<TASK>
dpll_pin_ref_sync_set+0x19f/0x4a0
dpll_nl_pin_set_doit+0x17d/0x840
genl_family_rcv_msg_doit+0xd6/0x130
genl_rcv_msg+0x181/0x2b0
netlink_rcv_skb+0x55/0x100
genl_rcv+0x23/0x30
netlink_unicast+0x24d/0x370
netlink_sendmsg+0x1e2/0x420
__sys_sendto+0x1db/0x1f0
__x64_sys_sendto+0x1f/0x30
do_syscall_64+0xe1/0x490
Commit d2e914a4a0 ("dpll: fix NULL pointer dereference in
dpll_msg_add_pin_ref_sync()") added the same guard to the read side, which
the kernel walks into by itself because the delete notification is emitted
from inside the unregister; the write side needs a pin-set to land in the
window and was left alone. Test the priv rather than look up pin_refs
directly, so that the two halves key off the same condition.
Fixes: 58256a26bf ("dpll: add reference sync get/set")
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Reviewed-by: Ivan Vecera <ivecera@redhat.com>
Link: https://patch.msgid.link/20260915213047.1352286-1-kuba@kernel.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
The per‑processor buffering scheme is supported only if the
number of pools (nrxqs * 2) does not exceed MVPP2_BM_MAX_POOLS (8).
This is already checked in mvpp2_probe() during the initial
activation of percpu_pools.
However, mvpp2_change_mtu() may later call
mvpp2_bm_switch_buffers(priv, true) without this check, which can
lead to an out-of-bounds access in the priv->page_pool array in
mvpp2_bm_init(). The array is sized to hold MVPP2_PORT_MAX_RXQ
entries, and mvpp2_get_nrxqs() may return exactly that value. The
per-CPU scheme then doubles it to nrxqs * 2, exceeding the array
bounds.
Check that the hardware version is MVPP22 or newer and that the
number of pools (nrxqs * 2) does not exceed MVPP2_BM_MAX_POOLS
before switching to per-CPU mode.
Found by Linux Verification Center (linuxtesting.org) with SVACE.
Fixes: 7d04b0b13b ("mvpp2: percpu buffers")
Signed-off-by: Dmitriy Okunev <dokunevdmitriy@gmail.com>
Link: https://patch.msgid.link/20260914091557.71769-1-dokunevdmitriy@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
PTP_SYS_OFFSET_EXTENDED returns system timestamps that do not correctly
bracket the PHC register read on MACB/GEM. On a Raspberry Pi 5, the
returned interval can be as short as 37 ns, while an ordered register
read takes approximately 1 us. This biases the midpoint used by phc2sys,
causing CLOCK_REALTIME to run approximately 0.5 us ahead when synchronized
to the PHC.
gem_tsu_get_time() reads the nanoseconds register using the driver's
relaxed MMIO accessor. On weakly ordered systems, the subsequent system
timestamp can be taken before the register read completes. The internal
smp_rmb() in the pre-timestamp path also does not guarantee ordering
against the subsequent MMIO read.
Add rmb() before and after the bracketed nanoseconds read in both the
normal and seconds rollover paths so the system timestamps bracket the
PHC read. Adding the post-read barrier increases the minimum interval on
the same Raspberry Pi 5 to approximately 1 us.
Fixes: e51bb5c278 ("net: macb: ptp: Switch to gettimex64() interface")
Tested-by: Nicolai Buchwitz <nb@tipi-net.de> # Raspberry Pi CM5, min bracket 37 ns -> 981 ns
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Reviewed-by: Théo Lebrun <theo.lebrun@bootlin.com>
Assisted-by: LLM
Signed-off-by: James Clark <jjc@jclark.com>
Link: https://patch.msgid.link/20260915045823.76100-1-jjc@jclark.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
A relatively large pile of fixes here, a lot of driver specific stuff
that's broadly unremarkable plus a few core fixes from Richard that fix
issues where SoundWire systems with multiple CODECs on the same link
would configure the CODECs to use the same bus slots leading to broken
audio.
-----BEGIN PGP SIGNATURE-----
iQEzBAABCgAdFiEEreZoqmdXGLWf4p/qJNaLcl1Uh9AFAmqrCNoACgkQJNaLcl1U
h9DHWgf8D3dIL06bqj6IoyMLCFNrcQ8BYbWUeWNu5YE0vP29ybdpYidTxJFjqF2t
TUB8fTO2u3LvfKIIgOSVyXN84i7/4EwtDjBz1iVzGhm0/2ZfEOitO2LtUvhCHiZi
+JnEOXdwa7wM9jv0On6B81r8+vXj7FaNmq/TnLbUU3R/DeaRx571k913lazZSRb0
cfPj1FGMUvpfBZ7DC011yEufDD4C8qVaktV6IpRqeBAxks1vmXQX7Lt78gJCxvQt
w8Uu2T8FpiTuYvObI7KW7IZr01IQtPpJ4N9ekUMuBfOqkjvG+XOETM2aEWg7q9Jk
Qu5GfyB2LIav4/On4pCxwJGaJiUavQ==
=IMNO
-----END PGP SIGNATURE-----
Merge tag 'asoc-fix-v7.3-rc3' of https://git.kernel.org/pub/scm/linux/kernel/git/broonie/sound into for-linus
ASoC: Fixes for v7.3
A relatively large pile of fixes here, a lot of driver specific stuff
that's broadly unremarkable plus a few core fixes from Richard that fix
issues where SoundWire systems with multiple CODECs on the same link
would configure the CODECs to use the same bus slots leading to broken
audio.
Daniel Zahka says:
====================
net: psp: avoid conflicts with skb->decrypted and sk_validate_xmit_skb()
Sashiko's review of commit da630d1da2b1 ("netdevsim: psp: drop tx key
ops") [1] showed that there is a hazard between PSP and offloaded TLS,
where both can clobber what the other set in the sk_validate_xmit_skb
callback.
It was discussed further on the mailing list [2], and it was pointed out
that there are conflicts with PSP and TLS ULP both using the
skb->decrypted bit.
The simplest fix is to make psp and tls mutually exclusive. This series
goes a bit further and makes psp exclusive with all TCP ULPs. The PSP
implementation that we have is not designed to be used with any TCP ULP,
so don't allow a socket to have state for both.
I will send a subsequent series to net-next which will remove the
ability to perform the rx-assoc and tx-assoc psp netlink calls on
sockets that are not in the TCP_ESTABLISHED state. This will close the
remaining quirk that a sk_clone() on a listen socket with psp tx-assoc
state will leave a stale sk->sk_validate_xmit_skb call back on a new,
non-psp socket. I do not believe that change needs to be regarded as a
fix, because it only stands to add unecessary validation code in the tx
path.
[1]: https://sashiko.dev/#/patchset/20260903-psp-prep-v1-0-d47e9c4c375d%40gmail.com
[2]: https://lore.kernel.org/netdev/20260903-psp-prep-v1-0-d47e9c4c375d@gmail.com/
====================
Link: https://patch.msgid.link/20260915-psp-ktls-fix-v2-0-0eedc3b148ec@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>