Commit Graph

1482584 Commits

Author SHA1 Message Date
Daniel Zahka
a41f24c612 net: psp: avoid conflicts with skb->decrypted and sk_validate_xmit_skb()
PSP conflicts with TLS ULP in its usage of both skb->decrypted and
sk->sk_validate_xmit_skb().

Make PSP mutually exclusive with TLS ULP, the only other user of either
of these. As other users of skb->decrypted come along, they can be added
to sk_has_decrypt_user(). It would make sense to also assert that
sk->sk_validate_xmit_skb() is also NULL in both of these setup paths for
similar future proofing, but the PSP listener/sk_clone() path is still
broken and it could be seen as a regression to not allow rx assoc to run
on a child of a listener socket with PSP tx assoc state.

Include all TCP ULPs in the sk_has_decrypt_user() check, even though TLS
is the only one that conflicts with PSP via the decrypted bit. This is
intentional because PSP was not designed to be used with ULPs. It is
best to close off surface area that may make bugs reachable, until
someone wishes to design and test an actual user of PSP with ULPs.

Fixes: 6b46ca260e ("net: psp: add socket security association code")
Signed-off-by: Daniel Zahka <daniel.zahka@gmail.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260915-psp-ktls-fix-v2-1-0eedc3b148ec@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 19:18:24 -07:00
Jakub Kicinski
dd56c0bc48 Merge branch 'net-stmmac-restore-previous-state-if-tc_setup_dwmac510_mqprio-fails'
Lorenzo Bianconi says:

====================
net: stmmac: restore previous state if tc_setup_dwmac510_mqprio() fails

Restore previous mqprio qdisc configuration if
tc_setup_dwmac510_mqprio() fails running the following configuration:

  $tc qdisc add dev eth0 root handle 1: mqprio queues 2@0 2@2
  $tc qdisc replace dev eth0 root handle 2: mqprio queues 2@0 2@2 fp E P

Propagate FPE preemption-class mapping errors in
tc_setup_dwmac510_mqprio() and tc_taprio_configure().
====================

Link: https://patch.msgid.link/20260911-stmmac-tc_setup_dwmac510_mqprio-error-path-v3-0-a76b1e2547c1@oss.qualcomm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 19:07:17 -07:00
Lorenzo Bianconi
02fffd1939 net: stmmac: preserve real_num_tx_queues on mqprio setup failure
With the FPE preemption-class mapping error now propagated from
stmmac_fpe_map_preemption_class(), tc_setup_dwmac510_mqprio() can fail
on the mapping step. The error path used to call stmmac_reset_tc_mqprio(),
which resets the number of real TX queues to priv->plat->tx_queues_to_use
(the platform maximum), overwriting the value that was active before the
offload was attempted (for example a lower count left over from a previous
mqprio configuration).

The issue can be triggered using the following configuration:

  # First mqprio config lowers the hw queue count below the platform
  # default (e.g. 8 TX queues).
  $tc qdisc add dev eth0 root handle 1: mqprio queues 2@0 2@2

  # Replace mqprio configuration with a second one that fails FPE
  # preemption-class mapping. stmmac driver resets the real_num_tx_queues
  # to the platform maximum, losing the previous configuration.
  $tc qdisc replace dev eth0 root handle 2: mqprio queues 2@0 2@2 fp E P

Save ndev->real_num_tx_queues before lowering it and restore it,
together with the TC-to-queue and priority-to-TC mappings, when the FPE
preemption-class mapping fails, instead of resetting the queue count to
the platform maximum.

Note that a failed setup makes the qdisc layer run mqprio_destroy() on
the new qdisc. Because priv->hw_offload is only assigned after
ndo_setup_tc() succeeds, mqprio_destroy() calls netdev_set_num_tc(dev, 0),
so dev->num_tc ends up 0 regardless of the driver-side restore and the
previous qdisc is not reactivated. The restore is still needed to keep
real_num_tx_queues and to avoid leaving the failed configuration's
TC-to-queue and priority-to-TC mappings in place.

Fixes: 195e4f409a ("net: stmmac: support fp parameter of tc-mqprio")
Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Link: https://patch.msgid.link/20260911-stmmac-tc_setup_dwmac510_mqprio-error-path-v3-2-a76b1e2547c1@oss.qualcomm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 19:07:13 -07:00
Lorenzo Bianconi
90e4b849df net: stmmac: propagate FPE preemption-class mapping errors
stmmac_fpe_map_preemption_class() dispatches through the
stmmac_do_void_callback() helper, which forces the callback's return
value to 0 whenever the op pointer is populated. As a result the
-EINVAL returned by dwmac5_fpe_map_preemption_class() (e.g. when a
preemptible TC owns more than one TXQ under SP scheduling) is silently
swallowed by every caller.

Switch the dispatch macro to stmmac_do_callback() so the callback's real
result is propagated, and honour it in the taprio and mqprio qdisc
offload.

Note that the taprio "if (ret)" check in tc_taprio_configure() used to
be dead code and now becomes live: a preemptible TC spanning more than
one TXQ under SP scheduling cannot be programmed in hardware, so a
taprio or mqprio configuration that previously returned success while
leaving the preemption-class register unprogrammed now fails with
-EINVAL. For taprio, the failure also runs the disable path, tearing
down the schedule that was just installed; this is the intended
behaviour.

Fixes: 195e4f409a ("net: stmmac: support fp parameter of tc-mqprio")
Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Link: https://patch.msgid.link/20260911-stmmac-tc_setup_dwmac510_mqprio-error-path-v3-1-a76b1e2547c1@oss.qualcomm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 19:07:13 -07:00
Guanglei Zhu
c7ead97042 net: wwan: t7xx: validate the netif index in t7xx_ccmni_recv_skb()
The netif index carried in the DPMAIF PIT header is five bits wide,
but ccmni_inst[] only has room for NIC_DEV_MAX (21) entries.
t7xx_ccmni_recv_skb() indexes the array without a bounds check, so
indexes 21 to 31 read past it.  The out-of-bounds value lands in the
callback table that follows the array, which is never NULL, so the
existing !ccmni check does not catch it and the driver dereferences
whatever sits there as a struct t7xx_ccmni.

Drop the skb when the index is out of range.

Fixes: 05d19bf500 ("net: wwan: t7xx: Add WWAN network interface")
Cc: stable@vger.kernel.org
Signed-off-by: Guanglei Zhu <zhugl3@xiaopeng.com>

Verified in a QEMU guest with a fault injector setting the netif
index to 25: the unpatched driver reads a value past ccmni_inst[],
which lands in the callback table, and dereferences it far enough to
queue the skb.  With this check the packet is dropped.  Well-formed
traffic on index 0 is unaffected.

Changes in v2: none.

Link: https://patch.msgid.link/20260911021734.1396599-3-zhugl3@xiaopeng.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 18:58:45 -07:00
Guanglei Zhu
31550d5855 net: wwan: mhi_wwan_mbim: check skb_copy_bits() return value
mhi_mbim_rx() ignores the return value of skb_copy_bits() when it
copies each datagram out of the NTB.  The datagram offset and length
come from the DPE, which is only checked to lie within the NTB
itself, so a modem can point a datagram outside the received skb.
The copy then fails and the freshly allocated skbn is passed to
netif_rx() with its uninitialized contents still in place, leaking
kernel heap memory into the network stack.

Free the skb and account an error when the copy fails.

Fixes: aa730a9905 ("net: wwan: Add MHI MBIM network driver")
Cc: stable@vger.kernel.org
Suggested-by: Loic Poulain <loic.poulain@oss.qualcomm.com>
Signed-off-by: Guanglei Zhu <zhugl3@xiaopeng.com>

Verified in a QEMU guest with a fault injector pointing a DPE
outside the received NTB: the copy fails, and the unpatched driver
hands the uninitialized skbn to the network stack (observed as
"unknown protocol" on bytes that were never written).  With this
check the failed datagram is dropped and counted as an rx error.

Changes in v2: factor the free-and-count sequence out into
mhi_mbim_rx_drop(), shared with the unknown-protocol path, as
suggested by Loic Poulain.

Link: https://patch.msgid.link/20260911021734.1396599-2-zhugl3@xiaopeng.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 18:58:45 -07:00
Guanglei Zhu
5d063822ac net: wwan: mhi_wwan_mbim: guard against a cyclic NDP chain
The NDP traversal in mhi_mbim_rx() only stops when wNextNdpIndex is
zero.  Nothing requires the offsets to advance, so a modem that
points an NDP at itself, or at an earlier NDP, keeps the loop
spinning forever on one CPU.

Break out when the next NDP offset is not larger than the current
one.

Fixes: aa730a9905 ("net: wwan: Add MHI MBIM network driver")
Cc: stable@vger.kernel.org
Suggested-by: Loic Poulain <loic.poulain@oss.qualcomm.com>
Signed-off-by: Guanglei Zhu <zhugl3@xiaopeng.com>

Verified in a QEMU guest with a fault injector feeding the driver's
receive callback an NTB whose single NDP points at itself: the
unpatched driver spins in mhi_mbim_rx() with one CPU pinned at 100%
and the thread never returns.  With this check the loop terminates
within one iteration.

Changes in v2: move the non-increasing check to the wNextNdpIndex
retrieval site, as suggested by Loic Poulain, instead of tracking
the previous offset in a separate variable.

Link: https://patch.msgid.link/20260911021734.1396599-1-zhugl3@xiaopeng.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 18:58:45 -07:00
Linus Walleij
1dd85662fe net: ethernet: cortina: Ack RX overrun interrupt correctly
The RX overrun interrupt is reported in interrupt status register 4, but
gmac_irq() acknowledges it using the RX descriptor error bit from status
register 0. For GMAC0 this writes the GMAC1 overrun bit, while for GMAC1
the shift leaves no bit in the 32-bit register.

Acknowledge the same per-port RX overrun bit that was detected.

Fixes: 4d5ae32f5e ("net: ethernet: Add a driver for Gemini gigabit ethernet")
Signed-off-by: Linus Walleij <linusw@kernel.org>
Link: https://patch.msgid.link/20260914-b4-gemini-ethernet-fixes-2-v2-1-5ab39a047b90@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 17:44:40 -07:00
Eric Dumazet
9ed55f3dbe net: lock the socket in sock_gettstamp()
sk->sk_flags must only be changed while holding the socket lock,
because sock_set_flag() and sock_reset_flag() use non atomic
operations (__set_bit() and __clear_bit()).

sock_gettstamp() is one of the last places where a bit of sk->sk_flags
is changed from a syscall without owning the socket lock, through
sock_enable_timestamp(sk, SOCK_TIMESTAMP).

sk_set_memalloc() and sk_clear_memalloc() also change sk->sk_flags
without the socket lock, but their callers (nbd, iscsi_tcp, nvme-tcp,
sunrpc, wireguard) need a careful audit, this will be addressed in a
separate patch.

Jungwoo Lee and Wongi Lee reported an UDP socket use-after-free
caused by this bug: a SIOCGSTAMPNS_NEW ioctl racing with bind()
can cancel the SOCK_RCU_FREE bit that udp_lib_get_port() just set,
because both threads perform a read-modify-write on the same word.

  CPU 0 (bind)                        CPU 1 (SIOCGSTAMPNS_NEW)
  --------------------------------    ----------------------------
  read sk_flags = F                   read sk_flags = F
  compute F | BIT(SOCK_RCU_FREE)      compute F | BIT(SOCK_TIMESTAMP)
  store F | BIT(SOCK_RCU_FREE)
  sk_add_node_rcu(sk, ...)
                                      store F | BIT(SOCK_TIMESTAMP)

After the lost update, SOCK_RCU_FREE is clear while the socket is
visible to lockless UDP receive lookups. sk_destruct() then frees
the socket immediately instead of waiting for a RCU grace period,
while the receive path still holds a reference-less pointer to it:

 BUG: KASAN: slab-use-after-free in ipv4_pktinfo_prepare+0x30/0x410
 Read of size 8 at addr ffff888008806610 by task exploit/207
 CPU: 0 UID: 1000 PID: 207 Comm: exploit Not tainted 6.12.95+ #1
  ipv4_pktinfo_prepare+0x30/0x410
  udp_queue_rcv_one_skb+0x51c/0x1180
  udp_unicast_rcv_skb+0x109/0x350
  ip_protocol_deliver_rcu+0x14b/0x310
  ip_local_deliver_finish+0x29d/0x390
  ip_local_deliver+0x24d/0x2a0

Only grab the socket lock when SOCK_TIMESTAMP has to be set,
to keep the common case lockless.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Reported-by: Jungwoo Lee <jwlee2217@gmail.com>
Reported-by: Wongi Lee <qw3rtyp0@gmail.com>
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260915043055.3441600-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 17:34:56 -07:00
Jakub Kicinski
490599ab23 eth: fbnic: ring the doorbell if a burst ends in a drop
fbnic_tx_map() skips the doorbell write, and the completion request,
for every packet handed to it with xmit_more set, counting on the
packet which ends the burst to publish them all. When that packet is
dropped instead - skb_put_padto(), skb_cow_head() or a DMA mapping
failure - nothing rings. The descriptors of the preceding packets stay
invisible to the HW until the next transmit on that queue, which for a
burst-then-idle workload may never come.

Remember the meta descriptor of the last packet left without a doorbell
and flush it from the error paths. The completion request has to be set
on that descriptor rather than simply writing the tail, otherwise the HW
would transmit the packets but never report a head, and the ring would
fill up and stall for good.

This is very similar to Joe's recent series of fixes for bnxt.
Not seen in real life, reproduced under QEMU with failure injection.

Fixes: 9a57bacd57 ("eth: fbnic: Add basic Tx handling")
Reviewed-by: Alexander Duyck <alexanderduyck@fb.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260915022327.913218-1-kuba@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 17:33:06 -07:00
Yige Jiang
5ae916fabc net: netsec: fix device_node reference leak on phy_np
netsec_of_probe() takes a reference on the PHY device_node with
of_parse_phandle() and stores it in priv->phy_np, but the driver never
drops it.  One device_node reference is leaked per probe, on the success
path as well as on every error path reached after netsec_of_probe().

Neither consumer takes ownership.  of_mdio_parse_addr() is a static
inline taking a const struct device_node * that only reads the "reg"
property.  of_phy_connect() borrows as well: of_phy_get_and_connect() in
drivers/net/mdio/of_mdio.c brackets its own call with of_node_get() at
:364 and of_node_put() at :373, which would be a double put if
of_phy_connect() consumed the reference.

The node is still in use at netsec_netdev_open() time, where it is
passed to of_phy_connect(), so it has device lifetime.  Release it at
the probe error label, which every failure path after the acquire
funnels through, and in netsec_remove().  Both releases precede
free_netdev(), since priv is netdev_priv(ndev).  The ACPI probe path
leaves priv->phy_np NULL and of_node_put(NULL) is a no-op.

There is no end-user visible symptom on currently supported platforms:
a device_node is only freed once OF_DYNAMIC is enabled and the node has
been detached, so on a static device tree the imbalance is inert.  It is
observable as a refcount that grows across bind/unbind cycles, and would
matter under device tree overlays.

Found by static analysis of reference acquire/release pairing rather
than from a runtime report.  No reproducer was produced and the change
has not been runtime tested; it is compile-tested only (arm64,
CONFIG_SNI_NETSEC=m via COMPILE_TEST).

Fixes: 533dd11a12 ("net: socionext: Add Synquacer NetSec driver")
Signed-off-by: Yige Jiang <yigejiang86@gmail.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260913064102.37452-1-yigejiang86@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 17:32:15 -07:00
Farhad Alemi
150dba2c69 net: remove WARN_ON_ONCE() from the dev_fill_forward_path() loop check
ipip_fill_forward_path() and ip6_tnl_fill_forward_path() look up the
route to the tunnel's remote endpoint and set ctx->dev to its device,
which is the tunnel itself when that route resolves back to the tunnel.
dev_fill_forward_path() then makes no progress and trips
WARN_ON_ONCE(last_dev == ctx->dev) as soon as a flowtable tries to
offload a flow through the tunnel. That routing loop is a configuration
any CAP_NET_ADMIN user can set up, and ip_tunnel_xmit() and
ip6_tnl_xmit() already treat it as a tx error, so remove the warning and
just fail the walk, as commit 008e7a7c29 ("net: remove WARN_ON_ONCE
when accessing forward path array") did for the path stack overflow.

Fixes: ab427db178 ("netfilter: flowtable: Add IPIP rx sw acceleration")
Fixes: d98103575d ("netfilter: flowtable: Add IP6IP6 rx sw acceleration")
Closes: https://lore.kernel.org/all/CA+0ovCgaRvbd0Udj70b2xxG8Cx3CaCpNhnf1V4RWQuDveZYZhA@mail.gmail.com/
Suggested-by: Pablo Neira Ayuso <pablo@netfilter.org>
Signed-off-by: Farhad Alemi <farhad.alemi@berkeley.edu>
Reviewed-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Link: https://patch.msgid.link/CA+0ovCgKDOk+Bg6Gh5Lwx94u_jJjQ30-vY1JcY2BYfhnWJJbPA@mail.gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 17:30:23 -07:00
Mark Amirkan
37213e6112 net/packet: avoid truncating TPACKET_V3 private size
tpacket_req3.tp_sizeof_priv is an unsigned int, and packet_set_ring()
validates the full value against the block size.  init_prb_bdqc() then
stores it in the unsigned short blk_sizeof_priv field.

Commit 2b6867c2ce ("net/packet: fix overflow in check for priv area
size") fixed the validation arithmetic, but an accepted value above
USHRT_MAX still narrows when it is stored.

For a 131072-byte block, tp_sizeof_priv=65536 is valid.  The narrowing
makes offset_to_first_pkt 48 instead of 65584, so packet records can be
placed in the private area that userspace asked the kernel to preserve.

blk_sizeof_priv is internal state, so widen it to hold the validated
UAPI value.

Fixes: f6fb8f100b ("af-packet: TPACKET_V3 flexible buffer implementation.")
Cc: stable@vger.kernel.org
Signed-off-by: Mark Amirkan <markdamirkan@gmail.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260913-b4-send-packet-private-v1-1-925eab2cd388@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 17:29:11 -07:00
Mark Amirkan
60404266ef mptcp: return sk_wait_data() errors from recvmsg()
Commit 5813022985 ("mptcp: error out earlier on disconnect") made
mptcp_recvmsg() stop when sk_wait_data() returns an error.  The error is
stored in err, but the function then jumps to a path which returns
copied.  When no data was copied, recvmsg() therefore returns zero and
reports a false EOF.

Store the result in copied, which is the value returned by the function.
This also keeps the usual partial-read result when data was copied before
the error.

A recvmsg() blocked in one thread reproduces the issue when another
thread disconnects the same MPTCP socket with connect(AF_UNSPEC).
Before this change recvmsg() returns zero; afterwards it returns -EPIPE.

Fixes: 5813022985 ("mptcp: error out earlier on disconnect")
Cc: stable@vger.kernel.org
Signed-off-by: Mark Amirkan <markdamirkan@gmail.com>
Reviewed-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Link: https://patch.msgid.link/20260913-b4-send-mptcp-recv-error-v1-1-4eaa3684a8b8@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 17:28:06 -07:00
Mark Amirkan
33ff111d7b net/packet: clear RX owner on VNET header error
Commit 61fad6816f ("net/packet: tpacket_rcv: avoid a producer race
condition") added rx_owner_map and made tpacket_rcv() claim a V1 or V2
ring slot before converting the virtio-net header.  If the conversion
fails, the drop path leaves the slot claimed.

With a one-frame TPACKET_V2 ring, an unsupported UDP GSO packet leaves
the only slot unavailable, so the ring also drops the next valid packet.

Clear the ownership bit on this error path.  TPACKET_V3 already clears
its block state here.

Fixes: 61fad6816f ("net/packet: tpacket_rcv: avoid a producer race condition")
Cc: stable@vger.kernel.org
Signed-off-by: Mark Amirkan <markdamirkan@gmail.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260913-b4-send-packet-vnet-v1-1-5545ffb528ae@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 17:27:35 -07:00
Mark Amirkan
a9ce4053dc net: lan743x: fix RX checksum use-after-free
lan743x_rx_process_buffer() adds each non-first receive buffer to the
head skb's frag_list.  On the last descriptor, lan743x_rx_trim_skb()
linearizes the head and frees the fragment skb metadata.

The checksum-success path then writes ip_summed through the local skb
pointer, which still points to the final fragment.  This causes a
use-after-free write when a packet spans more than one receive buffer.

Set ip_summed on the surviving head skb instead.  Multi-buffer receive
can occur after a live MTU increase because existing ring entries keep
their old buffer size until they are replenished.

A KUnit test invoking lan743x_rx_process_buffer() with a two-buffer
packet produced a one-byte KASAN use-after-free write before this change.
The same test passed after the change.  The driver object also builds
with W=1.  This was not tested on physical LAN743x hardware.

Fixes: cd6910501c ("net: lan743x: Add support for Rx IP & TCP checksum offload")
Cc: stable@vger.kernel.org
Signed-off-by: Mark Amirkan <markdamirkan@gmail.com>
Reviewed-by: Chenguang Zhao <zhaochenguang@kylinos.cn>
Link: https://patch.msgid.link/20260913-b4-send-lan743x-uaf-v1-1-73d563d08ba9@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 17:26:40 -07:00
Jamal Hadi Salim
f6fb2ac5e1 selftests/tc-testing: add codel/fq_codel interval boundary cases
Add tdc cases locking the codel/fq_codel small-interval uAPI after
the dropping-loop bound (previous patch): sub-tick and two-tick
intervals are ACCEPTED (the loop bound makes them safe), the
1024us boundary is accepted, and a sub-tick target sojourn delay is
accepted (it does not participate in the control law):

  codel:     6e44/a8c3/a695/9793 - interval 1us/3us/1024us and
             target 1us accepted (rendered 0us/2us/1.02ms/0us by tc)
  fq_codel:  1b4d/3540/49c5/3e0f - interval 1us/3us/1024us and
             target 1us accepted

The positive cases match the full rendered qdisc line (tc renders
interval 1us as 0us, 3us as 2us, 1024us as 1.02ms), mirroring the
existing tests in these files.

These cases do not test the dropping-loop bound itself: tdc cannot
observe per-dequeue drop counts. c797 (fq_codel target 1 interval 1)
passes unmodified on the patched kernel, which is the uAPI evidence
for the previous patch.

Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-1L5H.v1.20260912080102@mojatatu.com.2
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 17:15:29 -07:00
Jamal Hadi Salim
7f4a5ec625 net/sched: codel: bound the dropping loop per dequeue call
The CoDel control law schedules the next drop one interval/sqrt(count)
after the previous drop, using the configured interval
(codel_params.interval). For very small intervals the scheduled step
rounds down to zero, so the dropping loop in codel_dequeue() never
advances and drains the entire backlog under the qdisc lock in one
call - an unprivileged user can trigger a soft lockup this way.

Fix in the shared codel code used by both codel and fq_codel:

1. Make the control-law step at least 1 tick so the dropping loop
   always moves forward.

2. Cap the dropping loop at CODEL_MAX_DROPS_PER_DEQUEUE (256) drops
   per codel_dequeue() call, resyncing drop_next to now when the cap
   is hit: the catch-up owed to the loop grows with the idle gap and
   the backlog, which no interval threshold can bound. This is a
   deliberate behaviour change after long idle gaps.

The cap applies to fq_codel (4b549a2ef4) and the mac80211 TXQ path
(fixed interval, cap only).

The target sojourn delay (codel_params.target) is not validated: it
does not feed the control law, so a sub-tick value is aggressive
rather than deadlock-prone.

Conditions to recreate the bug:
  - tc qdisc add dev lo root handle 1: tbf rate 1kbit burst 2kb limit 1000000
  - tc qdisc add dev lo parent 1:1 handle 10: codel interval 2us target 1ms noecn limit 1000000 (same for fq_codel)
  - unpatched kernel: tc accepts it; a UDP flood under the 1kbit tbf
    soft-lockups (watchdog: BUG: soft lockup) while one
    codel_dequeue() call drops the backlog under the qdisc lock
  - patched kernel: same setup, at most 256 drops per dequeue call,
    no soft lockup

Testing: claim reproducer and interval 2us/3us variants run clean;
tdc qdisc category passes (see the selftests patch).

Fixes: 76e3cc126b ("codel: Controlled Delay AQM")
Reported-by: Vega <vega@nebusec.ai>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Reviewed-by: Toke Høiland-Jørgensen <toke@toke.dk>
Link: https://patch.msgid.link/QDISC-1L5H.v1.20260912080102@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 17:15:29 -07:00
Jakub Kicinski
fefaac1176 Many fixes:
- mac80211: S1G TIM bitmap fix
  - ath12k: remove undocumented DT ABI implementation
  - various firmware API and over-the-air hardening changes
  - fixes for most cfg80211/mac80211 syzbot reports
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEpeA8sTs3M8SN2hR410qiO8sPaAAFAmqqVJYACgkQ10qiO8sP
 aACbjQ//VK16MorAA6e+zC44Cq68erE/0kCXjkwHjNPFrg8P2gD3qVHtPMUi5GLG
 ldWD5Hm3SlmGlpXIljJPQpTaMHGwVG+1j+n4TgwonomXarsaftSriSPKMww3/O2c
 kk+LAujowo0OImPNxY4noWRVcEcSo/tsTvbYDcAKjc91yHs/gU6TSeT3E/YFAfbt
 HSNbMImaSdTBM6TnvwepGZG2RisKWMgiYhpPFzO26TvoYnXBCgN48kOVz+x/6M0P
 o7i3coZiRO1o7ec9SsxgKUZUDNVMNoiEDWPQBJpaq2vOV1eFCrwuJr4/KKYk738F
 zSGsIK8z1R6j7RRB2qk3qYRrrwuNS1FZWeo4S0iGWvLwqbL7nsrrvkFc5O9uFure
 RV/Uf2okycaIZICe1rSalTDtjWgp6beRQSj3Ep75MSa/iqj6Rtggo9CxM7+aPkoa
 z0Q9MKqmuY0yXZCtI0EOayXcOpdoHWot95NJQ5nkR0ge5WJ2tu+myzepesQ9a5Vb
 JEOw8N4gJ9Hdst94gDqiGzDQb0xRYeMSgET3we9gA9KMu3NXBrbRLhRbkk+16fdA
 DoqbwD6wrH6nXcgihLBEILRLQWHe8pC7PoI2LvuTfzP3xiD50NvRVfdwNDuWLg/+
 2uP1Yl4x67vXNUqa619taRwtwFuJn+M31uwZ57eMw3U1edNlfug=
 =/zWX
 -----END PGP SIGNATURE-----

Merge tag 'wireless-2026-09-16' of https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wireless

Johannes Berg says:

====================
Many fixes:
 - mac80211: S1G TIM bitmap fix
 - ath12k: remove undocumented DT ABI implementation
 - various firmware API and over-the-air hardening changes
 - fixes for most cfg80211/mac80211 syzbot reports

* tag 'wireless-2026-09-16' of https://git.kernel.org/pub/scm/linux/kernel/git/wireless/wireless: (67 commits)
  wifi: brcmsmac: fix UAF in brcms_free_timer()
  wifi: brcmfmac: fix lost 802.1x TX completion wakeup
  wifi: ath11k: cleanup arsta in ath11k_mac_peer_cleanup_all()
  wifi: wcn36xx: Fix potential use-after-free in TX ack timer teardown
  wifi: ath12k: ahb: Revert undocumented ABI and dead code
  wifi: mac80211: refuse to make a monitor active when it has no queue
  wifi: libipw: reject TKIP frames without a full MIC
  wifi: virt_wifi: don't transfer operstate before register
  wifi: cfg80211: check if AP has been started or joined a mesh before adding new station
  wifi: cfg80211: move link_id validation earlier in nl80211_new_station()
  wifi: cfg80211: do not support direct add of station to AP_VLAN interfaces
  wifi: cfg80211: verify if AP_VLAN belongs to the correct AP
  wifi: mac80211: set up the TX info early to fix failure paths
  wifi: mac80211: mesh: release the channel if start fails
  wifi: mac80211: mesh: reset the CSA state when leaving
  wifi: mac80211: add HE 6 GHz capability in the scan elems len
  wifi: mac80211: don't access the TSF of a down interface
  wifi: mac80211: don't RCU-dereference the mesh CSA settings we just set
  wifi: mac80211: don't allow link changes when iface is down
  wifi: mac80211: require a peer station for TDLS setup confirm
  ...
====================

Link: https://patch.msgid.link/20260916083642.110609-3-johannes@sipsolutions.net
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 15:54:56 -07:00
Jakub Kicinski
7c7d5e9d7e ipsec-2026-09-16
-----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEH7ZpcWbFyOOp6OJbrB3Eaf9PW7cFAmqqagcACgkQrB3Eaf9P
 W7ervw//ValStkQ37oXUHXQkevKs+rAFhleUHogbxG+C/45nx7ToVgtL0oRFs/ss
 WmxjClmthosPdcgz9uLyIi/xiIWmwa9iYsmq/A7FF5bewz8ha0/jzkxi6EZvGFxK
 oHJWIe11pcDV3NoEjA33Z6k63vZcVhD3QVlIk4mg4b4leEDgTw6y8z/k5sv6/QTc
 xP7lykI6VeKPyBydSAdwHonJZ3BWLE0/Y2uiQtVVZAFT9Ln2cpqAPpwYGDFKpgM/
 H8NYqB4X3fhdxc/rQhqdgS0mdhEwukulyYg/znwUI/DYb5Dt4Eh51y0fpIcjFX2y
 Gp31bMXih7j66SuOwEXx48dSNgxHEXwbUdUoeQReVmFam0epbOyPoBgT5ZLskWfo
 JbdeZR7kAugCnX/XpTmI8pk3G52i6LvCY1GDPyIYTFGcVHg9Y9plw1WEs88pqGNC
 9+pkRapF04gdaxjl4FioKdtoKXTyYiVxEB/nWkBfUE9Q5iS16/G9Lrcd4Dw3aN8f
 imsTf1yW6kJ6RbF9QKS4kuzVJnPFwxg/EmmHFQtpZOkH23BwZEU7obqIwpNYgnoG
 /A9vnwQdAxGU4ar6n2wSPXABzFYlsJAwyMlOGrOkzWC9t/TAA5NHjkyskO5pY8Mt
 0bhWCiNfRYzbSF94UEltYsKZKrx9kV6c2lCltKCfR89KhY5Dorw=
 =skHa
 -----END PGP SIGNATURE-----

Merge tag 'ipsec-2026-09-16' of git://git.kernel.org/pub/scm/linux/kernel/git/klassert/ipsec

Steffen Klassert says:

====================
pull request (net): ipsec 2026-09-16

1) xfrm: iptfs: fix stack OOB read in iptfs_skb_reset_frag_walk()
   Add the up-front nr_frags guard iptfs_skb_add_frags() already has,
   so an out-of-range offset can't walk past the on-stack frags[] array.

2) xfrm: serialize state GC with device state flush
   Serialize xfrm_state destruction against the deferred-device pass
   with a dedicated mutex, since the device GC list doesn't hold a state
   reference and the two paths could free the same state.

3) xfrm: add missing RCU read lock in xfrm_send_migrate_state()
   Hold the RCU read lock around xfrm_nlmsg_multicast() so the
   rcu_dereference() of net->xfrm.nlsk doesn't warn.

4) xfrm: iptfs: fix runt reassembly panic from short inner tot_len
   Require the runt length to cover at least the minimum IP header,
   so a tot_len in [6, 19] (IPv4) can't write past the declared length
   and trip skb_over_panic().

5) ipv6: xfrm: use full sockets in local error paths
   Use skb_to_full_sk() in xfrm6_local_rxpmtu() and xfrm6_local_error()
   and bail out without a full socket, so a TCP_NEW_SYN_RECV request_sock
   isn't miscast as a full inet/IPv6 socket.

6) xfrm: fix compat ALLOCSPI request use-after-free
   Drop the redundant alloc_compat() in xfrm_alloc_userspi() so the
   compat translator no longer reads past the payload and publishes a
   child a multicast clone can still see after xfrm_user_rcv_msg() frees.

7) xfrm: add missing rcu_read_lock(), skb_dst_force() and dev_hold() for xfrm_trans_reinject()
   Force the dst before queuing, hold dev across the workqueue deferral,
   and take rcu_read_lock() around the finish() loop, so transport-mode
   reinjection doesn't deref non-refcounted dst/dev under workqueue.

8) xfrm: use hlist_del_init_rcu for state_cache and state_cache_input
   Switch to hlist_del_init_rcu() so a second __xfrm_state_delete() is
   a no-op instead of writing through LIST_POISON2, closing the UAFs.

9) esp: downgrade zerocopy managed frags before mutating skb frags
   Call skb_zcopy_downgrade_managed() before ESP rewrites the skb frag
   array, so per-frag unrefs in esp_ssg_unref() and skb_release_data()
   stay balanced for ubuf-owned managed frags.

10) xfrm: hold net_device reference under RCU in bundle creation
    Read dst->dev via dst_dev_rcu() and keep RCU active through
    xfrm_fill_dst(), so a concurrent RTM_DELLINK can't free dev
    under bundle creation.

11) xfrm: save input state data before secpath resets
    Save the state protocol on the stack while it's still valid and
    use the saved address family for transport_finish(), so post-reset
    dereferences (VTI, XFRM if, MAX_DEPTH error) can't UAF the state.

12) net: xfrm: reject unrepresentable espintcp transport headers
    Use the careful transport-header helper and drop the skb through
    the XFRM error path when the offset can't be represented, instead
    of silently truncating it.

* tag 'ipsec-2026-09-16' of git://git.kernel.org/pub/scm/linux/kernel/git/klassert/ipsec:
  net: xfrm: reject unrepresentable espintcp transport headers
  xfrm: save input state data before secpath resets
  xfrm: hold net_device reference under RCU in bundle creation
  esp: downgrade zerocopy managed frags before mutating skb frags
  xfrm: use hlist_del_init_rcu for state_cache and state_cache_input
  xfrm: add missing rcu_read_lock(), skb_dst_force() and dev_hold() for xfrm_trans_reinject()
  xfrm: fix compat ALLOCSPI request use-after-free
  ipv6: xfrm: use full sockets in local error paths
  xfrm: iptfs: fix runt reassembly panic from short inner tot_len
  xfrm: add missing RCU read lock in xfrm_send_migrate_state()
  xfrm: serialize state GC with device state flush
  xfrm: iptfs: fix stack OOB read in iptfs_skb_reset_frag_walk()
====================

Link: https://patch.msgid.link/20260916101938.118628-1-steffen.klassert@secunet.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 15:54:19 -07:00
Eric Dumazet
ceac0de741 netlink: do not free nlk->groups while lockless readers can use it
netlink_realloc_groups() uses krealloc() under netlink_table_grab().
Whenever NLGRPSZ(groups) lands in a different kmalloc bucket, the old
bitmap is freed immediately.

Two readers of nlk->groups / nlk->ngroups do not hold the netlink
table lock:

1) sk_diag_dump_groups(). Hashed (bound) sockets are dumped from the
   rhashtable walk in __netlink_diag_dump(), which only holds RCU.
   Only the mc_list part of the dump takes nl_table_lock.

2) netlink_native_seq_show() (/proc/net/netlink), whose walk has been
   lockless since commit 21e4902aea ("netlink: Lockless lookup with
   RCU grace period in socket release").

Both can read a freed buffer, and sk_diag_dump_groups() can also read
past the end of the old (smaller) buffer if it happens to load the old
@groups pointer together with the new @ngroups value, copying the
result into a NETLINK_DIAG_GROUPS attribute.

This is the same class of bug that commit f773608026 ("netlink:
access nlk groups safely in netlink bind and getname") fixed for bind()
and getname(); these two readers were missed. Simply grabbing the table
lock in sk_diag_dump_groups() is not an option, because it is also
called with nl_table_lock already held from the mc_list section of the
dump.

Make the lockless readers safe instead:

- Allocate a new bitmap and free the old one after an RCU grace period,
  instead of relying on the implicit kfree() done by krealloc().

- Publish @groups before @ngroups, both with release semantics, and have
  the lockless readers load @ngroups first. A reader can then never pair
  the new (bigger) size with the old (smaller) buffer, and a reader
  picking up the new pointer while still seeing the old size is
  guaranteed to see the initialized bitmap.

netlink_realloc_groups() is called from process context (bind() and
setsockopt()), so kfree_rcu_mightsleep() can be used, once the table
has been released.

Fixes: 21e4902aea ("netlink: Lockless lookup with RCU grace period in socket release")
Fixes: ad20207432 ("netlink: Use rhashtable walk interface in diag dump")
Reported-by: James Burton <jamesburton@meta.com>
Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260911160804.917099-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 18:44:02 -07:00
Nikolay Aleksandrov
2842ce397d net: bridge: vlan: fix bugs caused by switchdev deletion errors
Allowing switchdev to prevent vlan deletion and error out in __vlan_del
could cause multiple different issues - inconsistent state, memory leaks
when flushing, NULL pointer dereference on bridge error when flushing.
It doesn't make sense to allow it to stop __vlan_del, so log the error
and continue with software vlan deletion. This is also consistent with
8021q behaviour.

Suggested-by: Ido Schimmel <idosch@nvidia.com>
Fixes: bf361ad381 ("net: bridge: check __vlan_vid_del for error")
Fixes: 5454f5c28e ("net: bridge: vlan: check for errors from __vlan_del in __vlan_flush")
Fixes: 2594e9064a ("bridge: vlan: add per-vlan struct and move to rhashtables")
Fixes: 9c86ce2c1a ("net: bridge: Notify about bridge VLANs")
Signed-off-by: Nikolay Aleksandrov <razor@blackwall.org>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260914105258.3436918-1-razor@blackwall.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 18:31:57 -07:00
Lorenzo Bianconi
f0ef4b1eae net: stmmac: do not overwrite phc_index when no PTP clock is registered
stmmac_get_ts_info() reports phc_index as 0 when hardware timestamping
is supported but no PTP clock has been registered yet (e.g. while the
interface is down). Zero is a valid PHC index and would make userspace
resolve the wrong clock; the absence of a clock should be reported as
-1.

The ethtool core already initializes phc_index to -1 before invoking
the get_ts_info callback (ethtool_init_tsinfo()), so just drop the
erroneous assignment.

Fixes: 9364fa7fcf ("net: stmmac: Remove setting of RX software timestamp")
Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Reviewed-by: Rahul Rameshbabu <rrameshbabu@nvidia.com>
Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Reviewed-by: Gal Pressman <gal@nvidia.com>
Link: https://patch.msgid.link/20260914-stmmac-fix-phc_index-v2-1-bf3d90373fe4@oss.qualcomm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 18:25:06 -07:00
Zhiling Zou
3f118c8217 openvswitch: avoid reallocating confirmed conntrack labels
ovs_ct_get_conn_labels() adds the labels extension when a conntrack
entry does not have one.  Confirmed conntracks can be read locklessly,
so adding an extension may reallocate and free the extension block
while another CPU accesses it.

Only add the extension for unconfirmed conntracks.  A confirmed
conntrack without labels now fails the caller's label operation instead
of reallocating its extension storage.

Fixes: c2ac667358 ("openvswitch: Allow matching on conntrack label")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Signed-off-by: Zhiling Zou <zhilinz@nebusec.ai>
Reviewed-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Link: https://patch.msgid.link/372fbb062b40ae6723684f55484be86ff0064f8e.1789218015.git.zhilinz@nebusec.ai
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 18:00:46 -07:00
Jakub Kicinski
ad77dba64d Merge branch 'net-drop_monitor-fix-concurrency-issues-preemption-warning-and-buffer-overrun'
Eric Dumazet says:

====================
net: drop_monitor: fix concurrency issues, preemption warning, and buffer overrun

This series addresses several issues discovered in the drop_monitor subsystem:

Patch 1 adds missing tracepoint unregistration synchronization to the
net_dm_trace_on_set() error unwind path, preventing in-flight probes
from scheduling work after the module reference has been dropped.

Patch 2 resolves a race condition during monitoring teardown where per-CPU
timers can be re-armed after deletion if a concurrent worker encounters a
memory allocation failure, switching to timer_shutdown_sync().

Patch 3 fixes a CONFIG_DEBUG_PREEMPT warning reported by syzbot when
kfree_skb() is invoked from preemptible process context, using raw_cpu_ptr()
since each per-CPU queue is safely protected by its own spinlock.

Patch 4 fixes an out-of-bounds write in reset_per_cpu_data() where memset()
overwrote the allocated SKB tailroom by sizeof(struct nlattr) bytes.
====================

Link: https://patch.msgid.link/20260910204612.3762015-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 17:58:37 -07:00
Eric Dumazet
439f392084 drop_monitor: fix out-of-bounds write in reset_per_cpu_data()
In reset_per_cpu_data(), al is computed as:

    al = sizeof(struct net_dm_alert_msg);
    al += dm_hit_limit * sizeof(struct net_dm_drop_point);
    al += sizeof(struct nlattr);

    skb = genlmsg_new(al, GFP_KERNEL);
    ...
    nla = nla_reserve(skb, NLA_UNSPEC, sizeof(struct net_dm_alert_msg));
    ...
    msg = nla_data(nla);
    memset(msg, 0, al);

Because al includes sizeof(struct nlattr) (the 4-byte attribute header),
genlmsg_new() allocates al bytes of tailroom starting at nla.
However, msg points to nla_data(nla), which is located
sizeof(struct nlattr) bytes past nla. Calling memset(msg, 0, al)
therefore writes al bytes starting from msg, exceeding the allocated
buffer by sizeof(struct nlattr) (4 bytes) and corrupting
skb_shared_info.

Fix this by letting al represent only the payload length, allocating
the skb with genlmsg_new(nla_total_size(al), GFP_KERNEL), and zeroing
al bytes from msg.

Fixes: 683703a26e ("drop_monitor: Update netlink protocol to include netlink attribute header in alert message")
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Link: https://patch.msgid.link/20260910204612.3762015-5-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 17:58:35 -07:00
Eric Dumazet
c19b7d3508 drop_monitor: use raw_cpu_ptr() in tracepoint probes
syzbot reported a preemption warning in sk_skb_reason_drop():

 BUG: using smp_processor_id() in preemptible [00000000] code: syz.0.17/5917
 caller is net_dm_packet_trace_kfree_skb_hit+0x119/0x350 net/core/drop_monitor.c:519

In net_dm_packet_trace_kfree_skb_hit(), data = this_cpu_ptr(&dm_cpu_data)
is evaluated before spin_lock_irqsave(&data->drop_queue.lock, flags).
When kfree_skb() is called from preemptible context (e.g. process context
during close() on /dev/net/tun), preemption is enabled, triggering the
CONFIG_DEBUG_PREEMPT warning in smp_processor_id().

The same pattern exists in net_dm_hw_trap_summary_probe() and
net_dm_hw_trap_packet_probe() for dm_hw_cpu_data.

This is a false positive because each per-cpu structure is protected
by its own spinlock. If the task migrates to another CPU right after
reading the per-cpu pointer, the lock still safely synchronizes
access to that queue.

Use raw_cpu_ptr() instead of this_cpu_ptr() to silence
CONFIG_DEBUG_PREEMPT without disturbing interrupt state or breaking
PREEMPT_RT locking semantics.

Fixes: ca30707dee ("drop_monitor: Add packet alert mode")
Fixes: 5855357cd4 ("drop_monitor: Prepare probe functions for devlink tracepoint")
Reported-by: syzbot+dc57fd6722deb17e92af@syzkaller.appspotmail.com
Closes: https://lore.kernel.org/netdev/6aa316b2.f81106d8.2ab401.0014.GAE@google.com/
Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260910204612.3762015-4-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 17:58:35 -07:00
Eric Dumazet
c391a40f71 drop_monitor: use timer_shutdown_sync() to prevent timer rearming during teardown
In drop_monitor teardown paths (net_dm_trace_off_set(),
net_dm_hw_monitor_stop(), and error unwind paths in net_dm_trace_on_set()
and net_dm_hw_monitor_start()), per-CPU timers are stopped using
timer_delete_sync() followed by cancel_work_sync().

However, there is a circular dependency between send_timer and
dm_alert_work:
1) sched_send_work() (timer callback) schedules dm_alert_work.
2) send_dm_alert() / net_dm_hw_summary_work() calls reset_per_cpu_data()
   or net_dm_hw_reset_per_cpu_data().
3) If memory allocation fails under memory pressure in the reset
   function, it re-arms the timer via mod_timer(&data->send_timer, ...).

If dm_alert_work is running concurrently while timer_delete_sync()
executes on another CPU, an allocation failure in the worker will
re-arm the timer after timer_delete_sync() has already returned.
Once cancel_work_sync() completes and module_put() is called, the timer
remains active in the timer wheel. If the module is then unloaded, the
timer will fire and execute sched_send_work() in freed memory,
triggering a kernel panic / use-after-free.

Switch from timer_delete_sync() to timer_shutdown_sync(). This guarantees
that any in-flight timer handler has finished and prevents subsequent
re-arming attempts from running workers from succeeding. When monitoring
is restarted later, timer_setup() is invoked, which cleanly
re-initializes the timer.

Fixes: 9398e9c0b1 ("drop_monitor: Perform cleanup upon probe registration failure")
Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260910204612.3762015-3-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 17:58:34 -07:00
Eric Dumazet
6a038ef2b5 drop_monitor: synchronize tracepoint unregistration on error path
If register_trace_napi_poll() fails in net_dm_trace_on_set(),
unregister_trace_kfree_skb() is called to roll back the kfree_skb
tracepoint registration.

However, tracepoint_synchronize_unregister() is omitted before calling
cancel_work_sync() and module_put(). An in-flight probe executing
concurrently on another CPU could call schedule_work() after
cancel_work_sync() has already returned, leaving a pending work item
scheduled after the module reference is dropped. If the module is then
unloaded, executing the work item triggers a kernel panic.

Add tracepoint_synchronize_unregister() after unregister_trace_kfree_skb()
in the error path, matching net_dm_trace_off_set() and
net_dm_hw_probe_unregister().

Fixes: 7c747838a5 ("drop_monitor: Split tracing enable / disable to different functions")
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Link: https://patch.msgid.link/20260910204612.3762015-2-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 17:58:34 -07:00
Gris Ge
455ebeadf7 net: ip_tunnel: initialize options_len before referencing options
The following command triggers a kernel panic:

  ip link add d0 type dummy; ip link set d0 up
  ip route add 10.30.0.0/16 \
    encap ip id 300 geneve_opts 4660:66:11223344 dev d0

  memcpy: detected buffer overflow: 4 byte write of buffer size 0
  kernel BUG at lib/string_helpers.c:1044!
  ...
  ip_tun_parse_opts.part.0.cold+0x10/0x10
  ip_tun_build_state+0x116/0x2a0

On kernels built with GCC 15+ and `CONFIG_FORTIFY_SOURCE`, the fortified
`memcpy()` got 0 sized destination with request of 4 bytes length:

  static int ip_tun_parse_opts_geneve(...)
  {
      ...
      attr = tb[LWTUNNEL_IP_OPT_GENEVE_DATA];
      data_len = nla_len(attr); /* == 4 */

      struct geneve_opt *opt = ip_tunnel_info_opts(info) + opts_len;
      memcpy(opt->opt_data, nla_data(attr), data_len);
      /*     ^^^^^^^^^^^^^ 0 since options_len is assigned afterwards */

Fixed by initializing the counter before the options are referenced.
Matching what `tunnel_key_opts_set()` already does.

Fixes: bb5e62f2d5 ("net: Add options as a flexible array to struct ip_tunnel_info")
Cc: stable@vger.kernel.org
Signed-off-by: Gris Ge <cnfourt@gmail.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Reviewed-by: Gustavo A. R. Silva <gustavoars@kernel.org>
Link: https://patch.msgid.link/20260913090851.468216-1-cnfourt@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 17:34:31 -07:00
Eric Dumazet
ecc7253683 pppoatm: ensure a writable skb header and linear data
In pppoatm_send(), LLC encapsulation checks whether there is sufficient
headroom for the 4-byte LLC header, but does not ensure that the skb header
is writable.

Normal transmit packets passing through ppp_start_xmit() have their header
unshared via skb_cow_head(). However, packets can also reach pppoatm_send()
via PPP channel bridging (PPPIOCBRIDGECHAN) without going through
ppp_start_xmit().

Use skb_cow_head() to ensure both sufficient headroom and a writable
header before pushing the LLC header.

While at it:
- Call pskb_may_pull(skb, 1) before inspecting skb->data[0] to prevent
  out-of-bounds reads on zero-length or non-linear frames (e.g. from
  bridging).
- Defer SC_COMP_PROT protocol compression until after pppoatm_may_send()
  succeeds. This eliminates the temporary skb allocation on admission failure
  and completely removes the fragile "undo" heuristic at the nospace label,
  avoiding any risk of reading uninitialized headroom or performing an
  unbalanced skb_push().

Fixes: 4cf476ced4 ("ppp: add PPPIOCBRIDGECHAN and PPPIOCUNBRIDGECHAN ioctls")
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260912233048.3977192-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 17:07:01 -07:00
Jakub Kicinski
562219874c Merge branch 'net-sched-fix-action-batch-deletion-cleanup'
Xuanqiang Luo says:

====================
net/sched: fix action batch deletion cleanup

Batched RTM_DELACTION requests can leak references to unprocessed actions
when deletion stops at a filter-bound action.

Patch 1 fixes the failure cleanup.

Patch 2 adds tc-testing regression coverage.

Failure reproduction (key output excerpts):

  python3 tdc.py -f /root/tc-testing/batch-delete.json

not ok 1 d710 - Release tail references after first action deletion fails
	Could not match regex pattern. Verify command output:
[...]
	 index 2 ref 2 bind 0
[...]
	 index 3 ref 2 bind 0

not ok 2 d711 - Release tail references after middle action deletion fails
	Could not match regex pattern. Verify command output:
[...]
	 index 3 ref 2 bind 0

not ok 3 d713 - Delete a tail action once after a failed batch
	Could not match regex pattern. Verify command output:
total acts 2
[...]
	 index 2 ref 1 bind 0
====================

Link: https://patch.msgid.link/20260910093413.34509-1-xuanqiang.luo@linux.dev
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 17:03:14 -07:00
Xuanqiang Luo
14c5eb685c selftests: tc-testing: test action batch deletion failure cleanup
Add tests for cleanup after a batched RTM_DELACTION request fails at
a gact action bound to a filter. Check that subsequent actions retain
their original reference counts and that earlier successful deletions
are preserved.

Cover failures at the first and middle entries. Verify that a remaining
unbound action can be removed with one subsequent delete.

Signed-off-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Link: https://patch.msgid.link/20260910093413.34509-3-xuanqiang.luo@linux.dev
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 17:03:09 -07:00
Xuanqiang Luo
6e05e46fa8 net/sched: act_api: release tail references on DELACTION failure
A batched RTM_DELACTION request takes a temporary reference on each
action before attempting any deletion. tcf_action_delete() clears
each processed slot and drops its temporary reference before attempting
the deletion. If deletion fails, tca_action_gd() calls
tcf_action_put_many() to release the remaining references, but its
tcf_act_for_each_action() iterator stops at the first NULL slot.

When a batch stops at an action bound to a filter, this leaks a
reference on each subsequent action. A later delete of an unbound
action can then return success without removing it from the IDR.

Walk the full array in tcf_action_put_many() and skip NULL slots to
release the references held on the unprocessed actions.

Fixes: a0e947c9cc ("net/sched: act_api: avoid non-contiguous action array")
Cc: stable@vger.kernel.org
Signed-off-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Link: https://patch.msgid.link/20260910093413.34509-2-xuanqiang.luo@linux.dev
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 17:03:09 -07:00
Lorenzo Bianconi
15989abd74 net: stmmac: fix TSO header length truncation
stmmac_tso_xmit() stores the protocol header length returned by
stmmac_tso_header_size() in a u8. stmmac_tso_valid_packet() admits
headers up to 1023 bytes, so a header longer than 255 bytes wraps modulo
256 (486 becomes 230, 256 becomes 0).

A TCP over IPv6 socket carrying a few hundred bytes of sticky
destination/hop-by-hop options makes skb_tcp_all_headers() exceed 255
while staying below the 1023-byte limit, so such an skb reaches
stmmac_tso_xmit().

Widen proto_hdr_len to unsigned int, which is sufficient since the value
is bounded by the hardware limit, and adjust the debug print specifier
accordingly.

Fixes: 9edfa7dab8 ("net: stmmac: enable TSO for IPv6")
Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/20260911-stmmac-fix-header-length-v1-1-8fc103334327@oss.qualcomm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 16:55:52 -07:00
Jakub Kicinski
433cfc3025 Merge branch 'af_unix-fix-inconsistent-scc_index'
Kuniyuki Iwashima says:

====================
af_unix: Fix inconsistent scc_index.

James Burton reported that a single SCC could have multiple
scc_index and unix_vertex_dead() fails to detect a dead SCC.

Patch 1 fixes it and Patch 2 adds a test case.
====================

Link: https://patch.msgid.link/20260912030852.1467872-1-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 16:45:07 -07:00
Kuniyuki Iwashima
b645ccd410 selftest: af_unix: Add test case with mixed lowpoint in scm_rights.c.
The new test case creates two SCCs so that each of them
has multiple scc_index.

Without patch, GC cannot free the sockets and the test fails.

  #  RUN           scm_rights.dgram.mixed_lowpoints ...
  # scm_rights.c:176:mixed_lowpoints:Expected 0 (0) == ret (12)
  # mixed_lowpoints: Test terminated by assertion
  #          FAIL  scm_rights.dgram.mixed_lowpoints
  not ok 5 scm_rights.dgram.mixed_lowpoints
  ...
  # FAILED: 45 / 50 tests passed.
  # Totals: pass:45 fail:5 xfail:0 xpass:0 skip:0 error:0

With the patch, all tests pass.

  # PASSED: 50 / 50 tests passed.
  # Totals: pass:50 fail:0 xfail:0 xpass:0 skip:0 error:0

Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260912030852.1467872-3-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 16:45:05 -07:00
Kuniyuki Iwashima
4a4263dfea af_unix: Unify scc_index when finalising SCC in __unix_walk_scc().
Commit bfdb01283e ("af_unix: Assign a unique index to SCC.")
changed Tarjan's algorithm to update lowlink with lowlink,
which is called lowpoint (unix_vertex.scc_index).

unix_vertex_dead() assumes all vertices in an SCC share the same
lowpoint, but this is not always true if an SCC has two or more
back edges, depending on the order of DFS.

For example, the graph below has two back edges from B to A
and from C to B.

  A --> B --> C
  ^    | ^    |
  `----' `----'

If DFS walks through A -> B -> C -> B (-> C -> B) -> A (-> B -> A),
each index and scc_index will be updated as follows.

  A --> B --> C    C = (3, 3)  (index, scc_index)
                   B = (2, 2)
                   A = (1, 1)

  A ... B ... C    C = (3, 2)<-.
         ^    |    B = (2, 2) -'
         `----'    A = (1, 1)

  A ... B ... C    C = (3, 2)
  ^    | .    .    B = (2, 1)<-.
  `----'  ....     A = (1, 1) -'

Then, unix_vertex_dead() thinks that B is passed to another
SCC with scc_index 2, and the SCC is not garbage-collected.

This does not happen if DFS walks in a different order below
or starts from B.

    1      3
  A --> B --> C
  ^    | ^    |
  `----' `----'
     2      4

Let's unify scc_index across the SCC when finalising it.

Note that updating v->index was previously done in unix_scc_dead(),
when called from __unix_walk_scc(), just to save one loop.  Since
__unix_walk_scc() now iterates over the SCC anyway, the update is
moved back to __unix_walk_scc() and 'fast' argument is dropped.

Fixes: 4090fa373f ("af_unix: Replace garbage collection algorithm.")
Reported-by: James Burton <jamesburton@meta.com>
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260912030852.1467872-2-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 16:45:05 -07:00
Jakub Kicinski
c5e367a8a3 bluetooth pull request for net:
Core:
 
  - hci: put the peer's on-air address on air when we cannot resolve
  - hci: keep dst_type with dst when reusing an LE connection
  - hci_core: Fix queuing tx_work after workqueue is drained
  - hci_sync: Serialize local codec list cleanup
  - hci_codec: validate vendor codec count length
  - eir: validate service data length before reading UUID
  - RFCOMM: avoid socket lock inversion in listener cleanup
  - ISO: Fix parent socket leak in iso_conn_ready()
  - ISO: set BT_LISTEN before requesting a BIG sync
  - coredump: Quiesce dump work on unregister
 
 Drivers:
 
  - btintel_pcie: validate TX skb length in send_sync
  - btmtk: fix wrong status for short WMT FUNC_CTRL events
  - btmtksdio, btmtkuart: validate WMT event length before struct access
  - hci_qca: Do not write to the serial port after it is closed
  - btusb: fix NXP IW610 composite device handling
  - btintel_pcie: fix off-by-one bounds check in RX submit
  - btmtksdio: Fix PM runtime reference leak in shutdown
 -----BEGIN PGP SIGNATURE-----
 
 iQJNBAABCgA3FiEE7E6oRXp8w05ovYr/9JCA4xAyCykFAmqpmw8ZHGx1aXoudm9u
 LmRlbnR6QGludGVsLmNvbQAKCRD0kIDjEDILKZC2D/kBM2/ZEFhSwp6Lu7orXKHn
 AQOwoh0uv4TBZaR8V6ogSLdWmdv4rb1Ayjj2SFkNNNwmuUkPYhG30DySw73C7QpV
 uza4SAPO9OF3RRzXvBiVpxF4TmlWYMbApAuMtaa/NmZfZ38pv6T66cFsxhc+lfje
 qNUWkCyI9/4hY7rIxL2dcBQ3yLQwMYwKCwdMwj2/rMX4BN7IXVsMqZ4HBHmCx3Ht
 M2GGnng4DNFgCN/ZcUMLASRR61At3Iop39E8NLnnnVqJgn6p2DOx+G+gnYpdnUOy
 zdpm/yLSJD2maDVRlxVWlf071jpUVCjCY2HvjnvnBTKDgqf9GD/3SVTrv3f/3gxt
 /NVHwnHjZkhMt2sY//NVo4dGD0s8PJjEAdWzm/c7lZ97/8AHsSFMtBe53Jc/xy9l
 lvmTV3FrTkzf08VT5QyrTjj9dQjOcJfeOqfOHJrRQSfjqvNCR7OZjjDLavgv+3q/
 Op5bE0RY/47qa7yWE7gyjyVFeAoK2C4XsXDIc8cHOdJlozlqKV0jVQUvs9+wbl8r
 0ULFdcRAs3/OIybYBW4qhOGjZjS/ilK4CjCNGjEPMD1IEz6SutJPxtp1Cmh9KFZT
 3u5vNsbXB/Oa5yM+m5KZwxoRhYG6aCivCLruEWVe8tco8nWlcz1iPRPBs+bB2cvX
 5tJwh5QCT35FUBo3j+0PvQ==
 =dGOw
 -----END PGP SIGNATURE-----

Merge tag 'for-net-2026-09-15' of git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth

Luiz Augusto von Dentz says:

====================
bluetooth pull request for net:

Core:

 - hci: put the peer's on-air address on air when we cannot resolve
 - hci: keep dst_type with dst when reusing an LE connection
 - hci_core: Fix queuing tx_work after workqueue is drained
 - hci_sync: Serialize local codec list cleanup
 - hci_codec: validate vendor codec count length
 - eir: validate service data length before reading UUID
 - RFCOMM: avoid socket lock inversion in listener cleanup
 - ISO: Fix parent socket leak in iso_conn_ready()
 - ISO: set BT_LISTEN before requesting a BIG sync
 - coredump: Quiesce dump work on unregister

Drivers:

 - btintel_pcie: validate TX skb length in send_sync
 - btmtk: fix wrong status for short WMT FUNC_CTRL events
 - btmtksdio, btmtkuart: validate WMT event length before struct access
 - hci_qca: Do not write to the serial port after it is closed
 - btusb: fix NXP IW610 composite device handling
 - btintel_pcie: fix off-by-one bounds check in RX submit
 - btmtksdio: Fix PM runtime reference leak in shutdown

* tag 'for-net-2026-09-15' of git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth:
  Bluetooth: RFCOMM: avoid socket lock inversion in listener cleanup
  Bluetooth: keep dst_type with dst when reusing an LE connection
  Bluetooth: btintel_pcie: fix off-by-one bounds check in RX submit
  Bluetooth: btmtksdio: Fix PM runtime reference leak in shutdown
  Bluetooth: btmtksdio, btmtkuart: validate WMT event length before struct access
  Bluetooth: btmtk: fix wrong status for short WMT FUNC_CTRL events
  Bluetooth: ISO: set BT_LISTEN before requesting a BIG sync
  Bluetooth: ISO: Fix parent socket leak in iso_conn_ready()
  Bluetooth: hci_sync: Serialize local codec list cleanup
  Bluetooth: hci_qca: Do not write to the serial port after it is closed
  Bluetooth: hci_codec: validate vendor codec count length
  Bluetooth: put the peer's on-air address on air when we cannot resolve
  Bluetooth: coredump: Quiesce dump work on unregister
  Bluetooth: btintel_pcie: validate TX skb length in send_sync
  Bluetooth: hci_core: Fix queuing tx_work after workqueue is drained
  Bluetooth: eir: validate service data length before reading UUID
  Bluetooth: btusb: fix NXP IW610 composite device handling
====================

Link: https://patch.msgid.link/20260915192441.1130583-1-luiz.dentz@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 16:40:55 -07:00
Juan Perdomo
801fb950ca Bluetooth: RFCOMM: avoid socket lock inversion in listener cleanup
rfcomm_sock_cleanup_listen() closes unaccepted child sockets through
rfcomm_sock_close(), which takes the child socket lock before
rfcomm_dlc_close() acquires rfcomm_mutex. The RFCOMM worker takes these
locks in reverse order while handling connections and DLC state changes,
so lockdep reports a possible deadlock.

Close dequeued children without taking their socket lock. The accept queue
owns a reference to each child, and bt_accept_dequeue() locks the child
while unlinking it and clearing its parent pointer.

Dropping the child lock makes it important to prevent a concurrent
rfcomm_connect_ind() from enqueueing a new child after cleanup observes an
empty queue. Set a listening socket to BT_CLOSED while its lock is still
held, before dropping the lock and draining the queue. The state check in
rfcomm_connect_ind() then rejects new children once cleanup starts.

Reported-by: syzbot+0cece8fa7d83523f47a3@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=0cece8fa7d83523f47a3
Fixes: b7ce436a5d ("Bluetooth: switch to lock_sock in RFCOMM")
Signed-off-by: Juan Perdomo <jcperdomo100@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:55:41 -04:00
Radek Podgorny
555cd2bd86 Bluetooth: keep dst_type with dst when reusing an LE connection
hci_connect_le() swaps the caller's identity address for the peer's
cached RPA when one is known, and stamps the matching
ADDR_LE_DEV_RANDOM on the local dst_type. On the conn-reuse path only
the address is copied into the connection:

  if (conn) {
          bacpy(&conn->dst, dst);

so conn->dst ends up holding an RPA while conn->dst_type still names the
identity it was resolved from, and hci_le_create_conn_sync() puts that
pair on air unchanged. An RPA declared as a public address is not
something any peer can answer.

Measured on a CYW43438 against a peer advertising an RPA the host holds
the IRK for, connecting to the identity address over a raw L2CAP socket.
The first attempt creates the connection, the second takes the reuse
path:

  LE Create Connection  3C:78:95:78:37:C3  type public
  LE Create Connection  5B:75:A2:26:D6:18  type public
  LE Connection Complete: Unknown Connection Identifier (0x02)

The second address is the peer's RPA. btmon annotates it with an OUI
lookup rather than "(Resolvable)" precisely because the command declares
it public; the same bit pattern annotates as resolvable once the type is
right.

The mistyped pair is also why nothing downstream repairs it.
hci_bdaddr_is_rpa() tests the type before the address, so an RPA carrying
a public type is not recognised as one, and hci_find_irk_by_addr() then
searches for an identity address that does not match it either.

Copy the type along with the address.

The assignment used to be unconditional just below this block and covered
both paths; it moved into hci_conn_add_unset(), which the reuse path does
not go through.

Cc: stable@vger.kernel.org
Fixes: 14b06c3a88 ("Bluetooth: HCI: Always use the identity address when initializing a connection")
Assisted-by: Claude:claude-opus-5
Signed-off-by: Radek Podgorny <radek@podgorny.cz>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:55:32 -04:00
Sai Teja Aluvala
2ea5a87a5a Bluetooth: btintel_pcie: fix off-by-one bounds check in RX submit
btintel_pcie_submit_rx() used frbd_index > rxq->count to guard the
FRBD array access, allowing frbd_index == rxq->count to pass through
and index one element past the end of the array. Change the check to
>= rxq->count so every out-of-range index is rejected.

This issue was reported by Claude Mythos.

Fixes: c2b636b3f7 (Bluetooth: btintel_pcie: Add support for PCIe transport)
Signed-off-by: Sai Teja Aluvala <aluvala.sai.teja@intel.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:55:27 -04:00
Tzung-Bi Shih
7b60ee5f46 Bluetooth: btmtksdio: Fix PM runtime reference leak in shutdown
In btmtksdio_shutdown(), pm_runtime_get_sync() is called at the
beginning of the function.  However, if sending the WMT function
control command fails later, the driver returns early.

It bypasses the corresponding pm_runtime_put_noidle() and
pm_runtime_disable() calls, leaking the PM usage counter and leaving PM
runtime enabled indefinitely.

Fall through to execute the PM runtime cleanup block even if WMT errors.

Fixes: 7f3c563c57 ("Bluetooth: btmtksdio: Add runtime PM support to SDIO based Bluetooth")
Signed-off-by: Tzung-Bi Shih <tzungbi@kernel.org>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:55:01 -04:00
Chris Lu
8879e3e0a8 Bluetooth: btmtksdio, btmtkuart: validate WMT event length before struct access
btmtksdio.c and btmtkuart.c cast a received WMT event straight to
struct btmtk_hci_wmt_evt and read its op/flag fields without checking
the event is long enough to contain them, unlike btmtk.c. The
FUNC_CTRL case then further casts to struct btmtk_hci_wmt_evt_funcc
and reads its 2-byte status field, again without a length check.
Firmware that sends a short or malformed WMT event makes both drivers
read past the end of the received SKB.

Mirror btmtk.c: validate the base WMT header with skb_pull_data()
before touching any of its fields, and when a FUNC_CTRL event turns
out to be the short, header-only form (a plain enable/disable ack
with no status word), decode the result from the header's own flag
byte instead (0 = success, otherwise failure).

Verified setup on MT7920, MT7921, MT7922 and MT7925: no regression.

Fixes: 9aebfd4a22 ("Bluetooth: mediatek: add support for MediaTek MT7663S and MT7668S SDIO devices")
Fixes: e0b67035a9 ("Bluetooth: mediatek: update the common setup between MT7622 and other devices")
Assisted-by: Claude:claude-opus-5
Signed-off-by: Chris Lu <chris.lu@mediatek.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:54:54 -04:00
Chris Lu
78b6abd6c7 Bluetooth: btmtk: fix wrong status for short WMT FUNC_CTRL events
A too-short BTMTK_WMT_FUNC_CTRL event (WMT header only, no trailing
2-byte status word) is always treated as BTMTK_WMT_ON_UNDONE. This
short form is how firmware acks a plain enable/disable request, and
the actual result is carried in the header's own flag byte (0 =
success), not a separate status word. Decode it from there instead of
assuming failure.

Verified setup on MT7920, MT7921, MT7922 and MT7925: no regression.

Fixes: e3ac0d9f1a ("Bluetooth: btmtk: accept too short WMT FUNC_CTRL events")
Assisted-by: Claude:claude-opus-5
Signed-off-by: Chris Lu <chris.lu@mediatek.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:54:48 -04:00
Luiz Augusto von Dentz
296e7f3c50 Bluetooth: ISO: set BT_LISTEN before requesting a BIG sync
A BIS connection is matched to its parent socket by looking for a
socket in BT_LISTEN state with the same BIG handle:

  iso_conn_ready()
    if (test_bit(HCI_CONN_BIG_SYNC, &hcon->flags))
            parent = iso_get_sock(hdev, &hcon->src, &hcon->dst,
                                  BT_LISTEN, iso_match_big_hcon, hcon);

The socket was only moved to BT_LISTEN after iso_conn_big_sync()
returned, while the LE BIG Create Sync command has already been queued
by then. If the BIG sync is established before the state is updated,
which is easy to hit with an emulated controller as the command may
complete in a few hundred microseconds, no parent is found and the BIS
connections are never notified to the listening socket.

The user space is then left waiting for connections that never arrive,
e.g. bluetoothd never completes a MediaTransport1.Acquire of a
Broadcast Sink transport.

Move the socket to BT_LISTEN before requesting the BIG sync, so the
state is visible by the time the command is queued, and restore the
previous state if the request could not be started. Since the socket is
briefly visible as a listening socket, child sockets may have been
queued in the meantime, so drain the accept queue before restoring the
state: the cleanup paths of BT_CONNECT2/BT_CONNECTED don't do it and the
children would be left with a dangling parent pointer.

Fixes: fbdc4bc472 ("Bluetooth: ISO: Use defer setup to separate PA sync and BIG sync")
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:53:12 -04:00
Luiz Augusto von Dentz
ca18ee413a Bluetooth: ISO: Fix parent socket leak in iso_conn_ready()
iso_get_sock() returns the parent socket with a reference held, which is
dropped by sock_put() once the child socket has been set up. The error
path taken when iso_sock_alloc() fails only calls release_sock() and
returns, leaking the reference and thus the parent socket itself.

Drop the reference on that path as well.

Fixes: fa224d0c09 ("Bluetooth: ISO: Reassociate a socket with an active BIS")
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:53:03 -04:00
Chengfeng Ye
9a10987a2f Bluetooth: hci_sync: Serialize local codec list cleanup
hci_dev_close_sync() clears hdev->local_codecs after releasing hdev->lock.
Codec list additions and both traversals in sco_sock_getsockopt() use that
lock, but the close path does not. A close and BT_CODEC query can therefore
interleave as follows:

  hci_dev_close_sync()          sco_sock_getsockopt()
                                hci_dev_lock()
                                fetch codec entry
  hci_codec_list_clear()
    kfree(entry)
                                read entry->id

The reader then accesses an entry which the close path has freed. KASAN
reported:

  BUG: KASAN: slab-use-after-free in sco_sock_getsockopt+0xfa0/0xfe0
  Read of size 1 at addr ffff8881001c3450
  Call Trace:
   sco_sock_getsockopt+0xfa0/0xfe0
   do_sock_getsockopt+0x537/0x7b0
   __sys_getsockopt+0xf2/0x170
  Allocated by task 92:
   hci_codec_list_add.isra.0+0x2c/0x440
   hci_read_codec_capabilities+0x224/0x590
   hci_read_supported_codecs+0x2c2/0x640
  Freed by task 92:
   kfree+0x131/0x3c0
   hci_codec_list_clear+0xd8/0x160
   hci_dev_close_sync+0x92a/0xfa0

Take hdev->lock around the clear operation at its existing point in the
close path. This makes the clear wait for active readers and prevents a new
traversal until the list is empty without changing teardown ordering.

Fixes: b938790e70 ("Bluetooth: hci_codec: Fix leaking content of local_codecs")
Cc: stable@vger.kernel.org
Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:52:31 -04:00
Ibrahim Abdelkader
4e93c65f87 Bluetooth: hci_qca: Do not write to the serial port after it is closed
hci_uart_close() closes the serdev port if HCI_QUIRK_NON_PERSISTENT_SETUP
is set (for example, for the WCN399x family). A failed hci_dev_open_sync()
following a successful qca_setup() calls hdev->close() but not
hdev->shutdown(), so the port is closed while power->vregs_on is left true.
qca_serdev_remove() then passes its power->vregs_on test and calls
qca_power_off(), which writes to the closed port unconditionally.

Seen on a WCN3988 by unbinding the driver after a controller failure. The
trace below is from a 7.0.0 based kernel, where qca_power_off() was still
named qca_power_shutdown():

  Unable to handle kernel NULL pointer dereference at virtual address
  0000000000000038
  Call trace:
   tty_set_termios+0x50/0x238 (P)
   ttyport_set_baudrate+0x84/0xc0
   serdev_device_set_baudrate+0x24/0x40
   qca_power_shutdown+0x158/0x1fc [hci_uart]
   qca_serdev_remove+0x54/0x68 [hci_uart]
   serdev_drv_remove+0x1c/0x2c
   device_remove+0x4c/0x80
   device_release_driver_internal+0x1cc/0x224
   device_driver_detach+0x18/0x24
   unbind_store+0xb4/0xc0

Check HCI_UART_PROTO_READY, which hci_uart_close() clears in the same place
it closes the port, before writing to it. The regulator disable is left
unconditional so the controller is still powered down.

The dangling serport->tty that turns this into a use-after-free is
addressed in a separate patch.

Fixes: fa9ad876b8 ("Bluetooth: hci_qca: Add support for Qualcomm Bluetooth chip wcn3990")
Signed-off-by: Ibrahim Abdelkader <iabdelka@qti.qualcomm.com>
Reviewed-by: Hans de Goede <johannes.goede@oss.qualcomm.com>
Signed-off-by: Hans de Goede <johannes.goede@oss.qualcomm.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:52:19 -04:00
Laxman Acharya Padhya
d0795cfd6f Bluetooth: hci_codec: validate vendor codec count length
The Read Local Supported Codecs parsers consume the variable-sized
standard codec array before parsing the vendor codec count.  Although the
initial reply-size check includes a vendor count byte in the fixed layout,
it does not guarantee that the byte remains after the standard codec array.

If a controller reply ends immediately after that array, calculating the
vendor codec array size reads vnd_codecs->num beyond the skb data.  Use
skb_pull_data() to validate and consume each codec header before using its
count in both command variants.

Fixes: 8961987f3f ("Bluetooth: Enumerate local supported codec and cache details")
Fixes: 9ae664028a ("Bluetooth: Add support for Read Local Supported Codecs V2")
Cc: stable@vger.kernel.org
Suggested-by: Luiz Augusto von Dentz <luiz.dentz@gmail.com>
Signed-off-by: Laxman Acharya Padhya <acharyalaxman8848@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:52:12 -04:00