Commit Graph

1472353 Commits

Author SHA1 Message Date
Hyunwoo Kim
03a9d10ecf sctp: drop a chunk if its transport was removed
sctp_rcv() resolves the transport once per packet and leaves it in
chunk->transport. The lookup reference, or the one sctp_add_backlog() takes
if the socket is owned by userspace, keeps it around until the chunk has
been processed.

An authenticated ASCONF DEL-IP can remove it in the meantime.
sctp_assoc_rm_peer() takes the transport out of the association and calls
sctp_transport_free(), which tags it dead and drops the reference the
association held. There is a window on both paths: the packet can sit on
the socket backlog, and on the direct path the lookup completes before
bh_lock_sock().

The DATA chunk in that packet puts the removed transport back into
asoc->peer.last_data_from. Once the packet is done that reference goes
away and the transport is freed by RCU, so the next delayed SACK carries
the pointer into the SACK chunk and sctp_outq_select_transport() reads the
freed transport's state.

Drop the chunk in sctp_inq_push(), next to the existing rcvr->dead check.
Both paths reach it with the association's socket lock held. The peer
retransmits it.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Cc: stable@vger.kernel.org
Signed-off-by: Hyunwoo Kim <imv4bel@gmail.com>
Acked-by: Xin Long <lucien.xin@gmail.com>
Link: https://patch.msgid.link/aoUJHQmxL0LFIMCw@v4bel
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:31:09 -07:00
Triet Hoang
25b863cd6d tools: ynl: handle calloc failure in ynl_ntf_parse
Check the return value of calloc() before dereferencing the allocated
response structure in ynl_ntf_parse().

Signed-off-by: Triet Hoang <triet.hoang.dev@gmail.com>
Link: https://patch.msgid.link/20260818132739.469624-1-triet.hoang.dev@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:29:36 -07:00
Jamal Hadi Salim
4c660ee8c8 net: sched: fix 32-bit backlog wrap in gred, bfifo and plug enqueue
gred_enqueue(), bfifo_enqueue() and plug_enqueue() admit a packet when the
current backlog plus the packet length fits within the queue limit:

  sch->qstats.backlog + qdisc_pkt_len(skb) <= sch->limit (gred default VQ)
  gred_backlog+qdisc_pkt_len(skb) <= q->limit  (gred configured VQ)
  sch->qstats.backlog + qdisc_pkt_len(skb) <= sch->limit (bfifo)
  sch->qstats.backlog + skb->len <= q->limit             (plug)

sch->qstats.backlog and q->backlog are u32, and qdisc_pkt_len()/skb->len
are unsigned int, so all sums are computed in 32 bits and wrap at 2^32.
Once the true backlog exceeds 4 GiB the wrapped sum becomes small and
admission keeps succeeding, so the queue grows without bound and the kernel
can be driven to OOM.

Promote the sums to u64 so admission stops once the true backlog exceeds
the limit.  The limit is u32, so the bounded queue stays below 2^32 and
the stored u32 backlog never wraps.

The bug can only be reproduced as root (albeit with ridiculous setup):
 attach a gred (or bfifo/plug) qdisc with a limit near 4 GiB,
 leaving the default VQ unconfigured (for gred), and drive >4 GiB of
 queued traffic (e.g. via a size table / stab to inflate qdisc_pkt_len,
 or sustained high-rate traffic). The u32 backlog+len sum wraps at 2^32,
 admission keeps succeeding, and the queue grows unboundedly to OOM.

Fixes: a3eb95f891 ("net_sched: gred: add TCA_GRED_LIMIT attribute")
Reported-by: vega@nebusec.ai
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260818095927.15901-1-jhs@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:28:45 -07:00
Mina Almasry
d9c56501c7 net: tcp: block mixing readable and unreadable frags
Protect tcp_sendmsg_locked() from mistakenly mixing readable and
unreadable page fragments in the same SKB.

Check that the devmem binding matches the existing SKB's readability.
If a mismatch is detected, avoid collapsing and create a new segment.

Fixes: bd61848900 ("net: devmem: Implement TX path")
Suggested-by: Eric Dumazet <edumazet@google.com>
Cc: Pavel Begunkov <asml.silence@gmail.com>
Cc: Stanislav Fomichev <sdf@fomichev.me>
Cc: Bobby Eshleman <bobbyeshleman@gmail.com>
Signed-off-by: Mina Almasry <almasrymina@google.com>
Link: https://patch.msgid.link/20260814191336.187243-2-almasrymina@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:22:14 -07:00
Mina Almasry
68d8c65326 net: core: propagate unreadable flag in skb_zerocopy
skb_zerocopy() fails to propagate the unreadable flag when copying
unreadable fragments, causing target skbs to appear as readable memory.

This patch fixes the flag propagation. Additionally, it returns -EFAULT
if readable fragments are mixed with unreadable fragments during
extraction, and returns -EFAULT in openvswitch queue_userspace_packet().

Fixes: 65249feb6b ("net: add support for skbs with unreadable frags")
Cc: Stanislav Fomichev <sdf@fomichev.me>
Cc: Bobby Eshleman <bobbyeshleman@gmail.com>
Cc: Florian Westphal <fw@strlen.de>
Cc: Aaron Conole <aconole@redhat.com>
Cc: Eelco Chaudron <echaudro@redhat.com>
Cc: Willem de Bruijn <willemb@google.com>
Signed-off-by: Mina Almasry <almasrymina@google.com>
Reviewed-by: Pavel Begunkov <asml.silence@gmail.com>
Reviewed-by: Ilya Maximets <i.maximets@ovn.org>
Link: https://patch.msgid.link/20260814191336.187243-1-almasrymina@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:22:14 -07:00
Kyle Zeng
992cc9f94c net/packet: defer vmalloc TX_RING free until skbs finish
AF_PACKET TX_RING skbs keep a raw pointer to their ring frame. The skb
page references preserve page-backed ring blocks after pg_vec is freed,
but they do not preserve a vmalloc mapping.

tpacket_destruct_skb() currently drops the pending reference before
writing the timestamp and TP_STATUS_AVAILABLE to the frame. Move the
decrement after those stores. The smp_wmb() in __packet_set_status()
orders the frame stores before the decrement.

Also recheck pending TX frames under pg_vec_lock before non-closing
ring replacement, so a racing send cannot add a pending skb between
the initial check and the ring swap.

Ring allocation can produce a mixture of page-backed and vmalloc-backed
blocks. Allocate deferred-work storage during TX ring setup when the
first vmalloc-backed block is encountered, and keep its pointer in the
pg_vec allocation header. If allocation fails, return -ENOMEM from ring
setup. On socket close, a non-NULL pointer identifies a vmalloc-backed
vector without a scan. If TX skbs remain, defer the whole vector to
system_long_wq.

After pg_vec is detached, a late destructor can skip the pending
decrement. Use socket write-memory accounting as the deferred lifetime
gate instead: an skb remains charged through its final sock_wfree(),
after all ring-frame accesses. The delayed work retains a socket
reference and reschedules itself until no TX skbs remain.

Move pending_refcnt release to packet_sock_destruct() so late skb
destructors and deferred cleanup can safely use it after
packet_release(). Page-backed teardown remains synchronous, and no lock
is added to the TX completion hot path.

Fixes: b013840810 ("packet: use percpu mmap tx frame pending refcount")
Cc: stable@vger.kernel.org
Link: https://lore.kernel.org/netdev/20260721015824.45829-1-kylebot@openai.com/
Suggested-by: Eric Dumazet <edumazet@google.com>
Suggested-by: Willem de Bruijn <willemdebruijn.kernel@gmail.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Signed-off-by: Kyle Zeng <kylebot@openai.com>
Link: https://patch.msgid.link/20260816235646.76500-1-kylebot@openai.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:18:46 -07:00
Jakub Kicinski
51a26bb843 Merge branch 'ionic_rcq_shared' of https://github.com/abhijitG-xlnx/linux
Abhijit Gangurde says:

====================
Extend the net/ionic firmware identity structure to expose
the rcq_sign_bit field from the RDMA LIF identity.

* 'ionic_rcq_shared' of https://github.com/abhijitG-xlnx/linux:
  net: ionic: Fetch RCQ sign bit from firmware
====================

Link: https://patch.msgid.link/
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:07:01 -07:00
Eric Dumazet
447cbe95eb vlan: fix skb_under_panic and races when toggling HW VLAN offload
Toggling hardware VLAN TX offload (NETIF_F_HW_VLAN_CTAG_TX or
NETIF_F_HW_VLAN_STAG_TX) on a lower device invokes vlan_transfer_features(),
which dynamically changed vlandev->hard_header_len.

This causes two issues:
1. Lockless TX paths (e.g. packet_snd in af_packet.c, ip6_finish_output2)
   read dev->hard_header_len without holding RTNL lock. Mutating
   hard_header_len dynamically under RTNL creates a data race where upper
   layers reserve insufficient headroom based on a stale hard_header_len,
   resulting in skb_under_panic when vlan_dev_hard_header() is called.
2. In addition, vlan_transfer_features() updated hard_header_len without
   updating header_ops, causing a mismatch between allocated headroom
   and header creation.

Always setting dev->hard_header_len = real_dev->hard_header_len and
dev->needed_headroom = real_dev->needed_headroom + VLAN_HLEN unconditionally
ensures:
- dev->hard_header_len remains 100% static and immutable at real_dev->hard_header_len,
  eliminating all dynamic runtime updates and data races on hard_header_len.
- Upper layers allocating skbs via LL_RESERVED_SPACE() will always reserve
  sufficient headroom for software VLAN tag insertion (real_dev->hard_header_len +
  real_dev->needed_headroom + VLAN_HLEN).
- vlandev inherits real_dev->needed_tailroom so underlying trailer/padding/ICV
  requirements are honored.
- AF_PACKET SOCK_RAW network header offsets remain correctly aligned at
  real_dev->hard_header_len.
- vlan_header_ops is used unconditionally.

Note to stable teams: Make sure to backport these commits:

e16e960d55 ("ipvlan: inherit needed_headroom and needed_tailroom from phy_dev")
cef51860be ("macvlan: inherit needed_headroom and needed_tailroom from lowerdev")

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Reported-by: Tangxin Xie <xietangxin@h-partners.com>
Closes: https://lore.kernel.org/netdev/99d678ae-c7b2-4b44-b534-b8320679deb3@h-partners.com/
Cc: <stable@vger.kernel.org> # 3.19: e16e960d55a4: ipvlan: inherit needed_headroom and needed_tailroom from phy_dev
Cc: <stable@vger.kernel.org> # 3.19: cef51860becd: macvlan: inherit needed_headroom and needed_tailroom from lowerdev
Cc: <stable@vger.kernel.org> # 3.19
Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260811085246.2267779-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:05:43 -07:00
Zhan Xusheng
d302d7109f ethtool: remove unused __ETHTOOL_LINK_MODE_MASK_NWORDS
From: Zhan Xusheng <zhanxusheng@xiaomi.com>

Added by commit f625aa9be8 ("ethtool: provide link mode information with
LINKMODES_GET request") and never used.  The same count is computed as
__ETHTOOL_LINK_MODE_MASK_NU32 in net/ethtool/ioctl.c.

Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Link: https://patch.msgid.link/20260818023704.125721-1-zhanxusheng@xiaomi.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:04:44 -07:00
Kyle Zeng
f12c2de4f5 batman-adv: reject unrepresentable multicast TVLV offsets
The network and transport header fields in struct sk_buff are 16-bit
offsets from skb->head, and U16_MAX is reserved as the unset transport
header value. batadv_tvlv_call_handler() sets both fields from a received
multicast TVLV without checking whether the TVLV end is representable.

If the end offset exceeds the field's range, skb_set_transport_header()
truncates it so that the transport header precedes the network header.
The negative difference is then returned by skb_network_header_len() as
a large u32. batadv_mcast_forw_packet() consequently accepts an oversized
multicast tracker and accesses memory beyond the skb data.

Add skb_set_transport_header_careful(), an offset-aware counterpart to
skb_reset_transport_header_careful(), which validates the final
head-relative offset before assigning it. Use the new helper in
batadv_tvlv_call_handler() and reject unrepresentable TVLVs before
setting the network header.

Fixes: 07afe1ba28 ("batman-adv: mcast: implement multicast packet reception and forwarding")
Cc: stable@vger.kernel.org
Signed-off-by: Kyle Zeng <kylebot@openai.com>
Co-developed-by: David Lee <david.lee@trailofbits.com>
Signed-off-by: David Lee <david.lee@trailofbits.com>
Acked-by: Sven Eckelmann <sven@narfation.org>
Link: https://patch.msgid.link/20260817084955.944189-1-david.lee@trailofbits.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:58:33 -07:00
Kyle Zeng
44930446dd ipv6: seg6: clear IPv4 control block on IPIP decapsulation
End.DX4 and End.DT4 decapsulate an IPv4 packet through
decap_and_validate() and send it directly to IPv4 routing. The inner
packet therefore bypasses ip_rcv_core(), which normally clears IPCB
before IPv4 interprets skb->cb.

The skb instead retains IP6CB data from the outer packet. IP6CB and
IPCB use the same skb->cb storage, so IP6CB(skb)->lastopt overlaps
IPCB(skb)->opt.optlen and srr, while IP6CB(skb)->nhoff overlaps rr and
ts.

The sender can make the stale optlen byte nonzero with a valid outer
extension-header chain. The reproducers put an eight-byte Destination
Options header immediately after the 40-byte IPv6 header and before the
Segment Routing Header. ipv6_destopt_rcv() records the sender-controlled
Destination Options offset in both lastopt and nhoff, setting them to
40. On the reproduced little-endian x86-64 kernel, IPv4 therefore sees
optlen = 40 and rr = 40.

Both tcp_v4_save_options() and __ip_options_echo() skip option copying
when optlen is zero. Here optlen is 40, so the TCP SYN path allocates
room for 40 bytes of option data and calls __ip_options_echo(). The
stale rr value makes that function read inner packet byte 41 as the
Record Route option length. The reproducers set that sender-controlled
byte to 255, so __ip_options_echo() copies 255 bytes into the 40-byte
option-data area.

Separate End.DX4 and End.DT4 reproducers on the unpatched v7.2-rc5
kernel both produced:

  BUG: KASAN: slab-out-of-bounds in __ip_options_echo()
  Write of size 255

The relevant End.DX4 call path is:

  __ip_options_echo
  tcp_v4_route_req
  tcp_conn_request
  tcp_v4_conn_request
  tcp_rcv_state_process
  tcp_v4_do_rcv
  tcp_v4_rcv
  ip_protocol_deliver_rcu
  ip_local_deliver_finish
  ip_local_deliver
  input_action_end_dx4_finish
  input_action_end_dx4

The relevant End.DT4 call path is:

  __ip_options_echo
  tcp_v4_route_req
  tcp_conn_request
  tcp_v4_conn_request
  tcp_rcv_state_process
  tcp_v4_do_rcv
  tcp_v4_rcv
  ip_protocol_deliver_rcu
  ip_local_deliver_finish
  ip_local_deliver
  input_action_end_dt4

tcp_v4_save_options() is inlined into the tcp_v4_route_req() path, so
it does not appear as a separate frame.

When decap_and_validate() handles IPPROTO_IPIP, save the ingress
interface from IP6CB, clear IPCB, and restore the saved value. Doing
this in the common decapsulation path covers End.DX4, End.DT4, and
End.DT46's IPv4 arm.

Use IP6CB(skb)->iif rather than skb->skb_iif. These actions run after
l3mdev processing, which can replace skb_iif with the L3 master;
IP6CB iif still records the receiving interface set at IPv6 ingress.

Fixes: 891ef8dd2a ("ipv6: sr: implement additional seg6local actions")
Cc: stable@vger.kernel.org
Suggested-by: Andrea Mayer <andrea.mayer@uniroma2.it>
Signed-off-by: Kyle Zeng <kylebot@openai.com>
Co-developed-by: David Lee <david.lee@trailofbits.com>
Signed-off-by: David Lee <david.lee@trailofbits.com>
Reviewed-by: Andrea Mayer <andrea.mayer@uniroma2.it>
Link: https://patch.msgid.link/20260817085839.946321-1-david.lee@trailofbits.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:57:24 -07:00
Eric Dumazet
2ee66e9487 inetpeer: randomize RB-tree node comparison using SipHash
The inetpeer rate limiting system stores peer entries in a Red-Black tree
keyed deterministically on the remote IP address. Because tree lookups walk
the RB-tree using standard lexicographical comparisons (inetpeer_addr_cmp),
an off-path adversary can predict the exact topology of the tree and the
sequence of nodes traversed during lookups (the gc_stack candidate list).

By combining deterministic tree traversal with aggressive garbage collection
(triggered when tree size exceeds inet_peer_threshold), an attacker can
selectively force the eviction of targeted inet_peer nodes. When an evicted
node is subsequently re-created upon receiving a new packet, its rate-limiting
token bucket (rate_tokens, rate_last) is reset to full capacity. This creates
a side-channel primitive allowing off-path attackers to bypass IP-keyed ICMP
rate limits and infer open UDP ports (similar to SAD DNS style attacks).

Mitigate this by randomizing the RB-tree node comparison logic using SipHash
with a secret key (inetpeer_hash_key) initialized via net_get_random_once().
Nodes are ordered in the tree by SipHash(addr, key) rather than raw IP
addresses. Because the secret key is unknown to external entities, the tree
layout and lookup traversal paths are unpredictable to off-path adversaries,
breaking the deterministic eviction gadget.

Cache the computed 64-bit SipHash (hash) in struct inet_peer and compute the
target hash (dhash) once at the beginning of inet_getpeer() to avoid recomputing
SipHash at every step of the RB-tree walk.

Fixes: b145425f26 ("inetpeer: remove AVL implementation in favor of RB tree")
Reported-by: Michael Blunt <michaelbblunt@gmail.com>
Suggested-by: Michael Blunt <michaelbblunt@gmail.com>
Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260818151213.3953963-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:55:31 -07:00
Eric Dumazet
235b42b586 ip6mr: do not clone dst in ip6mr_cache_report()
IPv6 input attaches a non-refcounted (NOREF) dst to skbs under RCU.
When an ingress multicast packet misses MFC lookup,
ip6mr_cache_unresolved() places the skb onto the unresolved queue,
escaping the receive-side RCU grace period.

If the underlying route is deleted and freed, and the MFC queue is later
resolved with a wrong parent interface, ip6_mr_forward() invokes
ip6mr_cache_report(..., MRT6MSG_WRONGMIF), which executes
dst_clone(skb_dst(pkt)) on the freed dst entry, triggering a slab
use-after-free.

Report packets queued to mroute6_sk (a raw socket) and netlink
notifications do not require an attached dst entry.

Fix this by:
1. Removing dst_clone() in ip6mr_cache_report() and ensuring report skbs
   do not hold a dst.
2. Dropping skb_dst before queuing unresolved skbs in
   ip6mr_cache_unresolved(), matching the fact that multicast
   forwarding resolves outgoing routes anew via ip6_route_output().

Fixes: 67f415dd29 ("ipv6: convert rx data path to not take refcnt on dst")
Reported-by: Zero Day Initiative <zdi-disclosures@trendmicro.com>
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Link: https://patch.msgid.link/20260818172755.4083692-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:54:29 -07:00
Harshit Varu
b878dfdd12 mptcp: fix uninitialized local_id in syncookie MP_JOIN reconstruction
mptcp_token_join_cookie_init_state() restores remote_nonce, local_nonce,
backup, join_id, token and msk from the saved cookie entry when rebuilding
the request socket for a MP_JOIN 4th-ACK handled under SYN cookies, but it
does not restore local_id, even though the SYN path saved it.
subflow_ulp_clone() then reads that uninitialized field and stores it as
the joined subflow's address-ID. Because the request-sock slab is
SLAB_TYPESAFE_BY_RCU and not zeroed on allocation, the value is the stale
byte of a previously freed request socket, which an off-path peer can
influence by sending concurrent MP_JOIN SYNs. This corrupts the path
manager's id-based subflow bookkeeping for the connection.

Restore subflow_req->local_id from the cookie entry, as done for the other
fields.

Fixes: 9466a1cceb ("mptcp: enable JOIN requests even if cookies are in use")
Cc: stable@vger.kernel.org
Signed-off-by: Harshit Varu <harshitvaru666@gmail.com>
Reviewed-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Link: https://patch.msgid.link/20260815115205.197151-1-harshitvaru666@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:53:14 -07:00
Pengpeng Hou
5e8076e4e4 net: qlcnic: validate unified ROM sections before loading
The unified ROM parser reads directory, product, and data-descriptor fields
from the firmware file.  Existing validation forms table and data ends with
unchecked additions and multiplications.  Malformed values can wrap before
they are compared with the firmware size.  The parser also dereferences
typed pointers at firmware-controlled offsets.

Valid descriptor extents alone are insufficient for the consumers.  The
loader reads a fixed-size bootloader regardless of its declared size, the
version parser assumes a 17-byte tail, and a partial final firmware word is
read as a full u64.  A truncated image can therefore make the driver read
beyond the firmware allocation during validation or loading.

Replace the pointer-returning parser with bounded range helpers.  Validate
table entry sizes, descriptor indices, section ranges, the fixed
bootloader load length, and the version tail before exposing any section.
Read all file fields with unaligned little-endian accessors and assemble a
partial final word from only the bytes that remain.  Apply the same range
checks to the legacy image before reading its fixed fields.

Fixes: af19b49152 ("qlcnic: Qlogic ethernet driver for CNA devices")
Signed-off-by: Pengpeng Hou <pengpeng@iscas.ac.cn>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260816052109.4607-1-pengpeng@iscas.ac.cn
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:52:11 -07:00
Tetsuo Handa
f85dc137aa net: add missing ref_tracker_dir_exit() to net_passive_dec()
I found that trying to read /sys/kernel/debug/ref_tracker/* causes NULL
pointer dereference crash when alloc_netdev_mqs() via unshare() returned
NULL, for commit 9ba74e6c9e ("net: add networking namespace refcount
tracker") added ref_tracker_dir_exit(&net->refcnt_tracker) to only
__put_net() path whereas commit 65b584f536 ("ref_tracker: automatically
register a file in debugfs for a ref_tracker_dir") added
ref_tracker_dir_debugfs() to ref_tracker_dir_init() path.

Since preinit_net() calls ref_tracker_dir_init(&net->refcnt_tracker) and
ref_tracker_dir_init(&net->notrefcnt_tracker), we need to make sure that
both ref_tracker_dir_exit(&net->refcnt_tracker) and
ref_tracker_dir_exit(&net->notrefcnt_tracker) are called before
net_passive_dec() schedules for kmem_cache_free() via net_complete_free().

ref_tracker_dir_exit(&net->refcnt_tracker) is called via put_net() when
ns_ref_put() returned true. But put_net() is not called when copy_net_ns()
fails. Therefore, call ref_tracker_dir_exit() from net_passive_dec() if
put_net() is not yet called.

Link: https://sashiko.dev/#/patchset/b06ce35d-e7bc-47a5-8e0a-e82be7e4dd08%40I-love.SAKURA.ne.jp
Fixes: 9ba74e6c9e ("net: add networking namespace refcount tracker")
Reviewed-by: Eric Dumazet <edumazet@google.com>
Signed-off-by: Tetsuo Handa <penguin-kernel@I-love.SAKURA.ne.jp>
Link: https://patch.msgid.link/64254d80-9248-466c-8108-95f43bd71117@I-love.SAKURA.ne.jp
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:44:12 -07:00
Jiangshan Yi
d2796ffe38 bnx2x: fix double free in bnx2x_init_firmware() error path
bnx2x_init_firmware() frees bp->init_ops, bp->init_data and
bp->init_ops_offsets in its error path without setting them to NULL.
The cleanup function bnx2x_release_firmware() frees the same three
pointers unconditionally, so if init_firmware fails and
release_firmware is later called (e.g. from __bnx2x_remove or through
the function state machine), all three are freed a second time.

Set each pointer to NULL after kfree() in the error path so that the
subsequent kfree(NULL) in bnx2x_release_firmware() is a safe no-op.

Fixes: 94a78b79cb ("bnx2x: Separated FW from the source.")
Cc: stable@vger.kernel.org
Signed-off-by: Jiangshan Yi <yijiangshan@kylinos.cn>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260815122149.951215-1-yijiangshan@kylinos.cn
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:39:53 -07:00
Cen Zhang (Microsoft)
d2c26c2911 ipv6: avoid divide by zero in rt6_multipath_rebalance
rt6_multipath_rebalance() calculates the total eligible nexthop weight
in one pass and programs upper bounds in a second pass. Since
RTM_NEWROUTE is RTNL-free, a concurrent
ignore_routes_with_linkdown update can make the first pass return zero
while the second sees an eligible nexthop, causing
rt6_upper_bound_set() to divide by zero.

UBSAN: division-overflow in net/ipv6/route.c:4845:17
Oops: divide error: 0000 [#1] SMP KASAN NOPTI
  rt6_upper_bound_set() net/ipv6/route.c:4845
  rt6_multipath_rebalance()
  fib6_add_rt2node()
  ip6_route_multipath_add()
  inet6_rtm_newroute()

Skip upper-bound calculation when the first pass reports a zero total.
This respects the lock-free performance considerations here and solves
insecure scenarios.

Fixes: bd11ff421d ("ipv6: Get rid of RTNL for SIOCDELRT and RTM_DELROUTE.")
Reported-by: AutonomousCodeSecurity@microsoft.com
Reported-by: Xiang Mei (Microsoft) <xmei5@asu.edu>
Reported-by: Cen Zhang (Microsoft) <blbllhy@gmail.com>
Signed-off-by: Cen Zhang (Microsoft) <blbllhy@gmail.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260817013237.2797-1-blbllhy@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:38:59 -07:00
Eric Dumazet
07e98a4d5e netdevsim: update queue NAPI association on queue reset
In netdevsim, receive queues (struct nsim_rq) embed their own struct
napi_struct. When queue reset is performed (e.g. via queue_reset
debugfs), nsim_queue_start() swaps in a newly allocated struct nsim_rq,
and nsim_queue_mem_free() later deletes and frees the old one.

However, nsim_queue_start() failed to update the queue-to-NAPI mapping
via netif_queue_set_napi(). As a result, dev->_rx[idx].napi continued to
point to the old NAPI struct. After the old queue was freed, a subsequent
queue dump via Netlink (NETDEV_CMD_QUEUE_GET) triggered a KASAN
slab-use-after-free read in nla_put_napi_id() when accessing
rxq->napi->napi_id.

Fix this by calling netif_queue_set_napi() in nsim_queue_start() to
associate the new NAPI with the RX queue, and clear the association
with netif_queue_set_napi(..., NULL) in nsim_del_napi() during teardown.

Fixes: 5bc8e8dbef ("netdevsim: add queue management API support")
Reported-by: syzbot+483a6efbc4882c1201ee@syzkaller.appspotmail.com
Closes: https://lore.kernel.org/netdev/6a82c3d4.f7a79266.2f965f.0024.GAE@google.com/T/#u
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Vadim Fedorenko <vadim.fedorenko@linux.dev>
Link: https://patch.msgid.link/20260817082511.2300402-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:35:57 -07:00
Aldo Ariel Panzardo
408da1df18 net: mctp: hold a reference to the route device in mctp_route_lookup()
mctp_route_lookup() uses rt->dev without holding a reference on it.
mctp_route_lookup_single() returns the route under RCU only, so the
route's device can be torn down concurrently: mctp_dev_put() drops the
last reference and synchronously kfree()s mdev->addrs.  mctp_dev_saddr()
then reads rt->dev->addrs[0], giving a use-after-free reachable by an
unprivileged local AF_MCTP user on the receive/forwarding path (no
CAP_NET_RAW required):

  BUG: KASAN: slab-use-after-free in mctp_route_lookup
  Read of size 1 at addr ... by task mctp_uaf/...
   mctp_route_lookup
   mctp_pkttype_receive
  Freed by task ...:
   kfree
   mctp_dev_put
   mctp_dev_notify

In the same window mctp_dst_from_route() -> mctp_dev_hold() also
increments a refcount that has already reached zero
("refcount_t: addition on 0 ... mctp_dev_hold").

This reintroduces the use-after-free class of CVE-2023-3439: the source
address lookup was moved ahead of the point where the destination takes
its device reference.

Take a reference with refcount_inc_not_zero() before touching rt->dev,
skip a device that is already dead, and drop the reference once the
destination has taken its own.

Fixes: 22cb45afd2 ("net: mctp: perform source address lookups when we populate our dst")
Cc: stable@vger.kernel.org
Signed-off-by: Aldo Ariel Panzardo <qwe.aldo@gmail.com>
Link: https://patch.msgid.link/20260813022102.2792032-1-qwe.aldo@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:27:40 -07:00
Ruoyu Wang
6b9eaa61ff net: ipa: balance runtime PM reference on remove error
ipa_remove() takes a runtime PM reference before accessing IPA hardware
during teardown. If a concurrent modem start or stop keeps
ipa_modem_stop() busy across both attempts, the callback intentionally
returns without releasing the remaining resources because proceeding
with teardown could crash. That return also skips the matching
pm_runtime_put_noidle(), leaving the callback's usage-count reference
held.

Drop only this runtime PM reference before returning.
pm_runtime_put_noidle() does not request an idle transition, so the
hardware and resources retained on this exceptional path remain
untouched while the usage count stays balanced.

This issue was found by a static analysis checker and confirmed by
manual source review.

Fixes: 923a6b6984 ("net: ipa: get clock in ipa_probe()")
Signed-off-by: Ruoyu Wang <ruoyuw560@gmail.com>
Reviewed-by: Alex Elder <elder@riscstar.com>
Link: https://patch.msgid.link/20260815151737.3758320-1-ruoyuw560@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:26:24 -07:00
Jakub Kicinski
a4d7fc3255 Merge branch 'forcedeth-two-register-window-bounds-fixes'
Marek Czernohous says:

====================
forcedeth: two register-window bounds fixes

Two bounds fixes in forcedeth, both in the same shape: a loop that walks
the register window one step too far. They are independent of each other
and touch different functions.

1/2 nv_suspend() and nv_resume() save and restore the non-PCI config
    space with i <= register_size/sizeof(u32). On a VER3 device that is
    exactly the length of saved_config_space[], so the last iteration
    reads and writes one element past the array, and on resume it
    writel()s that element one dword past the length the driver mapped.
    UBSAN catches it.

2/2 nv_tx_timeout() dumps the window in rows of eight dwords but only
    bounds the row's starting offset, so the final row reads between 12
    and 28 bytes past register_size, on every one of the three supported
    window sizes.

Neither is a regression. Both are long standing, and 1/2 in particular
is not new to the list:

  - The identical off-by-one in nv_get_regs() was fixed by commit
    ba9aa13428 ("forcedeth: fix buffer overflow") in 2012. The two
    loops in this patch were missed at the time.
  - The suspend and resume side was then reported on LKML in September
    2013 by Marc Weber, with the same analysis and the same
    one-character fix. Sergei Shtylyov replied asking for the patch
    inline rather than attached, and the thread ended there.

So this is not a new discovery. It is the same bug at the two sites the
2012 fix did not reach, finally sent in the form the list asks for.

How bad is it, stated plainly

  1/2 writes one u32 past the end of a declared array, on a suspend
  path, on every suspend of a VER3 device. That is an out-of-bounds
  store, it is what UBSAN reports, and with CONFIG_UBSAN_TRAP=y it is a
  trap that aborts the running kernel code. That is the stable case, and
  I think it stands on its own: memory safety, reproduced on hardware,
  one character to fix, no behavioural change for anyone else.

  What I will not claim is drama beyond that. The element it lands in is
  np->name_rx, a scratch string that nv_request_irq() rewrites with
  sprintf() before it is ever used, so on a kernel without UBSAN_TRAP
  nothing observable is corrupted. The patch says which member and why,
  so you can judge the severity yourself instead of taking my word.

  The MMIO side of both patches is milder still. ioremap() rounds the
  requested length up to page granularity, so these accesses stay inside
  the page the CPU has mapped and no fault is expected on any
  architecture with PAGE_SIZE >= 4K. What they leave is the window the
  driver asked for. 2/2 is only that, and carries no stable tag.

Behaviour change in 2/2, so it is not buried in the patch

  The partial trailing row of the debug dump is no longer printed: 16
  bytes for VER1, 20 for VER2, 4 for VER3. That is a deliberate trade
  against open-coding a second, narrower dump in a debug-only path. If
  you would rather keep those registers, a short remainder loop on top
  is the obvious follow-up.

Testing

  Reference hardware: Apple Macmini3,1 (MCP79 chipset), forcedeth
  driving the onboard NIC.

  1/2 is reproduced and fixed on that machine. One point of method
  first: UBSAN reports each source location only once per module load,
  so a quiet second suspend proves nothing. Both runs below are the
  first S3 cycle after a fresh load of the module in question.

    stock module,   first S3 after load:  2 splats, one per loop
    patched module, first S3 after load:  none

  The patched module was built, stripped, installed and reloaded, with
  the md5 of the running module checked against the installed one. The
  link came back, the DHCP lease was restored and ping showed no loss.
  That measurement was taken on 2026-08-04 on a 7.1.6 based kernel. The
  stock half has since been reproduced again on 7.1.8, most recently on
  2026-08-13, reporting line 6225 from pci_pm_suspend and line 6240
  from pci_pm_resume.

  I have not repeated the patched half on net/main itself. The runtime
  measurements come from a distro kernel on the reference hardware,
  which is the only machine I have with this NIC; the series itself is
  based on and built against net/main.

  2/2 has no runtime test. Its path sits behind the debug_tx_timeout
  module parameter and needs a genuine TX timeout, which I cannot force
  safely on this machine. It rests on the arithmetic in the patch and
  on the build below.

  Build: allmodconfig with W=1 on x86_64, whole tree, zero compiler
  warnings and zero errors; forcedeth.c specifically produces none.
  That took about 30 hours on the two cores I have, which is why I say
  it plainly rather than in passing.

  I have not run allyesconfig. If you want that too, say so and I will
  queue it before reposting rather than claim a build I did not do.

Two checkpatch notes on 1/2, both deliberate

  "Prefer a maximum 75 chars per line" fires on a line that is quoted
  UBSAN output. The splat is trimmed, and 1/2 says what was cut, but I
  did not rewrap the lines that remain: reflowing diagnostic output to
  satisfy a heuristic makes it harder to match against a real log.

  Two "spaces preferred around that '/'" CHECKs fire on
  register_size/sizeof(u32). That spacing is what the file already uses,
  including in nv_get_regs(), which is otherwise the same loop. Adding
  spaces would leave the two lines I touch inconsistent with their
  neighbourhood, so I kept the change to the one character that is
  wrong. Happy to do it the other way round if you prefer.

AI assistance

  Per Documentation/process/coding-assistants.rst: this work is AI
  assisted. I use Claude (claude-opus-5) as a coding and analysis
  assistant. Both patches carry an Assisted-by trailer accordingly, and
  no Signed-off-by is added by the tool.

  Nature of the assistance: the assistant did the code archaeology and
  most of the drafting. I described the symptom, asked for the mechanism
  to be traced in the source rather than guessed, and asked for each
  claim to be backed by a file and a line. The UBSAN output and the S3
  measurements are from the machine, not model output.

  It is also what found the 2012 fix and the 2013 report above, on a
  second pass over an earlier draft of this posting that claimed the bug
  had never been reported. That claim was wrong and would have wasted
  your time, so it seems worth saying that the checking pass is part of
  the process here and not a flourish. I reviewed the result, I
  understand the code, and I take responsibility for it.
====================

Link: https://patch.msgid.link/178682367884.3748309.5288746298966501007@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:25:00 -07:00
Marek Czernohous
cfa9178ce2 forcedeth: stop the tx_timeout register dump past the requested window
nv_tx_timeout() dumps the register window in rows of eight dwords:

	for (i = 0; i <= np->register_size; i += 32) {
		netdev_info(dev, "%3x: %08x ... %08x\n", i,
			    readl(base + i + 0), ..., readl(base + i + 28));

The loop bound only checks the row's starting offset, so the final row
reads a full 32 bytes from a position that is below the end of the window
but too close to it. base is mapped with exactly that length:

	np->base = ioremap(addr, np->register_size);

so the tail of that row is read from beyond the length the driver asked
for. Per variant, the last iteration reads past register_size by:

	NV_PCI_REGSZ_VER1 (0x270): row 0x260 reads to 0x27f, 16 bytes over
	NV_PCI_REGSZ_VER2 (0x2d4): row 0x2c0 reads to 0x2df, 12 bytes over
	NV_PCI_REGSZ_VER3 (0x604): row 0x600 reads to 0x61f, 28 bytes over

This happens on every supported device, not just one of them. Note that
it is not a consequence of the sizes being odd: with i <= register_size
the offending row is reached whatever the size, and a size that were a
multiple of 32 would overrun by a full row rather than by a remainder.

To be precise about the severity: the reads stay inside the BAR. Memory
BAR sizes are powers of two, the driver only accepts a region with
pci_resource_len() >= register_size (forcedeth.c:5757-5762), and the
next power of two at or above each register_size already covers the
offending row: 0x400 for 0x270 and 0x2d4, 0x800 for 0x604. ioremap()
also rounds the mapped length up to page granularity, so the reads land
inside the mapping the CPU has as well. What they leave is the window
the driver asked for, not the BAR and not the mapping. That is still a
driver reading registers it did not ask for, and it is trivial to
avoid, but nobody should expect a fault from it.

Changing <= to < is not enough: register_size is a length and every size
above is larger than its last row start, so i still reaches the offending
row. Check that the whole row fits instead.

The trade-off is that a partial trailing row is no longer dumped: 16 bytes
for VER1, 20 for VER2, 4 for VER3. That seemed preferable to reading
outside the requested window, and to open-coding a second, narrower dump
for the remainder in what is a debug-only path. Extending the dump to
cover the tail can be done on top if anyone misses those registers.

Only reachable with the debug_tx_timeout module parameter, which defaults
to false. It has not been observed at runtime: forcing a genuine TX
timeout on the reference machine is not something I can do safely, so this
rests on the arithmetic above and on a build test, not on a reproduction.
UBSAN does not catch it either, since these are MMIO reads rather than an
array access. It was found by reading the function while fixing the
saved_config_space off-by-one in nv_suspend() and nv_resume().

The dump was introduced with a fixed 0x400 bound while ioremap() mapped
only NV_PCI_REGSZ (0x270), so it read about 0x190 bytes too far from the
start. Commit 86a0f04387 ("[PATCH] forcedeth: fix initialization")
later replaced 0x400 with np->register_size, which shrank the overrun to
the remainder but did not remove it.

Fixes: c2dba06dae ("[PATCH] forcedeth: rewritten tx irq handling")
Signed-off-by: Marek Czernohous <marek@czernohous.de>
Reviewed-by: Simon Horman <horms@kernel.org>
Reviewed-by: Zhu Yanjun <yanjun.zhu@linux.dev>
Link: https://patch.msgid.link/178682367886.3748309.6978554332066826294@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:24:44 -07:00
Marek Czernohous
9393f1d656 forcedeth: fix off-by-one when saving/restoring non-PCI config space
nv_suspend() and nv_resume() walk the non-PCI configuration space with

	for (i = 0; i <= np->register_size/sizeof(u32); i++)

which runs one iteration too many. saved_config_space is declared as

	u32 saved_config_space[NV_PCI_REGSZ_MAX/4];

and NV_PCI_REGSZ_VER3 is equal to NV_PCI_REGSZ_MAX (0x604), so on a VER3
device register_size/sizeof(u32) is exactly the array length and the last
iteration addresses one element past the end.

The element it lands on is np->name_rx[0..3]: saved_config_space[] is
followed immediately by char name_rx[IFNAMSIZ + 3], and char needs no
padding. Nothing observable is corrupted by that, because nv_request_irq()
rewrites name_rx with sprintf() before it is ever passed to request_irq().
The bug is the out-of-bounds access itself, which UBSAN reports and which
CONFIG_UBSAN_TRAP=y turns into a trap that aborts the running kernel code,
plus an MMIO read and, on resume, an MMIO writel() to base + 0x604, one
dword past the range the driver mapped:

	np->base = ioremap(addr, np->register_size);

VER1 and VER2 devices stay inside the array, but they too get the stray
read and the stray write one dword past their own window.

Caught by UBSAN on an Apple Macmini3,1 (MCP79) during a deep S3 cycle.
The splat below is trimmed: the build path in the file name, the CPU
and taint lines, the Workqueue line, the "?" hint frames, and the
frames below device_suspend are all cut. The kernel was tainted, with
an out-of-tree nouveau and CPU_OUT_OF_SPEC; forcedeth itself was the
stock module.

  UBSAN: array-index-out-of-bounds in drivers/net/ethernet/nvidia/forcedeth.c:6225:25
  index 385 is out of range for type 'u32 [385]'
  Call Trace:
   dump_stack_lvl+0x5d/0x80
   ubsan_epilogue+0x5/0x2b
   __ubsan_handle_out_of_bounds.cold+0x54/0x59
   __this_module+0xe398c/0xe9010 [forcedeth]
   pci_pm_suspend+0x80/0x170
   dpm_run_callback+0x51/0x160
   device_suspend+0x1a2/0x4a0
   ...

Both loops are hit. UBSAN reports each source location only once per module
load (__ubsan_handle_out_of_bounds() calls suppress_report(), which does
test_and_set_bit(REPORTED_BIT, ...) on the struct source_location), so the
two splats land in the first S3 cycle after the module is loaded and later
cycles are silent even though the access still runs off the end every time.
In that first cycle line 6225 is reported from pci_pm_suspend and line 6240
from pci_pm_resume.

The same off-by-one was fixed in nv_get_regs() by commit ba9aa13428
("forcedeth: fix buffer overflow") in 2012; these two loops were missed.
The suspend and resume side was reported on LKML in September 2013 by Marc
Weber, with the same analysis and the same one-character fix, but the patch
was attached rather than sent inline and the thread ended there.

Use < instead of <=, which saves and restores exactly register_size bytes.

Fixes: 1a1ca86158 ("[netdrvr] forcedeth: save/restore device configuration space")
Cc: stable@vger.kernel.org
Signed-off-by: Marek Czernohous <marek@czernohous.de>
Reviewed-by: Simon Horman <horms@kernel.org>
Reviewed-by: Zhu Yanjun <yanjun.zhu@linux.dev>
Link: https://patch.msgid.link/178682367885.3748309.10595890901761762683@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:24:44 -07:00
Jakub Kicinski
ac9d95c71a Merge branch 'net-mlx5-preserve-speed-and-state-across-vport-modify-commands'
Tariq Toukan says:

====================
net/mlx5: Preserve speed and state across vport modify commands

The firmware vport modify command bundles both admin state and max tx
speed in a single operation, which requires each side to preserve the
other field when it only intends to change one.

When modifying max tx speed, the driver already queries the current
admin state and passes it back to avoid overwriting it. However, this
query and the subsequent modify were not atomic, a state change
between the two could cause the modify to overwrite the new state with
a stale value. The fix holds esw->state_lock across the query-modify
sequence.

When support for setting max tx speed via the vport modify command was
introduced, the existing admin state modify path was not updated to
preserve the current speed. As a result, the firmware interprets the
zero speed field as an intentional reset. The fix adds a speed query
before the state modify and passes the result back in the command.

To support that, mlx5_query_vport_max_tx_speed() had to be fixed first:
it was returning zero whenever the vport was DOWN, which was correct
for the query_port_speed verb but would defeat the purpose of querying
before a state modify. The DOWN-to-zero logic is moved to the
verb-layer caller so the function returns the raw firmware value.

Patch #1  holds esw->state_lock across the state query and modify in
          the speed modify path
Patch #2  moves the vport DOWN zero mapping to the verb-layer caller
          so the query returns the raw firmware value
Patch #3  queries current max tx speed before modifying vport state to
          preserve it
====================

Link: https://patch.msgid.link/20260816065015.3280733-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:22:44 -07:00
Or Har-Toov
ad0ae7aefa net/mlx5: E-Switch, preserve max tx speed on vport state modification
When modifying vport state, the firmware interprets a zero in the max tx
speed field as an intentional reset, which can overwrite previously set
values. This patch attempts to fix this by querying the current max tx
speed from firmware before modifying the vport state and passing it back
in the modification command. If the query fails, fall back to the cached
agg_max_tx_speed value to avoid inadvertently resetting the speed.

Fixes: 50f1d188c5 ("net/mlx5: Propagate LAG effective max_tx_speed to vports")
Signed-off-by: Or Har-Toov <ohartoov@nvidia.com>
Reviewed-by: Mark Bloch <mbloch@nvidia.com>
Reviewed-by: Shay Drori <shayd@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260816065015.3280733-4-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:22:42 -07:00
Or Har-Toov
20f11b5cfa net/mlx5: Move vport DOWN state check out of mlx5_query_vport_max_tx_speed()
mlx5_query_vport_max_tx_speed() was introduced to serve the
query_port_speed path, which uses max_tx_speed == 0 when port is down.

This is incorrect for callers that need the actual configured speed
regardless of vport state, such as modify-vport-state helpers
that must preserve the speed across state transitions.

Move this logic to the caller function in the verb flow and let
mlx5_query_vport_max_tx_speed() return the raw firmware value
unconditionally.

Fixes: aaecff5e13 ("RDMA/mlx5: Implement query_port_speed callback")
Signed-off-by: Or Har-Toov <ohartoov@nvidia.com>
Reviewed-by: Shay Drori <shayd@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260816065015.3280733-3-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:22:42 -07:00
Mark Bloch
ff0f9b7aa1 net/mlx5: E-Switch, use state lock for vport state changes
Protect vport admin state modifications and vport iteration with the
eswitch state_lock mutex to ensure proper serialization of concurrent
vport state changes.

Currently, calls to mlx5_modify_vport_admin_state() and loops iterating
over eswitch vports can race with each other, potentially leading to
inconsistent vport state. Fix this by acquiring esw->state_lock

Fixes: 7d0314b11c ("net/mlx5e: Modify uplink state on interface up/down")
Signed-off-by: Mark Bloch <mbloch@nvidia.com>
Reviewed-by: Shay Drori <shayd@nvidia.com>
Reviewed-by: Or Har-Toov <ohartoov@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260816065015.3280733-2-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:22:42 -07:00
Fan Ye
c5ae83ee02 net: thunderbolt: Count delivered packets in rx_packets and rx_bytes
tbnet_poll() increments rx_packets once per received frame because that is
the NAPI work unit, and then adds the same number to stats.rx_packets. An
skb is handed to the stack only when the last frame of a packet arrives,
so once the MTU exceeds TBNET_MAX_PAYLOAD_SIZE the statistic reports
frames. tx_packets is bumped once per skb, so the two ends of a link
disagree: at MTU 65330 the receiver reports 16 times the packets its
sender sent.

rx_bytes has the matching problem: frames of a packet that is later
dropped mid-assembly are already accounted, so it does not correspond to
rx_packets as documented. Account for both where the packet is completed,
and leave the NAPI work counter alone.

Fixes: e69b6c02b4 ("net: Add support for networking over Thunderbolt cable")
Signed-off-by: Fan Ye <fy15309206903@gmail.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Acked-by: Mika Westerberg <westeri@kernel.org>
Link: https://patch.msgid.link/20260815-tbnet-rx-stats-v1-1-8da375c2cd09@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:20:06 -07:00
Andrea Mayer
f826df9533 ipv6: rpl: fix NULL dereference of idev in ipv6_rpl_srh_rcv()
ipv6_rpl_srh_rcv() dereferences idev from __in6_dev_get() without a NULL
check when reading idev->cnf.rpl_seg_enabled.

When the device's MTU drops below IPV6_MIN_MTU, addrconf_ifdown() clears
dev->ip6_ptr through RCU_INIT_POINTER(). A packet that passed the idev
check in ip6_rcv_core() can then reach ipv6_rpl_srh_rcv() with
dev->ip6_ptr already NULL.

Reproduced by flooding the receiving interface with ping6 traffic while
flapping its MTU between 1500 and 1200:

 BUG: KASAN: null-ptr-deref in ipv6_rpl_srh_rcv+0xb3/0x1070
 Read of size 4 at addr 00000000000006b4 by task ping6/394

 CPU: 2 UID: 0 PID: 394 Comm: ping6 Not tainted 7.2.0-rc7-micro-vm-dev-00095-g24ef02f934ee #240 PREEMPT(full)
 Call Trace:
  <IRQ>
  kasan_report+0xc6/0x100
  ipv6_rpl_srh_rcv+0xb3/0x1070
  ip6_protocol_deliver_rcu+0x759/0x9a0
  ip6_input_finish+0xa8/0x1b0
  ip6_input+0xe1/0x490
  ipv6_rcv+0x33d/0x460
  __netif_receive_skb_one_core+0xd6/0x130
  process_backlog+0x2cc/0xa00
  __napi_poll.constprop.0+0x56/0x270
  net_rx_action+0x327/0x730
  handle_softirqs+0x11e/0x630
  do_softirq+0xb3/0xf0
  </IRQ>

Both ipv6_rpl_srh_rcv() and ipv6_srh_rcv() are called only from
ipv6_rthdr_rcv(), which already has an idev lookup.

Fix the NULL dereference on the RPL path by checking idev in
ipv6_rthdr_rcv(), before it calls either function. The callees take idev as
an argument and no longer call __in6_dev_get(), so the packet is now
dropped in one place, with SKB_DROP_REASON_IPV6DISABLED on both paths.

Fixes: 8610c7c6e3 ("net: ipv6: add support for rpl sr exthdr")
Cc: stable@vger.kernel.org
Signed-off-by: Andrea Mayer <andrea.mayer@uniroma2.it>
Tested-by: Xiang Mei <xmei5@asu.edu>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260817132644.2223-1-andrea.mayer@uniroma2.it
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:18:58 -07:00
Hyunwoo Kim
da4471557f net/tcp-ao: fix use-after-free of current_key on reconnect to another peer
tcp_inbound_ao_hash() is called before bh_lock_sock_nested() is taken,
with only rcu_read_lock() held. On the fast path for established
sockets, if the rnext_keyid sent by the peer differs from
current_key->sndid, the key the peer asked for is looked up and stored
in current_key. The lookup is inside the RCU read side, but current_key
outlives it.

When the socket is disconnected and connect() is called again for
another peer, tcp_ao_connect_init() unlinks every key that does not
match the new peer and frees it with call_rcu(). If current_key points
at such a key, it is cleared to NULL.

The fast path reads sk_state only once on entry, so a softirq that got
into it while the socket was still established can update current_key
after that loop has already run. The update is inside the RCU read side,
so it comes before the call_rcu() callback, and once the callback frees
the key, current_key is left pointing at freed memory.

The next transmission picks that pointer up in tcp_get_current_key().
tcp_ao_transmit_skb() then reads the traffic key from the freed object,
which is the use-after-free.

Wait for one grace period before unlinking, and only if a key is going
to be removed. By the time tcp_connect() runs the socket is already in
TCP_SYN_SENT, and TCP_AO_ESTABLISHED does not contain TCPF_SYN_SENT, so
a softirq entering after the wait cannot reach the fast path, and the
ones already in it have finished. The existing NULL handling in the loop
is then enough.

Fixes: 0a3a809089 ("net/tcp: Verify inbound TCP-AO signed segments")
Cc: stable@vger.kernel.org
Signed-off-by: Hyunwoo Kim <imv4bel@gmail.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Acked-by: Paolo Abeni <pabeni@redhat.com>
Link: https://patch.msgid.link/aoIriv3pHDgII2YR@v4bel
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:18:12 -07:00
Victor Nogueira
8e2efb3f45 net/sched: add get_fill_size callbacks for actions missing them
Several tc actions - act_police, act_bpf, act_pedit, act_ife, act_sample,
act_ct, act_ctinfo and act_tunnel_key among them - provide no
get_fill_size() callback, so tcf_action_fill_size() falls back to
tcf_action_shared_attrs_size() which does not account for the
action-specific netlink attributes emitted inside TCA_ACT_OPTIONS by
their dump functions.

When an RTM_NEWACTION request with NLM_F_ECHO (or an RTNLGRP_TC
listener) creates several actions, tcf_add_notify_msg() allocates the
echo skb from this underestimated size. When this happens, the act_api
code fails to add all of the fields to the netlink message and, thus,
fails to send it. Issue is that, when that happens, this failure doesn't
stop the action instances from being added. So any user watching these
events will be under the false impression that no actions were created at
all.

For example, act_pedit overruns with 32 actions of four munge keys each,
act_police with 32 policers once the optional rate/peakrate/result/avrate
attributes are present.

To fix this, add the missing get_fill_size callbacks returning the
worst-case size of each action's dump attributes, following the pattern
used by act_gact/act_skbedit/act_vlan. Also widen the TCA_GACT_TM
accounting in tcf_action_shared_attrs_size() to nla_total_size_64bit(),
since actions dump their tcf_t with nla_put_64bit(), which may be
preceded by an NLA_PAD attribute.

Note: We only provided fixes for the actions we reproduced this bug with
as of today. We can send a separate hardening patch for the remaining
actions to net-next later. The other pre-existing issues, pointed out by
Clashiko [1], will be fixed in upcoming patches.

[1] https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260810164357.1653956-1-victor%40mojatatu.com

Fixes: 4e76e75d6a ("net sched actions: calculate add/delete event message size")
Reported-by: Vega <vega@nebusec.ai>
Acked-by: Jamal Hadi Salim <jhs@mojatatu.com>
Signed-off-by: Victor Nogueira <victor@mojatatu.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260816201327.2435335-1-victor@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:16:43 -07:00
Eric Dumazet
e5c8e301b4 selftests: net: packetdrill: add tests for advertised MSS with PMTU exceptions
Add packetdrill tests for IPv4 and IPv6 to verify that the advertised
MSS in SYN-ACK is derived from the configured interface/route MTU,
and is not shrunk by learned Path MTU exceptions from previous
outbound connections.

Signed-off-by: Eric Dumazet <edumazet@google.com>
Signed-off-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Link: https://patch.msgid.link/20260815071532.301908-1-jiayuan.chen@linux.dev
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:14:23 -07:00
Jiayuan Chen
2640e64195 net: advertise TCP MSS from the configured MTU, not the learned PMTU
The MSS a host puts in its SYN tells the peer how big a segment it may
send us. Right now we can shrink it with a PMTU we learned on our own
send path, which is the wrong direction entirely.

On asymmetric paths this bites - think DSR load balancers, where the
request side goes through a smaller-MTU overlay. We learn a small PMTU
going out, then advertise a small MSS, and the peer stays capped for the
whole connection even though its path back to us is wide. MSS only shows
up in the SYN and never grows back.

On symmetric paths we lose nothing by dropping it either: the peer runs
its own PMTU discovery and usually already knows the real path MTU.

So work out the advertised MSS from the configured route or device MTU
and ignore the learned PMTU. Our send side is unchanged, still clamped by
tcp_current_mss(). Add ip_dst_mtu_configured()/ip6_dst_mtu_configured()
and use them from the two default_advmss() paths.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Fixes: 164a5e7ad5 ("ipv4: ipv4_default_advmss() should use route mtu")
Cc: stable@vger.kernel.org
Signed-off-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260815070413.294559-1-jiayuan.chen@linux.dev
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:14:23 -07:00
Théo Lebrun
ec518a7c4b net: macb: drop CONFIG_OF #if block
Fix -Wimplicit-function-declaration error on CONFIG_OF=n builds:

   drivers/net/ethernet/cadence/macb_main.c: In function ‘macb_probe’:
   drivers/net/ethernet/cadence/macb_main.c:5951:15: error: implicit
   declaration of function ‘macb_alloc_tieoff’ [...]
    5951 |         err = macb_alloc_tieoff(bp);
         |               ^~~~~~~~~~~~~~~~~
   drivers/net/ethernet/cadence/macb_main.c:5973:9: error: implicit
   declaration of function ‘macb_free_tieoff’ [...]
    5973 |         macb_free_tieoff(bp);
         |         ^~~~~~~~~~~~~~~~

Error got introduced because functions are mistakenly declared in a
`#if defined(CONFIG_OF)` block. Instead of moving functions around,
avoid any future mistake and drop the block entirely.

Change the module content slightly on CONFIG_OF=n. Previously match
tables were ignored. Now they appear in the resulting build. This is
considered trivial in size by most and is the common case:

   ⟩ 18 out of 254 OF net drivers reference CONFIG_OF
   ⟩ rg -lF 'MODULE_DEVICE_TABLE(of,' drivers/net/ | tee /tmp/a | wc -l
   254
   ⟩ xargs -a /tmp/a rg -l CONFIG_OF | wc -l
   18

Tangent: no, of_match_ptr() does not imply that the compiler can
optimize out match tables, because MODULE_DEVICE_TABLE(of, ...)
unconditionally puts the match tables in the binary. It is only meant
to avoid undefined declaration issues when match tables are hidden
behind a #ifdef, as was done before. We therefore drop the macro call.

Fixes: 5262eab946 ("net: macb: allocate tieoff descriptor once across device lifetime")
Reported-by: Nathan Chancellor <nathan@kernel.org>
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Acked-by: Conor Dooley <conor.dooley@microchip.com>
Link: https://patch.msgid.link/20260820-macb-fix-x86-v1-1-b2e7c902104e@bootlin.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:13:32 -07:00
Jakub Kicinski
50720728b1 ipsec-2026-08-18
-----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEH7ZpcWbFyOOp6OJbrB3Eaf9PW7cFAmqEJIUACgkQrB3Eaf9P
 W7eUHRAAhUaCftYnbSKKcvB8DgDysrRRFOieJ5ucqyYWCc51/O8bQWspzvFd2fiP
 cq7KLubREyGD8FqMNwl94J2zTW7awrGWyNkiA0TwNouOWIM5yu4eg7aZ1+edOMrx
 FF15HM8Q4DNgfHGdNYZKzRzP+72qLNEY92o6nbDYQUZmB33tFjic44+7Vphhjwb3
 t/GulrwfA8M/98oDgmzqwxSIz+/5E+kXSqLouD/vCMXPbdDv0m1xW2iNPHkU+Bom
 Kk6WNlcPwJWmpM5mfaWP4C2T1reJnyi99MorBco69PrGFhCVxBftQO08qGaE5EeR
 YbNNrvPKs7mcCqnwfhDObKz8GdkPIvt79p/UKQjardN1ts/aU5N8CD4bwSzMHWep
 dmz3j9sydtQom+YXYxAgr50DKpyZKOS7abQou4jTmwTz5/fAHVSPzfzE7aSJpE6o
 Df9gW7cGmAs4KSeQaHotEBOR790AedwG1bHdn7C/KqOdd4e8IwzC+6ZLNjzlrC/f
 ZtwN64Ct8uChIs6A+SAnzD+C7SEP8k0A/MFOwbf+Ov5kYLkzFYL/JudB4eK437kQ
 K+VcZj/jrs3aBqZm5Y/O5PlK4/Bpa04XJamK2cG9la5RGked4VWSffSd9NQIg07t
 2Lct15Qe7K9aJrLkXA92Qbjkmhq92RKqVCJ7ylgkwW1cRm5FxDc=
 =wXnr
 -----END PGP SIGNATURE-----

Merge tag 'ipsec-2026-08-18' of git://git.kernel.org/pub/scm/linux/kernel/git/klassert/ipsec

Steffen Klassert says:

====================
pull request (net): ipsec 2026-08-18

1) xfrm6: fix out-of-bounds write in xfrm6_input_addr() when secpath is full
   Tighten the secpath-depth check so a full chain can't write
   past xvec[].

2) Add and revert "esp: do not unref managed frag pages in esp_ssg_unref()"
   The patch does not fully fully resolve the issue, a corrected version
   will follow.

3) xfrm: espintcp: fix UAF during close
   Synchronize espintcp close with the xfrm_trans_reinject work
   queue so the freed socket message isn't dereferenced again.

4) xfrm: drop ESP-in-TCP packets with no ingress device
   Drop queued ESP-in-TCP records whose saved ingress device has
   gone away, avoiding a NULL device deref in the XFRM input path.

5) xfrm: avoid lock inversion in nat keepalive work
   Split the NAT keepalive walk into a reference-collection phase
   and a per-state lock phase to break the AB-BA with state removal.
   This patch has some issues that are fixed with a followup patch.

6) xfrm: Fix skb double-free in xfrm_dev_direct_output()
   Stop freeing the skb unconditionally in xfrm_dev_direct_output(),
   letting local_out()'s result indicate when ownership has moved on.

7) xfrm: ah6: validate routing header segments_left
   Validate the segments_left/hdrlen invariant before rearranging
   the routing-header addresses, avoiding an OOB memmove on
   malformed HDRINCL packets.

8) xfrm: fix xfrm_state_construct() auth-trunc leak
   Detect an already-attached auth-trunc allocation by the pointer
   rather than inferring it from the algorithm id, so a prior
   attach isn't overwritten and lost.

9) xfrm: bound nat keepalive state collection
   Replace the per-state allocation in the NAT keepalive walk
   with a fixed-size batch that drains under BH-disabled locking
   and resumes from the cursor, bounding the worker's memory.

* tag 'ipsec-2026-08-18' of git://git.kernel.org/pub/scm/linux/kernel/git/klassert/ipsec:
  xfrm: bound nat keepalive state collection
  Revert "esp: do not unref managed frag pages in esp_ssg_unref()"
  xfrm: fix xfrm_state_construct() auth-trunc leak
  xfrm: ah6: validate routing header segments_left
  xfrm: Fix skb double-free in xfrm_dev_direct_output()
  xfrm: avoid lock inversion in nat keepalive work
  xfrm: drop ESP-in-TCP packets with no ingress device
  xfrm: espintcp: fix UAF during close
  esp: do not unref managed frag pages in esp_ssg_unref()
  xfrm6: fix out-of-bounds write in xfrm6_input_addr() when secpath is full
====================

Link: https://patch.msgid.link/20260818092920.653034-1-steffen.klassert@secunet.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 11:38:14 -07:00
Jakub Kicinski
066ae87fe9 netfilter pull request 26-08-18
-----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEjF9xRqF1emXiQiqU1w0aZmrPKyEFAmqDlHcACgkQ1w0aZmrP
 KyES3w/8C0DTvpbwLOr/QDgbgbo39u78Ih/OxXBqxtXth8bc84DFGXgFxJdHlujI
 g18zRn9/r4Jo2K56nkJDlUiVVJKOsxuhW5WFCG49qgc3cq4DDwhh85DNg3XFgQLs
 q+UPf6UHADChfDbBzLezuHKY/Cot8BvpfirQhcQT2kfgv8XmHp3UNidyu2GM3Zqp
 E31MhRm/ei+IR2R3hhNOs14/cEP3tAY4jTgN4Z3hs63SKw6nZpGDpVtdEHLHWAn0
 1b4EawsPlhGp8fjGGNq2hYCeyVKTQrfe3jS/roJ3IfU3Wq/k5XKufSUcG2WNQgRX
 KifhAXlabcW+tNabay6IWoClIzQPuV+SxcNtswvx+XAnrQ0h9qZb2YDYQlpAE2KP
 uZ4CWMyzhognmUtpfB957RV3/Q8qG3pleqwz+LpVeGGm6VX/cT20g460X7eyq9+e
 Ux27fCfSNkC4UHhR+iqM8BAqpsI47OS/X+PC2enCzmSh+JIfEsBeOcUylHVT+p14
 mF4USETDzWaXN9yHK1bnbANzG/sRo8KT/pS7taN7+vk5Czje50tckqjgZ0bX5ZAK
 tOxZGKBajDEjntOWLEEpMvFknkxsI7H55YGU9Glhsbfrf/gfxv61+XSCe8kQ2+a8
 QbRuMepR45Vfz04ZiHJoalTi7B9FTsGC3dXCGuL9Es8xf8xObPw=
 =qmiO
 -----END PGP SIGNATURE-----

Merge tag 'nf-next-26-08-18' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf-next

Pablo Neira Ayuso says:

====================
Netfilter/IPVS fixes for net-next

This contains fixes for nf_tables, revisit issues with expectation
infra updates reported by sashiko, an ipset fix for deletions in the
hash:net type and tne fix for the IPVS FTP helper.

1) Validate layer 4 header mangling done via nfnetlink_queue and
   nft_payload, this is a follow up to recent similar validation
   at layer 3. From Zhiling Zou.

2) Do not allocate memory on delete operations in ipset hash:net
   type, delete operation must always succeed. From Florian Westphal.

3) Deliver nft_obj overquota packet path notification directly via
   nfnetlink, do not use the control plane batch logic.
   From Fourie Zhang.

4) Follow up to controlidate check for reinserted dead expectations,
   to cover the nf_conntrack_expect_related_pair() function too.

5) Do not expose expectation dead flag to userspace via ctnetlink.

6) Make commit set_update_list per-netns to prepare to publish
   set clone earlier.

7) Publish the set clone earlier from commit path to address set
   lookup failures during table re-creation, this is targetting
   the rbtree and pipapo set backends.

8) Fix an integer overflow in the IPVS FTP helper. A similar fix
   was already proposed for the conntrack FTP helper months ago.
   From Joas Antonio dos Santos.

* tag 'nf-next-26-08-18' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf-next:
  ipvs: fix integer overflow in ftp helper port/address parsing
  netfilter: nf_tables: call set ops .commit when building new ruleset blob
  netfilter: nf_tables: move set_update_list to nftables per-netns
  netfilter: ctnetlink: do not expose expectation DEAD flag
  netfilter: nf_conntrack_expect: consolidate check for insertion of dead expectation
  netfilter: nf_tables: don't queue packet path object notifications
  netfilter: ipset: remove need to allocate memory on delete operations
  netfilter: validate L4 headers after userspace packet writes
====================

Link: https://patch.msgid.link/20260817232957.1281637-1-pablo@netfilter.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 11:36:57 -07:00
Yuyang Huang
47cdab0d51 ipv6: use RCU iterator to dump route exceptions
rt6_nh_dump_exceptions() uses hlist_for_each_entry() to iterate over
RCU-protected exception lists. The caller holds rcu_read_lock(), but does
not hold rt6_exception_lock, so rt6_insert_exception() can concurrently
add an entry with hlist_add_head_rcu().

KCSAN reports this race (irrelevant details omitted):

  ==================================================================
  BUG: KCSAN: data-race in rt6_insert_exception / rt6_nh_dump_exceptions

  write (marked) to 0xffff8a7c44c59620 of 8 bytes by interrupt on cpu 5:
    rt6_insert_exception+0x3bb/0x760
    __ip6_rt_update_pmtu+0x4fe/0x750
    ip6_sk_update_pmtu+0x19a/0x3b0
    udpv6_err+0x3ff/0x800
    icmpv6_notify+0x1e1/0x440
    icmpv6_rcv+0x8c0/0xab0
    ip6_protocol_deliver_rcu+0x616/0x840
    ip6_input_finish+0xb9/0x160
    ...
    entry_SYSCALL_64_after_hwframe+0x77/0x7f

  read to 0xffff8a7c44c59620 of 8 bytes by task 549 on cpu 14:
    rt6_nh_dump_exceptions+0xb3/0x260
    rt6_dump_route+0x53e/0x5f0
    fib6_dump_node+0x6d/0xf0
    fib6_walk_continue+0x290/0x2d0
    fib6_dump_table+0x28d/0x360
    inet6_dump_fib+0x37d/0x620
    rtnl_dumpit+0x7b/0xd0
    netlink_dump+0x3ae/0x7e0
    ...
    entry_SYSCALL_64_after_hwframe+0x77/0x7f

  4 locks held by dumper/549:
    ...
    #1: (rcu_read_lock){....}-{1:3}, at: inet6_dump_fib+0x88/0x620
    #2: (&tb->tb6_lock){+.-.}-{3:3}, at: fib6_dump_table+0x1e9/0x360
    #3: (rcu_read_lock){....}-{1:3}, at: rt6_dump_route+0x483/0x5f0

  value changed: 0xffff8a7c44e05700 -> 0xffff8a7c45d60100

  Reported by Kernel Concurrency Sanitizer on:
  CPU: 14 UID: 0 PID: 549 Comm: dumper Not tainted
  7.2.0-rc7-virtme #38 PREEMPT(lazy)
  ...

Use hlist_for_each_entry_rcu() to safely iterate over the exception list.

Fixes: 1e47b4837f ("ipv6: Dump route exceptions if requested")
Cc: stable@vger.kernel.org
Signed-off-by: Yuyang Huang <sigefriedhyy@gmail.com>
Reviewed-by: Stefano Brivio <sbrivio@redhat.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260815084651.69477-1-sigefriedhyy@gmail.com
Signed-off-by: David S. Miller <davem@davemloft.net>
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 11:34:25 -07:00
Jorijn van der Graaf
3cbfd627ee net: ipa: fix stalled modem TX queue after runtime resume
ipa_start_xmit() unconditionally stops the TX queue before calling
pm_runtime_get(), relying on the wake scheduled by runtime resume
(ipa_modem_wake_queue_work()) to restart it once power is ACTIVE.
But that work is queued from within the runtime resume callback,
before the device's power state reaches RPM_ACTIVE, so it can run
while the device is still RPM_RESUMING.  The wake is then consumed
too early: the transmit it restarts stops the queue again,
pm_runtime_get() returns -EINPROGRESS without arranging any future
wake (deferred_resume exists only for RPM_SUSPENDING), and after the
resume completes nothing is left to wake the queue.  Transmit stalls
permanently: packets pile up in the qdisc behind the stopped queue,
the device runtime-suspends, and since the netdev registers no
ndo_tx_timeout the watchdog never fires.  Observed on SM7635
(Fairphone 6) as the cellular data path going permanently deaf
within hours, RX included, since nothing resumes the suspended
endpoints.

Close the window by making the wake work wait for the resume to
complete (pm_runtime_get_sync()) before waking the queue.  Every
queue stop is then guaranteed a later wake that happens while power
is ACTIVE; a transmit racing a new suspend/resume cycle re-schedules
the work.  If the device could not be resumed, wake the queue anyway
so pending packets are dropped by the transmit path rather than
stranded.

The STARTED power flag used to narrow this window: a wake running
before the transmit path's stop suppressed that stop, but only once,
as the flag was cleared by the first stop it absorbed.  Removing the
flag made a single transmit during an in-flight resume sufficient to
strand the queue, which is the form observed.

With an accelerated reproducer (autosuspend delay shortened to 5 ms,
~20 packets/s of TX), an unpatched kernel stalled three times in
230 s / 4380 packets; with this patch the same test ran 3601 s /
70298 packets without a stall.

Fixes: 688de12f08 ("net: ipa: kill the STARTED IPA power flag")
Cc: stable@vger.kernel.org
Signed-off-by: Jorijn van der Graaf <jorijnvdgraaf@catcrafts.net>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260815040302.653650-1-jorijnvdgraaf@catcrafts.net
Signed-off-by: David S. Miller <davem@davemloft.net>
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 11:33:39 -07:00
Ruoyu Wang
b74a072d8f net: bridge: Reject descending VLAN tunnel ranges
A pair of descending VLAN and tunnel IDs can pass the tunnel range span
check. The VLAN subtraction produces a negative int, which is converted
to unsigned when compared with the u32 tunnel ID subtraction. It can
therefore equal the wrapped tunnel ID delta.

The range loop then performs no iterations. Since the batched
notification handling added a post-loop error check, this leaves err
uninitialized and makes the request's return value unpredictable.

Reject descending VLAN ranges before comparing the spans. Valid
ascending and single-entry ranges remain unchanged, while malformed
descending ranges consistently return -EINVAL.

This issue was found by a static analysis checker and confirmed by
manual source review.

Fixes: 9433944368 ("net: bridge: notify on vlan tunnel changes done via the old api")
Signed-off-by: Ruoyu Wang <ruoyuw560@gmail.com>
Acked-by: Nikolay Aleksandrov <razor@blackwall.org>
Link: https://patch.msgid.link/20260814134053.1387275-1-ruoyuw560@gmail.com
Signed-off-by: David S. Miller <davem@davemloft.net>
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 11:32:44 -07:00
Cen Zhang (Microsoft)
e37b2abca8 xsk: fix NULL pointer dereference in __xsk_rcv()
In the __xsk_rcv() multi-buffer path, xsk_buff_alloc() is called in a
loop without checking its return value. xsk_buff_can_alloc() only
counts fill queue entries without validating their addresses, so it
can succeed while xsk_buff_alloc() rejects all remaining entries and
returns NULL.

  Oops: general protection fault, probably for non-canonical address
   0xdffffc0000000000
  KASAN: null-ptr-deref in range
   [0x0000000000000000-0x0000000000000007]
  RIP: 0010:__xsk_rcv+0x426/0xc20 (net/xdp/xsk.c:350)
  Call Trace:
   xsk_generic_rcv+0x26d/0x5f0
   xdp_do_generic_redirect+0x3c5/0xcf0
   do_xdp_generic+0x92f/0xe70
   __netif_receive_skb_core.constprop.0+0xf7e/0x2b30

Fix this with a two-stage transaction. First allocate and stage all
buffers required for the packet, recycling all staged buffers with
xsk_buff_free() if any allocation fails. Only after this stage
succeeds, copy the data, reserve the RX descriptors, and release the
buffers in an error-free loop.

Fixes: 804627751b ("xsk: add support for AF_XDP multi-buffer on Rx path")
Reported-by: AutonomousCodeSecurity@microsoft.com
Signed-off-by: Cen Zhang (Microsoft) <blbllhy@gmail.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Reviewed-by: Jason Xing <kerneljasonxing@gmail.com>
Link: https://patch.msgid.link/20260813215328.99311-1-blbllhy@gmail.com
Signed-off-by: David S. Miller <davem@davemloft.net>
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 11:31:35 -07:00
Linus Torvalds
91ec203513 Networking changes for 7.3.
Core & protocols
 ----------------
 
  - A few steps lowering rtnl_lock dependence:
    - per-netns netdev unregistration for select SW drivers
      (e.g. veth, ipvlan, tunnels)
    - rtnl_lock-less FIB rule changes (RTM_NEWRULE and RTM_DELRULE)
    - prepare software drivers and TC qdiscs for rtnl_lock-less GET
 
  - Support BIG TCP (>64kB TSO) in UDP tunnels (vxlan, geneve).
 
  - Support buffers larger than PAGE_SIZE in devmem zero-copy API.
 
  - Improve MPTCP handling of extreme memory pressure handling,
    when out-of-order queue had to be pruned.
 
  - Report the per-group user count via RTM_GETMULTICAST.
 
  - Expose the route deletion reason in RTM_DELROUTE.
 
  - Add a SO_RIGHTS_NOTRUNC option to UNIX sockets to enable more useful
    handling of LSM denials when receiving SCM_RIGHTS messages: instead
    of truncating the message at the first blocked fd, keep every fd slot
    and store the LSM errno in the blocked slot.
 
  - IPv6 Segment Routing - support looking up the post-encap SID
    (address) in a different/specified routing table.
 
  - Support PRP RedBox (interlink) creation.
 
  - Support per-nexthop UDP dst port in VXLAN.
 
  - Continue converting getsockopt callbacks in a number of protocols
    to iov_iter.
 
 Ethernet
 --------
 
  - Marge initial CXL support for AMD/Solarflare NICs (shared branch
    with the CXL tree).
 
  - New drivers:
    - ADIN1140 10BASE-T1S MACPHY
    - Initial skeleton of Intel iXD and ZTE Dinghai drivers.
 
  - High-speed NICs:
    - AMD/Pensando:
      - support firmware flashing
    - Cisco (enic):
      - SR-IOV V2 admin channel and MBOX protocol
    - Huawei (hns3):
      - support for ethtool pfc_prevention_tout
    - nVidia/Mellanox:
      - support sharing bandwidth control across interfaces of
        the same device
    - Marvell (octeontx2-pf):
      - link RQ page pools to netdev for Netlink stats
    - Google vNIC:
      - XDP metadata support for DQ RDA
    - Microsoft vNIC:
      - support forcing full-page RX buffers
 
  - Other NICs:
    - Synopsys IP:
      - eic7700: support for eth1
    - Microchip (lan743x):
      - support for RMII interface
    - Wangxun:
      - support for ethtool -G and -C for VFs
      - add Tx timeout and PCIe error handling
    - Intel (igb/igc):
      - RSS key get/set support
      - support for forcing link speed without auto-negotiation
 
  - Switches:
    - NXP (dpaa2):
      - support bonding/LAG offload
    - Mediatek:
      - mt7530: EN7528 support
      - initial support for MT7628
    - Micrel (ksz8/9):
      - refactoring work to move towards library model
      - PTP support for KSZ8463
    - nVidia/Mellanox:
      - support rtnl-lock-less ethtool callbacks
    - Realtek:
      - rtl8366rb: use generic RTL83xx code
      - support SGMII and HSGMII for RTL8367S
 
  - PHYs:
    - Airoha:
      - EcoNet EN7528 PHY support
    - DAPU Telecom
      - DAPU Telecom DAP8211R(I) Gigabit PHY support
    - Realtek:
      - support RTL8261C_CG
      - support RTL8261D
 
 Wireless
 --------
 
  - nl80211: per-link statistics support for multi-link operation
 
  - mac80211: AQL/airtime-fairness support for multicast
 
  - Merge Peripheral Authentication Service (PAS) / TEE support
    for ath12k (shared branch with the firmware/qcom tree).
 
  - New drivers:
    - mm81x for Morse Micro Long-Range S1G devices
    - nxpwifi for NXP devices (mostly forked off from mwifiex)
 
  - Driver changes:
    - Broadcom (brcmfmac):
      - DPP support, some Cypress part update
    - MediaTek (mt76):
      - mt7928 support
      - mt7925 NAN support
      - mt7996 AP powersave improvements
    - Qualcomm (ath12k):
      - much kernel infrastructure integration work
      - AHB platform MultiPD support
    - Realtek (rt89):
      - LED support
      - RTL8922DE support
      - dual-BT coex for RTL8922D
    - Intel:
      - new FW version support
 
 Bluetooth
 ---------
 
  - HCI: add support for Shorter Connection Interval (SCI) feature.
 
  - af_bluetooth: add minimal context analysis annotations.
 
  - Driver changes:
    - Intel:
      - add Bluetooth SAR revision 2 support
      - add vendor_reset PCI sysfs for PLDR
    - Mediatek:
      - add USB IDs for MT7902 and MT7922 devices
    - Realtek:
      - add USB IDs for 8761CU and 8852BE devices
    - NXP:
      - add M.2 Bluetooth device support using pwrseq
 
 Misc
 ----
 
  - DPLL support for manual/numerical oscillator control (NCO)
    (implement in zl3073x).
 
  - MCTP support for MCTP over USB v1.1 (DMTF DSP0283).
 
  - Power-over-Ethernet: support Realtek PSE controllers.
 
  - Remove the IBM EHEA driver.
 
  - Remove tulip/xircom_cb driver.
 
 Signed-off-by: Jakub Kicinski <kuba@kernel.org>
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEE6jPA+I1ugmIBA4hXMUZtbf5SIrsFAmqEwJ4ACgkQMUZtbf5S
 Irsegw//fmHJae525nxg3DHoXhrUz8EDDOVoLH6oyWyLQnh5bmbReAY/+oWA4m54
 3KKKO0b2rtgRvmY/7rnjAt3bjecYgCSjvZT7I+NosB0QbbBYc14PtHfYig9HffYm
 uCXfNJOk+aJ2QK4ncEvU2SjgE89Ya7cC+yARFBAwYx4zi/Qx24RB+ziOyvkQ8ksX
 atvMOZrnhwqvYUFOwnOLNHTpvdxB/ZsNwWY6iXcx6EYp9xrtPusbh3FlushWkwxH
 8cI/dNla44TcIKXAzRn0znRdgiEVmCMyHvOv7LKaOfy8P3I+knmuIf/mScYQqOEF
 T143HdXhVSBZFRtLtFKXIja/KsvCjX9lCeMn/2ak0brQDUREcacXxYbuZKDsNAAK
 zXt/+5qAcm/mO8W1gKR9Ulfli5bhFN4HKXgXMLjo5ucPtzfPxFN7HGxTiC3Cxv1v
 lSXexKaj74pNBVFmADrb5jWbq7oG+GzIdjzx3ycvm2q39Fr4nJ2SzrSPPNwc/ItQ
 IHv3tGLQKXlr8dl0+p2mDkRInmHXrawVNsB1UgN8E/jtcwT2QMwyWOV6s5G3uEDl
 a+0U/XsrPvDYBTUCRs/KaOJQGB90QkzLe9DATt159mf+rPzAX2/oCDo8xIEe+kWV
 aivP+YutFfMH/CSC9PMuvdLE2KmoPY4mibAeE4/4AYLKtJnc/yU=
 =zDto
 -----END PGP SIGNATURE-----

Merge tag 'net-next-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net-next

Pull networking updates from Jakub Kicinski:
 "One of the 'small improvements all over the place' releases for us.

  It's hard to draw any direct comparisons because summer vacations
  disrupted our patch processing (and presumably - generation) quite a
  bit.

  Quick and dirty count suggests we (Paolo and I) merged a very similar
  number of net (632) and net-next (648) patches. This is not telling
  the full story either because 1/3 to 1/2 of the net-next patches also
  *seem* like AI-driven low priority fixes, cleanups and clarifications.

  We are completely overwhelmed, of course. The glimmer of hope is that
  we secured sufficient LLM budget and access (thank you Meta!) to run
  reviews with multiple frontier models on each patch. This eliminates
  some hallucinations. That said, in terms of review, the LLMs can only
  do so much.

  The sad truth is that our APIs (especially for rare events like PCIe
  errors, timeouts etc) have always been racy, and now LLMs don't let us
  ignore that. I expect our direction for the next release will be to
  tweak the reviews a little bit more, but start shifting focus to
  letting the LLMs take care of the busy work - managing patchwork,
  automating common process complaints, editing commit messages, and
  maybe applying patches which already got "reviewed-by" tags from
  people we trust...

  Core & protocols:

   - A few steps lowering rtnl_lock dependence:
      - per-netns netdev unregistration for select SW drivers (e.g.
        veth, ipvlan, tunnels)
      - rtnl_lock-less FIB rule changes (RTM_NEWRULE and RTM_DELRULE)
      - prepare software drivers and TC qdiscs for rtnl_lock-less GET

   - Support BIG TCP (>64kB TSO) in UDP tunnels (vxlan, geneve)

   - Support buffers larger than PAGE_SIZE in devmem zero-copy API

   - Improve MPTCP handling of extreme memory pressure handling, when
     out-of-order queue had to be pruned

   - Report the per-group user count via RTM_GETMULTICAST

   - Expose the route deletion reason in RTM_DELROUTE

   - Add a SO_RIGHTS_NOTRUNC option to UNIX sockets to enable more
     useful handling of LSM denials when receiving SCM_RIGHTS messages:
     instead of truncating the message at the first blocked fd, keep
     every fd slot and store the LSM errno in the blocked slot

   - IPv6 Segment Routing - support looking up the post-encap SID
     (address) in a different/specified routing table

   - Support PRP RedBox (interlink) creation

   - Support per-nexthop UDP dst port in VXLAN

   - Continue converting getsockopt callbacks in a number of protocols
     to iov_iter

  Ethernet:

   - Merge initial CXL support for AMD/Solarflare NICs (shared branch
     with the CXL tree)

   - New drivers:
      - ADIN1140 10BASE-T1S MACPHY
      - Initial skeleton of Intel iXD and ZTE Dinghai drivers

   - High-speed NICs:
      - AMD/Pensando:
         - support firmware flashing
      - Cisco (enic):
         - SR-IOV V2 admin channel and MBOX protocol
      - Huawei (hns3):
         - support for ethtool pfc_prevention_tout
      - nVidia/Mellanox:
         - support sharing bandwidth control across interfaces
           of the same device
      - Marvell (octeontx2-pf):
         - link RQ page pools to netdev for Netlink stats
      - Google vNIC:
         - XDP metadata support for DQ RDA
      - Microsoft vNIC:
         - support forcing full-page RX buffers

   - Other NICs:
      - Synopsys IP:
         - eic7700: support for eth1
      - Microchip (lan743x):
         - support for RMII interface
      - Wangxun:
         - support for ethtool -G and -C for VFs
         - add Tx timeout and PCIe error handling
      - Intel (igb/igc):
         - RSS key get/set support
         - support for forcing link speed without auto-negotiation

   - Switches:
      - NXP (dpaa2):
         - support bonding/LAG offload
      - Mediatek:
         - mt7530: EN7528 support
         - initial support for MT7628
      - Micrel (ksz8/9):
         - refactoring work to move towards library model
         - PTP support for KSZ8463
      - nVidia/Mellanox:
         - support rtnl-lock-less ethtool callbacks
      - Realtek:
         - rtl8366rb: use generic RTL83xx code
         - support SGMII and HSGMII for RTL8367S

   - PHYs:
      - Airoha:
         - EcoNet EN7528 PHY support
      - DAPU Telecom
         - DAPU Telecom DAP8211R(I) Gigabit PHY support
      - Realtek:
         - support RTL8261C_CG
         - support RTL8261D

  Wireless:

   - nl80211: per-link statistics support for multi-link operation

   - mac80211: AQL/airtime-fairness support for multicast

   - Merge Peripheral Authentication Service (PAS) / TEE support for
     ath12k (shared branch with the firmware/qcom tree)

   - New drivers:
      - mm81x for Morse Micro Long-Range S1G devices
      - nxpwifi for NXP devices (mostly forked off from mwifiex)

   - Driver changes:
      - Broadcom (brcmfmac):
         - DPP support, some Cypress part update
      - MediaTek (mt76):
         - mt7928 support
         - mt7925 NAN support
         - mt7996 AP powersave improvements
      - Qualcomm (ath12k):
         - much kernel infrastructure integration work
         - AHB platform MultiPD support
      - Realtek (rt89):
         - LED support
         - RTL8922DE support
         - dual-BT coex for RTL8922D
      - Intel:
         - new FW version support

  Bluetooth:

   - HCI: add support for Shorter Connection Interval (SCI) feature

   - af_bluetooth: add minimal context analysis annotations

   - Driver changes:
      - Intel:
         - add Bluetooth SAR revision 2 support
         - add vendor_reset PCI sysfs for PLDR
      - Mediatek:
         - add USB IDs for MT7902 and MT7922 devices
      - Realtek:
         - add USB IDs for 8761CU and 8852BE devices
      - NXP:
         - add M.2 Bluetooth device support using pwrseq

  Misc:

   - DPLL support for manual/numerical oscillator control (NCO)
     (implement in zl3073x)

   - MCTP support for MCTP over USB v1.1 (DMTF DSP0283)

   - Power-over-Ethernet: support Realtek PSE controllers

   - Remove the IBM EHEA driver

   - Remove tulip/xircom_cb driver"

* tag 'net-next-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net-next: (1433 commits)
  net/mlx5e: do not HW-GRO coalesce small frames
  net: openvswitch: fix nf_connlabels leak in ovs_ct_init
  net: add missing ref_tracker_dir_exit() to alloc_netdev_mqs()
  net: openvswitch: fix flow mask use-after-free on flow deletion
  sctp: stop processing a packet once its association is deleted
  dpll: zl3073x: add PTP clock support
  dpll: zl3073x: add channel ToD, phase step and TIE operations
  dpll: zl3073x: scale poll interval proportionally to timeout
  ptp: vmclock: prevent read-only mappings from becoming writable
  ipv4: reject undersized MTUs in ip_do_fragment()
  bonding: initialize err for empty target lists
  net: dsa: initial support for MT7628 embedded switch
  net: dsa: initial MT7628 tagging driver
  net: phy: mediatek: add phy driver for MT7628 built-in Fast Ethernet PHYs
  dt-bindings: net: dsa: add MT7628 ESW
  net: pse-pd: realtek-pse-mcu: add UART transport
  net: pse-pd: realtek-pse-mcu: add I2C transport
  net: pse-pd: add Realtek PSE MCU core
  dt-bindings: net: pse-pd: add bindings for Realtek PSE MCU
  vsock: use sock_error() to consume sk_err after a failed connect
  ...
2026-08-20 08:16:04 -07:00
Linus Torvalds
5a8cd539ac Major changes:
- Redesign the verifier error reporting: failures now carry source and
   instruction annotations along with the causal event history that led
   to them, making program rejections far easier to debug and repair
   (Kumar Kartikeya Dwivedi)
 
 - Add arena argument support to kfuncs and struct_ops through the new
   __arena and __arena__nullable suffixes (Tejun Heo, Puranjay Mohan,
   Kumar Kartikeya Dwivedi, Ihor Solodrai)
 
 - Signed BPF program loader rework to accommodate both BPF and security
   community needs where the kernel runs the signature verification at
   BPF_PROG_LOAD time before the LSM admission hook (Daniel Borkmann)
 
 - Add a set of ksock kfuncs which let BPF LSM and syscall programs
   create, connect and send on UDP sockets in order to emit telemetry
   data (Mahe Tardy)
 
 - Unify helper and kfunc call argument verification and classify kfunc
   arguments purely from BTF into a generated bpf_func_proto which is
   computed once at add-call time (Amery Hung)
 
 Other features and fixes:
 
 - Enable EXECMEM_ROX_CACHE for BPF allocations on x86 (Mike Rapoport)
 
 - Add bidirectional VLAN support to bpf_fib_lookup() through the new
   BPF_FIB_LOOKUP_VLAN and BPF_FIB_LOOKUP_VLAN_INPUT flags
   (Avinash Duduskar)
 
 - Infer zext_dst from static register liveness analysis to fix 32-bit
   zero-extension semantics, and remove the artificial limitations on
   pointer types eligible for spilling (Eduard Zingerman)
 
 - Inline the numeric open-coded iterator kfuncs so that bpf_for() loops
   no longer pay a kfunc call on every iteration (Puranjay Mohan)
 
 - Add an arena-based bitmap data structure to libarena along with
   serial and parallel selftests (Emil Tsalapatis)
 
 - Teach resolve_btfids to discover kfuncs from the kernel's BTF ID sets
   and to emit kfunc BTF decl tags, reducing the kernel build's
   dependency on pahole features (Ihor Solodrai)
 
 - Add BPF_F_ADJ_ROOM_DECAP_* flags to bpf_skb_adjust_room() so that
   tunnel decapsulation can update the GSO and encapsulation state of
   the skb (Nick Hudson)
 
 - Fix the ring buffer pending_pos walk and the available-data
   accounting on 32-bit position wrap (Israel Téllez García)
 
 - Add memory usage accounting for arena maps and fix an mmap_lock
   deadlock on arena lock failure (Jiayuan Chen)
 
 - Add tracing_multi link info support to the kernel UAPI and bpftool,
   and refactor the stack map code to run with preemption disabled
   (Jiri Olsa)
 
 - Support BPF_F_EGRESS in bpf_redirect_peer() to emit the skb in the
   egress direction of the target's peer device (Jordan Rife)
 
 - Add a KF_SPINLOCK_SAFE kfunc flag so that providers, in particular
   modules, can declare kfuncs safe to call under bpf_spin_lock instead
   of relying on the verifier's hard-coded allowlist (Kaitao Cheng)
 
 - Introduce global percpu data for BPF programs with libbpf probing
   and bpftool skeleton support, and stop exposing uninitialized kernel
   heap memory when copying per-CPU map values (Leon Hwang)
 
 - Add s390 JIT support for load-acquire and store-release instructions
   (Maxim Khmelevskii)
 
 - Fix a CFI mismatch in the task work callback and an arm64 KASAN
   false positive after bpf_throw() (Mykyta Yatsenko)
 
 - Reject writes through untrusted BTF pointers and bound the
   rdonly/rdwr_buf_size kfunc arguments (Nicholas Dudar)
 
 - Invalidate RCU pointers only after the final spin unlock and account
   for preempt and IRQ disabled regions as overlapping RCU protection
   (Ning Ding)
 
 - Support mixing bpf2bpf calls and tail calls on RV64, add signed
   operations and 32-bit atomics to the RV32 JIT, and add timed may_goto
   support (Pu Lehui, Kuan-Wei Chiu, Feng Jiang)
 
 - Fix a use-after-free on mm_struct in bpf_find_vma() for foreign tasks
   and an mmap_lock leak in the irq_work path (Sanghyun Park)
 
 - Populate mmap-able BPF array map memory lazily which makes mmap() O(1)
   instead of proportional to the map size (Song Liu)
 
 - Introduce a jit_required flag and reject programs with inlined
   helpers when no JIT is available, where the interpreter would
   otherwise jump into an invalid address (Tiezhu Yang)
 
 - Fix the x86 JIT per-CPU address resolution into an extended register
   where the REX prefix dropped the high destination register bit
   (Vineet Gupta)
 
 - Reject MEM_ALLOC BTF accesses past object bounds, arena frees below
   the arena base, and mixed arena and ordinary atomic paths
   (Yiyang Chen)
 
 - Fix the trampoline handling of 128-bit arguments and of return values
   larger than 8 bytes (Yonghong Song)
 
 - Ensure that any fault prone load is rewritten with exception table
   handling, and fix the arena load-acquire and atomic fetch handling
   in the x86, arm64, riscv and s390 JITs (Daniel Borkmann)
 
 - Many more fixes and cleanups across the verifier, arena, trampolines,
   sockmap, cgroup, ring buffer, x86/arm64/riscv/s390 JITs, libbpf,
   bpftool, resolve_btfids and selftests.
 
 Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
 -----BEGIN PGP SIGNATURE-----
 
 iIsEABYKADMWIQTFp0I1jqZrAX+hPRXbK58LschIgwUCaoNzBBUcZGFuaWVsQGlv
 Z2VhcmJveC5uZXQACgkQ2yufC7HISIOb3QEAy5cyrLXY+VWofhsC9wULkHyETOdj
 oTkdohQomZp4VhEA/1RZXdHVS1ANFgreWv0fMorUOHEKv2ZuNokfk3LWgW4L
 =VRyL
 -----END PGP SIGNATURE-----

Merge tag 'bpf-next-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf-next

Pull bpf updates from Daniel Borkmann:
 "Major changes:

   - Redesign the verifier error reporting: failures now carry source
     and instruction annotations along with the causal event history
     that led to them, making program rejections far easier to debug and
     repair (Kumar Kartikeya Dwivedi)

   - Add arena argument support to kfuncs and struct_ops through the new
     __arena and __arena__nullable suffixes (Tejun Heo, Puranjay Mohan,
     Kumar Kartikeya Dwivedi, Ihor Solodrai)

   - Signed BPF program loader rework to accommodate both BPF and
     security community needs where the kernel runs the signature
     verification at BPF_PROG_LOAD time before the LSM admission hook
     (Daniel Borkmann)

   - Add a set of ksock kfuncs which let BPF LSM and syscall programs
     create, connect and send on UDP sockets in order to emit telemetry
     data (Mahe Tardy)

   - Unify helper and kfunc call argument verification and classify
     kfunc arguments purely from BTF into a generated bpf_func_proto
     which is computed once at add-call time (Amery Hung)

  Other features and fixes:

   - Enable EXECMEM_ROX_CACHE for BPF allocations on x86 (Mike Rapoport)

   - Add bidirectional VLAN support to bpf_fib_lookup() through the new
     BPF_FIB_LOOKUP_VLAN and BPF_FIB_LOOKUP_VLAN_INPUT flags (Avinash
     Duduskar)

   - Infer zext_dst from static register liveness analysis to fix 32-bit
     zero-extension semantics, and remove the artificial limitations on
     pointer types eligible for spilling (Eduard Zingerman)

   - Inline the numeric open-coded iterator kfuncs so that bpf_for()
     loops no longer pay a kfunc call on every iteration (Puranjay
     Mohan)

   - Add an arena-based bitmap data structure to libarena along with
     serial and parallel selftests (Emil Tsalapatis)

   - Teach resolve_btfids to discover kfuncs from the kernel's BTF ID
     sets and to emit kfunc BTF decl tags, reducing the kernel build's
     dependency on pahole features (Ihor Solodrai)

   - Add BPF_F_ADJ_ROOM_DECAP_* flags to bpf_skb_adjust_room() so that
     tunnel decapsulation can update the GSO and encapsulation state of
     the skb (Nick Hudson)

   - Fix the ring buffer pending_pos walk and the available-data
     accounting on 32-bit position wrap (Israel Téllez García)

   - Add memory usage accounting for arena maps and fix an mmap_lock
     deadlock on arena lock failure (Jiayuan Chen)

   - Add tracing_multi link info support to the kernel UAPI and bpftool,
     and refactor the stack map code to run with preemption disabled
     (Jiri Olsa)

   - Support BPF_F_EGRESS in bpf_redirect_peer() to emit the skb in the
     egress direction of the target's peer device (Jordan Rife)

   - Add a KF_SPINLOCK_SAFE kfunc flag so that providers, in particular
     modules, can declare kfuncs safe to call under bpf_spin_lock
     instead of relying on the verifier's hard-coded allowlist (Kaitao
     Cheng)

   - Introduce global percpu data for BPF programs with libbpf probing
     and bpftool skeleton support, and stop exposing uninitialized
     kernel heap memory when copying per-CPU map values (Leon Hwang)

   - Add s390 JIT support for load-acquire and store-release
     instructions (Maxim Khmelevskii)

   - Fix a CFI mismatch in the task work callback and an arm64 KASAN
     false positive after bpf_throw() (Mykyta Yatsenko)

   - Reject writes through untrusted BTF pointers and bound the
     rdonly/rdwr_buf_size kfunc arguments (Nicholas Dudar)

   - Invalidate RCU pointers only after the final spin unlock and
     account for preempt and IRQ disabled regions as overlapping RCU
     protection (Ning Ding)

   - Support mixing bpf2bpf calls and tail calls on RV64, add signed
     operations and 32-bit atomics to the RV32 JIT, and add timed
     may_goto support (Pu Lehui, Kuan-Wei Chiu, Feng Jiang)

   - Fix a use-after-free on mm_struct in bpf_find_vma() for foreign
     tasks and an mmap_lock leak in the irq_work path (Sanghyun Park)

   - Populate mmap-able BPF array map memory lazily which makes mmap()
     O(1) instead of proportional to the map size (Song Liu)

   - Introduce a jit_required flag and reject programs with inlined
     helpers when no JIT is available, where the interpreter would
     otherwise jump into an invalid address (Tiezhu Yang)

   - Fix the x86 JIT per-CPU address resolution into an extended
     register where the REX prefix dropped the high destination register
     bit (Vineet Gupta)

   - Reject MEM_ALLOC BTF accesses past object bounds, arena frees below
     the arena base, and mixed arena and ordinary atomic paths (Yiyang
     Chen)

   - Fix the trampoline handling of 128-bit arguments and of return
     values larger than 8 bytes (Yonghong Song)

   - Ensure that any fault prone load is rewritten with exception table
     handling, and fix the arena load-acquire and atomic fetch handling
     in the x86, arm64, riscv and s390 JITs (Daniel Borkmann)

   - Many more fixes and cleanups across the verifier, arena,
     trampolines, sockmap, cgroup, ring buffer, x86/arm64/riscv/s390
     JITs, libbpf, bpftool, resolve_btfids and selftests"

* tag 'bpf-next-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf-next: (373 commits)
  selftests/bpf: Add tests for a store on a fault prone qdisc pointer
  selftests/bpf: Add tests for fault prone loads out of RCU pointers
  selftests/bpf: Add tests for pointer type merge at a shared load
  selftests/bpf: Remove duplicate copies of the arena spinlock qnodes
  selftests/bpf: Retry stat generation in cgroup_iter_memcg
  selftests/bpf: Test pseudo-function policy diagnostics
  bpf: Distinguish function references in policy diagnostics
  bpf: Preserve source attribution without source text
  selftests/bpf: Test kfunc argument diagnostics
  bpf: Correct kfunc argument diagnostics
  bpf: Use canonical stack argument names in diagnostics
  bpf: Preserve R0 lineage across helper calls
  selftests/bpf: Exercise negative optlen in cgroup getsockopt hook
  bpf: Reject negative optlen in cgroup getsockopt hook
  selftests/bpf: tc_tunnel - validate decap GSO and encapsulation state
  bpf: Clear decap state on skb_adjust_room shrink path
  bpf: Allow new DECAP flags and add guard rails
  bpf: Add BPF_F_ADJ_ROOM_DECAP_* flags for tunnel decapsulation
  bpf: Refactor masks for ADJ_ROOM flags and encap validation
  bpf: Name the enum for BPF_FUNC_skb_adjust_room flags
  ...
2026-08-20 07:36:20 -07:00
Linus Torvalds
a4ff2be345 This update includes the following changes:
API:
 
 - Add af_alg_restrict sysctl and white list.
 - Fix potential suspend/resume races in hwrng.
 
 Algorithms:
 
 - Optimize vli additive operations using compiler builtins in ecc.
 
 Drivers:
 
 - Remove unsafe/deprecated algorithms from qce.
 - Mark qce as BROKEN.
 - Add runtime PM and interconnect bandwidth scaling support to qce.
 - Remove crypto_rng from qcom, sun8i and caam.
 - Fix SG list issues in iaa.
 - Fix SEV init path bugs in ccp.
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEn51F/lCuNhUwmDeSxycdCkmxi6cFAmqEC3gACgkQxycdCkmx
 i6fOCw/9EzjhD0xotKv0Kylk/ukG+UYhh2j1zwFbMUuRl0GQCklSfM19h0DQ53vS
 FgJqf+69Q9sfn2HxihxQJGkW+NmiqwlHG9veu2PRBXRnMDjFQ8LDHAVEShvL4Jzv
 PW9daF3KsOjlcFuOVHum9SdQ2tdsoClEtBv8W9ndOxxGxGj3827etOWSTTp4DprP
 y2bpcE3R+CjmOgATmAQfOiKdLtghv8SspSRUmwVmj8lVijkjTiH9UELwQ4Tp407q
 vZU027gBHaWKb4VBLPX3NUg0UaJFsieKGrty2EZqX0nXF4f2jaPYB1iP9Y+92qcP
 WfvidNpVeLpRlJpf5QH1ZqiH7qf9I1YdOXcNe8IL+3b+9SYiyqDY3vuCeBE05dOJ
 Oty9m8pIV7IwZmhUIhZ0PdIl58urzxYPvCdD0IdAsA0sdNQEBcbTnuVuQjxt4AG0
 GYPqiZdtRs5r3MkdRpymV49TBJ+vMY3Wo1lnCcnpCgxugTAx7tkFtMNBJSHF7DbC
 N1vwb2EYaTyzmH5Vr3dHLWeONskyIRa0WNhnszmIBie3M0AYF2sOkNw3iXVPQxXv
 a5XbXh/kAx0nc417dp1B8lZclHH2bWvEKHYalpT33GX4qsGuI4bH4uKaOrL0pXp1
 WyNzGiCiWJBp2o9wNj/TxKCp7oGeD4ZRv0HeRrEHezSInoTrt9w=
 =f9oK
 -----END PGP SIGNATURE-----

Merge tag 'v7.3-p1' of git://git.kernel.org/pub/scm/linux/kernel/git/herbert/crypto-2.6

Pull crypto update from Herbert Xu:
 "API:
   - Add af_alg_restrict sysctl and white list
   - Fix potential suspend/resume races in hwrng

  Algorithms:
   - Optimize vli additive operations using compiler builtins in ecc

  Drivers:
   - Remove unsafe/deprecated algorithms from qce
   - Mark qce as BROKEN
   - Add runtime PM and interconnect bandwidth scaling support to qce
   - Remove crypto_rng from qcom, sun8i and caam
   - Fix SG list issues in iaa
   - Fix SEV init path bugs in ccp"

* tag 'v7.3-p1' of git://git.kernel.org/pub/scm/linux/kernel/git/herbert/crypto-2.6: (122 commits)
  crypto: lskcipher - propagate errors from unaligned crypt
  crypto: keembay - use crypto_memneq() to compare CCM AEAD tags
  crypto: keembay - use crypto_memneq() to compare GCM AEAD tags
  crypto: sa2ul - use crypto_memneq() to compare AEAD tag
  hwrng: drivers - use named initializers for acpi_device_id
  crypto: qce - fix CCM AAD buffer underallocation
  crypto: iaa - unmap dst before software fallback on decompress
  crypto: iaa - use bounce buffer for multi-sg decompress input
  crypto: iaa - avoid counting fallback decompression bytes
  crypto: iaa - fall back to software for multi-entry scatterlists
  hwrng: core - Stop/start hwrng_fillfn() kthread before/after suspend-resume
  crypto: hisilicon/sec2 - fix CCM algorithm long packet failure
  crypto: eip93 - use struct_size() and flexible array for ring allocation
  crypto: krb5 - use kfree_sensitive() for derived key buffers
  crypto: af_alg - Stop after finding name in allowlist
  crypto: af_alg - Replace 'bool privileged' with flags
  crypto: af_alg - Make cbc(paes) privileged-only
  hwrng: imx-rngc - Disable clock on registration failure
  crypto: qat - remove dead ADF_HEX code
  crypto: qce - simplify qce_handle_request
  ...
2026-08-19 17:25:42 -07:00
Linus Torvalds
a51ec5e8e5 integrity-v7.3
-----BEGIN PGP SIGNATURE-----
 
 iIoEABYKADIWIQQdXVVFGN5XqKr1Hj7LwZzRsCrn5QUCaoTvgxQcem9oYXJAbGlu
 dXguaWJtLmNvbQAKCRDLwZzRsCrn5dC1AP9Byw49KACrNM8vvic/MTB4i3azMbFU
 ZNbcXgj9xRwyngD/fjwktpYUwEvgMIDybIu0Tr+CXUwoI8vrmPZpsvkeUwI=
 =eSwF
 -----END PGP SIGNATURE-----

Merge tag 'integrity-v7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/zohar/linux-integrity

Pull integrity updates from Mimi Zohar:

 - TPM initialization is sometimes delayed until deferred_probe_initcall

   Since ordering is not guaranteed within the same initcall level, IMA
   may initialize before the TPM and fall back to TPM-bypass mode. A new
   config option, CONFIG_IMA_INIT_LATE_SYNC, allows those building the
   kernel to defer IMA initialization to late_initcall_sync, accepting
   the integrity risk of missing early measurements in exchange for
   avoiding TPM-bypass mode.

 - The raw policy rules are now measured, as well as the complete
   policy, closing a gap in integrity measurement coverage

* tag 'integrity-v7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/zohar/linux-integrity:
  ima: measure userspace policy writes before parsing
  ima: add critical data measurement for loaded policy
  security: ima: rename boot_aggregate when ima is initialised at late_sync
  security: ima: introduce IMA_INIT_LATE_SYNC option
  security: lsm: allow LSMs to register for late_initcall_sync init
2026-08-19 17:17:17 -07:00
Linus Torvalds
cbad8981fa Patches for v7.3
-----BEGIN PGP SIGNATURE-----
 
 iQJLBAABCAA1FiEEC+9tH1YyUwIQzUIeOKUVfIxDyBEFAmp+ClEXHGNhc2V5QHNj
 aGF1Zmxlci1jYS5jb20ACgkQOKUVfIxDyBG85g/9GPmiweiDUjMq3dlrAzJ4vi4f
 h9+Eu4y7Relg7wMvXnOuc3ZB/QxGKSHBlzxwfCIoqtAC9wFbbqP4hhJ0c6yladCL
 M0fmJT/JRZcIZdHzwu2G5skNxwtms9sFyyBi4Pq6YkA/fszHig5vZJquVIP0mIZV
 DF3dsOC2fw8g/kSe4OcZF+0qZABaY/LHrZTS/Mit9GQa8AVS0CLFubrbE2pOSQJ5
 TwLXtqOZQVWq814kzFmdOomQNdzR35HH30EXgan5Q06oN25i6YbL1RZkOXKhLn4t
 eGDT33wKiEt9M7x0R4MZOFnkAKgs8hDtBvBd1/lvkvU7rZkG7FWH+weh9CiFtyoT
 LkbPxVmdFI1IZCOaxqB6Pl/ZKy7QQbm0Hy3TEL7vizhSbh5qOZhjwTjESvOQgGA5
 ynUkJIuECYfyXyFJSiA1QRGR9lOwgXWa8qJYl71GrTCgGu/CDraDARkOLgM+afih
 WU4TZsTlwOJSP0UhfmhAzMxNs7K2YjYj2tvDhV0aKMz7WLLd+ivJIYNsupfftqO6
 iCWjshncvF4NrULumErLcm9BBy9SjUH4JansLSfh52TI4XfKP9sEQSkCCLbOK5J4
 uGAotKWUIyAL+cfOrMCmViFMopCe7LcrDFsF5ocT84wRDT5K8ljnMm8lpKzrfrh2
 N1OA5EHi88OfmVnules=
 =ucTi
 -----END PGP SIGNATURE-----

Merge tag 'Smack-for-7.3' of https://github.com/cschaufler/smack-next

Pull smack updates from Casey Schaufler:

 - Spelling fix

 - Code optimization in smackfs

 - Fix credential mis-uses

 - Place limits on two of the smackfs interfaces

* tag 'Smack-for-7.3' of https://github.com/cschaufler/smack-next:
  smack: fix cred UAF in smack_file_send_sigiotask()
  smack: restrict smackfs/{direct,mapped} values to 0-255
  smack: deduplicate smackfs/{direct,mapped} file_operations
  smack: show msgrcv() subject task in audit
  smack: fix incorrect task context in smack_msg_queue_msgrcv
  security: smack: fix spelling mistake
  smack: simplify write handlers of sysfs entries
  Smack: Fix error in capability bypass
2026-08-19 16:51:39 -07:00
Linus Torvalds
09005a6398 lsm/stable-7.3 PR 20260814
-----BEGIN PGP SIGNATURE-----
 
 iQJIBAABCgAyFiEES0KozwfymdVUl37v6iDy2pc3iXMFAmp/iWkUHHBhdWxAcGF1
 bC1tb29yZS5jb20ACgkQ6iDy2pc3iXPvbQ/8DT62doPc0ECTXQXcTNjBOddjNspB
 pk2mK238UQP50aU7/su4RdgmGG+spVoPc7oqeavnm2J+c552t0eHI61itYMe7nkY
 uOIjShLN93g9pjG4IqhmCDvGTpsQp9Oiec5F/6C++7OUT5oUqm/faXAZtwFLFgmx
 pTNf91w+u5s3DJjqG5zqEdteRrQMzDNozdq4YbNkzeIiofsUJvq6IJ6rV66kyOa/
 zk2hC6zI3mL0Vuin3WdxKqd1mD7mhYWxjl2nt/0TTIRMKmycxovlvbfADXXi2hlp
 4tALFgbEmeOgXr7HWvOqClNZ002/gG84Ty2F/8D/p0EGlzOc0Ga/q1dSAM/JJ5RZ
 bArXY3qgioU4/zOp998kXIYgRV3i6ZBBEzqTavcVtgCtNTgc2MH+KftD9HJaVjx8
 keRClwB85r72mNqKVcTVgQ8MJvA+Fq03JHc/JH3npv2tSPELK4V4vkW0nCGQsLXV
 /oX87E9/Wh6wWchvUTYToX5j6eNRUjD8xQzIAbPWTilusPcIp9pW58I6gaDKAu6A
 dNox9I064JfLK+LEmn8Le8AatTT0g/mj1wTD0KeRfo0zKyjSYoT0ECnXBeAfg9QI
 Vr6ZTXPM0fMpKEHqJvYLJAqVbm3kJKZqxr1z4nmz14elms1SmwdmOSq66Qm8fRCS
 P5fGVxFzNbDVZvs=
 =eToU
 -----END PGP SIGNATURE-----

Merge tag 'lsm-pr-20260814' of git://git.kernel.org/pub/scm/linux/kernel/git/pcmoore/lsm

Pull LSM updates from Paul Moore:

 - Remove task_euid()

   The task_euid(), and Rust counterpart, was never widely used, for
   good reason, and now that the only user is gone we're removing it to
   rid ourselves of both dead and funky code.

 - Documentation improvements

   Correct some of the kdoc comments for security_task_prctl() and
   clarify the rust comments on task UID accessors.

 - Fix a memory leak in the LSM syscall selftests

* tag 'lsm-pr-20260814' of git://git.kernel.org/pub/scm/linux/kernel/git/pcmoore/lsm:
  selftests/lsm: Fix memory leak in attr_lsm_count
  cred: delete task_euid()
  rust: task: clarify comments on task UID accessors
  lsm: clarify security_task_prctl() hook documentation
2026-08-19 16:28:23 -07:00
Linus Torvalds
4253eb09d2 selinux/stable-7.3 PR 20260814
-----BEGIN PGP SIGNATURE-----
 
 iQJIBAABCgAyFiEES0KozwfymdVUl37v6iDy2pc3iXMFAmp/iXIUHHBhdWxAcGF1
 bC1tb29yZS5jb20ACgkQ6iDy2pc3iXN+Wg/7B8/owEGutN2OwRQKfo/qIjYsCp2B
 a4mhT/K5JTVYYTr5PsbVb8OkRrIdpZhLTOGzRN7e16o7HrV77szJ2zeiizlHp7q2
 t7iuwWtLbF9q2qlHEg2/p+VZvfq82bBPgLHwzfuXh+2TMVUHWHw//gw90Aqjg7np
 BNxagAdTM5zodv4OqNnwUkve9Y90VTBD3pbsFrn15lgh3efovg9ya2Xy/sJlSVQG
 8yGJtFxj+CodsLBW76G3/jjOiTC1/XMGacZYgHWIaZsxr2EHjmCZdU6jdoZxqXOi
 iG45yaf/HOIMnuSoVq6R5PG4I3paxa9Po75z3NWBdxQQqHOiAguYG8pHdxPxrDv2
 QOkyHai04/x9XCqiyA91TflTflGZOdio5M9YHbUv2FjNyMsenEdjuT1prUpsLZ/Y
 1dl0iLDuotS3l/NRtm0Yc+epl8N6GeCB0UCJK1AHSAZJt+twOy0jJfRvUIVfpw9B
 cQgkyegEKlm7PX2nzre/A727H6CrnZsneYjQeE1/+0szmRDKfYEOghSuw3FmKyqr
 1gAm851G27ySF+75h9SiNLmdJgkHeOGa5pxO+qa2ZtZFXfK9ie3317i2t7Aoaiec
 6dfHTF5JsfdHcqjESuD+ibVbPpvGC4SFkWd3RbogiNhREbt4Sy2pw3j4WxUND9bv
 f1hVBz4/SyV9M+4=
 =6vH+
 -----END PGP SIGNATURE-----

Merge tag 'selinux-pr-20260814' of git://git.kernel.org/pub/scm/linux/kernel/git/pcmoore/selinux

Pull selinux updates from Paul Moore:

 - Convert a __get_free_page() call into a kmalloc() call

   We had some very old code that called out to __get_free_page() for
   allocating a pathname. There is no reason this couldn't be done with
   a call to kmalloc() so we've done the conversion and now there is one
   less __get_free_page() caller in the kernel.

 - Limit the number of retired/unknown DCCP netlink messages

   While DCCP is gone from the kernel, there are still userspace tools
   which try to talk to the kernel about DCCP sockets which were
   generating SELinux related log noise (unrecognized netlink message).
   This pull request both limits the log messages to just the first
   instance and also explains to the user that DCCP support has been
   removed.

 - Convert the SELinux strlcat() calls to seq_buf_XXX() calls

   As part of the effort to drop the strlcat() API from the kernel, the
   SELinux/IMA code was converted over to using seq_buf_XXX() calls.

 - Only calculate the SELinux IMA configuration string length once

   Previously each call to generate a SELinux configuration string for
   IMA would have to calculate the length of the string. While the
   contents of the string will likely change over the lifetime of the
   system, the length of the string will not. Calculate the string
   length once at boot and reuse the length value throughout the
   lifetime of the system.

 - Further validation of the SELinux policy at policy load time

   Perform additional sanity checks on the policy constraints and types.

 - Proper cleanup and error handling for selinuxfs init failures

   We were not properly cleaning up some state in the case where
   selinuxfs fails to initialize properly. It's somewhat of an academic
   exercise as a failure to initialize selinuxfs will cause the system
   to fail on boot, but it's arguably better to make sure we do things
   the proper way.

 - Various code cleanups

   Convert integer flags to boolean types and drop an uncessary goto
   from the SELinux code.

* tag 'selinux-pr-20260814' of git://git.kernel.org/pub/scm/linux/kernel/git/pcmoore/selinux:
  selinux: validate constraint expression attr and op at load time
  selinux: compute the IMA configuration settings string length once at boot
  selinux: replace strlcat() with seq_buf in selinux_ima_collect_state()
  selinux: suppress warning flood for retired DCCP netlink messages
  selinux: tighten type validation during policy load
  selinux: drop unnecessary goto and label from avc_alloc_node()
  selinux: convert int flags to bool flags in ss/services.c
  selinux: clean up selinuxfs resources on init failure
  selinux: hooks: use kmalloc() to allocate path buffer
2026-08-19 16:24:46 -07:00
Linus Torvalds
83453b6f51 audit/stable-7.3 PR 20260814
-----BEGIN PGP SIGNATURE-----
 
 iQJIBAABCgAyFiEES0KozwfymdVUl37v6iDy2pc3iXMFAmp/iW4UHHBhdWxAcGF1
 bC1tb29yZS5jb20ACgkQ6iDy2pc3iXP5gw/9FZSIJurmLZ9s+GPWczZFvkOB5aA9
 jcBy7qCcRLrzlCIzrb9X8yBvgfRuGZcXgUiY9yCLYLJfeo9CECfYtSaqSN+3lBgg
 0rTFjRmFijxc2m/xcimCxeh+5jMymWs/h7eqI8uPH5mrK65Ox2s2x9dCyHYHlvJ2
 /Gl9igndDJ8I8OfHN4lEljSWXai2tONnWe4BFRrkcFUm6MwI6IKpRVVke1Pi6yKz
 cEij/A3VIpVXuH+AYCnctBNrz/voKcU7VjK+opuaBG5Tx/R2g6pWsC8jGjsu0uiP
 VTMhPmwaFdoTmCnt8zrqrrBaNwRqKypIKMdWKd2g0EcnLE7qv4HNTTPCm9WQya7t
 UjBtBArytTbPg7TaIl5KP4/I18ZjFTHMVuAOjyjZUvWm/Sl4lf+V2/x8Hkh7nwbB
 ffOuqnMS1+f4L/GKUNgBG5eHtOkNa+f2ZbtMdvHU8D55dP+k5cK/2lGDPPWtsFoQ
 jdsgcBG9sGp6pWytacZ/se4vd3wRFeCMbRsntBRYaGJN4zNOf55fZPgmGhtx+30O
 r0K0SiXmc/mqNake/f8rqwUar3Pqd+lj3rmEi1uNzqAq+VPlkuGwhqRmf1mS+/26
 4n1wQsTp4yuR6wQ7PYyA+/TzeIYtTxRvMLA2lGdGMSe5SY8+D4I40yy7KIvK/yaC
 wHh9rRffxFx04Ek=
 =9fzH
 -----END PGP SIGNATURE-----

Merge tag 'audit-pr-20260814' of git://git.kernel.org/pub/scm/linux/kernel/git/pcmoore/audit

Pull audit updates from Paul Moore:

 - Drop BUG_ON() assertions from two functions

   While I don't recall any bug reports from either of these assertions
   in recent memory, neither of these checks warrant the kernel panic
   that could result from BUG_ON(). One of the BUG_ON() calls is
   converted to a WARN_ON_ONCE() and the other to a lockdep assertion.

 - Fix an audit tree reference counting problem

   Fix a corner case where audit could end up unintentionally dropping
   the last reference to an audit tree while the tree was still in use.

   We should probably revisit the audit tree handling code in full, but
   this patch works, and should be easy to backport to stable trees and
   downstream kernels.

 - Update the audit syscall classification tables

   Add some missing syscalls to the PERM class

* tag 'audit-pr-20260814' of git://git.kernel.org/pub/scm/linux/kernel/git/pcmoore/audit:
  audit: avoid dropping live tree ref on fsnotify rule autoremove
  audit: drop BUG_ON() from audit_signal_info_syscall()
  audit: drop BUG_ON() from audit_add_to_parent()
  audit: add missing syscalls to PERM class tables
2026-08-19 16:21:32 -07:00
Linus Torvalds
cb8a75eec0 ring-buffer updates for 7.3:
- Remove unneeded semicolon
 
   A macro ended with a semicolon that wasn't needed.
 
 - Fix freeing cpu_buffer extra subbuffer with order greater than zero
 
   When the cpu_buffer was being freed, its "free" page, was using
   free_page() to free it when it could be more than one page.
 
 - Hold the cpu_buffer lock when resizing the subbuffer
 
   The freeing of the "free" page of the cpu_buffer was done without locking.
   The order of the data was being saved and then the "free" page was set to
   NULL. But there is a race that the "free" page could have been updated
   between those two operations. Add locking around it to prevent the race.
 
 - Save the order of the data along with the data in the free page
 
   The cpu_buffer would store just the data portion of the subbuffer page in
   its descriptor. But it did not store the order of the data pages. The order
   was being saved in the global buffer descriptor. But this leads to races.
 
   Have the cpu_buffer save the subbuf data along with its metadata (which
   includes the order of the page) to make sure when it frees it, it frees
   the correct order along with it.
 
 - Remove the subbuf_size and use the order directly when needed
 
   Having a size field for the size of the subbufer along with its order
   allowed for races to have them get out of sync. Remove the subbuf_size and
   use the order from the subbuf meta data directly under locks.
 
   Use the subbuf_order for other calculations in the ring buffer.
 
 - Remove the useless "cpus" field of trace_buffer
 
   The code has been restructured and the "cpus" field is no longer used.
   Remove it.
 
 - Remove the "mapped" field of the ring buffer and use a helper function instead.
 
   The "mapped" field has become a bit overused and made the code come
   complex in using a counter for what is denoted as being mapped or not.
   There are other fields that are set when the ring buffer is considered
   mapped. Add a helper function to check those fields and use that instead
   of keeping track of a counter.
 -----BEGIN PGP SIGNATURE-----
 
 iIoEABYKADIWIQRRSw7ePDh/lE+zeZMp5XQQmuv6qgUCaoC9jxQccm9zdGVkdEBn
 b29kbWlzLm9yZwAKCRAp5XQQmuv6qlJwAQDG/FAI6Zr4f2jUIoEWPL7KGkhmHeuv
 rP1bIJVeIoy+RgEA+vjq6PNNGvN2DO0qnotu5UAhHxywM1KaUKQjOCDaJQI=
 =f8Ym
 -----END PGP SIGNATURE-----

Merge tag 'trace-ringbuffer-v7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace

Pull ring-buffer updates from Steven Rostedt:

 - Remove unneeded semicolon

   A macro ended with a semicolon that wasn't needed.

 - Fix freeing cpu_buffer extra subbuffer with order greater than zero

   When the cpu_buffer was being freed, its "free" page, was using
   free_page() to free it when it could be more than one page.

 - Hold the cpu_buffer lock when resizing the subbuffer

   The freeing of the "free" page of the cpu_buffer was done without
   locking. The order of the data was being saved and then the "free"
   page was set to NULL. But there is a race that the "free" page could
   have been updated between those two operations. Add locking around it
   to prevent the race.

 - Save the order of the data along with the data in the free page

   The cpu_buffer would store just the data portion of the subbuffer
   page in its descriptor. But it did not store the order of the data
   pages. The order was being saved in the global buffer descriptor. But
   this leads to races.

   Have the cpu_buffer save the subbuf data along with its metadata
   (which includes the order of the page) to make sure when it frees it,
   it frees the correct order along with it.

 - Remove the subbuf_size and use the order directly when needed

   Having a size field for the size of the subbufer along with its order
   allowed for races to have them get out of sync. Remove the
   subbuf_size and use the order from the subbuf meta data directly
   under locks.

   Use the subbuf_order for other calculations in the ring buffer.

 - Remove the useless "cpus" field of trace_buffer

   The code has been restructured and the "cpus" field is no longer
   used. Remove it.

 - Remove the "mapped" field of the ring buffer and use a helper
   function instead.

   The "mapped" field has become a bit overused and made the code come
   complex in using a counter for what is denoted as being mapped or
   not. There are other fields that are set when the ring buffer is
   considered mapped. Add a helper function to check those fields and
   use that instead of keeping track of a counter.

* tag 'trace-ringbuffer-v7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace:
  ring-buffer: Remove ring_buffer_per_cpu::mapped
  ring-buffer: Remove trace_buffer::cpus
  ring-buffer: Dynamically calculate max_data_size
  ring-buffer: Fix subbuf resize race with ring_buffer_alloc_read_page()
  ring-buffer: Fix subbuf resize race with ring buffer readers
  ring-buffer: Make cpu_buffer::free_page a buffer_data_read_page
  ring-buffer: Hold cpu_buffer::lock when resizing a subbuf
  ring-buffer: Free cpu_buffer::free_page with subbuf_order
  ring-buffer: drop unneeded semicolon
2026-08-19 14:22:07 -07:00