Commit Graph

1481676 Commits

Author SHA1 Message Date
Jakub Kicinski
cdb719f4b8 net: dsa: bcm_sf2: bound the CFP rule dump by the caller's buffer size
bcm_sf2_cfp_rule_get_all() walks the whole cfp.unique bitmap into
rule_locs[] without consulting nfc->rule_cnt, which is how many entries
the caller had room for.  ETHTOOL_GRXCLSRLALL requires no CAP_NET_ADMIN
and the ioctl sizes the buffer from the rule_cnt userspace passes in, so
once an admin has installed CFP rules any user can ask for fewer slots
than there are rules and run off the end of the allocation.  A rule_cnt
of 0 leaves the buffer pointer NULL and the walk dereferences it.

Fixes: 7318166cac ("net: dsa: bcm_sf2: Add support for ethtool::rxnfc")
Reviewed-by: Jonas Gorski <jonas.gorski@gmail.com>
Reviewed-by: Florian Fainelli <florian.fainelli@broadcom.com>
Reviewed-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260903032611.3000029-2-kuba@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:19:57 -07:00
Eric Dumazet
1746ef2e2d bonding: use skb_cow_head() in bond_do_alb_xmit() and rlb_arp_xmit()
In bond_do_alb_xmit() and rlb_arp_xmit(), make sure to unclone
skb head via skb_cow_head() before modifying the source MAC address
(Ethernet header and ARP payload) to avoid silent corruption if
the skb is shared or cloned. Avoid caching the header pointers
across skb_cow_head().

In rlb_arp_xmit(), only modify arp->mac_src if it differs from
tx_slave->dev->dev_addr to avoid an unnecessary copy and head
reallocation.

Also, we should not assume mac header is set in output path.

Use skb_eth_hdr() instead of eth_hdr() to fix the issue,
and remove now redundant skb_reset_mac_header() calls.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Cc: Jay Vosburgh <jv@jvosburgh.net>
Reviewed-by: Nikolay Aleksandrov <razor@blackwall.org>
Link: https://patch.msgid.link/20260903143940.1180513-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:18:41 -07:00
Ido Schimmel
5bd9e4e7cd nexthop: Initialize extack in remove_nh_grp_entry()
remove_nh_grp_entry() prints the extack message when a listener fails
to replace the reduced nexthop group. However, extack is not
initialized and listeners are not required to set a message when
returning an error. Neither netdevsim nor mlxsw do so when an
allocation fails, resulting in the dereference of an uninitialized
stack pointer.

Fix by zero-initializing extack, as was done in commit 6347c5314c
("nexthop: initialize extack in nh_res_bucket_migrate()").

Fixes: 833a1065ee ("nexthop: Emit a notification when a nexthop group is reduced")
Signed-off-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260903080259.10378-1-idosch@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:17:41 -07:00
Jakub Kicinski
fb3088dc58 Merge tag 'ieee802154-for-net-2026-09-03' of git://git.kernel.org/pub/scm/linux/kernel/git/wpan/wpan
Stefan Schmidt says:

====================
pull-request: ieee802154 for net 2026-09-03

Zhiling Zou fixed a NULL deref when coming from a TUN device.

Fan Wu fixed a UAF in the cc2520 driver.

Chenguang Zhao fixed up some out of date comments in 6lowpan.

David Carlier fixed a potential double free in the hwsim driver.

Ibrahim Hashimov reworked the queuing in the RX path to fix a UAF on beacon
and MAC frames.

* tag 'ieee802154-for-net-2026-09-03' of git://git.kernel.org/pub/scm/linux/kernel/git/wpan/wpan:
  mac802154: fix use-after-free of sdata via queued RX frames
  ieee802154: hwsim: serialize pib updates to fix double-free
  ieee802154: 6lowpan: fix NULL dereference in lowpan_newlink
  ieee802154: cc2520: fix FIFOP work use-after-free
  net: 6lowpan: fix mismatched comments
====================

Link: https://patch.msgid.link/20260903093012.4032586-1-stefan@datenfreihafen.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:16:28 -07:00
Jakub Kicinski
641d03105c Merge branch 'enic-fix-v2-vf-mailbox-reply-matching-and-carrier-reopen'
Satish Kharat says:

====================
enic: fix V2 VF mailbox reply matching and carrier reopen

Preserve ENIC V2 VF carrier state across netdev reopen and match mailbox
replies to their originating requests.

A V2 VF receives carrier state from PF MBOX notifications. enic_stop()
forces carrier off, but enic_open() does not request another notification
or restore the previous state. An ordinary netdev close/open can therefore
leave the VF without carrier until the PF sends another link-state
notification.

The first patch preserves the last PF-reported link state across an
ordinary netdev close/open. Internal reset paths invalidate the saved state
before reopening the datapath, so carrier remains off until the PF provides
a new notification.

Mailbox messages carry a message number that replies and acknowledgments
echo. ENIC currently assigns a new number to outgoing replies and accepts
VF replies by message type alone. After a request times out, a delayed
reply can therefore satisfy a later request of the same type.

The second patch makes PF replies and the VF link-state acknowledgment echo
the initiating message number. The VF accepts a reply only when both its
type and message number match the pending request. Reply handling and
timeout cleanup are protected by the same lock, and message numbering
remains monotonic across admin-channel reopen.

Validation:

- The VF module with this series applied passed multiple netdev close/open
  and ENIC module unload/reload cycles, as well as guest reboot testing,
  with carrier restored as expected after each operation.
- With this series folded into the full SR-IOV development stack, VFIO
  guests passed bidirectional cross-DUT PF-to-VF, VF-to-PF, and VF-to-VF
  traffic using standard-sized and 8972-byte jumbo ICMP packets.
====================

Link: https://patch.msgid.link/20260830-b4-enic-v2-mbox-fixes-net-v1-0-23adf9bfd426@cisco.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 19:10:20 -07:00
Satish Kharat
8972d252f4 enic: match mailbox replies to request numbers
The version-1 VF mailbox protocol identifies every message with a message
number, and a reply or acknowledgment echoes the number of the message it
answers.  ENIC instead generates a new number for outgoing replies and
accepts a VF reply by message type alone.

If a request times out, a delayed reply can therefore satisfy a subsequent
request of the same type and cause the VF to consume the result of the old
request.

Allow replies to reuse the initiating message number. Make the in-tree PF
handlers and the VF link-state acknowledgment echo that number. Record the
expected reply type and message number on the VF, and require both values
to match before accepting a reply.

Protect expected-reply state with a lock so reply acceptance and timeout
invalidation cannot race. Keep message numbers monotonic across an admin-
channel reopen so a delayed reply from an earlier channel generation cannot
match a new request.

Reply-number echo is part of the established version-1 protocol, so this
remains compatible with deployed V2-capable PF implementations that already
echo msg_num.

Fixes: 72b65c9405 ("enic: add MBOX VF handlers for capability, register and link state")
Signed-off-by: Satish Kharat <satishkh@cisco.com>
Link: https://patch.msgid.link/20260830-b4-enic-v2-mbox-fixes-net-v1-2-23adf9bfd426@cisco.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 19:10:18 -07:00
Satish Kharat
b752e041d5 enic: preserve V2 VF carrier across netdev reopen
A V2 VF receives carrier state only from PF MBOX notifications.
enic_stop() forces carrier off, but enic_open() does not request a fresh
notification or restore the previous one.  An ordinary down/up cycle
therefore leaves the VF in NO-CARRIER and unable to pass traffic until the
PF repeats the link-state command, even when the physical link remained
up.

Cache each valid PF link-state notification.  Serialize updates with the V2
VF datapath running state.  Keep carrier off while the netdev is stopped.
Restore the cached state after an ordinary open.  Before either internal
reset reopens the datapath, invalidate the cache.  Carrier then remains off
until re-registration receives a fresh PF link-state notification.

Fixes: 72b65c9405 ("enic: add MBOX VF handlers for capability, register and link state")
Signed-off-by: Satish Kharat <satishkh@cisco.com>
Link: https://patch.msgid.link/20260830-b4-enic-v2-mbox-fixes-net-v1-1-23adf9bfd426@cisco.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 19:10:18 -07:00
Joe Damato
39b23c1c40 bnxt_en: Prevent queue stop with deferred completions
When the driver receives a burst of packets, it can mark a BD with the
NO_CMPL bit to defer completions. The expectation is that the last
packet in the ring will have this bit unset and the completion generated
by that packet will cleanup that packet and the ones preceding it. This
helps to reduce the number of completions fired.

The suppressed completions are controlled by the driver and the number
of packets with suppressed completions scales with the size of the ring.
SW USO packets, on the other hand, have an upper bound on the maximum
number of BDs which can be consumed which does not scale with the ring
size.

So, for small rings it is possible that: a burst of packets is handed to
the driver, the driver defers completions for all of the packets because
the number of free descriptors stays above the threshold in the driver.
Then, a USO packet arrives, but the number of BDs available is not
enough and the USO code exits early.

In this case, you end up in a state where the ring is full of packets
with their completions suppressed, which can cause the queue to stop and
never be restarted.

Assuming default CONFIG_MAX_SKB_FRAGS, this is only possible for small
rings (<= 457 descriptors, below the driver default value) when
a burst of packets fills the ring, followed by a large USO packet that
can't fit. For larger rings, the delta between the completion
suppression threshold and the BDs required for SW USO is large enough
that completions will fire and this case is unreachable.

This issue was pointed out by Sashiko and while it seems fairly unlikely
given that the queue size must be small to trigger this, it is indeed
possible.

Fix this by tracking the last BD which deferred completions and
centralizing the logic for deciding when to ring the doorbell. The NO_CMPL
bit is now cleared in bnxt_txr_db_kick(), so every doorbell site is
covered, including the SW USO early exit. This guarantees the ring always
ends in a BD which generates a completion to clean it and wake the queue.

Fixes: cc5d90667d ("net: bnxt: Implement software USO")
Cc: <stable@vger.kernel.org> # v7.1+: 4e15e89faac9: net: bnxt: ring the doorbell when SW USO exits early
Signed-off-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260902213956.4160615-1-joe@dama.to
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 17:49:57 -07:00
Lorenzo Bianconi
a09ceadff9 mailmap: add entries for Lorenzo Bianconi
Add the active email address for Lorenzo Bianconi and map the old,
no-longer-used addresses to it, so that git can attribute his
contributions to a single identity.
This is done to avoid bouncing emails sent to email addresses that are
no longer active.

Signed-off-by: Lorenzo Bianconi <lorenzo@kernel.org>
Link: https://patch.msgid.link/20260901-lorenzo-mailmap-v2-1-0ee832de0caf@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 17:47:52 -07:00
Ido Schimmel
b58d749633 tunnels: Drop stale dst when building an ICMP error for PMTUD
Bridged UDP tunnels such as VXLAN and GENEVE build an ICMP error packet
around an overlay packet if the packet is going to exceed the underlay
path MTU. The ICMP error packet is then injected back into the Rx path
with the source and destination addresses swapped, so that it will be
delivered to the overlay source.

If the overlay packet was routed to the UDP tunnel or locally generated,
then it is already carrying a valid dst entry and this entry is not
dropped when transforming the packet to an ICMP error packet. This
causes the IP layer to reuse the dst entry, leading to the ICMP error
packet being dropped or routed out of the UDP tunnel interface in case
of forwarding.

Prior to the blamed commit this could not happen, as
skb_tunnel_check_pmtu() did not build ICMP errors for PACKET_HOST
packets. Such packets were instead encapsulated and, unless the DF bit
was set in the outer header, fragmented by the underlay.

Fix this by making sure that the ICMP error packet does not have a valid
dst entry, thereby forcing the IP layer to perform a route lookup.

Adjust the bridged PMTU exception selftests accordingly. When the
local sender in ns_a pings the overlay destination with a deadline
(-w), ping exits on the first socket error before any reply is
received and returns a non-zero exit code. The test therefore only
passed because the ICMP error was never delivered. Use a packet count
(-c) like the ns_c line above it, so that the ICMP error counts
against the packet budget and the exit code depends on whether echo
replies were received. This passes with and without the fix.

Fixes: 8930424777 ("tunnels: Accept PACKET_HOST in skb_tunnel_check_pmtu().")
Cc: stable@vger.kernel.org
Reported-by: Laika Price <laikabcprice@gmail.com>
Closes: https://lore.kernel.org/netdev/20260614-master-v3-1-9f5060ba1ed1@gmail.com/
Reported-by: Yaroslav Dudkov <aroslavdudkov622@gmail.com>
Closes: https://lore.kernel.org/netdev/20260901081825.287173-1-aroslavdudkov622@gmail.com/
Reported-by: Charles Bordet <rough.rock3059@datachamp.fr>
Closes: https://lore.kernel.org/netdev/aHVhQLPJIhq-SYPM@eldamar.lan/
Signed-off-by: Ido Schimmel <idosch@nvidia.com>
Tested-by: Yaroslav Dudkov <aroslavdudkov622@gmail.com>
Reviewed-by: David Ahern <dsahern@kernel.org>
Reviewed-by: Stefano Brivio <sbrivio@redhat.com>
Reviewed-by: Guillaume Nault <gnault@redhat.com>
Link: https://patch.msgid.link/20260902190112.4126199-1-idosch@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 17:40:55 -07:00
XingWang Xiang
6a1094c34d genetlink: pin family module during policy dump
The generic netlink controller's policy dump keeps pointers to the target
family's operation and policy tables in its callback state.  A dump may be
split across multiple skbs and remain pending after the initial request.

Netlink pins the module which owns the dump callback, but in this case
that is the controller's owner rather than the target family's owner.  The
target family can consequently be unregistered and its module unloaded
while a policy dump is pending.  Advancing the dump then dereferences
policy memory from the unloaded module.

Take a reference to the target family's module when the dump starts.
Drop it from the error and done paths.  This matches the lifetime for which
the dump context retains the family and policy pointers.

Fixes: d07dcf9aad ("netlink: add infrastructure to expose policies to userspace")
Cc: stable@vger.kernel.org
Signed-off-by: XingWang Xiang <v3rdant.xiang@gmail.com>
Link: https://patch.msgid.link/20260902084317.4092542-1-v3rdant.xiang@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 17:21:35 -07:00
Taylor Bates
2ac174dfcd mlxsw: spectrum_ptp: Fix napi_gro_receive() call from GC workqueue context
Currently mlxsw_sp1_ptp_ht_gc_collect() is run from the PTP
garbage-collection workqueue, rather than the NAPI poll context. For any
unmatched PTP entries carrying an SKB, it calls
mlxsw_sp1_ptp_unmatched_finish() -> mlxsw_sp1_ptp_packet_finish(). For
ingress packets, this calls mlxsw_sp_rx_listener_no_mark_func(). The end
of that function is the following:

    skb->protocol = eth_type_trans(skb, skb->dev);
    napi_gro_receive(mlxsw_skb_cb(skb)->rx_md_info.napi, skb);

The napi pointer is one that was placed in the SKB control block when the
trapped packet was received in the NAPI context. Later, when the GC reaps
the unmatched entry (up to MLXSW_SP1_PTP_HT_GC_TIMEOUT later), the call to
napi_gro_receive() mutates the NAPI instance's GRO list, which is unsafe
if the poll is running concurrently on another CPU.

In mlxsw_sp1_ptp_ht_gc_collect(), local_bh_disable() is called to prevent
softirq processing, but this only applies to the local CPU. Additionally,
its comment is stale. It states that mlxsw_sp1_ptp_unmatched_finish()
invokes netif_receive_skb(). This has not been accurate since the
referenced commit; this patch makes that comment accurate again.
mlxsw_pci_napi_devs_init() calls netif_threaded_enable() on the NAPI RX
net_device without any conditions. The NAPI instance's poll, which may be
running concurrent to the GC, is running as an independently-scheduled
kthread which may be on a different CPU. The call to local_bh_disable()
does not guard against this.

If a tx-timestamp timeout produces an unmatched entry (which can be easily
reproduced by running ptp4l and waiting for a port to reach the
UNCALIBRATED/SLAVE state) while the owning NAPI thread is in the middle of
a poll on another CPU, both sides mutate the GRO list concurrently, as
shown below:

  [39.846] port 1 (swp1): MASTER to UNCALIBRATED on RS_SLAVE
  list_add corruption. next->prev should be prev (ffff8d620faf4138), but was ffff8d624150f700. (next=ffff8d620faf4138).
  kernel BUG at lib/list_debug.c:29!
  Oops: invalid opcode: 0000 [#1] SMP PTI
  CPU: 1 UID: 0 PID: 539 Comm: napi/mlxsw_rx-0 Not tainted 6.18.48 #1-NixOS PREEMPT(lazy)
  Hardware name: Mellanox Technologies Ltd. MSN2410/VMOD0001, BIOS 4.6.5 09/13/2018
  RIP: 0010:__list_add_valid_or_report+0x79/0xb0
  RSP: 0018:ffffcdf8c0f27c08 EFLAGS: 00010246
  RAX: 0000000000000075 RBX: ffff8d624150fd00 RCX: 0000000000000000
  RDX: 0000000000000000 RSI: 0000000000000001 RDI: ffff8d6315d1e540
  RBP: ffff8d620faf4070 R08: 0000000000000000 R09: 00000000ffffdfff
  R10: ffffffffa5c60fe0 R11: ffffcdf8c0f27ab8 R12: 0000000000000003
  R13: 000000000000003d R14: 00000000000001bc R15: 0000000000000001
  FS:  0000000000000000(0000) GS:ffff8d636f63f000(0000) knlGS:0000000000000000
  CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
  CR2: 0000562689a60c24 CR3: 000000015f224004 CR4: 00000000001726f0
  Call Trace:
  <TASK>
  gro_receive_skb+0xee/0x230
  mlxsw_sp1_ptp_got_packet+0x61/0x140 [mlxsw_spectrum]
  mlxsw_core_skb_receive+0xdf/0x1b0 [mlxsw_core]
  mlxsw_pci_napi_poll_cq_rx+0x780/0x9d0 [mlxsw_pci]
  __napi_poll+0x31/0x1e0
  napi_threaded_poll_loop+0x16b/0x1c0
  napi_threaded_poll+0x71/0xa0
  kthread+0xfb/0x260
  ret_from_fork+0x22d/0x260
  ret_from_fork_asm+0x1a/0x30
  </TASK>
  Kernel panic - not syncing: Fatal exception in interrupt

The machinery that leads to this kernel panic has not been changed between
6.18.48 and mainline.

This patch adds an ingress-delivery helper for the PTP packet_finish()
path that calls netif_receive_skb() instead of napi_gro_receive().
netif_receive_skb(), unlike napi_gro_receive(), can be called from outside
of the NAPI instance's poll context, which can occur at the call site for
this path. RX stats accounting and the skb->dev assignment are still
preserved; the only change is the delivery call itself.

This removes GRO batching for any PTP event traffic received by the mlxsw
trap, but given the relatively low volume of traffic characteristic of the
protocol, and impact limited to only Spectrum-1 ASICs, this is an
acceptable solution.

Fixes: 1ba06ca96c ("mlxsw: Switch to napi_gro_receive()")
Signed-off-by: Taylor Bates <tmbates12@gmail.com>
Reviewed-by: Petr Machata <petrm@nvidia.com>
Link: https://patch.msgid.link/20260902024949.2273997-1-tmbates12@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 17:01:03 -07:00
Zihan Xi
efdfb1e27a ipv4: fib: bound automatic table ID allocation
fib_empty_table() probes every table ID from 1 until it finds a
free one.  IPv4 tables are stored in a 256-bucket hash table, so a
dense set of IDs makes each probe walk a growing hash chain while
RTNL is held.

Automatic table assignment ("ip rule ... table 0") is an IPv4-only
legacy path.  Bound the automatically allocated ID to 4096 so the
RTNL hold stays bounded, without changing lookups of explicitly
specified table IDs.

This changes user-visible behavior.  A table-0 rule previously
received the lowest free ID in 1..RT_TABLE_MAX (0xFFFFFFFF).  After
this patch the search stops at 4096 and the rule add fails with
ENOBUFS if that range is fully occupied.  Explicit table IDs above
4096 remain usable.

The automatic path is unused in practice: it is IPv4-only, not
documented by ip-rule, uncovered by kernel selftests, and both
NetworkManager and systemd refuse table 0.

Fixes: b801f54917 ("[NET]: Increate RT_TABLE_MAX to 2^32")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Suggested-by: Ido Schimmel <idosch@nvidia.com>
Signed-off-by: Zihan Xi <zihanx@nebusec.ai>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Petr Vorel <pvorel@suse.cz>
Link: https://patch.msgid.link/6f2f2a7a136aee005512a2e1ac8ede62ac8c7bb6.1788258884.git.zihanx@nebusec.ai
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 16:58:32 -07:00
Jakub Kicinski
b6dd676d32 Merge branch 'net-bcmasp-fix-tx-ring-accounting-bugs'
Danesh Petigara says:

====================
net: bcmasp: fix TX ring accounting bugs

Two fixes for TX descriptor ring handling in the bcmasp driver:

  - tx_spb_ring_full() re-initialized next_index from
    intf->tx_spb_index on every loop iteration instead of advancing
    it, so it only ever checked a single descriptor slot regardless
    of cnt. This let bcmasp_xmit() proceed even when the ring didn't
    actually have enough free slots for the SKB's fragments.

  - bcmasp_xmit() only set txcb->last for the final fragment of an
    SKB, leaving stale true values in reused descriptor slots from a
    prior transmission. Combined with the ring-full miscount above,
    this could cause bcmasp_tx_reclaim() to treat a mid-SKB
    descriptor as the last one and free the sk_buff while later
    fragments were still in flight.

Patch 1 clears txcb->last unconditionally before it is set, and
patch 2 fixes the ring-full slot check to advance through each
candidate slot.
====================

Link: https://patch.msgid.link/20260831184235.4133351-1-danesh.petigara@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 16:55:02 -07:00
Justin Chen
0c5cf62e72 net: bcmasp: fix tx_spb_ring_full() checking same slot cnt times
The loop initialised next_index from intf->tx_spb_index on every
iteration, so incr_ring() always produced the same result and only
one slot was ever tested.  Move the initialisation before the loop
so each iteration advances next_index and the function correctly
checks that cnt consecutive descriptor slots are available before
allowing a new transmission.

Fixes: 490cb41200 ("net: bcmasp: Add support for ASP2.0 Ethernet controller")
Signed-off-by: Justin Chen <justin.chen@broadcom.com>
Signed-off-by: Danesh Petigara <danesh.petigara@broadcom.com>
Reviewed-by: Florian Fainelli <florian.fainelli@broadcom.com>
Link: https://patch.msgid.link/20260831184235.4133351-3-danesh.petigara@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 16:54:57 -07:00
Justin Chen
18e5e0ec0e net: bcmasp: clear txcb->last before writing each descriptor
bcmasp_xmit() only wrote txcb->last = true for the final fragment
of an SKB; non-final fragments left the field untouched.  If a
descriptor slot was reused while it still held a stale true from
a previous SKB (possible when tx_spb_ring_full() underreported
fullness), bcmasp_tx_reclaim() would see last == true mid-SKB and
call dev_consume_skb_any() prematurely, freeing the sk_buff while
its remaining fragments were still in flight.

Unconditionally clear txcb->last before the conditional set so every
descriptor slot starts from a known false state regardless of what a
prior transmission left behind.

Fixes: 490cb41200 ("net: bcmasp: Add support for ASP2.0 Ethernet controller")
Signed-off-by: Justin Chen <justin.chen@broadcom.com>
Signed-off-by: Danesh Petigara <danesh.petigara@broadcom.com>
Reviewed-by: Florian Fainelli <florian.fainelli@broadcom.com>
Link: https://patch.msgid.link/20260831184235.4133351-2-danesh.petigara@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 16:54:57 -07:00
Linus Torvalds
adf50c47a4 Including fixes from bluetooth.
Previous releases - regressions:
 
   - page_pool: keep frag_offset aligned for odd-sized requests
 
   - sched: fix u32 duplicate handle when node ID pool is exhausted
 
   - udp: create exceptions before socket matching
 
   - igmp: convert struct ip_sf_list to RCU
 
   - ip6_gre: check tunnel info before xmit in ip6gre_tunnel_xmit
 
   - rds: acquire the fastpath locks in rds_conn_shutdown()
 
   - tipc:
     - protect node reset trace dump with node lock
     - fix NULL deref in tipc_named_node_up() on empty publication list
 
   - bluetooth:
       L2CAP: fix out-of-bounds write in l2cap_ecred_connect
       hci_core: fix race condition during device registration
 
   - eth: mlx5e: prevent stale XSK buffer release on refill retries
 
   - eth: bridge: don't truncate the port group walk on teardown
 
 Previous releases - always broken:
 
   - gro: fix nesting of TCP GSO SKBs in skb_gro_receive_list()
 
   - sched: fix skb sizing and action leak on reoffload delete
 
   - tcp: fix use-after-free in do_tcp_getsockopt()
 
   - af_packet: don't cast tpacket_hdr.tp_len to int in tpacket_parse_header().
 
   - sctp: fix soft lockup from unpadded ASCONF-ACK parameter iteration
 
   - iptunnel: fix stale transport header during tunnel decapsulation
 
   - eth: vxlan: fix use-after-free in vxlan_mdb_remote_src_del()
 
   - eth: bonding: fix uninitialized transport header access in alb_determine_nd()
 
 Signed-off-by: Paolo Abeni <pabeni@redhat.com>
 -----BEGIN PGP SIGNATURE-----
 
 iQJGBAABCgAwFiEEg1AjqC77wbdLX2LbKSR5jcyPE6QFAmqZp4kSHHBhYmVuaUBy
 ZWRoYXQuY29tAAoJECkkeY3MjxOkEjkQALMGg903vZ4TfGlzKayFzyhcd5ZC8G4F
 R4M+UgTGRfNuas/1YwjpyOpvOYyFgGZ9xBmYNFdsW0YCzZwu8PxpXgu6WTZ+F4gu
 sDtWoAbN5V6CfY3fdC7IbXTp8t4CX+shQAsVvEp39Y8SJF4AZeMn8N0+Lnu4DlD3
 DAPo/lYSSfvv7RK/5Jvr9FWo7vyoEylfG+LekzGASmWGwhC3h7kWGB4RB4PhJmyq
 vRIj2ZjnzdDxu4N7ZGh+EEu5SBCcLP0e/dIMCDg++HAghDqPJ+7pzbWC1kFtQ0ss
 qOSyws/xMW3D0Rb68tkiikYWRwgvXUsfEL7Jdf2lhC1xDI8ZpxrrnPYYSZS4rsjb
 hjeBwtzRZhv5R0PnNlaZyNpFICIW3XwqP0bYqH/Z/CgwwkKYd+Rp6Tm7hkLpafNx
 Py607x2Ff/L2Aydp8csJEyqFP33QOHAfeW+X/YCo4jTc0zBTMSsOblG/EsPPBdNX
 fvhVkx4NdqAvIdLYm65cdvhe5dZtIhOAhwAMcrGiMIia4vCIsXq3fWbAe6Phtu8T
 KHpQ7Esg/if6blNPpflBuVPsoU+5N6mL7a+zsurqvilJBC5DgYaxwbMlgRSUmJzg
 2rR6IlrOsb0FFITNUv+xgPEGzMGV1rL4wnaMSmVN0tjmFQvcu24XQ17D+ug27+eG
 rD+E2lrX5rNn
 =0sv8
 -----END PGP SIGNATURE-----

Merge tag 'net-7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net

Pull networking fixes from Paolo Abeni:
 "Including fixes from bluetooth.

  Previous releases - regressions:

    - page_pool: keep frag_offset aligned for odd-sized requests

    - sched: fix u32 duplicate handle when node ID pool is exhausted

    - udp: create exceptions before socket matching

    - igmp: convert struct ip_sf_list to RCU

    - ip6_gre: check tunnel info before xmit in ip6gre_tunnel_xmit

    - rds: acquire the fastpath locks in rds_conn_shutdown()

    - tipc:
        - protect node reset trace dump with node lock
        - fix NULL deref in tipc_named_node_up() on empty publication
          list

    - bluetooth:
        - L2CAP: fix out-of-bounds write in l2cap_ecred_connect
        - hci_core: fix race condition during device registration

    - eth:
        - mlx5e: prevent stale XSK buffer release on refill retries
        - bridge: don't truncate the port group walk on teardown

  Previous releases - always broken:

    - gro: fix nesting of TCP GSO SKBs in skb_gro_receive_list()

    - sched: fix skb sizing and action leak on reoffload delete

    - tcp: fix use-after-free in do_tcp_getsockopt()

    - af_packet: don't cast tpacket_hdr.tp_len to int in
      tpacket_parse_header()

    - sctp: fix soft lockup from unpadded ASCONF-ACK parameter iteration

    - iptunnel: fix stale transport header during tunnel decapsulation

    - eth:
        - vxlan: fix use-after-free in vxlan_mdb_remote_src_del()
        - bonding: fix uninitialized transport header access in
          alb_determine_nd()"

* tag 'net-7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (83 commits)
  net: gro: Fix nesting of TCP GSO SKBs in skb_gro_receive_list()
  net: stmmac: reconfigure RX packet parser table in stmmac_hw_setup() after reset
  net: airoha: enable RX_DONE interrupt for RX queue 31
  net/rds: don't let rds_conn_shutdown() consume a concurrent drop
  net/rds: acquire the fastpath locks in rds_conn_shutdown()
  net/rds: acquire RDS_IN_XMIT in rds_tcp_reset_callbacks()
  net/rds: tcp: don't force RDS_CONN_RESETTING over a concurrent shutdown
  net/rds: clear cp_flags bits individually in rds_conn_path_reset()
  net/rds: use clear_bit_unlock() in release_refill()
  net/rds: use wq_has_sleeper() in release_in_xmit()
  net: usb: qmi_wwan: add Compal EXM-G1x support
  net: macb: exclude software FCS from TX byte statistics
  net: Remove conflicting altnames for dying netns in __dev_change_net_namespace().
  net: bridge: mcast: don't truncate the port group walk on teardown
  bonding: do not clear curr_active_slave prematurely when releasing all slaves
  net: qrtr: Send HELLO message on endpoint register
  octeontx2-af: Fix limiting SRIOV VF count logic
  bonding: alb: fix uninitialized transport header access in alb_determine_nd()
  s390/ctcm: Prevent XID null dereference
  net: psp: do not inherit the Rx association on clone
  ...
2026-09-03 10:18:12 -07:00
Linus Torvalds
8ab1afb2eb - dm-crypt: fix a race condition that could make errors not being reported
- dm-cache: fix rwsem being locked and unlocked from different processes
 
 - dm-integrity: set the 'stable writes' flag
 
 - dm-integrity: fix a buffer overflow introduced in this merge window
 
 - dm-integrity: fix an infinite loop if tag size is greater than 64
 
 - dm-cache: fix demotion statistics
 
 - dm-integrity: fix NULL pointer dereference in the data-recovery mode
 
 - dm-ebs: remove a bogus restriction on the starting sector offset
 -----BEGIN PGP SIGNATURE-----
 
 iIoEABYIADIWIQRnH8MwLyZDhyYfesYTAyx9YGnhbQUCapmHXxQcbXBhdG9ja2FA
 cmVkaGF0LmNvbQAKCRATAyx9YGnhbXwlAP9TGvtNybtjhdQJmo3t528XRdwYrYuJ
 TZTRcQ+KMXEAkgEAxG1l4vvECTwSkUhZXub2own7+IypiY8N5cFOWCkdVwc=
 =hVzb
 -----END PGP SIGNATURE-----

Merge tag 'for-7.3/dm-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/device-mapper/linux-dm

Pull device mapper fixes from Mikulas Patocka:

 - fix a dm-crypt race condition that could make errors not being reported

 - dm-cache:
    - fix rwsem being locked and unlocked from different processes
    - fix demotion statistics

 - dm-integrity:
    - set the 'stable writes' flag
    - fix a buffer overflow introduced in this merge window
    - fix an infinite loop if tag size is greater than 64

 - fix NULL pointer dereference in dm-integrity data-recovery mode

 - remove a bogus restriction on the dm-ebs starting sector offset

* tag 'for-7.3/dm-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/device-mapper/linux-dm:
  dm-ebs: fix incorrect device offset check in ebs_ctr()
  dm-integrity: fix NULL pointer dereference when the 'R' flag is used
  dm cache: fix demotion stats in passthrough mode
  dm-integrity: fix infinite loop on discard with large tag size
  dm-integrity: fix buffer overflow with keyed discard
  dm-integrity: require stable writes for internal hash modes
  dm cache: fix issue with background work locking
  dm-crypt: fix a tiny race condition in crypt_dec_pending
2026-09-03 08:30:45 -07:00
Linus Torvalds
97be98b94d - Serialize truncate, fallocate, and mmap fault paths with
invalidate_lock, avoiding mmap failures during concurrent size changes
    and exposure of uninitialized data during allocation.
 
  - Correct fallocate signal and zeroing error handling.
 
  - Fix FITRIM range alignment to prevent discard requests from extending
    into allocated clusters.
 
  - Fix free-cluster accounting when cluster-freeing rollback or bitmap
    clearing fails.
 
  - Keep volumes marked dirty when ntfs errors have been recorded.
 
  - Compute bi_sector in 512-byte units, preventing silent corruption on
    4Kn devices.
 
  - Validate sectors_per_cluster values and prevent undefined shifts when
    parsing MFT and index record sizes.
 
  - Bound $AttrDef traversal to the loaded table size.
 
  - Fix MFT record resizing, memmove overlap, and kmap_local cleanup issues.
 
  - Improve error propagation across attribute, EA, and reparse operations,
    including returning -ERANGE for undersized xattr buffers.
 
  - Avoid modifying the HasEA flag when setxattr fails and return
    DT_UNKNOWN when directory inode lookup fails.
 
  - Reduce contention in WOF decompression by performing block reads outside
    the decompression lock.
 -----BEGIN PGP SIGNATURE-----
 
 iQJKBAABCgA0FiEE6NzKS6Uv/XAAGHgyZwv7A1FEIQgFAmqZGasWHGxpbmtpbmpl
 b25Aa2VybmVsLm9yZwAKCRBnC/sDUUQhCIdnD/9OEohX3GvIqwHT90GubLIGunJr
 S2D1MSJz0AwNF0sNQhTawAc3fbjwI77B2H3mI/Xghkd4IvgtzcY/L/jYfaZ3M7sn
 Grctto0BypHI5DuBbArfjTQdW/NkPR0IpXGyBLQ8sO6aYVUPGAG0lvL9tT1Zm52N
 JQU1mtjEihE5ZpD79gx8PexuDJHIg0uuok4EANk9Vu+Ub68bDBsnl/Zyxm4spIEA
 976QAdboGDvo+71IdpPSaMuSAMytOf7LDJqxECqZXN5aUOoz9wrJnjELVg+xRE6c
 AFM9hHZ4tZ0zs5A0EpR835URaB/bxGWpbGdkCyDDBm+QMHiTGNp5nGFl/Sd/NO5B
 NcSaj0Tc2+7DbcTLU2hk1FhUsEk8eTwBZK05gxE6OajAIfAyXqxTNeB1th5smLKG
 PjKWjQh9F2okIB71D6jkdntAs/0RPyuu37bTl0EtJeuRoWYZooHkgeA+tV350BQz
 vjOwQuVUnDoNRQ0z1egrqAZalgjoNG7xLu0fI+n7eXZ5a4XZvQpYlKWKsodXYVwX
 TBwnQhst8zEx49fe5dIBGnLiZhhQMe2zxmh7lxAOu6VcVPE4Xv6jSwzcilyiPtQF
 YrdZGOEgtByOLsD+m1PRYZvukvQDDpn+NX6w4dJ+YL3MjGB7X1e2c7imqonJJYM1
 0hByGPAn3g+5XdGSig==
 =gAcy
 -----END PGP SIGNATURE-----

Merge tag 'ntfs-for-7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/ntfs

Pull ntfs fixes from Namjae Jeon:

 - Serialize truncate, fallocate, and mmap fault paths with
   invalidate_lock, avoiding mmap failures during concurrent size
   changes and exposure of uninitialized data during allocation

 - Correct fallocate signal and zeroing error handling

 - Fix FITRIM range alignment to prevent discard requests from extending
   into allocated clusters

 - Fix free-cluster accounting when cluster-freeing rollback or bitmap
   clearing fails

 - Keep volumes marked dirty when ntfs errors have been recorded

 - Compute bi_sector in 512-byte units, preventing silent corruption on
   4Kn devices

 - Validate sectors_per_cluster values and prevent undefined shifts when
   parsing MFT and index record sizes

 - Bound $AttrDef traversal to the loaded table size

 - Fix MFT record resizing, memmove overlap, and kmap_local cleanup
   issues

 - Improve error propagation across attribute, EA, and reparse
   operations, including returning -ERANGE for undersized xattr buffers

 - Avoid modifying the HasEA flag when setxattr fails and return
   DT_UNKNOWN when directory inode lookup fails

 - Reduce contention in WOF decompression by performing block reads
   outside the decompression lock

* tag 'ntfs-for-7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/ntfs: (23 commits)
  ntfs: take invalidate_lock in ntfs_filemap_page_mkwrite()
  ntfs: take invalidate_lock in ntfs_setattr_size()
  ntfs: handle signal interruption in fallocate
  ntfs: fix FITRIM range alignment
  ntfs: read WOF chunks outside the decompression lock
  ntfs: leave HasEA flag untouched on setxattr failure
  ntfs: fix race between fallocate and mmap reads
  ntfs: fix memmove overlap in ntfs_new_attr_flags
  ntfs: compute bi_sector in 512-byte units
  ntfs: reject invalid sectors_per_cluster in the boot sector
  ntfs: bound $AttrDef table walk to the loaded table size
  ntfs: fix undefined behavior in mft/index record size calculation
  ntfs: treat any nonzero dio zero-range return as an error
  ntfs: fix incorrect MFT record pointer passed to ntfs_attr_record_resize
  ntfs: do not mark the volume clean in sync_fs when errors were recorded
  ntfs: skip free cluster decrement when rollback fails
  ntfs: only count successfully cleared runs when freeing clusters
  ntfs: fix kmap_local leak in write_mft_record_nolock() error paths
  ntfs: return real error from ntfs_non_resident_attr_record_add()
  ntfs: preserve error code in ntfs_resident_attr_record_add()
  ...
2026-09-03 08:10:04 -07:00
HW He
66817a9794 net: gro: Fix nesting of TCP GSO SKBs in skb_gro_receive_list()
Fraglist GRO and hardware GRO can create an fraglist of
HW-GRO packets. This cannot be segmented back into
the original form on TCP tethering scenario.

Avoid constructing such a GSO packet, by flushing an already
built fraglist GRO packet if a hardware GRO packet arrives.

Scenario (Tethering/Forwarding):
1.Driver submits a single TCP packet, P1. P1 is kept in the
gro_list as the first packet.

2. The driver submits a TCP GSO skb, P2. P2 has already aggregated
multiple TCP packets by HW_GRO, and its non-linear data is stored in
frags[].

3. P1 and P2 match the GRO rules, and since there is no local socket,
they are aggregated by skb_gro_receive_list(). The resulting skb,
P3, has a frag_list entry that still contains frags[]:
P3: [ Linear Data ] -> frag_list -> [ Linear Data ]
                                    [ frag[1] ]
                                    [ frag[2] ]
                                    ...
4. Later, tcp4_gso_segment() or tcp6_gso_segment() calls
skb_segment_list() to segment P3. However, skb_segment_list() only
segments the entries in frag_list. It does not segment the frags[]
inside P2, so P3 is not restored to the original packets, which leads
to IP fragmentation or packet drop in the following path.

Check skb_is_gso(skb) and current GRO method, make sure fraglist GRO
applies to consecutive non-GSO skb, others adopt regular GRO path.

Fixes: 8d95dc474f ("net: add code for TCP fraglist GRO")
Signed-off-by: Zhaoping Shu <zhaoping.shu@mediatek.com>
Signed-off-by: HW He <hw.he@mediatek.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260901082312.14596-1-zhaoping.shu@mediatek.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-03 12:20:38 +02:00
Lorenzo Bianconi
6b8fed2675 net: stmmac: reconfigure RX packet parser table in stmmac_hw_setup() after reset
The core software reset issued in stmmac_init_dma_engine() during
ndo_open() callback clears the MTL RX packet parser registers, but
stmmac_rxp_config() is only invoked from the cls_u32 add/delete paths.
After an ifdown/ifup cycle the hardware therefore runs with the default
all-pass table while priv->tc_entries still reports the filters as
installed. Re-apply the RX packet parser table from priv->tc_entries in
stmmac_hw_setup(), right after the software reset, so the filters are
restored when the interface is brought up again.

Fixes: 4dbbe8dde8 ("net: stmmac: Add support for U32 TC filter using Flexible RX Parser")
Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Link: https://patch.msgid.link/20260831-stmmac_tc_cls32_reconfigure-v1-1-21cb459e64ae@oss.qualcomm.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-03 11:45:30 +02:00
Lorenzo Bianconi
7db28abbea net: airoha: enable RX_DONE interrupt for RX queue 31
RX queue 31 has always been allocated and filled by airoha_qdma_init_rx()
since RX_DONE_INT_MASK spans queues 0-31, but none of the RX_IRQ*
_BANK_PIN_MASK values covered BIT(31). As a consequence the RX_DONE
interrupt for queue 31 was never enabled, airoha_qdma_rx_process() never
ran on that queue and its buffers were never reaped.

Route RX queue 31's RX_DONE interrupt to IRQ bank 1 so that the queue
is drained and its buffers returned to the page pool.

Fixes: f252493e18 ("net: airoha: Enable multiple IRQ lines support in airoha_eth driver.")
Signed-off-by: Lorenzo Bianconi <lorenzo@kernel.org>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260830-airoha-rxdone-rxq31-v1-1-830a91503f2f@kernel.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-03 11:40:51 +02:00
Ibrahim Hashimov
2f37fba846 mac802154: fix use-after-free of sdata via queued RX frames
The RX softirq producer ieee802154_subif_frame() queues received beacon
and MAC-command frames onto local->rx_beacon_list / rx_mac_cmd_list and
schedules a process-context worker, storing a raw mac_pkt->sdata (and
skb->dev == sdata->dev) with neither a reference nor any locking:

 - the lists have no lock: the softirq producer list_add_tail()s while the
   mac_wq worker list_del()s, so sibling interfaces on the same phy corrupt
   the list;

 - the workers dereference the interface after it may have been freed.
   mac802154_rx_mac_cmd_worker() touches mac_pkt->sdata directly, and
   mac802154_rx_beacon_worker() -> mac802154_process_beacon() dereferences
   skb->dev (== sdata->dev). Removing an interface frees its sdata
   (netdev_priv) while a queued frame still points at it, so a later worker
   run is a use-after-free.

Reproduced under KASAN by flooding a victim interface with MAC command
frames and removing it (the beacon path is the same class via skb->dev):

  BUG: KASAN: slab-use-after-free in mac802154_rx_mac_cmd_worker+0x463/0x630 [mac802154]
  Read of size 4 at addr ffff888002f9ea18 by task kworker/u8:1/31
  Workqueue: phy0-mac-cmds mac802154_rx_mac_cmd_worker [mac802154]
  Call Trace:
   mac802154_rx_mac_cmd_worker+0x463/0x630 [mac802154]
   process_one_work+0x611/0xe80
   worker_thread+0x52e/0xdc0
   kthread+0x30c/0x630
   ret_from_fork+0x2fd/0x3e0

Fix both lists together:

 - add local->rx_lock and take it around every list access: the softirq
   producer (plain spin_lock, softirq context) and the workers and flush
   (spin_lock_bh, process context);

 - pin the interface for the lifetime of a queued frame with
   netdev_hold()/netdev_put(), so the worker can safely dereference sdata /
   skb->dev even while the interface is being removed;

 - dequeue under the lock at the head and loop-drain the whole list in the
   workers (they previously processed one frame per run and relied on a
   later enqueue to drain the rest);

 - drop not-yet-started frames of an interface before it is unregistered,
   from ieee802154_if_remove() (after the RCU grace period) and from the
   ieee802154_remove_interfaces() loop -- the latter is the whole-phy
   teardown path, which does not go through ieee802154_if_remove().

An in-flight worker that already dequeued a frame keeps its own netdev
reference; unregister_netdevice() then waits it out in netdev_run_todo(),
which runs at rtnl_unlock() (rtnl released) and after the interface has
been closed, so it does not pin rtnl. A worker blocked in an association
TX only delays that one interface's unregister (the usual "waiting for %s
to become free"), it does not hold rtnl. netdev_hold() is used for this
reason instead of a cancel_work_sync() under rtnl, which would block on
the worker's unbounded MLME TX wait via ieee802154_sync_queue().

The mac-command worker additionally skips processing for a stopped
interface (ieee802154_sdata_running()), avoiding a needless association
response during teardown.

Fixes: 57588c7117 ("mac802154: Handle passive scanning")
Cc: stable@vger.kernel.org
Signed-off-by: Ibrahim Hashimov <security@auditcode.ai>
Assisted-by: AuditCode-AI:2026.07
Reviewed-by: Miquel Raynal <miquel.raynal@bootlin.com>
Link: https://lore.kernel.org/20260725135154.99876-1-security@auditcode.ai
Signed-off-by: Stefan Schmidt <stefan@datenfreihafen.org>
2026-09-03 11:00:54 +02:00
Jakub Kicinski
2f38e26a57 Merge branch 'net-rds-own-the-fastpath-locks-across-connection-teardown'
Allison Henderson says:

====================
net/rds: own the fastpath locks across connection teardown

This is v5 of the follow-up set to "net/rds: Bug fix ports, part 2"
[1] (v1 at [2], v2 at [3], v3 at [4], v4 at [5]).  During review of part 2,
the later half of that series needed more work than a respin, so it was
split off into this set together with the companion fixes identified
along the way.  As discussed on the v2 thread, it is targeted at net.

RDS connection teardown quiesces the transmit and receive-refill fast
paths by waiting for the RDS_IN_XMIT/RDS_RECV_REFILL bits to be
sampled clear.  Sampling a bit clear is not owning it: the fast path
can re-take its bit right after the wait returns and then run
concurrently with the transport shutdown and the send-state reset.
Oracle UEK closed this by making teardown acquire the bits as locks
("rds: Make sure transmit path and connection tear-down does not run
concurrently"); patches 5 and 6 do the same for the two
rds_send_path_reset() call sites upstream.

Making teardown block on the bits as locks promotes several latent
ordering bugs from rare to load-bearing, so they are fixed first:

  Patches 1 and 2 fix the release side of the two bit locks.
  release_in_xmit() and release_refill() both clear their bit and then
  test for waiters, but the barrier is on the wrong side of the clear
  to order the critical section's stores before the release, and the
  waiter check does not order against the clear.  Once teardown blocks
  on these bits as locks (uninterruptible and untimed), a lost wake-up
  or a store observed out of order stops mattering only in theory.
  Use clear_bit_unlock() and wq_has_sleeper(), the pattern already
  half-present in release_in_xmit().

  Patch 3: rds_conn_path_reset() wipes the whole cp_flags word with a
  plain store.  Once teardown owns bits in that word across the reset,
  a blanket store would end lock ownership early - and it already
  races atomic RMWs on the same word today.  Clear the bits the reset
  is responsible for individually, as Oracle UEK also does.

  Patch 4: rds_tcp_reset_callbacks() stores RDS_CONN_RESETTING
  unconditionally, which can overwrite the RDS_CONN_ERROR or
  RDS_CONN_DISCONNECTING of a shutdown already in progress on the same
  path and send that shutdown through an extra drop cycle.  Once the
  accept path can park for the duration of a teardown (patch 6) that
  window widens, so make the transition conditional first, as Oracle
  UEK does.

With those in place, patch 5 converts rds_tcp_reset_callbacks() from
waiting on RDS_IN_XMIT to acquiring it, holding it across the socket
swap and rds_send_path_reset(), and patch 6 has rds_conn_shutdown()
hold both bit locks across the transport shutdown and path reset.

Patch 7 fixes a pre-existing teardown-state hole that this series
makes easier to hit but did not introduce.  Since commit
e97656d03c the final transition in rds_conn_shutdown() accepts
RDS_CONN_ERROR as well as RDS_CONN_DISCONNECTING, so that a FIN
processed during the teardown does not derail the shutdown.  But
consuming that RDS_CONN_ERROR also consumes the shutdown pass that a
concurrent rds_conn_path_drop() queued along with it.  For a FIN that
is harmless; for rds_tcp_accept_one() it is not.  A drop can race the
accept's DOWN -> CONNECTING path claim, the accept then installs the
freshly accepted socket while the drop's teardown - which sampled
tc->t_sock before that socket existed - is still running,
rds_connect_path_complete() fails and drops the path again, and if the
in-flight shutdown's final transition then swallows that
RDS_CONN_ERROR, the pass that should reap the just-installed socket
finds the path already RDS_CONN_DOWN and does nothing.  The socket is
leaked with its callbacks armed and its rds_tcp_connection still on
rds_tcp_tc_list, the peer sees an established connection that nothing
reads, and the path wedges in RDS_CONN_DOWN.  Make the final
transition DISCONNECTING -> DOWN only and leave a racing drop's
RDS_CONN_ERROR alone, so the pass it queued runs and tears down
whatever attached to the path; the branch quiesces the reconnect
timer itself, since a pending destroy can suppress that pass (see the
changes below).

This surfaced while re-reviewing v3: whether the
release-then-transition ordering in patch 6 could let a woken waiter
install a socket that the teardown then strands.  Chasing that down,
the reachable form of the leak turned out to be the accept-vs-drop
race above rather than the parked-waiter path (a path mid-teardown is
never handed to rds_tcp_reset_callbacks(): rds_tcp_accept_one_path()
only claims a path it can move DOWN -> CONNECTING), and it predates
this series.  It reproduces on an instrumented kernel - a test-only
drop injected into the accept window plus a widened teardown-to-tail
window - as an ESTABLISHED socket with an ever-growing receive queue
on a path stuck down; the same kernel runs clean with patch 7.

The set was built per-commit, run through the rds selftests (tcp and
rdma/rxe), and exercised with a connection/netns churn load and
module load/unload cycles; the patch 7 destroy-window fix was
additionally verified against an instrumented kernel that reproduces
the timer-left-armed WARN deterministically (fires on every destroyed
path unfixed, silent with the fix).
====================

Link: https://patch.msgid.link/20260828223921.202913-1-achender@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-02 19:42:26 -07:00
Allison Henderson
260c6308fe net/rds: don't let rds_conn_shutdown() consume a concurrent drop
rds_conn_shutdown() finishes by moving the path from
RDS_CONN_DISCONNECTING to RDS_CONN_DOWN, and also accepts
RDS_CONN_ERROR as the starting state of that final transition, so that
a FIN processed in softirq context during the teardown does not derail
the shutdown into a noisy error path.

But consuming that RDS_CONN_ERROR also consumes the shutdown pass that
came with it: rds_conn_path_drop() sets RDS_CONN_ERROR and then queues
cp_down_w, and a pass that starts on a path already in RDS_CONN_DOWN
is a no-op.  For the FIN case that is harmless - the socket the FIN
arrived on is the very socket the teardown just released.  It is not
harmless for a dropper that attached something to the path first.

rds_tcp_accept_one() is such a dropper.  Its path claim in
rds_tcp_accept_one_path() transitions RDS_CONN_DOWN ->
RDS_CONN_CONNECTING, and a concurrent drop - a FIN on a previous
socket in softirq context, an administrative reset - can put the path
into RDS_CONN_ERROR between that claim and the state check that
follows, which accepts RDS_CONN_ERROR.  The accept then installs the
freshly accepted socket with rds_tcp_set_callbacks() while the queued
teardown - which sampled tc->t_sock before this socket existed - is
still running.  rds_connect_path_complete() fails its transition to
RDS_CONN_UP and drops the path again, queueing the pass that should
reap the socket it just installed.  If the in-flight shutdown's final
transition consumes that drop's RDS_CONN_ERROR, the queued pass finds
the path in RDS_CONN_DOWN and does nothing.  The installed socket is
never torn down: it sits established with its callbacks armed and its
rds_tcp_connection on rds_tcp_tc_list, the peer sees a connection that
nothing ever reads, and the path is wedged in RDS_CONN_DOWN until some
later event drops it again.  Reproduced with widened race windows as
an ever-growing receive queue on a socket owned by a path stuck in
RDS_CONN_DOWN, with the peer's send path wedged behind it.

Make the final transition only DISCONNECTING -> DOWN.  If it fails
because the path is in RDS_CONN_ERROR, a drop raced the teardown:
cancel the reconnect timer and clear RDS_RECONNECT_PENDING - the one
piece of the skipped tail that must not be left behind - and return,
letting the pass the drop queued finish the job: it tears down
whatever attached to the path in the meantime, completes the
transition to RDS_CONN_DOWN, and re-arms the reconnect from its own
tail.

The timer quiesce in that branch matters because the racing drop does
not always queue that pass: rds_conn_path_drop() returns without
queueing when a destroy is pending - exactly the situation during a
netns teardown or module unload, when a FIN on the dying socket is
processed while rds_conn_path_destroy() flushes cp_down_w.  If the
flushed pass is the one that takes this return, no later pass exists,
and rds_conn_path_destroy() would find cp_conn_w still armed
(WARN_ON) and then free a path whose reconnect timer can still fire.
With the cancel in the branch, every exit of a shutdown pass leaves
the timer quiesced no matter which pass completes the transition.

The FIN case keeps making progress, one pass later and still without
noisy logging.  Any other state keeps today's rds_conn_path_error()
handling; no current cp_state writer can leave a DISCONNECTING path
in anything but RDS_CONN_ERROR (every other writer is a cmpxchg from
a non-DISCONNECTING state), so that branch is defensive.

On kernels without the preceding patches the same hazard exists with
the sample-based quiesce; the fix applies there equally.

Fixes: e97656d03c ("rds: tcp: allow progress of rds_conn_shutdown if the rds_connection is marked ERROR by an intervening FIN")
Signed-off-by: Allison Henderson <achender@kernel.org>
Link: https://patch.msgid.link/20260828223921.202913-8-achender@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-02 19:42:24 -07:00
Håkon Bugge
813f3582ac net/rds: acquire the fastpath locks in rds_conn_shutdown()
rds_conn_shutdown() quiesces the transmit and receive-refill paths by
waiting for RDS_IN_XMIT and RDS_RECV_REFILL to be sampled clear, and
then runs the transport shutdown and rds_conn_path_reset().  Sampling
the bits clear is not the same as owning them: the moment after the
wait_event() returns, rds_send_xmit() can re-acquire RDS_IN_XMIT (or
rds_ib_recv_refill() can re-acquire RDS_RECV_REFILL) and run
concurrently with the teardown.

The sender does recheck the connection state after taking the lock,
but that recheck is a classic store-buffering pattern: teardown writes
the state and reads the bit while the sender writes the bit and reads
the state.  acquire_in_xmit() is only an acquire operation, so on
weakly ordered architectures both sides can miss each other's write,
and the transmit path then runs while the transport zeroes its rings
(e.g. rds_ib_ring_init()) and rds_send_path_reset() rewrites the
transmit state under it.

Oracle UEK fixed the same class of crashes - a 14-year tail of
BUG_ON()s in rds_ib_sub_signaled(), unexpected op-codes and NULL
dereferences in rds_ib_send_cqe_handler() during failover testing -
by making the teardown path *acquire* the fastpath bit locks instead
of testing them ("rds: Make sure transmit path and connection
tear-down does not run concurrently").  Ownership of a single word is
decided by RMW atomicity, so no cross-variable ordering is needed.

Do the same here: take both locks before calling the transport
shutdown, hold them across rds_conn_path_reset(), and release them
explicitly with a wake-up afterwards.  Both are released with
clear_bit_unlock(), so that the ring re-initialization done by the
transport shutdown and the transmit state rewritten by
rds_send_path_reset() are ordered before either bit is seen clear by
the next acquire_in_xmit() or acquire_refill().

The fastpath users of these bits - rds_send_xmit() and
rds_ib_recv_refill() - are trylock style and back off while teardown
owns the locks, so no new lock dependency is introduced for them.
rds_tcp_reset_callbacks() is different: since the previous patch it
acquires RDS_IN_XMIT as well, and it blocks doing so, so its wait now
spans the teardown instead of at most one send batch.  That waiter
runs from rds_tcp_accept_one() on the single-threaded krdsd workqueue
and holds rds_tcp_accept_lock and t_conn_path_lock while it waits, so
a duelling SYN accepted while its path is being torn down parks
accept processing for the duration of the teardown - for TCP bounded
by the (up to 5 s) drain loop in rds_tcp_conn_path_shutdown().  An IB
path's drain in rds_ib_conn_path_shutdown() has no round cap, but no
blocking waiter either: rds_tcp_reset_callbacks() is the only blocking
acquirer of these bits and waits only on its own TCP path, and the
fastpaths are trylock-and-back-off on both transports, so a long IB
drain lengthens only that path's own quiesce.  The
window is narrow: the accept-side state check has to pass before the
teardown moves the path to RDS_CONN_DISCONNECTING.

Because krdsd is a single global workqueue, everything else queued
there - accept processing for other connections and network
namespaces, and the flush_workqueue(rds_wq) in rds_tcp_listen_stop()
during namespace teardown - waits behind the parked accept worker for
that time.  It cannot deadlock, although the waits do point at each
other: the teardown blocks until the bit's holder releases it, and
the holder may be that krdsd accept worker.  The holder finishes
without needing anything the teardown owns: the sync cancels
rds_tcp_reset_callbacks() issues target cp_send_w and cp_recv_w on
the path's ordered cp_wq, whose only execution slot is occupied by
the blocked cp_down_w itself, so they are pending at most and cancel
without flushing - a reliance on cp_wq being ordered that is now
noted next to those cancels (on the allocation-failure fallback where
a path shares rds_wq, the work items simply serialize).
Nor is the blocking wait itself new: rds_tcp_reset_callbacks() has
waited on RDS_IN_XMIT from the krdsd work item since
commit 335b48d980 ("RDS: TCP: Add/use rds_tcp_reset_callbacks to
reset tcp socket safely"); this patch stretches its worst case from
a sender's batch to the teardown's drain.  The alternative to parking
is the accept path racing the teardown, which is what these patches
close; making the teardown itself non-blocking is a separate item.

One observable side effect: the SENDING flag reported by rds-info has
always mirrored RDS_IN_XMIT, so it now also covers the window where
teardown owns the bit.

The comments that describe the old sample-based handshake or name
rds_send_xmit() as the only other holder of these bits - in
rds_send_xmit(), above rds_conn_path_reset(), in rds_ib_recv_refill()
and in rds_tcp_reset_callbacks() - are updated to match.

For anyone backporting this patch standalone: it depends on
"net/rds: clear cp_flags bits individually in rds_conn_path_reset()"
and "net/rds: acquire RDS_IN_XMIT in rds_tcp_reset_callbacks()"
earlier in this series.  Without the former, the blanket cp_flags
clear in rds_conn_path_reset() would drop both held bits in the middle
of the teardown; without the latter, rds_tcp_reset_callbacks() would
still sample t_sock without owning RDS_IN_XMIT.  "net/rds: use
clear_bit_unlock() in release_refill()" is needed for the refill
side's release to pair with the acquire added here, and the follow-up
"net/rds: don't let rds_conn_shutdown() consume a concurrent drop"
completes the teardown-state handling for the waiter this patch
parks; a backport should carry all four.

Fixes: 0f4b1c7e89 ("rds: fix rds_send_xmit() serialization")
Signed-off-by: Håkon Bugge <haakon.bugge@oracle.com>
[achender: reimplement for net-next shutdown path: acquire the existing
 RDS_IN_XMIT/RDS_RECV_REFILL bit locks in rds_conn_shutdown() and release
 after teardown; update comments and commit message]
Signed-off-by: Allison Henderson <achender@kernel.org>
Link: https://patch.msgid.link/20260828223921.202913-7-achender@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-02 19:42:23 -07:00
Allison Henderson
02c5f9dc2e net/rds: acquire RDS_IN_XMIT in rds_tcp_reset_callbacks()
rds_tcp_reset_callbacks() quiesces the transmit path by setting the
path state to RDS_CONN_RESETTING and then waiting for RDS_IN_XMIT to
be sampled clear before swapping the underlying socket and calling
rds_send_path_reset().

Sampling the bit clear is not the same as owning it: rds_send_xmit()
can re-acquire RDS_IN_XMIT right after the wait_event() returns.  Its
state recheck after taking the lock is a store-buffering pattern (the
resetter writes the state and reads the bit, the sender writes the
bit and reads the state) and acquire_in_xmit() is only an acquire
operation, so on weakly ordered architectures both sides can miss
each other's write and the transmit path then runs concurrently with
rds_send_path_reset() rewriting cp_xmit_* state - which is exactly
what the comment above rds_send_path_reset() tells its callers to
prevent.

Take the lock instead, hold it across the socket swap and
rds_send_path_reset(), and release it with a wake-up at the end.  The
lock-ordering constraint documented above the wait still holds: the
lock is acquired before lock_sock(), so a sender inside tcp_sendmsg()
can never be waited on while we hold the socket lock.

Two details of the old code go away with the same change:

 - t_sock is now read only after the lock is acquired.  The old code
   cached it before waiting; the teardown in rds_conn_shutdown()
   releases that socket and clears t_sock, so a pointer cached before
   the wait can be stale by the time the accept path resumes.  Reading
   it under RDS_IN_XMIT is what makes the exclusion complete once the
   teardown owns the same lock, which the next patch arranges; until
   then the teardown still only samples the bit, and the two paths
   remain as exposed to each other as they are today.

 - The old !osock early path called rds_send_path_reset() with no
   serialization at all.  It now runs under the lock like the normal
   path.  The conditional RDS_CONN_RESETTING transition of the
   previous patch happens before the socket check either way: a path
   found without a socket is either still connecting (its reconnect
   worker blocked on t_conn_path_lock) and legitimately goes
   RESETTING -> UP on the new socket, or it has been torn down
   meanwhile and is dropped.

The in-function comment describing the old wait-based quiesce is
rewritten to describe the lock-based one, and the stale block comment
above the function (which still described a return value and an
incomplete list of t_sock writers) is refreshed to name all four
writers - the connect, accept, teardown and swap paths - and what
serializes each of them.

Fixes: 335b48d980 ("RDS: TCP: Add/use rds_tcp_reset_callbacks to reset tcp socket safely")
Signed-off-by: Allison Henderson <achender@kernel.org>
Link: https://patch.msgid.link/20260828223921.202913-6-achender@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-02 19:42:23 -07:00
Gerd Rausch
e8e60d74fe net/rds: tcp: don't force RDS_CONN_RESETTING over a concurrent shutdown
rds_tcp_reset_callbacks() resolves a duelling SYN by storing
RDS_CONN_RESETTING into cp_state unconditionally.  Nothing serializes
that store against the shutdown path: rds_tcp_accept_one() checks
for RDS_CONN_CONNECTING or RDS_CONN_ERROR under t_conn_path_lock, but
neither rds_conn_path_drop(), which forces RDS_CONN_ERROR, nor
rds_conn_shutdown(), which moves the path to RDS_CONN_DISCONNECTING
under cp_cm_lock, takes that lock.  The store can therefore land on
top of a shutdown that is already in progress, or that gets queued
right after the accept-side check.

When it does, the shutdown worker's final DISCONNECTING -> DOWN
transition fails and the path goes through rds_conn_path_error() and
a second drop/shutdown cycle instead of a clean reconnect, tearing
down the socket the accept path has just installed.  Before commit
ad22d24be6 ("net/rds: No shortcut out of RDS_CONN_ERROR") a path
found in RDS_CONN_RESETTING even made rds_conn_shutdown() bail out
altogether.

Make the transition conditional: move CONNECTING -> RESETTING (or
stay in RESETTING from an earlier duel), and drop the path in any
other state.  The drop has side effects of its own: it replaces the
shutdown's RDS_CONN_DISCONNECTING (or RDS_CONN_ERROR) with
RDS_CONN_ERROR and queues one more cp_down_w run.  The difference is
that rds_conn_shutdown() accepts RDS_CONN_ERROR in its final
transition to RDS_CONN_DOWN, so the shutdown in flight completes
normally instead of through rds_conn_path_error(); the extra
down-work pass then finds the path already down and falls through to
the reconnect check, or catches a reconnect that has already started
and restarts it.  The accept path still installs the new socket,
rds_connect_path_complete() then fails its RESETTING -> UP transition
and drops it: the raced socket ends up torn down as it does today.
The comment at that call site, which promised that
rds_connect_path_complete() marks the path RDS_CONN_UP, is updated to
name this outcome as well.

The state can change again between the failed transitions and the
drop.  That is inherent to rds_conn_path_drop(), which the socket
state-change callbacks also call unconditionally, and costs at most
one extra drop/reconnect cycle.

Based on Oracle UEK commit "net/rds: Don't force state
RDS_CONN_RESETTING" by Gerd Rausch.

Fixes: 9c79440e2c ("RDS: TCP: fix race windows in send-path quiescence by rds_tcp_accept_one()")
Signed-off-by: Gerd Rausch <gerd.rausch@oracle.com>
[achender: port to net-next: use the two-argument
 rds_conn_path_transition()/rds_conn_path_drop() and rewrite the
 changelog for the upstream shutdown path]
Signed-off-by: Allison Henderson <achender@kernel.org>
Link: https://patch.msgid.link/20260828223921.202913-5-achender@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-02 19:42:23 -07:00
Allison Henderson
103c4b13c4 net/rds: clear cp_flags bits individually in rds_conn_path_reset()
rds_conn_path_reset() wipes the whole flag word with a plain
cp->cp_flags = 0 store.  Every other accessor of that word uses
atomic bitops, and some of them can run concurrently with the reset:
RDS_LL_SEND_FULL is set from rds_send_xmit() and cleared from the
transport completion paths, neither of which holds anything that
excludes the shutdown worker.  A plain store racing an atomic
read-modify-write on the same word is a data race, and whichever
side loses has its update silently discarded.

Clear the two bits the reset is actually responsible for instead.
RDS_IN_XMIT and RDS_RECV_REFILL need no store at all here: they
belong to the caller, rds_conn_shutdown(), which waits for both to be
clear before calling the transport shutdown and this reset.

This also gives every bit in cp_flags a single well-defined writer
discipline, which the following patches rely on when they turn
RDS_IN_XMIT and RDS_RECV_REFILL into bit locks held across the
teardown: a blanket store mid-teardown would destroy lock ownership
that an atomic clear preserves.

Oracle UEK carries the same conversion ("net/rds: Preserve essential
connection state flags"), motivated by its asynchronous shutdown
state machine, whose progress and destroy flags must survive the
reset.  UEK's variant also clears RDS_IN_XMIT and RDS_RECV_REFILL
because there the reset runs as the final step of a teardown that
owns both bits, making those clears its unlock.  Upstream that
release belongs in rds_conn_shutdown(): once a later patch in this
series turns the two bits into locks held across the teardown, ending
ownership needs release semantics and a wake-up that a plain clear
inside the reset would not provide.

Based on Oracle UEK commit "net/rds: Preserve essential connection
state flags" by Gerd Rausch.

Fixes: 00e0f34c61 ("RDS: Connection handling")
Signed-off-by: Allison Henderson <achender@kernel.org>
Link: https://patch.msgid.link/20260828223921.202913-4-achender@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-02 19:42:23 -07:00
Allison Henderson
17c4476dbb net/rds: use clear_bit_unlock() in release_refill()
release_refill() drops the RDS_RECV_REFILL bit with a plain
clear_bit().  clear_bit() has no ordering semantics, and the
smp_mb__after_atomic() that follows it sits on the wrong side for a
lock release: it orders the clear against the waitqueue_active() load
below it, but does nothing to order the refill critical section's ring
and descriptor stores before the clear itself.

That matters once connection teardown owns RDS_RECV_REFILL as a lock
across the transport shutdown and path reset, rather than sampling it
clear, which "net/rds: acquire the fastpath locks in
rds_conn_shutdown()" later in this series arranges: on a weakly
ordered architecture the teardown can win the bit and start the
shutdown and reset while some of the refill's stores are not yet
visible to it.  The same gap existed under the sample-based scheme - a
waiter that saw the bit clear had no guarantee it also observed the
refill's stores - but taking the bit as a lock makes the missing
release pairing load-bearing.

Switch to clear_bit_unlock(), which orders the critical section before
the release, and replace the open-coded barrier-plus-waitqueue_active()
with wq_has_sleeper(), whose internal full barrier keeps the
store-buffering guarantee between clearing the bit and checking for
sleepers.  This mirrors what "net/rds: use wq_has_sleeper() in
release_in_xmit()" does for RDS_IN_XMIT.

The fast-path acquire side, acquire_refill(), uses test_and_set_bit(),
a full-barrier RMW that pairs with this release.  The teardown at this
point in the series still samples the bit, so on its own this change
is release-side hardening; the shutdown-conversion patch named above
makes the teardown acquire the bit with the same RMW, completing the
pairing at the end of the series.

Fixes: 73ce4317bf ("RDS: make sure we post recv buffers")
Signed-off-by: Allison Henderson <achender@kernel.org>
Link: https://patch.msgid.link/20260828223921.202913-3-achender@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-02 19:42:23 -07:00
Allison Henderson
6d0c8b7073 net/rds: use wq_has_sleeper() in release_in_xmit()
release_in_xmit() clears RDS_IN_XMIT with clear_bit_unlock() and then
checks waitqueue_active() to decide whether anyone needs waking.
clear_bit_unlock() is only a release operation: it orders the
critical section before the bit clear, but does not order the
subsequent plain load of the wait queue head after it.  The waiter
side does the mirror image - it adds itself to the wait queue and
then tests the bit.  That is the classic store-buffering pattern: the
releasing CPU can read the wait queue as empty while the waiting CPU
still reads the bit as set, so the sleeper is never woken.

The waiters are rds_conn_shutdown() and rds_tcp_reset_callbacks(),
both in uninterruptible wait_event() with no timeout.  A lost wake-up
strands the shutdown worker on its single-threaded workqueue until
some other sender releases the bit again - and on a connection that
is being torn down precisely because it failed, there may never be
another sender.

The barrier used to be there: release_in_xmit() did clear_bit()
followed by smp_mb__after_atomic() until commit 1422f28826 ("rds:
introduce acquire/release ordering in acquire/release_in_xmit()")
folded both into clear_bit_unlock(), which strengthened the lock
hand-off but silently dropped the full barrier the wake-up check
depends on.  The refill counterpart, release_refill() in
net/rds/ib_recv.c, still carries its smp_mb__after_atomic() for
exactly this reason.

Use wq_has_sleeper(), which is waitqueue_active() preceded by the
required full barrier.

Fixes: 1422f28826 ("rds: introduce acquire/release ordering in acquire/release_in_xmit()")
Signed-off-by: Allison Henderson <achender@kernel.org>
Link: https://patch.msgid.link/20260828223921.202913-2-achender@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-02 19:42:23 -07:00
Ian Lin
08710f033e net: usb: qmi_wwan: add Compal EXM-G1x support
The Compal EXM-G1x is a Qualcomm SDX12-based LTE modem. Add support for
its QMI WWAN interface 8 using the DTR quirk.

Tested on a Compal EXM-G1x modem.

Signed-off-by: Ian Lin <jisayme@gmail.com>
Link: https://patch.msgid.link/20260831084124.65074-1-jisayme@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-02 19:24:38 -07:00
Nicolai Buchwitz
d85f521a9a net: macb: exclude software FCS from TX byte statistics
Frames for which macb_pad_and_fcs() supplies the FCS have four FCS
bytes appended, and TX completion then accounts the grown skb->len.
tx_bytes is defined to exclude the FCS, so these frames are reported
four bytes too large.

Track only the number of FCS bytes appended in software, 0 or
ETH_FCS_LEN, and subtract that from skb->len at completion. skb->len
already reflects the padded length by then, so there is nothing else
to store. macb_pad_and_fcs() already returns 0 on every non-error
path. Return the FCS length from there instead, rather than
recomputing the same check in the caller. BQL stays on the padded
skb->len that netdev_tx_sent_queue() saw.

Fixes: 653e92a917 ("net: macb: add support for padding and fcs computation")
Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
Link: https://patch.msgid.link/20260831113128.1678674-1-nb@tipi-net.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-02 19:22:53 -07:00
Kuniyuki Iwashima
debac3a20d net: Remove conflicting altnames for dying netns in __dev_change_net_namespace().
syzbot reported the warning in cfg80211_pernet_exit(). [0]

The repro does the following:

  1. create two device in root netns and non-root netns
  2. assign the same altname for the two devices
  3. remove the non-root netns

Since commit 7663d52209 ("net: check for altname conflicts
when changing netdev's netns"), cfg80211_switch_netns() and
cfg802154_switch_netns() fail if init_net has a device with the
conflicting altname.

default_device_exit_net() had the same issue and commit d09486a04f
("net: fix removing a namespace with conflicting altnames") fixed it.

cfg80211_pernet_exit() and cfg802154_pernet_exit() need the same fix.

Let's generalise the fix by removing conflicting altnames for dying
netns in __dev_change_net_namespace().

[0]:
cfg80211_switch_netns(rdev, &init_net)
WARNING: net/wireless/core.c:1871 at cfg80211_pernet_exit+0xd5/0x120 net/wireless/core.c:1871, CPU#1: kworker/u8:9/1160
Modules linked in:
CPU: 1 UID: 0 PID: 1160 Comm: kworker/u8:9 Not tainted syzkaller #0 PREEMPT(full)
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 07/24/2026
Workqueue: netns cleanup_net
RIP: 0010:cfg80211_pernet_exit+0xd5/0x120 net/wireless/core.c:1871
Code: e8 03 42 80 3c 20 00 74 08 4c 89 f7 e8 b4 ef 0e f7 4d 8b 36 49 81 fe 20 10 4a 90 74 12 e8 03 3d 9f f6 eb 85 e8 fc 3c 9f f6 90 <0f> 0b 90 eb cc e8 f1 3c 9f f6 eb 05 e8 ea 3c 9f f6 5b 41 5c 41 5e
RSP: 0018:ffffc900057a78f0 EFLAGS: 00010293
RAX: ffffffff8b287154 RBX: ffff88807ba72780 RCX: ffff8880213e8000
RDX: 0000000000000000 RSI: 00000000ffffffef RDI: 0000000000000000
RBP: 00000000ffffffef R08: ffffffff9024cc67 R09: 0000000000000000
R10: fffff52000af4eb0 R11: fffffbfff204998d R12: dffffc0000000000
R13: ffffffff904a1080 R14: ffff888144ed0008 R15: ffff888144ed0e20
FS:  0000000000000000(0000) GS:ffff888124de6000(0000) knlGS:0000000000000000
CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 00005642de0a8a70 CR3: 000000007a40c000 CR4: 00000000003526f0
Call Trace:
 <TASK>
 ops_exit_list net/core/net_namespace.c:200 [inline]
 ops_undo_list+0x43d/0x8d0 net/core/net_namespace.c:253
 cleanup_net+0x572/0x810 net/core/net_namespace.c:706
 process_one_work kernel/workqueue.c:3387 [inline]
 process_scheduled_works+0xc3d/0x1630 kernel/workqueue.c:3470
 worker_thread+0xa47/0xfb0 kernel/workqueue.c:3551
 kthread+0x38b/0x480 kernel/kthread.c:436
 ret_from_fork+0x514/0xb70 arch/x86/kernel/process.c:158
 ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
 </TASK>

Fixes: 36fbf1e52b ("net: rtnetlink: add linkprop commands to add and delete alternative ifnames")
Reported-by: syzbot+74f338e09f1ef3ee6457@syzkaller.appspotmail.com
Closes: https://lore.kernel.org/all/6a96219e.04428c52.29b18.0001.GAE@google.com/T/
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260901005550.2042357-1-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-02 18:30:00 -07:00
Jun Yang
5a3f7a683a net: bridge: mcast: don't truncate the port group walk on teardown
__br_multicast_disable_port_ctx() and br_multicast_del_port() walk
port->mglist with hlist_for_each_entry_safe(). However,
br_multicast_find_del_pg() can also delete other entries from the same
list through br_multicast_fwd_src_remove() or __fwd_del_star_excl().

If such an entry is the iterator's saved next node, hlist_del_init()
clears its ->next and terminates the walk early. The reproducer triggers
this in both teardown walks, leaving port groups in the bridge mdb with
a dangling ->key.port after del_nbp() frees the port:

  BUG: KASAN: slab-use-after-free in __mdb_fill_info+0x1191/0x1320
   __mdb_fill_info+0x1191/0x1320
   br_mdb_dump+0x594/0xe40
   rtnl_mdb_dump+0x1cf/0x5d0

Use hlist_del_init_rcu() to unlink the group while preserving ->next.
br_multicast_del_pg() and the teardown walks run under
br->multicast_lock. The GC worker must acquire the same lock before
detaching the group for destruction, so the node remains alive while
the walk uses the preserved pointer.

Preserving ->next means a walk can now reach a group that an earlier
iteration already deleted as a side effect. That group is off mp->ports,
so br_multicast_find_del_pg() would fall through its port scan and hit
the trailing WARN_ON(1). Skip such groups at the top of that helper: a
port group is put on port->mglist when it is created and only unlinked
when it is deleted, so hlist_unhashed() identifies exactly this case.

Fixes: b08123684b ("net: bridge: mcast: install S,G entries automatically based on reports")
Cc: stable@vger.kernel.org
Suggested-by: Nikolay Aleksandrov <razor@blackwall.org>
Reported-by: TencentOS Corvus AI <corvus@tencent.com>
Signed-off-by: Jun Yang <junvyyang@tencent.com>
Reviewed-by: Nikolay Aleksandrov <razor@blackwall.org>
Link: https://patch.msgid.link/20260831111330.199543-1-junvyyang@tencent.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-02 18:25:35 -07:00
Eric Dumazet
af602c7aa5 bonding: do not clear curr_active_slave prematurely when releasing all slaves
When releasing all slaves during bond destruction (all == true),
__bond_release_one() unconditionally clears bond->curr_active_slave to
NULL in every iteration.

If a backup slave is released before the active slave,
bond_alb_deinit_slave() triggers rlb_teach_disabled_mac_on_primary(),
which increments the active slave dev promiscuity counter and sets
bond_info->primary_is_promisc = 1.

Because bond->curr_active_slave was prematurely cleared to NULL when
releasing the backup slave, the subsequent iteration releasing the active
slave evaluates oldcurrent as NULL, so bond_change_active_slave(bond, NULL)
is skipped. Consequently, bond_alb_handle_active_change() is never called
to decrement the promiscuity counter, permanently leaking promiscuous
mode on the physical device after bond teardown.

When oldcurrent == slave, bond_change_active_slave(bond, NULL) already sets
bond->curr_active_slave to NULL. We only need to avoid selecting a new
active slave when all == true. Replace the if (all) branch with
if (!all && oldcurrent == slave).

Fixes: 0896341a44 ("bonding: fix bond_release_all inconsistencies")
Signed-off-by: Eric Dumazet <edumazet@google.com>
Acked-by: Jay Vosburgh <jv@jvosburgh.net>
Reviewed-by: Nikolay Aleksandrov <razor@blackwall.org>
Link: https://patch.msgid.link/20260831203042.164466-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-02 18:19:19 -07:00
Linus Torvalds
940de590b8 hardening fix for v7.3-rc2
- Default randstruct off with rust for better allmodconfig coverage
   (Mark Brown)
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQRSPkdeREjth1dHnSE2KwveOeQkuwUCaphPQAAKCRA2KwveOeQk
 u+JmAP9tcRZkdQkz5oBGNN58SB1eeJ/AQOuOXvqr+fe9Hw0wnQD+OylI5IR5cN9K
 7dCHFDbAd3prFT648fnmMlPhBrTpwAw=
 =x6kf
 -----END PGP SIGNATURE-----

Merge tag 'hardening-v7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/kees/linux

Pull hardening fix from Kees Cook:

 - Default randstruct off with rust for better allmodconfig coverage
   (Mark Brown)

* tag 'hardening-v7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/kees/linux:
  hardening: Default randstruct off with rust for better allmodconfig support
2026-09-02 16:02:02 -07:00
Genjian Zhang
7ac81e2d22 dm-ebs: fix incorrect device offset check in ebs_ctr()
<offset> is a backing-device sector offset; ti->len is the virtual
target length. Comparing them rejects valid tables, e.g.:

  dmsetup create ebs0 --table "0 1048576 ebs /dev/sda 2097152 1 8"
  -> ebs: Invalid device offset sector (-EINVAL)

Drop the check. Bounds against the backing device are already
enforced later by device_area_is_invalid() via ebs_iterate_devices().

Cc: stable@vger.kernel.org
Fixes: d3c7b35c20 ("dm: add emulated block size target")
Signed-off-by: Genjian Zhang <zhanggenjian@kylinos.cn>
Signed-off-by: Mikulas Patocka <mpatocka@redhat.com>
2026-09-02 17:01:13 +02:00
Mikulas Patocka
7d4d4f3b66 dm-integrity: fix NULL pointer dereference when the 'R' flag is used
If the dm-integrity device has the SB_FLAG_DIRTY_BITMAP flag set and the
user activates the device in the 'R' mode, a crash in dm_integrity_resume
happens because the function attempts to read the journal containing the
bitmap.

This patch makes dm-integrity skip any writes to the device in
dm_integrity_resume if the device is activated in the 'R' mode.

Signed-off-by: Mikulas Patocka <mpatocka@redhat.com>
Fixes: 468dfca38b ("dm integrity: add a bitmap mode")
Cc: stable@vger.kernel.org
2026-09-02 16:36:26 +02:00
Chris Lew
544d85de4d net: qrtr: Send HELLO message on endpoint register
HELLO is currently handled entirely by the name server (NS): it is
sent once as a broadcast when the NS initializes, and again as a
reply whenever the NS receives an inbound HELLO from a remote.

Some remote QRTR endpoints (e.g. an external WLAN chipset attached
over MHI) operate in a slave role: they only ever send a HELLO in
response to one they receive, and never initiate. Since the host cannot
tell in advance which remotes behave this way, if the host also only
replies, both sides wait on the other to speak first and no HELLO is
ever exchanged, stalling further communication.

To fix this:
- Transfer HELLO handshake ownership to the core layer. A HELLO is
  now sent once, per endpoint, at registration time.
- Schedule a delayed work item on endpoint registration to send a
  HELLO once the name server is bound. The work reschedules itself
  with a 100ms backoff if the name server socket is not yet bound or
  if allocating the control packet fails, so a transient startup
  condition does not abandon the handshake permanently.
- Enforce HELLO-first ordering by dropping non-HELLO packets and
  returning -EAGAIN until the HELLO is confirmed sent, using bool
  hello_sent guarded by ep_lock to make the gate check atomic with
  xmit().
- Skip nodes with nid == QRTR_EP_NID_AUTO in bcast_enqueue(), to avoid
  broadcasting control packets with QRTR_EP_NID_AUTO as the destination
  node ID.
- Remove say_hello() from the name server's ctrl_cmd_hello() handler
  and from qrtr_ns_init(); the core layer is now the sole sender of
  the outbound HELLO. This removes the NS's reply-on-receive
  behaviour without a replacement.

Signed-off-by: Chris Lew <christopher.lew@oss.qualcomm.com>
Co-developed-by: Deepak Kumar Singh <deepak.singh@oss.qualcomm.com>
Signed-off-by: Deepak Kumar Singh <deepak.singh@oss.qualcomm.com>
Co-developed-by: Pranav Mahesh Phansalkar <pranav.phansalkar@oss.qualcomm.com>
Signed-off-by: Pranav Mahesh Phansalkar <pranav.phansalkar@oss.qualcomm.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
2026-09-02 12:14:32 +01:00
Sunil Goutham
f695390ea6 octeontx2-af: Fix limiting SRIOV VF count logic
When RVU PF0/AF's VFs are SDP instead of LBK, limiting the VF count
based on the LBK channel count is incorrect.

Apply LBK channel-based VF limits only when the VF device ID matches
the LBK RVU AFVF device.

Fixes: 9bd6caf335 ("octeontx2-af: Enable sriov on AF to create VFs")
Signed-off-by: Sunil Goutham <sgoutham@marvell.com>
Signed-off-by: Nitin Shetty J <nshettyj@marvell.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
2026-09-02 09:49:27 +01:00
David Carlier
979d5b8de8 ieee802154: hwsim: serialize pib updates to fix double-free
hwsim_update_pib() does an unserialized read-swap-free of phy->pib:

	pib_old = rtnl_dereference(phy->pib);
	...
	rcu_assign_pointer(phy->pib, pib);
	kfree_rcu(pib_old, rcu);

It assumes the RTNL is held, but ->set_channel is not always called
under it: the mac802154 scan worker changes channels via
drv_set_channel() without the RTNL. Such an update can race an
RTNL-held one on the same phy; both read the same pib_old and both
kfree_rcu() it, double-freeing the object. With SLUB percpu sheaves
batching kfree_rcu(), this surfaces as a KASAN invalid-free in
rcu_free_sheaf().

struct hwsim_phy has no lock for pib. Add one and make the swap atomic
with rcu_replace_pointer() under it, dropping the misleading
rtnl_dereference().

Reported-by: syzbot+60332fd095f8bb2946ad@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=60332fd095f8bb2946ad
Fixes: f25da51fdc ("ieee802154: hwsim: add replacement for fakelb")
Signed-off-by: David Carlier <devnexen@gmail.com>
Cc: stable@vger.kernel.org
Link: https://lore.kernel.org/20260709221858.158063-1-devnexen@gmail.com
Signed-off-by: Stefan Schmidt <stefan@datenfreihafen.org>
2026-09-02 10:26:47 +02:00
Zhiling Zou
bf79662bc8 ieee802154: 6lowpan: fix NULL dereference in lowpan_newlink
TUNSETLINK allows a TUN device to change its link-layer type to
ARPHRD_IEEE802154 without initializing ieee802154_ptr. lowpan_newlink()
checks only the device type before dereferencing the pointer, so an
RTM_NEWLINK request can trigger a NULL pointer dereference.

Reject devices without ieee802154_ptr along with devices of the wrong type.

Fixes: 51e0e5d812 ("ieee802154: 6lowpan: remove multiple lowpan per wpan support")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Signed-off-by: Zhiling Zou <zhilinz@nebusec.ai>
Link: https://lore.kernel.org/0b715da69bd15a86ddc47dad5cf12da648211050.1787997209.git.zhilinz@nebusec.ai
Signed-off-by: Stefan Schmidt <stefan@datenfreihafen.org>
2026-09-02 09:54:27 +02:00
Fan Wu
ff5891b266 ieee802154: cc2520: fix FIFOP work use-after-free
The FIFOP interrupt handler queues cc2520_fifop_irqwork.  On removal,
cc2520_remove() only flushes the work.  The devm-managed FIFOP IRQ
remains active until after ->remove() returns and can queue the work
again after that flush, allowing it to run after the private data is
released.

Disable the work with disable_work_sync() instead of flushing it, so
the handler can no longer queue it once removal begins.  Destroy the
buffer mutex last, since the worker and the stop callback invoked
through ieee802154_unregister_hw() both take it.

Found by an in-house static analysis tool.

Fixes: 0da6bc8cc3 ("ieee802154: cc2520: adds driver for TI CC2520 radio")
Cc: stable@vger.kernel.org # v6.10+
Suggested-by: Miquel Raynal <miquel.raynal@bootlin.com>
Reviewed-by: Miquel Raynal <miquel.raynal@bootlin.com>
Assisted-by: Codex:gpt-5.6
Signed-off-by: Fan Wu <fanwu01@zju.edu.cn>
Link: https://lore.kernel.org/20260812061714.175966-1-fanwu01@zju.edu.cn
Signed-off-by: Stefan Schmidt <stefan@datenfreihafen.org>
2026-09-02 09:20:31 +02:00
Chenguang Zhao
9012da455a net: 6lowpan: fix mismatched comments
Rename @_nexthdrlen to @_hdrlen and drop stale @nhc from
lowpan_nhc_do_uncompression docs

Signed-off-by: Chenguang Zhao <zhaochenguang@kylinos.cn>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://lore.kernel.org/20260806091045.1701326-1-chenguang.zhao@linux.dev
Signed-off-by: Stefan Schmidt <stefan@datenfreihafen.org>
2026-09-02 09:03:06 +02:00
Eric Dumazet
70f3995830 bonding: alb: fix uninitialized transport header access in alb_determine_nd()
alb_determine_nd() uses icmp6_hdr(skb) to inspect ICMPv6 headers.
However, in xmit paths (e.g. packets sent via AF_PACKET / raw sockets
or forwarded packets), skb->transport_header is not guaranteed to be
initialized. While pskb_network_may_pull() ensures the packet data is
linear starting from the network header, it does not set or adjust the
transport header offset.

Dereferencing icmp6_hdr(skb) can therefore access out-of-bounds memory.

Fetch the icmp6hdr directly after ipv6hdr following pskb_network_may_pull(),
and reload ipv6hdr in case pskb_may_pull() reallocated skb->head.
Also remove the unused bond argument from alb_determine_nd().

Fixes: 0da8aa00bf ("net: bonding: Add support for IPV6 ns/na to balance-alb/balance-tlb mode")
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260831194626.119371-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-01 16:59:21 -07:00
Aswin Karuvally
b264d84227 s390/ctcm: Prevent XID null dereference
The mpc_validate_xid() function sets grp->saved_xid2->xid2_flag2 to 0x40
to signal XID validation error. If peer XID is NULL or r/w channel
pairing mismatch happens, grp->saved_xid2 is never initialized. An
attempt to set the flag in such case leads to NULL dereference.

Fix this by using the always available priv->xid->xid2_flag2 instead of
grp->saved_xid2->xid2_flag2 for validation errors.

Fixes: 293d984f0e ("ctcm: infrastructure for replaced ctc driver")
Cc: stable@vger.kernel.org
Signed-off-by: Aswin Karuvally <aswin@linux.ibm.com>
Link: https://patch.msgid.link/20260827063408.2168914-1-aswin@linux.ibm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-01 16:55:04 -07:00
Mark Brown
2625480a1b hardening: Default randstruct off with rust for better allmodconfig support
Currently randstruct does not support rust so we have Kconfig dependencies
which prevent rust being enabled when randstruct is. Unfortunately this
prevents rust being enabled in allmodconfig, our standard coverage build.
randstruct gets turned on by default, then the dependency on !RANDSTRUCT
causes rust to get disabled.

Work around this by disabling randstruct by default if we have a usable
rust toolchain and rust support for the architecture, circular
dependencies prevent us directly depending on !RUST. This means we might
end up with a configuration that disables both rust and randstruct but
hopefully it's more likely go give the expected result.

Signed-off-by: Mark Brown <broonie@kernel.org>
Acked-by: Miguel Ojeda <ojeda@kernel.org>
Link: https://patch.msgid.link/20260901-rust-reverse-randstruct-dep-v4-1-3bfa19efe1fa@kernel.org
Signed-off-by: Kees Cook <kees@kernel.org>
2026-09-01 16:01:35 -07:00
Linus Torvalds
89a312991d SMB client fixes for v7.3-rc2
A batch of bug fixes for the SMB client:
 
  - Fixes for fallocate range operations (insert, collapse, zero, punch
    hole): the insert range implementation copied overlapping chunks in
    the wrong direction, corrupting file data on every server except
    Windows.  Several related issues in the same area are also
    addressed — stale page cache and FS-Cache readback, an integer
    truncation on large files, missing RLIMIT_FSIZE validation and
    missing sparse file marking.
 
  - Data corruption fixes in the O_TRUNC open path: one where i_size
    was zeroed before the server confirmed the truncate and another
    where the lack of locking allowed concurrent buffered writes to be
    silently discarded.
 
  - Heap overflow fixes in legacy SMB1 paths: one in extended attribute
    writes and one in POSIX ACL handling, both exploitable via
    unprivileged setxattr(2).
 
  - Fix for multiuser mount with krb5 failing because the username
    option was not propagated to new per-user connections.
 
  - Fix for split debug message in __release_mid() after a printk
    conversion.
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQTcqRusfSdYROJQwGkpVtNKoQNdYwUCapcv6wAKCRApVtNKoQNd
 Y0dgAQDvlnpdCsg1SZZN7T/wSy08fP7GEl2lUoCb8m6LQHlDlwEA/RG+FeY8sRkb
 iJexqIGT85a48SHmpSavzBO3qkuqIQo=
 =+Qbr
 -----END PGP SIGNATURE-----

Merge tag 'cifs-fixes-7.3-rc2' of https://git.manguebit.org/linux

Pull smb client fixes from Paulo Alcantara:

 - Fixes for fallocate range operations (insert, collapse, zero, punch
   hole)

   The insert range implementation copied overlapping chunks in the
   wrong direction, corrupting file data on every server except Windows.

   Several related issues in the same area are also addressed — stale
   page cache and FS-Cache readback, an integer truncation on large
   files, missing RLIMIT_FSIZE validation and missing sparse file
   marking.

 - Data corruption fixes in the O_TRUNC open path: one where i_size was
   zeroed before the server confirmed the truncate and another where the
   lack of locking allowed concurrent buffered writes to be silently
   discarded

 - Heap overflow fixes in legacy SMB1 paths: one in extended attribute
   writes and one in POSIX ACL handling, both exploitable via
   unprivileged setxattr(2)

 - Fix for multiuser mount with krb5 failing because the username option
   was not propagated to new per-user connections

 - Fix for split debug message in __release_mid() after a printk
   conversion

* tag 'cifs-fixes-7.3-rc2' of https://git.manguebit.org/linux:
  smb: client: reject SetEA requests that do not fit the request buffer
  smb: client: fix data corruption with concurrent writes and O_TRUNC
  cifs: don't update i_size in cifs_do_truncate without a cached handle
  smb: client: fix heap overflow in cifs_do_set_acl()
  smb: client: fix multiuser mount with krb5
  smb: client: transport: Fix debug printing in __release_mid()
  smb/client: invalidate fscache for fallocate range operations
  smb/client: fix stale page cache in insert/collapse range
  smb/client: fix integer truncation in collapse range
  smb/client: fix data corruption in emulated insert range
  smb/client: mark file sparse before emulating insert range
  smb/client: validate new EOF for zero range
  smb/client: validate new EOF for insert range
  cifs: add revalidation on FSCTL failure in smb2_duplicate_extents()
2026-09-01 13:37:14 -07:00
Linus Torvalds
9a58da8005 - Prevent unintended data exposure by clearing pipe compound padding and
the response buffer.
 
  - Initialize missing fields in FS_OBJECT_ID_INFORMATION,
    FS_CONTROL_INFORMATION, and FS_POSIX_INFORMATION.
 
  - Propagate DACL parsing and allocation failures so malformed security
    descriptors are rejected.
 
  - Rate-limit errors for unmapped SIDs to prevent kernel log flooding.
 
  - Drain multichannel sessions during LOGOFF, wake deferred locks and
    cancellable requests, and ensure cancellation callbacks run only once.
 
  - Fix listener kthread reference handling and teardown ordering during
    netdevice events.
 
  - Validate normalized-name and IPC share configuration response lengths.
 
  - Update the KSMBD MAINTAINERS entry and add Paulo Alcantara as
    an SMBDIRECT co-maintainer.
 -----BEGIN PGP SIGNATURE-----
 
 iQJKBAABCgA0FiEE6NzKS6Uv/XAAGHgyZwv7A1FEIQgFAmqWnz0WHGxpbmtpbmpl
 b25Aa2VybmVsLm9yZwAKCRBnC/sDUUQhCJK8EACCE2K2p9CH6kiy9VnMjEqTbIBF
 ZRCmxrspoPAMuTbK6529dXHUVTsXlUdJ/FVzGwNLtvXwEIVjNaQDqBEFWCdPElE+
 8grKsC1S3gH3t8Z1wT6eNh5cpDoA+rWJDbNK4DsmHdoVagyjd9dd7fkMi7nq0WJS
 NO7BTHaTuTaZDul8UXc1gqkVLviZZWkrtkGVVnsJV1z5cFls6P81cVmtzP0836cU
 kVDYSI0EZnX+1P5CtOxL3r5LDBex6lRHU+rj1ypJRJDM2nR+bYIeJk+XMjylKCHT
 liPj7dwI/ptVzp+n3dbcTyhLZayDhZ0/GeJanX2/midtiNSKhao9h94BymPU91jV
 JugPlkAO8Vqwo7xojWRqudz4Kg/vgr66NexQ/3W2tuRXXFN4kEWmQG0N5+kH0K3d
 sJ5xA9uLj24+d29fjylkdSGpuRLR8XcR01he2CaqLRopXZrCxFChwzZwbads1rI/
 kXtYrORB0u99ScwTRQeW90dzeZ+1R3aHOyf8H86zyJ07l2NxG8t5L/49vuaiaiEZ
 5r4hhPVumlmDQdPoOcugOmkJL68+W4TzS7UfcOSgq43W31BE1dYsaX7u4MHNk/Nq
 UiWJJPArJ3ry8e4GLqQXx4ylZJnykGS9s676gFCO1GAI+eXoQ4R0k7RCKy/TW0ap
 t3oa0Yj+Y4QAZCD3Rw==
 =UL/J
 -----END PGP SIGNATURE-----

Merge tag 'ksmbd-for-7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/smb

Pull smb server fixes from Namjae Jeon:

 - Prevent unintended data exposure by clearing pipe compound padding
   and the response buffer

 - Initialize missing fields in FS_OBJECT_ID_INFORMATION,
   FS_CONTROL_INFORMATION, and FS_POSIX_INFORMATION

 - Propagate DACL parsing and allocation failures so malformed security
   descriptors are rejected

 - Rate-limit errors for unmapped SIDs to prevent kernel log flooding

 - Drain multichannel sessions during LOGOFF, wake deferred locks and
   cancellable requests, and ensure cancellation callbacks run only once

 - Fix listener kthread reference handling and teardown ordering during
   netdevice events

 - Validate normalized-name and IPC share configuration response lengths

 - Update the KSMBD MAINTAINERS entry and add Paulo Alcantara as an
   SMBDIRECT co-maintainer

* tag 'ksmbd-for-7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/smb:
  ksmbd: validate normalized name response length
  ksmbd: fix listener task lifetime on netdev events
  ksmbd: prevent out-of-bounds reads in share config responses
  ksmbd: rate limit unmapped SID errors
  ksmbd: propagate DACL parsing errors
  ksmbd: zero pipe read compound padding
  ksmbd: safely drain sessions during logoff
  MAINTAINERS: Update the KSMBD entry
  MAINTAINERS: Add Paulo Alcantara as an SMBDIRECT co-maintainer
  ksmbd: fill in FileSysIdentifier in FS_POSIX_INFORMATION
  ksmbd: initialize FileSystemControlFlags in FS_CONTROL_INFORMATION
  ksmbd: zero the FS_OBJECT_ID_INFORMATION buffer before filling it in
2026-09-01 08:17:01 -07:00