Commit Graph

1483181 Commits

Author SHA1 Message Date
Weiming Shi
c82b797abe net/sched: reject IDR error pointers when deleting actions
tcf_action_delete() drops the reference held by its lookup before calling
tcf_idr_delete_index() with the saved action index.  An unlocked
classifier can remove that action and reserve the same IDR slot with
ERR_PTR(-EBUSY) in between.

tcf_idr_delete_index() only checks the lookup result for NULL.  It
therefore treats the reservation as a tc_action and dereferences
tcfa_bindcnt.  A hardware execution breakpoint was used to schedule the
interleaving without changing the kernel source.  KASAN reported this
decoded trace:

  BUG: KASAN: null-ptr-deref in tca_action_gd+0x5b9/0x1010
  Read of size 4 at addr 0000000000000010 by task poc/150
  Oops: general protection fault, probably for non-canonical address 0xdffffc0000000002
  RIP: tca_action_gd+0x5c0/0x1010:
    arch_atomic_read at arch/x86/include/asm/atomic.h:23
    raw_atomic_read at include/linux/atomic/atomic-arch-fallback.h:457
    atomic_read at include/linux/atomic/atomic-instrumented.h:33
    tcf_idr_delete_index at net/sched/act_api.c:766
    tcf_action_delete at net/sched/act_api.c:1859
    tcf_del_notify at net/sched/act_api.c:2014
    tca_action_gd at net/sched/act_api.c:2064
  R13: 0000000000000010 R15: fffffffffffffff0
  Kernel panic - not syncing: Fatal exception

R15 contains ERR_PTR(-EBUSY), and adding the tcfa_bindcnt offset produces
the address in R13.  With the guard applied, the same reproducer returned
-ENOENT without a KASAN report or panic.  Treat error pointers as absent
and return -ENOENT.

Fixes: 0190c1d452 ("net: sched: atomically check-allocate action")
Cc: stable@vger.kernel.org
Reported-by: Xiang Mei <xmei5@asu.edu>
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Link: https://patch.msgid.link/20260914065123.4109709-2-bestswngs@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-19 14:48:44 -07:00
Quentin Armitage
6c096bb08d net: allow IFLA_INET_CONF messages when NLA_F_NESTED unset
Commit fa8fca8871 ("ipv4: validate IPV4_DEVCONF attributes properly")
added validation of IFLA_INET_CONF attributes, and in the process
changed the call of nla_for_each_nested() to nla_parse_nested(). A
side effect of this change is that the IFLA_INET_CONF option is now
tested for NLA_F_NESTED being set, and fails if it is not. Prior to the
commit there was no check of NLA_F_NESTED.

Change nla_parse_nested() to nla_parse(). This restores the previous
functionality of not checking NLA_F_NESTED, thereby allowing code that
(incorrectly) doesn't set NLA_F_NESTED to continue to work.

This issue was identified because keepalived started logging errors when
it was configuring macvlans that it created.

Fixes: fa8fca8871 ("ipv4: validate IPV4_DEVCONF attributes properly")
Signed-off-by: Quentin Armitage <quentin@armitage.org.uk>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260915213320.1527029-2-quentin@armitage.org.uk
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-18 18:22:59 -07:00
Ralf Lici
4581c3d2ad net/mlx5e: advertise MACsec offload only when supported
Commit 339ccec8d4 ("net/mlx5: Enable MACsec offload feature for VLAN
interface") added NETIF_F_HW_MACSEC unconditionally to vlan_features so
that VLAN devices could inherit MACsec offload support.

mlx5e_build_nic_netdev subsequently copies vlan_features into
hw_features and features. As a result, all mlx5e NIC netdevices
advertise MACsec hardware offload, even when the firmware does not
support it and the driver does not install macsec_ops.

Set the MACsec feature bits in mlx5e_macsec_build_netdev, after device
capabilities have been validated. This preserves MACsec-over-VLAN
support and the ethtool feature control on capable devices, without
advertising either on unsupported hardware.

Fixes: 339ccec8d4 ("net/mlx5: Enable MACsec offload feature for VLAN interface")
Cc: stable@vger.kernel.org
Reviewed-by: Tariq Toukan <tariqt@nvidia.com>
Signed-off-by: Ralf Lici <ralf@mandelbit.com>
Link: https://patch.msgid.link/20260917122724.654639-1-ralf@mandelbit.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-18 18:16:33 -07:00
Yiqi Sun
d2c31b8374 sctp: avoid livelock while updating retransmit path
sctp_assoc_update_retran_path() can loop forever when every remaining
transport, including retran_path, is SCTP_UNCONFIRMED: the state check
runs before the wraparound test, so the loop cannot observe that it has
completed a full pass.

Fix this by considering a transport only when it is not UNCONFIRMED,
then checking whether the walk has returned to retran_path. This makes
the full-pass termination independent of the transport state while
preserving the existing fallback selection semantics.

Also restore the NULL guard around the retran_path assignment. In the
all-UNCONFIRMED case there is no eligible replacement transport, and
installing NULL would leave later retransmit-path users and the debug
print with a NULL path.

Fixes: 4c47af4d5e ("net: sctp: rework multihoming retransmission path selection to rfc4960")
Signed-off-by: Yiqi Sun <sunyiqixm@gmail.com>
Acked-by: Xin Long <lucien.xin@gmail.com>
Link: https://patch.msgid.link/20260915095017.942213-1-sunyiqixm@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-18 17:24:10 -07:00
Jakub Kicinski
6c01564da3 Merge branch 'eth-fbnic-a-collection-of-fixes'
Alexander Duyck says:

====================
eth: fbnic: a collection of fixes

This series collects a handful of independent fbnic fixes for issues on
released kernels, plus one core ethtool fix needed by the fbnic offline
self test.

The first patch keeps rtnl_lock held on the ethtool ioctl path for the self
test. Since the ioctl path became rtnl-optional for ops-locked drivers,
fbnic's offline self test (which brings the interface down and up via
netif_close()/netif_open()) runs holding only the instance lock, tripping a
lockdep splat / ASSERT_RTNL and reconfiguring the device without the lock
it requires. A similar issue was found with Broadcom drivers so we expanded
the scope for v2 to just have the rtnl lock held for all selftest calls.

The second addresses a comparison issue in that we were limiting the
maximum number of standalone Tx queues to one less than the maximum number
of Tx queues. To resolve this it was just a matter of replacing a "<" with
a "<=".

The third addresses an indexing issue with netdev queues on fbnic in which
the NAPI vector was assumed to be findable as the Rx index modulo the
number of NAPI vectors. However this is actually not the case for if Tx
only and Rx only queues are setup. To resolve this we make use of the
cached NAPI pointer in the netdev Rx queues themselves.

The fourth patch fixes a NULL pointer dereference on unbind after a failed
PCIe error recovery: fbnic_pm_suspend() frees the napi vectors via a direct
ndo_stop() while leaving netif_running() true, and when slot_reset ->
resume fails the data path is never re-allocated. To prevent the panic we
reset num_napi to 0 before we free the IRQs which prevents walking the
unallocated napi vectors when we unbind the interface later.

The last two patches address the FW mailbox. One sets AW_FLUSH_MODE
alongside AW_FLUSH when tearing down the Rx ring, so the write pipeline
actually drains the staged requests instead of hanging on the BME halt.
The other handles completions flagged with FW_ERR on both mailboxes, which
the driver previously ignored. This resulted in us parsing a stale Rx page,
and spinning the capabilities poll to a timeout on a healthy ring.
====================

Link: https://patch.msgid.link/178941996343.7700.9376081102002673062.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-18 17:20:17 -07:00
Alexander Duyck
1b97a269a5 eth: fbnic: Handle FW mailbox completions flagged with an error
The firmware can complete a mailbox descriptor while also setting FW_ERR
to indicate it could not process the request, for example on a mailbox
DMA error. The completion carries no valid data.

The driver did not check FW_ERR. On the Rx mailbox it would sync and
parse the stale page as a normal message, and on the Tx mailbox it
silently freed the request. If the initial capabilities exchange in
fbnic_mbx_poll_tx_ready() hit FW_ERR -- on the Tx request or on the Rx
response descriptor -- no response was parsed and the poll spun until it
timed out even though the ring was healthy.

Check FW_ERR on both mailboxes. Count it per-mailbox in
fbnic_fw_mbx.resp_error, which is also shown in debugfs, warn (rate
limited, since the bit is firmware controlled), and drop the Rx page
instead of parsing it.

In fbnic_mbx_poll_tx_ready() re-issue the capabilities request when
either the Tx or the Rx resp_error counter advances, so a FW_ERR on the
request or on its response triggers a retry rather than a timeout. A
valid capabilities response is honored before the retry check, so a
response parsed in the same poll as an unrelated FW_ERR is not discarded.
The counters are mailbox-wide rather than keyed to the capabilities
request; that is sufficient here because the exchange runs during
bring-up before any other mailbox traffic, and any spurious retry is
bounded by the existing 10s timeout.

Fixes: da3cde0820 ("eth: fbnic: Add FW communication mechanism")
Signed-off-by: Alexander Duyck <alexanderduyck@fb.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942023343.7700.9423398932961964439.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-18 17:20:16 -07:00
Alexander Duyck
8947f13e43 eth: fbnic: Set AW_FLUSH_MODE alongside AW_FLUSH when flushing the mailbox
When tearing down the FW mailbox Rx ring, fbnic_mbx_reset_desc_ring()
writes AW_CFG with FLUSH set and everything else, BME included, cleared.
Clearing BME halts the device's writes to the host but leaves the staged
requests parked in the PUL write pipeline rather than draining them, so
on the write path FLUSH alone never terminates the outstanding requests
and the flush the firmware waits on never completes.

Add the FLUSH_MODE definition and set both bits so the staged writes
drain out of the pipeline on their own. BME stays cleared, so nothing
lands on the host; it is restored later in fbnic_mbx_init_desc_ring()
when the ring is rebuilt, once the outstanding writes are gone.

The read path is unaffected. AR_CFG has no equivalent mode bit and
AR_FLUSH terminates the outstanding reads by itself, so it is left as
is.

Both writes remain plain stores rather than read-modify-writes. That is
deliberate: the matching write in fbnic_mbx_init_desc_ring() restores
BME and the TLP attributes, and clears both flush bits as a side effect.

Fixes: 3b12f00ddd ("fbnic: Gate AXI read/write enabling on FW mailbox")
Signed-off-by: Alexander Duyck <alexanderduyck@fb.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942022583.7700.11050671998277309744.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-18 17:20:16 -07:00
Alexander Duyck
4bcc4a92c6 eth: fbnic: reset num_napi when the napi vectors are freed
fbn->num_napi is the count of live napi vectors, each of which owns an
IRQ.  The PM path had freed them without clearing the count.
fbnic_pm_suspend() tears the datapath down via ndo_stop() and frees the
IRQs, but leaves netif_running() true so resume knows to re-open.  Resume
rebuilds the datapath in __fbnic_pm_resume() and fbnic_reset_queues() sets
num_napi and __fbnic_open() re-allocates the vectors.

When the datapath is torn down but never rebuilt, num_napi is left
pointing at freed vectors under 2 different scenarios:
 - a PCIe error recovery that fails (fbnic_err_slot_reset() ->
   __fbnic_pm_resume() returns an error -> PCI_ERS_RESULT_DISCONNECT), so
   .resume never runs; or
 - an __fbnic_open() that fails partway on resume and unwinds, freeing
   the vectors after fbnic_reset_queues() has already set num_napi.

The netdev is then running with num_napi > 0 but napi[] freed, and the
eventual remove/unbind close re-enters fbnic_down() -> fbnic_dbg_down()
and dereferences the freed vectors:
  BUG: kernel NULL pointer dereference, address: 0000000000000210
  RIP: fbnic_dbg_down+0x28

Clear num_napi when the vectors are freed: in the suspend teardown (a
good resume re-establishes it before __fbnic_open()) and on the resume
open failure.  A redundant ndo_stop() then walks an empty napi[].  The
normal ndo_stop() down/up cycle is untouched and keeps num_napi for the
next ndo_open().

Fixes: bc6107771b ("eth: fbnic: Allocate a netdevice and napi vectors with queues")
Signed-off-by: Alexander Duyck <alexanderduyck@fb.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942021809.7700.10804028989308077839.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-18 17:20:16 -07:00
Alexander Duyck
b5d9e9d4d0 eth: fbnic: use the Rx queue napi pointer to find the napi vector
The queue management ndos pick the napi vector for an Rx queue with:
	nv = fbn->napi[idx % fbn->num_napi];

The issue is this is only correct in the cases where there are no
standalone Tx vectors. In those cases we were allocating the Tx vectors
first and then the Rx so the queues would be pointing to Tx NAPI vectors
instead of the Rx ones.

The mapping the ndos want is already recorded.  fbnic_set_netif_napi()
publishes it with netif_queue_set_napi(), which stores the napi pointer
in netdev_rx_queue.napi, and fbnic_reset_netif_napi() clears it again.
Both run under the netdev instance lock that the queue management ndos
also hold, so the pointer can be read directly.

Use it and drop the divide.  The pointer is NULL exactly while the
datapath is down, so fbnic_queue_mem_alloc() can reject that case rather
than reaching into freed state: netdev_rx_queue_restart() calls it
before it tests netif_running(), and fbnic_pm_suspend() leaves
netif_running() true across a PCIe recovery that never completes, so a
queue restart can arrive after fbnic_stop() has freed the rings and the
vectors.  fbnic_stop() clears the association in
fbnic_reset_netif_queues() before fbnic_free_napi_vectors(), so the
NULL is always published first.  fbnic_queue_start() and
fbnic_queue_stop() need no check of their own, as
netdev_rx_queue_reconfig() only reaches them once fbnic_queue_mem_alloc()
has succeeded under the same instance lock.

Fixes: da43127a8e ("eth: fbnic: support queue ops / zero-copy Rx")
Signed-off-by: Alexander Duyck <alexanderduyck@fb.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942021136.7700.4391219358260544104.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-18 17:20:16 -07:00
Björn Töpel
1f4c73064a eth: fbnic: Handle maximum standalone channels
Standalone channels use one NAPI vector for each Tx and Rx queue.
fbnic's allocation path excludes FBNIC_MAX_TXQS from that layout. A
64-Tx/64-Rx configuration therefore records 128 vectors but allocates
only 64, leaving NULL entries that resource setup dereferences.

Include the maximum vector count in standalone allocation.

Fixes: bc6107771b ("eth: fbnic: Allocate a netdevice and napi vectors with queues")
Signed-off-by: Björn Töpel <bjorn@kernel.org>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942020457.7700.13129750616387075931.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-18 17:19:58 -07:00
Alexander Duyck
1b82958f3f net: ethtool: keep rtnl_lock for the ioctl self test
An offline self test that brings the interface down and back up with
netif_close() / netif_open() requires rtnl_lock for both. Since the
ethtool IOCTL path became rtnl-optional for ops-locked drivers, the
ETHTOOL_TEST ioctl runs holding only the netdev instance lock, so on an
ops-locked driver the self test now tears the device down without
rtnl_lock.

With lockdep this reproduces deterministically on every offline self
test on such a driver; note the sole lock held is the instance lock, not
rtnl:

  WARNING: suspicious RCU usage
  net/core/netpoll.c:207 suspicious rcu_dereference_protected() usage!
  1 lock held by ethtool/107:
   #0: (&dev->lock){+.+.}, at: dev_ethtool
  Call Trace:
   netpoll_poll_disable
   __dev_close_many
   netif_close_many
   netif_close
   fbnic_self_test
   dev_ethtool_locked
   dev_ethtool
   dev_ioctl
   sock_ioctl
   __x64_sys_ioctl

Without lockdep the same condition trips ASSERT_RTNL() in
__dev_close_many() / __dev_open(); that check only samples the global
rtnl state, so it can be masked by a concurrent rtnl holder, but the
device is still being reconfigured without the lock it requires.

The ethtool self_test is a legacy ioctl-only command, so an ETHTOOL_TEST
case is only needed on the ioctl path. Add an opt-in bit for drivers whose
self test needs rtnl_lock and set it on the ops-locked drivers whose
offline self test tears the interface down and up:

  - fbnic (ops-locked via queue_mgmt_ops): fbnic_self_test() offline path
    uses netif_close() / netif_open().
  - bnxt (ops-locked via queue_mgmt_ops): bnxt_self_test() offline path
    goes through bnxt_close_nic() / bnxt_half_open_nic() /
    bnxt_half_close_nic() / bnxt_open_nic(), which close and reopen the
    device.

Fixes: f994752b11 ("net: ethtool: optionally skip rtnl_lock on IOCTL path")
Signed-off-by: Alexander Duyck <alexanderduyck@fb.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942019771.7700.338431553546884773.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-18 17:19:15 -07:00
Abhishek Ojha
95c4d54ed0 net: phy: micrel: Advance register data pointer in write loop
lanphy_write_reg_data() does not advance the data pointer while iterating
over the register table. As a result, it writes the first entry num times
and leaves the remaining errata registers unconfigured.

Single-entry tables are unaffected, but tables with multiple entries
leave every entry after the first unapplied.

Advance the data pointer after each successful write so every table entry
is applied in order.

Fixes: c8732e9339 ("net: phy: micrel: lan8842 errata")
Cc: stable@vger.kernel.org
Signed-off-by: Abhishek Ojha <abhishek.ojha@savoirfairelinux.com>
Reviewed-by: Andrew Lunn <andrew@lunn.ch>
Link: https://patch.msgid.link/20260916231928.1336305-1-abhishek.ojha@savoirfairelinux.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-18 17:07:41 -07:00
Kuniyuki Iwashima
dd47bcf279 ip6_gre: Call ip6erspan_tunnel_unlink_md() in ip6erspan_changelink().
The cited commit accidentally added ip6gre_tunnel_unlink_md()
in ip6erspan_changelink().

Let's correct it to ip6erspan_tunnel_unlink_md().

Fixes: b80d0b93b9 ("net: ip6_gre: fix tunnel metadata device sharing.")
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260916230927.378957-1-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-18 17:02:59 -07:00
Norbert Szetei
ee319bd3a0 ipv6: do not let ipv6_find_hdr() return an offset past the packet end
ipv6_find_hdr() walks the extension header chain, skipping each header by
the length that header itself declares.  ipv6_optlen() returns up to 2048,
and the skip is never checked against skb->len, so the offset stored in
*offset can point past the end of the packet.

openvswitch installs that offset as the transport header, and
update_ipv6_checksum() then reads and writes the transport checksum field
out of bounds:

  BUG: KASAN: slab-use-after-free in inet_proto_csum_replace16+0x445/0x470
  Read of size 2 at addr ffff88810b754b06 by task ovs_ipv6_oob/629
  CPU: 4 UID: 1000 PID: 629 Comm: ovs_ipv6_oob Tainted: G N 7.3.0-rc3+ #348
  Call Trace:
   inet_proto_csum_replace16+0x445/0x470
   set_ipv6_addr+0x3dd/0x460
   do_execute_actions+0x6a3d/0x7c40
   ovs_execute_actions+0xfd/0x480
   ovs_packet_cmd_execute+0xc38/0xf20
   genl_rcv_msg+0x59e/0x870
   netlink_rcv_skb+0x18b/0x450
   genl_rcv+0x2d/0x40
   netlink_unicast+0x6bc/0xa20

  The buggy address belongs to the object at ffff88810b754980
   which belongs to the cache skbuff_small_head of size 704
  The buggy address is located 390 bytes inside of
   freed 704-byte region [ffff88810b754980, ffff88810b754c40)

Other callers use that offset too, so bound it here rather than in one
caller.

Reject a header whose declared length does not fit in the packet.
ipv6_find_hdr() already fails with -EBADMSG on a malformed chain, so this
adds no new failure mode.

Fixes: f8f626754e ("ipv6: Move ipv6_find_hdr() out of Netfilter code.")
Suggested-by: Ilya Maximets <i.maximets@ovn.org>
Suggested-by: Eric Dumazet <edumazet@google.com>
Cc: stable@vger.kernel.org
Signed-off-by: Norbert Szetei <norbert@doyensec.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Ilya Maximets <i.maximets@ovn.org>
Link: https://patch.msgid.link/8F80BA1A-DDFD-432D-9075-242A3435FEB5@doyensec.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-18 16:53:22 -07:00
Andy Moreton
24fedc7a56 sfc: add X4D PF support
X4D is an X4 controller instance as an IP block in an SoC.
It has the same feature set as X4.

Signed-off-by: Andy Moreton <andy.moreton@amd.com>
Reviewed-by: Pieter Jansen van Vuuren <pieter.jansen-van-vuuren@amd.com>
Reviewed-by: Alejandro Lucero <alucerop@amd.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260916125641.12238-1-alucerop@amd.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-18 16:52:12 -07:00
Lorenzo Bianconi
310d1ac61a net: ethernet: mtk_eth_soc: unregister net_devices in case of probe failure
If register_netdev() fails for one of the MTK_MAX_DEVS devices in
mtk_probe(), the error path jumps to err_deinit_ppe, skipping
mtk_unreg_dev(). The previously registered net_devices are then freed by
mtk_free_dev() while still in NETREG_REGISTERED state, hitting the
BUG_ON(dev->reg_state != NETREG_UNREGISTERED).

Route the register_netdev() failure to err_unreg_netdev so the net_devices
registered so far are properly unregistered before being freed.

Fixes: 8a8a9e89f8 ("net: ethernet: mediatek: cleanup error path inside mtk_hw_init")
Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Link: https://patch.msgid.link/20260916-mtk_eth_soc-netdev-fix-v1-1-5dac50eb65b1@oss.qualcomm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-18 16:51:28 -07:00
Jérémy Jean
2566866fc3 net: gue: reject invalid REMCSUM offsets
The REMCSUM option carries an absolute checksum start and checksum field
offset. gue_remcsum() passes them to skb_remcsum_process(), whose
partial path stores offset - start in the u16 skb->csum_offset variable.
If offset is less than start, this underflows.

A forwarded packet can retain CHECKSUM_PARTIAL and reach a NETIF_F_HW_CSUM
driver which trusts the metadata, leading skb_copy_and_csum_dev() to write
two bytes about 64 KiB beyond the destination buffer.

Reject reversed tuples in validate_gue_flags(), after the existing length
validation, so all GUE parsers enforce the ordering in one place.

Fixes: fe881ef11c ("gue: Use checksum partial with remote checksum offload")
Signed-off-by: Jérémy Jean <Jeremy.Jean@oss.cyber.gouv.fr>
Link: https://patch.msgid.link/20260915124806.2852293-2-Jeremy.Jean@oss.cyber.gouv.fr
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-18 16:44:08 -07:00
Nguyen Ngoc Thang
47abe7a5c4 net/sched: act_ct: don't WARN on benign flow_offload_alloc() failure
flow_offload_alloc() returns NULL when the conntrack entry is dying
(e.g. raced with a conntrack flush) or when the GFP_ATOMIC allocation
fails; both are expected under load and neither is a kernel bug. This
path runs from softirq on every committed packet, so with
panic_on_warn=1 an unprivileged user can panic the box just by racing
a conntrack flush against a `tc ... action ct commit` classifier.

Reproduced with a custom repro under QEMU: a small, fixed set of UDP
flows through `tc filter ... action ct commit` on lo, raced against
threads flooding bare ctnetlink CT_DELETE (flush) requests. Hits
WARNING: net/sched/act_ct.c:437 (tcf_ct_flow_table_add(), inlined
into tcf_ct_act() in this build) within ~15s on the unpatched kernel;
same setup is clean on the patched kernel. The fix itself is
behavior-preserving: both branches already did `goto err_alloc`
before and after, only the WARN is removed.

Fixes: 64ff70b80f ("net/sched: act_ct: Offload established connections to flow table")
Reported-by: syzbot+6cc37aba98dac721c415@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=6cc37aba98dac721c415
Signed-off-by: Nguyen Ngoc Thang <ngocthang2710.1999@gmail.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260915150816.36487-1-ngocthang2710.1999@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-18 16:35:19 -07:00
Eric Dumazet
99cc2a62e0 tipc: reject invalid and unexpected GRP_ACK_MSG to prevent bc_ackers underflow
Commit 48a5fe3877 ("tipc: fix bc_ackers underflow on duplicate
GRP_ACK_MSG") rejected duplicate/stale ACKs in tipc_group_proto_rcv()
by returning early when less_eq(acked, m->bc_acked).

However, that check remains incomplete in two ways:

1. When grp->bc_ackers is zero (e.g. on a quiet group, when replicast
   ACKs were not requested, or after all expected members have already
   acknowledged), an unexpected GRP_ACK_MSG with acked > m->bc_acked
   passes less_eq() and unconditionally decrements grp->bc_ackers.
   Because bc_ackers is a u16, this wraps to 65535, causing
   tipc_group_bc_cong() to permanently report congestion and blocking
   all future group broadcasts on the socket.

2. During an active broadcast round (grp->bc_ackers > 0), the sender
   transmits packet S and advances grp->bc_snd_nxt to S + 1. Receivers
   increment their expected counter to S + 1 upon consuming packet S,
   so the only valid ACK value for the current round is strictly
   acked == grp->bc_snd_nxt.

   However, tipc_group_update_bc_members() initializes each member's
   m->bc_acked to prev = grp->bc_snd_nxt - 1 (S - 1 before increment).
   This leaves a 2-sequence gap (S - 1 to S + 1) in sequence space.
   An incoming ACK is therefore neither rejected as duplicate nor
   prevented from decrementing grp->bc_ackers if an unexpected or stale
   value (such as S) is received. A member sending acked = S followed
   by acked = S + 1 could decrement grp->bc_ackers twice in the same
   round, prematurely clearing bc_ackers or underflowing it.

Fix this by:
- Dropping GRP_ACK_MSG immediately if grp->bc_ackers is zero.
- Requiring acked == grp->bc_snd_nxt and rejecting duplicates where
  m->bc_acked == acked. Because replicast broadcast rounds are strictly
  sequential, only grp->bc_snd_nxt can be acknowledged, and each member
  can acknowledge at most once per round.

Note that a related pre-existing issue in tipc_group_delete_member()
(where grp->bc_ackers decrementing to zero upon member departure does
not restore *grp->open or trigger a socket wakeup) will be addressed
in a separate patch.

Fixes: 48a5fe3877 ("tipc: fix bc_ackers underflow on duplicate GRP_ACK_MSG")
Fixes: 2f487712b8 ("tipc: guarantee that group broadcast doesn't bypass group unicast")
Reported-by: James Burton <jamesburton@meta.com>
Cc: stable@vger.kernel.org
Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260913044233.193927-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-18 15:42:47 -07:00
Linkui Xiao
46bc52d135 ipv4: fib: fix data-race and stale genid check around nh->nh_saddr
fib_select_multipath() compares nexthop_nh->nh_saddr against the flow
source address with no lock held, while fib_info_update_nhc_saddr()
stores a new value from another CPU as soon as the preferred source
address of the egress device changes.

Commit 195374d893 ("ipv4: fib: annotate races around nh->nh_saddr_genid
and nh->nh_saddr") added WRITE_ONCE() on the store side and READ_ONCE()
in fib_result_prefsrc() after syzbot reported

	BUG: KCSAN: data-race in fib_select_path / fib_select_path

but it only covered that reader. fib_select_multipath(), reached from
fib_select_path(), is a second lockless reader of nh->nh_saddr and was
left bare.

Moreover, nh_saddr is only meaningful when nh_saddr_genid matches
dev_addr_genid, as established by commit 436c3b66ec ("ipv4: Invalidate
nexthop cache nh_saddr more correctly."). fib_select_multipath()
skips that validation, so it can score a nexthop using a stale source
address and skew the ECMP selection.

Annotate both reads with READ_ONCE() and refresh the cached source
address via fib_info_update_nhc_saddr() when the genid does not match,
mirroring fib_result_prefsrc().

Fixes: 32607a332c ("ipv4: prefer multipath nexthop that matches source address")
Signed-off-by: Linkui Xiao <xiaolinkui@kylinos.cn>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260916125316.988044-1-xiaolinkui@126.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17 19:19:04 -07:00
Zhang Yunfei
651010592b net: txgbe: fix FDIR filter restore for VF rules
txgbe_fdir_filter_restore() reprograms every filter from
txgbe->fdir_filter_list after a reset. It extracts the ring part of
filter->action with ethtool_get_flow_spec_ring() and maps it onto a
PF rx ring, silently dropping the VF part of the cookie that
txgbe_add_ethtool_fdir_entry() stores there (input->action =
fsp->ring_cookie).

For a rule directed at a VF, restore therefore reprograms the filter
to the PF queue with the same ring index: after any down/up or
txgbe_reinit_locked(), traffic matching the rule is steered to the
PF instead of the VF.

Handle VF rules the same way txgbe_add_ethtool_fdir_entry() does:
validate vf against wx->num_vfs and ring against
wx->num_rx_queues_per_pool, and map the ring onto the absolute
queue index ((vf - 1) * wx->num_rx_queues_per_pool) + ring.

Fixes: 7a91722e0d ("net: txgbe: Support the FDIR rules assigned to VFs")
Cc: stable@vger.kernel.org
Signed-off-by: Zhang Yunfei <zhangyunfei1@kylinos.cn>
Link: https://patch.msgid.link/20260911091123.798931-1-zhangyunfei1@kylinos.cn
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17 19:16:48 -07:00
Aldo Ariel Panzardo
2ec28c09b3 vsock: ignore empty child namespace mode writes
__vsock_net_mode_string() returns success without updating new_mode when
the transfer length is zero. Its caller then reads the uninitialized enum
and may permanently store a stack-derived value in the write-once child
mode.

Return before calling __vsock_net_mode_string() when *lenp is zero so
that the helper is never invoked with nothing to parse and new_mode is
never read uninitialized. This also prevents an empty write from
locking the current mode.

Fixes: eafb64f40c ("vsock: add netns to vsock core")
Cc: stable@vger.kernel.org
Reviewed-by: Luigi Leonardi <leonardi@redhat.com>
Signed-off-by: Aldo Ariel Panzardo <qwe.aldo@gmail.com>
Reviewed-by: Stefano Garzarella <sgarzare@redhat.com>
Reviewed-by: Bobby Eshleman <bobbyeshleman@meta.com>
Link: https://patch.msgid.link/20260915173050.3176344-1-qwe.aldo@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17 19:06:31 -07:00
Jakub Kicinski
b73bcf7c1f Merge branch 'net-mlx5-sd-lag-and-devcom-stability-fixes'
Tariq Toukan says:

====================
net/mlx5: SD LAG and devcom stability fixes

This series by Shay fixes four bugs in the Socket Direct LAG and devcom
subsystems, all related to initialization/teardown ordering and
concurrent access to the LAG device.
====================

Link: https://patch.msgid.link/20260915113459.3934760-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17 17:54:36 -07:00
Shay Drory
bae23d1ae6 net/mlx5: LAG, reload IB reps of LAG master before the rest
In a shared-FDB LAG the master device creates the bond IB device; the
other LAG members do not create their own, they populate a port inside
the master's IB device. mlx5_lag_reload_ib_reps_unlocked() reloaded the
members' IB reps in iteration order, with no guarantee the master is
reloaded first. When a non-master member is reloaded before the master,
it tries to populate its port in an IB device that has not been
recreated yet.

Hence, reload the master's IB reps first, then every other member.

Fixes: 2b204cdb12 ("net/mlx5: LAG, use xa_alloc to manage LAG device indices")
Signed-off-by: Shay Drory <shayd@nvidia.com>
Reviewed-by: Akiva Goldberger <agoldberger@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260915113459.3934760-4-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17 17:54:31 -07:00
Shay Drory
e1e29ada2b net/mlx5: SD, unload reps on shared FDB create error path
mlx5_lag_shared_fdb_create() sets sd_fdb_active on every group member
before reloading the representors, so mlx5_lag_is_active() is already
true and the guard in mlx5_esw_offloads_rep_load() does not skip the
VF/SF reps. If the reload then fails, the error path clears
sd_fdb_active and destroys the shared FDB, leaving the reps loaded
while SD LAG is inactive - the state cited commit was written
to prevent.

Unload the reps in the error path as well.

Fixes: 68c2dd59a6 ("net/mlx5: E-Switch, Tie rep load/unload to SD LAG state")
Signed-off-by: Shay Drory <shayd@nvidia.com>
Reviewed-by: Akiva Goldberger <agoldberger@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260915113459.3934760-3-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17 17:54:31 -07:00
Shay Drory
d09e8f6465 net/mlx5: devcom, Base component size on linked devices
mlx5_devcom_comp_get_size() returns the component's kref count. That
kref is bumped in mlx5_devcom_register_component() under comp_list_lock,
before the comp_dev is linked onto comp_dev_list_head under comp->sem.
The event broadcast (mlx5_devcom_locked_send_event()) walks that list.

Hence, a caller can read the expected size, but send_event won't be sent
to all peers. In the SD group registration path, this lets a member
broadcast its role-election event over an incomplete list, electing a
primary that never completes the group, is never marked ready, and
leaves the group with a stale primary.

Track the number of linked comp_devs in a dedicated counter, maintained
under comp->sem together with the list add/remove, and return it from
mlx5_devcom_comp_get_size().

Fixes: 9bb1ac8073 ("net/mlx5: devcom, Add component size getter")
Signed-off-by: Shay Drory <shayd@nvidia.com>
Reviewed-by: Akiva Goldberger <agoldberger@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260915113459.3934760-2-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17 17:54:31 -07:00
Heyang Tan
39c6580765 octeontx2-af: use seq_file for rsrc_alloc debugfs
The rsrc_alloc debugfs reader writes rows directly to userspace without
respecting the caller's read count. It also uses the current row length as
the userspace stride, which can corrupt output when rows have different
widths.

Use seq_file to handle userspace buffer sizes, offsets, and partial reads,
and write output columns directly to the seq_file buffer.

Fixes: 23205e6d06 ("octeontx2-af: Dump current resource provisioning status")
Signed-off-by: Heyang Tan <thy15333007817@163.com>
Reviewed-by: Ratheesh Kannoth <rkannoth@marvell.com>
Link: https://patch.msgid.link/20260914020521.146-1-thy15333007817@163.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17 17:49:27 -07:00
Kyle Hendry
daf677c2c6 net: pcs: rzn1-miic: Fix config array initialization
Fix memset parameters to initialize the entire DT value array

Fixes: f39e968dc1 ("net: pcs: rzn1-miic: Move configuration data to SoC-specific struct")
Reviewed-by: Geert Uytterhoeven <geert+renesas@glider.be>
Signed-off-by: Kyle Hendry <khendry@reliablecontrols.com>
Reviewed-by: Lad Prabhakar <prabhakar.mahadev-lad.rj@bp.renesas.com>
Link: https://patch.msgid.link/20260915-rzn1-miic-fix-array-v5-1-b7173fd5b97d@reliablecontrols.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17 17:07:43 -07:00
Jamal Hadi Salim
960ab631f3 selftests/tc-testing: add u32 manual table handle IDR tests
35fc: create a manual table with handle 801:, then add an auto-allocated
table. Before the fix, the auto allocation reuses id 1 and hands out the
same handle 0x80100000, aliasing the manual table; the test requires the
manual 801: handle to keep exactly one entry in the dump.

a6e8: with a live u32 table keeping the tc_u_common alive, add and delete
a manual table with handle 901:, then re-add it. Unpatched, the delete
leaks the raw-keyed IDR entry and the re-add fails with -ENOSPC; the
test requires the re-add to succeed.

Reviewed-by: Victor Nogueira <victor@mojatatu.com>
Tested-by: hybris <hybris@mojatatu.ai>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/QDISC-LQFE.v1.20260911041746.2@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17 17:00:28 -07:00
Jamal Hadi Salim
0a5f5d9e94 net/sched: cls_u32: fix manual hash table handle IDR aliasing
A u32 hash table created with an explicit handle ('tc filter add ...
handle 801: u32 divisor N') keys its IDR entry on the raw handle, while
the destroy paths free it under handle2id(handle). The two key domains
disagree for handles in the 0x800..0xFFF htid range:
handle2id() folds them back into the auto-allocated id space (1..0x7FF).

A manual table therefore leaves its raw-keyed IDR entry unreachable on
delete (a permanent leak), and its delete can drop the idr entry of an
unrelated live auto table. A later auto allocation can then hand out a
handle that aliases the live manual table; u32_lookup_ht() first-match
routes lookups and TCA_U32_LINK for that htid to the wrong table.

Key the divisor-path alloc on handle2id(handle) so allocation and
removal share one key domain. A manual handle that maps onto an id
already in use is rejected with -ENOSPC, and auto allocation skips ids
held by live manual tables.

Conditions to recreate:
  ip link add test0 type dummy
  tc qdisc add dev test0 clsact
  tc filter add dev test0 ingress protocol ip pref 1 \
          handle 801: u32 divisor 16
  tc filter add dev test0 ingress protocol ip pref 2 u32 divisor 16
  tc -d filter show dev test0 ingress | grep 'fh 801:'
  # unpatched: two live tables with handle 0x80100000 (the pref 2 root
  # hnode is auto-allocated id 1); patched: the auto hnode takes id 2.

Also tested with a poc with a live u32 table on the block, add/delete a manual
table 'handle 901: u32 divisor 1' twice; unpatched, the re-add fails with
-ENOSPC because the raw key leaked on the first delete.

Fixes: 73af53d820 ("net: sched: cls_u32: Fix u32's systematic failure to free IDR entries for hnodes.")
Reported-by: Sashiko (gemini + nipa) <sashiko-bot@kernel.org>
Closes: https://sashiko.dev/#/patchset/20260822222049.114526-1-jhs@mojatatu.com
Reviewed-by: Victor Nogueira <victor@mojatatu.com>
Tested-by: hybris <hybris@mojatatu.ai>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/QDISC-LQFE.v1.20260911041746.1@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17 17:00:28 -07:00
Björn Töpel
8e0b235bd9 eth: fbnic: Fix payload page pool error cleanup
The payload page pool pointer contains an error pointer when its
allocation fails. The cleanup path passes that error pointer to
page_pool_destroy() instead of destroying the header page pool. This
can dereference the error pointer and leave the header page pool
allocated.

Destroy the header page pool instead.

Fixes: 8a11010fdd ("eth: fbnic: allocate unreadable page pool for the payloads")
Reported-by: Sashiko <netdev-bot+sashiko@kernel.org>
Link: https://lore.kernel.org/netdev/178915061000.219967.7726187707862333281@kernel.org/
Signed-off-by: Björn Töpel <bjorn@kernel.org>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260915104917.3978113-1-bjorn@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17 16:02:43 -07:00
Wei Wang
c0078f4d8c mailmap: add entry for Wei Wang
My Meta email address is no longer active. Map it to my current address
so that git and get_maintainer.pl stop pointing at a dead address for my
contributions.

Signed-off-by: Wei Wang <weiwan@google.com>
Link: https://patch.msgid.link/20260915201412.2201757-1-weiwan@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17 14:43:38 -07:00
Linus Torvalds
b5a051f6b8 Including fixes from Netfilter, Bluetooth, IPSec and WiFi.
Previous releases - regressions:
 
   - netfilter: hold reference on ct until flow is released
 
   - bridge:
     - move switchdev call outside rcu
     - vlan: fix bugs caused by switchdev deletion errors
 
   - wifi:
     - mac80211: reset state when starting AP fails
     - cfg80211: don't free driver-owned scan requests
 
   - tcp: don't call skb_clone_and_charge_r() for close()d listener in tcp_v6_do_rcv().
 
   - mptcp: return sk_wait_data() errors from recvmsg()
 
   - xfrm: serialize state GC with device state flush
 
   - drop_monitor: synchronize tracepoint unregistration on error path
 
   - bluetooth:
     - eir: validate service data length before reading UUID
     - hci_sync: serialize local codec list cleanup
     - RFCOMM: avoid socket lock inversion in listener cleanup
 
   - eth: lan743x: fix RX checksum use-after-free
 
   - eth: mvpp2: prevent buffer overflow in page_pool allocation
 
 Previous releases - always broken:
 
   - core: lock the socket in sock_gettstamp()
 
   - neighbour: enforce min/max to NDTPA_INTERVAL_PROBE_TIME_MS.
 
   - sched: codel: bound the dropping loop per dequeue call
 
   - wifi: mac80211: include TIM bitmap control for buffered S1G mcast traffic
 
   - psp: avoid conflicts with skb->decrypted and sk_validate_xmit_skb()
 
   - xfrm: fix stack OOB read in iptfs_skb_reset_frag_walk()
 
   - bluetooth: hci_qca: do not write to the serial port after it is closed
 
   - dsa: mxl862xx: disable the stats poll on teardown
 
   - eth: stmmac: fix TSO header length truncation
 
   - eth: ip_tunnel: initialize `options_len` before referencing options
 
 Signed-off-by: Paolo Abeni <pabeni@redhat.com>
 -----BEGIN PGP SIGNATURE-----
 
 iQJKBAABCgA0FiEEg1AjqC77wbdLX2LbKSR5jcyPE6QFAmqsEOsWHHBhb2xvLmFi
 ZW5pQGdtYWlsLmNvbQAKCRApJHmNzI8TpKMnD/4vxx/YloxnyDssUB95WcTfiHau
 XD+YDfTTf4Qy3DwKnMiesZ4w2547FXG+LwJZGpAGWpRSE99OawFHvx5gyn3Vf4IU
 HEsOuZEYMwhSeeMJxGdCg8K6McNEQx+aAD7D8gLoisnJr/DCKBJtxFgKVEjaKUfP
 aCG+FmbBqdS7hwBVbYREwmqwQSMnNzRWLd7/10/oYMcLfvgEsIkZisRErAW2BgZj
 mzGw/IcG+QRldDPGDwLwLzfEG9o1a2JctSRrQ3uLHZ3VOdmpnSkmf25s2IaEFAn6
 xfIhGPMgYIBDrakn/Ci4fAF0L98FtBq4Sa21HlvPMBst7rcce5x49ddSlIUWxRVZ
 Fbvs/0IMN0cEpYGbVJxE8iF3yo+t8XvsMdGS1JXd/ycaL9lpF+9gCzNvChSxba3N
 ik7BGAmlg76rqMuzeeMbWqMCmOcCBhQsb7iZXjNStJoiVY+UT2QfgvJ1TqF8Qrar
 Eu/xFkaqk8i/7jrx4ujceg9XpRt1Y3Y3Pq0sgdzlzTm162AV/kg1fvwHc6ORIML7
 71XpBSGvjyTHoh95ob/w3/5fec/ekzvWlCBL/cYIwBJ4R6n7Vk5e0Y1yZDffPSo3
 xQR0r4YdLewkjQdtnGEih6FSyWsu6nYpVSuwHbof1bQSzvlIM6TAabJmtJckFTld
 1bR2C9UVw5mQYMy8PA==
 =Gw4S
 -----END PGP SIGNATURE-----

Merge tag 'net-7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net

Pull networking fixes from Paolo Abeni:
 "Including fixes from Netfilter, Bluetooth, IPSec and WiFi.

  Previous releases - regressions:

   - netfilter: hold reference on ct until flow is released

   - bridge:
      - move switchdev call outside rcu
      - vlan: fix bugs caused by switchdev deletion errors

   - wifi:
      - mac80211: reset state when starting AP fails
      - cfg80211: don't free driver-owned scan requests

   - tcp: don't call skb_clone_and_charge_r() for close()d listener in
     tcp_v6_do_rcv()

   - mptcp: return sk_wait_data() errors from recvmsg()

   - xfrm: serialize state GC with device state flush

   - drop_monitor: synchronize tracepoint unregistration on error path

   - bluetooth:
      - eir: validate service data length before reading UUID
      - hci_sync: serialize local codec list cleanup
      - RFCOMM: avoid socket lock inversion in listener cleanup

   - eth:
      - lan743x: fix RX checksum use-after-free
      - mvpp2: prevent buffer overflow in page_pool allocation

  Previous releases - always broken:

   - core: lock the socket in sock_gettstamp()

   - neighbour: enforce min/max to NDTPA_INTERVAL_PROBE_TIME_MS.

   - sched: codel: bound the dropping loop per dequeue call

   - wifi: mac80211: include TIM bitmap control for buffered S1G mcast
     traffic

   - psp: avoid conflicts with skb->decrypted and sk_validate_xmit_skb()

   - xfrm: fix stack OOB read in iptfs_skb_reset_frag_walk()

   - bluetooth: hci_qca: do not write to the serial port after it is
     closed

   - dsa: mxl862xx: disable the stats poll on teardown

   - eth:
      - stmmac: fix TSO header length truncation
      - ip_tunnel: initialize `options_len` before referencing options"

* tag 'net-7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (159 commits)
  mptcp: fix bad accounting in __mptcp_subflow_push_pending()
  mptcp: close race between scheduler and state change
  mptcp: avoid unneeded actions on subflow reset
  net: skbuff: do not leave stale header offsets after pskb_carve()
  selftests: net: packetdrill: test exclusion of old ACK from TCP fast path
  tcp: exclude old ACKs from tcp fast path
  dpll: reject a reference sync pin which is not on the pin's dpll
  net: mvpp2: prevent buffer overflow in page_pool allocation
  net: macb: fix ordering around PTP timestamp read
  selftests: drv-net: psp: test PSP and TCP ULP mutual exclusion
  net: psp: avoid conflicts with skb->decrypted and sk_validate_xmit_skb()
  net: stmmac: preserve real_num_tx_queues on mqprio setup failure
  net: stmmac: propagate FPE preemption-class mapping errors
  net: wwan: t7xx: validate the netif index in t7xx_ccmni_recv_skb()
  net: wwan: mhi_wwan_mbim: check skb_copy_bits() return value
  net: wwan: mhi_wwan_mbim: guard against a cyclic NDP chain
  net: ethernet: cortina: Ack RX overrun interrupt correctly
  net: lock the socket in sock_gettstamp()
  eth: fbnic: ring the doorbell if a burst ends in a drop
  net: netsec: fix device_node reference leak on phy_np
  ...
2026-09-17 10:40:48 -07:00
Linus Torvalds
4982d3552a sound fixes for 7.3-rc4
A collection of small fixes.  Most of them are device-specific fixes
 while there are a few core fixes.  The continued flux, but not too
 scaring yet.  Some highlights below.
 
 ALSA Core:
 - Fix potential UAF after asynchronous card release
 - Fix a race condition in PCM timer initialization order
 
 USB-Audio:
 - Hardening fixes for issues reported by fuzzer for 6fire, bcd2000,
   and implicit FB packets
 - Fix double list addition in implicit FB handling
 - Quirks for AVerMedia GC553Pro and Behringer FCA1616
 
 HD-Audio:
 - Quirks / fixes for HP OmniBook 7, OMEN 15, and Victus 15 laptops
 
 ASoC:
 - Support for DAI link codec channel mask to avoid mismatches
 - Fix HDMI-codec channel status change report
 - Fixes for various codecs and platforms: Realtek rt712/rt721
   (calibration, reset fixes), Cirrus Logic (empty EFI variable
   validation, capture channel fixup), AMD ACP SoundWire (bounds
   checks, refactorings), ADAU1977 (OF match table support, SPI
   cleanups), ES8336 (Huawei Matebook B3-420 quirk), UX500 (macro fix)
 -----BEGIN PGP SIGNATURE-----
 
 iQJCBAABCAAsFiEEIXTw5fNLNI7mMiVaLtJE4w1nLE8FAmqrwIYOHHRpd2FpQHN1
 c2UuZGUACgkQLtJE4w1nLE8WeQ//eN0ufLXXqy5U6X7kri8eM4vjRDZa5v1z8oD/
 oHBdkgPmSAOJOddCVHuyPyx5BP4reBgKbukxUTBjuJvVRD4iOYfvXSW1LzVCVcly
 f60rJl1/Ck3RcXfVabNV9eeGCFtAwR00U/3oH7z0eDhjwisX2F4ucAMmgxyJUKTh
 pLdMPFsI/XW5CWeVbPVObGlOB8Bf/c79wy5NAnpixAkR1WaXUSLKoF5djO9qIQex
 eTIknPmMYLLSfzFfO0TWY1PPRPz5qJHDr6Acer0VMTHyZF0yg2bR+gRInbM1at2H
 c/lg8m899ZbobSCiHFEJdPH/W5x3iHqFi3hKVtKtQ5niot+gWTirQpjUCZ2HxNOp
 5fYEyieSJ0W/t2NfNW0SP+DUOvKaIxY9VUuGCW7z/YgjW5DvSxipd2FOC4Av0s6F
 l67vYazzPzf9d34NIM333FHeSZ4WMXVKKTfB34CQD93lVB1GladNA6lFFer5zuqG
 Sc8q+YGUF65OZEnbslANIDvqPG0eMNlqBn/iyqXQO+P/C6oXWIlqFm4DUtM5a8wB
 3Vf/e9L4mzXwKZfoW2xsJLgfpAO0faYiTAO6fr7wjBPrRl7VoenOlsCshN5GLTmM
 b4bQ29LZrpk9e3e8BdGdZzT/kfQEZ31O5sawaj6f+a2JG09xCi9qxklxV5wFvZ5E
 FTO+HgE=
 =PnHM
 -----END PGP SIGNATURE-----

Merge tag 'sound-7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/tiwai/sound

Pull sound fixes from Takashi Iwai:
 "A collection of small fixes. Most of them are device-specific fixes
  while there are a few core fixes. The continued flux, but not too
  scaring yet. Some highlights below.

  ALSA Core:
   - Fix potential UAF after asynchronous card release
   - Fix a race condition in PCM timer initialization order

  USB-Audio:
   - Hardening fixes for issues reported by fuzzer for 6fire, bcd2000,
     and implicit FB packets
   - Fix double list addition in implicit FB handling
   - Quirks for AVerMedia GC553Pro and Behringer FCA1616

  HD-Audio:
   - Quirks / fixes for HP OmniBook 7, OMEN 15, and Victus 15 laptops

  ASoC:
   - Support for DAI link codec channel mask to avoid mismatches
   - Fix HDMI-codec channel status change report
   - Fixes for various codecs and platforms: Realtek rt712/rt721
     (calibration, reset fixes), Cirrus Logic (empty EFI variable
     validation, capture channel fixup), AMD ACP SoundWire (bounds
     checks, refactorings), ADAU1977 (OF match table support, SPI
     cleanups), ES8336 (Huawei Matebook B3-420 quirk), UX500 (macro
     fix)"

* tag 'sound-7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/tiwai/sound: (33 commits)
  ASoC: adau1977-i2c: add OF match table for I2C
  ASoC: adau1977-spi: drop __maybe_unused and of_match_ptr()
  ASoC: adau1977: make the Kconfig symbols user selectable
  ASoC: amd: acp: fix card name length warning in SOF SoundWire machine driver
  ASoC: amd: acp: fix ffs() operator precedence for SoundWire link ID
  ASoC: amd: acp: refactor codec config count in SOF SoundWire machine driver
  ASoC: amd: acp: bounds-check SoundWire link ID in machine drivers
  ASoC: cs-amp-lib: Prevent NULL pointer if efi variable is zero length
  ASoC: codecs: rt712-sdca-dmic: fix uninitialized stream_config->type
  ASoC: hdmi-codec: Report a change when the channel status moves
  ASoC: ux500: Parenthesize MSP_{RX,TX}_CLKPOL_BIT() arguments
  ASoC: rt721: Reset codec to fix abnormal sound
  ALSA: usb-audio: fix list_add double-add in push_back_to_ready_list
  ALSA: hda: trace PCM open only after assigning a stream
  ALSA: usb-audio: skip the broken mute control on AVerMedia GC553Pro
  ALSA: hda/realtek: Enable mute LEDs on HP OmniBook 7 17-dc0xxx
  ALSA: 6fire: fix OOB write from device-reported iso length
  ALSA: usb-audio: Add capture quirk for Behringer FCA1616
  ALSA: hda/realtek: Add mute LED quirk for HP OMEN 15-ax
  ASoC: Intel: sof_es8336: Add a quirk for Huawei Matebook B3-420
  ...
2026-09-17 09:57:09 -07:00
Linus Torvalds
f143ea21cf power sequencing fixes for v7.3-rc4
- fix kconfig issue in pwrseq-thread-gpu
 - fix error path logic in pwrseq_unit_enable()
 - fix two NULL-pointer dereference bugs in power sequencing core
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEkeUTLeW1Rh17omX8BZ0uy/82hMMFAmqrtNgACgkQBZ0uy/82
 hMOupg/+Psfi2riTDG4/j6pzCWlLmOnL3Rfh1QLWMwGRKM7XHC8yN8CUHQS8ub2O
 l+z8ZZgJ663noQTLkjGbt4uOXyywykNaDnwQqO1DebWba4WgovIU6rVAbzaBOBqk
 0XlAr0mSUu2P75JpJwUS1X3VCJ08r44U4kd1NeseNMrtAPSobB4BLmgN68Iku/tY
 AJ439da77hdZCxLEJQef1NU56r7q8toXQpq4thHanUrM66WT78vXwpiPqWnFSOuk
 eQkM033zsq2oc8Olc7XYr01r0KsdyhQKrC+L4H6pIsdzlGOKfxM7fsTGiyw+vu3/
 IY+aHX+B2D5Zl9jX5bvkoQrgDjkBHQROXHvpQNgnTRL9gpNNUb1gdGPJkntptJ5w
 NCF+xg0TCRzaQZ5iVtN7QFlFUpDq5PQKQGAbh32PjnItJSvdfwE/+2BAGP2rAKui
 mosOAyQ2bQNQyje3vWSg+6I72fKKizGfH1rrOIKY2tK9vdnq9MYHpG+2tFkSJ+KZ
 3yviZidGMgWT5t2XfUXq8+ZjiaCIyqV7l1NBAbuAUxWf6RUtNX1mLblYJfEvGIsg
 6CfGmYKdn1dWC4FIGBoyYPb2ZQ1hCwuxZtdLRl8ARVdoW+yfTpcZEVoBqbunUsf9
 2JFrax/SSJhT35v8drSRaDSs8YTNU1NrD3U8mUJcezMN91ZlCXs=
 =RnVa
 -----END PGP SIGNATURE-----

Merge tag 'pwrseq-fixes-for-v7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/brgl/linux

Pull power sequencing fixes from Bartosz Golaszewski:

 - fix kconfig issue in pwrseq-thread-gpu

 - fix error path logic in pwrseq_unit_enable()

 - fix two NULL-pointer dereference bugs in power sequencing core

* tag 'pwrseq-fixes-for-v7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/brgl/linux:
  power: sequencing: fix NULL-pointer dereference in pwrseq_device_register()
  power: sequencing: fix NULL-pointer dereference in pwrseq_unit_new()
  power: sequencing: don't call .post_enable() if pwrseq_unit_enable() failed
  power: sequencing: Fix build issue with COMPILE_TEST
2026-09-17 09:40:25 -07:00
Linus Torvalds
61cc777ca7 gpio fixes for v7.3-rc4
- fix fwnode reference leak on failure in shared GPIO handling
 - fix regression in OF_POPULATED logic after the unification of GPIO hog
   handling between OF, ACPI and machine variants
 - don't call free_irq() if no IRQ is installed in gpio-virtuser
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEkeUTLeW1Rh17omX8BZ0uy/82hMMFAmqrs8YACgkQBZ0uy/82
 hMPTsw/+Jm+8Z0tuCuryhBsDwMiqS4Gnr8TahCO3aB3UyMqApkiTS+hlqWGMRWbF
 Np07I4uu3ca6ohutKN/RN6K5hrjlJz6FhU1kliJ9RrK1a3bjLH4rbYoXBVsYeIpC
 mjLx9lyt6RmS4RaHPV75xPEmsAdxWbMyar6SQfiZT2t96Czsrph/VggX9kMbnXt3
 SgKTM5SyHrKmw4DgnQmZ4OcWt8p2edW+5DO+jxRmPlWUvYE/q91yemaedw5wBEoz
 ftrvbuIr+JRKKOSugjbswwBbJ0pVUkMm+hwkyAfWSPp83aF+sm2rUB46Gkr6tfQ7
 oJVqVewAV6vqW/XoAnB8vr2KO3As5HFEx8xLYZpEf9RlOIDsu8R3HOooFkvO0flP
 EsQSFccdX4WEHZoSc81iJl/TJjoM2gJtBZqqOvkHr2RZ7zdKPcnW81VHJGkSyZiW
 o70PMzcQ+FAu7I/o5Q+cqsw1eG68wHhRYZG0q38mJCRznjlEop9yBqCAx3yi+1mF
 cIhB37FlfRkFYGXByrqUFhi9V8gHWdDTChIQFJBMx1c99mBpT/eHD2hbGmac42LP
 QI4sbywoN6L3m+UJpuJwS7KoO4NtJaVN444nqnS5C8bxBGRdXV3JQHelzwhSSED9
 cq+MyMQm5hjS+i6SrADGGg+PytUF9w5Z7yRq2LFNUFwcS0TaHpk=
 =NiRo
 -----END PGP SIGNATURE-----

Merge tag 'gpio-fixes-for-v7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/brgl/linux

Pull gpio fixes from Bartosz Golaszewski:

 - fix fwnode reference leak on failure in shared GPIO handling

 - fix regression in OF_POPULATED logic after the unification of GPIO
   hog handling between OF, ACPI and machine variants

 - don't call free_irq() if no IRQ is installed in gpio-virtuser

* tag 'gpio-fixes-for-v7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/brgl/linux:
  gpio: virtuser: skip free_irq when no IRQ is installed
  gpiolib: of: don't mark hog nodes OF_POPULATED before a chip is found
  gpiolib: Put fwnode reference on failure
2026-09-17 09:08:20 -07:00
Jakub Kicinski
3b95a04eb5 Merge branch 'mptcp-misc-fixes-for-v7-3-rc4'
Matthieu Baerts says:

====================
mptcp: misc fixes for v7.3-rc4

Here are two unrelated fixes:

- Patch 1: avoid unneeded actions on subflow reset. A fix for another
  fix introduced in v6.12 and targeting a commit from v5.7.

- Patch 2: close a possible race when scheduling a closing path. A fix
  for another fix introduced in v6.0 and targeting v5.10.

- Patch 3: fix bad accounting when __subflow_push_pending returns an
  error. A fix for v6.6.
====================

Link: https://patch.msgid.link/20260917-net-mptcp-misc-fixes-7-3-rc4-v2-0-0cf5c72667c8@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17 08:14:39 -07:00
Paolo Abeni
f3ef033573 mptcp: fix bad accounting in __mptcp_subflow_push_pending()
If __subflow_push_pending() errors out we should avoid updating the
copied byte counters, to avoid mismatch push call later on.

Fixes: 0fa1b3783a ("mptcp: use get_send wrapper")
Cc: stable@vger.kernel.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Reviewed-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Link: https://patch.msgid.link/20260917-net-mptcp-misc-fixes-7-3-rc4-v2-3-0cf5c72667c8@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17 08:14:33 -07:00
Paolo Abeni
42064de57f mptcp: close race between scheduler and state change
The mptcp scheduler may race with subflow sockets state change: data
transmission on the selected socket may fail and a later release could
try to use mss_now reset to 0 for a divide operation.

Address the issue by explicitly checking for the critical scenario.

Fixes: c886d70286 ("mptcp: do not queue data on closed subflows")
Cc: stable@vger.kernel.org
Reported-by: Shardul Bankar <shardul.b@mpiricsoftware.com>
Reported-by: Xinyang Ge <xinyang@anthropic.com>
Closes: https://lore.kernel.org/20260525194828.1137119-1-shardul.b@mpiricsoftware.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Reviewed-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Link: https://patch.msgid.link/20260917-net-mptcp-misc-fixes-7-3-rc4-v2-2-0cf5c72667c8@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17 08:14:33 -07:00
Paolo Abeni
2b0f561f21 mptcp: avoid unneeded actions on subflow reset
Once in a blue moon, the mptcp receive path can recursively call
mptcp_data_ready() via state change under unlucky error conditions, and
then try to hold the data lock again.

Break the recursion loop explicitly checking for the exceptional
condition.

Add a new flag instead of using an existing one like 'closing', to exit
early in subflow_state_change(), and explicitly flush the RX queue at
reset time.

This avoids unneeded processing to check for available data -- calling
get_mapping_status() and more on a dying subflow -- but also in error
reporting and worker scheduling.

Note that we must consume the currently peeked skb before invoking
mptcp_dss_corruption to avoid consuming it again after the eventual
reset has freed it.

Fixes: e32d262c89 ("mptcp: handle consistently DSS corruption")
Cc: stable@vger.kernel.org
Reported-by: Xinyang Ge <xinyang@anthropic.com>
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Reviewed-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Link: https://patch.msgid.link/20260917-net-mptcp-misc-fixes-7-3-rc4-v2-1-0cf5c72667c8@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-17 08:14:33 -07:00
Linus Torvalds
4aec9ad1c6 dma-mapping fixes for Linux 7.3
A few fixes for the DMA-mapping code:
 - resolved regression in accessing encrypted memory by IOMMU-backed
 devices (Aneesh Kumar K.V),
 - improved failure handling and removed rare bug in swiotlb/highmem
 (Donggeun Yoo).
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQSrngzkoBtlA8uaaJ+Jp1EFxbsSRAUCaquwMwAKCRCJp1EFxbsS
 RMT6AP0elpdaZXNY0KwUBTwU95H604J+donqriepHABIBhIDEQD9GWZqNf/m1gEI
 tR5lHQ3+NGs0Q7Vd2ed1vSe82HQSsgU=
 =QxUC
 -----END PGP SIGNATURE-----

Merge tag 'dma-mapping-7.3-2026-09-17' of git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux

Pull dma-mapping fixes from Marek Szyprowski:
 "A few fixes for the DMA-mapping code:

   - resolved regression in accessing encrypted memory by IOMMU-backed
     devices (Aneesh Kumar K.V)

   - improved failure handling and removed rare bug in swiotlb/highmem
     (Donggeun Yoo)"

* tag 'dma-mapping-7.3-2026-09-17' of git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux:
  x86/mm: Don't force unencrypted DMA for IOMMU-backed devices
  dma-mapping: don't trace the DMA address when the allocation fails
  swiotlb: use the adjusted address for the highmem page lookup
  dma-coherent: report a failed reserved memory assignment
2026-09-17 08:03:37 -07:00
Eric Dumazet
a5117e1ecc net: skbuff: do not leave stale header offsets after pskb_carve()
pskb_carve_inside_header() and pskb_carve_inside_nonlinear() remove
the first bytes of a packet and reallocate skb->head.

All the headers that were present before the operation are gone,
but both functions call skb_headers_offset_update(skb, 0), which
is a no-op : skb->mac_header, skb->network_header,
skb->transport_header and skb->csum_start keep their old values and
now describe bytes which are no longer there.

Both helpers size the new head from the old skb_end_offset(), so the
stale offsets still land inside the new allocation. They point past
skb_tail_pointer() though, to bytes that were never initialized.

pskb_carve_inside_nonlinear() is the worst case, because it leaves a
zombie skb with an empty linear part (skb->data ==
skb_tail_pointer(skb), skb_headlen(skb) == 0), while
skb_mac_header_was_set() is still true and skb->mac_header is way
ahead of skb->data.

The only user of pskb_extract() is rds_tcp_data_recv(), and the
carved skb is queued on tinc->ti_skb_list. When the RDS incoming
message is released, rds_tcp_inc_free() calls skb_queue_purge(),
which frees the skbs with SKB_DROP_REASON_QUEUE_PURGE. This is
visible from drop_monitor, which then tries to pull back to the
(bogus) mac header :

skbuff: __skb_pull(len=234)
skb len=6968 data_len=6968 headroom=0 headlen=0 tailroom=0
end-tail=384 mac=(234,14) mac_len=14 net=(248,40) trans=288
shinfo(txflags=0 nr_frags=1 gso(size=1428 type=16 segs=5))
csum(0x100120 start=288 offset=16 ip_summed=3 complete_sw=0 valid=1 level=0)
hash(0x7b446c6c sw=0 l4=1) proto=0x86dd pkttype=0 iif=60
kernel BUG at ./include/linux/skbuff.h:2847!

Add skb_carve_reset_headers() to mark the mac and transport headers
as not set, reset the network header, clear skb->mac_len, and drop
a now meaningless CHECKSUM_PARTIAL (csum_start no longer describes
anything).

Invalidate the inner offsets as well. Unlike mac_header and
transport_header they have no "unset" sentinel, so a leftover
non-zero value still looks like a real header. Zero
skb->inner_mac_header, skb->inner_network_header,
skb->inner_transport_header, skb->inner_protocol and
skb->encapsulation, so that all the header state is invalidated in
one place.

v2: fixed an inaccurate changelog. The stale offsets stay inside the
    new skb->head, which is never smaller than the old one, they
    simply point past skb_tail_pointer() to bytes that are gone.
    Thanks to Xuanqiang Luo for insisting on this.
    Also invalidate the inner header state, as suggested by the
    netdev AI review :
    https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260911114922.621937-1-edumazet%40google.com

Fixes: 6fa01ccd88 ("skbuff: Add pskb_extract() helper function")
Reported-by: syzbot+586af68eb819833c2d91@syzkaller.appspotmail.com
Closes: https://lore.kernel.org/netdev/6aa3e9d3.f2639fcc.29487d.0028.GAE@google.com/
Cc: Xuanqiang Luo <xuanqiang.luo@linux.dev>
Cc: Allison Henderson <achender@kernel.org>
Cc: rds-devel@oss.oracle.com
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Link: https://patch.msgid.link/20260915130423.3956471-1-edumazet@google.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-17 16:05:02 +02:00
Paolo Abeni
ad9c65b8f9 Merge branch 'tcp-exclude-old-acks-from-fast-path'
Inbal Schussheim says:

====================
tcp: exclude old ACKs from fast path

Exclude ACKs outside [SND.UNA, SND.NXT] from TCP header prediction so
that they fall through to the slow path, where ACK
validation is applied.

Add a packetdrill test for a data segment carrying an
excessively old ACK. The test fails on the unpatched kernel and passes
with the fix.

v2: https://lore.kernel.org/netdev/20260909075644.1408171-1-inbal.lipshtat@mail.huji.ac.il/
v1: https://lore.kernel.org/netdev/20260906123151.1391349-1-inbal.lipshtat@mail.huji.ac.il/T/#u
====================

Link: https://patch.msgid.link/20260914090408.1435080-1-inbal.lipshtat@mail.huji.ac.il
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-17 15:18:49 +02:00
Inbal Schussheim
d841cd7513 selftests: net: packetdrill: test exclusion of old ACK from TCP fast path
Add a packetdrill test for an in-sequence data segment carrying an
excessively old ACK.

Verify that the segment falls through from the TCP fast path to the slow
path, where the existing ACK validation rejects it and sends a challenge
ACK. The payload is not accepted and RCV.NXT remains unchanged.

Based on the reproducer from Commit 3d501dd326
("tcp: do not accept ACK of bytes we never sent").

Signed-off-by: Inbal Schussheim <inbal.lipshtat@mail.huji.ac.il>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260914090408.1435080-3-inbal.lipshtat@mail.huji.ac.il
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-17 15:17:58 +02:00
Inbal Schussheim
f81e6c3fb0 tcp: exclude old ACKs from tcp fast path
Exclude old ACKs before SND.UNA from the tcp fast path
as well as ACKs after SND.NXT.

Such ACKs will fall through to the slow path, where tcp_ack()
performs the appropriate validation and challenge ACK handling
according to RFC5961 and Commit 3d501dd326 ("tcp: do not
accept ACK of bytes we never sent").

This prevents old ACKs from being accepted
or modifying connection state as part of the fast path before
appropriate ACK validation is applied.
In particular, this prevents payload carried by a segment with
an excessively old ACK from advancing RCV.NXT before the ACK
is rejected.

Fixes: 31770e34e4 ("tcp: Revert "tcp: remove header prediction"")
Reported-by: Amit Klein <amit.klein@mail.huji.ac.il>
Reported-by: Tamir Shahar <tamir.shahar1@mail.huji.ac.il>
Reported-by: Inbal Schussheim <inbal.lipshtat@mail.huji.ac.il>
Suggested-by: Eric Dumazet <edumazet@google.com>
Cc: stable@vger.kernel.org
Signed-off-by: Inbal Schussheim <inbal.lipshtat@mail.huji.ac.il>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260914090408.1435080-2-inbal.lipshtat@mail.huji.ac.il
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-17 15:17:58 +02:00
Jakub Kicinski
d798162eb3 dpll: reject a reference sync pin which is not on the pin's dpll
dpll_pin_ref_sync_state_set() resolves the partner's driver private data
with dpll_pin_on_dpll_priv() and passes the result to ref_sync_get() and
ref_sync_set() without looking at it. The helper returns NULL when the
partner holds no ref on that dpll. Of the two drivers implementing the
feature only zl3073x dereferences the pointer (sync_pin->id); ice ignores
it, so ice cannot fault here.

The NULL is a teardown race, not a steady state - zl3073x registers every
input pin with every channel, so the partner is normally present on the
dpll the base pin resolves to. zl3073x_dev_stop() unregisters pins one at
a time, taking and dropping dpll_lock for each, and between the partner's
turn and the base pin's the partner is out of that dpll's pin_refs while
still registered with the channels not yet torn down, so
dpll_pin_available() keeps passing. That path is not only driver removal:
devlink reload and devlink dev flash both run zl3073x_dev_stop().

Reproduced by holding that state open with a mock dpll device, which is
where the frame name comes from:

 BUG: kernel NULL pointer dereference, address: 0000000000000000
 Oops: Oops: 0000 [#1] SMP NOPTI
 RIP: 0010:mock_ref_sync_get+0x5/0x30
 Call Trace:
  <TASK>
  dpll_pin_ref_sync_set+0x19f/0x4a0
  dpll_nl_pin_set_doit+0x17d/0x840
  genl_family_rcv_msg_doit+0xd6/0x130
  genl_rcv_msg+0x181/0x2b0
  netlink_rcv_skb+0x55/0x100
  genl_rcv+0x23/0x30
  netlink_unicast+0x24d/0x370
  netlink_sendmsg+0x1e2/0x420
  __sys_sendto+0x1db/0x1f0
  __x64_sys_sendto+0x1f/0x30
  do_syscall_64+0xe1/0x490

Commit d2e914a4a0 ("dpll: fix NULL pointer dereference in
dpll_msg_add_pin_ref_sync()") added the same guard to the read side, which
the kernel walks into by itself because the delete notification is emitted
from inside the unregister; the write side needs a pin-set to land in the
window and was left alone. Test the priv rather than look up pin_refs
directly, so that the two halves key off the same condition.

Fixes: 58256a26bf ("dpll: add reference sync get/set")
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Reviewed-by: Ivan Vecera <ivecera@redhat.com>
Link: https://patch.msgid.link/20260915213047.1352286-1-kuba@kernel.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-17 14:51:23 +02:00
Dmitriy Okunev
14cb1e7702 net: mvpp2: prevent buffer overflow in page_pool allocation
The per‑processor buffering scheme is supported only if the
number of pools (nrxqs * 2) does not exceed MVPP2_BM_MAX_POOLS (8).
This is already checked in mvpp2_probe() during the initial
activation of percpu_pools.

However, mvpp2_change_mtu() may later call
mvpp2_bm_switch_buffers(priv, true) without this check, which can
lead to an out-of-bounds access in the priv->page_pool array in
mvpp2_bm_init(). The array is sized to hold MVPP2_PORT_MAX_RXQ
entries, and mvpp2_get_nrxqs() may return exactly that value. The
per-CPU scheme then doubles it to nrxqs * 2, exceeding the array
bounds.

Check that the hardware version is MVPP22 or newer and that the
number of pools (nrxqs * 2) does not exceed MVPP2_BM_MAX_POOLS
before switching to per-CPU mode.

Found by Linux Verification Center (linuxtesting.org) with SVACE.

Fixes: 7d04b0b13b ("mvpp2: percpu buffers")
Signed-off-by: Dmitriy Okunev <dokunevdmitriy@gmail.com>
Link: https://patch.msgid.link/20260914091557.71769-1-dokunevdmitriy@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-17 14:43:29 +02:00
James Clark
9ca4ba2425 net: macb: fix ordering around PTP timestamp read
PTP_SYS_OFFSET_EXTENDED returns system timestamps that do not correctly
bracket the PHC register read on MACB/GEM. On a Raspberry Pi 5, the
returned interval can be as short as 37 ns, while an ordered register
read takes approximately 1 us. This biases the midpoint used by phc2sys,
causing CLOCK_REALTIME to run approximately 0.5 us ahead when synchronized
to the PHC.

gem_tsu_get_time() reads the nanoseconds register using the driver's
relaxed MMIO accessor. On weakly ordered systems, the subsequent system
timestamp can be taken before the register read completes. The internal
smp_rmb() in the pre-timestamp path also does not guarantee ordering
against the subsequent MMIO read.

Add rmb() before and after the bracketed nanoseconds read in both the
normal and seconds rollover paths so the system timestamps bracket the
PHC read. Adding the post-read barrier increases the minimum interval on
the same Raspberry Pi 5 to approximately 1 us.

Fixes: e51bb5c278 ("net: macb: ptp: Switch to gettimex64() interface")
Tested-by: Nicolai Buchwitz <nb@tipi-net.de> # Raspberry Pi CM5, min bracket 37 ns -> 981 ns
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Reviewed-by: Théo Lebrun <theo.lebrun@bootlin.com>
Assisted-by: LLM
Signed-off-by: James Clark <jjc@jclark.com>
Link: https://patch.msgid.link/20260915045823.76100-1-jjc@jclark.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-17 14:25:52 +02:00
Takashi Iwai
546b928da0 ASoC: Fixes for v7.3
A relatively large pile of fixes here, a lot of driver specific stuff
 that's broadly unremarkable plus a few core fixes from Richard that fix
 issues where SoundWire systems with multiple CODECs on the same link
 would configure the CODECs to use the same bus slots leading to broken
 audio.
 -----BEGIN PGP SIGNATURE-----
 
 iQEzBAABCgAdFiEEreZoqmdXGLWf4p/qJNaLcl1Uh9AFAmqrCNoACgkQJNaLcl1U
 h9DHWgf8D3dIL06bqj6IoyMLCFNrcQ8BYbWUeWNu5YE0vP29ybdpYidTxJFjqF2t
 TUB8fTO2u3LvfKIIgOSVyXN84i7/4EwtDjBz1iVzGhm0/2ZfEOitO2LtUvhCHiZi
 +JnEOXdwa7wM9jv0On6B81r8+vXj7FaNmq/TnLbUU3R/DeaRx571k913lazZSRb0
 cfPj1FGMUvpfBZ7DC011yEufDD4C8qVaktV6IpRqeBAxks1vmXQX7Lt78gJCxvQt
 w8Uu2T8FpiTuYvObI7KW7IZr01IQtPpJ4N9ekUMuBfOqkjvG+XOETM2aEWg7q9Jk
 Qu5GfyB2LIav4/On4pCxwJGaJiUavQ==
 =IMNO
 -----END PGP SIGNATURE-----

Merge tag 'asoc-fix-v7.3-rc3' of https://git.kernel.org/pub/scm/linux/kernel/git/broonie/sound into for-linus

ASoC: Fixes for v7.3

A relatively large pile of fixes here, a lot of driver specific stuff
that's broadly unremarkable plus a few core fixes from Richard that fix
issues where SoundWire systems with multiple CODECs on the same link
would configure the CODECs to use the same bus slots leading to broken
audio.
2026-09-17 08:15:32 +02:00
Jakub Kicinski
c9151088f1 Merge branch 'net-psp-avoid-conflicts-with-skb-decrypted-and-sk_validate_xmit_skb'
Daniel Zahka says:

====================
net: psp: avoid conflicts with skb->decrypted and sk_validate_xmit_skb()

Sashiko's review of commit da630d1da2b1 ("netdevsim: psp: drop tx key
ops") [1] showed that there is a hazard between PSP and offloaded TLS,
where both can clobber what the other set in the sk_validate_xmit_skb
callback.

It was discussed further on the mailing list [2], and it was pointed out
that there are conflicts with PSP and TLS ULP both using the
skb->decrypted bit.

The simplest fix is to make psp and tls mutually exclusive. This series
goes a bit further and makes psp exclusive with all TCP ULPs. The PSP
implementation that we have is not designed to be used with any TCP ULP,
so don't allow a socket to have state for both.

I will send a subsequent series to net-next which will remove the
ability to perform the rx-assoc and tx-assoc psp netlink calls on
sockets that are not in the TCP_ESTABLISHED state. This will close the
remaining quirk that a sk_clone() on a listen socket with psp tx-assoc
state will leave a stale sk->sk_validate_xmit_skb call back on a new,
non-psp socket. I do not believe that change needs to be regarded as a
fix, because it only stands to add unecessary validation code in the tx
path.

[1]: https://sashiko.dev/#/patchset/20260903-psp-prep-v1-0-d47e9c4c375d%40gmail.com
[2]: https://lore.kernel.org/netdev/20260903-psp-prep-v1-0-d47e9c4c375d@gmail.com/
====================

Link: https://patch.msgid.link/20260915-psp-ktls-fix-v2-0-0eedc3b148ec@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 19:18:26 -07:00