Commit Graph

1481711 Commits

Author SHA1 Message Date
Jamal Hadi Salim
4864f58c53 net/sched: fq_pie: clamp quantum in change path
fq_pie_change() accepts any quantum value from userspace, including 1.
With a crafted size table qdisc_pkt_len reaches ~2 GiB, so quantum=1
makes the deficit-refill loop spin ~2^31 times under the qdisc lock
(a soft lockup / denial of service).

Add max(256U, ...) matching fq_codel_change().

Conditions to recreate the bug:
  CONFIG_NET_SCH_FQ_PIE=y. Requires CAP_NET_ADMIN (namespace-local via
  unshare -Urn suffices).

  tc qdisc add dev dummy0 root fq_pie
  tc qdisc change dev dummy0 root fq_pie quantum 1 stab data 32768 size_log 15 cell_log 0

Fixes: ec97ecf1eb ("net: sched: add Flow Queue PIE packet scheduler")
Reported-by: Vega <vega@nebusec.ai>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-0CFC.v3.20260901204856@mojatatu.com.3
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:47:29 -07:00
Jamal Hadi Salim
094cc07f98 net/sched: fq: clamp quantum and initial_quantum in change path
The fq change path accepts TCA_FQ_QUANTUM in [1, INT_MAX] and
TCA_FQ_INITIAL_QUANTUM up to INT_MAX, while fq_init() already clamps to
[1, 1<<20]. A user can override the init clamp via tc qdisc change,
restoring the small-quantum deficit spin that the init clamp prevents.

Narrow iq_range.max to 1<<20 so TCA_FQ_INITIAL_QUANTUM is rejected at
parse time. Clamp TCA_FQ_QUANTUM to [256, 1<<20] in fq_change() and
fq_init() quantum to [256, 1<<20] for tiny-MTU devices.

Conditions to recreate the bug:
  CONFIG_NET_SCH_FQ=y. Requires CAP_NET_ADMIN (namespace-local via
  unshare -Urn suffices).

  tc qdisc add dev dummy0 root fq
  tc qdisc change dev dummy0 root fq quantum 1 stab data 32768 size_log 15 cell_log 0

Fixes: 709f34f7c2 ("net/sched: fq: add overflow bounds to quantum and initial quantum")
Reported-by: Vega <vega@nebusec.ai>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-0CFC.v3.20260901204856@mojatatu.com.2
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:47:29 -07:00
Carolina Jubran
df99553f84 net/mlx5e: Keep HW timestamp stats monotonic across reconfiguration
`mlx5e_stats_ts_get()` currently selects either DMA or port timestamp
counters based on `tx_ptp_opened`. This flag is intentionally kept set
once the PTP TX queues have been opened so their statistics remain
available after queue teardown. As a result, DMA timestamps are no
longer reported after switching from port timestamping back to DMA
timestamping.

The function also reads statistics only from the currently active
channels and TCs. Reducing the number of channels or TCs can therefore
drop previously accumulated timestamp counters from the reported value.

Read the persistent channel statistics instead and always include DMA
timestamp counters. Once the PTP TX queues have been opened, also
include the port timestamp counters.

This also drops state_lock. It previously protected live channel/PTP
pointers, the new code only reads persistent channel_stats and
ptp_stats via mlx5e_stats_nch_read(), which is already safe for
lockless stats access.

Fixes: 3579032c08 ("net/mlx5e: Implement ethtool hardware timestamping statistics")
Signed-off-by: Carolina Jubran <cjubran@nvidia.com>
Reviewed-by: Shahar Shitrit <shshitrit@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902193731.3668958-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:23:37 -07:00
Lama Kayal
c0c6f4ba8a net/mlx5: E-Switch, prevent mc_list repopulation during vport disable
In mlx5_esw_vport_disable(), move esw_apply_vport_rx_mode() ahead
of esw_vport_change_handle_locked() so vport->allmulti_rule is
NULL before the change handler observes it.

During FW-fatal recovery the disable runs while dev->state ==
INTERNAL_ERROR. The promisc query inside esw_update_vport_rx_mode()
fails and returns early, leaving vport->allmulti_rule intact, so
esw_update_vport_mc_promisc() runs and adds MLX5_ACTION_ADD entries
to vport->mc_list whose flow rules are then installed in the FDB
by esw_add_mc_addr(). esw_destroy_legacy_table() tears down the
FDB with those refs still held, corrupting the sub-tree and
leaving dangling flow_rule pointers in vport->mc_list.

Two-stage failure on `echo 1 > /sys/bus/pci/devices/<bdf>/reset`:

  refcount_t: underflow; use-after-free.
   tree_put_node+0xef/0x110 [mlx5_core]
   clean_tree+0x44/0xd0 [mlx5_core] (x5)
   mlx5_fs_core_cleanup+0x57/0x1c0 [mlx5_core]
   mlx5_unload+0x65/0xd0 [mlx5_core]
   ... mlx5_health_try_recover

  BUG: unable to handle page fault for address: 0000000003000055
   down_write+0x1c/0x60
   mlx5_del_flow_rules+0x33/0x1f0 [mlx5_core]
   esw_del_mc_addr+0x7b/0x170 [mlx5_core]
   esw_apply_vport_addr_list+0x56/0xf0 [mlx5_core]
   esw_vport_change_handle_locked+0x28b/0x310 [mlx5_core]
   mlx5_esw_vport_enable+0x270/0x4a0 [mlx5_core]
   ... mlx5_load ... mlx5_health_try_recover

esw_apply_vport_rx_mode(false, false) clears vport->allmulti_rule
via its local state machine even when the FW del fails. With the
rule NULL the !IS_ERR_OR_NULL(allmulti_rule) gate in the change
handler closes, no rules are installed during disable, and the
reload starts with a clean mc_list.

Fixes: 922f56e9a7 ("net/mlx5: Fix steering rules cleanup")
Signed-off-by: Lama Kayal <lkayal@nvidia.com>
Reviewed-by: Cosmin Ratiu <cratiu@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902193854.3669035-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:21:43 -07:00
Yael Chemla
7ee07f601f net/mlx5: E-Switch: fix use-after-free in mlx5_eswitch_termtbl_put
In mlx5_eswitch_termtbl_put(), the zero-ref cleanup check reads
tt->ref_count after termtbl_mutex has been released.  Two concurrent
callers on the same mlx5_termtbl_handle race: one decrements ref_count
to zero, removes the hash entry, and calls kfree(tt) while the other
has already dropped the mutex and is about to evaluate
if (!tt->ref_count), producing a use-after-free.

Fix this by capturing the result of the decrement into a stack-local
last variable before dropping the mutex.  The cleanup decision is now
made entirely under termtbl_mutex, and tt is not touched after
kfree.

Fixes: 10caabdaad ("net/mlx5e: Use termination table for VLAN push actions")
Signed-off-by: Yael Chemla <ychemla@nvidia.com>
Reviewed-by: Dragos Tatulea <dtatulea@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902193514.3668880-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:20:26 -07:00
Carolina Jubran
af3aef0245 net/mlx5e: Fix use-after-free race in sample_restore_put()
Concurrent teardown of TC sample rules sharing the same restore
context may re-read restore->count after dropping restore_lock.
At that point another thread may already have completed cleanup and
freed the restore object.

Use the result of the refcount decrement while holding restore_lock to
determine whether cleanup is needed.

Fixes: 36a3196256 ("net/mlx5e: TC, Add sampler restore handle API")
Signed-off-by: Carolina Jubran <cjubran@nvidia.com>
Reviewed-by: Shahar Shitrit <shshitrit@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902193341.3668809-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:19:22 -07:00
Carolina Jubran
e7ee897408 net/mlx5e: Fix ETS zero BW reporting when one TC holds 100%
When ETS TCs with zero bandwidth are configured, the driver programs the
firmware using an alternate representation. On get, it needs
to recognize that representation so those TCs can be translated back and
reported as 0% bandwidth.

The existing detection relied on the programmed bandwidth because it was
enough to identify this representation. However, when a single ETS TC
owns 100% of the bandwidth, its firmware representation becomes the
same as a strict-priority TC, causing zero-bandwidth ETS TCs to be
reported with non-zero bandwidth values.

Use the cached TSA instead to distinguish the ETS and strict-priority
cases.

Fixes: be0f161ef1 ("net/mlx5e: DCBNL, Implement tc with ets type and zero bandwidth")
Signed-off-by: Carolina Jubran <cjubran@nvidia.com>
Reviewed-by: Alex Lazar <alazar@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902193224.3668743-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:18:47 -07:00
Akiva Goldberger
b3c79dee50 net/mlx5: LAG, use local tracker to update active ports
The CREATE_LAG command is handled asynchronously by queuing a work,
which stores a local copy of ldev->tracker. When the work is processed,
it is possible that the values of the local copy and ldev->tracker have
diverged.

A single CREATE_LAG command programs two related fields into the
firmware: the v2p (virtual-to-physical) map, which selects the physical
egress port for each hash bucket, and the active_port bitmask, which
tells the firmware which physical ports are currently up so it can
redirect QP/TIS away from inactive ports. For the firmware to steer
traffic correctly, both must be derived from the same view of the ports'
link state.

The v2p map is computed by mlx5_infer_tx_affinity_mapping() from the
local tracker snapshot, but lag_active_port_bits() called
mlx5_infer_tx_enabled() on the live ldev->tracker instead. If
ldev->tracker changed between the snapshot and command execution, the
two fields reflect different port states: the v2p map may steer a bucket
to a port that the active_port mask marks as inactive (or vice versa).
The firmware then receives a self-contradictory configuration and can
redirect or drop traffic on a port the mapping still points at, until a
later event happens to reconcile the state.

Update lag_active_port_bits so that it receives the local version of the
tracker from when the work was queued, effectively closing the window
for injecting an inconsistency.

Fixes: c5c13b456c ("net/mlx5: Lag, set active ports if support bypass port select flow table")
Signed-off-by: Akiva Goldberger <agoldberger@nvidia.com>
Reviewed-by: Shay Drori <shayd@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902192740.3665435-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:16:01 -07:00
Jakub Kicinski
502381cf73 Merge branch 'net-mlx5e-rs-fec-variant-fixes'
Tariq Toukan says:

====================
net/mlx5e: RS FEC variant fixes

This series by Shahar fixes three related bugs in the RS FEC handling
for mlx5e, all stemming from incomplete coverage of the three RS FEC
hardware variants: RS_528_514 (bit 2), RS_544_514_INTERLEAVED_QUAD
(bit 4), and RS_544_514 (bit 7).
====================

Link: https://patch.msgid.link/20260902164634.3657606-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:14:51 -07:00
Shahar Shitrit
c84ce45a7a net/mlx5e: Fix reporting support for all RS FEC variants
get_fec_supported_advertised() populates the FEC modes reported as
supported to userspace. The MLX5E_ADVERTISE_SUPPORTED_FEC macro only
checked MLX5E_FEC_RS_528_514, causing devices that support only the
other RS variants (RS_544_514_INTERLEAVED_QUAD or RS_544_514) to not
advertise RS as supported to ethtool at all.

Introduce MLX5E_FEC_RS_MASK covering all three RS bit positions,
update the macro to accept a bitmask directly rather than a single
enum value, and pass MLX5E_FEC_RS_MASK for the RS entry.

Fixes: b5ede32d33 ("net/mlx5e: Add support for FEC modes based on 50G per lane links")
Fixes: 4e343c11ef ("net/mlx5e: Support FEC settings for 200G per lane link modes")
Signed-off-by: Shahar Shitrit <shshitrit@nvidia.com>
Reviewed-by: Dragos Tatulea <dtatulea@nvidia.com>
Reviewed-by: Yael Chemla <ychemla@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902164634.3657606-4-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:14:48 -07:00
Shahar Shitrit
b9d755c5a3 net/mlx5e: Fix setting RS FEC after remapping
When a user sets a FEC mode via ethtool, the driver maps the ethtool
FEC type to the lowest mlx5 bit of that type. For RS FEC, this is
MLX5E_FEC_RS_528_514 (bit 2). The driver then checks whether this
bit is supported by at least one link mode by inspecting the
fec_override_cap fields via mlx5e_fec_in_caps(), and returns
-EOPNOTSUPP if not.

This check is incorrect. RS FEC has three supported hardware variants:
RS_528_514 (bit 2), RS_544_514_INTERLEAVED_QUAD (bit 4), and
RS_544_514 (bit 7). mlx5e_remap_fec_conf_mode() already remaps bit 2
to the appropriate RS variant per link mode when writing the admin
fields, but the early capability check is done against the raw
unmapped bit. As a result, a device that supports RS_544_514 or
RS_544_514_INTERLEAVED_QUAD but not RS_528_514 will incorrectly reject
the user's RS FEC request.

Remove the early support check from mlx5e_set_fec_mode() and fold
it into the existing write loop, checking caps against the remapped
policy per link mode. Return -EOPNOTSUPP before the final register
write if no link mode accepted the policy.

Fixes: 2608a2f831 ("net/mlx5e: Fix return status when setting unsupported FEC mode")
Signed-off-by: Shahar Shitrit <shshitrit@nvidia.com>
Reviewed-by: Dragos Tatulea <dtatulea@nvidia.com>
Reviewed-by: Yael Chemla <ychemla@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902164634.3657606-3-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:14:48 -07:00
Shahar Shitrit
802eedcc0b net/mlx5e: Fix missing FEC mode mapping for RS_544_514_INTERLEAVED_QUAD
MLX5E_FEC_RS_544_514_INTERLEAVED_QUAD is missing from
pplm_fec_2_ethtool_linkmodes[], leaving index 4 zero-initialized.
As a result, when this FEC mode is active, find_first_bit() returns
index 4, causing __set_bit() to set bit 0
(ETHTOOL_LINK_MODE_10baseT_Half_BIT) instead of
ETHTOOL_LINK_MODE_FEC_RS_BIT. Consequently, ethtool reports:

  Advertised FEC modes: Not reported

Add the missing mapping to ETHTOOL_LINK_MODE_FEC_RS_BIT.

Fixes: 4e343c11ef ("net/mlx5e: Support FEC settings for 200G per lane link modes")
Signed-off-by: Shahar Shitrit <shshitrit@nvidia.com>
Reviewed-by: Dragos Tatulea <dtatulea@nvidia.com>
Reviewed-by: Yael Chemla <ychemla@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902164634.3657606-2-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:14:48 -07:00
Weiming Shi
e6662f2100 net/sched: defer qdisc freeing after failed creation
An RTM_NEWQDISC request can make clsact bind a populated shared ingress
block during ->init(), publishing an embedded mini_Qdisc to lockless
readers.  If the same request has an invalid TCA_RATE, estimator setup
fails after ->init(); the unwind removes the pointer but synchronously
frees its containing qdisc while tc_run() may still hold it.

Retire failed qdiscs through the same RCU helper as normal destruction.
Inline the synchronous free into the callback now that no direct callers
remain.

Fixes: 51ab2994c3 ("net: sched: allow ingress and clsact qdiscs to share filter blocks")
Reported-by: Xiang Mei <xmei5@asu.edu>
Link: https://lore.kernel.org/netdev/20260805102505.740806-1-david.lee@trailofbits.com/
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Link: https://patch.msgid.link/20260902155231.2149915-2-bestswngs@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 12:40:05 -07:00
XingWang Xiang
2b4707a149 net: mctp: i3c: serialize probe with bus removal
mctp_i3c_probe() drops busdevs_lock after finding the matching bus. A
concurrent I3C_NOTIFY_BUS_REMOVE can then unregister and free the bus
netdev before probe passes its private data to mctp_i3c_add_device().
The latter consequently adds a list node through a freed mbus pointer.

Keep busdevs_lock held until the device has been added. This also
satisfies the __must_hold annotation on mctp_i3c_add_device().

Fixes: c8755b29b5 ("mctp i3c: MCTP I3C driver")
Signed-off-by: XingWang Xiang <v3rdant.xiang@gmail.com>
Acked-by: Matt Johnston <matt@codeconstruct.com.au>
Signed-off-by: David S. Miller <davem@davemloft.net>
2026-09-05 09:39:00 +01:00
Sahil Chandna
80dd7e754b net: mana: Reserve extra CQ slot for the fence completion CQE
The RX completion queue is sized to hold exactly one CQE per posted RX WQE.
MANA_FENCE_RQ makes hardware post an additional CQE_RX_OBJECT_FENCE after
the packet CQEs. The current sizing reserves no extra slot for it and in
rare cases, CQ has no guaranteed slot for the fence CQE when it is full of
packet CQEs. This can lead to dropping the fence completion while the
driver waits holding RTNL lock throughout the timeout duration.
Reserve one extra CQE slot for CQE_RX_OBJECT_FENCE. mana_gd_alloc_memory()
requires queue_size to be a power-of-two and at least MANA_PAGE_SIZE;
the reservation pushes cq_size past a power-of-two, so round up the CQ size
in mana_create_rxq().

Cc: stable@vger.kernel.org
Fixes: 6cc74443a7 ("net: mana: Add RX fencing")
Signed-off-by: Sahil Chandna <sahilchandna@linux.microsoft.com>
Reviewed-by: Haiyang Zhang <haiyangz@microsoft.com>
Link: https://patch.msgid.link/20260901121837.3503240-1-sahilchandna@linux.microsoft.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 18:36:38 -07:00
Jakub Kicinski
db88216424 Merge branch 'pds_core-fixes-for-the-pci-reset-path'
Nikhil P. Rao says:

====================
pds_core: fixes for the PCI reset path [part]

Patch 1 is the v2 patch with the pdsc_core_init() and
pdsc_identify_ver() checks removed. Those checks are dead code. commit
cd09971dcc ("pds_core: keep the health thread stopped during reset")
disables health_work across the reset, so the health thread can no
longer reach pdsc_setup() with the BARs unmapped. The only other callers
are probe and pdsc_reset_done(), and both map the BARs earlier in the
same call, so cmd_regs cannot be NULL by the time they get there.

The pdsc_core_init() check is also worse than what it replaces. Its
bail-out jumps to err_out_uninit, which ends up in pdsc_intr_free() and
writes to pdsc->intr_ctrl, also NULL at that point.

On v2 I said I would convert pdsc_identify() and pdsc_core_init() to
pdsc_devcmd_with_data() once the PLDM series landed. Dropping that: the
helper has no read-back path and both callers need one, and giving them
an -ENXIO return means hardening pdsc_intr_free() against a NULL
intr_ctrl on the err_out_uninit path. That is a lot of churn to
deduplicate two call sites.

Patch 2 is the VF pci_release_regions() fix, older than the cmd_regs
race, so it carries its own Fixes tag.

The v2 changelog claim that pdsc_unmap_bars() clears db_pages was wrong.
It clears info_regs, cmd_regs, intr_status and intr_ctrl; db_pages is
never mapped.

v2: https://lore.kernel.org/20260804235946.177762-1-nikhil.rao@amd.com
v1: https://lore.kernel.org/20260729055258.1416225-1-nikhil.rao@amd.com
====================

Link: https://patch.msgid.link/20260901044219.1361466-1-nikhil.rao@amd.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 18:07:45 -07:00
Nikhil P. Rao
73608de7e5 pds_core: don't release PCI regions for VFs on reset
pdsc_reset_prepare() called pci_release_regions() unconditionally, but
only PFs call pci_request_regions() (pdsc_init_pf). On a VF FLR this
makes the kernel warn "Trying to free nonexistent resource".

Fixes: ffa5585833 ("pds_core: implement pci reset handlers")
Reported-by: sashiko-bot <sashiko-bot@kernel.org>
Closes: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260804235946.177762-1-nikhil.rao%40amd.com
Signed-off-by: Nikhil P. Rao <nikhil.rao@amd.com>
Link: https://patch.msgid.link/20260901044219.1361466-3-nikhil.rao@amd.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 18:07:36 -07:00
Nikhil P. Rao
7980325b2f pds_core: fix cmd_regs access racing BAR unmap on reset
pdsc_reset_prepare() and pdsc_reset_done()'s pdsc_map_bars() error path
clear/iounmap cmd_regs without devcmd_lock, and
pdsc_legacy_firmware_update()'s download loop derefs cmd_regs after
dropping and retaking the lock without re-checking. An FLR concurrent
with a devlink flash can unmap cmd_regs under an in-flight devcmd,
causing a NULL deref or a write to unmapped MMIO.

Take devcmd_lock across the BAR unmap/remap, and re-check cmd_regs in
the download loop. Only the PF maps cmd_regs and runs devcmd, so skip
the unmap on a VF, as pdsc_remove() and pdsc_reset_done() already do.

A reset that completes entirely within the unlocked window is not a
correctness problem for the image: the device clears its update session,
so a resumed download is rejected, and it verifies the staged image
before writing a flash slot, reporting PDS_RC_BAD_FW rather than
activating it.

pdsc_unmap_bars() also clears info_regs, intr_status and intr_ctrl. The
interrupt and start/stop readers of those are quiesced before the unmap
by pdsc_fw_down(), which frees the interrupts and tears down the queues.
The debugfs readers are not, since those files outlive a reset; that is
pre-existing and out of scope here.

Fixes: e96094c1d1 ("pds_core: Clear BARs on reset")
Reported-by: sashiko-bot <sashiko-bot@kernel.org>
Closes: https://sashiko.dev/#/patchset/20260708212222.296202-1-nikhil.rao%40amd.com?part=3
Signed-off-by: Nikhil P. Rao <nikhil.rao@amd.com>
Link: https://patch.msgid.link/20260901044219.1361466-2-nikhil.rao@amd.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 18:07:36 -07:00
Jakub Kicinski
6262acad9d Merge branch 'net-cap-tx_queue_len-at-s16_max-to-prevent-oversized-ring-allocations'
Jamal Hadi Salim says:

====================
net: cap tx_queue_len at S16_MAX to prevent oversized ring allocations

An unprivileged user (via unshare -Urn) can set a huge tx_queue_len
and exhaust global memory through ring allocations sized from it
(pfifo_fast skb_arrays, tun/tap ptr_rings).
The reproducer from vega@nebusec.ai set the following params for
illustration: txqlen of 500000 -> ~32 GiB/ring attempts, 1.6 GB tun,
~960 MB tap. Gets worse when you consider qdiscs like mq.

What we fix: every path an unprivileged user can use to install
an oversized tx_queue_len is rejected with -ERANGE before any ring is
allocated; per-ring memory is bounded at 256 KiB.

This is for you sashikos: What we deliberately _do not fix_
bound the NUMBER of rings. With the cap in place the worst case moves
from "one knob" to the aggregate of ring x queues x devices, example:

  ip link add v0 numtxqueues 4096 txqueuelen 32767 type veth
  tc qdisc add dev v0 root mq
    -> 4096 * 3 * 32767 * 8 = ~3.0 GiB (one command)
  50 tun devices x 256 queues x 32767 x 8 = ~3.1 GiB

Unfortunately tx_queue_len is a bit ambigious in meaning:
In some cases it means a ring size (which is pre-allocated, ex:
tun, tap, and pfifo_fast); a cap of 4096 seems reasonable here.
but in other cases it is used to indicate a queue limit ex:
the qdisc consumers that allocate nothing (pfifo/bfifo/gred/plug/sfb,
htb direct_qlen, qfq, teql). 32767 is a legitimate high-BDP queue
length, so we are going to keep that value.

Getting back to you sashikos, after this is merged and shows up
in net-next we will send followup patches as follows:
this series is not misread as "closes the OOM class"):

a) Per-site ring limits at six identified locations
    - pfifo_fast init/resize,
    - tun attach/resize,
    - tap minor/resize)

   if you can spot more in your review we will take care of those as well.

b) memcg accounting (GFP_KERNEL_ACCOUNT) for those ring
   allocations: contains a memcg-limited container's ring memory.
   Not GFP_KERNEL_ACCOUNT has no effect on the unshare attacker
   but will protect against containers  (memory.max in its cgroup)
====================

Link: https://patch.msgid.link/QDISC-2899.v2.20260901233641@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 16:30:53 -07:00
Jamal Hadi Salim
0a7252d7f8 selftests: tc-testing: add tx_queue_len cap regression tests
Add nine test cases for the S16_MAX tx_queue_len cap to the
pfifo_fast suite. Netlink cases exercise the ifla_policy bound
(2/3); the two new sysfs cases exercise the netif_change_tx_queue_len()
choke point that 1/3 owns (SIOCSIFTXQLEN shares it; the ioctl is not
portably reachable from tdc):

- dbe3: set txqueuelen 32767 (S16_MAX) - accepted, pins the exact
  boundary value.
- b50e: set txqueuelen 32768 - rejected with -ERANGE.
- 40f8: write 32768 to /sys/class/net/*/tx_queue_len - rejected
  (covers patch 1/3 directly; netlink cannot reach this path).
- 4b6e: write 32767 via sysfs - accepted, boundary positive control
  for the patch-1 path.
- b90d: create a dummy with txqueuelen 32767 - accepted.
- 57ab: create a dummy with txqueuelen 32768 - rejected at netlink
  parse time.
- e777: create a dummy with txqueuelen 500000 - rejected (the v1
  bypass path flagged by review).
- 31ac: create a veth with an oversized txqueuelen on the peer nest -
  rejected (the peer nest is parsed against ifla_policy too).
- b567: create a veth with txqueuelen on both ends within the cap -
  accepted (positive control for the peer nest).

The three negative-creation verifies assert device absence
("ip -o link show" must not contain the device), not merely absence
of a qlen pattern - the device does not exist when creation fails, so
the exit code carries the signal and the verify adds content.

The v1 04b5 "resize rollback" case is dropped: with the cap checked
first, netif_change_tx_queue_len() returns -ERANGE before the write,
the notifier or any qdisc resize, so the case exercised no resize and
no rollback. It was also nondeterministic: pre-patch, the resize
issues three ~11 MB kvmallocs for qlen 500000 which normally succeed,
so the case passed on an unfixed kernel only under memory pressure -
its outcome depended on the test host's free memory.

Test commands run inside the netns, but nsPlugin creates the veth
peer in the root namespace, so the teardown deletes the in-ns end
only; deleting the peer via the pair is implicit.

Note: iproute2 treats "txqueuelen" appearing after "type X" as a
link-type attribute and silently drops it, so the creation cases
place it before "type" to actually reach the kernel.

Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-2899.v2.20260901233641@mojatatu.com.3
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 16:30:49 -07:00
Jamal Hadi Salim
1aa9e143bf net: reject oversized tx_queue_len at netlink parse time
rtnl_create_link() assigns IFLA_TXQLEN directly to dev->tx_queue_len
without going through netif_change_tx_queue_len(), so a device created
with "ip link add ... txqueuelen 500000" bypasses the S16_MAX cap and
still triggers the oversized ring allocations in pfifo_fast, tun and
tap. The veth peer nest (rtnl_nla_parse_ifinfomsg()) and the
RTM_NEWLINK-on-existing-device path reach the same sinks.

Enforce the cap in ifla_policy instead: IFLA_TXQLEN becomes
NLA_POLICY_FULL_RANGE(NLA_U32, &txqlen_range) with
txqlen_range = { .min = 0, .max = S16_MAX }. All netlink consumers
parse against this policy - rtnl_setlink(), rtnl_newlink() (create
and change), and the veth peer nest - so every netlink path is capped
at parse time and rejects the attribute with -ERANGE plus a proper
"integer out of range" extack message before any device state is
modified (the RTM_SETLINK half-application wart is gone with it).

Document the bound in the rt-link.yaml netlink spec.

Conditions to recreate the bug:
- CONFIG_NET_SCHED=y, CONFIG_VETH=y, CONFIG_USER_NS=y, CONFIG_NET_NS=y.
- Unprivileged user in a fresh user+net namespace (unshare -Urn):
  ip link add v0 txqueuelen 500000 type veth peer name v1
  -> on the fixed kernel this is rejected with -ERANGE ("integer out
  of range" extack) instead of installing an oversized tx_queue_len
  that later inflates pfifo_fast/tun/tap ring allocations.
- ip link set v0 txqueuelen 500000 is likewise rejected at parse time.

Fixes: 38f7b870d4 ("[RTNETLINK]: Link creation API")
Reported-by: Vega <vega@nebusec.ai>
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-2899.v2.20260901233641@mojatatu.com.2
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 16:30:49 -07:00
Jamal Hadi Salim
66ab4c59b7 net: cap tx_queue_len at S16_MAX to prevent oversized ring allocations
Several subsystems allocate ring buffers sized by dev->tx_queue_len
with no upper bound. An unprivileged user (via unshare -Urn) can set a
huge tx_queue_len and exhaust global memory with ring allocations:

- pfifo_fast: pfifo_fast_init() and pfifo_fast_change_tx_queue_len()
  allocate 3 skb_array rings of tx_queue_len entries each.
- tun: tun_queue_resize() and the queue-attach path resize ptr_rings
  to tx_queue_len on the NETDEV_CHANGE_TX_QUEUE_LEN notifier.
- tap (macvtap/ipvtap): tap_queue_resize() and tap_init() resize/init
  ptr_rings to tx_queue_len on the same notifier.

netif_change_tx_queue_len() is the single entry point for IFLA_TXQLEN,
sysfs, and the SIOCSIFTXQLEN ioctl. Cap new_len at S16_MAX (32767)
there so the oversized value is rejected at set time. This takes
effect whether the device is up or down, before dev->tx_queue_len is
written, before any notifier fires, and before any ring is allocated.
The "> S16_MAX" check also subsumes the previous unsigned-long
truncation test, and a negative ifr_qlen from the ioctl lands far
above the cap after conversion, so both old failure modes are covered
by the one comparison.

tx_queue_len is ambigious: both a per-ring sizing multiplier and a
default queue-length/limit knob for consumers that allocate
nothing at set time (pfifo/bfifo/gred/plug/sfb limits, htb
direct_qlen, qfq max_classes, teql). 32767 is chosen as the largest
value NLA_POLICY_FULL_RANGE can express for the u32 IFLA_TXQLEN
policy in patch 2/3 while staying a legitimate queue length on
high-BDP paths; the ring-memory trade-off of a shared knob is
disclosed below.

Conditions to recreate the bug:
- CONFIG_NET_SCHED=y, CONFIG_VETH=y, CONFIG_USER_NS=y, CONFIG_NET_NS=y.
- Unprivileged user in a fresh user+net namespace (unshare -Urn).
- pfifo_fast: create veth pairs, set tx_queue_len to 500000, attach
  mq+pfifo_fast. ~28 iterations OOMs a 2GB guest.
- tun: create 50 tun devices with IFF_MULTI_QUEUE, set tx_queue_len to
  500000, open 8 queues each. ~1.6GB of ptr_ring allocations OOMs a
  512MB guest.
- tap: same as tun with IFF_TAP. ~960MB OOMs a 512MB guest.
- On the fixed kernel the oversized tx_queue_len is rejected with
  -ERANGE at set time (all four paths: RTM_SETLINK, RTM_NEWLINK
  create, sysfs, ioctl - the latter two via this check, the former
  two via this check and the 2/3 parse policy respectively).

Fixes: 6a643ddb56 ("net: introduce helper dev_change_tx_queue_len()")
Reported-by: Vega <vega@nebusec.ai>
Closes: https://lore.kernel.org/netdev/20260828121902.66837-1-jhs@mojatatu.com/
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-2899.v2.20260901233641@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 16:30:48 -07:00
Seungwon Bae
98fc57d167 vxlan: reject dynamic fdb entries that reference a nexthop id
The commit cited in the Fixes tag allowed VXLAN FDB entries to point to
FDB nexthops so that overlay traffic could be load balanced across
multiple VTEPs. Such entries can only be configured from user space,
cannot be learned and cannot roam. They only make sense with a user space
control plane such as E-VPN where data plane learning is disabled.

Despite that, the VXLAN driver does not currently prevent such entries
from being configured with the "dynamic" flag. The per-nexthop FDB list
is only protected by the per-device hash lock, which is not sufficient
when two VXLAN devices point to the same FDB nexthop and therefore share
the list. Aging runs in softirq context without RTNL, so an entry deleted
by one device can race with an addition or deletion from the other,
leading to list corruption:

  list_del corruption. next->prev should be ffff8881069d9548, but was
  dead000000000122. (next=ffff8881069d9448)
  WARNING: CPU: 0 PID: 90 at lib/list_debug.c:65
  __list_del_entry_valid_or_report+0x1aa/0x210
  ...
   vxlan_fdb_destroy+0x5b8/0xad0
   vxlan_cleanup+0x328/0x450
   call_timer_fn+0x2a/0x1c0
   run_timer_softirq+0x18c/0x210
  BUG: KASAN: slab-use-after-free in vxlan_fdb_destroy

Fix this by rejecting the bogus configuration of dynamic FDB entries that
point to FDB nexthops, both when created and when an existing entry is
updated. As such, the per-nexthop FDB list is only ever mutated under the
RTNL lock. Add test cases to make sure that this does not regress in the
future.

Fixes: 1274e1cc42 ("vxlan: ecmp support for mac fdb entries")
Suggested-by: Ido Schimmel <idosch@nvidia.com>
Signed-off-by: Seungwon Bae <qotmddnjs@ajou.ac.kr>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260902155956.296699-1-qotmddnjs@ajou.ac.kr
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 16:14:52 -07:00
Alexandra Winter
907a56ab3e s390/ism: folio_put() after error
dmb->cpu_addr was allocated via folio_alloc(). Use folio_put() instead of
kfree() in the error exit of ism_alloc_dmb() to avoid slab allocator
corruption.

While at it, reset dmb->cpu_addr after folio_put to avoid unintentional UAF
by future callers.

Fixes: 83781384a9 ("s390/ism: Properly fix receive message buffer allocation")
Signed-off-by: Alexandra Winter <wintera@linux.ibm.com>
Reviewed-by: Gerd Bayer <gbayer@linux.ibm.com>
Link: https://patch.msgid.link/20260902143733.433574-1-wintera@linux.ibm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 16:12:36 -07:00
Alexandra Winter
1668a31e3b dibs: Unregister dibs_class after error
In case dibs_loopback_init() fails, e.g. because of -ENOMEM, dibs_init()
must unregister dibs_class. Otherwise dibs_class and /sys/class/dibs exist
even though the functionality is not available. A retry to load the module
fails with -EEXIST.

Unregister dibs_class in the error path of dibs_init.

Note that before
commit ad3dfa80be ("dibs: change dibs_class to a const struct")
class_destroy(dibs_class) is required instead of
class_unregister(&dibs_class).

Fixes: 8047373498 ("dibs: Create class dibs")
Signed-off-by: Alexandra Winter <wintera@linux.ibm.com>
Link: https://patch.msgid.link/20260902143438.426664-1-wintera@linux.ibm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 16:12:24 -07:00
Jason Winter
5d50e90add net: usb: cx82310_eth: drop URB after 0xffff reboot sentinel to prevent partial_data heap overflow
The 0xffff length sentinel detects a router reboot and schedules
re-enabling of ethernet mode, but then falls through to the rest
of the loop body.  The next check is

	} else if (len > CX82310_MTU) {

which is the else of the just-matched if -- it never fires for
len == 0xffff.  The MTU bound that normally caps the
incomplete-packet save path is silently bypassed.

With 0xffff > skb->len always true (rx_urb_size is 4096), the
incomplete-packet branch saves dev->partial_len = skb->len bytes
into dev->partial_data.  partial_data is kmalloc(hard_mtu) =
kmalloc(CX82310_MTU + 2) = 1516 bytes, but skb->len after the
2-byte header pull can be up to 4094.  A device that sends a
4096-byte URB starting with [0xff 0xff] therefore copies 4094
device-provided bytes into a buffer allocated for 1516 bytes,
exceeding its requested size by 2578 bytes.

The next URB then reads dev->partial_len (4094) back from the same
1516-byte buffer and dev->partial_rem (65535 - 4094 = 61441) from
the new URB's ~4KB skb, both well past their allocations, and
delivers the spliced result as a 64KB "frame" to the network
stack.

Bail out of rx_fixup after scheduling the re-enable work; the
remainder of a reboot-marker URB is not meaningful packet data.
This restores the invariant that partial_len < CX82310_MTU + 2 on
the save path, since every other route there has already passed
the MTU check.

Fixes: ca139d76b0 ("cx82310_eth: re-enable ethernet mode after router reboot")
Signed-off-by: Jason Winter <jjx@live.nl>
Link: https://patch.msgid.link/BESP194MB283265DDDC63B6B78D8D34FBB8B72@BESP194MB2832.EURP194.PROD.OUTLOOK.COM
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 16:05:48 -07:00
Fourie Zhang
78a86d75a7 net: mpls: clear inner_protocol when the last label is popped
skb_mpls_push() records the pre-encapsulation network header once, gated
on !skb->inner_protocol. skb_mpls_pop() never clears that record, so it
outlives the encapsulation it describes.

Open vSwitch can then re-push MPLS onto a packet whose
inner_network_header still points at the older, deeper offset: push a
label, pop every label, recirculate (ovs_flow_key_update() re-derives
key->eth.type and resets network_header, but leaves inner_*), then push
again. ovs_fragment() trusts the record:

	skb->network_header = skb->inner_network_header;

so skb_network_offset() goes negative. The bound check is signed:

	if (skb_network_offset(skb) > MAX_L2_LEN)

a negative offset passes it, and prepare_frag() widens the value:

	unsigned int hlen = skb_network_offset(skb);
	memcpy(&data->l2_data, skb->data, hlen);

which is a ~4GiB memcpy out of a 30-byte per-CPU buffer.

Reproduced on v7.3-rc1. RDX is the truncated length, (unsigned int)(-8):

  BUG: unable to handle page fault for address: ffffe8ffffc16000
  #PF: supervisor write access in kernel mode
  Oops: 0002 [#1] SMP KASAN NOPTI
  RIP: 0010:memcpy+0x8/0x20
  RDX: 00000000fffffff8 RSI: ffff888105d732db RDI: ffffe8ffffc16000
   prepare_frag+0x3df/0x4e0
   ovs_fragment+0x589/0x7e0
   do_output+0x4ce/0x5e0
   do_execute_actions+0x55d2/0x7b30
   ovs_execute_actions+0xea/0x450

Same root-cause shape as commit 975b5b067f ("ipv6: sr: restore network
header before routing and forwarding"): a stale network header offset
reaching a consumer that widens it. Here it originates in the MPLS
push/pop path.

Clear inner_protocol once the packet is no longer MPLS, so a later push
re-records the current header. net/sched/act_mpls.c is the only other
skb_mpls_pop() caller and gets the same fix; sch_frag.c saves and
restores inner_protocol around fragmentation in the same way OVS does.

Fixes: 48d2ab609b ("net: mpls: Fixups for GSO")
Cc: stable@vger.kernel.org
Signed-off-by: Fourie Zhang <fouriezhang@tencent.com>
Acked-by: Jiri Benc <jbenc@redhat.com>
Link: https://patch.msgid.link/20260902092719.2874481-1-fouriezhang@tencent.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 16:03:53 -07:00
Nikhil P. Rao
c91b4d6e5c ionic: use netif_txq_maybe_stop() in ionic_tx()
Commit 061b9bedbe ("ionic: Rework Tx start/stop flow") replaced
ionic_maybe_stop_tx() with netif_txq_maybe_stop() to get the memory
barriers around the stop/start bits right, but did not cover the stop
in ionic_tx() added by commit 138506ab24 ("ionic: Check stop no
restart"). Convert the remaining site.

netif_txq_maybe_stop() requires the ring indexes to be updated before
it is invoked, so the post has to come first. But ring_dbell comes
from __netdev_tx_sent_queue(), which runs after that and reads the
stop bit, so it is not known in time to pass to ionic_txq_post(). Post
without the doorbell and ring it separately.

The stop condition is unchanged. The re-check only clears the stop bit
when space has become available, so the doorbell starvation fixed by
commit 138506ab24 ("ionic: Check stop no restart") cannot recur.

Fixes: 138506ab24 ("ionic: Check stop no restart")
Signed-off-by: Nikhil P. Rao <nikhil.rao@amd.com>
Reviewed-by: Brett Creeley <brett.creeley@amd.com>
Link: https://patch.msgid.link/20260901055627.1373129-1-nikhil.rao@amd.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:50:51 -07:00
Viswajith Murali
1f29543126 octeontx2-af: mcs: Clear stale X2P calibration state before calibration
Some firmware versions leave MCSX_MIL_GLOBAL bit 5 set on boot.
If the bit is already set when the driver attempts X2P calibration,
the hardware sees no rising edge and calibration never triggers.
Clear the bit and wait briefly before starting calibration to ensure
a clean rising edge.

Fixes: ca7f49ff88 ("octeontx2-af: cn10k: Introduce driver for macsec block.")
Signed-off-by: Nitin Shetty J <nshettyj@marvell.com>
Signed-off-by: Viswajith Murali <viswajithm@marvell.com>
Link: https://patch.msgid.link/20260901094318.1395356-1-nshettyj@marvell.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:44:59 -07:00
Nikolay Aleksandrov
4b772869a1 net: bridge: mcast: properly convert mglist to rcu
Sashiko reported a bug [1] that br_multicast_del_port_group unlists the
port group not using proper rcu helper that preserves the next pointer and
after that immediately frees the port group without waiting for rcu grace
period. The only rcu walker of mglist is br_multicast_list_adjacent() and
it turns out that function has always been buggy because mglist was never
properly converted to RCU. Fix it by converting it to rcu and moving its
initialization after eth_addr's. Initializing p->next can use
RCU_INIT_POINTER because we have a barrier from the hlist_add_head_rcu call
later, besides we're initializing an unpublished structure anyway.

[1] https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260826014200.362304-1-littleddfu%40gmail.com

Fixes: 07f8ac4a1e ("bridge: add export of multicast database adjacent to net_dev")
Signed-off-by: Nikolay Aleksandrov <razor@blackwall.org>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260903093851.1494297-1-razor@blackwall.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:29:47 -07:00
Jakub Kicinski
742b94967d Merge branch 'eth-fix-bugs-in-ntuple-filter-reporting'
Jakub Kicinski says:

====================
eth: fix bugs in ntuple filter reporting

Looking thru some reports prompted by:
  Add new way to add BPF LSM hooks
  https://lore.kernel.org/20260831110934.241898-1-a.s.protopopov@gmail.com

I/Claude noticed 3 drivers with buggy n-tuple filter dump.
Fix these drivers, add a hopefully clearer mention in the doc.
====================

Link: https://patch.msgid.link/20260903032611.3000029-1-kuba@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:20:00 -07:00
Jakub Kicinski
47a582b2b0 ethtool: document that GRXCLSRLALL rule_cnt is a caller-provided limit
Three drivers have shipped a get_rxnfc() which dumps its entire rule
table into rule_locs, reading rule_cnt as "how many rules do I have"
rather than "how many entries did the caller allocate".  Nothing in the
callback's documentation contradicted that reading.  The distinction only
matters because the ioctl lets an unprivileged caller pick rule_cnt
directly, so getting it wrong is a heap overflow rather than a truncated
dump.

Reviewed-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260903032611.3000029-6-kuba@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:19:58 -07:00
Jakub Kicinski
b1fffc2731 net: dsa: mv88e6xxx: bound the policy rule dump by the caller's buffer size
mv88e6xxx_get_rxnfc() uses rxnfc->rule_cnt as the write index while
dumping the policy IDR, clobbering the input value before it has been
looked at.  That input is the number of entries the caller had room for.
ETHTOOL_GRXCLSRLALL requires no CAP_NET_ADMIN and the ioctl sizes the
buffer from the rule_cnt userspace passes in, so once an admin has
installed policy rules any user can ask for fewer slots than there are
rules and run off the end of the allocation.  A rule_cnt of 0 leaves the
buffer pointer NULL and the walk dereferences it.

Count into a local so the caller's limit survives the walk, and stop with
-EMSGSIZE once it is reached.

Fixes: da7dc87553 ("net: dsa: mv88e6xxx: add RXNFC support")
Reviewed-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260903032611.3000029-5-kuba@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:19:57 -07:00
Jakub Kicinski
108bb2142e eth: nfp: drop the replaced rule from the list when reprogramming fails
nfp_net_fs_add() replaces an existing rule by deleting it from the
hardware, decrementing nn->fs.count and programming the new one.  If
nfp_net_fs_add_hw() fails the old entry stays on nn->fs.list - only the
success path reaches list_replace() - so the list is one longer than
nn->fs.count, and it advertises a rule whose hardware entry has already
been torn down.

nn->fs.count is what ETHTOOL_GRXCLSRLCNT reports, so userspace then sizes
its buffer one entry short of what the GRXCLSRLALL walk wants to write.
That used to overwrite one u32 past the allocation; since the walk is
bounded it is a permanent -EMSGSIZE instead, as nothing ever resyncs the
counter.

Fixes: 9eb03bb1c0 ("nfp: add ethtool flow steering callbacks")
Reviewed-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260903032611.3000029-4-kuba@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:19:57 -07:00
Jakub Kicinski
f1986bf87b eth: nfp: bound the ntuple rule dump by the caller's buffer size
nfp_net_get_fs_loc() dumps every entry of nn->fs.list into rule_locs[]
without consulting cmd->rule_cnt, which is how many entries the caller
had room for.  ETHTOOL_GRXCLSRLALL requires no CAP_NET_ADMIN and the
ioctl sizes the buffer from the rule_cnt userspace passes in, so once an
admin has installed flow steering rules any user can ask for fewer slots
than there are rules and run off the end of the allocation.  A rule_cnt
of 0 leaves the buffer pointer NULL and the walk dereferences it.

Bail out with -EMSGSIZE when the buffer fills up, the way the other
ntuple capable drivers do, and report how many locations were filled so
a shrinking rule list does not leave the caller reading stale slots.

Reported-by: VEGA <vega@nebusec.ai>
Fixes: 9eb03bb1c0 ("nfp: add ethtool flow steering callbacks")
Reviewed-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260903032611.3000029-3-kuba@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:19:57 -07:00
Jakub Kicinski
cdb719f4b8 net: dsa: bcm_sf2: bound the CFP rule dump by the caller's buffer size
bcm_sf2_cfp_rule_get_all() walks the whole cfp.unique bitmap into
rule_locs[] without consulting nfc->rule_cnt, which is how many entries
the caller had room for.  ETHTOOL_GRXCLSRLALL requires no CAP_NET_ADMIN
and the ioctl sizes the buffer from the rule_cnt userspace passes in, so
once an admin has installed CFP rules any user can ask for fewer slots
than there are rules and run off the end of the allocation.  A rule_cnt
of 0 leaves the buffer pointer NULL and the walk dereferences it.

Fixes: 7318166cac ("net: dsa: bcm_sf2: Add support for ethtool::rxnfc")
Reviewed-by: Jonas Gorski <jonas.gorski@gmail.com>
Reviewed-by: Florian Fainelli <florian.fainelli@broadcom.com>
Reviewed-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260903032611.3000029-2-kuba@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:19:57 -07:00
Eric Dumazet
1746ef2e2d bonding: use skb_cow_head() in bond_do_alb_xmit() and rlb_arp_xmit()
In bond_do_alb_xmit() and rlb_arp_xmit(), make sure to unclone
skb head via skb_cow_head() before modifying the source MAC address
(Ethernet header and ARP payload) to avoid silent corruption if
the skb is shared or cloned. Avoid caching the header pointers
across skb_cow_head().

In rlb_arp_xmit(), only modify arp->mac_src if it differs from
tx_slave->dev->dev_addr to avoid an unnecessary copy and head
reallocation.

Also, we should not assume mac header is set in output path.

Use skb_eth_hdr() instead of eth_hdr() to fix the issue,
and remove now redundant skb_reset_mac_header() calls.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Cc: Jay Vosburgh <jv@jvosburgh.net>
Reviewed-by: Nikolay Aleksandrov <razor@blackwall.org>
Link: https://patch.msgid.link/20260903143940.1180513-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:18:41 -07:00
Ido Schimmel
5bd9e4e7cd nexthop: Initialize extack in remove_nh_grp_entry()
remove_nh_grp_entry() prints the extack message when a listener fails
to replace the reduced nexthop group. However, extack is not
initialized and listeners are not required to set a message when
returning an error. Neither netdevsim nor mlxsw do so when an
allocation fails, resulting in the dereference of an uninitialized
stack pointer.

Fix by zero-initializing extack, as was done in commit 6347c5314c
("nexthop: initialize extack in nh_res_bucket_migrate()").

Fixes: 833a1065ee ("nexthop: Emit a notification when a nexthop group is reduced")
Signed-off-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260903080259.10378-1-idosch@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:17:41 -07:00
Jakub Kicinski
fb3088dc58 Merge tag 'ieee802154-for-net-2026-09-03' of git://git.kernel.org/pub/scm/linux/kernel/git/wpan/wpan
Stefan Schmidt says:

====================
pull-request: ieee802154 for net 2026-09-03

Zhiling Zou fixed a NULL deref when coming from a TUN device.

Fan Wu fixed a UAF in the cc2520 driver.

Chenguang Zhao fixed up some out of date comments in 6lowpan.

David Carlier fixed a potential double free in the hwsim driver.

Ibrahim Hashimov reworked the queuing in the RX path to fix a UAF on beacon
and MAC frames.

* tag 'ieee802154-for-net-2026-09-03' of git://git.kernel.org/pub/scm/linux/kernel/git/wpan/wpan:
  mac802154: fix use-after-free of sdata via queued RX frames
  ieee802154: hwsim: serialize pib updates to fix double-free
  ieee802154: 6lowpan: fix NULL dereference in lowpan_newlink
  ieee802154: cc2520: fix FIFOP work use-after-free
  net: 6lowpan: fix mismatched comments
====================

Link: https://patch.msgid.link/20260903093012.4032586-1-stefan@datenfreihafen.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:16:28 -07:00
Jakub Kicinski
641d03105c Merge branch 'enic-fix-v2-vf-mailbox-reply-matching-and-carrier-reopen'
Satish Kharat says:

====================
enic: fix V2 VF mailbox reply matching and carrier reopen

Preserve ENIC V2 VF carrier state across netdev reopen and match mailbox
replies to their originating requests.

A V2 VF receives carrier state from PF MBOX notifications. enic_stop()
forces carrier off, but enic_open() does not request another notification
or restore the previous state. An ordinary netdev close/open can therefore
leave the VF without carrier until the PF sends another link-state
notification.

The first patch preserves the last PF-reported link state across an
ordinary netdev close/open. Internal reset paths invalidate the saved state
before reopening the datapath, so carrier remains off until the PF provides
a new notification.

Mailbox messages carry a message number that replies and acknowledgments
echo. ENIC currently assigns a new number to outgoing replies and accepts
VF replies by message type alone. After a request times out, a delayed
reply can therefore satisfy a later request of the same type.

The second patch makes PF replies and the VF link-state acknowledgment echo
the initiating message number. The VF accepts a reply only when both its
type and message number match the pending request. Reply handling and
timeout cleanup are protected by the same lock, and message numbering
remains monotonic across admin-channel reopen.

Validation:

- The VF module with this series applied passed multiple netdev close/open
  and ENIC module unload/reload cycles, as well as guest reboot testing,
  with carrier restored as expected after each operation.
- With this series folded into the full SR-IOV development stack, VFIO
  guests passed bidirectional cross-DUT PF-to-VF, VF-to-PF, and VF-to-VF
  traffic using standard-sized and 8972-byte jumbo ICMP packets.
====================

Link: https://patch.msgid.link/20260830-b4-enic-v2-mbox-fixes-net-v1-0-23adf9bfd426@cisco.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 19:10:20 -07:00
Satish Kharat
8972d252f4 enic: match mailbox replies to request numbers
The version-1 VF mailbox protocol identifies every message with a message
number, and a reply or acknowledgment echoes the number of the message it
answers.  ENIC instead generates a new number for outgoing replies and
accepts a VF reply by message type alone.

If a request times out, a delayed reply can therefore satisfy a subsequent
request of the same type and cause the VF to consume the result of the old
request.

Allow replies to reuse the initiating message number. Make the in-tree PF
handlers and the VF link-state acknowledgment echo that number. Record the
expected reply type and message number on the VF, and require both values
to match before accepting a reply.

Protect expected-reply state with a lock so reply acceptance and timeout
invalidation cannot race. Keep message numbers monotonic across an admin-
channel reopen so a delayed reply from an earlier channel generation cannot
match a new request.

Reply-number echo is part of the established version-1 protocol, so this
remains compatible with deployed V2-capable PF implementations that already
echo msg_num.

Fixes: 72b65c9405 ("enic: add MBOX VF handlers for capability, register and link state")
Signed-off-by: Satish Kharat <satishkh@cisco.com>
Link: https://patch.msgid.link/20260830-b4-enic-v2-mbox-fixes-net-v1-2-23adf9bfd426@cisco.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 19:10:18 -07:00
Satish Kharat
b752e041d5 enic: preserve V2 VF carrier across netdev reopen
A V2 VF receives carrier state only from PF MBOX notifications.
enic_stop() forces carrier off, but enic_open() does not request a fresh
notification or restore the previous one.  An ordinary down/up cycle
therefore leaves the VF in NO-CARRIER and unable to pass traffic until the
PF repeats the link-state command, even when the physical link remained
up.

Cache each valid PF link-state notification.  Serialize updates with the V2
VF datapath running state.  Keep carrier off while the netdev is stopped.
Restore the cached state after an ordinary open.  Before either internal
reset reopens the datapath, invalidate the cache.  Carrier then remains off
until re-registration receives a fresh PF link-state notification.

Fixes: 72b65c9405 ("enic: add MBOX VF handlers for capability, register and link state")
Signed-off-by: Satish Kharat <satishkh@cisco.com>
Link: https://patch.msgid.link/20260830-b4-enic-v2-mbox-fixes-net-v1-1-23adf9bfd426@cisco.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 19:10:18 -07:00
Joe Damato
39b23c1c40 bnxt_en: Prevent queue stop with deferred completions
When the driver receives a burst of packets, it can mark a BD with the
NO_CMPL bit to defer completions. The expectation is that the last
packet in the ring will have this bit unset and the completion generated
by that packet will cleanup that packet and the ones preceding it. This
helps to reduce the number of completions fired.

The suppressed completions are controlled by the driver and the number
of packets with suppressed completions scales with the size of the ring.
SW USO packets, on the other hand, have an upper bound on the maximum
number of BDs which can be consumed which does not scale with the ring
size.

So, for small rings it is possible that: a burst of packets is handed to
the driver, the driver defers completions for all of the packets because
the number of free descriptors stays above the threshold in the driver.
Then, a USO packet arrives, but the number of BDs available is not
enough and the USO code exits early.

In this case, you end up in a state where the ring is full of packets
with their completions suppressed, which can cause the queue to stop and
never be restarted.

Assuming default CONFIG_MAX_SKB_FRAGS, this is only possible for small
rings (<= 457 descriptors, below the driver default value) when
a burst of packets fills the ring, followed by a large USO packet that
can't fit. For larger rings, the delta between the completion
suppression threshold and the BDs required for SW USO is large enough
that completions will fire and this case is unreachable.

This issue was pointed out by Sashiko and while it seems fairly unlikely
given that the queue size must be small to trigger this, it is indeed
possible.

Fix this by tracking the last BD which deferred completions and
centralizing the logic for deciding when to ring the doorbell. The NO_CMPL
bit is now cleared in bnxt_txr_db_kick(), so every doorbell site is
covered, including the SW USO early exit. This guarantees the ring always
ends in a BD which generates a completion to clean it and wake the queue.

Fixes: cc5d90667d ("net: bnxt: Implement software USO")
Cc: <stable@vger.kernel.org> # v7.1+: 4e15e89faac9: net: bnxt: ring the doorbell when SW USO exits early
Signed-off-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260902213956.4160615-1-joe@dama.to
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 17:49:57 -07:00
Lorenzo Bianconi
a09ceadff9 mailmap: add entries for Lorenzo Bianconi
Add the active email address for Lorenzo Bianconi and map the old,
no-longer-used addresses to it, so that git can attribute his
contributions to a single identity.
This is done to avoid bouncing emails sent to email addresses that are
no longer active.

Signed-off-by: Lorenzo Bianconi <lorenzo@kernel.org>
Link: https://patch.msgid.link/20260901-lorenzo-mailmap-v2-1-0ee832de0caf@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 17:47:52 -07:00
Ido Schimmel
b58d749633 tunnels: Drop stale dst when building an ICMP error for PMTUD
Bridged UDP tunnels such as VXLAN and GENEVE build an ICMP error packet
around an overlay packet if the packet is going to exceed the underlay
path MTU. The ICMP error packet is then injected back into the Rx path
with the source and destination addresses swapped, so that it will be
delivered to the overlay source.

If the overlay packet was routed to the UDP tunnel or locally generated,
then it is already carrying a valid dst entry and this entry is not
dropped when transforming the packet to an ICMP error packet. This
causes the IP layer to reuse the dst entry, leading to the ICMP error
packet being dropped or routed out of the UDP tunnel interface in case
of forwarding.

Prior to the blamed commit this could not happen, as
skb_tunnel_check_pmtu() did not build ICMP errors for PACKET_HOST
packets. Such packets were instead encapsulated and, unless the DF bit
was set in the outer header, fragmented by the underlay.

Fix this by making sure that the ICMP error packet does not have a valid
dst entry, thereby forcing the IP layer to perform a route lookup.

Adjust the bridged PMTU exception selftests accordingly. When the
local sender in ns_a pings the overlay destination with a deadline
(-w), ping exits on the first socket error before any reply is
received and returns a non-zero exit code. The test therefore only
passed because the ICMP error was never delivered. Use a packet count
(-c) like the ns_c line above it, so that the ICMP error counts
against the packet budget and the exit code depends on whether echo
replies were received. This passes with and without the fix.

Fixes: 8930424777 ("tunnels: Accept PACKET_HOST in skb_tunnel_check_pmtu().")
Cc: stable@vger.kernel.org
Reported-by: Laika Price <laikabcprice@gmail.com>
Closes: https://lore.kernel.org/netdev/20260614-master-v3-1-9f5060ba1ed1@gmail.com/
Reported-by: Yaroslav Dudkov <aroslavdudkov622@gmail.com>
Closes: https://lore.kernel.org/netdev/20260901081825.287173-1-aroslavdudkov622@gmail.com/
Reported-by: Charles Bordet <rough.rock3059@datachamp.fr>
Closes: https://lore.kernel.org/netdev/aHVhQLPJIhq-SYPM@eldamar.lan/
Signed-off-by: Ido Schimmel <idosch@nvidia.com>
Tested-by: Yaroslav Dudkov <aroslavdudkov622@gmail.com>
Reviewed-by: David Ahern <dsahern@kernel.org>
Reviewed-by: Stefano Brivio <sbrivio@redhat.com>
Reviewed-by: Guillaume Nault <gnault@redhat.com>
Link: https://patch.msgid.link/20260902190112.4126199-1-idosch@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 17:40:55 -07:00
XingWang Xiang
6a1094c34d genetlink: pin family module during policy dump
The generic netlink controller's policy dump keeps pointers to the target
family's operation and policy tables in its callback state.  A dump may be
split across multiple skbs and remain pending after the initial request.

Netlink pins the module which owns the dump callback, but in this case
that is the controller's owner rather than the target family's owner.  The
target family can consequently be unregistered and its module unloaded
while a policy dump is pending.  Advancing the dump then dereferences
policy memory from the unloaded module.

Take a reference to the target family's module when the dump starts.
Drop it from the error and done paths.  This matches the lifetime for which
the dump context retains the family and policy pointers.

Fixes: d07dcf9aad ("netlink: add infrastructure to expose policies to userspace")
Cc: stable@vger.kernel.org
Signed-off-by: XingWang Xiang <v3rdant.xiang@gmail.com>
Link: https://patch.msgid.link/20260902084317.4092542-1-v3rdant.xiang@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 17:21:35 -07:00
Taylor Bates
2ac174dfcd mlxsw: spectrum_ptp: Fix napi_gro_receive() call from GC workqueue context
Currently mlxsw_sp1_ptp_ht_gc_collect() is run from the PTP
garbage-collection workqueue, rather than the NAPI poll context. For any
unmatched PTP entries carrying an SKB, it calls
mlxsw_sp1_ptp_unmatched_finish() -> mlxsw_sp1_ptp_packet_finish(). For
ingress packets, this calls mlxsw_sp_rx_listener_no_mark_func(). The end
of that function is the following:

    skb->protocol = eth_type_trans(skb, skb->dev);
    napi_gro_receive(mlxsw_skb_cb(skb)->rx_md_info.napi, skb);

The napi pointer is one that was placed in the SKB control block when the
trapped packet was received in the NAPI context. Later, when the GC reaps
the unmatched entry (up to MLXSW_SP1_PTP_HT_GC_TIMEOUT later), the call to
napi_gro_receive() mutates the NAPI instance's GRO list, which is unsafe
if the poll is running concurrently on another CPU.

In mlxsw_sp1_ptp_ht_gc_collect(), local_bh_disable() is called to prevent
softirq processing, but this only applies to the local CPU. Additionally,
its comment is stale. It states that mlxsw_sp1_ptp_unmatched_finish()
invokes netif_receive_skb(). This has not been accurate since the
referenced commit; this patch makes that comment accurate again.
mlxsw_pci_napi_devs_init() calls netif_threaded_enable() on the NAPI RX
net_device without any conditions. The NAPI instance's poll, which may be
running concurrent to the GC, is running as an independently-scheduled
kthread which may be on a different CPU. The call to local_bh_disable()
does not guard against this.

If a tx-timestamp timeout produces an unmatched entry (which can be easily
reproduced by running ptp4l and waiting for a port to reach the
UNCALIBRATED/SLAVE state) while the owning NAPI thread is in the middle of
a poll on another CPU, both sides mutate the GRO list concurrently, as
shown below:

  [39.846] port 1 (swp1): MASTER to UNCALIBRATED on RS_SLAVE
  list_add corruption. next->prev should be prev (ffff8d620faf4138), but was ffff8d624150f700. (next=ffff8d620faf4138).
  kernel BUG at lib/list_debug.c:29!
  Oops: invalid opcode: 0000 [#1] SMP PTI
  CPU: 1 UID: 0 PID: 539 Comm: napi/mlxsw_rx-0 Not tainted 6.18.48 #1-NixOS PREEMPT(lazy)
  Hardware name: Mellanox Technologies Ltd. MSN2410/VMOD0001, BIOS 4.6.5 09/13/2018
  RIP: 0010:__list_add_valid_or_report+0x79/0xb0
  RSP: 0018:ffffcdf8c0f27c08 EFLAGS: 00010246
  RAX: 0000000000000075 RBX: ffff8d624150fd00 RCX: 0000000000000000
  RDX: 0000000000000000 RSI: 0000000000000001 RDI: ffff8d6315d1e540
  RBP: ffff8d620faf4070 R08: 0000000000000000 R09: 00000000ffffdfff
  R10: ffffffffa5c60fe0 R11: ffffcdf8c0f27ab8 R12: 0000000000000003
  R13: 000000000000003d R14: 00000000000001bc R15: 0000000000000001
  FS:  0000000000000000(0000) GS:ffff8d636f63f000(0000) knlGS:0000000000000000
  CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
  CR2: 0000562689a60c24 CR3: 000000015f224004 CR4: 00000000001726f0
  Call Trace:
  <TASK>
  gro_receive_skb+0xee/0x230
  mlxsw_sp1_ptp_got_packet+0x61/0x140 [mlxsw_spectrum]
  mlxsw_core_skb_receive+0xdf/0x1b0 [mlxsw_core]
  mlxsw_pci_napi_poll_cq_rx+0x780/0x9d0 [mlxsw_pci]
  __napi_poll+0x31/0x1e0
  napi_threaded_poll_loop+0x16b/0x1c0
  napi_threaded_poll+0x71/0xa0
  kthread+0xfb/0x260
  ret_from_fork+0x22d/0x260
  ret_from_fork_asm+0x1a/0x30
  </TASK>
  Kernel panic - not syncing: Fatal exception in interrupt

The machinery that leads to this kernel panic has not been changed between
6.18.48 and mainline.

This patch adds an ingress-delivery helper for the PTP packet_finish()
path that calls netif_receive_skb() instead of napi_gro_receive().
netif_receive_skb(), unlike napi_gro_receive(), can be called from outside
of the NAPI instance's poll context, which can occur at the call site for
this path. RX stats accounting and the skb->dev assignment are still
preserved; the only change is the delivery call itself.

This removes GRO batching for any PTP event traffic received by the mlxsw
trap, but given the relatively low volume of traffic characteristic of the
protocol, and impact limited to only Spectrum-1 ASICs, this is an
acceptable solution.

Fixes: 1ba06ca96c ("mlxsw: Switch to napi_gro_receive()")
Signed-off-by: Taylor Bates <tmbates12@gmail.com>
Reviewed-by: Petr Machata <petrm@nvidia.com>
Link: https://patch.msgid.link/20260902024949.2273997-1-tmbates12@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 17:01:03 -07:00
Zihan Xi
efdfb1e27a ipv4: fib: bound automatic table ID allocation
fib_empty_table() probes every table ID from 1 until it finds a
free one.  IPv4 tables are stored in a 256-bucket hash table, so a
dense set of IDs makes each probe walk a growing hash chain while
RTNL is held.

Automatic table assignment ("ip rule ... table 0") is an IPv4-only
legacy path.  Bound the automatically allocated ID to 4096 so the
RTNL hold stays bounded, without changing lookups of explicitly
specified table IDs.

This changes user-visible behavior.  A table-0 rule previously
received the lowest free ID in 1..RT_TABLE_MAX (0xFFFFFFFF).  After
this patch the search stops at 4096 and the rule add fails with
ENOBUFS if that range is fully occupied.  Explicit table IDs above
4096 remain usable.

The automatic path is unused in practice: it is IPv4-only, not
documented by ip-rule, uncovered by kernel selftests, and both
NetworkManager and systemd refuse table 0.

Fixes: b801f54917 ("[NET]: Increate RT_TABLE_MAX to 2^32")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Suggested-by: Ido Schimmel <idosch@nvidia.com>
Signed-off-by: Zihan Xi <zihanx@nebusec.ai>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Petr Vorel <pvorel@suse.cz>
Link: https://patch.msgid.link/6f2f2a7a136aee005512a2e1ac8ede62ac8c7bb6.1788258884.git.zihanx@nebusec.ai
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 16:58:32 -07:00
Jakub Kicinski
b6dd676d32 Merge branch 'net-bcmasp-fix-tx-ring-accounting-bugs'
Danesh Petigara says:

====================
net: bcmasp: fix TX ring accounting bugs

Two fixes for TX descriptor ring handling in the bcmasp driver:

  - tx_spb_ring_full() re-initialized next_index from
    intf->tx_spb_index on every loop iteration instead of advancing
    it, so it only ever checked a single descriptor slot regardless
    of cnt. This let bcmasp_xmit() proceed even when the ring didn't
    actually have enough free slots for the SKB's fragments.

  - bcmasp_xmit() only set txcb->last for the final fragment of an
    SKB, leaving stale true values in reused descriptor slots from a
    prior transmission. Combined with the ring-full miscount above,
    this could cause bcmasp_tx_reclaim() to treat a mid-SKB
    descriptor as the last one and free the sk_buff while later
    fragments were still in flight.

Patch 1 clears txcb->last unconditionally before it is set, and
patch 2 fixes the ring-full slot check to advance through each
candidate slot.
====================

Link: https://patch.msgid.link/20260831184235.4133351-1-danesh.petigara@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 16:55:02 -07:00
Justin Chen
0c5cf62e72 net: bcmasp: fix tx_spb_ring_full() checking same slot cnt times
The loop initialised next_index from intf->tx_spb_index on every
iteration, so incr_ring() always produced the same result and only
one slot was ever tested.  Move the initialisation before the loop
so each iteration advances next_index and the function correctly
checks that cnt consecutive descriptor slots are available before
allowing a new transmission.

Fixes: 490cb41200 ("net: bcmasp: Add support for ASP2.0 Ethernet controller")
Signed-off-by: Justin Chen <justin.chen@broadcom.com>
Signed-off-by: Danesh Petigara <danesh.petigara@broadcom.com>
Reviewed-by: Florian Fainelli <florian.fainelli@broadcom.com>
Link: https://patch.msgid.link/20260831184235.4133351-3-danesh.petigara@broadcom.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-03 16:54:57 -07:00