Commit Graph

1481724 Commits

Author SHA1 Message Date
Kuniyuki Iwashima
ca0b0a8687 selftest: af_unix: Add zero-buffer test for msg_oob.c
The previous patches fixed two issues related to zero-length
buffer with MSG_PEEK for MSG_OOB skb.

Let's add corresponding tests in msg_oob.c.

Without this series:

  # FAILED: 50 / 60 tests passed.
  # Totals: pass:50 fail:10 xfail:0 xpass:0 skip:0 error:0

With this series:

  # PASSED: 60 / 60 tests passed.
  # Totals: pass:60 fail:0 xfail:0 xpass:0 skip:0 error:0

Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260902202202.892676-4-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-07 16:58:42 -07:00
Kuniyuki Iwashima
6e5ee08eb5 af_unix: Return immediately when manage_oob() returns NULL for 0-length buffer.
Fahad Alharbi reported that recv(0, MSG_PEEK) triggers busy-wait
in unix_stream_read_generic() if recv() is blocking and the last
skb in the queue is MSG_OOB skb.

In such a situation, TCP returns 0 immediately regardless of
blocking or non-blocking.

Let's follow the behaviour.

Fixes: 314001f0bf ("af_unix: Add OOB support")
Reported-by: Fahad Alharbi <fahad@codepure.com>
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260902202202.892676-3-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-07 16:58:42 -07:00
Kuniyuki Iwashima
94fd4debd2 af_unix: Update last skb marker in manage_oob().
Fahad Alharbi reported that blocking recv(MSG_PEEK) could hog CPU
due to OOB skb.

In the following cases, manage_oob() skips OOB skb(s) and returns
NULL for the last recv(MSG_PEEK):

  socketpair(AF_UNIX, SOCK_STREAM, 0, sk);

  1) skb -> OOB skb -> NULL
     send(sk[0], "ab", 2, MSG_OOB);
     recv(sk[1], buf, 0, MSG_PEEK);

  2) skb -> consumed OOB skb -> NULL
     send(sk[0], "ab", 2, MSG_OOB);
     recv(sk[1], buf, 1, MSG_OOB);
     recv(sk[1], buf, 0, MSG_PEEK);

  3) consumed OOB skb -> OOB skb -> NULL
     send(sk[0], "a", 1, MSG_OOB);
     recv(sk[1], buf, 0, MSG_OOB);
     send(sk[0], "b", 1, MSG_OOB);
     recv(sk[1], buf, 1, MSG_PEEK);

Then, @copied is 0 in unix_stream_read_generic() (zero-length buffer,
or non-OOB skb is not yet consumed), and unix_stream_data_wait() is
called.

However, it returns immediately because @last is not updated in
unix_stream_read_generic(), and the thread busy-waits for a new skb.

Let's update @last in manage_oob().

For MSG_PEEK, @last is updated with the skipped OOB, and for the
non-peek case, @last matches the returned value (when !copied)
because OOB is unlinked.

Note that manage_oob() is inlined and no stack canary is added.

Fixes: 22dd70eb2c ("af_unix: Don't peek OOB data without MSG_OOB.")
Reported-by: Fahad Alharbi <fahad@codepure.com>
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260902202202.892676-2-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-07 16:58:42 -07:00
Nagamani PV
74f27fc864 s390/qeth: allow bridgeport queries despite OS_MISMATCH
When HiperSockets interfaces on the same VCHID span different OS
families, reads of the sysfs attributes bridge_role and bridge_state
fail with -EPERM if bridge port ownership belongs to another OS family.

As a result, userspace tools such as 'lszdev -ii' cannot retrieve
bridge_role and bridge_state, even though firmware returns valid bridge
port data for QUERY_BRIDGE_PORTS requests.

The firmware reports IPA_RC_SBP_IQD_OS_MISMATCH (0x0010) to indicate
that bridge port ownership belongs to a different OS family. For
QUERY_BRIDGE_PORTS operations, firmware still returns valid bridge port
data (role=none, state=inactive) together with a primary return code of
0x0000 (success).

Allow QUERY_BRIDGE_PORTS requests to return the bridge port data
provided by the firmware despite OS_MISMATCH. To make the OS family
mismatch visible to userspace, represent the firmware-reported role
"none" as "none (OS family mismatch)" while preserving the reported
bridge_state.

The behavior for non-QUERY bridge port commands is unchanged; SET
operations continue to return -EPERM when another OS family owns the
bridge port.

This restores readability of bridge_role and bridge_state.

Fixes: 1b05cf6285 ("qeth: Include error message for "OS Mismatch"")
Cc: stable@vger.kernel.org
Suggested-by: Halil Pasic <pasic@linux.ibm.com>
Reviewed-by: Alexandra Winter <wintera@linux.ibm.com>
Signed-off-by: Nagamani PV <nagamani@linux.ibm.com>
Link: https://patch.msgid.link/20260901155344.3561483-1-nagamani@linux.ibm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-07 16:39:26 -07:00
Vineeth Karumanchi
38b6be1010 net: macb: fix NULL pointer dereference on unbind with fixed-link
When the device tree describes a fixed-link and has no "mdio" child
node, macb_mii_init() returns early without allocating the MDIO bus,
leaving bp->mii_bus as NULL.

Two cleanup paths then dereference this NULL bus:

1. On driver unbind, macb_remove() unconditionally calls
   mdiobus_unregister(bp->mii_bus), which oopses:

  Unable to handle kernel NULL pointer dereference at virtual address 00000000000004a8
  pc : mdiobus_unregister+0x14/0xa4
  lr : macb_remove+0x38/0xa4
  Call trace:
   mdiobus_unregister+0x14/0xa4 (P)
   macb_remove+0x38/0xa4
   platform_remove+0x20/0x30
   device_release_driver_internal+0x1c8/0x224
   unbind_store+0xb4/0xbc

2. On the probe error path in macb_probe(), reached when
   macb_mii_init() has succeeded but a subsequent step fails, the
   err_out_unregister_mdio label runs the same unconditional cleanup.

mdiobus_unregister() and mdiobus_free() do not guard against a NULL
bus, so guard the calls in both macb_remove() and the probe error
path.

Fixes: d0c3601f2c ("net: macb: Avoid 20s boot delay by skipping MDIO bus registration for fixed-link PHY")
Signed-off-by: Vineeth Karumanchi <vineeth.karumanchi@amd.com>
Reviewed-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Link: https://patch.msgid.link/20260902102836.2019355-1-vineeth.karumanchi@amd.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:58:52 -07:00
Jakub Kicinski
e7c93ad4bd Merge branch 'net-sched-clamp-quantum-psched_mtu-in-change-paths'
Jamal Hadi Salim says:

====================
net/sched: clamp quantum/psched_mtu in change paths

This is a followup to commit 709f34f7c2 ("net/sched: fq: add overflow
bounds to quantum and initial quantum").

The quantum_backlog_overflow series and the five siblings that followed
clamped the init-path quantum in fq, fq_codel, fq_pie, hhf, sfq. The
change() paths were not clamped but it is the same pattern, same writer
of q->quantum, same privilege level (CAP_NET_ADMIN in a user namespace).
A user can override the init clamp via tc qdisc change, restoring the
small-quantum deficit spin that the init clamp was meant to prevent.

This series also covers two siblings that were missed entirely by the
original series: sch_dualpi2 and sch_pie call psched_mtu() without any
clamp at all. With a crafted size table qdisc_pkt_len reaches ~2 GiB,
so quantum=1 (or a zero psched_mtu on a headerless device) makes the
deficit-refill loop spin ~2^31 times under the qdisc lock (a soft
lockup / denial of service).

Each patch fixes one qdisc with its own Fixes: tag so they can be
backported independently - the commits they fix shift differently in
the git tree.

Patch 1: fq - clamp TCA_FQ_QUANTUM and TCA_FQ_INITIAL_QUANTUM in change
Patch 2: fq_pie - clamp quantum in change path
Patch 3: sfq - clamp quantum and reject > 1<<20 in change path
Patch 4: hhf - clamp quantum in change and init paths
Patch 5: dualpi2 - clamp psched_mtu at all 3 call sites
Patch 6: pie - clamp psched_mtu in pie_drop_early
Patch 7: drr - clamp quantum in change class
Patch 8: ets - clamp quantum in parse and fallback paths
Patch 9: selftests - update ETS test 41f5 for clamped quanta

Conditions to recreate (applies to all): create the qdisc, then
tc qdisc change ... quantum 1 with a STAB size table inflating
qdisc_pkt_len. Requires CAP_NET_ADMIN in a user namespace (unshare -Urn).
====================

Link: https://patch.msgid.link/QDISC-0CFC.v3.20260901204856@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:48:08 -07:00
Jamal Hadi Salim
8f0229bef3 selftests: tc-testing: update ETS test 41f5 for clamped quanta
Commit "net/sched: ets: clamp quantum in parse and fallback paths"
moved the quantum floor into ets_quantum_parse(), so every explicitly
configured quantum is now clamped to [256, 1 << 20], not just the
psched_mtu() fallback.

Test 41f5 passes "quanta 4294967294 1 1" and matches the values back
verbatim, so all three bands now differ from what it expects:

  before: bands 3 quanta 4294967294 1 1
  after:  bands 3 quanta 1048576 256 256

Update the match pattern accordingly.

Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-0CFC.v3.20260901204856@mojatatu.com.10
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:48:08 -07:00
Jamal Hadi Salim
1c38487f46 net/sched: ets: clamp quantum in parse and fallback paths
ets_qdisc_change() falls back to psched_mtu() with no floor for bands
without an explicit quantum. With a crafted size table qdisc_pkt_len
reaches ~2 GiB, so a zero psched_mtu on a headerless device makes the
deficit-refill loop spin under the qdisc lock.

Move the floor into ets_quantum_parse() so explicitly configured quanta
are also clamped to [256, 1<<20], not just the fallback path.

Conditions to recreate the bug:
  CONFIG_NET_SCH_ETS=y. Requires CAP_NET_ADMIN (namespace-local via
  unshare -Urn suffices).

  tc qdisc add dev dummy0 root ets bands 3 strict 2 quanta 1 1

Fixes: dcc68b4d80 ("net: sch_ets: Add a new Qdisc")
Reported-by: Vega <vega@nebusec.ai>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-0CFC.v3.20260901204856@mojatatu.com.9
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:48:08 -07:00
Jamal Hadi Salim
8382abec0f net/sched: drr: clamp quantum in change class
drr_change_class() rejects explicit quantum==0 but falls back to
psched_mtu() with no floor. With a crafted size table qdisc_pkt_len
reaches ~2 GiB, so quantum=1 (or a zero psched_mtu on a headerless
device) makes the deficit-refill loop spin under the qdisc lock.

Add clamp_t(u32, quantum, 256, 1<<20) after the zero reject and on the
fallback path. The explicit-zero reject is preserved.

Conditions to recreate the bug:
  CONFIG_NET_SCH_DRR=y. Requires CAP_NET_ADMIN (namespace-local via
  unshare -Urn suffices).

  tc qdisc add dev dummy0 root drr
  tc class add dev dummy0 parent 1: classid 1:1 drr quantum 1

Fixes: 13d2a1d2b0 ("pkt_sched: add DRR scheduler")
Reported-by: Vega <vega@nebusec.ai>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-0CFC.v3.20260901204856@mojatatu.com.8
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:48:08 -07:00
Jamal Hadi Salim
54370e44c0 net/sched: pie: clamp psched_mtu in pie_drop_early
pie_drop_early() calls psched_mtu() with no clamp. With mtu=0x80000000
the bytemode divide silently zeroes the drop probability, disabling AQM.
Clamp to [1, 1<<20].

Conditions to recreate the bug:
  CONFIG_NET_SCH_PIE=y. Requires CAP_NET_ADMIN (namespace-local via
  unshare -Urn suffices).

  tc qdisc add dev dummy0 root pie
  tc qdisc change dev dummy0 root pie stab data 32768 size_log 15 cell_log 0

Fixes: d4b36210c2 ("net: pkt_sched: PIE AQM scheme")
Reported-by: Vega <vega@nebusec.ai>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-0CFC.v3.20260901204856@mojatatu.com.7
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:48:08 -07:00
Jamal Hadi Salim
3c01f1ca5d net/sched: dualpi2: clamp psched_mtu at all call sites
dualpi2_calculate_c_protection(), must_drop(), and get_memory_limit()
call psched_mtu() with no clamp. A huge MTU makes (s32)psched_mtu()
overflow in the signed multiply for c_protection_init, and 2 *
psched_mtu() wraps in get_memory_limit(). With a crafted size table
qdisc_pkt_len reaches ~2 GiB, causing a soft lockup / denial of service.

Clamp psched_mtu() to [1, 1<<20] at all three call sites.

Conditions to recreate the bug:
  CONFIG_NET_SCH_DUALPI2=y. Requires CAP_NET_ADMIN (namespace-local via
  unshare -Urn suffices).

  tc qdisc add dev dummy0 root dualpi2
  tc qdisc change dev dummy0 root dualpi2 stab data 32768 size_log 15 cell_log 0

Fixes: 320d031ad6 ("sched: Struct definition and parsing of dualpi2 qdisc")
Reported-by: Vega <vega@nebusec.ai>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-0CFC.v3.20260901204856@mojatatu.com.6
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:48:02 -07:00
Jamal Hadi Salim
eb56a495f5 net/sched: hhf: clamp quantum in change and init paths
hhf_change() accepts any quantum from userspace, including 1. With a
crafted size table qdisc_pkt_len reaches ~2 GiB, so quantum=1 makes
the deficit-refill loop spin ~2^31 times under the qdisc lock
(a soft lockup / denial of service).

Add max(256U, ...) in hhf_change() matching fq_codel_change(). Clamp
hhf_init() to [256, 1<<20] matching the siblings, and remove the old
fallback that only set quantum=256 on overflow.

Conditions to recreate the bug:
  CONFIG_NET_SCH_HHF=y. Requires CAP_NET_ADMIN (namespace-local via
  unshare -Urn suffices).

  tc qdisc add dev dummy0 root hhf
  tc qdisc change dev dummy0 root hhf quantum 1 stab data 32768 size_log 15 cell_log 0

Fixes: 10239edf86 ("net-qdisc-hhf: Heavy-Hitter Filter (HHF) qdisc")
Reported-by: Vega <vega@nebusec.ai>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-0CFC.v3.20260901204856@mojatatu.com.5
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:47:30 -07:00
Jamal Hadi Salim
fb9f88a33c net/sched: sfq: clamp quantum in change path
sfq_change() accepts any non-negative quantum (only rejects
(int)ctl->quantum < 0). With a crafted size table qdisc_pkt_len reaches
~2 GiB, so quantum=1 makes the deficit-refill loop spin ~2^31 times
under the qdisc lock (a soft lockup / denial of service).

Add max(256U, ...) matching fq_codel_change(). Reject quantum > 1<<20
with -EINVAL, matching fq_codel_change() and the init clamp.

Conditions to recreate the bug:
  CONFIG_NET_SCH_SFQ=y. Requires CAP_NET_ADMIN (namespace-local via
  unshare -Urn suffices).

  tc qdisc add dev dummy0 root sfq
  tc qdisc change dev dummy0 root sfq quantum 1 stab data 32768 size_log 15 cell_log 0

Fixes: e4650d7ae4 ("net_sched: sch_sfq: handle bigger packets")
Reported-by: Vega <vega@nebusec.ai>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-0CFC.v3.20260901204856@mojatatu.com.4
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:47:30 -07:00
Jamal Hadi Salim
4864f58c53 net/sched: fq_pie: clamp quantum in change path
fq_pie_change() accepts any quantum value from userspace, including 1.
With a crafted size table qdisc_pkt_len reaches ~2 GiB, so quantum=1
makes the deficit-refill loop spin ~2^31 times under the qdisc lock
(a soft lockup / denial of service).

Add max(256U, ...) matching fq_codel_change().

Conditions to recreate the bug:
  CONFIG_NET_SCH_FQ_PIE=y. Requires CAP_NET_ADMIN (namespace-local via
  unshare -Urn suffices).

  tc qdisc add dev dummy0 root fq_pie
  tc qdisc change dev dummy0 root fq_pie quantum 1 stab data 32768 size_log 15 cell_log 0

Fixes: ec97ecf1eb ("net: sched: add Flow Queue PIE packet scheduler")
Reported-by: Vega <vega@nebusec.ai>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-0CFC.v3.20260901204856@mojatatu.com.3
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:47:29 -07:00
Jamal Hadi Salim
094cc07f98 net/sched: fq: clamp quantum and initial_quantum in change path
The fq change path accepts TCA_FQ_QUANTUM in [1, INT_MAX] and
TCA_FQ_INITIAL_QUANTUM up to INT_MAX, while fq_init() already clamps to
[1, 1<<20]. A user can override the init clamp via tc qdisc change,
restoring the small-quantum deficit spin that the init clamp prevents.

Narrow iq_range.max to 1<<20 so TCA_FQ_INITIAL_QUANTUM is rejected at
parse time. Clamp TCA_FQ_QUANTUM to [256, 1<<20] in fq_change() and
fq_init() quantum to [256, 1<<20] for tiny-MTU devices.

Conditions to recreate the bug:
  CONFIG_NET_SCH_FQ=y. Requires CAP_NET_ADMIN (namespace-local via
  unshare -Urn suffices).

  tc qdisc add dev dummy0 root fq
  tc qdisc change dev dummy0 root fq quantum 1 stab data 32768 size_log 15 cell_log 0

Fixes: 709f34f7c2 ("net/sched: fq: add overflow bounds to quantum and initial quantum")
Reported-by: Vega <vega@nebusec.ai>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-0CFC.v3.20260901204856@mojatatu.com.2
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:47:29 -07:00
Carolina Jubran
df99553f84 net/mlx5e: Keep HW timestamp stats monotonic across reconfiguration
`mlx5e_stats_ts_get()` currently selects either DMA or port timestamp
counters based on `tx_ptp_opened`. This flag is intentionally kept set
once the PTP TX queues have been opened so their statistics remain
available after queue teardown. As a result, DMA timestamps are no
longer reported after switching from port timestamping back to DMA
timestamping.

The function also reads statistics only from the currently active
channels and TCs. Reducing the number of channels or TCs can therefore
drop previously accumulated timestamp counters from the reported value.

Read the persistent channel statistics instead and always include DMA
timestamp counters. Once the PTP TX queues have been opened, also
include the port timestamp counters.

This also drops state_lock. It previously protected live channel/PTP
pointers, the new code only reads persistent channel_stats and
ptp_stats via mlx5e_stats_nch_read(), which is already safe for
lockless stats access.

Fixes: 3579032c08 ("net/mlx5e: Implement ethtool hardware timestamping statistics")
Signed-off-by: Carolina Jubran <cjubran@nvidia.com>
Reviewed-by: Shahar Shitrit <shshitrit@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902193731.3668958-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:23:37 -07:00
Lama Kayal
c0c6f4ba8a net/mlx5: E-Switch, prevent mc_list repopulation during vport disable
In mlx5_esw_vport_disable(), move esw_apply_vport_rx_mode() ahead
of esw_vport_change_handle_locked() so vport->allmulti_rule is
NULL before the change handler observes it.

During FW-fatal recovery the disable runs while dev->state ==
INTERNAL_ERROR. The promisc query inside esw_update_vport_rx_mode()
fails and returns early, leaving vport->allmulti_rule intact, so
esw_update_vport_mc_promisc() runs and adds MLX5_ACTION_ADD entries
to vport->mc_list whose flow rules are then installed in the FDB
by esw_add_mc_addr(). esw_destroy_legacy_table() tears down the
FDB with those refs still held, corrupting the sub-tree and
leaving dangling flow_rule pointers in vport->mc_list.

Two-stage failure on `echo 1 > /sys/bus/pci/devices/<bdf>/reset`:

  refcount_t: underflow; use-after-free.
   tree_put_node+0xef/0x110 [mlx5_core]
   clean_tree+0x44/0xd0 [mlx5_core] (x5)
   mlx5_fs_core_cleanup+0x57/0x1c0 [mlx5_core]
   mlx5_unload+0x65/0xd0 [mlx5_core]
   ... mlx5_health_try_recover

  BUG: unable to handle page fault for address: 0000000003000055
   down_write+0x1c/0x60
   mlx5_del_flow_rules+0x33/0x1f0 [mlx5_core]
   esw_del_mc_addr+0x7b/0x170 [mlx5_core]
   esw_apply_vport_addr_list+0x56/0xf0 [mlx5_core]
   esw_vport_change_handle_locked+0x28b/0x310 [mlx5_core]
   mlx5_esw_vport_enable+0x270/0x4a0 [mlx5_core]
   ... mlx5_load ... mlx5_health_try_recover

esw_apply_vport_rx_mode(false, false) clears vport->allmulti_rule
via its local state machine even when the FW del fails. With the
rule NULL the !IS_ERR_OR_NULL(allmulti_rule) gate in the change
handler closes, no rules are installed during disable, and the
reload starts with a clean mc_list.

Fixes: 922f56e9a7 ("net/mlx5: Fix steering rules cleanup")
Signed-off-by: Lama Kayal <lkayal@nvidia.com>
Reviewed-by: Cosmin Ratiu <cratiu@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902193854.3669035-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:21:43 -07:00
Yael Chemla
7ee07f601f net/mlx5: E-Switch: fix use-after-free in mlx5_eswitch_termtbl_put
In mlx5_eswitch_termtbl_put(), the zero-ref cleanup check reads
tt->ref_count after termtbl_mutex has been released.  Two concurrent
callers on the same mlx5_termtbl_handle race: one decrements ref_count
to zero, removes the hash entry, and calls kfree(tt) while the other
has already dropped the mutex and is about to evaluate
if (!tt->ref_count), producing a use-after-free.

Fix this by capturing the result of the decrement into a stack-local
last variable before dropping the mutex.  The cleanup decision is now
made entirely under termtbl_mutex, and tt is not touched after
kfree.

Fixes: 10caabdaad ("net/mlx5e: Use termination table for VLAN push actions")
Signed-off-by: Yael Chemla <ychemla@nvidia.com>
Reviewed-by: Dragos Tatulea <dtatulea@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902193514.3668880-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:20:26 -07:00
Carolina Jubran
af3aef0245 net/mlx5e: Fix use-after-free race in sample_restore_put()
Concurrent teardown of TC sample rules sharing the same restore
context may re-read restore->count after dropping restore_lock.
At that point another thread may already have completed cleanup and
freed the restore object.

Use the result of the refcount decrement while holding restore_lock to
determine whether cleanup is needed.

Fixes: 36a3196256 ("net/mlx5e: TC, Add sampler restore handle API")
Signed-off-by: Carolina Jubran <cjubran@nvidia.com>
Reviewed-by: Shahar Shitrit <shshitrit@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902193341.3668809-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:19:22 -07:00
Carolina Jubran
e7ee897408 net/mlx5e: Fix ETS zero BW reporting when one TC holds 100%
When ETS TCs with zero bandwidth are configured, the driver programs the
firmware using an alternate representation. On get, it needs
to recognize that representation so those TCs can be translated back and
reported as 0% bandwidth.

The existing detection relied on the programmed bandwidth because it was
enough to identify this representation. However, when a single ETS TC
owns 100% of the bandwidth, its firmware representation becomes the
same as a strict-priority TC, causing zero-bandwidth ETS TCs to be
reported with non-zero bandwidth values.

Use the cached TSA instead to distinguish the ETS and strict-priority
cases.

Fixes: be0f161ef1 ("net/mlx5e: DCBNL, Implement tc with ets type and zero bandwidth")
Signed-off-by: Carolina Jubran <cjubran@nvidia.com>
Reviewed-by: Alex Lazar <alazar@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902193224.3668743-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:18:47 -07:00
Akiva Goldberger
b3c79dee50 net/mlx5: LAG, use local tracker to update active ports
The CREATE_LAG command is handled asynchronously by queuing a work,
which stores a local copy of ldev->tracker. When the work is processed,
it is possible that the values of the local copy and ldev->tracker have
diverged.

A single CREATE_LAG command programs two related fields into the
firmware: the v2p (virtual-to-physical) map, which selects the physical
egress port for each hash bucket, and the active_port bitmask, which
tells the firmware which physical ports are currently up so it can
redirect QP/TIS away from inactive ports. For the firmware to steer
traffic correctly, both must be derived from the same view of the ports'
link state.

The v2p map is computed by mlx5_infer_tx_affinity_mapping() from the
local tracker snapshot, but lag_active_port_bits() called
mlx5_infer_tx_enabled() on the live ldev->tracker instead. If
ldev->tracker changed between the snapshot and command execution, the
two fields reflect different port states: the v2p map may steer a bucket
to a port that the active_port mask marks as inactive (or vice versa).
The firmware then receives a self-contradictory configuration and can
redirect or drop traffic on a port the mapping still points at, until a
later event happens to reconcile the state.

Update lag_active_port_bits so that it receives the local version of the
tracker from when the work was queued, effectively closing the window
for injecting an inconsistency.

Fixes: c5c13b456c ("net/mlx5: Lag, set active ports if support bypass port select flow table")
Signed-off-by: Akiva Goldberger <agoldberger@nvidia.com>
Reviewed-by: Shay Drori <shayd@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902192740.3665435-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:16:01 -07:00
Jakub Kicinski
502381cf73 Merge branch 'net-mlx5e-rs-fec-variant-fixes'
Tariq Toukan says:

====================
net/mlx5e: RS FEC variant fixes

This series by Shahar fixes three related bugs in the RS FEC handling
for mlx5e, all stemming from incomplete coverage of the three RS FEC
hardware variants: RS_528_514 (bit 2), RS_544_514_INTERLEAVED_QUAD
(bit 4), and RS_544_514 (bit 7).
====================

Link: https://patch.msgid.link/20260902164634.3657606-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:14:51 -07:00
Shahar Shitrit
c84ce45a7a net/mlx5e: Fix reporting support for all RS FEC variants
get_fec_supported_advertised() populates the FEC modes reported as
supported to userspace. The MLX5E_ADVERTISE_SUPPORTED_FEC macro only
checked MLX5E_FEC_RS_528_514, causing devices that support only the
other RS variants (RS_544_514_INTERLEAVED_QUAD or RS_544_514) to not
advertise RS as supported to ethtool at all.

Introduce MLX5E_FEC_RS_MASK covering all three RS bit positions,
update the macro to accept a bitmask directly rather than a single
enum value, and pass MLX5E_FEC_RS_MASK for the RS entry.

Fixes: b5ede32d33 ("net/mlx5e: Add support for FEC modes based on 50G per lane links")
Fixes: 4e343c11ef ("net/mlx5e: Support FEC settings for 200G per lane link modes")
Signed-off-by: Shahar Shitrit <shshitrit@nvidia.com>
Reviewed-by: Dragos Tatulea <dtatulea@nvidia.com>
Reviewed-by: Yael Chemla <ychemla@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902164634.3657606-4-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:14:48 -07:00
Shahar Shitrit
b9d755c5a3 net/mlx5e: Fix setting RS FEC after remapping
When a user sets a FEC mode via ethtool, the driver maps the ethtool
FEC type to the lowest mlx5 bit of that type. For RS FEC, this is
MLX5E_FEC_RS_528_514 (bit 2). The driver then checks whether this
bit is supported by at least one link mode by inspecting the
fec_override_cap fields via mlx5e_fec_in_caps(), and returns
-EOPNOTSUPP if not.

This check is incorrect. RS FEC has three supported hardware variants:
RS_528_514 (bit 2), RS_544_514_INTERLEAVED_QUAD (bit 4), and
RS_544_514 (bit 7). mlx5e_remap_fec_conf_mode() already remaps bit 2
to the appropriate RS variant per link mode when writing the admin
fields, but the early capability check is done against the raw
unmapped bit. As a result, a device that supports RS_544_514 or
RS_544_514_INTERLEAVED_QUAD but not RS_528_514 will incorrectly reject
the user's RS FEC request.

Remove the early support check from mlx5e_set_fec_mode() and fold
it into the existing write loop, checking caps against the remapped
policy per link mode. Return -EOPNOTSUPP before the final register
write if no link mode accepted the policy.

Fixes: 2608a2f831 ("net/mlx5e: Fix return status when setting unsupported FEC mode")
Signed-off-by: Shahar Shitrit <shshitrit@nvidia.com>
Reviewed-by: Dragos Tatulea <dtatulea@nvidia.com>
Reviewed-by: Yael Chemla <ychemla@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902164634.3657606-3-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:14:48 -07:00
Shahar Shitrit
802eedcc0b net/mlx5e: Fix missing FEC mode mapping for RS_544_514_INTERLEAVED_QUAD
MLX5E_FEC_RS_544_514_INTERLEAVED_QUAD is missing from
pplm_fec_2_ethtool_linkmodes[], leaving index 4 zero-initialized.
As a result, when this FEC mode is active, find_first_bit() returns
index 4, causing __set_bit() to set bit 0
(ETHTOOL_LINK_MODE_10baseT_Half_BIT) instead of
ETHTOOL_LINK_MODE_FEC_RS_BIT. Consequently, ethtool reports:

  Advertised FEC modes: Not reported

Add the missing mapping to ETHTOOL_LINK_MODE_FEC_RS_BIT.

Fixes: 4e343c11ef ("net/mlx5e: Support FEC settings for 200G per lane link modes")
Signed-off-by: Shahar Shitrit <shshitrit@nvidia.com>
Reviewed-by: Dragos Tatulea <dtatulea@nvidia.com>
Reviewed-by: Yael Chemla <ychemla@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902164634.3657606-2-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:14:48 -07:00
Weiming Shi
e6662f2100 net/sched: defer qdisc freeing after failed creation
An RTM_NEWQDISC request can make clsact bind a populated shared ingress
block during ->init(), publishing an embedded mini_Qdisc to lockless
readers.  If the same request has an invalid TCA_RATE, estimator setup
fails after ->init(); the unwind removes the pointer but synchronously
frees its containing qdisc while tc_run() may still hold it.

Retire failed qdiscs through the same RCU helper as normal destruction.
Inline the synchronous free into the callback now that no direct callers
remain.

Fixes: 51ab2994c3 ("net: sched: allow ingress and clsact qdiscs to share filter blocks")
Reported-by: Xiang Mei <xmei5@asu.edu>
Link: https://lore.kernel.org/netdev/20260805102505.740806-1-david.lee@trailofbits.com/
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Link: https://patch.msgid.link/20260902155231.2149915-2-bestswngs@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 12:40:05 -07:00
XingWang Xiang
2b4707a149 net: mctp: i3c: serialize probe with bus removal
mctp_i3c_probe() drops busdevs_lock after finding the matching bus. A
concurrent I3C_NOTIFY_BUS_REMOVE can then unregister and free the bus
netdev before probe passes its private data to mctp_i3c_add_device().
The latter consequently adds a list node through a freed mbus pointer.

Keep busdevs_lock held until the device has been added. This also
satisfies the __must_hold annotation on mctp_i3c_add_device().

Fixes: c8755b29b5 ("mctp i3c: MCTP I3C driver")
Signed-off-by: XingWang Xiang <v3rdant.xiang@gmail.com>
Acked-by: Matt Johnston <matt@codeconstruct.com.au>
Signed-off-by: David S. Miller <davem@davemloft.net>
2026-09-05 09:39:00 +01:00
Sahil Chandna
80dd7e754b net: mana: Reserve extra CQ slot for the fence completion CQE
The RX completion queue is sized to hold exactly one CQE per posted RX WQE.
MANA_FENCE_RQ makes hardware post an additional CQE_RX_OBJECT_FENCE after
the packet CQEs. The current sizing reserves no extra slot for it and in
rare cases, CQ has no guaranteed slot for the fence CQE when it is full of
packet CQEs. This can lead to dropping the fence completion while the
driver waits holding RTNL lock throughout the timeout duration.
Reserve one extra CQE slot for CQE_RX_OBJECT_FENCE. mana_gd_alloc_memory()
requires queue_size to be a power-of-two and at least MANA_PAGE_SIZE;
the reservation pushes cq_size past a power-of-two, so round up the CQ size
in mana_create_rxq().

Cc: stable@vger.kernel.org
Fixes: 6cc74443a7 ("net: mana: Add RX fencing")
Signed-off-by: Sahil Chandna <sahilchandna@linux.microsoft.com>
Reviewed-by: Haiyang Zhang <haiyangz@microsoft.com>
Link: https://patch.msgid.link/20260901121837.3503240-1-sahilchandna@linux.microsoft.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 18:36:38 -07:00
Jakub Kicinski
db88216424 Merge branch 'pds_core-fixes-for-the-pci-reset-path'
Nikhil P. Rao says:

====================
pds_core: fixes for the PCI reset path [part]

Patch 1 is the v2 patch with the pdsc_core_init() and
pdsc_identify_ver() checks removed. Those checks are dead code. commit
cd09971dcc ("pds_core: keep the health thread stopped during reset")
disables health_work across the reset, so the health thread can no
longer reach pdsc_setup() with the BARs unmapped. The only other callers
are probe and pdsc_reset_done(), and both map the BARs earlier in the
same call, so cmd_regs cannot be NULL by the time they get there.

The pdsc_core_init() check is also worse than what it replaces. Its
bail-out jumps to err_out_uninit, which ends up in pdsc_intr_free() and
writes to pdsc->intr_ctrl, also NULL at that point.

On v2 I said I would convert pdsc_identify() and pdsc_core_init() to
pdsc_devcmd_with_data() once the PLDM series landed. Dropping that: the
helper has no read-back path and both callers need one, and giving them
an -ENXIO return means hardening pdsc_intr_free() against a NULL
intr_ctrl on the err_out_uninit path. That is a lot of churn to
deduplicate two call sites.

Patch 2 is the VF pci_release_regions() fix, older than the cmd_regs
race, so it carries its own Fixes tag.

The v2 changelog claim that pdsc_unmap_bars() clears db_pages was wrong.
It clears info_regs, cmd_regs, intr_status and intr_ctrl; db_pages is
never mapped.

v2: https://lore.kernel.org/20260804235946.177762-1-nikhil.rao@amd.com
v1: https://lore.kernel.org/20260729055258.1416225-1-nikhil.rao@amd.com
====================

Link: https://patch.msgid.link/20260901044219.1361466-1-nikhil.rao@amd.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 18:07:45 -07:00
Nikhil P. Rao
73608de7e5 pds_core: don't release PCI regions for VFs on reset
pdsc_reset_prepare() called pci_release_regions() unconditionally, but
only PFs call pci_request_regions() (pdsc_init_pf). On a VF FLR this
makes the kernel warn "Trying to free nonexistent resource".

Fixes: ffa5585833 ("pds_core: implement pci reset handlers")
Reported-by: sashiko-bot <sashiko-bot@kernel.org>
Closes: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260804235946.177762-1-nikhil.rao%40amd.com
Signed-off-by: Nikhil P. Rao <nikhil.rao@amd.com>
Link: https://patch.msgid.link/20260901044219.1361466-3-nikhil.rao@amd.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 18:07:36 -07:00
Nikhil P. Rao
7980325b2f pds_core: fix cmd_regs access racing BAR unmap on reset
pdsc_reset_prepare() and pdsc_reset_done()'s pdsc_map_bars() error path
clear/iounmap cmd_regs without devcmd_lock, and
pdsc_legacy_firmware_update()'s download loop derefs cmd_regs after
dropping and retaking the lock without re-checking. An FLR concurrent
with a devlink flash can unmap cmd_regs under an in-flight devcmd,
causing a NULL deref or a write to unmapped MMIO.

Take devcmd_lock across the BAR unmap/remap, and re-check cmd_regs in
the download loop. Only the PF maps cmd_regs and runs devcmd, so skip
the unmap on a VF, as pdsc_remove() and pdsc_reset_done() already do.

A reset that completes entirely within the unlocked window is not a
correctness problem for the image: the device clears its update session,
so a resumed download is rejected, and it verifies the staged image
before writing a flash slot, reporting PDS_RC_BAD_FW rather than
activating it.

pdsc_unmap_bars() also clears info_regs, intr_status and intr_ctrl. The
interrupt and start/stop readers of those are quiesced before the unmap
by pdsc_fw_down(), which frees the interrupts and tears down the queues.
The debugfs readers are not, since those files outlive a reset; that is
pre-existing and out of scope here.

Fixes: e96094c1d1 ("pds_core: Clear BARs on reset")
Reported-by: sashiko-bot <sashiko-bot@kernel.org>
Closes: https://sashiko.dev/#/patchset/20260708212222.296202-1-nikhil.rao%40amd.com?part=3
Signed-off-by: Nikhil P. Rao <nikhil.rao@amd.com>
Link: https://patch.msgid.link/20260901044219.1361466-2-nikhil.rao@amd.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 18:07:36 -07:00
Jakub Kicinski
6262acad9d Merge branch 'net-cap-tx_queue_len-at-s16_max-to-prevent-oversized-ring-allocations'
Jamal Hadi Salim says:

====================
net: cap tx_queue_len at S16_MAX to prevent oversized ring allocations

An unprivileged user (via unshare -Urn) can set a huge tx_queue_len
and exhaust global memory through ring allocations sized from it
(pfifo_fast skb_arrays, tun/tap ptr_rings).
The reproducer from vega@nebusec.ai set the following params for
illustration: txqlen of 500000 -> ~32 GiB/ring attempts, 1.6 GB tun,
~960 MB tap. Gets worse when you consider qdiscs like mq.

What we fix: every path an unprivileged user can use to install
an oversized tx_queue_len is rejected with -ERANGE before any ring is
allocated; per-ring memory is bounded at 256 KiB.

This is for you sashikos: What we deliberately _do not fix_
bound the NUMBER of rings. With the cap in place the worst case moves
from "one knob" to the aggregate of ring x queues x devices, example:

  ip link add v0 numtxqueues 4096 txqueuelen 32767 type veth
  tc qdisc add dev v0 root mq
    -> 4096 * 3 * 32767 * 8 = ~3.0 GiB (one command)
  50 tun devices x 256 queues x 32767 x 8 = ~3.1 GiB

Unfortunately tx_queue_len is a bit ambigious in meaning:
In some cases it means a ring size (which is pre-allocated, ex:
tun, tap, and pfifo_fast); a cap of 4096 seems reasonable here.
but in other cases it is used to indicate a queue limit ex:
the qdisc consumers that allocate nothing (pfifo/bfifo/gred/plug/sfb,
htb direct_qlen, qfq, teql). 32767 is a legitimate high-BDP queue
length, so we are going to keep that value.

Getting back to you sashikos, after this is merged and shows up
in net-next we will send followup patches as follows:
this series is not misread as "closes the OOM class"):

a) Per-site ring limits at six identified locations
    - pfifo_fast init/resize,
    - tun attach/resize,
    - tap minor/resize)

   if you can spot more in your review we will take care of those as well.

b) memcg accounting (GFP_KERNEL_ACCOUNT) for those ring
   allocations: contains a memcg-limited container's ring memory.
   Not GFP_KERNEL_ACCOUNT has no effect on the unshare attacker
   but will protect against containers  (memory.max in its cgroup)
====================

Link: https://patch.msgid.link/QDISC-2899.v2.20260901233641@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 16:30:53 -07:00
Jamal Hadi Salim
0a7252d7f8 selftests: tc-testing: add tx_queue_len cap regression tests
Add nine test cases for the S16_MAX tx_queue_len cap to the
pfifo_fast suite. Netlink cases exercise the ifla_policy bound
(2/3); the two new sysfs cases exercise the netif_change_tx_queue_len()
choke point that 1/3 owns (SIOCSIFTXQLEN shares it; the ioctl is not
portably reachable from tdc):

- dbe3: set txqueuelen 32767 (S16_MAX) - accepted, pins the exact
  boundary value.
- b50e: set txqueuelen 32768 - rejected with -ERANGE.
- 40f8: write 32768 to /sys/class/net/*/tx_queue_len - rejected
  (covers patch 1/3 directly; netlink cannot reach this path).
- 4b6e: write 32767 via sysfs - accepted, boundary positive control
  for the patch-1 path.
- b90d: create a dummy with txqueuelen 32767 - accepted.
- 57ab: create a dummy with txqueuelen 32768 - rejected at netlink
  parse time.
- e777: create a dummy with txqueuelen 500000 - rejected (the v1
  bypass path flagged by review).
- 31ac: create a veth with an oversized txqueuelen on the peer nest -
  rejected (the peer nest is parsed against ifla_policy too).
- b567: create a veth with txqueuelen on both ends within the cap -
  accepted (positive control for the peer nest).

The three negative-creation verifies assert device absence
("ip -o link show" must not contain the device), not merely absence
of a qlen pattern - the device does not exist when creation fails, so
the exit code carries the signal and the verify adds content.

The v1 04b5 "resize rollback" case is dropped: with the cap checked
first, netif_change_tx_queue_len() returns -ERANGE before the write,
the notifier or any qdisc resize, so the case exercised no resize and
no rollback. It was also nondeterministic: pre-patch, the resize
issues three ~11 MB kvmallocs for qlen 500000 which normally succeed,
so the case passed on an unfixed kernel only under memory pressure -
its outcome depended on the test host's free memory.

Test commands run inside the netns, but nsPlugin creates the veth
peer in the root namespace, so the teardown deletes the in-ns end
only; deleting the peer via the pair is implicit.

Note: iproute2 treats "txqueuelen" appearing after "type X" as a
link-type attribute and silently drops it, so the creation cases
place it before "type" to actually reach the kernel.

Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-2899.v2.20260901233641@mojatatu.com.3
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 16:30:49 -07:00
Jamal Hadi Salim
1aa9e143bf net: reject oversized tx_queue_len at netlink parse time
rtnl_create_link() assigns IFLA_TXQLEN directly to dev->tx_queue_len
without going through netif_change_tx_queue_len(), so a device created
with "ip link add ... txqueuelen 500000" bypasses the S16_MAX cap and
still triggers the oversized ring allocations in pfifo_fast, tun and
tap. The veth peer nest (rtnl_nla_parse_ifinfomsg()) and the
RTM_NEWLINK-on-existing-device path reach the same sinks.

Enforce the cap in ifla_policy instead: IFLA_TXQLEN becomes
NLA_POLICY_FULL_RANGE(NLA_U32, &txqlen_range) with
txqlen_range = { .min = 0, .max = S16_MAX }. All netlink consumers
parse against this policy - rtnl_setlink(), rtnl_newlink() (create
and change), and the veth peer nest - so every netlink path is capped
at parse time and rejects the attribute with -ERANGE plus a proper
"integer out of range" extack message before any device state is
modified (the RTM_SETLINK half-application wart is gone with it).

Document the bound in the rt-link.yaml netlink spec.

Conditions to recreate the bug:
- CONFIG_NET_SCHED=y, CONFIG_VETH=y, CONFIG_USER_NS=y, CONFIG_NET_NS=y.
- Unprivileged user in a fresh user+net namespace (unshare -Urn):
  ip link add v0 txqueuelen 500000 type veth peer name v1
  -> on the fixed kernel this is rejected with -ERANGE ("integer out
  of range" extack) instead of installing an oversized tx_queue_len
  that later inflates pfifo_fast/tun/tap ring allocations.
- ip link set v0 txqueuelen 500000 is likewise rejected at parse time.

Fixes: 38f7b870d4 ("[RTNETLINK]: Link creation API")
Reported-by: Vega <vega@nebusec.ai>
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-2899.v2.20260901233641@mojatatu.com.2
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 16:30:49 -07:00
Jamal Hadi Salim
66ab4c59b7 net: cap tx_queue_len at S16_MAX to prevent oversized ring allocations
Several subsystems allocate ring buffers sized by dev->tx_queue_len
with no upper bound. An unprivileged user (via unshare -Urn) can set a
huge tx_queue_len and exhaust global memory with ring allocations:

- pfifo_fast: pfifo_fast_init() and pfifo_fast_change_tx_queue_len()
  allocate 3 skb_array rings of tx_queue_len entries each.
- tun: tun_queue_resize() and the queue-attach path resize ptr_rings
  to tx_queue_len on the NETDEV_CHANGE_TX_QUEUE_LEN notifier.
- tap (macvtap/ipvtap): tap_queue_resize() and tap_init() resize/init
  ptr_rings to tx_queue_len on the same notifier.

netif_change_tx_queue_len() is the single entry point for IFLA_TXQLEN,
sysfs, and the SIOCSIFTXQLEN ioctl. Cap new_len at S16_MAX (32767)
there so the oversized value is rejected at set time. This takes
effect whether the device is up or down, before dev->tx_queue_len is
written, before any notifier fires, and before any ring is allocated.
The "> S16_MAX" check also subsumes the previous unsigned-long
truncation test, and a negative ifr_qlen from the ioctl lands far
above the cap after conversion, so both old failure modes are covered
by the one comparison.

tx_queue_len is ambigious: both a per-ring sizing multiplier and a
default queue-length/limit knob for consumers that allocate
nothing at set time (pfifo/bfifo/gred/plug/sfb limits, htb
direct_qlen, qfq max_classes, teql). 32767 is chosen as the largest
value NLA_POLICY_FULL_RANGE can express for the u32 IFLA_TXQLEN
policy in patch 2/3 while staying a legitimate queue length on
high-BDP paths; the ring-memory trade-off of a shared knob is
disclosed below.

Conditions to recreate the bug:
- CONFIG_NET_SCHED=y, CONFIG_VETH=y, CONFIG_USER_NS=y, CONFIG_NET_NS=y.
- Unprivileged user in a fresh user+net namespace (unshare -Urn).
- pfifo_fast: create veth pairs, set tx_queue_len to 500000, attach
  mq+pfifo_fast. ~28 iterations OOMs a 2GB guest.
- tun: create 50 tun devices with IFF_MULTI_QUEUE, set tx_queue_len to
  500000, open 8 queues each. ~1.6GB of ptr_ring allocations OOMs a
  512MB guest.
- tap: same as tun with IFF_TAP. ~960MB OOMs a 512MB guest.
- On the fixed kernel the oversized tx_queue_len is rejected with
  -ERANGE at set time (all four paths: RTM_SETLINK, RTM_NEWLINK
  create, sysfs, ioctl - the latter two via this check, the former
  two via this check and the 2/3 parse policy respectively).

Fixes: 6a643ddb56 ("net: introduce helper dev_change_tx_queue_len()")
Reported-by: Vega <vega@nebusec.ai>
Closes: https://lore.kernel.org/netdev/20260828121902.66837-1-jhs@mojatatu.com/
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-2899.v2.20260901233641@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 16:30:48 -07:00
Seungwon Bae
98fc57d167 vxlan: reject dynamic fdb entries that reference a nexthop id
The commit cited in the Fixes tag allowed VXLAN FDB entries to point to
FDB nexthops so that overlay traffic could be load balanced across
multiple VTEPs. Such entries can only be configured from user space,
cannot be learned and cannot roam. They only make sense with a user space
control plane such as E-VPN where data plane learning is disabled.

Despite that, the VXLAN driver does not currently prevent such entries
from being configured with the "dynamic" flag. The per-nexthop FDB list
is only protected by the per-device hash lock, which is not sufficient
when two VXLAN devices point to the same FDB nexthop and therefore share
the list. Aging runs in softirq context without RTNL, so an entry deleted
by one device can race with an addition or deletion from the other,
leading to list corruption:

  list_del corruption. next->prev should be ffff8881069d9548, but was
  dead000000000122. (next=ffff8881069d9448)
  WARNING: CPU: 0 PID: 90 at lib/list_debug.c:65
  __list_del_entry_valid_or_report+0x1aa/0x210
  ...
   vxlan_fdb_destroy+0x5b8/0xad0
   vxlan_cleanup+0x328/0x450
   call_timer_fn+0x2a/0x1c0
   run_timer_softirq+0x18c/0x210
  BUG: KASAN: slab-use-after-free in vxlan_fdb_destroy

Fix this by rejecting the bogus configuration of dynamic FDB entries that
point to FDB nexthops, both when created and when an existing entry is
updated. As such, the per-nexthop FDB list is only ever mutated under the
RTNL lock. Add test cases to make sure that this does not regress in the
future.

Fixes: 1274e1cc42 ("vxlan: ecmp support for mac fdb entries")
Suggested-by: Ido Schimmel <idosch@nvidia.com>
Signed-off-by: Seungwon Bae <qotmddnjs@ajou.ac.kr>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260902155956.296699-1-qotmddnjs@ajou.ac.kr
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 16:14:52 -07:00
Alexandra Winter
907a56ab3e s390/ism: folio_put() after error
dmb->cpu_addr was allocated via folio_alloc(). Use folio_put() instead of
kfree() in the error exit of ism_alloc_dmb() to avoid slab allocator
corruption.

While at it, reset dmb->cpu_addr after folio_put to avoid unintentional UAF
by future callers.

Fixes: 83781384a9 ("s390/ism: Properly fix receive message buffer allocation")
Signed-off-by: Alexandra Winter <wintera@linux.ibm.com>
Reviewed-by: Gerd Bayer <gbayer@linux.ibm.com>
Link: https://patch.msgid.link/20260902143733.433574-1-wintera@linux.ibm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 16:12:36 -07:00
Alexandra Winter
1668a31e3b dibs: Unregister dibs_class after error
In case dibs_loopback_init() fails, e.g. because of -ENOMEM, dibs_init()
must unregister dibs_class. Otherwise dibs_class and /sys/class/dibs exist
even though the functionality is not available. A retry to load the module
fails with -EEXIST.

Unregister dibs_class in the error path of dibs_init.

Note that before
commit ad3dfa80be ("dibs: change dibs_class to a const struct")
class_destroy(dibs_class) is required instead of
class_unregister(&dibs_class).

Fixes: 8047373498 ("dibs: Create class dibs")
Signed-off-by: Alexandra Winter <wintera@linux.ibm.com>
Link: https://patch.msgid.link/20260902143438.426664-1-wintera@linux.ibm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 16:12:24 -07:00
Jason Winter
5d50e90add net: usb: cx82310_eth: drop URB after 0xffff reboot sentinel to prevent partial_data heap overflow
The 0xffff length sentinel detects a router reboot and schedules
re-enabling of ethernet mode, but then falls through to the rest
of the loop body.  The next check is

	} else if (len > CX82310_MTU) {

which is the else of the just-matched if -- it never fires for
len == 0xffff.  The MTU bound that normally caps the
incomplete-packet save path is silently bypassed.

With 0xffff > skb->len always true (rx_urb_size is 4096), the
incomplete-packet branch saves dev->partial_len = skb->len bytes
into dev->partial_data.  partial_data is kmalloc(hard_mtu) =
kmalloc(CX82310_MTU + 2) = 1516 bytes, but skb->len after the
2-byte header pull can be up to 4094.  A device that sends a
4096-byte URB starting with [0xff 0xff] therefore copies 4094
device-provided bytes into a buffer allocated for 1516 bytes,
exceeding its requested size by 2578 bytes.

The next URB then reads dev->partial_len (4094) back from the same
1516-byte buffer and dev->partial_rem (65535 - 4094 = 61441) from
the new URB's ~4KB skb, both well past their allocations, and
delivers the spliced result as a 64KB "frame" to the network
stack.

Bail out of rx_fixup after scheduling the re-enable work; the
remainder of a reboot-marker URB is not meaningful packet data.
This restores the invariant that partial_len < CX82310_MTU + 2 on
the save path, since every other route there has already passed
the MTU check.

Fixes: ca139d76b0 ("cx82310_eth: re-enable ethernet mode after router reboot")
Signed-off-by: Jason Winter <jjx@live.nl>
Link: https://patch.msgid.link/BESP194MB283265DDDC63B6B78D8D34FBB8B72@BESP194MB2832.EURP194.PROD.OUTLOOK.COM
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 16:05:48 -07:00
Fourie Zhang
78a86d75a7 net: mpls: clear inner_protocol when the last label is popped
skb_mpls_push() records the pre-encapsulation network header once, gated
on !skb->inner_protocol. skb_mpls_pop() never clears that record, so it
outlives the encapsulation it describes.

Open vSwitch can then re-push MPLS onto a packet whose
inner_network_header still points at the older, deeper offset: push a
label, pop every label, recirculate (ovs_flow_key_update() re-derives
key->eth.type and resets network_header, but leaves inner_*), then push
again. ovs_fragment() trusts the record:

	skb->network_header = skb->inner_network_header;

so skb_network_offset() goes negative. The bound check is signed:

	if (skb_network_offset(skb) > MAX_L2_LEN)

a negative offset passes it, and prepare_frag() widens the value:

	unsigned int hlen = skb_network_offset(skb);
	memcpy(&data->l2_data, skb->data, hlen);

which is a ~4GiB memcpy out of a 30-byte per-CPU buffer.

Reproduced on v7.3-rc1. RDX is the truncated length, (unsigned int)(-8):

  BUG: unable to handle page fault for address: ffffe8ffffc16000
  #PF: supervisor write access in kernel mode
  Oops: 0002 [#1] SMP KASAN NOPTI
  RIP: 0010:memcpy+0x8/0x20
  RDX: 00000000fffffff8 RSI: ffff888105d732db RDI: ffffe8ffffc16000
   prepare_frag+0x3df/0x4e0
   ovs_fragment+0x589/0x7e0
   do_output+0x4ce/0x5e0
   do_execute_actions+0x55d2/0x7b30
   ovs_execute_actions+0xea/0x450

Same root-cause shape as commit 975b5b067f ("ipv6: sr: restore network
header before routing and forwarding"): a stale network header offset
reaching a consumer that widens it. Here it originates in the MPLS
push/pop path.

Clear inner_protocol once the packet is no longer MPLS, so a later push
re-records the current header. net/sched/act_mpls.c is the only other
skb_mpls_pop() caller and gets the same fix; sch_frag.c saves and
restores inner_protocol around fragmentation in the same way OVS does.

Fixes: 48d2ab609b ("net: mpls: Fixups for GSO")
Cc: stable@vger.kernel.org
Signed-off-by: Fourie Zhang <fouriezhang@tencent.com>
Acked-by: Jiri Benc <jbenc@redhat.com>
Link: https://patch.msgid.link/20260902092719.2874481-1-fouriezhang@tencent.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 16:03:53 -07:00
Nikhil P. Rao
c91b4d6e5c ionic: use netif_txq_maybe_stop() in ionic_tx()
Commit 061b9bedbe ("ionic: Rework Tx start/stop flow") replaced
ionic_maybe_stop_tx() with netif_txq_maybe_stop() to get the memory
barriers around the stop/start bits right, but did not cover the stop
in ionic_tx() added by commit 138506ab24 ("ionic: Check stop no
restart"). Convert the remaining site.

netif_txq_maybe_stop() requires the ring indexes to be updated before
it is invoked, so the post has to come first. But ring_dbell comes
from __netdev_tx_sent_queue(), which runs after that and reads the
stop bit, so it is not known in time to pass to ionic_txq_post(). Post
without the doorbell and ring it separately.

The stop condition is unchanged. The re-check only clears the stop bit
when space has become available, so the doorbell starvation fixed by
commit 138506ab24 ("ionic: Check stop no restart") cannot recur.

Fixes: 138506ab24 ("ionic: Check stop no restart")
Signed-off-by: Nikhil P. Rao <nikhil.rao@amd.com>
Reviewed-by: Brett Creeley <brett.creeley@amd.com>
Link: https://patch.msgid.link/20260901055627.1373129-1-nikhil.rao@amd.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:50:51 -07:00
Viswajith Murali
1f29543126 octeontx2-af: mcs: Clear stale X2P calibration state before calibration
Some firmware versions leave MCSX_MIL_GLOBAL bit 5 set on boot.
If the bit is already set when the driver attempts X2P calibration,
the hardware sees no rising edge and calibration never triggers.
Clear the bit and wait briefly before starting calibration to ensure
a clean rising edge.

Fixes: ca7f49ff88 ("octeontx2-af: cn10k: Introduce driver for macsec block.")
Signed-off-by: Nitin Shetty J <nshettyj@marvell.com>
Signed-off-by: Viswajith Murali <viswajithm@marvell.com>
Link: https://patch.msgid.link/20260901094318.1395356-1-nshettyj@marvell.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:44:59 -07:00
Nikolay Aleksandrov
4b772869a1 net: bridge: mcast: properly convert mglist to rcu
Sashiko reported a bug [1] that br_multicast_del_port_group unlists the
port group not using proper rcu helper that preserves the next pointer and
after that immediately frees the port group without waiting for rcu grace
period. The only rcu walker of mglist is br_multicast_list_adjacent() and
it turns out that function has always been buggy because mglist was never
properly converted to RCU. Fix it by converting it to rcu and moving its
initialization after eth_addr's. Initializing p->next can use
RCU_INIT_POINTER because we have a barrier from the hlist_add_head_rcu call
later, besides we're initializing an unpublished structure anyway.

[1] https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260826014200.362304-1-littleddfu%40gmail.com

Fixes: 07f8ac4a1e ("bridge: add export of multicast database adjacent to net_dev")
Signed-off-by: Nikolay Aleksandrov <razor@blackwall.org>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260903093851.1494297-1-razor@blackwall.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:29:47 -07:00
Jakub Kicinski
742b94967d Merge branch 'eth-fix-bugs-in-ntuple-filter-reporting'
Jakub Kicinski says:

====================
eth: fix bugs in ntuple filter reporting

Looking thru some reports prompted by:
  Add new way to add BPF LSM hooks
  https://lore.kernel.org/20260831110934.241898-1-a.s.protopopov@gmail.com

I/Claude noticed 3 drivers with buggy n-tuple filter dump.
Fix these drivers, add a hopefully clearer mention in the doc.
====================

Link: https://patch.msgid.link/20260903032611.3000029-1-kuba@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:20:00 -07:00
Jakub Kicinski
47a582b2b0 ethtool: document that GRXCLSRLALL rule_cnt is a caller-provided limit
Three drivers have shipped a get_rxnfc() which dumps its entire rule
table into rule_locs, reading rule_cnt as "how many rules do I have"
rather than "how many entries did the caller allocate".  Nothing in the
callback's documentation contradicted that reading.  The distinction only
matters because the ioctl lets an unprivileged caller pick rule_cnt
directly, so getting it wrong is a heap overflow rather than a truncated
dump.

Reviewed-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260903032611.3000029-6-kuba@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:19:58 -07:00
Jakub Kicinski
b1fffc2731 net: dsa: mv88e6xxx: bound the policy rule dump by the caller's buffer size
mv88e6xxx_get_rxnfc() uses rxnfc->rule_cnt as the write index while
dumping the policy IDR, clobbering the input value before it has been
looked at.  That input is the number of entries the caller had room for.
ETHTOOL_GRXCLSRLALL requires no CAP_NET_ADMIN and the ioctl sizes the
buffer from the rule_cnt userspace passes in, so once an admin has
installed policy rules any user can ask for fewer slots than there are
rules and run off the end of the allocation.  A rule_cnt of 0 leaves the
buffer pointer NULL and the walk dereferences it.

Count into a local so the caller's limit survives the walk, and stop with
-EMSGSIZE once it is reached.

Fixes: da7dc87553 ("net: dsa: mv88e6xxx: add RXNFC support")
Reviewed-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260903032611.3000029-5-kuba@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:19:57 -07:00
Jakub Kicinski
108bb2142e eth: nfp: drop the replaced rule from the list when reprogramming fails
nfp_net_fs_add() replaces an existing rule by deleting it from the
hardware, decrementing nn->fs.count and programming the new one.  If
nfp_net_fs_add_hw() fails the old entry stays on nn->fs.list - only the
success path reaches list_replace() - so the list is one longer than
nn->fs.count, and it advertises a rule whose hardware entry has already
been torn down.

nn->fs.count is what ETHTOOL_GRXCLSRLCNT reports, so userspace then sizes
its buffer one entry short of what the GRXCLSRLALL walk wants to write.
That used to overwrite one u32 past the allocation; since the walk is
bounded it is a permanent -EMSGSIZE instead, as nothing ever resyncs the
counter.

Fixes: 9eb03bb1c0 ("nfp: add ethtool flow steering callbacks")
Reviewed-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260903032611.3000029-4-kuba@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:19:57 -07:00
Jakub Kicinski
f1986bf87b eth: nfp: bound the ntuple rule dump by the caller's buffer size
nfp_net_get_fs_loc() dumps every entry of nn->fs.list into rule_locs[]
without consulting cmd->rule_cnt, which is how many entries the caller
had room for.  ETHTOOL_GRXCLSRLALL requires no CAP_NET_ADMIN and the
ioctl sizes the buffer from the rule_cnt userspace passes in, so once an
admin has installed flow steering rules any user can ask for fewer slots
than there are rules and run off the end of the allocation.  A rule_cnt
of 0 leaves the buffer pointer NULL and the walk dereferences it.

Bail out with -EMSGSIZE when the buffer fills up, the way the other
ntuple capable drivers do, and report how many locations were filled so
a shrinking rule list does not leave the caller reading stale slots.

Reported-by: VEGA <vega@nebusec.ai>
Fixes: 9eb03bb1c0 ("nfp: add ethtool flow steering callbacks")
Reviewed-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260903032611.3000029-3-kuba@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:19:57 -07:00
Jakub Kicinski
cdb719f4b8 net: dsa: bcm_sf2: bound the CFP rule dump by the caller's buffer size
bcm_sf2_cfp_rule_get_all() walks the whole cfp.unique bitmap into
rule_locs[] without consulting nfc->rule_cnt, which is how many entries
the caller had room for.  ETHTOOL_GRXCLSRLALL requires no CAP_NET_ADMIN
and the ioctl sizes the buffer from the rule_cnt userspace passes in, so
once an admin has installed CFP rules any user can ask for fewer slots
than there are rules and run off the end of the allocation.  A rule_cnt
of 0 leaves the buffer pointer NULL and the walk dereferences it.

Fixes: 7318166cac ("net: dsa: bcm_sf2: Add support for ethtool::rxnfc")
Reviewed-by: Jonas Gorski <jonas.gorski@gmail.com>
Reviewed-by: Florian Fainelli <florian.fainelli@broadcom.com>
Reviewed-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260903032611.3000029-2-kuba@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:19:57 -07:00
Eric Dumazet
1746ef2e2d bonding: use skb_cow_head() in bond_do_alb_xmit() and rlb_arp_xmit()
In bond_do_alb_xmit() and rlb_arp_xmit(), make sure to unclone
skb head via skb_cow_head() before modifying the source MAC address
(Ethernet header and ARP payload) to avoid silent corruption if
the skb is shared or cloned. Avoid caching the header pointers
across skb_cow_head().

In rlb_arp_xmit(), only modify arp->mac_src if it differs from
tx_slave->dev->dev_addr to avoid an unnecessary copy and head
reallocation.

Also, we should not assume mac header is set in output path.

Use skb_eth_hdr() instead of eth_hdr() to fix the issue,
and remove now redundant skb_reset_mac_header() calls.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Cc: Jay Vosburgh <jv@jvosburgh.net>
Reviewed-by: Nikolay Aleksandrov <razor@blackwall.org>
Link: https://patch.msgid.link/20260903143940.1180513-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 15:18:41 -07:00