Commit Graph

1481759 Commits

Author SHA1 Message Date
Long Li
f6d61fe4c1 net: mana: Clear RDMA teardown and suspend state in mana_rdma_probe()
mana_rdma_remove() sets gd->rdma_teardown to stop
mana_rdma_service_handle() from acting on servicing events, but nothing
ever clears it. A hardware service reset (GDMA_EQE_HWC_RESET_REQUEST)
goes through mana_gd_suspend() -> mana_rdma_remove() and mana_gd_resume()
-> mana_rdma_probe(), so from the first reset onwards every
GDMA_EQE_HWC_SOC_SERVICE event returns early and RDMA suspend/resume
servicing is silently dropped for the life of the device.

gd->is_suspended has the same problem: it is set when servicing removes
the adev and is cleared only by a matching resume. A reset while RDMA is
suspended re-adds the adev but leaves is_suspended set, so a later resume
event calls add_adev() on top of a live gd->adev and leaks it. This is
currently masked by the rdma_teardown bug.

Clear both in mana_rdma_probe(). On the reset path mana_rdma_remove()
has closed the gate and drained the service workqueue, so clear
is_suspended first and re-open the gate with smp_store_release(), paired
with smp_load_acquire() in the handler, so the handler cannot observe an
open gate with a stale is_suspended. On the initial probe path the gate
was never closed and both flags are already clear.

This does not order gd->adev, which add_adev() publishes afterwards. A
servicing event arriving in that window is still dropped, as it is in
mainline today on the initial probe path; closing it needs probe and the
handler to be serialized and is left to a separate change.

Fixes: 505cc26bca ("net: mana: Add support for auxiliary device servicing events")
Signed-off-by: Long Li <longli@microsoft.com>
Link: https://patch.msgid.link/20260902175153.3410560-1-longli@microsoft.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-08 15:53:29 -07:00
Jakub Kicinski
1b8e56030d netfilter pull request 26-09-07
-----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEjF9xRqF1emXiQiqU1w0aZmrPKyEFAmqe60oACgkQ1w0aZmrP
 KyF4xw//THfqJRAWHoVGyQcfRQs1/FdmE5XcdN8pDdS55XvYljDx+A8UyJlMoBaq
 tXr76jA74nngwkwkh8dmaCfGurHe1GtVkqE38AMfOEUbuImqvznFG9fp1mb3u3Kz
 jmoeOhxjZGXzBw9ng1xs+Ip0opU8GgulVdGSD0NzzHFRGdj29MA/q5Y1Eo2KqYQ/
 NlKgivExotll20YirWgOMHrAt5uGqtJLWcZWCj3G1ASdlzNHcM2HF+1vW4TvMFEy
 jKSfN8u+YdI+k+5TLQtPFvodMqWgsZhZ6llqbxUav7uO/NkZLha0wJf9Lwky9s+F
 4/iRsnXos9imEa4pm8tkVg12xl9P3rLMrFYqfrIoNL94nUDgXyLZ+44qBBROWnaN
 97KJ4fok7Ny38cIqz4CbwEWncB71obBokmBhY+byoyhwAmMkxBnfj6C5emU4NXvp
 6EiB0xh4PXBYaCfZO4JnDPXHNNFkSi+MGBhpKlV0K2ZRNAKB4jYuWw3jSGdwTNIA
 bxKsH/j3nYvpNCDBeEhxQyokNJ9yWxhNRz41Ejym4BKN0PndkQlU8dmhS+kDWiyD
 Gswqmi6WJhflxxI6mq3FROl10vmZnSQfJk4nKhT4TYxmyvB8HbA5FWV7Gd+77yxg
 i7zmLrE3TzKu3cbSuM+o5Qei9tcTDCdMa8HH3dZbPbQlFvocGVc=
 =BBrp
 -----END PGP SIGNATURE-----

Merge tag 'nf-26-09-07' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf

Pablo Neira Ayuso says:

====================
Netfilter/IPVS fixes for net

The following patchset contains Netfilter/IPVS fixes for net:

1) Reject malformed messages in IPVS sync, from Kyle Zeng.

2) Fix possible stale infoleak in IPVS sync, also from Kyle Zeng.

3) Out-of-bound read in the SIP conntrack helper, from
   Joas Antonio dos Santos.

4) UaF on cttimeout module removal, from Chengfeng Ye.

5) Unregister nf_loggers before netns teardown to fix UaF,
   also from Chengfeng Ye.

6) Fix race in nfnetlink_log due to concurrent instance destruction,
   from Florian Westphal.

7) Remove arp_table 32bit compat interface, this is already off in
    many distributions, from Florian Westphal.

8) Set IP6T_F_PROTO flag is e->ipv6.proto is set on to deal with
    insufficient validation of xtables extensions when used from
    legacy ip6tables, from Florian.

9) Set on the NLM_F_DUMP_FILTERED flag when all is filtering out
   in ctnetlink, from Ilya Maximets.

* tag 'nf-26-09-07' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf:
  netfilter: report NLM_F_DUMP_FILTERED when all is filtered out
  netfilter: ip6_tables: set F_PROTO when proto value is nonzero
  netfilter: arp_tables: remove the 32bit compat interface
  netfilter: nfnetlink_log: cope with concurrent instance destruction
  netfilter: nf_log: unregister loggers before per-net teardown
  netfilter: cttimeout: prevent UAF during module unload
  netfilter: nf_conntrack_sip: fix OOB read in sip_skip_whitespace()
  ipvs: fix reversed sequence option serialization
  ipvs: reject invalid states in connection template sync records
====================

Link: https://patch.msgid.link/20260907171732.1407739-1-pablo@netfilter.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-08 13:53:17 -07:00
Paolo Abeni
e0554c6276 Merge branch 'net-ethernet-cortina-fix-rx-budget-accounting'
Linus Walleij says:

====================
net: ethernet: cortina: Fix RX budget accounting

Finish RX updates before releasing NAPI ownership, report actual NAPI
work, charge dropped frames to the poll budget, and drive free-queue
refills from consumed RX descriptors.

Track RX drop state across descriptor chains so discarded frames are
counted exactly once.

Tested on the D-Link DIR-685.

Hi Sashiko, yes there are more latent issues I will get to them, but
my LLM thinks those are on the top of the list.

Assisted-by: LLM
Signed-off-by: Linus Walleij <linusw@kernel.org>
====================

Link: https://patch.msgid.link/20260903-gemini-ethernet-fixes-v2-0-2bbbd598ca6e@kernel.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-08 12:35:24 +02:00
Linus Walleij
e89e88ad41 net: ethernet: cortina: Count RX descriptors for freeq refill
The software free queue provides one buffer fragment for every descriptor
moved to an RX queue. The refill heuristic instead advances by NAPI work,
which counts frames. A fragmented or discarded frame can consume several
queue entries while adding only one to the refill count.

Count the RX descriptors as they are consumed and report that separately
from NAPI work. Use the descriptor count to drive free queue refills.

Fixes: 4d5ae32f5e ("net: ethernet: Add a driver for Gemini gigabit ethernet")
Assisted-by: LLM
Reviewed-by: Joe Damato <joe@dama.to>
Signed-off-by: Linus Walleij <linusw@kernel.org>
Link: https://patch.msgid.link/20260903-gemini-ethernet-fixes-v2-5-2bbbd598ca6e@kernel.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-08 12:35:22 +02:00
Linus Walleij
6520198c43 net: ethernet: cortina: Count RX drops once per frame
The absence of a partial skb means either that the driver is not
assembling a frame or that the current frame was already dropped.
Consequently, repeated descriptor errors can increment rx_dropped more
than once, while an orphaned descriptor chain can reach EOF without being
counted at all.

Track the dropping state across NAPI polls. Clear it at frame boundaries
and route mapping failures and orphaned continuations through the common
drop path so each discarded frame is counted exactly once.

Fixes: 4d5ae32f5e ("net: ethernet: Add a driver for Gemini gigabit ethernet")
Reported-by: Joe Damato <joe@dama.to>
Closes: https://lore.kernel.org/netdev/apdK5aMmvYssz35F@devvm20253.cco0.facebook.com/
Assisted-by: LLM
Signed-off-by: Linus Walleij <linusw@kernel.org>
Link: https://patch.msgid.link/20260903-gemini-ethernet-fixes-v2-4-2bbbd598ca6e@kernel.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-08 12:35:22 +02:00
Linus Walleij
b856c552f5 net: ethernet: cortina: Count dropped frames as NAPI work
The RX loop only consumes budget when it successfully delivers a frame.
Error paths keep consuming descriptors without reducing the budget, so a
stream of bad frames can process the entire receive ring in one poll.

Move the budget accounting to a common end-of-frame path. This counts
each completed frame as NAPI work whether it was delivered or dropped,
matching the behavior of the vendor driver.

Fixes: 4d5ae32f5e ("net: ethernet: Add a driver for Gemini gigabit ethernet")
Assisted-by: LLM
Signed-off-by: Linus Walleij <linusw@kernel.org>
Link: https://patch.msgid.link/20260903-gemini-ethernet-fixes-v2-3-2bbbd598ca6e@kernel.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-08 12:35:21 +02:00
Linus Walleij
baa26841cb net: ethernet: cortina: Finish RX updates before NAPI completion
napi_complete_done() releases ownership of the NAPI instance, but the
Gemini poll keeps the RX statistics writer section open and updates the
free queue after calling it. A new poll can therefore start while the old
writer is still active.

Finish the statistics and free queue updates before releasing ownership.
Only re-enable RX interrupts when napi_complete_done() reports successful
completion.

Fixes: 4d5ae32f5e ("net: ethernet: Add a driver for Gemini gigabit ethernet")
Suggested-by: Joe Damato <joe@dama.to>
Assisted-by: LLM
Signed-off-by: Linus Walleij <linusw@kernel.org>
Link: https://patch.msgid.link/20260903-gemini-ethernet-fixes-v2-2-2bbbd598ca6e@kernel.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-08 12:35:21 +02:00
Linus Walleij
a0de06d0da net: ethernet: cortina: Fix budget accounting
The gmac_rx() function returns the remaining NAPI budget, but its
caller treats the return value as the number of packets received. An
idle poll therefore reports a full budget and remains scheduled.

Return the number of received packets instead. Preserve the existing
free queue refill accounting by adding that count directly; continuing
to subtract it from the budget would invert the refill behavior.

Fixes: 4d5ae32f5e ("net: ethernet: Add a driver for Gemini gigabit ethernet")
Link: https://lore.kernel.org/r/20260509-gemini-ethernet-fixes-v1-4-6c5d20ddc35b@kernel.org
Link: https://lore.kernel.org/r/20260512131456.189452-1-pabeni@redhat.com
Assisted-by: LLM
Reviewed-by: Joe Damato <joe@dama.to>
Signed-off-by: Linus Walleij <linusw@kernel.org>
Link: https://patch.msgid.link/20260903-gemini-ethernet-fixes-v2-1-2bbbd598ca6e@kernel.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-08 12:35:21 +02:00
Paolo Abeni
c51fe22812 Merge branch 'net-macb-fix-the-link-speed-the-taprio-setup-reads'
Aleksei Sviridkin says:

====================
net: macb: fix the link speed the taprio setup reads

Two small fixes in macb_taprio_setup_replace(), both in how it obtains
the link speed it scales the schedule with.

The first: it hands phylink_ethtool_ksettings_get() a stack variable
it never zeroed, while phylink fills only what the link mode provides
and even reads one field back from the caller. The second: the speed
check is written as "<= 0" on a u32, so SPEED_UNKNOWN passes it and
turns into a 1 ns hardware limit that every entry then exceeds.

Compile-tested against net; the driver has no test surface, and no
macb board here.
====================

Link: https://patch.msgid.link/20260903123652.23900-1-f@lex.la
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-08 11:46:36 +02:00
Aleksei Sviridkin
2b6c0e25a3 net: macb: reject an unknown link speed in the taprio setup
speed is a u32, so SPEED_UNKNOWN arrives as 0xffffffff and passes the
"speed <= 0" check, which only ever catches zero. That is what an
autonegotiating link reports while it is down: the limit derived from
the speed collapses to a nanosecond at most and the first entry fails
with a misleading "exceeds hardware limit". Zero stays covered, it is
what an interface that was never opened reports, and
enst_max_hw_interval() divides by it. Say which case it was in the
error.

Fixes: 89934dbf16 ("net: macb: Add TAPRIO traffic scheduling support")
Assisted-by: LLM
Signed-off-by: Aleksei Sviridkin <f@lex.la>
Link: https://patch.msgid.link/20260903123652.23900-3-f@lex.la
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-08 11:46:34 +02:00
Aleksei Sviridkin
0523d5c52a net: macb: zero the link settings taprio reads back
macb_taprio_setup_replace() calls phylink_ethtool_ksettings_get() with
an uninitialised kset, and kset is not only an out-parameter. On a
fixed link, or an in-band link with no PHY, phylink writes speed and
duplex only if kset->base.rate_matching already reads RATE_MATCH_NONE,
a field it never writes itself; in PHY mode before the PHY is attached
it writes port and supported and nothing more. Either way the speed
read back afterwards can be stack garbage. The ethtool core zeroes the
structure on every path into the op, which is why its callers never
see this; taprio is the only in-kernel caller passing its own variable.

Fixes: 89934dbf16 ("net: macb: Add TAPRIO traffic scheduling support")
Assisted-by: LLM
Signed-off-by: Aleksei Sviridkin <f@lex.la>
Link: https://patch.msgid.link/20260903123652.23900-2-f@lex.la
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-08 11:46:34 +02:00
Paolo Abeni
ae20d47d26 Merge branch 'fix-a-variety-of-tpa-bugs'
Joe Damato says:

====================
Fix a variety of TPA bugs

I am sending this series as an extension to my v4 [1] which was just 1 patch.

Note that patch 5 of this series can now cause the device to fail closed if
memory is tight; bnxt_init_nic propagates an error that was previously
swallowed and fails closed instead of succeeding in a degraded state. If the
maintainers want the device to come up with a partially populated rx_tpa[],
then patch 5 can be dropped and this series can still be applied
and will otherwise work as intended.

This series addresses a variety of bugs orbiting the TPA code in the bnxt
driver that Sashiko (or Clashiko or whatever) pointed out and the series ends
with the patch from the v4 [1].

A lot of the noise generated by the AIs while reviewing my v4 are unrelated
bugs with different fixes tags that, IMHO, distract a bit from the crash at
boot that is currently occurring with Thor2 hardware on recent kernels.

That said, I've tried to wrangle this series together which I hope will solve
most of the important bugs the AIs are feeling something about.

I do not know what other rabbit holes the AIs will find when I submit this
series, but if there is some reasonable stop-gap that we can get applied to
fix the crashes on Thor2 (while I iterate on the rest of the bugs at the
pleasure of the AIs) that would be excellent.

I boot tested this on a Thor1 and a Thor2 machine and there were no crashes at
boot.

[1]: https://lore.kernel.org/all/20260828190900.1767611-1-joe@dama.to/
====================

Link: https://patch.msgid.link/20260902015652.2421609-1-joe@dama.to
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-08 10:47:07 +02:00
Joe Damato
c0aceaf65b bnxt_en: Bound SW TPA IDs to prevent crashes
FW supports up to 1024 concurrent TPAs, so the FW TPA ID is in the range
0..1023 (see commit ec4d8e7cf0 ("bnxt_en: Add TPA ID mapping logic for
57500 chips.")). bnxt_alloc_agg_idx is intended to wrap the FW ID down to a
software ID which is used to index rxr->rx_tpa, and to generate a mapping
between FW IDs and the wrapped software ID.

On a 57608 with firmware version 233, the firmware advertises 32
concurrent TPAs. As of the commit under fixes, bp->max_tpa on this NIC
is set to 32.

If the software ID from bnxt_alloc_agg_idx is above 31, this results in
an invalid address being loaded on this line:

  tpa_info = &rxr->rx_tpa[agg_id];

because rx_tpa is allocated with only bp->max_tpa (32) entries. Writes
to tpa_info later in the code are out of bounds.

This bug results in a crash at boot:

Oops: general protection fault, kernel NULL pointer dereference 0x8: 0000 [#1] SMP NOPTI
RIP: 0010:bnxt_rx_pkt+0xc0/0x1560
RSP: 0018:ffffc900009b8c78 EFLAGS: 00010246
RAX: 0000000000000000 RBX: 0000000000000048 RCX: 0000000206682516
RDX: ffffc900009b8db4 RSI: 0000000000000000 RDI: 01ffffff038fe1c0
RBP: ffffc9006e687480 R08: ffffc9006e687000 R09: 0000000000003048
R10: 0000000000000480 R11: ffff8881c6083900 R12: 0000000006682516
R13: ffff8881c6095400 R14: 0000000000000016 R15: ffff8881c6b66680
FS:  0000000000000000(0000) GS:ffff88fef3c77000(0000) knlGS:0000000000000000
CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 00007fc8bda40584 CR3: 000000807c812001 CR4: 0000000008772ef0
PKRU: 55555554
Call Trace:
 <IRQ>
 ? __netif_receive_skb_list_core+0x1ca/0x250
 __bnxt_poll_work+0x152/0x280
 bnxt_poll_p5+0x1cd/0x480
 __napi_poll+0x30/0x180
 net_rx_action+0x20b/0x3b0
 ? note_gp_changes+0x53/0xe0
 ? tick_setup_sched_timer+0x180/0x180
 ? __napi_schedule+0x9a/0xb0
 ? bnxt_msix+0x24/0x30
 handle_softirqs+0xdd/0x2c0
 __irq_exit_rcu.llvm.3171231171502365008+0x47/0xf0
 common_interrupt+0x85/0x90
 </IRQ>
 <TASK>
 asm_common_interrupt+0x22/0x40

This stack trace is from a crash triggered when an out of bounds rx_tpa
is dereferenced. The invalid write mentioned above is silent in this
particular crash.

Fix this by allocating rx_tpa with bp->max_tpa rounded up to the next
power of 2 (bp->max_tpa_roundup_size) entries and masking the FW TPA ID
with that size, so the wrapped ID can never index past the end of the
array.

Fixes: 54c28fab2f ("bnxt_en: Set bp->max_tpa according to what the FW supports")
Reported-by: Raphael Cardoso Fernandes <raphaelcf@meta.com>
Suggested-by: Michael Chan <michael.chan@broadcom.com>
Cc: stable@vger.kernel.org
Signed-off-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260902015652.2421609-7-joe@dama.to
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-08 10:47:03 +02:00
Joe Damato
8e6a850c07 bnxt_en: Propagate RX ring init failures in bnxt_init_nic()
bnxt_init_rx_rings() returns an error when bnxt_alloc_one_rx_ring()
fails, but bnxt_init_nic() discards that return value and calls
bnxt_init_chip(), which enables TPA.

If an allocation fails, this could leave rxr->rx_tpa[] partially zeroed
and TPA would be enabled over an array with zeroed entries. This would
lead to a zeroed DMA address being handed out if the agg_idx is
translated to a SW index at a zeroed entry.

Fix this by propagating the error out of bnxt_init_nic(). Both callers
already check its return value and unwind with bnxt_free_skbs() and
bnxt_free_mem(), which tolerate a partially initialized RX ring.

Fixes: c0c050c58d ("bnxt_en: New Broadcom ethernet driver.")
Reported-by: Sashiko <sashiko-bot+sashiko@kernel.org>
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260828190900.1767611-1-joe%40dama.to
Cc: stable@vger.kernel.org
Signed-off-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260902015652.2421609-6-joe@dama.to
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-08 10:47:03 +02:00
Joe Damato
961e2a17c5 bnxt_en: Handle buffer allocation failure in bnxt_rx_ring_reset()
bnxt_rx_ring_reset() frees the ring buffers and then reallocates them,
ignoring the result.

bnxt_alloc_one_rx_ring() can fail in bnxt_alloc_one_tpa_info_data(), which
returns -ENOMEM on the first failed allocation and leaves the remaining
rxr->rx_tpa[] entries zeroed.

The error isn't propagated up, so the loop in bnxt_rx_ring_reset
continues and at the end the code re-enables TPA with partially
unallocated rx_tpa array.

This means that when the agg_id from hardware is mapped to a SW index in
rxr->rx_tpa[], an uninitialized slot can be chosen which would hand a
zero DMA address to the device.

Fix this by falling back to a global reset, which is what the existing
code already does when other functions fail, but unlike the other
failure cases this particular failure has to return because TPA can't
be re-enabled since the allocation failed.

Fixes: 8fbf58e17d ("bnxt_en: Implement RX ring reset in response to buffer errors.")
Reported-by: Sashiko <sashiko-bot+sashiko@kernel.org>
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260828190900.1767611-1-joe%40dama.to
Cc: stable@vger.kernel.org
Signed-off-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260902015652.2421609-5-joe@dama.to
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-08 10:47:03 +02:00
Joe Damato
b814dfbfeb bnxt_en: Propagate TPA buffer allocation failures in bnxt_queue_mem_alloc()
bnxt_alloc_one_tpa_info_data() returns -ENOMEM as soon as one allocation
fails. This leaves the remaining rxr->rx_tpa[] entries zeroed.

bnxt_queue_mem_alloc() discards that return value, so the partially
initialized ring is installed by bnxt_queue_start().

Since the agg_id is picked by the hardware and bnxt_alloc_agg_idx maps
it to a SW index in rxr->rx_tpa[], it is possible that an uninitialized
slot can be chosen which would hand a zero DMA address to the device.

Fix this by checking the return value of bnxt_alloc_one_tpa_info_data
and unwinding, freeing the ring buffers.

Fixes: bd649c5cc9 ("bnxt_en: handle tpa_info in queue API implementation")
Reported-by: Sashiko <sashiko-bot+sashiko@kernel.org>
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260828190900.1767611-1-joe%40dama.to
Cc: stable@vger.kernel.org
Signed-off-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260902015652.2421609-4-joe@dama.to
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-08 10:47:02 +02:00
Joe Damato
5ce7f36c33 bnxt_en: Don't free the live ring's TPA state on queue restart failure
bnxt_queue_mem_alloc() shallow copies the live RX ring into the clone:

  memcpy(clone, rxr, sizeof(*rxr));

the code currently clears pointers that the clone owns (such as
rx_agg_bmap), but rx_tpa and rx_tpa_idx_map are left pointing at memory
of the live ring that was cloned.

If an allocation failure happens later and the err_free_tpa_info label
is taken, the live ring's memory can be freed while still in use.

Fix this by initializing the clone's pointers to NULL to prevent live
ring state from being freed inadvertently.

Fixes: bd649c5cc9 ("bnxt_en: handle tpa_info in queue API implementation")
Reported-by: Sashiko <sashiko-bot+sashiko@kernel.org>
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260828190900.1767611-1-joe%40dama.to
Cc: stable@vger.kernel.org
Signed-off-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260902015652.2421609-3-joe@dama.to
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-08 10:47:02 +02:00
Joe Damato
4e17b5007b bnxt_en: Only restore LRO if the device supports TPA
With a P5+ device with firmware that reports max_aggs_supported == 0, it is
possible to make LRO settable by attaching and detaching an XDP program
even though the device does not support TPA.

Fix this by testing BNXT_SUPPORTS_TPA before restoring the feature bit.

Fixes: f0aa6a37a3 ("eth: bnxt: always recalculate features after XDP clearing, fix null-deref")
Reported-by: Sashiko <sashiko-bot+sashiko@kernel.org>
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260828190900.1767611-1-joe%40dama.to
Cc: stable@vger.kernel.org
Signed-off-by: Joe Damato <joe@dama.to>
Link: https://patch.msgid.link/20260902015652.2421609-2-joe@dama.to
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-08 10:47:02 +02:00
Zhiling Zou
8d6cd18850 ipv6: flowlabel: cap duplicate leases per socket
ipv6_flowlabel_get() allocates an ipv6_fl_socklist entry for every
successful GET. The recheck path for a compatible existing flowlabel
links another lease without applying any lease admission check. Repeated
GET requests for one shareable label can therefore grow a socket's lease
list without bound.

Reject a new unprivileged lease once the socket already holds
FL_MAX_PER_SOCK leases. Check this on the shared recheck path so reuse
of a globally interned label, including the fl_intern() collision path,
is covered as well. New-label admission remains under the existing
mem_check() policy.

Use capable(CAP_NET_ADMIN) rather than ns_capable(), matching
mem_check(). An unprivileged user must not bypass the cap by creating a
user namespace and a netns where they have CAP_NET_ADMIN, which would
still consume host memory.

Check the capability only when the socket reaches the limit, so
successful unprivileged GET requests below the cap do not generate a
capability audit. Do the admission check before updating linger and
expires so a rejected GET does not refresh the shared label, matching
the existing socket-list allocation failure path.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Suggested-by: Ido Schimmel <idosch@nvidia.com>
Signed-off-by: Zhiling Zou <zhilinz@nebusec.ai>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/83f8535972ff6e3741548476a1d50dec24c758be.1788415194.git.zhilinz@nebusec.ai
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-07 21:24:05 -07:00
Sebastian Sjoholm
4ff75f130d net: usb: qmi_wwan: add Quectel RG660QB
Add support for the Quectel RG660QB 5G module (USB ID 2c7c:013d).
Its QMI interface (interface 4) uses class/subclass/protocol ff/ff/ff
like the other recent Quectel modules, so match it the same way.

The remaining interfaces are handled by the option driver.

Tested with an early sample of the module on a Quectel 5G EVB connected
over USB 3 to a Raspberry Pi 5: qmicli talks to the module via
/dev/cdc-wdm0.

Signed-off-by: Sebastian Sjoholm <sebastian.sjoholm@gmail.com>
Link: https://patch.msgid.link/20260903180044.6179-1-sebastian.sjoholm@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-07 17:28:59 -07:00
Jakub Kicinski
7473a66d3a Merge branch 'fix-udp-length-overflow-in-edge-cases'
Alice Mikityanska says:

====================
Fix UDP length overflow in edge cases

These are fixes for rare edge cases of 16-bit UDP length field overflow
that might happen on netdevs with MTU >= 64k.

Exposed by the new WARN added to udp_set_len_short, reported by syzbot.
====================

Link: https://patch.msgid.link/20260901195714.673548-1-alice.kernel@fastmail.im
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-07 17:13:58 -07:00
Alice Mikityanska
199271ebc7 net: ipv6: Clamp to IP6_MAX_MTU in ip6_dst_mtu_maybe_forward
Commit 427faee167 ("net: ipv6: introduce ip6_dst_mtu_maybe_forward")
dropped the IP6_MAX_MTU clamp that used to be present in ip6_mtu(). A
similar IPv4 commit ac6627a28d ("net: ipv4: Consolidate ipv4_mtu and
ip_dst_mtu_maybe_forward") preserves the IP_MAX_MTU clamp.

Restore the upper bound in the IPv6 flow to avoid potential 16-bit
overflows in forwarding paths.

Fixes: 427faee167 ("net: ipv6: introduce ip6_dst_mtu_maybe_forward")
Signed-off-by: Alice Mikityanska <alice@isovalent.com>
Suggested-by: Willem de Bruijn <willemdebruijn.kernel@gmail.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260901195714.673548-5-alice.kernel@fastmail.im
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-07 17:13:55 -07:00
Alice Mikityanska
18a9a43421 selftests: net: Test UDP length overflow with PMTU discover and big MTU
Two previous commits fixed overflow of UDP length when setsockopt
IP(V6)_MTU_DISCOVER is set to IPV6_PMTUDISC_DO or IP(V6)_PMTUDISC_PROBE,
and a large packet is sent over a netdev with an unusually large MTU.

This commit adds the selftests that replicate the described steps to
reproduce for IPv6 and IPv4, and also one more test that ensures that
sending UDP jumbograms over a raw socket is still possible after the
fix.

Signed-off-by: Alice Mikityanska <alice@isovalent.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260901195714.673548-4-alice.kernel@fastmail.im
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-07 17:13:55 -07:00
Alice Mikityanska
0ae10b6be4 net: ipv6: Fix UDP length overflow with PMTU discover and big MTU
This commit bounds cork->base.fragsize to IP6_MAX_MTU for UDP sockets to
avoid a possible overflow of UDP length that triggers a WARN in
udp_set_len_short when setsockopt IPV6_MTU_DISCOVER is set to
IPV6_PMTUDISC_DO or IPV6_PMTUDISC_PROBE, and a large packet is sent over
a netdev with an unusually large MTU.

Steps to reproduce (included in the new selftest):

1. Set device MTU bigger than IP6_MAX_MTU. cork->base.fragsize will be
   set to that MTU in ip6_setup_cork.
2. Set IPV6_MTU_DISCOVER to IPV6_PMTUDISC_PROBE or IPV6_PMTUDISC_DO. It
   lets maxnonfragsize be set to device MTU (cork->fragsize) in
   __ip6_append_data, rather than to IP6_MAX_MTU.
3. Send 65528 bytes of payload (+8 bytes of UDP header, +40 bytes of
   IPv6 header). Device MTU allows it (it's only one byte bigger than
   IP6_MAX_MTU, and the device MTU is bigger than that).
4. The UDP length in the built packet is 65536, which overflows the
   16-bit length field and triggers the WARN in udp_set_len_short.

To avoid breaking sending UDP jumbograms over raw IPv6 sockets, limit
the change to UDP sockets only.

The original overflow bug with IPv6 and IPV6_PMTUDISC_DO seems to
predate git history (verified reproduction on 2.6.21), was fixed later,
and then reappeared in commit 427faee167 ("net: ipv6: introduce
ip6_dst_mtu_maybe_forward"), which is chosen as the Fixes tag here. The
overflow with IPV6_PMTUDISC_PROBE reproduces since its introduction in
commit 628a5c5618 ("[INET]: Add IP(V6)_PMTUDISC_RPOBE").

Fixes: 427faee167 ("net: ipv6: introduce ip6_dst_mtu_maybe_forward")
Reported-by: syzbot+ce13c07d96d04716eaa2@syzkaller.appspotmail.com
Closes: https://lore.kernel.org/netdev/6a6a966c.86abc875.e5c3d.0054.GAE@google.com/
Signed-off-by: Alice Mikityanska <alice@isovalent.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260901195714.673548-3-alice.kernel@fastmail.im
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-07 17:13:55 -07:00
Alice Mikityanska
b83641e0ab net: ipv4: Fix UDP length overflow with PMTU discover and big MTU
This commit bounds cork->base.fragsize to IP_MAX_MTU to avoid a
possible overflow of UDP length that triggers a WARN in
udp_set_len_short when setsockopt IP_MTU_DISCOVER is set to
IP_PMTUDISC_PROBE, and a large packet is sent over a netdev with an
unusually large MTU.

Steps to reproduce:

1. Set device MTU bigger than IP_MAX_MTU + 20. cork->base.fragsize will
   be set to that MTU in ip_setup_cork.
2. Set IP_MTU_DISCOVER to IP_PMTUDISC_PROBE. It lets maxnonfragsize be
   set to device MTU (cork->fragsize) in __ip_append_data, rather than
   to IP_MAX_MTU.
3. Send 65528 bytes of payload (+8 bytes of UDP header, +20 bytes of
   IPv4 header). Device MTU allows it (it's only one byte bigger than
   IP_MAX_MTU + IPv4 header, and the device MTU is bigger than that).
4. The UDP length in the built packet is 65536, which overflows the
   16-bit length field and triggers the WARN in udp_set_len_short.

Note: IP_PMTUDISC_DO with IPv4 is safe, because ip_dst_mtu_maybe_forward
always clamps at IP_MAX_MTU, unlike ip6_dst_mtu_maybe_forward.

The Fixes tag points at the first commit where I could reproduce the
overflow with IPv4 and IP_PMTUDISC_PROBE.

Fixes: daba287b29 ("ipv4: fix DO and PROBE pmtu mode regarding local fragmentation with UFO/CORK")
Reported-by: syzbot+ce13c07d96d04716eaa2@syzkaller.appspotmail.com
Closes: https://lore.kernel.org/netdev/6a6a966c.86abc875.e5c3d.0054.GAE@google.com/
Signed-off-by: Alice Mikityanska <alice@isovalent.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260901195714.673548-2-alice.kernel@fastmail.im
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-07 17:13:55 -07:00
Jakub Kicinski
dead0c41db Merge branch 'af_unix-minor-fixes-for-msg_oob-and-msg_peek'
Kuniyuki Iwashima says:

====================
af_unix: Minor fixes for MSG_OOB and MSG_PEEK.

Fahad Alharbi reported blocking recv(MSG_PEEK) could hog CPU
due to OOB skb.

Patch 1 and 2 fixes the issues and Patch 3 adds tests.
====================

Link: https://patch.msgid.link/20260902202202.892676-1-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-07 16:58:46 -07:00
Kuniyuki Iwashima
ca0b0a8687 selftest: af_unix: Add zero-buffer test for msg_oob.c
The previous patches fixed two issues related to zero-length
buffer with MSG_PEEK for MSG_OOB skb.

Let's add corresponding tests in msg_oob.c.

Without this series:

  # FAILED: 50 / 60 tests passed.
  # Totals: pass:50 fail:10 xfail:0 xpass:0 skip:0 error:0

With this series:

  # PASSED: 60 / 60 tests passed.
  # Totals: pass:60 fail:0 xfail:0 xpass:0 skip:0 error:0

Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260902202202.892676-4-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-07 16:58:42 -07:00
Kuniyuki Iwashima
6e5ee08eb5 af_unix: Return immediately when manage_oob() returns NULL for 0-length buffer.
Fahad Alharbi reported that recv(0, MSG_PEEK) triggers busy-wait
in unix_stream_read_generic() if recv() is blocking and the last
skb in the queue is MSG_OOB skb.

In such a situation, TCP returns 0 immediately regardless of
blocking or non-blocking.

Let's follow the behaviour.

Fixes: 314001f0bf ("af_unix: Add OOB support")
Reported-by: Fahad Alharbi <fahad@codepure.com>
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260902202202.892676-3-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-07 16:58:42 -07:00
Kuniyuki Iwashima
94fd4debd2 af_unix: Update last skb marker in manage_oob().
Fahad Alharbi reported that blocking recv(MSG_PEEK) could hog CPU
due to OOB skb.

In the following cases, manage_oob() skips OOB skb(s) and returns
NULL for the last recv(MSG_PEEK):

  socketpair(AF_UNIX, SOCK_STREAM, 0, sk);

  1) skb -> OOB skb -> NULL
     send(sk[0], "ab", 2, MSG_OOB);
     recv(sk[1], buf, 0, MSG_PEEK);

  2) skb -> consumed OOB skb -> NULL
     send(sk[0], "ab", 2, MSG_OOB);
     recv(sk[1], buf, 1, MSG_OOB);
     recv(sk[1], buf, 0, MSG_PEEK);

  3) consumed OOB skb -> OOB skb -> NULL
     send(sk[0], "a", 1, MSG_OOB);
     recv(sk[1], buf, 0, MSG_OOB);
     send(sk[0], "b", 1, MSG_OOB);
     recv(sk[1], buf, 1, MSG_PEEK);

Then, @copied is 0 in unix_stream_read_generic() (zero-length buffer,
or non-OOB skb is not yet consumed), and unix_stream_data_wait() is
called.

However, it returns immediately because @last is not updated in
unix_stream_read_generic(), and the thread busy-waits for a new skb.

Let's update @last in manage_oob().

For MSG_PEEK, @last is updated with the skipped OOB, and for the
non-peek case, @last matches the returned value (when !copied)
because OOB is unlinked.

Note that manage_oob() is inlined and no stack canary is added.

Fixes: 22dd70eb2c ("af_unix: Don't peek OOB data without MSG_OOB.")
Reported-by: Fahad Alharbi <fahad@codepure.com>
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260902202202.892676-2-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-07 16:58:42 -07:00
Nagamani PV
74f27fc864 s390/qeth: allow bridgeport queries despite OS_MISMATCH
When HiperSockets interfaces on the same VCHID span different OS
families, reads of the sysfs attributes bridge_role and bridge_state
fail with -EPERM if bridge port ownership belongs to another OS family.

As a result, userspace tools such as 'lszdev -ii' cannot retrieve
bridge_role and bridge_state, even though firmware returns valid bridge
port data for QUERY_BRIDGE_PORTS requests.

The firmware reports IPA_RC_SBP_IQD_OS_MISMATCH (0x0010) to indicate
that bridge port ownership belongs to a different OS family. For
QUERY_BRIDGE_PORTS operations, firmware still returns valid bridge port
data (role=none, state=inactive) together with a primary return code of
0x0000 (success).

Allow QUERY_BRIDGE_PORTS requests to return the bridge port data
provided by the firmware despite OS_MISMATCH. To make the OS family
mismatch visible to userspace, represent the firmware-reported role
"none" as "none (OS family mismatch)" while preserving the reported
bridge_state.

The behavior for non-QUERY bridge port commands is unchanged; SET
operations continue to return -EPERM when another OS family owns the
bridge port.

This restores readability of bridge_role and bridge_state.

Fixes: 1b05cf6285 ("qeth: Include error message for "OS Mismatch"")
Cc: stable@vger.kernel.org
Suggested-by: Halil Pasic <pasic@linux.ibm.com>
Reviewed-by: Alexandra Winter <wintera@linux.ibm.com>
Signed-off-by: Nagamani PV <nagamani@linux.ibm.com>
Link: https://patch.msgid.link/20260901155344.3561483-1-nagamani@linux.ibm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-07 16:39:26 -07:00
Ilya Maximets
7a099b347f netfilter: report NLM_F_DUMP_FILTERED when all is filtered out
NLM_F_DUMP_FILTERED is only set on data elements in the conntrack dump.
But when everything is filtered out it is confusing for the user space,
since the flag is not reported anymore and it looks like the table was
empty, which may or may not be the case.

'answer_flags' were introduced precisely for this use case, and the
conntrack dump should set the flag in there in case the filtering was
applied.

This is important, for example, to be able to tell if the filters are
supported or not by the kernel without modifying the kernel state.

With the proper reporting of NLM_F_DUMP_FILTERED on NLMSG_DONE, an
application in user space can just try and dump with an arbitrary
filter without worrying that there could be no matching entry.  The
reported flag will signal that the filtering was applied and therefore
supported.

Fixes: cb8aa9a3af ("netfilter: ctnetlink: add kernel side filtering for dump")
Cc: stable@vger.kernel.org
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Florian Westphal <fw@strlen.de>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
2026-09-07 18:48:56 +02:00
Florian Westphal
da4afc5a95 netfilter: ip6_tables: set F_PROTO when proto value is nonzero
The ip6tables traverser doesn't search the extension header chain unless
userspace did set the IP6T_F_PROTO flag.

This also means that userspace that sets the e->ipv6.proto flag can bypass
the protocol check for the rule by not setting this flag.

That in turn means that all ip6_tables modules and targets that want to
reject rules without '-p' flag MUST also check for that flag.

Not all do, likely because they got copied from iptables which lacks
this flag (no extension headers).

Instead of fixing up all the relevant targets, emulate ip6tables behaviour
in the kernel (like nft_compat.c) and set the flag if the protocol is set.

Reported-by: Zhiling Zou <zhilinz@nebusec.ai>
Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Signed-off-by: Florian Westphal <fw@strlen.de>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
2026-09-07 18:48:56 +02:00
Florian Westphal
0bd7ed1a32 netfilter: arp_tables: remove the 32bit compat interface
This feature is required to use 32bit arptables binary on 64bit kernels.
It's already off in many distributions including Debian and Fedora for
many years.

Zap arptables first, it's the most esoteric of the 4 flavors.

Signed-off-by: Florian Westphal <fw@strlen.de>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
2026-09-07 18:48:52 +02:00
Florian Westphal
387d744fa7 netfilter: nfnetlink_log: cope with concurrent instance destruction
Instances are refcounted. However, only memory release happens on the
1 -> 0 transition; the unlink from hashes can occur with any refcount.

Uncooperative userspace can force a situation where a queue is pending
for destruction from netlink event while a different socket with same
portid processes an UNBIND request.

With right timing, this will unhash the instance again:

Oops: general protection fault, [..]
Call Trace:
 <TASK>
 nfulnl_recv_config+0x31a/0xd50
 nfnetlink_rcv_msg+0x7c2/0xeb0

Fixes: 0597f2680d ("[NETFILTER]: Add new "nfnetlink_log" userspace packet logging facility")
Reported-by: Eulgyu Kim <eulgyukim@snu.ac.kr>
Reported-by: Jaeyoung Chung <jjy600901@snu.ac.kr>
Signed-off-by: Florian Westphal <fw@strlen.de>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
2026-09-07 18:43:16 +02:00
Vineeth Karumanchi
38b6be1010 net: macb: fix NULL pointer dereference on unbind with fixed-link
When the device tree describes a fixed-link and has no "mdio" child
node, macb_mii_init() returns early without allocating the MDIO bus,
leaving bp->mii_bus as NULL.

Two cleanup paths then dereference this NULL bus:

1. On driver unbind, macb_remove() unconditionally calls
   mdiobus_unregister(bp->mii_bus), which oopses:

  Unable to handle kernel NULL pointer dereference at virtual address 00000000000004a8
  pc : mdiobus_unregister+0x14/0xa4
  lr : macb_remove+0x38/0xa4
  Call trace:
   mdiobus_unregister+0x14/0xa4 (P)
   macb_remove+0x38/0xa4
   platform_remove+0x20/0x30
   device_release_driver_internal+0x1c8/0x224
   unbind_store+0xb4/0xbc

2. On the probe error path in macb_probe(), reached when
   macb_mii_init() has succeeded but a subsequent step fails, the
   err_out_unregister_mdio label runs the same unconditional cleanup.

mdiobus_unregister() and mdiobus_free() do not guard against a NULL
bus, so guard the calls in both macb_remove() and the probe error
path.

Fixes: d0c3601f2c ("net: macb: Avoid 20s boot delay by skipping MDIO bus registration for fixed-link PHY")
Signed-off-by: Vineeth Karumanchi <vineeth.karumanchi@amd.com>
Reviewed-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Link: https://patch.msgid.link/20260902102836.2019355-1-vineeth.karumanchi@amd.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:58:52 -07:00
Jakub Kicinski
e7c93ad4bd Merge branch 'net-sched-clamp-quantum-psched_mtu-in-change-paths'
Jamal Hadi Salim says:

====================
net/sched: clamp quantum/psched_mtu in change paths

This is a followup to commit 709f34f7c2 ("net/sched: fq: add overflow
bounds to quantum and initial quantum").

The quantum_backlog_overflow series and the five siblings that followed
clamped the init-path quantum in fq, fq_codel, fq_pie, hhf, sfq. The
change() paths were not clamped but it is the same pattern, same writer
of q->quantum, same privilege level (CAP_NET_ADMIN in a user namespace).
A user can override the init clamp via tc qdisc change, restoring the
small-quantum deficit spin that the init clamp was meant to prevent.

This series also covers two siblings that were missed entirely by the
original series: sch_dualpi2 and sch_pie call psched_mtu() without any
clamp at all. With a crafted size table qdisc_pkt_len reaches ~2 GiB,
so quantum=1 (or a zero psched_mtu on a headerless device) makes the
deficit-refill loop spin ~2^31 times under the qdisc lock (a soft
lockup / denial of service).

Each patch fixes one qdisc with its own Fixes: tag so they can be
backported independently - the commits they fix shift differently in
the git tree.

Patch 1: fq - clamp TCA_FQ_QUANTUM and TCA_FQ_INITIAL_QUANTUM in change
Patch 2: fq_pie - clamp quantum in change path
Patch 3: sfq - clamp quantum and reject > 1<<20 in change path
Patch 4: hhf - clamp quantum in change and init paths
Patch 5: dualpi2 - clamp psched_mtu at all 3 call sites
Patch 6: pie - clamp psched_mtu in pie_drop_early
Patch 7: drr - clamp quantum in change class
Patch 8: ets - clamp quantum in parse and fallback paths
Patch 9: selftests - update ETS test 41f5 for clamped quanta

Conditions to recreate (applies to all): create the qdisc, then
tc qdisc change ... quantum 1 with a STAB size table inflating
qdisc_pkt_len. Requires CAP_NET_ADMIN in a user namespace (unshare -Urn).
====================

Link: https://patch.msgid.link/QDISC-0CFC.v3.20260901204856@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:48:08 -07:00
Jamal Hadi Salim
8f0229bef3 selftests: tc-testing: update ETS test 41f5 for clamped quanta
Commit "net/sched: ets: clamp quantum in parse and fallback paths"
moved the quantum floor into ets_quantum_parse(), so every explicitly
configured quantum is now clamped to [256, 1 << 20], not just the
psched_mtu() fallback.

Test 41f5 passes "quanta 4294967294 1 1" and matches the values back
verbatim, so all three bands now differ from what it expects:

  before: bands 3 quanta 4294967294 1 1
  after:  bands 3 quanta 1048576 256 256

Update the match pattern accordingly.

Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-0CFC.v3.20260901204856@mojatatu.com.10
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:48:08 -07:00
Jamal Hadi Salim
1c38487f46 net/sched: ets: clamp quantum in parse and fallback paths
ets_qdisc_change() falls back to psched_mtu() with no floor for bands
without an explicit quantum. With a crafted size table qdisc_pkt_len
reaches ~2 GiB, so a zero psched_mtu on a headerless device makes the
deficit-refill loop spin under the qdisc lock.

Move the floor into ets_quantum_parse() so explicitly configured quanta
are also clamped to [256, 1<<20], not just the fallback path.

Conditions to recreate the bug:
  CONFIG_NET_SCH_ETS=y. Requires CAP_NET_ADMIN (namespace-local via
  unshare -Urn suffices).

  tc qdisc add dev dummy0 root ets bands 3 strict 2 quanta 1 1

Fixes: dcc68b4d80 ("net: sch_ets: Add a new Qdisc")
Reported-by: Vega <vega@nebusec.ai>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-0CFC.v3.20260901204856@mojatatu.com.9
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:48:08 -07:00
Jamal Hadi Salim
8382abec0f net/sched: drr: clamp quantum in change class
drr_change_class() rejects explicit quantum==0 but falls back to
psched_mtu() with no floor. With a crafted size table qdisc_pkt_len
reaches ~2 GiB, so quantum=1 (or a zero psched_mtu on a headerless
device) makes the deficit-refill loop spin under the qdisc lock.

Add clamp_t(u32, quantum, 256, 1<<20) after the zero reject and on the
fallback path. The explicit-zero reject is preserved.

Conditions to recreate the bug:
  CONFIG_NET_SCH_DRR=y. Requires CAP_NET_ADMIN (namespace-local via
  unshare -Urn suffices).

  tc qdisc add dev dummy0 root drr
  tc class add dev dummy0 parent 1: classid 1:1 drr quantum 1

Fixes: 13d2a1d2b0 ("pkt_sched: add DRR scheduler")
Reported-by: Vega <vega@nebusec.ai>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-0CFC.v3.20260901204856@mojatatu.com.8
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:48:08 -07:00
Jamal Hadi Salim
54370e44c0 net/sched: pie: clamp psched_mtu in pie_drop_early
pie_drop_early() calls psched_mtu() with no clamp. With mtu=0x80000000
the bytemode divide silently zeroes the drop probability, disabling AQM.
Clamp to [1, 1<<20].

Conditions to recreate the bug:
  CONFIG_NET_SCH_PIE=y. Requires CAP_NET_ADMIN (namespace-local via
  unshare -Urn suffices).

  tc qdisc add dev dummy0 root pie
  tc qdisc change dev dummy0 root pie stab data 32768 size_log 15 cell_log 0

Fixes: d4b36210c2 ("net: pkt_sched: PIE AQM scheme")
Reported-by: Vega <vega@nebusec.ai>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-0CFC.v3.20260901204856@mojatatu.com.7
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:48:08 -07:00
Jamal Hadi Salim
3c01f1ca5d net/sched: dualpi2: clamp psched_mtu at all call sites
dualpi2_calculate_c_protection(), must_drop(), and get_memory_limit()
call psched_mtu() with no clamp. A huge MTU makes (s32)psched_mtu()
overflow in the signed multiply for c_protection_init, and 2 *
psched_mtu() wraps in get_memory_limit(). With a crafted size table
qdisc_pkt_len reaches ~2 GiB, causing a soft lockup / denial of service.

Clamp psched_mtu() to [1, 1<<20] at all three call sites.

Conditions to recreate the bug:
  CONFIG_NET_SCH_DUALPI2=y. Requires CAP_NET_ADMIN (namespace-local via
  unshare -Urn suffices).

  tc qdisc add dev dummy0 root dualpi2
  tc qdisc change dev dummy0 root dualpi2 stab data 32768 size_log 15 cell_log 0

Fixes: 320d031ad6 ("sched: Struct definition and parsing of dualpi2 qdisc")
Reported-by: Vega <vega@nebusec.ai>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-0CFC.v3.20260901204856@mojatatu.com.6
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:48:02 -07:00
Jamal Hadi Salim
eb56a495f5 net/sched: hhf: clamp quantum in change and init paths
hhf_change() accepts any quantum from userspace, including 1. With a
crafted size table qdisc_pkt_len reaches ~2 GiB, so quantum=1 makes
the deficit-refill loop spin ~2^31 times under the qdisc lock
(a soft lockup / denial of service).

Add max(256U, ...) in hhf_change() matching fq_codel_change(). Clamp
hhf_init() to [256, 1<<20] matching the siblings, and remove the old
fallback that only set quantum=256 on overflow.

Conditions to recreate the bug:
  CONFIG_NET_SCH_HHF=y. Requires CAP_NET_ADMIN (namespace-local via
  unshare -Urn suffices).

  tc qdisc add dev dummy0 root hhf
  tc qdisc change dev dummy0 root hhf quantum 1 stab data 32768 size_log 15 cell_log 0

Fixes: 10239edf86 ("net-qdisc-hhf: Heavy-Hitter Filter (HHF) qdisc")
Reported-by: Vega <vega@nebusec.ai>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-0CFC.v3.20260901204856@mojatatu.com.5
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:47:30 -07:00
Jamal Hadi Salim
fb9f88a33c net/sched: sfq: clamp quantum in change path
sfq_change() accepts any non-negative quantum (only rejects
(int)ctl->quantum < 0). With a crafted size table qdisc_pkt_len reaches
~2 GiB, so quantum=1 makes the deficit-refill loop spin ~2^31 times
under the qdisc lock (a soft lockup / denial of service).

Add max(256U, ...) matching fq_codel_change(). Reject quantum > 1<<20
with -EINVAL, matching fq_codel_change() and the init clamp.

Conditions to recreate the bug:
  CONFIG_NET_SCH_SFQ=y. Requires CAP_NET_ADMIN (namespace-local via
  unshare -Urn suffices).

  tc qdisc add dev dummy0 root sfq
  tc qdisc change dev dummy0 root sfq quantum 1 stab data 32768 size_log 15 cell_log 0

Fixes: e4650d7ae4 ("net_sched: sch_sfq: handle bigger packets")
Reported-by: Vega <vega@nebusec.ai>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-0CFC.v3.20260901204856@mojatatu.com.4
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:47:30 -07:00
Jamal Hadi Salim
4864f58c53 net/sched: fq_pie: clamp quantum in change path
fq_pie_change() accepts any quantum value from userspace, including 1.
With a crafted size table qdisc_pkt_len reaches ~2 GiB, so quantum=1
makes the deficit-refill loop spin ~2^31 times under the qdisc lock
(a soft lockup / denial of service).

Add max(256U, ...) matching fq_codel_change().

Conditions to recreate the bug:
  CONFIG_NET_SCH_FQ_PIE=y. Requires CAP_NET_ADMIN (namespace-local via
  unshare -Urn suffices).

  tc qdisc add dev dummy0 root fq_pie
  tc qdisc change dev dummy0 root fq_pie quantum 1 stab data 32768 size_log 15 cell_log 0

Fixes: ec97ecf1eb ("net: sched: add Flow Queue PIE packet scheduler")
Reported-by: Vega <vega@nebusec.ai>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-0CFC.v3.20260901204856@mojatatu.com.3
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:47:29 -07:00
Jamal Hadi Salim
094cc07f98 net/sched: fq: clamp quantum and initial_quantum in change path
The fq change path accepts TCA_FQ_QUANTUM in [1, INT_MAX] and
TCA_FQ_INITIAL_QUANTUM up to INT_MAX, while fq_init() already clamps to
[1, 1<<20]. A user can override the init clamp via tc qdisc change,
restoring the small-quantum deficit spin that the init clamp prevents.

Narrow iq_range.max to 1<<20 so TCA_FQ_INITIAL_QUANTUM is rejected at
parse time. Clamp TCA_FQ_QUANTUM to [256, 1<<20] in fq_change() and
fq_init() quantum to [256, 1<<20] for tiny-MTU devices.

Conditions to recreate the bug:
  CONFIG_NET_SCH_FQ=y. Requires CAP_NET_ADMIN (namespace-local via
  unshare -Urn suffices).

  tc qdisc add dev dummy0 root fq
  tc qdisc change dev dummy0 root fq quantum 1 stab data 32768 size_log 15 cell_log 0

Fixes: 709f34f7c2 ("net/sched: fq: add overflow bounds to quantum and initial quantum")
Reported-by: Vega <vega@nebusec.ai>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-0CFC.v3.20260901204856@mojatatu.com.2
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:47:29 -07:00
Carolina Jubran
df99553f84 net/mlx5e: Keep HW timestamp stats monotonic across reconfiguration
`mlx5e_stats_ts_get()` currently selects either DMA or port timestamp
counters based on `tx_ptp_opened`. This flag is intentionally kept set
once the PTP TX queues have been opened so their statistics remain
available after queue teardown. As a result, DMA timestamps are no
longer reported after switching from port timestamping back to DMA
timestamping.

The function also reads statistics only from the currently active
channels and TCs. Reducing the number of channels or TCs can therefore
drop previously accumulated timestamp counters from the reported value.

Read the persistent channel statistics instead and always include DMA
timestamp counters. Once the PTP TX queues have been opened, also
include the port timestamp counters.

This also drops state_lock. It previously protected live channel/PTP
pointers, the new code only reads persistent channel_stats and
ptp_stats via mlx5e_stats_nch_read(), which is already safe for
lockless stats access.

Fixes: 3579032c08 ("net/mlx5e: Implement ethtool hardware timestamping statistics")
Signed-off-by: Carolina Jubran <cjubran@nvidia.com>
Reviewed-by: Shahar Shitrit <shshitrit@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902193731.3668958-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:23:37 -07:00
Lama Kayal
c0c6f4ba8a net/mlx5: E-Switch, prevent mc_list repopulation during vport disable
In mlx5_esw_vport_disable(), move esw_apply_vport_rx_mode() ahead
of esw_vport_change_handle_locked() so vport->allmulti_rule is
NULL before the change handler observes it.

During FW-fatal recovery the disable runs while dev->state ==
INTERNAL_ERROR. The promisc query inside esw_update_vport_rx_mode()
fails and returns early, leaving vport->allmulti_rule intact, so
esw_update_vport_mc_promisc() runs and adds MLX5_ACTION_ADD entries
to vport->mc_list whose flow rules are then installed in the FDB
by esw_add_mc_addr(). esw_destroy_legacy_table() tears down the
FDB with those refs still held, corrupting the sub-tree and
leaving dangling flow_rule pointers in vport->mc_list.

Two-stage failure on `echo 1 > /sys/bus/pci/devices/<bdf>/reset`:

  refcount_t: underflow; use-after-free.
   tree_put_node+0xef/0x110 [mlx5_core]
   clean_tree+0x44/0xd0 [mlx5_core] (x5)
   mlx5_fs_core_cleanup+0x57/0x1c0 [mlx5_core]
   mlx5_unload+0x65/0xd0 [mlx5_core]
   ... mlx5_health_try_recover

  BUG: unable to handle page fault for address: 0000000003000055
   down_write+0x1c/0x60
   mlx5_del_flow_rules+0x33/0x1f0 [mlx5_core]
   esw_del_mc_addr+0x7b/0x170 [mlx5_core]
   esw_apply_vport_addr_list+0x56/0xf0 [mlx5_core]
   esw_vport_change_handle_locked+0x28b/0x310 [mlx5_core]
   mlx5_esw_vport_enable+0x270/0x4a0 [mlx5_core]
   ... mlx5_load ... mlx5_health_try_recover

esw_apply_vport_rx_mode(false, false) clears vport->allmulti_rule
via its local state machine even when the FW del fails. With the
rule NULL the !IS_ERR_OR_NULL(allmulti_rule) gate in the change
handler closes, no rules are installed during disable, and the
reload starts with a clean mc_list.

Fixes: 922f56e9a7 ("net/mlx5: Fix steering rules cleanup")
Signed-off-by: Lama Kayal <lkayal@nvidia.com>
Reviewed-by: Cosmin Ratiu <cratiu@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902193854.3669035-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:21:43 -07:00
Yael Chemla
7ee07f601f net/mlx5: E-Switch: fix use-after-free in mlx5_eswitch_termtbl_put
In mlx5_eswitch_termtbl_put(), the zero-ref cleanup check reads
tt->ref_count after termtbl_mutex has been released.  Two concurrent
callers on the same mlx5_termtbl_handle race: one decrements ref_count
to zero, removes the hash entry, and calls kfree(tt) while the other
has already dropped the mutex and is about to evaluate
if (!tt->ref_count), producing a use-after-free.

Fix this by capturing the result of the decrement into a stack-local
last variable before dropping the mutex.  The cleanup decision is now
made entirely under termtbl_mutex, and tt is not touched after
kfree.

Fixes: 10caabdaad ("net/mlx5e: Use termination table for VLAN push actions")
Signed-off-by: Yael Chemla <ychemla@nvidia.com>
Reviewed-by: Dragos Tatulea <dtatulea@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902193514.3668880-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:20:26 -07:00
Carolina Jubran
af3aef0245 net/mlx5e: Fix use-after-free race in sample_restore_put()
Concurrent teardown of TC sample rules sharing the same restore
context may re-read restore->count after dropping restore_lock.
At that point another thread may already have completed cleanup and
freed the restore object.

Use the result of the refcount decrement while holding restore_lock to
determine whether cleanup is needed.

Fixes: 36a3196256 ("net/mlx5e: TC, Add sampler restore handle API")
Signed-off-by: Carolina Jubran <cjubran@nvidia.com>
Reviewed-by: Shahar Shitrit <shshitrit@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902193341.3668809-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:19:22 -07:00
Carolina Jubran
e7ee897408 net/mlx5e: Fix ETS zero BW reporting when one TC holds 100%
When ETS TCs with zero bandwidth are configured, the driver programs the
firmware using an alternate representation. On get, it needs
to recognize that representation so those TCs can be translated back and
reported as 0% bandwidth.

The existing detection relied on the programmed bandwidth because it was
enough to identify this representation. However, when a single ETS TC
owns 100% of the bandwidth, its firmware representation becomes the
same as a strict-priority TC, causing zero-bandwidth ETS TCs to be
reported with non-zero bandwidth values.

Use the cached TSA instead to distinguish the ETS and strict-priority
cases.

Fixes: be0f161ef1 ("net/mlx5e: DCBNL, Implement tc with ets type and zero bandwidth")
Signed-off-by: Carolina Jubran <cjubran@nvidia.com>
Reviewed-by: Alex Lazar <alazar@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Link: https://patch.msgid.link/20260902193224.3668743-1-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-05 13:18:47 -07:00