Commit Graph

1482472 Commits

Author SHA1 Message Date
Xuanqiang Luo
14c5eb685c selftests: tc-testing: test action batch deletion failure cleanup
Add tests for cleanup after a batched RTM_DELACTION request fails at
a gact action bound to a filter. Check that subsequent actions retain
their original reference counts and that earlier successful deletions
are preserved.

Cover failures at the first and middle entries. Verify that a remaining
unbound action can be removed with one subsequent delete.

Signed-off-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Link: https://patch.msgid.link/20260910093413.34509-3-xuanqiang.luo@linux.dev
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 17:03:09 -07:00
Xuanqiang Luo
6e05e46fa8 net/sched: act_api: release tail references on DELACTION failure
A batched RTM_DELACTION request takes a temporary reference on each
action before attempting any deletion. tcf_action_delete() clears
each processed slot and drops its temporary reference before attempting
the deletion. If deletion fails, tca_action_gd() calls
tcf_action_put_many() to release the remaining references, but its
tcf_act_for_each_action() iterator stops at the first NULL slot.

When a batch stops at an action bound to a filter, this leaks a
reference on each subsequent action. A later delete of an unbound
action can then return success without removing it from the IDR.

Walk the full array in tcf_action_put_many() and skip NULL slots to
release the references held on the unprocessed actions.

Fixes: a0e947c9cc ("net/sched: act_api: avoid non-contiguous action array")
Cc: stable@vger.kernel.org
Signed-off-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Link: https://patch.msgid.link/20260910093413.34509-2-xuanqiang.luo@linux.dev
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 17:03:09 -07:00
Lorenzo Bianconi
15989abd74 net: stmmac: fix TSO header length truncation
stmmac_tso_xmit() stores the protocol header length returned by
stmmac_tso_header_size() in a u8. stmmac_tso_valid_packet() admits
headers up to 1023 bytes, so a header longer than 255 bytes wraps modulo
256 (486 becomes 230, 256 becomes 0).

A TCP over IPv6 socket carrying a few hundred bytes of sticky
destination/hop-by-hop options makes skb_tcp_all_headers() exceed 255
while staying below the 1023-byte limit, so such an skb reaches
stmmac_tso_xmit().

Widen proto_hdr_len to unsigned int, which is sufficient since the value
is bounded by the hardware limit, and adjust the debug print specifier
accordingly.

Fixes: 9edfa7dab8 ("net: stmmac: enable TSO for IPv6")
Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/20260911-stmmac-fix-header-length-v1-1-8fc103334327@oss.qualcomm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 16:55:52 -07:00
Jakub Kicinski
433cfc3025 Merge branch 'af_unix-fix-inconsistent-scc_index'
Kuniyuki Iwashima says:

====================
af_unix: Fix inconsistent scc_index.

James Burton reported that a single SCC could have multiple
scc_index and unix_vertex_dead() fails to detect a dead SCC.

Patch 1 fixes it and Patch 2 adds a test case.
====================

Link: https://patch.msgid.link/20260912030852.1467872-1-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 16:45:07 -07:00
Kuniyuki Iwashima
b645ccd410 selftest: af_unix: Add test case with mixed lowpoint in scm_rights.c.
The new test case creates two SCCs so that each of them
has multiple scc_index.

Without patch, GC cannot free the sockets and the test fails.

  #  RUN           scm_rights.dgram.mixed_lowpoints ...
  # scm_rights.c:176:mixed_lowpoints:Expected 0 (0) == ret (12)
  # mixed_lowpoints: Test terminated by assertion
  #          FAIL  scm_rights.dgram.mixed_lowpoints
  not ok 5 scm_rights.dgram.mixed_lowpoints
  ...
  # FAILED: 45 / 50 tests passed.
  # Totals: pass:45 fail:5 xfail:0 xpass:0 skip:0 error:0

With the patch, all tests pass.

  # PASSED: 50 / 50 tests passed.
  # Totals: pass:50 fail:0 xfail:0 xpass:0 skip:0 error:0

Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260912030852.1467872-3-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 16:45:05 -07:00
Kuniyuki Iwashima
4a4263dfea af_unix: Unify scc_index when finalising SCC in __unix_walk_scc().
Commit bfdb01283e ("af_unix: Assign a unique index to SCC.")
changed Tarjan's algorithm to update lowlink with lowlink,
which is called lowpoint (unix_vertex.scc_index).

unix_vertex_dead() assumes all vertices in an SCC share the same
lowpoint, but this is not always true if an SCC has two or more
back edges, depending on the order of DFS.

For example, the graph below has two back edges from B to A
and from C to B.

  A --> B --> C
  ^    | ^    |
  `----' `----'

If DFS walks through A -> B -> C -> B (-> C -> B) -> A (-> B -> A),
each index and scc_index will be updated as follows.

  A --> B --> C    C = (3, 3)  (index, scc_index)
                   B = (2, 2)
                   A = (1, 1)

  A ... B ... C    C = (3, 2)<-.
         ^    |    B = (2, 2) -'
         `----'    A = (1, 1)

  A ... B ... C    C = (3, 2)
  ^    | .    .    B = (2, 1)<-.
  `----'  ....     A = (1, 1) -'

Then, unix_vertex_dead() thinks that B is passed to another
SCC with scc_index 2, and the SCC is not garbage-collected.

This does not happen if DFS walks in a different order below
or starts from B.

    1      3
  A --> B --> C
  ^    | ^    |
  `----' `----'
     2      4

Let's unify scc_index across the SCC when finalising it.

Note that updating v->index was previously done in unix_scc_dead(),
when called from __unix_walk_scc(), just to save one loop.  Since
__unix_walk_scc() now iterates over the SCC anyway, the update is
moved back to __unix_walk_scc() and 'fast' argument is dropped.

Fixes: 4090fa373f ("af_unix: Replace garbage collection algorithm.")
Reported-by: James Burton <jamesburton@meta.com>
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260912030852.1467872-2-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 16:45:05 -07:00
Jakub Kicinski
c5e367a8a3 bluetooth pull request for net:
Core:
 
  - hci: put the peer's on-air address on air when we cannot resolve
  - hci: keep dst_type with dst when reusing an LE connection
  - hci_core: Fix queuing tx_work after workqueue is drained
  - hci_sync: Serialize local codec list cleanup
  - hci_codec: validate vendor codec count length
  - eir: validate service data length before reading UUID
  - RFCOMM: avoid socket lock inversion in listener cleanup
  - ISO: Fix parent socket leak in iso_conn_ready()
  - ISO: set BT_LISTEN before requesting a BIG sync
  - coredump: Quiesce dump work on unregister
 
 Drivers:
 
  - btintel_pcie: validate TX skb length in send_sync
  - btmtk: fix wrong status for short WMT FUNC_CTRL events
  - btmtksdio, btmtkuart: validate WMT event length before struct access
  - hci_qca: Do not write to the serial port after it is closed
  - btusb: fix NXP IW610 composite device handling
  - btintel_pcie: fix off-by-one bounds check in RX submit
  - btmtksdio: Fix PM runtime reference leak in shutdown
 -----BEGIN PGP SIGNATURE-----
 
 iQJNBAABCgA3FiEE7E6oRXp8w05ovYr/9JCA4xAyCykFAmqpmw8ZHGx1aXoudm9u
 LmRlbnR6QGludGVsLmNvbQAKCRD0kIDjEDILKZC2D/kBM2/ZEFhSwp6Lu7orXKHn
 AQOwoh0uv4TBZaR8V6ogSLdWmdv4rb1Ayjj2SFkNNNwmuUkPYhG30DySw73C7QpV
 uza4SAPO9OF3RRzXvBiVpxF4TmlWYMbApAuMtaa/NmZfZ38pv6T66cFsxhc+lfje
 qNUWkCyI9/4hY7rIxL2dcBQ3yLQwMYwKCwdMwj2/rMX4BN7IXVsMqZ4HBHmCx3Ht
 M2GGnng4DNFgCN/ZcUMLASRR61At3Iop39E8NLnnnVqJgn6p2DOx+G+gnYpdnUOy
 zdpm/yLSJD2maDVRlxVWlf071jpUVCjCY2HvjnvnBTKDgqf9GD/3SVTrv3f/3gxt
 /NVHwnHjZkhMt2sY//NVo4dGD0s8PJjEAdWzm/c7lZ97/8AHsSFMtBe53Jc/xy9l
 lvmTV3FrTkzf08VT5QyrTjj9dQjOcJfeOqfOHJrRQSfjqvNCR7OZjjDLavgv+3q/
 Op5bE0RY/47qa7yWE7gyjyVFeAoK2C4XsXDIc8cHOdJlozlqKV0jVQUvs9+wbl8r
 0ULFdcRAs3/OIybYBW4qhOGjZjS/ilK4CjCNGjEPMD1IEz6SutJPxtp1Cmh9KFZT
 3u5vNsbXB/Oa5yM+m5KZwxoRhYG6aCivCLruEWVe8tco8nWlcz1iPRPBs+bB2cvX
 5tJwh5QCT35FUBo3j+0PvQ==
 =dGOw
 -----END PGP SIGNATURE-----

Merge tag 'for-net-2026-09-15' of git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth

Luiz Augusto von Dentz says:

====================
bluetooth pull request for net:

Core:

 - hci: put the peer's on-air address on air when we cannot resolve
 - hci: keep dst_type with dst when reusing an LE connection
 - hci_core: Fix queuing tx_work after workqueue is drained
 - hci_sync: Serialize local codec list cleanup
 - hci_codec: validate vendor codec count length
 - eir: validate service data length before reading UUID
 - RFCOMM: avoid socket lock inversion in listener cleanup
 - ISO: Fix parent socket leak in iso_conn_ready()
 - ISO: set BT_LISTEN before requesting a BIG sync
 - coredump: Quiesce dump work on unregister

Drivers:

 - btintel_pcie: validate TX skb length in send_sync
 - btmtk: fix wrong status for short WMT FUNC_CTRL events
 - btmtksdio, btmtkuart: validate WMT event length before struct access
 - hci_qca: Do not write to the serial port after it is closed
 - btusb: fix NXP IW610 composite device handling
 - btintel_pcie: fix off-by-one bounds check in RX submit
 - btmtksdio: Fix PM runtime reference leak in shutdown

* tag 'for-net-2026-09-15' of git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth:
  Bluetooth: RFCOMM: avoid socket lock inversion in listener cleanup
  Bluetooth: keep dst_type with dst when reusing an LE connection
  Bluetooth: btintel_pcie: fix off-by-one bounds check in RX submit
  Bluetooth: btmtksdio: Fix PM runtime reference leak in shutdown
  Bluetooth: btmtksdio, btmtkuart: validate WMT event length before struct access
  Bluetooth: btmtk: fix wrong status for short WMT FUNC_CTRL events
  Bluetooth: ISO: set BT_LISTEN before requesting a BIG sync
  Bluetooth: ISO: Fix parent socket leak in iso_conn_ready()
  Bluetooth: hci_sync: Serialize local codec list cleanup
  Bluetooth: hci_qca: Do not write to the serial port after it is closed
  Bluetooth: hci_codec: validate vendor codec count length
  Bluetooth: put the peer's on-air address on air when we cannot resolve
  Bluetooth: coredump: Quiesce dump work on unregister
  Bluetooth: btintel_pcie: validate TX skb length in send_sync
  Bluetooth: hci_core: Fix queuing tx_work after workqueue is drained
  Bluetooth: eir: validate service data length before reading UUID
  Bluetooth: btusb: fix NXP IW610 composite device handling
====================

Link: https://patch.msgid.link/20260915192441.1130583-1-luiz.dentz@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 16:40:55 -07:00
Juan Perdomo
801fb950ca Bluetooth: RFCOMM: avoid socket lock inversion in listener cleanup
rfcomm_sock_cleanup_listen() closes unaccepted child sockets through
rfcomm_sock_close(), which takes the child socket lock before
rfcomm_dlc_close() acquires rfcomm_mutex. The RFCOMM worker takes these
locks in reverse order while handling connections and DLC state changes,
so lockdep reports a possible deadlock.

Close dequeued children without taking their socket lock. The accept queue
owns a reference to each child, and bt_accept_dequeue() locks the child
while unlinking it and clearing its parent pointer.

Dropping the child lock makes it important to prevent a concurrent
rfcomm_connect_ind() from enqueueing a new child after cleanup observes an
empty queue. Set a listening socket to BT_CLOSED while its lock is still
held, before dropping the lock and draining the queue. The state check in
rfcomm_connect_ind() then rejects new children once cleanup starts.

Reported-by: syzbot+0cece8fa7d83523f47a3@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=0cece8fa7d83523f47a3
Fixes: b7ce436a5d ("Bluetooth: switch to lock_sock in RFCOMM")
Signed-off-by: Juan Perdomo <jcperdomo100@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:55:41 -04:00
Radek Podgorny
555cd2bd86 Bluetooth: keep dst_type with dst when reusing an LE connection
hci_connect_le() swaps the caller's identity address for the peer's
cached RPA when one is known, and stamps the matching
ADDR_LE_DEV_RANDOM on the local dst_type. On the conn-reuse path only
the address is copied into the connection:

  if (conn) {
          bacpy(&conn->dst, dst);

so conn->dst ends up holding an RPA while conn->dst_type still names the
identity it was resolved from, and hci_le_create_conn_sync() puts that
pair on air unchanged. An RPA declared as a public address is not
something any peer can answer.

Measured on a CYW43438 against a peer advertising an RPA the host holds
the IRK for, connecting to the identity address over a raw L2CAP socket.
The first attempt creates the connection, the second takes the reuse
path:

  LE Create Connection  3C:78:95:78:37:C3  type public
  LE Create Connection  5B:75:A2:26:D6:18  type public
  LE Connection Complete: Unknown Connection Identifier (0x02)

The second address is the peer's RPA. btmon annotates it with an OUI
lookup rather than "(Resolvable)" precisely because the command declares
it public; the same bit pattern annotates as resolvable once the type is
right.

The mistyped pair is also why nothing downstream repairs it.
hci_bdaddr_is_rpa() tests the type before the address, so an RPA carrying
a public type is not recognised as one, and hci_find_irk_by_addr() then
searches for an identity address that does not match it either.

Copy the type along with the address.

The assignment used to be unconditional just below this block and covered
both paths; it moved into hci_conn_add_unset(), which the reuse path does
not go through.

Cc: stable@vger.kernel.org
Fixes: 14b06c3a88 ("Bluetooth: HCI: Always use the identity address when initializing a connection")
Assisted-by: Claude:claude-opus-5
Signed-off-by: Radek Podgorny <radek@podgorny.cz>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:55:32 -04:00
Sai Teja Aluvala
2ea5a87a5a Bluetooth: btintel_pcie: fix off-by-one bounds check in RX submit
btintel_pcie_submit_rx() used frbd_index > rxq->count to guard the
FRBD array access, allowing frbd_index == rxq->count to pass through
and index one element past the end of the array. Change the check to
>= rxq->count so every out-of-range index is rejected.

This issue was reported by Claude Mythos.

Fixes: c2b636b3f7 (Bluetooth: btintel_pcie: Add support for PCIe transport)
Signed-off-by: Sai Teja Aluvala <aluvala.sai.teja@intel.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:55:27 -04:00
Tzung-Bi Shih
7b60ee5f46 Bluetooth: btmtksdio: Fix PM runtime reference leak in shutdown
In btmtksdio_shutdown(), pm_runtime_get_sync() is called at the
beginning of the function.  However, if sending the WMT function
control command fails later, the driver returns early.

It bypasses the corresponding pm_runtime_put_noidle() and
pm_runtime_disable() calls, leaking the PM usage counter and leaving PM
runtime enabled indefinitely.

Fall through to execute the PM runtime cleanup block even if WMT errors.

Fixes: 7f3c563c57 ("Bluetooth: btmtksdio: Add runtime PM support to SDIO based Bluetooth")
Signed-off-by: Tzung-Bi Shih <tzungbi@kernel.org>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:55:01 -04:00
Chris Lu
8879e3e0a8 Bluetooth: btmtksdio, btmtkuart: validate WMT event length before struct access
btmtksdio.c and btmtkuart.c cast a received WMT event straight to
struct btmtk_hci_wmt_evt and read its op/flag fields without checking
the event is long enough to contain them, unlike btmtk.c. The
FUNC_CTRL case then further casts to struct btmtk_hci_wmt_evt_funcc
and reads its 2-byte status field, again without a length check.
Firmware that sends a short or malformed WMT event makes both drivers
read past the end of the received SKB.

Mirror btmtk.c: validate the base WMT header with skb_pull_data()
before touching any of its fields, and when a FUNC_CTRL event turns
out to be the short, header-only form (a plain enable/disable ack
with no status word), decode the result from the header's own flag
byte instead (0 = success, otherwise failure).

Verified setup on MT7920, MT7921, MT7922 and MT7925: no regression.

Fixes: 9aebfd4a22 ("Bluetooth: mediatek: add support for MediaTek MT7663S and MT7668S SDIO devices")
Fixes: e0b67035a9 ("Bluetooth: mediatek: update the common setup between MT7622 and other devices")
Assisted-by: Claude:claude-opus-5
Signed-off-by: Chris Lu <chris.lu@mediatek.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:54:54 -04:00
Chris Lu
78b6abd6c7 Bluetooth: btmtk: fix wrong status for short WMT FUNC_CTRL events
A too-short BTMTK_WMT_FUNC_CTRL event (WMT header only, no trailing
2-byte status word) is always treated as BTMTK_WMT_ON_UNDONE. This
short form is how firmware acks a plain enable/disable request, and
the actual result is carried in the header's own flag byte (0 =
success), not a separate status word. Decode it from there instead of
assuming failure.

Verified setup on MT7920, MT7921, MT7922 and MT7925: no regression.

Fixes: e3ac0d9f1a ("Bluetooth: btmtk: accept too short WMT FUNC_CTRL events")
Assisted-by: Claude:claude-opus-5
Signed-off-by: Chris Lu <chris.lu@mediatek.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:54:48 -04:00
Luiz Augusto von Dentz
296e7f3c50 Bluetooth: ISO: set BT_LISTEN before requesting a BIG sync
A BIS connection is matched to its parent socket by looking for a
socket in BT_LISTEN state with the same BIG handle:

  iso_conn_ready()
    if (test_bit(HCI_CONN_BIG_SYNC, &hcon->flags))
            parent = iso_get_sock(hdev, &hcon->src, &hcon->dst,
                                  BT_LISTEN, iso_match_big_hcon, hcon);

The socket was only moved to BT_LISTEN after iso_conn_big_sync()
returned, while the LE BIG Create Sync command has already been queued
by then. If the BIG sync is established before the state is updated,
which is easy to hit with an emulated controller as the command may
complete in a few hundred microseconds, no parent is found and the BIS
connections are never notified to the listening socket.

The user space is then left waiting for connections that never arrive,
e.g. bluetoothd never completes a MediaTransport1.Acquire of a
Broadcast Sink transport.

Move the socket to BT_LISTEN before requesting the BIG sync, so the
state is visible by the time the command is queued, and restore the
previous state if the request could not be started. Since the socket is
briefly visible as a listening socket, child sockets may have been
queued in the meantime, so drain the accept queue before restoring the
state: the cleanup paths of BT_CONNECT2/BT_CONNECTED don't do it and the
children would be left with a dangling parent pointer.

Fixes: fbdc4bc472 ("Bluetooth: ISO: Use defer setup to separate PA sync and BIG sync")
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:53:12 -04:00
Luiz Augusto von Dentz
ca18ee413a Bluetooth: ISO: Fix parent socket leak in iso_conn_ready()
iso_get_sock() returns the parent socket with a reference held, which is
dropped by sock_put() once the child socket has been set up. The error
path taken when iso_sock_alloc() fails only calls release_sock() and
returns, leaking the reference and thus the parent socket itself.

Drop the reference on that path as well.

Fixes: fa224d0c09 ("Bluetooth: ISO: Reassociate a socket with an active BIS")
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:53:03 -04:00
Chengfeng Ye
9a10987a2f Bluetooth: hci_sync: Serialize local codec list cleanup
hci_dev_close_sync() clears hdev->local_codecs after releasing hdev->lock.
Codec list additions and both traversals in sco_sock_getsockopt() use that
lock, but the close path does not. A close and BT_CODEC query can therefore
interleave as follows:

  hci_dev_close_sync()          sco_sock_getsockopt()
                                hci_dev_lock()
                                fetch codec entry
  hci_codec_list_clear()
    kfree(entry)
                                read entry->id

The reader then accesses an entry which the close path has freed. KASAN
reported:

  BUG: KASAN: slab-use-after-free in sco_sock_getsockopt+0xfa0/0xfe0
  Read of size 1 at addr ffff8881001c3450
  Call Trace:
   sco_sock_getsockopt+0xfa0/0xfe0
   do_sock_getsockopt+0x537/0x7b0
   __sys_getsockopt+0xf2/0x170
  Allocated by task 92:
   hci_codec_list_add.isra.0+0x2c/0x440
   hci_read_codec_capabilities+0x224/0x590
   hci_read_supported_codecs+0x2c2/0x640
  Freed by task 92:
   kfree+0x131/0x3c0
   hci_codec_list_clear+0xd8/0x160
   hci_dev_close_sync+0x92a/0xfa0

Take hdev->lock around the clear operation at its existing point in the
close path. This makes the clear wait for active readers and prevents a new
traversal until the list is empty without changing teardown ordering.

Fixes: b938790e70 ("Bluetooth: hci_codec: Fix leaking content of local_codecs")
Cc: stable@vger.kernel.org
Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:52:31 -04:00
Ibrahim Abdelkader
4e93c65f87 Bluetooth: hci_qca: Do not write to the serial port after it is closed
hci_uart_close() closes the serdev port if HCI_QUIRK_NON_PERSISTENT_SETUP
is set (for example, for the WCN399x family). A failed hci_dev_open_sync()
following a successful qca_setup() calls hdev->close() but not
hdev->shutdown(), so the port is closed while power->vregs_on is left true.
qca_serdev_remove() then passes its power->vregs_on test and calls
qca_power_off(), which writes to the closed port unconditionally.

Seen on a WCN3988 by unbinding the driver after a controller failure. The
trace below is from a 7.0.0 based kernel, where qca_power_off() was still
named qca_power_shutdown():

  Unable to handle kernel NULL pointer dereference at virtual address
  0000000000000038
  Call trace:
   tty_set_termios+0x50/0x238 (P)
   ttyport_set_baudrate+0x84/0xc0
   serdev_device_set_baudrate+0x24/0x40
   qca_power_shutdown+0x158/0x1fc [hci_uart]
   qca_serdev_remove+0x54/0x68 [hci_uart]
   serdev_drv_remove+0x1c/0x2c
   device_remove+0x4c/0x80
   device_release_driver_internal+0x1cc/0x224
   device_driver_detach+0x18/0x24
   unbind_store+0xb4/0xc0

Check HCI_UART_PROTO_READY, which hci_uart_close() clears in the same place
it closes the port, before writing to it. The regulator disable is left
unconditional so the controller is still powered down.

The dangling serport->tty that turns this into a use-after-free is
addressed in a separate patch.

Fixes: fa9ad876b8 ("Bluetooth: hci_qca: Add support for Qualcomm Bluetooth chip wcn3990")
Signed-off-by: Ibrahim Abdelkader <iabdelka@qti.qualcomm.com>
Reviewed-by: Hans de Goede <johannes.goede@oss.qualcomm.com>
Signed-off-by: Hans de Goede <johannes.goede@oss.qualcomm.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:52:19 -04:00
Laxman Acharya Padhya
d0795cfd6f Bluetooth: hci_codec: validate vendor codec count length
The Read Local Supported Codecs parsers consume the variable-sized
standard codec array before parsing the vendor codec count.  Although the
initial reply-size check includes a vendor count byte in the fixed layout,
it does not guarantee that the byte remains after the standard codec array.

If a controller reply ends immediately after that array, calculating the
vendor codec array size reads vnd_codecs->num beyond the skb data.  Use
skb_pull_data() to validate and consume each codec header before using its
count in both command variants.

Fixes: 8961987f3f ("Bluetooth: Enumerate local supported codec and cache details")
Fixes: 9ae664028a ("Bluetooth: Add support for Read Local Supported Codecs V2")
Cc: stable@vger.kernel.org
Suggested-by: Luiz Augusto von Dentz <luiz.dentz@gmail.com>
Signed-off-by: Laxman Acharya Padhya <acharyalaxman8848@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:52:12 -04:00
Radek Podgorny
4914c49989 Bluetooth: put the peer's on-air address on air when we cannot resolve
An identity address only reaches a peer that is advertising an RPA if the
controller resolves it on our behalf. Where it cannot, the host has to put
the peer's on-air address on air itself.

hci_connect_le() still swaps the caller's identity address for the peer's
cached RPA before creating the connection, but __hci_conn_add() resolves
the RPA back to the identity address when it stores it, so the identity is
what goes out. Storing the identity is right when the controller
translates it on the way to the radio; without LL Privacy, or with this
peer absent from the resolving list, nothing does.

A peer advertising an RPA cannot answer its identity address, so the
attempt burns a full create-connection timeout. That is not merely a slow
connect: a controller without extended scanning cannot scan while it is
initiating, so every dead attempt also takes the scanner off the air for
the whole timeout.

Measured on a CYW43438, which reports neither LL Privacy nor extended
advertising (LE features 3f 00 00 08 00 00 00 00), against a peer
advertising a resolvable private address the host holds the IRK for, with
the connection requested on the peer's identity address:

  before: LE Create Connection to the identity address, public type
          1.61s -> 22.07s, then LE Create Connection Cancel
          LE Connection Complete: Unknown Connection Identifier (0x02)
  after:  LE Create Connection to the peer's RPA, random type
          LE Connection Complete: Success

Advertising reports reaching the host per second, same window, same five
unrelated devices on the adapter:

  before   1s:2   [nothing from 2s through 21s]   22s:5  23s:3
  after    0s:11 1s:5 2s:2 3s:5 4s:3 5s:4 ... 21s:2 22s:1 23s:2

One dead connect costs twenty seconds of scanning for every device on the
adapter, not just the one being dialled.

Keep the RPA in conn->dst unless the controller will translate the
identity address: address resolution enabled and the peer's identity
actually programmed into the resolving list. Testing ll_privacy_capable()
alone would not be enough: it reports the feature bit, not whether
resolution is switched on and not whether this peer is in the list.
Resolution is cleared with the other volatile flags on power-off and
switched off again while suspend pauses scanning, and a peer's IRK is only
programmed along the accept list path, so a direct-connect target, a peer
without HCI_CONN_FLAG_ADDRESS_RESOLUTION, and one that did not fit in a
full list are all absent from it.

With the peer programmed, the identity address stays in conn->dst and the
controller translates it: measured on an Intel controller, the host dials
the identity and LE Enhanced Connection Complete reports Resolved Public
with the peer's RPA in the separate peer resolvable private address field.
With the peer absent from the list the same setup dials the RPA itself.

Everything downstream already copes with an RPA in conn->dst: it is what
every outgoing LE connection stored before 14b06c3a88, the connection
complete event names the address that was dialled, and
le_conn_complete_evt() resolves it back to the identity once the link is
up. ISO links keep the unconditional conversion: they are created from an
existing ACL or a periodic sync and never dial this address themselves.

Keeping the RPA is only right while the peer is still using it, which is
why the preceding patch drops the cached RPA as soon as the peer is seen
advertising its identity address. Without that, a peer that turns privacy
off would be dialled on the address it abandoned rather than the one it
is answering on.

Fixes: 14b06c3a88 ("Bluetooth: HCI: Always use the identity address when initializing a connection")
Assisted-by: Claude:claude-opus-5
Assisted-by: Claude:claude-fable-5
Signed-off-by: Radek Podgorny <radek@podgorny.cz>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:50:25 -04:00
Weiming Shi
d236517c26 Bluetooth: coredump: Quiesce dump work on unregister
hci_devcd_handle_pkt_init() arms dump_timeout and coredump producers
queue dump_rx without holding an hdev reference. Unregister leaves both
works live, so disconnecting during an active dump lets them access hdev
after hci_release_dev() frees it.

Shut down coredump processing during unregister. Close the producer gate
under dump_q.lock before disabling both works, then free the active buffer
and queued packets under hci_dev_lock. Serializing the gate with enqueue
prevents controller-specific workers from adding packets after the final
purge.

Fixes: 9695ef876f ("Bluetooth: Add support for hci devcoredump")
Reported-by: syzbot+b170dbf55520ebf5969a@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=b170dbf55520ebf5969a
Reported-by: Aby Sam Ross <abysamross@gmail.com>
Link: https://lore.kernel.org/r/20260322210849.68743-1-abysamross@gmail.com
Suggested-by: Aby Sam Ross <abysamross@gmail.com>
Reported-by: Tristan Madani <tristan@talencesecurity.com>
Link: https://lore.kernel.org/r/20260814231248.3096377-1-tristmd@gmail.com
Reported-by: Xiang Mei <xmei5@asu.edu>
Assisted-by: OpenAI Codex:gpt-5
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Reported-by: Xiang Mei <xmei5@asu.edu>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:49:41 -04:00
Chandrashekar Devegowda
4b837ebd0e Bluetooth: btintel_pcie: validate TX skb length in send_sync
btintel_pcie_prepare_tx() copies skb->len bytes into a fixed
BTINTEL_PCIE_BUFFER_SIZE (4096) DMA slot via an unchecked memcpy.
Oversized packets are currently rejected only in
btintel_pcie_send_frame(); any future caller of
btintel_pcie_send_sync() would silently overflow the DMA buffer.

Add the bounds check in btintel_pcie_send_sync() itself, right
before skb_push() and the DMA copy.

Assisted-by: Copilot:claude-sonnet-5 code-review code-generation
Fixes: 6e65a09f92 ("Bluetooth: btintel_pcie: Add *setup* function to download firmware")
Signed-off-by: Chandrashekar Devegowda <chandrashekar.devegowda@intel.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:49:34 -04:00
ThangNN99
6610c6fe4b Bluetooth: hci_core: Fix queuing tx_work after workqueue is drained
hci_send_acl(), hci_send_sco() and hci_send_iso() queue hdev->tx_work
unconditionally. They can run from the L2CAP/SCO/ISO socket send path
while hci_dev_close_sync() is draining hdev->workqueue (HCIDEVDOWN
racing with a socket write). Since that queue_work() is not chained
work from the tx_work worker itself, __queue_work() sees the queue
marked __WQ_DRAINING, warns "cannot queue %ps on wq %s", and drops
the work:

  WARNING: CPU: 1 PID: 5985 at kernel/workqueue.c:2352 __queue_work
  Call Trace:
   queue_work_on
   l2cap_chan_send
   l2cap_sock_sendmsg
   ...

hci_dev_close_sync() already sets HCI_CMD_DRAIN_WORKQUEUE before
draining, but only hci_cmd_work() and handle_cmd_cnt_and_timer()
check it before queuing. Route the tx_work producers through the
same guard via a shared hci_sched_tx() helper.

Fixes: 525daaea45 ("Bluetooth: hci_sync: Set HCI_CMD_DRAIN_WORKQUEUE during device close")
Reported-by: syzbot+b6919040d9958e2fc1ae@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=b6919040d9958e2fc1ae
Signed-off-by: ThangNN99 <ngocthang2710.1999@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:48:44 -04:00
Aamir Ahmed
e824176679 Bluetooth: eir: validate service data length before reading UUID
eir_get_service_data() reads a 16-bit UUID from the service data using
get_unaligned_le16() without first checking that the data is long enough
to hold a UUID16 (2 bytes). If a malformed EIR entry has a service data
field with only 1 byte of payload (field_len=2), eir_get_data() returns
dlen=1. The subsequent get_unaligned_le16() then reads 1 byte past the
field boundary.

Additionally, if the corrupted UUID happens to match, the length
calculation "dlen - 2" underflows to SIZE_MAX since dlen is size_t.
Current callers either pass NULL for the length parameter or bounds-check
the returned length, but future callers may not.

Add a check that dlen >= sizeof(u16) and skip fields that are too short
to contain a valid UUID16.

Fixes: 8f9ae5b3ae ("Bluetooth: eir: Add helpers for managing service data")
Cc: stable@vger.kernel.org
Signed-off-by: Aamir Ahmed <elb12345@hotmail.co.uk>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:30:14 -04:00
Nicolas Thibert
2b50adefed Bluetooth: btusb: fix NXP IW610 composite device handling
The NXP IW610 module exposes itself as a composite USB device
(0471:0215) with three interfaces: two real Bluetooth HCI interfaces
(class 0xe0) and one vendor-specific WiFi interface (class 0xff) used
by mwifiex-nxp.

The composite device's whole USB descriptor reports class 0xe0/01/01
(Bluetooth), so btusb_table's generic USB_DEVICE_INFO(0xe0, 0x01, 0x01)
entry matches every interface, not just the two real HCI ones -- btusb
ends up binding the WiFi interface too, and mwifiex-nxp never gets it.

Fix:
1. In btusb_table (the table the USB core actually matches against),
   explicitly ignore the WiFi interface via BTUSB_IGNORE, ahead of the
   generic entry.
2. In quirks_table, scope the existing BTUSB_MARVELL entry to the BT
   interface class instead of matching the whole device by VID/PID
   (harmless either way since quirks_table isn't consulted for initial
   binding, but keep it correct).

Not upstream anywhere: checked NXP's own i.MX kernel fork
(nxp-imx/linux-imx), no IW610 references in btusb.c on any branch --
their reference designs wire this chip differently (WiFi over SDIO
per their release notes), so they never hit this.

Signed-off-by: Nicolas Thibert <nithibert@gmail.com>
Cc: stable@vger.kernel.org
Assisted-by: LLM (Claude Sonnet 5, Anthropic)
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
2026-09-15 14:26:10 -04:00
Eric Dumazet
83a945a529 tcp: do not let tcp_rmem be set below 4096
We can hit a division by zero crash in tcp_rcvbuf_grow()
and tcp_rcv_space_adjust():

divide error: 0000 [#1] PREEMPT SMP
RIP: 0010:tcp_rcvbuf_grow+0x187/0x450 net/ipv4/tcp_input.c:939
...
grow = div_u64(((u64)rcvwin << 1) * (newval - oldval), oldval);

The division uses oldval = tp->rcvq_space.space as divisor.
When tp->rcvq_space.space is zero, this leads to a divide-by-zero
exception.

tp->rcvq_space.space is initialized in tcp_init_buffer_space():
    tp->rcvq_space.space = min3(tp->rcv_ssthresh, tp->rcv_wnd,
                                (u32)TCP_INIT_CWND * tp->advmss);

If tcp_rmem[1] is configured to very small values (such as 1),
sk->sk_rcvbuf is initialized to 1. Then tcp_full_space(sk), which
computes (sk->sk_rcvbuf * scaling_ratio) >> 8, truncates to 0.
This sets tp->window_clamp = 0, tp->rcv_ssthresh = 0, and
tp->rcvq_space.space = 0. Later, when data arrives and DRS is invoked,
tcp_rcvbuf_grow() divides by oldval == 0.

Back in 2015, commit b1cb59cf2e ("net: sysctl_net_core: check SNDBUF
and RCVBUF for min length") ensured that net.core.rmem_default and
net.core.rmem_max cannot be set below SOCK_MIN_RCVBUF. Similarly,
SO_RCVBUF setsockopt enforces max_t(int, val * 2, SOCK_MIN_RCVBUF).

However, net.ipv4.tcp_rmem still had .extra1 = SYSCTL_ONE, allowing
arbitrarily small values.

Because SOCK_MIN_RCVBUF depends on sizeof(struct sk_buff) and cacheline
alignment, its value varies across architectures and configuration options.
Using a fixed constant of 4096 ensures a predictable, architecture-
independent lower bound that is safely above SOCK_MIN_RCVBUF everywhere
and matches the documented 4K default.

Fix this by setting tcp_rmem.extra1 to 4096 and updating the documentation.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260912144848.3448026-1-edumazet@google.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-15 15:21:17 +02:00
Kuniyuki Iwashima
8e759cd1f6 tcp: Don't call skb_clone_and_charge_r() for close()d listener in tcp_v6_do_rcv().
tcp_v6_do_rcv() no longer calls skb_clone_and_charge_r() for
TCP_LISTEN since commit 073d89808c ("net: fix data-races around
sk->sk_forward_alloc").

However, there is still a small race window between tcp_v6_rcv()
and tcp_v6_do_rcv(), where concurrent close() changes TCP_LISTEN
to TCP_CLOSE, causing skb_clone_and_charge_r() to be called
locklessly and resulting in the splat below. [0]

Let's avoid calling skb_clone_and_charge_r() for TCP_CLOSE as well.

This is fine for non-listeners because tcp_rcv_state_process()
drops skb for TCP_CLOSE and opt_skb was freed immediately anyway.

[0]:
sk->sk_forward_alloc
WARNING: net/ipv4/af_inet.c:162 at inet_sock_destruct+0x64d/0x810 net/ipv4/af_inet.c:162, CPU#1: ksoftirqd/1/28
Modules linked in:
CPU: 1 UID: 0 PID: 28 Comm: ksoftirqd/1 Not tainted 7.2.0 #17 PREEMPT(full)
Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.17.0-debian-1.17.0-1 04/01/2014
RIP: 0010:inet_sock_destruct+0x64d/0x810 net/ipv4/af_inet.c:162
Code: 3d 49 ff e9 06 fd ff ff e8 d0 5b 83 f8 90 0f 0b 90 e9 35 fe ff ff e8 c2 5b 83 f8 90 0f 0b 90 e9 c5 fe ff ff e8 b4 5b 83 f8 90 <0f> 0b 90 e9 04 ff ff ff e8 a6 5b 83 f8 90 0f 0b 90 e9 65 fe ff ff
RSP: 0018:ffffc90000677bb8 EFLAGS: 00010246
RAX: 0000000000000000 RBX: ffff8880117bde80 RCX: ffffffff8957eb41
RDX: ffff88801dad5d00 RSI: ffffffff8957ec3c RDI: 0000000000000005
RBP: 00000000fffff000 R08: ffffffff8957eb41 R09: 00000000fffff000
R10: 0000000000000005 R11: 0000000000000000 R12: dffffc0000000000
R13: ffff8880117bdf10 R14: ffffffff81c08eb7 R15: 0000000000000003
FS:  0000000000000000(0000) GS:ffff8880d7ae5000(0000) knlGS:0000000000000000
CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 00007f93a1021138 CR3: 00000000207a9000 CR4: 0000000000350ef0
Call Trace:
 <TASK>
 __sk_destruct+0x82/0xae0 net/core/sock.c:2356
 rcu_do_batch kernel/rcu/tree.c:2645 [inline]
 rcu_core+0x59c/0x1100 kernel/rcu/tree.c:2897
 handle_softirqs+0x1e4/0x9b0 kernel/softirq.c:622
 run_ksoftirqd kernel/softirq.c:1076 [inline]
 run_ksoftirqd+0x38/0x60 kernel/softirq.c:1068
 smpboot_thread_fn+0x458/0xc80 kernel/smpboot.c:160
 kthread+0x396/0x4a0 kernel/kthread.c:436
 ret_from_fork+0x8e0/0xe40 arch/x86/kernel/process.c:158
 ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
 </TASK>

Fixes: e994b2f0fb ("tcp: do not lock listener to process SYN packets")
Reported-by: Taras Madan <tarasmadan@google.com>
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260914011420.115556-1-kuniyu@google.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-15 15:13:22 +02:00
Jamal Hadi Salim
0654f4dba1 selftests/tc-testing: add hhf hh_limit cap tests
Cover the new TCA_HHF_HH_FLOWS_LIMIT bound: values above 2*HH_FLOWS_CNT
(4294967295, 65536, 2049) are rejected with the configured limit left
untouched on both the change and the add path, the boundary value 2048 is
accepted (installed at 100 first so the boundary change is load-bearing),
and an add-time hh_limit 500 is preserved instead of being clobbered by
the default.

Reviewed-by: Victor Nogueira <victor@mojatatu.com>
Tested-by: hybris <hybris@mojatatu.ai>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/QDISC-B855.v1.20260911153152@mojatatu.com.2
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-15 13:30:53 +02:00
Jamal Hadi Salim
2cef2588c9 net/sched: hhf: cap hh_flows_limit at change time
hhf_change() stores TCA_HHF_HH_FLOWS_LIMIT with no upper bound. A huge
hh_flows_limit lets each new heavy-hitter flow pass the
hh_flows_current_cnt check in alloc_new_hh() and forces a fixed-size
kzalloc(GFP_ATOMIC) per flow under spoofed traffic, for unbounded memory
growth.

Bound the attribute with NLA_POLICY_MAX() at 2*HH_FLOWS_CNT (the
hhf_init() default) and report the rejected value via extack. The
deprecated nested parse is kept: legacy tc does not set NLA_F_NESTED on
TCA_OPTIONS. Configs relying on hh_limit above the default were relying
on unbounded, unsafe behaviour and are not supported going forward.

hhf_init() also ran hhf_change() before setting the default
hh_flows_limit, so a user-supplied hh_limit at add time was clobbered
back to 2048. Set the default before hhf_change() so the configured
value sticks.

This is a follow-up to commit eb56a495f5 ("net/sched: hhf: clamp
quantum in change and init paths"), which bounded the quantum of the
same qdisc; the hh_flows_limit bound is the remaining unbounded knob of
that series' scope.

Conditions to recreate the bug: CAP_NET_ADMIN in a user namespace;
tc qdisc change dev X root hhf hh_limit 4294967295 succeeds and the
value is echoed by tc qdisc show, unbounding heavy-hitter flow
allocations; also tc qdisc add dev X root hhf hh_limit 500 stores 2048
instead of 500.

Fixes: 10239edf86 ("net-qdisc-hhf: Heavy-Hitter Filter (HHF) qdisc")
Cc: stable@vger.kernel.org
Reported-by: Sashiko (gemini) <sashiko-bot@kernel.org>
Closes: https://sashiko.dev/#/patchset/20260822195509.112717-1-jhs@mojatatu.com
Reviewed-by: Victor Nogueira <victor@mojatatu.com>
Tested-by: hybris <hybris@mojatatu.ai>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/QDISC-B855.v1.20260911153152@mojatatu.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-15 13:30:53 +02:00
Nikolay Aleksandrov
18a6fe05fb net: bridge: mst: move switchdev call outside rcu
This is a follow-up of one of sashiko's pre-existing bug reports.
br_mst_set_state() calls switchdev_port_attr_set() for nonzero MSTIs
while holding rcu_read_lock() which invokes the blocking switchdev
notifier chain and may sleep. Nonzero MSTI changes come from netlink
with rtnl held. Move the switchdev call before entering the rcu section and
assert that rtnl is held.

The call cannot be deferred because netlink needs its error and extack.
Also DSA reads the old bridge MST state during the callback and checks it.
A deferred callback will be late and will see the updated state.

Fixes: 3a7c1661ae ("net: bridge: mst: fix vlan use-after-free")
Signed-off-by: Nikolay Aleksandrov <razor@blackwall.org>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260911105021.1385934-1-razor@blackwall.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-15 12:27:41 +02:00
Hohyun Sim
7c8810c2e6 net: fddi: skfp: fix NULL deref when setting the MAC address while down
skfp_ctl_set_mac_address() calls ResetAdapter() unconditionally, without
checking netif_running(). ResetAdapter() first calls card_stop(), which
sets smc->hw.hw_state to STOPPED, and then mac_drv_clear_tx_queue(),
which walks the two transmit queues:

	for (i = QUEUE_S; i <= QUEUE_A0; i++) {
		queue = smc->hw.fp.tx[i] ;
		...
		t = queue->tx_curr_get ;

smc->hw.fp.tx[] is only populated by init_tx(), which is reached from
skfp_open() through init_smt() -> init_fddi_driver() -> init_fplus() ->
init_mac() -> init_tx(). The private area is allocated and zeroed by
alloc_fddidev(), so on an interface that has never been brought up both
queue pointers are still NULL. The hw_state test at the top of
mac_drv_clear_tx_queue() does not catch this, because card_stop() has
just set STOPPED; the function proceeds into the loop and dereferences
NULL. ResetAdapter() does call init_smt() itself, but only after the
queues have been cleared.

Setting the MAC address on a down interface therefore oopses:

  ip link set dev fddi0 address 02:00:00:00:00:01

  BUG: KASAN: null-ptr-deref in mac_drv_clear_tx_queue+0x68/0x2c0 [skfp]
  Read of size 8 at addr 0000000000000010 by task ip/302
  Call Trace:
   <TASK>
   mac_drv_clear_tx_queue+0x68/0x2c0 [skfp 6c01d4bab63c36978bd0a7d7e90837adb44cc37b]
   ResetAdapter+0x29/0x100 [skfp 6c01d4bab63c36978bd0a7d7e90837adb44cc37b]
   skfp_ctl_set_mac_address+0x57/0x80 [skfp 6c01d4bab63c36978bd0a7d7e90837adb44cc37b]
   netif_set_mac_address+0x1e4/0x2c0
   do_setlink+0x684/0x2680
   </TASK>

Address 0x10 is the offset of tx_curr_get, the third pointer in
struct s_smt_tx_queue, on 64-bit. mac_drv_clear_rx_queue(), which
ResetAdapter() calls immediately afterwards, dereferences
smc->hw.fp.rx[QUEUE_R1] in the same way behind the same ineffective
hw_state test; the transmit queue merely crashes first. Both are
covered by the guard below.

Skip the adapter reset when the interface is down. dev_addr_set() is
left unconditional, so the new address is still recorded in
dev->dev_addr. Nothing is lost by not resetting the adapter here:
skfp_open() deliberately re-reads the factory address on every open,

	read_address(smc, NULL);
	eth_hw_addr_set(dev, smc->hw.fddi_canon_addr.a);

and the comment above it states this is done to discard exactly such an
address override across a close/open cycle. An address set while the
interface is down could not have survived the following open even
before this change, so the guard removes no working behaviour. Guarding
the hardware side of ndo_set_mac_address() with netif_running() is
established practice; skge_set_mac_address() has done so since commit
2eb3e621c4 ("skge: set mac address bonding fix").

Guarding the reset as a whole, rather than NULL-checking the queues, is
also what the rest of the driver expects. After a previous open/close
the queue pointers are stale but non-NULL, so there is no crash, yet
ResetAdapter() goes on to call smt_online() and STI_FBI() ("Enable
Board Interrupts") while skfp_close() has already called free_irq() -
the adapter would be brought back online with no handler installed. The
only other ResetAdapter() caller is skfp_interrupt(), which by
construction runs only while the device is open.

Found by automated driver testing against an emulated SysKonnect FDDI
adapter under a KASAN-enabled 7.0.0 kernel. Triggering it requires
CAP_NET_ADMIN.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Assisted-by: LLM KASAN
Signed-off-by: Hohyun Sim <tlaghgus0425@korea.ac.kr>
Link: https://patch.msgid.link/20260910063743.110747-1-tlaghgus0425@korea.ac.kr
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-15 10:31:54 +02:00
Dong Chenchen
2998147b59 ipv4: icmp: reject RTN_UNREACHABLE input routes in icmp_route_lookup
When the forward output route cannot be used in icmp_route_lookup(),
it enters the "reverse path" and calls ip_route_input() on fl4_dec.daddr,
the original packet's source address.

ip_route_input() only returns an error for truly invalid packets. For
unreachable addresses it will succeed and return an input route whose
dst.output is set to ip_rt_bug(). The existing check only rejects
RTN_LOCAL routes, so the RTN_UNREACHABLE route types can still be returned
and later used for output, syzkaller triggering a WARN_ON_ONCE()
in ip_rt_bug() as bellow:

 ------------[ cut here ]------------
 WARNING: net/ipv4/route.c:1273 at ip_rt_bug+0x14/0x20
 RIP: 0010:ip_rt_bug+0x14/0x20
 Call Trace:
  ip_push_pending_frames+0xfa/0x100
  __icmp_send+0x905/0xf10
  ip_options_compile+0xc0/0xd0
  ip_rcv_finish_core+0x321/0xae0
  ip_rcv+0x1de/0x260
  __netif_receive_skb_one_core+0x11a/0x130
  netif_receive_skb+0x7b/0x260
  tun_get_user+0x11bf/0x1c10
 ------------[ cut here ]------------

Reject input route that is RTN_UNREACHABLE to fix it. The net warning
is only printed for RTN_LOCAL, as RTN_UNREACHABLE is not the result of
a race condition.

Fixes: 8b7817f3a9 ("[IPSEC]: Add ICMP host relookup support")
Suggested-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Signed-off-by: Dong Chenchen <dongchenchen2@huawei.com>
Link: https://patch.msgid.link/20260910140042.1880242-1-dongchenchen2@huawei.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-15 10:20:38 +02:00
Nicolai Buchwitz
23ca4ddc4f net: bcmgenet: restore the hardware filters on open
bcmgenet_hfb_init() runs INIT_LIST_HEAD() on priv->rxnfc_list, which drops
every rule off the list, and bcmgenet_open() calls it on each ifup. Every
rule the user configured is silently lost:

  # ethtool -N eth0 flow-type ether dst $MAC action 0
  Added rule with ID 0
  # ethtool -n eth0 | grep -c Filter:
  1
  # ip link set eth0 down && ip link set eth0 up
  # ethtool -n eth0 | grep -c Filter:
  0

Initialise the lists once at probe and restore the rules on open, as
bcmgenet_resume() already does.

Fixes: 3e37095228 ("net: bcmgenet: add support for ethtool rxnfc flows")
Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
Reviewed-by: Justin Chen <justin.chen@broadcom.com>
Reviewed-by: Florian Fainelli <florian.fainelli@broadcom.com>
Link: https://patch.msgid.link/20260913190052.939955-1-nb@tipi-net.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-14 19:21:58 -07:00
Daniel Golle
9e92ad4630 net: dsa: mxl862xx: disable the stats poll on teardown
mxl862xx_setup() arms the stats poll before mxl862xx_setup_mdio(), and
nothing stops it until dsa_register_switch() has returned an error to
mxl862xx_probe(). DSA frees the dsa_port list before it returns, so a
poll that fires once .setup or a later step of dsa_tree_setup() has
failed walks freed ports. On shutdown the user ports stay registered,
and the WORK_STOPPED flag test in mxl862xx_get_stats64() is not atomic
with the cancel in mxl862xx_shutdown(), so a re-arm that read the flag
before it was set queues the poll after cancel_delayed_work_sync() has
returned.

Arm the poll once .setup has succeeded and stop it from a .teardown op,
which DSA calls on unregister and after a failed registration, in both
cases before it frees the ports. Use disable_delayed_work_sync() there
and in shutdown(): it drains a running poll as the cancel did and turns
every later attempt to queue the work into a no-op, so the re-arm
cannot bring the poll back. remove() and the probe error path only set
WORK_STOPPED, which crc_err_work tests before it walks the ports.

Fixes: a21d33a526 ("net: dsa: mxl862xx: implement .get_stats64")
Signed-off-by: Daniel Golle <daniel@makrotopia.org>
Link: https://patch.msgid.link/1eb6f7fc1789b67e4b11e3f4d5ff080d0b6f7cbb.1789045590.git.daniel@makrotopia.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-14 19:05:46 -07:00
Andrea Mayer
7616242a2b seg6: set IPSKB_L3SLAVE from IP6SKB_L3SLAVE on IPIP decapsulation
When an SRv6 packet arrives on an interface enslaved to a VRF,
vrf_ip6_rcv() sets IP6SKB_L3SLAVE in IP6CB, but decap_and_validate()
has never set IPSKB_L3SLAVE in IPCB. The bit stayed clear in the
common case, and with CONFIG_IPV6_MIP6 the leftover frag_max_size of
a reassembled outer packet could even set it, with no VRF involved.
Commit 44930446dd ("ipv6: seg6: clear IPv4 control block on IPIP
decapsulation") then made the unreliable bit reliably clear.

The effect of the missing flag is visible with End.DX4 when a
delivery to a local address of the node reaches the socket lookup.
For example, a UDP socket bound to the enslaved ingress interface
does not receive any of the decapsulated packets, while an unbound
socket outside the VRF does.
This contradicts Documentation/networking/vrf.rst: by default the
scope of an unbound UDP or TCP socket is limited to the default VRF.

Set IPSKB_L3SLAVE for IPv4 in decap_and_validate(), which already does
the same for IPv6. The socket lookup then matches the decapsulated
packet like any other packet received on that enslaved interface. Such
a packet matches an unbound UDP or TCP socket only when
udp_l3mdev_accept or tcp_l3mdev_accept is set.

Fixes: 891ef8dd2a ("ipv6: sr: implement additional seg6local actions")
Signed-off-by: Andrea Mayer <andrea.mayer@uniroma2.it>
Reviewed-by: David Ahern <dsahern@kernel.org>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Link: https://patch.msgid.link/20260913194421.31-1-andrea.mayer@uniroma2.it
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-14 19:03:01 -07:00
Ahmed Naseef
bde5212360 net: phy: mediatek: do not report link and per-speed LED rules together
mtk_phy_led_hw_ctrl_get() reports TRIGGER_NETDEV_LINK whenever any of the
speed bits in on_set is on, and in addition reports every individual
TRIGGER_NETDEV_LINK_* bit that is set. The netdev trigger refuses that
combination: netdev_led_attr_store() rejects TRIGGER_NETDEV_LINK together
with any per-speed rule, and it validates the whole resulting mode rather
than just the bit being written. Once the hardware has any link bit
programmed, every write to the trigger attributes of that LED therefore
fails with -EINVAL and the LED can no longer be configured.

The rules are also fed back into the hardware: the trigger stores what is
read back, and a later write of device_name programs it again, expanding
TRIGGER_NETDEV_LINK to every speed in on_set. An LED configured for a
single speed is thereby silently widened to "on at any link speed".

Both are easy to see on the EcoNet EN7528, whose four PHYs share one LED
block. The first LED programs the block correctly, the second reads those
rules back and rewrites them widened, and the remaining two then read the
widened value, so an LED configured for "link_10 link_100" ends up lit on a
1000 Mbps link.

on_set holds every speed the LED can indicate and is exactly what
mtk_phy_led_hw_ctrl_set() programs for TRIGGER_NETDEV_LINK, so report the
speed independent rule only when all of them are on, and the individual
speeds otherwise. The mapping is then the inverse of the one used when
programming the LED and round trips without changing the register.

Fixes: c66937b0f8 ("net: phy: mediatek-ge-soc: support PHY LEDs")
Cc: stable@vger.kernel.org
Signed-off-by: Ahmed Naseef <naseefkm@gmail.com>
Reviewed-by: Andrew Lunn <andrew@lunn.ch>
Link: https://patch.msgid.link/20260912134306.3544329-1-naseefkm@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-14 19:01:51 -07:00
Chunfeng Song
6fb0a9d907 rust: net: phy: fix off-by-one bit positions in device status accessors
The hand-written bitfield offsets in is_link_up(), is_autoneg_enabled()
and is_autoneg_completed() were correct when the abstraction was
merged: at that time autoneg, link, and autoneg_complete were at bits
13, 14, and 15 of struct phy_device's first bitfield unit. Commit
2796ff1e3d ("net: phy: add flag is_genphy_driven to struct phy_device")
later inserted is_genphy_driven just before autoneg, shifting the three
fields up by one, so the accessors now read:

  is_link_up()           reads bit 14 = autoneg
  is_autoneg_enabled()   reads bit 13 = is_genphy_driven
  is_autoneg_completed() reads bit 15 = link

The official ax88796b Rust driver uses all three accessors in its
read_status() implementation, so it inherits the bug.
phy_attach_direct() sets is_genphy_driven only when it falls back to
the generic driver, and ax88796b has a real driver, so
is_genphy_driven stays 0. The broken is_autoneg_enabled() therefore
reads bit 13 as 0, compares it against AUTONEG_ENABLE (1), and always
returns false, so read_status() never reaches the
resolve_aneg_linkmode() call.

The ordinary bindgen accessors take &self. Calling them through
(*phydev).link() would create a shared reference to the complete
bindings::phy_device, which is not appropriate for an object wrapped in
Opaque.

Use the bindgen-generated raw accessors (link_raw(), autoneg_raw(),
and autoneg_complete_raw()) instead. They retain the bit positions and
endianness handling generated from the C layout without creating a Rust
reference to the complete phy_device. Drop the hand-written numbers
together with the TODO comment that marked them as a stopgap.

The raw accessors are only emitted by bindgen 0.71 and later, and were
added at the Rust-for-Linux project's request, so this fix can only be
backported to stable branches whose minimum bindgen version is at least
that, hence the scope on the Cc: stable line below.

Found by a static equivalence audit (C2RustDrv, a C-to-Rust driver
migration tool) that compares hand-written bitfield offsets against
the bindgen layout of struct phy_device. Verified by building the
bindings and checking the generated accessors; no runtime testing was
possible without PHY hardware.

Fixes: 2796ff1e3d ("net: phy: add flag is_genphy_driven to struct phy_device")
Cc: stable@vger.kernel.org # Only 7.1.y and later (requires bindgen's raw pointer accessors).
Link: https://github.com/rust-lang/rust-bindgen/issues/2674
Signed-off-by: Chunfeng Song <springbreeze@stu.pku.edu.cn>
Reviewed-by: FUJITA Tomonori <fujita.tomonori@gmail.com>
Link: https://patch.msgid.link/20260910055110.167110-1-springbreeze@stu.pku.edu.cn
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-14 18:56:16 -07:00
Aohan Mei
f97d8c7bab rds: ib: use rds_conn_drop() on protocol version mismatch
rds_ib_cm_connect_complete() runs from the RDMA-CM event handler with
conn->c_cm_lock held.  When the peer negotiates a protocol version
older than RDS_PROTOCOL_COMPAT_VERSION, the handler calls
rds_conn_destroy(), which is only safe in the rmmod path: it
synchronously tears the connection down and flush_work()es the
shutdown work cp_down_w.

That shutdown work (rds_conn_shutdown()) needs cp_cm_lock, which is
the very lock the event handler still holds, so the flush never
completes: the two workers wait on each other and the RDS connection
workqueues stall for good.

All other RDMA-CM failure paths (REJECTED, CONNECT_ERROR,
DISCONNECTED) use rds_conn_drop(), which marks the connection
RDS_CONN_ERROR and schedules the shutdown work asynchronously.  Use
it here as well.

Fixes: f147dd9eca ("RDS/IB: Disallow connections less than RDS 3.1")
Reported-by: TencentOS Corvus AI <corvus@tencent.com>
Cc: stable@vger.kernel.org
Reviewed-by: Allison Henderson <achender@kernel.org>
Signed-off-by: Aohan Mei <henrymei@tencent.com>
Link: https://patch.msgid.link/20260911073436.3542080-1-ljp1205831794@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-14 18:47:52 -07:00
Jakub Kicinski
6c21ebc813 netfilter pull request 26-09-11
-----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEjF9xRqF1emXiQiqU1w0aZmrPKyEFAmqj4z4ACgkQ1w0aZmrP
 KyFvYw//RFZnFWU4PPLMWFiHCKfBQXaDwCM6ADS2phgVLTLbs0OEQ8Mw+XvkOt2v
 ardyCu5Ujy7udvF2kaha0ACNw+Z3DYOVu+VkzhKJ4TrBaKhr9kob6Vuczdw3QQ10
 Kx6UAh3PYWV2BHqB+IoNzd40PBVTHs2ckeSntUNJ1nkGQb2f+kkjWLj3tY+OwaUv
 fUlD2e34GXdQDEAvYaRHn7Lkv+PMXx4r40EQ238yH9+wTtqBIgCRogTbLHQ3d1Lu
 r9RHMqHBOkJu5C+gDcVD1r7TXf6S1wBcohEGSvNO+VwF5KDHNylQYS/ONrnUgvRL
 9Mr9GmYXgwLN9FxNjYxlo7cjipXHvpONMJ4B0+iz1VE0SL0Z4BNqfpBISqsWW9xP
 Ox3l4IjnV0Ver+tBcNml6JcwGip1cggHrIi9lb1Zf4cMk+ndq3O+sSSW+EFpKpDn
 szu9mUrAaHpVOEHUeFqTYjC/Xvq1x1W7QkZ5BqNo4jycROhnRWeWvvUSebW1ruL+
 exWYMVA0w3aXPN6tfjGlHMCUuCgMLpOkSvXBgPKIILqnH7rueT5Qhh6wDqpD9NkW
 xaY1J56MEuhzn2Cf8YkVDdtSgXHc9dteRmg/dLLhnubycewEjDYZTm0vGnLDBYVb
 0xvVzNBQm8e7oEoczZQI9Z8Q83Pqi+OhVopZRc7SRaqJCH3pAUU=
 =H3Kp
 -----END PGP SIGNATURE-----

Merge tag 'nf-26-09-11' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf

Pablo Neira Ayuso says:

====================
Netfilter fixes for net

1) Fix KMSAN reports an uninit-value in nf_nat_setup_info() for netmap,
   from Theodor Arsenij Larionov Trichkine.

2) Restrict deletion of netdevice in basechain and flowtable to exact
   matching only, from Fernando F. Mancera.

3) Fix nf_nat_register_fn() error path allowing for a memleak.

4) Hold reference on ct until flow is released to address, otherwise
   access to release ct->ext or different ct due to typesafe RCU
   semantics.

* tag 'nf-26-09-11' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf:
  netfilter: flowtable: hold reference on ct until flow is released
  netfilter: nf_nat: unregister and release hooks on error
  netfilter: nf_tables: fix device name and prefix match in hook lookup
  netfilter: nft_nat: fully initialise new_addr in netmap setup
====================

Link: https://patch.msgid.link/20260913205447.1889203-1-pablo@netfilter.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-14 17:06:52 -07:00
Jakub Kicinski
e6b6078ea1 Merge branch 'neighbour-small-fixes-for-rtm_-get-set-neightbl'
Kuniyuki Iwashima says:

====================
neighbour: Small fixes for RTM_{GET,SET}NEIGHTBL.

While working on the follow-up suggested here,

  https://lore.kernel.org/20260902143023.GA3966681@shredder

I found a few bugs in RTM_GETNEIGHTBL and RTM_SETNEIGHTBL,
which this series fixes.

Patch 1, 3, 4 will conflict with net-next in neightbl_dump_info()
due to removal of net_eq() below:

	p = list_next_entry(&tbl->parms, list);
	list_for_each_entry_from_rcu(p, &tbl->parms_list, list) {
		if (!net_eq(neigh_parms_net(p), net))
			continue;

Note also that currently neigh_proc_dointvec_ms_jiffies_positive()
is buggy and does not enforce min/max, and it needs this fix:

  https://lore.kernel.org/20260905233819.1064529-2-kuniyu@google.com
====================

Link: https://patch.msgid.link/20260909233143.2401847-1-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-11 17:25:19 -07:00
Kuniyuki Iwashima
979aabdad8 neighbour: Skip default parms when resumed in neightbl_dump_info().
neightbl_dump_info() calls neightbl_fill_info() in each loop
to render the default parms.

If there are many devices and neightbl_fill_param_info() failed,
neightbl_fill_info() is called again when the dump resumes:

  # ynl --family rt-neigh --dump getneightbl --output-json |
    jq '.[] | {name: .name, ifindex: .parms.ifindex}'
  ...
  {
    "name": "ndisc_cache",
    "ifindex": null
  }
  ...
  {
    "name": "ndisc_cache",
    "ifindex": 6
  }
  {
    "name": "ndisc_cache",
    "ifindex": null
  }
  {
    "name": "ndisc_cache",
    "ifindex": 5
  }

Let's skip neightbl_fill_info() if it is already called in
neightbl_dump_info().

Note that we cannot use !neigh_skip instead of !default_skip
because default_skip == 1 && neigh_skip == 0 could be true
if the first neightbl_fill_param_info() fails.

Also, nidx must be cleared at the end of each table loop;
otherwise, if neightbl_fill_info() for a subsequent table
fails, the leftover nidx from the previous table would be
saved in cb->args[1], resulting in erroneously skipping parms
of the subsequent table in the next dump.

Fixes: c7fb64db00 ("[NETLINK]: Neighbour table configuration and statistics via rtnetlink")
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260909233143.2401847-5-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-11 17:25:14 -07:00
Kuniyuki Iwashima
7b430fcfc9 neighbour: Don't render blackhole_netdev via RTM_GETNEIGHTBL.
The cited commits started to initialise blackhole_netdev with
neigh_parms_alloc().

This is visible in init_net as the ifindex==0 entries via
RTM_GETNEIGHTBL:

  # ynl --family rt-neigh --dump getneightbl --output-json \
    | jq '.[] | select(.parms.ifindex == 0)
              | {name: .name, ifindex: .parms.ifindex}'
  {
    "name": "arp_cache",
    "ifindex": 0
  }
  {
    "name": "ndisc_cache",
    "ifindex": 0
  }

For RTM_SETNEIGHTBL, ifindex being 0 means wildcard.

Let's skip blackhole_netdev's parms in neightbl_dump_info().

Note that lookup_neigh_parms() does not need the same change
because the default parms is always the first entry and matches
with ifindex == 0.

Fixes: e5f80fcf86 ("ipv6: give an IPv6 dev to blackhole_netdev")
Fixes: 22600596b6 ("ipv4: give an IPv4 dev to blackhole_netdev")
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260909233143.2401847-4-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-11 17:25:13 -07:00
Kuniyuki Iwashima
6d79b223ec neighbour: Enforce min/max to NDTPA_INTERVAL_PROBE_TIME_MS.
NDTPA_INTERVAL_PROBE_TIME_MS sets .type and .min but misses
.validation_type, so no validation is applied:

  # ynl --family rt-neigh --do setneightbl \
  --json '{"name": "arp_cache", "parms": {"interval-probe-time-ms": 0}}'

  # ynl --family rt-neigh --dump getneightbl --output-json | \
  jq '.[] | select(.name == "arp_cache" and has("config"))
          | .parms["interval-probe-time-ms"]'
  0

Moreover, nla_get_msecs() uses msecs_to_jiffies(), and u64 is
silently cast to u32, so a larger value can bypass the min check:

  e.g. 4294967296 == 0x100000000

  # ynl --family rt-neigh --do setneightbl \
  --json '{"name": "arp_cache", "parms": {"interval-probe-time-ms": 4294967296}}'

  # ynl --family rt-neigh --dump getneightbl --output-json | \
  jq '.[] | select(.name == "arp_cache" and has("config"))
          | .parms["interval-probe-time-ms"]'
  0

msecs_to_jiffies() returns MAX_JIFFY_OFFSET if the value is
larger than INT_MAX.  Also, INT_MAX ms overflows int NEIGH_VAR()
when HZ > 1000 (Alpha, MIPS), and passing a negative integer to
queue_delayed_work(unsigned long delay) causes sign extension,
which wraps around the expiry time to the past, resulting in it
being handled as 0 delay in the timer wheel.

Let's use NLA_POLICY_FULL_RANGE() and limit the max to 1 day.

The same max check is applied to sysctl as well.

Note that this controls the probe interval for NTF_MANAGED
entries, so the max of 1 day is unlikely to break any
deployments.

Fixes: 211da42eaa ("net, neigh: introduce interval_probe_time_ms for periodic probe")
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260909233143.2401847-3-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-11 17:25:13 -07:00
Kuniyuki Iwashima
764dcebb03 neighbour: Add missing RCU annotation for neightbl_dump_info().
neightbl_dump_info() fetches the first non-default neigh_parms
with list_next_entry(&tbl->parms, ...) and iterates through the
list with list_for_each_entry_from_rcu().

However, list_next_entry() does not use RCU helper.

Let's use list_for_each_entry_rcu() and skip the default parms.

Fixes: 4ae34be500 ("neighbour: Convert RTM_GETNEIGHTBL to RCU.")
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260909233143.2401847-2-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-11 17:25:13 -07:00
Pablo Neira Ayuso
e75a9fa1d4 netfilter: flowtable: hold reference on ct until flow is released
nf_ct_put() releases the ct->ext area inmediately, the rcu typesafe
semantics also allow to refer to the wrong conntrack from the flowtable
datapath. Hold reference on ct until flow is released after rcu grace
period.

Add rcu_barrier() on module exit path, to ensure pending flow entries
are release before module goes away.

Fixes: 0ff90b6c20 ("netfilter: nf_flow_offload: fix use-after-free and a resource leak")
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
2026-09-11 13:04:15 +02:00
Pablo Neira Ayuso
cbdd39ce42 netfilter: nf_nat: unregister and release hooks on error
If nf_hook_entries_insert_raw() fails, the NAT hooks get never released,
resulting in a memleak.

Postpone setting nat_proto_net->nat_hook_ops when the hooks are
registered to simplify the error path to decide whether the nat hooks
need unwinding.

Fixes: 1cd472bf03 ("netfilter: nf_nat: add nat hook register functions to nf_nat")
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
2026-09-11 13:04:15 +02:00
Fernando Fernandez Mancera
444e4c88c9 netfilter: nf_tables: fix device name and prefix match in hook lookup
Currently, a netdev chain or flowtable hooked to a device prefix can be
unintentionally deleted by a control-plane request targeting an exact
device name or even a shorter one due to the usage of min() to calculate
the length to match.

Fix this by making sure an exact device match never matches a prefix and
that both the target and the candidate have the same length during
delete operation. The add and update paths retain the existing overlap
matching to prevent a single device from matching multiple hooks.

Reported-by: Wei Fang <void0red@gmail.com>
Closes: https://lore.kernel.org/netfilter-devel/CANE+tVrDeNCHQVmsqkV2ozeBqyE3GtRDMhZgsg1bhw10yGNTRQ@mail.gmail.com/
Fixes: 6d07a28950 ("netfilter: nf_tables: Support wildcard netdev hook specs")
Signed-off-by: Fernando Fernandez Mancera <fmancera@suse.de>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
2026-09-11 13:04:15 +02:00
Theodor Arsenij Larionov Trichkine
d313499df6 netfilter: nft_nat: fully initialise new_addr in netmap setup
nft_nat_setup_netmap() builds the mapped address in an on-stack
union nf_inet_addr. For an IPv4 mapping it writes only the 4-byte .ip
member and the loop runs a single 32-bit iteration, but it then copies
the whole 16-byte union into range->min_addr and range->max_addr, so the
upper 12 bytes reach nf_nat_setup_info() uninitialised.

KMSAN reports an uninit-value in nf_nat_setup_info() reached from
nft_nat_eval(). The IPv6 path fills all 16 bytes and is not affected.

Zero-initialise new_addr.

Fixes: 3ff7ddb135 ("netfilter: nft_nat: add netmap support")
Signed-off-by: Theodor Arsenij Larionov Trichkine <theodorlarionov@gmail.com>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
2026-09-11 13:04:14 +02:00
Linus Torvalds
7844502343 Nothing too exciting, usual stream of fixes.
Including fixes from Netfilter, Bluetooth and WPAN.
 
 Current release - new code bugs:
 
  - Bluetooth: hci_sync: fix not setting CE length properly
 
  - eth: enic: match mailbox replies to request numbers
 
 Previous releases - regressions:
 
  - tunnels: drop stale dst when building an ICMP error for PMTUD
 
  - ipv6: null-check fib6_node before accessing in __ip6_del_rt_siblings()
    (bug in the rtnl_lock -> RCU conversion)
 
  - eth: bnxt_en: fix crashes on Thor2 due to OOB coalescing buffer accesses
 
  - eth: bnxt_en: prevent queue stop with deferred completions
 
 Previous releases - always broken:
 
  - eth: ice: don't dereference pointers from TP_printk()
 
  - eth: fix OOB writes on ethtool flow rule dump in 3 drivers
 
  - eth: mlx5: fix FEC configuration with RS_544_514_INTERLEAVED_QUAD
 
  - dsa: tag_brcm: legacy FCS: request needed tailroom
 
 Misc:
 
  - net: cap tx_queue_len at S16_MAX to prevent oversized ring alloc
 
  - ipv6: flowlabel: cap duplicate leases per socket
 
 Signed-off-by: Jakub Kicinski <kuba@kernel.org>
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEE6jPA+I1ugmIBA4hXMUZtbf5SIrsFAmqi3pQACgkQMUZtbf5S
 IrvUIg//X9nIxY2F5PzJ5jD9p5ccXrQMLe7kT5AW2tP5PDC8d5PIv4Q5XzFKQPU7
 XElAUKvBxmofwU2lqILYGi8AeUpqHZtKPY7XzKeqd6i72KOD6mGzzYNijqttBXcM
 vFVtIeKExXjAwvNc2as1SeXVEAAAkBtrCFuMNHMq0C56yK4md/XVkCDHaJkomNit
 geke1U8gut3rZddWKxp4WDbL8Wmx9yM0uDMBznO/+cwITObA0Hme3IgRndglzz7n
 n4Ih+EG4tRrD3kUf6oePzKQ47cd+qnSVlVTCZUwB5E/HKqWJFXxSN4Sv+mez0sAS
 rrI5hl+luNKUYrZ8/jiNlvajgAL4+AYpCKPDJbXrOW+z+x4BC2VYZBAHLoUr5ZAq
 Z5OYU9SgD1oGntqkI8mAEiRTEu+4gjhIEhjENHEzqdjUogaBIp7MWwCrNBAnFWvs
 2McmNfZZMVhxKpyYnndUStsVQySVPASb0CXeqTIO6PJsAp/HBjoMYKYKPhhwk0Gp
 lE8zHjEnPVofRfXfT+oZnbS8is2nC9FjBy9ksIGcC7vyTOdPsoIBoB8JY0x/INRM
 SOJvyxrdnVkMjiBejkdOa5X9HbD1cA/NVyzT2WEaZBGPmIqfNBgUIOxnXD4CShuy
 9zX8qtHsUmmYxPteF30Uhfe0kyLQ9OnjUz2Bl++EruxpIE3+i48=
 =f9J9
 -----END PGP SIGNATURE-----

Merge tag 'net-7.3-rc3' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net

Pull networking fixes from Jakub Kicinski:
 "Nothing too exciting, usual stream of fixes. Including fixes from
  Netfilter, Bluetooth and WPAN.

  Current release - new code bugs:

   - Bluetooth: hci_sync: fix not setting CE length properly

   - eth: enic: match mailbox replies to request numbers

  Previous releases - regressions:

   - tunnels: drop stale dst when building an ICMP error for PMTUD

   - ipv6: null-check fib6_node before accessing in __ip6_del_rt_siblings()
     (bug in the rtnl_lock -> RCU conversion)

   - eth: bnxt_en:
       - fix crashes on Thor2 due to OOB coalescing buffer accesses
       - prevent queue stop with deferred completions

  Previous releases - always broken:

   - eth:
       - ice: don't dereference pointers from TP_printk()
       - fix OOB writes on ethtool flow rule dump in 3 drivers
       - mlx5: fix FEC configuration with RS_544_514_INTERLEAVED_QUAD

   - dsa: tag_brcm: legacy FCS: request needed tailroom

  Misc:

   - net: cap tx_queue_len at S16_MAX to prevent oversized ring alloc

   - ipv6: flowlabel: cap duplicate leases per socket"

* tag 'net-7.3-rc3' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (164 commits)
  selftests: tc-testing: test action batch failure cleanup
  net/sched: act_api: release all action references on NEWACTION failure
  openvswitch: fix wrong flag value in get_ipv6_ext_hdrs()
  ipmr: account multicast table and route memory
  net: phy: dp83td510: handle the active-high LED polarity mode
  net: macb: initialize PTP state before registering clock
  net: hsr: enable promiscuous mode on interlink port with fwd offload
  ipv6: fix fib6 walker UAF on seq stop
  net: stmmac: fix TX descriptor availability check for TSO traffic
  net/rds: fix tcp stream corruption with large pages
  net: mana: restore the XDP program pointer when pre-allocation fails
  net: phy: dp83867: handle the active-high LED polarity mode
  octeontx2-af: fix PF/CGX debugfs PCI bus lookup
  net: net_failover: Fix the deadlock in net_failover_slave_name_change()
  net: phy: mediatek-ge: disable EEE on the MT7530 PHY
  tcp: reject non zerocopy devmem tx
  net: ethernet: mtk_eth_soc: populate lpi_interfaces to fix EEE support
  net: dsa: mt7530: populate lpi_interfaces to fix EEE support
  net: hinic: fix mailbox segment buffer overflow
  net: sun4i-emac: fix missing of_node_put() for phy_node
  ...
2026-09-10 14:07:48 -07:00
Linus Torvalds
0a96d0d726 smb client fixes for v7.3-rc3
A batch of bug fixes for the smb client:
 
  - File type corruption fixes in reparse point handling: setting S_IFMT
    bits without clearing the existing type first corrupted the file mode
    (e.g. S_IFREG | S_IFCHR == S_IFLNK). Fixed in the WSL, POSIX and
    native symlink reparse parsers. Also fixes an uninitialized SID
    structure in the POSIX readdir path when parsing fails.
 
  - Ownership mapping fixes: forceuid/forcegid mount options were
    ignored in several code paths (SID-to-id mapping, WSL extended
    attributes, POSIX extensions getattr), allowing an untrusted server
    to dictate local file ownership despite explicit mount overrides.
 
  - Heap overflow and overflow fixes in DACL rewriting: replacing short
    SIDs with long ones could overflow the DACL buffer, and the u16
    accumulator for DACL size could wrap around with enough ACEs.
 
  - Reference count leak fixes in oplock break and deferred close:
    duplicate oplock breaks on a queued work item leaked a
    cifsFileInfo reference, and deferred close had a similar leak when
    requeueing a running work item. Both cause busy-inode oopses on
    unmount.
 
  - DFS superblock use-after-free fix: the iterator callback stored a
    raw superblock pointer without pinning it, racing with automount
    expiry.
 
  - One-byte slab OOB read in the native symlink parser when handling
    share-root relative paths.
 
  - Hardening of legacy SMB1 input: reject userspace-crafted
    cifs.idmap key descriptions that bypass kernel origin checks, and
    validate DataOffset in CIFSSMBRead() to prevent heap info
    disclosure from a malicious server.
 
  - DFS cache fix: defer metadata updates until target copying
    succeeds to prevent partial-state cache entries on allocation
    failure.
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQTcqRusfSdYROJQwGkpVtNKoQNdYwUCaqLMxAAKCRApVtNKoQNd
 YwSyAQDUDSxCnDMmJbRr4e22oF/YrGSN/snp8cqrZlZh2pb5/gD/c2G3xMJA85YP
 yL/G8auRWkpwDl0/Parqptjhx1c9YwE=
 =PjPM
 -----END PGP SIGNATURE-----

Merge tag 'cifs-fixes-7.3-rc3' of https://git.manguebit.org/linux

Pull smb client fixes from Paulo Alcantara:

 - File type corruption fixes in reparse point handling: setting S_IFMT
   bits without clearing the existing type first corrupted the file mode
   (e.g. S_IFREG | S_IFCHR == S_IFLNK). Fixed in the WSL, POSIX and
   native symlink reparse parsers. Also fixes an uninitialized SID
   structure in the POSIX readdir path when parsing fails.

 - Ownership mapping fixes: forceuid/forcegid mount options were
   ignored in several code paths (SID-to-id mapping, WSL extended
   attributes, POSIX extensions getattr), allowing an untrusted server
   to dictate local file ownership despite explicit mount overrides.

 - Heap overflow and overflow fixes in DACL rewriting: replacing short
   SIDs with long ones could overflow the DACL buffer, and the u16
   accumulator for DACL size could wrap around with enough ACEs.

 - Reference count leak fixes in oplock break and deferred close:
   duplicate oplock breaks on a queued work item leaked a
   cifsFileInfo reference, and deferred close had a similar leak when
   requeueing a running work item. Both cause busy-inode oopses on
   unmount.

 - DFS superblock use-after-free fix: the iterator callback stored a
   raw superblock pointer without pinning it, racing with automount
   expiry.

 - One-byte slab OOB read in the native symlink parser when handling
   share-root relative paths.

 - Hardening of legacy SMB1 input: reject userspace-crafted
   cifs.idmap key descriptions that bypass kernel origin checks, and
   validate DataOffset in CIFSSMBRead() to prevent heap info
   disclosure from a malicious server.

 - DFS cache fix: defer metadata updates until target copying
   succeeds to prevent partial-state cache entries on allocation
   failure.

* tag 'cifs-fixes-7.3-rc3' of https://git.manguebit.org/linux:
  smb: client: fix one-byte OOB read in smb2_parse_native_symlink()
  smb: client: fail DACL rewrite when the new DACL exceeds 64K
  smb: client: fix heap overflow in DACL owner/group rewrite
  smb: client: fix file type corruption in cifs_reparse_point_to_fattr()
  smb: client: fix file type corruption in posix_reparse_to_fattr()
  smb: client: fix file type corruption in wsl_to_fattr()
  smb: client: avoid using uninitialized SIDs in cifs_posix_to_fattr()
  smb: client: fix WSL reparse point uid/gid override
  smb: client: honor forceuid/forcegid when mapping SIDs to uid/gid
  smb: client: fix uid/gid override in getattr with posix extensions
  smb: client: fix cifsFileInfo reference leak in deferred close
  smb: client: avoid leaking refcount when cifs_sb_tlink() fails
  smb: client: avoid leaking refcount in cifs_queue_oplock_break()
  smb: client: fill cache fields after populating cache in copy_ref_data()
  smb: client: pin DFS superblock in iterator callback
  smb: client: reject userspace cifs.idmap descriptions
  smb: client: reject out-of-bounds DataOffset in CIFSSMBRead()
  smb: client: reject short READ responses in CIFSSMBRead()
2026-09-10 14:03:48 -07:00
Linus Torvalds
ad724d319c Summary
* Replace CONFIG_PROC_SYSCTL with CONFIG_SYSCTL
 
   CONFIG_SYSCTL is the config string that controls sysctl subsys.
 
 * Testing
 
   Ran through x86_64 selftest. Skipped linux-next for this trivial fix.
 -----BEGIN PGP SIGNATURE-----
 
 iQGzBAABCgAdFiEErkcJVyXmMSXOyyeQupfNUreWQU8FAmqerc0ACgkQupfNUreW
 QU/cjwv/TO4x+9L4pCbz9wrlELCjhGuUWq4QZl78n/UqcNFGZ6wXQS+9WHoAa9r3
 pfYHU0e7YfaYBJ+OaKHJq6IYRQ0M8zk+//K+fjdIj47poZo+Oqsv+bM5AnLtll3c
 ltHjEUZTOPwGMagokOJgZiIJuf6L1Ex2DOU/+MEqFEwSoGg4IorGeT/lKRUS2/RQ
 cdh+JQppSIvYeqMG1XsM7f73TuY48wyemu795sRZxtRgypv/RkN9JfVP8Qj50jBq
 G/07NHc+f2Zq+m/oq20be1pphJTpj4NOkg2oTF3LJtNuNreSKtpaQL5lWw3eL3Go
 qxDnZNhxCf3UAIMPpppQmvKqe/zeMALb1c6PixeCp1rezAa6Gu3AL3OV1tf2SKvc
 ArNK8yhY+8Cx6n/RTdNwM8HEOy41Hy+f16uN06/EezEXpFh7KwqT1vKlc4hwarPm
 3o3kJzUiam+Oz5hrnijDyDf8nPigHfqqg/hXR2p/x/6my7tA7jtjwX4DdZ2tzLcD
 9rtcNpiO
 =y3Vr
 -----END PGP SIGNATURE-----

Merge tag 'sysctl-7.03-fixes-rc3' of git://git.kernel.org/pub/scm/linux/kernel/git/sysctl/sysctl

Pull sysctl fix from Joel Granados:
 "This fell through the cracks during the latest merge window. There are
  no more CONFIG_PROC_SYSCTL uses after this fix:

   - Replace CONFIG_PROC_SYSCTL with CONFIG_SYSCTL

     CONFIG_SYSCTL is the config string that controls sysctl subsys"

* tag 'sysctl-7.03-fixes-rc3' of git://git.kernel.org/pub/scm/linux/kernel/git/sysctl/sysctl:
  syscall_user_dispatch: Use CONFIG_SYSCTL for sysctl guard
2026-09-10 09:36:56 -07:00