From f033482d76a9f18080c7a40c5f9c678bd7adc8f3 Mon Sep 17 00:00:00 2001 From: Christiano Amora Date: Wed, 16 Sep 2026 10:46:22 -0300 Subject: [PATCH 001/189] Bluetooth: SMP: reject Security Request over BR/EDR Bose QC Ultra Headphones (dual-mode, same public address on both transports) occasionally send an SMP Security Request on the BR/EDR SMP fixed channel right after the ACL link is encrypted. The kernel handles it as if it were an LE link: smp_cmd_security_req() has no transport check, smp_ltk_encrypt() looks up an LTK with the ACL connection's dst_type, and hci_find_ltk() matches the peer's LE LTK because the LE public address type is stored as ADDR_LE_DEV_PUBLIC (0), the same value as BDADDR_BREDR. HCI_OP_LE_START_ENC is then issued on the ACL handle, the controller rejects it with Invalid HCI Command Parameters, and hci_cs_le_start_enc() disconnects the link with HCI_ERROR_AUTH_FAILURE. The headphones drop within a second of connecting, before any profile is up; a manual reconnect works. btmon (MediaTek MT7922, kernel 7.0.12): > HCI Event: Encryption Change (0x08) plen 4 Status: Success (0x00) Handle: 50 Address: BC:87:FA:47:73:5E (Bose Corporation) Encryption: Enabled with AES-CCM (0x02) > ACL Data RX: Handle 50 flags 0x02 dlen 6 BR/EDR SMP: Security Request (0x0b) len 1 Authentication requirement: No bonding, No MITM, SC (0x08) < HCI Command: LE Start Encryption (0x08|0x0019) plen 28 Handle: 50 Address: BC:87:FA:47:73:5E (Bose Corporation) > HCI Event: Command Status (0x0f) plen 4 LE Start Encryption (0x08|0x0019) ncmd 1 Status: Invalid HCI Command Parameters (0x12) < HCI Command: Disconnect (0x01|0x0006) plen 3 Handle: 50 Address: BC:87:FA:47:73:5E (Bose Corporation) Reason: Authentication Failure (0x05) SMP over BR/EDR is limited to cross-transport key derivation; the Security Request procedure (Core Specification Vol 3, Part H, Section 2.4.6, PDU in Section 3.6.7) has no BR/EDR counterpart. Reply with Pairing Failed / Command Not Supported on a non-LE link, before the PDU is parsed, and keep the connection. The reply is sent directly rather than through smp_failure(): rejecting a command on the wrong transport is not an authentication failure, and MGMT_EV_AUTH_FAILED would make bluetoothd disconnect the device. Tested on the affected host (kernel 7.0.12, MediaTek MT7922, Bose QC Ultra) with the patched module built out of tree: 7 days and 49 reconnects without a drop, against 2 drops in the 3 days before the patch. Every disconnect in that week had a userspace or remote reason. Fixes: b5ae344d4c0f ("Bluetooth: Add full SMP BR/EDR support") Assisted-by: LLM Signed-off-by: Christiano Amora Signed-off-by: Luiz Augusto von Dentz --- net/bluetooth/smp.c | 17 +++++++++++++++++ 1 file changed, 17 insertions(+) diff --git a/net/bluetooth/smp.c b/net/bluetooth/smp.c index 6091c47cb002..d23f9d0729c4 100644 --- a/net/bluetooth/smp.c +++ b/net/bluetooth/smp.c @@ -2269,6 +2269,23 @@ static u8 smp_cmd_security_req(struct l2cap_conn *conn, struct sk_buff *skb) bt_dev_dbg(hdev, "conn %p", conn); + /* SMP over BR/EDR only covers cross-transport key derivation; the + * Security Request procedure has no BR/EDR counterpart. Reject it + * here, otherwise smp_ltk_encrypt() finds the peer's LE LTK + * (ADDR_LE_DEV_PUBLIC and BDADDR_BREDR are both 0) and issues + * HCI_OP_LE_START_ENC on the ACL handle, which the controller + * rejects and hci_cs_le_start_enc() turns into a disconnect. Reply + * without smp_failure(): this is not an authentication failure, and + * MGMT_EV_AUTH_FAILED would make bluetoothd drop the device. + */ + if (hcon->type != LE_LINK) { + u8 reason = SMP_CMD_NOTSUPP; + + smp_send_cmd(conn, SMP_CMD_PAIRING_FAIL, sizeof(reason), + &reason); + return 0; + } + if (skb->len < sizeof(*rp)) return SMP_INVALID_PARAMS; From f2bbb36426581045a8bf7793da5419b9375e4348 Mon Sep 17 00:00:00 2001 From: Zijun Hu Date: Tue, 15 Sep 2026 19:17:18 -0700 Subject: [PATCH 002/189] Bluetooth: btnxpuart: Fix skb leak in nxp_process_fw_dump() When CONFIG_DEV_COREDUMP=n, hci_devcd_append() returns -EOPNOTSUPP without freeing its skb argument. This leaks the cloned skb and also prevents nxp_set_ind_reset() from being called to perform recovery. Fix by guarding the hci_devcd_append(hdev, skb_clone(skb, GFP_ATOMIC)) call with IS_ENABLED(CONFIG_DEV_COREDUMP). Fixes: 998e447f443f ("Bluetooth: btnxpuart: Add support for HCI coredump feature") Signed-off-by: Zijun Hu Signed-off-by: Luiz Augusto von Dentz --- drivers/bluetooth/btnxpuart.c | 8 +++++--- 1 file changed, 5 insertions(+), 3 deletions(-) diff --git a/drivers/bluetooth/btnxpuart.c b/drivers/bluetooth/btnxpuart.c index 25e7b41b349f..4e23d71d00e8 100644 --- a/drivers/bluetooth/btnxpuart.c +++ b/drivers/bluetooth/btnxpuart.c @@ -1388,9 +1388,11 @@ static int nxp_process_fw_dump(struct hci_dev *hdev, struct sk_buff *skb) msecs_to_jiffies(20000)); } - err = hci_devcd_append(hdev, skb_clone(skb, GFP_ATOMIC)); - if (err < 0) - goto free_skb; + if (IS_ENABLED(CONFIG_DEV_COREDUMP)) { + err = hci_devcd_append(hdev, skb_clone(skb, GFP_ATOMIC)); + if (err < 0) + goto free_skb; + } if (buf_len == 0) { bt_dev_warn(hdev, "==== FW dump complete ==="); From 71af682ba4692c2ed9ace4c3d4ca462ae368c029 Mon Sep 17 00:00:00 2001 From: Lee Jones Date: Tue, 15 Sep 2026 12:08:22 +0000 Subject: [PATCH 003/189] Bluetooth: mgmt: Dequeue pending mesh_send_sync entries on cancel In send_cancel(), pending mesh_tx objects are removed from the hdev->mesh_pending list and freed via mesh_send_complete(). However, if a mesh transmission was already queued onto hdev->cmd_sync_work_list via mesh_next(), the queued entry retains a raw pointer to mesh_tx. When hci_cmd_sync_work later processes the entry, it attempts to execute mesh_send_sync and its destroy callback mesh_send_start_complete using the already freed mesh_tx pointer, leading to a use-after-free. Fix this by invoking hci_cmd_sync_dequeue() for mesh_send_sync on the target mesh_tx before completing it. If the entry is found and dequeued, its destroy callback will complete and free the object; otherwise, mesh_send_complete() is called directly. Additionally, ensure the transmission queue advances after cancellation or errors. In mesh_send_start_complete(), call mesh_next() on error unless err is -ECANCELED, because hci_cmd_sync_dequeue() holds hdev->cmd_sync_work_lock and calling mesh_next() synchronously would deadlock. Instead, advance the queue in send_cancel() once the lock is released and if no transmission is in progress. Fixes: b338d91703fa ("Bluetooth: Implement support for Mesh") Signed-off-by: Lee Jones Signed-off-by: Luiz Augusto von Dentz --- net/bluetooth/mgmt.c | 19 +++++++++++++++---- 1 file changed, 15 insertions(+), 4 deletions(-) diff --git a/net/bluetooth/mgmt.c b/net/bluetooth/mgmt.c index ac4864e56ec7..f740e745ae78 100644 --- a/net/bluetooth/mgmt.c +++ b/net/bluetooth/mgmt.c @@ -2316,6 +2316,8 @@ static void mesh_send_start_complete(struct hci_dev *hdev, void *data, int err) hci_dev_clear_flag(hdev, HCI_MESH_SENDING); /* Send Complete Error Code for handle */ mesh_send_complete(hdev, mesh_tx, false); + if (err != -ECANCELED) + mesh_next(hdev, NULL, 0); return; } @@ -2425,19 +2427,28 @@ static int send_cancel(struct hci_dev *hdev, void *data) do { mesh_tx = mgmt_mesh_next(hdev, cmd->sk); - if (mesh_tx) - mesh_send_complete(hdev, mesh_tx, false); + if (mesh_tx) { + if (!hci_cmd_sync_dequeue(hdev, mesh_send_sync, + mesh_tx, NULL)) + mesh_send_complete(hdev, mesh_tx, false); + } } while (mesh_tx); } else { mesh_tx = mgmt_mesh_find(hdev, cancel->handle); - if (mesh_tx && mesh_tx->sk == cmd->sk) - mesh_send_complete(hdev, mesh_tx, false); + if (mesh_tx && mesh_tx->sk == cmd->sk) { + if (!hci_cmd_sync_dequeue(hdev, mesh_send_sync, + mesh_tx, NULL)) + mesh_send_complete(hdev, mesh_tx, false); + } } mgmt_cmd_complete(cmd->sk, hdev->id, MGMT_OP_MESH_SEND_CANCEL, 0, NULL, 0); + if (!hci_dev_test_flag(hdev, HCI_MESH_SENDING)) + mesh_next(hdev, NULL, 0); + return 0; } From e06d549fcd4a0ba381ed67ddf1ab3c7a6ca4314c Mon Sep 17 00:00:00 2001 From: Aldo Ariel Panzardo Date: Tue, 15 Sep 2026 13:04:29 -0300 Subject: [PATCH 004/189] Bluetooth: hci_conn: fix CIS hold ownership on reuse Commit 69997d50ec57 ("Bluetooth: ISO: handle bound CIS cleanup via hci_conn") made hci_bind_cis() and hci_connect_cis() return a connection with one hold for the ISO layer. hci_bind_cis() currently takes that hold only after configuring a CIS, so its BT_CONNECTED and matching BT_BOUND paths return a bare lookup result. Its configuration failure path can likewise call hci_conn_drop() before taking a hold. Take the hold before any state-dependent return or configuration error so every successful return follows the documented ownership contract and every error drop is balanced. hci_connect_cis() also assumes hci_conn_link() always takes a new CIS hold before dropping the one returned by hci_bind_cis(). However, the helper returns an existing link without taking another hold. In that case, preserve the CIS hold for the caller and drop the redundant LE hold because the existing link already owns its parent hold. Returning early also avoids changing an existing CIS back to BT_CONNECT. Fixes: 69997d50ec57 ("Bluetooth: ISO: handle bound CIS cleanup via hci_conn") Cc: stable@vger.kernel.org Signed-off-by: Aldo Ariel Panzardo Signed-off-by: Luiz Augusto von Dentz --- net/bluetooth/hci_conn.c | 9 ++++++++- 1 file changed, 8 insertions(+), 1 deletion(-) diff --git a/net/bluetooth/hci_conn.c b/net/bluetooth/hci_conn.c index fa72cf8aaa7a..e32bb9c342a5 100644 --- a/net/bluetooth/hci_conn.c +++ b/net/bluetooth/hci_conn.c @@ -2079,6 +2079,8 @@ struct hci_conn *hci_bind_cis(struct hci_dev *hdev, bdaddr_t *dst, cis->conn_timeout = timeout; } + hci_conn_hold(cis); + if (cis->state == BT_CONNECTED) return cis; @@ -2120,7 +2122,6 @@ struct hci_conn *hci_bind_cis(struct hci_dev *hdev, bdaddr_t *dst, return ERR_PTR(-EINVAL); } - hci_conn_hold(cis); cis->state = BT_BOUND; return cis; @@ -2497,6 +2498,12 @@ struct hci_conn *hci_connect_cis(struct hci_dev *hdev, bdaddr_t *dst, return cis; } + /* The existing link already owns the hold on its parent. */ + if (cis->link) { + hci_conn_drop(le); + return cis; + } + link = hci_conn_link(le, cis); hci_conn_drop(cis); if (!link) { From 0fcd4dad555c96e0bd3a1b8c569f989be85c7341 Mon Sep 17 00:00:00 2001 From: Aldo Ariel Panzardo Date: Tue, 15 Sep 2026 13:04:30 -0300 Subject: [PATCH 005/189] Bluetooth: ISO: release unused CIS holds after channel attach hci_bind_cis() and hci_connect_cis() return one hci_conn hold for the ISO layer. A new channel association consumes that hold, which is eventually released by iso_conn_free(). There are two cases where iso_chan_add() does not create an association: it returns success when the socket is already attached to the same iso_conn, and it returns -EBUSY when another socket is attached. The hold returned for the current call is unused in both cases. This occurs when deferred setup calls iso_connect_cis() again for its existing socket, or when another socket attempts to reuse the CIS. Detect the idempotent case while the connection is locked and release the unused hold after iso_chan_add(). Also release it on -EBUSY. Do not drop it for other errors: a newly allocated iso_conn releases the transferred hold when its last temporary reference is put. Fixes: 69997d50ec57 ("Bluetooth: ISO: handle bound CIS cleanup via hci_conn") Cc: stable@vger.kernel.org Signed-off-by: Aldo Ariel Panzardo Signed-off-by: Luiz Augusto von Dentz --- net/bluetooth/iso.c | 7 +++++++ 1 file changed, 7 insertions(+) diff --git a/net/bluetooth/iso.c b/net/bluetooth/iso.c index eb99653f33f9..7657c2a0abbf 100644 --- a/net/bluetooth/iso.c +++ b/net/bluetooth/iso.c @@ -496,6 +496,7 @@ static int iso_connect_cis(struct sock *sk) struct hci_dev *hdev; bdaddr_t src, dst; u8 src_type; + bool already_attached; int err; lock_sock(sk); @@ -568,8 +569,14 @@ static int iso_connect_cis(struct sock *sk) goto unlock; } + iso_conn_lock(conn); + already_attached = iso_pi(sk)->conn == conn && conn->sk == sk; + iso_conn_unlock(conn); + err = iso_chan_add(conn, sk, NULL); iso_conn_put(conn); + if (already_attached || err == -EBUSY) + hci_conn_drop(hcon); if (err) goto unlock; From e93fad891c72deb84cae49430163b384ebcc92b1 Mon Sep 17 00:00:00 2001 From: Aldo Ariel Panzardo Date: Tue, 15 Sep 2026 13:03:58 -0300 Subject: [PATCH 006/189] Bluetooth: hci_sock: reject out-of-range OCF values The raw HCI socket security filter has 128 OCF bits per supported OGF, but masks the 10-bit OCF with 127 before looking up the command. An unprivileged socket can therefore submit a reserved OCF that aliases an allowlisted command modulo 128. A conforming controller should reject reserved opcodes. Nevertheless, the security decision must apply to the opcode that will actually be sent, especially since controller-specific behavior is outside the host stack's control. Reject OCF values that cannot be represented by the security filter instead of aliasing them onto an unrelated command. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Cc: stable@vger.kernel.org Signed-off-by: Aldo Ariel Panzardo Signed-off-by: Luiz Augusto von Dentz --- net/bluetooth/hci_sock.c | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/net/bluetooth/hci_sock.c b/net/bluetooth/hci_sock.c index 070ca388f9ac..6413593eee51 100644 --- a/net/bluetooth/hci_sock.c +++ b/net/bluetooth/hci_sock.c @@ -1881,7 +1881,8 @@ static int hci_sock_sendmsg(struct socket *sock, struct msghdr *msg, u16 ocf = hci_opcode_ocf(opcode); if (((ogf > HCI_SFLT_MAX_OGF) || - !hci_test_bit(ocf & HCI_FLT_OCF_BITS, + (ocf > HCI_FLT_OCF_BITS) || + !hci_test_bit(ocf, &hci_sec_filter.ocf_mask[ogf])) && !capable(CAP_NET_RAW)) { err = -EPERM; From b0a6cf99afd57a39598b1beca0e86ef5004980de Mon Sep 17 00:00:00 2001 From: Aldo Ariel Panzardo Date: Tue, 15 Sep 2026 13:03:07 -0300 Subject: [PATCH 007/189] Bluetooth: hci_sock: validate event length before filtering is_filtered_packet() reads the event code from skb->data[0] without first checking that the skb is nonempty. When an opcode filter is configured, it also reads the command opcode at offsets 3 or 4 without checking that a Command Complete or Command Status event is long enough. hci_send_to_sock() invokes the filter before hci_event_packet() validates the event header. A malformed event supplied by a controller or a vhci device can therefore cause an out-of-bounds read. Keep the unmasked event code for the opcode checks. The masked value is needed for the 64-bit event bitmap, but using it to identify command events aliases event codes above 0x3f. In particular, Synchronous Train Complete (0x4f) was treated as Command Status (0x0f) even though its payload has no opcode. Reject actual command events that are too short for the field being inspected. A truncated command event cannot match a configured opcode. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Cc: stable@vger.kernel.org Signed-off-by: Aldo Ariel Panzardo Signed-off-by: Luiz Augusto von Dentz --- net/bluetooth/hci_sock.c | 17 ++++++++++++++--- 1 file changed, 14 insertions(+), 3 deletions(-) diff --git a/net/bluetooth/hci_sock.c b/net/bluetooth/hci_sock.c index 6413593eee51..6d56c77741e1 100644 --- a/net/bluetooth/hci_sock.c +++ b/net/bluetooth/hci_sock.c @@ -164,6 +164,7 @@ static bool is_filtered_packet(struct sock *sk, struct sk_buff *skb) { struct hci_filter *flt; int flt_type, flt_event; + u8 event; /* Apply filter */ flt = &hci_pi(sk)->filter; @@ -177,7 +178,11 @@ static bool is_filtered_packet(struct sock *sk, struct sk_buff *skb) if (hci_skb_pkt_type(skb) != HCI_EVENT_PKT) return false; - flt_event = (*(__u8 *)skb->data & HCI_FLT_EVENT_BITS); + if (skb->len < 1) + return true; + + event = *(__u8 *)skb->data; + flt_event = event & HCI_FLT_EVENT_BITS; if (!hci_test_bit(flt_event, &flt->event_mask)) return true; @@ -186,11 +191,17 @@ static bool is_filtered_packet(struct sock *sk, struct sk_buff *skb) if (!flt->opcode) return false; - if (flt_event == HCI_EV_CMD_COMPLETE && + if (event == HCI_EV_CMD_COMPLETE && skb->len < 5) + return true; + + if (event == HCI_EV_CMD_COMPLETE && flt->opcode != get_unaligned((__le16 *)(skb->data + 3))) return true; - if (flt_event == HCI_EV_CMD_STATUS && + if (event == HCI_EV_CMD_STATUS && skb->len < 6) + return true; + + if (event == HCI_EV_CMD_STATUS && flt->opcode != get_unaligned((__le16 *)(skb->data + 4))) return true; From 4c94557dd02569efa6c1072a0439addaef9a5224 Mon Sep 17 00:00:00 2001 From: Aldo Ariel Panzardo Date: Tue, 15 Sep 2026 13:03:32 -0300 Subject: [PATCH 008/189] Bluetooth: ISO: balance the parent hold in hci_bind_bis() hci_conn_link() takes a lifetime reference to its parent with hci_conn_get(), but only takes an operational hold on the child. hci_conn_unlink() later balances both a hold and a reference on the parent. The SCO and CIS paths pass a parent acquired from a connect helper, so it already has a hold. For an additional BIS, hci_bind_bis() obtains the parent from hci_conn_hash_lookup_big(), which returns a bare pointer. Unlinking the child then drops the parent's existing hold and can schedule it for disconnection while its socket is still using it. Take a hold on the parent before linking it and drop that hold if linking fails. A successful link transfers the hold to hci_conn_unlink(). Fixes: fa224d0c094a ("Bluetooth: ISO: Reassociate a socket with an active BIS") Cc: stable@vger.kernel.org Signed-off-by: Aldo Ariel Panzardo Signed-off-by: Luiz Augusto von Dentz --- net/bluetooth/hci_conn.c | 5 ++++- 1 file changed, 4 insertions(+), 1 deletion(-) diff --git a/net/bluetooth/hci_conn.c b/net/bluetooth/hci_conn.c index e32bb9c342a5..96195d2fd10f 100644 --- a/net/bluetooth/hci_conn.c +++ b/net/bluetooth/hci_conn.c @@ -2375,10 +2375,13 @@ struct hci_conn *hci_bind_bis(struct hci_dev *hdev, bdaddr_t *dst, __u8 sid, parent = hci_conn_hash_lookup_big(hdev, conn->iso_qos.bcast.big); if (parent && parent != conn) { + hci_conn_hold(parent); link = hci_conn_link(parent, conn); hci_conn_drop(conn); - if (!link) + if (!link) { + hci_conn_drop(parent); return ERR_PTR(-ENOLINK); + } } return conn; From 6c78a213d9070b610c7f418af2c25b66180b7e37 Mon Sep 17 00:00:00 2001 From: Aldo Ariel Panzardo Date: Tue, 15 Sep 2026 13:02:39 -0300 Subject: [PATCH 009/189] Bluetooth: L2CAP: validate frame length before control and FCS access l2cap_data_rcv() unpacks either a two-byte or four-byte control field without first ensuring that it is present. A short ERTM or streaming-mode frame can therefore cause an out-of-bounds read. There is a second short-frame case when CRC16 is enabled. After the control field is pulled, l2cap_check_fcs() subtracts two from skb->len without checking it. If fewer than two bytes remain, the subtraction wraps; skb_trim() leaves the buffer unchanged and the subsequent FCS load reads past the logical end of the frame. Validate that the frame contains both its control field and, when enabled, its FCS before either field is accessed. Fixes: 1c2acffb76d4 ("Bluetooth: Add initial support for ERTM packets transfers") Fixes: fcc203c30d72 ("Bluetooth: Add support for FCS option to L2CAP") Cc: stable@vger.kernel.org Signed-off-by: Aldo Ariel Panzardo Signed-off-by: Luiz Augusto von Dentz --- net/bluetooth/l2cap_core.c | 10 +++++++++- 1 file changed, 9 insertions(+), 1 deletion(-) diff --git a/net/bluetooth/l2cap_core.c b/net/bluetooth/l2cap_core.c index 644e31160d55..aaa2a1cd489a 100644 --- a/net/bluetooth/l2cap_core.c +++ b/net/bluetooth/l2cap_core.c @@ -6702,9 +6702,17 @@ static int l2cap_stream_rx(struct l2cap_chan *chan, struct l2cap_ctrl *control, static int l2cap_data_rcv(struct l2cap_chan *chan, struct sk_buff *skb) { struct l2cap_ctrl *control = &bt_cb(skb)->l2cap; - u16 len; + u16 len, min_len; u8 event; + min_len = test_bit(FLAG_EXT_CTRL, &chan->flags) ? + L2CAP_EXT_CTRL_SIZE : L2CAP_ENH_CTRL_SIZE; + if (chan->fcs == L2CAP_FCS_CRC16) + min_len += L2CAP_FCS_SIZE; + + if (skb->len < min_len) + goto drop; + __unpack_control(chan, skb); len = skb->len; From b5dbb41b212c50c095a4dbee3017a84fe94f033b Mon Sep 17 00:00:00 2001 From: Aldo Ariel Panzardo Date: Tue, 15 Sep 2026 12:59:52 -0300 Subject: [PATCH 010/189] Bluetooth: mgmt: fix race in read_unconf_index_list() read_unconf_index_list() counts unconfigured controllers before allocating its response, then checks the device flags again while filling it. hci_dev_list_lock stabilizes list membership, but it does not serialize the per-device flags. During asynchronous controller setup, the worker can set HCI_UNCONFIGURED and clear HCI_SETUP between the two passes. A controller omitted from the allocation count can then become eligible for the fill pass, causing an out-of-bounds write to rp->index[]. Allocate space for every device on hci_dev_list. Since list membership cannot change while hci_dev_list_lock is held, the response remains large enough regardless of flag transitions. The reported count and response length still include only eligible unconfigured controllers. Fixes: 73d1df2a7a10 ("Bluetooth: Add support for Read Unconfigured Index List command") Cc: stable@vger.kernel.org Signed-off-by: Aldo Ariel Panzardo Signed-off-by: Luiz Augusto von Dentz --- net/bluetooth/mgmt.c | 8 ++------ 1 file changed, 2 insertions(+), 6 deletions(-) diff --git a/net/bluetooth/mgmt.c b/net/bluetooth/mgmt.c index f740e745ae78..41956cdde982 100644 --- a/net/bluetooth/mgmt.c +++ b/net/bluetooth/mgmt.c @@ -496,13 +496,9 @@ static int read_unconf_index_list(struct sock *sk, struct hci_dev *hdev, read_lock(&hci_dev_list_lock); - count = 0; - list_for_each_entry(d, &hci_dev_list, list) { - if (hci_dev_test_flag(d, HCI_UNCONFIGURED)) - count++; - } + count = list_count_nodes(&hci_dev_list); - rp_len = sizeof(*rp) + (2 * count); + rp_len = sizeof(*rp) + (sizeof(__le16) * count); rp = kmalloc(rp_len, GFP_ATOMIC); if (!rp) { read_unlock(&hci_dev_list_lock); From c0078f4d8c8fcd8edfe8c53ffe77d48565c1957e Mon Sep 17 00:00:00 2001 From: Wei Wang Date: Tue, 15 Sep 2026 20:14:12 +0000 Subject: [PATCH 011/189] mailmap: add entry for Wei Wang My Meta email address is no longer active. Map it to my current address so that git and get_maintainer.pl stop pointing at a dead address for my contributions. Signed-off-by: Wei Wang Link: https://patch.msgid.link/20260915201412.2201757-1-weiwan@google.com Signed-off-by: Jakub Kicinski --- .mailmap | 1 + 1 file changed, 1 insertion(+) diff --git a/.mailmap b/.mailmap index 1f5540bc33f1..526217c88495 100644 --- a/.mailmap +++ b/.mailmap @@ -975,6 +975,7 @@ Vladimir Davydov Vlastimil Babka WangYuli WangYuli +Wei Wang Weiwen Hu WeiXiong Liao Wen Gong From 8e0b235bd918d06f54ba8fddd2c3ddc36ca59c15 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Bj=C3=B6rn=20T=C3=B6pel?= Date: Tue, 15 Sep 2026 12:49:15 +0200 Subject: [PATCH 012/189] eth: fbnic: Fix payload page pool error cleanup MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The payload page pool pointer contains an error pointer when its allocation fails. The cleanup path passes that error pointer to page_pool_destroy() instead of destroying the header page pool. This can dereference the error pointer and leave the header page pool allocated. Destroy the header page pool instead. Fixes: 8a11010fdd96 ("eth: fbnic: allocate unreadable page pool for the payloads") Reported-by: Sashiko Link: https://lore.kernel.org/netdev/178915061000.219967.7726187707862333281@kernel.org/ Signed-off-by: Björn Töpel Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260915104917.3978113-1-bjorn@kernel.org Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/meta/fbnic/fbnic_txrx.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c b/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c index e7918d3f6aba..661dee1661af 100644 --- a/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c +++ b/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c @@ -1622,7 +1622,7 @@ fbnic_alloc_qt_page_pools(struct fbnic_net *fbn, struct fbnic_q_triad *qt, return 0; err_destroy_sub0: - page_pool_destroy(pp); + page_pool_destroy(qt->sub0.page_pool); return PTR_ERR(pp); } From 0a5f5d9e94dead312d32c366b917c64e552b72f7 Mon Sep 17 00:00:00 2001 From: Jamal Hadi Salim Date: Wed, 16 Sep 2026 06:01:14 -0400 Subject: [PATCH 013/189] net/sched: cls_u32: fix manual hash table handle IDR aliasing A u32 hash table created with an explicit handle ('tc filter add ... handle 801: u32 divisor N') keys its IDR entry on the raw handle, while the destroy paths free it under handle2id(handle). The two key domains disagree for handles in the 0x800..0xFFF htid range: handle2id() folds them back into the auto-allocated id space (1..0x7FF). A manual table therefore leaves its raw-keyed IDR entry unreachable on delete (a permanent leak), and its delete can drop the idr entry of an unrelated live auto table. A later auto allocation can then hand out a handle that aliases the live manual table; u32_lookup_ht() first-match routes lookups and TCA_U32_LINK for that htid to the wrong table. Key the divisor-path alloc on handle2id(handle) so allocation and removal share one key domain. A manual handle that maps onto an id already in use is rejected with -ENOSPC, and auto allocation skips ids held by live manual tables. Conditions to recreate: ip link add test0 type dummy tc qdisc add dev test0 clsact tc filter add dev test0 ingress protocol ip pref 1 \ handle 801: u32 divisor 16 tc filter add dev test0 ingress protocol ip pref 2 u32 divisor 16 tc -d filter show dev test0 ingress | grep 'fh 801:' # unpatched: two live tables with handle 0x80100000 (the pref 2 root # hnode is auto-allocated id 1); patched: the auto hnode takes id 2. Also tested with a poc with a live u32 table on the block, add/delete a manual table 'handle 901: u32 divisor 1' twice; unpatched, the re-add fails with -ENOSPC because the raw key leaked on the first delete. Fixes: 73af53d82076 ("net: sched: cls_u32: Fix u32's systematic failure to free IDR entries for hnodes.") Reported-by: Sashiko (gemini + nipa) Closes: https://sashiko.dev/#/patchset/20260822222049.114526-1-jhs@mojatatu.com Reviewed-by: Victor Nogueira Tested-by: hybris Signed-off-by: Jamal Hadi Salim Reviewed-by: Simon Horman Link: https://patch.msgid.link/QDISC-LQFE.v1.20260911041746.1@mojatatu.com Signed-off-by: Jakub Kicinski --- net/sched/cls_u32.c | 12 ++++++++++-- 1 file changed, 10 insertions(+), 2 deletions(-) diff --git a/net/sched/cls_u32.c b/net/sched/cls_u32.c index a3e65c8cf29e..76ce2d124079 100644 --- a/net/sched/cls_u32.c +++ b/net/sched/cls_u32.c @@ -1003,8 +1003,16 @@ static int u32_change(struct net *net, struct sk_buff *in_skb, return -ENOMEM; } } else { - err = idr_alloc_u32(&tp_c->handle_idr, ht, &handle, - handle, GFP_KERNEL); + /* The IDR is keyed on the mapped id, and that is + * what the destroy paths remove. Ask for it here, + * so a manual handle colliding with the + * auto-allocated id space is rejected (-ENOSPC) + * instead of aliasing a future auto id. + */ + u32 id = handle2id(handle); + + err = idr_alloc_u32(&tp_c->handle_idr, ht, &id, id, + GFP_KERNEL); if (err) { kfree(ht); return err; From 960ab631f3d8789586db71b9914b361059ddcb6e Mon Sep 17 00:00:00 2001 From: Jamal Hadi Salim Date: Wed, 16 Sep 2026 06:01:15 -0400 Subject: [PATCH 014/189] selftests/tc-testing: add u32 manual table handle IDR tests 35fc: create a manual table with handle 801:, then add an auto-allocated table. Before the fix, the auto allocation reuses id 1 and hands out the same handle 0x80100000, aliasing the manual table; the test requires the manual 801: handle to keep exactly one entry in the dump. a6e8: with a live u32 table keeping the tc_u_common alive, add and delete a manual table with handle 901:, then re-add it. Unpatched, the delete leaks the raw-keyed IDR entry and the re-add fails with -ENOSPC; the test requires the re-add to succeed. Reviewed-by: Victor Nogueira Tested-by: hybris Signed-off-by: Jamal Hadi Salim Reviewed-by: Simon Horman Link: https://patch.msgid.link/QDISC-LQFE.v1.20260911041746.2@mojatatu.com Signed-off-by: Jakub Kicinski --- .../tc-testing/tc-tests/filters/u32.json | 48 +++++++++++++++++++ 1 file changed, 48 insertions(+) diff --git a/tools/testing/selftests/tc-testing/tc-tests/filters/u32.json b/tools/testing/selftests/tc-testing/tc-tests/filters/u32.json index e2b03f2b5e89..edc5148a8d97 100644 --- a/tools/testing/selftests/tc-testing/tc-tests/filters/u32.json +++ b/tools/testing/selftests/tc-testing/tc-tests/filters/u32.json @@ -376,5 +376,53 @@ "teardown": [ "$TC qdisc del dev $DUMMY clsact" ] + }, + { + "id": "35fc", + "name": "u32 manual table then auto table: auto allocation must not alias a live manual handle", + "category": [ + "filter", + "u32" + ], + "plugins": { + "requires": "nsPlugin" + }, + "setup": [ + "$TC qdisc add dev $DEV1 ingress", + "$TC filter add dev $DEV1 ingress protocol ip pref 1 handle 801: u32 divisor 16" + ], + "cmdUnderTest": "$TC filter add dev $DEV1 ingress protocol ip pref 2 u32 divisor 16", + "expExitCode": "0", + "verifyCmd": "$TC -d filter show dev $DEV1 ingress", + "matchPattern": "fh 801:", + "matchCount": "1", + "teardown": [ + "$TC qdisc del dev $DEV1 ingress" + ] + }, + { + "id": "a6e8", + "name": "u32 manual table add/del does not leak its idr entry (re-adding the same handle succeeds)", + "category": [ + "filter", + "u32" + ], + "plugins": { + "requires": "nsPlugin" + }, + "setup": [ + "$TC qdisc add dev $DEV1 ingress", + "$TC filter add dev $DEV1 ingress protocol ip pref 1 u32 divisor 16", + "$TC filter add dev $DEV1 ingress protocol ip pref 5 handle 901: u32 divisor 1", + "$TC filter del dev $DEV1 ingress protocol ip pref 5 handle 901: u32" + ], + "cmdUnderTest": "$TC filter add dev $DEV1 ingress protocol ip pref 6 handle 901: u32 divisor 1", + "expExitCode": "0", + "verifyCmd": "$TC -d filter show dev $DEV1 ingress", + "matchPattern": "fh 901: ht divisor 1", + "matchCount": "1", + "teardown": [ + "$TC qdisc del dev $DEV1 ingress" + ] } ] From daf677c2c6449011ee695d55b48b5b2977a36f88 Mon Sep 17 00:00:00 2001 From: Kyle Hendry Date: Tue, 15 Sep 2026 10:39:20 -0700 Subject: [PATCH 015/189] net: pcs: rzn1-miic: Fix config array initialization Fix memset parameters to initialize the entire DT value array Fixes: f39e968dc168a7bd ("net: pcs: rzn1-miic: Move configuration data to SoC-specific struct") Reviewed-by: Geert Uytterhoeven Signed-off-by: Kyle Hendry Reviewed-by: Lad Prabhakar Link: https://patch.msgid.link/20260915-rzn1-miic-fix-array-v5-1-b7173fd5b97d@reliablecontrols.com Signed-off-by: Jakub Kicinski --- drivers/net/pcs/pcs-rzn1-miic.c | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/drivers/net/pcs/pcs-rzn1-miic.c b/drivers/net/pcs/pcs-rzn1-miic.c index 2b72fa98ddf1..cb74861e823c 100644 --- a/drivers/net/pcs/pcs-rzn1-miic.c +++ b/drivers/net/pcs/pcs-rzn1-miic.c @@ -683,7 +683,8 @@ static int miic_parse_dt(struct miic *miic, u32 *mode_cfg) if (!dt_val) return -ENOMEM; - memset(dt_val, MIIC_MODCTRL_CONF_NONE, sizeof(*dt_val)); + memset(dt_val, MIIC_MODCTRL_CONF_NONE, + sizeof(*dt_val) * miic->of_data->conf_conv_count); if (of_property_read_u32(np, "renesas,miic-switch-portin", &conf) == 0) dt_val[0] = conf; From 39c6580765dad6477fb2637f6f616e0d276aae65 Mon Sep 17 00:00:00 2001 From: Heyang Tan Date: Mon, 14 Sep 2026 10:05:21 +0800 Subject: [PATCH 016/189] octeontx2-af: use seq_file for rsrc_alloc debugfs The rsrc_alloc debugfs reader writes rows directly to userspace without respecting the caller's read count. It also uses the current row length as the userspace stride, which can corrupt output when rows have different widths. Use seq_file to handle userspace buffer sizes, offsets, and partial reads, and write output columns directly to the seq_file buffer. Fixes: 23205e6d06d4 ("octeontx2-af: Dump current resource provisioning status") Signed-off-by: Heyang Tan Reviewed-by: Ratheesh Kannoth Link: https://patch.msgid.link/20260914020521.146-1-thy15333007817@163.com Signed-off-by: Jakub Kicinski --- .../marvell/octeontx2/af/rvu_debugfs.c | 104 ++++++------------ 1 file changed, 33 insertions(+), 71 deletions(-) diff --git a/drivers/net/ethernet/marvell/octeontx2/af/rvu_debugfs.c b/drivers/net/ethernet/marvell/octeontx2/af/rvu_debugfs.c index 904374baae6f..2927633465d9 100644 --- a/drivers/net/ethernet/marvell/octeontx2/af/rvu_debugfs.c +++ b/drivers/net/ethernet/marvell/octeontx2/af/rvu_debugfs.c @@ -714,110 +714,72 @@ static int get_max_column_width(struct rvu *rvu) } /* Dumps current provisioning status of all RVU block LFs */ -static ssize_t rvu_dbg_rsrc_attach_status(struct file *filp, - char __user *buffer, - size_t count, loff_t *ppos) +static int rvu_dbg_rsrc_attach_status(struct seq_file *filp, void *unused) { - int index, off = 0, flag = 0, len = 0, i = 0; - struct rvu *rvu = filp->private_data; - int bytes_not_copied = 0; + struct rvu *rvu = filp->private; + int index, pf, vf, pcifunc; struct rvu_block block; - int pf, vf, pcifunc; - int buf_size = 2048; int lf_str_size; char *lfs; - char *buf; - /* don't allow partial reads */ - if (*ppos != 0) - return 0; - - buf = kzalloc(buf_size, GFP_KERNEL); - if (!buf) - return -ENOMEM; - - /* Get the maximum width of a column */ lf_str_size = get_max_column_width(rvu); + if (lf_str_size < 0) + return lf_str_size; lfs = kzalloc(lf_str_size, GFP_KERNEL); - if (!lfs) { - kfree(buf); + if (!lfs) return -ENOMEM; - } - off += scnprintf(&buf[off], buf_size - 1 - off, "%-*s", lf_str_size, - "pcifunc"); + + seq_printf(filp, "%-*s", lf_str_size, "pcifunc"); for (index = 0; index < BLK_COUNT; index++) - if (strlen(rvu->hw->block[index].name)) { - off += scnprintf(&buf[off], buf_size - 1 - off, - "%-*s", lf_str_size, - rvu->hw->block[index].name); - } + if (strlen(rvu->hw->block[index].name)) + seq_printf(filp, "%-*s", lf_str_size, + rvu->hw->block[index].name); - off += scnprintf(&buf[off], buf_size - 1 - off, "\n"); - bytes_not_copied = copy_to_user(buffer + (i * off), buf, off); - if (bytes_not_copied) - goto out; - - i++; - *ppos += off; + seq_putc(filp, '\n'); for (pf = 0; pf < rvu->hw->total_pfs; pf++) { for (vf = 0; vf <= rvu->hw->total_vfs; vf++) { - off = 0; - flag = 0; pcifunc = rvu_make_pcifunc(rvu->pdev, pf, vf); if (!pcifunc) continue; - if (vf) { - sprintf(lfs, "PF%d:VF%d", pf, vf - 1); - off = scnprintf(&buf[off], - buf_size - 1 - off, - "%-*s", lf_str_size, lfs); - } else { - sprintf(lfs, "PF%d", pf); - off = scnprintf(&buf[off], - buf_size - 1 - off, - "%-*s", lf_str_size, lfs); - } - for (index = 0; index < BLK_COUNT; index++) { block = rvu->hw->block[index]; if (!strlen(block.name)) continue; - len = 0; - lfs[len] = '\0'; + lfs[0] = '\0'; get_lf_str_list(&block, pcifunc, lfs); if (strlen(lfs)) - flag = 1; - - off += scnprintf(&buf[off], buf_size - 1 - off, - "%-*s", lf_str_size, lfs); + break; } - if (flag) { - off += scnprintf(&buf[off], - buf_size - 1 - off, "\n"); - bytes_not_copied = copy_to_user(buffer + - (i * off), - buf, off); - if (bytes_not_copied) - goto out; + if (index == BLK_COUNT) + continue; - i++; - *ppos += off; + if (vf) + sprintf(lfs, "PF%d:VF%d", pf, vf - 1); + else + sprintf(lfs, "PF%d", pf); + seq_printf(filp, "%-*s", lf_str_size, lfs); + + for (index = 0; index < BLK_COUNT; index++) { + block = rvu->hw->block[index]; + if (!strlen(block.name)) + continue; + + lfs[0] = '\0'; + get_lf_str_list(&block, pcifunc, lfs); + seq_printf(filp, "%-*s", lf_str_size, lfs); } + seq_putc(filp, '\n'); } } -out: kfree(lfs); - kfree(buf); - if (bytes_not_copied) - return -EFAULT; - return *ppos; + return 0; } -RVU_DEBUG_FOPS(rsrc_status, rsrc_attach_status, NULL); +RVU_DEBUG_SEQ_FOPS(rsrc_status, rsrc_attach_status, NULL); static int rvu_dbg_rvu_pf_cgx_map_display(struct seq_file *filp, void *unused) { From d09e8f64653c93da5793c16be19330968f2a32e6 Mon Sep 17 00:00:00 2001 From: Shay Drory Date: Tue, 15 Sep 2026 14:34:57 +0300 Subject: [PATCH 017/189] net/mlx5: devcom, Base component size on linked devices mlx5_devcom_comp_get_size() returns the component's kref count. That kref is bumped in mlx5_devcom_register_component() under comp_list_lock, before the comp_dev is linked onto comp_dev_list_head under comp->sem. The event broadcast (mlx5_devcom_locked_send_event()) walks that list. Hence, a caller can read the expected size, but send_event won't be sent to all peers. In the SD group registration path, this lets a member broadcast its role-election event over an incomplete list, electing a primary that never completes the group, is never marked ready, and leaves the group with a stale primary. Track the number of linked comp_devs in a dedicated counter, maintained under comp->sem together with the list add/remove, and return it from mlx5_devcom_comp_get_size(). Fixes: 9bb1ac80738a ("net/mlx5: devcom, Add component size getter") Signed-off-by: Shay Drory Reviewed-by: Akiva Goldberger Signed-off-by: Tariq Toukan Link: https://patch.msgid.link/20260915113459.3934760-2-tariqt@nvidia.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/mellanox/mlx5/core/lib/devcom.c | 5 ++++- 1 file changed, 4 insertions(+), 1 deletion(-) diff --git a/drivers/net/ethernet/mellanox/mlx5/core/lib/devcom.c b/drivers/net/ethernet/mellanox/mlx5/core/lib/devcom.c index 64f92427602d..75855481522b 100644 --- a/drivers/net/ethernet/mellanox/mlx5/core/lib/devcom.c +++ b/drivers/net/ethernet/mellanox/mlx5/core/lib/devcom.c @@ -37,6 +37,7 @@ struct mlx5_devcom_comp { struct mlx5_devcom_key key; mlx5_devcom_event_handler_t handler; struct kref ref; + int nr_devs; bool ready; struct rw_semaphore sem; struct lock_class_key lock_key; @@ -170,6 +171,7 @@ devcom_alloc_comp_dev(struct mlx5_devcom_dev *devc, down_write(&comp->sem); list_add_tail(&devcom->list, &comp->comp_dev_list_head); + WRITE_ONCE(comp->nr_devs, comp->nr_devs + 1); up_write(&comp->sem); return devcom; @@ -182,6 +184,7 @@ devcom_free_comp_dev(struct mlx5_devcom_comp_dev *devcom) down_write(&comp->sem); list_del(&devcom->list); + WRITE_ONCE(comp->nr_devs, comp->nr_devs - 1); up_write(&comp->sem); kref_put(&devcom->devc->ref, mlx5_devcom_dev_release); @@ -284,7 +287,7 @@ int mlx5_devcom_comp_get_size(struct mlx5_devcom_comp_dev *devcom) { struct mlx5_devcom_comp *comp = devcom->comp; - return kref_read(&comp->ref); + return READ_ONCE(comp->nr_devs); } int mlx5_devcom_locked_send_event(struct mlx5_devcom_comp_dev *devcom, From e1e29ada2b938b13ba689a06a8bd8604564da2b3 Mon Sep 17 00:00:00 2001 From: Shay Drory Date: Tue, 15 Sep 2026 14:34:58 +0300 Subject: [PATCH 018/189] net/mlx5: SD, unload reps on shared FDB create error path mlx5_lag_shared_fdb_create() sets sd_fdb_active on every group member before reloading the representors, so mlx5_lag_is_active() is already true and the guard in mlx5_esw_offloads_rep_load() does not skip the VF/SF reps. If the reload then fails, the error path clears sd_fdb_active and destroys the shared FDB, leaving the reps loaded while SD LAG is inactive - the state cited commit was written to prevent. Unload the reps in the error path as well. Fixes: 68c2dd59a6c7 ("net/mlx5: E-Switch, Tie rep load/unload to SD LAG state") Signed-off-by: Shay Drory Reviewed-by: Akiva Goldberger Signed-off-by: Tariq Toukan Link: https://patch.msgid.link/20260915113459.3934760-3-tariqt@nvidia.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/mellanox/mlx5/core/lag/shared_fdb.c | 1 + 1 file changed, 1 insertion(+) diff --git a/drivers/net/ethernet/mellanox/mlx5/core/lag/shared_fdb.c b/drivers/net/ethernet/mellanox/mlx5/core/lag/shared_fdb.c index 6b4ad3c53f2f..424040918fa3 100644 --- a/drivers/net/ethernet/mellanox/mlx5/core/lag/shared_fdb.c +++ b/drivers/net/ethernet/mellanox/mlx5/core/lag/shared_fdb.c @@ -270,6 +270,7 @@ int mlx5_lag_shared_fdb_create(struct mlx5_lag *ldev, pf->sd_fdb_active = false; } mlx5_lag_destroy_single_fdb_filter(ldev, group_id); + mlx5_lag_unload_reps_from_locked(ldev, filter); } err_add_devices: mlx5_lag_add_devices_filter(ldev, filter); From bae23d1ae62092c7f0ec6d5f7e1be5d164822638 Mon Sep 17 00:00:00 2001 From: Shay Drory Date: Tue, 15 Sep 2026 14:34:59 +0300 Subject: [PATCH 019/189] net/mlx5: LAG, reload IB reps of LAG master before the rest In a shared-FDB LAG the master device creates the bond IB device; the other LAG members do not create their own, they populate a port inside the master's IB device. mlx5_lag_reload_ib_reps_unlocked() reloaded the members' IB reps in iteration order, with no guarantee the master is reloaded first. When a non-master member is reloaded before the master, it tries to populate its port in an IB device that has not been recreated yet. Hence, reload the master's IB reps first, then every other member. Fixes: 2b204cdb1206 ("net/mlx5: LAG, use xa_alloc to manage LAG device indices") Signed-off-by: Shay Drory Reviewed-by: Akiva Goldberger Signed-off-by: Tariq Toukan Link: https://patch.msgid.link/20260915113459.3934760-4-tariqt@nvidia.com Signed-off-by: Jakub Kicinski --- .../net/ethernet/mellanox/mlx5/core/lag/lag.c | 44 ++++++++++++++----- 1 file changed, 32 insertions(+), 12 deletions(-) diff --git a/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c b/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c index c655f6e32e9b..dd14cdc378de 100644 --- a/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c +++ b/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c @@ -1266,25 +1266,45 @@ void mlx5_lag_remove_devices(struct mlx5_lag *ldev) mlx5_lag_remove_devices_filter(ldev, MLX5_LAG_FILTER_PORTS); } +static int mlx5_lag_reload_ib_reps_idx(struct mlx5_lag *ldev, int idx, + u32 flags) +{ + struct lag_func *pf = mlx5_lag_pf(ldev, idx); + struct mlx5_eswitch *esw; + int ret; + + if (pf->dev->priv.flags & flags) + return 0; + + esw = pf->dev->priv.eswitch; + mlx5_esw_reps_block(esw); + ret = mlx5_eswitch_reload_ib_reps(esw); + mlx5_esw_reps_unblock(esw); + + return ret; +} + static int mlx5_lag_reload_ib_reps_unlocked(struct mlx5_lag *ldev, u32 flags, u32 filter, bool cont_on_fail) { - struct lag_func *pf; + int master_idx = mlx5_lag_get_dev_index_by_seq_filter(ldev, MLX5_LAG_P1, + filter); int ret; int i; - mlx5_lag_for_each(i, 0, ldev, filter) { - pf = mlx5_lag_pf(ldev, i); - if (!(pf->dev->priv.flags & flags)) { - struct mlx5_eswitch *esw; + if (master_idx < 0) + return -EINVAL; - esw = pf->dev->priv.eswitch; - mlx5_esw_reps_block(esw); - ret = mlx5_eswitch_reload_ib_reps(esw); - mlx5_esw_reps_unblock(esw); - if (ret && !cont_on_fail) - return ret; - } + ret = mlx5_lag_reload_ib_reps_idx(ldev, master_idx, flags); + if (ret && !cont_on_fail) + return ret; + + mlx5_lag_for_each(i, 0, ldev, filter) { + if (i == master_idx) + continue; + ret = mlx5_lag_reload_ib_reps_idx(ldev, i, flags); + if (ret && !cont_on_fail) + return ret; } return 0; From 2ec28c09b320ba241bea8a70ee5cb9ccf4a099e8 Mon Sep 17 00:00:00 2001 From: Aldo Ariel Panzardo Date: Tue, 15 Sep 2026 14:30:50 -0300 Subject: [PATCH 020/189] vsock: ignore empty child namespace mode writes __vsock_net_mode_string() returns success without updating new_mode when the transfer length is zero. Its caller then reads the uninitialized enum and may permanently store a stack-derived value in the write-once child mode. Return before calling __vsock_net_mode_string() when *lenp is zero so that the helper is never invoked with nothing to parse and new_mode is never read uninitialized. This also prevents an empty write from locking the current mode. Fixes: eafb64f40ca4 ("vsock: add netns to vsock core") Cc: stable@vger.kernel.org Reviewed-by: Luigi Leonardi Signed-off-by: Aldo Ariel Panzardo Reviewed-by: Stefano Garzarella Reviewed-by: Bobby Eshleman Link: https://patch.msgid.link/20260915173050.3176344-1-qwe.aldo@gmail.com Signed-off-by: Jakub Kicinski --- net/vmw_vsock/af_vsock.c | 3 +++ 1 file changed, 3 insertions(+) diff --git a/net/vmw_vsock/af_vsock.c b/net/vmw_vsock/af_vsock.c index f840498b58af..9b71479a2b29 100644 --- a/net/vmw_vsock/af_vsock.c +++ b/net/vmw_vsock/af_vsock.c @@ -2889,6 +2889,9 @@ static int vsock_net_child_mode_string(const struct ctl_table *table, int write, net = container_of(table->data, struct net, vsock.child_ns_mode); + if (!*lenp) + return 0; + ret = __vsock_net_mode_string(table, write, buffer, lenp, ppos, vsock_net_child_mode(net), &new_mode); if (ret) From 651010592bdce7005c1179498327e51bfc4fe1a5 Mon Sep 17 00:00:00 2001 From: Zhang Yunfei Date: Fri, 11 Sep 2026 17:11:23 +0800 Subject: [PATCH 021/189] net: txgbe: fix FDIR filter restore for VF rules txgbe_fdir_filter_restore() reprograms every filter from txgbe->fdir_filter_list after a reset. It extracts the ring part of filter->action with ethtool_get_flow_spec_ring() and maps it onto a PF rx ring, silently dropping the VF part of the cookie that txgbe_add_ethtool_fdir_entry() stores there (input->action = fsp->ring_cookie). For a rule directed at a VF, restore therefore reprograms the filter to the PF queue with the same ring index: after any down/up or txgbe_reinit_locked(), traffic matching the rule is steered to the PF instead of the VF. Handle VF rules the same way txgbe_add_ethtool_fdir_entry() does: validate vf against wx->num_vfs and ring against wx->num_rx_queues_per_pool, and map the ring onto the absolute queue index ((vf - 1) * wx->num_rx_queues_per_pool) + ring. Fixes: 7a91722e0dd4 ("net: txgbe: Support the FDIR rules assigned to VFs") Cc: stable@vger.kernel.org Signed-off-by: Zhang Yunfei Link: https://patch.msgid.link/20260911091123.798931-1-zhangyunfei1@kylinos.cn Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/wangxun/txgbe/txgbe_fdir.c | 13 +++++++++++-- 1 file changed, 11 insertions(+), 2 deletions(-) diff --git a/drivers/net/ethernet/wangxun/txgbe/txgbe_fdir.c b/drivers/net/ethernet/wangxun/txgbe/txgbe_fdir.c index a84010828551..59a47532618c 100644 --- a/drivers/net/ethernet/wangxun/txgbe/txgbe_fdir.c +++ b/drivers/net/ethernet/wangxun/txgbe/txgbe_fdir.c @@ -591,15 +591,24 @@ static void txgbe_fdir_filter_restore(struct wx *wx) queue = TXGBE_RDB_FDIR_DROP_QUEUE; } else { u32 ring = ethtool_get_flow_spec_ring(filter->action); + u8 vf = ethtool_get_flow_spec_ring_vf(filter->action); - if (ring >= wx->num_rx_queues) { + if (!vf && ring >= wx->num_rx_queues) { wx_err(wx, "FDIR restore failed, ring:%u\n", ring); continue; + } else if (vf && (vf > wx->num_vfs || + ring >= wx->num_rx_queues_per_pool)) { + wx_err(wx, "FDIR restore failed, vf:%u, ring:%u\n", + vf, ring); + continue; } /* Map the ring onto the absolute queue index */ - queue = wx->rx_ring[ring]->reg_idx; + if (!vf) + queue = wx->rx_ring[ring]->reg_idx; + else + queue = ((vf - 1) * wx->num_rx_queues_per_pool) + ring; } ret = txgbe_fdir_write_perfect_filter(wx, From 46bc52d13594848023e681860df8700c8db14354 Mon Sep 17 00:00:00 2001 From: Linkui Xiao Date: Wed, 16 Sep 2026 20:53:16 +0800 Subject: [PATCH 022/189] ipv4: fib: fix data-race and stale genid check around nh->nh_saddr fib_select_multipath() compares nexthop_nh->nh_saddr against the flow source address with no lock held, while fib_info_update_nhc_saddr() stores a new value from another CPU as soon as the preferred source address of the egress device changes. Commit 195374d89368 ("ipv4: fib: annotate races around nh->nh_saddr_genid and nh->nh_saddr") added WRITE_ONCE() on the store side and READ_ONCE() in fib_result_prefsrc() after syzbot reported BUG: KCSAN: data-race in fib_select_path / fib_select_path but it only covered that reader. fib_select_multipath(), reached from fib_select_path(), is a second lockless reader of nh->nh_saddr and was left bare. Moreover, nh_saddr is only meaningful when nh_saddr_genid matches dev_addr_genid, as established by commit 436c3b66ec98 ("ipv4: Invalidate nexthop cache nh_saddr more correctly."). fib_select_multipath() skips that validation, so it can score a nexthop using a stale source address and skew the ECMP selection. Annotate both reads with READ_ONCE() and refresh the cached source address via fib_info_update_nhc_saddr() when the genid does not match, mirroring fib_result_prefsrc(). Fixes: 32607a332cfe ("ipv4: prefer multipath nexthop that matches source address") Signed-off-by: Linkui Xiao Reviewed-by: Ido Schimmel Reviewed-by: Eric Dumazet Link: https://patch.msgid.link/20260916125316.988044-1-xiaolinkui@126.com Signed-off-by: Jakub Kicinski --- net/ipv4/fib_semantics.c | 13 ++++++++++++- 1 file changed, 12 insertions(+), 1 deletion(-) diff --git a/net/ipv4/fib_semantics.c b/net/ipv4/fib_semantics.c index 7a362f2e2c2b..5c9021ea3a79 100644 --- a/net/ipv4/fib_semantics.c +++ b/net/ipv4/fib_semantics.c @@ -2176,6 +2176,15 @@ static bool fib_good_nh(const struct fib_nh *nh) return !!(state & NUD_VALID); } +static __be32 fib_nh_saddr(struct net *net, const struct fib_info *fi, + struct fib_nh *nh, int genid) +{ + if (READ_ONCE(nh->nh_saddr_genid) == genid) + return READ_ONCE(nh->nh_saddr); + + return fib_info_update_nhc_saddr(net, &nh->nh_common, fi->fib_scope); +} + void fib_select_multipath(struct fib_result *res, int hash, const struct flowi4 *fl4) { @@ -2184,6 +2193,7 @@ void fib_select_multipath(struct fib_result *res, int hash, bool use_neigh; int score = -1; __be32 saddr; + int genid; if (unlikely(res->fi->nh)) { nexthop_path_fib_result(res, hash); @@ -2192,6 +2202,7 @@ void fib_select_multipath(struct fib_result *res, int hash, use_neigh = READ_ONCE(net->ipv4.sysctl_fib_multipath_use_neigh); saddr = fl4 ? fl4->saddr : 0; + genid = saddr ? atomic_read(&net->ipv4.dev_addr_genid) : 0; change_nexthops(fi) { int nh_upper_bound, nh_score = 0; @@ -2204,7 +2215,7 @@ void fib_select_multipath(struct fib_result *res, int hash, (use_neigh && !fib_good_nh(nexthop_nh))) continue; - if (saddr && nexthop_nh->nh_saddr == saddr) + if (saddr && fib_nh_saddr(net, fi, nexthop_nh, genid) == saddr) nh_score += 2; if (hash <= nh_upper_bound) nh_score++; From d644b23afe1ef509c9961a6d84a093c2587edf02 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?J=C3=A9r=C3=A9my=20Jean?= Date: Tue, 18 Aug 2026 20:00:15 +0000 Subject: [PATCH 023/189] netfilter: flowtable: publish HW_DEAD after worker is done MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit flow_offload_work_del() sets NF_FLOW_HW_DEAD before the work handler clears NF_FLOW_HW_PENDING. Once a flow is both HW_DYING and HW_DEAD, a concurrent garbage collection pass can remove it and schedule it for RCU freeing. The offload worker holds neither an RCU read lock nor a reference to the flow. If it is preempted after publishing HW_DEAD, the RCU callback can free the flow before the worker resumes and clears HW_PENDING, resulting in a use-after-free. Move HW_DEAD publication to the common worker epilogue after the pending bit is cleared, making it the final flow access by destroy work. Order all preceding flow accesses before publishing the bit that allows garbage collection to free the object. Fixes: 2c8897953f3b ("netfilter: flowtable: Add pending bit for offload work") Assisted-by: Codex:gpt-5 Signed-off-by: Jérémy Jean Signed-off-by: Pablo Neira Ayuso --- net/netfilter/nf_flow_table_offload.c | 7 ++++++- 1 file changed, 6 insertions(+), 1 deletion(-) diff --git a/net/netfilter/nf_flow_table_offload.c b/net/netfilter/nf_flow_table_offload.c index 801a3dd9ceea..6757fd89c1f1 100644 --- a/net/netfilter/nf_flow_table_offload.c +++ b/net/netfilter/nf_flow_table_offload.c @@ -995,7 +995,6 @@ static void flow_offload_work_del(struct flow_offload_work *offload) flow_offload_tuple_del(offload, FLOW_OFFLOAD_DIR_ORIGINAL); if (test_bit(NF_FLOW_HW_BIDIRECTIONAL, &offload->flow->flags)) flow_offload_tuple_del(offload, FLOW_OFFLOAD_DIR_REPLY); - set_bit(NF_FLOW_HW_DEAD, &offload->flow->flags); } static void flow_offload_tuple_stats(struct flow_offload_work *offload, @@ -1059,6 +1058,12 @@ static void flow_offload_work_handler(struct work_struct *work) } clear_bit(NF_FLOW_HW_PENDING, &offload->flow->flags); + if (offload->cmd == FLOW_CLS_DESTROY) { + /* Publish after the worker's last flow access. */ + smp_mb__before_atomic(); + set_bit(NF_FLOW_HW_DEAD, &offload->flow->flags); + } + kfree(offload); } From 9461613afc59acef44a0071b0dd5075f6e993ffe Mon Sep 17 00:00:00 2001 From: Florian Westphal Date: Thu, 3 Sep 2026 02:41:46 +0200 Subject: [PATCH 024/189] netfilter: nfnetlink_queue: hold nfnl mutex in event notifier We must serialize the release notifier and the config netlink function. A concurrent thread can issue close() which can call the release function while unrelated socket processes UNBIND request for same portid: Oops: general protection fault, [..] RIP: 0010:__instance_destroy+0x60/0x210 [nfnetlink_queue] Call Trace: nfqnl_recv_config+0x9b0/0xdc0 [nfnetlink_queue] nfnetlink_rcv_msg+0x7c2/0xeb0 ? __pfx_nfnetlink_rcv_msg+0x10/0x10 After this, parallel UNBIND and URELEASE events are impossible. This change isn't nice, but its the shortest fix given instances are not refcounted and the nfnetlink config callback drops the rcu read lock early due to need for sleeping allocations. Fixes: 7af4cc3fa158 ("[NETFILTER]: Add "nfnetlink_queue" netfilter queue handler over nfnetlink") Signed-off-by: Florian Westphal Signed-off-by: Pablo Neira Ayuso --- net/netfilter/nfnetlink_queue.c | 8 +++++--- 1 file changed, 5 insertions(+), 3 deletions(-) diff --git a/net/netfilter/nfnetlink_queue.c b/net/netfilter/nfnetlink_queue.c index c727668b0c5b..a3bc00280051 100644 --- a/net/netfilter/nfnetlink_queue.c +++ b/net/netfilter/nfnetlink_queue.c @@ -1593,6 +1593,7 @@ nfqnl_rcv_nl_event(struct notifier_block *this, if (event == NETLINK_URELEASE && n->protocol == NETLINK_NETFILTER) { int i; + nfnl_lock(NFNL_SUBSYS_QUEUE); /* destroy all instances for this portid */ spin_lock(&q->instances_lock); for (i = 0; i < INSTANCE_BUCKETS; i++) { @@ -1606,6 +1607,7 @@ nfqnl_rcv_nl_event(struct notifier_block *this, } } spin_unlock(&q->instances_lock); + nfnl_unlock(NFNL_SUBSYS_QUEUE); } return NOTIFY_DONE; } @@ -1925,9 +1927,9 @@ static int nfqnl_recv_config(struct sk_buff *skb, const struct nfnl_info *info, /* Lookup queue under RCU. After peer_portid check (or for new queue * in BIND case), the queue is owned by the socket sending this message. - * A socket cannot simultaneously send a message and close, so while - * processing this CONFIG message, nfqnl_rcv_nl_event() (triggered by - * socket close) cannot destroy this queue. Safe to use without RCU. + * nfqnl_rcv_nl_event() will block on the nfnl subsys mutex that is + * held by the caller, so the queue cannot be destroyed in parallel, + * even after we drop the RCU read lock. */ rcu_read_lock(); queue = instance_lookup(q, queue_num); From 1b9b5323725e458906c7620a3bc10398b51ad954 Mon Sep 17 00:00:00 2001 From: Weiming Shi Date: Sun, 6 Sep 2026 16:44:10 +0800 Subject: [PATCH 025/189] netfilter: ip6t_rpfilter: reject routes without inet6_dev ip6_route_lookup() can return an error-free route whose rt6i_idev is NULL. Lowering an external nexthop device's MTU below IPV6_MIN_MTU tears down its inet6_dev while fib6_ifdown() leaves routes using nexthop objects in the FIB. An unprivileged user can construct this state with rtnetlink in a private user and network namespace, then trigger a NULL dereference through an IPv6 rpfilter lookup: Oops: general protection fault, probably for non-canonical address 0xdffffc0000000000 KASAN: null-ptr-deref in range [0x0000000000000000-0x0000000000000007] RIP: rpfilter_mt (net/ipv6/netfilter/ip6t_rpfilter.c:75) Call Trace: ip6t_do_table (net/ipv6/netfilter/ip6_tables.c:316) nf_hook_slow (net/netfilter/core.c:619) ipv6_rcv (net/ipv6/ip6_input.c:351) __netif_receive_skb_one_core (net/core/dev.c:6216) process_backlog (net/core/dev.c:6680) __napi_poll (net/core/dev.c:7739) net_rx_action (net/core/dev.c:7959) handle_softirqs (kernel/softirq.c:622) do_softirq.part.0 (kernel/softirq.c:523) __local_bh_enable_ip (kernel/softirq.c:450) __dev_queue_xmit (net/core/dev.c:4913) packet_sendmsg (net/packet/af_packet.c:3139) __sys_sendto (net/socket.c:2252) __x64_sys_sendto (net/socket.c:2259) do_syscall_64 (arch/x86/entry/syscall_64.c:94) entry_SYSCALL_64_after_hwframe (arch/x86/entry/entry_64.S:121) Kernel panic - not syncing: Fatal exception in interrupt Reject routes without an inet6_dev immediately after lookup. Such routes are not eligible for reverse-path filtering, and the check protects all later rt6i_idev dereferences. Fixes: e26f9a480fb6 ("netfilter: add ipv6 reverse path filter match") Reported-by: co+459f67f4d8af8ce6@bugs.sh Closes: https://lore.kernel.org/all/VtWUkE8QzJt5CroTj2V2v3ZQ0gwbXZ7nq7I3@bugs.sh/ Suggested-by: Florian Westphal Assisted-by: Claude:gpt-5 Cc: stable@vger.kernel.org Signed-off-by: Weiming Shi Signed-off-by: Pablo Neira Ayuso --- net/ipv6/netfilter/ip6t_rpfilter.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/net/ipv6/netfilter/ip6t_rpfilter.c b/net/ipv6/netfilter/ip6t_rpfilter.c index 67c87a88cde4..b5def30c3127 100644 --- a/net/ipv6/netfilter/ip6t_rpfilter.c +++ b/net/ipv6/netfilter/ip6t_rpfilter.c @@ -61,7 +61,7 @@ static bool rpfilter_lookup_reverse6(struct net *net, const struct sk_buff *skb, fl6.flowi6_oif = dev->ifindex; rt = (void *)ip6_route_lookup(net, &fl6, skb, lookup_flags); - if (rt->dst.error) + if (rt->dst.error || !rt->rt6i_idev) goto out; if (rt->rt6i_flags & (RTF_REJECT|RTF_ANYCAST)) From 82313c169eddc02b1bf5ba6b427803e272d3ec42 Mon Sep 17 00:00:00 2001 From: Luxiao Xu Date: Sun, 6 Sep 2026 21:29:55 +0800 Subject: [PATCH 026/189] netfilter: ip6t_rt: fix zero-address non-strict match out-of-bounds read rt_mt6_check() permits rules to be configured with rtinfo->addrnr == 0 even when address matching (IP6T_RT_FST_MASK) is requested. In the IP6T_RT_FST_NSTRICT path, rt_mt6() evaluates packet routing addresses against rtinfo->addrs[i] and terminates backwards at the bottom of the loop: if (ipv6_addr_equal(ap, &rtinfo->addrs[i])) { i++; } if (i == rtinfo->addrnr) break; When addrnr is 0, if the first packet address matches rtinfo->addrs[0], i is incremented to 1. Because i is now strictly greater than addrnr (0), the loop termination condition (i == rtinfo->addrnr) is bypassed and will never be satisfied. If a crafted IPv6 packet contains matching routing addresses, i will advance past IP6T_RT_HOPS (16). The subsequent call to ipv6_addr_equal() reads beyond struct ip6t_rt, triggering UBSAN/KASAN out-of-bounds warnings or kernel panics. Fix this by: 1. Rejecting rules in rt_mt6_check() where IP6T_RT_FST_MASK is set but rtinfo->addrnr is zero. 2. In rt_mt6(), moving the termination condition (i < rtinfo->addrnr) into the for-loop header condition and removing the backwards break at the end of the loop body. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Cc: stable@vger.kernel.org Reported-by: Vega Suggested-by: Florian Westphal Assisted-by: LLM Signed-off-by: Luxiao Xu Signed-off-by: Ren Wei Signed-off-by: Pablo Neira Ayuso --- net/ipv6/netfilter/ip6t_rt.c | 11 ++++++++--- 1 file changed, 8 insertions(+), 3 deletions(-) diff --git a/net/ipv6/netfilter/ip6t_rt.c b/net/ipv6/netfilter/ip6t_rt.c index 8051425213dd..9880faf3cc7d 100644 --- a/net/ipv6/netfilter/ip6t_rt.c +++ b/net/ipv6/netfilter/ip6t_rt.c @@ -96,7 +96,8 @@ static bool rt_mt6(const struct sk_buff *skb, struct xt_action_param *par) unsigned int i = 0; for (temp = 0; - temp < (unsigned int)((hdrlen - 8) / 16); + temp < (unsigned int)((hdrlen - 8) / 16) && + i < rtinfo->addrnr; temp++) { ap = skb_header_pointer(skb, ptr @@ -112,8 +113,6 @@ static bool rt_mt6(const struct sk_buff *skb, struct xt_action_param *par) if (ipv6_addr_equal(ap, &rtinfo->addrs[i])) i++; - if (i == rtinfo->addrnr) - break; } if (i == rtinfo->addrnr) return ret; @@ -162,6 +161,12 @@ static int rt_mt6_check(const struct xt_mtchk_param *par) pr_info_ratelimited("too many addresses specified\n"); return -EINVAL; } + + if ((rtinfo->flags & IP6T_RT_FST_MASK) && !rtinfo->addrnr) { + pr_info_ratelimited("address list match requested but addrnr is 0\n"); + return -EINVAL; + } + if ((rtinfo->flags & (IP6T_RT_RES | IP6T_RT_FST_MASK)) && (!(rtinfo->flags & IP6T_RT_TYP) || (rtinfo->rt_type != 0) || From a311a898172743558b82f6035ef2aa8c310a4223 Mon Sep 17 00:00:00 2001 From: Karl Mehltretter Date: Thu, 10 Sep 2026 22:02:28 +0200 Subject: [PATCH 027/189] netfilter: nft_synproxy: use the family-aware checksum helper nft_synproxy_do_eval() verifies the TCP checksum before it switches on skb->protocol. It uses nf_ip_checksum(), which constructs an IPv4 pseudo header and relies on the IPv4 header checksum when folding the whole skb. Neither operation is valid for an IPv6 packet. A correctly checksummed IPv6 segment can therefore fail verification when it reaches the hook as CHECKSUM_NONE or, at NF_INET_LOCAL_IN, CHECKSUM_COMPLETE. nft_synproxy_do_eval() returns NF_DROP before nft_synproxy_eval_v6() can send a SYN-ACK. nft_synproxy_validate() deliberately admits NFPROTO_IPV6 and NFPROTO_INET, and the xtables counterpart ip6t_SYNPROXY.c already calls nf_ip6_checksum(). Use nf_checksum() with nft_pf() so the checksum helper dispatches to the packet family's implementation. Fixes: ad49d86e07a4 ("netfilter: nf_tables: Add synproxy support") Assisted-by: LLM Signed-off-by: Karl Mehltretter Signed-off-by: Pablo Neira Ayuso --- net/netfilter/nft_synproxy.c | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/net/netfilter/nft_synproxy.c b/net/netfilter/nft_synproxy.c index 9ed288c9d168..554a96a000f4 100644 --- a/net/netfilter/nft_synproxy.c +++ b/net/netfilter/nft_synproxy.c @@ -118,7 +118,8 @@ static void nft_synproxy_do_eval(const struct nft_synproxy *priv, return; } - if (nf_ip_checksum(skb, nft_hook(pkt), thoff, IPPROTO_TCP)) { + if (nf_checksum(skb, nft_hook(pkt), thoff, IPPROTO_TCP, + nft_pf(pkt))) { regs->verdict.code = NF_DROP; return; } From e290145564886d6a3038810c621f738c1fe9fa51 Mon Sep 17 00:00:00 2001 From: Julian Anastasov Date: Fri, 11 Sep 2026 14:43:15 +0300 Subject: [PATCH 028/189] ipvs: revalidate ihl before icmp_send While the outer IP header is already pulled into the skb head, we must be careful and revalidate the embedded headers after reading them from the skb frags to prevent possible out-of-bounds access. One such place reported by Sashiko is ip_vs_in_icmp() where local process can change the ihl field and after pskb_may_pull() we can see larger value. Even if icmp_send() has checks to prevent out-of-bounds access, play safe and add check to drop the packet if the ihl field is changed. As the outer headers are pulled, make sure the transport header is updated too, it was used before commit 7fcc2fe39fed ("net: icmp: avoid invalid transport header access in icmp_send tracepoint") Fixes: f2edb9f7706d ("ipvs: implement passive PMTUD for IPIP packets") Link: https://sashiko.dev/#/patchset/20260806105211.34622-1-ja%40ssi.bg Signed-off-by: Julian Anastasov Signed-off-by: Pablo Neira Ayuso --- net/netfilter/ipvs/ip_vs_core.c | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/net/netfilter/ipvs/ip_vs_core.c b/net/netfilter/ipvs/ip_vs_core.c index ba0957798bad..fd503f0efb57 100644 --- a/net/netfilter/ipvs/ip_vs_core.c +++ b/net/netfilter/ipvs/ip_vs_core.c @@ -1960,6 +1960,12 @@ ip_vs_in_icmp(struct netns_ipvs *ipvs, struct sk_buff *skb, int *related, /* Ensure the IP header is present in headroom */ if (!pskb_may_pull(skb, hlen_orig)) goto ignore_tunnel; + skb_set_transport_header(skb, hlen_orig); + /* Before now we may used ihl from skb frag, revalidate it after + * copying it into skb head to prevent out-of-bounds access + */ + if (ip_hdr(skb)->ihl * 4 != hlen_orig) + goto ignore_tunnel; IP_VS_DBG(12, "Sending ICMP for %pI4->%pI4: t=%u, c=%u, i=%u\n", &ip_hdr(skb)->saddr, &ip_hdr(skb)->daddr, type, code, ntohl(info)); From 207d591c353201f3bd3e0c89bb7d44a849c8fd59 Mon Sep 17 00:00:00 2001 From: Naman Gulati Date: Sat, 12 Sep 2026 01:10:51 +0000 Subject: [PATCH 029/189] netfilter: ctnetlink: fix suspicious RCU usage in expect_iter_name expect_iter_name() is invoked by nf_ct_expect_iterate_net() under spin_lock_bh(&nf_conntrack_expect_lock). It does not hold rcu_read_lock(). When accessing exp->helper with rcu_dereference() in syzbot's report, lockdep warns: ============================= WARNING: suspicious RCU usage syzkaller #0 Not tainted ----------------------------- net/netfilter/nf_conntrack_netlink.c:3393 suspicious rcu_dereference_check() usage! locks held by syz-executor381/5628: 2, last CPU#1: #0: ffffffff9aee42a0 (nfnl_subsys_ctnetlink_exp){+.+.}-{4:4}, at: nfnetlink_rcv_msg+0xa69/0x12b0 #1: ffffffff8ea74d58 (nf_conntrack_expect_lock){+...}-{3:3}, at: nf_ct_expect_iterate_net+0x38/0x180 Call Trace: dump_stack_lvl+0xe8/0x150 lockdep_rcu_suspicious+0x140/0x1d0 expect_iter_name+0xfb/0x100 nf_ct_expect_iterate_net+0xf2/0x180 ctnetlink_del_expect+0x45d/0x640 nfnetlink_rcv_msg+0xcc2/0x12b0 netlink_rcv_skb+0x226/0x4a0 nfnetlink_rcv+0x2b9/0x28c0 netlink_unicast+0x7bd/0x940 netlink_sendmsg+0x813/0xb40 ____sys_sendmsg+0x54e/0x850 ___sys_sendmsg+0x2a5/0x360 __sys_sendmsg+0x2a5/0x360 do_syscall_64+0x166/0x520 entry_SYSCALL_64_after_hwframe+0x77/0x7f Use rcu_dereference_protected() with lockdep_is_held() on nf_conntrack_expect_lock instead, similar to expect_iter_me() in nf_conntrack_helper.c. Fixes: f01794106042 ("netfilter: nf_conntrack_expect: use expect->helper") Reported-by: syzbot+4bd730aede2791e40bdf@syzkaller.appspotmail.com Closes: https://lore.kernel.org/netdev/6aa4a377.f81106d8.2ab401.0024.GAE@google.com/T/#u Signed-off-by: Naman Gulati Signed-off-by: Pablo Neira Ayuso --- net/netfilter/nf_conntrack_netlink.c | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/net/netfilter/nf_conntrack_netlink.c b/net/netfilter/nf_conntrack_netlink.c index 579ada063b1b..4e5d7c701436 100644 --- a/net/netfilter/nf_conntrack_netlink.c +++ b/net/netfilter/nf_conntrack_netlink.c @@ -3392,7 +3392,8 @@ static bool expect_iter_name(struct nf_conntrack_expect *exp, void *data) struct nf_conntrack_helper *helper; const char *name = data; - helper = rcu_dereference(exp->helper); + helper = rcu_dereference_protected(exp->helper, + lockdep_is_held(&nf_conntrack_expect_lock)); if (!helper) return false; From 70194dc37670bd08e44b471389861cc01bd3a3c9 Mon Sep 17 00:00:00 2001 From: Aohan Mei Date: Mon, 14 Sep 2026 19:51:47 +0800 Subject: [PATCH 030/189] netfilter: nf_tables: skip expired catchall elements on insert and delete nft_setelem_catchall_insert() looks up duplicates with nft_set_elem_active() only, while nft_set_catchall_lookup() and the dump path additionally skip expired elements. Once a catchall element with a timeout expires, this predicate drift makes it invisible to userspace dumps, yet it still blocks re-insertion: with NLM_F_EXCL the request fails with -EEXIST, and without it the request reports success but silently inserts nothing. The stale entry only goes away when the (user-tunable) gc interval elapses, so the catchall rule may silently stop matching for an arbitrarily long time after its first expiration. The delete path shows the same drift: nft_setelem_catchall_deactivate() picks the first active-next entry in the catchall list, so with an expired entry still pending GC it retires the stale entry instead of the fresh one, and it deactivates an element that userspace no longer sees instead of failing with -ENOENT. Align both walks with the lookup and dump predicates: only an element that is active and not expired counts as a duplicate or delete candidate, using the per-netns timestamp taken at transaction start, in line with the set backend .insert/.deactivate and catchall GC sync paths. Reported-by: TencentOS Corvus AI Cc: stable@vger.kernel.org Fixes: aaa31047a6d2 ("netfilter: nftables: add catch-all set element support") Assisted-by: CodeBuddy:Kimi-K3 Signed-off-by: Aohan Mei Signed-off-by: Pablo Neira Ayuso --- net/netfilter/nf_tables_api.c | 10 ++++++++-- 1 file changed, 8 insertions(+), 2 deletions(-) diff --git a/net/netfilter/nf_tables_api.c b/net/netfilter/nf_tables_api.c index c0b754a2d45b..b59628e6240c 100644 --- a/net/netfilter/nf_tables_api.c +++ b/net/netfilter/nf_tables_api.c @@ -6995,11 +6995,14 @@ static int nft_setelem_catchall_insert(const struct net *net, { struct nft_set_elem_catchall *catchall; u8 genmask = nft_genmask_next(net); + u64 tstamp = nft_net_tstamp(net); struct nft_set_ext *ext; list_for_each_entry(catchall, &set->catchall_list, list) { ext = nft_set_elem_ext(set, catchall->elem); - if (nft_set_elem_active(ext, genmask)) { + if (nft_set_elem_active(ext, genmask) && + !__nft_set_elem_expired(ext, tstamp) && + !nft_set_elem_is_dead(ext)) { *priv = catchall->elem; return -EEXIST; } @@ -7092,11 +7095,14 @@ static int nft_setelem_catchall_deactivate(const struct net *net, struct nft_set_elem *elem) { struct nft_set_elem_catchall *catchall; + u64 tstamp = nft_net_tstamp(net); struct nft_set_ext *ext; list_for_each_entry(catchall, &set->catchall_list, list) { ext = nft_set_elem_ext(set, catchall->elem); - if (!nft_is_active_next(net, ext)) + if (!nft_is_active_next(net, ext) || + __nft_set_elem_expired(ext, tstamp) || + nft_set_elem_is_dead(ext)) continue; kfree(elem->priv); From 99cc2a62e07a44a22254d7beca9ef1f8ad886d0d Mon Sep 17 00:00:00 2001 From: Eric Dumazet Date: Sun, 13 Sep 2026 04:42:33 +0000 Subject: [PATCH 031/189] tipc: reject invalid and unexpected GRP_ACK_MSG to prevent bc_ackers underflow Commit 48a5fe38772b ("tipc: fix bc_ackers underflow on duplicate GRP_ACK_MSG") rejected duplicate/stale ACKs in tipc_group_proto_rcv() by returning early when less_eq(acked, m->bc_acked). However, that check remains incomplete in two ways: 1. When grp->bc_ackers is zero (e.g. on a quiet group, when replicast ACKs were not requested, or after all expected members have already acknowledged), an unexpected GRP_ACK_MSG with acked > m->bc_acked passes less_eq() and unconditionally decrements grp->bc_ackers. Because bc_ackers is a u16, this wraps to 65535, causing tipc_group_bc_cong() to permanently report congestion and blocking all future group broadcasts on the socket. 2. During an active broadcast round (grp->bc_ackers > 0), the sender transmits packet S and advances grp->bc_snd_nxt to S + 1. Receivers increment their expected counter to S + 1 upon consuming packet S, so the only valid ACK value for the current round is strictly acked == grp->bc_snd_nxt. However, tipc_group_update_bc_members() initializes each member's m->bc_acked to prev = grp->bc_snd_nxt - 1 (S - 1 before increment). This leaves a 2-sequence gap (S - 1 to S + 1) in sequence space. An incoming ACK is therefore neither rejected as duplicate nor prevented from decrementing grp->bc_ackers if an unexpected or stale value (such as S) is received. A member sending acked = S followed by acked = S + 1 could decrement grp->bc_ackers twice in the same round, prematurely clearing bc_ackers or underflowing it. Fix this by: - Dropping GRP_ACK_MSG immediately if grp->bc_ackers is zero. - Requiring acked == grp->bc_snd_nxt and rejecting duplicates where m->bc_acked == acked. Because replicast broadcast rounds are strictly sequential, only grp->bc_snd_nxt can be acknowledged, and each member can acknowledge at most once per round. Note that a related pre-existing issue in tipc_group_delete_member() (where grp->bc_ackers decrementing to zero upon member departure does not restore *grp->open or trigger a socket wakeup) will be addressed in a separate patch. Fixes: 48a5fe38772b ("tipc: fix bc_ackers underflow on duplicate GRP_ACK_MSG") Fixes: 2f487712b893 ("tipc: guarantee that group broadcast doesn't bypass group unicast") Reported-by: James Burton Cc: stable@vger.kernel.org Signed-off-by: Eric Dumazet Link: https://patch.msgid.link/20260913044233.193927-1-edumazet@google.com Signed-off-by: Jakub Kicinski --- net/tipc/group.c | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/net/tipc/group.c b/net/tipc/group.c index 14e6732624e2..74f6d3dac078 100644 --- a/net/tipc/group.c +++ b/net/tipc/group.c @@ -797,10 +797,10 @@ void tipc_group_proto_rcv(struct tipc_group *grp, bool *usr_wakeup, tipc_group_open(m, usr_wakeup); return; case GRP_ACK_MSG: - if (!m) + if (!m || !grp->bc_ackers) return; acked = msg_grp_bc_acked(hdr); - if (less_eq(acked, m->bc_acked)) + if (acked != grp->bc_snd_nxt || m->bc_acked == acked) return; m->bc_acked = acked; if (--grp->bc_ackers) From 47abe7a5c4eb53269aca3506446f851572a059a3 Mon Sep 17 00:00:00 2001 From: Nguyen Ngoc Thang Date: Tue, 15 Sep 2026 22:08:16 +0700 Subject: [PATCH 032/189] net/sched: act_ct: don't WARN on benign flow_offload_alloc() failure flow_offload_alloc() returns NULL when the conntrack entry is dying (e.g. raced with a conntrack flush) or when the GFP_ATOMIC allocation fails; both are expected under load and neither is a kernel bug. This path runs from softirq on every committed packet, so with panic_on_warn=1 an unprivileged user can panic the box just by racing a conntrack flush against a `tc ... action ct commit` classifier. Reproduced with a custom repro under QEMU: a small, fixed set of UDP flows through `tc filter ... action ct commit` on lo, raced against threads flooding bare ctnetlink CT_DELETE (flush) requests. Hits WARNING: net/sched/act_ct.c:437 (tcf_ct_flow_table_add(), inlined into tcf_ct_act() in this build) within ~15s on the unpatched kernel; same setup is clean on the patched kernel. The fix itself is behavior-preserving: both branches already did `goto err_alloc` before and after, only the WARN is removed. Fixes: 64ff70b80fd4 ("net/sched: act_ct: Offload established connections to flow table") Reported-by: syzbot+6cc37aba98dac721c415@syzkaller.appspotmail.com Closes: https://syzkaller.appspot.com/bug?extid=6cc37aba98dac721c415 Signed-off-by: Nguyen Ngoc Thang Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260915150816.36487-1-ngocthang2710.1999@gmail.com Signed-off-by: Jakub Kicinski --- net/sched/act_ct.c | 5 ++--- 1 file changed, 2 insertions(+), 3 deletions(-) diff --git a/net/sched/act_ct.c b/net/sched/act_ct.c index 9080cb386c16..55f3521edb4c 100644 --- a/net/sched/act_ct.c +++ b/net/sched/act_ct.c @@ -432,11 +432,10 @@ static void tcf_ct_flow_table_add(struct tcf_ct_flow_table *ct_ft, if (test_and_set_bit(IPS_OFFLOAD_BIT, &ct->status)) return; + /* NULL if ct is dying (raced flush) or the atomic alloc failed. */ entry = flow_offload_alloc(ct); - if (!entry) { - WARN_ON_ONCE(1); + if (!entry) goto err_alloc; - } if (tcp) { ct->proto.tcp.seen[0].flags |= IP_CT_TCP_FLAG_BE_LIBERAL; From 2566866fc30965d915d0b52b5c3323b362619f0e Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?J=C3=A9r=C3=A9my=20Jean?= Date: Tue, 15 Sep 2026 12:48:07 +0000 Subject: [PATCH 033/189] net: gue: reject invalid REMCSUM offsets MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The REMCSUM option carries an absolute checksum start and checksum field offset. gue_remcsum() passes them to skb_remcsum_process(), whose partial path stores offset - start in the u16 skb->csum_offset variable. If offset is less than start, this underflows. A forwarded packet can retain CHECKSUM_PARTIAL and reach a NETIF_F_HW_CSUM driver which trusts the metadata, leading skb_copy_and_csum_dev() to write two bytes about 64 KiB beyond the destination buffer. Reject reversed tuples in validate_gue_flags(), after the existing length validation, so all GUE parsers enforce the ordering in one place. Fixes: fe881ef11cf0 ("gue: Use checksum partial with remote checksum offload") Signed-off-by: Jérémy Jean Link: https://patch.msgid.link/20260915124806.2852293-2-Jeremy.Jean@oss.cyber.gouv.fr Signed-off-by: Jakub Kicinski --- include/net/gue.h | 19 +++++++++++++++---- 1 file changed, 15 insertions(+), 4 deletions(-) diff --git a/include/net/gue.h b/include/net/gue.h index caefd6da8693..d377155fd0b3 100644 --- a/include/net/gue.h +++ b/include/net/gue.h @@ -84,8 +84,9 @@ static inline size_t guehdr_priv_flags_len(__be32 flags) } /* Validate standard and private flags. Returns non-zero (meaning invalid) - * if there is an unknown standard or private flags, or the options length for - * the flags exceeds the options length specific in hlen of the GUE header. + * if there is an unknown standard or private flags, if the options length for + * the flags exceeds the options length specified in hlen of the GUE header, or + * if a private option contains invalid data. */ static inline int validate_gue_flags(struct guehdr *guehdr, size_t optlen) { @@ -103,8 +104,8 @@ static inline int validate_gue_flags(struct guehdr *guehdr, size_t optlen) /* Private flags are last four bytes accounted in * guehdr_flags_len */ - __be32 pflags = *(__be32 *)((void *)&guehdr[1] + - len - GUE_LEN_PRIV); + void *data = (void *)&guehdr[1] + len; + __be32 pflags = *(__be32 *)(data - GUE_LEN_PRIV); if (pflags & ~GUE_PFLAGS_ALL) return 1; @@ -112,6 +113,16 @@ static inline int validate_gue_flags(struct guehdr *guehdr, size_t optlen) len += guehdr_priv_flags_len(pflags); if (len > optlen) return 1; + + if (pflags & GUE_PFLAG_REMCSUM) { + __be16 *pd = data; + + /* The field offset pd[1] must not be less + * than the start pd[0]. + */ + if (ntohs(pd[1]) < ntohs(pd[0])) + return 1; + } } return 0; From 310d1ac61a4d5a2ca8356a3a48d263acf54503ce Mon Sep 17 00:00:00 2001 From: Lorenzo Bianconi Date: Wed, 16 Sep 2026 15:30:13 +0200 Subject: [PATCH 034/189] net: ethernet: mtk_eth_soc: unregister net_devices in case of probe failure If register_netdev() fails for one of the MTK_MAX_DEVS devices in mtk_probe(), the error path jumps to err_deinit_ppe, skipping mtk_unreg_dev(). The previously registered net_devices are then freed by mtk_free_dev() while still in NETREG_REGISTERED state, hitting the BUG_ON(dev->reg_state != NETREG_UNREGISTERED). Route the register_netdev() failure to err_unreg_netdev so the net_devices registered so far are properly unregistered before being freed. Fixes: 8a8a9e89f801 ("net: ethernet: mediatek: cleanup error path inside mtk_hw_init") Signed-off-by: Lorenzo Bianconi Link: https://patch.msgid.link/20260916-mtk_eth_soc-netdev-fix-v1-1-5dac50eb65b1@oss.qualcomm.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/mediatek/mtk_eth_soc.c | 6 +++++- 1 file changed, 5 insertions(+), 1 deletion(-) diff --git a/drivers/net/ethernet/mediatek/mtk_eth_soc.c b/drivers/net/ethernet/mediatek/mtk_eth_soc.c index fd7a49ae88d0..2ea5dfe85539 100644 --- a/drivers/net/ethernet/mediatek/mtk_eth_soc.c +++ b/drivers/net/ethernet/mediatek/mtk_eth_soc.c @@ -4509,6 +4509,10 @@ static int mtk_unreg_dev(struct mtk_eth *eth) mac = netdev_priv(eth->netdev[i]); if (MTK_HAS_CAPS(eth->soc->caps, MTK_QDMA)) unregister_netdevice_notifier(&mac->device_notifier); + + if (eth->netdev[i]->reg_state != NETREG_REGISTERED) + continue; + unregister_netdev(eth->netdev[i]); } @@ -5344,7 +5348,7 @@ static int mtk_probe(struct platform_device *pdev) err = register_netdev(eth->netdev[i]); if (err) { dev_err(eth->dev, "error bringing up device\n"); - goto err_deinit_ppe; + goto err_unreg_netdev; } else netif_info(eth, probe, eth->netdev[i], "mediatek frame engine at 0x%08lx, irq %d\n", From 24fedc7a569bce181728f0dd616504bef9f1513b Mon Sep 17 00:00:00 2001 From: Andy Moreton Date: Wed, 16 Sep 2026 13:56:41 +0100 Subject: [PATCH 035/189] sfc: add X4D PF support X4D is an X4 controller instance as an IP block in an SoC. It has the same feature set as X4. Signed-off-by: Andy Moreton Reviewed-by: Pieter Jansen van Vuuren Reviewed-by: Alejandro Lucero Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260916125641.12238-1-alucerop@amd.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/sfc/efx.c | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/drivers/net/ethernet/sfc/efx.c b/drivers/net/ethernet/sfc/efx.c index 3806cd3dd7f4..953f1d90b9b1 100644 --- a/drivers/net/ethernet/sfc/efx.c +++ b/drivers/net/ethernet/sfc/efx.c @@ -904,6 +904,10 @@ static const struct pci_device_id efx_pci_table[] = { .driver_data = (unsigned long)&efx_x4_nic_type}, {PCI_DEVICE(PCI_VENDOR_ID_SOLARFLARE, 0x2c03), /* X4 PF (FF only) */ .driver_data = (unsigned long)&efx_x4_nic_type}, + {PCI_DEVICE(PCI_VENDOR_ID_SOLARFLARE, 0x8c03), /* X4D PF (FF/LL) */ + .driver_data = (unsigned long)&efx_x4_nic_type}, + {PCI_DEVICE(PCI_VENDOR_ID_SOLARFLARE, 0xac03), /* X4D PF (FF only) */ + .driver_data = (unsigned long)&efx_x4_nic_type}, {0} /* end of list */ }; From ee319bd3a0e976af5087cbe59ebc50a66f31d202 Mon Sep 17 00:00:00 2001 From: Norbert Szetei Date: Wed, 16 Sep 2026 21:57:53 +0200 Subject: [PATCH 036/189] ipv6: do not let ipv6_find_hdr() return an offset past the packet end ipv6_find_hdr() walks the extension header chain, skipping each header by the length that header itself declares. ipv6_optlen() returns up to 2048, and the skip is never checked against skb->len, so the offset stored in *offset can point past the end of the packet. openvswitch installs that offset as the transport header, and update_ipv6_checksum() then reads and writes the transport checksum field out of bounds: BUG: KASAN: slab-use-after-free in inet_proto_csum_replace16+0x445/0x470 Read of size 2 at addr ffff88810b754b06 by task ovs_ipv6_oob/629 CPU: 4 UID: 1000 PID: 629 Comm: ovs_ipv6_oob Tainted: G N 7.3.0-rc3+ #348 Call Trace: inet_proto_csum_replace16+0x445/0x470 set_ipv6_addr+0x3dd/0x460 do_execute_actions+0x6a3d/0x7c40 ovs_execute_actions+0xfd/0x480 ovs_packet_cmd_execute+0xc38/0xf20 genl_rcv_msg+0x59e/0x870 netlink_rcv_skb+0x18b/0x450 genl_rcv+0x2d/0x40 netlink_unicast+0x6bc/0xa20 The buggy address belongs to the object at ffff88810b754980 which belongs to the cache skbuff_small_head of size 704 The buggy address is located 390 bytes inside of freed 704-byte region [ffff88810b754980, ffff88810b754c40) Other callers use that offset too, so bound it here rather than in one caller. Reject a header whose declared length does not fit in the packet. ipv6_find_hdr() already fails with -EBADMSG on a malformed chain, so this adds no new failure mode. Fixes: f8f626754ebe ("ipv6: Move ipv6_find_hdr() out of Netfilter code.") Suggested-by: Ilya Maximets Suggested-by: Eric Dumazet Cc: stable@vger.kernel.org Signed-off-by: Norbert Szetei Reviewed-by: Ido Schimmel Reviewed-by: Ilya Maximets Link: https://patch.msgid.link/8F80BA1A-DDFD-432D-9075-242A3435FEB5@doyensec.com Signed-off-by: Jakub Kicinski --- net/ipv6/exthdrs_core.c | 3 +++ 1 file changed, 3 insertions(+) diff --git a/net/ipv6/exthdrs_core.c b/net/ipv6/exthdrs_core.c index 9d06d487e8b1..4a9748338cf4 100644 --- a/net/ipv6/exthdrs_core.c +++ b/net/ipv6/exthdrs_core.c @@ -278,6 +278,9 @@ int ipv6_find_hdr(const struct sk_buff *skb, unsigned int *offset, hdrlen = ipv6_optlen(hp); if (!found) { + if (skb->len - start < hdrlen) + return -EBADMSG; + nexthdr = hp->nexthdr; start += hdrlen; } From dd47bcf279f1083f09bf5266890b26263361022b Mon Sep 17 00:00:00 2001 From: Kuniyuki Iwashima Date: Wed, 16 Sep 2026 23:09:24 +0000 Subject: [PATCH 037/189] ip6_gre: Call ip6erspan_tunnel_unlink_md() in ip6erspan_changelink(). The cited commit accidentally added ip6gre_tunnel_unlink_md() in ip6erspan_changelink(). Let's correct it to ip6erspan_tunnel_unlink_md(). Fixes: b80d0b93b991 ("net: ip6_gre: fix tunnel metadata device sharing.") Signed-off-by: Kuniyuki Iwashima Reviewed-by: Xuanqiang Luo Reviewed-by: Ido Schimmel Link: https://patch.msgid.link/20260916230927.378957-1-kuniyu@google.com Signed-off-by: Jakub Kicinski --- net/ipv6/ip6_gre.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/net/ipv6/ip6_gre.c b/net/ipv6/ip6_gre.c index 8ebda0b6a78b..e61cb10b50dc 100644 --- a/net/ipv6/ip6_gre.c +++ b/net/ipv6/ip6_gre.c @@ -2279,7 +2279,7 @@ static int ip6erspan_changelink(struct net_device *dev, struct nlattr *tb[], return PTR_ERR(t); ip6erspan_set_version(data, &p); - ip6gre_tunnel_unlink_md(ign, t); + ip6erspan_tunnel_unlink_md(ign, t); ip6gre_tunnel_unlink(ign, t); ip6erspan_tnl_change(t, &p, !tb[IFLA_MTU]); ip6erspan_tunnel_link_md(ign, t); From 95c4d54ed02283e9a09e8cd7360e384daa67a741 Mon Sep 17 00:00:00 2001 From: Abhishek Ojha Date: Wed, 16 Sep 2026 19:19:28 -0400 Subject: [PATCH 038/189] net: phy: micrel: Advance register data pointer in write loop lanphy_write_reg_data() does not advance the data pointer while iterating over the register table. As a result, it writes the first entry num times and leaves the remaining errata registers unconfigured. Single-entry tables are unaffected, but tables with multiple entries leave every entry after the first unapplied. Advance the data pointer after each successful write so every table entry is applied in order. Fixes: c8732e933925 ("net: phy: micrel: lan8842 errata") Cc: stable@vger.kernel.org Signed-off-by: Abhishek Ojha Reviewed-by: Andrew Lunn Link: https://patch.msgid.link/20260916231928.1336305-1-abhishek.ojha@savoirfairelinux.com Signed-off-by: Jakub Kicinski --- drivers/net/phy/micrel.c | 1 + 1 file changed, 1 insertion(+) diff --git a/drivers/net/phy/micrel.c b/drivers/net/phy/micrel.c index ae830781824b..5c8461db7b4b 100644 --- a/drivers/net/phy/micrel.c +++ b/drivers/net/phy/micrel.c @@ -6344,6 +6344,7 @@ static int lanphy_write_reg_data(struct phy_device *phydev, data->val); if (ret) break; + data++; } return ret; From 1b82958f3f035df5ccaab5430a2302f08a5d5351 Mon Sep 17 00:00:00 2001 From: Alexander Duyck Date: Mon, 14 Sep 2026 14:09:57 -0700 Subject: [PATCH 039/189] net: ethtool: keep rtnl_lock for the ioctl self test An offline self test that brings the interface down and back up with netif_close() / netif_open() requires rtnl_lock for both. Since the ethtool IOCTL path became rtnl-optional for ops-locked drivers, the ETHTOOL_TEST ioctl runs holding only the netdev instance lock, so on an ops-locked driver the self test now tears the device down without rtnl_lock. With lockdep this reproduces deterministically on every offline self test on such a driver; note the sole lock held is the instance lock, not rtnl: WARNING: suspicious RCU usage net/core/netpoll.c:207 suspicious rcu_dereference_protected() usage! 1 lock held by ethtool/107: #0: (&dev->lock){+.+.}, at: dev_ethtool Call Trace: netpoll_poll_disable __dev_close_many netif_close_many netif_close fbnic_self_test dev_ethtool_locked dev_ethtool dev_ioctl sock_ioctl __x64_sys_ioctl Without lockdep the same condition trips ASSERT_RTNL() in __dev_close_many() / __dev_open(); that check only samples the global rtnl state, so it can be masked by a concurrent rtnl holder, but the device is still being reconfigured without the lock it requires. The ethtool self_test is a legacy ioctl-only command, so an ETHTOOL_TEST case is only needed on the ioctl path. Add an opt-in bit for drivers whose self test needs rtnl_lock and set it on the ops-locked drivers whose offline self test tears the interface down and up: - fbnic (ops-locked via queue_mgmt_ops): fbnic_self_test() offline path uses netif_close() / netif_open(). - bnxt (ops-locked via queue_mgmt_ops): bnxt_self_test() offline path goes through bnxt_close_nic() / bnxt_half_open_nic() / bnxt_half_close_nic() / bnxt_open_nic(), which close and reopen the device. Fixes: f994752b1127 ("net: ethtool: optionally skip rtnl_lock on IOCTL path") Signed-off-by: Alexander Duyck Reviewed-by: Simon Horman Link: https://patch.msgid.link/178942019771.7700.338431553546884773.stgit@ahduyck-xeon-server.home.arpa Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/broadcom/bnxt/bnxt_ethtool.c | 3 ++- drivers/net/ethernet/meta/fbnic/fbnic_ethtool.c | 3 ++- include/linux/ethtool.h | 2 ++ net/ethtool/common.h | 2 ++ 4 files changed, 8 insertions(+), 2 deletions(-) diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt_ethtool.c b/drivers/net/ethernet/broadcom/bnxt/bnxt_ethtool.c index 62bc9cae613c..622e89587e5d 100644 --- a/drivers/net/ethernet/broadcom/bnxt/bnxt_ethtool.c +++ b/drivers/net/ethernet/broadcom/bnxt/bnxt_ethtool.c @@ -5733,7 +5733,8 @@ const struct ethtool_ops bnxt_ethtool_ops = { .op_needs_rtnl = ETHTOOL_OP_NEEDS_RTNL_SCHANNELS | ETHTOOL_OP_NEEDS_RTNL_SRINGPARAM | ETHTOOL_OP_NEEDS_RTNL_SCOALESCE | - ETHTOOL_OP_NEEDS_RTNL_RSS, + ETHTOOL_OP_NEEDS_RTNL_RSS | + ETHTOOL_OP_NEEDS_RTNL_TEST, .supported_coalesce_params = ETHTOOL_COALESCE_USECS | ETHTOOL_COALESCE_MAX_FRAMES | ETHTOOL_COALESCE_USECS_IRQ | diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_ethtool.c b/drivers/net/ethernet/meta/fbnic/fbnic_ethtool.c index 0e47088ec44b..423f179c9d47 100644 --- a/drivers/net/ethernet/meta/fbnic/fbnic_ethtool.c +++ b/drivers/net/ethernet/meta/fbnic/fbnic_ethtool.c @@ -2025,7 +2025,8 @@ static const struct ethtool_ops fbnic_ethtool_ops = { ETHTOOL_OP_NEEDS_RTNL_SPAUSEPARAM | ETHTOOL_OP_NEEDS_RTNL_SCHANNELS | ETHTOOL_OP_NEEDS_RTNL_SRINGPARAM | - ETHTOOL_OP_NEEDS_RTNL_GLINK, + ETHTOOL_OP_NEEDS_RTNL_GLINK | + ETHTOOL_OP_NEEDS_RTNL_TEST, .get_drvinfo = fbnic_get_drvinfo, .get_regs_len = fbnic_get_regs_len, .get_regs = fbnic_get_regs, diff --git a/include/linux/ethtool.h b/include/linux/ethtool.h index 253600c0eccd..c4c9ce038611 100644 --- a/include/linux/ethtool.h +++ b/include/linux/ethtool.h @@ -944,6 +944,7 @@ struct kernel_ethtool_ts_info { #define ETHTOOL_OP_NEEDS_RTNL_SPAUSEPARAM BIT(6) #define ETHTOOL_OP_NEEDS_RTNL_RSS BIT(7) #define ETHTOOL_OP_NEEDS_RTNL_GLINK BIT(8) +#define ETHTOOL_OP_NEEDS_RTNL_TEST BIT(9) /** * struct ethtool_ops - optional netdev operations @@ -981,6 +982,7 @@ struct kernel_ethtool_ts_info { * - netdev_update_features() * - netif_set_real_num_tx_queues() * - ethtool_op_get_link() (syncs link watch under rtnl_lock) + * - netif_open() / netif_close() (used by @self_test) * * @get_drvinfo: Report driver/device information. Modern drivers no * longer have to implement this callback. Most fields are diff --git a/net/ethtool/common.h b/net/ethtool/common.h index 4e5356e26f40..ae32e7fdb563 100644 --- a/net/ethtool/common.h +++ b/net/ethtool/common.h @@ -163,6 +163,8 @@ ethtool_ioctl_needs_rtnl(const struct net_device *dev, u32 ethcmd) return ops->op_needs_rtnl & ETHTOOL_OP_NEEDS_RTNL_RSS; case ETHTOOL_GLINK: return ops->op_needs_rtnl & ETHTOOL_OP_NEEDS_RTNL_GLINK; + case ETHTOOL_TEST: + return ops->op_needs_rtnl & ETHTOOL_OP_NEEDS_RTNL_TEST; } return false; } From 1f4c73064a50f53d596c6f1d06d2d700f43c4b32 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Bj=C3=B6rn=20T=C3=B6pel?= Date: Mon, 14 Sep 2026 14:10:04 -0700 Subject: [PATCH 040/189] eth: fbnic: Handle maximum standalone channels MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Standalone channels use one NAPI vector for each Tx and Rx queue. fbnic's allocation path excludes FBNIC_MAX_TXQS from that layout. A 64-Tx/64-Rx configuration therefore records 128 vectors but allocates only 64, leaving NULL entries that resource setup dereferences. Include the maximum vector count in standalone allocation. Fixes: bc6107771bb4 ("eth: fbnic: Allocate a netdevice and napi vectors with queues") Signed-off-by: Björn Töpel Reviewed-by: Simon Horman Link: https://patch.msgid.link/178942020457.7700.13129750616387075931.stgit@ahduyck-xeon-server.home.arpa Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/meta/fbnic/fbnic_txrx.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c b/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c index 661dee1661af..a30aa4450848 100644 --- a/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c +++ b/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c @@ -1791,7 +1791,7 @@ int fbnic_alloc_napi_vectors(struct fbnic_net *fbn) int err; /* Allocate 1 Tx queue per napi vector */ - if (num_napi < FBNIC_MAX_TXQS && num_napi == num_tx + num_rx) { + if (num_napi <= FBNIC_MAX_TXQS && num_napi == num_tx + num_rx) { while (num_tx) { err = fbnic_alloc_napi_vector(fbd, fbn, num_napi, v_idx, From b5d9e9d4d0c13bc8b60d8d97e7a07fb25fea639e Mon Sep 17 00:00:00 2001 From: Alexander Duyck Date: Mon, 14 Sep 2026 14:10:11 -0700 Subject: [PATCH 041/189] eth: fbnic: use the Rx queue napi pointer to find the napi vector The queue management ndos pick the napi vector for an Rx queue with: nv = fbn->napi[idx % fbn->num_napi]; The issue is this is only correct in the cases where there are no standalone Tx vectors. In those cases we were allocating the Tx vectors first and then the Rx so the queues would be pointing to Tx NAPI vectors instead of the Rx ones. The mapping the ndos want is already recorded. fbnic_set_netif_napi() publishes it with netif_queue_set_napi(), which stores the napi pointer in netdev_rx_queue.napi, and fbnic_reset_netif_napi() clears it again. Both run under the netdev instance lock that the queue management ndos also hold, so the pointer can be read directly. Use it and drop the divide. The pointer is NULL exactly while the datapath is down, so fbnic_queue_mem_alloc() can reject that case rather than reaching into freed state: netdev_rx_queue_restart() calls it before it tests netif_running(), and fbnic_pm_suspend() leaves netif_running() true across a PCIe recovery that never completes, so a queue restart can arrive after fbnic_stop() has freed the rings and the vectors. fbnic_stop() clears the association in fbnic_reset_netif_queues() before fbnic_free_napi_vectors(), so the NULL is always published first. fbnic_queue_start() and fbnic_queue_stop() need no check of their own, as netdev_rx_queue_reconfig() only reaches them once fbnic_queue_mem_alloc() has succeeded under the same instance lock. Fixes: da43127a8edc ("eth: fbnic: support queue ops / zero-copy Rx") Signed-off-by: Alexander Duyck Reviewed-by: Simon Horman Link: https://patch.msgid.link/178942021136.7700.4391219358260544104.stgit@ahduyck-xeon-server.home.arpa Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/meta/fbnic/fbnic_txrx.c | 26 +++++++++++++++++--- 1 file changed, 23 insertions(+), 3 deletions(-) diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c b/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c index a30aa4450848..10caacffee0f 100644 --- a/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c +++ b/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c @@ -7,6 +7,7 @@ #include #include #include +#include #include #include #include @@ -2853,6 +2854,17 @@ void fbnic_napi_depletion_check(struct net_device *netdev) fbnic_wrfl(fbd); } +/* Returns the napi vector servicing an Rx queue, or NULL if the datapath + * is torn down. The association is published by fbnic_set_netif_napi() + * and cleared by fbnic_reset_netif_napi(), both under the instance lock. + */ +static struct fbnic_napi_vector *fbnic_rxq_nv(struct net_device *dev, int idx) +{ + struct napi_struct *napi = __netif_get_rx_queue(dev, idx)->napi; + + return napi ? container_of(napi, struct fbnic_napi_vector, napi) : NULL; +} + static int fbnic_queue_mem_alloc(struct net_device *dev, struct netdev_queue_config *qcfg, void *qmem, int idx) @@ -2865,8 +2877,16 @@ static int fbnic_queue_mem_alloc(struct net_device *dev, if (!netif_running(dev)) return fbnic_alloc_qt_page_pools(fbn, qt, idx); + /* A failed PCIe recovery or resume can leave the datapath torn down + * while netif_running() is still true. This ndo runs before + * netdev_rx_queue_restart() checks netif_running(), so bail out + * rather than touching rings and vectors that are already freed. + */ + nv = fbnic_rxq_nv(dev, idx); + if (!nv) + return -ENETDOWN; + real = container_of(fbn->rx[idx], struct fbnic_q_triad, cmpl); - nv = fbn->napi[idx % fbn->num_napi]; fbnic_ring_init(&qt->sub0, real->sub0.doorbell, real->sub0.q_idx, real->sub0.flags); @@ -2917,7 +2937,7 @@ static int fbnic_queue_start(struct net_device *dev, struct fbnic_q_triad *real; real = container_of(fbn->rx[idx], struct fbnic_q_triad, cmpl); - nv = fbn->napi[idx % fbn->num_napi]; + nv = fbnic_rxq_nv(dev, idx); fbnic_aggregate_ring_bdq_counters(fbn, &real->sub0); fbnic_aggregate_ring_bdq_counters(fbn, &real->sub1); @@ -2939,7 +2959,7 @@ static int fbnic_queue_stop(struct net_device *dev, void *qmem, int idx) int err; real = container_of(fbn->rx[idx], struct fbnic_q_triad, cmpl); - nv = fbn->napi[idx % fbn->num_napi]; + nv = fbnic_rxq_nv(dev, idx); fbnic_dbg_nv_exit(nv); napi_disable_locked(&nv->napi); From 4bcc4a92c603fe7f062cea22e20da2e0ad6b12c3 Mon Sep 17 00:00:00 2001 From: Alexander Duyck Date: Mon, 14 Sep 2026 14:10:18 -0700 Subject: [PATCH 042/189] eth: fbnic: reset num_napi when the napi vectors are freed fbn->num_napi is the count of live napi vectors, each of which owns an IRQ. The PM path had freed them without clearing the count. fbnic_pm_suspend() tears the datapath down via ndo_stop() and frees the IRQs, but leaves netif_running() true so resume knows to re-open. Resume rebuilds the datapath in __fbnic_pm_resume() and fbnic_reset_queues() sets num_napi and __fbnic_open() re-allocates the vectors. When the datapath is torn down but never rebuilt, num_napi is left pointing at freed vectors under 2 different scenarios: - a PCIe error recovery that fails (fbnic_err_slot_reset() -> __fbnic_pm_resume() returns an error -> PCI_ERS_RESULT_DISCONNECT), so .resume never runs; or - an __fbnic_open() that fails partway on resume and unwinds, freeing the vectors after fbnic_reset_queues() has already set num_napi. The netdev is then running with num_napi > 0 but napi[] freed, and the eventual remove/unbind close re-enters fbnic_down() -> fbnic_dbg_down() and dereferences the freed vectors: BUG: kernel NULL pointer dereference, address: 0000000000000210 RIP: fbnic_dbg_down+0x28 Clear num_napi when the vectors are freed: in the suspend teardown (a good resume re-establishes it before __fbnic_open()) and on the resume open failure. A redundant ndo_stop() then walks an empty napi[]. The normal ndo_stop() down/up cycle is untouched and keeps num_napi for the next ndo_open(). Fixes: bc6107771bb4 ("eth: fbnic: Allocate a netdevice and napi vectors with queues") Signed-off-by: Alexander Duyck Reviewed-by: Simon Horman Link: https://patch.msgid.link/178942021809.7700.10804028989308077839.stgit@ahduyck-xeon-server.home.arpa Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/meta/fbnic/fbnic_pci.c | 18 ++++++++++++++---- 1 file changed, 14 insertions(+), 4 deletions(-) diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_pci.c b/drivers/net/ethernet/meta/fbnic/fbnic_pci.c index 8b9bc9e8ea56..c6698e3002a1 100644 --- a/drivers/net/ethernet/meta/fbnic/fbnic_pci.c +++ b/drivers/net/ethernet/meta/fbnic/fbnic_pci.c @@ -434,6 +434,7 @@ static int fbnic_pm_suspend(struct device *dev) { struct fbnic_dev *fbd = dev_get_drvdata(dev); struct net_device *netdev = fbd->netdev; + struct fbnic_net *fbn; if (fbnic_init_failure(fbd)) goto null_uc_addr; @@ -441,11 +442,16 @@ static int fbnic_pm_suspend(struct device *dev) rtnl_lock(); netdev_lock(netdev); + fbn = netdev_priv(netdev); + netif_device_detach(netdev); if (netif_running(netdev)) netdev->netdev_ops->ndo_stop(netdev); + /* The IRQs are about to be freed, so drop the napi vector count */ + fbn->num_napi = 0; + netdev_unlock(netdev); rtnl_unlock(); @@ -508,16 +514,20 @@ static int __fbnic_pm_resume(struct device *dev) if (fbnic_init_failure(fbd)) return 0; + rtnl_lock(); + netdev_lock(netdev); + fbn = netdev_priv(netdev); /* Reset the queues if needed */ fbnic_reset_queues(fbn, fbn->num_tx_queues, fbn->num_rx_queues); - rtnl_lock(); - netdev_lock(netdev); - - if (netif_running(netdev)) + if (netif_running(netdev)) { err = __fbnic_open(fbn); + /* On failure the vectors are freed, so drop the count */ + if (err) + fbn->num_napi = 0; + } netdev_unlock(netdev); rtnl_unlock(); From 8947f13e436a4ff5eed9f8f019b2865a07af4bb2 Mon Sep 17 00:00:00 2001 From: Alexander Duyck Date: Mon, 14 Sep 2026 14:10:25 -0700 Subject: [PATCH 043/189] eth: fbnic: Set AW_FLUSH_MODE alongside AW_FLUSH when flushing the mailbox When tearing down the FW mailbox Rx ring, fbnic_mbx_reset_desc_ring() writes AW_CFG with FLUSH set and everything else, BME included, cleared. Clearing BME halts the device's writes to the host but leaves the staged requests parked in the PUL write pipeline rather than draining them, so on the write path FLUSH alone never terminates the outstanding requests and the flush the firmware waits on never completes. Add the FLUSH_MODE definition and set both bits so the staged writes drain out of the pipeline on their own. BME stays cleared, so nothing lands on the host; it is restored later in fbnic_mbx_init_desc_ring() when the ring is rebuilt, once the outstanding writes are gone. The read path is unaffected. AR_CFG has no equivalent mode bit and AR_FLUSH terminates the outstanding reads by itself, so it is left as is. Both writes remain plain stores rather than read-modify-writes. That is deliberate: the matching write in fbnic_mbx_init_desc_ring() restores BME and the TLP attributes, and clears both flush bits as a side effect. Fixes: 3b12f00ddd08 ("fbnic: Gate AXI read/write enabling on FW mailbox") Signed-off-by: Alexander Duyck Reviewed-by: Simon Horman Link: https://patch.msgid.link/178942022583.7700.11050671998277309744.stgit@ahduyck-xeon-server.home.arpa Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/meta/fbnic/fbnic_csr.h | 1 + drivers/net/ethernet/meta/fbnic/fbnic_fw.c | 9 ++++++++- 2 files changed, 9 insertions(+), 1 deletion(-) diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_csr.h b/drivers/net/ethernet/meta/fbnic/fbnic_csr.h index 64b958df7774..14af30e189d6 100644 --- a/drivers/net/ethernet/meta/fbnic/fbnic_csr.h +++ b/drivers/net/ethernet/meta/fbnic/fbnic_csr.h @@ -974,6 +974,7 @@ enum { /* PUL User Registers */ #define FBNIC_CSR_START_PUL_USER 0x31000 /* CSR section delimiter */ #define FBNIC_PUL_OB_TLP_HDR_AW_CFG 0x3103d /* 0xc40f4 */ +#define FBNIC_PUL_OB_TLP_HDR_AW_CFG_FLUSH_MODE CSR_BIT(20) #define FBNIC_PUL_OB_TLP_HDR_AW_CFG_FLUSH CSR_BIT(19) #define FBNIC_PUL_OB_TLP_HDR_AW_CFG_BME CSR_BIT(18) #define FBNIC_PUL_OB_TLP_HDR_AW_CFG_RDE_ATTR CSR_GENMASK(17, 15) diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_fw.c b/drivers/net/ethernet/meta/fbnic/fbnic_fw.c index 283d25fae79e..59aa879798b9 100644 --- a/drivers/net/ethernet/meta/fbnic/fbnic_fw.c +++ b/drivers/net/ethernet/meta/fbnic/fbnic_fw.c @@ -60,8 +60,15 @@ static void fbnic_mbx_reset_desc_ring(struct fbnic_dev *fbd, int mbx_idx) */ switch (mbx_idx) { case FBNIC_IPC_MBX_RX_IDX: + /* Clearing BME blocks the device from writing to the host + * but leaves the requests parked in the write pipeline. The + * write path only clears outstanding requests when both FLUSH + * and FLUSH_MODE are set; FLUSH_MODE lets them drain without + * landing on the host. + */ wr32(fbd, FBNIC_PUL_OB_TLP_HDR_AW_CFG, - FBNIC_PUL_OB_TLP_HDR_AW_CFG_FLUSH); + FBNIC_PUL_OB_TLP_HDR_AW_CFG_FLUSH | + FBNIC_PUL_OB_TLP_HDR_AW_CFG_FLUSH_MODE); break; case FBNIC_IPC_MBX_TX_IDX: wr32(fbd, FBNIC_PUL_OB_TLP_HDR_AR_CFG, From 1b97a269a5bdde20d4e69511f27649c9cb82b7c7 Mon Sep 17 00:00:00 2001 From: Alexander Duyck Date: Mon, 14 Sep 2026 14:10:33 -0700 Subject: [PATCH 044/189] eth: fbnic: Handle FW mailbox completions flagged with an error The firmware can complete a mailbox descriptor while also setting FW_ERR to indicate it could not process the request, for example on a mailbox DMA error. The completion carries no valid data. The driver did not check FW_ERR. On the Rx mailbox it would sync and parse the stale page as a normal message, and on the Tx mailbox it silently freed the request. If the initial capabilities exchange in fbnic_mbx_poll_tx_ready() hit FW_ERR -- on the Tx request or on the Rx response descriptor -- no response was parsed and the poll spun until it timed out even though the ring was healthy. Check FW_ERR on both mailboxes. Count it per-mailbox in fbnic_fw_mbx.resp_error, which is also shown in debugfs, warn (rate limited, since the bit is firmware controlled), and drop the Rx page instead of parsing it. In fbnic_mbx_poll_tx_ready() re-issue the capabilities request when either the Tx or the Rx resp_error counter advances, so a FW_ERR on the request or on its response triggers a retry rather than a timeout. A valid capabilities response is honored before the retry check, so a response parsed in the same poll as an unrelated FW_ERR is not discarded. The counters are mailbox-wide rather than keyed to the capabilities request; that is sufficient here because the exchange runs during bring-up before any other mailbox traffic, and any spurious retry is bounded by the existing 10s timeout. Fixes: da3cde08209e ("eth: fbnic: Add FW communication mechanism") Signed-off-by: Alexander Duyck Reviewed-by: Simon Horman Link: https://patch.msgid.link/178942023343.7700.9423398932961964439.stgit@ahduyck-xeon-server.home.arpa Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/meta/fbnic/fbnic_csr.h | 4 ++ .../net/ethernet/meta/fbnic/fbnic_debugfs.c | 4 +- drivers/net/ethernet/meta/fbnic/fbnic_fw.c | 38 ++++++++++++++++++- drivers/net/ethernet/meta/fbnic/fbnic_fw.h | 1 + 4 files changed, 44 insertions(+), 3 deletions(-) diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_csr.h b/drivers/net/ethernet/meta/fbnic/fbnic_csr.h index 14af30e189d6..baba3471bf5a 100644 --- a/drivers/net/ethernet/meta/fbnic/fbnic_csr.h +++ b/drivers/net/ethernet/meta/fbnic/fbnic_csr.h @@ -1216,6 +1216,10 @@ enum { #define FBNIC_IPC_MBX_DESC_LEN_MASK DESC_GENMASK(63, 48) #define FBNIC_IPC_MBX_DESC_EOM DESC_BIT(46) #define FBNIC_IPC_MBX_DESC_ADDR_MASK DESC_GENMASK(45, 3) +/* Set with FW_CMPL when the FW completed a descriptor without successfully + * processing it (e.g. a mailbox DMA error); the completion has no valid data. + */ +#define FBNIC_IPC_MBX_DESC_FW_ERR DESC_BIT(2) #define FBNIC_IPC_MBX_DESC_FW_CMPL DESC_BIT(1) #define FBNIC_IPC_MBX_DESC_HOST_CMPL DESC_BIT(0) diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_debugfs.c b/drivers/net/ethernet/meta/fbnic/fbnic_debugfs.c index 3c4563c8f403..6edfa0aa69f1 100644 --- a/drivers/net/ethernet/meta/fbnic/fbnic_debugfs.c +++ b/drivers/net/ethernet/meta/fbnic/fbnic_debugfs.c @@ -539,8 +539,8 @@ static void fbnic_dbg_fw_mbx_display(struct seq_file *s, /* Generate header */ seq_puts(s, mbx_idx == FBNIC_IPC_MBX_RX_IDX ? "Rx\n" : "Tx\n"); - seq_printf(s, "Rdy: %d Head: %d Tail: %d\n", - mbx->ready, mbx->head, mbx->tail); + seq_printf(s, "Rdy: %d Head: %d Tail: %d resp_error: %llu\n", + mbx->ready, mbx->head, mbx->tail, mbx->resp_error); snprintf(hdr, sizeof(hdr), "%3s %-4s %s %-12s %s %-3s %-16s\n", "Idx", "Len", "E", "Addr", "F", "H", "Raw"); diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_fw.c b/drivers/net/ethernet/meta/fbnic/fbnic_fw.c index 59aa879798b9..6d7eb8479edf 100644 --- a/drivers/net/ethernet/meta/fbnic/fbnic_fw.c +++ b/drivers/net/ethernet/meta/fbnic/fbnic_fw.c @@ -292,6 +292,12 @@ static void fbnic_mbx_process_tx_msgs(struct fbnic_dev *fbd) if (!(desc & FBNIC_IPC_MBX_DESC_FW_CMPL)) break; + if (desc & FBNIC_IPC_MBX_DESC_FW_ERR) { + tx_mbx->resp_error++; + dev_warn_ratelimited(fbd->dev, + "FW completed a Tx mailbox request with an error\n"); + } + fbnic_mbx_unmap_and_free_msg(fbd, FBNIC_IPC_MBX_TX_IDX, head); head++; @@ -1673,6 +1679,13 @@ static void fbnic_mbx_process_rx_msgs(struct fbnic_dev *fbd) if (!(desc & FBNIC_IPC_MBX_DESC_FW_CMPL)) break; + if (desc & FBNIC_IPC_MBX_DESC_FW_ERR) { + rx_mbx->resp_error++; + dev_warn_ratelimited(fbd->dev, + "FW reported an error on an Rx mailbox message; dropping\n"); + goto next_page; + } + dma_sync_single_for_cpu(fbd->dev, rx_mbx->buf_info[head].addr, FBNIC_RX_PAGE_SIZE, DMA_FROM_DEVICE); @@ -1740,7 +1753,9 @@ void fbnic_mbx_poll(struct fbnic_dev *fbd) int fbnic_mbx_poll_tx_ready(struct fbnic_dev *fbd) { struct fbnic_fw_mbx *tx_mbx = &fbd->mbx[FBNIC_IPC_MBX_TX_IDX]; + struct fbnic_fw_mbx *rx_mbx = &fbd->mbx[FBNIC_IPC_MBX_RX_IDX]; unsigned long timeout = jiffies + 10 * HZ + 1; + u64 tx_resp_error, rx_resp_error; int err, i; do { @@ -1771,6 +1786,9 @@ int fbnic_mbx_poll_tx_ready(struct fbnic_dev *fbd) * mgmt.version once we get the actual version from the firmware * in the capabilities request message. */ +send_cap_req: + tx_resp_error = tx_mbx->resp_error; + rx_resp_error = rx_mbx->resp_error; err = fbnic_fw_xmit_simple_msg(fbd, FBNIC_TLV_MSG_ID_HOST_CAP_REQ); if (err) goto clean_mbx; @@ -1788,9 +1806,27 @@ int fbnic_mbx_poll_tx_ready(struct fbnic_dev *fbd) msleep(20); fbnic_mbx_poll(fbd); + /* A valid capabilities response ends the poll. Check it + * before the FW_ERR retry below so a response parsed in the + * same poll as an unrelated FW_ERR is not discarded. + */ + if (fbd->fw_cap.running.mgmt.version >= MIN_FW_VER_CODE) + break; + /* set err, but wait till mgmt.version check to report it */ - if (!time_is_after_jiffies(timeout)) + if (!time_is_after_jiffies(timeout)) { err = -ETIMEDOUT; + continue; + } + + /* The FW can flag our capabilities request (Tx) or its + * response (Rx) with FW_ERR, in which case it produced no + * usable response. The ring is not wedged, so re-issue the + * request instead of spinning until the timeout. + */ + if (tx_mbx->resp_error != tx_resp_error || + rx_mbx->resp_error != rx_resp_error) + goto send_cap_req; } return 0; diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_fw.h b/drivers/net/ethernet/meta/fbnic/fbnic_fw.h index d84723e4cfa3..5f9969247e30 100644 --- a/drivers/net/ethernet/meta/fbnic/fbnic_fw.h +++ b/drivers/net/ethernet/meta/fbnic/fbnic_fw.h @@ -13,6 +13,7 @@ struct fbnic_tlv_msg; struct fbnic_fw_mbx { u8 ready, head, tail; + u64 resp_error; struct { struct fbnic_tlv_msg *msg; dma_addr_t addr; From d2c31b837406395e576afeb25958c98e9938f3f6 Mon Sep 17 00:00:00 2001 From: Yiqi Sun Date: Tue, 15 Sep 2026 17:50:17 +0800 Subject: [PATCH 045/189] sctp: avoid livelock while updating retransmit path sctp_assoc_update_retran_path() can loop forever when every remaining transport, including retran_path, is SCTP_UNCONFIRMED: the state check runs before the wraparound test, so the loop cannot observe that it has completed a full pass. Fix this by considering a transport only when it is not UNCONFIRMED, then checking whether the walk has returned to retran_path. This makes the full-pass termination independent of the transport state while preserving the existing fallback selection semantics. Also restore the NULL guard around the retran_path assignment. In the all-UNCONFIRMED case there is no eligible replacement transport, and installing NULL would leave later retransmit-path users and the debug print with a NULL path. Fixes: 4c47af4d5eb2 ("net: sctp: rework multihoming retransmission path selection to rfc4960") Signed-off-by: Yiqi Sun Acked-by: Xin Long Link: https://patch.msgid.link/20260915095017.942213-1-sunyiqixm@gmail.com Signed-off-by: Jakub Kicinski --- net/sctp/associola.c | 15 ++++++++------- 1 file changed, 8 insertions(+), 7 deletions(-) diff --git a/net/sctp/associola.c b/net/sctp/associola.c index c0512c827d0f..4521be3bd85a 100644 --- a/net/sctp/associola.c +++ b/net/sctp/associola.c @@ -1289,18 +1289,19 @@ void sctp_assoc_update_retran_path(struct sctp_association *asoc) /* Manually skip the head element. */ if (&trans->transports == &asoc->peer.transport_addr_list) continue; - if (trans->state == SCTP_UNCONFIRMED) - continue; - trans_next = sctp_trans_elect_best(trans, trans_next); - /* Active is good enough for immediate return. */ - if (trans_next->state == SCTP_ACTIVE) - break; + if (trans->state != SCTP_UNCONFIRMED) { + trans_next = sctp_trans_elect_best(trans, trans_next); + /* Active is good enough for immediate return. */ + if (trans_next->state == SCTP_ACTIVE) + break; + } /* We've reached the end, time to update path. */ if (trans == asoc->peer.retran_path) break; } - asoc->peer.retran_path = trans_next; + if (trans_next) + asoc->peer.retran_path = trans_next; pr_debug("%s: association:%p updated new path to addr:%pISpc\n", __func__, asoc, &asoc->peer.retran_path->ipaddr.sa); From 4581c3d2adc3c73a019bc38db64ca11f28bbd7fd Mon Sep 17 00:00:00 2001 From: Ralf Lici Date: Thu, 17 Sep 2026 14:27:23 +0200 Subject: [PATCH 046/189] net/mlx5e: advertise MACsec offload only when supported Commit 339ccec8d43d ("net/mlx5: Enable MACsec offload feature for VLAN interface") added NETIF_F_HW_MACSEC unconditionally to vlan_features so that VLAN devices could inherit MACsec offload support. mlx5e_build_nic_netdev subsequently copies vlan_features into hw_features and features. As a result, all mlx5e NIC netdevices advertise MACsec hardware offload, even when the firmware does not support it and the driver does not install macsec_ops. Set the MACsec feature bits in mlx5e_macsec_build_netdev, after device capabilities have been validated. This preserves MACsec-over-VLAN support and the ethtool feature control on capable devices, without advertising either on unsupported hardware. Fixes: 339ccec8d43d ("net/mlx5: Enable MACsec offload feature for VLAN interface") Cc: stable@vger.kernel.org Reviewed-by: Tariq Toukan Signed-off-by: Ralf Lici Link: https://patch.msgid.link/20260917122724.654639-1-ralf@mandelbit.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/mellanox/mlx5/core/en_accel/macsec.c | 2 ++ drivers/net/ethernet/mellanox/mlx5/core/en_main.c | 1 - 2 files changed, 2 insertions(+), 1 deletion(-) diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en_accel/macsec.c b/drivers/net/ethernet/mellanox/mlx5/core/en_accel/macsec.c index daff53ba7d09..38a3415acf7a 100644 --- a/drivers/net/ethernet/mellanox/mlx5/core/en_accel/macsec.c +++ b/drivers/net/ethernet/mellanox/mlx5/core/en_accel/macsec.c @@ -1724,6 +1724,8 @@ void mlx5e_macsec_build_netdev(struct mlx5e_priv *priv) mlx5_core_dbg(priv->mdev, "mlx5e: MACsec acceleration enabled\n"); netdev->macsec_ops = &macsec_offload_ops; netdev->features |= NETIF_F_HW_MACSEC; + netdev->hw_features |= NETIF_F_HW_MACSEC; + netdev->vlan_features |= NETIF_F_HW_MACSEC; netif_keep_dst(netdev); } diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en_main.c b/drivers/net/ethernet/mellanox/mlx5/core/en_main.c index fc110a7d16e8..4c6060d54fbe 100644 --- a/drivers/net/ethernet/mellanox/mlx5/core/en_main.c +++ b/drivers/net/ethernet/mellanox/mlx5/core/en_main.c @@ -5851,7 +5851,6 @@ static void mlx5e_build_nic_netdev(struct net_device *netdev) netdev->vlan_features |= NETIF_F_SG; netdev->vlan_features |= NETIF_F_HW_CSUM; - netdev->vlan_features |= NETIF_F_HW_MACSEC; netdev->vlan_features |= NETIF_F_GRO; netdev->vlan_features |= NETIF_F_TSO; netdev->vlan_features |= NETIF_F_TSO6; From 6c096bb08de97cdca051fecddad22cac6a1fd275 Mon Sep 17 00:00:00 2001 From: Quentin Armitage Date: Tue, 15 Sep 2026 22:33:21 +0100 Subject: [PATCH 047/189] net: allow IFLA_INET_CONF messages when NLA_F_NESTED unset Commit fa8fca88714c ("ipv4: validate IPV4_DEVCONF attributes properly") added validation of IFLA_INET_CONF attributes, and in the process changed the call of nla_for_each_nested() to nla_parse_nested(). A side effect of this change is that the IFLA_INET_CONF option is now tested for NLA_F_NESTED being set, and fails if it is not. Prior to the commit there was no check of NLA_F_NESTED. Change nla_parse_nested() to nla_parse(). This restores the previous functionality of not checking NLA_F_NESTED, thereby allowing code that (incorrectly) doesn't set NLA_F_NESTED to continue to work. This issue was identified because keepalived started logging errors when it was configuring macvlans that it created. Fixes: fa8fca88714c ("ipv4: validate IPV4_DEVCONF attributes properly") Signed-off-by: Quentin Armitage Reviewed-by: Ido Schimmel Link: https://patch.msgid.link/20260915213320.1527029-2-quentin@armitage.org.uk Signed-off-by: Jakub Kicinski --- net/ipv4/devinet.c | 7 ++++--- 1 file changed, 4 insertions(+), 3 deletions(-) diff --git a/net/ipv4/devinet.c b/net/ipv4/devinet.c index a90be57c63be..5b6b11c943e4 100644 --- a/net/ipv4/devinet.c +++ b/net/ipv4/devinet.c @@ -2117,9 +2117,10 @@ static int inet_validate_link_af(const struct net_device *dev, return err; if (tb[IFLA_INET_CONF]) { - err = nla_parse_nested(nested_tb, IPV4_DEVCONF_MAX, - tb[IFLA_INET_CONF], inet_devconf_policy, - extack); + err = nla_parse(nested_tb, IPV4_DEVCONF_MAX, + nla_data(tb[IFLA_INET_CONF]), + nla_len(tb[IFLA_INET_CONF]), + inet_devconf_policy, extack); if (err < 0) return err; From c82b797abe668d0b668601a93ba2c0b071a63574 Mon Sep 17 00:00:00 2001 From: Weiming Shi Date: Mon, 14 Sep 2026 14:51:23 +0800 Subject: [PATCH 048/189] net/sched: reject IDR error pointers when deleting actions tcf_action_delete() drops the reference held by its lookup before calling tcf_idr_delete_index() with the saved action index. An unlocked classifier can remove that action and reserve the same IDR slot with ERR_PTR(-EBUSY) in between. tcf_idr_delete_index() only checks the lookup result for NULL. It therefore treats the reservation as a tc_action and dereferences tcfa_bindcnt. A hardware execution breakpoint was used to schedule the interleaving without changing the kernel source. KASAN reported this decoded trace: BUG: KASAN: null-ptr-deref in tca_action_gd+0x5b9/0x1010 Read of size 4 at addr 0000000000000010 by task poc/150 Oops: general protection fault, probably for non-canonical address 0xdffffc0000000002 RIP: tca_action_gd+0x5c0/0x1010: arch_atomic_read at arch/x86/include/asm/atomic.h:23 raw_atomic_read at include/linux/atomic/atomic-arch-fallback.h:457 atomic_read at include/linux/atomic/atomic-instrumented.h:33 tcf_idr_delete_index at net/sched/act_api.c:766 tcf_action_delete at net/sched/act_api.c:1859 tcf_del_notify at net/sched/act_api.c:2014 tca_action_gd at net/sched/act_api.c:2064 R13: 0000000000000010 R15: fffffffffffffff0 Kernel panic - not syncing: Fatal exception R15 contains ERR_PTR(-EBUSY), and adding the tcfa_bindcnt offset produces the address in R13. With the guard applied, the same reproducer returned -ENOENT without a KASAN report or panic. Treat error pointers as absent and return -ENOENT. Fixes: 0190c1d452a9 ("net: sched: atomically check-allocate action") Cc: stable@vger.kernel.org Reported-by: Xiang Mei Signed-off-by: Weiming Shi Link: https://patch.msgid.link/20260914065123.4109709-2-bestswngs@gmail.com Signed-off-by: Jakub Kicinski --- net/sched/act_api.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/net/sched/act_api.c b/net/sched/act_api.c index 3f653721c45f..e45a63be397c 100644 --- a/net/sched/act_api.c +++ b/net/sched/act_api.c @@ -758,7 +758,7 @@ static int tcf_idr_delete_index(struct tcf_idrinfo *idrinfo, u32 index) mutex_lock(&idrinfo->lock); p = idr_find(&idrinfo->action_idr, index); - if (!p) { + if (IS_ERR_OR_NULL(p)) { mutex_unlock(&idrinfo->lock); return -ENOENT; } From 9d565b6b72fe3f41fd43636e143072848105189f Mon Sep 17 00:00:00 2001 From: Aamir Ahmed Date: Tue, 15 Sep 2026 00:06:58 +0100 Subject: [PATCH 049/189] net: usb: catc: bound the RX packet length in catc_rx_done() catc_rx_done() walks a multi-packet URB, reading a two-byte length from each packet header. Its bound, pkt_len > urb->actual_length, ignores the header offset and compares against the whole transfer rather than the bytes left from pkt_start, so a crafted packet header makes skb_copy_to_linear_data() read past the buffer. A length below ETH_HLEN is also accepted, including zero, and eth_type_trans() then reads a MAC header from the uninitialised tailroom of a shorter skb. The is_f5u011 branch takes its length straight from the transfer, so a zero-length URB reaches the same path. Track the bytes remaining from the current packet, and reject a header that does not fit, a length past what is left, and a length below an Ethernet header. A transfer shorter than an Ethernet header, including a zero-length one, previously became a runt skb passed to netif_rx() and counted as received; it is now counted in rx_length_errors and ends the walk. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Signed-off-by: Aamir Ahmed Reviewed-by: Simon Horman Link: https://patch.msgid.link/AS8P251MB00015FD7716F38C345619B56C8BB2@AS8P251MB0001.EURP251.PROD.OUTLOOK.COM Signed-off-by: Jakub Kicinski --- drivers/net/usb/catc.c | 15 ++++++++++++--- 1 file changed, 12 insertions(+), 3 deletions(-) diff --git a/drivers/net/usb/catc.c b/drivers/net/usb/catc.c index 96e82f94edcf..39b678f175dd 100644 --- a/drivers/net/usb/catc.c +++ b/drivers/net/usb/catc.c @@ -233,17 +233,26 @@ static void catc_rx_done(struct urb *urb) } do { - if(!catc->is_f5u011) { - pkt_len = le16_to_cpup((__le16*)pkt_start); - if (pkt_len > urb->actual_length) { + int remaining = urb->actual_length - + (pkt_start - (u8 *)urb->transfer_buffer); + + if (!catc->is_f5u011) { + if (remaining < pkt_offset) { catc->netdev->stats.rx_length_errors++; catc->netdev->stats.rx_errors++; break; } + pkt_len = le16_to_cpup((__le16 *)pkt_start); } else { pkt_len = urb->actual_length; } + if (pkt_len < ETH_HLEN || pkt_len + pkt_offset > remaining) { + catc->netdev->stats.rx_length_errors++; + catc->netdev->stats.rx_errors++; + break; + } + if (!(skb = dev_alloc_skb(pkt_len))) return; From ab888242fce4f16f6c4d4c6ec53939ad36aa3b3a Mon Sep 17 00:00:00 2001 From: Xiang Mei Date: Tue, 15 Sep 2026 01:31:52 -0700 Subject: [PATCH 050/189] vlan: require the MAC header to be present in __vlan_insert_inner_tag() __vlan_insert_inner_tag() only guarantees head room via skb_cow_head(), never that mac_len bytes of MAC header are present. Its ETH_HLEN wrappers - __vlan_insert_tag() under skb_vlan_push(), and vlan_insert_tag() under validate_xmit_vlan() on the generic transmit path - therefore rewrite the first 16 bytes at skb->data: a 12-byte memmove plus two 2-byte stores at +12 and +14. No caller supplies the bound, while the pop helpers use skb_ensure_writable()/pskb_may_pull(). An IFF_TUN device has hard_header_len == 0, so packet_snd() accepts a one-byte AF_PACKET/SOCK_RAW frame. The first vlan push only sets a hwaccel tag; the next - clsact "action vlan push" or bpf_skb_vlan_push() - enters the helper with skb->len still 1. The head comes from skbuff_small_head without __GFP_ZERO, so each push drags bytes from beyond skb->tail into the frame. After three the one-byte send leaves as 13 bytes carrying 11 bytes of uninitialised slab: 0000: 5a b3 62 12 80 88 ff ff 00 b3 62 12 81 `------------------------------' only 0x5a was sent; the rest is slab, here the top 56 bits of a linear-map address Require the MAC header the helper rewrites to be present, so such a frame is dropped rather than transmitted. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Cc: stable@vger.kernel.org Reported-by: co+0ea1ac045375cf05@bugs.sh Signed-off-by: Xiang Mei Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260915083152.705309-1-xmei5@asu.edu Signed-off-by: Jakub Kicinski --- include/linux/if_vlan.h | 3 +++ 1 file changed, 3 insertions(+) diff --git a/include/linux/if_vlan.h b/include/linux/if_vlan.h index 20cc16ea4e5a..4846032bf4ff 100644 --- a/include/linux/if_vlan.h +++ b/include/linux/if_vlan.h @@ -365,6 +365,9 @@ static inline int __vlan_insert_inner_tag(struct sk_buff *skb, const u8 meta_len = mac_len > ETH_TLEN ? skb_metadata_len(skb) : 0; struct vlan_ethhdr *veth; + if (unlikely(!pskb_may_pull(skb, mac_len))) + return -EINVAL; + if (skb_cow_head(skb, meta_len + VLAN_HLEN) < 0) return -ENOMEM; From 8a60ade2277e1f0e0d0578d565354e52292fa46d Mon Sep 17 00:00:00 2001 From: Jamal Hadi Salim Date: Thu, 17 Sep 2026 06:57:32 -0400 Subject: [PATCH 051/189] net/sched: sch_hfsc: bound the classify inner-filter walk with a drift budget MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit hfsc_classify() applies the "filter may only point downwards" level check only when the filter result carries no bound class. A filter created with a flowid gets res.class set once at bind time, so the check never runs for it during classification. hfsc_adjust_levels() can later raise a class's level without revalidating existing bindings, leaving two binds that were each legal at bind time pointing at each other; the classify walk then bounces between two interior classes forever with the qdisc lock held and BH disabled — a soft lockup from a single packet. The stuck walk trips the watchdog: watchdog: BUG: soft lockup - CPU#3 stuck for 13s! [ping:444] RIP: 0010:u32_classify+0x542/0x17f0 ... tcf_classify+0x66/0xa0 hfsc_enqueue+0x166/0xdf0 Bound the traversal with a budget of non-descending hops, the only way a configured walk can move without descending the class tree once levels drift after bind time. The budget is cumulative over the whole walk and is deliberately not reset on a descending hop: a chain that alternates a descent with a lateral hop would return the budget every lap and never trip. Descending hops never decrement it, so legitimately deep trees are unaffected and a terminating lateral chain still classifies normally. Drop the packet with a rate-limited warning once the budget is exhausted, mirroring the merged HTB fix. This is a follow-up to commit 729c4896ab82 ("net/sched: sch_htb: limit htb_classify inner-class filter hops"), which bounded the same classify loop on the HTB side but left the HFSC walk unbounded. Conditions to recreate the bug: - CONFIG_NET_SCHED, CONFIG_NET_SCH_HFSC, CONFIG_NET_CLS_U32, CONFIG_LOCKUP_DETECTOR. - Build a cycle with two legal-at-bind-time flowid binds and a level drift: class X 1:1 (child of root) with leaf child 1:10; class Y 1:2 (sibling of X) with children 1:20 and 1:200; root u32 filter flowid 1:1; filter on X flowid 1:2 (legal when Y is a leaf); after Y's level rises to 2, filter on Y flowid 1:1 (legal then). Send one packet (ping on the device). Unfixed kernel: classify spins with the qdisc lock held; with softlockup_panic=1 it panics. - Reachable from unprivileged user via unshare -Urn (CAP_NET_ADMIN). Fixes: a2f79227138c ("net_sched: sch_hfsc: fix classification loops") Reported-by: Sashiko (gemini + nipa) Closes: https://lore.kernel.org/netdev/QDISC-CTUU.v2.20260913192614@mojatatu.com/ Link: https://sashiko.dev/#/patchset/QDISC-CTUU.v2.20260913192614@mojatatu.com Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/QDISC-CTUU.v2.20260913192614%40mojatatu.com Reviewed-by: Victor Nogueira Tested-by: hybris Signed-off-by: Jamal Hadi Salim Link: https://patch.msgid.link/QDISC-CTUU.v3.20260916184908@mojatatu.com Signed-off-by: Jakub Kicinski --- net/sched/sch_hfsc.c | 22 ++++++++++++++++++++++ 1 file changed, 22 insertions(+) diff --git a/net/sched/sch_hfsc.c b/net/sched/sch_hfsc.c index e87f5021a199..284490fd6ca9 100644 --- a/net/sched/sch_hfsc.c +++ b/net/sched/sch_hfsc.c @@ -386,6 +386,15 @@ cftree_update(struct hfsc_class *cl) #define SM_MASK ((1ULL << SM_SHIFT) - 1) #define ISM_MASK ((1ULL << ISM_SHIFT) - 1) +/* + * Cap on the non-descending hops a classify walk may take before its + * filter chain is treated as misconfigured. A flowid binding that was + * legal at bind time can become lateral once hfsc_adjust_levels() + * raises a class level; a few such hops are legitimate, an unbounded + * run means the chain cycles. + */ +#define HFSC_CLASSIFY_MAX_DRIFT 8 + static inline u64 seg_x2y(u64 x, u64 sm) { @@ -1133,6 +1142,7 @@ hfsc_classify(struct sk_buff *skb, struct Qdisc *sch, int *qerr) struct hfsc_class *head, *cl; struct tcf_result res; struct tcf_proto *tcf; + unsigned int drift; int result; if (TC_H_MAJ(skb->priority ^ sch->handle) == 0 && @@ -1142,6 +1152,7 @@ hfsc_classify(struct sk_buff *skb, struct Qdisc *sch, int *qerr) *qerr = NET_XMIT_SUCCESS | __NET_XMIT_BYPASS; head = &q->root; + drift = HFSC_CLASSIFY_MAX_DRIFT; tcf = rcu_dereference_bh(q->root.filter_list); while (tcf && (result = tcf_classify_qdisc(skb, tcf, &res, false)) >= 0) { #ifdef CONFIG_NET_CLS_ACT @@ -1167,6 +1178,17 @@ hfsc_classify(struct sk_buff *skb, struct Qdisc *sch, int *qerr) if (cl->level == 0) return cl; /* hit leaf class */ + /* + * flowid binds skip the level check above (res.class is set + * at bind time and levels drift after), so a walk can follow + * lateral hops without descending; a bounded number of them + * is legal, more means the chain cycles. + */ + if (cl->level >= head->level && drift-- == 0) { + pr_warn_ratelimited("hfsc: classify hop budget exhausted, dropping packet\n"); + return NULL; + } + /* apply inner filter chain */ tcf = rcu_dereference_bh(cl->filter_list); head = cl; From 1e24c4f2ee44be0eee94092b5d13cbdb4bdf0d60 Mon Sep 17 00:00:00 2001 From: Jamal Hadi Salim Date: Thu, 17 Sep 2026 06:57:33 -0400 Subject: [PATCH 052/189] selftests: tc-testing: add a lateral-drift hfsc classify-walk test The classify-loop fix bounds a walk's non-descending hops, so the guard must not misfire on a legal walk that reaches its leaf through a level-drift lateral chain. Add a case that builds exactly that chain and asserts traffic still reaches the chain's own leaf. A lateral hop can only exist because a bind was legal when it was made and a later class add raised the target's level, so the setup binds each hop while the target is still a leaf and only then deepens it: bind 1:1 -> 1:2 while 1:2 is a leaf, add 1:20 under 1:2, add 1:3 and bind 1:2 -> 1:3 while 1:3 is a leaf, then add 1:30 and 1:31 under 1:3 and bind 1:3 -> 1:31. The walk root -> 1:1 -> 1:2 -> 1:3 -> 1:31 then takes two lateral hops and must reach leaf 1:31. The default class is 1:30, distinct from the asserted leaf, and the verify pattern is anchored to the 1:31 stats line, so neither a fall-through to the default nor a nonzero count on another class can satisfy the check. On the patched kernel the test passes; with the bound forced to zero the walk falls to the default and 1:31 stays idle, so the test fails. Reviewed-by: Victor Nogueira Tested-by: hybris Signed-off-by: Jamal Hadi Salim Link: https://patch.msgid.link/QDISC-CTUU.v3.20260916184908@mojatatu.com.2 Signed-off-by: Jakub Kicinski --- .../tc-testing/tc-tests/qdiscs/hfsc.json | 34 +++++++++++++++++++ 1 file changed, 34 insertions(+) diff --git a/tools/testing/selftests/tc-testing/tc-tests/qdiscs/hfsc.json b/tools/testing/selftests/tc-testing/tc-tests/qdiscs/hfsc.json index c98c339424d4..4f6bbb8b57f9 100644 --- a/tools/testing/selftests/tc-testing/tc-tests/qdiscs/hfsc.json +++ b/tools/testing/selftests/tc-testing/tc-tests/qdiscs/hfsc.json @@ -169,5 +169,39 @@ "teardown": [ "$TC qdisc del dev $DUMMY handle 1: root" ] + }, + { + "id": "8c39", + "name": "HFSC classify walk still reaches leaf after lateral drift", + "category": [ + "qdisc", + "hfsc" + ], + "plugins": { + "requires": "nsPlugin" + }, + "setup": [ + "ip link set lo up", + "$TC qdisc add dev lo handle 1: root hfsc default 30", + "$TC class add dev lo parent 1: classid 1:1 hfsc rt m2 100kbit", + "$TC class add dev lo parent 1:1 classid 1:10 hfsc rt m2 50kbit", + "$TC class add dev lo parent 1: classid 1:2 hfsc rt m2 100kbit", + "$TC filter add dev lo parent 1: protocol ip prio 1 u32 match u8 0 0 at 0 flowid 1:1", + "$TC filter add dev lo parent 1:1 protocol ip prio 1 u32 match u8 0 0 at 0 flowid 1:2", + "$TC class add dev lo parent 1:2 classid 1:20 hfsc rt m2 10kbit", + "$TC class add dev lo parent 1: classid 1:3 hfsc rt m2 100kbit", + "$TC filter add dev lo parent 1:2 protocol ip prio 1 u32 match u8 0 0 at 0 flowid 1:3", + "$TC class add dev lo parent 1:3 classid 1:30 hfsc rt m2 10kbit", + "$TC class add dev lo parent 1:3 classid 1:31 hfsc rt m2 100kbit", + "$TC filter add dev lo parent 1:3 protocol ip prio 1 u32 match u8 0 0 at 0 flowid 1:31" + ], + "cmdUnderTest": "ping -n -c 10 -W 1 127.0.0.1", + "expExitCode": "0", + "verifyCmd": "$TC -s class show dev lo", + "matchPattern": "class hfsc 1:31 parent 1:3 rt[^\\n]*\\n Sent [0-9]+ bytes [1-9][0-9]* pkt", + "matchCount": "1", + "teardown": [ + "$TC qdisc del dev lo handle 1: root" + ] } ] From 7a6d08ee0f0e30023d18779bb314db8fd9a3b6d4 Mon Sep 17 00:00:00 2001 From: Ralf Lici Date: Fri, 28 Aug 2026 16:50:22 +0200 Subject: [PATCH 053/189] ovpn: preserve IPv6 scope id for netlink peer endpoints ovpn accepts OVPN_A_PEER_REMOTE_IPV6_SCOPE_ID and reports bind->remote.in6.sin6_scope_id in peer dumps, but the netlink endpoint parser never copied the attribute into the sockaddr_in6 used to create or update the peer bind. As a result, an IPv6 link-local remote endpoint configured through netlink loses its interface scope, unlike on the peer float path where ipv6_iface_scope_id populates the field. The UDPv6 output path then builds a flow with flowi6_oif set to zero and route lookup can fail or select the wrong interface. Copy the scope id when parsing non-v4-mapped IPv6 remote endpoints. The existing precheck already rejects the scope-id attribute for IPv4 and v4-mapped IPv6 remotes. Fixes: 1d36a36f6d53 ("ovpn: implement peer add/get/dump/delete via netlink") Signed-off-by: Ralf Lici Signed-off-by: Antonio Quartulli --- drivers/net/ovpn/netlink.c | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/drivers/net/ovpn/netlink.c b/drivers/net/ovpn/netlink.c index 4dad85294198..2ba762082acc 100644 --- a/drivers/net/ovpn/netlink.c +++ b/drivers/net/ovpn/netlink.c @@ -100,6 +100,8 @@ static bool ovpn_nl_attr_sockaddr_remote(struct nlattr **attrs, struct sockaddr_in6 *sin6; struct sockaddr_in *sin; struct in6_addr *in6; + struct nlattr *scope; + u32 scope_id = 0; __be16 port = 0; __be32 *in; @@ -114,6 +116,9 @@ static bool ovpn_nl_attr_sockaddr_remote(struct nlattr **attrs, } else if (attrs[OVPN_A_PEER_REMOTE_IPV6]) { ss->ss_family = AF_INET6; in6 = nla_data(attrs[OVPN_A_PEER_REMOTE_IPV6]); + scope = attrs[OVPN_A_PEER_REMOTE_IPV6_SCOPE_ID]; + if (scope) + scope_id = nla_get_u32(scope); } else { return false; } @@ -126,6 +131,7 @@ static bool ovpn_nl_attr_sockaddr_remote(struct nlattr **attrs, if (!ipv6_addr_v4mapped(in6)) { sin6 = (struct sockaddr_in6 *)ss; sin6->sin6_port = port; + sin6->sin6_scope_id = scope_id; memcpy(&sin6->sin6_addr, in6, sizeof(*in6)); break; } From 77393b4d72dfeb764b2af2b848acc659f6fcfd0a Mon Sep 17 00:00:00 2001 From: Ralf Lici Date: Fri, 28 Aug 2026 16:50:23 +0200 Subject: [PATCH 054/189] ovpn: skip UDP source validation for unspecified addresses ovpn validates the cached local UDP source address before reusing or refreshing a peer dst cache. This is only meaningful when a concrete source address is selected. For IPv6, calling ipv6_chk_addr with :: checks whether the unspecified address itself is configured on the host. A peer may legitimately have bind->local.ipv6 set to :: when no local endpoint was configured or after a stale learned address was cleared. In that case the source should be left unspecified and selected by ip6_dst_lookup_flow(). For IPv4, inet_confirm_addr(..., local = 0, ...) asks for local address autoselection rather than validating a chosen source. Skip the precheck there as well and let ip_route_output_flow select or reject the source. Only validate non-zero/non-any source addresses. Fixes: 08857b5ec5d9 ("ovpn: implement basic TX path (UDP)") Signed-off-by: Ralf Lici Signed-off-by: Antonio Quartulli --- drivers/net/ovpn/udp.c | 7 ++++--- 1 file changed, 4 insertions(+), 3 deletions(-) diff --git a/drivers/net/ovpn/udp.c b/drivers/net/ovpn/udp.c index 7f69e8890b5b..df4750dabd1e 100644 --- a/drivers/net/ovpn/udp.c +++ b/drivers/net/ovpn/udp.c @@ -161,8 +161,8 @@ static int ovpn_udp4_output(struct ovpn_peer *peer, struct ovpn_bind *bind, if (rt) goto transmit; - if (unlikely(!inet_confirm_addr(sock_net(sk), NULL, 0, fl.saddr, - RT_SCOPE_HOST))) { + if (fl.saddr && unlikely(!inet_confirm_addr(sock_net(sk), NULL, 0, + fl.saddr, RT_SCOPE_HOST))) { /* we may end up here when the cached address is not usable * anymore. In this case we reset address/cache and perform a * new look up @@ -238,7 +238,8 @@ static int ovpn_udp6_output(struct ovpn_peer *peer, struct ovpn_bind *bind, if (dst) goto transmit; - if (unlikely(!ipv6_chk_addr(sock_net(sk), &fl.saddr, NULL, 0))) { + if (!ipv6_addr_any(&fl.saddr) && + unlikely(!ipv6_chk_addr(sock_net(sk), &fl.saddr, NULL, 0))) { /* we may end up here when the cached address is not usable * anymore. In this case we reset address/cache and perform a * new look up From 7c66b7a4ae80a9309e6dc1d24b7b6b897e6348eb Mon Sep 17 00:00:00 2001 From: Ralf Lici Date: Fri, 28 Aug 2026 16:50:24 +0200 Subject: [PATCH 055/189] ovpn: track UDP socket route key for peer dst cache ovpn stores the route used to transmit UDP packets in a per-peer dst cache. A cached dst is only valid for the route lookup inputs used when it was resolved. Some of those inputs are mutable while userspace still owns the UDP socket. In particular, changes to the socket mark or UDP source port do not invalidate ovpn's peer dst cache, so ovpn can keep using a route selected with an old socket route key. Replace the cached mark with a route key containing the socket-owned lookup inputs currently used by ovpn, and reset the peer dst cache when the key changes. Before storing a newly looked-up dst, recheck the route key under the peer lock so a dst resolved for stale socket state is not published. Fixes: 08857b5ec5d9 ("ovpn: implement basic TX path (UDP)") Signed-off-by: Ralf Lici Signed-off-by: Antonio Quartulli --- drivers/net/ovpn/peer.c | 1 + drivers/net/ovpn/peer.h | 19 ++++++++- drivers/net/ovpn/udp.c | 90 +++++++++++++++++++++++++++++++++++------ 3 files changed, 95 insertions(+), 15 deletions(-) diff --git a/drivers/net/ovpn/peer.c b/drivers/net/ovpn/peer.c index c95656ca7c35..b400783c2efa 100644 --- a/drivers/net/ovpn/peer.c +++ b/drivers/net/ovpn/peer.c @@ -113,6 +113,7 @@ struct ovpn_peer *ovpn_peer_new(struct ovpn_priv *ovpn, u32 id) RCU_INIT_POINTER(peer->bind, NULL); ovpn_crypto_state_init(&peer->crypto); spin_lock_init(&peer->lock); + seqcount_spinlock_init(&peer->route_key_seq, &peer->lock); kref_init(&peer->refcount); ovpn_peer_stats_init(&peer->vpn_stats); ovpn_peer_stats_init(&peer->link_stats); diff --git a/drivers/net/ovpn/peer.h b/drivers/net/ovpn/peer.h index dfa5c0037e02..063535699ecd 100644 --- a/drivers/net/ovpn/peer.h +++ b/drivers/net/ovpn/peer.h @@ -10,6 +10,7 @@ #ifndef _NET_OVPN_OVPNPEER_H_ #define _NET_OVPN_OVPNPEER_H_ +#include #include #include @@ -17,6 +18,16 @@ #include "socket.h" #include "stats.h" +/** + * struct ovpn_route_key - route key used for the peer dst cache + * @mark: fwmark used for route lookup + * @sport: UDP source port used for route lookup + */ +struct ovpn_route_key { + u32 mark; + __be16 sport; +}; + /** * struct ovpn_peer - the main remote peer object * @ovpn: main openvpn instance this peer belongs to @@ -45,6 +56,8 @@ * @tcp.sk_cb.ops: pointer to the original prot_ops object (TCP only) * @crypto: the crypto configuration (ciphers, keys, etc..) * @dst_cache: cache for dst_entry used to send to peer + * @route_key: route key matching the current dst cache contents + * @route_key_seq: seqcount protecting lockless route_key reads * @bind: remote peer binding * @keepalive_interval: seconds after which a new keepalive should be sent * @keepalive_xmit_exp: future timestamp when next keepalive should be sent @@ -55,7 +68,7 @@ * @vpn_stats: per-peer in-VPN TX/RX stats * @link_stats: per-peer link/transport TX/RX stats * @delete_reason: why peer was deleted (i.e. timeout, transport error, ..) - * @lock: protects binding to peer (bind) and keepalive* fields + * @lock: protects binding to peer (bind), route_key and keepalive* fields * @refcount: reference counter * @rcu: used to free peer in an RCU safe way * @release_entry: entry for the socket release list @@ -99,6 +112,8 @@ struct ovpn_peer { } tcp; struct ovpn_crypto_state crypto; struct dst_cache dst_cache; + struct ovpn_route_key route_key; + seqcount_spinlock_t route_key_seq; struct ovpn_bind __rcu *bind; unsigned long keepalive_interval; unsigned long keepalive_xmit_exp; @@ -109,7 +124,7 @@ struct ovpn_peer { struct ovpn_peer_stats vpn_stats; struct ovpn_peer_stats link_stats; enum ovpn_del_peer_reason delete_reason; - spinlock_t lock; /* protects bind and keepalive* */ + spinlock_t lock; /* protects bind, route_key and keepalive* */ struct kref refcount; struct rcu_head rcu; struct llist_node release_entry; diff --git a/drivers/net/ovpn/udp.c b/drivers/net/ovpn/udp.c index df4750dabd1e..c6d591cb7ff4 100644 --- a/drivers/net/ovpn/udp.c +++ b/drivers/net/ovpn/udp.c @@ -131,6 +131,48 @@ static int ovpn_udp_encap_recv(struct sock *sk, struct sk_buff *skb) return 0; } +static bool ovpn_route_key_equal(const struct ovpn_route_key *a, + const struct ovpn_route_key *b) +{ + return a->mark == b->mark && a->sport == b->sport; +} + +/** + * ovpn_dst_cache_check_key - reset peer dst cache after key changes + * @peer: the peer owning the dst cache + * @cache: the cache that might need to be reset + * @key: the route key for the packet being transmitted + * + * Reset the peer dst cache if it was populated for a different route key. + */ +static void ovpn_dst_cache_check_key(struct ovpn_peer *peer, + struct dst_cache *cache, + const struct ovpn_route_key *key) +{ + struct ovpn_route_key old_key; + unsigned int seq; + + /* snapshot the saved key before deciding whether the cache matches */ + do { + seq = read_seqcount_begin(&peer->route_key_seq); + old_key = peer->route_key; + } while (read_seqcount_retry(&peer->route_key_seq, seq)); + + /* nothing changed: the current cache can be reused */ + if (likely(ovpn_route_key_equal(&old_key, key))) + return; + + /* recheck under lock because another path may have updated the key */ + spin_lock_bh(&peer->lock); + if (!ovpn_route_key_equal(&peer->route_key, key)) { + write_seqcount_begin(&peer->route_key_seq); + peer->route_key = *key; + dst_cache_reset(cache); + write_seqcount_end(&peer->route_key_seq); + } + spin_unlock_bh(&peer->lock); +} + /** * ovpn_udp4_output - send IPv4 packet over udp socket * @peer: the destination peer @@ -138,21 +180,23 @@ static int ovpn_udp_encap_recv(struct sock *sk, struct sk_buff *skb) * @cache: dst cache * @sk: the socket to send the packet over * @skb: the packet to send + * @key: the route key snapshot used for cache validation and flow lookup * * Return: 0 on success or a negative error code otherwise */ static int ovpn_udp4_output(struct ovpn_peer *peer, struct ovpn_bind *bind, struct dst_cache *cache, struct sock *sk, - struct sk_buff *skb) + struct sk_buff *skb, + const struct ovpn_route_key *key) { struct rtable *rt; struct flowi4 fl = { .saddr = bind->local.ipv4.s_addr, .daddr = bind->remote.in4.sin_addr.s_addr, - .fl4_sport = inet_sk(sk)->inet_sport, + .fl4_sport = key->sport, .fl4_dport = bind->remote.in4.sin_port, .flowi4_proto = sk->sk_protocol, - .flowi4_mark = sk->sk_mark, + .flowi4_mark = key->mark, }; int ret; @@ -193,7 +237,12 @@ static int ovpn_udp4_output(struct ovpn_peer *peer, struct ovpn_bind *bind, ret); goto err; } - dst_cache_set_ip4(cache, &rt->dst, fl.saddr); + + /* avoid storing a stale cache */ + spin_lock_bh(&peer->lock); + if (likely(ovpn_route_key_equal(key, &peer->route_key))) + dst_cache_set_ip4(cache, &rt->dst, fl.saddr); + spin_unlock_bh(&peer->lock); transmit: udp_tunnel_xmit_skb(rt, sk, skb, fl.saddr, fl.daddr, 0, @@ -213,12 +262,14 @@ static int ovpn_udp4_output(struct ovpn_peer *peer, struct ovpn_bind *bind, * @cache: dst cache * @sk: the socket to send the packet over * @skb: the packet to send + * @key: the route key snapshot used for cache validation and flow lookup * * Return: 0 on success or a negative error code otherwise */ static int ovpn_udp6_output(struct ovpn_peer *peer, struct ovpn_bind *bind, struct dst_cache *cache, struct sock *sk, - struct sk_buff *skb) + struct sk_buff *skb, + const struct ovpn_route_key *key) { struct dst_entry *dst; int ret; @@ -226,10 +277,10 @@ static int ovpn_udp6_output(struct ovpn_peer *peer, struct ovpn_bind *bind, struct flowi6 fl = { .saddr = bind->local.ipv6, .daddr = bind->remote.in6.sin6_addr, - .fl6_sport = inet_sk(sk)->inet_sport, + .fl6_sport = key->sport, .fl6_dport = bind->remote.in6.sin6_port, .flowi6_proto = sk->sk_protocol, - .flowi6_mark = sk->sk_mark, + .flowi6_mark = key->mark, .flowi6_oif = bind->remote.in6.sin6_scope_id, }; @@ -259,7 +310,12 @@ static int ovpn_udp6_output(struct ovpn_peer *peer, struct ovpn_bind *bind, &bind->remote.in6, ret); goto err; } - dst_cache_set_ip6(cache, dst, &fl.saddr); + + /* avoid storing a stale cache */ + spin_lock_bh(&peer->lock); + if (likely(ovpn_route_key_equal(key, &peer->route_key))) + dst_cache_set_ip6(cache, dst, &fl.saddr); + spin_unlock_bh(&peer->lock); transmit: /* user IPv6 packets may be larger than the transport interface @@ -288,6 +344,7 @@ static int ovpn_udp6_output(struct ovpn_peer *peer, struct ovpn_bind *bind, * @cache: dst cache * @sk: the socket to send the packet over * @skb: the packet to send + * @key: route key snapshot used for cache validation and flow lookup * * rcu_read_lock should be held on entry. * On return, the skb is consumed. @@ -295,7 +352,8 @@ static int ovpn_udp6_output(struct ovpn_peer *peer, struct ovpn_bind *bind, * Return: 0 on success or a negative error code otherwise */ static int ovpn_udp_output(struct ovpn_peer *peer, struct dst_cache *cache, - struct sock *sk, struct sk_buff *skb) + struct sock *sk, struct sk_buff *skb, + struct ovpn_route_key *key) { struct ovpn_bind *bind; int ret; @@ -315,11 +373,11 @@ static int ovpn_udp_output(struct ovpn_peer *peer, struct dst_cache *cache, switch (bind->remote.in4.sin_family) { case AF_INET: - ret = ovpn_udp4_output(peer, bind, cache, sk, skb); + ret = ovpn_udp4_output(peer, bind, cache, sk, skb, key); break; #if IS_ENABLED(CONFIG_IPV6) case AF_INET6: - ret = ovpn_udp6_output(peer, bind, cache, sk, skb); + ret = ovpn_udp6_output(peer, bind, cache, sk, skb, key); break; #endif default: @@ -341,15 +399,21 @@ static int ovpn_udp_output(struct ovpn_peer *peer, struct dst_cache *cache, void ovpn_udp_send_skb(struct ovpn_peer *peer, struct sock *sk, struct sk_buff *skb) { + struct ovpn_route_key key = { + .mark = READ_ONCE(sk->sk_mark), + .sport = READ_ONCE(inet_sk(sk)->inet_sport), + }; int ret; skb->dev = peer->ovpn->dev; - skb->mark = READ_ONCE(sk->sk_mark); + skb->mark = key.mark; /* no checksum performed at this layer */ skb->ip_summed = CHECKSUM_NONE; + ovpn_dst_cache_check_key(peer, &peer->dst_cache, &key); + /* crypto layer -> transport (UDP) */ - ret = ovpn_udp_output(peer, &peer->dst_cache, sk, skb); + ret = ovpn_udp_output(peer, &peer->dst_cache, sk, skb, &key); if (unlikely(ret < 0)) kfree_skb(skb); } From fa603710bdb9aea33c0d9cc2c05ed24d84f58753 Mon Sep 17 00:00:00 2001 From: Ralf Lici Date: Fri, 28 Aug 2026 16:50:25 +0200 Subject: [PATCH 056/189] ovpn: validate peer state before caching UDP dst UDP route lookup runs without peer->lock while the bind is protected by RCU. The route key is snapshotted separately. Either can change while the lookup is in progress. The TX path currently checks only the route key before publishing the looked-up dst. If the bind changes but the route key does not, a dst resolved from the old endpoint can be installed in the cache after the bind replacement. Compare both the bind pointer and the route key under peer->lock before updating the cache. The RCU read-side critical section keeps the old bind alive throughout the lookup, so pointer identity is sufficient to detect a replacement. Fixes: f0281c1d3732 ("ovpn: add support for updating local or remote UDP endpoint") Signed-off-by: Ralf Lici Signed-off-by: Antonio Quartulli --- drivers/net/ovpn/udp.c | 33 +++++++++++++++++++++++++++++++-- 1 file changed, 31 insertions(+), 2 deletions(-) diff --git a/drivers/net/ovpn/udp.c b/drivers/net/ovpn/udp.c index c6d591cb7ff4..eeef4a7229f5 100644 --- a/drivers/net/ovpn/udp.c +++ b/drivers/net/ovpn/udp.c @@ -173,6 +173,35 @@ static void ovpn_dst_cache_check_key(struct ovpn_peer *peer, spin_unlock_bh(&peer->lock); } +/** + * ovpn_dst_cache_current - check whether a route lookup matches peer state + * @peer: the peer owning the bind and dst cache + * @bind: the RCU bind used for the route lookup + * @key: the route key used for the route lookup + * + * Check that @bind is still the current peer bind and that @key still matches + * the peer route key. The caller must hold @peer->lock. The TX path keeps + * @bind inside an RCU read-side critical section, so pointer identity is enough + * to detect whether the bind was replaced while the route lookup was running. + * + * Return: true if the lookup result still matches the current peer state and + * may update the dst cache. + */ +static bool ovpn_dst_cache_current(const struct ovpn_peer *peer, + const struct ovpn_bind *bind, + const struct ovpn_route_key *key) +{ + const struct ovpn_bind *curr_bind; + + lockdep_assert_held(&peer->lock); + + curr_bind = rcu_dereference_protected(peer->bind, + lockdep_is_held(&peer->lock)); + + return curr_bind == bind && + ovpn_route_key_equal(key, &peer->route_key); +} + /** * ovpn_udp4_output - send IPv4 packet over udp socket * @peer: the destination peer @@ -240,7 +269,7 @@ static int ovpn_udp4_output(struct ovpn_peer *peer, struct ovpn_bind *bind, /* avoid storing a stale cache */ spin_lock_bh(&peer->lock); - if (likely(ovpn_route_key_equal(key, &peer->route_key))) + if (likely(ovpn_dst_cache_current(peer, bind, key))) dst_cache_set_ip4(cache, &rt->dst, fl.saddr); spin_unlock_bh(&peer->lock); @@ -313,7 +342,7 @@ static int ovpn_udp6_output(struct ovpn_peer *peer, struct ovpn_bind *bind, /* avoid storing a stale cache */ spin_lock_bh(&peer->lock); - if (likely(ovpn_route_key_equal(key, &peer->route_key))) + if (likely(ovpn_dst_cache_current(peer, bind, key))) dst_cache_set_ip6(cache, dst, &fl.saddr); spin_unlock_bh(&peer->lock); From aea934a221ec6a867221e5b765f65f1857befd53 Mon Sep 17 00:00:00 2001 From: Ralf Lici Date: Fri, 28 Aug 2026 16:50:26 +0200 Subject: [PATCH 057/189] ovpn: replace bind when learning local endpoint struct ovpn_bind is published through peer->bind with RCU, but local endpoint learning updates bind->local in place under peer->lock. UDP TX reads the field without that lock. In particular, a concurrent IPv6 update can therefore result in a torn address read. Use ovpn_peer_reset_sockaddr to publish a replacement bind when learning a new local endpoint, just as a remote endpoint change does. Preserve the current remote address and reset the dst cache only after the new bind has been published successfully. Track remote endpoint changes separately so that float notification and transport-address rehashing remain limited to actual peer floats. Fixes: f0281c1d3732 ("ovpn: add support for updating local or remote UDP endpoint") Signed-off-by: Ralf Lici Signed-off-by: Antonio Quartulli --- drivers/net/ovpn/peer.c | 43 ++++++++++++++++++++++------------------- 1 file changed, 23 insertions(+), 20 deletions(-) diff --git a/drivers/net/ovpn/peer.c b/drivers/net/ovpn/peer.c index b400783c2efa..430c6cd48db8 100644 --- a/drivers/net/ovpn/peer.c +++ b/drivers/net/ovpn/peer.c @@ -200,13 +200,12 @@ static void __ovpn_peer_hash_transp_addr(struct ovpn_peer *peer, */ void ovpn_peer_endpoints_update(struct ovpn_peer *peer, struct sk_buff *skb) { + const void *local_ip = NULL; struct sockaddr_storage ss; struct sockaddr_in6 *sa6; - bool reset_cache = false; struct sockaddr_in *sa; struct ovpn_bind *bind; - const void *local_ip; - size_t salen = 0; + bool floated = false; spin_lock_bh(&peer->lock); bind = rcu_dereference_protected(peer->bind, @@ -233,8 +232,7 @@ void ovpn_peer_endpoints_update(struct ovpn_peer *peer, struct sk_buff *skb) .sin_addr.s_addr = ip_hdr(skb)->saddr, .sin_port = udp_hdr(skb)->source, }; - salen = sizeof(*sa); - reset_cache = true; + floated = true; break; } @@ -246,10 +244,12 @@ void ovpn_peer_endpoints_update(struct ovpn_peer *peer, struct sk_buff *skb) netdev_name(peer->ovpn->dev), peer->id, &bind->local.ipv4.s_addr, &ip_hdr(skb)->daddr); - bind->local.ipv4.s_addr = ip_hdr(skb)->daddr; - reset_cache = true; + local_ip = &ip_hdr(skb)->daddr; + memcpy(&ss, &bind->remote, sizeof(struct sockaddr_in)); + break; } - break; + /* nothing changed */ + goto unlock; case htons(ETH_P_IPV6): /* float check */ if (unlikely(!ovpn_bind_skb_src_match(bind, skb))) { @@ -271,8 +271,7 @@ void ovpn_peer_endpoints_update(struct ovpn_peer *peer, struct sk_buff *skb) ipv6_iface_scope_id(&ipv6_hdr(skb)->saddr, skb->skb_iif), }; - salen = sizeof(*sa6); - reset_cache = true; + floated = true; break; } @@ -285,26 +284,30 @@ void ovpn_peer_endpoints_update(struct ovpn_peer *peer, struct sk_buff *skb) netdev_name(peer->ovpn->dev), peer->id, &bind->local.ipv6, &ipv6_hdr(skb)->daddr); - bind->local.ipv6 = ipv6_hdr(skb)->daddr; - reset_cache = true; + local_ip = &ipv6_hdr(skb)->daddr; + memcpy(&ss, &bind->remote, sizeof(struct sockaddr_in6)); + break; } - break; + /* nothing changed */ + goto unlock; default: goto unlock; } - if (unlikely(reset_cache)) - dst_cache_reset(&peer->dst_cache); - - /* if the peer did not float, we can bail out now */ - if (likely(!salen)) - goto unlock; - if (unlikely(ovpn_peer_reset_sockaddr(peer, (struct sockaddr_storage *)&ss, local_ip) < 0)) goto unlock; + /* reset the cache only after a successful bind update to avoid useless + * cache misses on concurrent TX + */ + dst_cache_reset(&peer->dst_cache); + + /* if only the local address changed, bail out now */ + if (!floated) + goto unlock; + net_dbg_ratelimited("%s: peer %d floated to %pIScp", netdev_name(peer->ovpn->dev), peer->id, &ss); From 7d8104988f423572df1f3347ce578037b1043f34 Mon Sep 17 00:00:00 2001 From: Ralf Lici Date: Fri, 28 Aug 2026 16:50:27 +0200 Subject: [PATCH 058/189] ovpn: replace bind when clearing stale local source The UDP output fallback clears bind->local in place when the remembered source address is no longer usable. The bind is RCU-published and read locklessly by concurrent TX, so an IPv6 reader can observe a torn address. Retry the route lookup with source address autoselection without modifying the bind. After a successful lookup, revalidate the bind and route key under peer->lock, reset the dst cache, and best-effort publish a replacement bind with a wildcard local address. Do not cache the resolved dst when clearing the local source. Replacing the source invalidates all per-CPU cache entries, while dst_cache_set_ip4 and dst_cache_set_ip6 update only the current CPU slot. The current packet can still use the resolved route; if bind allocation fails, a later cache miss retries the repair. Fixes: 08857b5ec5d9 ("ovpn: implement basic TX path (UDP)") Signed-off-by: Ralf Lici Signed-off-by: Antonio Quartulli --- drivers/net/ovpn/udp.c | 81 +++++++++++++++++++++++++++++------------- 1 file changed, 56 insertions(+), 25 deletions(-) diff --git a/drivers/net/ovpn/udp.c b/drivers/net/ovpn/udp.c index eeef4a7229f5..055cdb1bee13 100644 --- a/drivers/net/ovpn/udp.c +++ b/drivers/net/ovpn/udp.c @@ -185,7 +185,7 @@ static void ovpn_dst_cache_check_key(struct ovpn_peer *peer, * to detect whether the bind was replaced while the route lookup was running. * * Return: true if the lookup result still matches the current peer state and - * may update the dst cache. + * may update the dst cache or replace the bind. */ static bool ovpn_dst_cache_current(const struct ovpn_peer *peer, const struct ovpn_bind *bind, @@ -218,6 +218,9 @@ static int ovpn_udp4_output(struct ovpn_peer *peer, struct ovpn_bind *bind, struct sk_buff *skb, const struct ovpn_route_key *key) { + struct sockaddr_storage remote; + struct in_addr local = {}; + bool reset_local = false; struct rtable *rt; struct flowi4 fl = { .saddr = bind->local.ipv4.s_addr, @@ -236,24 +239,17 @@ static int ovpn_udp4_output(struct ovpn_peer *peer, struct ovpn_bind *bind, if (fl.saddr && unlikely(!inet_confirm_addr(sock_net(sk), NULL, 0, fl.saddr, RT_SCOPE_HOST))) { - /* we may end up here when the cached address is not usable - * anymore. In this case we reset address/cache and perform a - * new look up + /* The learned local address is not usable anymore. + * Retry with source address autoselection. */ fl.saddr = 0; - spin_lock_bh(&peer->lock); - bind->local.ipv4.s_addr = 0; - spin_unlock_bh(&peer->lock); - dst_cache_reset(cache); + reset_local = true; } rt = ip_route_output_flow(sock_net(sk), &fl, sk); if (IS_ERR(rt) && PTR_ERR(rt) == -EINVAL) { fl.saddr = 0; - spin_lock_bh(&peer->lock); - bind->local.ipv4.s_addr = 0; - spin_unlock_bh(&peer->lock); - dst_cache_reset(cache); + reset_local = true; rt = ip_route_output_flow(sock_net(sk), &fl, sk); } @@ -267,10 +263,28 @@ static int ovpn_udp4_output(struct ovpn_peer *peer, struct ovpn_bind *bind, goto err; } - /* avoid storing a stale cache */ + /* avoid storing a stale cache or local address */ spin_lock_bh(&peer->lock); - if (likely(ovpn_dst_cache_current(peer, bind, key))) - dst_cache_set_ip4(cache, &rt->dst, fl.saddr); + if (likely(ovpn_dst_cache_current(peer, bind, key))) { + if (!reset_local) { + dst_cache_set_ip4(cache, &rt->dst, fl.saddr); + spin_unlock_bh(&peer->lock); + goto transmit; + } + + /* invalidate per-CPU dst entries that may still carry + * the stale source + */ + dst_cache_reset(cache); + + /* preserve the current remote */ + memcpy(&remote, &bind->remote, sizeof(struct sockaddr_in)); + /* The current packet already has a valid wildcard-source route. + * If replacing the bind fails, leave the stale local in place; + * a later cache miss will retry the repair. + */ + ovpn_peer_reset_sockaddr(peer, &remote, &local); + } spin_unlock_bh(&peer->lock); transmit: @@ -300,6 +314,9 @@ static int ovpn_udp6_output(struct ovpn_peer *peer, struct ovpn_bind *bind, struct sk_buff *skb, const struct ovpn_route_key *key) { + struct in6_addr local = in6addr_any; + struct sockaddr_storage remote; + bool reset_local = false; struct dst_entry *dst; int ret; @@ -320,15 +337,11 @@ static int ovpn_udp6_output(struct ovpn_peer *peer, struct ovpn_bind *bind, if (!ipv6_addr_any(&fl.saddr) && unlikely(!ipv6_chk_addr(sock_net(sk), &fl.saddr, NULL, 0))) { - /* we may end up here when the cached address is not usable - * anymore. In this case we reset address/cache and perform a - * new look up + /* The learned local address is not usable anymore. + * Retry with source address autoselection. */ fl.saddr = in6addr_any; - spin_lock_bh(&peer->lock); - bind->local.ipv6 = in6addr_any; - spin_unlock_bh(&peer->lock); - dst_cache_reset(cache); + reset_local = true; } dst = ip6_dst_lookup_flow(sock_net(sk), sk, &fl, NULL); @@ -340,10 +353,28 @@ static int ovpn_udp6_output(struct ovpn_peer *peer, struct ovpn_bind *bind, goto err; } - /* avoid storing a stale cache */ + /* avoid storing a stale cache or local address */ spin_lock_bh(&peer->lock); - if (likely(ovpn_dst_cache_current(peer, bind, key))) - dst_cache_set_ip6(cache, dst, &fl.saddr); + if (likely(ovpn_dst_cache_current(peer, bind, key))) { + if (!reset_local) { + dst_cache_set_ip6(cache, dst, &fl.saddr); + spin_unlock_bh(&peer->lock); + goto transmit; + } + + /* invalidate per-CPU dst entries that may still carry + * the stale source + */ + dst_cache_reset(cache); + + /* preserve the current remote */ + memcpy(&remote, &bind->remote, sizeof(struct sockaddr_in6)); + /* The current packet already has a valid wildcard-source route. + * If replacing the bind fails, leave the stale local in place; + * a later cache miss will retry the repair. + */ + ovpn_peer_reset_sockaddr(peer, &remote, &local); + } spin_unlock_bh(&peer->lock); transmit: From b43beccb3713fafada57814b0a652f4a876eb75f Mon Sep 17 00:00:00 2001 From: Ralf Lici Date: Fri, 28 Aug 2026 15:00:06 +0200 Subject: [PATCH 059/189] ovpn: always unhash old VPN addresses before rehashing ovpn_peer_hash_vpn_ip updates the per-peer VPN address hash entries after userspace changes a peer VPN address. The current code removes an old hash entry only when the new address for that family is not the unspecified address. When an address is cleared to 0.0.0.0 or ::, its hash node therefore remains linked in the bucket selected by the old address. The address comparison performed during lookup prevents the old address from matching, but the table retains a stale entry until the peer is removed or another address is configured for that family. Always remove both old VPN address hash entries before conditionally adding the currently configured addresses back. This ensures that a cleared address leaves its hash node unhashed. Fixes: 1d36a36f6d53 ("ovpn: implement peer add/get/dump/delete via netlink") Signed-off-by: Ralf Lici Signed-off-by: Antonio Quartulli --- drivers/net/ovpn/peer.c | 10 ++++------ 1 file changed, 4 insertions(+), 6 deletions(-) diff --git a/drivers/net/ovpn/peer.c b/drivers/net/ovpn/peer.c index 430c6cd48db8..bbd9e17fb0bf 100644 --- a/drivers/net/ovpn/peer.c +++ b/drivers/net/ovpn/peer.c @@ -994,10 +994,11 @@ void ovpn_peer_hash_vpn_ip(struct ovpn_peer *peer) if (hlist_unhashed(&peer->hash_entry_id)) return; - if (peer->vpn_addrs.ipv4.s_addr != htonl(INADDR_ANY)) { - /* remove potential old hashing */ - hlist_nulls_del_init_rcu(&peer->hash_entry_addr4); + /* remove potential old hashing */ + hlist_nulls_del_init_rcu(&peer->hash_entry_addr4); + hlist_nulls_del_init_rcu(&peer->hash_entry_addr6); + if (peer->vpn_addrs.ipv4.s_addr != htonl(INADDR_ANY)) { nhead = ovpn_get_hash_head(peer->ovpn->peers->by_vpn_addr4, &peer->vpn_addrs.ipv4, sizeof(peer->vpn_addrs.ipv4)); @@ -1005,9 +1006,6 @@ void ovpn_peer_hash_vpn_ip(struct ovpn_peer *peer) } if (!ipv6_addr_any(&peer->vpn_addrs.ipv6)) { - /* remove potential old hashing */ - hlist_nulls_del_init_rcu(&peer->hash_entry_addr6); - nhead = ovpn_get_hash_head(peer->ovpn->peers->by_vpn_addr6, &peer->vpn_addrs.ipv6, sizeof(peer->vpn_addrs.ipv6)); From d25e885b31a0f2808d936f95c9a558a8a792669b Mon Sep 17 00:00:00 2001 From: Ralf Lici Date: Fri, 28 Aug 2026 15:00:07 +0200 Subject: [PATCH 060/189] ovpn: reject duplicate peer VPN addresses In MP mode, ovpn uses the peer VPN addresses as lookup keys for selecting the peer that should receive an outgoing tunnel packet. However, the netlink peer configuration path does not currently reject duplicate VPN addresses. If two peers are configured with the same VPN address, both can be inserted in the VPN address hash table and lookups return whichever peer is found first. This makes peer selection ambiguous and dependent on hash insertion order. Reject peer creation or update when the resulting VPN address is already assigned to another peer. Ignore unspecified addresses because those are not inserted in the VPN address hash tables. This changes such configurations from being accepted to being rejected, but they have never worked reliably because peer selection is ambiguous. Fixes: 1d36a36f6d53 ("ovpn: implement peer add/get/dump/delete via netlink") Signed-off-by: Ralf Lici Signed-off-by: Antonio Quartulli --- drivers/net/ovpn/netlink.c | 37 ++++++++++++++++----- drivers/net/ovpn/peer.c | 67 +++++++++++++++++++++++++++++++++++++- drivers/net/ovpn/peer.h | 6 ++++ 3 files changed, 101 insertions(+), 9 deletions(-) diff --git a/drivers/net/ovpn/netlink.c b/drivers/net/ovpn/netlink.c index 2ba762082acc..e23f7d1f49e0 100644 --- a/drivers/net/ovpn/netlink.c +++ b/drivers/net/ovpn/netlink.c @@ -480,8 +480,10 @@ int ovpn_nl_peer_new_doit(struct sk_buff *skb, struct genl_info *info) int ovpn_nl_peer_set_doit(struct sk_buff *skb, struct genl_info *info) { - struct nlattr *attrs[OVPN_A_PEER_MAX + 1]; struct ovpn_priv *ovpn = info->user_ptr[0]; + struct nlattr *attrs[OVPN_A_PEER_MAX + 1]; + struct in6_addr vpn_addr6; + struct in_addr vpn_addr4; struct ovpn_socket *sock; struct ovpn_peer *peer; u32 peer_id; @@ -528,28 +530,47 @@ int ovpn_nl_peer_set_doit(struct sk_buff *skb, struct genl_info *info) rcu_read_unlock(); spin_lock_bh(&ovpn->lock); - ret = ovpn_nl_peer_modify(peer, info, attrs); - if (ret < 0) { - spin_unlock_bh(&ovpn->lock); - ovpn_peer_put(peer); - return ret; + + /* reject peer with conflicting VPN address */ + if (attrs[OVPN_A_PEER_VPN_IPV4]) { + vpn_addr4.s_addr = nla_get_in_addr(attrs[OVPN_A_PEER_VPN_IPV4]); + if (ovpn_peer_vpn_addr_conflict4(ovpn, peer, &vpn_addr4)) + goto addr_conflict; } + if (attrs[OVPN_A_PEER_VPN_IPV6]) { + vpn_addr6 = nla_get_in6_addr(attrs[OVPN_A_PEER_VPN_IPV6]); + if (ovpn_peer_vpn_addr_conflict6(ovpn, peer, &vpn_addr6)) + goto addr_conflict; + } + + ret = ovpn_nl_peer_modify(peer, info, attrs); + if (ret < 0) + goto unlock; /* ret == 1 means that VPN IPv4/6 has been modified and rehashing * is required */ - if (ret > 0) + if (ret > 0) { ovpn_peer_hash_vpn_ip(peer); + ret = 0; + } /* if the remote endpoint was updated, the by_transp_addr hash bucket * also needs to be refreshed, otherwise incoming packets from the new * remote address would fail the lockless lookup */ if (attrs[OVPN_A_PEER_REMOTE_IPV4] || attrs[OVPN_A_PEER_REMOTE_IPV6]) ovpn_peer_hash_transp_addr(peer); + +unlock: spin_unlock_bh(&ovpn->lock); ovpn_peer_put(peer); - return 0; + return ret; +addr_conflict: + NL_SET_ERR_MSG_FMT_MOD(info->extack, + "VPN IP is already assigned to another peer"); + ret = -EADDRINUSE; + goto unlock; } static int ovpn_nl_send_peer(struct sk_buff *skb, const struct genl_info *info, diff --git a/drivers/net/ovpn/peer.c b/drivers/net/ovpn/peer.c index bbd9e17fb0bf..2067825bb5b6 100644 --- a/drivers/net/ovpn/peer.c +++ b/drivers/net/ovpn/peer.c @@ -488,7 +488,7 @@ static struct ovpn_peer *ovpn_peer_get_by_vpn_addr4(struct ovpn_priv *ovpn, * Return: the peer if found or NULL otherwise */ static struct ovpn_peer *ovpn_peer_get_by_vpn_addr6(struct ovpn_priv *ovpn, - struct in6_addr *addr) + const struct in6_addr *addr) { struct hlist_nulls_head *nhead; struct hlist_nulls_node *ntmp; @@ -513,6 +513,64 @@ static struct ovpn_peer *ovpn_peer_get_by_vpn_addr6(struct ovpn_priv *ovpn, return NULL; } +/** + * ovpn_peer_vpn_addr_conflict4 - check if the VPN v4 address is already in use + * @ovpn: the openvpn instance to search + * @peer: peer being added or updated, or NULL + * @addr: VPN IPv4 address to check + * + * Check whether @addr is already assigned to another peer. @peer is ignored + * when found, allowing peer updates that keep an existing address. + * Unspecified addresses are ignored. + * + * Note: the caller must hold @ovpn->lock. + * + * Return: true on conflict, false otherwise. + */ +bool ovpn_peer_vpn_addr_conflict4(struct ovpn_priv *ovpn, + const struct ovpn_peer *peer, + const struct in_addr *addr) +{ + struct ovpn_peer *tmp = NULL; + + lockdep_assert_held(&ovpn->lock); + + /* we don't hash INADDR_ANY, no conflict in that case */ + if (addr->s_addr != htonl(INADDR_ANY)) + tmp = ovpn_peer_get_by_vpn_addr4(ovpn, addr->s_addr); + + return tmp && tmp != peer; +} + +/** + * ovpn_peer_vpn_addr_conflict6 - check if the VPN v6 address is already in use + * @ovpn: the openvpn instance to search + * @peer: peer being added or updated, or NULL + * @addr: VPN IPv6 address to check + * + * Check whether @addr is already assigned to another peer. @peer is ignored + * when found, allowing peer updates that keep an existing address. + * Unspecified addresses are ignored. + * + * Note: the caller must hold @ovpn->lock. + * + * Return: true on conflict, false otherwise. + */ +bool ovpn_peer_vpn_addr_conflict6(struct ovpn_priv *ovpn, + const struct ovpn_peer *peer, + const struct in6_addr *addr) +{ + struct ovpn_peer *tmp = NULL; + + lockdep_assert_held(&ovpn->lock); + + /* we don't hash ::, no conflict in that case */ + if (!ipv6_addr_any(addr)) + tmp = ovpn_peer_get_by_vpn_addr6(ovpn, addr); + + return tmp && tmp != peer; +} + /** * ovpn_peer_transp_match - check if sockaddr and peer binding match * @peer: the peer to get the binding from @@ -1040,6 +1098,13 @@ static int ovpn_peer_add_mp(struct ovpn_priv *ovpn, struct ovpn_peer *peer) goto out; } + /* reject peer with conflicting VPN address */ + if (ovpn_peer_vpn_addr_conflict4(ovpn, NULL, &peer->vpn_addrs.ipv4) || + ovpn_peer_vpn_addr_conflict6(ovpn, NULL, &peer->vpn_addrs.ipv6)) { + ret = -EADDRINUSE; + goto out; + } + bind = rcu_dereference_protected(peer->bind, true); /* peers connected via TCP have bind == NULL */ if (bind) { diff --git a/drivers/net/ovpn/peer.h b/drivers/net/ovpn/peer.h index 063535699ecd..1879bfb76992 100644 --- a/drivers/net/ovpn/peer.h +++ b/drivers/net/ovpn/peer.h @@ -164,6 +164,12 @@ struct ovpn_peer *ovpn_peer_get_by_transp_addr(struct ovpn_priv *ovpn, struct ovpn_peer *ovpn_peer_get_by_id(struct ovpn_priv *ovpn, u32 peer_id); struct ovpn_peer *ovpn_peer_get_by_dst(struct ovpn_priv *ovpn, struct sk_buff *skb); +bool ovpn_peer_vpn_addr_conflict4(struct ovpn_priv *ovpn, + const struct ovpn_peer *peer, + const struct in_addr *addr); +bool ovpn_peer_vpn_addr_conflict6(struct ovpn_priv *ovpn, + const struct ovpn_peer *peer, + const struct in6_addr *addr); void ovpn_peer_hash_vpn_ip(struct ovpn_peer *peer); void ovpn_peer_hash_transp_addr(struct ovpn_peer *peer); bool ovpn_peer_check_by_src(struct ovpn_priv *ovpn, struct sk_buff *skb, From 025af3a0a892514f9f27f186338ba3d44365547a Mon Sep 17 00:00:00 2001 From: Ralf Lici Date: Fri, 28 Aug 2026 15:00:08 +0200 Subject: [PATCH 061/189] ovpn: reject multipeer peers without VPN addresses In MP mode, ovpn uses the peer VPN addresses to select the peer for outgoing tunnel packets. Peer creation currently requires a VPN IPv4 or IPv6 attribute, but it only checks for the presence of the attribute and not for a usable address value. This allows userspace to create an MP peer with only unspecified VPN addresses, or to update an existing peer so that both VPN address families become unspecified. Such a peer cannot be selected through the VPN address hash tables. Reject MP peer creation or update when the resulting peer would not have at least one VPN address configured. This changes such configurations from being accepted to being rejected, but they have never been usable because the peer cannot be selected through the VPN address hash tables. Fixes: 1d36a36f6d53 ("ovpn: implement peer add/get/dump/delete via netlink") Signed-off-by: Ralf Lici Signed-off-by: Antonio Quartulli --- drivers/net/ovpn/netlink.c | 36 ++++++++++++++++++++++++++++++------ 1 file changed, 30 insertions(+), 6 deletions(-) diff --git a/drivers/net/ovpn/netlink.c b/drivers/net/ovpn/netlink.c index e23f7d1f49e0..e9e0f75e0443 100644 --- a/drivers/net/ovpn/netlink.c +++ b/drivers/net/ovpn/netlink.c @@ -352,8 +352,10 @@ static int ovpn_nl_peer_modify(struct ovpn_peer *peer, struct genl_info *info, int ovpn_nl_peer_new_doit(struct sk_buff *skb, struct genl_info *info) { - struct nlattr *attrs[OVPN_A_PEER_MAX + 1]; + struct in_addr vpn_addr4 = { .s_addr = htonl(INADDR_ANY) }; + struct in6_addr vpn_addr6 = IN6ADDR_ANY_INIT; struct ovpn_priv *ovpn = info->user_ptr[0]; + struct nlattr *attrs[OVPN_A_PEER_MAX + 1]; struct ovpn_socket *ovpn_sock; struct socket *sock = NULL; struct ovpn_peer *peer; @@ -377,11 +379,20 @@ int ovpn_nl_peer_new_doit(struct sk_buff *skb, struct genl_info *info) return -EINVAL; /* in MP mode VPN IPs are required for selecting the right peer */ - if (ovpn->mode == OVPN_MODE_MP && !attrs[OVPN_A_PEER_VPN_IPV4] && - !attrs[OVPN_A_PEER_VPN_IPV6]) { - NL_SET_ERR_MSG_FMT_MOD(info->extack, - "VPN IP must be provided in MP mode"); - return -EINVAL; + if (ovpn->mode == OVPN_MODE_MP) { + if (attrs[OVPN_A_PEER_VPN_IPV4]) + vpn_addr4.s_addr = + nla_get_in_addr(attrs[OVPN_A_PEER_VPN_IPV4]); + if (attrs[OVPN_A_PEER_VPN_IPV6]) + vpn_addr6 = + nla_get_in6_addr(attrs[OVPN_A_PEER_VPN_IPV6]); + + if (vpn_addr4.s_addr == htonl(INADDR_ANY) && + ipv6_addr_any(&vpn_addr6)) { + NL_SET_ERR_MSG_FMT_MOD(info->extack, + "at least one VPN IP must be configured in MP mode"); + return -EINVAL; + } } peer_id = nla_get_u32(attrs[OVPN_A_PEER_ID]); @@ -531,6 +542,9 @@ int ovpn_nl_peer_set_doit(struct sk_buff *skb, struct genl_info *info) spin_lock_bh(&ovpn->lock); + vpn_addr4 = peer->vpn_addrs.ipv4; + vpn_addr6 = peer->vpn_addrs.ipv6; + /* reject peer with conflicting VPN address */ if (attrs[OVPN_A_PEER_VPN_IPV4]) { vpn_addr4.s_addr = nla_get_in_addr(attrs[OVPN_A_PEER_VPN_IPV4]); @@ -543,6 +557,16 @@ int ovpn_nl_peer_set_doit(struct sk_buff *skb, struct genl_info *info) goto addr_conflict; } + /* in MP mode VPN IPs are required for selecting the right peer */ + if (ovpn->mode == OVPN_MODE_MP && + vpn_addr4.s_addr == htonl(INADDR_ANY) && + ipv6_addr_any(&vpn_addr6)) { + NL_SET_ERR_MSG_FMT_MOD(info->extack, + "at least one VPN IP must be configured in MP mode"); + ret = -EINVAL; + goto unlock; + } + ret = ovpn_nl_peer_modify(peer, info, attrs); if (ret < 0) goto unlock; From 5940f3407b78062442cb01f541ef6eed709fc380 Mon Sep 17 00:00:00 2001 From: Ralf Lici Date: Fri, 28 Aug 2026 15:00:09 +0200 Subject: [PATCH 062/189] ovpn: reject invalid peer VPN addresses In MP mode, ovpn uses peer VPN addresses as lookup keys for selecting the peer that should receive outgoing tunnel packets. The netlink configuration path currently accepts address values that cannot sensibly identify a VPN peer, such as multicast, broadcast or loopback addresses. Reject invalid peer VPN addresses when creating or updating an MP peer. Keep accepting the unspecified address as the internal unset value, provided that at least one VPN address family remains configured. Fixes: 1d36a36f6d53 ("ovpn: implement peer add/get/dump/delete via netlink") Signed-off-by: Ralf Lici Signed-off-by: Antonio Quartulli --- drivers/net/ovpn/netlink.c | 55 +++++++++++++++++++++++++++++--------- 1 file changed, 42 insertions(+), 13 deletions(-) diff --git a/drivers/net/ovpn/netlink.c b/drivers/net/ovpn/netlink.c index e9e0f75e0443..5432bc2eb8e8 100644 --- a/drivers/net/ovpn/netlink.c +++ b/drivers/net/ovpn/netlink.c @@ -185,6 +185,39 @@ static sa_family_t ovpn_nl_family_get(struct nlattr *addr4, return AF_UNSPEC; } +static int ovpn_nl_peer_check_vpn_addrs(const struct in_addr *addr4, + const struct in6_addr *addr6, + struct genl_info *info) +{ + int addr6_type; + + if (addr4->s_addr == htonl(INADDR_ANY) && ipv6_addr_any(addr6)) { + NL_SET_ERR_MSG_MOD(info->extack, + "at least one VPN IP must be configured in MP mode"); + return -EINVAL; + } + + if (ipv4_is_multicast(addr4->s_addr) || ipv4_is_lbcast(addr4->s_addr) || + ipv4_is_loopback(addr4->s_addr)) { + NL_SET_ERR_MSG_MOD(info->extack, + "VPN IPv4 address must be valid unicast or any"); + return -EADDRNOTAVAIL; + } + + if (!ipv6_addr_any(addr6)) { + addr6_type = ipv6_addr_type(addr6); + + if (!(addr6_type & IPV6_ADDR_UNICAST) || + (addr6_type & (IPV6_ADDR_LOOPBACK | IPV6_ADDR_COMPATv4))) { + NL_SET_ERR_MSG_MOD(info->extack, + "VPN IPv6 address must be valid unicast or any"); + return -EADDRNOTAVAIL; + } + } + + return 0; +} + static int ovpn_nl_peer_precheck(struct ovpn_priv *ovpn, struct genl_info *info, struct nlattr **attrs) @@ -387,12 +420,10 @@ int ovpn_nl_peer_new_doit(struct sk_buff *skb, struct genl_info *info) vpn_addr6 = nla_get_in6_addr(attrs[OVPN_A_PEER_VPN_IPV6]); - if (vpn_addr4.s_addr == htonl(INADDR_ANY) && - ipv6_addr_any(&vpn_addr6)) { - NL_SET_ERR_MSG_FMT_MOD(info->extack, - "at least one VPN IP must be configured in MP mode"); - return -EINVAL; - } + ret = ovpn_nl_peer_check_vpn_addrs(&vpn_addr4, &vpn_addr6, + info); + if (ret < 0) + return ret; } peer_id = nla_get_u32(attrs[OVPN_A_PEER_ID]); @@ -558,13 +589,11 @@ int ovpn_nl_peer_set_doit(struct sk_buff *skb, struct genl_info *info) } /* in MP mode VPN IPs are required for selecting the right peer */ - if (ovpn->mode == OVPN_MODE_MP && - vpn_addr4.s_addr == htonl(INADDR_ANY) && - ipv6_addr_any(&vpn_addr6)) { - NL_SET_ERR_MSG_FMT_MOD(info->extack, - "at least one VPN IP must be configured in MP mode"); - ret = -EINVAL; - goto unlock; + if (ovpn->mode == OVPN_MODE_MP) { + ret = ovpn_nl_peer_check_vpn_addrs(&vpn_addr4, &vpn_addr6, + info); + if (ret < 0) + goto unlock; } ret = ovpn_nl_peer_modify(peer, info, attrs); From 006208026819d5e9e5ec07b3e73d95960a059327 Mon Sep 17 00:00:00 2001 From: Ralf Lici Date: Fri, 28 Aug 2026 15:00:10 +0200 Subject: [PATCH 063/189] selftests: ovpn: validate peer VPN addresses Exercise peer VPN address validation through both peer creation and update. Check missing, unspecified, duplicate, multicast, broadcast, loopback, IPv4-compatible and IPv4-mapped addresses. Temporarily configure a peer with both address families to verify that either family can be cleared while the other remains configured, then restore the original addresses before running the existing traffic tests. Extend ovpn-cli peer updates with an optional VPN address and preserve peer creation errors so the negative tests can observe rejected netlink requests. Signed-off-by: Ralf Lici Signed-off-by: Antonio Quartulli --- tools/testing/selftests/net/ovpn/common.sh | 13 ++++ tools/testing/selftests/net/ovpn/ovpn-cli.c | 54 ++++++++++----- tools/testing/selftests/net/ovpn/test.sh | 75 ++++++++++++++++++++- 3 files changed, 123 insertions(+), 19 deletions(-) diff --git a/tools/testing/selftests/net/ovpn/common.sh b/tools/testing/selftests/net/ovpn/common.sh index 2d844eb3aa6e..5e9c81e885e6 100644 --- a/tools/testing/selftests/net/ovpn/common.sh +++ b/tools/testing/selftests/net/ovpn/common.sh @@ -136,6 +136,19 @@ ovpn_create_ns() { ip netns add "ovpn_peer${1}" } +ovpn_peer_vpn_addr() { + local peer="$1" + local file + + if [ "${OVPN_PROTO}" == "UDP" ]; then + file="${OVPN_UDP_PEERS_FILE}" + else + file="${OVPN_TCP_PEERS_FILE}" + fi + + awk -v peer="${peer}" '$1 == peer {print $NF; exit}' "${file}" +} + ovpn_setup_ns() { local peer="ovpn_peer${1}" local server_ns="ovpn_peer0" diff --git a/tools/testing/selftests/net/ovpn/ovpn-cli.c b/tools/testing/selftests/net/ovpn/ovpn-cli.c index f4effa7580c0..3b612a8a18fe 100644 --- a/tools/testing/selftests/net/ovpn/ovpn-cli.c +++ b/tools/testing/selftests/net/ovpn/ovpn-cli.c @@ -650,6 +650,26 @@ static int ovpn_connect(struct ovpn_ctx *ovpn) return ret; } +static int ovpn_nl_put_vpn_addr(struct nl_msg *msg, + const struct ovpn_ctx *ovpn) +{ + if (!ovpn->peer_ip_set) + return 0; + + switch (ovpn->peer_ip.in4.sin_family) { + case AF_INET: + return nla_put_u32(msg, OVPN_A_PEER_VPN_IPV4, + ovpn->peer_ip.in4.sin_addr.s_addr); + case AF_INET6: + return nla_put(msg, OVPN_A_PEER_VPN_IPV6, + sizeof(struct in6_addr), + &ovpn->peer_ip.in6.sin6_addr); + default: + fprintf(stderr, "Invalid family for peer address\n"); + return -EAFNOSUPPORT; + } +} + static int ovpn_new_peer(struct ovpn_ctx *ovpn, bool is_tcp) { struct nlattr *attr; @@ -691,22 +711,9 @@ static int ovpn_new_peer(struct ovpn_ctx *ovpn, bool is_tcp) } } - if (ovpn->peer_ip_set) { - switch (ovpn->peer_ip.in4.sin_family) { - case AF_INET: - NLA_PUT_U32(ctx->nl_msg, OVPN_A_PEER_VPN_IPV4, - ovpn->peer_ip.in4.sin_addr.s_addr); - break; - case AF_INET6: - NLA_PUT(ctx->nl_msg, OVPN_A_PEER_VPN_IPV6, - sizeof(struct in6_addr), - &ovpn->peer_ip.in6.sin6_addr); - break; - default: - fprintf(stderr, "Invalid family for peer address\n"); - goto nla_put_failure; - } - } + ret = ovpn_nl_put_vpn_addr(ctx->nl_msg, ovpn); + if (ret) + goto nla_put_failure; nla_nest_end(ctx->nl_msg, attr); @@ -732,6 +739,10 @@ static int ovpn_set_peer(struct ovpn_ctx *ovpn) ovpn->keepalive_interval); NLA_PUT_U32(ctx->nl_msg, OVPN_A_PEER_KEEPALIVE_TIMEOUT, ovpn->keepalive_timeout); + + ret = ovpn_nl_put_vpn_addr(ctx->nl_msg, ovpn); + if (ret) + goto nla_put_failure; nla_nest_end(ctx->nl_msg, attr); ret = ovpn_nl_msg_send(ctx, NULL); @@ -1730,13 +1741,14 @@ static void usage(const char *cmd) fprintf(stderr, "\tmark: socket FW mark value\n"); fprintf(stderr, - "* set_peer : set peer attributes\n"); + "* set_peer [vpnaddr]: set peer attributes\n"); fprintf(stderr, "\tiface: ovpn interface name\n"); fprintf(stderr, "\tpeer_id: peer ID of the peer to modify\n"); fprintf(stderr, "\tkeepalive_interval: interval for sending ping messages\n"); fprintf(stderr, "\tkeepalive_timeout: time after which a peer is timed out\n"); + fprintf(stderr, "\tvpnaddr: peer VPN IP\n"); fprintf(stderr, "* del_peer : delete peer\n"); fprintf(stderr, "\tiface: ovpn interface name\n"); @@ -2090,6 +2102,8 @@ static int ovpn_run_cmd(struct ovpn_ctx *ovpn) return ret; ret = ovpn_new_peer(ovpn, false); + if (ret < 0) + return ret; ovpn_waitbg(); break; case CMD_NEW_MULTI_PEER: @@ -2331,6 +2345,12 @@ static int ovpn_parse_cmd_args(struct ovpn_ctx *ovpn, int argc, char *argv[]) "keepalive interval value out of range\n"); return -1; } + + if (argc > 6) { + ret = ovpn_parse_remote(ovpn, NULL, NULL, argv[6]); + if (ret < 0) + return -1; + } break; case CMD_DEL_PEER: if (argc < 4) diff --git a/tools/testing/selftests/net/ovpn/test.sh b/tools/testing/selftests/net/ovpn/test.sh index 9b5610837032..392109d5e14e 100755 --- a/tools/testing/selftests/net/ovpn/test.sh +++ b/tools/testing/selftests/net/ovpn/test.sh @@ -56,6 +56,76 @@ ovpn_prepare_network() { done } +ovpn_new_test_peer() { + local peer_id="$1" + + shift + ip netns exec ovpn_peer0 "${OVPN_CLI}" new_peer tun0 \ + "${peer_id}" none 65000 10.10.1.2 1 "$@" +} + +ovpn_set_peer_vpn_addr() { + ip netns exec ovpn_peer0 "${OVPN_CLI}" set_peer tun0 \ + "$1" 60 120 "$2" +} + +ovpn_run_vpn_addr_validation() { + local addr + local peer1_addr4 + local test_peer_id=$((OVPN_NUM_PEERS + 1)) + local test_peer_addr6="2001:db8::2" + # Do not include 0.0.0.0 or :: here. They are invalid on creation, but + # clear one address family on update and are valid if the other remains. + local -a invalid_addrs=( + "127.0.0.1" + "224.0.0.1" + "255.255.255.255" + "::1" + "::192.0.2.1" + "::ffff:192.0.2.1" + "ff02::1" + ) + + peer1_addr4=$(ovpn_peer_vpn_addr 1) + + ovpn_cmd_fail "reject peer without VPN address" \ + ovpn_new_test_peer "${test_peer_id}" + + for addr in "0.0.0.0" "::" "${invalid_addrs[@]}"; do + ovpn_cmd_fail "reject new peer VPN address ${addr}" \ + ovpn_new_test_peer "${test_peer_id}" "${addr}" + done + + ovpn_cmd_fail "reject duplicate IPv4 address on peer creation" \ + ovpn_new_test_peer "${test_peer_id}" "${peer1_addr4}" + ovpn_cmd_fail "reject clearing the last peer VPN address" \ + ovpn_set_peer_vpn_addr 1 0.0.0.0 + + for addr in "${invalid_addrs[@]}"; do + ovpn_cmd_fail "reject updated peer VPN address ${addr}" \ + ovpn_set_peer_vpn_addr 1 "${addr}" + done + + ovpn_cmd_fail "reject duplicate IPv4 address on peer update" \ + ovpn_set_peer_vpn_addr 2 "${peer1_addr4}" + + ovpn_cmd_ok "add peer IPv6 address" \ + ovpn_set_peer_vpn_addr 1 "${test_peer_addr6}" + ovpn_cmd_fail "reject duplicate IPv6 address on peer creation" \ + ovpn_new_test_peer "${test_peer_id}" "${test_peer_addr6}" + ovpn_cmd_fail "reject duplicate IPv6 address on peer update" \ + ovpn_set_peer_vpn_addr 2 "${test_peer_addr6}" + + ovpn_cmd_ok "clear peer IPv4 address" \ + ovpn_set_peer_vpn_addr 1 0.0.0.0 + ovpn_cmd_fail "reject clearing the remaining peer IPv6 address" \ + ovpn_set_peer_vpn_addr 1 :: + ovpn_cmd_ok "restore peer IPv4 address" \ + ovpn_set_peer_vpn_addr 1 "${peer1_addr4}" + ovpn_cmd_ok "clear peer IPv6 address" \ + ovpn_set_peer_vpn_addr 1 :: +} + ovpn_run_basic_traffic() { local p local header1 @@ -293,15 +363,16 @@ trap ovpn_stage_err ERR ktap_print_header if [ "${OVPN_FLOAT}" == "1" ]; then - ktap_set_plan 13 + ktap_set_plan 14 else - ktap_set_plan 12 + ktap_set_plan 13 fi ovpn_cleanup modprobe -q ovpn || true ovpn_run_stage "setup network topology" ovpn_prepare_network +ovpn_run_stage "validate peer VPN addresses" ovpn_run_vpn_addr_validation ovpn_run_stage "run baseline data traffic" ovpn_run_basic_traffic ovpn_run_stage "run LAN traffic behind peer1" ovpn_run_lan_traffic [ "${OVPN_FLOAT}" == "1" ] && ovpn_run_stage "run floating peer checks" \ From f0ca020cbb9bb7f3f4ea8ba1dfcf30a282aec91e Mon Sep 17 00:00:00 2001 From: Hui Peng Date: Sat, 19 Sep 2026 22:17:38 +0000 Subject: [PATCH 064/189] Bluetooth: bnep: fix out-of-bounds reads on short RX/TX frames and control fallthrough Fix multiple out-of-bounds reads in Bluetooth BNEP frame processing: 1. In bnep_rx_frame() and bnep_ctrl_frame() (net/bluetooth/bnep/core.c), use pskb_may_pull() to verify the BNEP header, control type byte, filter count, and extension headers exist before reading them, and return 0 after handling BNEP_CONTROL instead of falling through to Ethernet frame submission when no extension headers follow. 2. In bnep_net_xmit() (net/bluetooth/bnep/netdev.c), verify skb->len >= ETH_HLEN with pskb_may_pull() before reading the 14-byte Ethernet header to prevent an out-of-bounds heap read and infoleak on short AF_PACKET TX frames. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Assisted-by: LLM Signed-off-by: Hui Peng Signed-off-by: Luiz Augusto von Dentz --- net/bluetooth/bnep/core.c | 17 ++++++++++++++++- net/bluetooth/bnep/netdev.c | 8 +++++++- 2 files changed, 23 insertions(+), 2 deletions(-) diff --git a/net/bluetooth/bnep/core.c b/net/bluetooth/bnep/core.c index f7d88c33e23e..ad24d2486665 100644 --- a/net/bluetooth/bnep/core.c +++ b/net/bluetooth/bnep/core.c @@ -270,9 +270,14 @@ static int bnep_rx_extension(struct bnep_session *s, struct sk_buff *skb) BT_DBG("type 0x%x len %u", h->type, h->len); + if (skb->len < h->len) { + err = -EILSEQ; + break; + } + switch (h->type & BNEP_TYPE_MASK) { case BNEP_EXT_CONTROL: - bnep_rx_control(s, skb->data, skb->len); + bnep_rx_control(s, skb->data, h->len); break; default: @@ -373,6 +378,11 @@ static int bnep_rx_frame(struct bnep_session *s, struct sk_buff *skb) goto badframe; } + if ((type & BNEP_TYPE_MASK) == BNEP_CONTROL) { + kfree_skb(skb); + return 0; + } + /* Strip 802.1p header */ if (ntohs(s->eh.h_proto) == ETH_P_8021Q) { if (!skb_pull(skb, 4)) @@ -451,6 +461,11 @@ static int bnep_tx_frame(struct bnep_session *s, struct sk_buff *skb) goto send; } + if (skb->len < ETH_HLEN) { + kfree_skb(skb); + return 0; + } + iv[il++] = (struct kvec) { &type, 1 }; len++; diff --git a/net/bluetooth/bnep/netdev.c b/net/bluetooth/bnep/netdev.c index ee1e39a3daff..b451ef457741 100644 --- a/net/bluetooth/bnep/netdev.c +++ b/net/bluetooth/bnep/netdev.c @@ -166,6 +166,12 @@ static netdev_tx_t bnep_net_xmit(struct sk_buff *skb, BT_DBG("skb %p, dev %p", skb, dev); + if (!pskb_may_pull(skb, ETH_HLEN)) { + dev->stats.tx_dropped++; + kfree_skb(skb); + return NETDEV_TX_OK; + } + #ifdef CONFIG_BT_BNEP_MC_FILTER if (bnep_net_mc_filter(skb, s)) { kfree_skb(skb); @@ -218,7 +224,7 @@ void bnep_net_setup(struct net_device *dev) dev->addr_len = ETH_ALEN; ether_setup(dev); - dev->min_mtu = 0; + dev->min_mtu = ETH_MIN_MTU; dev->max_mtu = ETH_MAX_MTU; dev->priv_flags &= ~IFF_TX_SKB_SHARING; dev->netdev_ops = &bnep_netdev_ops; From 37a11129345337efd6eef8e62b03b6348cd0dd8b Mon Sep 17 00:00:00 2001 From: Ravindra Date: Tue, 15 Sep 2026 10:42:15 +0530 Subject: [PATCH 065/189] Bluetooth: btintel_pcie: validate device-supplied DMA indices In btintel_pcie_msix_rx_handle(), the driver processes RX completion descriptors (urbd1) written by the PCIe device into DMA-coherent memory. urbd1->frbd_tag (a 16-bit field fully controlled by the device firmware via DMA) is used directly as an array index into rxq->bufs[] without any bounds check. rxq->bufs[] has only BTINTEL_PCIE_RX_DESCS_COUNT (64) entries, while frbd_tag can be any value 0-65535. A malicious or malfunctioning device can write an out-of-range frbd_tag, causing the driver to dereference an out-of-bounds data_buf pointer. Additionally, cr_hia is read from a DMA-shared index array also writable by the device; if the device sets cr_hia >= rxq->count, the while-loop never terminates because cr_tia is wrapped via modulo rxq->count and can never equal an out-of-range cr_hia. Add bounds validation for cr_hia and frbd_tag in the RX path, and cr_hia in the TX path. Log invalid values with bt_dev_err before returning. Fixes: c2b636b3f788 ("Bluetooth: btintel_pcie: Add support for PCIe transport") Signed-off-by: Ravindra Signed-off-by: Luiz Augusto von Dentz --- drivers/bluetooth/btintel_pcie.c | 16 ++++++++++++++++ 1 file changed, 16 insertions(+) diff --git a/drivers/bluetooth/btintel_pcie.c b/drivers/bluetooth/btintel_pcie.c index 6e6e2b19815c..2819e001797b 100644 --- a/drivers/bluetooth/btintel_pcie.c +++ b/drivers/bluetooth/btintel_pcie.c @@ -1099,6 +1099,11 @@ static void btintel_pcie_msix_tx_handle(struct btintel_pcie_data *data) txq = &data->txq; + if (cr_hia >= txq->count) { + bt_dev_err(data->hdev, "TXQ: invalid cr_hia %u", cr_hia); + return; + } + while (cr_tia != cr_hia) { data->tx_wait_done = true; wake_up(&data->tx_wait_q); @@ -1650,6 +1655,11 @@ static void btintel_pcie_msix_rx_handle(struct btintel_pcie_data *data) rxq = &data->rxq; + if (cr_hia >= rxq->count) { + bt_dev_err(hdev, "RXQ: invalid cr_hia %u", cr_hia); + return; + } + /* The firmware sends multiple CD in a single MSI-X and it needs to * process all received CDs in this interrupt. */ @@ -1657,6 +1667,12 @@ static void btintel_pcie_msix_rx_handle(struct btintel_pcie_data *data) urbd1 = &rxq->urbd1s[cr_tia]; ipc_print_urbd1(data->hdev, urbd1, cr_tia); + if (urbd1->frbd_tag >= rxq->count) { + bt_dev_err(hdev, "RXQ: invalid frbd_tag %u", + urbd1->frbd_tag); + return; + } + buf = &rxq->bufs[urbd1->frbd_tag]; if (!buf) { bt_dev_err(hdev, "RXQ: failed to get the DMA buffer for %d", From 46f8ffd0a1f1eb6cbc94946a92c11ef601e228a1 Mon Sep 17 00:00:00 2001 From: Hui Peng Date: Sat, 19 Sep 2026 11:25:18 +0000 Subject: [PATCH 066/189] Bluetooth: RFCOMM: fix NULL dereference of dlc->session in RFCOMM_CONNINFO The RFCOMM_CONNINFO getsockopt handler accepts a socket that is not connected as long as deferred setup is enabled: if (sk->sk_state != BT_CONNECTED && !rfcomm_pi(sk)->dlc->defer_setup) { err = -ENOTCONN; break; } l2cap_sk = rfcomm_pi(sk)->dlc->session->sock->sk; dlc->defer_setup is set in rfcomm_sock_init() when rfcomm_connect_ind() creates a child socket for an incoming connection on a listening socket that has BT_DEFER_SETUP enabled. It is never cleared afterwards. The session, however, can go away underneath it. rfcomm_recv_disc() forces the dlc state before tearing it down: d->state = BT_CLOSED; __rfcomm_dlc_close(d, err); The RFCOMM_DEFER_SETUP early return in __rfcomm_dlc_close() only covers BT_CONNECT, BT_CONFIG, BT_OPEN and BT_CONNECT2, so with the state already BT_CLOSED that switch does not match and the function falls through to rfcomm_dlc_unlink(), which sets d->session = NULL, while d->defer_setup stays 1. A getsockopt(SOL_RFCOMM, RFCOMM_CONNINFO) on the accepted socket after that point therefore skips the -ENOTCONN path -- sk->sk_state is BT_CLOSED, but dlc->defer_setup is still set -- and dereferences the NULL session. No race is needed: once the DISC has been processed, the dereference is unconditional. Reproduced on a KASAN kernel under QEMU with a BR/EDR peer emulated over /dev/vhci: the peer brings up an ACL link, opens L2CAP on the RFCOMM PSM, starts a session and sends SABM for a channel bound with BT_DEFER_SETUP, and sends DISC for that dlci after the socket has been accepted. getsockopt(SOL_RFCOMM, RFCOMM_CONNINFO) on the accepted socket then hits: Oops: general protection fault, probably for non-canonical address 0xdffffc0000000002: 0000 [#1] SMP KASAN PTI KASAN: null-ptr-deref in range [0x0000000000000010-0x0000000000000017] CPU: 1 UID: 0 PID: 150 Comm: init Tainted: G B 7.3.0-rc3-g5dd1818b15d9 Hardware name: QEMU Standard PC (i440FX + PIIX, 1996) RIP: 0010:rfcomm_sock_getsockopt+0x529/0x780 Call Trace: do_sock_getsockopt+0x3ad/0x7d0 __sys_getsockopt+0x10e/0x1b0 __x64_sys_getsockopt+0xc2/0x160 do_syscall_64+0xda/0x4b0 entry_SYSCALL_64_after_hwframe+0x77/0x7f 0x10 is the offset of sock in struct rfcomm_session; rfcomm_sock_getsockopt_old() is inlined into rfcomm_sock_getsockopt(). Commit 43a556b2fd43 ("Bluetooth: RFCOMM: take rfcomm_mutex for the deferred setup accept") fixed the same "a remote DISC clears the session while deferred setup is still flagged" problem in rfcomm_dlc_accept(); this is the remaining instance of it, in the getsockopt path. Deferred setup only leaves a socket usable here once it has reached BT_CONNECT2, so restrict the exception to that state and check that a session is actually present before following it. Fixes: bb23c0ab8246 ("Bluetooth: Add support for deferring RFCOMM connection setup") Cc: stable@vger.kernel.org Assisted-by: LLM Signed-off-by: Hui Peng Signed-off-by: Luiz Augusto von Dentz --- net/bluetooth/rfcomm/sock.c | 6 ++++-- 1 file changed, 4 insertions(+), 2 deletions(-) diff --git a/net/bluetooth/rfcomm/sock.c b/net/bluetooth/rfcomm/sock.c index e2486bc11cbc..fb924d0e34ec 100644 --- a/net/bluetooth/rfcomm/sock.c +++ b/net/bluetooth/rfcomm/sock.c @@ -786,8 +786,10 @@ static int rfcomm_sock_getsockopt_old(struct socket *sock, int optname, break; case RFCOMM_CONNINFO: - if (sk->sk_state != BT_CONNECTED && - !rfcomm_pi(sk)->dlc->defer_setup) { + if ((sk->sk_state != BT_CONNECTED && + !(sk->sk_state == BT_CONNECT2 && + rfcomm_pi(sk)->dlc->defer_setup)) || + !rfcomm_pi(sk)->dlc->session) { err = -ENOTCONN; break; } From 6d91041bb38b97e2feb625123cc0529d7b83a0e1 Mon Sep 17 00:00:00 2001 From: Hui Peng Date: Sat, 19 Sep 2026 11:25:14 +0000 Subject: [PATCH 067/189] Bluetooth: RFCOMM: Reject short EA=0 frames in rfcomm_recv_frame() While rfcomm_recv_frame() verifies that skb->len is at least sizeof(*hdr) + 1 (4 bytes: 3-byte header + 1-byte FCS), an RFCOMM frame with an extended 2-byte length field (!__test_ea(hdr->len)) has a 4-byte header plus a 1-byte FCS (5 bytes minimum, sizeof(*hdr) + 2). When a 4-byte RFCOMM frame with EA == 0 arrives: 1. The initial skb->len < sizeof(*hdr) + 1 check passes (4 < 4 is false). 2. Trimming the FCS byte decrements skb->len to 3. 3. If __check_fcs() succeeds, skb_pull(skb, 4) fails (4 > 3) and returns NULL without advancing skb->data. 4. Because the return value of skb_pull() is ignored, the un-pulled 3-byte struct rfcomm_hdr remains at skb->data and is either queued as application payload via rfcomm_recv_data() or parsed as a multiplexer control command via rfcomm_recv_mcc() on DLCI 0. Fix this by extending the length check in rfcomm_recv_frame() to also require skb->len >= sizeof(*hdr) + 2 when !__test_ea(hdr->len). Fixes: b230e5bf501c ("Bluetooth: RFCOMM: validate skb length in rfcomm_recv_frame") Assisted-by: LLM Signed-off-by: Hui Peng Signed-off-by: Luiz Augusto von Dentz --- net/bluetooth/rfcomm/core.c | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/net/bluetooth/rfcomm/core.c b/net/bluetooth/rfcomm/core.c index f7463f092283..d91e2a6ee26c 100644 --- a/net/bluetooth/rfcomm/core.c +++ b/net/bluetooth/rfcomm/core.c @@ -1817,7 +1817,8 @@ static struct rfcomm_session *rfcomm_recv_frame(struct rfcomm_session *s, return s; } - if (skb->len < sizeof(*hdr) + 1) { + if (skb->len < sizeof(*hdr) + 1 || + (!__test_ea(hdr->len) && skb->len < sizeof(*hdr) + 2)) { kfree_skb(skb); return s; } From d06f2ebf67ff2962fe00d687e4f0d4703eb41a12 Mon Sep 17 00:00:00 2001 From: Ratheesh Kannoth Date: Wed, 16 Sep 2026 07:51:11 +0530 Subject: [PATCH 068/189] octeontx2-af: Fix memory scaling limitation in SR-IOV mode The original code used DMA_ATTR_FORCE_CONTIGUOUS, which could exhaust the CMA pool when a large number of VFs were requested. Fix this by switching to the DMA streaming API. This is equivalent on Octeon platforms, which provide full I/O coherency via the SMMU. Cc: Leon Romanovsky Fixes: 73d33dbc0723 ("octeontx2-af: Use DMA_ATTR_FORCE_CONTIGUOUS attribute in DMA alloc") Signed-off-by: Ratheesh Kannoth Reviewed-by: Leon Romanovsky Link: https://patch.msgid.link/20260916022111.1083017-1-rkannoth@marvell.com Signed-off-by: Jakub Kicinski --- .../ethernet/marvell/octeontx2/af/common.h | 45 ++++++++++++++++--- 1 file changed, 39 insertions(+), 6 deletions(-) diff --git a/drivers/net/ethernet/marvell/octeontx2/af/common.h b/drivers/net/ethernet/marvell/octeontx2/af/common.h index 779413a383b7..78e42549d990 100644 --- a/drivers/net/ethernet/marvell/octeontx2/af/common.h +++ b/drivers/net/ethernet/marvell/octeontx2/af/common.h @@ -7,6 +7,10 @@ #ifndef COMMON_H #define COMMON_H +#include +#include +#include + #include "rvu_struct.h" #define OTX2_ALIGN 128 /* Align to cacheline */ @@ -44,6 +48,33 @@ struct qmem { u32 qsize; }; +static inline void *otx2_dma_alloc_coherent(struct device *dev, size_t size, + dma_addr_t *dma_handle) +{ + dma_addr_t dma_addr; + void *vaddr; + + vaddr = kzalloc(size, GFP_KERNEL); + if (!vaddr) + return NULL; + + dma_addr = dma_map_single(dev, vaddr, size, DMA_BIDIRECTIONAL); + if (dma_mapping_error(dev, dma_addr)) { + kfree(vaddr); + return NULL; + } + + *dma_handle = dma_addr; + return vaddr; +} + +static inline void otx2_dma_free_coherent(struct device *dev, size_t size, + void *vaddr, dma_addr_t dma_handle) +{ + dma_unmap_single(dev, dma_handle, size, DMA_BIDIRECTIONAL); + kfree(vaddr); +} + static inline int qmem_alloc(struct device *dev, struct qmem **q, int qsize, int entry_sz) { @@ -60,8 +91,11 @@ static inline int qmem_alloc(struct device *dev, struct qmem **q, qmem->entry_sz = entry_sz; qmem->alloc_sz = (qsize * entry_sz) + OTX2_ALIGN; - qmem->base = dma_alloc_attrs(dev, qmem->alloc_sz, &qmem->iova, - GFP_KERNEL, DMA_ATTR_FORCE_CONTIGUOUS); + + if (get_order(PAGE_ALIGN(qmem->alloc_sz)) > MAX_PAGE_ORDER) + return -ENOMEM; + + qmem->base = otx2_dma_alloc_coherent(dev, qmem->alloc_sz, &qmem->iova); if (!qmem->base) return -ENOMEM; @@ -80,10 +114,9 @@ static inline void qmem_free(struct device *dev, struct qmem *qmem) return; if (qmem->base) - dma_free_attrs(dev, qmem->alloc_sz, - qmem->base - qmem->align, - qmem->iova - qmem->align, - DMA_ATTR_FORCE_CONTIGUOUS); + otx2_dma_free_coherent(dev, qmem->alloc_sz, + qmem->base - qmem->align, + qmem->iova - qmem->align); devm_kfree(dev, qmem); } From 0346ec2f080b40d95ed05b853bb9226289e75212 Mon Sep 17 00:00:00 2001 From: Kuniyuki Iwashima Date: Fri, 18 Sep 2026 08:22:05 +0000 Subject: [PATCH 069/189] ipv6: Prevent rt6_insert_exception() for dying fib6_info. Before the cited commit, fib6_nh_flush_exceptions() always set from->exception_bucket_flushed = 1 under rt6_exception_lock to prevent rt6_insert_exception() from inserting a new exception for a dying fib6_info. The flag was replaced with the FIB6_EXCEPTION_BUCKET_FLUSHED bit stored in nh->rt6i_exception_bucket. The problem is that now the bit is only set when the bucket is not NULL and fib6_nh_flush_exceptions() is called from fib6_nh_release() after fib6_ref has already reached zero. If rt6_insert_exception() is called while the target fib6_info is being removed via fib6_purge_rt(), a new exception could be created successfully because rt6_flush_exceptions() no longer sets the bit. This creates a reference cycle between the fib6_info and the exception route, leaking the fib6_info, its nexthop device, and all per-CPU routes in fib6_nh->rt6i_pcpu, which stalls netdev unregistration. [ 34.680602] unregister_netdevice: waiting for gre6 to become free. Usage count = 68 [ 44.920675] unregister_netdevice: waiting for gre6 to become free. Usage count = 68 [ 55.176582] unregister_netdevice: waiting for gre6 to become free. Usage count = 68 Let's call fib6_drop_pcpu_from() before rt6_flush_exceptions(), to set fib6_destroying before rt6_exception_lock, and check f6i->fib6_destroying in rt6_insert_exception(). Note that FIB6_EXCEPTION_BUCKET_FLUSHED logic is dead and we can clean it up in net-next. Fixes: cc5c073a693f ("ipv6: Move exception bucket to fib6_nh") Signed-off-by: Kuniyuki Iwashima Reviewed-by: Ido Schimmel Link: https://patch.msgid.link/20260918082209.2853582-1-kuniyu@google.com Signed-off-by: Jakub Kicinski --- net/ipv6/ip6_fib.c | 2 +- net/ipv6/route.c | 5 +++++ 2 files changed, 6 insertions(+), 1 deletion(-) diff --git a/net/ipv6/ip6_fib.c b/net/ipv6/ip6_fib.c index 9ea75703b38d..9ff761962b45 100644 --- a/net/ipv6/ip6_fib.c +++ b/net/ipv6/ip6_fib.c @@ -1043,8 +1043,8 @@ static void fib6_purge_rt(struct fib6_info *rt, struct fib6_node *fn, struct fib6_table *table = rt->fib6_table; /* Flush all cached dst in exception table */ - rt6_flush_exceptions(rt); fib6_drop_pcpu_from(rt); + rt6_flush_exceptions(rt); if (rt->nh) { spin_lock(&rt->nh->lock); diff --git a/net/ipv6/route.c b/net/ipv6/route.c index 08bd68f1b5bb..884d9ab0d50d 100644 --- a/net/ipv6/route.c +++ b/net/ipv6/route.c @@ -1729,6 +1729,11 @@ static int rt6_insert_exception(struct rt6_info *nrt, spin_lock_bh(&rt6_exception_lock); + if (f6i->fib6_destroying) { + err = -ENOENT; + goto out; + } + bucket = rcu_dereference_protected(nh->rt6i_exception_bucket, lockdep_is_held(&rt6_exception_lock)); if (!bucket) { From 31995571219c8ac30913d9c0dccad033fbb0b3da Mon Sep 17 00:00:00 2001 From: Nicolai Buchwitz Date: Fri, 18 Sep 2026 11:55:40 +0200 Subject: [PATCH 070/189] net: don't require the hwtstamp NDOs when a PHY provides timestamping Removing the legacy ioctl fallback made both hwtstamp NDOs mandatory. A device that only timestamps in its PHY implements neither, so SIOCSHWTSTAMP fails with EOPNOTSUPP before anything looks at the PHY and PTP stops working there. The check only ever picked the legacy path. That path is gone, so drop it and test where the NDOs are actually called. SIOCGHWTSTAMP is new here, not restored. The old path went through phy_mii_ioctl(), which only handled SIOCSHWTSTAMP. Such a device now returns -ENODEV while absent instead of -EOPNOTSUPP, like the ones that do implement the NDOs. Fixes: 5062245a5a7f ("net: remove legacy way to get/set HW timestamp config") Signed-off-by: Nicolai Buchwitz Reviewed-by: Kory Maincent Link: https://patch.msgid.link/20260918095540.34286-1-nb@tipi-net.de Signed-off-by: Jakub Kicinski --- net/core/dev_ioctl.c | 25 +++++++++---------------- 1 file changed, 9 insertions(+), 16 deletions(-) diff --git a/net/core/dev_ioctl.c b/net/core/dev_ioctl.c index a320e264eaaf..164643140a52 100644 --- a/net/core/dev_ioctl.c +++ b/net/core/dev_ioctl.c @@ -276,19 +276,18 @@ int dev_get_hwtstamp_phylib(struct net_device *dev, if (phy_is_default_hwtstamp(dev->phydev)) return phy_hwtstamp_get(dev->phydev, cfg); + if (!dev->netdev_ops->ndo_hwtstamp_get) + return -EOPNOTSUPP; + return dev->netdev_ops->ndo_hwtstamp_get(dev, cfg); } static int dev_get_hwtstamp(struct net_device *dev, struct ifreq *ifr) { - const struct net_device_ops *ops = dev->netdev_ops; struct kernel_hwtstamp_config kernel_cfg = {}; struct hwtstamp_config cfg; int err; - if (!ops->ndo_hwtstamp_get) - return -EOPNOTSUPP; - if (!netif_device_present(dev)) return -ENODEV; @@ -359,12 +358,18 @@ int dev_set_hwtstamp_phylib(struct net_device *dev, cfg->source = phy_ts ? HWTSTAMP_SOURCE_PHYLIB : HWTSTAMP_SOURCE_NETDEV; if (phy_ts && dev->see_all_hwtstamp_requests) { + if (!ops->ndo_hwtstamp_get) + return -EOPNOTSUPP; + err = ops->ndo_hwtstamp_get(dev, &old_cfg); if (err) return err; } if (!phy_ts || dev->see_all_hwtstamp_requests) { + if (!ops->ndo_hwtstamp_set) + return -EOPNOTSUPP; + err = ops->ndo_hwtstamp_set(dev, cfg, extack); if (err) { if (extack->_msg) @@ -390,7 +395,6 @@ int dev_set_hwtstamp_phylib(struct net_device *dev, static int dev_set_hwtstamp(struct net_device *dev, struct ifreq *ifr) { - const struct net_device_ops *ops = dev->netdev_ops; struct kernel_hwtstamp_config kernel_cfg = {}; struct netlink_ext_ack extack = {}; struct hwtstamp_config cfg; @@ -413,9 +417,6 @@ static int dev_set_hwtstamp(struct net_device *dev, struct ifreq *ifr) return err; } - if (!ops->ndo_hwtstamp_set) - return -EOPNOTSUPP; - if (!netif_device_present(dev)) return -ENODEV; @@ -441,15 +442,11 @@ static int dev_set_hwtstamp(struct net_device *dev, struct ifreq *ifr) int generic_hwtstamp_get_lower(struct net_device *dev, struct kernel_hwtstamp_config *kernel_cfg) { - const struct net_device_ops *ops = dev->netdev_ops; int err; if (!netif_device_present(dev)) return -ENODEV; - if (!ops->ndo_hwtstamp_get) - return -EOPNOTSUPP; - netdev_lock_ops(dev); err = dev_get_hwtstamp_phylib(dev, kernel_cfg); netdev_unlock_ops(dev); @@ -462,15 +459,11 @@ int generic_hwtstamp_set_lower(struct net_device *dev, struct kernel_hwtstamp_config *kernel_cfg, struct netlink_ext_ack *extack) { - const struct net_device_ops *ops = dev->netdev_ops; int err; if (!netif_device_present(dev)) return -ENODEV; - if (!ops->ndo_hwtstamp_set) - return -EOPNOTSUPP; - netdev_lock_ops(dev); err = dev_set_hwtstamp_phylib(dev, kernel_cfg, extack); netdev_unlock_ops(dev); From 7cce782d8327b7291334c4a304cf3fd909a74d9d Mon Sep 17 00:00:00 2001 From: Ivan Vecera Date: Thu, 17 Sep 2026 16:37:36 +0200 Subject: [PATCH 071/189] dpll: use exact lookup for reference sync pin id dpll_pin_ref_sync_state_set() looks up the reference sync pin in the pin->ref_sync_pins xarray, which is keyed by the sync pin's id (see dpll_pin_ref_sync_pair_add() using xa_insert() with ref_sync_pin->id). The pin id to operate on is supplied by userspace via DPLL_A_PIN_ID. The lookup however used xa_find() with a ULONG_MAX limit, which returns the first present entry with an index greater than or equal to the requested id, not the entry stored exactly at that id. If userspace passes an id that is not paired as a reference sync pin, but another pin with a higher id is present in the xarray, xa_find() silently returns that wrong pin and the subsequent ref_sync_set() operates on it. The request only fails when the given id is larger than every present key. Use xa_load() for an exact-key lookup instead, mirroring the deletion path in dpll_pin_ref_sync_pair_del(). Fixes: 58256a26bfb3 ("dpll: add reference sync get/set") Signed-off-by: Ivan Vecera Link: https://patch.msgid.link/20260917143736.526221-1-ivecera@redhat.com Signed-off-by: Jakub Kicinski --- drivers/dpll/dpll_netlink.c | 3 +-- 1 file changed, 1 insertion(+), 2 deletions(-) diff --git a/drivers/dpll/dpll_netlink.c b/drivers/dpll/dpll_netlink.c index 45365214fbef..fb24fd53f2e1 100644 --- a/drivers/dpll/dpll_netlink.c +++ b/drivers/dpll/dpll_netlink.c @@ -1210,8 +1210,7 @@ dpll_pin_ref_sync_state_set(struct dpll_pin *pin, struct dpll_device *dpll; int ret; - ref_sync_pin = xa_find(&pin->ref_sync_pins, &ref_sync_pin_idx, - ULONG_MAX, XA_PRESENT); + ref_sync_pin = xa_load(&pin->ref_sync_pins, ref_sync_pin_idx); if (!ref_sync_pin) { NL_SET_ERR_MSG(extack, "reference sync pin not found"); return -EINVAL; From c06bde80ae7a7b595732f7cabcb92cf08db9d56a Mon Sep 17 00:00:00 2001 From: Pengpeng Hou Date: Sun, 20 Sep 2026 11:47:45 +0800 Subject: [PATCH 072/189] net: usb: sr9700: include receive overhead in the length check The receive fixup subtracts the Ethernet CRC from the reported packet length, but compares that payload length against the whole remaining receive buffer. The following copy starts after the three-byte header, and the cursor advance consumes both that header and the four-byte CRC. Require the payload to fit after SR_RX_OVERHEAD before copying it or advancing to the next packet. The loop already ensures that the remaining buffer is larger than the overhead, so the subtraction is safe. The issue was found by our static-analysis tool. Fixes: c9b37458e956 ("USB2NET : SR9700 : One chip USB 1.1 USB2NET SR9700Device Driver Support") Reviewed-by: Ethan Nelson-Moore Tested-by: Ethan Nelson-Moore Signed-off-by: Pengpeng Hou Link: https://patch.msgid.link/20260920034745.18468-1-hppiscas@163.com Signed-off-by: Jakub Kicinski --- drivers/net/usb/sr9700.c | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/drivers/net/usb/sr9700.c b/drivers/net/usb/sr9700.c index 937e6fef3ac6..50981a28376a 100644 --- a/drivers/net/usb/sr9700.c +++ b/drivers/net/usb/sr9700.c @@ -355,7 +355,8 @@ static int sr9700_rx_fixup(struct usbnet *dev, struct sk_buff *skb) /* ignore the CRC length */ len = (skb->data[1] | (skb->data[2] << 8)) - 4; - if (len > ETH_FRAME_LEN || len > skb->len || len < 0) + if (len > ETH_FRAME_LEN || len < 0 || + len > skb->len - SR_RX_OVERHEAD) return 0; /* the last packet of current skb */ From be581d6635579489eadff9cea4caee142ec6d8bc Mon Sep 17 00:00:00 2001 From: Yuya Kusakabe Date: Fri, 18 Sep 2026 22:49:30 +0900 Subject: [PATCH 073/189] selftests: net: fix CONFIG_SYSCTL sort order in configs Commit 8d75c338f0bc ("sysctl: remove CONFIG_PROC_SYSCTL, it just mirrors CONFIG_SYSCTL") renamed CONFIG_PROC_SYSCTL to CONFIG_SYSCTL in place, which left the entry out of alphabetical order in the net and packetdrill configs. The netdev CI check for sorted selftest configs now fails for every patch that touches either file. Fixes: 8d75c338f0bc ("sysctl: remove CONFIG_PROC_SYSCTL, it just mirrors CONFIG_SYSCTL") Signed-off-by: Yuya Kusakabe Reviewed-by: Joel Granados Link: https://patch.msgid.link/20260918-selftests-net-config-sort-v1-1-968ea6e8c1b7@gmail.com Signed-off-by: Jakub Kicinski --- tools/testing/selftests/net/config | 2 +- tools/testing/selftests/net/packetdrill/config | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/tools/testing/selftests/net/config b/tools/testing/selftests/net/config index 30d5fcb09a83..737e7e6327b3 100644 --- a/tools/testing/selftests/net/config +++ b/tools/testing/selftests/net/config @@ -118,10 +118,10 @@ CONFIG_NFT_NAT=m CONFIG_NUMA=y CONFIG_OPENVSWITCH=m CONFIG_PAGE_POOL_STATS=y -CONFIG_SYSCTL=y CONFIG_PSAMPLE=m CONFIG_RPS=y CONFIG_SYN_COOKIES=y +CONFIG_SYSCTL=y CONFIG_SYSFS=y CONFIG_TAP=m CONFIG_TCP_CONG_DCTCP=y diff --git a/tools/testing/selftests/net/packetdrill/config b/tools/testing/selftests/net/packetdrill/config index 83dde525c53c..b7df8bf30920 100644 --- a/tools/testing/selftests/net/packetdrill/config +++ b/tools/testing/selftests/net/packetdrill/config @@ -4,8 +4,8 @@ CONFIG_IPV6=y CONFIG_NET_NS=y CONFIG_NET_SCH_FIFO=y CONFIG_NET_SCH_FQ=y -CONFIG_SYSCTL=y CONFIG_SYN_COOKIES=y +CONFIG_SYSCTL=y CONFIG_TCP_CONG_CUBIC=y CONFIG_TCP_MD5SIG=y CONFIG_TUN=y From 10de7ed8ef4840da9ca21de4c29578657ac367db Mon Sep 17 00:00:00 2001 From: Andrea Parri Date: Thu, 17 Sep 2026 13:55:42 +0200 Subject: [PATCH 074/189] net/mlx5e: fix swapped IPv6 IPsec policy masks IPv6 XFRM policies may use different source and destination prefix lengths. mlx5e_ipsec_policy_mask() builds the corresponding masks independently, but setup_fte_addr6() installs each mask in the opposite address field. When the prefix lengths differ, this makes the source match use the destination prefix and the destination match use the source prefix. The resulting hardware rule can both miss traffic covered by the policy and match traffic outside it. Install each mask in its corresponding match field. Fixes: ca7992f52c2c ("net/mlx5e: Properly match IPsec subnet addresses") Cc: stable@vger.kernel.org Signed-off-by: Andrea Parri Reviewed-by: Tariq Toukan Link: https://patch.msgid.link/20260917115542.177675-1-parri.andrea@gmail.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/mellanox/mlx5/core/en_accel/ipsec_fs.c | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ipsec_fs.c b/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ipsec_fs.c index 329608c59313..8ffa8068e90a 100644 --- a/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ipsec_fs.c +++ b/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ipsec_fs.c @@ -1564,14 +1564,14 @@ static void setup_fte_addr6(struct mlx5_flow_spec *spec, memcpy(MLX5_ADDR_OF(fte_match_param, spec->match_value, outer_headers.src_ipv4_src_ipv6.ipv6_layout.ipv6), saddr, 16); memcpy(MLX5_ADDR_OF(fte_match_param, spec->match_criteria, - outer_headers.src_ipv4_src_ipv6.ipv6_layout.ipv6), dmask, 16); + outer_headers.src_ipv4_src_ipv6.ipv6_layout.ipv6), smask, 16); } if (!addr6_all_zero(daddr)) { memcpy(MLX5_ADDR_OF(fte_match_param, spec->match_value, outer_headers.dst_ipv4_dst_ipv6.ipv6_layout.ipv6), daddr, 16); memcpy(MLX5_ADDR_OF(fte_match_param, spec->match_criteria, - outer_headers.dst_ipv4_dst_ipv6.ipv6_layout.ipv6), smask, 16); + outer_headers.dst_ipv4_dst_ipv6.ipv6_layout.ipv6), dmask, 16); } } From be31fe6333f534155e6b408f1ef6d77974bb41aa Mon Sep 17 00:00:00 2001 From: Kuniyuki Iwashima Date: Sun, 20 Sep 2026 19:14:32 +0000 Subject: [PATCH 075/189] ipv6: Fix dst leak for uncached routes. ip6_route_output_flags(), ip6_rt_put_flags(), and ip6_dst_check() detect an uncached route by list_empty(&rt->dst.rt_uncached), which replaced the static DST_NOCACHE flag check in commit a4c2fd7f7891 ("net: remove DST_NOCACHE flag"). When a device is unregistered, rt6_uncached_list_flush_dev() unlinks uncached routes tied to the device from rt6_uncached_list. Previously, they were moved to another list with list_move() (__list_del_entry() + list_add()), and since commit 98aa546af5e4 ("inet: remove (struct uncached_list)->quarantine"), the routes are just unlinked with list_del_init(). If list_del_init() runs concurrently, list_empty() evaluates to true; ip6_route_output_flags() calls dst_hold_safe() incorrectly and ip6_rt_put_flags() skips ip6_rt_put(), leaking dst, and thus dev tied via rt->from as well. The same race is partially fixed by commit 9a6f0c4d5796 ("dst: fix races in rt6_uncached_list_del() and rt_del_uncached_list()"). Let's check rt6->dst.rt_uncached_list instead. Note that IPv4 does not have the same issue. Fixes: 98aa546af5e4 ("inet: remove (struct uncached_list)->quarantine") Signed-off-by: Kuniyuki Iwashima Reviewed-by: Hangbin Liu Reviewed-by: Xuanqiang Luo Reviewed-by: Ido Schimmel Reviewed-by: Eric Dumazet Link: https://patch.msgid.link/20260920191558.2990636-1-kuniyu@google.com Signed-off-by: Jakub Kicinski --- include/net/ip6_route.h | 4 ++-- net/ipv6/route.c | 7 ++++--- 2 files changed, 6 insertions(+), 5 deletions(-) diff --git a/include/net/ip6_route.h b/include/net/ip6_route.h index b9e8d2b759e9..0f9b7a260d25 100644 --- a/include/net/ip6_route.h +++ b/include/net/ip6_route.h @@ -101,12 +101,12 @@ static inline struct dst_entry *ip6_route_output(struct net *net, } /* Only conditionally release dst if flags indicates - * !RT6_LOOKUP_F_DST_NOREF or dst is in uncached_list. + * !RT6_LOOKUP_F_DST_NOREF or dst is uncached. */ static inline void ip6_rt_put_flags(struct rt6_info *rt, int flags) { if (!(flags & RT6_LOOKUP_F_DST_NOREF) || - !list_empty(&rt->dst.rt_uncached)) + rt->dst.rt_uncached_list) ip6_rt_put(rt); } diff --git a/net/ipv6/route.c b/net/ipv6/route.c index 884d9ab0d50d..153ce16628c1 100644 --- a/net/ipv6/route.c +++ b/net/ipv6/route.c @@ -139,6 +139,7 @@ void rt6_uncached_list_add(struct rt6_info *rt) { struct uncached_list *ul = raw_cpu_ptr(&rt6_uncached_list); + /* Set once and never cleared: non-NULL marks an uncached route. */ rt->dst.rt_uncached_list = ul; spin_lock_bh(&ul->lock); @@ -2726,8 +2727,8 @@ struct dst_entry *ip6_route_output_flags(struct net *net, rcu_read_lock(); dst = ip6_route_output_flags_noref(net, sk, fl6, flags); rt6 = dst_rt6_info(dst); - /* For dst cached in uncached_list, refcnt is already taken. */ - if (list_empty(&rt6->dst.rt_uncached) && !dst_hold_safe(dst)) { + /* For an uncached dst, refcnt is already taken. */ + if (!rt6->dst.rt_uncached_list && !dst_hold_safe(dst)) { dst = &net->ipv6.ip6_null_entry->dst; dst_hold(dst); } @@ -2836,7 +2837,7 @@ INDIRECT_CALLABLE_SCOPE struct dst_entry *ip6_dst_check(struct dst_entry *dst, from = rcu_dereference(rt->from); if (from && (rt->rt6i_flags & RTF_PCPU || - unlikely(!list_empty(&rt->dst.rt_uncached)))) + unlikely(rt->dst.rt_uncached_list))) dst_ret = rt6_dst_from_check(rt, from, cookie); else dst_ret = rt6_check(rt, from, cookie); From 5fd0783b99d4af98f65cd58b56ec203d1d426104 Mon Sep 17 00:00:00 2001 From: Shardul Bankar Date: Thu, 17 Sep 2026 14:55:32 +0530 Subject: [PATCH 076/189] udp: relocate a connected socket in the 4-tuple hash table on re-connect A connected UDP socket that connects again to a different peer is not re-filed in the 4-tuple hash table: sk binds to 127.0.0.1:21001 sk connects to 127.0.0.2:20001 // filed under hash(sk, peer1) sk connects to 127.0.0.3:20002 // still filed under hash(sk, peer1) packet from 127.0.0.3:20002 // hash(sk, peer2) misses, so the // lookup falls back to scoring the // hash2 chain for this address // and port udp_lib_hash4() returns early when the socket is already hashed, assuming ->rehash() relocates it. ->rehash() runs from __ip{4,6}_datagram_connect() only while the receive address is unset, which a second connect never is: the first connect assigns it, whether the socket was bound to a specific address or to the wildcard. commit 644f9108f3a5 ("udp: Make rehash4 independent in udp_lib_rehash()") added that early return and named connect(AF_UNSPEC) as the way around it. That workaround does not help a socket with both SOCK_BINDADDR_LOCK and SOCK_BINDPORT_LOCK set, because __udp_disconnect() skips ->rehash() for the first and ->unhash() for the second. Delivery is correct either way. Relocate the socket when the hash it is filed under differs from the one requested, which is what commit 78c91ae2c6de ("ipv4/udp: Add 4-tuple hash for connected socket") did before the early return became unconditional. It is done here under hslot->lock, which that version did not take, to match udp_lib_rehash() and udp_lib_unhash(). hslot2 is unchanged, so hash4_cnt needs no adjustment, as in udp_lib_rehash(). A first connect is unaffected, and IPv6 shares the code. With 500 sockets on the port, a re-connected socket measured 522,553 pps without this change and 2,055,078 with it. The UDP side was noted as remaining work in [1]. Link: https://lore.kernel.org/netdev/apnHqmYZQ4yzOP4N@v4bel/ [1] Fixes: 644f9108f3a5 ("udp: Make rehash4 independent in udp_lib_rehash()") Assisted-by: LLM Signed-off-by: Shardul Bankar Reviewed-by: Kuniyuki Iwashima Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-1-718891af0d7a@mpiricsoftware.com Signed-off-by: Paolo Abeni --- net/ipv4/udp.c | 21 +++++++++++++++------ 1 file changed, 15 insertions(+), 6 deletions(-) diff --git a/net/ipv4/udp.c b/net/ipv4/udp.c index bb8cfc62cb00..0fa3cdbdcc21 100644 --- a/net/ipv4/udp.c +++ b/net/ipv4/udp.c @@ -617,14 +617,23 @@ void udp_lib_hash4(struct sock *sk, u16 hash) struct net *net = sock_net(sk); struct udp_table *udptable; - /* Connected udp socket can re-connect to another remote address, which - * will be handled by rehash. Thus no need to redo hash4 here. - */ - if (udp_hashed4(sk)) - return; - udptable = net->ipv4.udp_table; hslot = udp_hashslot(udptable, net, udp_sk(sk)->udp_port_hash); + + /* A connected socket can re-connect to another address. rehash() + * relocates it, but only runs when the local address changes, so a + * socket bound to a specific address would stay filed under the + * previous peer's hash. Move it here. + */ + if (udp_hashed4(sk)) { + if (udp_sk(sk)->udp_lrpa_hash != hash) { + spin_lock_bh(&hslot->lock); + udp_rehash4(udptable, sk, hash); + spin_unlock_bh(&hslot->lock); + } + return; + } + hslot2 = udp_hashslot2(udptable, udp_sk(sk)->udp_portaddr_hash); hslot4 = udp_hashslot4(udptable, hash); udp_sk(sk)->udp_lrpa_hash = hash; From 9e95b1a94c9c49b4ba722251bbca2759b9c51737 Mon Sep 17 00:00:00 2001 From: Shardul Bankar Date: Thu, 17 Sep 2026 14:55:33 +0530 Subject: [PATCH 077/189] udp: remove a disconnected socket from the 4-tuple hash table A UDP socket bound to a specific address and port keeps its entry in the 4-tuple hash table after it is disconnected: sk binds to 127.0.0.1:21001 sk connects to 127.0.0.2:20001 // filed in the 4-tuple table sk disconnects, connect(AF_UNSPEC) // still filed, peer now 0.0.0.0:0 __udp_disconnect() takes a socket out of that table only as a side effect of ->rehash() or ->unhash(), and it skips ->rehash() when SOCK_BINDADDR_LOCK is set and ->unhash() when SOCK_BINDPORT_LOCK is set. commit 6996a2d2d0a6 ("udp: Unhash auto-bound connected sk from 4-tuple hash table when disconnected.") fixed the same end state for a wildcard-bound socket, by a path this one does not take. The entry is counted whether or not anything hits it. hash4_cnt on the hash2 slot stays raised for as long as the socket lives, so udp_has_hash4() keeps sending every packet for that address and port through the 4-tuple lookup first. On IPv6 it can also be hit. __udp_disconnect() does not clear sk_v6_daddr, so udp_v6_rehash() files the entry under the peer the socket was connected to with a zero dport, and inet6_match() compares that same field: a datagram from the former peer with a zero source port matches, and source port zero is accepted on receive. On IPv4 the peer is cleared, so a match would need a zero source address as well, which the routing layer rejects as martian. The stale sk_v6_daddr is a separate defect, not addressed here; removing the entry closes this path either way. The entry can also be relocated. __udp_disconnect() clears sk_bound_dev_if, so a subsequent SO_BINDTODEVICE calls ->rehash(), and because the receive address is still specific udp_lib_rehash() moves the entry instead of removing it, into the bucket that (rcv_saddr, num, 0, 0) hashes to -- a pure function of the address and port, so every socket reaching this state on one address and port collects in one bucket. The bucket cannot be chosen from outside, as udp_ehashfn() is seeded with a per-boot secret. This last one became reachable only with commit 644f9108f3a5 ("udp: Make rehash4 independent in udp_lib_rehash()"), which moved the hash4 handling out of a branch a disconnected socket does not take; the stale entry itself dates from the commit in Fixes. Take the socket out of the table before __udp_disconnect() runs, while it still matches how it was filed. This also reaches the wildcard case ahead of udp_lib_rehash()'s udp_unhash4() branch, leaving that branch unreachable from udp_disconnect(); removing it belongs in net-next. udp_disconnect() and udp_abort() are the only UDP entries into __udp_disconnect(), which is shared with raw, ping and l2tp sockets that are not struct udp_sock: ping_prot.obj_size is sizeof(struct inet_sock), so udp_hashed4() on one would read past the allocation. Fixes: 78c91ae2c6de ("ipv4/udp: Add 4-tuple hash for connected socket") Assisted-by: LLM Signed-off-by: Shardul Bankar Reviewed-by: Kuniyuki Iwashima Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-2-718891af0d7a@mpiricsoftware.com Signed-off-by: Paolo Abeni --- net/ipv4/udp.c | 23 +++++++++++++++++++++++ 1 file changed, 23 insertions(+) diff --git a/net/ipv4/udp.c b/net/ipv4/udp.c index 0fa3cdbdcc21..b090bd1f59e8 100644 --- a/net/ipv4/udp.c +++ b/net/ipv4/udp.c @@ -2206,9 +2206,31 @@ int __udp_disconnect(struct sock *sk, int flags) } EXPORT_SYMBOL(__udp_disconnect); +/* __udp_disconnect() takes a socket out of the 4-tuple hash table only via + * ->rehash() or ->unhash(), and neither runs for a socket bound to a + * specific address and port. Remove it here, before its peer is cleared. + */ +static void udp_unhash4_on_disconnect(struct sock *sk) +{ + struct net *net = sock_net(sk); + struct udp_table *udptable; + struct udp_hslot *hslot; + + if (!udp_hashed4(sk)) + return; + + udptable = net->ipv4.udp_table; + hslot = udp_hashslot(udptable, net, udp_sk(sk)->udp_port_hash); + + spin_lock_bh(&hslot->lock); + udp_unhash4(udptable, sk); + spin_unlock_bh(&hslot->lock); +} + int udp_disconnect(struct sock *sk, int flags) { lock_sock(sk); + udp_unhash4_on_disconnect(sk); __udp_disconnect(sk, flags); release_sock(sk); return 0; @@ -3140,6 +3162,7 @@ int udp_abort(struct sock *sk, int err) sk->sk_err = err; sk_error_report(sk); + udp_unhash4_on_disconnect(sk); __udp_disconnect(sk, 0); out: From a92e1a412c53dc0d9ad639e7abf8b3fc70a5b6ad Mon Sep 17 00:00:00 2001 From: Myeonghun Pak Date: Thu, 17 Sep 2026 14:33:36 -0400 Subject: [PATCH 078/189] tg3: clean up PHYLIB resources on probe failure tg3_get_invariants() can register an MDIO bus and connect a PHY for USE_PHYLIB devices. If tg3_init_one() later fails, its common error path releases the mappings and netdev without undoing those PHYLIB resources. Disconnect the PHY and unregister the MDIO bus before the remaining teardown. Guard PHY cleanup with USE_PHYLIB to match tg3_phy_init(), and call tg3_mdio_fini() unconditionally to match tg3_mdio_init(). The existing IS_CONNECTED and MDIOBUS_INITED flags make both helpers safe when initialization only completed partially. This issue was identified during our ongoing static-analysis research while reviewing kernel code. Fixes: 158d7abdae85 ("tg3: Add mdio bus registration") Assisted-by: OpenAI:GPT-5.6 Co-developed-by: Ijae Kim Signed-off-by: Ijae Kim Signed-off-by: Myeonghun Pak Link: https://patch.msgid.link/20260917183336.36239-1-mhun512@gmail.com Signed-off-by: Paolo Abeni --- drivers/net/ethernet/broadcom/tg3.c | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/drivers/net/ethernet/broadcom/tg3.c b/drivers/net/ethernet/broadcom/tg3.c index 73a4b569b03e..caa7a6caa6e2 100644 --- a/drivers/net/ethernet/broadcom/tg3.c +++ b/drivers/net/ethernet/broadcom/tg3.c @@ -18047,6 +18047,10 @@ static int tg3_init_one(struct pci_dev *pdev, return 0; err_out_apeunmap: + if (tg3_flag(tp, USE_PHYLIB)) + tg3_phy_fini(tp); + tg3_mdio_fini(tp); + if (tp->aperegs) { iounmap(tp->aperegs); tp->aperegs = NULL; From a644f09b2090ad22a13fbcf9d141084f573108ef Mon Sep 17 00:00:00 2001 From: Wentao Liang Date: Thu, 17 Sep 2026 11:01:35 +0000 Subject: [PATCH 079/189] fsl/fman: Fix clk reference leak in read_dts_node() of_clk_get() returns a clock with its reference count incremented, but read_dts_node() only uses it to read the rate and never calls clk_put(). The clock is not stored anywhere, so the reference cannot be released later either. Release the clock once its rate has been read, which also covers the error path taken when the rate is zero. Fixes: 414fd46e7762 ("fsl/fman: Add FMan support") Cc: stable@vger.kernel.org Signed-off-by: Wentao Liang Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260917110135.2148068-1-vulab@iscas.ac.cn Signed-off-by: Paolo Abeni --- drivers/net/ethernet/freescale/fman/fman.c | 1 + 1 file changed, 1 insertion(+) diff --git a/drivers/net/ethernet/freescale/fman/fman.c b/drivers/net/ethernet/freescale/fman/fman.c index 299bab043175..46cc28895e56 100644 --- a/drivers/net/ethernet/freescale/fman/fman.c +++ b/drivers/net/ethernet/freescale/fman/fman.c @@ -2734,6 +2734,7 @@ static struct fman *read_dts_node(struct platform_device *of_dev) } clk_rate = clk_get_rate(clk); + clk_put(clk); if (!clk_rate) { err = -EINVAL; dev_err(&of_dev->dev, "%s: Failed to determine FM%d clock rate\n", From 2d14720beb58870b15a52b236c3ab0be0e06e915 Mon Sep 17 00:00:00 2001 From: Muhammad Bilal Date: Sun, 20 Sep 2026 00:19:37 +0500 Subject: [PATCH 080/189] net: spacemit: clear TX descriptor on fragment mapping failure emac_tx_mem_map() writes TX_DESC_0_OWN into the ring descriptor for every slot beyond old_head as soon as that slot's memset()'d local copy is committed with "*tx_desc_addr = tx_desc", i.e. before the buffers for that slot have necessarily all been mapped successfully. If emac_tx_map_frag() then fails on a later fragment, the err_free_skb path calls emac_free_tx_buf() to unmap and drop the skb, but leaves the already-written descriptor memory untouched, and tx_ring->head is never advanced past old_head (the "tx_ring->head = head" store is skipped by the goto). So a slot between old_head and the rolled-back head can be left with TX_DESC_0_OWN set and buffer_addr_{1,2} pointing at DMA mappings that emac_free_tx_buf() just tore down, while software considers that slot free again. The next successful emac_tx_mem_map() call only rebuilds old_head itself; if the DMA engine auto-advances into the following descriptor once it finishes old_head's packet, it will fetch that stale, already-unmapped address. emac_tx_clean_desc() already treats emac_free_tx_buf() and clearing the descriptor as a pair when reclaiming completed descriptors; do the same in the mapping failure path. Fixes: bfec6d7f2001 ("net: spacemit: Add K1 Ethernet MAC") Signed-off-by: Muhammad Bilal Reviewed-by: Vivian Wang Reviewed-by: Troy Mitchell Link: https://patch.msgid.link/20260919191937.271202-1-meatuni001@gmail.com Signed-off-by: Paolo Abeni --- drivers/net/ethernet/spacemit/k1_emac.c | 3 +++ 1 file changed, 3 insertions(+) diff --git a/drivers/net/ethernet/spacemit/k1_emac.c b/drivers/net/ethernet/spacemit/k1_emac.c index f7f16397a2c2..d641ac26a1e8 100644 --- a/drivers/net/ethernet/spacemit/k1_emac.c +++ b/drivers/net/ethernet/spacemit/k1_emac.c @@ -803,6 +803,9 @@ static void emac_tx_mem_map(struct emac_priv *priv, struct sk_buff *skb) while (i != head) { emac_free_tx_buf(priv, i); + tx_desc_addr = &((struct emac_desc *)tx_ring->desc_addr)[i]; + memset(tx_desc_addr, 0, sizeof(*tx_desc_addr)); + if (++i == tx_ring->total_cnt) i = 0; } From ac4334522e4ba4a3b6710dd5d4cc98092824b8ca Mon Sep 17 00:00:00 2001 From: bui duc phuc Date: Fri, 18 Sep 2026 11:28:04 +0700 Subject: [PATCH 081/189] net: ethernet: ti: netcp: fix pm_runtime usage counter leak on error pm_runtime_get_sync() leaves the runtime PM usage counter incremented even when it fails, but the error path in netcp_probe() does not call pm_runtime_put_noidle() to balance it, leaking a reference each time resume fails. Use pm_runtime_resume_and_get() instead, which automatically drops the usage counter on failure, fixing the leak. Fixes: 84640e27f230 ("net: netcp: Add Keystone NetCP core ethernet driver") Signed-off-by: bui duc phuc Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260918042804.13101-1-phucduc.bui@gmail.com Signed-off-by: Paolo Abeni --- drivers/net/ethernet/ti/netcp_core.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/drivers/net/ethernet/ti/netcp_core.c b/drivers/net/ethernet/ti/netcp_core.c index eb8fc2ed05f4..4f9e20468bbb 100644 --- a/drivers/net/ethernet/ti/netcp_core.c +++ b/drivers/net/ethernet/ti/netcp_core.c @@ -2225,7 +2225,7 @@ static int netcp_probe(struct platform_device *pdev) return -ENOMEM; pm_runtime_enable(&pdev->dev); - ret = pm_runtime_get_sync(&pdev->dev); + ret = pm_runtime_resume_and_get(&pdev->dev); if (ret < 0) { dev_err(dev, "Failed to enable NETCP power-domain\n"); pm_runtime_disable(&pdev->dev); From 999e8295bc41d6ce45b8e54f88150efa96f3f01e Mon Sep 17 00:00:00 2001 From: Wentao Liang Date: Thu, 17 Sep 2026 11:08:28 +0000 Subject: [PATCH 082/189] net: hisilicon: hns_dsaf_mac: fix mdio device leak in hns_mac_register_phy() hns_dsaf_find_platform_device() returns the mdio platform device with its reference count incremented. hns_mac_register_phy() never drops that reference, so the mdio device can not be released. Release the reference on both the deferred probe and the normal path. Fixes: 1d1afa2ebf82 ("net: hns: register phy device in each mac initial sequence") Cc: stable@vger.kernel.org Signed-off-by: Wentao Liang Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260917110828.2148390-1-vulab@iscas.ac.cn Signed-off-by: Paolo Abeni --- drivers/net/ethernet/hisilicon/hns/hns_dsaf_mac.c | 3 +++ 1 file changed, 3 insertions(+) diff --git a/drivers/net/ethernet/hisilicon/hns/hns_dsaf_mac.c b/drivers/net/ethernet/hisilicon/hns/hns_dsaf_mac.c index bc6b269be299..f1cb6d56e4b9 100644 --- a/drivers/net/ethernet/hisilicon/hns/hns_dsaf_mac.c +++ b/drivers/net/ethernet/hisilicon/hns/hns_dsaf_mac.c @@ -793,6 +793,7 @@ static int hns_mac_register_phy(struct hns_mac_cb *mac_cb) dev_err(mac_cb->dev, "mac%d mdio is NULL, dsaf will probe again later\n", mac_cb->mac_id); + put_device(&pdev->dev); return -EPROBE_DEFER; } @@ -801,6 +802,8 @@ static int hns_mac_register_phy(struct hns_mac_cb *mac_cb) dev_dbg(mac_cb->dev, "mac%d register phy addr:%d\n", mac_cb->mac_id, addr); + put_device(&pdev->dev); + return rc; } From 23d42b9a3bcd55b17d3b371544fedc708c2397e9 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Th=C3=A9o=20Lebrun?= Date: Fri, 18 Sep 2026 21:53:52 +0200 Subject: [PATCH 083/189] net: macb: fix dma_alloc_coherent() leak on macb_alloc() error paths MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Fix 3 leaks in macb_alloc() error paths: - Tx buffer allocated but crossing a 4G boundary: Tx leaked. - Rx buffer allocation fails: Tx leaked. - Rx buffer allocated but crossing a 4G boundary: Tx & Rx leaked. This is because our error handling calls macb_free(bp) which in turn frees the buffers stored in bp->queues[0], but nothing has been stored in there. Fix by storing allocated buffers into bp->queues[0] ASAP. Fixes: 78d901897b3c ("net: macb: single dma_alloc_coherent() for DMA descriptors") Cc: stable@vger.kernel.org Signed-off-by: Théo Lebrun Reviewed-by: Nicolai Buchwitz Link: https://patch.msgid.link/20260918-macb-alloc-leak-v1-1-aba9a3d4f6e3@bootlin.com Signed-off-by: Paolo Abeni --- drivers/net/ethernet/cadence/macb_main.c | 14 ++++++++++++-- 1 file changed, 12 insertions(+), 2 deletions(-) diff --git a/drivers/net/ethernet/cadence/macb_main.c b/drivers/net/ethernet/cadence/macb_main.c index b8234ac4b602..8e5c034dc3a4 100644 --- a/drivers/net/ethernet/cadence/macb_main.c +++ b/drivers/net/ethernet/cadence/macb_main.c @@ -2749,14 +2749,24 @@ static int macb_alloc(struct macb *bp) size = bp->num_queues * macb_tx_ring_size_per_queue(bp); tx = dma_alloc_coherent(dev, size, &tx_dma, GFP_KERNEL); - if (!tx || upper_32_bits(tx_dma) != upper_32_bits(tx_dma + size - 1)) + if (!tx) + goto out_err; + /* Record the buffer so that the error path frees it. */ + bp->queues[0].tx_ring = tx; + bp->queues[0].tx_ring_dma = tx_dma; + if (upper_32_bits(tx_dma) != upper_32_bits(tx_dma + size - 1)) goto out_err; netdev_dbg(bp->netdev, "Allocated %zu bytes for %u TX rings at %08lx (mapped %p)\n", size, bp->num_queues, (unsigned long)tx_dma, tx); size = bp->num_queues * macb_rx_ring_size_per_queue(bp); rx = dma_alloc_coherent(dev, size, &rx_dma, GFP_KERNEL); - if (!rx || upper_32_bits(rx_dma) != upper_32_bits(rx_dma + size - 1)) + if (!rx) + goto out_err; + /* Record the buffer so that the error path frees it. */ + bp->queues[0].rx_ring = rx; + bp->queues[0].rx_ring_dma = rx_dma; + if (upper_32_bits(rx_dma) != upper_32_bits(rx_dma + size - 1)) goto out_err; netdev_dbg(bp->netdev, "Allocated %zu bytes for %u RX rings at %08lx (mapped %p)\n", size, bp->num_queues, (unsigned long)rx_dma, rx); From 261e8a37ecbaf462cdf9c336d2b2f5056088401a Mon Sep 17 00:00:00 2001 From: Jakub Kicinski Date: Fri, 18 Sep 2026 15:29:48 -0700 Subject: [PATCH 084/189] genetlink: report the real command id for dump-only ops in policy dumps The op-to-policy map a CTRL_CMD_GETPOLICY dump returns is the only way for userspace to find out which policy index belongs to which command. ctrl_dumppolicy_put_op() tags the nest with doit->cmd, but an op which only has a dumpit has no doit and every path which fills the split ops in zeroes it out, so those entries all claim to be command 0. nlctrl's own CTRL_CMD_GETPOLICY and NETDEV_CMD_QSTATS_GET are both in that group: [{'family-id': 16, 'op-policy': {'do': 0, 'dump': 0, 'op-id': 3}}, {'family-id': 16, 'op-policy': {'dump': 1, 'op-id': 0}}, ctrl_fill_info() gets this right - it uses the iterator's cmd for CTRL_ATTR_OP_ID - so the two introspection interfaces of the same family contradict each other today. Pass the command in rather than reconstructing it from doit->cmd | dumpit->cmd inside the helper, both callers already have it. Fixes: 26588edbef60 ("genetlink: support split policies in ctrl_dumppolicy_put_op()") Signed-off-by: Jakub Kicinski Link: https://patch.msgid.link/20260918222949.4190284-1-kuba@kernel.org Signed-off-by: Paolo Abeni --- net/netlink/genetlink.c | 8 +++++--- 1 file changed, 5 insertions(+), 3 deletions(-) diff --git a/net/netlink/genetlink.c b/net/netlink/genetlink.c index 41d37442f186..5cc1037d4917 100644 --- a/net/netlink/genetlink.c +++ b/net/netlink/genetlink.c @@ -1656,7 +1656,7 @@ static void *ctrl_dumppolicy_prep(struct sk_buff *skb, } static int ctrl_dumppolicy_put_op(struct sk_buff *skb, - struct netlink_callback *cb, + struct netlink_callback *cb, u32 cmd, struct genl_split_ops *doit, struct genl_split_ops *dumpit) { @@ -1677,7 +1677,7 @@ static int ctrl_dumppolicy_put_op(struct sk_buff *skb, if (!nest_pol) goto err; - nest_op = nla_nest_start(skb, doit->cmd); + nest_op = nla_nest_start(skb, cmd); if (!nest_op) goto err; @@ -1721,7 +1721,8 @@ static int ctrl_dumppolicy(struct sk_buff *skb, struct netlink_callback *cb) &doit, &dumpit))) return -ENOENT; - if (ctrl_dumppolicy_put_op(skb, cb, &doit, &dumpit)) + if (ctrl_dumppolicy_put_op(skb, cb, ctx->op, + &doit, &dumpit)) return skb->len; /* done with the per-op policy index list */ @@ -1730,6 +1731,7 @@ static int ctrl_dumppolicy(struct sk_buff *skb, struct netlink_callback *cb) while (ctx->dump_map) { if (ctrl_dumppolicy_put_op(skb, cb, + ctx->op_iter->cmd, &ctx->op_iter->doit, &ctx->op_iter->dumpit)) return skb->len; From a87529034b9c3c3f3bb71c33e2d6d94658e06383 Mon Sep 17 00:00:00 2001 From: Jakub Kicinski Date: Fri, 18 Sep 2026 15:29:49 -0700 Subject: [PATCH 085/189] selftests: net: nl_nlctrl: check the op ids in the policy map Validate that the op map in the policy dump is correct. Signed-off-by: Jakub Kicinski Link: https://patch.msgid.link/20260918222949.4190284-2-kuba@kernel.org Signed-off-by: Paolo Abeni --- tools/testing/selftests/net/nl_nlctrl.py | 116 +++++++++++++++++++---- 1 file changed, 99 insertions(+), 17 deletions(-) diff --git a/tools/testing/selftests/net/nl_nlctrl.py b/tools/testing/selftests/net/nl_nlctrl.py index fe1f66dc9435..237b3d273260 100755 --- a/tools/testing/selftests/net/nl_nlctrl.py +++ b/tools/testing/selftests/net/nl_nlctrl.py @@ -9,40 +9,86 @@ from lib.py import ksft_run, ksft_exit from lib.py import ksft_eq, ksft_ge, ksft_true, ksft_in, ksft_not_in from lib.py import NetdevFamily, EthtoolFamily, NlctrlFamily +# Families we can expect to always be around, and which between them +# cover ops with a do, with a dump, and with both. +FAMILIES = ('nlctrl', 'netdev') -def getfamily_do(ctrl) -> None: - """Query a single family by name and validate its ops.""" - fam = ctrl.getfamily({'family-name': 'netdev'}) - ksft_eq(fam['family-name'], 'netdev') + +def _get_ops(ctrl, name): + """Get the ops of a family, keyed by command id.""" + fam = ctrl.getfamily({'family-name': name}) + ksft_eq(fam['family-name'], name) ksft_true(fam['family-id'] > 0) # The format of ops is quite odd, [{$idx: {"id"...}}, {$idx: {"id"...}}] # Discard the indices and re-key by command id. ops_by_id = {v['id']: v for op in fam['ops'] for v in op.values()} - ksft_eq(len(ops_by_id), len(fam['ops'])) + ksft_eq(len(ops_by_id), len(fam['ops']), + comment=f"{name} lists a command twice") + return ops_by_id - # All ops should have a policy (either do or dump has one) - for op in ops_by_id.values(): - ksft_in('cmd-cap-haspol', op['flags'], - comment=f"op {op['id']} missing haspol") + +def _get_policy_map(ctrl, req): + """ + The policy map in the Netlink replies looks like this: + + [{'family-id': 16, 'op-policy': {'do': 0, 'dump': 0, 'op-id': 3}}, + {'family-id': 16, 'op-policy': {'dump': 1, 'op-id': 4}}, ...] + + Return the mapping: + + {3:{'do','dump'}, 4:{'dump'}} + + The policy itself is discarded here, only return which command has policy. + """ + pol_map = {} + for msg in ctrl.getpolicy(req, dump=True): + if 'op-policy' not in msg: + continue + modes = dict(msg['op-policy']) + cmd = modes.pop('op-id') + ksft_not_in(cmd, pol_map, comment=f"command {cmd} reported twice") + pol_map[cmd] = set(modes.keys()) + return pol_map + + +def getfamily_do(ctrl) -> None: + """Query single families by name and validate their ops.""" + ops = {name: _get_ops(ctrl, name) for name in FAMILIES} + + for name, ops_by_id in ops.items(): + for op in ops_by_id.values(): + # All ops in nlctrl and netdev have a policy + ksft_in('cmd-cap-haspol', op['flags'], + comment=f"{name} op {op['id']} missing haspol") + ksft_true(op['flags'] & {'cmd-cap-do', 'cmd-cap-dump'}, + comment=f"{name} op {op['id']} has no handler") + + # nlctrl getfamily (id 3) does both, getpolicy (id 10) is dump-only + ksft_in('cmd-cap-do', ops['nlctrl'][3]['flags']) + ksft_in('cmd-cap-dump', ops['nlctrl'][3]['flags']) + ksft_not_in('cmd-cap-do', ops['nlctrl'][10]['flags']) + ksft_in('cmd-cap-dump', ops['nlctrl'][10]['flags']) + + netdev = ops['netdev'] # dev-get (id 1) should support both do and dump - ksft_in('cmd-cap-do', ops_by_id[1]['flags']) - ksft_in('cmd-cap-dump', ops_by_id[1]['flags']) + ksft_in('cmd-cap-do', netdev[1]['flags']) + ksft_in('cmd-cap-dump', netdev[1]['flags']) # qstats-get (id 12) is dump-only - ksft_not_in('cmd-cap-do', ops_by_id[12]['flags']) - ksft_in('cmd-cap-dump', ops_by_id[12]['flags']) + ksft_not_in('cmd-cap-do', netdev[12]['flags']) + ksft_in('cmd-cap-dump', netdev[12]['flags']) # napi-set (id 14) is do-only and requires admin - ksft_in('cmd-cap-do', ops_by_id[14]['flags']) - ksft_not_in('cmd-cap-dump', ops_by_id[14]['flags']) - ksft_in('admin-perm', ops_by_id[14]['flags']) + ksft_in('cmd-cap-do', netdev[14]['flags']) + ksft_not_in('cmd-cap-dump', netdev[14]['flags']) + ksft_in('admin-perm', netdev[14]['flags']) # Notification-only commands (dev-add/del/change-ntf etc.) must # not appear in the ops list since they have no do/dump handlers. for ntf_id in [2, 3, 4, 6, 7, 8]: - ksft_not_in(ntf_id, ops_by_id, + ksft_not_in(ntf_id, netdev, comment=f"ntf-only cmd {ntf_id} should not be in ops") @@ -103,6 +149,41 @@ def getpolicy_dump(_ctrl) -> None: comment="linkinfo-set should not have a dump policy") +def getpolicy_op_map(ctrl) -> None: + """Check the op-to-policy map consistency. Each op with 'haspol' flag + has to have a policy. The policy back-references must name only + real ops that exist, have given modes (do vs dump) and have 'haspol'. + """ + for name in FAMILIES: + ops_by_id = _get_ops(ctrl, name) + haspol = {cmd for cmd, op in ops_by_id.items() + if 'cmd-cap-haspol' in op['flags']} + + pol_map = _get_policy_map(ctrl, {'family-name': name}) + ksft_eq(set(pol_map), haspol, + comment=f"{name} policy map does not match the op list") + + # Walk the op list rather than the map, the map may be missing + # the very op we are after. Asking for a command the family does + # not have is an error, so it must not come from the map either. + for cmd in sorted(haspol): + modes = pol_map.get(cmd, set()) + + # The kernel only reports a mode the op actually has. + if 'do' in modes: + ksft_in('cmd-cap-do', ops_by_id[cmd]['flags'], + comment=f"{name} cmd {cmd} has no do") + if 'dump' in modes: + ksft_in('cmd-cap-dump', ops_by_id[cmd]['flags'], + comment=f"{name} cmd {cmd} has no dump") + + # Asking for one op builds the map in a different place in + # the kernel, it has to report what the full dump did. + single = _get_policy_map(ctrl, {'family-name': name, 'op': cmd}) + ksft_eq(single, {cmd: modes}, + comment=f"{name} cmd {cmd} policy differs from the dump") + + def getpolicy_by_op(_ctrl) -> None: """Query policy for specific ops, check attr names are resolved.""" ndev = NetdevFamily() @@ -122,6 +203,7 @@ def main() -> None: ksft_run([getfamily_do, getfamily_dump, getpolicy_dump, + getpolicy_op_map, getpolicy_by_op], args=(ctrl, )) ksft_exit() From 9892d71cf0ce3ff3d4fed2d9a3968fd4feb1c918 Mon Sep 17 00:00:00 2001 From: Coia Prant Date: Sun, 20 Sep 2026 01:20:21 +0800 Subject: [PATCH 086/189] net: pcs: xpcs: fix clock reference leak on xpcs_init_clks failure xpcs_init_clks() takes references with clk_bulk_get_optional() and then enables them with clk_bulk_prepare_enable(). If the enable step fails, the function returns without dropping the references. xpcs_create() handles the failure through out_free_data, which calls xpcs_free_data() but never xpcs_clear_clks(), so the clk references are leaked. Add the missing clk_bulk_put() on the enable failure path. The prepare/enable side is already rolled back by clk_bulk_prepare_enable() itself. Fixes: f6bb3e9d98c2 ("net: pcs: xpcs: Add Synopsys DW xPCS platform device driver") Signed-off-by: Coia Prant Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260919172021.2336748-1-coiaprant@gmail.com Signed-off-by: Paolo Abeni --- drivers/net/pcs/pcs-xpcs.c | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/drivers/net/pcs/pcs-xpcs.c b/drivers/net/pcs/pcs-xpcs.c index 0337e2bcc012..b415b93d77c1 100644 --- a/drivers/net/pcs/pcs-xpcs.c +++ b/drivers/net/pcs/pcs-xpcs.c @@ -1545,8 +1545,10 @@ static int xpcs_init_clks(struct dw_xpcs *xpcs) return dev_err_probe(dev, ret, "Failed to get clocks\n"); ret = clk_bulk_prepare_enable(DW_XPCS_NUM_CLKS, xpcs->clks); - if (ret) + if (ret) { + clk_bulk_put(DW_XPCS_NUM_CLKS, xpcs->clks); return dev_err_probe(dev, ret, "Failed to enable clocks\n"); + } return 0; } From 17741334d00bf5ebd37f8c1c36bc9c146a351deb Mon Sep 17 00:00:00 2001 From: Wentao Liang Date: Thu, 17 Sep 2026 11:58:11 +0000 Subject: [PATCH 087/189] net: usb: lan78xx: Fix URB reference leak in lan78xx_submit_deferred_urbs() usb_get_from_anchor() hands over a reference to the URB, which the caller must release. lan78xx_submit_deferred_urbs() never does, so every deferred Tx URB keeps an extra reference: the counter grows on each suspend/resume cycle and the URBs are never freed when the buffers are released. Drop the reference after submitting, and on the path that drops the packet instead of submitting it. Fixes: 5f4cc6e25148 ("lan78xx: Fix race conditions in suspend/resume handling") Cc: stable@vger.kernel.org Signed-off-by: Wentao Liang Link: https://patch.msgid.link/20260917115811.2150119-1-vulab@iscas.ac.cn Signed-off-by: Paolo Abeni --- drivers/net/usb/lan78xx.c | 2 ++ 1 file changed, 2 insertions(+) diff --git a/drivers/net/usb/lan78xx.c b/drivers/net/usb/lan78xx.c index cb782d81d84f..5655941f1478 100644 --- a/drivers/net/usb/lan78xx.c +++ b/drivers/net/usb/lan78xx.c @@ -5239,10 +5239,12 @@ static bool lan78xx_submit_deferred_urbs(struct lan78xx_net *dev) !netif_carrier_ok(dev->net) || pipe_halted) { lan78xx_release_tx_buf(dev, skb); + usb_put_urb(urb); continue; } ret = usb_submit_urb(urb, GFP_ATOMIC); + usb_put_urb(urb); if (ret == 0) { netif_trans_update(dev->net); From 35e6f970f553954d92ba20afa885139a8e7dd0d7 Mon Sep 17 00:00:00 2001 From: Bernardo Soares Date: Fri, 18 Sep 2026 10:59:30 +0100 Subject: [PATCH 088/189] net/mlx5: Bridge, don't fail switchdev events of sibling eswitch ports mlx5 registers the bridge offload switchdev notifiers once per eswitch instance, but the notifier chains are global, so every instance sees every event and must filter out the ones that aren't its own. The existing filter, mlx5_esw_bridge_dev_same_hw(), only checks that the event netdevice sits on the same HCA - intentional for merged eswitch, where one bridge can span representors of several eswitches on one HCA - but same-HCA doesn't mean the instance actually has that port: peer ports are only created reactively from NETDEV_CHANGEUPPER, so an instance brought up after a sibling PF's port was already enslaved has none. The port object and attribute handlers claim the event anyway once same-HW passes, then fail the port lookup and return -EINVAL, which gets reported to user space even though the owning instance already handled it (e.g. "bridge vlan add ... RTNETLINK answers: Invalid argument"). Fix by filtering on the tracked port instead. The same gap exists in the generic recursive lower-device walk used by attribute changes on a bridge with more than one representor enslaved directly: mlx5_esw_bridge_lower_rep_vport_num_vhca_id_get() is entered with the bridge master netdevice, falls through to its generic netdev_for_each_lower_dev() loop, and returns as soon as the recursion into any one lower device yields a non-NULL rep - the underlying base case, mlx5_esw_bridge_rep_vport_num_vhca_id_get(), only checks mlx5_esw_bridge_dev_same_hw(), not ownership by the calling instance's br_offloads. mlx5_esw_bridge_lag_rep_get(), used for the LAG-master case, already filters on mlx5_esw_bridge_dev_same_esw() per candidate and so cannot select a sibling's rep; it is not the source of this bug. On a merged-eswitch HCA with a bridge spanning representors of more than one eswitch instance directly, the walk can return a sibling's rep instead of continuing to the one the calling instance actually owns, so the attribute change fails the same way as above. Fix by checking mlx5_esw_bridge_port_exists() at the point each rep is picked, same as the previous fix did for the notifier filter. Fixes: c358ea1741bc ("net/mlx5: Bridge, allow merged eswitch connectivity") Signed-off-by: Bernardo Soares Cc: Vlad Buslov Cc: Saeed Mahameed Reviewed-by: Mark Bloch Link: https://patch.msgid.link/20260918095931.29792-2-bsoares.it@gmail.com Signed-off-by: Jakub Kicinski --- .../mellanox/mlx5/core/en/rep/bridge.c | 45 +++++++++++++++---- .../ethernet/mellanox/mlx5/core/esw/bridge.c | 6 +++ .../ethernet/mellanox/mlx5/core/esw/bridge.h | 2 + 3 files changed, 44 insertions(+), 9 deletions(-) diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en/rep/bridge.c b/drivers/net/ethernet/mellanox/mlx5/core/en/rep/bridge.c index baac38bece14..4b7b0a0fc2b2 100644 --- a/drivers/net/ethernet/mellanox/mlx5/core/en/rep/bridge.c +++ b/drivers/net/ethernet/mellanox/mlx5/core/en/rep/bridge.c @@ -85,9 +85,16 @@ mlx5_esw_bridge_lower_rep_vport_num_vhca_id_get(struct net_device *dev, struct m struct net_device *lower_dev; struct list_head *iter; - if (netif_is_lag_master(dev) || mlx5e_eswitch_rep(dev)) - return mlx5_esw_bridge_rep_vport_num_vhca_id_get(dev, esw, vport_num, - esw_owner_vhca_id); + if (netif_is_lag_master(dev) || mlx5e_eswitch_rep(dev)) { + struct net_device *rep; + + rep = mlx5_esw_bridge_rep_vport_num_vhca_id_get(dev, esw, vport_num, + esw_owner_vhca_id); + if (rep && !mlx5_esw_bridge_port_exists(*vport_num, *esw_owner_vhca_id, + esw->br_offloads)) + return NULL; + return rep; + } netdev_for_each_lower_dev(dev, lower_dev, iter) { struct net_device *rep; @@ -104,6 +111,28 @@ mlx5_esw_bridge_lower_rep_vport_num_vhca_id_get(struct net_device *dev, struct m return NULL; } +static bool mlx5_esw_bridge_rep_port_lookup(struct net_device *dev, + struct mlx5_esw_bridge_offloads *br_offloads, + u16 *vport_num, u16 *esw_owner_vhca_id) +{ + if (!mlx5_esw_bridge_rep_vport_num_vhca_id_get(dev, br_offloads->esw, vport_num, + esw_owner_vhca_id)) + return false; + + return mlx5_esw_bridge_port_exists(*vport_num, *esw_owner_vhca_id, br_offloads); +} + +static bool mlx5_esw_bridge_lower_rep_port_lookup(struct net_device *dev, + struct mlx5_esw_bridge_offloads *br_offloads, + u16 *vport_num, u16 *esw_owner_vhca_id) +{ + if (!mlx5_esw_bridge_lower_rep_vport_num_vhca_id_get(dev, br_offloads->esw, vport_num, + esw_owner_vhca_id)) + return false; + + return mlx5_esw_bridge_port_exists(*vport_num, *esw_owner_vhca_id, br_offloads); +} + static bool mlx5_esw_bridge_is_local(struct net_device *dev, struct net_device *rep, struct mlx5_eswitch *esw) { @@ -218,8 +247,7 @@ mlx5_esw_bridge_port_obj_add(struct net_device *dev, u16 vport_num, esw_owner_vhca_id; int err; - if (!mlx5_esw_bridge_rep_vport_num_vhca_id_get(dev, br_offloads->esw, &vport_num, - &esw_owner_vhca_id)) + if (!mlx5_esw_bridge_rep_port_lookup(dev, br_offloads, &vport_num, &esw_owner_vhca_id)) return 0; port_obj_info->handled = true; @@ -251,8 +279,7 @@ mlx5_esw_bridge_port_obj_del(struct net_device *dev, const struct switchdev_obj_port_mdb *mdb; u16 vport_num, esw_owner_vhca_id; - if (!mlx5_esw_bridge_rep_vport_num_vhca_id_get(dev, br_offloads->esw, &vport_num, - &esw_owner_vhca_id)) + if (!mlx5_esw_bridge_rep_port_lookup(dev, br_offloads, &vport_num, &esw_owner_vhca_id)) return 0; port_obj_info->handled = true; @@ -283,8 +310,8 @@ mlx5_esw_bridge_port_obj_attr_set(struct net_device *dev, u16 vport_num, esw_owner_vhca_id; int err = 0; - if (!mlx5_esw_bridge_lower_rep_vport_num_vhca_id_get(dev, br_offloads->esw, &vport_num, - &esw_owner_vhca_id)) + if (!mlx5_esw_bridge_lower_rep_port_lookup(dev, br_offloads, &vport_num, + &esw_owner_vhca_id)) return 0; port_attr_info->handled = true; diff --git a/drivers/net/ethernet/mellanox/mlx5/core/esw/bridge.c b/drivers/net/ethernet/mellanox/mlx5/core/esw/bridge.c index 87b5fd349594..ac90ccda1272 100644 --- a/drivers/net/ethernet/mellanox/mlx5/core/esw/bridge.c +++ b/drivers/net/ethernet/mellanox/mlx5/core/esw/bridge.c @@ -1686,6 +1686,12 @@ int mlx5_esw_bridge_vport_peer_unlink(struct net_device *br_netdev, u16 vport_nu extack); } +bool mlx5_esw_bridge_port_exists(u16 vport_num, u16 esw_owner_vhca_id, + struct mlx5_esw_bridge_offloads *br_offloads) +{ + return mlx5_esw_bridge_port_lookup(vport_num, esw_owner_vhca_id, br_offloads); +} + int mlx5_esw_bridge_port_vlan_add(u16 vport_num, u16 esw_owner_vhca_id, u16 vid, u16 flags, struct mlx5_esw_bridge_offloads *br_offloads, struct netlink_ext_ack *extack) diff --git a/drivers/net/ethernet/mellanox/mlx5/core/esw/bridge.h b/drivers/net/ethernet/mellanox/mlx5/core/esw/bridge.h index d6f539161993..a4e59cc21089 100644 --- a/drivers/net/ethernet/mellanox/mlx5/core/esw/bridge.h +++ b/drivers/net/ethernet/mellanox/mlx5/core/esw/bridge.h @@ -80,6 +80,8 @@ int mlx5_esw_bridge_vlan_proto_set(u16 vport_num, u16 esw_owner_vhca_id, u16 pro struct mlx5_esw_bridge_offloads *br_offloads); int mlx5_esw_bridge_mcast_set(u16 vport_num, u16 esw_owner_vhca_id, bool enable, struct mlx5_esw_bridge_offloads *br_offloads); +bool mlx5_esw_bridge_port_exists(u16 vport_num, u16 esw_owner_vhca_id, + struct mlx5_esw_bridge_offloads *br_offloads); int mlx5_esw_bridge_port_vlan_add(u16 vport_num, u16 esw_owner_vhca_id, u16 vid, u16 flags, struct mlx5_esw_bridge_offloads *br_offloads, struct netlink_ext_ack *extack); From 2e51097c982b7b22382fda3202ba29f0ea33e8c0 Mon Sep 17 00:00:00 2001 From: Bernardo Soares Date: Fri, 18 Sep 2026 10:59:31 +0100 Subject: [PATCH 089/189] net/mlx5: Bridge, don't fail unlink of untracked/unsupported peer ports mlx5_esw_bridge_vport_unlink() returns -EINVAL when the port isn't tracked by this instance's br_offloads. This is reachable on a sibling instance that registered its notifier after the port was already enslaved: it never saw the NETDEV_CHANGEUPPER link event, so peer_link() never created a peer port for it, but it does see the later unlink event and fails. Return 0 instead, and give mlx5_esw_bridge_vport_peer_unlink() the same merged_eswitch capability guard peer_link() already has, since without it peer_link() likewise never creates a port to unlink. This also matters beyond the -EINVAL itself: mlx5_esw_bridge_switchdev_port_event() runs on the per-netns netdev_chain, and notifier_from_errno(-EINVAL) sets NOTIFY_STOP_MASK, which call_netdevice_notifiers_info() checks to stop calling further listeners on that chain - so the old -EINVAL silently dropped the event for any listener registered later on the same chain, even though none of it was visible to user space since __netdev_upper_dev_unlink() discards the return value. Fixes: c358ea1741bc ("net/mlx5: Bridge, allow merged eswitch connectivity") Signed-off-by: Bernardo Soares Reviewed-by: Mark Bloch Link: https://patch.msgid.link/20260918095931.29792-3-bsoares.it@gmail.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/mellanox/mlx5/core/esw/bridge.c | 9 +++++---- 1 file changed, 5 insertions(+), 4 deletions(-) diff --git a/drivers/net/ethernet/mellanox/mlx5/core/esw/bridge.c b/drivers/net/ethernet/mellanox/mlx5/core/esw/bridge.c index ac90ccda1272..b4cf3c5ac0dd 100644 --- a/drivers/net/ethernet/mellanox/mlx5/core/esw/bridge.c +++ b/drivers/net/ethernet/mellanox/mlx5/core/esw/bridge.c @@ -1649,10 +1649,8 @@ int mlx5_esw_bridge_vport_unlink(struct net_device *br_netdev, u16 vport_num, int err; port = mlx5_esw_bridge_port_lookup(vport_num, esw_owner_vhca_id, br_offloads); - if (!port) { - NL_SET_ERR_MSG_MOD(extack, "Port is not attached to any bridge"); - return -EINVAL; - } + if (!port) + return 0; if (port->bridge->ifindex != br_netdev->ifindex) { NL_SET_ERR_MSG_MOD(extack, "Port is attached to another bridge"); return -EINVAL; @@ -1682,6 +1680,9 @@ int mlx5_esw_bridge_vport_peer_unlink(struct net_device *br_netdev, u16 vport_nu struct mlx5_esw_bridge_offloads *br_offloads, struct netlink_ext_ack *extack) { + if (!MLX5_CAP_ESW(br_offloads->esw->dev, merged_eswitch)) + return 0; + return mlx5_esw_bridge_vport_unlink(br_netdev, vport_num, esw_owner_vhca_id, br_offloads, extack); } From 0bf6bb567f0edaa771e7dd208ef98da50e6a4485 Mon Sep 17 00:00:00 2001 From: Wentao Liang Date: Thu, 17 Sep 2026 11:31:31 +0000 Subject: [PATCH 090/189] net/mlx5: Fix rev_entry reference leak in mlx5_tc_ct_shared_counter_get() When the reverse entry is found but its counter is already being released, refcount_inc_not_zero() fails and the reference taken by mlx5_tc_ct_entry_get() is never dropped before falling through to create_counter. Drop it so the reverse entry is not kept alive forever by a shared counter lookup that did not use it. Fixes: 1edae2335adf ("net/mlx5e: CT: Use the same counter for both directions") Cc: stable@vger.kernel.org Signed-off-by: Wentao Liang Reviewed-by: Tariq Toukan Link: https://patch.msgid.link/20260917113131.2149024-1-vulab@iscas.ac.cn Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/mellanox/mlx5/core/en/tc_ct.c | 3 +++ 1 file changed, 3 insertions(+) diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en/tc_ct.c b/drivers/net/ethernet/mellanox/mlx5/core/en/tc_ct.c index 6c87a1c7db09..c1841ce74d9a 100644 --- a/drivers/net/ethernet/mellanox/mlx5/core/en/tc_ct.c +++ b/drivers/net/ethernet/mellanox/mlx5/core/en/tc_ct.c @@ -1082,6 +1082,9 @@ mlx5_tc_ct_shared_counter_get(struct mlx5_tc_ct_priv *ct_priv, spin_unlock_bh(&ct_priv->ht_lock); + if (rev_entry) + mlx5_tc_ct_entry_put(rev_entry); + create_counter: shared_counter = mlx5_tc_ct_counter_create(ct_priv); From 4498467a8af06cfa3d71cb04bd7c4170dec8f449 Mon Sep 17 00:00:00 2001 From: Aohan Mei Date: Mon, 21 Sep 2026 17:37:04 +0800 Subject: [PATCH 091/189] sctp: discard the rest of the packet on a stale-cookie error When an association is in COOKIE-ECHOED state and the peer sends a bundled [ERROR(Stale Cookie)][DATA] packet from one of its non-primary addresses, processing the ERROR chunk takes the non-fatal stale-cookie retry path sctp_sf_do_5_2_6_stale(), which queues SCTP_CMD_DEL_NON_PRIMARY while keeping the association alive. sctp_cmd_del_non_primary() removes every non-primary transport - including the very transport this packet arrived on, which is still referenced by the receive lookup and shared by all chunks of the packet via chunk->transport. sctp_assoc_rm_peer() does redirect asoc->peer.last_data_from away from the removed transport, but right afterwards the bundled DATA chunk makes sctp_assoc_bh_rcv() re-register asoc->peer.last_data_from = chunk->transport unconditionally, undoing the redirection with the just-removed transport. Once the packet is done, the receive reference is dropped and the transport is RCU-freed, while the surviving association keeps the dangling last_data_from. A later FWD-TSN (or the delayed SACK timer) makes sctp_gen_sack() dereference it (->param_flags and friends), and sctp_make_sack()/sctp_outq_select_transport() may write to the freed object and link it into the live transport list. This is a use-after-free triggerable by any malicious SCTP peer (or a local unprivileged user acting as one) with no capabilities required: BUG: KASAN: slab-use-after-free in sctp_do_sm+0x498a/0x5660 Read of size 4 at addr ffff88800e1e356c by task poc/115 Call Trace: sctp_do_sm <- sctp_assoc_bh_rcv <- sctp_inq_push <- sctp_rcv <- ip_protocol_deliver_rcu <- ip_rcv Allocated: sctp_transport_new <- sctp_assoc_add_peer <- sctp_process_init (INIT-ACK processing) Freed: kfree <- sctp_transport_destroy_rcu <- rcu_core (call_rcu queued by sctp_transport_put at end of sctp_rcv) The buggy address is located 364 bytes inside of freed 1024-byte region [ffff88800e1e3400, ffff88800e1e3800), cache kmalloc-1k Note that commit 03a9d10ecf71 ("sctp: drop a chunk if its transport was removed") only covers the window between the receive lookup and the chunk processing (e.g. an ASCONF DEL-IP racing the socket backlog); here the transport is removed *while* the packet is being processed, by an earlier chunk of the same packet, so the drop in sctp_inq_push() does not reach this path. Verified with the bundled [ERROR(Stale Cookie)][DATA] + FWD-TSN reproducer: the KASAN report above still fires with that commit applied, and is gone with this patch on top. Fix it by discarding the rest of the packet on this path, as suggested by Xin. After the stale-cookie ERROR has sent the association back to COOKIE-WAIT and removed the non-primary transports, the remaining chunks of the packet can only run against the restarted handshake while referencing the removed arrival transport through chunk->transport: besides the last_data_from registration above, sctp_cmd_setup_t2() and the sctp_make_*() reply builders would also copy that pointer into association-lifetime state that sctp_assoc_rm_peer() has already sanitized. Let the peer retransmit them, in line with what sctp_inq_push() does for chunks whose transport was removed before processing. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Suggested-by: Xin Long Reported-by: TencentOS Corvus AI Cc: stable@vger.kernel.org Signed-off-by: Aohan Mei Acked-by: Xin Long Link: https://patch.msgid.link/20260921093707.1432184-1-ljp1205831794@gmail.com Signed-off-by: Jakub Kicinski --- net/sctp/sm_statefuns.c | 2 ++ 1 file changed, 2 insertions(+) diff --git a/net/sctp/sm_statefuns.c b/net/sctp/sm_statefuns.c index c701cff6ea67..8dc65498763c 100644 --- a/net/sctp/sm_statefuns.c +++ b/net/sctp/sm_statefuns.c @@ -2654,6 +2654,8 @@ static enum sctp_disposition sctp_sf_do_5_2_6_stale( sctp_add_cmd_sf(commands, SCTP_CMD_REPLY, SCTP_CHUNK(reply)); + sctp_add_cmd_sf(commands, SCTP_CMD_DISCARD_PACKET, SCTP_NULL()); + return SCTP_DISPOSITION_CONSUME; nomem: From 07e1a9408b6c2f9d0cfb757b67dabb52da7a32b2 Mon Sep 17 00:00:00 2001 From: Willem de Bruijn Date: Fri, 18 Sep 2026 20:47:28 -0400 Subject: [PATCH 092/189] virtio_net: copy zerocopy frags in start_xmit without NAPI Virtio-net without NAPI frees completed skbs lazily on the next start_xmit. Senders waiting for in-flight zerocopy buffers can deadlock if they cannot transmit more packets, as then no completed packets will be freed. When !use_napi, virtio-net already calls skb_orphan to avoid waiting up for transmitted skbs to be freed. For zerocopy packets that require deep copying on orphan (i.e. those that do not set SKBFL_DONT_ORPHAN, such as PACKET_TX_RING), call skb_orphan_frags before orphaning to release the buffers. This fixes the tpacket_snd slot reuse bug on skb_orphan for virtio-net, and prevents PACKET_TX_RING from running out of slots. This fix also touches vhost_net zerocopy packets, which also do not set SKBFL_DONT_ORPHAN. This is fine: vhost_net packets only encounter virtio-net in nested virtualization, and only if napi_tx is explicitly disabled (it has been default-enabled since Linux 4.12). In that rare case, copying the frags is desirable anyway to prevent holding guest descriptors pinned across unbounded intervals. This is a prerequisite for the next patch, which converts PACKET_TX_RING to standard zerocopy completion. Without this patch first, a bounded ring sender can stall indefinitely behind a virtio-net virtqueue that cannot reclaim. Fixes: 5cd8d46ea156 ("packet: copy user buffers before orphan or clone") Cc: stable@vger.kernel.org Cc: mst@redhat.com Cc: jasowangio@gmail.com Signed-off-by: Willem de Bruijn Link: https://patch.msgid.link/20260919004748.1463985-2-willemdebruijn.kernel@gmail.com Signed-off-by: Jakub Kicinski --- drivers/net/virtio_net.c | 9 +++++++++ 1 file changed, 9 insertions(+) diff --git a/drivers/net/virtio_net.c b/drivers/net/virtio_net.c index e34c52d059d3..bf82ef9874ab 100644 --- a/drivers/net/virtio_net.c +++ b/drivers/net/virtio_net.c @@ -3349,6 +3349,14 @@ static netdev_tx_t start_xmit(struct sk_buff *skb, struct net_device *dev) else virtqueue_disable_cb(sq->vq); + if (!use_napi && + unlikely(skb_orphan_frags(skb, GFP_ATOMIC))) { + DEV_STATS_INC(dev, tx_dropped); + dev_kfree_skb_any(skb); + kick = !xmit_more || netif_xmit_stopped(txq); + goto kick_vq; + } + /* timestamp packet in software */ skb_tx_timestamp(skb); @@ -3381,6 +3389,7 @@ static netdev_tx_t start_xmit(struct sk_buff *skb, struct net_device *dev) kick = use_napi ? __netdev_tx_sent_queue(txq, skb->len, xmit_more) : !xmit_more || netif_xmit_stopped(txq); +kick_vq: if (kick) { if (virtqueue_kick_prepare(sq->vq) && virtqueue_notify(sq->vq)) { u64_stats_update_begin(&sq->stats.syncp); From 9518405613863d0bf0700927a21367f5942cf058 Mon Sep 17 00:00:00 2001 From: Willem de Bruijn Date: Fri, 18 Sep 2026 20:47:29 -0400 Subject: [PATCH 093/189] packet: use ubuf_info completion for TX_RING packets tpacket_snd sends skbs with frags pointing into its ring slots. Slots are released when skb->destructor is called. A call to skb_orphan calls skb->destructor before the skb is freed. This can cause the slot to be reused while still linked into the skb. Switch to standard zerocopy completion (ubuf_info) so the slot is only released once all references to the payload are freed or copied. Restore skb->destructor to standard sock_wfree. The ubuf_info completion callback can be called with a NULL skb, but only from net_zcopy_put and related API, used by zerocopy implementations that hold their own reference on the uarg, such as MSG_ZEROCOPY. This uarg is only ever completed from skb_zcopy_clear, so skb is always set. To prevent userspace from aliasing in-flight state on shared ring slots, allocate tpacket_uarg per packet, rather than per slot. This adds a small allocation to the transmit path. Use standard kmalloc to allow backporting to stable kernels. The uarg holds an sk_wmem_alloc reference, rather than an sk_refcnt reference. packet_free_tx_ring waits on sk_wmem_alloc before freeing the ring pages. Always allocate vec->deferred for tx_ring so page-backed rings also wait on sk_wmem_alloc when skb_copy_ubufs drops page refs before calling tpacket_ubuf_complete. Drop the tx_ring.pg_vec test that tpacket_destruct_skb performed before accessing the slot. The sk_wmem_alloc reference now guarantees that the slot is valid. The test is also not sufficient by itself, as it reads pg_vec without pg_vec_lock, so it can race with packet_set_ring. As a result a slot is released when its payload is copied, which can be before transmission (e.g., in skb_orphan_frags_rx). If copied before skb_tx_timestamp() is called, no slot timestamp is recorded, similar to when skb_orphan() was called early in the datapath before this patch. Revert the now unused previous skb_zcopy_.._nouarg infra. Depends on commit 992cc9f94ca9 ("net/packet: defer vmalloc TX_RING free until skbs finish"). Reported-by: Katherine Leaver Reported-by: Bjoern Doebel Closes: https://lore.kernel.org/netdev/20260909085542.3370986-1-doebel@amazon.de/ Fixes: 5cd8d46ea156 ("packet: copy user buffers before orphan or clone") Cc: stable@vger.kernel.org Signed-off-by: Willem de Bruijn Link: https://patch.msgid.link/20260919004748.1463985-3-willemdebruijn.kernel@gmail.com Signed-off-by: Jakub Kicinski --- include/linux/skbuff.h | 19 +------- net/packet/af_packet.c | 102 ++++++++++++++++++++++++++--------------- 2 files changed, 65 insertions(+), 56 deletions(-) diff --git a/include/linux/skbuff.h b/include/linux/skbuff.h index 421f6fc45451..b14d6be7370b 100644 --- a/include/linux/skbuff.h +++ b/include/linux/skbuff.h @@ -1834,22 +1834,6 @@ static inline void skb_zcopy_set(struct sk_buff *skb, struct ubuf_info *uarg, } } -static inline void skb_zcopy_set_nouarg(struct sk_buff *skb, void *val) -{ - skb_shinfo(skb)->destructor_arg = (void *)((uintptr_t) val | 0x1UL); - skb_shinfo(skb)->flags |= SKBFL_ZEROCOPY_FRAG; -} - -static inline bool skb_zcopy_is_nouarg(struct sk_buff *skb) -{ - return (uintptr_t) skb_shinfo(skb)->destructor_arg & 0x1UL; -} - -static inline void *skb_zcopy_get_nouarg(struct sk_buff *skb) -{ - return (void *)((uintptr_t) skb_shinfo(skb)->destructor_arg & ~0x1UL); -} - static inline void net_zcopy_put(struct ubuf_info *uarg) { if (uarg) @@ -1872,8 +1856,7 @@ static inline void skb_zcopy_clear(struct sk_buff *skb, bool zerocopy_success) struct ubuf_info *uarg = skb_zcopy(skb); if (uarg) { - if (!skb_zcopy_is_nouarg(skb)) - uarg->ops->complete(skb, uarg, zerocopy_success); + uarg->ops->complete(skb, uarg, zerocopy_success); skb_shinfo(skb)->flags &= ~SKBFL_ALL_ZEROCOPY; } diff --git a/net/packet/af_packet.c b/net/packet/af_packet.c index 50cae32ae269..64b501db660a 100644 --- a/net/packet/af_packet.c +++ b/net/packet/af_packet.c @@ -2530,26 +2530,6 @@ static int tpacket_rcv(struct sk_buff *skb, struct net_device *dev, goto drop_n_restore; } -static void tpacket_destruct_skb(struct sk_buff *skb) -{ - struct packet_sock *po = pkt_sk(skb->sk); - - if (likely(po->tx_ring.pg_vec)) { - void *ph; - __u32 ts; - - ph = skb_zcopy_get_nouarg(skb); - - ts = __packet_set_timestamp(po, ph, skb); - __packet_set_status(po, ph, TP_STATUS_AVAILABLE | ts); - - packet_dec_pending(&po->tx_ring); - complete(&po->skb_completion); - } - - sock_wfree(skb); -} - static int __packet_snd_vnet_parse(struct virtio_net_hdr *vnet_hdr, size_t len) { if ((vnet_hdr->flags & VIRTIO_NET_HDR_F_NEEDS_CSUM) && @@ -2589,27 +2569,56 @@ static int packet_snd_vnet_parse(struct msghdr *msg, size_t *len, return 0; } +struct tpacket_uarg { + struct ubuf_info ubuf; + struct packet_sock *po; + void *ph; +}; + +static void tpacket_ubuf_complete(struct sk_buff *skb, struct ubuf_info *uarg, + bool success) +{ + struct tpacket_uarg *tu = container_of(uarg, struct tpacket_uarg, ubuf); + struct packet_sock *po = tu->po; + void *ph = tu->ph; + __u32 ts; + + DEBUG_NET_WARN_ON_ONCE(!skb); + + if (!refcount_dec_and_test(&uarg->refcnt)) + return; + + ts = __packet_set_timestamp(po, ph, skb); + __packet_set_status(po, ph, TP_STATUS_AVAILABLE | ts); + + packet_dec_pending(&po->tx_ring); + complete(&po->skb_completion); + + kfree(tu); + sk_free(&po->sk); +} + +static const struct ubuf_info_ops tpacket_ubuf_ops = { + .complete = tpacket_ubuf_complete, +}; + static int tpacket_fill_skb(struct packet_sock *po, struct sk_buff *skb, - void *frame, struct net_device *dev, void *data, int tp_len, + struct net_device *dev, void *data, int tp_len, __be16 proto, unsigned char *addr, int hlen, int copylen, int hard_header_len, const struct sockcm_cookie *sockc) { - union tpacket_uhdr ph; int to_write, offset, len, nr_frags, len_max; struct socket *sock = po->sk.sk_socket; struct page *page; int err; - ph.raw = frame; - skb->protocol = proto; skb->dev = dev; skb->priority = sockc->priority; skb->mark = sockc->mark; skb_set_delivery_type_by_clockid(skb, sockc->transmit_time, po->sk.sk_clockid); skb_setup_tx_timestamp(skb, sockc); - skb_zcopy_set_nouarg(skb, ph.raw); skb_reserve(skb, hlen); skb_reset_network_header(skb); @@ -2749,6 +2758,7 @@ static int tpacket_snd(struct packet_sock *po, struct msghdr *msg) struct virtio_net_hdr vnet_hdr; bool has_vnet_hdr = false; struct sockcm_cookie sockc; + struct tpacket_uarg *uarg; __be16 proto; int err, reserve = 0; void *ph; @@ -2876,7 +2886,7 @@ static int tpacket_snd(struct packet_sock *po, struct msghdr *msg) err = len_sum; goto out_status; } - tp_len = tpacket_fill_skb(po, skb, ph, dev, data, tp_len, proto, + tp_len = tpacket_fill_skb(po, skb, dev, data, tp_len, proto, addr, hlen, copylen, hard_header_len, &sockc); if (likely(tp_len >= 0) && @@ -2908,7 +2918,24 @@ static int tpacket_snd(struct packet_sock *po, struct msghdr *msg) virtio_net_hdr_set_proto(skb, &vnet_hdr); } - skb->destructor = tpacket_destruct_skb; + uarg = kmalloc(sizeof(*uarg), GFP_KERNEL); + if (unlikely(!uarg)) { + if (likely(len_sum > 0)) + err = len_sum; + else + err = -ENOMEM; + goto out_status; + } + uarg->po = po; + uarg->ph = ph; + uarg->ubuf.ops = &tpacket_ubuf_ops; + uarg->ubuf.flags = SKBFL_ZEROCOPY_FRAG; + refcount_set(&uarg->ubuf.refcnt, 1); + + /* Hold a sk_wmem_alloc reference until completion */ + refcount_inc(&po->sk.sk_wmem_alloc); + skb_zcopy_init(skb, &uarg->ubuf); + __packet_set_status(po, ph, TP_STATUS_SENDING); packet_inc_pending(&po->tx_ring); @@ -4486,21 +4513,20 @@ static struct pgv *alloc_pg_vec(struct tpacket_req *req, int order, bool tx_ring vec->len = block_nr; pg_vec = vec->pg_vec; + if (tx_ring) { + vec->deferred = kzalloc_obj(*vec->deferred, + GFP_KERNEL | __GFP_NOWARN); + if (!vec->deferred) + goto out_free_pgvec; + vec->deferred->vec = vec; + INIT_DELAYED_WORK(&vec->deferred->work, + packet_free_pg_vec_work); + } + for (i = 0; i < block_nr; i++) { pg_vec[i].buffer = alloc_one_pg_vec_page(order); if (unlikely(!pg_vec[i].buffer)) goto out_free_pgvec; - - if (tx_ring && !vec->deferred && - is_vmalloc_addr(pg_vec[i].buffer)) { - vec->deferred = kzalloc_obj(*vec->deferred, - GFP_KERNEL | __GFP_NOWARN); - if (!vec->deferred) - goto out_free_pgvec; - vec->deferred->vec = vec; - INIT_DELAYED_WORK(&vec->deferred->work, - packet_free_pg_vec_work); - } } out: From cae23ae3f7887a1cf8a75da38edcebeef695040f Mon Sep 17 00:00:00 2001 From: Xin Long Date: Mon, 21 Sep 2026 14:03:45 -0400 Subject: [PATCH 094/189] sctp: hold asoc or transport before mod_timer() in timer handlers Take the association or transport reference before rearming a timer in the timer handlers. The existing code calls mod_timer() before taking the reference needed by the rearmed timer without holding the sock lock. This creates a race with timer cleanup: if the timer is deleted after mod_timer() returns but before the reference is taken, the cleanup path can drop the timer's reference and destroy the transport or association. The timer handler then takes a reference on the already freed object and eventually drops it, causing a refcount underflow. Hold the object before mod_timer() and drop the reference if mod_timer() reports that the timer was already pending in timer handlers. Apply the same ordering to the proto-unreachable path, which can rearm a transport timer outside the timer handlers without holding the sock lock. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Reported-by: Tangxin Xie Signed-off-by: Xin Long Link: https://patch.msgid.link/c31b5e3ee2b7274e804f5eba2f21e2412e7eef7a.1790013825.git.lucien.xin@gmail.com Signed-off-by: Jakub Kicinski --- net/sctp/input.c | 7 ++++--- net/sctp/sm_sideeffect.c | 37 ++++++++++++++++++++++--------------- 2 files changed, 26 insertions(+), 18 deletions(-) diff --git a/net/sctp/input.c b/net/sctp/input.c index 864741fae418..9494cfa51106 100644 --- a/net/sctp/input.c +++ b/net/sctp/input.c @@ -436,9 +436,10 @@ void sctp_icmp_proto_unreachable(struct sock *sk, if (timer_pending(&t->proto_unreach_timer)) return; else { - if (!mod_timer(&t->proto_unreach_timer, - jiffies + (HZ/20))) - sctp_transport_hold(t); + sctp_transport_hold(t); + if (mod_timer(&t->proto_unreach_timer, + jiffies + (HZ / 20))) + sctp_transport_put(t); } } else { struct net *net = sock_net(sk); diff --git a/net/sctp/sm_sideeffect.c b/net/sctp/sm_sideeffect.c index 0d99b7e8c082..35f540fb15fc 100644 --- a/net/sctp/sm_sideeffect.c +++ b/net/sctp/sm_sideeffect.c @@ -244,8 +244,9 @@ void sctp_generate_t3_rtx_event(struct timer_list *t) pr_debug("%s: sock is busy\n", __func__); /* Try again later. */ - if (!mod_timer(&transport->T3_rtx_timer, jiffies + (HZ/20))) - sctp_transport_hold(transport); + sctp_transport_hold(transport); + if (mod_timer(&transport->T3_rtx_timer, jiffies + (HZ / 20))) + sctp_transport_put(transport); goto out_unlock; } @@ -280,8 +281,9 @@ static void sctp_generate_timeout_event(struct sctp_association *asoc, timeout_type); /* Try again later. */ - if (!mod_timer(&asoc->timers[timeout_type], jiffies + (HZ/20))) - sctp_association_hold(asoc); + sctp_association_hold(asoc); + if (mod_timer(&asoc->timers[timeout_type], jiffies + (HZ / 20))) + sctp_association_put(asoc); goto out_unlock; } @@ -378,8 +380,9 @@ void sctp_generate_heartbeat_event(struct timer_list *t) pr_debug("%s: sock is busy\n", __func__); /* Try again later. */ - if (!mod_timer(&transport->hb_timer, jiffies + (HZ/20))) - sctp_transport_hold(transport); + sctp_transport_hold(transport); + if (mod_timer(&transport->hb_timer, jiffies + (HZ / 20))) + sctp_transport_put(transport); goto out_unlock; } @@ -388,8 +391,9 @@ void sctp_generate_heartbeat_event(struct timer_list *t) timeout = sctp_transport_timeout(transport); if (elapsed < timeout) { elapsed = timeout - elapsed; - if (!mod_timer(&transport->hb_timer, jiffies + elapsed)) - sctp_transport_hold(transport); + sctp_transport_hold(transport); + if (mod_timer(&transport->hb_timer, jiffies + elapsed)) + sctp_transport_put(transport); goto out_unlock; } @@ -422,9 +426,10 @@ void sctp_generate_proto_unreach_event(struct timer_list *t) pr_debug("%s: sock is busy\n", __func__); /* Try again later. */ - if (!mod_timer(&transport->proto_unreach_timer, - jiffies + (HZ/20))) - sctp_transport_hold(transport); + sctp_transport_hold(transport); + if (mod_timer(&transport->proto_unreach_timer, + jiffies + (HZ / 20))) + sctp_transport_put(transport); goto out_unlock; } @@ -458,8 +463,9 @@ void sctp_generate_reconf_event(struct timer_list *t) pr_debug("%s: sock is busy\n", __func__); /* Try again later. */ - if (!mod_timer(&transport->reconf_timer, jiffies + (HZ / 20))) - sctp_transport_hold(transport); + sctp_transport_hold(transport); + if (mod_timer(&transport->reconf_timer, jiffies + (HZ / 20))) + sctp_transport_put(transport); goto out_unlock; } @@ -495,8 +501,9 @@ void sctp_generate_probe_event(struct timer_list *t) pr_debug("%s: sock is busy\n", __func__); /* Try again later. */ - if (!mod_timer(&transport->probe_timer, jiffies + (HZ / 20))) - sctp_transport_hold(transport); + sctp_transport_hold(transport); + if (mod_timer(&transport->probe_timer, jiffies + (HZ / 20))) + sctp_transport_put(transport); goto out_unlock; } From 73b3fc67a493e9b9a8e94a38839f6638d12fe7a6 Mon Sep 17 00:00:00 2001 From: Mina Almasry Date: Mon, 21 Sep 2026 19:54:12 +0000 Subject: [PATCH 095/189] net: devmem: document that bind-tx is unprivileged by design Unlike bind-rx, which configures shared NIC RX queues to steer incoming traffic into the caller's dmabuf and requires CAP_NET_ADMIN (uns-admin-perm), bind-tx only DMA-maps the caller's dmabuf so the caller can transmit from it on their own sockets without affecting other traffic or device configuration. Add a comment in netdev.yaml and above netdev_nl_bind_tx_doit() to make it explicit that NETDEV_CMD_BIND_TX is unprivileged by design. Signed-off-by: Mina Almasry Acked-by: Stanislav Fomichev Acked-by: Daniel Borkmann Link: https://patch.msgid.link/20260921195545.493253-1-almasrymina@google.com Signed-off-by: Jakub Kicinski --- Documentation/netlink/specs/netdev.yaml | 2 ++ net/core/netdev-genl.c | 6 ++++++ 2 files changed, 8 insertions(+) diff --git a/Documentation/netlink/specs/netdev.yaml b/Documentation/netlink/specs/netdev.yaml index 3e3f03bd5c29..0562ba1f567a 100644 --- a/Documentation/netlink/specs/netdev.yaml +++ b/Documentation/netlink/specs/netdev.yaml @@ -851,6 +851,8 @@ operations: name: bind-tx doc: Bind dmabuf to netdev for TX attribute-set: dmabuf + # Intentionally unprivileged (no admin-perm / uns-admin-perm); see + # comment above netdev_nl_bind_tx_doit(). do: request: attributes: diff --git a/net/core/netdev-genl.c b/net/core/netdev-genl.c index cb18db681640..33b9f4eb9565 100644 --- a/net/core/netdev-genl.c +++ b/net/core/netdev-genl.c @@ -1168,6 +1168,12 @@ netdev_find_netmem_tx_dev(struct net_device *dev) return NULL; } +/* Note: NETDEV_CMD_BIND_TX is intentionally unprivileged (no + * GENL_ADMIN_PERM / GENL_UNS_ADMIN_PERM). Unlike bind-rx, which configures + * shared NIC RX queues, bind-tx only DMA-maps the caller's dmabuf so they can + * transmit from it on their own sockets without affecting other traffic or + * device state. + */ int netdev_nl_bind_tx_doit(struct sk_buff *skb, struct genl_info *info) { struct net_devmem_dmabuf_binding *binding; From d68acbf93531abdb5b02b21994cd4c15a3c95b42 Mon Sep 17 00:00:00 2001 From: Maxime Chevallier Date: Thu, 17 Sep 2026 23:53:32 +0200 Subject: [PATCH 096/189] net: stmmac: selftests: Support running selftests on DSA conduits Most stmmac selftests rely on dev_add_pack() to add custom handlers, that validate the packets sent to ourselves through MAC loopback. However, when the stmmac-driven interface is a DSA CPU conduit, all frames that are received have ETH_P_XDSA as a protocol, even though they don't actually contain any tag as they come from the loopback and not the switch. This will prevent any incoming packet to match our packet handlers. Let's register a ETH_P_ALL packet handler when we detect that we're a DSA conduit, and use a proxy packet handler to filter the h_proto. As this allows external frames to be received through our .func(), the packet handler is added after the dev->addr field is populated in our selftest attributes. Note that we may still receive incoming packets from the switch, but these frames shouldn't interfere with the very specific frames used for selftests, and stmmac selftests in general aren't safe against external traffic interferences. This was validated on a WPQ864 devkit for IPQ8064, that has the SoC connected to a QCA8k switch. The ARP offload's packet handler is left alone, this feature is just not implemented in stmmac and due for removal. Fixes: 091810dbded9 ("net: stmmac: Introduce selftests support") Reviewed-by: Nicolai Buchwitz Signed-off-by: Maxime Chevallier Link: https://patch.msgid.link/20260917215339.2022523-2-maxime.chevallier@bootlin.com Signed-off-by: Jakub Kicinski --- .../stmicro/stmmac/stmmac_selftests.c | 89 +++++++++++++++---- 1 file changed, 71 insertions(+), 18 deletions(-) diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c index 6372ec7c3f31..614b5995dec5 100644 --- a/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c +++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c @@ -12,6 +12,7 @@ #include #include #include +#include #include #include #include @@ -237,6 +238,9 @@ struct stmmac_test_priv { struct stmmac_packet_attrs *packet; struct packet_type pt; struct completion comp; + __be16 packet_type; + int (*func)(struct sk_buff *skb, struct net_device *ndev, + struct packet_type *pt, struct net_device *orig_ndev); int double_vlan; int vlan_id; int ok; @@ -316,6 +320,50 @@ static int stmmac_test_loopback_validate(struct sk_buff *skb, return 0; } +static int stmmac_sft_filter(struct sk_buff *skb, struct net_device *ndev, + struct packet_type *pt, + struct net_device *orig_ndev) +{ + struct stmmac_test_priv *tpriv = pt->af_packet_priv; + struct ethhdr *hdr = eth_hdr(skb); + int ret = 0; + + if (hdr->h_proto == tpriv->packet_type) { + struct sk_buff *nskb = skb_clone(skb, GFP_ATOMIC); + + if (nskb) + ret = tpriv->func(nskb, ndev, pt, orig_ndev); + } + + kfree_skb(skb); + return ret; +} + +static void stmmac_sft_add_pack(struct packet_type *pt) +{ + struct stmmac_test_priv *tpriv = pt->af_packet_priv; + + if (netdev_uses_dsa(tpriv->pt.dev)) { + tpriv->packet_type = tpriv->pt.type; + tpriv->func = tpriv->pt.func; + + /* DSA conduit will report ETH_P_XDSA, so our packet handler + * won't match. Let's register a ETH_P_ALL match and filter + * manually in stmmac_sft_filter. + */ + tpriv->pt.type = htons(ETH_P_ALL); + tpriv->pt.func = stmmac_sft_filter; + tpriv->pt.ignore_outgoing = true; + } + + dev_add_pack(pt); +} + +static void stmmac_sft_remove_pack(struct packet_type *pt) +{ + dev_remove_pack(pt); +} + static int __stmmac_test_loopback(struct stmmac_priv *priv, struct stmmac_packet_attrs *attr) { @@ -337,7 +385,7 @@ static int __stmmac_test_loopback(struct stmmac_priv *priv, tpriv->packet = attr; if (!attr->dont_wait) - dev_add_pack(&tpriv->pt); + stmmac_sft_add_pack(&tpriv->pt); skb = stmmac_test_get_udp_skb(priv, attr); if (!skb) { @@ -360,7 +408,7 @@ static int __stmmac_test_loopback(struct stmmac_priv *priv, cleanup: if (!attr->dont_wait) - dev_remove_pack(&tpriv->pt); + stmmac_sft_remove_pack(&tpriv->pt); kfree(tpriv); return ret; } @@ -767,7 +815,7 @@ static int stmmac_test_flowctrl(struct stmmac_priv *priv) tpriv->pt.func = stmmac_test_flowctrl_validate; tpriv->pt.dev = priv->dev; tpriv->pt.af_packet_priv = tpriv; - dev_add_pack(&tpriv->pt); + stmmac_sft_add_pack(&tpriv->pt); /* Compute minimum number of packets to make FIFO full */ pkt_count = rx_fifo_size; @@ -823,7 +871,7 @@ static int stmmac_test_flowctrl(struct stmmac_priv *priv) cleanup: dev_mc_del(priv->dev, paddr); dev_set_promiscuity(priv->dev, -1); - dev_remove_pack(&tpriv->pt); + stmmac_sft_remove_pack(&tpriv->pt); kfree(tpriv); return ret; } @@ -928,18 +976,20 @@ static int __stmmac_test_vlanfilt(struct stmmac_priv *priv) * HASH values. */ tpriv->vlan_id = 0x123; - dev_add_pack(&tpriv->pt); ret = vlan_vid_add(priv->dev, htons(ETH_P_8021Q), tpriv->vlan_id); if (ret) goto cleanup; + attr.vlan = 1; + attr.dst = priv->dev->dev_addr; + attr.sport = 9; + attr.dport = 9; + + stmmac_sft_add_pack(&tpriv->pt); + for (i = 0; i < 4; i++) { - attr.vlan = 1; attr.vlan_id_out = tpriv->vlan_id + i; - attr.dst = priv->dev->dev_addr; - attr.sport = 9; - attr.dport = 9; skb = stmmac_test_get_udp_skb(priv, &attr); if (!skb) { @@ -966,9 +1016,9 @@ static int __stmmac_test_vlanfilt(struct stmmac_priv *priv) } vlan_del: + stmmac_sft_remove_pack(&tpriv->pt); vlan_vid_del(priv->dev, htons(ETH_P_8021Q), tpriv->vlan_id); cleanup: - dev_remove_pack(&tpriv->pt); kfree(tpriv); return ret; } @@ -1022,18 +1072,20 @@ static int __stmmac_test_dvlanfilt(struct stmmac_priv *priv) * HASH values. */ tpriv->vlan_id = 0x123; - dev_add_pack(&tpriv->pt); ret = vlan_vid_add(priv->dev, htons(ETH_P_8021AD), tpriv->vlan_id); if (ret) goto cleanup; + attr.vlan = 2; + attr.dst = priv->dev->dev_addr; + attr.sport = 9; + attr.dport = 9; + + stmmac_sft_add_pack(&tpriv->pt); + for (i = 0; i < 4; i++) { - attr.vlan = 2; attr.vlan_id_out = tpriv->vlan_id + i; - attr.dst = priv->dev->dev_addr; - attr.sport = 9; - attr.dport = 9; skb = stmmac_test_get_udp_skb(priv, &attr); if (!skb) { @@ -1060,9 +1112,9 @@ static int __stmmac_test_dvlanfilt(struct stmmac_priv *priv) } vlan_del: + stmmac_sft_remove_pack(&tpriv->pt); vlan_vid_del(priv->dev, htons(ETH_P_8021AD), tpriv->vlan_id); cleanup: - dev_remove_pack(&tpriv->pt); kfree(tpriv); return ret; } @@ -1293,7 +1345,6 @@ static int stmmac_test_vlanoff_common(struct stmmac_priv *priv, bool svlan) tpriv->pt.af_packet_priv = tpriv; tpriv->packet = &attr; tpriv->vlan_id = 0x123; - dev_add_pack(&tpriv->pt); ret = vlan_vid_add(priv->dev, htons(proto), tpriv->vlan_id); if (ret) @@ -1301,6 +1352,8 @@ static int stmmac_test_vlanoff_common(struct stmmac_priv *priv, bool svlan) attr.dst = priv->dev->dev_addr; + stmmac_sft_add_pack(&tpriv->pt); + skb = stmmac_test_get_udp_skb(priv, &attr); if (!skb) { ret = -ENOMEM; @@ -1318,9 +1371,9 @@ static int stmmac_test_vlanoff_common(struct stmmac_priv *priv, bool svlan) ret = tpriv->ok ? 0 : -ETIMEDOUT; vlan_del: + stmmac_sft_remove_pack(&tpriv->pt); vlan_vid_del(priv->dev, htons(proto), tpriv->vlan_id); cleanup: - dev_remove_pack(&tpriv->pt); kfree(tpriv); return ret; } From c8c1795aa8106293020a836d01e127e98f442925 Mon Sep 17 00:00:00 2001 From: Maxime Chevallier Date: Thu, 17 Sep 2026 23:53:33 +0200 Subject: [PATCH 097/189] net: stmmac: selftests: Validate EEE based on the actual LPI timer value The EEE selftest is a 2-step test : - It validates that we enter in LPI mode with the irq_tx_path_in_lpi_mode_n counter - It then validates that we exit LPI when sending a frame, with the irq_tx_path_exit_lpi_mode_n counter. The current state of the test lacks 2 main things : - We don't know exactly when was the previous frame sent (it's from the previous selftest) - The timeout is hardcoded, while the LPI is entered after a user-configurable delay. On top of that, the timeout loop uses a pre-decrement iterator (--retries) that actually only iterate nine times, so 900ms while the default LPI value is 1 second. Let's therefore make it more deterministic : - Send a frame at the beginning of the test - Wait for more than the lpi timer value, we timeout after about twice the value, - Then send another frame, and verify that we do go out of LPI, also with a timeout. As LPI timer can get pretty high, bail out if LPI timer is over 5 seconds. Note that the test's goal isn't to validate the LPI timer value itself, only that we enter/leave LPI mode. Fixes: 091810dbded9 ("net: stmmac: Introduce selftests support") Reviewed-by: Nicolai Buchwitz Signed-off-by: Maxime Chevallier Link: https://patch.msgid.link/20260917215339.2022523-3-maxime.chevallier@bootlin.com Signed-off-by: Jakub Kicinski --- .../stmicro/stmmac/stmmac_selftests.c | 44 ++++++++++++++++--- 1 file changed, 37 insertions(+), 7 deletions(-) diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c index 614b5995dec5..2f9f7746c40a 100644 --- a/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c +++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c @@ -30,6 +30,7 @@ struct stmmachdr { sizeof(struct stmmachdr)) #define STMMAC_TEST_PKT_MAGIC 0xdeadcafecafedeadULL #define STMMAC_LB_TIMEOUT msecs_to_jiffies(200) +#define STMMAC_SFT_MAX_LPI (5 * USEC_PER_SEC) struct stmmac_packet_attrs { int vlan; @@ -462,12 +463,16 @@ static int stmmac_test_mmc(struct stmmac_priv *priv) static int stmmac_test_eee(struct stmmac_priv *priv) { struct stmmac_extra_stats *initial, *final; - int retries = 10; + unsigned long timeout, max_duration; int ret; if (!priv->dma_cap.eee || !priv->eee_active) return -EOPNOTSUPP; + /* Bail out if the configured LPI timer is too long */ + if (priv->tx_lpi_timer > STMMAC_SFT_MAX_LPI) + return -EOPNOTSUPP; + initial = kzalloc_obj(*initial); if (!initial) return -ENOMEM; @@ -478,14 +483,21 @@ static int stmmac_test_eee(struct stmmac_priv *priv) goto out_free_initial; } + /* Snapshot stats, we want to count the in_lpi events. We may enter + * LPI just after the packet was sent. + */ memcpy(initial, &priv->xstats, sizeof(*initial)); + /* Send a frame, then wait to enter LPI */ ret = stmmac_test_mac_loopback(priv); if (ret) goto out_free_final; + max_duration = usecs_to_jiffies(2 * priv->tx_lpi_timer); + /* We have no traffic in the line so, sooner or later it will go LPI */ - while (--retries) { + timeout = jiffies + max_duration; + while (!time_after(jiffies, timeout)) { memcpy(final, &priv->xstats, sizeof(*final)); if (final->irq_tx_path_in_lpi_mode_n > @@ -494,20 +506,38 @@ static int stmmac_test_eee(struct stmmac_priv *priv) msleep(100); } - if (!retries) { + memcpy(final, &priv->xstats, sizeof(*final)); + if (final->irq_tx_path_in_lpi_mode_n <= + initial->irq_tx_path_in_lpi_mode_n) { ret = -ETIMEDOUT; goto out_free_final; } - if (final->irq_tx_path_in_lpi_mode_n <= - initial->irq_tx_path_in_lpi_mode_n) { - ret = -EINVAL; + /* Re-snapshot, as we want to measure exit_lpi events. We should be + * in LPI right now. + */ + memcpy(initial, &priv->xstats, sizeof(*initial)); + + /* TX something so we go out of LPI */ + ret = stmmac_test_mac_loopback(priv); + if (ret) goto out_free_final; + + /* Wait for the exit LPI interrupt */ + timeout = jiffies + max_duration; + while (!time_after(jiffies, timeout)) { + memcpy(final, &priv->xstats, sizeof(*final)); + + if (final->irq_tx_path_exit_lpi_mode_n > + initial->irq_tx_path_exit_lpi_mode_n) + break; + msleep(100); } + memcpy(final, &priv->xstats, sizeof(*final)); if (final->irq_tx_path_exit_lpi_mode_n <= initial->irq_tx_path_exit_lpi_mode_n) { - ret = -EINVAL; + ret = -ETIMEDOUT; goto out_free_final; } From ba804b23d76d278ee475b8427fa7c5623ce5e270 Mon Sep 17 00:00:00 2001 From: Maxime Chevallier Date: Thu, 17 Sep 2026 23:53:34 +0200 Subject: [PATCH 098/189] net: stmmac: selftests: Check the dev->features for S-TAG offload testing The S-TAG offload insertion incorrectly checks the dvlan (double vlan) DMA cap, which is different than S-TAG support. Use NETIF_F_HW_VLAN_STAG_TX to check if the feature is supported instead. Note that this flag isn't set in stmmac yet, but contrary to ARP offload, this is a feature that has a chance to get there eventually so let's leave the selftest here for now. It'll report -EOPNOTSUPP in the meantime. Fixes: 091810dbded9 ("net: stmmac: Introduce selftests support") Reviewed-by: Nicolai Buchwitz Signed-off-by: Maxime Chevallier Link: https://patch.msgid.link/20260917215339.2022523-4-maxime.chevallier@bootlin.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c index 2f9f7746c40a..de02c0da56dc 100644 --- a/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c +++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c @@ -1415,7 +1415,7 @@ static int stmmac_test_vlanoff(struct stmmac_priv *priv) static int stmmac_test_svlanoff(struct stmmac_priv *priv) { - if (!priv->dma_cap.dvlan) + if (!(priv->dev->features & NETIF_F_HW_VLAN_STAG_TX)) return -EOPNOTSUPP; return stmmac_test_vlanoff_common(priv, true); } From 960db6f65788c21249ea04a947d5c01e38d19294 Mon Sep 17 00:00:00 2001 From: Maxime Chevallier Date: Thu, 17 Sep 2026 23:53:35 +0200 Subject: [PATCH 099/189] net: stmmac: selftests: Capture all packets for vlan checks While we use vlan_vid_add to trigger the tag filtering machinery in the driver, there's no netdev associated to the VLAN. This causes the skb to arrive with empty skb->vlan_tci fields, as the packet is marked OTHERHOST in __netif_receive_skb_core(), and we fail our validation. Let's use the proxy mechanism introduced for DSA, that registers a ETH_P_ALL packet handler that runs earlier, before the vlan netdev lookup, then filters for the correct ethertype before passing an skb clone to our validation function. As we may receive external frames with the right tag from the outside, let's move the address check in the vlan validation function earlier. Fixes: 091810dbded9 ("net: stmmac: Introduce selftests support") Reviewed-by: Nicolai Buchwitz Signed-off-by: Maxime Chevallier Link: https://patch.msgid.link/20260917215339.2022523-5-maxime.chevallier@bootlin.com Signed-off-by: Jakub Kicinski --- .../stmicro/stmmac/stmmac_selftests.c | 19 +++++++++++++------ 1 file changed, 13 insertions(+), 6 deletions(-) diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c index de02c0da56dc..43b8411c5112 100644 --- a/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c +++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c @@ -242,6 +242,7 @@ struct stmmac_test_priv { __be16 packet_type; int (*func)(struct sk_buff *skb, struct net_device *ndev, struct packet_type *pt, struct net_device *orig_ndev); + bool capture_all; int double_vlan; int vlan_id; int ok; @@ -344,13 +345,15 @@ static void stmmac_sft_add_pack(struct packet_type *pt) { struct stmmac_test_priv *tpriv = pt->af_packet_priv; - if (netdev_uses_dsa(tpriv->pt.dev)) { + if (netdev_uses_dsa(tpriv->pt.dev) || tpriv->capture_all) { tpriv->packet_type = tpriv->pt.type; tpriv->func = tpriv->pt.func; /* DSA conduit will report ETH_P_XDSA, so our packet handler * won't match. Let's register a ETH_P_ALL match and filter - * manually in stmmac_sft_filter. + * manually in stmmac_sft_filter. This is also useful for + * VLAN tests, to capture packets otherwise marked as + * OTHERHOST. */ tpriv->pt.type = htons(ETH_P_ALL); tpriv->pt.func = stmmac_sft_filter; @@ -943,6 +946,11 @@ static int stmmac_test_vlan_validate(struct sk_buff *skb, goto out; if (skb_headlen(skb) < (STMMAC_TEST_PKT_SIZE - ETH_HLEN)) goto out; + + ehdr = (struct ethhdr *)skb_mac_header(skb); + if (!ether_addr_equal_unaligned(ehdr->h_dest, tpriv->packet->dst)) + goto out; + if (tpriv->vlan_id) { if (skb->vlan_proto != htons(proto)) goto out; @@ -954,10 +962,6 @@ static int stmmac_test_vlan_validate(struct sk_buff *skb, } } - ehdr = (struct ethhdr *)skb_mac_header(skb); - if (!ether_addr_equal_unaligned(ehdr->h_dest, tpriv->packet->dst)) - goto out; - ihdr = ip_hdr(skb); if (tpriv->double_vlan) ihdr = (struct iphdr *)(skb_network_header(skb) + 4); @@ -999,6 +1003,7 @@ static int __stmmac_test_vlanfilt(struct stmmac_priv *priv) tpriv->pt.dev = priv->dev; tpriv->pt.af_packet_priv = tpriv; tpriv->packet = &attr; + tpriv->capture_all = true; /* * As we use HASH filtering, false positives may appear. This is a @@ -1095,6 +1100,7 @@ static int __stmmac_test_dvlanfilt(struct stmmac_priv *priv) tpriv->pt.dev = priv->dev; tpriv->pt.af_packet_priv = tpriv; tpriv->packet = &attr; + tpriv->capture_all = true; /* * As we use HASH filtering, false positives may appear. This is a @@ -1375,6 +1381,7 @@ static int stmmac_test_vlanoff_common(struct stmmac_priv *priv, bool svlan) tpriv->pt.af_packet_priv = tpriv; tpriv->packet = &attr; tpriv->vlan_id = 0x123; + tpriv->capture_all = true; ret = vlan_vid_add(priv->dev, htons(proto), tpriv->vlan_id); if (ret) From b42e7012773a0e81e97a2dda6ef907f5147a6658 Mon Sep 17 00:00:00 2001 From: Maxime Chevallier Date: Thu, 17 Sep 2026 23:53:36 +0200 Subject: [PATCH 100/189] net: stmmac: dwmac4: Use the correct bufzise when the len is exactly 8K DMA bufsize selection isn't made on the MTU but the actual frame length, so including the L2 header. On DWMAC4, if the len is exactly BUF_SIZE_8KiB, the next larger size is incorrectly selected. Lets fix the comparison and while at it, rename the parameter from len to mtu. Fixes: c3efed5ad1b0 ("net: stmmac: Enable dwmac4 jumbo frame more than 8KiB"). Signed-off-by: Maxime Chevallier Reviewed-by: Nicolai Buchwitz Link: https://patch.msgid.link/20260917215339.2022523-6-maxime.chevallier@bootlin.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/stmicro/stmmac/dwmac4_descs.c | 4 ++-- drivers/net/ethernet/stmicro/stmmac/hwif.h | 2 +- drivers/net/ethernet/stmicro/stmmac/ring_mode.c | 4 ++-- 3 files changed, 5 insertions(+), 5 deletions(-) diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac4_descs.c b/drivers/net/ethernet/stmicro/stmmac/dwmac4_descs.c index 2994df41ec2c..c6a8f8d73501 100644 --- a/drivers/net/ethernet/stmicro/stmmac/dwmac4_descs.c +++ b/drivers/net/ethernet/stmicro/stmmac/dwmac4_descs.c @@ -474,11 +474,11 @@ static void dwmac4_set_sarc(struct dma_desc *p, u32 sarc_type) sarc_type)); } -static int set_16kib_bfsize(int mtu) +static int set_16kib_bfsize(int len) { int ret = 0; - if (unlikely(mtu >= BUF_SIZE_8KiB)) + if (unlikely(len > BUF_SIZE_8KiB)) ret = BUF_SIZE_16KiB; return ret; } diff --git a/drivers/net/ethernet/stmicro/stmmac/hwif.h b/drivers/net/ethernet/stmicro/stmmac/hwif.h index 9314bcb85c22..857f7562c6c6 100644 --- a/drivers/net/ethernet/stmicro/stmmac/hwif.h +++ b/drivers/net/ethernet/stmicro/stmmac/hwif.h @@ -540,7 +540,7 @@ struct stmmac_mode_ops { bool (*is_jumbo_frm)(unsigned int len, bool enh_desc); int (*jumbo_frm)(struct stmmac_tx_queue *tx_q, struct sk_buff *skb, int csum); - int (*set_16kib_bfsize)(int mtu); + int (*set_16kib_bfsize)(int len); void (*init_desc3)(struct dma_desc *p); void (*refill_desc3)(struct stmmac_rx_queue *rx_q, struct dma_desc *p); void (*clean_desc3)(struct stmmac_tx_queue *tx_q, struct dma_desc *p); diff --git a/drivers/net/ethernet/stmicro/stmmac/ring_mode.c b/drivers/net/ethernet/stmicro/stmmac/ring_mode.c index f7949419eb9f..d2f0c321661d 100644 --- a/drivers/net/ethernet/stmicro/stmmac/ring_mode.c +++ b/drivers/net/ethernet/stmicro/stmmac/ring_mode.c @@ -124,10 +124,10 @@ static void clean_desc3(struct stmmac_tx_queue *tx_q, struct dma_desc *p) p->des3 = 0; } -static int set_16kib_bfsize(int mtu) +static int set_16kib_bfsize(int len) { int ret = 0; - if (unlikely(mtu > BUF_SIZE_8KiB)) + if (unlikely(len > BUF_SIZE_8KiB)) ret = BUF_SIZE_16KiB; return ret; } From b8a26d46c0a4254f8bfe143681adb9d18d020299 Mon Sep 17 00:00:00 2001 From: Maxime Chevallier Date: Thu, 17 Sep 2026 23:53:37 +0200 Subject: [PATCH 101/189] net: stmmac: size the RX buffers from the frame length, not the MTU When picking the buffsize to use based on the MTU, we shouldn't check only the MTU value, but also : - ETH_HLEN for the L2 header, - up to 2 VLAN tags, - the FCS, The default bufsize is 1536 bytes, which is enough to contain all the above so this hasn't surfaced before, but the addition of NET_IP_ALIGN to the start of buffer address tripped the Jumbo selftest, leading to this discovery. With that, we don't need the '>=' checks on the buffer len, we can use more consistent comparison operators in stmmac_set_bfsize. Fixes: 286a83721720 ("stmmac: add CHAINED descriptor mode support (V4)") Reviewed-by: Nicolai Buchwitz Signed-off-by: Maxime Chevallier Link: https://patch.msgid.link/20260917215339.2022523-7-maxime.chevallier@bootlin.com Signed-off-by: Jakub Kicinski --- .../net/ethernet/stmicro/stmmac/stmmac_main.c | 20 ++++++++++--------- 1 file changed, 11 insertions(+), 9 deletions(-) diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c index 1fb5f804ea23..d5a984ad864f 100644 --- a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c +++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c @@ -1536,17 +1536,17 @@ static unsigned int stmmac_rx_offset(struct stmmac_priv *priv) return NET_SKB_PAD + NET_IP_ALIGN; } -static int stmmac_set_bfsize(int mtu) +static int stmmac_set_bfsize(int len) { int ret; - if (mtu >= BUF_SIZE_8KiB) + if (len > BUF_SIZE_8KiB) ret = BUF_SIZE_16KiB; - else if (mtu >= BUF_SIZE_4KiB) + else if (len > BUF_SIZE_4KiB) ret = BUF_SIZE_8KiB; - else if (mtu >= BUF_SIZE_2KiB) + else if (len > BUF_SIZE_2KiB) ret = BUF_SIZE_4KiB; - else if (mtu > DEFAULT_BUFSIZE) + else if (len > DEFAULT_BUFSIZE) ret = BUF_SIZE_2KiB; else ret = DEFAULT_BUFSIZE; @@ -4063,7 +4063,7 @@ static struct stmmac_dma_conf * stmmac_setup_dma_desc(struct stmmac_priv *priv, unsigned int mtu) { struct stmmac_dma_conf *dma_conf; - int bfsize, ret; + int bfsize, len, ret; u8 chan; dma_conf = kzalloc_obj(*dma_conf); @@ -4073,13 +4073,15 @@ stmmac_setup_dma_desc(struct stmmac_priv *priv, unsigned int mtu) return ERR_PTR(-ENOMEM); } - /* Returns 0 or BUF_SIZE_16KiB if mtu > 8KiB and dwmac4 or ring mode */ - bfsize = stmmac_set_16kib_bfsize(priv, mtu); + len = mtu + ETH_HLEN + 2 * VLAN_HLEN + ETH_FCS_LEN; + + /* Returns 0 or BUF_SIZE_16KiB if len > 8KiB and dwmac4 or ring mode */ + bfsize = stmmac_set_16kib_bfsize(priv, len); if (bfsize < 0) bfsize = 0; if (bfsize < BUF_SIZE_16KiB) - bfsize = stmmac_set_bfsize(mtu); + bfsize = stmmac_set_bfsize(len); dma_conf->dma_buf_sz = bfsize; /* Chose the tx/rx size from the already defined one in the From c4ac6e94eb9423126bda907a7f2933284f7450ff Mon Sep 17 00:00:00 2001 From: Maxime Chevallier Date: Thu, 17 Sep 2026 23:53:38 +0200 Subject: [PATCH 102/189] net: stmmac: selftests: Account for alignment shift on dwmac1000 for Jumbo test On dwmac1000, we currently only support single-descriptor frames. The Jumbo test started failing when NET_IP_ALIGN was added to align the IP header, as this tests tries to send the biggest possible frame. On dwmac1000 the DMA transfer is aligned on 4-bytes, so adding a 2-byte shift at the start-of-buffer address means it takes a whole extra 4-byte DMA burst to receive the Jumbo packet, causing it to spill over the next descriptor. This doesn't seem to happen on dwmac4 and xgmac that appear to correctly handle unaligned xfers (only tested on dwmac4) Let's account for that in the Jumbo test, reduce the size of our big packet by the align size. Fixes: 23680bf5f8c6 ("net: stmmac: restore NET_IP_ALIGN in the RX DMA offset") Reviewed-by: Nicolai Buchwitz Signed-off-by: Maxime Chevallier Link: https://patch.msgid.link/20260917215339.2022523-8-maxime.chevallier@bootlin.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c | 3 +++ 1 file changed, 3 insertions(+) diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c index 43b8411c5112..c25dc9f89270 100644 --- a/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c +++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c @@ -1789,6 +1789,9 @@ static int __stmmac_test_jumbo(struct stmmac_priv *priv, u16 queue) struct stmmac_packet_attrs attr = { }; int size = priv->dma_conf.dma_buf_sz; + if (!dwmac_is_xmac(priv->plat->core_type)) + size -= NET_IP_ALIGN; + attr.dst = priv->dev->dev_addr; attr.max_size = size - ETH_FCS_LEN; attr.queue_mapping = queue; From 0160953d8eec75c3c55562158c46442ff1fd410b Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Bj=C3=B6rn=20T=C3=B6pel?= Date: Fri, 18 Sep 2026 13:46:40 +0200 Subject: [PATCH 103/189] eth: fbnic: Avoid rounding zero ring sizes MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit roundup_pow_of_two() is undefined for zero. ethtool permits a zero ring size to reach the driver, where the minimum-size check should reject it. Leave zero unchanged while rounding nonzero ring sizes. The minimum-size check then rejects zero deterministically without changing the established behavior for other values. Fixes: 6cbf18a05c06 ("eth: fbnic: support ring size configuration") Reported-by: Sashiko Link: https://lore.kernel.org/netdev/178971206933.22033.236948278674126701@kernel.org/ Suggested-by: Alexander Duyck Signed-off-by: Björn Töpel Reviewed-by: Joe Damato Link: https://patch.msgid.link/20260918114641.1281172-1-bjorn@kernel.org Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/meta/fbnic/fbnic_ethtool.c | 13 +++++++++---- 1 file changed, 9 insertions(+), 4 deletions(-) diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_ethtool.c b/drivers/net/ethernet/meta/fbnic/fbnic_ethtool.c index 423f179c9d47..76e9a545bb16 100644 --- a/drivers/net/ethernet/meta/fbnic/fbnic_ethtool.c +++ b/drivers/net/ethernet/meta/fbnic/fbnic_ethtool.c @@ -313,6 +313,11 @@ fbnic_get_ringparam(struct net_device *netdev, struct ethtool_ringparam *ring, kernel_ring->hds_thresh = fbn->hds_thresh; } +static u32 fbnic_ring_size_pow2(u32 size) +{ + return size ? roundup_pow_of_two(size) : 0; +} + static void fbnic_set_rings(struct fbnic_net *fbn, struct ethtool_ringparam *ring, struct kernel_ethtool_ringparam *kernel_ring) @@ -334,10 +339,10 @@ fbnic_set_ringparam(struct net_device *netdev, struct ethtool_ringparam *ring, struct fbnic_net *clone; int err; - ring->rx_pending = roundup_pow_of_two(ring->rx_pending); - ring->rx_mini_pending = roundup_pow_of_two(ring->rx_mini_pending); - ring->rx_jumbo_pending = roundup_pow_of_two(ring->rx_jumbo_pending); - ring->tx_pending = roundup_pow_of_two(ring->tx_pending); + ring->rx_pending = fbnic_ring_size_pow2(ring->rx_pending); + ring->rx_mini_pending = fbnic_ring_size_pow2(ring->rx_mini_pending); + ring->rx_jumbo_pending = fbnic_ring_size_pow2(ring->rx_jumbo_pending); + ring->tx_pending = fbnic_ring_size_pow2(ring->tx_pending); /* These are absolute minimums allowing the device and driver to operate * but not necessarily guarantee reasonable performance. Settings below From 9c572a83037a7dcd653ba3a9cc468c16b857d0c9 Mon Sep 17 00:00:00 2001 From: Bernard Ladenthin Date: Fri, 18 Sep 2026 15:39:37 +0200 Subject: [PATCH 104/189] net/sched: fix potential stack infoleak in em_text_dump() em_text_dump() allocates struct tcf_em_text on the stack without zeroing it. strscpy() writes the algorithm name and a NUL terminator into conf.algo[], leaving the remaining bytes uninitialised. nla_put_nohdr() then copies the full struct to the netlink response. KMSAN on Linux 7.2-rc6 reports two kernel-infoleak splats from this path, one triggered via "tc filter show" and one via a raw RTM_GETTFILTER dump: BUG: KMSAN: kernel-infoleak in _copy_to_iter+0x1c9/0x2620 nla_put_nohdr+0x83/0x130 em_text_dump+0x291/0x550 Local variable conf created at: em_text_dump+0x5d/0x550 Bytes 168-179 of 199 are uninitialized I am not certain whether this constitutes a real security problem in practice: the test was conducted in a controlled KMSAN environment and the leaked stack bytes may or may not carry sensitive data on actual production kernels. I am reporting it because KMSAN flagged it as a kernel-infoleak and the fix is straightforward. I can provide a userspace reproducer on request. The original code used strncpy() which zero-pads to the destination size. Commit b04202d6065c ("net/sched: replace strncpy with strscpy") replaced it with strscpy(), which does not pad, creating this condition. Zero-initialising the struct closes it. Fixes: b04202d6065c ("net/sched: replace strncpy with strscpy") Link: https://lore.kernel.org/netdev/20250327143733.187438-1-richard120310@gmail.com/ Assisted-by: Claude:claude-sonnet-4-6 [KMSAN] Signed-off-by: Bernard Ladenthin Acked-by: Jamal Hadi Salim Link: https://patch.msgid.link/20260918133953.12494-1-bernard.ladenthin@gmail.com Signed-off-by: Jakub Kicinski --- net/sched/em_text.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/net/sched/em_text.c b/net/sched/em_text.c index 343f1aebeec2..4132f8c3c5fc 100644 --- a/net/sched/em_text.c +++ b/net/sched/em_text.c @@ -113,7 +113,7 @@ static void em_text_destroy(struct tcf_ematch *m) static int em_text_dump(struct sk_buff *skb, struct tcf_ematch *m) { struct text_match *tm = EM_TEXT_PRIV(m); - struct tcf_em_text conf; + struct tcf_em_text conf = {}; strscpy(conf.algo, tm->config->ops->name); conf.from_offset = tm->from_offset; From 686f942332b1667f13f3b8d6a2f50bcfbf42e277 Mon Sep 17 00:00:00 2001 From: Pengpeng Hou Date: Wed, 15 Jul 2026 16:43:25 +0800 Subject: [PATCH 105/189] nfc: nfcmrvl: validate helper command length before pull The firmware download receive path removes the NCI data header and reads the helper command before validating the remaining packet length. A short frame can therefore reach the data access before the malformed packet is rejected. Validate the complete helper command length before stripping the NCI data header. Fixes: 3194c6870158 ("NFC: nfcmrvl: add firmware download support") Signed-off-by: Pengpeng Hou Link: https://patch.msgid.link/20260715084325.40276-1-pengpeng@iscas.ac.cn Signed-off-by: David Heidelberg --- drivers/nfc/nfcmrvl/fw_dnld.c | 11 ++++++++--- 1 file changed, 8 insertions(+), 3 deletions(-) diff --git a/drivers/nfc/nfcmrvl/fw_dnld.c b/drivers/nfc/nfcmrvl/fw_dnld.c index 2b8f401d8fd7..8b9d5257320d 100644 --- a/drivers/nfc/nfcmrvl/fw_dnld.c +++ b/drivers/nfc/nfcmrvl/fw_dnld.c @@ -263,9 +263,14 @@ static int process_state_fw_dnld(struct nfcmrvl_private *priv, * B8..N: payload */ - /* Remove NCI HDR */ - skb_pull(skb, 3); - if (skb->data[0] != HELPER_CMD_PACKET_FORMAT || skb->len != 5) { + if (skb->len != NCI_DATA_HDR_SIZE + 5) { + nfc_err(priv->dev, "bad command"); + return -EINVAL; + } + + /* Remove NCI header */ + skb_pull(skb, NCI_DATA_HDR_SIZE); + if (skb->data[0] != HELPER_CMD_PACKET_FORMAT) { nfc_err(priv->dev, "bad command"); return -EINVAL; } From a653c01ce447f10c36b901646888c0330363af4f Mon Sep 17 00:00:00 2001 From: Pengpeng Hou Date: Wed, 15 Jul 2026 16:44:05 +0800 Subject: [PATCH 106/189] nfc: st21nfca: validate received frame size st21nfca_hci_i2c_repack() trims a received frame at its EOF marker before removing byte stuffing. It then assumes the truncated frame contains the LLC header and two CRC bytes, and it unconditionally reads the byte after an escape marker. A malformed frame can place EOF immediately after the start marker or can end its data portion with an escape marker. The former leaves too few bytes for check_crc(), while the latter makes the unstuffing loop read past the current skb length. Require the minimum framing bytes both before and after unstuffing. Use separate input and output cursors while removing byte stuffing, and reject an escape marker without its encoded byte. This keeps malformed frames within the received frame boundary before CRC processing. Fixes: 3096e25a3e40 ("NFC: st21nfca: Fix incorrect byte stuffing revocation") Signed-off-by: Pengpeng Hou Link: https://patch.msgid.link/20260715084405.41546-1-pengpeng@iscas.ac.cn Signed-off-by: David Heidelberg --- drivers/nfc/st21nfca/i2c.c | 29 +++++++++++++++++++---------- 1 file changed, 19 insertions(+), 10 deletions(-) diff --git a/drivers/nfc/st21nfca/i2c.c b/drivers/nfc/st21nfca/i2c.c index a4c93ff7c5b0..0f44c783bd04 100644 --- a/drivers/nfc/st21nfca/i2c.c +++ b/drivers/nfc/st21nfca/i2c.c @@ -289,27 +289,36 @@ static int check_crc(u8 *buf, int buflen) */ static int st21nfca_hci_i2c_repack(struct sk_buff *skb) { - int i, j, r, size; + int read, write, r, size; - if (skb->len < 1 || (skb->len > 1 && skb->data[1] != 0)) + if (skb->len < ST21NFCA_FRAME_HEADROOM || + !IS_START_OF_FRAME(skb->data)) return -EBADMSG; size = get_frame_size(skb->data, skb->len); if (size > 0) { + if (size < ST21NFCA_FRAME_HEADROOM + 2) + return -EBADMSG; + skb_trim(skb, size); /* remove ST21NFCA byte stuffing for upper layer */ - for (i = 1, j = 0; i < skb->len; i++) { - if (skb->data[i + j] == + for (read = 1, write = 1; read < skb->len;) { + if (skb->data[read] == (u8) ST21NFCA_ESCAPE_BYTE_STUFFING) { - skb->data[i] = skb->data[i + j + 1] - | ST21NFCA_BYTE_STUFFING_MASK; - i++; - j++; + if (read + 1 == skb->len) + return -EBADMSG; + + skb->data[write++] = skb->data[read + 1] + | ST21NFCA_BYTE_STUFFING_MASK; + read += 2; + } else { + skb->data[write++] = skb->data[read++]; } - skb->data[i] = skb->data[i + j]; } /* remove byte stuffing useless byte */ - skb_trim(skb, i - j); + skb_trim(skb, write); + if (skb->len < ST21NFCA_FRAME_HEADROOM + 2) + return -EBADMSG; /* remove ST21NFCA_SOF_EOF from head */ skb_pull(skb, 1); From bf1460acdf8cf5a07c819f59785d40f20d113099 Mon Sep 17 00:00:00 2001 From: Aldo Ariel Panzardo Date: Thu, 16 Jul 2026 20:26:57 -0300 Subject: [PATCH 107/189] nfc: llcp: Fix list corruption / refcount desync in nfc_llcp_recv_dm() nfc_llcp_recv_dm() handles DM(NOBOUND)/DM(REJ) for a socket that is still linked on local->connecting_sockets: it looks the socket up with nfc_llcp_connecting_sock_get(), sets sk->sk_state = LLCP_CLOSED and returns, without taking the socket lock and without unlinking the socket from the connecting_sockets list. llcp_sock_release() selects the list to unlink from by sk_state: a socket in LLCP_CONNECTING is unlinked from connecting_sockets, otherwise from the sockets list. Because recv_dm left the socket physically on connecting_sockets but in the LLCP_CLOSED state, release() takes the else branch and calls nfc_llcp_sock_unlink(&local->sockets, sk). That runs sk_del_node_init() while holding sockets.lock, i.e. it removes the socket from the connecting_sockets hlist under the wrong lock. A concurrent connect() linking another socket onto connecting_sockets under connecting_sockets.lock then mutates the same hlist unserialized, which corrupts the list and desyncs the sk_add_node()/sk_del_node_init() sock_hold()/__sock_put() pairing. An unprivileged local process holding LLCP sockets, with the DM supplied by the remote peer over an established LLCP link, can drive this to leak kernel sockets without bound (the mis-decrement goes through the non-freeing __sock_put() path, so the object is never released), leading to memory exhaustion / DoS. This is the same class of bug that was fixed in the sibling handler nfc_llcp_recv_cc() by commit b493ea2765cc ("nfc: llcp: Fix use-after-free race in nfc_llcp_recv_cc()"); recv_dm did not receive the equivalent fix. Fix it the same way: take lock_sock(), re-check that the socket is still hashed (release() may have won the race), and for the NOBOUND/REJ case unlink it from connecting_sockets before moving it to LLCP_CLOSED. The unlink drops the connecting_sockets membership reference via sk_del_node_init(), leaving the socket unhashed, so the later nfc_llcp_sock_unlink() in llcp_sock_release() becomes a no-op and no double put occurs. Fixes: a69f32af86e3 ("NFC: Socket linked list") Signed-off-by: Aldo Ariel Panzardo Link: https://patch.msgid.link/20260716232657.203145-1-qwe.aldo@gmail.com Signed-off-by: David Heidelberg --- net/nfc/llcp_core.c | 25 +++++++++++++++++++++++++ 1 file changed, 25 insertions(+) diff --git a/net/nfc/llcp_core.c b/net/nfc/llcp_core.c index cac1b5487064..bd6361e2efa4 100644 --- a/net/nfc/llcp_core.c +++ b/net/nfc/llcp_core.c @@ -1251,6 +1251,7 @@ static void nfc_llcp_recv_dm(struct nfc_llcp_local *local, struct nfc_llcp_sock *llcp_sock; struct sock *sk; u8 dsap, ssap, reason; + bool connecting = false; dsap = nfc_llcp_dsap(skb); ssap = nfc_llcp_ssap(skb); @@ -1262,6 +1263,7 @@ static void nfc_llcp_recv_dm(struct nfc_llcp_local *local, case LLCP_DM_NOBOUND: case LLCP_DM_REJ: llcp_sock = nfc_llcp_connecting_sock_get(local, dsap); + connecting = true; break; default: @@ -1276,10 +1278,33 @@ static void nfc_llcp_recv_dm(struct nfc_llcp_local *local, sk = &llcp_sock->sk; + lock_sock(sk); + + /* Check if socket was destroyed whilst waiting for the lock */ + if (!sk_hashed(sk)) { + release_sock(sk); + nfc_llcp_sock_put(llcp_sock); + return; + } + + /* + * For DM(NOBOUND)/DM(REJ) the socket is still linked on the + * connecting_sockets list. Unlink it here, under the socket lock, + * before moving it to LLCP_CLOSED: llcp_sock_release() selects the + * list to unlink from by sk_state, so leaving a connecting socket + * in the CLOSED state would make it unlink from the wrong list and + * corrupt the connecting_sockets list / desync the socket refcount. + * This mirrors nfc_llcp_recv_cc(). + */ + if (connecting) + nfc_llcp_sock_unlink(&local->connecting_sockets, sk); + sk->sk_err = ENXIO; sk->sk_state = LLCP_CLOSED; sk->sk_state_change(sk); + release_sock(sk); + nfc_llcp_sock_put(llcp_sock); } From 092c6a605cbd6414ef499834c2e0da69c2c3388e Mon Sep 17 00:00:00 2001 From: Doruk Tan Ozturk Date: Sat, 11 Jul 2026 14:36:51 +0200 Subject: [PATCH 108/189] nfc: port100: reject frames whose declared length exceeds the received data port100_recv_response() passes the URB transfer buffer to port100_rx_frame_is_valid(), which checksums le16_to_cpu(frame->datalen) bytes of frame->data. datalen is a 16-bit field supplied by the device and is never checked against the number of bytes actually received (urb->actual_length), so a device reporting a datalen larger than the received frame makes port100_data_checksum() read out of bounds past the transfer buffer. Reject a response whose declared frame size does not fit the received length before validating it. Found by 0sec (https://0sec.ai) using automated source analysis; the missing bound is evident from source. Compile-tested. Fixes: 562d4d59b8a1 ("NFC: Sony Port-100 Series driver") Cc: stable@vger.kernel.org Assisted-by: 0sec:claude-opus-4-8 Signed-off-by: Doruk Tan Ozturk Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260711123651.32595-1-doruk@0sec.ai Signed-off-by: David Heidelberg --- drivers/nfc/port100.c | 7 +++++++ 1 file changed, 7 insertions(+) diff --git a/drivers/nfc/port100.c b/drivers/nfc/port100.c index b613f5e2fd57..e769a30b8b5c 100644 --- a/drivers/nfc/port100.c +++ b/drivers/nfc/port100.c @@ -636,6 +636,13 @@ static void port100_recv_response(struct urb *urb) in_frame = dev->in_urb->transfer_buffer; + if (urb->actual_length < PORT100_FRAME_HEADER_LEN || + urb->actual_length < port100_rx_frame_size(in_frame)) { + nfc_err(&dev->interface->dev, "Received a truncated frame\n"); + cmd->status = -EIO; + goto sched_wq; + } + if (!port100_rx_frame_is_valid(in_frame)) { nfc_err(&dev->interface->dev, "Received an invalid frame\n"); cmd->status = -EIO; From 3d8afc5243ea2ee803d98e69eb4a01748167ac1b Mon Sep 17 00:00:00 2001 From: Lei Zhu Date: Wed, 29 Jul 2026 15:24:26 +0800 Subject: [PATCH 109/189] selftests: nci: Correct pthread_create return value check The pthread_create() functions returns 0 on success and a positive value on failure. Modify the return value check to correctly detect failure cases. Fixes: 72696bd8a09d ("selftests: nci: Extract the start/stop discovery function") Signed-off-by: Lei Zhu Link: https://patch.msgid.link/20260729072426.303484-1-zhulei_szu@163.com Signed-off-by: David Heidelberg --- tools/testing/selftests/nci/nci_dev.c | 12 +++++++----- 1 file changed, 7 insertions(+), 5 deletions(-) diff --git a/tools/testing/selftests/nci/nci_dev.c b/tools/testing/selftests/nci/nci_dev.c index 312f84ee0444..c053f5cf2745 100644 --- a/tools/testing/selftests/nci/nci_dev.c +++ b/tools/testing/selftests/nci/nci_dev.c @@ -438,7 +438,7 @@ FIXTURE_SETUP(NCI) else rc = pthread_create(&thread_t, NULL, virtual_dev_open, (void *)&self->virtual_nci_fd); - ASSERT_GT(rc, -1); + ASSERT_EQ(rc, 0); rc = send_cmd_with_idx(self->sd, self->fid, self->pid, NFC_CMD_DEV_UP, self->dev_idex); @@ -509,7 +509,7 @@ FIXTURE_TEARDOWN(NCI) rc = pthread_create(&thread_t, NULL, virtual_deinit, (void *)&self->virtual_nci_fd); - ASSERT_GT(rc, -1); + ASSERT_EQ(rc, 0); rc = send_cmd_with_idx(self->sd, self->fid, self->pid, NFC_CMD_DEV_DOWN, self->dev_idex); EXPECT_EQ(rc, 0); @@ -590,7 +590,7 @@ int start_polling(int dev_idx, int proto, int virtual_fd, int sd, int fid, int p rc = pthread_create(&thread_t, NULL, virtual_poll_start, (void *)&virtual_fd); - if (rc < 0) + if (rc) return rc; rc = send_cmd_mt_nla(sd, fid, pid, NFC_CMD_START_POLL, 2, nla_start_poll_type, @@ -610,7 +610,7 @@ int stop_polling(int dev_idx, int virtual_fd, int sd, int fid, int pid) rc = pthread_create(&thread_t, NULL, virtual_poll_stop, (void *)&virtual_fd); - if (rc < 0) + if (rc) return rc; rc = send_cmd_with_idx(sd, fid, pid, @@ -830,6 +830,8 @@ int disconnect_tag(int nfc_sock, int virtual_fd) status = pthread_create(&thread_t, NULL, virtual_deactivate_proc, (void *)&virtual_fd); + if (status) + return status; close(nfc_sock); pthread_join(thread_t, (void **)&status); @@ -874,7 +876,7 @@ TEST_F(NCI, deinit) else rc = pthread_create(&thread_t, NULL, virtual_deinit, (void *)&self->virtual_nci_fd); - ASSERT_GT(rc, -1); + ASSERT_EQ(rc, 0); rc = send_cmd_with_idx(self->sd, self->fid, self->pid, NFC_CMD_DEV_DOWN, self->dev_idex); From c3eef2f988a3db9690369d7cef9a3344dd9788d3 Mon Sep 17 00:00:00 2001 From: Lee Jones Date: Wed, 2 Sep 2026 12:30:31 +0000 Subject: [PATCH 110/189] nfc: llcp: Fix race condition in accept_queue lifecycle In nfc_llcp_socket_release(), sockets and listener accept queues are walked under the local sockets rwlock and bh_lock_sock(). However, bh_lock_sock() does not synchronise against process-context lock_sock() held by nfc_llcp_accept_dequeue() during accept(). Because socket_release() does not check sock_owned_by_user(), both paths can concurrently unlink and release the same child socket, resulting in use-after-free or a NULL pointer dereference of child->parent in nfc_llcp_accept_unlink(). Fix this synchronisation race by having nfc_llcp_socket_release() use process-context lock_sock() instead of bh_lock_sock(): 1. Pop sockets from the local sockets list under the write lock using nfc_llcp_sock_list_pop() so lock_sock() can be acquired without holding the rwlock. 2. Because lock_sock() can sleep, defer the final release of the nfc_llcp_local structure to a workqueue (release_work). This avoids a sleeping-in-atomic bug when the last local reference is dropped from softirq context. Additionally, hold a single device reference on local from registration until final destruction. 3. In nfc_llcp_local_get(), use kref_get_unless_zero() to prevent resurrecting a local object whose teardown has been scheduled. 4. In llcp_sock_accept(), verify that the listener socket state is still LLCP_LISTEN after waking from schedule_timeout() to prevent hangs if the listener is closed concurrently. 5. When unlinking unaccepted child sockets during listener release, unlink them from local->sockets, call sock_orphan(), and drop their initial sk_alloc creation reference via sock_put(). 6. Make nfc_llcp_accept_unlink() idempotent by guarding parent access with a NULL check. Fixes: 50b78b2a6500 ("NFC: Fix sleeping in atomic when releasing socket") Signed-off-by: Lee Jones Link: https://patch.msgid.link/20260902123033.1169067-1-lee@kernel.org Signed-off-by: David Heidelberg --- net/nfc/llcp.h | 1 + net/nfc/llcp_core.c | 123 +++++++++++++++++++++++++++----------------- net/nfc/llcp_sock.c | 51 +++++++++++++----- 3 files changed, 116 insertions(+), 59 deletions(-) diff --git a/net/nfc/llcp.h b/net/nfc/llcp.h index d8345ed57c95..23ae7a0112d3 100644 --- a/net/nfc/llcp.h +++ b/net/nfc/llcp.h @@ -91,6 +91,7 @@ struct nfc_llcp_local { struct hlist_head pending_sdreqs; struct timer_list sdreq_timer; struct work_struct sdreq_timeout_work; + struct work_struct release_work; u8 sdreq_next_tid; /* sockets array */ diff --git a/net/nfc/llcp_core.c b/net/nfc/llcp_core.c index bd6361e2efa4..23553e7426ec 100644 --- a/net/nfc/llcp_core.c +++ b/net/nfc/llcp_core.c @@ -20,6 +20,8 @@ static LIST_HEAD(llcp_devices); /* Protects llcp_devices list */ static DEFINE_SPINLOCK(llcp_devices_lock); +static struct workqueue_struct *llcp_wq; + static void nfc_llcp_rx_skb(struct nfc_llcp_local *local, struct sk_buff *skb); void nfc_llcp_sock_link(struct llcp_sock_list *l, struct sock *sk) @@ -63,21 +65,33 @@ static void nfc_llcp_socket_purge(struct nfc_llcp_sock *sock) } } +static struct sock *nfc_llcp_sock_list_pop(struct llcp_sock_list *l) +{ + struct sock *sk; + + write_lock(&l->lock); + sk = sk_head(&l->head); + if (sk) { + sock_hold(sk); + sk_del_node_init(sk); + } + write_unlock(&l->lock); + + return sk; +} + static void nfc_llcp_socket_release(struct nfc_llcp_local *local, bool device, int err) { struct sock *sk; - struct hlist_node *tmp; struct nfc_llcp_sock *llcp_sock; skb_queue_purge(&local->tx_queue); - write_lock(&local->sockets.lock); - - sk_for_each_safe(sk, tmp, &local->sockets.head) { + while ((sk = nfc_llcp_sock_list_pop(&local->sockets))) { llcp_sock = nfc_llcp_sock(sk); - bh_lock_sock(sk); + lock_sock(sk); nfc_llcp_socket_purge(llcp_sock); @@ -91,17 +105,27 @@ static void nfc_llcp_socket_release(struct nfc_llcp_local *local, bool device, list_for_each_entry_safe(lsk, n, &llcp_sock->accept_queue, accept_queue) { + bool put_creation = false; + accept_sk = &lsk->sk; - bh_lock_sock(accept_sk); + lock_sock_nested(accept_sk, + SINGLE_DEPTH_NESTING); - nfc_llcp_accept_unlink(accept_sk); + if (nfc_llcp_sock(accept_sk)->parent == sk) { + nfc_llcp_accept_unlink(accept_sk); + nfc_llcp_sock_unlink(&local->sockets, accept_sk); - if (err) - accept_sk->sk_err = err; - accept_sk->sk_state = LLCP_CLOSED; - accept_sk->sk_state_change(sk); + if (err) + accept_sk->sk_err = err; + accept_sk->sk_state = LLCP_CLOSED; + accept_sk->sk_state_change(accept_sk); + sock_orphan(accept_sk); + put_creation = true; + } - bh_unlock_sock(accept_sk); + release_sock(accept_sk); + if (put_creation) + sock_put(accept_sk); /* creation ref */ } } @@ -110,23 +134,18 @@ static void nfc_llcp_socket_release(struct nfc_llcp_local *local, bool device, sk->sk_state = LLCP_CLOSED; sk->sk_state_change(sk); - bh_unlock_sock(sk); - - sk_del_node_init(sk); + release_sock(sk); + sock_put(sk); } - write_unlock(&local->sockets.lock); - /* If we still have a device, we keep the RAW sockets alive */ if (device == true) return; - write_lock(&local->raw_sockets.lock); - - sk_for_each_safe(sk, tmp, &local->raw_sockets.head) { + while ((sk = nfc_llcp_sock_list_pop(&local->raw_sockets))) { llcp_sock = nfc_llcp_sock(sk); - bh_lock_sock(sk); + lock_sock(sk); nfc_llcp_socket_purge(llcp_sock); @@ -135,26 +154,20 @@ static void nfc_llcp_socket_release(struct nfc_llcp_local *local, bool device, sk->sk_state = LLCP_CLOSED; sk->sk_state_change(sk); - bh_unlock_sock(sk); - - sk_del_node_init(sk); + release_sock(sk); + sock_put(sk); } - - write_unlock(&local->raw_sockets.lock); } static struct nfc_llcp_local *nfc_llcp_local_get(struct nfc_llcp_local *local) { - /* Since using nfc_llcp_local may result in usage of nfc_dev, whenever - * we hold a reference to local, we also need to hold a reference to - * the device to avoid UAF. - */ - if (!nfc_get_device(local->dev->idx)) + if (!local) return NULL; - kref_get(&local->ref); + if (kref_get_unless_zero(&local->ref)) + return local; - return local; + return NULL; } static void local_cleanup(struct nfc_llcp_local *local) @@ -172,30 +185,34 @@ static void local_cleanup(struct nfc_llcp_local *local) nfc_llcp_free_sdp_tlv_list(&local->pending_sdreqs); } +static void local_release_work(struct work_struct *work) +{ + struct nfc_llcp_local *local; + struct nfc_dev *dev; + + local = container_of(work, struct nfc_llcp_local, release_work); + dev = local->dev; + + local_cleanup(local); + kfree(local); + nfc_put_device(dev); +} + static void local_release(struct kref *ref) { struct nfc_llcp_local *local; local = container_of(ref, struct nfc_llcp_local, ref); - local_cleanup(local); - kfree(local); + queue_work(llcp_wq, &local->release_work); } int nfc_llcp_local_put(struct nfc_llcp_local *local) { - struct nfc_dev *dev; - int ret; - - if (local == NULL) + if (!local) return 0; - dev = local->dev; - - ret = kref_put(&local->ref, local_release); - nfc_put_device(dev); - - return ret; + return kref_put(&local->ref, local_release); } static struct nfc_llcp_sock *nfc_llcp_sock_get(struct nfc_llcp_local *local, @@ -1705,6 +1722,7 @@ int nfc_llcp_register_device(struct nfc_dev *ndev) INIT_WORK(&local->rx_work, nfc_llcp_rx_work); INIT_WORK(&local->timeout_work, nfc_llcp_timeout_work); + INIT_WORK(&local->release_work, local_release_work); rwlock_init(&local->sockets.lock); rwlock_init(&local->connecting_sockets.lock); @@ -1748,10 +1766,23 @@ void nfc_llcp_unregister_device(struct nfc_dev *dev) int __init nfc_llcp_init(void) { - return nfc_llcp_sock_init(); + int ret; + + llcp_wq = alloc_workqueue("nfc_llcp_wq", WQ_UNBOUND, 0); + if (!llcp_wq) + return -ENOMEM; + + ret = nfc_llcp_sock_init(); + if (ret) { + destroy_workqueue(llcp_wq); + return ret; + } + + return 0; } void nfc_llcp_exit(void) { nfc_llcp_sock_exit(); + destroy_workqueue(llcp_wq); } diff --git a/net/nfc/llcp_sock.c b/net/nfc/llcp_sock.c index 5558d8a4d48b..ce6875eb58fb 100644 --- a/net/nfc/llcp_sock.c +++ b/net/nfc/llcp_sock.c @@ -392,11 +392,12 @@ void nfc_llcp_accept_unlink(struct sock *sk) pr_debug("state %d\n", sk->sk_state); - list_del_init(&llcp_sock->accept_queue); - sk_acceptq_removed(llcp_sock->parent); - llcp_sock->parent = NULL; - - sock_put(sk); + if (llcp_sock->parent) { + list_del_init(&llcp_sock->accept_queue); + sk_acceptq_removed(llcp_sock->parent); + llcp_sock->parent = NULL; + sock_put(sk); + } } void nfc_llcp_accept_enqueue(struct sock *parent, struct sock *sk) @@ -423,12 +424,20 @@ struct sock *nfc_llcp_accept_dequeue(struct sock *parent, list_for_each_entry_safe(lsk, n, &llcp_parent->accept_queue, accept_queue) { + struct nfc_llcp_local *local; + sk = &lsk->sk; - lock_sock(sk); + lock_sock_nested(sk, SINGLE_DEPTH_NESTING); if (sk->sk_state == LLCP_CLOSED) { - release_sock(sk); + local = nfc_llcp_sock(sk)->local; + nfc_llcp_accept_unlink(sk); + if (local) + nfc_llcp_sock_unlink(&local->sockets, sk); + sock_orphan(sk); + release_sock(sk); + sock_put(sk); continue; } @@ -464,7 +473,7 @@ static int llcp_sock_accept(struct socket *sock, struct socket *newsock, pr_debug("parent %p\n", sk); - lock_sock_nested(sk, SINGLE_DEPTH_NESTING); + lock_sock(sk); if (sk->sk_state != LLCP_LISTEN) { ret = -EBADFD; @@ -490,7 +499,12 @@ static int llcp_sock_accept(struct socket *sock, struct socket *newsock, release_sock(sk); timeo = schedule_timeout(timeo); - lock_sock_nested(sk, SINGLE_DEPTH_NESTING); + lock_sock(sk); + + if (sk->sk_state != LLCP_LISTEN) { + ret = -EBADFD; + break; + } } __set_current_state(TASK_RUNNING); remove_wait_queue(sk_sleep(sk), &wait); @@ -629,13 +643,24 @@ static int llcp_sock_release(struct socket *sock) list_for_each_entry_safe(lsk, n, &llcp_sock->accept_queue, accept_queue) { - accept_sk = &lsk->sk; - lock_sock(accept_sk); + bool put_creation = false; - nfc_llcp_send_disconnect(lsk); - nfc_llcp_accept_unlink(accept_sk); + accept_sk = &lsk->sk; + lock_sock_nested(accept_sk, SINGLE_DEPTH_NESTING); + + if (nfc_llcp_sock(accept_sk)->parent == sk) { + nfc_llcp_send_disconnect(lsk); + nfc_llcp_accept_unlink(accept_sk); + nfc_llcp_sock_unlink(&local->sockets, accept_sk); + + accept_sk->sk_state = LLCP_CLOSED; + sock_orphan(accept_sk); + put_creation = true; + } release_sock(accept_sk); + if (put_creation) + sock_put(accept_sk); /* creation ref */ } } From eda518d2cdb6074a0bcdfa06af291616bcb5c421 Mon Sep 17 00:00:00 2001 From: Chaithanya Lagisetty Date: Tue, 1 Sep 2026 07:06:18 +0000 Subject: [PATCH 111/189] selftests: nci: Fix uninitialized family ID on missing attribute get_family_id() walks the generic netlink CTRL_CMD_GETFAMILY reply looking for the CTRL_ATTR_FAMILY_ID attribute and returns the parsed value in the local variable "id". If the reply does not carry that attribute, the parsing loop never assigns "id" and the function returns an indeterminate stack value, which the caller stores in self->fid and uses for subsequent netlink requests. Initialize "id" to 0 so a missing attribute yields a deterministic (invalid) family ID instead of a garbage value. Fixes: f595cf1242f3 ("selftests: Add nci suite") Signed-off-by: Chaithanya Lagisetty Reviewed-by: Hangbin Liu Link: https://patch.msgid.link/20260901070618.3299012-1-nagachaithanya9911@gmail.com Signed-off-by: David Heidelberg --- tools/testing/selftests/nci/nci_dev.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tools/testing/selftests/nci/nci_dev.c b/tools/testing/selftests/nci/nci_dev.c index c053f5cf2745..23fd38acfcf4 100644 --- a/tools/testing/selftests/nci/nci_dev.c +++ b/tools/testing/selftests/nci/nci_dev.c @@ -182,7 +182,7 @@ static int get_family_id(int sd, __u32 pid, __u32 *event_group) } ans; struct nlattr *na; int resp_len; - __u16 id; + __u16 id = 0; int len; int rc; From 6be581aeffc215bfc77939cd59902b0dbc4af23e Mon Sep 17 00:00:00 2001 From: Chris Gellermann Date: Fri, 4 Sep 2026 11:59:15 +0200 Subject: [PATCH 112/189] selftests/nci: Fix out-of-bounds store on thread join The NCI test collects the exit status of its helper threads by passing the address of an int to pthread_join(): int status; ... pthread_join(thread_t, (void **) &status); pthread_join() stores a void pointer to the memory location. On 64-bit systems, a void pointer is wider than an int, so the store overruns the 4 bytes of space allocated on the stack for the integer and corrupts the adjacent stack. On our CHERI system, this caused a fault due to a capability bounds violation. Fix this by introducing a helper that joins a thread through a void pointer and converts the result back to an integer, which is what the helper threads return. While here, also fix the logic in disconnect_tag() if the helper thread creation failed. Previously, it would have joined a thread that was never created when pthread_create() failed. Fixes: f595cf1242f3 ("selftests: Add nci suite") Signed-off-by: Chris Gellermann Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260904095915.3372241-1-christian.gellermann@codasip.com Signed-off-by: David Heidelberg --- tools/testing/selftests/nci/nci_dev.c | 31 +++++++++++++++++---------- 1 file changed, 20 insertions(+), 11 deletions(-) diff --git a/tools/testing/selftests/nci/nci_dev.c b/tools/testing/selftests/nci/nci_dev.c index 23fd38acfcf4..07427fa42888 100644 --- a/tools/testing/selftests/nci/nci_dev.c +++ b/tools/testing/selftests/nci/nci_dev.c @@ -8,6 +8,7 @@ #include #include +#include #include #include #include @@ -87,6 +88,16 @@ struct msgtemplate { char buf[MAX_MSG_SIZE]; }; +static int join_thread_status(pthread_t thread) +{ + void *thread_ret = NULL; + + if (pthread_join(thread, &thread_ret)) + return -1; + + return (int)(intptr_t)thread_ret; +} + static int create_nl_socket(void) { int fd; @@ -444,7 +455,7 @@ FIXTURE_SETUP(NCI) NFC_CMD_DEV_UP, self->dev_idex); EXPECT_EQ(rc, 0); - pthread_join(thread_t, (void **)&status); + status = join_thread_status(thread_t); ASSERT_EQ(status, 0); self->open_state = true; } @@ -514,7 +525,7 @@ FIXTURE_TEARDOWN(NCI) NFC_CMD_DEV_DOWN, self->dev_idex); EXPECT_EQ(rc, 0); - pthread_join(thread_t, (void **)&status); + status = join_thread_status(thread_t); ASSERT_EQ(status, 0); } @@ -585,7 +596,6 @@ int start_polling(int dev_idx, int proto, int virtual_fd, int sd, int fid, int p void *nla_start_poll_data[2] = {&dev_idx, &proto}; int nla_start_poll_len[2] = {4, 4}; pthread_t thread_t; - int status; int rc; rc = pthread_create(&thread_t, NULL, virtual_poll_start, @@ -598,14 +608,12 @@ int start_polling(int dev_idx, int proto, int virtual_fd, int sd, int fid, int p if (rc != 0) return rc; - pthread_join(thread_t, (void **)&status); - return status; + return join_thread_status(thread_t); } int stop_polling(int dev_idx, int virtual_fd, int sd, int fid, int pid) { pthread_t thread_t; - int status; int rc; rc = pthread_create(&thread_t, NULL, virtual_poll_stop, @@ -618,8 +626,7 @@ int stop_polling(int dev_idx, int virtual_fd, int sd, int fid, int pid) if (rc != 0) return rc; - pthread_join(thread_t, (void **)&status); - return status; + return join_thread_status(thread_t); } TEST_F(NCI, start_poll) @@ -834,8 +841,10 @@ int disconnect_tag(int nfc_sock, int virtual_fd) return status; close(nfc_sock); - pthread_join(thread_t, (void **)&status); - return status; + if (status) + return -1; + + return join_thread_status(thread_t); } TEST_F(NCI, t4t_tag_read) @@ -882,7 +891,7 @@ TEST_F(NCI, deinit) NFC_CMD_DEV_DOWN, self->dev_idex); EXPECT_EQ(rc, 0); - pthread_join(thread_t, (void **)&status); + status = join_thread_status(thread_t); self->open_state = 0; ASSERT_EQ(status, 0); From 51814683e28fc64eceb415962376956c3cfc75a7 Mon Sep 17 00:00:00 2001 From: Chris Gellermann Date: Fri, 4 Sep 2026 18:42:52 +0200 Subject: [PATCH 113/189] nfc: virtual_ncidev: Add missing ioctl compat handler The compat handler for ioctls to the virtual nci device is missing. So, nci-specific ioctls of a compat task return with -1 and errno set to ENOTTY. Add a handler. The handling of an ioctl() call of a compat task to get the index of virtual nci device (IOCTL_GET_NCIDEV_IDX) lands in the default case of the ioctl compat handler (see fs/ioctl.c): COMPAT_SYSCALL_DEFINE3(ioctl, ...) { ... default: error = do_vfs_ioctl(fd_file(f), fd, cmd, ...); if (error != -ENOIOCTLCMD) break; if (fd_file(f)->f_op->compat_ioctl) error = fd_file(f)->f_op->compat_ioctl(fd_file(f), cmd, arg); if (error == -ENOIOCTLCMD) error = -ENOTTY; ... } There, do_vfs_ioctl() returns -ENOIOCTLCMD and compat_ioctl is not set for virtual_ncidev_fops, i.e. f_op->compat_ioctl == NULL. So, the ioctl() syscall returns with -1 and errno set to ENOTTY to the compat task. To fix this, use the compat_ptr_ioctl helper for compat handling here. It shall be used for ioctls that "either ignore the argument or pass a pointer to a compatible data type". The driver's sole ioctl takes a user void pointer and copies nfc_dev->idx to it, a 4-byte integer across all ABIs. This issue has been found by running the nci_dev kernel selftest as rv64 binary on top of a CHERI kernel, where the ioctl() ends up in the ioctl compat handler, similar to a 32-bit application on top of a 64-bit kernel. Fixes: e624e6c3e777 ("nfc: Add a virtual nci device driver") Signed-off-by: Chris Gellermann Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260904164252.18351-1-christian.gellermann@codasip.com Signed-off-by: David Heidelberg --- drivers/nfc/virtual_ncidev.c | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/drivers/nfc/virtual_ncidev.c b/drivers/nfc/virtual_ncidev.c index 8eeb447ac96e..e51c647b27eb 100644 --- a/drivers/nfc/virtual_ncidev.c +++ b/drivers/nfc/virtual_ncidev.c @@ -195,7 +195,8 @@ static const struct file_operations virtual_ncidev_fops = { .write = virtual_ncidev_write, .open = virtual_ncidev_open, .release = virtual_ncidev_close, - .unlocked_ioctl = virtual_ncidev_ioctl + .unlocked_ioctl = virtual_ncidev_ioctl, + .compat_ioctl = compat_ptr_ioctl, }; static struct miscdevice miscdev = { From 273f9d667cde649f8de9d72b1303cc2f4b658c50 Mon Sep 17 00:00:00 2001 From: Aamir Ahmed Date: Tue, 15 Sep 2026 19:54:27 +0100 Subject: [PATCH 114/189] nfc: llcp: drop truncated I/RR/RNR PDUs in nfc_llcp_recv_hdlc() nfc_llcp_recv_hdlc() reads the sequence byte skb->data[2], via nfc_llcp_ns()/nfc_llcp_nr(), before any length check. The receive path only guarantees the two-byte LLCP header -- __nfc_llcp_recv() checks it with pskb_may_pull() and nfc_llcp_recv_agf() admits two-byte inner PDUs -- so a two-byte I, RR or RNR PDU reads one byte of uninitialised skb tailroom. The byte becomes N(R)/N(S); a peer can already set those with a well-formed PDU, so this is acting on uninitialised memory, not new peer control. Guard the read with pskb_may_pull(), as commit 95674f506c63 ("nfc: llcp: reject PDUs shorter than the LLCP header") did for the two-byte header, so the sequence byte is present and linear before it is read. RR and RNR PDUs are LLCP_HEADER_SIZE + LLCP_SEQUENCE_SIZE bytes and an I PDU is longer, so no valid frame is rejected; a truncated PDU is malformed, so return without a DM reply. Fixes: d646960f7986 ("NFC: Initial LLCP support") Cc: stable@vger.kernel.org Assisted-by: LLM Signed-off-by: Aamir Ahmed Reviewed-by: Simon Horman Link: https://patch.msgid.link/AS8P251MB0001789BBF04B72745C7D96BC8BA2@AS8P251MB0001.EURP251.PROD.OUTLOOK.COM Signed-off-by: David Heidelberg --- net/nfc/llcp_core.c | 3 +++ 1 file changed, 3 insertions(+) diff --git a/net/nfc/llcp_core.c b/net/nfc/llcp_core.c index 23553e7426ec..5017f6aa57a0 100644 --- a/net/nfc/llcp_core.c +++ b/net/nfc/llcp_core.c @@ -1091,6 +1091,9 @@ static void nfc_llcp_recv_hdlc(struct nfc_llcp_local *local, struct sock *sk; u8 dsap, ssap, ptype, ns, nr; + if (!pskb_may_pull(skb, LLCP_HEADER_SIZE + LLCP_SEQUENCE_SIZE)) + return; + ptype = nfc_llcp_ptype(skb); dsap = nfc_llcp_dsap(skb); ssap = nfc_llcp_ssap(skb); From 66f4300206b82b0b143ef0d9be90cd8d29f23a47 Mon Sep 17 00:00:00 2001 From: Cong Nguyen Date: Mon, 14 Sep 2026 19:11:29 +0700 Subject: [PATCH 115/189] nfc: llcp: fix sdreq TLV list leak on parse/alloc/send failure nfc_genl_llc_sdreq() builds a list of TLV nodes while walking nested netlink attrs, but 3 error paths (nested-attr parse failure, TLV alloc ENOMEM, nfc_llcp_send_snl_sdreq() failure) all skip freeing what was already queued. Route them through a new free_list label, mirroring the SDRES path in the same file which already does this. Harmless on the success path too -- send_snl_sdreq() drains the list as it moves nodes, so it's already empty by the time free_list runs. Fixes: d9b8d8e19b07 ("NFC: llcp: Service Name Lookup netlink interface") Assisted-by: Claude:claude-opus-4 Signed-off-by: Cong Nguyen Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260914121129.2098606-1-congnt264@gmail.com Signed-off-by: David Heidelberg --- net/nfc/netlink.c | 7 +++++-- 1 file changed, 5 insertions(+), 2 deletions(-) diff --git a/net/nfc/netlink.c b/net/nfc/netlink.c index 0c58824cb150..224bdfa2dd0d 100644 --- a/net/nfc/netlink.c +++ b/net/nfc/netlink.c @@ -1181,7 +1181,7 @@ static int nfc_genl_llc_sdreq(struct sk_buff *skb, struct genl_info *info) if (rc != 0) { rc = -EINVAL; - goto put_local; + goto free_list; } if (!sdp_attrs[NFC_SDP_ATTR_URI]) @@ -1200,7 +1200,7 @@ static int nfc_genl_llc_sdreq(struct sk_buff *skb, struct genl_info *info) sdreq = nfc_llcp_build_sdreq_tlv(tid, uri, uri_len); if (sdreq == NULL) { rc = -ENOMEM; - goto put_local; + goto free_list; } tlvs_len += sdreq->tlv_len; @@ -1215,6 +1215,9 @@ static int nfc_genl_llc_sdreq(struct sk_buff *skb, struct genl_info *info) rc = nfc_llcp_send_snl_sdreq(local, &sdreq_list, tlvs_len); +free_list: + nfc_llcp_free_sdp_tlv_list(&sdreq_list); + put_local: nfc_llcp_local_put(local); From d2acbde7e67df44efa8f0963462d1192e7694ffc Mon Sep 17 00:00:00 2001 From: Myeonghun Pak Date: Sun, 13 Sep 2026 00:26:25 -0400 Subject: [PATCH 116/189] nfc: trf7970a: power down on startup RX gain failure trf7970a_startup() powers up the device before applying the optional RX gain reduction. If the register read or write fails, it returns without undoing that power-up. Probe's unwind only drops the separate regulator references acquired by probe, leaving the additional VIN enable from startup unbalanced. The system resume caller also has no power-down on this error. Call trf7970a_power_down() before returning the RX gain error to deassert the enable GPIOs, release the startup VIN reference and restore the powered-off state. Runtime PM has not been enabled yet, so the full shutdown helper is not appropriate here. Preserve the original SPI error. This issue was identified during our ongoing static-analysis research while reviewing kernel code. Fixes: 5d69351820ea ("NFC: trf7970a: Create device-tree parameter for RX gain reduction") Cc: stable@vger.kernel.org Assisted-by: OpenAI:GPT-5.6 Co-developed-by: Ijae Kim Signed-off-by: Ijae Kim Signed-off-by: Myeonghun Pak Reviewed-by: Paul Geurts Link: https://patch.msgid.link/20260913042625.31296-1-mhun512@gmail.com Signed-off-by: David Heidelberg --- drivers/nfc/trf7970a.c | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/drivers/nfc/trf7970a.c b/drivers/nfc/trf7970a.c index 60883001fa5d..ddfc58c29964 100644 --- a/drivers/nfc/trf7970a.c +++ b/drivers/nfc/trf7970a.c @@ -1997,8 +1997,10 @@ static int trf7970a_startup(struct trf7970a *trf) return ret; ret = trf7970a_update_rx_gain_reduction(trf); - if (ret) + if (ret) { + trf7970a_power_down(trf); return ret; + } pm_runtime_set_active(trf->dev); pm_runtime_enable(trf->dev); From dcab71a7011918f6fdba7adcec02d217dcb84b8d Mon Sep 17 00:00:00 2001 From: Luxiao Xu Date: Wed, 9 Sep 2026 13:19:24 +0800 Subject: [PATCH 117/189] nfc: fix use-after-free in nfc_get_local_general_bytes Commit 6709d4b7bc2e ("net: nfc: Fix use-after-free caused by nfc_llcp_find_local") attempted to fix a use-after-free (UAF) issue by invoking nfc_llcp_local_put(local) after accessing local->gb. However, if the reference count drops to zero, local is freed immediately, leading to a use-after-free when callers access the returned pointer. Alternative approaches using dynamic allocation (e.g. kmemdup) introduced memory leaks because callers consistently treat the returned pointer as borrowed memory. Fix this properly by refactoring nfc_llcp_general_bytes() and nfc_get_local_general_bytes() to accept a caller-provided output buffer (out_gb) and its maximum length (gb_max_len). The general bytes are safely copied into out_gb before calling nfc_llcp_local_put(local), ensuring safe lifetime management without ownership transfer complications. Update all callers across drivers (microread, pn533, pn544, st21nfca, digital_dep, and nci) to provide their own destination buffers and pass them to nfc_get_local_general_bytes(). Fixes: 6709d4b7bc2e ("net: nfc: Fix use-after-free caused by nfc_llcp_find_local") Cc: stable@vger.kernel.org Reported-by: Vega Assisted-by: LLM Signed-off-by: Luxiao Xu Signed-off-by: Ren Wei Reviewed-by: Simon Horman Link: https://patch.msgid.link/3cbaac3bee23f8ff3a3284ed32d347696eb1d208.1788841683.git.rakukuip@gmail.com Signed-off-by: David Heidelberg --- drivers/nfc/microread/microread.c | 6 +++--- drivers/nfc/pn533/pn533.c | 14 ++++++++------ drivers/nfc/pn533/pn533.h | 4 +++- drivers/nfc/pn544/pn544.c | 7 +++---- drivers/nfc/st21nfca/core.c | 8 ++++---- include/net/nfc/hci.h | 2 +- include/net/nfc/nfc.h | 3 ++- net/nfc/core.c | 15 +++++++-------- net/nfc/digital_dep.c | 8 ++++---- net/nfc/llcp_core.c | 17 +++++++++++++---- net/nfc/nci/core.c | 10 +++++----- net/nfc/nfc.h | 3 ++- 12 files changed, 55 insertions(+), 42 deletions(-) diff --git a/drivers/nfc/microread/microread.c b/drivers/nfc/microread/microread.c index dfa2490db545..2bfafa94e83d 100644 --- a/drivers/nfc/microread/microread.c +++ b/drivers/nfc/microread/microread.c @@ -251,9 +251,9 @@ static int microread_start_poll(struct nfc_hci_dev *hdev, param[1] |= (1 << 1); if ((im_protocols | tm_protocols) & NFC_PROTO_NFC_DEP_MASK) { - hdev->gb = nfc_get_local_general_bytes(hdev->ndev, - &hdev->gb_len); - if (hdev->gb == NULL || hdev->gb_len == 0) { + nfc_get_local_general_bytes(hdev->ndev, hdev->gb, + sizeof(hdev->gb), &hdev->gb_len); + if (hdev->gb_len == 0) { im_protocols &= ~NFC_PROTO_NFC_DEP_MASK; tm_protocols &= ~NFC_PROTO_NFC_DEP_MASK; } diff --git a/drivers/nfc/pn533/pn533.c b/drivers/nfc/pn533/pn533.c index f5a6a7c20d5a..b0133e51dce9 100644 --- a/drivers/nfc/pn533/pn533.c +++ b/drivers/nfc/pn533/pn533.c @@ -1357,10 +1357,11 @@ static int pn533_poll_dep(struct nfc_dev *nfc_dev) u8 *next, nfcid3[NFC_NFCID3_MAXSIZE]; u8 passive_data[PASSIVE_DATA_LEN] = {0x00, 0xff, 0xff, 0x00, 0x3}; - if (!dev->gb) { - dev->gb = nfc_get_local_general_bytes(nfc_dev, &dev->gb_len); - - if (!dev->gb || !dev->gb_len) { + if (!dev->gb_len) { + nfc_get_local_general_bytes(nfc_dev, dev->gb, + sizeof(dev->gb), + &dev->gb_len); + if (!dev->gb_len) { dev->poll_dep = 0; queue_work(dev->wq, &dev->rf_work); } @@ -1658,8 +1659,9 @@ static int pn533_start_poll(struct nfc_dev *nfc_dev, } if (tm_protocols) { - dev->gb = nfc_get_local_general_bytes(nfc_dev, &dev->gb_len); - if (dev->gb == NULL) + nfc_get_local_general_bytes(nfc_dev, dev->gb, + sizeof(dev->gb), &dev->gb_len); + if (dev->gb_len == 0) tm_protocols = 0; } diff --git a/drivers/nfc/pn533/pn533.h b/drivers/nfc/pn533/pn533.h index 09e35b8693f5..5ab668e05121 100644 --- a/drivers/nfc/pn533/pn533.h +++ b/drivers/nfc/pn533/pn533.h @@ -6,6 +6,8 @@ * Copyright (C) 2012-2013 Tieto Poland */ +#include + #define PN533_DEVICE_STD 0x1 #define PN533_DEVICE_PASORI 0x2 #define PN533_DEVICE_ACR122U 0x3 @@ -166,7 +168,7 @@ struct pn533 { struct timer_list listen_timer; int cancel_listen; - u8 *gb; + u8 gb[NFC_MAX_GT_LEN]; size_t gb_len; u8 tgt_available_prots; diff --git a/drivers/nfc/pn544/pn544.c b/drivers/nfc/pn544/pn544.c index 9d0a16ac465e..c4fa70e45c14 100644 --- a/drivers/nfc/pn544/pn544.c +++ b/drivers/nfc/pn544/pn544.c @@ -377,10 +377,9 @@ static int pn544_hci_start_poll(struct nfc_hci_dev *hdev, return r; if ((im_protocols | tm_protocols) & NFC_PROTO_NFC_DEP_MASK) { - hdev->gb = nfc_get_local_general_bytes(hdev->ndev, - &hdev->gb_len); - pr_debug("generate local bytes %p\n", hdev->gb); - if (hdev->gb == NULL || hdev->gb_len == 0) { + nfc_get_local_general_bytes(hdev->ndev, hdev->gb, + sizeof(hdev->gb), &hdev->gb_len); + if (hdev->gb_len == 0) { im_protocols &= ~NFC_PROTO_NFC_DEP_MASK; tm_protocols &= ~NFC_PROTO_NFC_DEP_MASK; } diff --git a/drivers/nfc/st21nfca/core.c b/drivers/nfc/st21nfca/core.c index fd39a05c9622..6bfeb8e7ed89 100644 --- a/drivers/nfc/st21nfca/core.c +++ b/drivers/nfc/st21nfca/core.c @@ -351,10 +351,10 @@ static int st21nfca_hci_start_poll(struct nfc_hci_dev *hdev, if (r < 0) return r; } else { - hdev->gb = nfc_get_local_general_bytes(hdev->ndev, - &hdev->gb_len); - - if (hdev->gb == NULL || hdev->gb_len == 0) { + nfc_get_local_general_bytes(hdev->ndev, hdev->gb, + sizeof(hdev->gb), + &hdev->gb_len); + if (hdev->gb_len == 0) { im_protocols &= ~NFC_PROTO_NFC_DEP_MASK; tm_protocols &= ~NFC_PROTO_NFC_DEP_MASK; } diff --git a/include/net/nfc/hci.h b/include/net/nfc/hci.h index 756c11084f65..86ed63e5d533 100644 --- a/include/net/nfc/hci.h +++ b/include/net/nfc/hci.h @@ -144,7 +144,7 @@ struct nfc_hci_dev { data_exchange_cb_t async_cb; void *async_cb_context; - u8 *gb; + u8 gb[NFC_MAX_GT_LEN]; size_t gb_len; unsigned long quirks; diff --git a/include/net/nfc/nfc.h b/include/net/nfc/nfc.h index c54df042db6b..bcafab5c53e5 100644 --- a/include/net/nfc/nfc.h +++ b/include/net/nfc/nfc.h @@ -273,7 +273,8 @@ struct sk_buff *nfc_alloc_recv_skb(unsigned int size, gfp_t gfp); int nfc_set_remote_general_bytes(struct nfc_dev *dev, const u8 *gt, u8 gt_len); -u8 *nfc_get_local_general_bytes(struct nfc_dev *dev, size_t *gb_len); +u8 *nfc_get_local_general_bytes(struct nfc_dev *dev, u8 *out_gb, + size_t gb_max_len, size_t *gb_len); int nfc_fw_download_done(struct nfc_dev *dev, const char *firmware_name, u32 result); diff --git a/net/nfc/core.c b/net/nfc/core.c index a92a6566e6a0..f521669293f0 100644 --- a/net/nfc/core.c +++ b/net/nfc/core.c @@ -279,10 +279,10 @@ static struct nfc_target *nfc_find_target(struct nfc_dev *dev, u32 target_idx) int nfc_dep_link_up(struct nfc_dev *dev, int target_index, u8 comm_mode) { - int rc = 0; - u8 *gb; - size_t gb_len; struct nfc_target *target; + u8 gb[NFC_MAX_GT_LEN]; + size_t gb_len = 0; + int rc = 0; pr_debug("dev_name=%s comm %d\n", dev_name(&dev->dev), comm_mode); @@ -301,7 +301,7 @@ int nfc_dep_link_up(struct nfc_dev *dev, int target_index, u8 comm_mode) goto error; } - gb = nfc_llcp_general_bytes(dev, &gb_len); + nfc_get_local_general_bytes(dev, gb, sizeof(gb), &gb_len); if (gb_len > NFC_MAX_GT_LEN) { rc = -EINVAL; goto error; @@ -644,11 +644,10 @@ int nfc_set_remote_general_bytes(struct nfc_dev *dev, const u8 *gb, u8 gb_len) } EXPORT_SYMBOL(nfc_set_remote_general_bytes); -u8 *nfc_get_local_general_bytes(struct nfc_dev *dev, size_t *gb_len) +u8 *nfc_get_local_general_bytes(struct nfc_dev *dev, u8 *out_gb, + size_t gb_max_len, size_t *gb_len) { - pr_debug("dev_name=%s\n", dev_name(&dev->dev)); - - return nfc_llcp_general_bytes(dev, gb_len); + return nfc_llcp_general_bytes(dev, out_gb, gb_max_len, gb_len); } EXPORT_SYMBOL(nfc_get_local_general_bytes); diff --git a/net/nfc/digital_dep.c b/net/nfc/digital_dep.c index 3982fa084737..968547c306a5 100644 --- a/net/nfc/digital_dep.c +++ b/net/nfc/digital_dep.c @@ -1490,14 +1490,14 @@ static int digital_tg_send_atr_res(struct nfc_digital_dev *ddev, struct digital_atr_req *atr_req) { struct digital_atr_res *atr_res; + u8 gb[NFC_MAX_GT_LEN]; struct sk_buff *skb; - u8 *gb, payload_bits; + u8 payload_bits; size_t gb_len; int rc; - gb = nfc_get_local_general_bytes(ddev->nfc_dev, &gb_len); - if (!gb) - gb_len = 0; + nfc_get_local_general_bytes(ddev->nfc_dev, gb, sizeof(gb), + &gb_len); skb = digital_skb_alloc(ddev, sizeof(struct digital_atr_res) + gb_len); if (!skb) diff --git a/net/nfc/llcp_core.c b/net/nfc/llcp_core.c index 5017f6aa57a0..5ce3ce64baf2 100644 --- a/net/nfc/llcp_core.c +++ b/net/nfc/llcp_core.c @@ -652,23 +652,32 @@ static int nfc_llcp_build_gb(struct nfc_llcp_local *local) return ret; } -u8 *nfc_llcp_general_bytes(struct nfc_dev *dev, size_t *general_bytes_len) +u8 *nfc_llcp_general_bytes(struct nfc_dev *dev, u8 *out_gb, size_t gb_max_len, + size_t *general_bytes_len) { struct nfc_llcp_local *local; + if (!out_gb || !general_bytes_len) + return NULL; + local = nfc_llcp_find_local(dev); - if (local == NULL) { + if (!local) { *general_bytes_len = 0; return NULL; } nfc_llcp_build_gb(local); - *general_bytes_len = local->gb_len; + if (local->gb_len) { + *general_bytes_len = min_t(size_t, local->gb_len, gb_max_len); + memcpy(out_gb, local->gb, *general_bytes_len); + } else { + *general_bytes_len = 0; + } nfc_llcp_local_put(local); - return local->gb; + return out_gb; } int nfc_llcp_set_remote_gb(struct nfc_dev *dev, const u8 *gb, u8 gb_len) diff --git a/net/nfc/nci/core.c b/net/nfc/nci/core.c index 5f46c4b5720f..73e3a96470ac 100644 --- a/net/nfc/nci/core.c +++ b/net/nfc/nci/core.c @@ -780,15 +780,15 @@ static int nci_set_local_general_bytes(struct nfc_dev *nfc_dev) { struct nci_dev *ndev = nfc_get_drvdata(nfc_dev); struct nci_set_config_param param; + u8 gb[NFC_MAX_GT_LEN]; int rc; - param.val = nfc_get_local_general_bytes(nfc_dev, ¶m.len); - if ((param.val == NULL) || (param.len == 0)) + nfc_get_local_general_bytes(nfc_dev, gb, sizeof(gb), + ¶m.len); + if (param.len == 0) return 0; - if (param.len > NFC_MAX_GT_LEN) - return -EINVAL; - + param.val = gb; param.id = NCI_PN_ATR_REQ_GEN_BYTES; rc = nci_request(ndev, nci_set_config_req, ¶m, diff --git a/net/nfc/nfc.h b/net/nfc/nfc.h index 0b1e6466f4fb..82c5dfdad10e 100644 --- a/net/nfc/nfc.h +++ b/net/nfc/nfc.h @@ -49,7 +49,8 @@ void nfc_llcp_mac_is_up(struct nfc_dev *dev, u32 target_idx, int nfc_llcp_register_device(struct nfc_dev *dev); void nfc_llcp_unregister_device(struct nfc_dev *dev); int nfc_llcp_set_remote_gb(struct nfc_dev *dev, const u8 *gb, u8 gb_len); -u8 *nfc_llcp_general_bytes(struct nfc_dev *dev, size_t *general_bytes_len); +u8 *nfc_llcp_general_bytes(struct nfc_dev *dev, u8 *out_gb, size_t gb_max_len, + size_t *general_bytes_len); int nfc_llcp_data_received(struct nfc_dev *dev, struct sk_buff *skb); struct nfc_llcp_local *nfc_llcp_find_local(struct nfc_dev *dev); int nfc_llcp_local_put(struct nfc_llcp_local *local); From 7f2ea5ed588c03d481f0301e6c3d4240132383fb Mon Sep 17 00:00:00 2001 From: Pengpeng Hou Date: Sun, 30 Aug 2026 21:29:58 +0800 Subject: [PATCH 118/189] nfc: st21nfca: validate ISO15693 inventory length The ISO15693 inventory helper removes a two-byte prefix without checking that it exists, then accepts a one-byte remainder before reading data[1] as the DSFID. Require the prefix and at least two remaining bytes before copying the UID data and reading the DSFID. Fixes: 7974728094d3 ("NFC: st21nfca: Add ISO15693 Reader/Writer support") Signed-off-by: Pengpeng Hou Link: https://patch.msgid.link/20260830132958.6397-1-pengpeng@iscas.ac.cn Signed-off-by: David Heidelberg --- drivers/nfc/st21nfca/core.c | 4 +--- 1 file changed, 1 insertion(+), 3 deletions(-) diff --git a/drivers/nfc/st21nfca/core.c b/drivers/nfc/st21nfca/core.c index 6bfeb8e7ed89..b5c1ca3acfbe 100644 --- a/drivers/nfc/st21nfca/core.c +++ b/drivers/nfc/st21nfca/core.c @@ -577,9 +577,7 @@ static int st21nfca_get_iso15693_inventory(struct nfc_hci_dev *hdev, if (r < 0) goto exit; - skb_pull(inventory_skb, 2); - - if (inventory_skb->len == 0 || + if (!skb_pull(inventory_skb, 2) || inventory_skb->len < 2 || inventory_skb->len > NFC_ISO15693_UID_MAXSIZE) { r = -EPROTO; goto exit; From c04981e42d94f39c1dba965cc462a046e946a6c5 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?=C3=96mer=20Mete=20Kaya?= Date: Wed, 9 Sep 2026 15:16:22 +0300 Subject: [PATCH 119/189] nfc: llcp: fix -ENOMEM on connect with zero-length service name MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit When service_name_len is 0, kmemdup() returns ZERO_SIZE_PTR which passes the NULL check, causing nfc_llcp_send_connect() to attempt building a zero-length service name TLV and fail with -ENOMEM. Fix by setting service_name to NULL directly when service_name_len is 0. Fixes: d646960f7986 ("NFC: Initial LLCP support") Signed-off-by: Ömer Mete Kaya Link: https://patch.msgid.link/20260909122029.34081-1-omermetekaya0@gmail.com Signed-off-by: David Heidelberg --- net/nfc/llcp_sock.c | 16 ++++++++++------ 1 file changed, 10 insertions(+), 6 deletions(-) diff --git a/net/nfc/llcp_sock.c b/net/nfc/llcp_sock.c index ce6875eb58fb..1e5ee4bcde68 100644 --- a/net/nfc/llcp_sock.c +++ b/net/nfc/llcp_sock.c @@ -759,12 +759,16 @@ static int llcp_sock_connect(struct socket *sock, struct sockaddr_unsized *_addr llcp_sock->service_name_len = min_t(unsigned int, addr->service_name_len, NFC_LLCP_MAX_SERVICE_NAME); - llcp_sock->service_name = kmemdup(addr->service_name, - llcp_sock->service_name_len, - GFP_KERNEL); - if (!llcp_sock->service_name) { - ret = -ENOMEM; - goto sock_llcp_release; + if (llcp_sock->service_name_len == 0) { + llcp_sock->service_name = NULL; + } else { + llcp_sock->service_name = kmemdup(addr->service_name, + llcp_sock->service_name_len, + GFP_KERNEL); + if (!llcp_sock->service_name) { + ret = -ENOMEM; + goto sock_llcp_release; + } } nfc_llcp_sock_link(&local->connecting_sockets, sk); From 408cff6bd60636df201274d320edfdfde9ed41db Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?=C3=96mer=20Mete=20Kaya?= Date: Wed, 9 Sep 2026 15:14:31 +0300 Subject: [PATCH 120/189] nfc: llcp: fix WKS SAP hijacking via prefix match in nfc_llcp_wks_sap() MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit nfc_llcp_wks_sap() compares only service_name_len bytes, so a short service_name like "u" matches longer WKS strings like "urn:nfc:sn:snep". Fix by requiring exact length match before strncmp(). Fixes: d646960f7986 ("NFC: Initial LLCP support") Signed-off-by: Ömer Mete Kaya Link: https://patch.msgid.link/20260909121437.33744-1-omermetekaya0@gmail.com Signed-off-by: David Heidelberg --- net/nfc/llcp_core.c | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/net/nfc/llcp_core.c b/net/nfc/llcp_core.c index 5ce3ce64baf2..def712e1b2b2 100644 --- a/net/nfc/llcp_core.c +++ b/net/nfc/llcp_core.c @@ -369,7 +369,8 @@ static int nfc_llcp_wks_sap(const char *service_name, size_t service_name_len) if (wks[sap] == NULL) continue; - if (strncmp(wks[sap], service_name, service_name_len) == 0) + if (strlen(wks[sap]) == service_name_len && + !strncmp(wks[sap], service_name, service_name_len)) return sap; } From 7dcf371a35632f035baf77bcf2c129165f772ce4 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?=C3=96mer=20Mete=20Kaya?= Date: Tue, 8 Sep 2026 19:18:01 +0300 Subject: [PATCH 121/189] nfc: llcp: fix slab-out-of-bounds reads when logging service names MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit nfc_llcp_wks_sap() and nfc_llcp_build_sdreq_tlv() pass non-null- terminated strings to pr_debug() using the %s format specifier. The buffers are allocated via kmemdup() or come from netlink attributes and are not guaranteed to be null-terminated, causing __dynamic_pr_debug() to read beyond the allocated region: KASAN: slab-out-of-bounds Read in __dynamic_pr_debug Fix both call sites by using %.*s with the explicit length to limit the output to the actual length of the string. Fixes: d9b8d8e19b07 ("NFC: llcp: Service Name Lookup netlink interface") Reported-by: syzbot+1e3df0852e82c21ca418@syzkaller.appspotmail.com Closes: https://syzkaller.appspot.com/bug?extid=1e3df0852e82c21ca418 Signed-off-by: Ömer Mete Kaya Link: https://patch.msgid.link/20260908161952.731468-1-omermetekaya0@gmail.com Signed-off-by: David Heidelberg --- net/nfc/llcp_commands.c | 2 +- net/nfc/llcp_core.c | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/net/nfc/llcp_commands.c b/net/nfc/llcp_commands.c index ca89fe967d6a..80a00938c869 100644 --- a/net/nfc/llcp_commands.c +++ b/net/nfc/llcp_commands.c @@ -135,7 +135,7 @@ struct nfc_llcp_sdp_tlv *nfc_llcp_build_sdreq_tlv(u8 tid, const char *uri, { struct nfc_llcp_sdp_tlv *sdreq; - pr_debug("uri: %s, len: %zu\n", uri, uri_len); + pr_debug("uri: %.*s, len: %zu\n", (int)uri_len, uri, uri_len); /* sdreq->tlv_len is u8, takes uri_len, + 3 for header, + 1 for NULL */ if (WARN_ON_ONCE(uri_len > U8_MAX - 4)) diff --git a/net/nfc/llcp_core.c b/net/nfc/llcp_core.c index def712e1b2b2..74bf817007cf 100644 --- a/net/nfc/llcp_core.c +++ b/net/nfc/llcp_core.c @@ -358,7 +358,7 @@ static int nfc_llcp_wks_sap(const char *service_name, size_t service_name_len) { int sap, num_wks; - pr_debug("%s\n", service_name); + pr_debug("%.*s\n", (int)service_name_len, service_name); if (service_name == NULL) return -EINVAL; From b61732f47316d45f27706db7812950145d3327b5 Mon Sep 17 00:00:00 2001 From: Deepanshu Kartikey Date: Wed, 23 Sep 2026 09:26:27 +0530 Subject: [PATCH 122/189] nfc: pn533: fix OOB read in pn533_acr122_is_rx_frame_valid() frame->ccid.datalen is read directly from the USB response frame and used, unchecked, as an index into frame->data[]. A malicious or malfunctioning device can set this field to an arbitrary value, causing the driver to read far outside the received buffer. Bound ccid.datalen against the maximum possible ACR122 frame size before using it. This replaces the existing datalen == 0 check, since datalen < 2 already covers that case and additionally rejects datalen == 1, which would still underflow the "datalen - 2" offset used below. Fixes: 9815c7cf22da ("NFC: pn533: Separate physical layer from the core implementation") Reported-by: syzbot+1853daab1a47603d4678@syzkaller.appspotmail.com Closes: https://syzkaller.appspot.com/bug?extid=1853daab1a47603d4678 Tested-by: syzbot+1853daab1a47603d4678@syzkaller.appspotmail.com Assisted-by: LLM Signed-off-by: Deepanshu Kartikey Link: https://patch.msgid.link/20260923035627.6210-1-kartikey406@gmail.com Signed-off-by: David Heidelberg --- drivers/nfc/pn533/usb.c | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/drivers/nfc/pn533/usb.c b/drivers/nfc/pn533/usb.c index efb07f944fce..972eaac09e59 100644 --- a/drivers/nfc/pn533/usb.c +++ b/drivers/nfc/pn533/usb.c @@ -319,7 +319,9 @@ static bool pn533_acr122_is_rx_frame_valid(void *_frame, struct pn533 *dev) if (frame->ccid.type != 0x83) return false; - if (!frame->ccid.datalen) + if (frame->ccid.datalen < 2 || + frame->ccid.datalen > PN533_ACR122_FRAME_MAX_PAYLOAD_LEN + + PN533_ACR122_RX_FRAME_TAIL_LEN) return false; if (frame->data[frame->ccid.datalen - 2] == 0x63) From cfa165cbfbed9d0f4bbc22fef4309f595a3ab187 Mon Sep 17 00:00:00 2001 From: Victor Nogueira Date: Sun, 20 Sep 2026 14:07:01 -0300 Subject: [PATCH 123/189] net/sched: act_gate: budget the per-entry list in get_fill_size tcf_gate_get_fill_size returns only the TCA_GATE_PARMS size, but tcf_gate_dump also emits three 64-bit timestamps, the clock id, flags, priority and the variable-length TCA_GATE_ENTRY_LIST nest. The per-entry nest is unbounded: parse_gate_list places no cap on the number of sched-entries, so a gate with many entries can push the real dump well past the skb that tca_get_fill allocates from this size. RTM_NEWACTION then fails the add-notify with -EINVAL while the action is already committed to the IDR, and a subsequent RTM_GETACTION on the installed gate also returns -EINVAL because its dump no longer fits. Fix this by accounting for the missing fields in tcf_gate_get_fill_size along with all elements in the entries list. Note that sizing the reply from the action lets an oversized gate install cleanly for the first time: with the input unbounded by parse_gate_list, the sized skb can now grow well above NLMSG_GOODSIZE per netlink request (a transient GFP_KERNEL allocation reachable only with namespace-local CAP_NET_ADMIN). Overload from a malicious netns admin is hardening material, not net, per the discussion at https://lore.kernel.org/netdev/20260914191108.55a1a4f1@kernel.org/; a follow-up patch for net-next will cap the sched-entry count. Fixes: 4e76e75d6aba ("net sched actions: calculate add/delete event message size") Reported-by: Sashiko Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260824153903.4143642-1-victor@mojatatu.com Tested-by: hybris Co-developed-by: Jamal Hadi Salim Signed-off-by: Jamal Hadi Salim Signed-off-by: Victor Nogueira Link: https://patch.msgid.link/QDISC-3BLH.v1.20260914203033@mojatatu.com Signed-off-by: Jakub Kicinski --- net/sched/act_gate.c | 30 +++++++++++++++++++++++++++++- 1 file changed, 29 insertions(+), 1 deletion(-) diff --git a/net/sched/act_gate.c b/net/sched/act_gate.c index 5d228a402204..6d6d45e03c07 100644 --- a/net/sched/act_gate.c +++ b/net/sched/act_gate.c @@ -681,7 +681,35 @@ static void tcf_gate_stats_update(struct tc_action *a, u64 bytes, u64 packets, static size_t tcf_gate_get_fill_size(const struct tc_action *act) { - return nla_total_size(sizeof(struct tc_gate)); + struct tcf_gate *gact = to_gate(act); + const struct tcf_gate_params *p; + struct tcfg_gate_entry *entry; + size_t size = nla_total_size(sizeof(struct tc_gate)) /* TCA_GATE_PARMS */ + + 3 * nla_total_size_64bit(sizeof(u64)) /* TCA_GATE_BASE_TIME + * TCA_GATE_CYCLE_TIME + * TCA_GATE_CYCLE_TIME_EXT + */ + + nla_total_size(sizeof(s32)) /* TCA_GATE_CLOCKID */ + + nla_total_size(sizeof(u32)) /* TCA_GATE_FLAGS */ + + nla_total_size(sizeof(s32)) /* TCA_GATE_PRIORITY */ + + nla_total_size(0); /* TCA_GATE_ENTRY_LIST */ + /* TCA_GATE_TM is budgeted by tcf_action_shared_attrs_size() */ + + rcu_read_lock(); + p = rcu_dereference(gact->param); + if (p) { + list_for_each_entry_rcu(entry, &p->entries, list) + /* TCA_GATE_ONE_ENTRY nest and its attributes */ + size += nla_total_size(0) + + nla_total_size(sizeof(u32)) /* TCA_GATE_ENTRY_INDEX */ + + nla_total_size(0) /* TCA_GATE_ENTRY_GATE */ + + nla_total_size(sizeof(u32)) /* TCA_GATE_ENTRY_INTERVAL */ + + nla_total_size(sizeof(s32)) /* TCA_GATE_ENTRY_MAX_OCTETS */ + + nla_total_size(sizeof(s32)); /* TCA_GATE_ENTRY_IPV */ + } + rcu_read_unlock(); + + return size; } static void tcf_gate_entry_destructor(void *priv) From ab1404ac81154a89fb61ac50ae9a04cd8d4834dc Mon Sep 17 00:00:00 2001 From: Fourie Zhang Date: Sun, 20 Sep 2026 19:08:43 +0800 Subject: [PATCH 124/189] net: bridge: mdb: restart port group walk after deletion br_mdb_flush_pgs() keeps a pointer-to-pointer cursor while walking mp->ports. br_multicast_del_pg() can re-enter the same MDB entry through br_multicast_sg_del_exclude_ports() and unlink other port groups. If the cursor points into one of those groups, the next iteration dereferences a stale cursor and can leave mp->ports pointing at freed memory. A following RTM_GETMDB exposes the dangling pointer: BUG: KASAN: slab-use-after-free in br_mdb_dump Read of size 8 br_mdb_dump rtnl_mdb_dump rtnl_dumpit netlink_dump Reset the cursor to mp->ports after every deletion. The deletion removes at least the selected group, so the restarted walk always makes progress. Fixes: a6acb535afb2 ("bridge: mdb: Add MDB bulk deletion support") Cc: stable@vger.kernel.org Signed-off-by: Fourie Zhang Acked-by: Nikolay Aleksandrov Link: https://patch.msgid.link/20260920110852.60293-1-fouriezhang@tencent.com Signed-off-by: Jakub Kicinski --- net/bridge/br_mdb.c | 2 ++ 1 file changed, 2 insertions(+) diff --git a/net/bridge/br_mdb.c b/net/bridge/br_mdb.c index e0c7020b12f5..a01bd280c722 100644 --- a/net/bridge/br_mdb.c +++ b/net/bridge/br_mdb.c @@ -1523,6 +1523,8 @@ static void br_mdb_flush_pgs(struct net_bridge *br, } br_multicast_del_pg(mp, p, pp); + /* br_multicast_del_pg() can remove other groups from this list. */ + pp = &mp->ports; } } From fdfec06ac1eb5cdbc18d556c70a7b26837208218 Mon Sep 17 00:00:00 2001 From: Nicolai Buchwitz Date: Tue, 22 Sep 2026 09:31:40 +0200 Subject: [PATCH 125/189] MAINTAINERS: add Nicolai Buchwitz as GENET maintainer I have been contributing to and reviewing the GENET driver for a while now. Florian asked if I would like to formalize this commitment, so add myself as a maintainer. Signed-off-by: Nicolai Buchwitz Acked-by: Florian Fainelli Acked-by: Justin Chen Link: https://patch.msgid.link/20260922073140.1471858-1-nb@tipi-net.de Signed-off-by: Jakub Kicinski --- MAINTAINERS | 1 + 1 file changed, 1 insertion(+) diff --git a/MAINTAINERS b/MAINTAINERS index 3b2eb2a7a89a..7200905501b8 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -5454,6 +5454,7 @@ F: include/linux/brcmphy.h BROADCOM GENET ETHERNET DRIVER M: Doug Berger M: Florian Fainelli +M: Nicolai Buchwitz R: Broadcom internal kernel review list L: netdev@vger.kernel.org S: Maintained From 7e87508b5c4d81210d0a736ed01962e52f5c4c56 Mon Sep 17 00:00:00 2001 From: Nicolai Buchwitz Date: Tue, 22 Sep 2026 15:06:38 +0200 Subject: [PATCH 126/189] net: bcmgenet: stop Tx NAPI before disabling the queues bcmgenet_netif_stop() and the Wake-on-LAN branch of bcmgenet_suspend() both disable the Tx queues first and stop Tx NAPI several steps later. A completion in flight calls netif_tx_wake_queue() in between, and nothing stops the queue again, so a transmit can reach the rings after they have been freed. Close is safe because dev_deactivate_many() stops the qdisc first. bcmgenet_suspend() does not, so stop Tx NAPI before the queues on both paths. KASAN on a Raspberry Pi CM4, driven from an MTU change because suspend freezes user space before the callback runs: BUG: KASAN: use-after-free in bcmgenet_xmit+0x17f8/0x2258 Write of size 8 at addr ffffff8055844a68 by task ksoftirqd/0/14 bcmgenet_xmit+0x17f8/0x2258 dev_hard_start_xmit+0x13c/0x588 sch_direct_xmit+0x108/0x340 __dev_queue_xmit+0x1190/0x3848 Fixes: 254f3239dd07 ("net: bcmgenet: revise suspend/resume") Signed-off-by: Nicolai Buchwitz Link: https://patch.msgid.link/20260922130639.1660797-1-nb@tipi-net.de Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/broadcom/genet/bcmgenet.c | 6 ++++-- 1 file changed, 4 insertions(+), 2 deletions(-) diff --git a/drivers/net/ethernet/broadcom/genet/bcmgenet.c b/drivers/net/ethernet/broadcom/genet/bcmgenet.c index b916080f4ff1..ca62041efecd 100644 --- a/drivers/net/ethernet/broadcom/genet/bcmgenet.c +++ b/drivers/net/ethernet/broadcom/genet/bcmgenet.c @@ -3441,6 +3441,8 @@ static void bcmgenet_netif_stop(struct net_device *dev, bool stop_phy) { struct bcmgenet_priv *priv = netdev_priv(dev); + /* Stop completion polling before it can wake a stopped queue */ + bcmgenet_disable_tx_napi(priv); netif_tx_disable(dev); /* Disable MAC receive */ @@ -3455,7 +3457,6 @@ static void bcmgenet_netif_stop(struct net_device *dev, bool stop_phy) /* Disable MAC transmit. TX DMA disabled must be done before this */ umac_enable_set(priv, CMD_TX_EN, false); - bcmgenet_disable_tx_napi(priv); bcmgenet_disable_rx_napi(priv); bcmgenet_intr_disable(priv); @@ -4320,6 +4321,8 @@ static int bcmgenet_suspend(struct device *d) netif_device_detach(dev); if (device_may_wakeup(d) && priv->wolopts) { + /* Stop completion polling before it can wake a stopped queue */ + bcmgenet_disable_tx_napi(priv); netif_tx_disable(dev); /* Suspend non-wake Rx data flows */ @@ -4348,7 +4351,6 @@ static int bcmgenet_suspend(struct device *d) netdev_warn(priv->dev, "Timed out while disabling TX DMA\n"); - bcmgenet_disable_tx_napi(priv); bcmgenet_disable_rx_napi(priv); disable_irq(priv->irq1); bcmgenet_tx_reclaim_all(dev); From 7104a370714346b667712913dc16abf14bbc97ed Mon Sep 17 00:00:00 2001 From: Jakub Kicinski Date: Mon, 21 Sep 2026 16:18:56 -0700 Subject: [PATCH 127/189] veth: manage XDP program pointers during channel resize veth_set_channels() tears down XDP resources for removed RX queues without clearing rq->xdp_prog. If the program is then detached or replaced, those queues keep the old pointer after bpf_prog_put(). A later channel increase can re-enable NAPI and run the freed program. BUG: unable to handle page fault for address: ffffc90000256048 Oops: Oops: 0000 [#1] SMP KASAN NOPTI RIP: veth_xdp_rcv_skb (include/linux/filter.h:779 include/net/xdp.h:696 drivers/net/veth.c:820) Call Trace: veth_xdp_rcv (drivers/net/veth.c:941) veth_poll (drivers/net/veth.c:986) __napi_poll (net/core/dev.c:7787) net_rx_action (net/core/dev.c:7850 net/core/dev.c:8007) handle_softirqs (kernel/softirq.c:645) Kernel panic - not syncing: Fatal exception in interrupt Fixes: 4752eeb3d891 ("veth: implement support for set_channel ethtool op") Signed-off-by: Weiming Shi Acked-by: Stanislav Fomichev Reviewed-by: Jiayuan Chen Reviewed-by: Jason Xing Link: https://patch.msgid.link/20260921231856.1798630-1-kuba@kernel.org Signed-off-by: Jakub Kicinski --- drivers/net/veth.c | 2 ++ 1 file changed, 2 insertions(+) diff --git a/drivers/net/veth.c b/drivers/net/veth.c index 6ed3ee81153f..71227d0389aa 100644 --- a/drivers/net/veth.c +++ b/drivers/net/veth.c @@ -1054,6 +1054,7 @@ static int __veth_napi_enable_range(struct net_device *dev, int start, int end) for (i = start; i < end; i++) { struct veth_rq *rq = &priv->rq[i]; + rcu_assign_pointer(rq->xdp_prog, priv->_xdp_prog); napi_enable(&rq->xdp_napi); rcu_assign_pointer(priv->rq[i].napi, &priv->rq[i].xdp_napi); } @@ -1088,6 +1089,7 @@ static void veth_napi_del_range(struct net_device *dev, int start, int end) rcu_assign_pointer(priv->rq[i].napi, NULL); napi_disable(&rq->xdp_napi); + rcu_assign_pointer(rq->xdp_prog, NULL); __netif_napi_del(&rq->xdp_napi); } synchronize_net(); From 3b4e0b0c008a8c1b474730248cd5b873026c74bd Mon Sep 17 00:00:00 2001 From: Shihuang Liu Date: Sat, 19 Sep 2026 21:36:04 +0800 Subject: [PATCH 128/189] net: skbuff: fix pull-bound underflow in skb_checksum_setup_ipv6() skb_maybe_pull_tail() subtracts skb_headlen(skb) from the unsigned max argument and passes the result to __pskb_pull_tail() as a signed int. The function does not ensure that max is at least skb_headlen(skb). This can happen while parsing IPv6 extension headers when an skb already has a linear area larger than MAX_IPV6_HDR_LEN. Once the parser needs data beyond the linear area, max - skb_headlen(skb) wraps and is converted to a negative delta. __pskb_pull_tail() then passes that negative length to skb_copy_bits(), where it can become a very large copy length. Pass the requested length itself as the pull bound at the three extension-header call sites, so the delta can no longer go negative. Fixes: 1431fb31ecba ("xen-netback: fix fragment detection in checksum setup") Suggested-by: Eric Dumazet Signed-off-by: Shihuang Liu Link: https://patch.msgid.link/20260919133604.50948-1-shlomojune6@gmail.com Signed-off-by: Jakub Kicinski --- net/core/skbuff.c | 9 ++++++--- 1 file changed, 6 insertions(+), 3 deletions(-) diff --git a/net/core/skbuff.c b/net/core/skbuff.c index 609f2c7f4a47..b4edbd06655e 100644 --- a/net/core/skbuff.c +++ b/net/core/skbuff.c @@ -5977,7 +5977,8 @@ static int skb_checksum_setup_ipv6(struct sk_buff *skb, bool recalculate) err = skb_maybe_pull_tail(skb, off + sizeof(struct ipv6_opt_hdr), - MAX_IPV6_HDR_LEN); + off + + sizeof(struct ipv6_opt_hdr)); if (err < 0) goto out; @@ -5992,7 +5993,8 @@ static int skb_checksum_setup_ipv6(struct sk_buff *skb, bool recalculate) err = skb_maybe_pull_tail(skb, off + sizeof(struct ip_auth_hdr), - MAX_IPV6_HDR_LEN); + off + + sizeof(struct ip_auth_hdr)); if (err < 0) goto out; @@ -6007,7 +6009,8 @@ static int skb_checksum_setup_ipv6(struct sk_buff *skb, bool recalculate) err = skb_maybe_pull_tail(skb, off + sizeof(struct frag_hdr), - MAX_IPV6_HDR_LEN); + off + + sizeof(struct frag_hdr)); if (err < 0) goto out; From ab7aa05c06ae340e5c7530bb78fa8d23794e460b Mon Sep 17 00:00:00 2001 From: Ido Schimmel Date: Tue, 22 Sep 2026 16:12:39 +0300 Subject: [PATCH 129/189] vrf: Stop corrupting skb->csum when capturing CHECKSUM_COMPLETE packets The VRF device is an Ethernet device but it can have non-Ethernet ports such as IP tunnels. Before the cited commit, capturing packets from such ports on the VRF device resulted in these packets being detected as malformed since they lack an Ethernet header. The cited commit fixed it by pushing a dummy Ethernet header to such packets before the capture and pulling it afterwards. In the case of CHECKSUM_COMPLETE packets it also updated skb->csum with the checksum of the dummy Ethernet header. This is wrong as skb->csum should not include the checksum of the Ethernet header ("checksum of the _whole_ packet as seen by netif_rx()"). This also means that L4 protocols receive a corrupted skb->csum and potentially drop the packet, as is the case with UDP packets whose checksum was completed by software. Fix by removing the unnecessary call to skb_postpush_rcsum(). Fixes: 048939088220 ("vrf: add mac header for tunneled packets when sniffer is attached") Reported-by: Stefano Sasso Closes: https://lore.kernel.org/netdev/CALtE316UtL3x7LL6uxfXzx8rW6AbzYPeDOb478hqJCr_-dj=Wg@mail.gmail.com/ Signed-off-by: Ido Schimmel Reviewed-by: David Ahern Reviewed-by: Eric Dumazet Reviewed-by: Andrea Mayer Link: https://patch.msgid.link/20260922131239.2509494-1-idosch@nvidia.com Signed-off-by: Jakub Kicinski --- drivers/net/vrf.c | 2 -- 1 file changed, 2 deletions(-) diff --git a/drivers/net/vrf.c b/drivers/net/vrf.c index a0557a3a7026..d4dc6d690a75 100644 --- a/drivers/net/vrf.c +++ b/drivers/net/vrf.c @@ -1175,8 +1175,6 @@ static int vrf_prepare_mac_header(struct sk_buff *skb, skb->protocol = eth->h_proto; skb->pkt_type = PACKET_HOST; - skb_postpush_rcsum(skb, skb->data, ETH_HLEN); - skb_pull_inline(skb, ETH_HLEN); return 0; From 2d959c75c27f90e9ec489d18ce5ee6b852ad4741 Mon Sep 17 00:00:00 2001 From: Hui Peng Date: Mon, 21 Sep 2026 04:40:25 +0000 Subject: [PATCH 130/189] ipv6: sr: enforce exact attribute length for SEG6_ATTR_DST In seg6_genl_policy, SEG6_ATTR_DST is defined with .type = NLA_BINARY and .len = sizeof(struct in6_addr). For NLA_BINARY, .len only enforces the maximum payload length and permits shorter payloads (e.g., 0 bytes). When seg6_genl_set_tunsrc() copies sizeof(struct in6_addr) bytes via kmemdup(val, sizeof(*val), GFP_KERNEL), a short SEG6_ATTR_DST attribute triggers a 16-byte out-of-bounds read past skb->tail into uninitialized skb->head memory, which is stored in sdata->tun_src and leaked back to userspace via SEG6_CMD_GET_TUNSRC. Switch SEG6_ATTR_DST in seg6_genl_policy to NLA_POLICY_EXACT_LEN(sizeof(struct in6_addr)) so that generic netlink validation rejects any attribute whose length is not exactly sizeof(struct in6_addr) with -ERANGE. Tested in QEMU against Linux 7.3.0-rc3 by sending a SEG6_CMD_SET_TUNSRC Generic Netlink message with a 0-byte SEG6_ATTR_DST attribute followed by SEG6_CMD_GET_TUNSRC. On the unfixed kernel, SEG6_CMD_SET_TUNSRC succeeds (err = 0) and SEG6_CMD_GET_TUNSRC leaks 16 bytes of uninitialized kernel heap memory (tun_src = 836a61ecc4d25a1042a8d60411cfb378); with this patch applied, SEG6_CMD_SET_TUNSRC is rejected by netlink policy validation with -ERANGE (-34) and tun_src remains zeroed. Fixes: 915d7e5e5930 ("ipv6: sr: add code base for control plane support of SR-IPv6") Cc: stable@vger.kernel.org Signed-off-by: Hui Peng Reviewed-by: Hangbin Liu Reviewed-by: Justin Iurman Reviewed-by: Andrea Mayer Link: https://patch.msgid.link/20260921044025.1535982-1-benquike@gmail.com Signed-off-by: Jakub Kicinski --- net/ipv6/seg6.c | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/net/ipv6/seg6.c b/net/ipv6/seg6.c index 62a7eb779202..8c2b156c227a 100644 --- a/net/ipv6/seg6.c +++ b/net/ipv6/seg6.c @@ -138,8 +138,8 @@ void seg6_icmp_srh(struct sk_buff *skb, struct inet6_skb_parm *opt) static struct genl_family seg6_genl_family; static const struct nla_policy seg6_genl_policy[SEG6_ATTR_MAX + 1] = { - [SEG6_ATTR_DST] = { .type = NLA_BINARY, - .len = sizeof(struct in6_addr) }, + [SEG6_ATTR_DST] = + NLA_POLICY_EXACT_LEN(sizeof(struct in6_addr)), [SEG6_ATTR_DSTLEN] = { .type = NLA_S32, }, [SEG6_ATTR_HMACKEYID] = { .type = NLA_U32, }, [SEG6_ATTR_SECRET] = { .type = NLA_BINARY, }, From 6b491af01aa5c0633a580e3b11bc7277adc903b2 Mon Sep 17 00:00:00 2001 From: Nicolo Giuliani Date: Mon, 21 Sep 2026 05:56:29 +0200 Subject: [PATCH 131/189] net: dsa: mv88e6xxx: 88E6191X and 88E6193X have no PTP The 88E6191X and 88E6193X are 6393 family devices that share mv88e6393x_ops with the 88E6393X and are marked as ptp_support. Marvell's UMSD driver describes both as parts without AVB (88E6193X: "BGA package - No AVB, No Routing, No Cut-through"), and the register access confirms it on an 88E6193X: the whole indirect AVB register space behind Global 2 registers 0x16 and 0x17 reads zero, for every port, block and address, with the 6390 and with the 6352 command encoding. Writes to the TAI registers, including the clock period register and the TAI global configuration register, read back as zero. Since commit 7e3c18097a70 ("net: dsa: mv88e6xxx: read cycle counter period from hardware") the PTP setup reads the TAI clock period, so the switch fails to probe: mv88e6xxx ...: unexpected cycle counter period of 0 ps Add mv88e6191x_ops, a copy of mv88e6393x_ops without avb_ops and ptp_ops, use it for the 88E6191X and the 88E6193X and stop setting ptp_support for them. The 88E6393X is unchanged. Tested on an 88E6193X (Sophos XGS 107w): the switch probes and the ports work. I do not have an 88E6191X, it is changed because UMSD describes it the same way. Fixes: de776d0d316f ("net: dsa: mv88e6xxx: add support for mv88e6393x family") Suggested-by: Andrew Lunn Signed-off-by: Nicolo Giuliani Reviewed-by: Andrew Lunn Link: https://patch.msgid.link/20260921-send-net-v2-1-031ad720f140@studio.unibo.it Signed-off-by: Jakub Kicinski --- drivers/net/dsa/mv88e6xxx/chip.c | 68 ++++++++++++++++++++++++++++++-- 1 file changed, 64 insertions(+), 4 deletions(-) diff --git a/drivers/net/dsa/mv88e6xxx/chip.c b/drivers/net/dsa/mv88e6xxx/chip.c index 7f68a0c55802..a4a8c7e11bf4 100644 --- a/drivers/net/dsa/mv88e6xxx/chip.c +++ b/drivers/net/dsa/mv88e6xxx/chip.c @@ -5639,6 +5639,68 @@ static const struct mv88e6xxx_ops mv88e6390x_ops = { .pcs_ops = &mv88e6390_pcs_ops, }; +static const struct mv88e6xxx_ops mv88e6191x_ops = { + /* MV88E6XXX_FAMILY_6393 without AVB and PTP: 6191X and 6193X */ + .irl_init_all = mv88e6390_g2_irl_init_all, + .get_eeprom = mv88e6xxx_g2_get_eeprom8, + .set_eeprom = mv88e6xxx_g2_set_eeprom8, + .set_switch_mac = mv88e6xxx_g2_set_switch_mac, + .phy_read = mv88e6xxx_g2_smi_phy_read_c22, + .phy_write = mv88e6xxx_g2_smi_phy_write_c22, + .phy_read_c45 = mv88e6xxx_g2_smi_phy_read_c45, + .phy_write_c45 = mv88e6xxx_g2_smi_phy_write_c45, + .port_set_link = mv88e6xxx_port_set_link, + .port_sync_link = mv88e6xxx_port_sync_link, + .port_set_rgmii_delay = mv88e6390_port_set_rgmii_delay, + .port_set_speed_duplex = mv88e6393x_port_set_speed_duplex, + .port_tag_remap = mv88e6390_port_tag_remap, + .port_set_policy = mv88e6393x_port_set_policy, + .port_set_frame_mode = mv88e6351_port_set_frame_mode, + .port_set_ucast_flood = mv88e6352_port_set_ucast_flood, + .port_set_mcast_flood = mv88e6352_port_set_mcast_flood, + .port_set_ether_type = mv88e6393x_port_set_ether_type, + .port_set_jumbo_size = mv88e6165_port_set_jumbo_size, + .port_egress_rate_limiting = mv88e6097_port_egress_rate_limiting, + .port_pause_limit = mv88e6390_port_pause_limit, + .port_disable_learn_limit = mv88e6xxx_port_disable_learn_limit, + .port_disable_pri_override = mv88e6xxx_port_disable_pri_override, + .port_get_cmode = mv88e6352_port_get_cmode, + .port_set_cmode = mv88e6393x_port_set_cmode, + .port_setup_message_port = mv88e6xxx_setup_message_port, + .port_set_upstream_port = mv88e6393x_port_set_upstream_port, + .port_enable_tcam = mv88e6xxx_port_enable_tcam, + .stats_snapshot = mv88e6390_g1_stats_snapshot, + .stats_set_histogram = mv88e6390_g1_stats_set_histogram, + .stats_get_sset_count = mv88e6320_stats_get_sset_count, + .stats_get_strings = mv88e6320_stats_get_strings, + .stats_get_stat = mv88e6390_stats_get_stat, + /* .set_cpu_port is missing because this family does not support a global + * CPU port, only per port CPU port which is set via + * .port_set_upstream_port method. + */ + .set_egress_port = mv88e6393x_set_egress_port, + .watchdog_ops = &mv88e6393x_watchdog_ops, + .mgmt_rsvd2cpu = mv88e6393x_port_mgmt_rsvd2cpu, + .pot_clear = mv88e6xxx_g2_pot_clear, + .hardware_reset_pre = mv88e6xxx_g2_eeprom_wait, + .hardware_reset_post = mv88e6xxx_g2_eeprom_wait, + .reset = mv88e6352_g1_reset, + .rmu_disable = mv88e6390_g1_rmu_disable, + .atu_get_hash = mv88e6165_g1_atu_get_hash, + .atu_set_hash = mv88e6165_g1_atu_set_hash, + .vtu_getnext = mv88e6390_g1_vtu_getnext, + .vtu_loadpurge = mv88e6390_g1_vtu_loadpurge, + .stu_getnext = mv88e6390_g1_stu_getnext, + .stu_loadpurge = mv88e6390_g1_stu_loadpurge, + .serdes_get_lane = mv88e6393x_serdes_get_lane, + .serdes_irq_mapping = mv88e6390_serdes_irq_mapping, + /* TODO: serdes stats */ + .gpio_ops = &mv88e6352_gpio_ops, + .phylink_get_caps = mv88e6393x_phylink_get_caps, + .pcs_ops = &mv88e6393x_pcs_ops, + .tcam_ops = &mv88e6393_tcam_ops, +}; + static const struct mv88e6xxx_ops mv88e6393x_ops = { /* MV88E6XXX_FAMILY_6393 */ .irl_init_all = mv88e6390_g2_irl_init_all, @@ -6163,8 +6225,7 @@ static const struct mv88e6xxx_info mv88e6xxx_table[] = { .atu_move_port_mask = 0x1f, .pvt = true, .multi_chip = true, - .ptp_support = true, - .ops = &mv88e6393x_ops, + .ops = &mv88e6191x_ops, }, [MV88E6193X] = { @@ -6190,8 +6251,7 @@ static const struct mv88e6xxx_info mv88e6xxx_table[] = { .atu_move_port_mask = 0x1f, .pvt = true, .multi_chip = true, - .ptp_support = true, - .ops = &mv88e6393x_ops, + .ops = &mv88e6191x_ops, }, [MV88E6220] = { From d22609f3d13fc5baacd92c222731b03c593401db Mon Sep 17 00:00:00 2001 From: Hui Peng Date: Mon, 21 Sep 2026 04:59:20 +0000 Subject: [PATCH 132/189] fou: reject omitted FOU_ATTR_IPPROTO on FOU_ENCAP_DIRECT Commit 7a9bc9e3f423 ("fou: Don't allow 0 for FOU_ATTR_IPPROTO.") added NLA_POLICY_MIN(NLA_U8, 1) to fou_nl_policy[FOU_ATTR_IPPROTO], which rejects an explicitly supplied FOU_ATTR_IPPROTO == 0 attribute with -ERANGE. However, FOU_ATTR_IPPROTO is an optional netlink attribute. When a user sends FOU_CMD_ADD with FOU_ATTR_TYPE set to FOU_ENCAP_DIRECT and omits FOU_ATTR_IPPROTO entirely, nla_policy validation succeeds and parse_nl_config() leaves cfg->protocol as 0 (from memset(cfg, 0, sizeof(*cfg))). fou_create() then creates a FOU_ENCAP_DIRECT socket with fou->protocol == 0. In fou_udp_recv(), returning -fou->protocol to udp_queue_rcv_one_skb() triggers IP protocol resubmission when fou->protocol > 0, whereas returning 0 tells the UDP tunnel layer that the skb was consumed without freeing it. When fou->protocol == 0, every packet received on the socket returns 0 from fou_udp_recv() and leaks the sk_buff. Reject FOU_ENCAP_DIRECT when !cfg->protocol in fou_create() so that creating a direct encapsulation port without FOU_ATTR_IPPROTO fails with -EINVAL while leaving FOU_CMD_DEL and FOU_CMD_GET (which share parse_nl_config()) unaffected. Tested in QEMU against Linux 7.3.0-rc3 by sending a FOU_CMD_ADD Generic Netlink request with FOU_ATTR_PORT = 5555 and FOU_ATTR_TYPE = FOU_ENCAP_DIRECT while omitting FOU_ATTR_IPPROTO. On the unfixed kernel, FOU_CMD_ADD succeeds (err = 0), FOU_CMD_GET reports fou->type = 1 and fou->protocol = 0, and sending 4000 UDP packets to 127.0.0.1:5555 leaks all 4000 sk_buffs (SUnreclaim in /proc/meminfo grows from 41456 kB to 59008 kB, +17552 kB); with this patch applied, FOU_CMD_ADD is rejected with -EINVAL (-22). Fixes: 23461551c006 ("fou: Support for foo-over-udp RX path") Fixes: 7a9bc9e3f423 ("fou: Don't allow 0 for FOU_ATTR_IPPROTO.") Cc: stable@vger.kernel.org Signed-off-by: Hui Peng Reviewed-by: Hangbin Liu Link: https://patch.msgid.link/20260921045920.1613098-1-benquike@gmail.com Signed-off-by: Jakub Kicinski --- net/ipv4/fou_core.c | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/net/ipv4/fou_core.c b/net/ipv4/fou_core.c index 5e867f1b5c1d..076fdca44f51 100644 --- a/net/ipv4/fou_core.c +++ b/net/ipv4/fou_core.c @@ -600,6 +600,10 @@ static int fou_create(struct net *net, struct fou_cfg *cfg, /* Initial for fou type */ switch (cfg->type) { case FOU_ENCAP_DIRECT: + if (!cfg->protocol) { + err = -EINVAL; + goto error; + } tunnel_cfg.encap_rcv = fou_udp_recv; tunnel_cfg.gro_receive = fou_gro_receive; tunnel_cfg.gro_complete = fou_gro_complete; From 89a8a1eef2d441b7825a6c0116ce817235a7ffe6 Mon Sep 17 00:00:00 2001 From: Jonas Jelonek Date: Fri, 18 Sep 2026 21:19:55 +0000 Subject: [PATCH 133/189] net: mdio: realtek-rtl9300: fix RTL931x C22 extended page selection The RTL931x indirect access engine has a separate nine-bit extended page field. The driver leaves it at zero, and otto_emdio_run_cmd() therefore programs extended page zero for every Clause 22 transaction. This overrides page selection made through PHY register 30, causing accesses to private PHY pages to hit extended page zero instead. Set the field to its 0x1ff "do not change" value for RTL931x Clause 22 reads and writes. This preserves extended page selection made through PHY register 30 and restores access to its private register pages. Fixes: 5ebdcac59aff ("net: mdio: realtek-rtl9300: Add support for RTL931x") Signed-off-by: Jonas Jelonek Acked-by: Markus Stockhausen Link: https://patch.msgid.link/20260918211955.3955777-1-jonas@jonasjelonek.de Signed-off-by: Jakub Kicinski --- drivers/net/mdio/mdio-realtek-rtl9300.c | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/drivers/net/mdio/mdio-realtek-rtl9300.c b/drivers/net/mdio/mdio-realtek-rtl9300.c index afd52a1cd7f8..9ce2b7807532 100644 --- a/drivers/net/mdio/mdio-realtek-rtl9300.c +++ b/drivers/net/mdio/mdio-realtek-rtl9300.c @@ -88,6 +88,8 @@ #define RTL9310_SMI_INDRT_ACCESS_BC_PHYID_CTRL 0x0c14 #define RTL9310_BC_PORT_ID GENMASK(10, 5) #define RTL9310_SMI_INDRT_ACCESS_CTRL_1 0x0c04 +#define RTL9310_SMI_INDRT_EXT_PAGE GENMASK(8, 0) +#define RTL9310_SMI_INDRT_EXT_PAGE_NO_CHANGE 0x1ff #define RTL9310_SMI_INDRT_ACCESS_CTRL_2_LOW 0x0c08 #define RTL9310_SMI_INDRT_ACCESS_CTRL_2_HIGH 0x0c0c #define RTL9310_SMI_INDRT_ACCESS_CTRL_3 0x0c10 /* I/O fields flipped */ @@ -325,6 +327,8 @@ static int otto_emdio_9310_read_c22(struct mii_bus *bus, int port, int regnum, u .broadcast = FIELD_PREP(RTL9310_BC_PORT_ID, port), .c22_data = FIELD_PREP(RTL9310_PHY_CTRL_REG_ADDR, regnum) | FIELD_PREP(RTL9310_PHY_CTRL_MAIN_PAGE, RAW_PAGE(priv)), + .ext_page = FIELD_PREP(RTL9310_SMI_INDRT_EXT_PAGE, + RTL9310_SMI_INDRT_EXT_PAGE_NO_CHANGE), }; return otto_emdio_read_cmd(bus, RTL9310_PHY_CTRL_TYPE_C22, &cmd_data, @@ -337,6 +341,8 @@ static int otto_emdio_9310_write_c22(struct mii_bus *bus, int port, int regnum, struct otto_emdio_cmd_regs cmd_data = { .c22_data = FIELD_PREP(RTL9310_PHY_CTRL_REG_ADDR, regnum) | FIELD_PREP(RTL9310_PHY_CTRL_MAIN_PAGE, RAW_PAGE(priv)), + .ext_page = FIELD_PREP(RTL9310_SMI_INDRT_EXT_PAGE, + RTL9310_SMI_INDRT_EXT_PAGE_NO_CHANGE), .io_data = FIELD_PREP(RTL9310_PHY_CTRL_INDATA, value), .port_mask_high = (u32)(BIT_ULL(port) >> 32), .port_mask_low = (u32)(BIT_ULL(port)), From d8b6529e80bcb4fb8177121404cbb3377acaebd2 Mon Sep 17 00:00:00 2001 From: Zijie Huang Date: Mon, 21 Sep 2026 01:36:19 +0800 Subject: [PATCH 134/189] net: arp: terminate device name before lookup The ARP ioctl copies a user-provided struct arpreq into a stack object. Its arp_dev field may contain IFNAMSIZ bytes without a NUL terminator. Such input is passed to dev_get_by_name_rcu() or __dev_get_by_name(), where strcmp() can read past the end of the stack object when a matching alternative interface name exists. Terminate the field before the lookup to prevent the out-of-bounds read. Fixes: 36fbf1e52bd3 ("net: rtnetlink: add linkprop commands to add and delete alternative ifnames") Cc: stable@vger.kernel.org Reported-by: Vega Signed-off-by: Zijie Huang Signed-off-by: Ren Wei Reviewed-by: Ido Schimmel Link: https://patch.msgid.link/fabf02a70787d17299e4b3153eadffaf20d154b3.1789910973.git.milkory@outlook.com Signed-off-by: Jakub Kicinski --- net/ipv4/arp.c | 1 + 1 file changed, 1 insertion(+) diff --git a/net/ipv4/arp.c b/net/ipv4/arp.c index d409f606aec0..60009d92e071 100644 --- a/net/ipv4/arp.c +++ b/net/ipv4/arp.c @@ -1278,6 +1278,7 @@ int arp_ioctl(struct net *net, unsigned int cmd, void __user *arg) err = copy_from_user(&r, arg, sizeof(struct arpreq)); if (err) return -EFAULT; + r.arp_dev[IFNAMSIZ - 1] = '\0'; break; default: return -EINVAL; From 4eb3f195ef08c5acaed87958297e41cc49588dde Mon Sep 17 00:00:00 2001 From: Ivan Delalande Date: Fri, 18 Sep 2026 15:47:15 -0700 Subject: [PATCH 135/189] tg3: use random MAC address when tg3_get_device_address fails Some of the tg3 NICs we use (BCM57762) reset the SRAM MAC address to the placeholder address on link flaps, tg3_chip_reset, etc. We've typically fixed it from userspace, but since e4c00ba7274b ("tg3: replace placeholder MAC address with device property") was merged, tg3 just fails probe as we don't have a way to get it through the generic device_get_mac_address infrastructure as fallback on our systems. Make the driver assign a random address in this condition instead of being fatal for probe. Set deferred_probe_reason through dev_warn_probe if the address isn't yet available from the provider. Fixes: e4c00ba7274b ("tg3: replace placeholder MAC address with device property") Suggested-by: Jakub Kicinski Link: https://lore.kernel.org/netdev/20260909191751.651aa5c4@kernel.org/ Signed-off-by: Ivan Delalande Link: https://patch.msgid.link/20260918224715.GA654128@visor Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/broadcom/tg3.c | 11 +++++++---- 1 file changed, 7 insertions(+), 4 deletions(-) diff --git a/drivers/net/ethernet/broadcom/tg3.c b/drivers/net/ethernet/broadcom/tg3.c index caa7a6caa6e2..8b6806a79edf 100644 --- a/drivers/net/ethernet/broadcom/tg3.c +++ b/drivers/net/ethernet/broadcom/tg3.c @@ -17915,11 +17915,14 @@ static int tg3_init_one(struct pci_dev *pdev, err = tg3_get_device_address(tp, addr); if (err) { - dev_err(&pdev->dev, - "Could not obtain valid ethernet address, aborting\n"); - goto err_out_apeunmap; + dev_warn_probe(&pdev->dev, err, + "Could not obtain a valid ethernet address\n"); + if (err == -EPROBE_DEFER) + goto err_out_apeunmap; + eth_hw_addr_random(dev); + } else { + eth_hw_addr_set(dev, addr); } - eth_hw_addr_set(dev, addr); intmbx = MAILBOX_INTERRUPT_0 + TG3_64BIT_REG_LOW; rcvmbx = MAILBOX_RCVRET_CON_IDX_0 + TG3_64BIT_REG_LOW; From 0a7822e34a0bfde31b194ac3da3253e5032b44cc Mon Sep 17 00:00:00 2001 From: Lorenzo Bianconi Date: Mon, 21 Sep 2026 16:46:18 +0200 Subject: [PATCH 136/189] net: stmmac: clear stale buf->page after recycling on skb build failure In stmmac_rx(), when napi_build_skb() fails the descriptor page is recycled back to the page pool with page_pool_recycle_direct(), but buf->page is left pointing at the recycled page, unlike every other consumption site in the function which clears the pointer after handing the page away. With the stale pointer stmmac_rx_refill() skips the replacement allocation and programs the already-recycled page back into the RX descriptor. Clear buf->page on the napi_build_skb() failure path to keep the buffer lifecycle consistent with the other consumption sites. Fixes: df542f669307 ("net: stmmac: Switch to zero-copy in non-XDP RX path") Signed-off-by: Lorenzo Bianconi Reviewed-by: Maxime Chevallier Link: https://patch.msgid.link/20260921-stmmac-fix-napi-build-skb-error-v1-1-3d54bf6d9bb6@oss.qualcomm.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/stmicro/stmmac/stmmac_main.c | 1 + 1 file changed, 1 insertion(+) diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c index d5a984ad864f..3f34d491c959 100644 --- a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c +++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c @@ -5889,6 +5889,7 @@ static int stmmac_rx(struct stmmac_priv *priv, int limit, u32 queue) if (!skb) { page_pool_recycle_direct(rx_q->page_pool, buf->page); + buf->page = NULL; rx_dropped++; count++; goto drain_data; From 87cd6b717e4069dca34e0866e57eaeb44d3b173e Mon Sep 17 00:00:00 2001 From: Yuya Kusakabe Date: Tue, 22 Sep 2026 05:49:56 +0900 Subject: [PATCH 137/189] net: ipv6: keep room for the mac header in dst_dev_overhead() The seg6, ioam6 and rpl lwtunnels size their skb_cow_head() request as the length they are about to push plus dst_dev_overhead(), then push the new headers and rebuild the mac header below them with skb_mac_header_rebuild(). That rebuild needs skb->mac_len of headroom, but dst_dev_overhead() leaves LL_RESERVED_SPACE() of the egress device, 16 bytes for plain Ethernet. Where the mac header is longer than that, as it is on ingress through a VLAN device with reorder_hdr off, the rebuild runs out of room: skb_set_mac_header(skb, -skb->mac_len) computes a negative offset, stores it unchecked in the u16 skb->mac_header, and the memmove that follows writes skb->mac_len bytes about 64 KB past skb->head. Forwarding plain ping6 traffic through such a device reproduces it on all five seg6 encapsulation modes and on the rpl and ioam6 inline paths; skb->mac_header comes back as 65534 on a 704-byte head. Return the larger of the two. The helper already returns skb->mac_len when it has no dst, so this only makes the other branch agree, and it covers every caller rather than each call site in turn. Fixes: 40475b63761a ("net: ipv6: seg6_iptunnel: mitigate 2-realloc issue") Fixes: dce525185bc9 ("net: ipv6: ioam6_iptunnel: mitigate 2-realloc issue") Fixes: 985ec6f5e623 ("net: ipv6: rpl_iptunnel: mitigate 2-realloc issue") Suggested-by: Andrea Mayer Signed-off-by: Yuya Kusakabe Reviewed-by: Justin Iurman Reviewed-by: Gabriel Goller Reviewed-by: Eric Dumazet Reviewed-by: Andrea Mayer Link: https://patch.msgid.link/20260922-seg6-maclen-headroom-v3-1-7b2f982ef79d@gmail.com Signed-off-by: Jakub Kicinski --- include/net/dst.h | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/include/net/dst.h b/include/net/dst.h index 307073eae7f8..dbedfe72e1fd 100644 --- a/include/net/dst.h +++ b/include/net/dst.h @@ -455,7 +455,8 @@ static inline unsigned int dst_dev_overhead(struct dst_entry *dst, struct sk_buff *skb) { if (likely(dst)) - return LL_RESERVED_SPACE(dst->dev); + return max_t(unsigned int, skb->mac_len, + LL_RESERVED_SPACE(dst->dev)); return skb->mac_len; } From 0e2bec77ea62895416600c90588f593516572bca Mon Sep 17 00:00:00 2001 From: Florian Fainelli Date: Mon, 21 Sep 2026 15:00:17 -0700 Subject: [PATCH 138/189] net: bcmgenet: fix 64-bit RTNL stats reading in ethtool on 32-bit systems When bcmgenet was converted to 64-bit statistics, STAT_RTNL members were switched to point into struct rtnl_link_stats64, whose fields are 64-bit (__u64) regardless of architecture. However, bcmgenet_get_ethtool_stats() retained a legacy check: if (sizeof(unsigned long) != sizeof(u32) && s->stat_sizeof == sizeof(unsigned long)) On 32-bit systems, sizeof(unsigned long) == sizeof(u32), causing this condition to evaluate to false. As a result, 64-bit RTNL stats fields were read via *(u32 *)p. On 32-bit Big-Endian systems (such as MIPS BE), this reads the high 32 bits and returns 0 until the counter exceeds 4GB; on 32-bit Little-Endian systems (such as 32-bit ARM), the value is truncated to 32 bits. Fix this by checking if s->stat_sizeof == sizeof(u64) so 64-bit fields are always read as 64-bit values. Fixes: 59aa6e3072aa ("net: bcmgenet: switch to use 64bit statistics") Reviewed-by: Nicolai Buchwitz Signed-off-by: Florian Fainelli Link: https://patch.msgid.link/20260921220021.281418-2-florian.fainelli@broadcom.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/broadcom/genet/bcmgenet.c | 5 ++--- 1 file changed, 2 insertions(+), 3 deletions(-) diff --git a/drivers/net/ethernet/broadcom/genet/bcmgenet.c b/drivers/net/ethernet/broadcom/genet/bcmgenet.c index ca62041efecd..47a6c073c8d9 100644 --- a/drivers/net/ethernet/broadcom/genet/bcmgenet.c +++ b/drivers/net/ethernet/broadcom/genet/bcmgenet.c @@ -1346,9 +1346,8 @@ static void bcmgenet_get_ethtool_stats(struct net_device *dev, p = (char *)&stats64; p += s->stat_offset; - if (sizeof(unsigned long) != sizeof(u32) && - s->stat_sizeof == sizeof(unsigned long)) - data[i] = *(unsigned long *)p; + if (s->stat_sizeof == sizeof(u64)) + data[i] = *(u64 *)p; else data[i] = *(u32 *)p; } From 3aeaa609fda19c09d5298c9fedaaa3b6229601b5 Mon Sep 17 00:00:00 2001 From: Florian Fainelli Date: Mon, 21 Sep 2026 15:00:18 -0700 Subject: [PATCH 139/189] net: bcmgenet: initialize u64 stats seq counter for all queues bcmgenet_gstrings_stats statically defines ethtool statistics for queues 0 through GENET_MAX_MQ_CNT (4). However, bcmgenet_probe() only initialized the u64_stats_sync seq counter up to priv->hw_params->rx_queues and priv->hw_params->tx_queues. Since priv->hw_params->rx_queues is 0 across all hardware versions (and priv->hw_params->tx_queues is 0 on GENET V1), rings 1..4 have uninitialized u64_stats_sync structures. When ethtool -S is run on 32-bit kernels, bcmgenet_get_ethtool_stats() reads stats from rx_rings[1..4], causing lockdep warnings due to the uninitialized sequence counters. Initialize the sequence counters for all GENET_MAX_MQ_CNT + 1 queues. Fixes: ffc2c8c4a714 ("net: bcmgenet: Initialize u64 stats seq counter") Reviewed-by: Nicolai Buchwitz Signed-off-by: Florian Fainelli Link: https://patch.msgid.link/20260921220021.281418-3-florian.fainelli@broadcom.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/broadcom/genet/bcmgenet.c | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/drivers/net/ethernet/broadcom/genet/bcmgenet.c b/drivers/net/ethernet/broadcom/genet/bcmgenet.c index 47a6c073c8d9..2ab37bb031ed 100644 --- a/drivers/net/ethernet/broadcom/genet/bcmgenet.c +++ b/drivers/net/ethernet/broadcom/genet/bcmgenet.c @@ -4135,10 +4135,10 @@ static int bcmgenet_probe(struct platform_device *pdev) priv->rx_rings[i].rx_max_coalesced_frames = 1; /* Initialize u64 stats seq counter for 32bit machines */ - for (i = 0; i <= priv->hw_params->rx_queues; i++) + for (i = 0; i <= GENET_MAX_MQ_CNT; i++) { u64_stats_init(&priv->rx_rings[i].stats64.syncp); - for (i = 0; i <= priv->hw_params->tx_queues; i++) u64_stats_init(&priv->tx_rings[i].stats64.syncp); + } /* libphy will determine the link state */ netif_carrier_off(dev); From cbbc1aee7776c7fa1d89e6cb963a23e58c495dca Mon Sep 17 00:00:00 2001 From: Florian Fainelli Date: Mon, 21 Sep 2026 15:00:19 -0700 Subject: [PATCH 140/189] net: bcmgenet: do not skip WoL power up on GENET V1 bcmgenet_power_up() had an early check for bcmgenet_has_ext(priv) before dispatching by power mode. GENET V1 does not have the EXT block (unlike GENET V2+), which causes bcmgenet_power_up() to immediately return 0. As a consequence, when waking up from GENET_POWER_WOL_MAGIC on GENET V1, bcmgenet_wol_power_up_cfg() is never invoked to disable the WoL clock, clear wake event masks, and restore normal PHY and MAC operations. Move the bcmgenet_has_ext() checks to the GENET_POWER_PASSIVE and GENET_POWER_CABLE_SENSE cases where the EXT registers are actually accessed, allowing GENET_POWER_WOL_MAGIC cleanup to execute on all hardware versions. Fixes: c3ae64ae0c08 ("net: bcmgenet: handle GENET_POWER_WOL_MAGIC") Reviewed-by: Nicolai Buchwitz Signed-off-by: Florian Fainelli Link: https://patch.msgid.link/20260921220021.281418-4-florian.fainelli@broadcom.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/broadcom/genet/bcmgenet.c | 13 ++++++++----- 1 file changed, 8 insertions(+), 5 deletions(-) diff --git a/drivers/net/ethernet/broadcom/genet/bcmgenet.c b/drivers/net/ethernet/broadcom/genet/bcmgenet.c index 2ab37bb031ed..07032e4193ea 100644 --- a/drivers/net/ethernet/broadcom/genet/bcmgenet.c +++ b/drivers/net/ethernet/broadcom/genet/bcmgenet.c @@ -1762,13 +1762,12 @@ static int bcmgenet_power_up(struct bcmgenet_priv *priv, int ret = 0; u32 reg; - if (!bcmgenet_has_ext(priv)) - return ret; - - reg = bcmgenet_ext_readl(priv, EXT_EXT_PWR_MGMT); - switch (mode) { case GENET_POWER_PASSIVE: + if (!bcmgenet_has_ext(priv)) + break; + + reg = bcmgenet_ext_readl(priv, EXT_EXT_PWR_MGMT); reg &= ~(EXT_PWR_DOWN_DLL | EXT_PWR_DOWN_BIAS | EXT_ENERGY_DET_MASK); if (GENET_IS_V5(priv) && !bcmgenet_has_ephy_16nm(priv)) { @@ -1792,8 +1791,12 @@ static int bcmgenet_power_up(struct bcmgenet_priv *priv, break; case GENET_POWER_CABLE_SENSE: + if (!bcmgenet_has_ext(priv)) + break; + /* enable APD */ if (!GENET_IS_V5(priv)) { + reg = bcmgenet_ext_readl(priv, EXT_EXT_PWR_MGMT); reg |= EXT_PWR_DN_EN_LD; bcmgenet_ext_writel(priv, reg, EXT_EXT_PWR_MGMT); } From 273941c85fc2632cd3e56ddff737b9245de7697d Mon Sep 17 00:00:00 2001 From: Florian Fainelli Date: Mon, 21 Sep 2026 15:00:20 -0700 Subject: [PATCH 141/189] net: bcmgenet: validate Ethernet address in bcmgenet_set_mac_addr bcmgenet_set_mac_addr() did not check whether the provided MAC address is a valid Ethernet address before applying it. Userspace could configure an invalid address (such as all zeroes or a multicast address) while the interface is down. Add a call to is_valid_ether_addr() and return -EADDRNOTAVAIL if the MAC address is not valid. Fixes: 1c1008c793fa ("net: bcmgenet: add main driver file") Reviewed-by: Nicolai Buchwitz Signed-off-by: Florian Fainelli Link: https://patch.msgid.link/20260921220021.281418-5-florian.fainelli@broadcom.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/broadcom/genet/bcmgenet.c | 3 +++ 1 file changed, 3 insertions(+) diff --git a/drivers/net/ethernet/broadcom/genet/bcmgenet.c b/drivers/net/ethernet/broadcom/genet/bcmgenet.c index 07032e4193ea..011376a1678a 100644 --- a/drivers/net/ethernet/broadcom/genet/bcmgenet.c +++ b/drivers/net/ethernet/broadcom/genet/bcmgenet.c @@ -3635,6 +3635,9 @@ static int bcmgenet_set_mac_addr(struct net_device *dev, void *p) if (netif_running(dev)) return -EBUSY; + if (!is_valid_ether_addr(addr->sa_data)) + return -EADDRNOTAVAIL; + eth_hw_addr_set(dev, addr->sa_data); return 0; From d64e277b955be4506931802837499b62c8f3968a Mon Sep 17 00:00:00 2001 From: Florian Fainelli Date: Mon, 21 Sep 2026 15:00:21 -0700 Subject: [PATCH 142/189] net: bcmgenet: mask DMA_TIMEOUT_MASK when reading DMA_RING0_TIMEOUT bcmgenet_get_coalesce() reads DMA_RING0_TIMEOUT to calculate rx_coalesce_usecs without masking out bits outside DMA_TIMEOUT_MASK (16 bits). If upper bits are non-zero or contain status/flags, the computed value of rx_coalesce_usecs returned to userspace via ethtool becomes corrupted. Mask the register read with DMA_TIMEOUT_MASK before computing the timeout in microseconds. Fixes: 4a29645bfe6c ("net: bcmgenet: Implement RX coalescing control knobs") Reviewed-by: Nicolai Buchwitz Signed-off-by: Florian Fainelli Link: https://patch.msgid.link/20260921220021.281418-6-florian.fainelli@broadcom.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/broadcom/genet/bcmgenet.c | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/drivers/net/ethernet/broadcom/genet/bcmgenet.c b/drivers/net/ethernet/broadcom/genet/bcmgenet.c index 011376a1678a..21668e41b696 100644 --- a/drivers/net/ethernet/broadcom/genet/bcmgenet.c +++ b/drivers/net/ethernet/broadcom/genet/bcmgenet.c @@ -852,7 +852,8 @@ static int bcmgenet_get_coalesce(struct net_device *dev, ec->rx_max_coalesced_frames = bcmgenet_rdma_ring_readl(priv, 0, DMA_MBUF_DONE_THRESH); ec->rx_coalesce_usecs = - bcmgenet_rdma_readl(priv, DMA_RING0_TIMEOUT) * 8192 / 1000; + (bcmgenet_rdma_readl(priv, DMA_RING0_TIMEOUT) & + DMA_TIMEOUT_MASK) * 8192 / 1000; for (i = 0; i <= priv->hw_params->rx_queues; i++) { ring = &priv->rx_rings[i]; From 4bdee8060d1e4581624e68fbd369b1afb14df4bc Mon Sep 17 00:00:00 2001 From: Myeonghun Pak Date: Mon, 21 Sep 2026 20:09:14 -0400 Subject: [PATCH 143/189] net: airoha: npu: cancel wdt_work after releasing the WDT IRQ airoha_npu_remove() calls cancel_work_sync() on each core's wdt_work, but the watchdog IRQ that queues it is requested with devm_request_irq() and is freed only after .remove() returns. airoha_npu_wdt_handler() can therefore schedule_work() again once the cancel has returned. struct airoha_npu, which contains the work, is devm_kzalloc()'d and is freed in that same unwind, so the late work dereferences freed memory. Register the work with devm_work_autocancel() before devm_request_irq() and drop .remove(). Devres runs in reverse order, so the IRQ is freed before cancel_work_sync(), including when probe fails. A cancel left in .remove() cannot get that order. Initializing the work first also stops a pending watchdog interrupt from queuing an uninitialized work item. Probe currently calls INIT_WORK() only after devm_request_irq(). This issue was identified during our ongoing static-analysis research while reviewing kernel code. Fixes: 23290c7bc190 ("net: airoha: Introduce Airoha NPU support") Cc: stable@vger.kernel.org # 6.15+ Co-developed-by: Ijae Kim Signed-off-by: Ijae Kim Signed-off-by: Myeonghun Pak Acked-by: Lorenzo Bianconi Link: https://patch.msgid.link/20260922000914.542068-1-mhun512@gmail.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/airoha/airoha_npu.c | 18 ++++++------------ 1 file changed, 6 insertions(+), 12 deletions(-) diff --git a/drivers/net/ethernet/airoha/airoha_npu.c b/drivers/net/ethernet/airoha/airoha_npu.c index 5bb4817a898d..4d3195eb00f7 100644 --- a/drivers/net/ethernet/airoha/airoha_npu.c +++ b/drivers/net/ethernet/airoha/airoha_npu.c @@ -5,6 +5,7 @@ */ #include +#include #include #include #include @@ -751,12 +752,15 @@ static int airoha_npu_probe(struct platform_device *pdev) if (irq < 0) return irq; + err = devm_work_autocancel(dev, &core->wdt_work, + airoha_npu_wdt_work); + if (err) + return err; + err = devm_request_irq(dev, irq, airoha_npu_wdt_handler, IRQF_SHARED, "airoha-npu-wdt", core); if (err) return err; - - INIT_WORK(&core->wdt_work, airoha_npu_wdt_work); } /* wlan IRQ lines */ @@ -803,18 +807,8 @@ static int airoha_npu_probe(struct platform_device *pdev) return 0; } -static void airoha_npu_remove(struct platform_device *pdev) -{ - struct airoha_npu *npu = platform_get_drvdata(pdev); - int i; - - for (i = 0; i < ARRAY_SIZE(npu->cores); i++) - cancel_work_sync(&npu->cores[i].wdt_work); -} - static struct platform_driver airoha_npu_driver = { .probe = airoha_npu_probe, - .remove = airoha_npu_remove, .driver = { .name = "airoha-npu", .of_match_table = of_airoha_npu_match, From a3f315be9d30eeb6938d11fa17fd4b32d52f7c42 Mon Sep 17 00:00:00 2001 From: Xuanqiang Luo Date: Mon, 21 Sep 2026 11:18:59 +0800 Subject: [PATCH 144/189] ip_gre: Reject enabling collect metadata through changelink ipgre_netlink_parms() can enable collect_md on an existing GRE, GRETAP or ERSPAN device. Unlike newlink, changelink does not enforce metadata tunnel uniqueness. Converting a non-metadata device can therefore replace the metadata receive entry for another device of the same type in the same netns. Deleting either device then clears the shared entry, breaking metadata receive lookup for the surviving device. If parameter validation fails after collect_md is set, deleting the modified device can also clear an entry it never owned. Reject enabling metadata mode in both changelink callbacks before any encapsulation or tunnel parameters are modified. Allow requests that repeat the metadata attribute on an existing metadata device. Fixes: 2e15ea390e6f ("ip_gre: Add support to collect tunnel metadata.") Signed-off-by: Xuanqiang Luo Reviewed-by: Ido Schimmel Reviewed-by: Hangbin Liu Link: https://patch.msgid.link/20260921031859.9283-1-xuanqiang.luo@linux.dev Signed-off-by: Jakub Kicinski --- net/ipv4/ip_gre.c | 12 ++++++++++++ 1 file changed, 12 insertions(+) diff --git a/net/ipv4/ip_gre.c b/net/ipv4/ip_gre.c index 82309efd417e..e4878e9aa636 100644 --- a/net/ipv4/ip_gre.c +++ b/net/ipv4/ip_gre.c @@ -1464,6 +1464,12 @@ static int ipgre_changelink(struct net_device *dev, struct nlattr *tb[], if (!rtnl_dev_link_net_capable(dev, t->net)) return -EPERM; + if (data && data[IFLA_GRE_COLLECT_METADATA] && !t->collect_md) { + NL_SET_ERR_MSG(extack, + "Enabling collect_md on an existing device is not supported"); + return -EOPNOTSUPP; + } + err = ipgre_newlink_encap_setup(dev, data); if (err) return err; @@ -1496,6 +1502,12 @@ static int erspan_changelink(struct net_device *dev, struct nlattr *tb[], if (!rtnl_dev_link_net_capable(dev, t->net)) return -EPERM; + if (data && data[IFLA_GRE_COLLECT_METADATA] && !t->collect_md) { + NL_SET_ERR_MSG(extack, + "Enabling collect_md on an existing device is not supported"); + return -EOPNOTSUPP; + } + err = ipgre_newlink_encap_setup(dev, data); if (err) return err; From 0f2fd31f63c65c409bd336af3a42d478a9207c1f Mon Sep 17 00:00:00 2001 From: Gilberto Conde Date: Mon, 21 Sep 2026 10:34:26 +0100 Subject: [PATCH 145/189] net: usb: qmi_wwan: add Quectel EG120K-EA Add support for the Quectel EG120K-EA LTE Cat.12 module (USB ID 2c7c:030b). Its QMI interface (interface 4) uses class/subclass/protocol ff/ff/ff like the other recent Quectel modules, so match it the same way. The product ID is shared with the EM060K, which the option driver already knows and uses to claim the serial interfaces. Without a qmi_wwan entry the data interface is left unbound. Tested on a GL.iNet GL-X2000, where the module is soldered down and enumerates at SuperSpeed: cdc-wdm0 and wwan0 appear and ModemManager brings up a data session. T: Bus=02 Lev=01 Prnt=01 Port=00 Cnt=01 Dev#= 2 Spd=5000 MxCh= 0 D: Ver= 3.10 Cls=00(>ifc ) Sub=00 Prot=00 MxPS= 9 #Cfgs= 1 P: Vendor=2c7c ProdID=030b Rev= 5.04 S: Manufacturer=Quectel S: Product=EG120K-EA S: SerialNumber=45546267 C:* #Ifs= 6 Cfg#= 1 Atr=a0 MxPwr=896mA I:* If#= 0 Alt= 0 #EPs= 2 Cls=ff(vend.) Sub=ff Prot=30 Driver=option E: Ad=01(O) Atr=02(Bulk) MxPS=1024 Ivl=0ms E: Ad=81(I) Atr=02(Bulk) MxPS=1024 Ivl=0ms I:* If#= 1 Alt= 0 #EPs= 3 Cls=ff(vend.) Sub=00 Prot=40 Driver=option E: Ad=83(I) Atr=03(Int.) MxPS= 10 Ivl=32ms E: Ad=82(I) Atr=02(Bulk) MxPS=1024 Ivl=0ms E: Ad=02(O) Atr=02(Bulk) MxPS=1024 Ivl=0ms I:* If#= 2 Alt= 0 #EPs= 3 Cls=ff(vend.) Sub=ff Prot=40 Driver=option E: Ad=85(I) Atr=03(Int.) MxPS= 10 Ivl=32ms E: Ad=84(I) Atr=02(Bulk) MxPS=1024 Ivl=0ms E: Ad=03(O) Atr=02(Bulk) MxPS=1024 Ivl=0ms I:* If#= 3 Alt= 0 #EPs= 3 Cls=ff(vend.) Sub=ff Prot=40 Driver=option E: Ad=87(I) Atr=03(Int.) MxPS= 10 Ivl=32ms E: Ad=86(I) Atr=02(Bulk) MxPS=1024 Ivl=0ms E: Ad=04(O) Atr=02(Bulk) MxPS=1024 Ivl=0ms I:* If#= 4 Alt= 0 #EPs= 3 Cls=ff(vend.) Sub=ff Prot=ff Driver=qmi_wwan E: Ad=88(I) Atr=03(Int.) MxPS= 8 Ivl=32ms E: Ad=8e(I) Atr=02(Bulk) MxPS=1024 Ivl=0ms E: Ad=0f(O) Atr=02(Bulk) MxPS=1024 Ivl=0ms I:* If#=12 Alt= 0 #EPs= 1 Cls=ff(vend.) Sub=ff Prot=70 Driver=(none) E: Ad=89(I) Atr=02(Bulk) MxPS=1024 Ivl=0ms Signed-off-by: Gilberto Conde Link: https://patch.msgid.link/20260921093426.2870266-1-gilbertorconde@gmail.com Signed-off-by: Jakub Kicinski --- drivers/net/usb/qmi_wwan.c | 1 + 1 file changed, 1 insertion(+) diff --git a/drivers/net/usb/qmi_wwan.c b/drivers/net/usb/qmi_wwan.c index f51cf9cb9421..0e2ab567dd40 100644 --- a/drivers/net/usb/qmi_wwan.c +++ b/drivers/net/usb/qmi_wwan.c @@ -1086,6 +1086,7 @@ static const struct usb_device_id products[] = { {QMI_MATCH_FF_FF_FF(0x2c7c, 0x0125)}, /* Quectel EC25, EC20 R2.0 Mini PCIe */ {QMI_MATCH_FF_FF_FF(0x2c7c, 0x013d)}, /* Quectel RG660QB */ {QMI_MATCH_FF_FF_FF(0x2c7c, 0x0306)}, /* Quectel EP06/EG06/EM06 */ + {QMI_MATCH_FF_FF_FF(0x2c7c, 0x030b)}, /* Quectel EM060K/EG120K-EA */ {QMI_MATCH_FF_FF_FF(0x2c7c, 0x0512)}, /* Quectel EG12/EM12 */ {QMI_MATCH_FF_FF_FF(0x2c7c, 0x0620)}, /* Quectel EM160R-GL */ {QMI_MATCH_FF_FF_FF(0x2c7c, 0x0800)}, /* Quectel RM500Q-GL */ From 481506a756dcd828ef42391cb08f38d8d96d38fc Mon Sep 17 00:00:00 2001 From: Sanghyun Park Date: Fri, 18 Sep 2026 12:26:58 +0900 Subject: [PATCH 146/189] vxlan: use one headroom snapshot for neighbour replies vxlan_na_create() samples LL_RESERVED_SPACE() to size the reply skb and then samples it again to reserve headroom. A concurrent vxlan_changelink() can update needed_headroom between the two reads, creating a TOCTOU race. The second value can exceed the allocation and make the Ethernet header write out of bounds. The race is reproducible on the unpatched kernel. It occurred when vxlan_na_create() generated a neighbour reply while vxlan_changelink() changed the link headroom. KASAN caught a four-byte write two bytes beyond a 704-byte skbuff_small_head allocation. Snapshot the headroom once and use that value for both allocation and reservation. Fixes: 4b29dba9c085 ("vxlan: fix nonfunctional neigh_reduce()") Signed-off-by: Sanghyun Park Link: https://patch.msgid.link/20260918032842.502409-2-sanghyun.park.cnu@gmail.com Signed-off-by: Jakub Kicinski --- drivers/net/vxlan/vxlan_core.c | 6 ++++-- 1 file changed, 4 insertions(+), 2 deletions(-) diff --git a/drivers/net/vxlan/vxlan_core.c b/drivers/net/vxlan/vxlan_core.c index c1d54339fa2b..3390e341d1e5 100644 --- a/drivers/net/vxlan/vxlan_core.c +++ b/drivers/net/vxlan/vxlan_core.c @@ -1958,13 +1958,15 @@ static struct sk_buff *vxlan_na_create(struct sk_buff *request, struct ipv6hdr *pip6; u8 *daddr; int na_olen = 8; /* opt hdr + ETH_ALEN for target */ + int headroom; int ns_olen; int i, len; if (dev == NULL || !pskb_may_pull(request, request->len)) return NULL; - len = LL_RESERVED_SPACE(dev) + sizeof(struct ipv6hdr) + + headroom = LL_RESERVED_SPACE(dev); + len = headroom + sizeof(struct ipv6hdr) + sizeof(*na) + na_olen + dev->needed_tailroom; reply = alloc_skb(len, GFP_ATOMIC); if (reply == NULL) @@ -1972,7 +1974,7 @@ static struct sk_buff *vxlan_na_create(struct sk_buff *request, reply->protocol = htons(ETH_P_IPV6); reply->dev = dev; - skb_reserve(reply, LL_RESERVED_SPACE(request->dev)); + skb_reserve(reply, headroom); skb_push(reply, sizeof(struct ethhdr)); skb_reset_mac_header(reply); From 4da3b7b8b50f3e2fde54a4c18a82a8e3f6223910 Mon Sep 17 00:00:00 2001 From: Norbert Szetei Date: Mon, 21 Sep 2026 17:03:57 +0200 Subject: [PATCH 147/189] net: xps: reject an out of range traffic class Only the entries below dev->num_tc are valid in dev->tc_to_txq[], and dev->prio_tc_map[] may only name classes below it. netdev_set_num_tc() lowers dev->num_tc without touching either array. netdev_txq_to_tc() walks all TC_MAX_QUEUE slots and netdev_get_prio_tc_map() returns the entry as it stands, so a leftover entry is handed out as a traffic class >= dev->num_tc. Taking that class from netdev_txq_to_tc(), __netif_set_xps_queue() rejects only a negative one and indexes an XPS map sized for dev->num_tc classes: tci = j * num_tc + tc; RCU_INIT_POINTER(new_dev_maps->attr_map[tci], map); attr_map[] holds nr_ids * num_tc entries and j runs over the ids named in the mask, so a class that is not below num_tc pushes tci past the end of the map for the last ids and the store overruns it. Any caller that lowers num_tc leaves such entries behind, and mqprio_destroy() tears down with netdev_set_num_tc(dev, 0) rather than netdev_reset_tc(). After mqprio with 8 classes then 1, tc_to_txq[1..7] still describe txq 1..7. The splat is from an XPS write to txq 2 on a veth with 8 rx queues: attr_map[] has 8 * 1 entries, tci = j + 2, and j == 6 stores one past the end of the 88-byte map: BUG: KASAN: slab-out-of-bounds in __netif_set_xps_queue (net/core/dev.c:2954) Write of size 8 at addr ffff88813016bc58 by task xps_oob/634 __netif_set_xps_queue (net/core/dev.c:2954) xps_rxqs_store (net/core/net-sysfs.c:1880) netdev_queue_attr_store (net/core/net-sysfs.c:1390) Allocated by task 634: __kmalloc_noprof (mm/slub.c:5439) __netif_set_xps_queue (net/core/dev.c:2937) The buggy address is located 0 bytes to the right of allocated 88-byte region [ffff88813016bc00, ffff88813016bc58) Reject a class the map has no room for. Fixes: 184c449f91fe ("net: Add support for XPS with QoS via traffic classes") Signed-off-by: Norbert Szetei Link: https://patch.msgid.link/162DD16F-54C6-444A-9E09-0B8CB3D591F2@doyensec.com Signed-off-by: Jakub Kicinski --- net/core/dev.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/net/core/dev.c b/net/core/dev.c index c67900354fa6..0292a16e16c2 100644 --- a/net/core/dev.c +++ b/net/core/dev.c @@ -2901,7 +2901,7 @@ int __netif_set_xps_queue(struct net_device *dev, const unsigned long *mask, dev = netdev_get_tx_queue(dev, index)->sb_dev ? : dev; tc = netdev_txq_to_tc(dev, index); - if (tc < 0) + if (tc < 0 || tc >= num_tc) return -EINVAL; } From e47a1958e12abc3a17b5231a4f21c8f1bf662e08 Mon Sep 17 00:00:00 2001 From: Yuqi Xu Date: Sat, 19 Sep 2026 16:45:27 +0800 Subject: [PATCH 148/189] net: ipconfig: bound DHCP option construction ic_dhcp_init_options() appends the hostname (option 12), vendor-class (option 60) and client-ID (option 61) options into the fixed 312-byte bootp_pkt.exten[] buffer. Only the client-ID branch checked the remaining space; the hostname and vendor-class writes were unbounded. A 64-byte hostname together with the maximum 252-byte dhcpclass= identifier needs 18 + (2 + 64) + (2 + 252) = 338 of the 312 available bytes even before the terminating END marker, so the vendor-class memcpy runs past the end of exten[]. With CONFIG_FORTIFY_SOURCE this is reported as a field-spanning write and, when the kernel is booted with panic_on_warn=1, aborts boot with a panic. Route the optional options through a common helper that makes sure the option, its 2-byte header and the END marker all fit and drops an option that would not. Configurations with short options keep sending exactly the same bytes as before. Fixes: 130c0f47fdf9 ("ipconfig: send host-name in DHCP requests") Cc: stable@vger.kernel.org Reported-by: Vega Assisted-by: LLM Signed-off-by: Yuqi Xu Reviewed-by: Ren Wei Reviewed-by: Simon Horman Link: https://patch.msgid.link/7808dfbfa2162dfd0b19f59aff5742d6e0db2abb.1789798023.git.xuyuqiabc@gmail.com Signed-off-by: Paolo Abeni --- net/ipv4/ipconfig.c | 45 ++++++++++++++++++++++++++------------------- 1 file changed, 26 insertions(+), 19 deletions(-) diff --git a/net/ipv4/ipconfig.c b/net/ipv4/ipconfig.c index 155db067eaec..1b8585404a41 100644 --- a/net/ipv4/ipconfig.c +++ b/net/ipv4/ipconfig.c @@ -676,6 +676,24 @@ static const u8 ic_bootp_cookie[4] = { 99, 130, 83, 99 }; #ifdef IPCONFIG_DHCP +static bool __init +ic_dhcp_add_option(u8 **options, const u8 *end, u8 type, const void *value, + int len) +{ + u8 *e = *options; + + /* leave room for the option header and the END marker */ + if (len > U8_MAX || end - e < len + 3) + return false; + + *e++ = type; + *e++ = len; + memcpy(e, value, len); + *options = e + len; + + return true; +} + static void __init ic_dhcp_init_options(u8 *options, struct ic_device *d) { @@ -691,6 +709,7 @@ ic_dhcp_init_options(u8 *options, struct ic_device *d) 42, /* NTP servers */ }; u8 mt = (ic_servaddr == NONE) ? DHCPDISCOVER : DHCPREQUEST; + u8 *end = options + sizeof(((struct bootp_pkt *)0)->exten); u8 *e = options; int len; @@ -721,31 +740,19 @@ ic_dhcp_init_options(u8 *options, struct ic_device *d) e += sizeof(ic_req_params); if (ic_host_name_set) { - *e++ = 12; /* host-name */ len = strlen(utsname()->nodename); - *e++ = len; - memcpy(e, utsname()->nodename, len); - e += len; + ic_dhcp_add_option(&e, end, 12, utsname()->nodename, len); } if (*vendor_class_identifier) { - pr_info("DHCP: sending class identifier \"%s\"\n", - vendor_class_identifier); - *e++ = 60; /* Class-identifier */ len = strlen(vendor_class_identifier); - *e++ = len; - memcpy(e, vendor_class_identifier, len); - e += len; + if (ic_dhcp_add_option(&e, end, 60, vendor_class_identifier, len)) + pr_info("DHCP: sending class identifier \"%s\"\n", + vendor_class_identifier); } len = strlen(dhcp_client_identifier + 1); - /* the minimum length of identifier is 2, include 1 byte type, - * and can not be larger than the length of options - */ - if (len >= 1 && len < 312 - (e - options) - 1) { - *e++ = 61; - *e++ = len + 1; - memcpy(e, dhcp_client_identifier, len + 1); - e += len + 1; - } + /* the minimum length of identifier is 2, include 1 byte type */ + if (len >= 1) + ic_dhcp_add_option(&e, end, 61, dhcp_client_identifier, len + 1); *e++ = 255; /* End of the list */ } From 00efbbd40bd5fd92c67b7cf1aab8904fa59a96f6 Mon Sep 17 00:00:00 2001 From: David Dai Date: Fri, 18 Sep 2026 16:11:55 -0500 Subject: [PATCH 149/189] bonding: crypto offload enabled, non-offload slave failover, rekey failed Create a bonding device (i.e. bond0) in active-backup mode, 2 slaves. Active slave: offload capable interface (i.e. eth1), primary interface. Backup slave: non-offload capable interface(i.e. eth2). Configure strongswan service swantl.conf child SA "hw_offload = crypto" Start strongswan service IPSec Crytpo Offload is enabled on top of bond0. i.e. ip xfrm state |grep offload crypto offload parameters: dev bond0 dir out mode crypto crypto offload parameters: dev bond0 dir in mode crypto Active slave eth1 takes adavantage of IPSec Crypto Offload capability. If active slave eth1 is down for any reason (i.e. eth1 link down): ip link set down dev eth1 non-offload capable interface eth2 failover to becomes active slave. The existing SAs can continue use software IPsec after failover. Traffic still keeps going properly. However if eth1 link had not recovered yet, strongswan service does new child SA rekey, or uses swanctl command to do new child SA rekey, it will fail because active slave eth2 doesn't support crypto offload. In bond_ipsec_add_sa routine, it returns -EINVAL now, which is treated as fatal error by xfrm_dev_state_add routine in kernel xfrm. To make the non-offload active slave survive the child SA rekey, need to make bond_ipsec_add_sa routine returns -EOPNOTSUPP instead when active slave doesn't support IPsec Crypto offload, the xfrm will gracefully fallback to create new SA using Software IPsec. Network traffic can keep going. After offload capable interface eth1 link is up, becomes active slave, next time strongswan child SA rekey will create a new SA which enables crypto offload again. Fixes: 18cb261afd7b ("bonding: support hardware encryption offload to slaves") Signed-off-by: David Dai Reviewed-by: Hangbin Liu Link: https://patch.msgid.link/20260918211155.1664493-1-zdai@linux.ibm.com Signed-off-by: Paolo Abeni --- drivers/net/bonding/bond_main.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/drivers/net/bonding/bond_main.c b/drivers/net/bonding/bond_main.c index a9bff7663eec..de2489c3d9bf 100644 --- a/drivers/net/bonding/bond_main.c +++ b/drivers/net/bonding/bond_main.c @@ -490,7 +490,7 @@ static int bond_ipsec_add_sa(struct net_device *bond_dev, !real_dev->xfrmdev_ops->xdo_dev_state_add || netif_is_bond_master(real_dev)) { NL_SET_ERR_MSG_MOD(extack, "Slave does not support ipsec offload"); - err = -EINVAL; + err = -EOPNOTSUPP; goto out; } From 36c2009d90f2210ef92e6f4f2850e8b57b09e754 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Gajdos=20Tam=C3=A1s?= Date: Mon, 21 Sep 2026 11:13:32 +0200 Subject: [PATCH 150/189] net: atl1c: fix soft lockup on out-of-range tpd_cons read MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The hardware can report an out-of-range tpd_cons (seen as 0xffff) while the PCIe link/MAC is resetting. An out-of-range value can never be reached and the loop below would spin forever. To avoid a soft lockup treat it as "nothing new to clean" instead. Reproduced on two machines, same NIC (Qualcomm Atheros AR8151 v2.0, 4-port), triggered by rebooting a Mikrotik CCR2004 PCIe card that the ports are directly linked to: - Ubuntu 26.04.1 LTS, kernel 7.0.0-31-generic. The link-flap precursor, before the lockup was captured with a full trace elsewhere: atl1c 0000:05:00.0 enp5s0f0: NETDEV WATCHDOG: CPU: 4: transmit queue 2 timed out 489984 ms atl1c 0000:05:00.0: MAC state machine can't be idle since disabled for 10ms second atl1c 0000:05:00.0: atl1c: enp5s0f0 NIC Link is Up<65535 Mbps Full Duplex> 65535 (0xffff) here is the same value tpd_cons reads back once the loop below gets stuck. - Proxmox VE, kernel 7.0.14-11-pve. Same NIC/trigger, this time caught by the soft lockup watchdog with a full stack trace: watchdog: BUG: soft lockup - CPU#12 stuck for 354s! [napi/eth%d-0:329] CPU: 12 UID: 0 PID: 329 Comm: napi/eth%d-0 Tainted: P O L 7.0.14-11-pve #1 PREEMPT(lazy) RIP: 0010:atl1c_clean_tx+0x142/0x2d0 [atl1c] Call Trace: __napi_poll+0x32/0x1e0 napi_threaded_poll_loop+0x286/0x2e0 napi_threaded_poll+0xfd/0x140 kthread+0xf7/0x130 ret_from_fork+0x2da/0x3a0 ret_from_fork_asm+0x1a/0x30 Fixes: 43250ddd75a35d ("atl1c: Atheros L1C Gigabit Ethernet driver") Cc: stable@vger.kernel.org Signed-off-by: Gajdos Tamás Link: https://patch.msgid.link/20260921091334.3571525-2-tamas@rimpianto.com Signed-off-by: Paolo Abeni --- drivers/net/ethernet/atheros/atl1c/atl1c_main.c | 3 +++ 1 file changed, 3 insertions(+) diff --git a/drivers/net/ethernet/atheros/atl1c/atl1c_main.c b/drivers/net/ethernet/atheros/atl1c/atl1c_main.c index 7efa3fc257b3..e58f1d2c26bd 100644 --- a/drivers/net/ethernet/atheros/atl1c/atl1c_main.c +++ b/drivers/net/ethernet/atheros/atl1c/atl1c_main.c @@ -1602,6 +1602,9 @@ static int atl1c_clean_tx(struct napi_struct *napi, int budget) AT_READ_REGW(&adapter->hw, atl1c_qregs[tpd_ring->num].tpd_cons, &hw_next_to_clean); + if (unlikely(hw_next_to_clean >= tpd_ring->count)) + hw_next_to_clean = next_to_clean; + while (next_to_clean != hw_next_to_clean) { buffer_info = &tpd_ring->buffer_info[next_to_clean]; if (buffer_info->skb) { From 374bf9e4b90f979e052332c4faca2d745c491a12 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Gajdos=20Tam=C3=A1s?= Date: Mon, 21 Sep 2026 11:13:33 +0200 Subject: [PATCH 151/189] net: atl1e: fix soft lockup on out-of-range hw_next_to_clean read MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Same issue as atl1c (see the first commit in this series, "net: atl1c: fix soft lockup on out-of-range tpd_cons read"): the hardware can report an out-of-range hw_next_to_clean (seen as 0xffff) while the PCIe link/MAC is resetting. An out-of-range value can never be reached and the loop below would spin forever. Treat it as "nothing new to clean" instead. Fixes: a6a5325239c202 ("atl1e: Atheros L1E Gigabit Ethernet driver") Cc: stable@vger.kernel.org Signed-off-by: Gajdos Tamás Link: https://patch.msgid.link/20260921091334.3571525-3-tamas@rimpianto.com Signed-off-by: Paolo Abeni --- drivers/net/ethernet/atheros/atl1e/atl1e_main.c | 3 +++ 1 file changed, 3 insertions(+) diff --git a/drivers/net/ethernet/atheros/atl1e/atl1e_main.c b/drivers/net/ethernet/atheros/atl1e/atl1e_main.c index 40290028580b..437989edb780 100644 --- a/drivers/net/ethernet/atheros/atl1e/atl1e_main.c +++ b/drivers/net/ethernet/atheros/atl1e/atl1e_main.c @@ -1234,6 +1234,9 @@ static bool atl1e_clean_tx_irq(struct atl1e_adapter *adapter) u16 hw_next_to_clean = AT_READ_REGW(&adapter->hw, REG_TPD_CONS_IDX); u16 next_to_clean = atomic_read(&tx_ring->next_to_clean); + if (unlikely(hw_next_to_clean >= tx_ring->count)) + hw_next_to_clean = next_to_clean; + while (next_to_clean != hw_next_to_clean) { tx_buffer = &tx_ring->tx_buffer[next_to_clean]; if (tx_buffer->dma) { From 43e746821f5f5afbbf68e388bf9fbe221e03bfca Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Gajdos=20Tam=C3=A1s?= Date: Mon, 21 Sep 2026 11:13:34 +0200 Subject: [PATCH 152/189] net: atl1: fix soft lockup on out-of-range cmb_tpd_next_to_clean read MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Same issue as atl1c (see the first commit in this series, "net: atl1c: fix soft lockup on out-of-range tpd_cons read"): the hardware can report an out-of-range cmb_tpd_next_to_clean (seen as 0xffff) while the PCIe link/MAC is resetting. An out-of-range value can never be reached and the loop below would spin forever. Treat it as "nothing new to clean" instead. Fixes: f3cc28c797604f ("Add Attansic L1 ethernet driver.") Cc: stable@vger.kernel.org Signed-off-by: Gajdos Tamás Link: https://patch.msgid.link/20260921091334.3571525-4-tamas@rimpianto.com Signed-off-by: Paolo Abeni --- drivers/net/ethernet/atheros/atlx/atl1.c | 3 +++ 1 file changed, 3 insertions(+) diff --git a/drivers/net/ethernet/atheros/atlx/atl1.c b/drivers/net/ethernet/atheros/atlx/atl1.c index 98a4d089270e..957d5598dda5 100644 --- a/drivers/net/ethernet/atheros/atlx/atl1.c +++ b/drivers/net/ethernet/atheros/atlx/atl1.c @@ -2066,6 +2066,9 @@ static int atl1_intr_tx(struct atl1_adapter *adapter) sw_tpd_next_to_clean = atomic_read(&tpd_ring->next_to_clean); cmb_tpd_next_to_clean = le16_to_cpu(adapter->cmb.cmb->tpd_cons_idx); + if (unlikely(cmb_tpd_next_to_clean >= tpd_ring->count)) + cmb_tpd_next_to_clean = sw_tpd_next_to_clean; + while (cmb_tpd_next_to_clean != sw_tpd_next_to_clean) { buffer_info = &tpd_ring->buffer_info[sw_tpd_next_to_clean]; if (buffer_info->dma) { From 77b1718e39e5c9f6956fb60807326af20baf889d Mon Sep 17 00:00:00 2001 From: Myeonghun Pak Date: Mon, 21 Sep 2026 21:46:05 -0400 Subject: [PATCH 153/189] bna: prevent IOC timer rearm during teardown bna: prevent IOC timer rearm during teardown bnad_pci_remove() and the probe disable_ioceth path call timer_delete_sync() for ioc_timer, sem_timer and hb_timer, but not for iocpf_timer. bnad_iocpf_timeout() then takes bnad->bna_lock after free_netdev() has freed the struct bnad. Deleting iocpf_timer last does not fix this. sem_timer and iocpf_timer rearm each other: bnad_iocpf_sem_timeout() can arm iocpf_timer, and bnad_iocpf_timeout() arms sem_timer from bfa_ioc_hw_sem_get() when the semaphore is busy. timer_delete_sync() only waits out its own callback. bnad_ioceth_disable() can time out and leave that callback live. Shut all four IOC timers down with timer_shutdown_sync() on both paths, so a later mod_timer() is ignored. This issue was identified during our ongoing static-analysis research while reviewing kernel code. Fixes: 1d32f7696286 ("bna: IOC failure auto recovery fix") Cc: stable@vger.kernel.org Assisted-by: LLM Co-developed-by: Ijae Kim Signed-off-by: Ijae Kim Signed-off-by: Myeonghun Pak Link: https://patch.msgid.link/20260922014605.588040-1-mhun512@gmail.com Signed-off-by: Paolo Abeni --- drivers/net/ethernet/brocade/bna/bnad.c | 24 ++++++++++++++++++------ 1 file changed, 18 insertions(+), 6 deletions(-) diff --git a/drivers/net/ethernet/brocade/bna/bnad.c b/drivers/net/ethernet/brocade/bna/bnad.c index 8b75004ba7c9..55dfd4896785 100644 --- a/drivers/net/ethernet/brocade/bna/bnad.c +++ b/drivers/net/ethernet/brocade/bna/bnad.c @@ -2571,6 +2571,22 @@ bnad_ioceth_disable(struct bnad *bnad) return err; } +/* + * The IOC timers rearm one another, so deleting one cannot stop a + * sibling callback from arming it again. Shut them down so a later + * mod_timer() is ignored. + */ +static void +bnad_ioc_timers_shutdown(struct bnad *bnad) +{ + struct bfa_ioc *ioc = &bnad->bna.ioceth.ioc; + + timer_shutdown_sync(&ioc->ioc_timer); + timer_shutdown_sync(&ioc->sem_timer); + timer_shutdown_sync(&ioc->hb_timer); + timer_shutdown_sync(&ioc->iocpf_timer); +} + static int bnad_ioceth_enable(struct bnad *bnad) { @@ -3727,9 +3743,7 @@ bnad_pci_probe(struct pci_dev *pdev, bnad_res_free(bnad, &bnad->mod_res_info[0], BNA_MOD_RES_T_MAX); disable_ioceth: bnad_ioceth_disable(bnad); - timer_delete_sync(&bnad->bna.ioceth.ioc.ioc_timer); - timer_delete_sync(&bnad->bna.ioceth.ioc.sem_timer); - timer_delete_sync(&bnad->bna.ioceth.ioc.hb_timer); + bnad_ioc_timers_shutdown(bnad); spin_lock_irqsave(&bnad->bna_lock, flags); bna_uninit(bna); spin_unlock_irqrestore(&bnad->bna_lock, flags); @@ -3770,9 +3784,7 @@ bnad_pci_remove(struct pci_dev *pdev) mutex_lock(&bnad->conf_mutex); bnad_ioceth_disable(bnad); - timer_delete_sync(&bnad->bna.ioceth.ioc.ioc_timer); - timer_delete_sync(&bnad->bna.ioceth.ioc.sem_timer); - timer_delete_sync(&bnad->bna.ioceth.ioc.hb_timer); + bnad_ioc_timers_shutdown(bnad); spin_lock_irqsave(&bnad->bna_lock, flags); bna_uninit(bna); spin_unlock_irqrestore(&bnad->bna_lock, flags); From 58eb1b3325edac42dc6df72c80962bd53a3c8ca7 Mon Sep 17 00:00:00 2001 From: Dongliang Qin Date: Tue, 22 Sep 2026 11:15:45 +0800 Subject: [PATCH 154/189] rds: ib: Clear the sg list when mapping an MR fails rds_ib_map_frmr() stores the caller's scatterlist in the MR before DMA mapping and registration can fail. On failure, __rds_rdma_map() unpins the pages and frees the scatterlist, but rds_ib_free_frmr() can still return the MR to the pool with the stale pointer set. This leaves the pool with a dangling scatterlist and can lead to local privilege escalation. KASAN detects the resulting use-after-free when the MR is later torn down: BUG: KASAN: slab-use-after-free in __rds_ib_teardown_mr Read of size 8 Call Trace: __rds_ib_teardown_mr rds_ib_unreg_frmr rds_ib_flush_mr_pool rds_ib_flush_mrs rds_free_mr rds_setsockopt Store the scatterlist in the MR only after DMA mapping succeeds. If DMA mapping fails, return directly while the MR fields remain clear; the caller keeps ownership of the scatterlist and its pinned pages. If a later registration step fails, unmap the scatterlist and clear the MR fields before returning. Fixes: 1659185fb4d0 ("RDS: IB: Support Fastreg MR (FRMR) memory registration mode") Cc: stable@vger.kernel.org Signed-off-by: Dongliang Qin Reviewed-by: Allison Henderson Link: https://patch.msgid.link/20260922031546.3874605-1-cccccccccccc777777@gmail.com Signed-off-by: Paolo Abeni --- net/rds/ib_frmr.c | 11 +++++------ 1 file changed, 5 insertions(+), 6 deletions(-) diff --git a/net/rds/ib_frmr.c b/net/rds/ib_frmr.c index bd861191157b..8397aa4a17ac 100644 --- a/net/rds/ib_frmr.c +++ b/net/rds/ib_frmr.c @@ -204,19 +204,16 @@ static int rds_ib_map_frmr(struct rds_ib_device *rds_ibdev, */ rds_ib_teardown_mr(ibmr); - ibmr->sg = sg; - ibmr->sg_len = sg_len; - ibmr->sg_dma_len = 0; frmr->sg_byte_len = 0; - WARN_ON(ibmr->sg_dma_len); - ibmr->sg_dma_len = ib_dma_map_sg(dev, ibmr->sg, ibmr->sg_len, + ibmr->sg_dma_len = ib_dma_map_sg(dev, sg, sg_len, DMA_BIDIRECTIONAL); if (unlikely(!ibmr->sg_dma_len)) { pr_warn("RDS/IB: %s failed!\n", __func__); return -EBUSY; } - frmr->sg_byte_len = 0; + ibmr->sg = sg; + ibmr->sg_len = sg_len; frmr->dma_npages = 0; len = 0; @@ -264,6 +261,8 @@ static int rds_ib_map_frmr(struct rds_ib_device *rds_ibdev, ib_dma_unmap_sg(rds_ibdev->dev, ibmr->sg, ibmr->sg_len, DMA_BIDIRECTIONAL); ibmr->sg_dma_len = 0; + ibmr->sg = NULL; + ibmr->sg_len = 0; return ret; } From c2de369c5c5b8599ca10fd5ca8d11fcd845c1331 Mon Sep 17 00:00:00 2001 From: Haseeb Malik Date: Mon, 21 Sep 2026 16:40:30 -0400 Subject: [PATCH 155/189] macsec: initialize SecY before registering the netdevice Creating a MACsec device with MAC offload over an LRO-capable lower device triggers a warning in rtmsg_ifinfo_build_skb() when IPv4 forwarding is enabled by default. register_netdevice() invokes inetdev_init(), which disables LRO and emits a NETDEV_FEAT_CHANGE notification. This reaches macsec_fill_info() before macsec_add_dev() initializes the SecY. key_len is still zero, so macsec_fill_info() returns -EMSGSIZE and trips the WARN_ON in rtmsg_ifinfo_build_skb(), even though the skb has enough space. Even without the warning, notifications during registration can report uninitialized SecY attributes, including the SCI. This ordering has existed since the driver was introduced. Initialize the SecY and apply the new-link attributes before registration. Move MAC address inheritance into macsec_newlink() so the SCI can also be initialized before registration-time notifications report it. Move the per-CPU statistics and metadata destination allocation into ndo_init(), and release partial allocations on failure. Fixes: c09440f7dcb3 ("macsec: introduce IEEE 802.1AE driver") Reported-by: syzbot+f2f6312ad1b5a0bfe316@syzkaller.appspotmail.com Closes: https://syzkaller.appspot.com/bug?extid=f2f6312ad1b5a0bfe316 Suggested-by: Sabrina Dubroca Link: https://lists.openwall.net/linux-kernel/2026/08/19/552 Signed-off-by: Haseeb Malik Reviewed-by: Sabrina Dubroca Link: https://patch.msgid.link/20260921-fix-macsec-net-v3-1-accf94f93f5e@gmail.com Signed-off-by: Paolo Abeni --- drivers/net/macsec.c | 84 +++++++++++++++++++++++--------------------- 1 file changed, 43 insertions(+), 41 deletions(-) diff --git a/drivers/net/macsec.c b/drivers/net/macsec.c index 6f9f3aceffaa..78a19b134632 100644 --- a/drivers/net/macsec.c +++ b/drivers/net/macsec.c @@ -3539,6 +3539,22 @@ static int macsec_dev_init(struct net_device *dev) if (err) return err; + err = -ENOMEM; + macsec->stats = netdev_alloc_pcpu_stats(struct pcpu_secy_stats); + if (!macsec->stats) + goto destroy_gro_cells; + + macsec->secy.tx_sc.stats = + netdev_alloc_pcpu_stats(struct pcpu_tx_sc_stats); + if (!macsec->secy.tx_sc.stats) + goto free_secy_stats; + + macsec->secy.tx_sc.md_dst = metadata_dst_alloc(0, METADATA_MACSEC, + GFP_KERNEL); + if (!macsec->secy.tx_sc.md_dst) + goto free_tx_sc_stats; + macsec->secy.tx_sc.md_dst->u.macsec_info.sci = macsec->secy.sci; + macsec_inherit_tso_max(dev); dev->hw_features = real_dev->hw_features & MACSEC_OFFLOAD_FEATURES; @@ -3551,8 +3567,6 @@ static int macsec_dev_init(struct net_device *dev) macsec_set_head_tail_room(dev); - if (is_zero_ether_addr(dev->dev_addr)) - eth_hw_addr_inherit(dev, real_dev); if (is_zero_ether_addr(dev->broadcast)) memcpy(dev->broadcast, real_dev->broadcast, dev->addr_len); @@ -3560,6 +3574,14 @@ static int macsec_dev_init(struct net_device *dev) netdev_hold(real_dev, &macsec->dev_tracker, GFP_KERNEL); return 0; + +free_tx_sc_stats: + free_percpu(macsec->secy.tx_sc.stats); +free_secy_stats: + free_percpu(macsec->stats); +destroy_gro_cells: + gro_cells_destroy(&macsec->gro_cells); + return err; } static void macsec_dev_uninit(struct net_device *dev) @@ -4116,26 +4138,11 @@ static sci_t dev_to_sci(struct net_device *dev, __be16 port) return make_sci(dev->dev_addr, port); } -static int macsec_add_dev(struct net_device *dev, sci_t sci, u8 icv_len) +static void macsec_init_secy(struct net_device *dev, sci_t sci, u8 icv_len) { struct macsec_dev *macsec = macsec_priv(dev); struct macsec_secy *secy = &macsec->secy; - macsec->stats = netdev_alloc_pcpu_stats(struct pcpu_secy_stats); - if (!macsec->stats) - return -ENOMEM; - - secy->tx_sc.stats = netdev_alloc_pcpu_stats(struct pcpu_tx_sc_stats); - if (!secy->tx_sc.stats) - return -ENOMEM; - - secy->tx_sc.md_dst = metadata_dst_alloc(0, METADATA_MACSEC, GFP_KERNEL); - if (!secy->tx_sc.md_dst) - /* macsec and secy percpu stats will be freed when unregistering - * net_device in macsec_free_netdev() - */ - return -ENOMEM; - if (sci == MACSEC_UNDEF_SCI) sci = dev_to_sci(dev, MACSEC_PORT_ES); @@ -4149,15 +4156,12 @@ static int macsec_add_dev(struct net_device *dev, sci_t sci, u8 icv_len) secy->xpn = DEFAULT_XPN; secy->sci = sci; - secy->tx_sc.md_dst->u.macsec_info.sci = sci; secy->tx_sc.active = true; secy->tx_sc.encoding_sa = DEFAULT_ENCODING_SA; secy->tx_sc.encrypt = DEFAULT_ENCRYPT; secy->tx_sc.send_sci = DEFAULT_SEND_SCI; secy->tx_sc.end_station = false; secy->tx_sc.scb = false; - - return 0; } static struct lock_class_key macsec_netdev_addr_lock_key; @@ -4220,6 +4224,24 @@ static int macsec_newlink(struct net_device *dev, if (rx_handler && rx_handler != macsec_handle_frame) return -EBUSY; + if (is_zero_ether_addr(dev->dev_addr)) + eth_hw_addr_inherit(dev, real_dev); + + if (data && data[IFLA_MACSEC_SCI]) + sci = nla_get_sci(data[IFLA_MACSEC_SCI]); + else if (data && data[IFLA_MACSEC_PORT]) + sci = dev_to_sci(dev, nla_get_be16(data[IFLA_MACSEC_PORT])); + else + sci = dev_to_sci(dev, MACSEC_PORT_ES); + + /* Registration can notify listeners before returning. */ + macsec_init_secy(dev, sci, icv_len); + if (data) { + err = macsec_changelink_common(dev, data); + if (err) + return err; + } + err = register_netdevice(dev); if (err < 0) return err; @@ -4232,31 +4254,11 @@ static int macsec_newlink(struct net_device *dev, if (err < 0) goto unregister; - /* need to be already registered so that ->init has run and - * the MAC addr is set - */ - if (data && data[IFLA_MACSEC_SCI]) - sci = nla_get_sci(data[IFLA_MACSEC_SCI]); - else if (data && data[IFLA_MACSEC_PORT]) - sci = dev_to_sci(dev, nla_get_be16(data[IFLA_MACSEC_PORT])); - else - sci = dev_to_sci(dev, MACSEC_PORT_ES); - if (rx_handler && sci_exists(real_dev, sci)) { err = -EBUSY; goto unlink; } - err = macsec_add_dev(dev, sci, icv_len); - if (err) - goto unlink; - - if (data) { - err = macsec_changelink_common(dev, data); - if (err) - goto del_dev; - } - /* If h/w offloading is available, propagate to the device */ if (macsec_is_offloaded(macsec)) { const struct macsec_ops *ops; From d6ec384c87cc851cfd13bb18c99ce351ccee6192 Mon Sep 17 00:00:00 2001 From: Fang Xieyan Date: Mon, 21 Sep 2026 20:54:41 +0800 Subject: [PATCH 156/189] net/sched: act_ife: validate metadata length before decoding skbmark_decode(), skbprio_decode() and skbtcindex_decode() read fixed-size values from the TLV payload without validating its length. A malformed IFE frame can declare a shorter payload, causing the decoders to consume bytes beyond the declared metadata value: [TLV type=IFE_META_SKBMARK len=4] -> dlen == 0, but decode reads 4 bytes The decoder may therefore set skb metadata from unintended input. Validate the payload length before decoding and return -EINVAL for invalid lengths. Read the values with get_unaligned_be32() and get_unaligned_be16(), as TLV payloads are not guaranteed to be aligned. Teach tcf_ife_decode() to log a decoder error separately from an unknown metaid; both are counted as overlimits and decoding continues with the remaining metadata. The metadata length issue was found by an automated audit of the IFE decode path at v6.18-rc7 and reproduced with a userspace sanitizer model of the decode path. Compile-tested on x86_64 with defconfig and NET_ACT_IFE=y: act_ife.o and the three act_meta_*.o build warning-free. Fixes: 084e2f6566d2 ("Support to encoding decoding skb mark on IFE action") Fixes: 200e10f46936 ("Support to encoding decoding skb prio on IFE action") Fixes: 408fbc22ef1e ("net sched ife action: Introduce skb tcindex metadata encap decap") Assisted-by: Hawkeye:GLM-5.3-flash Assisted-by: Qoder:Qwen3.8-Max Signed-off-by: Fang Xieyan Link: https://patch.msgid.link/20260921125441.81459-1-fangxy@xiaopeng.com Signed-off-by: Paolo Abeni --- net/sched/act_ife.c | 17 ++++++++++++----- net/sched/act_meta_mark.c | 6 ++++-- net/sched/act_meta_skbprio.c | 6 ++++-- net/sched/act_meta_skbtcindex.c | 6 ++++-- 4 files changed, 24 insertions(+), 11 deletions(-) diff --git a/net/sched/act_ife.c b/net/sched/act_ife.c index 9cea71fc1db3..2afd68983ece 100644 --- a/net/sched/act_ife.c +++ b/net/sched/act_ife.c @@ -737,6 +737,7 @@ static int tcf_ife_decode(struct sk_buff *skb, const struct tc_action *a, u8 *curr_data; u16 mtype; u16 dlen; + int ret; curr_data = ife_tlv_meta_decode(tlv_data, ifehdr_end, &mtype, &dlen, NULL); @@ -745,13 +746,19 @@ static int tcf_ife_decode(struct sk_buff *skb, const struct tc_action *a, return TC_ACT_SHOT; } - if (find_decode_metaid(skb, p, mtype, dlen, curr_data)) { - /* abuse overlimits to count when we receive metadata - * but dont have an ops for it + ret = find_decode_metaid(skb, p, mtype, dlen, curr_data); + if (ret < 0) { + /* abuse overlimits to count metadata we cannot + * decode: no ops for it, or the decoder rejected it */ - pr_info_ratelimited("Unknown metaid %d dlen %d\n", - mtype, dlen); qstats_cpu_overlimit_inc(ife->common.cpu_qstats); + + if (ret == -ENOENT) + pr_info_ratelimited("Unknown metaid %d dlen %d\n", + mtype, dlen); + else + pr_info_ratelimited("Failed to decode metaid %d dlen %d err %d\n", + mtype, dlen, ret); } } diff --git a/net/sched/act_meta_mark.c b/net/sched/act_meta_mark.c index ea0573cb8b2d..e2f61b22bf0f 100644 --- a/net/sched/act_meta_mark.c +++ b/net/sched/act_meta_mark.c @@ -10,6 +10,7 @@ #include #include #include +#include #include #include #include @@ -28,9 +29,10 @@ static int skbmark_encode(struct sk_buff *skb, void *skbdata, static int skbmark_decode(struct sk_buff *skb, void *data, u16 len) { - u32 ifemark = *(u32 *)data; + if (len != sizeof(u32)) + return -EINVAL; - skb->mark = ntohl(ifemark); + skb->mark = get_unaligned_be32(data); return 0; } diff --git a/net/sched/act_meta_skbprio.c b/net/sched/act_meta_skbprio.c index 2df3133ce5ad..5cdb57931eab 100644 --- a/net/sched/act_meta_skbprio.c +++ b/net/sched/act_meta_skbprio.c @@ -10,6 +10,7 @@ #include #include #include +#include #include #include #include @@ -33,9 +34,10 @@ static int skbprio_encode(struct sk_buff *skb, void *skbdata, static int skbprio_decode(struct sk_buff *skb, void *data, u16 len) { - u32 ifeprio = *(u32 *)data; + if (len != sizeof(u32)) + return -EINVAL; - skb->priority = ntohl(ifeprio); + skb->priority = get_unaligned_be32(data); return 0; } diff --git a/net/sched/act_meta_skbtcindex.c b/net/sched/act_meta_skbtcindex.c index 44547caead46..8803710c0905 100644 --- a/net/sched/act_meta_skbtcindex.c +++ b/net/sched/act_meta_skbtcindex.c @@ -10,6 +10,7 @@ #include #include #include +#include #include #include #include @@ -28,9 +29,10 @@ static int skbtcindex_encode(struct sk_buff *skb, void *skbdata, static int skbtcindex_decode(struct sk_buff *skb, void *data, u16 len) { - u16 ifetc_index = *(u16 *)data; + if (len != sizeof(u16)) + return -EINVAL; - skb->tc_index = ntohs(ifetc_index); + skb->tc_index = get_unaligned_be16(data); return 0; } From 7c9f391ec89cb621d7af375ace2eb9a6248e5b9d Mon Sep 17 00:00:00 2001 From: Christian Lamparter Date: Mon, 21 Sep 2026 18:28:17 +0200 Subject: [PATCH 157/189] net: emac: move setting of netops to fix crash fixes the following crash on driver initialization: |BUG: Kernel NULL pointer dereference on read at 0x00000158 |Faulting instruction address: 0xc0566b40 |Oops: Kernel access of bad area, sig: 11 [#1] |BE PAGE_SIZE=4K PowerPC 44x Platform |Modules linked in: |CPU: 0 UID: 0 PID: 1 Comm: swapper/0 Tainted: GW 7.3.0-rc3+ #1 |Tainted: [W]=WARN |Hardware name: MyBook Live APM821XX 0x12c41c83 PowerPC 44x Platform |NIP: c0566b40 LR: c05648f8 CTR: c04c1e9c |REGS: c1053a20 TRAP: 0300 Tainted: GW (7.3.0-rc3+) |MSR: 0002b000 CR: 24008808 XER: 00000000 |DEAR: 00000158 ESR: 00000000 |GPR00: c05648f8 c1053b10 c1063600 c1030000 c5ab3000 00000000 [...] |GPR08: 00000002 00000000 00000000 c1053b40 84002808 00000000 [...] |GPR16: cfffd210 00000002 c0beafcc cfffc960 00000000 c1030644 [...] |GPR24: c0beafbc c1037000 00000000 0000000a 00000000 c1030000 [...] |NIP [c0566b40] phy_link_topo_add_phy+0x2c/0x1d0 |LR [c05648f8] phy_attach_direct+0x1a4/0x368 |Call Trace: |[c1053b10] [c0811e04] klist_put+0x54/0xb4 (unreliable) |[c1053b40] [c05648f8] phy_attach_direct+0x1a4/0x368 |[c1053b70] [c0564ae8] phy_connect_direct+0x2c/0x60 |[c1053b90] [c056c7a8] of_phy_connect+0x50/0x74 |[c1053bc0] [c0572a88] emac_probe+0xd50/0x119c |[c1053c90] [c04cb770] platform_probe+0x74/0xa4 |[c1053cb0] [c04c8f08] really_probe+0x120/0x2b0 |[c1053cd0] [c04c9254] __driver_probe_device+0x1bc/0x1fc |[c1053d00] [c04c9334] driver_probe_device+0x38/0xa8 |[c1053d30] [c04c9560] __driver_attach+0xf4/0x10c |[c1053d50] [c04c6af8] bus_for_each_dev+0x68/0xd0 |[c1053d90] [c04c7c34] bus_add_driver+0xcc/0x1ec |[c1053dc0] [c04c9f6c] driver_register+0xcc/0x110 |[c1053de0] [c0aa3f40] emac_init+0x1c4/0x200 This bug showed up starting with v7.3-rc1. At the NIP in phy_link_topo_add_phy() is a netdev_need_ops_lock() check. This was added by the following commit ded86da4bbb7 ("net: ethtool: relax ethnl_req_get_phydev() locking assertion") The bug shows up because at the time of_phy_connect() was called, the netdev_ops were *not yet* determined. My fix is to move the code that sets netdev_ops+commac.ops+ethtool_ops further up as emac_init_config() derives that by looking at the device-tree and sets the required dev->phy_mode accordingly. During review, the Sashiko bot's AI stated that the commac assignment became a dead store. Great catch! To keep the original behavior as-is, one mentioned option "set dev->commac.ops = &emac_commac_ops only in the non-gige case?" sounded like a great plan. So the dev->commac.ops assignment for the non-gige-case moves into the else block. This patch was tested on a WD MyBook Live (RGMII). the device now works again. Fixes: ded86da4bbb7 ("net: ethtool: relax ethnl_req_get_phydev() locking assertion") Signed-off-by: Christian Lamparter Reviewed-by: Nicolai Buchwitz Reviewed-by: Maxime Chevallier Link: https://patch.msgid.link/49cd7343bc0e507c022071f3e2b5662b053dca73.1790007431.git.chunkeey@gmail.com Signed-off-by: Paolo Abeni --- drivers/net/ethernet/ibm/emac/core.c | 16 +++++++++------- 1 file changed, 9 insertions(+), 7 deletions(-) diff --git a/drivers/net/ethernet/ibm/emac/core.c b/drivers/net/ethernet/ibm/emac/core.c index 1d46cf6c2c12..e7043523457c 100644 --- a/drivers/net/ethernet/ibm/emac/core.c +++ b/drivers/net/ethernet/ibm/emac/core.c @@ -3044,6 +3044,15 @@ static int emac_probe(struct platform_device *ofdev) if (err) goto err_gone; + if (emac_phy_supports_gige(dev->phy_mode)) { + ndev->netdev_ops = &emac_gige_netdev_ops; + dev->commac.ops = &emac_commac_sg_ops; + } else { + ndev->netdev_ops = &emac_netdev_ops; + dev->commac.ops = &emac_commac_ops; + } + ndev->ethtool_ops = &emac_ethtool_ops; + dev->emacp = devm_platform_ioremap_resource(ofdev, 0); if (IS_ERR(dev->emacp)) { err = PTR_ERR(dev->emacp); @@ -3076,7 +3085,6 @@ static int emac_probe(struct platform_device *ofdev) dev->mdio_instance = platform_get_drvdata(dev->mdio_dev); /* Register with MAL */ - dev->commac.ops = &emac_commac_ops; dev->commac.dev = dev; dev->commac.tx_chan_mask = MAL_CHAN_MASK(dev->mal_tx_chan); dev->commac.rx_chan_mask = MAL_CHAN_MASK(dev->mal_rx_chan); @@ -3144,12 +3152,6 @@ static int emac_probe(struct platform_device *ofdev) ndev->features |= ndev->hw_features | NETIF_F_RXCSUM; } ndev->watchdog_timeo = 5 * HZ; - if (emac_phy_supports_gige(dev->phy_mode)) { - ndev->netdev_ops = &emac_gige_netdev_ops; - dev->commac.ops = &emac_commac_sg_ops; - } else - ndev->netdev_ops = &emac_netdev_ops; - ndev->ethtool_ops = &emac_ethtool_ops; /* MTU range: 46 - 1500 or whatever is in OF */ ndev->min_mtu = EMAC_MIN_MTU; From c2cdef41e0b4d8ed23a5b41e6ad4e64594e055e4 Mon Sep 17 00:00:00 2001 From: Aleksei Sviridkin Date: Fri, 18 Sep 2026 04:50:19 +0300 Subject: [PATCH 158/189] net: dsa: mt7530: fix NULL dereference on unbind of MT7531 and MT7621 The core and io supplies are only requested for ID_MT7530: both the devm_regulator_get() in probe and the regulator_enable() in mt7530_setup() are guarded by the switch id, but mt7530_remove() disables them unconditionally. On an MT7621 or an MT7531 both pointers are still NULL from devm_kzalloc(), so rmmod or a sysfs unbind calls regulator_disable() on NULL. Fixes: ddda1ac116c8 ("net: dsa: mt7530: support the 7530 switch on the Mediatek MT7621 SoC") Signed-off-by: Aleksei Sviridkin Link: https://patch.msgid.link/20260918015020.2518315-2-f@lex.la Signed-off-by: Jakub Kicinski --- drivers/net/dsa/mt7530-mdio.c | 18 ++++++++++-------- 1 file changed, 10 insertions(+), 8 deletions(-) diff --git a/drivers/net/dsa/mt7530-mdio.c b/drivers/net/dsa/mt7530-mdio.c index 784dd58a7158..de42f70afcfa 100644 --- a/drivers/net/dsa/mt7530-mdio.c +++ b/drivers/net/dsa/mt7530-mdio.c @@ -227,15 +227,17 @@ mt7530_remove(struct mdio_device *mdiodev) if (!priv) return; - ret = regulator_disable(priv->core_pwr); - if (ret < 0) - dev_err(priv->dev, - "Failed to disable core power: %d\n", ret); + if (priv->id == ID_MT7530) { + ret = regulator_disable(priv->core_pwr); + if (ret < 0) + dev_err(priv->dev, + "Failed to disable core power: %d\n", ret); - ret = regulator_disable(priv->io_pwr); - if (ret < 0) - dev_err(priv->dev, "Failed to disable io pwr: %d\n", - ret); + ret = regulator_disable(priv->io_pwr); + if (ret < 0) + dev_err(priv->dev, "Failed to disable io pwr: %d\n", + ret); + } mt7530_remove_common(priv); From 0d80ba0a204c6a16bd7778b50de578dff107c0fe Mon Sep 17 00:00:00 2001 From: Aleksei Sviridkin Date: Fri, 18 Sep 2026 04:50:20 +0300 Subject: [PATCH 159/189] net: dsa: mt7530: leave the MDIO IRQ mappings to regmap-irq mt7530_remove_common() disposes the per-PHY interrupt mappings from .remove, but the regmap-irq chip that owns the domain is devm-registered, so its parent interrupt is only freed once .remove has returned. The switch's own regmap-irq thread can therefore still dispatch on a mapping that is already gone: irq_find_mapping() returns 0, irq_to_desc() returns NULL and handle_nested_irq() locks desc->lock without checking it. The attached PHYs have not given those interrupts back yet either, which the kernel warns about a moment before the fault. regmap_del_irq_chip() disposes the same mappings itself, after freeing the parent interrupt and before removing the domain, so there is nothing left for the driver to do here. Until it runs the descriptors stay alive, and a late dispatch on one of them is harmless: dsa_unregister_switch() has freed the PHY handlers by then, so handle_nested_irq() finds no action and returns. Fixes: 254f6b272e3b ("dsa: mt7530: Utilize REGMAP_IRQ for interrupt handling") Signed-off-by: Aleksei Sviridkin Link: https://patch.msgid.link/20260918015020.2518315-3-f@lex.la Signed-off-by: Jakub Kicinski --- drivers/net/dsa/mt7530.c | 3 --- 1 file changed, 3 deletions(-) diff --git a/drivers/net/dsa/mt7530.c b/drivers/net/dsa/mt7530.c index 3e61eb3c2b1e..96832852c65a 100644 --- a/drivers/net/dsa/mt7530.c +++ b/drivers/net/dsa/mt7530.c @@ -3593,9 +3593,6 @@ EXPORT_SYMBOL_GPL(mt7530_probe_common); void mt7530_remove_common(struct mt7530_priv *priv) { - if (priv->irq_domain) - mt7530_free_mdio_irq(priv); - dsa_unregister_switch(priv->ds); mutex_destroy(&priv->reg_mutex); From 26cc0e69cce062cd3aa6fae33074684669c35a71 Mon Sep 17 00:00:00 2001 From: Hui Peng Date: Mon, 21 Sep 2026 05:10:01 +0000 Subject: [PATCH 160/189] mctp: route: iterate socket tag list in mctp_lookup_prealloc_tag() When a socket transmits a packet with MCTP_TAG_PREALLOC set, mctp_lookup_prealloc_tag() iterates over the per-netns &mns->keys list and matches netid, req_tag, peer_addr, and manual_alloc, without checking whether tmp->sk == &msk->sk. This allows any MCTP socket in the same network namespace to use and consume another socket's preallocated tag. Iterate the socket's own tag list (&msk->keys via sklist) instead of the namespace-wide &mns->keys list in mctp_lookup_prealloc_tag(), ensuring that only tags allocated by msk are matched. Tested in QEMU against Linux 7.3.0-rc3 by allocating a manual tag (0x18) on socket A via SIOCMCTPALLOCTAG for peer EID 9 and sending a 4-byte message with MCTP_TAG_PREALLOC from socket B in the same network namespace. On the unfixed kernel, sendto(sock_b) using socket A's preallocated tag succeeds (ret = 4); with this patch applied, sendto(sock_b) fails with -ENOENT (errno = 2) while sendto(sock_a) succeeds (ret = 4). Fixes: 63ed1aab3d40 ("mctp: Add SIOCMCTP{ALLOC,DROP}TAG ioctls for tag control") Suggested-by: Jeremy Kerr Cc: stable@vger.kernel.org Signed-off-by: Hui Peng Link: https://patch.msgid.link/20260921051002.1656692-1-benquike@gmail.com Signed-off-by: Jakub Kicinski --- net/mctp/route.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/net/mctp/route.c b/net/mctp/route.c index b19c63a5691a..e7c95eeacb48 100644 --- a/net/mctp/route.c +++ b/net/mctp/route.c @@ -825,7 +825,7 @@ static struct mctp_sk_key *mctp_lookup_prealloc_tag(struct mctp_sock *msk, spin_lock_irqsave(&mns->keys_lock, flags); - hlist_for_each_entry(tmp, &mns->keys, hlist) { + hlist_for_each_entry(tmp, &msk->keys, sklist) { if (tmp->net != netid) continue; From fe99bbeee5c5dbd3abc30721a8079ced59649d97 Mon Sep 17 00:00:00 2001 From: Yilin Zhang Date: Thu, 24 Sep 2026 12:49:00 +0800 Subject: [PATCH 161/189] tcp: fix use-after-free of retransmit_skb_hint in tcp_send_synack() When tcp_send_synack() replaces the cloned SYN skb at the head of the retransmit queue with a copy, it frees the original with tcp_rtx_queue_unlink_and_free() and only repairs tp->highest_sack. tp->retransmit_skb_hint keeps pointing at the freed skbuff_fclone_cache object. The dangling hint is read in tcp_verify_retransmit_hint() and used as the root of the rbtree walk in tcp_xmit_retransmit_queue(). An unprivileged TFO client (sendmsg(MSG_FASTOPEN)) can arm the hint with an attacker-supplied ICMP fragmentation-needed message, after which a simultaneous open frees the armed SYN skb: BUG: KASAN: slab-use-after-free in tcp_mark_skb_lost (net/ipv4/tcp_input.c:1316) Read of size 4 at addr ffff88800604d928 by task swapper/1/0 Call Trace: tcp_mark_skb_lost (net/ipv4/tcp_input.c:1316) tcp_simple_retransmit (net/ipv4/tcp_input.c:3158) tcp_v4_err (net/ipv4/tcp_ipv4.c:587) Sync the hint to the copy. Fixes: c31b70c9968f ("tcp: Add logic to check for SYN w/ data in tcp_simple_retransmit") Reported-by: Kimi Security Team Tested-by: Weiming Shi Signed-off-by: Yilin Zhang Reviewed-by: Eric Dumazet Link: https://patch.msgid.link/8a9dff4063a2745653b7e88ceb745d75efa16e68.1790224474.git.yilinzhang@moonshot.ai Signed-off-by: Jakub Kicinski --- net/ipv4/tcp_output.c | 3 +++ 1 file changed, 3 insertions(+) diff --git a/net/ipv4/tcp_output.c b/net/ipv4/tcp_output.c index d960e3de7d50..e0c392e29de5 100644 --- a/net/ipv4/tcp_output.c +++ b/net/ipv4/tcp_output.c @@ -3886,6 +3886,7 @@ void tcp_send_active_reset(struct sock *sk, enum sk_rst_reason reason) */ int tcp_send_synack(struct sock *sk) { + struct tcp_sock *tp = tcp_sk(sk); struct sk_buff *skb; skb = tcp_rtx_queue_head(sk); @@ -3903,6 +3904,8 @@ int tcp_send_synack(struct sock *sk) if (!nskb) return -ENOMEM; INIT_LIST_HEAD(&nskb->tcp_tsorted_anchor); + if (skb == tp->retransmit_skb_hint) + tp->retransmit_skb_hint = nskb; tcp_highest_sack_replace(sk, skb, nskb); tcp_rtx_queue_unlink_and_free(skb, sk); __skb_header_release(nskb); From f75f21ef36285e5f56ee0c428bd2909ee81165b9 Mon Sep 17 00:00:00 2001 From: Ming Wang Date: Sun, 20 Sep 2026 15:44:59 +0800 Subject: [PATCH 162/189] net: usb: cdc_mbim: add MeiG Smart SRM821 to ZLP whitelist The MeiG Smart SRM821 5G module (0x2dee:0x4d53) crashes and drops off the USB bus when it receives a Zero Length Packet (ZLP) after sending or receiving an NTB of exactly 16384 bytes (tx_max). According to the MBIM specification, devices do not require a ZLP if the NTB size is exactly dwNtbOutMaxSize. However, the cdc_mbim driver defaults to sending ZLPs for devices not explicitly whitelisted to accommodate non-conformant hardware. This default behavior breaks the strictly conformant MeiG SRM821 module. Add this device to the ZLP conformance whitelist (cdc_mbim_info) so the driver will pad the NTB to avoid sending ZLPs, preventing the device firmware from crashing. Cc: stable@vger.kernel.org Signed-off-by: Ming Wang Link: https://patch.msgid.link/20260920074500.826121-1-wangming01@loongson.cn Signed-off-by: Jakub Kicinski --- drivers/net/usb/cdc_mbim.c | 5 +++++ 1 file changed, 5 insertions(+) diff --git a/drivers/net/usb/cdc_mbim.c b/drivers/net/usb/cdc_mbim.c index 877fb0ed7d3d..a7010a0664c7 100644 --- a/drivers/net/usb/cdc_mbim.c +++ b/drivers/net/usb/cdc_mbim.c @@ -635,6 +635,11 @@ static const struct usb_device_id mbim_devs[] = { .driver_info = (unsigned long)&cdc_mbim_info, }, + /* MeiG Smart SRM821 ZLP conformance */ + { USB_DEVICE_AND_INTERFACE_INFO(0x2dee, 0x4d53, USB_CLASS_COMM, USB_CDC_SUBCLASS_MBIM, USB_CDC_PROTO_NONE), + .driver_info = (unsigned long)&cdc_mbim_info, + }, + /* Some Huawei devices, ME906s-158 (12d1:15c1) and E3372 * (12d1:157d), are known to fail unless the NDP is placed * after the IP packets. Applying the quirk to all Huawei From a940003f44e7e441c228151dd212642152700ec8 Mon Sep 17 00:00:00 2001 From: Aleksei Sviridkin Date: Mon, 21 Sep 2026 01:20:44 +0300 Subject: [PATCH 163/189] net: phylink: record the PHY only once bringup cannot fail phylink_bringup_phy() stores the PHY in pl->phydev before its last fallible step: on a MAC whose phylink ops implement LPI, phy_eee_rx_clock_stop() can fail with a real MDIO error. The callers unwind with phy_detach(), which knows nothing about pl->phydev, so a pointer to a PHY that is no longer attached outlives the failed connect. What that costs depends on how the caller got here. phylink_connect_phy() goes through phylink_attach_phy(), which refuses to attach while pl->phydev is set, turning a transient MDIO error into a permanent -EBUSY. The SFP path is worse than that: sfp_sm_probe_phy() answers the failure with phy_device_remove() and phy_device_free(), and it assigns sfp->mod_phy only past that error return, so nothing clears pl->phydev and it is left pointing at a freed phy_device that phylink_resolve() and the ethtool helpers go on reading. phylink_fwnode_phy_connect() has no such check, so a later connect overwrites the stale pointer and hides the problem. A disconnect does not: phylink_disconnect_phy() hands that pointer to phy_disconnect(), and the second phy_detach() on the same PHY drops references the first one already released. Found while making a DSA port survive a PHY whose driver arrives after the switch probes: keeping the port across a failed connect and retrying is what makes this window reachable. Publish the pointer after the last call that can fail instead of unwinding it afterwards. Nothing between the two points reads pl->phydev, and the registration that follows cannot fail: phy_request_interrupt() falls back to polling on its own. The PHY-side state keeps the order it had, so no MDIO operation moves relative to another. Fixes: 03abf2a7c654 ("net: phylink: add EEE management") Signed-off-by: Aleksei Sviridkin Link: https://patch.msgid.link/20260920222044.1752860-1-f@lex.la Signed-off-by: Jakub Kicinski --- drivers/net/phy/phylink.c | 20 +++++++++++++++++--- 1 file changed, 17 insertions(+), 3 deletions(-) diff --git a/drivers/net/phy/phylink.c b/drivers/net/phy/phylink.c index a1458da8111b..1bbcf46c8356 100644 --- a/drivers/net/phy/phylink.c +++ b/drivers/net/phy/phylink.c @@ -2129,7 +2129,6 @@ static int phylink_bringup_phy(struct phylink *pl, struct phy_device *phy, mutex_lock(&pl->phydev_mutex); mutex_lock(&phy->lock); mutex_lock(&pl->state_mutex); - pl->phydev = phy; pl->phy_state.interface = interface; pl->phy_state.pause = MLO_PAUSE_NONE; pl->phy_state.speed = SPEED_UNKNOWN; @@ -2196,10 +2195,25 @@ static int phylink_bringup_phy(struct phylink *pl, struct phy_device *phy, ret = 0; } - if (ret == 0 && phy_interrupt_is_valid(phy)) + if (ret) + return ret; + + /* Nothing below can fail, so the PHY can be recorded now. Doing it + * here rather than above keeps a failed bringup from leaving + * pl->phydev pointing at a PHY the caller is about to detach. + */ + mutex_lock(&pl->phydev_mutex); + mutex_lock(&phy->lock); + mutex_lock(&pl->state_mutex); + pl->phydev = phy; + mutex_unlock(&pl->state_mutex); + mutex_unlock(&phy->lock); + mutex_unlock(&pl->phydev_mutex); + + if (phy_interrupt_is_valid(phy)) phy_request_interrupt(phy); - return ret; + return 0; } static int phylink_attach_phy(struct phylink *pl, struct phy_device *phy, From 3173cba1170131816972ed2b6185cb970ba8b747 Mon Sep 17 00:00:00 2001 From: Jiawen Wu Date: Mon, 21 Sep 2026 15:15:49 +0800 Subject: [PATCH 164/189] net: libwx: fix races in Tx timestamp handling wx->ptp_tx_skb is shared between the Tx path, the PTP auxiliary worker and the timestamp cleanup paths. The WX_STATE_PTP_TX_IN_PROGRESS bit prevents multiple Tx paths from submitting timestamp requests, but does not serialize the worker against cleanup. As a result, wx_ptp_clear_tx_timestamp() can free an skb after wx_ptp_tx_hwtstamp_work() has obtained its pointer. The worker may then pass the freed skb to skb_tstamp_tx() and release the same reference again. The cleanup path may also clear the in-progress bit while the worker is still processing the old skb. This allows the Tx path to publish a new skb which the worker can subsequently overwrite with NULL, leaking its reference. Add a dedicated spinlock to protect publication and consumption of the Tx timestamp skb. Detach the skb and clear the in-progress bit while holding the lock, then deliver the timestamp and release the skb after dropping it. Use the same locked cleanup in the quiesce path, but keep the detach there free of register accesses: quiesce runs during PCIe error recovery, where MMIO is not reliable, and it deliberately did not touch the device before. The lock is taken with interrupts disabled, because netpoll can call ndo_start_xmit() with hard interrupts already off. When handling a Tx DMA mapping failure, keep the transmit path reference until after comparing the skb under the lock. This prevents skb address reuse from making the error path mistake a newer timestamp request for the failed one. Fixes: 06e75161b9d4 ("net: wangxun: Add support for PTP clock") Reported-by: Sashiko Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/6C7EC12D69217315%2B20260818074721.45536-1-jiawenwu%40trustnetic.com Signed-off-by: Jiawen Wu Link: https://patch.msgid.link/77431AF9A0E369F3+20260921071549.1141804-1-jiawenwu@trustnetic.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/wangxun/libwx/wx_hw.c | 1 + drivers/net/ethernet/wangxun/libwx/wx_lib.c | 47 ++++-- drivers/net/ethernet/wangxun/libwx/wx_ptp.c | 144 ++++++++++++------- drivers/net/ethernet/wangxun/libwx/wx_type.h | 2 + 4 files changed, 133 insertions(+), 61 deletions(-) diff --git a/drivers/net/ethernet/wangxun/libwx/wx_hw.c b/drivers/net/ethernet/wangxun/libwx/wx_hw.c index 122c4952d203..113552586be7 100644 --- a/drivers/net/ethernet/wangxun/libwx/wx_hw.c +++ b/drivers/net/ethernet/wangxun/libwx/wx_hw.c @@ -2518,6 +2518,7 @@ int wx_sw_init(struct wx *wx) } spin_lock_init(&wx->hw_stats_lock); + spin_lock_init(&wx->ptp_tx_lock); mutex_init(&wx->reset_lock); bitmap_zero(wx->state, WX_STATE_NBITS); bitmap_zero(wx->flags, WX_PF_FLAGS_NBITS); diff --git a/drivers/net/ethernet/wangxun/libwx/wx_lib.c b/drivers/net/ethernet/wangxun/libwx/wx_lib.c index ed5aad7857bd..dcbf5811046e 100644 --- a/drivers/net/ethernet/wangxun/libwx/wx_lib.c +++ b/drivers/net/ethernet/wangxun/libwx/wx_lib.c @@ -1200,9 +1200,11 @@ static int wx_tx_map(struct wx_ring *tx_ring, i--; } - dev_kfree_skb_any(first->skb); - first->skb = NULL; - + /* first->skb is released by the caller, which keeps a reference on it + * until the PTP cleanup has compared it against wx->ptp_tx_skb. That + * prevents the address from being reused by a newer request while the + * comparison is pending. + */ tx_ring->next_to_use = i; return -ENOMEM; @@ -1649,9 +1651,11 @@ static netdev_tx_t wx_xmit_frame_ring(struct sk_buff *skb, if (unlikely(skb_shinfo(skb)->tx_flags & SKBTX_HW_TSTAMP) && wx->ptp_clock) { + unsigned long flags; + + spin_lock_irqsave(&wx->ptp_tx_lock, flags); if (wx->tstamp_config.tx_type == HWTSTAMP_TX_ON && - !test_and_set_bit_lock(WX_STATE_PTP_TX_IN_PROGRESS, - wx->state)) { + !test_and_set_bit(WX_STATE_PTP_TX_IN_PROGRESS, wx->state)) { skb_shinfo(skb)->tx_flags |= SKBTX_IN_PROGRESS; tx_flags |= WX_TX_FLAGS_TSTAMP; wx->ptp_tx_skb = skb_get(skb); @@ -1659,6 +1663,7 @@ static netdev_tx_t wx_xmit_frame_ring(struct sk_buff *skb, } else { wx->tx_hwtstamp_skipped++; } + spin_unlock_irqrestore(&wx->ptp_tx_lock, flags); } /* record initial flags and protocol */ @@ -1677,19 +1682,35 @@ static netdev_tx_t wx_xmit_frame_ring(struct sk_buff *skb, wx->atr(tx_ring, first, ptype); if (wx_tx_map(tx_ring, first, hdr_len)) - goto cleanup_tx_tstamp; + goto out_drop; return NETDEV_TX_OK; out_drop: + /* The frame never reached the hardware, so no timestamp will ever be + * reported for it and the request has to be cancelled. The slot is + * shared, though: wx_ptp_clear_tx_timestamp() or wx_ptp_tx_hang() may + * have dropped our request already, and a transmit on another queue + * can have claimed the slot since. Only cancel it while it is still + * ours, otherwise we would free somebody else's skb and release their + * in-progress bit. + */ + if (unlikely(tx_flags & WX_TX_FLAGS_TSTAMP)) { + struct sk_buff *ptp_tx_skb = NULL; + unsigned long flags; + + spin_lock_irqsave(&wx->ptp_tx_lock, flags); + if (wx->ptp_tx_skb == skb) { + ptp_tx_skb = wx->ptp_tx_skb; + wx->ptp_tx_skb = NULL; + clear_bit(WX_STATE_PTP_TX_IN_PROGRESS, wx->state); + wx->tx_hwtstamp_errors++; + } + spin_unlock_irqrestore(&wx->ptp_tx_lock, flags); + + dev_kfree_skb_any(ptp_tx_skb); + } dev_kfree_skb_any(first->skb); first->skb = NULL; -cleanup_tx_tstamp: - if (unlikely(tx_flags & WX_TX_FLAGS_TSTAMP)) { - dev_kfree_skb_any(wx->ptp_tx_skb); - wx->ptp_tx_skb = NULL; - wx->tx_hwtstamp_errors++; - clear_bit_unlock(WX_STATE_PTP_TX_IN_PROGRESS, wx->state); - } return NETDEV_TX_OK; } diff --git a/drivers/net/ethernet/wangxun/libwx/wx_ptp.c b/drivers/net/ethernet/wangxun/libwx/wx_ptp.c index 4708e7f3958f..65b8937f6e94 100644 --- a/drivers/net/ethernet/wangxun/libwx/wx_ptp.c +++ b/drivers/net/ethernet/wangxun/libwx/wx_ptp.c @@ -129,6 +129,34 @@ static int wx_ptp_settime64(struct ptp_clock_info *ptp, return 0; } +/** + * __wx_ptp_detach_tx_skb - detach the skb tracking the Tx timestamp request + * @wx: the private board structure + * + * Detach the skb of the outstanding request and release the in-progress bit, + * so that a new request can be submitted. + * + * This performs no register access. Callers that need a timestamp the hardware + * may have left latched must unlatch it themselves, while the device is known + * to be alive. wx_ptp_quiesce() runs during PCIe error recovery, where MMIO is + * not reliable, and therefore deliberately skips the unlatch. + * + * Context: Expects wx->ptp_tx_lock to be held by the caller. + * Return: the detached skb, or NULL if no request was outstanding. The caller + * owns the returned reference and must release it once the lock is dropped. + */ +static struct sk_buff *__wx_ptp_detach_tx_skb(struct wx *wx) +{ + struct sk_buff *skb = wx->ptp_tx_skb; + + lockdep_assert_held(&wx->ptp_tx_lock); + + wx->ptp_tx_skb = NULL; + clear_bit(WX_STATE_PTP_TX_IN_PROGRESS, wx->state); + + return skb; +} + /** * wx_ptp_clear_tx_timestamp - utility function to clear Tx timestamp state * @wx: the private board structure @@ -139,12 +167,16 @@ static int wx_ptp_settime64(struct ptp_clock_info *ptp, */ static void wx_ptp_clear_tx_timestamp(struct wx *wx) { + struct sk_buff *skb; + unsigned long flags; + + spin_lock_irqsave(&wx->ptp_tx_lock, flags); + /* Unlatch a timestamp the hardware may have left pending. */ rd32ptp(wx, WX_TSC_1588_STMPH); - if (wx->ptp_tx_skb) { - dev_kfree_skb_any(wx->ptp_tx_skb); - wx->ptp_tx_skb = NULL; - } - clear_bit_unlock(WX_STATE_PTP_TX_IN_PROGRESS, wx->state); + skb = __wx_ptp_detach_tx_skb(wx); + spin_unlock_irqrestore(&wx->ptp_tx_lock, flags); + + dev_kfree_skb_any(skb); } /** @@ -175,49 +207,54 @@ static void wx_ptp_convert_to_hwtstamp(struct wx *wx, } /** - * wx_ptp_tx_hwtstamp - utility function which checks for TX time stamp + * wx_ptp_tx_hwtstamp_work - check for a pending Tx time stamp * @wx: the private board struct * - * if the timestamp is valid, we convert it into the timecounter ns - * value, then store that result into the shhwtstamps structure which - * is passed up the network stack + * If a Tx timestamp request is outstanding and the hardware has latched a + * valid value, we convert it into the timecounter ns value, then store that + * result into the shhwtstamps structure which is passed up the network stack. + * + * Return: 0 when there is nothing left to poll for, -1 when the timestamp is + * not available yet and the caller should poll again. */ -static void wx_ptp_tx_hwtstamp(struct wx *wx) -{ - struct skb_shared_hwtstamps shhwtstamps; - struct sk_buff *skb = wx->ptp_tx_skb; - u64 regval = 0; - - regval |= (u64)rd32ptp(wx, WX_TSC_1588_STMPL); - regval |= (u64)rd32ptp(wx, WX_TSC_1588_STMPH) << 32; - - wx_ptp_convert_to_hwtstamp(wx, &shhwtstamps, regval); - - wx->ptp_tx_skb = NULL; - clear_bit_unlock(WX_STATE_PTP_TX_IN_PROGRESS, wx->state); - skb_tstamp_tx(skb, &shhwtstamps); - dev_kfree_skb_any(skb); - wx->tx_hwtstamp_pkts++; -} - static int wx_ptp_tx_hwtstamp_work(struct wx *wx) { + struct skb_shared_hwtstamps shhwtstamps; + unsigned long flags; + struct sk_buff *skb; u32 tsynctxctl; + u64 regval = 0; + + spin_lock_irqsave(&wx->ptp_tx_lock, flags); /* we have to have a valid skb to poll for a timestamp */ if (!wx->ptp_tx_skb) { - wx_ptp_clear_tx_timestamp(wx); + rd32ptp(wx, WX_TSC_1588_STMPH); + __wx_ptp_detach_tx_skb(wx); + spin_unlock_irqrestore(&wx->ptp_tx_lock, flags); return 0; } /* stop polling once we have a valid timestamp */ tsynctxctl = rd32ptp(wx, WX_TSC_1588_CTL); - if (tsynctxctl & WX_TSC_1588_CTL_VALID) { - wx_ptp_tx_hwtstamp(wx); - return 0; + if (!(tsynctxctl & WX_TSC_1588_CTL_VALID)) { + spin_unlock_irqrestore(&wx->ptp_tx_lock, flags); + return -1; } - return -1; + regval |= (u64)rd32ptp(wx, WX_TSC_1588_STMPL); + regval |= (u64)rd32ptp(wx, WX_TSC_1588_STMPH) << 32; + skb = wx->ptp_tx_skb; + wx->ptp_tx_skb = NULL; + clear_bit(WX_STATE_PTP_TX_IN_PROGRESS, wx->state); + spin_unlock_irqrestore(&wx->ptp_tx_lock, flags); + + wx_ptp_convert_to_hwtstamp(wx, &shhwtstamps, regval); + skb_tstamp_tx(skb, &shhwtstamps); + dev_kfree_skb_any(skb); + wx->tx_hwtstamp_pkts++; + + return 0; } /** @@ -296,24 +333,29 @@ static void wx_ptp_rx_hang(struct wx *wx) */ static void wx_ptp_tx_hang(struct wx *wx) { - bool timeout = time_is_before_jiffies(wx->ptp_tx_start + - WX_PTP_TX_TIMEOUT); + struct sk_buff *skb = NULL; + unsigned long flags; - if (!wx->ptp_tx_skb) - return; - - if (!test_bit(WX_STATE_PTP_TX_IN_PROGRESS, wx->state)) - return; + spin_lock_irqsave(&wx->ptp_tx_lock, flags); /* If we haven't received a timestamp within the timeout, it is * reasonable to assume that it will never occur, so we can unlock the * timestamp bit when this occurs. */ - if (timeout) { - wx_ptp_clear_tx_timestamp(wx); - wx->tx_hwtstamp_timeouts++; - dev_warn(&wx->pdev->dev, "clearing Tx timestamp hang\n"); + if (wx->ptp_tx_skb && + test_bit(WX_STATE_PTP_TX_IN_PROGRESS, wx->state) && + time_is_before_jiffies(wx->ptp_tx_start + WX_PTP_TX_TIMEOUT)) { + rd32ptp(wx, WX_TSC_1588_STMPH); + skb = __wx_ptp_detach_tx_skb(wx); } + spin_unlock_irqrestore(&wx->ptp_tx_lock, flags); + + if (!skb) + return; + + dev_kfree_skb_any(skb); + wx->tx_hwtstamp_timeouts++; + dev_warn(&wx->pdev->dev, "clearing Tx timestamp hang\n"); } static long wx_ptp_do_aux_work(struct ptp_clock_info *ptp) @@ -841,6 +883,9 @@ EXPORT_SYMBOL(wx_ptp_stop); void wx_ptp_quiesce(struct wx *wx) { + struct sk_buff *skb; + unsigned long flags; + if (!test_and_clear_bit(WX_STATE_PTP_RUNNING, wx->state)) return; @@ -849,11 +894,14 @@ void wx_ptp_quiesce(struct wx *wx) if (wx->ptp_clock) ptp_cancel_worker_sync(wx->ptp_clock); - if (wx->ptp_tx_skb) { - dev_kfree_skb_any(wx->ptp_tx_skb); - wx->ptp_tx_skb = NULL; - } - clear_bit_unlock(WX_STATE_PTP_TX_IN_PROGRESS, wx->state); + /* Drop a pending Tx timestamp request. Do not touch the registers + * here: quiesce runs during PCIe error recovery, where the device may + * already be gone and MMIO is not reliable. + */ + spin_lock_irqsave(&wx->ptp_tx_lock, flags); + skb = __wx_ptp_detach_tx_skb(wx); + spin_unlock_irqrestore(&wx->ptp_tx_lock, flags); + dev_kfree_skb_any(skb); if (wx->ptp_clock) { ptp_clock_unregister(wx->ptp_clock); diff --git a/drivers/net/ethernet/wangxun/libwx/wx_type.h b/drivers/net/ethernet/wangxun/libwx/wx_type.h index 9454e90258d8..afd980dbb793 100644 --- a/drivers/net/ethernet/wangxun/libwx/wx_type.h +++ b/drivers/net/ethernet/wangxun/libwx/wx_type.h @@ -1430,6 +1430,8 @@ struct wx { unsigned long last_overflow_check; unsigned long last_rx_ptp_check; unsigned long ptp_tx_start; + /* protects ptp_tx_skb, ptp_tx_start and the in-progress state bit */ + spinlock_t ptp_tx_lock; seqlock_t hw_tc_lock; /* seqlock for ptp */ struct cyclecounter hw_cc; struct timecounter hw_tc; From 26b2bd70d22457556e2fa01cbf1192cb1a94d619 Mon Sep 17 00:00:00 2001 From: Ilya Maximets Date: Mon, 21 Sep 2026 16:55:43 +0200 Subject: [PATCH 165/189] net: openvswitch: conntrack: avoid modifying shared unconfirmed ct entry In a case where skb with an unconfirmed ct entry gets cloned, we may end up committing both but with different sets of extensions. The series of events: 1. The first clone wants to commit and runs the helpers wiring up the extension pointer into the expectation list. 2. Then it looses the confirmation keeping the entry unconfirmed. 3. Second clone now wants to commit labels and adds the new extension for that breaking the pointer in the expectation list causing UAF on the destruction path later. While this is possible to trigger, there should be no practical network pipeline where committing both clones without modifications into the same zone is needed. So, let's just reset the entry in case for some reason we got an skb with a shared one during commit. This doesn't affect any known use cases, but avoids any potential problems with sharing and modification of the unconfirmed ct entry. The fixes tag points to the introduction of helpers, since that's the main UAF trigger for the sharing. Fixes: cae3a2627520 ("openvswitch: Allow attaching helpers to ct action") Cc: stable@vger.kernel.org Reported-by: Axel Mierczuk Signed-off-by: Ilya Maximets Reviewed-by: Aaron Conole Link: https://patch.msgid.link/20260921145655.3167436-2-i.maximets@ovn.org Signed-off-by: Jakub Kicinski --- include/net/netfilter/nf_conntrack.h | 5 +++++ net/openvswitch/conntrack.c | 12 ++++++++++++ 2 files changed, 17 insertions(+) diff --git a/include/net/netfilter/nf_conntrack.h b/include/net/netfilter/nf_conntrack.h index bc42dd0e10e6..c39425e54d87 100644 --- a/include/net/netfilter/nf_conntrack.h +++ b/include/net/netfilter/nf_conntrack.h @@ -185,6 +185,11 @@ static inline void nf_ct_put(struct nf_conn *ct) nf_ct_destroy(&ct->ct_general); } +static inline bool nf_ct_shared(const struct nf_conn *ct) +{ + return refcount_read(&ct->ct_general.use) > 1; +} + /* load module; enable/disable conntrack in this namespace */ int nf_ct_netns_get(struct net *net, u8 nfproto); void nf_ct_netns_put(struct net *net, u8 nfproto); diff --git a/net/openvswitch/conntrack.c b/net/openvswitch/conntrack.c index 0f433688e17b..a733029c28dd 100644 --- a/net/openvswitch/conntrack.c +++ b/net/openvswitch/conntrack.c @@ -734,6 +734,18 @@ static int __ovs_ct_lookup(struct net *net, struct sw_flow_key *key, enum ip_conntrack_info ctinfo; struct nf_conn *ct; + /* If the ct entry is not confirmed and shared with some other skb, + * e.g., a cloned one, we can't just modify it with the commit as we + * must not modify the extension set. Reset. + */ + if (cached && info->commit) { + ct = nf_ct_get(skb, &ctinfo); + if (ct && !nf_ct_is_confirmed(ct) && nf_ct_shared(ct)) { + nf_reset_ct(skb); + cached = false; + } + } + if (!cached) { struct nf_hook_state state = { .hook = NF_INET_PRE_ROUTING, From 5e6c14dd42a1c1fe938e573dc6c9098145b2b0c4 Mon Sep 17 00:00:00 2001 From: Ilya Maximets Date: Mon, 21 Sep 2026 16:55:44 +0200 Subject: [PATCH 166/189] net: openvswitch: conntrack: remove 'add_helper' dead code This variable can only become 'true' when the connection is not confirmed, but it is only checked when it is confirmed. So, it can be treated as being always false and just removed. Fixes: 3c1860543fcc ("openvswitch: add nf_ct_is_confirmed check before assigning the helper") Cc: stable@vger.kernel.org Signed-off-by: Ilya Maximets Reviewed-by: Aaron Conole Link: https://patch.msgid.link/20260921145655.3167436-3-i.maximets@ovn.org Signed-off-by: Jakub Kicinski --- net/openvswitch/conntrack.c | 10 ++-------- 1 file changed, 2 insertions(+), 8 deletions(-) diff --git a/net/openvswitch/conntrack.c b/net/openvswitch/conntrack.c index a733029c28dd..c20f096eef40 100644 --- a/net/openvswitch/conntrack.c +++ b/net/openvswitch/conntrack.c @@ -778,8 +778,6 @@ static int __ovs_ct_lookup(struct net *net, struct sw_flow_key *key, ct = nf_ct_get(skb, &ctinfo); if (ct) { - bool add_helper = false; - /* Packets starting a new connection must be NATted before the * helper, so that the helper knows about the NAT. We enforce * this by delaying both NAT and helper calls for unconfirmed @@ -811,7 +809,6 @@ static int __ovs_ct_lookup(struct net *net, struct sw_flow_key *key, GFP_ATOMIC); if (err) return err; - add_helper = true; /* helper installed, add seqadj if NAT is required */ if (info->nat && !nfct_seqadj(ct)) { @@ -821,13 +818,10 @@ static int __ovs_ct_lookup(struct net *net, struct sw_flow_key *key, } /* Call the helper only if: - * - nf_conntrack_in() was executed above ("!cached") or a - * helper was just attached ("add_helper") for a confirmed - * connection, or + * - nf_conntrack_in() was executed above ("!cached"), or * - When committing an unconfirmed connection. */ - if ((nf_ct_is_confirmed(ct) ? !cached || add_helper : - info->commit)) { + if ((nf_ct_is_confirmed(ct) ? !cached : info->commit)) { int err = nf_ct_helper(skb, ct, ctinfo, info->family); err = verdict_to_errno(err); From 1a4151e6be57b098b7a5ebfbde58585e83200cdc Mon Sep 17 00:00:00 2001 From: Ilya Maximets Date: Mon, 21 Sep 2026 16:55:45 +0200 Subject: [PATCH 167/189] net: openvswitch: conntrack: fix helper UAF due to extensions realloc While calling the helpers, a raw pointer to the extensions area is wired into expectations list: -> nf_ct_helper() -> helper->help() -> nf_ct_expect_related_report() -> nf_ct_expect_insert() -> hlist_add_head_rcu(&exp->lnode, &master_help->expectations) In case the connection is not confirmed yet, more extensions can be added afterwards with *_ext_add() calls reallocating the extension space and leaving the now invalid pointer in the expectations list that is later accessed while removing the expectation. Make sure that helpers are called at the end after all the other extensions are already added. Note that the helper rejection now leaves the mark and labels set, but that's not different from how the NAT was handled before or how the mark and the labels were handled on confirmation failure. And there are no atomicity guarantees provided by the API anyway. Fixes: cae3a2627520 ("openvswitch: Allow attaching helpers to ct action") Cc: stable@vger.kernel.org Reported-by: Axel Mierczuk Signed-off-by: Ilya Maximets Reviewed-by: Aaron Conole Link: https://patch.msgid.link/20260921145655.3167436-4-i.maximets@ovn.org Signed-off-by: Jakub Kicinski --- net/openvswitch/conntrack.c | 19 +++++++++++++++---- 1 file changed, 15 insertions(+), 4 deletions(-) diff --git a/net/openvswitch/conntrack.c b/net/openvswitch/conntrack.c index c20f096eef40..d3326edcabf7 100644 --- a/net/openvswitch/conntrack.c +++ b/net/openvswitch/conntrack.c @@ -817,11 +817,14 @@ static int __ovs_ct_lookup(struct net *net, struct sw_flow_key *key, } } - /* Call the helper only if: - * - nf_conntrack_in() was executed above ("!cached"), or - * - When committing an unconfirmed connection. + /* Call the helper only if nf_conntrack_in() was executed + * above ("!cached"). + * + * For unconfirmed connections it will be called later during + * commit as we need to have all the other extensions allocated + * before the call. */ - if ((nf_ct_is_confirmed(ct) ? !cached : info->commit)) { + if (nf_ct_is_confirmed(ct) && !cached) { int err = nf_ct_helper(skb, ct, ctinfo, info->family); err = verdict_to_errno(err); @@ -1025,6 +1028,14 @@ static int ovs_ct_commit(struct net *net, struct sw_flow_key *key, return err; nf_conn_act_ct_ext_add(skb, ct, ctinfo); + + /* Call the helpers now. We couldn't do this before as + * all the extensions must be allocated before the call. + */ + err = nf_ct_helper(skb, ct, ctinfo, info->family); + err = verdict_to_errno(err); + if (err) + return err; } else if (IS_ENABLED(CONFIG_NF_CONNTRACK_LABELS) && labels_nonzero(&info->labels.mask)) { err = ovs_ct_set_labels(ct, key, &info->labels.value, From f85009dfcd65e5969526b0db7a49b5413746e630 Mon Sep 17 00:00:00 2001 From: Ilya Maximets Date: Mon, 21 Sep 2026 16:55:46 +0200 Subject: [PATCH 168/189] net/sched: act_ct: avoid modifying shared unconfirmed ct entry In a case where skb with an unconfirmed ct entry gets cloned, we may end up processing both again but with different sets of extensions. The series of events: 1. The first clone wants to commit and runs the helpers wiring up the extension pointer into the expectation list. 2. Then it looses the confirmation keeping the entry unconfirmed. 3. Second clone now wants to commit labels or run NAT and adds the new extension for that breaking the pointer in the expectation list causing UAF on the destruction path later. While this is possible to trigger, there should be no practical network pipeline where we need to process both clones without modifications in the same zone. So, let's just reset the entry in case for some reason we got an skb with a shared one. This doesn't affect any known use cases, but avoids any potential problems with sharing and modification of the unconfirmed ct entry. Unlike openvswitch module, act_ct allows for NAT without commit. Changing that would be a uAPI break. So, act_ct needs to reset on NAT regardless of the commit flag to avoid reallocation of the extension space. This, however, doesn't really change the picture for sensible networking cases as there should be no need to run the same packet twice (before and after the clone) through conntrack without packet header or zone changes and without commit. The fixes tag points to the introduction of helpers, since that's the main UAF trigger for the sharing. Fixes: a21b06e73191 ("net: sched: add helper support in act_ct") Cc: stable@vger.kernel.org Reported-by: Axel Mierczuk Signed-off-by: Ilya Maximets Reviewed-by: Aaron Conole Reviewed-by: Xin Long Reviewed-by: Jamal Hadi Salim Link: https://patch.msgid.link/20260921145655.3167436-5-i.maximets@ovn.org Signed-off-by: Jakub Kicinski --- net/sched/act_ct.c | 18 ++++++++++++++++-- 1 file changed, 16 insertions(+), 2 deletions(-) diff --git a/net/sched/act_ct.c b/net/sched/act_ct.c index 55f3521edb4c..e72143d36b11 100644 --- a/net/sched/act_ct.c +++ b/net/sched/act_ct.c @@ -979,11 +979,11 @@ TC_INDIRECT_SCOPE int tcf_ct_act(struct sk_buff *skb, const struct tc_action *a, struct tcf_result *res) { struct net *net = dev_net(skb->dev); + bool cached, commit, clear, nat; enum ip_conntrack_info ctinfo; struct tcf_ct *c = to_ct(a); struct nf_conn *tmpl = NULL; struct nf_hook_state state; - bool cached, commit, clear; int nh_ofs, err, retval; struct tcf_ct_params *p; bool add_helper = false; @@ -998,6 +998,7 @@ TC_INDIRECT_SCOPE int tcf_ct_act(struct sk_buff *skb, const struct tc_action *a, retval = p->action; commit = p->ct_action & TCA_CT_ACT_COMMIT; clear = p->ct_action & TCA_CT_ACT_CLEAR; + nat = p->ct_action & TCA_CT_ACT_NAT; tmpl = p->tmpl; tcf_lastuse_update(&c->tcf_tm); @@ -1046,6 +1047,19 @@ TC_INDIRECT_SCOPE int tcf_ct_act(struct sk_buff *skb, const struct tc_action *a, * different zone. */ cached = tcf_ct_skb_nfct_cached(net, skb, p); + + /* If the ct entry is not confirmed and shared with some other skb, + * e.g., a cloned one, we can't just modify it with a commit or nat + * as we must not modify the extension set. Reset. + */ + if (cached && (commit || nat)) { + ct = nf_ct_get(skb, &ctinfo); + if (ct && !nf_ct_is_confirmed(ct) && nf_ct_shared(ct)) { + nf_reset_ct(skb); + cached = false; + } + } + if (!cached) { if (tcf_ct_flow_table_lookup(p, skb, family)) { skip_add = true; @@ -1083,7 +1097,7 @@ TC_INDIRECT_SCOPE int tcf_ct_act(struct sk_buff *skb, const struct tc_action *a, if (err) goto drop; add_helper = true; - if (p->ct_action & TCA_CT_ACT_NAT && !nfct_seqadj(ct)) { + if (nat && !nfct_seqadj(ct)) { if (!nfct_seqadj_ext_add(ct)) goto drop; } From 00df72e39f306e2f7adb68528a5c92109da6a0a9 Mon Sep 17 00:00:00 2001 From: Ilya Maximets Date: Mon, 21 Sep 2026 16:55:47 +0200 Subject: [PATCH 169/189] net/sched: act_ct: remove 'add_helper' dead code This variable can only become 'true' when the connection is not confirmed, but it is only checked when it is confirmed. So, it can be treated as being always false and just removed. Fixes: a21b06e73191 ("net: sched: add helper support in act_ct") Cc: stable@vger.kernel.org Signed-off-by: Ilya Maximets Reviewed-by: Aaron Conole Reviewed-by: Xin Long Reviewed-by: Jamal Hadi Salim Link: https://patch.msgid.link/20260921145655.3167436-6-i.maximets@ovn.org Signed-off-by: Jakub Kicinski --- net/sched/act_ct.c | 5 ++--- 1 file changed, 2 insertions(+), 3 deletions(-) diff --git a/net/sched/act_ct.c b/net/sched/act_ct.c index e72143d36b11..f62051ec9d57 100644 --- a/net/sched/act_ct.c +++ b/net/sched/act_ct.c @@ -986,7 +986,6 @@ TC_INDIRECT_SCOPE int tcf_ct_act(struct sk_buff *skb, const struct tc_action *a, struct nf_hook_state state; int nh_ofs, err, retval; struct tcf_ct_params *p; - bool add_helper = false; bool skb_is_ours = false; bool skip_add = false; bool defrag = false; @@ -1096,14 +1095,14 @@ TC_INDIRECT_SCOPE int tcf_ct_act(struct sk_buff *skb, const struct tc_action *a, err = __nf_ct_try_assign_helper(ct, p->tmpl, GFP_ATOMIC); if (err) goto drop; - add_helper = true; + if (nat && !nfct_seqadj(ct)) { if (!nfct_seqadj_ext_add(ct)) goto drop; } } - if (nf_ct_is_confirmed(ct) ? ((!cached && !skip_add) || add_helper) : commit) { + if (nf_ct_is_confirmed(ct) ? (!cached && !skip_add) : commit) { err = nf_ct_helper(skb, ct, ctinfo, family); if (err != NF_ACCEPT) goto nf_error; From dad19b59da050cb60d3f7023dac2a042a84bf0bd Mon Sep 17 00:00:00 2001 From: Ilya Maximets Date: Mon, 21 Sep 2026 16:55:48 +0200 Subject: [PATCH 170/189] net/sched: act_ct: fix helper UAF due to extensions realloc While calling the helpers, a raw pointer to the extensions area is wired into expectations list: -> nf_ct_helper() -> helper->help() -> nf_ct_expect_related_report() -> nf_ct_expect_insert() -> hlist_add_head_rcu(&exp->lnode, &master_help->expectations) In case the connection is not confirmed yet, more extensions can be added afterwards with *_ext_add() calls reallocating the extension space and leaving the now invalid pointer in the expectations list that is later accessed while removing the expectation. Make sure that helpers are called at the end after all the other extensions are already added. Note that the helper rejection now leaves the mark and labels set, but that's not different from how the NAT was handled before or how the mark and the labels were handled on confirmation failure. And there are no atomicity guarantees provided by the API anyway. Fixes: a21b06e73191 ("net: sched: add helper support in act_ct") Cc: stable@vger.kernel.org Reported-by: Axel Mierczuk Signed-off-by: Ilya Maximets Reviewed-by: Xin Long Reviewed-by: Jamal Hadi Salim Reviewed-by: Aaron Conole Link: https://patch.msgid.link/20260921145655.3167436-7-i.maximets@ovn.org Signed-off-by: Jakub Kicinski --- net/sched/act_ct.c | 18 ++++++++++++------ 1 file changed, 12 insertions(+), 6 deletions(-) diff --git a/net/sched/act_ct.c b/net/sched/act_ct.c index f62051ec9d57..411e3dd92d07 100644 --- a/net/sched/act_ct.c +++ b/net/sched/act_ct.c @@ -1102,6 +1102,18 @@ TC_INDIRECT_SCOPE int tcf_ct_act(struct sk_buff *skb, const struct tc_action *a, } } + if (commit) { + tcf_ct_act_set_mark(ct, p->mark, p->mark_mask); + tcf_ct_act_set_labels(ct, p->labels, p->labels_mask); + + if (!nf_ct_is_confirmed(ct)) + nf_conn_act_ct_ext_add(skb, ct, ctinfo); + } + + /* Run helpers for the connection if nf_conntrack_in() was executed + * or if we're about to commit. This has to be done after all the + * extensions are already added. + */ if (nf_ct_is_confirmed(ct) ? (!cached && !skip_add) : commit) { err = nf_ct_helper(skb, ct, ctinfo, family); if (err != NF_ACCEPT) @@ -1109,12 +1121,6 @@ TC_INDIRECT_SCOPE int tcf_ct_act(struct sk_buff *skb, const struct tc_action *a, } if (commit) { - tcf_ct_act_set_mark(ct, p->mark, p->mark_mask); - tcf_ct_act_set_labels(ct, p->labels, p->labels_mask); - - if (!nf_ct_is_confirmed(ct)) - nf_conn_act_ct_ext_add(skb, ct, ctinfo); - /* This will take care of sending queued events * even if the connection is already confirmed. */ From 0958ea4355e2e9220ad4e13da3b7d94f365ed34f Mon Sep 17 00:00:00 2001 From: Guangshuo Li Date: Mon, 21 Sep 2026 23:42:01 +0800 Subject: [PATCH 171/189] net: ena: fix PHC cleanup on probe failure ena_probe() initializes the PHC as part of ena_device_init(), but the probe failure path does not destroy it before freeing the PHC private data. The normal removal path calls ena_phc_destroy() through ena_destroy_device() before ena_phc_free(). However, if probe fails after ena_device_init() succeeds, the error path reaches ena_phc_free() without unregistering the PTP clock or destroying the device PHC resources. Call ena_phc_destroy() in the probe error path before freeing the PHC private data. This issue was found by manual code inspection. Cc: stable@vger.kernel.org tags and describe this as a consistency cleanup Fixes: e0ea34158ee8 ("net: ena: Add PHC support in the ENA driver") Cc: stable@vger.kernel.org Signed-off-by: Guangshuo Li Cc: stable Link: https://patch.msgid.link/20260921154202.471662-2-lgs201920130244@gmail.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/amazon/ena/ena_netdev.c | 1 + 1 file changed, 1 insertion(+) diff --git a/drivers/net/ethernet/amazon/ena/ena_netdev.c b/drivers/net/ethernet/amazon/ena/ena_netdev.c index ea89619039d8..5f0864d16dd3 100644 --- a/drivers/net/ethernet/amazon/ena/ena_netdev.c +++ b/drivers/net/ethernet/amazon/ena/ena_netdev.c @@ -4122,6 +4122,7 @@ static int ena_probe(struct pci_dev *pdev, const struct pci_device_id *ent) err_device_destroy: ena_com_delete_host_info(ena_dev); ena_com_admin_destroy(ena_dev); + ena_phc_destroy(adapter); ena_devlink_destroy: ena_devlink_free(devlink); err_metrics_destroy: From 9476b4468862927297c94c440863cd8ed1e7cc83 Mon Sep 17 00:00:00 2001 From: Guangshuo Li Date: Mon, 21 Sep 2026 23:42:02 +0800 Subject: [PATCH 172/189] net: ena: fix MMIO read buffer leak on probe failure ena_device_init() initializes the MMIO read mechanism with ena_com_mmio_reg_read_request_init(), which allocates a coherent DMA buffer for MMIO read responses. The normal removal path releases this buffer through ena_com_mmio_reg_read_request_destroy(). However, if ena_probe() fails after ena_device_init() succeeds, the error path destroys the admin resources and eventually frees ena_dev without destroying the MMIO read request, leaving the coherent DMA buffer allocated. Call ena_com_mmio_reg_read_request_destroy() in the probe error path before releasing the remaining device resources. This issue was found by manual code inspection. Fixes: 1738cd3ed342 ("net: ena: Add a driver for Amazon Elastic Network Adapters (ENA)") Cc: stable@vger.kernel.org Signed-off-by: Guangshuo Li Link: https://patch.msgid.link/20260921154202.471662-3-lgs201920130244@gmail.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/amazon/ena/ena_netdev.c | 1 + 1 file changed, 1 insertion(+) diff --git a/drivers/net/ethernet/amazon/ena/ena_netdev.c b/drivers/net/ethernet/amazon/ena/ena_netdev.c index 5f0864d16dd3..7eb6456ed0d5 100644 --- a/drivers/net/ethernet/amazon/ena/ena_netdev.c +++ b/drivers/net/ethernet/amazon/ena/ena_netdev.c @@ -4123,6 +4123,7 @@ static int ena_probe(struct pci_dev *pdev, const struct pci_device_id *ent) ena_com_delete_host_info(ena_dev); ena_com_admin_destroy(ena_dev); ena_phc_destroy(adapter); + ena_com_mmio_reg_read_request_destroy(ena_dev); ena_devlink_destroy: ena_devlink_free(devlink); err_metrics_destroy: From 1a983a4e14c635c40354be110cd9a1a5c94e01e6 Mon Sep 17 00:00:00 2001 From: Sang-Hoon Choi Date: Tue, 22 Sep 2026 03:15:59 +0900 Subject: [PATCH 173/189] nfp: hold IPsec RX state under the XArray lock nfp_net_ipsec_rx() drops the XArray lock before taking a reference to the xfrm_state it found. The delete path can erase the entry and drop the last state reference in that interval. RX can then try to increment a zero refcount after the state has been queued for destruction. The driver queues firmware invalidation asynchronously; the delete path does not wait for the command to complete or drain pending RX processing. The XFRM garbage collector waits for an RCU grace period before freeing the state. That delays reclamation but does not make acquiring a reference from zero valid. Take the xfrm_state reference before releasing the XArray lock so xa_erase() cannot run between lookup and reference acquisition. Fixes: 57f273adbcd4 ("nfp: add framework to support ipsec offloading") Reported-by: Changyul Lee Signed-off-by: Sang-Hoon Choi Link: https://patch.msgid.link/179001455912.44752.17153022439349797877.idr-bug-92@gmail.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/netronome/nfp/crypto/ipsec.c | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/drivers/net/ethernet/netronome/nfp/crypto/ipsec.c b/drivers/net/ethernet/netronome/nfp/crypto/ipsec.c index 9e7c285eaa6b..960d7513aa8d 100644 --- a/drivers/net/ethernet/netronome/nfp/crypto/ipsec.c +++ b/drivers/net/ethernet/netronome/nfp/crypto/ipsec.c @@ -625,11 +625,12 @@ int nfp_net_ipsec_rx(struct nfp_meta_parsed *meta, struct sk_buff *skb) xa_lock(&nn->xa_ipsec); x = xa_load(&nn->xa_ipsec, saidx); + if (x) + xfrm_state_hold(x); xa_unlock(&nn->xa_ipsec); if (!x) return -EINVAL; - xfrm_state_hold(x); sp->xvec[sp->len++] = x; sp->olen++; xo = xfrm_offload(skb); From ba3d1f480c7a3fba963e7867ad6cc557c197acbc Mon Sep 17 00:00:00 2001 From: Allison Henderson Date: Mon, 21 Sep 2026 14:50:27 -0700 Subject: [PATCH 174/189] net/rds: size a connection's path set by the transport it ends up with __rds_conn_create() computes npaths from the caller's transport before it decides whether a connection to one of the host's own addresses is to be handled by the loopback transport instead. That substitution is what an RDS/TCP socket sending to a local address gets, and after it the path init loop still runs for the TCP transport's RDS_MPATH_WORKERS paths and allocates an ordered workqueue for each, while rds_loop_conn_alloc() only ever provides transport data for path 0. rds_conn_destroy() sizes its teardown from c_trans, by then the loopback transport, so it visits path 0 only - and rds_conn_path_destroy() would skip the other paths anyway, since it returns before destroy_workqueue() for a path without transport data. kfree(c_path) then drops the last pointers to seven workqueues. That repeats for every such connection, on every netns teardown or module unload, and every distinct local destination address is a separate connection. Recompute npaths once the transport is final, so that creation and destruction agree on the set of paths. The c_path array stays sized for the caller's transport; the unused entries are freed with it. Fixes: 4716af3897e9 ("net/rds: Give each connection path its own workqueue") Signed-off-by: Allison Henderson Link: https://patch.msgid.link/20260921215027.174657-1-achender@kernel.org Signed-off-by: Jakub Kicinski --- net/rds/connection.c | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/net/rds/connection.c b/net/rds/connection.c index b6c4beb50eaf..c752a8623cfc 100644 --- a/net/rds/connection.c +++ b/net/rds/connection.c @@ -276,6 +276,12 @@ static struct rds_connection *__rds_conn_create(struct net *net, conn->c_trans = trans; + /* The transport may just have been swapped for loopback; size the + * set of paths - which is also what rds_conn_destroy() tears down + * again - by the transport the connection actually uses. + */ + npaths = (trans->t_mp_capable ? RDS_MPATH_WORKERS : 1); + init_waitqueue_head(&conn->c_hs_waitq); for (i = 0; i < npaths; i++) { __rds_conn_path_init(conn, &conn->c_path[i], From 61cb282fe97b3b0ba32ca09417a693162bf4ae3f Mon Sep 17 00:00:00 2001 From: Sidraya Jayagond Date: Tue, 22 Sep 2026 09:31:49 +0200 Subject: [PATCH 175/189] net/smc: fix UAF on lgr list traversal in smcr_port_err() smcr_port_err() traverses smc_lgr_list.list without holding smc_lgr_list.lock, allowing a concurrent smc_lgr_terminate_sched() to free an lgr while it is still being dereferenced. Hold smc_lgr_list.lock across the traversal. Update smc_ib_gid_check() to call smcr_port_err() after releasing the lock. Fixes: 541afa10c126 ("net/smc: add smcr_port_err() and smcr_link_down() processing") Reviewed-by: Mahanta Jambigi Signed-off-by: Sidraya Jayagond Reviewed-by: Dust Li Link: https://patch.msgid.link/20260922073149.474762-1-sidraya@linux.ibm.com Signed-off-by: Jakub Kicinski --- net/smc/smc_core.c | 2 ++ net/smc/smc_ib.c | 10 ++++++++-- 2 files changed, 10 insertions(+), 2 deletions(-) diff --git a/net/smc/smc_core.c b/net/smc/smc_core.c index 04aedd957543..9974149659c2 100644 --- a/net/smc/smc_core.c +++ b/net/smc/smc_core.c @@ -1849,6 +1849,7 @@ void smcr_port_err(struct smc_ib_device *smcibdev, u8 ibport) struct smc_link_group *lgr, *n; int i; + spin_lock_bh(&smc_lgr_list.lock); list_for_each_entry_safe(lgr, n, &smc_lgr_list.list, list) { if (strncmp(smcibdev->pnetid[ibport - 1], lgr->pnet_id, SMC_MAX_PNETID_LEN)) @@ -1863,6 +1864,7 @@ void smcr_port_err(struct smc_ib_device *smcibdev, u8 ibport) smcr_link_down_cond_sched(lnk); } } + spin_unlock_bh(&smc_lgr_list.lock); } static void smc_link_down_work(struct work_struct *work) diff --git a/net/smc/smc_ib.c b/net/smc/smc_ib.c index 9bb495707445..daaa8a72da90 100644 --- a/net/smc/smc_ib.c +++ b/net/smc/smc_ib.c @@ -333,6 +333,7 @@ static bool smc_ib_check_link_gid(u8 gid[SMC_GID_SIZE], bool smcrv2, static void smc_ib_gid_check(struct smc_ib_device *smcibdev, u8 ibport) { struct smc_link_group *lgr; + bool stale_gid = false; int i; spin_lock_bh(&smc_lgr_list.lock); @@ -348,11 +349,16 @@ static void smc_ib_gid_check(struct smc_ib_device *smcibdev, u8 ibport) continue; if (!smc_ib_check_link_gid(lgr->lnk[i].gid, lgr->smc_version == SMC_V2, - smcibdev, ibport)) - smcr_port_err(smcibdev, ibport); + smcibdev, ibport)) { + stale_gid = true; + goto out; + } } } +out: spin_unlock_bh(&smc_lgr_list.lock); + if (stale_gid) + smcr_port_err(smcibdev, ibport); } static int smc_ib_remember_port_attr(struct smc_ib_device *smcibdev, u8 ibport) From b94773dc4df7026a6f29b2c65e96e88d29cdb576 Mon Sep 17 00:00:00 2001 From: Alexander Sverdlin Date: Tue, 22 Sep 2026 09:52:46 +0200 Subject: [PATCH 176/189] net: phy: intel-xway: workaround 100BASE-TX Link-Up issue MaxLinear GSW12x/GSW14x Ethernet Switch Errata Sheet states: "An issue has been sporadically observed after device power-on on the first link-up attempt in 100BASE-TX mode resulting in either the link-up taking a long time, or failing to link-up altogether... Workaround: After power-on, enable Cable Diagnostic Mode for all ports and disable it..." Implement the proposed workaround unconditionally in the Intel XWAY driver (MaxLinear GSW1xx switches incorporate Intel XWAY PHYs) because the diagnostic bits have the same meaning even in older integral PHYs such as GPY111/PEF7071/PHY11G. So it's not clear how to distinguish the affected newer integrated PHYs, but the workaround should not hurt the older PHYs. Cc: stable@vger.kernel.org Fixes: 22335939ec90 ("net: dsa: add driver for MaxLinear GSW1xx switch family") Signed-off-by: Alexander Sverdlin Reviewed-by: Andrew Lunn Link: https://patch.msgid.link/20260922075251.23386-1-alexander.sverdlin@siemens.com Signed-off-by: Jakub Kicinski --- drivers/net/phy/intel-xway.c | 29 ++++++++++++++++++++++++++++- 1 file changed, 28 insertions(+), 1 deletion(-) diff --git a/drivers/net/phy/intel-xway.c b/drivers/net/phy/intel-xway.c index afbcec711744..3cee31bb931f 100644 --- a/drivers/net/phy/intel-xway.c +++ b/drivers/net/phy/intel-xway.c @@ -16,6 +16,11 @@ #define XWAY_MDIO_ISTAT 0x1A /* interrupt status */ #define XWAY_MDIO_LED 0x1B /* led control */ +#define XWAY_MDIO_GCTRL_TM_MASK GENMASK(15, 13) +#define XWAY_MDIO_GCTRL_TM(mode) FIELD_PREP(XWAY_MDIO_GCTRL_TM_MASK, (mode)) +#define XWAY_MDIO_GCTRL_TM_NOP XWAY_MDIO_GCTRL_TM(0) /* Normal operation */ +#define XWAY_MDIO_GCTRL_TM_CDIAG XWAY_MDIO_GCTRL_TM(6) /* Cable diagnostics */ + #define XWAY_MDIO_ERRCNT_SEL GENMASK(11, 8) #define XWAY_MDIO_ERRCNT_COUNT GENMASK(7, 0) #define XWAY_MDIO_ERRCNT_SEL_RXERR 0 @@ -326,6 +331,28 @@ static int xway_gphy_probe(struct phy_device *phydev) return 0; } +static int xway_11g_int_config_init(struct phy_device *phydev) +{ + int err; + + /* An issue has been sporadically observed after device power-on on the + * first link-up attempt in 100BASE-TX mode resulting in either the + * link-up taking a long time, or failing to link-up altogether. + * + * Workaround: + * After power-on, enable Cable Diagnostic Mode for all ports and + * disable it. + */ + err = phy_modify(phydev, MII_CTRL1000, XWAY_MDIO_GCTRL_TM_MASK, XWAY_MDIO_GCTRL_TM_CDIAG); + if (err) + return err; + err = phy_modify(phydev, MII_CTRL1000, XWAY_MDIO_GCTRL_TM_MASK, XWAY_MDIO_GCTRL_TM_NOP); + if (err) + return err; + + return xway_gphy_config_init(phydev); +} + static int xway_gphy14_config_aneg(struct phy_device *phydev) { int reg, err; @@ -735,7 +762,7 @@ static struct phy_driver xway_gphy[] = { .phy_id_mask = 0xffffffff, .name = "Intel XWAY PHY11G (xRX v1.2 integrated)", /* PHY_GBIT_FEATURES */ - .config_init = xway_gphy_config_init, + .config_init = xway_11g_int_config_init, .probe = xway_gphy_probe, .handle_interrupt = xway_gphy_handle_interrupt, .config_intr = xway_gphy_config_intr, From 8e1937fed6738460554ec123c64839e2445e7d53 Mon Sep 17 00:00:00 2001 From: Ginger Li Date: Tue, 22 Sep 2026 16:09:09 +0800 Subject: [PATCH 177/189] tipc: Fix a data race on mon->peer_cnt in mon_timeout() mon_timeout() evaluates dom_size(mon->peer_cnt) before it takes mon->lock, while mon->peer_cnt is updated under that lock by tipc_mon_add_peer() and tipc_mon_remove_peer(). The value can therefore be stale, and the decision whether the local domain has to be recomputed can be based on an outdated member count. Read mon->peer_cnt inside the write_lock_bh(&mon->lock) protected region. Fixes: 35c55c9877f8 ("tipc: add neighbor monitoring framework") Signed-off-by: Ginger Li Reviewed-by: Tung Nguyen Link: https://patch.msgid.link/20260922080909.21123-1-ginger.jzllee@gmail.com Signed-off-by: Jakub Kicinski --- net/tipc/monitor.c | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/net/tipc/monitor.c b/net/tipc/monitor.c index a94b9b36a700..1a438e312d66 100644 --- a/net/tipc/monitor.c +++ b/net/tipc/monitor.c @@ -632,9 +632,10 @@ static void mon_timeout(struct timer_list *t) { struct tipc_monitor *mon = timer_container_of(mon, t, timer); struct tipc_peer *self; - int best_member_cnt = dom_size(mon->peer_cnt) - 1; + int best_member_cnt; write_lock_bh(&mon->lock); + best_member_cnt = dom_size(mon->peer_cnt) - 1; self = mon->self; if (self && (best_member_cnt != self->applied)) { mon_update_local_domain(mon); From 56d82862a0a243ac14ba11b6d7b57ddc2d064b95 Mon Sep 17 00:00:00 2001 From: Dairui Zhang Date: Wed, 23 Sep 2026 13:01:01 +0800 Subject: [PATCH 178/189] af_packet: fix integer overflow in prb_calc_retire_blk_tmo() prb_calc_retire_blk_tmo() computes in 32-bit int arithmetic: mbits = (blk_size_in_bytes * 8) / (1024 * 1024); If I'm reading the validation right, tp_block_size is user controlled and packet_set_ring() only rejects values that are <= 0 as int or not page aligned, so a 256MiB block goes right through (and alloc_one_pg_vec_page() even has a vzalloc fallback for it). 0x10000000 * 8 wraps to INT_MIN, and on a NIC reporting 1 Gbps (div == 1) the function ends up returning -2047. The condition is actually (8 * size) mod 2^32 >= 2^31 && div == 1, so the trigger set is [256,512), [768,1024), [1280,1536) and [1792,2048) MiB. Other sizes wrap to non-negative values and faster links divide the unsigned value back below 2^31, which is why this doesn't blow up for everyone. What makes it fatal is what happens next in init_prb_bdqc(): p1->interval_ktime = ms_to_ktime(prb_calc_retire_blk_tmo(...)); hrtimer_start(&p1->retire_blk_timer, p1->interval_ktime, HRTIMER_MODE_REL_SOFT); A negative relative timeout expires immediately. The callback unconditionally returns HRTIMER_RESTART, and hrtimer_forward() turns the negative interval into hrtimer_resolution: if (interval < hrtimer_resolution) interval = hrtimer_resolution; So the SOFT timer re-fires at the maximum rate forever, holding sk_receive_queue.lock each pass. One CPU spins in softirq until the socket is closed. Repeat with more rings and the machine is gone. The overflow itself is ancient - it was introduced together with TPACKET_V3 in f6fb8f100b80 ("af-packet: TPACKET_V3 flexible buffer implementation."). Its effect prior to f7460d2989fa ("net: af_packet: Use hrtimer to do the retire operation", v6.18) was not as clear-cut, though: the return value was stored into an unsigned short retire_blk_tov, so a negative result was truncated, and a 0-jiffy delay loop could be programmed as well. Neither is nearly as detrimental as the immediate maximum-rate spin the hrtimer conversion turned it into. (Unrelated to CVE-2019-20812 - that one was the ethtool failure path returning 0, which now returns DEFAULT_PRB_RETIRE_TOV.) Reproducer, needs CAP_NET_RAW (a --network host container has it by default) and a 1 Gbps NIC (QEMU e1000 works): int fd = socket(AF_PACKET, SOCK_RAW, htons(ETH_P_ALL)); bind(fd, ...); int v = TPACKET_V3; setsockopt(fd, SOL_PACKET, PACKET_VERSION, &v, sizeof(v)); struct tpacket_req3 req = { .tp_block_size = 0x10000000, .tp_block_nr = 1, .tp_frame_size = 2048, .tp_frame_nr = 0x10000000 / 2048, .tp_retire_blk_tov = 0, }; setsockopt(fd, SOL_PACKET, PACKET_RX_RING, &req, sizeof(req)); Compute in 64 bits instead. The operands are already bounded by the existing validation, so nothing else changes. If you'd prefer a different fix, just say so and I'll respin. Fixes: f6fb8f100b80 ("af-packet: TPACKET_V3 flexible buffer implementation.") Cc: stable@vger.kernel.org Signed-off-by: Dairui Zhang Reviewed-by: Willem de Bruijn Link: https://patch.msgid.link/20260923050101.1510064-1-zhangdairui@gmail.com Signed-off-by: Jakub Kicinski --- net/packet/af_packet.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/net/packet/af_packet.c b/net/packet/af_packet.c index 64b501db660a..7c83e01526ed 100644 --- a/net/packet/af_packet.c +++ b/net/packet/af_packet.c @@ -617,7 +617,7 @@ static int prb_calc_retire_blk_tmo(struct packet_sock *po, return DEFAULT_PRB_RETIRE_TOV; div = ecmd.base.speed / 1000; - mbits = (blk_size_in_bytes * 8) / (1024 * 1024); + mbits = (u64)blk_size_in_bytes * 8 / (1024 * 1024); if (div) mbits /= div; From 8db67bb6a1fffa4df68fbbc22e39943aeeff9178 Mon Sep 17 00:00:00 2001 From: Coia Prant Date: Wed, 23 Sep 2026 20:37:13 +0800 Subject: [PATCH 179/189] net: ethernet: stmmac: dwmac-rk: fix bulk clock leak when the PHY clock fails gmac_clk_enable() enables the bulk clocks first and then the optional PHY clock. If clk_prepare_enable() on the PHY clock fails, the function returns without rolling back the bulk clocks, and bsp_priv->clk_enabled stays false, so the later gmac_clk_enable(bsp_priv, false) becomes a no-op and the bulk clock references are leaked. Add the missing clk_bulk_disable_unprepare() on that failure path. Fixes: ea449f7fa0bf ("net: ethernet: stmmac: dwmac-rk: rework optional clock handling") Reviewed-by: Maxime Chevallier Reviewed-by: Heiko Stuebner Acked-by: Lorenzo Bianconi Signed-off-by: Coia Prant Link: https://patch.msgid.link/20260923123713.3137146-1-coiaprant@gmail.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/stmicro/stmmac/dwmac-rk.c | 5 ++++- 1 file changed, 4 insertions(+), 1 deletion(-) diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac-rk.c b/drivers/net/ethernet/stmicro/stmmac/dwmac-rk.c index 8d7042e68926..72bdbcb5e863 100644 --- a/drivers/net/ethernet/stmicro/stmmac/dwmac-rk.c +++ b/drivers/net/ethernet/stmicro/stmmac/dwmac-rk.c @@ -1162,8 +1162,11 @@ static int gmac_clk_enable(struct rk_priv_data *bsp_priv, bool enable) return ret; ret = clk_prepare_enable(bsp_priv->clk_phy); - if (ret) + if (ret) { + clk_bulk_disable_unprepare(bsp_priv->num_clks, + bsp_priv->clks); return ret; + } rk_configure_io_clksel(bsp_priv); rk_ungate_rmii_clock(bsp_priv); From 06e3f54e8b22040ada01a28343badf1990b4dd7d Mon Sep 17 00:00:00 2001 From: Eric Dumazet Date: Wed, 23 Sep 2026 13:03:18 +0000 Subject: [PATCH 180/189] net: flush skb_defer_nodes in dev_cpu_dead() When a CPU goes offline, dev_cpu_dead() drains its softnet queues (completion_queue, output_queue, poll_list, process_queue, and input_pkt_queue), but leaves net_hotdata.skb_defer_nodes untouched. If oldcpu goes offline while holding pending skbs in its skb_defer_nodes lists (e.g. below the sysctl_skb_defer_max >> 1 IPI threshold, or if the IPI races with CPU teardown), those skbs remain stranded until oldcpu is brought back online. If any of these skbs hold page_pool fragments, page_pool_destroy() will stall indefinitely waiting for inflight pages to be returned when a netdev or driver is torn down while oldcpu is offline. Additionally, if smp_call_function_single_async() fails in kick_defer_list_purge() because the target CPU went offline, reset defer_ipi_scheduled to 0 so future IPI kicks are not blocked when the CPU comes back online. Also, if oldcpu was the last online CPU on its NUMA node, drain that node's slot across all CPUs so no skbs deferred from that node remain stranded on idle remote CPUs (or if the node itself is subsequently offlined). Finally, in skb_attempt_defer_free(), re-check cpu_online(cpu) and whether the caller migrated CPUs after llist_add(), flushing the node list if so, to close the preemption TOCTOU race against CPU/node teardown. Fixes: 68822bdf76f1 ("net: generalize skb freeing deferral to per-cpu lists") Fixes: 5628f3fe3b16 ("net: add NUMA awareness to skb_attempt_defer_free()") Closes: https://lore.kernel.org/netdev/20260916003430.3612956-1-kris.pan@intel.com/ Signed-off-by: Eric Dumazet Cc: Kris Pan Link: https://patch.msgid.link/20260923130318.607255-1-edumazet@google.com Signed-off-by: Jakub Kicinski --- net/core/dev.c | 47 +++++++++++++++++++++++++++++++++++------------ net/core/dev.h | 2 ++ net/core/skbuff.c | 12 +++++++++--- 3 files changed, 46 insertions(+), 15 deletions(-) diff --git a/net/core/dev.c b/net/core/dev.c index 0292a16e16c2..f660fccfc0db 100644 --- a/net/core/dev.c +++ b/net/core/dev.c @@ -5376,7 +5376,8 @@ void kick_defer_list_purge(unsigned int cpu) backlog_unlock_irq_restore(sd, flags); } else if (!cmpxchg(&sd->defer_ipi_scheduled, 0, 1)) { - smp_call_function_single_async(cpu, &sd->defer_csd); + if (smp_call_function_single_async(cpu, &sd->defer_csd)) + WRITE_ONCE(sd->defer_ipi_scheduled, 0); } } @@ -6900,25 +6901,35 @@ bool napi_complete_done(struct napi_struct *n, int work_done) } EXPORT_SYMBOL(napi_complete_done); -static void skb_defer_free_flush(void) +static void __skb_defer_free_flush(struct skb_defer_node *sdn, int budget) { struct llist_node *free_list; struct sk_buff *skb, *next; + + if (llist_empty(&sdn->defer_list)) + return; + atomic_long_set(&sdn->defer_count, 0); + free_list = llist_del_all(&sdn->defer_list); + + llist_for_each_entry_safe(skb, next, free_list, ll_node) { + prefetch(next); + napi_consume_skb(skb, budget); + } +} + +void skb_defer_node_flush(struct skb_defer_node *sdn) +{ + __skb_defer_free_flush(sdn, 0); +} + +static void skb_defer_free_flush(void) +{ struct skb_defer_node *sdn; int node; for_each_node(node) { sdn = this_cpu_ptr(net_hotdata.skb_defer_nodes) + node; - - if (llist_empty(&sdn->defer_list)) - continue; - atomic_long_set(&sdn->defer_count, 0); - free_list = llist_del_all(&sdn->defer_list); - - llist_for_each_entry_safe(skb, next, free_list, ll_node) { - prefetch(next); - napi_consume_skb(skb, 1); - } + __skb_defer_free_flush(sdn, 1); } } @@ -12897,6 +12908,7 @@ static int dev_cpu_dead(unsigned int oldcpu) struct sk_buff **list_skb; struct sk_buff *skb; unsigned int cpu; + int node; struct softnet_data *sd, *oldsd, *remsd = NULL; local_irq_disable(); @@ -12957,6 +12969,17 @@ static int dev_cpu_dead(unsigned int oldcpu) rps_input_queue_head_incr(oldsd); } + for_each_node(node) + skb_defer_node_flush(per_cpu_ptr(net_hotdata.skb_defer_nodes, + oldcpu) + node); + node = cpu_to_node(oldcpu); + if (node_possible(node) && + !cpumask_intersects(cpumask_of_node(node), cpu_online_mask)) { + for_each_possible_cpu(cpu) + skb_defer_node_flush(per_cpu_ptr(net_hotdata.skb_defer_nodes, + cpu) + node); + } + return 0; } diff --git a/net/core/dev.h b/net/core/dev.h index b757faead4d1..04fb0e9a571e 100644 --- a/net/core/dev.h +++ b/net/core/dev.h @@ -399,6 +399,8 @@ static inline void napi_assert_will_not_race(const struct napi_struct *napi) WARN_ON(READ_ONCE(napi->list_owner) != -1); } +struct skb_defer_node; +void skb_defer_node_flush(struct skb_defer_node *sdn); void kick_defer_list_purge(unsigned int cpu); int dev_set_hwtstamp_phylib(struct net_device *dev, diff --git a/net/core/skbuff.c b/net/core/skbuff.c index b4edbd06655e..c3042d822afa 100644 --- a/net/core/skbuff.c +++ b/net/core/skbuff.c @@ -7360,8 +7360,8 @@ void skb_attempt_defer_free(struct sk_buff *skb) struct skb_defer_node *sdn; unsigned long defer_count; unsigned int defer_max; + int cpu, my_cpu; bool kick; - int cpu; if (static_branch_unlikely(&skb_defer_disable_key)) goto nodefer; @@ -7371,7 +7371,8 @@ void skb_attempt_defer_free(struct sk_buff *skb) goto nodefer; cpu = skb->alloc_cpu; - if (cpu == raw_smp_processor_id() || + my_cpu = raw_smp_processor_id(); + if (cpu == my_cpu || WARN_ON_ONCE(cpu >= nr_cpu_ids) || !cpu_online(cpu)) { nodefer: kfree_skb_napi_cache(skb); @@ -7382,7 +7383,7 @@ nodefer: kfree_skb_napi_cache(skb); DEBUG_NET_WARN_ON_ONCE(skb->destructor); DEBUG_NET_WARN_ON_ONCE(skb_nfct(skb)); - sdn = per_cpu_ptr(net_hotdata.skb_defer_nodes, cpu) + numa_node_id(); + sdn = per_cpu_ptr(net_hotdata.skb_defer_nodes, cpu) + cpu_to_node(my_cpu); defer_max = READ_ONCE(net_hotdata.sysctl_skb_defer_max); defer_count = atomic_long_inc_return(&sdn->defer_count); @@ -7392,6 +7393,11 @@ nodefer: kfree_skb_napi_cache(skb); llist_add(&skb->ll_node, &sdn->defer_list); + if (unlikely(!cpu_online(cpu) || my_cpu != raw_smp_processor_id())) { + skb_defer_node_flush(sdn); + return; + } + /* Send an IPI every time queue reaches half capacity. */ kick = (defer_count - 1) == (defer_max >> 1); From 83769c23fb1879edc916a526ba424285033baf2d Mon Sep 17 00:00:00 2001 From: Eric Dumazet Date: Wed, 23 Sep 2026 14:59:42 +0000 Subject: [PATCH 181/189] gve: DQO: fix header length used by gve_can_send_tso() for UDP GSO gve_can_send_tso() computes how many buffers each segment of a GSO packet would span, and for this it needs the length of the headers that the device replicates in front of every segment. It unconditionally uses skb_tcp_all_headers(), which reads the doff field of the TCP header. SKB_GSO_UDP_L4 packets have no TCP header: tcp_hdrlen() then reads one byte of the UDP payload, and header_len can be anything in [0, 60] instead of the transport offset plus the eight bytes of the UDP header that gve_prep_tso() programs into the TSO context descriptor. A wrong header length shifts all the segment boundaries computed in the loop, so the number of buffers per segment can be over or under estimated. In the first case, GSO is needlessly disabled for this packet by gve_features_check_dqo() and the stack has to segment it. In the second case, the driver hands the device a packet whose segments span more than GVE_TX_MAX_DATA_DESCS buffers. Use the UDP header length for SKB_GSO_UDP_L4 packets, matching what gve_prep_tso() does. Fixes: 014c607f86ab ("gve: add support for UDP GSO for DQO format") Closes: https://lore.kernel.org/netdev/CANn89i+MS4L60sFQ49=-f-mibeveUfcrpVkD5X+Qy6SOnEpd6w@mail.gmail.com/ Signed-off-by: Eric Dumazet Cc: Ankit Garg Cc: Harshitha Ramamurthy Cc: Joshua Washington Cc: Willem de Bruijn Reviewed-by: Ankit Garg Reviewed-by: Harshitha Ramamurthy Link: https://patch.msgid.link/20260923145942.731365-1-edumazet@google.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/google/gve/gve_tx_dqo.c | 8 +++++++- 1 file changed, 7 insertions(+), 1 deletion(-) diff --git a/drivers/net/ethernet/google/gve/gve_tx_dqo.c b/drivers/net/ethernet/google/gve/gve_tx_dqo.c index 80ab0a449ff5..0f6f7c5dbb2e 100644 --- a/drivers/net/ethernet/google/gve/gve_tx_dqo.c +++ b/drivers/net/ethernet/google/gve/gve_tx_dqo.c @@ -918,13 +918,19 @@ static bool gve_can_send_tso(const struct sk_buff *skb) { const int max_bufs_per_seg = GVE_TX_MAX_DATA_DESCS - 1; const struct skb_shared_info *shinfo = skb_shinfo(skb); - const int header_len = skb_tcp_all_headers(skb); const int gso_size = shinfo->gso_size; int cur_seg_num_bufs; int prev_frag_size; int cur_seg_size; + int header_len; int i; + /* Must match the header length programmed by gve_prep_tso(). */ + if (skb_is_gso_tcp(skb)) + header_len = skb_tcp_all_headers(skb); + else + header_len = skb_transport_offset(skb) + sizeof(struct udphdr); + cur_seg_size = skb_headlen(skb) - header_len; prev_frag_size = skb_headlen(skb); cur_seg_num_bufs = cur_seg_size > 0; From 3b430ea6234087957b0d3cd181e3116722b59819 Mon Sep 17 00:00:00 2001 From: Eddie Phillips Date: Thu, 24 Sep 2026 00:42:51 +0000 Subject: [PATCH 182/189] gve: fix TX drop when GSO MSS is too small for hw The device has a strict requirement that the minimum MSS (gso_size) for TSO/GSO packets must be at least 88 bytes. If a packet below this threshold is pushed to the hardware, it can cause hardware to silently drop the packet, leading to increased latency and retransmissions. Currently, this is validated too late in the transmit pipeline (gve_prep_tso), leading to silent drops. Fix this by moving the validation into the .ndo_features_check callback (gve_features_check_dqo). If we detect a GSO packet with a gso_size smaller than GVE_TX_MIN_TSO_MSS_DQO, we clear the GSO feature flags for this packet. Fixes: a57e5de476be ("gve: DQO: Add TX path") Signed-off-by: Eddie Phillips Signed-off-by: Eric Dumazet Reviewed-by: Harshitha Ramamurthy Link: https://patch.msgid.link/20260924004252.1196328-2-edumazet@google.com Signed-off-by: Jakub Kicinski --- drivers/net/ethernet/google/gve/gve_tx_dqo.c | 14 +++----------- 1 file changed, 3 insertions(+), 11 deletions(-) diff --git a/drivers/net/ethernet/google/gve/gve_tx_dqo.c b/drivers/net/ethernet/google/gve/gve_tx_dqo.c index 0f6f7c5dbb2e..e5fe17b04798 100644 --- a/drivers/net/ethernet/google/gve/gve_tx_dqo.c +++ b/drivers/net/ethernet/google/gve/gve_tx_dqo.c @@ -577,17 +577,6 @@ static int gve_prep_tso(struct sk_buff *skb) int header_len; int err; - /* Note: HW requires MSS (gso_size) to be <= 9728 and the total length - * of the TSO to be <= 262143. - * - * However, we don't validate these because: - * - Hypervisor enforces a limit of 9K MTU - * - Kernel will not produce a TSO larger than 64k - */ - - if (unlikely(shinfo->gso_size < GVE_TX_MIN_TSO_MSS_DQO)) - return -1; - /* Needed because we will modify header. */ err = skb_cow_head(skb, 0); if (err < 0) @@ -925,6 +914,9 @@ static bool gve_can_send_tso(const struct sk_buff *skb) int header_len; int i; + if (unlikely(gso_size < GVE_TX_MIN_TSO_MSS_DQO)) + return false; + /* Must match the header length programmed by gve_prep_tso(). */ if (skb_is_gso_tcp(skb)) header_len = skb_tcp_all_headers(skb); From 296c83b5ccc808c080865eb20fd7a477b0355bb7 Mon Sep 17 00:00:00 2001 From: Eric Dumazet Date: Thu, 24 Sep 2026 00:42:52 +0000 Subject: [PATCH 183/189] gve: DQO: reject TSO packets with an out of range MSS gve_prep_tso() notes that the device requires the MSS to be <= 9728, but does not enforce it, assuming the 9K MTU enforced by the hypervisor and the 64KB limit on TSO sizes are enough. This does not hold for packets that were not generated locally. A guest behind a tap, or any packet socket user, can provide an arbitrary gso_size in virtio_net_hdr. Layer 2 forwarding does not check the MTU for GSO packets (is_skb_forwardable()), and gso_features_check() only bounds skb->len and gso_segs, never gso_size. Such a packet reaches gve_tx_fill_tso_ctx_desc(), which puts gso_size into the mss field of the TSO context descriptor. This field is 14 bits wide, so a gso_size of 16384 is silently turned into an MSS of zero. Drop these packets from gve_prep_tso(), and make sure that gve_features_check_dqo() leaves their GSO bits alone: skb_segment() splits at gso_size regardless of the MTU, so falling back to software segmentation would give the device non TSO packets bigger than the 9728 bytes it supports. Note that the device can still be given oversized non TSO packets when the stack segments in software for other reasons, for instance after TSO has been disabled with ethtool. This is a generic issue, because the MTU check is skipped for GSO packets in the forwarding path, and is addressed separately. Fixes: a57e5de476be ("gve: DQO: Add TX path") Signed-off-by: Eric Dumazet Reviewed-by: Harshitha Ramamurthy Link: https://patch.msgid.link/20260924004252.1196328-3-edumazet@google.com Signed-off-by: Jakub Kicinski --- .../net/ethernet/google/gve/gve_desc_dqo.h | 5 ++++ drivers/net/ethernet/google/gve/gve_tx_dqo.c | 26 ++++++++++++++++++- 2 files changed, 30 insertions(+), 1 deletion(-) diff --git a/drivers/net/ethernet/google/gve/gve_desc_dqo.h b/drivers/net/ethernet/google/gve/gve_desc_dqo.h index f7786b03c744..d2c86c8eeae2 100644 --- a/drivers/net/ethernet/google/gve/gve_desc_dqo.h +++ b/drivers/net/ethernet/google/gve/gve_desc_dqo.h @@ -14,6 +14,11 @@ #define GVE_TX_MAX_HDR_SIZE_DQO 255 #define GVE_TX_MIN_TSO_MSS_DQO 88 +/* HW limit. This also has to fit in the 14 bits of the mss field of + * struct gve_tx_tso_context_desc_dqo. + */ +#define GVE_TX_MAX_TSO_MSS_DQO 9728 + #ifndef __LITTLE_ENDIAN_BITFIELD #error "Only little endian supported" #endif diff --git a/drivers/net/ethernet/google/gve/gve_tx_dqo.c b/drivers/net/ethernet/google/gve/gve_tx_dqo.c index e5fe17b04798..616c1921aebe 100644 --- a/drivers/net/ethernet/google/gve/gve_tx_dqo.c +++ b/drivers/net/ethernet/google/gve/gve_tx_dqo.c @@ -577,6 +577,20 @@ static int gve_prep_tso(struct sk_buff *skb) int header_len; int err; + /* Note: HW requires the total length of the TSO to be <= 262143, + * this is enforced by netif_set_tso_max_size(). + * + * MSS (gso_size) can not be trusted: packets forwarded from a tap or + * injected by a packet socket can carry an arbitrary value, while the + * mss field of the TSO context descriptor is only 14 bits wide. + * + * A too big MSS is dropped here instead of being rejected from + * gve_features_check_dqo(), because software segmentation would + * produce packets larger than the device can send. + */ + if (unlikely(shinfo->gso_size > GVE_TX_MAX_TSO_MSS_DQO)) + return -1; + /* Needed because we will modify header. */ err = skb_cow_head(skb, 0); if (err < 0) @@ -964,7 +978,17 @@ netdev_features_t gve_features_check_dqo(struct sk_buff *skb, struct net_device *dev, netdev_features_t features) { - if (skb_is_gso(skb) && !gve_can_send_tso(skb)) + if (!skb_is_gso(skb)) + return features; + + /* Keep the GSO bits for a too big MSS, so that gve_prep_tso() drops + * the packet: software segmentation would give packets larger than + * the device can send. + */ + if (skb_shinfo(skb)->gso_size > GVE_TX_MAX_TSO_MSS_DQO) + return features; + + if (!gve_can_send_tso(skb)) return features & ~NETIF_F_GSO_MASK; return features; From 72b5b9a28b996e09b8b5b944370c79851bb68f52 Mon Sep 17 00:00:00 2001 From: Zixuan Chai Date: Thu, 24 Sep 2026 09:26:05 +0800 Subject: [PATCH 184/189] llc: reserve device headroom for allocated frames llc_alloc_frame() reserves link-layer headroom using the device type. This is insufficient for stacked Ethernet devices such as VLAN devices, where vlan_dev_hard_header() pushes a VLAN header before the lower device's Ethernet header. An LLC response on such a device can therefore underflow skb headroom in eth_header(). Use LL_RESERVED_SPACE() to account for the device's actual required headroom while preserving the existing LLC device-type check. Fixes: bf9ae5386bca ("llc: use dev_hard_header") Cc: stable@vger.kernel.org Reported-by: VEGA Signed-off-by: Zixuan Chai Signed-off-by: Ren Wei Reviewed-by: Eric Dumazet Link: https://patch.msgid.link/20260924012613.2533934-1-weir@nebusec.ai Signed-off-by: Jakub Kicinski --- net/llc/llc_sap.c | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/net/llc/llc_sap.c b/net/llc/llc_sap.c index 1bd446a21092..3904a1b4ba84 100644 --- a/net/llc/llc_sap.c +++ b/net/llc/llc_sap.c @@ -19,12 +19,12 @@ #include #include -static int llc_mac_header_len(unsigned short devtype) +static int llc_mac_header_len(struct net_device *dev) { - switch (devtype) { + switch (dev->type) { case ARPHRD_ETHER: case ARPHRD_LOOPBACK: - return sizeof(struct ethhdr); + return LL_RESERVED_SPACE(dev); } return 0; } @@ -45,7 +45,7 @@ struct sk_buff *llc_alloc_frame(struct sock *sk, struct net_device *dev, int hlen = type == LLC_PDU_TYPE_U ? 3 : 4; struct sk_buff *skb; - hlen += llc_mac_header_len(dev->type); + hlen += llc_mac_header_len(dev); skb = alloc_skb(hlen + data_size, GFP_ATOMIC); if (skb) { From 72f9dd522f8d6c5a00be9695c7bb74631eb5069e Mon Sep 17 00:00:00 2001 From: Eric Dumazet Date: Thu, 24 Sep 2026 08:29:48 +0000 Subject: [PATCH 185/189] llc: fix skb UAF and leaks on llc_mac_hdr_init() failure In llc_conn_ac_resend_i_xxx_x_set_0_or_send_rr(), if llc_mac_hdr_init() fails, kfree_skb(skb) is called instead of kfree_skb(nskb). This leaks the newly allocated nskb, reads from the freed skb via LLC_I_GET_NR(pdu), and double-frees skb when llc_conn_state_process() drops its reference. In llc_sap_action_send_xid_r() and llc_sap_action_send_test_r(), nskb is leaked if llc_mac_hdr_init() returns an error. Free nskb in all three error paths. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Closes: https://lore.kernel.org/netdev/179022851638.2160803.1808206741379444999@kernel.org/ Signed-off-by: Eric Dumazet Link: https://patch.msgid.link/20260924082951.1599377-2-edumazet@google.com Signed-off-by: Jakub Kicinski --- net/llc/llc_c_ac.c | 2 +- net/llc/llc_s_ac.c | 4 ++++ 2 files changed, 5 insertions(+), 1 deletion(-) diff --git a/net/llc/llc_c_ac.c b/net/llc/llc_c_ac.c index 724ecd741d4c..1aa7fe28acdd 100644 --- a/net/llc/llc_c_ac.c +++ b/net/llc/llc_c_ac.c @@ -437,7 +437,7 @@ int llc_conn_ac_resend_i_xxx_x_set_0_or_send_rr(struct sock *sk, if (likely(!rc)) llc_conn_send_pdu(sk, nskb); else - kfree_skb(skb); + kfree_skb(nskb); } if (rc) { nr = LLC_I_GET_NR(pdu); diff --git a/net/llc/llc_s_ac.c b/net/llc/llc_s_ac.c index 98deee560373..831998211b52 100644 --- a/net/llc/llc_s_ac.c +++ b/net/llc/llc_s_ac.c @@ -121,6 +121,8 @@ int llc_sap_action_send_xid_r(struct llc_sap *sap, struct sk_buff *skb) rc = llc_mac_hdr_init(nskb, mac_sa, mac_da); if (likely(!rc)) rc = dev_queue_xmit(nskb); + else + kfree_skb(nskb); out: return rc; } @@ -170,6 +172,8 @@ int llc_sap_action_send_test_r(struct llc_sap *sap, struct sk_buff *skb) rc = llc_mac_hdr_init(nskb, mac_sa, mac_da); if (likely(!rc)) rc = dev_queue_xmit(nskb); + else + kfree_skb(nskb); out: return rc; } From ac704ff08e511c87643799c385f55ecd69b85e03 Mon Sep 17 00:00:00 2001 From: Eric Dumazet Date: Thu, 24 Sep 2026 08:29:49 +0000 Subject: [PATCH 186/189] bridge: check llc_mac_hdr_init() return value in br_send_bpdu() If llc_mac_hdr_init() fails (for instance if the port device type does not support LLC or dev_hard_header() fails), br_send_bpdu() should drop the skb instead of resetting the mac header to the LLC payload and transmitting a malformed frame. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Closes: https://lore.kernel.org/netdev/179022851638.2160803.1808206741379444999@kernel.org/ Cc: Nikolay Aleksandrov Cc: Ido Schimmel Cc: bridge@lists.linux.dev Signed-off-by: Eric Dumazet Acked-by: Nikolay Aleksandrov Link: https://patch.msgid.link/20260924082951.1599377-3-edumazet@google.com Signed-off-by: Jakub Kicinski --- net/bridge/br_stp_bpdu.c | 5 ++++- 1 file changed, 4 insertions(+), 1 deletion(-) diff --git a/net/bridge/br_stp_bpdu.c b/net/bridge/br_stp_bpdu.c index 74ec42ba1e7d..21d092f5acbb 100644 --- a/net/bridge/br_stp_bpdu.c +++ b/net/bridge/br_stp_bpdu.c @@ -52,7 +52,10 @@ static void br_send_bpdu(struct net_bridge_port *p, LLC_SAP_BSPAN, LLC_PDU_CMD); llc_pdu_init_as_ui_cmd(skb); - llc_mac_hdr_init(skb, p->dev->dev_addr, p->br->group_addr); + if (llc_mac_hdr_init(skb, p->dev->dev_addr, p->br->group_addr)) { + kfree_skb(skb); + return; + } skb_reset_mac_header(skb); From 907b978e82cb4c1c245fc2985bb27c5d5c88c8f6 Mon Sep 17 00:00:00 2001 From: Eric Dumazet Date: Thu, 24 Sep 2026 08:29:50 +0000 Subject: [PATCH 187/189] net/sched: sch_teql: fix shadowed err in __teql_resolve() __teql_resolve() declares an inner 'int err;' inside the 'if (neigh_event_send(n, skb_res) == 0)' block, shadowing the outer 'int err = 0;'. As a result, a negative return from dev_hard_header() is written to the inner variable and __teql_resolve() still returns 0. Remove the shadowed variable and set the outer err to -EINVAL when dev_hard_header() returns a negative error. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Closes: https://lore.kernel.org/netdev/179022851638.2160803.1808206741379444999@kernel.org/ Cc: Jamal Hadi Salim Cc: Jiri Pirko Signed-off-by: Eric Dumazet Link: https://patch.msgid.link/20260924082951.1599377-4-edumazet@google.com Signed-off-by: Jakub Kicinski --- net/sched/sch_teql.c | 7 ++----- 1 file changed, 2 insertions(+), 5 deletions(-) diff --git a/net/sched/sch_teql.c b/net/sched/sch_teql.c index 9e52afc2d980..409ce50cc0db 100644 --- a/net/sched/sch_teql.c +++ b/net/sched/sch_teql.c @@ -265,14 +265,11 @@ __teql_resolve(struct sk_buff *skb, struct sk_buff *skb_res, } if (neigh_event_send(n, skb_res) == 0) { - int err; char haddr[MAX_ADDR_LEN]; neigh_ha_snapshot(haddr, n, dev); - err = dev_hard_header(skb, dev, ntohs(skb_protocol(skb, false)), - haddr, NULL, skb->len); - - if (err < 0) + if (dev_hard_header(skb, dev, ntohs(skb_protocol(skb, false)), + haddr, NULL, skb->len) < 0) err = -EINVAL; } else { err = (skb_res == NULL) ? -EAGAIN : 1; From cd5dd68267c4238795fadaf02b3575ca3f8a6500 Mon Sep 17 00:00:00 2001 From: Eric Dumazet Date: Thu, 24 Sep 2026 08:29:51 +0000 Subject: [PATCH 188/189] vlan: ensure sufficient headroom in vlan_dev_hard_header() Callers that only reserve ETH_HLEN or less (such as llc_alloc_frame()), or skbs allocated before dynamic device/headroom changes (e.g. toggling VLAN_FLAG_REORDER_HDR or bonding/team switching slaves), can reach vlan_dev_hard_header() with insufficient headroom and trigger skb_under_panic(). Use skb_cow_head() in vlan_dev_hard_header() when VLAN_FLAG_REORDER_HDR is not set to ensure sufficient headroom for the VLAN header(s) and the underlying device hard header. Use READ_ONCE() to read dev->hard_header_len and dev->needed_headroom as they can be updated concurrently under RTNL (e.g. in vlan_transfer_features()) while vlan_dev_hard_header() runs locklessly on the transmit path. Also avoid LL_RESERVED_SPACE(dev) here so that the extra HH_DATA_MOD alignment padding does not trigger unnecessary pskb_expand_head() reallocations on inner stacked VLAN devices after the outer VLAN header has been pushed. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Reported-by: Zixuan Chai Closes: https://lore.kernel.org/netdev/cover.1789987105.git.petalzu987@gmail.com/ Link: https://lore.kernel.org/netdev/179022851638.2160803.1808206741379444999@kernel.org/ Cc: Hangbin Liu Signed-off-by: Eric Dumazet Link: https://patch.msgid.link/20260924082951.1599377-5-edumazet@google.com Signed-off-by: Jakub Kicinski --- net/8021q/vlan_dev.c | 5 +++++ 1 file changed, 5 insertions(+) diff --git a/net/8021q/vlan_dev.c b/net/8021q/vlan_dev.c index 2859cbac3f26..c949c6a82945 100644 --- a/net/8021q/vlan_dev.c +++ b/net/8021q/vlan_dev.c @@ -55,6 +55,11 @@ static int vlan_dev_hard_header(struct sk_buff *skb, struct net_device *dev, int rc; if (!(vlan->flags & VLAN_FLAG_REORDER_HDR)) { + unsigned int hlen = READ_ONCE(dev->hard_header_len) + + READ_ONCE(dev->needed_headroom); + + if (skb_cow_head(skb, hlen) < 0) + return -ENOMEM; vhdr = skb_push(skb, VLAN_HLEN); vlan_tci = vlan->vlan_id; From fc6d80eb504458d6416b75a94188b268c95c6533 Mon Sep 17 00:00:00 2001 From: Willem de Bruijn Date: Thu, 24 Sep 2026 11:44:12 -0400 Subject: [PATCH 189/189] tcp: prevent collapsing skbs across boundary in rtx queue tcp_write_collapse_fence() sets TCP_SKB_CB(skb)->eor = 1 on tcp_write_queue_tail(sk) to prevent skbs queued after a switch to device encryption from being collapsed into earlier skbs. The fence is a no-op if all earlier data has already been transmitted when the switch happens: sk->sk_write_queue is empty. The not yet acknowledged earlier skbs wait in sk->tcp_rtx_queue with eor 0. On a subsequent retransmit or SACK shift, tcp_retrans_try_collapse() or tcp_shift_skb_data() can then merge an skb queued after the switch into one queued before it. Both users of the fence are affected: - psp: devices only encrypt skbs with skb->decrypted set. The merged skb keeps decrypted = 0 from the earlier skb, so merged data sent after psp_sock_assoc_set_tx() is retransmitted in cleartext. - tls device offload: the merged skb straddles the start marker set in tls_set_device_offload(). The software fallback (fill_sg_in() returns -EINVAL) and the mlx5, nfp and funeth drivers cannot handle such an skb and drop it. Every retransmit rebuilds the same skb, so the connection stalls. Fix this in two places, for defense in depth: 1. Fall back to tcp_rtx_queue_tail(sk) in tcp_write_collapse_fence() when tcp_write_queue_tail(sk) is NULL. 2. Check !skb_cmp_decrypted(to, from) in tcp_skb_can_collapse(), as tcp_skb_can_collapse_rx() does on receive. skb_shift(), which both collapse paths call, already has a DEBUG_NET_WARN_ON_ONCE() for this condition. Fixes: e8f69799810c ("net/tls: Add generic NIC offload infrastructure") Cc: stable@vger.kernel.org Signed-off-by: Willem de Bruijn Reviewed-by: Eric Dumazet Reviewed-by: Daniel Zahka Link: https://patch.msgid.link/20260924154427.953800-1-willemdebruijn.kernel@gmail.com Signed-off-by: Jakub Kicinski --- include/net/tcp.h | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/include/net/tcp.h b/include/net/tcp.h index 436495ff2271..4416cdf9bf30 100644 --- a/include/net/tcp.h +++ b/include/net/tcp.h @@ -1232,9 +1232,9 @@ static inline bool tcp_skb_can_collapse_to(const struct sk_buff *skb) static inline bool tcp_skb_can_collapse(const struct sk_buff *to, const struct sk_buff *from) { - /* skb_cmp_decrypted() not needed, use tcp_write_collapse_fence() */ return likely(tcp_skb_can_collapse_to(to) && mptcp_skb_can_collapse(to, from) && + !skb_cmp_decrypted(to, from) && skb_pure_zcopy_same(to, from) && skb_frags_readable(to) == skb_frags_readable(from)); } @@ -2327,7 +2327,7 @@ static inline void tcp_rtx_queue_unlink_and_free(struct sk_buff *skb, struct soc static inline void tcp_write_collapse_fence(struct sock *sk) { - struct sk_buff *skb = tcp_write_queue_tail(sk); + struct sk_buff *skb = tcp_write_queue_tail(sk) ?: tcp_rtx_queue_tail(sk); if (skb) TCP_SKB_CB(skb)->eor = 1;