While calling the helpers, a raw pointer to the extensions area is
wired into expectations list:
-> nf_ct_helper()
-> helper->help()
-> nf_ct_expect_related_report()
-> nf_ct_expect_insert()
-> hlist_add_head_rcu(&exp->lnode, &master_help->expectations)
In case the connection is not confirmed yet, more extensions can be
added afterwards with *_ext_add() calls reallocating the extension
space and leaving the now invalid pointer in the expectations list
that is later accessed while removing the expectation.
Make sure that helpers are called at the end after all the other
extensions are already added.
Note that the helper rejection now leaves the mark and labels set,
but that's not different from how the NAT was handled before or how
the mark and the labels were handled on confirmation failure. And
there are no atomicity guarantees provided by the API anyway.
Fixes: a21b06e731 ("net: sched: add helper support in act_ct")
Cc: stable@vger.kernel.org
Reported-by: Axel Mierczuk <axel.mierczuk@1password.com>
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Xin Long <lucien.xin@gmail.com>
Reviewed-by: Jamal Hadi Salim <jhs@mojatatu.com>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Link: https://patch.msgid.link/20260921145655.3167436-7-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
This variable can only become 'true' when the connection is not
confirmed, but it is only checked when it is confirmed. So, it can be
treated as being always false and just removed.
Fixes: a21b06e731 ("net: sched: add helper support in act_ct")
Cc: stable@vger.kernel.org
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Reviewed-by: Xin Long <lucien.xin@gmail.com>
Reviewed-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/20260921145655.3167436-6-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
In a case where skb with an unconfirmed ct entry gets cloned, we may
end up processing both again but with different sets of extensions.
The series of events:
1. The first clone wants to commit and runs the helpers wiring up
the extension pointer into the expectation list.
2. Then it looses the confirmation keeping the entry unconfirmed.
3. Second clone now wants to commit labels or run NAT and adds the
new extension for that breaking the pointer in the expectation
list causing UAF on the destruction path later.
While this is possible to trigger, there should be no practical
network pipeline where we need to process both clones without
modifications in the same zone. So, let's just reset the entry in
case for some reason we got an skb with a shared one. This doesn't
affect any known use cases, but avoids any potential problems with
sharing and modification of the unconfirmed ct entry.
Unlike openvswitch module, act_ct allows for NAT without commit.
Changing that would be a uAPI break. So, act_ct needs to reset on NAT
regardless of the commit flag to avoid reallocation of the extension
space. This, however, doesn't really change the picture for sensible
networking cases as there should be no need to run the same packet
twice (before and after the clone) through conntrack without packet
header or zone changes and without commit.
The fixes tag points to the introduction of helpers, since that's the
main UAF trigger for the sharing.
Fixes: a21b06e731 ("net: sched: add helper support in act_ct")
Cc: stable@vger.kernel.org
Reported-by: Axel Mierczuk <axel.mierczuk@1password.com>
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Reviewed-by: Xin Long <lucien.xin@gmail.com>
Reviewed-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/20260921145655.3167436-5-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
While calling the helpers, a raw pointer to the extensions area is
wired into expectations list:
-> nf_ct_helper()
-> helper->help()
-> nf_ct_expect_related_report()
-> nf_ct_expect_insert()
-> hlist_add_head_rcu(&exp->lnode, &master_help->expectations)
In case the connection is not confirmed yet, more extensions can be
added afterwards with *_ext_add() calls reallocating the extension
space and leaving the now invalid pointer in the expectations list
that is later accessed while removing the expectation.
Make sure that helpers are called at the end after all the other
extensions are already added.
Note that the helper rejection now leaves the mark and labels set,
but that's not different from how the NAT was handled before or how
the mark and the labels were handled on confirmation failure. And
there are no atomicity guarantees provided by the API anyway.
Fixes: cae3a26275 ("openvswitch: Allow attaching helpers to ct action")
Cc: stable@vger.kernel.org
Reported-by: Axel Mierczuk <axel.mierczuk@1password.com>
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Link: https://patch.msgid.link/20260921145655.3167436-4-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
This variable can only become 'true' when the connection is not
confirmed, but it is only checked when it is confirmed. So, it can be
treated as being always false and just removed.
Fixes: 3c1860543f ("openvswitch: add nf_ct_is_confirmed check before assigning the helper")
Cc: stable@vger.kernel.org
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Link: https://patch.msgid.link/20260921145655.3167436-3-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
In a case where skb with an unconfirmed ct entry gets cloned, we may
end up committing both but with different sets of extensions.
The series of events:
1. The first clone wants to commit and runs the helpers wiring up
the extension pointer into the expectation list.
2. Then it looses the confirmation keeping the entry unconfirmed.
3. Second clone now wants to commit labels and adds the new extension
for that breaking the pointer in the expectation list causing
UAF on the destruction path later.
While this is possible to trigger, there should be no practical
network pipeline where committing both clones without modifications
into the same zone is needed. So, let's just reset the entry in case
for some reason we got an skb with a shared one during commit. This
doesn't affect any known use cases, but avoids any potential problems
with sharing and modification of the unconfirmed ct entry.
The fixes tag points to the introduction of helpers, since that's the
main UAF trigger for the sharing.
Fixes: cae3a26275 ("openvswitch: Allow attaching helpers to ct action")
Cc: stable@vger.kernel.org
Reported-by: Axel Mierczuk <axel.mierczuk@1password.com>
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Link: https://patch.msgid.link/20260921145655.3167436-2-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
When tcp_send_synack() replaces the cloned SYN skb at the head of the
retransmit queue with a copy, it frees the original with
tcp_rtx_queue_unlink_and_free() and only repairs tp->highest_sack.
tp->retransmit_skb_hint keeps pointing at the freed
skbuff_fclone_cache object.
The dangling hint is read in tcp_verify_retransmit_hint() and used as
the root of the rbtree walk in tcp_xmit_retransmit_queue(). An
unprivileged TFO client (sendmsg(MSG_FASTOPEN)) can arm the hint with
an attacker-supplied ICMP fragmentation-needed message, after which a
simultaneous open frees the armed SYN skb:
BUG: KASAN: slab-use-after-free in tcp_mark_skb_lost (net/ipv4/tcp_input.c:1316)
Read of size 4 at addr ffff88800604d928 by task swapper/1/0
Call Trace:
tcp_mark_skb_lost (net/ipv4/tcp_input.c:1316)
tcp_simple_retransmit (net/ipv4/tcp_input.c:3158)
tcp_v4_err (net/ipv4/tcp_ipv4.c:587)
Sync the hint to the copy.
Fixes: c31b70c996 ("tcp: Add logic to check for SYN w/ data in tcp_simple_retransmit")
Reported-by: Kimi Security Team <bug-report@moonshot.ai>
Tested-by: Weiming Shi <shiweiming@moonshot.ai>
Signed-off-by: Yilin Zhang <yilinzhang@moonshot.ai>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/8a9dff4063a2745653b7e88ceb745d75efa16e68.1790224474.git.yilinzhang@moonshot.ai
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
When a socket transmits a packet with MCTP_TAG_PREALLOC set,
mctp_lookup_prealloc_tag() iterates over the per-netns &mns->keys list
and matches netid, req_tag, peer_addr, and manual_alloc, without
checking whether tmp->sk == &msk->sk. This allows any MCTP socket in the
same network namespace to use and consume another socket's preallocated
tag.
Iterate the socket's own tag list (&msk->keys via sklist) instead of the
namespace-wide &mns->keys list in mctp_lookup_prealloc_tag(), ensuring
that only tags allocated by msk are matched.
Tested in QEMU against Linux 7.3.0-rc3 by allocating a manual tag
(0x18) on socket A via SIOCMCTPALLOCTAG for peer EID 9 and sending a
4-byte message with MCTP_TAG_PREALLOC from socket B in the same network
namespace. On the unfixed kernel, sendto(sock_b) using socket A's
preallocated tag succeeds (ret = 4); with this patch applied,
sendto(sock_b) fails with -ENOENT (errno = 2) while sendto(sock_a)
succeeds (ret = 4).
Fixes: 63ed1aab3d ("mctp: Add SIOCMCTP{ALLOC,DROP}TAG ioctls for tag control")
Suggested-by: Jeremy Kerr <jk@codeconstruct.com.au>
Cc: stable@vger.kernel.org
Signed-off-by: Hui Peng <benquike@gmail.com>
Link: https://patch.msgid.link/20260921051002.1656692-1-benquike@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
skbmark_decode(), skbprio_decode() and skbtcindex_decode() read fixed-size
values from the TLV payload without validating its length.
A malformed IFE frame can declare a shorter payload, causing the decoders
to consume bytes beyond the declared metadata value:
[TLV type=IFE_META_SKBMARK len=4]
-> dlen == 0, but decode reads 4 bytes
The decoder may therefore set skb metadata from unintended input.
Validate the payload length before decoding and return -EINVAL for
invalid lengths. Read the values with get_unaligned_be32() and
get_unaligned_be16(), as TLV payloads are not guaranteed to be
aligned. Teach tcf_ife_decode() to log a decoder error separately
from an unknown metaid; both are counted as overlimits and decoding
continues with the remaining metadata.
The metadata length issue was found by an automated audit of the IFE
decode path at v6.18-rc7 and reproduced with a userspace sanitizer
model of the decode path. Compile-tested on x86_64 with defconfig and
NET_ACT_IFE=y: act_ife.o and the three act_meta_*.o build
warning-free.
Fixes: 084e2f6566 ("Support to encoding decoding skb mark on IFE action")
Fixes: 200e10f469 ("Support to encoding decoding skb prio on IFE action")
Fixes: 408fbc22ef ("net sched ife action: Introduce skb tcindex metadata encap decap")
Assisted-by: Hawkeye:GLM-5.3-flash
Assisted-by: Qoder:Qwen3.8-Max
Signed-off-by: Fang Xieyan <fangxy@xiaopeng.com>
Link: https://patch.msgid.link/20260921125441.81459-1-fangxy@xiaopeng.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
rds_ib_map_frmr() stores the caller's scatterlist in the MR before DMA
mapping and registration can fail. On failure, __rds_rdma_map() unpins
the pages and frees the scatterlist, but rds_ib_free_frmr() can still
return the MR to the pool with the stale pointer set.
This leaves the pool with a dangling scatterlist and can lead to local
privilege escalation. KASAN detects the resulting use-after-free when the
MR is later torn down:
BUG: KASAN: slab-use-after-free in __rds_ib_teardown_mr
Read of size 8
Call Trace:
__rds_ib_teardown_mr
rds_ib_unreg_frmr
rds_ib_flush_mr_pool
rds_ib_flush_mrs
rds_free_mr
rds_setsockopt
Store the scatterlist in the MR only after DMA mapping succeeds. If DMA
mapping fails, return directly while the MR fields remain clear; the caller
keeps ownership of the scatterlist and its pinned pages. If a later
registration step fails, unmap the scatterlist and clear the MR fields
before returning.
Fixes: 1659185fb4 ("RDS: IB: Support Fastreg MR (FRMR) memory registration mode")
Cc: stable@vger.kernel.org
Signed-off-by: Dongliang Qin <cccccccccccc777777@gmail.com>
Reviewed-by: Allison Henderson <achender@kernel.org>
Link: https://patch.msgid.link/20260922031546.3874605-1-cccccccccccc777777@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
ic_dhcp_init_options() appends the hostname (option 12), vendor-class
(option 60) and client-ID (option 61) options into the fixed 312-byte
bootp_pkt.exten[] buffer. Only the client-ID branch checked the
remaining space; the hostname and vendor-class writes were unbounded.
A 64-byte hostname together with the maximum 252-byte dhcpclass=
identifier needs 18 + (2 + 64) + (2 + 252) = 338 of the 312 available
bytes even before the terminating END marker, so the vendor-class memcpy
runs past the end of exten[]. With CONFIG_FORTIFY_SOURCE this is
reported as a field-spanning write and, when the kernel is booted with
panic_on_warn=1, aborts boot with a panic.
Route the optional options through a common helper that makes sure the
option, its 2-byte header and the END marker all fit and drops an option
that would not. Configurations with short options keep sending exactly
the same bytes as before.
Fixes: 130c0f47fd ("ipconfig: send host-name in DHCP requests")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Assisted-by: LLM
Signed-off-by: Yuqi Xu <xuyuqiabc@gmail.com>
Reviewed-by: Ren Wei <weir@nebusec.ai>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/7808dfbfa2162dfd0b19f59aff5742d6e0db2abb.1789798023.git.xuyuqiabc@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Mostly security fixes.
fix use-after-free in nfc_get_local_general_bytes
llcp: Fix list corruption / refcount desync in nfc_llcp_recv_dm()
llcp: Fix race condition in accept_queue lifecycle
llcp: drop truncated I/RR/RNR PDUs in nfc_llcp_recv_hdlc()
llcp: fix -ENOMEM on connect with zero-length service name
llcp: fix WKS SAP hijacking via prefix match in nfc_llcp_wks_sap()
llcp: fix sdreq TLV list leak on parse/alloc/send failure
llcp: fix slab-out-of-bounds reads when logging service names
selftests/nci: Fix out-of-bounds store on thread join
selftests: nci: Correct pthread_create return value check
selftests: nci: Fix uninitialized family ID on missing attribute
virtual_ncidev: Add missing ioctl compat handler
nfcmrvl: validate helper command length before pull
pn533: fix OOB read in pn533_acr122_is_rx_frame_valid()
port100: reject frames whose declared length exceeds the received data
st21nfca: validate ISO15693 inventory length
st21nfca: validate received frame size
trf7970a: power down on startup RX gain failure
Signed-off-by: David Heidelberg <david@ixit.cz>
-----BEGIN PGP SIGNATURE-----
iQIzBAABCAAdFiEE13oJz+7cK71TpwR0YAI/xNNJIHIFAmq0M08ACgkQYAI/xNNJ
IHKXbw//cxnjUWiGiCWQKNGIrPQhO8xnxp29/adTfr5cpaCf9XS1Ah742JnsvV0J
jiKwrV2VdhIedF0gCU51pfMdgV4VWWup2KQJbl820m8EaRmaKY6Spao/HxznnXtF
H1JqYir2ISVADjX4hgQpqQIH2Vrz2o2QAkcsY+jjjydxxajmrrQh1SrJeEGom366
eDHAPmO9MTheJZ8gFXtuz4JdoipVWwY0Ta8D3mmVfgX/Jv5ESsPa4Ie4gQvQ1Vua
9bvYSgqrCGi4jtzEOxXNL1KfufjsbW9I64anwhxwjKJ+sL+/GJ0VfjN4Dgq9dnub
H4dUJs9vUCMLY+Ns6Wo+2pJQ49Ults2SdVjH7kvdJ2FiUn+FvoPDLrQY00P/lhaB
MPLrOCFaNVIzDKX98lxAzT/6c6xrNb5lTrkf17XHzcWa2H4eZj30pBaFfYrfiEH5
dlASR5aSy+GuXa8J/i7+jpQ+xl0GEg9tb0GJsvsWTDAftAl8/9CvcZ/TCbjKeb9L
oGeStPcl/i+MB9wGZLOga8lox0DX5PMNkIcp+RXIjVvPDYiCBh1a3lWHtz65Qx4m
9VlNcIpUzN14VUTkWm1i29DVHpW8V5c/eLmoDHzBUFAVSVr2M0CkFwhPYOByqIzt
lIcpLuTLXrAplzIFWMbi6gU8FSRAFX2I+vVr6rkgWIz2+QCPmYM=
=lO2R
-----END PGP SIGNATURE-----
Merge tag 'nfc-7.3-rc5' of https://codeberg.org/linux-nfc/linux
David Heidelberg says:
====================
NFC fixes for net 7.3-rc5
* tag 'nfc-7.3-rc5' of https://codeberg.org/linux-nfc/linux:
nfc: pn533: fix OOB read in pn533_acr122_is_rx_frame_valid()
nfc: llcp: fix slab-out-of-bounds reads when logging service names
nfc: llcp: fix WKS SAP hijacking via prefix match in nfc_llcp_wks_sap()
nfc: llcp: fix -ENOMEM on connect with zero-length service name
nfc: st21nfca: validate ISO15693 inventory length
nfc: fix use-after-free in nfc_get_local_general_bytes
nfc: trf7970a: power down on startup RX gain failure
nfc: llcp: fix sdreq TLV list leak on parse/alloc/send failure
nfc: llcp: drop truncated I/RR/RNR PDUs in nfc_llcp_recv_hdlc()
nfc: virtual_ncidev: Add missing ioctl compat handler
selftests/nci: Fix out-of-bounds store on thread join
selftests: nci: Fix uninitialized family ID on missing attribute
nfc: llcp: Fix race condition in accept_queue lifecycle
selftests: nci: Correct pthread_create return value check
nfc: port100: reject frames whose declared length exceeds the received data
nfc: llcp: Fix list corruption / refcount desync in nfc_llcp_recv_dm()
nfc: st21nfca: validate received frame size
nfc: nfcmrvl: validate helper command length before pull
====================
Link: https://patch.msgid.link/adeaccc1-cc04-4bb9-a28a-61a61d75ba14@ixit.cz
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Only the entries below dev->num_tc are valid in dev->tc_to_txq[], and
dev->prio_tc_map[] may only name classes below it. netdev_set_num_tc()
lowers dev->num_tc without touching either array.
netdev_txq_to_tc() walks all TC_MAX_QUEUE slots and
netdev_get_prio_tc_map() returns the entry as it stands, so a leftover
entry is handed out as a traffic class >= dev->num_tc. Taking that
class from netdev_txq_to_tc(), __netif_set_xps_queue() rejects only a
negative one and indexes an XPS map sized for dev->num_tc classes:
tci = j * num_tc + tc;
RCU_INIT_POINTER(new_dev_maps->attr_map[tci], map);
attr_map[] holds nr_ids * num_tc entries and j runs over the ids named
in the mask, so a class that is not below num_tc pushes tci past the end
of the map for the last ids and the store overruns it.
Any caller that lowers num_tc leaves such entries behind, and
mqprio_destroy() tears down with netdev_set_num_tc(dev, 0) rather than
netdev_reset_tc(). After mqprio with 8 classes then 1, tc_to_txq[1..7]
still describe txq 1..7. The splat is from an XPS write to txq 2 on a
veth with 8 rx queues: attr_map[] has 8 * 1 entries, tci = j + 2, and
j == 6 stores one past the end of the 88-byte map:
BUG: KASAN: slab-out-of-bounds in __netif_set_xps_queue (net/core/dev.c:2954)
Write of size 8 at addr ffff88813016bc58 by task xps_oob/634
__netif_set_xps_queue (net/core/dev.c:2954)
xps_rxqs_store (net/core/net-sysfs.c:1880)
netdev_queue_attr_store (net/core/net-sysfs.c:1390)
Allocated by task 634:
__kmalloc_noprof (mm/slub.c:5439)
__netif_set_xps_queue (net/core/dev.c:2937)
The buggy address is located 0 bytes to the right of
allocated 88-byte region [ffff88813016bc00, ffff88813016bc58)
Reject a class the map has no room for.
Fixes: 184c449f91 ("net: Add support for XPS with QoS via traffic classes")
Signed-off-by: Norbert Szetei <norbert@doyensec.com>
Link: https://patch.msgid.link/162DD16F-54C6-444A-9E09-0B8CB3D591F2@doyensec.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
ipgre_netlink_parms() can enable collect_md on an existing GRE, GRETAP
or ERSPAN device. Unlike newlink, changelink does not enforce metadata
tunnel uniqueness. Converting a non-metadata device can therefore
replace the metadata receive entry for another device of the same type
in the same netns. Deleting either device then clears the shared entry,
breaking metadata receive lookup for the surviving device.
If parameter validation fails after collect_md is set, deleting the
modified device can also clear an entry it never owned.
Reject enabling metadata mode in both changelink callbacks before any
encapsulation or tunnel parameters are modified. Allow requests that
repeat the metadata attribute on an existing metadata device.
Fixes: 2e15ea390e ("ip_gre: Add support to collect tunnel metadata.")
Signed-off-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Link: https://patch.msgid.link/20260921031859.9283-1-xuanqiang.luo@linux.dev
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The ARP ioctl copies a user-provided struct arpreq into a stack object. Its
arp_dev field may contain IFNAMSIZ bytes without a NUL terminator.
Such input is passed to dev_get_by_name_rcu() or __dev_get_by_name(), where
strcmp() can read past the end of the stack object when a matching
alternative interface name exists.
Terminate the field before the lookup to prevent the out-of-bounds read.
Fixes: 36fbf1e52b ("net: rtnetlink: add linkprop commands to add and delete alternative ifnames")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Signed-off-by: Zijie Huang <milkory@outlook.com>
Signed-off-by: Ren Wei <weir@nebusec.ai>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/fabf02a70787d17299e4b3153eadffaf20d154b3.1789910973.git.milkory@outlook.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Commit 7a9bc9e3f4 ("fou: Don't allow 0 for FOU_ATTR_IPPROTO.") added
NLA_POLICY_MIN(NLA_U8, 1) to fou_nl_policy[FOU_ATTR_IPPROTO], which
rejects an explicitly supplied FOU_ATTR_IPPROTO == 0 attribute with
-ERANGE.
However, FOU_ATTR_IPPROTO is an optional netlink attribute. When a user
sends FOU_CMD_ADD with FOU_ATTR_TYPE set to FOU_ENCAP_DIRECT and omits
FOU_ATTR_IPPROTO entirely, nla_policy validation succeeds and
parse_nl_config() leaves cfg->protocol as 0 (from memset(cfg, 0,
sizeof(*cfg))). fou_create() then creates a FOU_ENCAP_DIRECT socket with
fou->protocol == 0.
In fou_udp_recv(), returning -fou->protocol to udp_queue_rcv_one_skb()
triggers IP protocol resubmission when fou->protocol > 0, whereas
returning 0 tells the UDP tunnel layer that the skb was consumed without
freeing it. When fou->protocol == 0, every packet received on the socket
returns 0 from fou_udp_recv() and leaks the sk_buff.
Reject FOU_ENCAP_DIRECT when !cfg->protocol in fou_create() so that
creating a direct encapsulation port without FOU_ATTR_IPPROTO fails with
-EINVAL while leaving FOU_CMD_DEL and FOU_CMD_GET (which share
parse_nl_config()) unaffected.
Tested in QEMU against Linux 7.3.0-rc3 by sending a FOU_CMD_ADD Generic
Netlink request with FOU_ATTR_PORT = 5555 and FOU_ATTR_TYPE =
FOU_ENCAP_DIRECT while omitting FOU_ATTR_IPPROTO. On the unfixed kernel,
FOU_CMD_ADD succeeds (err = 0), FOU_CMD_GET reports fou->type = 1 and
fou->protocol = 0, and sending 4000 UDP packets to 127.0.0.1:5555 leaks
all 4000 sk_buffs (SUnreclaim in /proc/meminfo grows from 41456 kB to
59008 kB, +17552 kB); with this patch applied, FOU_CMD_ADD is rejected
with -EINVAL (-22).
Fixes: 23461551c0 ("fou: Support for foo-over-udp RX path")
Fixes: 7a9bc9e3f4 ("fou: Don't allow 0 for FOU_ATTR_IPPROTO.")
Cc: stable@vger.kernel.org
Signed-off-by: Hui Peng <benquike@gmail.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Link: https://patch.msgid.link/20260921045920.1613098-1-benquike@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
In seg6_genl_policy, SEG6_ATTR_DST is defined with .type = NLA_BINARY and
.len = sizeof(struct in6_addr). For NLA_BINARY, .len only enforces the
maximum payload length and permits shorter payloads (e.g., 0 bytes).
When seg6_genl_set_tunsrc() copies sizeof(struct in6_addr) bytes via
kmemdup(val, sizeof(*val), GFP_KERNEL), a short SEG6_ATTR_DST attribute
triggers a 16-byte out-of-bounds read past skb->tail into uninitialized
skb->head memory, which is stored in sdata->tun_src and leaked back to
userspace via SEG6_CMD_GET_TUNSRC.
Switch SEG6_ATTR_DST in seg6_genl_policy to
NLA_POLICY_EXACT_LEN(sizeof(struct in6_addr)) so that generic netlink
validation rejects any attribute whose length is not exactly
sizeof(struct in6_addr) with -ERANGE.
Tested in QEMU against Linux 7.3.0-rc3 by sending a SEG6_CMD_SET_TUNSRC
Generic Netlink message with a 0-byte SEG6_ATTR_DST attribute followed
by SEG6_CMD_GET_TUNSRC. On the unfixed kernel, SEG6_CMD_SET_TUNSRC
succeeds (err = 0) and SEG6_CMD_GET_TUNSRC leaks 16 bytes of
uninitialized kernel heap memory (tun_src =
836a61ecc4d25a1042a8d60411cfb378); with this patch applied,
SEG6_CMD_SET_TUNSRC is rejected by netlink policy validation with
-ERANGE (-34) and tun_src remains zeroed.
Fixes: 915d7e5e59 ("ipv6: sr: add code base for control plane support of SR-IPv6")
Cc: stable@vger.kernel.org
Signed-off-by: Hui Peng <benquike@gmail.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Reviewed-by: Justin Iurman <justin.iurman@gmail.com>
Reviewed-by: Andrea Mayer <andrea.mayer@uniroma2.it>
Link: https://patch.msgid.link/20260921044025.1535982-1-benquike@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
skb_maybe_pull_tail() subtracts skb_headlen(skb) from the unsigned max
argument and passes the result to __pskb_pull_tail() as a signed int. The
function does not ensure that max is at least skb_headlen(skb).
This can happen while parsing IPv6 extension headers when an skb already
has a linear area larger than MAX_IPV6_HDR_LEN. Once the parser needs data
beyond the linear area, max - skb_headlen(skb) wraps and is converted to a
negative delta. __pskb_pull_tail() then passes that negative length to
skb_copy_bits(), where it can become a very large copy length.
Pass the requested length itself as the pull bound at the three
extension-header call sites, so the delta can no longer go negative.
Fixes: 1431fb31ec ("xen-netback: fix fragment detection in checksum setup")
Suggested-by: Eric Dumazet <edumazet@google.com>
Signed-off-by: Shihuang Liu <shlomojune6@gmail.com>
Link: https://patch.msgid.link/20260919133604.50948-1-shlomojune6@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
br_mdb_flush_pgs() keeps a pointer-to-pointer cursor while walking
mp->ports. br_multicast_del_pg() can re-enter the same MDB entry through
br_multicast_sg_del_exclude_ports() and unlink other port groups. If the
cursor points into one of those groups, the next iteration dereferences a
stale cursor and can leave mp->ports pointing at freed memory.
A following RTM_GETMDB exposes the dangling pointer:
BUG: KASAN: slab-use-after-free in br_mdb_dump
Read of size 8
br_mdb_dump
rtnl_mdb_dump
rtnl_dumpit
netlink_dump
Reset the cursor to mp->ports after every deletion. The deletion removes at
least the selected group, so the restarted walk always makes progress.
Fixes: a6acb535af ("bridge: mdb: Add MDB bulk deletion support")
Cc: stable@vger.kernel.org
Signed-off-by: Fourie Zhang <fouriezhang@tencent.com>
Acked-by: Nikolay Aleksandrov <razor@blackwall.org>
Link: https://patch.msgid.link/20260920110852.60293-1-fouriezhang@tencent.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
tcf_gate_get_fill_size returns only the TCA_GATE_PARMS size, but
tcf_gate_dump also emits three 64-bit timestamps, the clock id, flags,
priority and the variable-length TCA_GATE_ENTRY_LIST nest. The per-entry
nest is unbounded: parse_gate_list places no cap on the number of
sched-entries, so a gate with many entries can push the real dump well
past the skb that tca_get_fill allocates from this size.
RTM_NEWACTION then fails the add-notify with -EINVAL while the action is
already committed to the IDR, and a subsequent RTM_GETACTION on the
installed gate also returns -EINVAL because its dump no longer fits.
Fix this by accounting for the missing fields in tcf_gate_get_fill_size
along with all elements in the entries list.
Note that sizing the reply from the action lets an oversized gate
install cleanly for the first time: with the input unbounded by
parse_gate_list, the sized skb can now grow well above
NLMSG_GOODSIZE per netlink request (a transient GFP_KERNEL allocation
reachable only with namespace-local CAP_NET_ADMIN). Overload from a
malicious netns admin is hardening material, not net, per the
discussion at
https://lore.kernel.org/netdev/20260914191108.55a1a4f1@kernel.org/;
a follow-up patch for net-next will cap the sched-entry count.
Fixes: 4e76e75d6a ("net sched actions: calculate add/delete event message size")
Reported-by: Sashiko <sashiko-bot@kernel.org>
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260824153903.4143642-1-victor@mojatatu.com
Tested-by: hybris <hybris@mojatatu.ai>
Co-developed-by: Jamal Hadi Salim <jhs@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Signed-off-by: Victor Nogueira <victor@mojatatu.com>
Link: https://patch.msgid.link/QDISC-3BLH.v1.20260914203033@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
nfc_llcp_wks_sap() and nfc_llcp_build_sdreq_tlv() pass non-null-
terminated strings to pr_debug() using the %s format specifier.
The buffers are allocated via kmemdup() or come from netlink
attributes and are not guaranteed to be null-terminated, causing
__dynamic_pr_debug() to read beyond the allocated region:
KASAN: slab-out-of-bounds Read in __dynamic_pr_debug
Fix both call sites by using %.*s with the explicit length to limit
the output to the actual length of the string.
Fixes: d9b8d8e19b ("NFC: llcp: Service Name Lookup netlink interface")
Reported-by: syzbot+1e3df0852e82c21ca418@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=1e3df0852e82c21ca418
Signed-off-by: Ömer Mete Kaya <omermetekaya0@gmail.com>
Link: https://patch.msgid.link/20260908161952.731468-1-omermetekaya0@gmail.com
Signed-off-by: David Heidelberg <david@ixit.cz>
nfc_llcp_wks_sap() compares only service_name_len bytes, so a short
service_name like "u" matches longer WKS strings like "urn:nfc:sn:snep".
Fix by requiring exact length match before strncmp().
Fixes: d646960f79 ("NFC: Initial LLCP support")
Signed-off-by: Ömer Mete Kaya <omermetekaya0@gmail.com>
Link: https://patch.msgid.link/20260909121437.33744-1-omermetekaya0@gmail.com
Signed-off-by: David Heidelberg <david@ixit.cz>
When service_name_len is 0, kmemdup() returns ZERO_SIZE_PTR which
passes the NULL check, causing nfc_llcp_send_connect() to attempt
building a zero-length service name TLV and fail with -ENOMEM.
Fix by setting service_name to NULL directly when service_name_len is 0.
Fixes: d646960f79 ("NFC: Initial LLCP support")
Signed-off-by: Ömer Mete Kaya <omermetekaya0@gmail.com>
Link: https://patch.msgid.link/20260909122029.34081-1-omermetekaya0@gmail.com
Signed-off-by: David Heidelberg <david@ixit.cz>
Commit 6709d4b7bc ("net: nfc: Fix use-after-free caused by
nfc_llcp_find_local") attempted to fix a use-after-free (UAF) issue by
invoking nfc_llcp_local_put(local) after accessing local->gb. However,
if the reference count drops to zero, local is freed immediately,
leading to a use-after-free when callers access the returned pointer.
Alternative approaches using dynamic allocation (e.g. kmemdup) introduced
memory leaks because callers consistently treat the returned pointer as
borrowed memory.
Fix this properly by refactoring nfc_llcp_general_bytes() and
nfc_get_local_general_bytes() to accept a caller-provided output buffer
(out_gb) and its maximum length (gb_max_len). The general bytes are
safely copied into out_gb before calling nfc_llcp_local_put(local),
ensuring safe lifetime management without ownership transfer complications.
Update all callers across drivers (microread, pn533, pn544, st21nfca,
digital_dep, and nci) to provide their own destination buffers and pass
them to nfc_get_local_general_bytes().
Fixes: 6709d4b7bc ("net: nfc: Fix use-after-free caused by nfc_llcp_find_local")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Assisted-by: LLM
Signed-off-by: Luxiao Xu <rakukuip@gmail.com>
Signed-off-by: Ren Wei <weir@nebusec.ai>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/3cbaac3bee23f8ff3a3284ed32d347696eb1d208.1788841683.git.rakukuip@gmail.com
Signed-off-by: David Heidelberg <david@ixit.cz>
nfc_genl_llc_sdreq() builds a list of TLV nodes while walking nested
netlink attrs, but 3 error paths (nested-attr parse failure, TLV alloc
ENOMEM, nfc_llcp_send_snl_sdreq() failure) all skip freeing what was
already queued.
Route them through a new free_list label, mirroring the SDRES path in
the same file which already does this. Harmless on the success path
too -- send_snl_sdreq() drains the list as it moves nodes, so it's
already empty by the time free_list runs.
Fixes: d9b8d8e19b ("NFC: llcp: Service Name Lookup netlink interface")
Assisted-by: Claude:claude-opus-4
Signed-off-by: Cong Nguyen <congnt264@gmail.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260914121129.2098606-1-congnt264@gmail.com
Signed-off-by: David Heidelberg <david@ixit.cz>
nfc_llcp_recv_hdlc() reads the sequence byte skb->data[2], via
nfc_llcp_ns()/nfc_llcp_nr(), before any length check. The receive path
only guarantees the two-byte LLCP header -- __nfc_llcp_recv() checks it
with pskb_may_pull() and nfc_llcp_recv_agf() admits two-byte inner PDUs
-- so a two-byte I, RR or RNR PDU reads one byte of uninitialised skb
tailroom. The byte becomes N(R)/N(S); a peer can already set those with
a well-formed PDU, so this is acting on uninitialised memory, not new
peer control.
Guard the read with pskb_may_pull(), as commit 95674f506c ("nfc: llcp:
reject PDUs shorter than the LLCP header") did for the two-byte header,
so the sequence byte is present and linear before it is read. RR and RNR
PDUs are LLCP_HEADER_SIZE + LLCP_SEQUENCE_SIZE bytes and an I PDU is
longer, so no valid frame is rejected; a truncated PDU is malformed, so
return without a DM reply.
Fixes: d646960f79 ("NFC: Initial LLCP support")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Aamir Ahmed <elb12345@hotmail.co.uk>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/AS8P251MB0001789BBF04B72745C7D96BC8BA2@AS8P251MB0001.EURP251.PROD.OUTLOOK.COM
Signed-off-by: David Heidelberg <david@ixit.cz>
In nfc_llcp_socket_release(), sockets and listener accept queues are
walked under the local sockets rwlock and bh_lock_sock(). However,
bh_lock_sock() does not synchronise against process-context lock_sock()
held by nfc_llcp_accept_dequeue() during accept(). Because
socket_release() does not check sock_owned_by_user(), both paths can
concurrently unlink and release the same child socket, resulting in
use-after-free or a NULL pointer dereference of child->parent in
nfc_llcp_accept_unlink().
Fix this synchronisation race by having nfc_llcp_socket_release() use
process-context lock_sock() instead of bh_lock_sock():
1. Pop sockets from the local sockets list under the write lock using
nfc_llcp_sock_list_pop() so lock_sock() can be acquired without
holding the rwlock.
2. Because lock_sock() can sleep, defer the final release of the
nfc_llcp_local structure to a workqueue (release_work). This avoids
a sleeping-in-atomic bug when the last local reference is dropped
from softirq context. Additionally, hold a single device reference
on local from registration until final destruction.
3. In nfc_llcp_local_get(), use kref_get_unless_zero() to prevent
resurrecting a local object whose teardown has been scheduled.
4. In llcp_sock_accept(), verify that the listener socket state is still
LLCP_LISTEN after waking from schedule_timeout() to prevent hangs if
the listener is closed concurrently.
5. When unlinking unaccepted child sockets during listener release,
unlink them from local->sockets, call sock_orphan(), and drop their
initial sk_alloc creation reference via sock_put().
6. Make nfc_llcp_accept_unlink() idempotent by guarding parent access with
a NULL check.
Fixes: 50b78b2a65 ("NFC: Fix sleeping in atomic when releasing socket")
Signed-off-by: Lee Jones <lee@kernel.org>
Link: https://patch.msgid.link/20260902123033.1169067-1-lee@kernel.org
Signed-off-by: David Heidelberg <david@ixit.cz>
nfc_llcp_recv_dm() handles DM(NOBOUND)/DM(REJ) for a socket that is still
linked on local->connecting_sockets: it looks the socket up with
nfc_llcp_connecting_sock_get(), sets sk->sk_state = LLCP_CLOSED and
returns, without taking the socket lock and without unlinking the socket
from the connecting_sockets list.
llcp_sock_release() selects the list to unlink from by sk_state: a socket
in LLCP_CONNECTING is unlinked from connecting_sockets, otherwise from the
sockets list. Because recv_dm left the socket physically on
connecting_sockets but in the LLCP_CLOSED state, release() takes the else
branch and calls nfc_llcp_sock_unlink(&local->sockets, sk). That runs
sk_del_node_init() while holding sockets.lock, i.e. it removes the socket
from the connecting_sockets hlist under the wrong lock. A concurrent
connect() linking another socket onto connecting_sockets under
connecting_sockets.lock then mutates the same hlist unserialized, which
corrupts the list and desyncs the sk_add_node()/sk_del_node_init()
sock_hold()/__sock_put() pairing. An unprivileged local process holding
LLCP sockets, with the DM supplied by the remote peer over an established
LLCP link, can drive this to leak kernel sockets without bound (the
mis-decrement goes through the non-freeing __sock_put() path, so the
object is never released), leading to memory exhaustion / DoS.
This is the same class of bug that was fixed in the sibling handler
nfc_llcp_recv_cc() by commit b493ea2765 ("nfc: llcp: Fix use-after-free
race in nfc_llcp_recv_cc()"); recv_dm did not receive the equivalent fix.
Fix it the same way: take lock_sock(), re-check that the socket is still
hashed (release() may have won the race), and for the NOBOUND/REJ case
unlink it from connecting_sockets before moving it to LLCP_CLOSED. The
unlink drops the connecting_sockets membership reference via
sk_del_node_init(), leaving the socket unhashed, so the later
nfc_llcp_sock_unlink() in llcp_sock_release() becomes a no-op and no
double put occurs.
Fixes: a69f32af86 ("NFC: Socket linked list")
Signed-off-by: Aldo Ariel Panzardo <qwe.aldo@gmail.com>
Link: https://patch.msgid.link/20260716232657.203145-1-qwe.aldo@gmail.com
Signed-off-by: David Heidelberg <david@ixit.cz>
em_text_dump() allocates struct tcf_em_text on the stack without zeroing
it. strscpy() writes the algorithm name and a NUL terminator into
conf.algo[], leaving the remaining bytes uninitialised. nla_put_nohdr()
then copies the full struct to the netlink response.
KMSAN on Linux 7.2-rc6 reports two kernel-infoleak splats from this path,
one triggered via "tc filter show" and one via a raw RTM_GETTFILTER dump:
BUG: KMSAN: kernel-infoleak in _copy_to_iter+0x1c9/0x2620
nla_put_nohdr+0x83/0x130
em_text_dump+0x291/0x550
Local variable conf created at: em_text_dump+0x5d/0x550
Bytes 168-179 of 199 are uninitialized
I am not certain whether this constitutes a real security problem in
practice: the test was conducted in a controlled KMSAN environment and
the leaked stack bytes may or may not carry sensitive data on actual
production kernels. I am reporting it because KMSAN flagged it as a
kernel-infoleak and the fix is straightforward. I can provide a
userspace reproducer on request.
The original code used strncpy() which zero-pads to the destination size.
Commit b04202d606 ("net/sched: replace strncpy with strscpy") replaced
it with strscpy(), which does not pad, creating this condition.
Zero-initialising the struct closes it.
Fixes: b04202d606 ("net/sched: replace strncpy with strscpy")
Link: https://lore.kernel.org/netdev/20250327143733.187438-1-richard120310@gmail.com/
Assisted-by: Claude:claude-sonnet-4-6 [KMSAN]
Signed-off-by: Bernard Ladenthin <bernard.ladenthin@gmail.com>
Acked-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/20260918133953.12494-1-bernard.ladenthin@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Unlike bind-rx, which configures shared NIC RX queues to steer incoming
traffic into the caller's dmabuf and requires CAP_NET_ADMIN
(uns-admin-perm), bind-tx only DMA-maps the caller's dmabuf so the caller
can transmit from it on their own sockets without affecting other traffic
or device configuration.
Add a comment in netdev.yaml and above netdev_nl_bind_tx_doit() to make it
explicit that NETDEV_CMD_BIND_TX is unprivileged by design.
Signed-off-by: Mina Almasry <almasrymina@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
Acked-by: Daniel Borkmann <daniel@iogearbox.net>
Link: https://patch.msgid.link/20260921195545.493253-1-almasrymina@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Take the association or transport reference before rearming a timer in the
timer handlers.
The existing code calls mod_timer() before taking the reference needed by
the rearmed timer without holding the sock lock. This creates a race with
timer cleanup: if the timer is deleted after mod_timer() returns but before
the reference is taken, the cleanup path can drop the timer's reference and
destroy the transport or association. The timer handler then takes a
reference on the already freed object and eventually drops it, causing a
refcount underflow.
Hold the object before mod_timer() and drop the reference if mod_timer()
reports that the timer was already pending in timer handlers. Apply the
same ordering to the proto-unreachable path, which can rearm a transport
timer outside the timer handlers without holding the sock lock.
Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Reported-by: Tangxin Xie <xietangxin@h-partners.com>
Signed-off-by: Xin Long <lucien.xin@gmail.com>
Link: https://patch.msgid.link/c31b5e3ee2b7274e804f5eba2f21e2412e7eef7a.1790013825.git.lucien.xin@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
tpacket_snd sends skbs with frags pointing into its ring slots. Slots
are released when skb->destructor is called.
A call to skb_orphan calls skb->destructor before the skb is freed.
This can cause the slot to be reused while still linked into the skb.
Switch to standard zerocopy completion (ubuf_info) so the slot is only
released once all references to the payload are freed or copied.
Restore skb->destructor to standard sock_wfree.
The ubuf_info completion callback can be called with a NULL skb, but
only from net_zcopy_put and related API, used by zerocopy implementations
that hold their own reference on the uarg, such as MSG_ZEROCOPY. This
uarg is only ever completed from skb_zcopy_clear, so skb is always set.
To prevent userspace from aliasing in-flight state on shared ring
slots, allocate tpacket_uarg per packet, rather than per slot. This
adds a small allocation to the transmit path. Use standard kmalloc to
allow backporting to stable kernels.
The uarg holds an sk_wmem_alloc reference, rather than an sk_refcnt
reference. packet_free_tx_ring waits on sk_wmem_alloc before freeing
the ring pages. Always allocate vec->deferred for tx_ring so page-backed
rings also wait on sk_wmem_alloc when skb_copy_ubufs drops page refs
before calling tpacket_ubuf_complete.
Drop the tx_ring.pg_vec test that tpacket_destruct_skb performed before
accessing the slot. The sk_wmem_alloc reference now guarantees that the
slot is valid. The test is also not sufficient by itself, as it reads
pg_vec without pg_vec_lock, so it can race with packet_set_ring.
As a result a slot is released when its payload is copied, which can
be before transmission (e.g., in skb_orphan_frags_rx). If copied
before skb_tx_timestamp() is called, no slot timestamp is recorded,
similar to when skb_orphan() was called early in the datapath before
this patch.
Revert the now unused previous skb_zcopy_.._nouarg infra.
Depends on commit 992cc9f94c ("net/packet: defer vmalloc TX_RING
free until skbs finish").
Reported-by: Katherine Leaver <kleaver@janestreet.com>
Reported-by: Bjoern Doebel <doebel@amazon.de>
Closes: https://lore.kernel.org/netdev/20260909085542.3370986-1-doebel@amazon.de/
Fixes: 5cd8d46ea1 ("packet: copy user buffers before orphan or clone")
Cc: stable@vger.kernel.org
Signed-off-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260919004748.1463985-3-willemdebruijn.kernel@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
When an association is in COOKIE-ECHOED state and the peer sends a
bundled [ERROR(Stale Cookie)][DATA] packet from one of its non-primary
addresses, processing the ERROR chunk takes the non-fatal stale-cookie
retry path sctp_sf_do_5_2_6_stale(), which queues
SCTP_CMD_DEL_NON_PRIMARY while keeping the association alive.
sctp_cmd_del_non_primary() removes every non-primary transport -
including the very transport this packet arrived on, which is still
referenced by the receive lookup and shared by all chunks of the
packet via chunk->transport.
sctp_assoc_rm_peer() does redirect asoc->peer.last_data_from away from
the removed transport, but right afterwards the bundled DATA chunk
makes sctp_assoc_bh_rcv() re-register
asoc->peer.last_data_from = chunk->transport unconditionally, undoing
the redirection with the just-removed transport.
Once the packet is done, the receive reference is dropped and the
transport is RCU-freed, while the surviving association keeps the
dangling last_data_from. A later FWD-TSN (or the delayed SACK timer)
makes sctp_gen_sack() dereference it (->param_flags and friends), and
sctp_make_sack()/sctp_outq_select_transport() may write to the freed
object and link it into the live transport list. This is a
use-after-free triggerable by any malicious SCTP peer (or a local
unprivileged user acting as one) with no capabilities required:
BUG: KASAN: slab-use-after-free in sctp_do_sm+0x498a/0x5660
Read of size 4 at addr ffff88800e1e356c by task poc/115
Call Trace: sctp_do_sm <- sctp_assoc_bh_rcv <- sctp_inq_push <-
sctp_rcv <- ip_protocol_deliver_rcu <- ip_rcv
Allocated: sctp_transport_new <- sctp_assoc_add_peer <-
sctp_process_init (INIT-ACK processing)
Freed: kfree <- sctp_transport_destroy_rcu <- rcu_core
(call_rcu queued by sctp_transport_put at end of sctp_rcv)
The buggy address is located 364 bytes inside of freed 1024-byte
region [ffff88800e1e3400, ffff88800e1e3800), cache kmalloc-1k
Note that commit 03a9d10ecf ("sctp: drop a chunk if its transport
was removed") only covers the window between the receive lookup and
the chunk processing (e.g. an ASCONF DEL-IP racing the socket backlog);
here the transport is removed *while* the packet is being processed,
by an earlier chunk of the same packet, so the drop in sctp_inq_push()
does not reach this path. Verified with the bundled [ERROR(Stale
Cookie)][DATA] + FWD-TSN reproducer: the KASAN report above still
fires with that commit applied, and is gone with this patch on top.
Fix it by discarding the rest of the packet on this path, as suggested
by Xin. After the stale-cookie ERROR has sent the association back to
COOKIE-WAIT and removed the non-primary transports, the remaining
chunks of the packet can only run against the restarted handshake
while referencing the removed arrival transport through
chunk->transport: besides the last_data_from registration above,
sctp_cmd_setup_t2() and the sctp_make_*() reply builders would also
copy that pointer into association-lifetime state that
sctp_assoc_rm_peer() has already sanitized. Let the peer retransmit
them, in line with what sctp_inq_push() does for chunks whose
transport was removed before processing.
Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Suggested-by: Xin Long <lucien.xin@gmail.com>
Reported-by: TencentOS Corvus AI <corvus@tencent.com>
Cc: stable@vger.kernel.org
Signed-off-by: Aohan Mei <henrymei@tencent.com>
Acked-by: Xin Long <lucien.xin@gmail.com>
Link: https://patch.msgid.link/20260921093707.1432184-1-ljp1205831794@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
The op-to-policy map a CTRL_CMD_GETPOLICY dump returns is the only way
for userspace to find out which policy index belongs to which command.
ctrl_dumppolicy_put_op() tags the nest with doit->cmd, but an op which
only has a dumpit has no doit and every path which fills the split ops
in zeroes it out, so those entries all claim to be command 0. nlctrl's
own CTRL_CMD_GETPOLICY and NETDEV_CMD_QSTATS_GET are both in that group:
[{'family-id': 16, 'op-policy': {'do': 0, 'dump': 0, 'op-id': 3}},
{'family-id': 16, 'op-policy': {'dump': 1, 'op-id': 0}},
ctrl_fill_info() gets this right - it uses the iterator's cmd for
CTRL_ATTR_OP_ID - so the two introspection interfaces of the same family
contradict each other today.
Pass the command in rather than reconstructing it from
doit->cmd | dumpit->cmd inside the helper, both callers already have it.
Fixes: 26588edbef ("genetlink: support split policies in ctrl_dumppolicy_put_op()")
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Link: https://patch.msgid.link/20260918222949.4190284-1-kuba@kernel.org
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
A UDP socket bound to a specific address and port keeps its entry in the
4-tuple hash table after it is disconnected:
sk binds to 127.0.0.1:21001
sk connects to 127.0.0.2:20001 // filed in the 4-tuple table
sk disconnects, connect(AF_UNSPEC) // still filed, peer now 0.0.0.0:0
__udp_disconnect() takes a socket out of that table only as a side effect
of ->rehash() or ->unhash(), and it skips ->rehash() when
SOCK_BINDADDR_LOCK is set and ->unhash() when SOCK_BINDPORT_LOCK is set.
commit 6996a2d2d0 ("udp: Unhash auto-bound connected sk from 4-tuple hash
table when disconnected.") fixed the same end state for a wildcard-bound
socket, by a path this one does not take.
The entry is counted whether or not anything hits it. hash4_cnt on the
hash2 slot stays raised for as long as the socket lives, so udp_has_hash4()
keeps sending every packet for that address and port through the 4-tuple
lookup first.
On IPv6 it can also be hit. __udp_disconnect() does not clear sk_v6_daddr,
so udp_v6_rehash() files the entry under the peer the socket was connected
to with a zero dport, and inet6_match() compares that same
field: a datagram from the former peer with a zero source port matches,
and source port zero is accepted on receive. On IPv4 the peer is cleared,
so a match would need a zero source address as well, which the routing
layer rejects as martian. The stale sk_v6_daddr is a separate defect, not
addressed here; removing the entry closes this path either way.
The entry can also be relocated. __udp_disconnect() clears sk_bound_dev_if,
so a subsequent SO_BINDTODEVICE calls ->rehash(), and because the receive
address is still specific udp_lib_rehash() moves the entry instead of
removing it, into the bucket that (rcv_saddr, num, 0, 0) hashes to -- a
pure function of the address and port, so every socket reaching this state
on one address and port collects in one bucket. The bucket cannot be chosen
from outside, as udp_ehashfn() is seeded with a per-boot secret. This last
one became reachable only with commit 644f9108f3 ("udp: Make rehash4
independent in udp_lib_rehash()"), which moved the hash4 handling out of a
branch a disconnected socket does not take; the stale entry itself dates
from the commit in Fixes.
Take the socket out of the table before __udp_disconnect() runs, while it
still matches how it was filed. This also reaches the wildcard case ahead
of udp_lib_rehash()'s udp_unhash4() branch, leaving that branch unreachable
from udp_disconnect(); removing it belongs in net-next. udp_disconnect()
and udp_abort() are the only UDP entries into __udp_disconnect(), which is
shared with raw, ping and l2tp sockets that are not struct udp_sock:
ping_prot.obj_size is sizeof(struct inet_sock), so udp_hashed4() on one
would read past the allocation.
Fixes: 78c91ae2c6 ("ipv4/udp: Add 4-tuple hash for connected socket")
Assisted-by: LLM
Signed-off-by: Shardul Bankar <shardul.b@mpiricsoftware.com>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-2-718891af0d7a@mpiricsoftware.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
A connected UDP socket that connects again to a different peer is not
re-filed in the 4-tuple hash table:
sk binds to 127.0.0.1:21001
sk connects to 127.0.0.2:20001 // filed under hash(sk, peer1)
sk connects to 127.0.0.3:20002 // still filed under hash(sk, peer1)
packet from 127.0.0.3:20002 // hash(sk, peer2) misses, so the
// lookup falls back to scoring the
// hash2 chain for this address
// and port
udp_lib_hash4() returns early when the socket is already hashed, assuming
->rehash() relocates it. ->rehash() runs from __ip{4,6}_datagram_connect()
only while the receive address is unset, which a second connect never is:
the first connect assigns it, whether the socket was bound to a specific
address or to the wildcard. commit 644f9108f3 ("udp: Make rehash4
independent in udp_lib_rehash()") added that early return and named
connect(AF_UNSPEC) as the way around it. That workaround does not help a
socket with both SOCK_BINDADDR_LOCK and SOCK_BINDPORT_LOCK set, because
__udp_disconnect() skips ->rehash() for the first and ->unhash() for the
second.
Delivery is correct either way.
Relocate the socket when the hash it is filed under differs from the one
requested, which is what commit 78c91ae2c6 ("ipv4/udp: Add 4-tuple hash
for connected socket") did before the early return became unconditional. It
is done here under hslot->lock, which that version did not take, to match
udp_lib_rehash() and udp_lib_unhash(). hslot2 is unchanged, so hash4_cnt
needs no adjustment, as in udp_lib_rehash(). A first connect is unaffected,
and IPv6 shares the code.
With 500 sockets on the port, a re-connected socket measured 522,553 pps
without this change and 2,055,078 with it. The UDP side was noted as
remaining work in [1].
Link: https://lore.kernel.org/netdev/apnHqmYZQ4yzOP4N@v4bel/ [1]
Fixes: 644f9108f3 ("udp: Make rehash4 independent in udp_lib_rehash()")
Assisted-by: LLM
Signed-off-by: Shardul Bankar <shardul.b@mpiricsoftware.com>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-1-718891af0d7a@mpiricsoftware.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Core:
- hci_conn: fix CIS hold ownership on reuse
- hci_sock: reject out-of-range OCF values
- hci_sock: validate event length before filtering
- L2CAP: validate frame length before control and FCS access
- RFCOMM: fix NULL dereference of dlc->session in RFCOMM_CONNINFO
- RFCOMM: Reject short EA=0 frames in rfcomm_recv_frame()
- ISO: release unused CIS holds after channel attach
- ISO: balance the parent hold in hci_bind_bis()
- SMP: reject Security Request over BR/EDR
- MGMT: fix race in read_unconf_index_list()
- MGMT: Dequeue pending mesh_send_sync entries on cancel
- BNEP: fix out-of-bounds reads on short RX/TX frames and control fallthrough
Drivers:
- btintel_pcie: validate device-supplied DMA indices
- btnxpuart: Fix skb leak in nxp_process_fw_dump()
-----BEGIN PGP SIGNATURE-----
iQJNBAABCgA3FiEE7E6oRXp8w05ovYr/9JCA4xAyCykFAmqxN44ZHGx1aXoudm9u
LmRlbnR6QGludGVsLmNvbQAKCRD0kIDjEDILKS7mD/43nt9IlCMp4fRi5eT3iV6J
phP/zJTiikgOMv87kTI0Q9OXY8Xl1nhIGrTiypXQIJNJGTW/OTHtMF+N55pA7pF4
exD0US01bKUcuopztHeP1Yk08CMKAtE9VTPLi/PdAQz6KTo7wcH7yFwfsO6avIk4
T/IXi9b8iKPBMKizjgUe8Uo6wrFoByD/o2VwTSQwtZOgWppUkCvVKKx10Y8EYJRG
UbDS8rpPtWN3t0fK5F4yVfkTP+9sA7uWb9bF2UoLcmov2ie33M50zf5giVofiNrX
pT6RWTCPfjkPdDro66l6wV4M8prEUVok0QhlqeYOUXSZJSOTqqXBTOorioABSeHz
i4QWH+KsqUFbaOJQuoF3v5jAccY8ZxKczRODK8Xt1+bY9O73lqpuWjQvoi98E/7i
vNCSU/bDpRS0MjYDTgQTG9c0rv9qxKI0GBC8gDVI78/nAzUIxdJ09/wDBm17aUpt
uxaBwYE1lhJwT0Lf+dwJrwVuTm5q5XotdRIsOwxZgfF25Ewk8ygUlEsQu/ioHmS+
q+9SL/mHb/wPI7j0YeVG2l7cHiHOZDEn0FW++WlgtF1Oj/HWUsz2UN/qdmpYZo1h
rGX7D2DdZQo8IOEmZZn7wqLNOxjPjFVkxgQtgWd5OhE0O+7/e2H4mNeDJyWs6rJp
lT1BJZnPQTFtwP41x7Hxng==
=cLTc
-----END PGP SIGNATURE-----
Merge tag 'for-net-2026-09-21' of git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth
Luiz Augusto von Dentz says:
====================
bluetooth pull request for net:
Core:
- hci_conn: fix CIS hold ownership on reuse
- hci_sock: reject out-of-range OCF values
- hci_sock: validate event length before filtering
- L2CAP: validate frame length before control and FCS access
- RFCOMM: fix NULL dereference of dlc->session in RFCOMM_CONNINFO
- RFCOMM: Reject short EA=0 frames in rfcomm_recv_frame()
- ISO: release unused CIS holds after channel attach
- ISO: balance the parent hold in hci_bind_bis()
- SMP: reject Security Request over BR/EDR
- MGMT: fix race in read_unconf_index_list()
- MGMT: Dequeue pending mesh_send_sync entries on cancel
- BNEP: fix out-of-bounds reads on short RX/TX frames and control fallthrough
Drivers:
- btintel_pcie: validate device-supplied DMA indices
- btnxpuart: Fix skb leak in nxp_process_fw_dump()
* tag 'for-net-2026-09-21' of git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth:
Bluetooth: RFCOMM: Reject short EA=0 frames in rfcomm_recv_frame()
Bluetooth: RFCOMM: fix NULL dereference of dlc->session in RFCOMM_CONNINFO
Bluetooth: btintel_pcie: validate device-supplied DMA indices
Bluetooth: bnep: fix out-of-bounds reads on short RX/TX frames and control fallthrough
Bluetooth: mgmt: fix race in read_unconf_index_list()
Bluetooth: L2CAP: validate frame length before control and FCS access
Bluetooth: ISO: balance the parent hold in hci_bind_bis()
Bluetooth: hci_sock: validate event length before filtering
Bluetooth: hci_sock: reject out-of-range OCF values
Bluetooth: ISO: release unused CIS holds after channel attach
Bluetooth: hci_conn: fix CIS hold ownership on reuse
Bluetooth: mgmt: Dequeue pending mesh_send_sync entries on cancel
Bluetooth: btnxpuart: Fix skb leak in nxp_process_fw_dump()
Bluetooth: SMP: reject Security Request over BR/EDR
====================
Link: https://patch.msgid.link/20260921135807.3459373-1-luiz.dentz@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
ip6_route_output_flags(), ip6_rt_put_flags(), and ip6_dst_check()
detect an uncached route by list_empty(&rt->dst.rt_uncached),
which replaced the static DST_NOCACHE flag check in commit
a4c2fd7f78 ("net: remove DST_NOCACHE flag").
When a device is unregistered, rt6_uncached_list_flush_dev()
unlinks uncached routes tied to the device from rt6_uncached_list.
Previously, they were moved to another list with list_move()
(__list_del_entry() + list_add()), and since commit 98aa546af5
("inet: remove (struct uncached_list)->quarantine"), the routes
are just unlinked with list_del_init().
If list_del_init() runs concurrently, list_empty() evaluates to
true; ip6_route_output_flags() calls dst_hold_safe() incorrectly
and ip6_rt_put_flags() skips ip6_rt_put(), leaking dst, and thus
dev tied via rt->from as well.
The same race is partially fixed by commit 9a6f0c4d57 ("dst:
fix races in rt6_uncached_list_del() and rt_del_uncached_list()").
Let's check rt6->dst.rt_uncached_list instead.
Note that IPv4 does not have the same issue.
Fixes: 98aa546af5 ("inet: remove (struct uncached_list)->quarantine")
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Reviewed-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260920191558.2990636-1-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Removing the legacy ioctl fallback made both hwtstamp NDOs mandatory. A
device that only timestamps in its PHY implements neither, so
SIOCSHWTSTAMP fails with EOPNOTSUPP before anything looks at the PHY and
PTP stops working there.
The check only ever picked the legacy path. That path is gone, so drop it
and test where the NDOs are actually called.
SIOCGHWTSTAMP is new here, not restored. The old path went through
phy_mii_ioctl(), which only handled SIOCSHWTSTAMP.
Such a device now returns -ENODEV while absent instead of -EOPNOTSUPP,
like the ones that do implement the NDOs.
Fixes: 5062245a5a ("net: remove legacy way to get/set HW timestamp config")
Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
Reviewed-by: Kory Maincent <kory.maincent@bootlin.com>
Link: https://patch.msgid.link/20260918095540.34286-1-nb@tipi-net.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Before the cited commit, fib6_nh_flush_exceptions() always set
from->exception_bucket_flushed = 1 under rt6_exception_lock to
prevent rt6_insert_exception() from inserting a new exception
for a dying fib6_info.
The flag was replaced with the FIB6_EXCEPTION_BUCKET_FLUSHED
bit stored in nh->rt6i_exception_bucket.
The problem is that now the bit is only set when the bucket
is not NULL and fib6_nh_flush_exceptions() is called from
fib6_nh_release() after fib6_ref has already reached zero.
If rt6_insert_exception() is called while the target fib6_info
is being removed via fib6_purge_rt(), a new exception could be
created successfully because rt6_flush_exceptions() no longer
sets the bit.
This creates a reference cycle between the fib6_info and the
exception route, leaking the fib6_info, its nexthop device,
and all per-CPU routes in fib6_nh->rt6i_pcpu, which stalls netdev
unregistration.
[ 34.680602] unregister_netdevice: waiting for gre6 to become free. Usage count = 68
[ 44.920675] unregister_netdevice: waiting for gre6 to become free. Usage count = 68
[ 55.176582] unregister_netdevice: waiting for gre6 to become free. Usage count = 68
Let's call fib6_drop_pcpu_from() before rt6_flush_exceptions(),
to set fib6_destroying before rt6_exception_lock, and check
f6i->fib6_destroying in rt6_insert_exception().
Note that FIB6_EXCEPTION_BUCKET_FLUSHED logic is dead and
we can clean it up in net-next.
Fixes: cc5c073a69 ("ipv6: Move exception bucket to fib6_nh")
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260918082209.2853582-1-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
-----BEGIN PGP SIGNATURE-----
iQIzBAABCgAdFiEEjF9xRqF1emXiQiqU1w0aZmrPKyEFAmqtHzMACgkQ1w0aZmrP
KyGtkQ//bMKQGEudQKCMCtQmPqaHyW1ajmAo17aumfczaE/nSjqrsGY5Ahw7OKlQ
4otcdPvI4qpV9sLTg41KFaIHIC5sozxt4Q3m3RNB2TbyCkGn9xpSpZxM5IpHybvE
83tVjSA0wpfIqxBEqKUqk8Z9AXtBLo/JocdfYry+6JUyj4PM76X2ViKpzaPbpoMU
1mndfLAYtADIIvs3805CmfdJmOkoSV6XCEsiNutPrJhiRfN4xJZ9leP9xb1zA0IQ
cnqiaw1xkTcFyWCicu4MqOkEALRknr9SL2yX1S9wx5Q6WHwU9JXUeQTlvfv7OoVP
uxuMlNr3WcbwHC9e1GfOHapzjrYgnvEe2Z79i2GFh51Ci+5L9Yr9XCQ/fc6G5NNZ
3W52kh35s3lXq32hll9Tkr7pf4cKLBA+IAJ19VNlRfMrPB0cz4EqbIZ6xNNuLqdh
DbEb3VgTT2dHwuGxEshJVmSfzfR+VeHBG2ZRlRmZElfhViHEwgPaAxkaJhNpPyub
qmHbZCXK0BVp/UrGHDm5rmHJtdkwprXY9YceZBRfW8Fr2Ler4rWQvy+uo9sRRFgN
oF9B6qzSl78THBv3UDB3U+aWuDv0I+VlDub0DKf0k9iWg8OoMlM9yw6OE9ULHt3m
LRvSAfF0LnF/z7byzumUYMVjpVuIVueDzasGrEZjWbHHE7MCsp0=
=RfSQ
-----END PGP SIGNATURE-----
Merge tag 'nf-26-09-18' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf
Pablo Neira Ayuso says:
====================
Netfilter/IPVS fixes for net
The following patchset contains Netfilter/IPVS fixes for net, they are:
1) Set on HW_DEAD after HW_PENDING is cleared in the flowtable offload
to ensure GC does not zap it, from Jérémy Jean.
2) Hold the nfnetlink_queue mutex while removing the queue instance
from the netlink notifier that handles NETLINK_URELEASE to fix a
possible race with the UNBIND command. From Florian Westphal.
3) Reject route with NULL rt6i_idev in ip6t_rpfilter. From Weiming Shi.
4) Reject rtinfo->addrnr set to zero from ip6t_rt .checkentry path.
This also fortifies the datapath loop as per Florian's request.
From Luxiao Xu.
5) Fix checksuming in nft_synproxy for IPv6, from Karl Mehltretter.
6) Revalidate ihl before calling icmp_send() in IPVS,
from Julian Anastasov.
7) Fix suspicious RCU usage splat in ctnetlink with expectations.
8) Check for expired catchall elements in the insert and deactivate
path. From Aohan Mei.
* tag 'nf-26-09-18' of git://git.kernel.org/pub/scm/linux/kernel/git/netfilter/nf:
netfilter: nf_tables: skip expired catchall elements on insert and delete
netfilter: ctnetlink: fix suspicious RCU usage in expect_iter_name
ipvs: revalidate ihl before icmp_send
netfilter: nft_synproxy: use the family-aware checksum helper
netfilter: ip6t_rt: fix zero-address non-strict match out-of-bounds read
netfilter: ip6t_rpfilter: reject routes without inet6_dev
netfilter: nfnetlink_queue: hold nfnl mutex in event notifier
netfilter: flowtable: publish HW_DEAD after worker is done
====================
Link: https://patch.msgid.link/20260918112844.194503-1-pablo@netfilter.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
While rfcomm_recv_frame() verifies that skb->len is at least
sizeof(*hdr) + 1 (4 bytes: 3-byte header + 1-byte FCS), an RFCOMM frame
with an extended 2-byte length field (!__test_ea(hdr->len)) has a 4-byte
header plus a 1-byte FCS (5 bytes minimum, sizeof(*hdr) + 2).
When a 4-byte RFCOMM frame with EA == 0 arrives:
1. The initial skb->len < sizeof(*hdr) + 1 check passes (4 < 4 is false).
2. Trimming the FCS byte decrements skb->len to 3.
3. If __check_fcs() succeeds, skb_pull(skb, 4) fails (4 > 3) and returns
NULL without advancing skb->data.
4. Because the return value of skb_pull() is ignored, the un-pulled
3-byte struct rfcomm_hdr remains at skb->data and is either queued as
application payload via rfcomm_recv_data() or parsed as a multiplexer
control command via rfcomm_recv_mcc() on DLCI 0.
Fix this by extending the length check in rfcomm_recv_frame() to also
require skb->len >= sizeof(*hdr) + 2 when !__test_ea(hdr->len).
Fixes: b230e5bf50 ("Bluetooth: RFCOMM: validate skb length in rfcomm_recv_frame")
Assisted-by: LLM
Signed-off-by: Hui Peng <benquike@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
The RFCOMM_CONNINFO getsockopt handler accepts a socket that is not
connected as long as deferred setup is enabled:
if (sk->sk_state != BT_CONNECTED &&
!rfcomm_pi(sk)->dlc->defer_setup) {
err = -ENOTCONN;
break;
}
l2cap_sk = rfcomm_pi(sk)->dlc->session->sock->sk;
dlc->defer_setup is set in rfcomm_sock_init() when rfcomm_connect_ind()
creates a child socket for an incoming connection on a listening socket
that has BT_DEFER_SETUP enabled. It is never cleared afterwards. The
session, however, can go away underneath it.
rfcomm_recv_disc() forces the dlc state before tearing it down:
d->state = BT_CLOSED;
__rfcomm_dlc_close(d, err);
The RFCOMM_DEFER_SETUP early return in __rfcomm_dlc_close() only covers
BT_CONNECT, BT_CONFIG, BT_OPEN and BT_CONNECT2, so with the state
already BT_CLOSED that switch does not match and the function falls
through to rfcomm_dlc_unlink(), which sets d->session = NULL, while
d->defer_setup stays 1.
A getsockopt(SOL_RFCOMM, RFCOMM_CONNINFO) on the accepted socket after
that point therefore skips the -ENOTCONN path -- sk->sk_state is
BT_CLOSED, but dlc->defer_setup is still set -- and dereferences the
NULL session. No race is needed: once the DISC has been processed, the
dereference is unconditional.
Reproduced on a KASAN kernel under QEMU with a BR/EDR peer emulated over
/dev/vhci: the peer brings up an ACL link, opens L2CAP on the RFCOMM
PSM, starts a session and sends SABM for a channel bound with
BT_DEFER_SETUP, and sends DISC for that dlci after the socket has been
accepted. getsockopt(SOL_RFCOMM, RFCOMM_CONNINFO) on the accepted
socket then hits:
Oops: general protection fault, probably for non-canonical address
0xdffffc0000000002: 0000 [#1] SMP KASAN PTI
KASAN: null-ptr-deref in range [0x0000000000000010-0x0000000000000017]
CPU: 1 UID: 0 PID: 150 Comm: init Tainted: G B 7.3.0-rc3-g5dd1818b15d9
Hardware name: QEMU Standard PC (i440FX + PIIX, 1996)
RIP: 0010:rfcomm_sock_getsockopt+0x529/0x780
Call Trace:
<TASK>
do_sock_getsockopt+0x3ad/0x7d0
__sys_getsockopt+0x10e/0x1b0
__x64_sys_getsockopt+0xc2/0x160
do_syscall_64+0xda/0x4b0
entry_SYSCALL_64_after_hwframe+0x77/0x7f
</TASK>
0x10 is the offset of sock in struct rfcomm_session;
rfcomm_sock_getsockopt_old() is inlined into rfcomm_sock_getsockopt().
Commit 43a556b2fd ("Bluetooth: RFCOMM: take rfcomm_mutex for the
deferred setup accept") fixed the same "a remote DISC clears the session
while deferred setup is still flagged" problem in rfcomm_dlc_accept();
this is the remaining instance of it, in the getsockopt path.
Deferred setup only leaves a socket usable here once it has reached
BT_CONNECT2, so restrict the exception to that state and check that a
session is actually present before following it.
Fixes: bb23c0ab82 ("Bluetooth: Add support for deferring RFCOMM connection setup")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Hui Peng <benquike@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
Fix multiple out-of-bounds reads in Bluetooth BNEP frame processing:
1. In bnep_rx_frame() and bnep_ctrl_frame() (net/bluetooth/bnep/core.c),
use pskb_may_pull() to verify the BNEP header, control type byte,
filter count, and extension headers exist before reading them, and
return 0 after handling BNEP_CONTROL instead of falling through to
Ethernet frame submission when no extension headers follow.
2. In bnep_net_xmit() (net/bluetooth/bnep/netdev.c), verify skb->len >=
ETH_HLEN with pskb_may_pull() before reading the 14-byte Ethernet
header to prevent an out-of-bounds heap read and infoleak on short
AF_PACKET TX frames.
Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Assisted-by: LLM
Signed-off-by: Hui Peng <benquike@gmail.com>
Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com>
hfsc_classify() applies the "filter may only point downwards" level check
only when the filter result carries no bound class. A filter created with
a flowid gets res.class set once at bind time, so the check never runs for
it during classification. hfsc_adjust_levels() can later raise a class's
level without revalidating existing bindings, leaving two binds that were
each legal at bind time pointing at each other; the classify walk then
bounces between two interior classes forever with the qdisc lock held and
BH disabled — a soft lockup from a single packet. The stuck walk trips
the watchdog:
watchdog: BUG: soft lockup - CPU#3 stuck for 13s! [ping:444]
RIP: 0010:u32_classify+0x542/0x17f0
...
tcf_classify+0x66/0xa0
hfsc_enqueue+0x166/0xdf0
Bound the traversal with a budget of non-descending hops, the only way a
configured walk can move without descending the class tree once levels
drift after bind time. The budget is cumulative over the whole walk and
is deliberately not reset on a descending hop: a chain that alternates a
descent with a lateral hop would return the budget every lap and never
trip. Descending hops never decrement it, so legitimately deep trees are
unaffected and a terminating lateral chain still classifies normally.
Drop the packet with a rate-limited warning once the budget is exhausted,
mirroring the merged HTB fix.
This is a follow-up to commit 729c4896ab ("net/sched: sch_htb: limit
htb_classify inner-class filter hops"), which bounded the same classify
loop on the HTB side but left the HFSC walk unbounded.
Conditions to recreate the bug:
- CONFIG_NET_SCHED, CONFIG_NET_SCH_HFSC, CONFIG_NET_CLS_U32,
CONFIG_LOCKUP_DETECTOR.
- Build a cycle with two legal-at-bind-time flowid binds and a level
drift: class X 1:1 (child of root) with leaf child 1:10; class Y 1:2
(sibling of X) with children 1:20 and 1:200; root u32 filter flowid
1:1; filter on X flowid 1:2 (legal when Y is a leaf); after Y's level
rises to 2, filter on Y flowid 1:1 (legal then). Send one packet (ping
on the device). Unfixed kernel: classify spins with the qdisc lock
held; with softlockup_panic=1 it panics.
- Reachable from unprivileged user via unshare -Urn (CAP_NET_ADMIN).
Fixes: a2f7922713 ("net_sched: sch_hfsc: fix classification loops")
Reported-by: Sashiko (gemini + nipa) <sashiko-bot@kernel.org>
Closes: https://lore.kernel.org/netdev/QDISC-CTUU.v2.20260913192614@mojatatu.com/
Link: https://sashiko.dev/#/patchset/QDISC-CTUU.v2.20260913192614@mojatatu.com
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/QDISC-CTUU.v2.20260913192614%40mojatatu.com
Reviewed-by: Victor Nogueira <victor@mojatatu.com>
Tested-by: hybris <hybris@mojatatu.ai>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-CTUU.v3.20260916184908@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
tcf_action_delete() drops the reference held by its lookup before calling
tcf_idr_delete_index() with the saved action index. An unlocked
classifier can remove that action and reserve the same IDR slot with
ERR_PTR(-EBUSY) in between.
tcf_idr_delete_index() only checks the lookup result for NULL. It
therefore treats the reservation as a tc_action and dereferences
tcfa_bindcnt. A hardware execution breakpoint was used to schedule the
interleaving without changing the kernel source. KASAN reported this
decoded trace:
BUG: KASAN: null-ptr-deref in tca_action_gd+0x5b9/0x1010
Read of size 4 at addr 0000000000000010 by task poc/150
Oops: general protection fault, probably for non-canonical address 0xdffffc0000000002
RIP: tca_action_gd+0x5c0/0x1010:
arch_atomic_read at arch/x86/include/asm/atomic.h:23
raw_atomic_read at include/linux/atomic/atomic-arch-fallback.h:457
atomic_read at include/linux/atomic/atomic-instrumented.h:33
tcf_idr_delete_index at net/sched/act_api.c:766
tcf_action_delete at net/sched/act_api.c:1859
tcf_del_notify at net/sched/act_api.c:2014
tca_action_gd at net/sched/act_api.c:2064
R13: 0000000000000010 R15: fffffffffffffff0
Kernel panic - not syncing: Fatal exception
R15 contains ERR_PTR(-EBUSY), and adding the tcfa_bindcnt offset produces
the address in R13. With the guard applied, the same reproducer returned
-ENOENT without a KASAN report or panic. Treat error pointers as absent
and return -ENOENT.
Fixes: 0190c1d452 ("net: sched: atomically check-allocate action")
Cc: stable@vger.kernel.org
Reported-by: Xiang Mei <xmei5@asu.edu>
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Link: https://patch.msgid.link/20260914065123.4109709-2-bestswngs@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Commit fa8fca8871 ("ipv4: validate IPV4_DEVCONF attributes properly")
added validation of IFLA_INET_CONF attributes, and in the process
changed the call of nla_for_each_nested() to nla_parse_nested(). A
side effect of this change is that the IFLA_INET_CONF option is now
tested for NLA_F_NESTED being set, and fails if it is not. Prior to the
commit there was no check of NLA_F_NESTED.
Change nla_parse_nested() to nla_parse(). This restores the previous
functionality of not checking NLA_F_NESTED, thereby allowing code that
(incorrectly) doesn't set NLA_F_NESTED to continue to work.
This issue was identified because keepalived started logging errors when
it was configuring macvlans that it created.
Fixes: fa8fca8871 ("ipv4: validate IPV4_DEVCONF attributes properly")
Signed-off-by: Quentin Armitage <quentin@armitage.org.uk>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260915213320.1527029-2-quentin@armitage.org.uk
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
sctp_assoc_update_retran_path() can loop forever when every remaining
transport, including retran_path, is SCTP_UNCONFIRMED: the state check
runs before the wraparound test, so the loop cannot observe that it has
completed a full pass.
Fix this by considering a transport only when it is not UNCONFIRMED,
then checking whether the walk has returned to retran_path. This makes
the full-pass termination independent of the transport state while
preserving the existing fallback selection semantics.
Also restore the NULL guard around the retran_path assignment. In the
all-UNCONFIRMED case there is no eligible replacement transport, and
installing NULL would leave later retransmit-path users and the debug
print with a NULL path.
Fixes: 4c47af4d5e ("net: sctp: rework multihoming retransmission path selection to rfc4960")
Signed-off-by: Yiqi Sun <sunyiqixm@gmail.com>
Acked-by: Xin Long <lucien.xin@gmail.com>
Link: https://patch.msgid.link/20260915095017.942213-1-sunyiqixm@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
An offline self test that brings the interface down and back up with
netif_close() / netif_open() requires rtnl_lock for both. Since the
ethtool IOCTL path became rtnl-optional for ops-locked drivers, the
ETHTOOL_TEST ioctl runs holding only the netdev instance lock, so on an
ops-locked driver the self test now tears the device down without
rtnl_lock.
With lockdep this reproduces deterministically on every offline self
test on such a driver; note the sole lock held is the instance lock, not
rtnl:
WARNING: suspicious RCU usage
net/core/netpoll.c:207 suspicious rcu_dereference_protected() usage!
1 lock held by ethtool/107:
#0: (&dev->lock){+.+.}, at: dev_ethtool
Call Trace:
netpoll_poll_disable
__dev_close_many
netif_close_many
netif_close
fbnic_self_test
dev_ethtool_locked
dev_ethtool
dev_ioctl
sock_ioctl
__x64_sys_ioctl
Without lockdep the same condition trips ASSERT_RTNL() in
__dev_close_many() / __dev_open(); that check only samples the global
rtnl state, so it can be masked by a concurrent rtnl holder, but the
device is still being reconfigured without the lock it requires.
The ethtool self_test is a legacy ioctl-only command, so an ETHTOOL_TEST
case is only needed on the ioctl path. Add an opt-in bit for drivers whose
self test needs rtnl_lock and set it on the ops-locked drivers whose
offline self test tears the interface down and up:
- fbnic (ops-locked via queue_mgmt_ops): fbnic_self_test() offline path
uses netif_close() / netif_open().
- bnxt (ops-locked via queue_mgmt_ops): bnxt_self_test() offline path
goes through bnxt_close_nic() / bnxt_half_open_nic() /
bnxt_half_close_nic() / bnxt_open_nic(), which close and reopen the
device.
Fixes: f994752b11 ("net: ethtool: optionally skip rtnl_lock on IOCTL path")
Signed-off-by: Alexander Duyck <alexanderduyck@fb.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/178942019771.7700.338431553546884773.stgit@ahduyck-xeon-server.home.arpa
Signed-off-by: Jakub Kicinski <kuba@kernel.org>