Commit Graph

10835 Commits

Author SHA1 Message Date
Norbert Szetei
4da3b7b8b5 net: xps: reject an out of range traffic class
Only the entries below dev->num_tc are valid in dev->tc_to_txq[], and
dev->prio_tc_map[] may only name classes below it. netdev_set_num_tc()
lowers dev->num_tc without touching either array.

netdev_txq_to_tc() walks all TC_MAX_QUEUE slots and
netdev_get_prio_tc_map() returns the entry as it stands, so a leftover
entry is handed out as a traffic class >= dev->num_tc. Taking that
class from netdev_txq_to_tc(), __netif_set_xps_queue() rejects only a
negative one and indexes an XPS map sized for dev->num_tc classes:

	tci = j * num_tc + tc;
	RCU_INIT_POINTER(new_dev_maps->attr_map[tci], map);

attr_map[] holds nr_ids * num_tc entries and j runs over the ids named
in the mask, so a class that is not below num_tc pushes tci past the end
of the map for the last ids and the store overruns it.

Any caller that lowers num_tc leaves such entries behind, and
mqprio_destroy() tears down with netdev_set_num_tc(dev, 0) rather than
netdev_reset_tc(). After mqprio with 8 classes then 1, tc_to_txq[1..7]
still describe txq 1..7. The splat is from an XPS write to txq 2 on a
veth with 8 rx queues: attr_map[] has 8 * 1 entries, tci = j + 2, and
j == 6 stores one past the end of the 88-byte map:

  BUG: KASAN: slab-out-of-bounds in __netif_set_xps_queue (net/core/dev.c:2954)
  Write of size 8 at addr ffff88813016bc58 by task xps_oob/634
   __netif_set_xps_queue (net/core/dev.c:2954)
   xps_rxqs_store (net/core/net-sysfs.c:1880)
   netdev_queue_attr_store (net/core/net-sysfs.c:1390)
  Allocated by task 634:
   __kmalloc_noprof (mm/slub.c:5439)
   __netif_set_xps_queue (net/core/dev.c:2937)
  The buggy address is located 0 bytes to the right of
   allocated 88-byte region [ffff88813016bc00, ffff88813016bc58)

Reject a class the map has no room for.

Fixes: 184c449f91 ("net: Add support for XPS with QoS via traffic classes")
Signed-off-by: Norbert Szetei <norbert@doyensec.com>
Link: https://patch.msgid.link/162DD16F-54C6-444A-9E09-0B8CB3D591F2@doyensec.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23 19:39:38 -07:00
Shihuang Liu
3b4e0b0c00 net: skbuff: fix pull-bound underflow in skb_checksum_setup_ipv6()
skb_maybe_pull_tail() subtracts skb_headlen(skb) from the unsigned max
argument and passes the result to __pskb_pull_tail() as a signed int.  The
function does not ensure that max is at least skb_headlen(skb).

This can happen while parsing IPv6 extension headers when an skb already
has a linear area larger than MAX_IPV6_HDR_LEN.  Once the parser needs data
beyond the linear area, max - skb_headlen(skb) wraps and is converted to a
negative delta.  __pskb_pull_tail() then passes that negative length to
skb_copy_bits(), where it can become a very large copy length.

Pass the requested length itself as the pull bound at the three
extension-header call sites, so the delta can no longer go negative.

Fixes: 1431fb31ec ("xen-netback: fix fragment detection in checksum setup")
Suggested-by: Eric Dumazet <edumazet@google.com>
Signed-off-by: Shihuang Liu <shlomojune6@gmail.com>
Link: https://patch.msgid.link/20260919133604.50948-1-shlomojune6@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-23 17:41:22 -07:00
Mina Almasry
73b3fc67a4 net: devmem: document that bind-tx is unprivileged by design
Unlike bind-rx, which configures shared NIC RX queues to steer incoming
traffic into the caller's dmabuf and requires CAP_NET_ADMIN
(uns-admin-perm), bind-tx only DMA-maps the caller's dmabuf so the caller
can transmit from it on their own sockets without affecting other traffic
or device configuration.

Add a comment in netdev.yaml and above netdev_nl_bind_tx_doit() to make it
explicit that NETDEV_CMD_BIND_TX is unprivileged by design.

Signed-off-by: Mina Almasry <almasrymina@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
Acked-by: Daniel Borkmann <daniel@iogearbox.net>
Link: https://patch.msgid.link/20260921195545.493253-1-almasrymina@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-22 18:40:15 -07:00
Nicolai Buchwitz
3199557121 net: don't require the hwtstamp NDOs when a PHY provides timestamping
Removing the legacy ioctl fallback made both hwtstamp NDOs mandatory. A
device that only timestamps in its PHY implements neither, so
SIOCSHWTSTAMP fails with EOPNOTSUPP before anything looks at the PHY and
PTP stops working there.

The check only ever picked the legacy path. That path is gone, so drop it
and test where the NDOs are actually called.

SIOCGHWTSTAMP is new here, not restored. The old path went through
phy_mii_ioctl(), which only handled SIOCSHWTSTAMP.

Such a device now returns -ENODEV while absent instead of -EOPNOTSUPP,
like the ones that do implement the NDOs.

Fixes: 5062245a5a ("net: remove legacy way to get/set HW timestamp config")
Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
Reviewed-by: Kory Maincent <kory.maincent@bootlin.com>
Link: https://patch.msgid.link/20260918095540.34286-1-nb@tipi-net.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-21 16:37:59 -07:00
Eric Dumazet
a5117e1ecc net: skbuff: do not leave stale header offsets after pskb_carve()
pskb_carve_inside_header() and pskb_carve_inside_nonlinear() remove
the first bytes of a packet and reallocate skb->head.

All the headers that were present before the operation are gone,
but both functions call skb_headers_offset_update(skb, 0), which
is a no-op : skb->mac_header, skb->network_header,
skb->transport_header and skb->csum_start keep their old values and
now describe bytes which are no longer there.

Both helpers size the new head from the old skb_end_offset(), so the
stale offsets still land inside the new allocation. They point past
skb_tail_pointer() though, to bytes that were never initialized.

pskb_carve_inside_nonlinear() is the worst case, because it leaves a
zombie skb with an empty linear part (skb->data ==
skb_tail_pointer(skb), skb_headlen(skb) == 0), while
skb_mac_header_was_set() is still true and skb->mac_header is way
ahead of skb->data.

The only user of pskb_extract() is rds_tcp_data_recv(), and the
carved skb is queued on tinc->ti_skb_list. When the RDS incoming
message is released, rds_tcp_inc_free() calls skb_queue_purge(),
which frees the skbs with SKB_DROP_REASON_QUEUE_PURGE. This is
visible from drop_monitor, which then tries to pull back to the
(bogus) mac header :

skbuff: __skb_pull(len=234)
skb len=6968 data_len=6968 headroom=0 headlen=0 tailroom=0
end-tail=384 mac=(234,14) mac_len=14 net=(248,40) trans=288
shinfo(txflags=0 nr_frags=1 gso(size=1428 type=16 segs=5))
csum(0x100120 start=288 offset=16 ip_summed=3 complete_sw=0 valid=1 level=0)
hash(0x7b446c6c sw=0 l4=1) proto=0x86dd pkttype=0 iif=60
kernel BUG at ./include/linux/skbuff.h:2847!

Add skb_carve_reset_headers() to mark the mac and transport headers
as not set, reset the network header, clear skb->mac_len, and drop
a now meaningless CHECKSUM_PARTIAL (csum_start no longer describes
anything).

Invalidate the inner offsets as well. Unlike mac_header and
transport_header they have no "unset" sentinel, so a leftover
non-zero value still looks like a real header. Zero
skb->inner_mac_header, skb->inner_network_header,
skb->inner_transport_header, skb->inner_protocol and
skb->encapsulation, so that all the header state is invalidated in
one place.

v2: fixed an inaccurate changelog. The stale offsets stay inside the
    new skb->head, which is never smaller than the old one, they
    simply point past skb_tail_pointer() to bytes that are gone.
    Thanks to Xuanqiang Luo for insisting on this.
    Also invalidate the inner header state, as suggested by the
    netdev AI review :
    https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260911114922.621937-1-edumazet%40google.com

Fixes: 6fa01ccd88 ("skbuff: Add pskb_extract() helper function")
Reported-by: syzbot+586af68eb819833c2d91@syzkaller.appspotmail.com
Closes: https://lore.kernel.org/netdev/6aa3e9d3.f2639fcc.29487d.0028.GAE@google.com/
Cc: Xuanqiang Luo <xuanqiang.luo@linux.dev>
Cc: Allison Henderson <achender@kernel.org>
Cc: rds-devel@oss.oracle.com
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Link: https://patch.msgid.link/20260915130423.3956471-1-edumazet@google.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-17 16:05:02 +02:00
Daniel Zahka
a41f24c612 net: psp: avoid conflicts with skb->decrypted and sk_validate_xmit_skb()
PSP conflicts with TLS ULP in its usage of both skb->decrypted and
sk->sk_validate_xmit_skb().

Make PSP mutually exclusive with TLS ULP, the only other user of either
of these. As other users of skb->decrypted come along, they can be added
to sk_has_decrypt_user(). It would make sense to also assert that
sk->sk_validate_xmit_skb() is also NULL in both of these setup paths for
similar future proofing, but the PSP listener/sk_clone() path is still
broken and it could be seen as a regression to not allow rx assoc to run
on a child of a listener socket with PSP tx assoc state.

Include all TCP ULPs in the sk_has_decrypt_user() check, even though TLS
is the only one that conflicts with PSP via the decrypted bit. This is
intentional because PSP was not designed to be used with ULPs. It is
best to close off surface area that may make bugs reachable, until
someone wishes to design and test an actual user of PSP with ULPs.

Fixes: 6b46ca260e ("net: psp: add socket security association code")
Signed-off-by: Daniel Zahka <daniel.zahka@gmail.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260915-psp-ktls-fix-v2-1-0eedc3b148ec@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 19:18:24 -07:00
Eric Dumazet
9ed55f3dbe net: lock the socket in sock_gettstamp()
sk->sk_flags must only be changed while holding the socket lock,
because sock_set_flag() and sock_reset_flag() use non atomic
operations (__set_bit() and __clear_bit()).

sock_gettstamp() is one of the last places where a bit of sk->sk_flags
is changed from a syscall without owning the socket lock, through
sock_enable_timestamp(sk, SOCK_TIMESTAMP).

sk_set_memalloc() and sk_clear_memalloc() also change sk->sk_flags
without the socket lock, but their callers (nbd, iscsi_tcp, nvme-tcp,
sunrpc, wireguard) need a careful audit, this will be addressed in a
separate patch.

Jungwoo Lee and Wongi Lee reported an UDP socket use-after-free
caused by this bug: a SIOCGSTAMPNS_NEW ioctl racing with bind()
can cancel the SOCK_RCU_FREE bit that udp_lib_get_port() just set,
because both threads perform a read-modify-write on the same word.

  CPU 0 (bind)                        CPU 1 (SIOCGSTAMPNS_NEW)
  --------------------------------    ----------------------------
  read sk_flags = F                   read sk_flags = F
  compute F | BIT(SOCK_RCU_FREE)      compute F | BIT(SOCK_TIMESTAMP)
  store F | BIT(SOCK_RCU_FREE)
  sk_add_node_rcu(sk, ...)
                                      store F | BIT(SOCK_TIMESTAMP)

After the lost update, SOCK_RCU_FREE is clear while the socket is
visible to lockless UDP receive lookups. sk_destruct() then frees
the socket immediately instead of waiting for a RCU grace period,
while the receive path still holds a reference-less pointer to it:

 BUG: KASAN: slab-use-after-free in ipv4_pktinfo_prepare+0x30/0x410
 Read of size 8 at addr ffff888008806610 by task exploit/207
 CPU: 0 UID: 1000 PID: 207 Comm: exploit Not tainted 6.12.95+ #1
  ipv4_pktinfo_prepare+0x30/0x410
  udp_queue_rcv_one_skb+0x51c/0x1180
  udp_unicast_rcv_skb+0x109/0x350
  ip_protocol_deliver_rcu+0x14b/0x310
  ip_local_deliver_finish+0x29d/0x390
  ip_local_deliver+0x24d/0x2a0

Only grab the socket lock when SOCK_TIMESTAMP has to be set,
to keep the common case lockless.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Reported-by: Jungwoo Lee <jwlee2217@gmail.com>
Reported-by: Wongi Lee <qw3rtyp0@gmail.com>
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260915043055.3441600-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 17:34:56 -07:00
Farhad Alemi
150dba2c69 net: remove WARN_ON_ONCE() from the dev_fill_forward_path() loop check
ipip_fill_forward_path() and ip6_tnl_fill_forward_path() look up the
route to the tunnel's remote endpoint and set ctx->dev to its device,
which is the tunnel itself when that route resolves back to the tunnel.
dev_fill_forward_path() then makes no progress and trips
WARN_ON_ONCE(last_dev == ctx->dev) as soon as a flowtable tries to
offload a flow through the tunnel. That routing loop is a configuration
any CAP_NET_ADMIN user can set up, and ip_tunnel_xmit() and
ip6_tnl_xmit() already treat it as a tx error, so remove the warning and
just fail the walk, as commit 008e7a7c29 ("net: remove WARN_ON_ONCE
when accessing forward path array") did for the path stack overflow.

Fixes: ab427db178 ("netfilter: flowtable: Add IPIP rx sw acceleration")
Fixes: d98103575d ("netfilter: flowtable: Add IP6IP6 rx sw acceleration")
Closes: https://lore.kernel.org/all/CA+0ovCgaRvbd0Udj70b2xxG8Cx3CaCpNhnf1V4RWQuDveZYZhA@mail.gmail.com/
Suggested-by: Pablo Neira Ayuso <pablo@netfilter.org>
Signed-off-by: Farhad Alemi <farhad.alemi@berkeley.edu>
Reviewed-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Link: https://patch.msgid.link/CA+0ovCgKDOk+Bg6Gh5Lwx94u_jJjQ30-vY1JcY2BYfhnWJJbPA@mail.gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-16 17:30:23 -07:00
Eric Dumazet
439f392084 drop_monitor: fix out-of-bounds write in reset_per_cpu_data()
In reset_per_cpu_data(), al is computed as:

    al = sizeof(struct net_dm_alert_msg);
    al += dm_hit_limit * sizeof(struct net_dm_drop_point);
    al += sizeof(struct nlattr);

    skb = genlmsg_new(al, GFP_KERNEL);
    ...
    nla = nla_reserve(skb, NLA_UNSPEC, sizeof(struct net_dm_alert_msg));
    ...
    msg = nla_data(nla);
    memset(msg, 0, al);

Because al includes sizeof(struct nlattr) (the 4-byte attribute header),
genlmsg_new() allocates al bytes of tailroom starting at nla.
However, msg points to nla_data(nla), which is located
sizeof(struct nlattr) bytes past nla. Calling memset(msg, 0, al)
therefore writes al bytes starting from msg, exceeding the allocated
buffer by sizeof(struct nlattr) (4 bytes) and corrupting
skb_shared_info.

Fix this by letting al represent only the payload length, allocating
the skb with genlmsg_new(nla_total_size(al), GFP_KERNEL), and zeroing
al bytes from msg.

Fixes: 683703a26e ("drop_monitor: Update netlink protocol to include netlink attribute header in alert message")
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Link: https://patch.msgid.link/20260910204612.3762015-5-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 17:58:35 -07:00
Eric Dumazet
c19b7d3508 drop_monitor: use raw_cpu_ptr() in tracepoint probes
syzbot reported a preemption warning in sk_skb_reason_drop():

 BUG: using smp_processor_id() in preemptible [00000000] code: syz.0.17/5917
 caller is net_dm_packet_trace_kfree_skb_hit+0x119/0x350 net/core/drop_monitor.c:519

In net_dm_packet_trace_kfree_skb_hit(), data = this_cpu_ptr(&dm_cpu_data)
is evaluated before spin_lock_irqsave(&data->drop_queue.lock, flags).
When kfree_skb() is called from preemptible context (e.g. process context
during close() on /dev/net/tun), preemption is enabled, triggering the
CONFIG_DEBUG_PREEMPT warning in smp_processor_id().

The same pattern exists in net_dm_hw_trap_summary_probe() and
net_dm_hw_trap_packet_probe() for dm_hw_cpu_data.

This is a false positive because each per-cpu structure is protected
by its own spinlock. If the task migrates to another CPU right after
reading the per-cpu pointer, the lock still safely synchronizes
access to that queue.

Use raw_cpu_ptr() instead of this_cpu_ptr() to silence
CONFIG_DEBUG_PREEMPT without disturbing interrupt state or breaking
PREEMPT_RT locking semantics.

Fixes: ca30707dee ("drop_monitor: Add packet alert mode")
Fixes: 5855357cd4 ("drop_monitor: Prepare probe functions for devlink tracepoint")
Reported-by: syzbot+dc57fd6722deb17e92af@syzkaller.appspotmail.com
Closes: https://lore.kernel.org/netdev/6aa316b2.f81106d8.2ab401.0014.GAE@google.com/
Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260910204612.3762015-4-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 17:58:35 -07:00
Eric Dumazet
c391a40f71 drop_monitor: use timer_shutdown_sync() to prevent timer rearming during teardown
In drop_monitor teardown paths (net_dm_trace_off_set(),
net_dm_hw_monitor_stop(), and error unwind paths in net_dm_trace_on_set()
and net_dm_hw_monitor_start()), per-CPU timers are stopped using
timer_delete_sync() followed by cancel_work_sync().

However, there is a circular dependency between send_timer and
dm_alert_work:
1) sched_send_work() (timer callback) schedules dm_alert_work.
2) send_dm_alert() / net_dm_hw_summary_work() calls reset_per_cpu_data()
   or net_dm_hw_reset_per_cpu_data().
3) If memory allocation fails under memory pressure in the reset
   function, it re-arms the timer via mod_timer(&data->send_timer, ...).

If dm_alert_work is running concurrently while timer_delete_sync()
executes on another CPU, an allocation failure in the worker will
re-arm the timer after timer_delete_sync() has already returned.
Once cancel_work_sync() completes and module_put() is called, the timer
remains active in the timer wheel. If the module is then unloaded, the
timer will fire and execute sched_send_work() in freed memory,
triggering a kernel panic / use-after-free.

Switch from timer_delete_sync() to timer_shutdown_sync(). This guarantees
that any in-flight timer handler has finished and prevents subsequent
re-arming attempts from running workers from succeeding. When monitoring
is restarted later, timer_setup() is invoked, which cleanly
re-initializes the timer.

Fixes: 9398e9c0b1 ("drop_monitor: Perform cleanup upon probe registration failure")
Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260910204612.3762015-3-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 17:58:34 -07:00
Eric Dumazet
6a038ef2b5 drop_monitor: synchronize tracepoint unregistration on error path
If register_trace_napi_poll() fails in net_dm_trace_on_set(),
unregister_trace_kfree_skb() is called to roll back the kfree_skb
tracepoint registration.

However, tracepoint_synchronize_unregister() is omitted before calling
cancel_work_sync() and module_put(). An in-flight probe executing
concurrently on another CPU could call schedule_work() after
cancel_work_sync() has already returned, leaving a pending work item
scheduled after the module reference is dropped. If the module is then
unloaded, executing the work item triggers a kernel panic.

Add tracepoint_synchronize_unregister() after unregister_trace_kfree_skb()
in the error path, matching net_dm_trace_off_set() and
net_dm_hw_probe_unregister().

Fixes: 7c747838a5 ("drop_monitor: Split tracing enable / disable to different functions")
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
Link: https://patch.msgid.link/20260910204612.3762015-2-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-15 17:58:34 -07:00
Kuniyuki Iwashima
979aabdad8 neighbour: Skip default parms when resumed in neightbl_dump_info().
neightbl_dump_info() calls neightbl_fill_info() in each loop
to render the default parms.

If there are many devices and neightbl_fill_param_info() failed,
neightbl_fill_info() is called again when the dump resumes:

  # ynl --family rt-neigh --dump getneightbl --output-json |
    jq '.[] | {name: .name, ifindex: .parms.ifindex}'
  ...
  {
    "name": "ndisc_cache",
    "ifindex": null
  }
  ...
  {
    "name": "ndisc_cache",
    "ifindex": 6
  }
  {
    "name": "ndisc_cache",
    "ifindex": null
  }
  {
    "name": "ndisc_cache",
    "ifindex": 5
  }

Let's skip neightbl_fill_info() if it is already called in
neightbl_dump_info().

Note that we cannot use !neigh_skip instead of !default_skip
because default_skip == 1 && neigh_skip == 0 could be true
if the first neightbl_fill_param_info() fails.

Also, nidx must be cleared at the end of each table loop;
otherwise, if neightbl_fill_info() for a subsequent table
fails, the leftover nidx from the previous table would be
saved in cb->args[1], resulting in erroneously skipping parms
of the subsequent table in the next dump.

Fixes: c7fb64db00 ("[NETLINK]: Neighbour table configuration and statistics via rtnetlink")
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260909233143.2401847-5-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-11 17:25:14 -07:00
Kuniyuki Iwashima
7b430fcfc9 neighbour: Don't render blackhole_netdev via RTM_GETNEIGHTBL.
The cited commits started to initialise blackhole_netdev with
neigh_parms_alloc().

This is visible in init_net as the ifindex==0 entries via
RTM_GETNEIGHTBL:

  # ynl --family rt-neigh --dump getneightbl --output-json \
    | jq '.[] | select(.parms.ifindex == 0)
              | {name: .name, ifindex: .parms.ifindex}'
  {
    "name": "arp_cache",
    "ifindex": 0
  }
  {
    "name": "ndisc_cache",
    "ifindex": 0
  }

For RTM_SETNEIGHTBL, ifindex being 0 means wildcard.

Let's skip blackhole_netdev's parms in neightbl_dump_info().

Note that lookup_neigh_parms() does not need the same change
because the default parms is always the first entry and matches
with ifindex == 0.

Fixes: e5f80fcf86 ("ipv6: give an IPv6 dev to blackhole_netdev")
Fixes: 22600596b6 ("ipv4: give an IPv4 dev to blackhole_netdev")
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260909233143.2401847-4-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-11 17:25:13 -07:00
Kuniyuki Iwashima
6d79b223ec neighbour: Enforce min/max to NDTPA_INTERVAL_PROBE_TIME_MS.
NDTPA_INTERVAL_PROBE_TIME_MS sets .type and .min but misses
.validation_type, so no validation is applied:

  # ynl --family rt-neigh --do setneightbl \
  --json '{"name": "arp_cache", "parms": {"interval-probe-time-ms": 0}}'

  # ynl --family rt-neigh --dump getneightbl --output-json | \
  jq '.[] | select(.name == "arp_cache" and has("config"))
          | .parms["interval-probe-time-ms"]'
  0

Moreover, nla_get_msecs() uses msecs_to_jiffies(), and u64 is
silently cast to u32, so a larger value can bypass the min check:

  e.g. 4294967296 == 0x100000000

  # ynl --family rt-neigh --do setneightbl \
  --json '{"name": "arp_cache", "parms": {"interval-probe-time-ms": 4294967296}}'

  # ynl --family rt-neigh --dump getneightbl --output-json | \
  jq '.[] | select(.name == "arp_cache" and has("config"))
          | .parms["interval-probe-time-ms"]'
  0

msecs_to_jiffies() returns MAX_JIFFY_OFFSET if the value is
larger than INT_MAX.  Also, INT_MAX ms overflows int NEIGH_VAR()
when HZ > 1000 (Alpha, MIPS), and passing a negative integer to
queue_delayed_work(unsigned long delay) causes sign extension,
which wraps around the expiry time to the past, resulting in it
being handled as 0 delay in the timer wheel.

Let's use NLA_POLICY_FULL_RANGE() and limit the max to 1 day.

The same max check is applied to sysctl as well.

Note that this controls the probe interval for NTF_MANAGED
entries, so the max of 1 day is unlikely to break any
deployments.

Fixes: 211da42eaa ("net, neigh: introduce interval_probe_time_ms for periodic probe")
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260909233143.2401847-3-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-11 17:25:13 -07:00
Kuniyuki Iwashima
764dcebb03 neighbour: Add missing RCU annotation for neightbl_dump_info().
neightbl_dump_info() fetches the first non-default neigh_parms
with list_next_entry(&tbl->parms, ...) and iterates through the
list with list_for_each_entry_from_rcu().

However, list_next_entry() does not use RCU helper.

Let's use list_for_each_entry_rcu() and skip the default parms.

Fixes: 4ae34be500 ("neighbour: Convert RTM_GETNEIGHTBL to RCU.")
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260909233143.2401847-2-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-11 17:25:13 -07:00
Jamal Hadi Salim
1aa9e143bf net: reject oversized tx_queue_len at netlink parse time
rtnl_create_link() assigns IFLA_TXQLEN directly to dev->tx_queue_len
without going through netif_change_tx_queue_len(), so a device created
with "ip link add ... txqueuelen 500000" bypasses the S16_MAX cap and
still triggers the oversized ring allocations in pfifo_fast, tun and
tap. The veth peer nest (rtnl_nla_parse_ifinfomsg()) and the
RTM_NEWLINK-on-existing-device path reach the same sinks.

Enforce the cap in ifla_policy instead: IFLA_TXQLEN becomes
NLA_POLICY_FULL_RANGE(NLA_U32, &txqlen_range) with
txqlen_range = { .min = 0, .max = S16_MAX }. All netlink consumers
parse against this policy - rtnl_setlink(), rtnl_newlink() (create
and change), and the veth peer nest - so every netlink path is capped
at parse time and rejects the attribute with -ERANGE plus a proper
"integer out of range" extack message before any device state is
modified (the RTM_SETLINK half-application wart is gone with it).

Document the bound in the rt-link.yaml netlink spec.

Conditions to recreate the bug:
- CONFIG_NET_SCHED=y, CONFIG_VETH=y, CONFIG_USER_NS=y, CONFIG_NET_NS=y.
- Unprivileged user in a fresh user+net namespace (unshare -Urn):
  ip link add v0 txqueuelen 500000 type veth peer name v1
  -> on the fixed kernel this is rejected with -ERANGE ("integer out
  of range" extack) instead of installing an oversized tx_queue_len
  that later inflates pfifo_fast/tun/tap ring allocations.
- ip link set v0 txqueuelen 500000 is likewise rejected at parse time.

Fixes: 38f7b870d4 ("[RTNETLINK]: Link creation API")
Reported-by: Vega <vega@nebusec.ai>
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-2899.v2.20260901233641@mojatatu.com.2
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 16:30:49 -07:00
Jamal Hadi Salim
66ab4c59b7 net: cap tx_queue_len at S16_MAX to prevent oversized ring allocations
Several subsystems allocate ring buffers sized by dev->tx_queue_len
with no upper bound. An unprivileged user (via unshare -Urn) can set a
huge tx_queue_len and exhaust global memory with ring allocations:

- pfifo_fast: pfifo_fast_init() and pfifo_fast_change_tx_queue_len()
  allocate 3 skb_array rings of tx_queue_len entries each.
- tun: tun_queue_resize() and the queue-attach path resize ptr_rings
  to tx_queue_len on the NETDEV_CHANGE_TX_QUEUE_LEN notifier.
- tap (macvtap/ipvtap): tap_queue_resize() and tap_init() resize/init
  ptr_rings to tx_queue_len on the same notifier.

netif_change_tx_queue_len() is the single entry point for IFLA_TXQLEN,
sysfs, and the SIOCSIFTXQLEN ioctl. Cap new_len at S16_MAX (32767)
there so the oversized value is rejected at set time. This takes
effect whether the device is up or down, before dev->tx_queue_len is
written, before any notifier fires, and before any ring is allocated.
The "> S16_MAX" check also subsumes the previous unsigned-long
truncation test, and a negative ifr_qlen from the ioctl lands far
above the cap after conversion, so both old failure modes are covered
by the one comparison.

tx_queue_len is ambigious: both a per-ring sizing multiplier and a
default queue-length/limit knob for consumers that allocate
nothing at set time (pfifo/bfifo/gred/plug/sfb limits, htb
direct_qlen, qfq max_classes, teql). 32767 is chosen as the largest
value NLA_POLICY_FULL_RANGE can express for the u32 IFLA_TXQLEN
policy in patch 2/3 while staying a legitimate queue length on
high-BDP paths; the ring-memory trade-off of a shared knob is
disclosed below.

Conditions to recreate the bug:
- CONFIG_NET_SCHED=y, CONFIG_VETH=y, CONFIG_USER_NS=y, CONFIG_NET_NS=y.
- Unprivileged user in a fresh user+net namespace (unshare -Urn).
- pfifo_fast: create veth pairs, set tx_queue_len to 500000, attach
  mq+pfifo_fast. ~28 iterations OOMs a 2GB guest.
- tun: create 50 tun devices with IFF_MULTI_QUEUE, set tx_queue_len to
  500000, open 8 queues each. ~1.6GB of ptr_ring allocations OOMs a
  512MB guest.
- tap: same as tun with IFF_TAP. ~960MB OOMs a 512MB guest.
- On the fixed kernel the oversized tx_queue_len is rejected with
  -ERANGE at set time (all four paths: RTM_SETLINK, RTM_NEWLINK
  create, sysfs, ioctl - the latter two via this check, the former
  two via this check and the 2/3 parse policy respectively).

Fixes: 6a643ddb56 ("net: introduce helper dev_change_tx_queue_len()")
Reported-by: Vega <vega@nebusec.ai>
Closes: https://lore.kernel.org/netdev/20260828121902.66837-1-jhs@mojatatu.com/
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/QDISC-2899.v2.20260901233641@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 16:30:48 -07:00
Fourie Zhang
78a86d75a7 net: mpls: clear inner_protocol when the last label is popped
skb_mpls_push() records the pre-encapsulation network header once, gated
on !skb->inner_protocol. skb_mpls_pop() never clears that record, so it
outlives the encapsulation it describes.

Open vSwitch can then re-push MPLS onto a packet whose
inner_network_header still points at the older, deeper offset: push a
label, pop every label, recirculate (ovs_flow_key_update() re-derives
key->eth.type and resets network_header, but leaves inner_*), then push
again. ovs_fragment() trusts the record:

	skb->network_header = skb->inner_network_header;

so skb_network_offset() goes negative. The bound check is signed:

	if (skb_network_offset(skb) > MAX_L2_LEN)

a negative offset passes it, and prepare_frag() widens the value:

	unsigned int hlen = skb_network_offset(skb);
	memcpy(&data->l2_data, skb->data, hlen);

which is a ~4GiB memcpy out of a 30-byte per-CPU buffer.

Reproduced on v7.3-rc1. RDX is the truncated length, (unsigned int)(-8):

  BUG: unable to handle page fault for address: ffffe8ffffc16000
  #PF: supervisor write access in kernel mode
  Oops: 0002 [#1] SMP KASAN NOPTI
  RIP: 0010:memcpy+0x8/0x20
  RDX: 00000000fffffff8 RSI: ffff888105d732db RDI: ffffe8ffffc16000
   prepare_frag+0x3df/0x4e0
   ovs_fragment+0x589/0x7e0
   do_output+0x4ce/0x5e0
   do_execute_actions+0x55d2/0x7b30
   ovs_execute_actions+0xea/0x450

Same root-cause shape as commit 975b5b067f ("ipv6: sr: restore network
header before routing and forwarding"): a stale network header offset
reaching a consumer that widens it. Here it originates in the MPLS
push/pop path.

Clear inner_protocol once the packet is no longer MPLS, so a later push
re-records the current header. net/sched/act_mpls.c is the only other
skb_mpls_pop() caller and gets the same fix; sch_frag.c saves and
restores inner_protocol around fragmentation in the same way OVS does.

Fixes: 48d2ab609b ("net: mpls: Fixups for GSO")
Cc: stable@vger.kernel.org
Signed-off-by: Fourie Zhang <fouriezhang@tencent.com>
Acked-by: Jiri Benc <jbenc@redhat.com>
Link: https://patch.msgid.link/20260902092719.2874481-1-fouriezhang@tencent.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-04 16:03:53 -07:00
Kuniyuki Iwashima
debac3a20d net: Remove conflicting altnames for dying netns in __dev_change_net_namespace().
syzbot reported the warning in cfg80211_pernet_exit(). [0]

The repro does the following:

  1. create two device in root netns and non-root netns
  2. assign the same altname for the two devices
  3. remove the non-root netns

Since commit 7663d52209 ("net: check for altname conflicts
when changing netdev's netns"), cfg80211_switch_netns() and
cfg802154_switch_netns() fail if init_net has a device with the
conflicting altname.

default_device_exit_net() had the same issue and commit d09486a04f
("net: fix removing a namespace with conflicting altnames") fixed it.

cfg80211_pernet_exit() and cfg802154_pernet_exit() need the same fix.

Let's generalise the fix by removing conflicting altnames for dying
netns in __dev_change_net_namespace().

[0]:
cfg80211_switch_netns(rdev, &init_net)
WARNING: net/wireless/core.c:1871 at cfg80211_pernet_exit+0xd5/0x120 net/wireless/core.c:1871, CPU#1: kworker/u8:9/1160
Modules linked in:
CPU: 1 UID: 0 PID: 1160 Comm: kworker/u8:9 Not tainted syzkaller #0 PREEMPT(full)
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 07/24/2026
Workqueue: netns cleanup_net
RIP: 0010:cfg80211_pernet_exit+0xd5/0x120 net/wireless/core.c:1871
Code: e8 03 42 80 3c 20 00 74 08 4c 89 f7 e8 b4 ef 0e f7 4d 8b 36 49 81 fe 20 10 4a 90 74 12 e8 03 3d 9f f6 eb 85 e8 fc 3c 9f f6 90 <0f> 0b 90 eb cc e8 f1 3c 9f f6 eb 05 e8 ea 3c 9f f6 5b 41 5c 41 5e
RSP: 0018:ffffc900057a78f0 EFLAGS: 00010293
RAX: ffffffff8b287154 RBX: ffff88807ba72780 RCX: ffff8880213e8000
RDX: 0000000000000000 RSI: 00000000ffffffef RDI: 0000000000000000
RBP: 00000000ffffffef R08: ffffffff9024cc67 R09: 0000000000000000
R10: fffff52000af4eb0 R11: fffffbfff204998d R12: dffffc0000000000
R13: ffffffff904a1080 R14: ffff888144ed0008 R15: ffff888144ed0e20
FS:  0000000000000000(0000) GS:ffff888124de6000(0000) knlGS:0000000000000000
CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 00005642de0a8a70 CR3: 000000007a40c000 CR4: 00000000003526f0
Call Trace:
 <TASK>
 ops_exit_list net/core/net_namespace.c:200 [inline]
 ops_undo_list+0x43d/0x8d0 net/core/net_namespace.c:253
 cleanup_net+0x572/0x810 net/core/net_namespace.c:706
 process_one_work kernel/workqueue.c:3387 [inline]
 process_scheduled_works+0xc3d/0x1630 kernel/workqueue.c:3470
 worker_thread+0xa47/0xfb0 kernel/workqueue.c:3551
 kthread+0x38b/0x480 kernel/kthread.c:436
 ret_from_fork+0x514/0xb70 arch/x86/kernel/process.c:158
 ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
 </TASK>

Fixes: 36fbf1e52b ("net: rtnetlink: add linkprop commands to add and delete alternative ifnames")
Reported-by: syzbot+74f338e09f1ef3ee6457@syzkaller.appspotmail.com
Closes: https://lore.kernel.org/all/6a96219e.04428c52.29b18.0001.GAE@google.com/T/
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://patch.msgid.link/20260901005550.2042357-1-kuniyu@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-02 18:30:00 -07:00
Norbert Szetei
1d2929d085 net: psp: do not inherit the Rx association on clone
sk->psp_assoc sits past sk_dontcopy_end, so sock_copy() copies it into
every socket accepted from a listener without taking a reference, while
inet_sock_destruct() puts for every inet socket. psp_twsk_init() does
refcount_inc() for the timewait socket, so a child closing through
TIME_WAIT cancels its own put and leaves the association with one
reference and N timewait sockets holding the same pointer. Closing the
listener frees it, and the timewait timers then put freed memory.

Rejecting the association on a listening socket is not sufficient: a socket
can acquire one while established and then be turned back into a listener,
because tcp_disconnect() leaves sk->psp_assoc in place.

  BUG: KASAN: slab-use-after-free in psp_twsk_assoc_free+0x6f/0xf0
  Write of size 4 at addr ffff888110f9255c by task swapper/7/0
   psp_twsk_assoc_free+0x6f/0xf0
   inet_twsk_put+0xda/0x1b0
   call_timer_fn+0x53/0x2e0
   __run_timers+0x764/0xa80
  Freed by task 99:
   kfree+0x1a7/0x500
   process_one_work+0x7ec/0x1100

An association carries a per-connection SPI and key, so a child must not
inherit the parent's. Clear it on clone.

Fixes: 6b46ca260e ("net: psp: add socket security association code")
Cc: stable@vger.kernel.org
Assisted-by: Claude:claude-opus-5
Signed-off-by: Norbert Szetei <norbert@doyensec.com>
Reviewed-by: Daniel Zahka <daniel.zahka@gmail.com>
Link: https://patch.msgid.link/BC10EB92-ABB3-41B2-AB16-266BEEBE18C0@doyensec.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-01 15:12:24 +02:00
Florian Schauer
dc0df5a0c6 page_pool: keep frag_offset aligned for odd-sized requests
page_pool_alloc_frag_netmem() rounds the requested fragment size with

	size = ALIGN(size, dma_get_cache_alignment());

dma_get_cache_alignment() returns 1 unless the architecture defines
ARCH_DMA_MINALIGN, which DMA-coherent architectures such as x86 do not.
There the ALIGN() is a no-op and pool->frag_offset advances by the raw,
unrounded size.

A single caller asking for an odd size then leaves frag_offset misaligned
for every fragment carved out of that page afterwards.  The pool is shared,
so the damage is not confined to the caller that caused it.

The per-cpu system_page_pool used by generic XDP hits this.
skb_pp_cow_data() allocates its fragments with the raw packet length:

	size = min_t(u32, len, PAGE_SIZE);
	truesize = size;
	page = page_pool_dev_alloc(pool, &page_off, &truesize);

leaving frag_offset odd for whatever is carved out of that page next.  Its
own head allocation is already aligned -- SKB_HEAD_ALIGN(size) plus the
XDP_PACKET_HEADROOM its callers pass -- so it is a later user of the shared
pool that pays: page_pool_dev_alloc_va() returns a misaligned buffer,
napi_build_skb() installs it as skb->head, and skb_shinfo(skb) ==
skb->head + skb->end is misaligned with it.

skb_shinfo()->dataref is a 4-byte atomic_t at offset 0x20, so the
atomic_inc() in __skb_clone() straddles a cache line.  On x86 with split
lock detection -- fatal for kernel split locks by default -- this panics
the machine:

  Oops: Split lock detected
  RIP: 0010:skb_clone+0x154/0x1e0
  Call Trace:
   <IRQ>
   raw_local_deliver+0x1ed/0x2c0
   ip_protocol_deliver_rcu+0x54/0x1c0
   ip_local_deliver_finish+0x85/0x100
   ip_local_deliver+0x67/0x100
   __netif_receive_skb_one_core+0x85/0xa0
   process_backlog+0x87/0x130

Reproduced by attaching any generic-mode XDP program to loopback and
opening a RAW IPPROTO_UDP socket, which makes raw_local_deliver() clone
every locally delivered UDP packet; ordinary DNS traffic then triggers it,
roughly once per 2500 clones.  Observed on 6.12.101 and 7.1.8.

Tracing page_pool_alloc_frag_netmem() over one such run shows the
amplification -- two odd-sized requests, nine misaligned offsets:

  requested size & 7:   0: 17035    5: 1    7: 1
  frag_offset & 7:      0: 17028    3: 1    4: 1    5: 1    6: 1    7: 5

and skb_pp_cow_data() returning heads that were aligned on entry:

  head 0xffff8f4c86aeac00 -> 0xffff8f4c53a9a9c4 (&7=4)
  head 0xffff8f4d6a8a42c0 -> 0xffff8f4c4f7b7a45 (&7=5)

Round the fragment size up to at least the alignment struct skb_shared_info
requires, so fragments are always suitably aligned for the objects callers
build on them.  Architectures needing a larger DMA alignment keep it.

This also makes the remainder computed in page_pool_alloc_netmem(),

	*size = max_size - *offset;

aligned, since max_size is a power of two -- which fixes the matching
misalignment of skb->end.

Verified with a controlled A/B under QEMU/KVM: same tree, same config,
same compiler, same rootfs and identical traffic, differing only by this
patch.  A SEC("xdp.frags") XDP_PASS program on lo plus UDP datagrams
larger than max_head_size drives skb_pp_cow_data()'s fragment loop, which
passes raw packet lengths to the pool.  Measured at the return of
skb_pp_cow_data():

                          unpatched   patched
  skb_pp_cow_data calls       40800     40800
  misaligned skb->head         1120         0
  dataref at line offset >60     80         0

The last row counts the accesses that actually fault:
skb_shinfo()->dataref sits at head+end+0x20 and is a 4-byte atomic, so
`lock incl` splits a 64-byte cache line only when that address lands at
offset 61..63.  All 80 occurrences were at offset 61; the panic reported
above was at offset 62.  Eliminating the misalignment removes every one
of them.

Same class of bug as commit 3bed3cc415 ("net: Do not allocate page
fragments that are not skb aligned"), which fixed the older
netdev_alloc_frag()/napi_alloc_frag() allocators.

Fixes: 53e0961da1 ("page_pool: add frag page recycling support in page pool")
Cc: stable@vger.kernel.org
Signed-off-by: Florian Schauer <florian@schauer.to>
Acked-by: Jesper Dangaard Brouer <hawk@kernel.org>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260828060822.2628276-1-florian@schauer.to
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-31 16:24:51 -07:00
Dong Chenchen
28a57fb2c5 net: iptunnel: fix stale transport header during tunnel decapsulation
Syzbot reported a crash in qdisc_pkt_len_segs_init() caused by a stale
transport_header offset after tunnel decapsulation.

BUG: unable to handle page fault for address: ffffed102091a42e
Oops: Oops: 0000 [#1] SMP KASAN NOPTI
CPU: 0 UID: 0 PID: 340 Comm: qdisc_uaf_repro Not tainted 7.2.0-rc4-00061-g248951ddc14d #256 PREEMPT(full)
Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
RIP: 0010:__asan_load2
<IRQ>
qdisc_pkt_len_segs_init (net/core/dev.c:4145)
__dev_queue_xmit (net/core/dev.c:4787)
br_dev_queue_push_xmit (net/bridge/br_forward.c:53)
br_handle_frame_finish (net/bridge/br_input.c:229)
br_handle_frame (net/bridge/br_input.c:315)
__netif_receive_skb_core.constprop.0 (net/core/dev.c:6099)
__netif_receive_skb_list_core (net/core/dev.c:6287)
netif_receive_skb_list_internal (net/core/dev.c:6445)
napi_complete_done (net/core/dev.c:6813)
gro_cell_poll (net/core/gro_cells.c:74)
__napi_poll (net/core/dev.c:7735)
net_rx_action (net/core/dev.c:7798 net/core/dev.c:7955)
handle_softirqs (kernel/softirq.c:622)
do_softirq (kernel/softirq.c:523  kernel/softirq.c:510 )
__local_bh_enable_ip (kernel/softirq.c:450)
tun_get_user (drivers/net/tun.c:1986 (discriminator 1))
tun_chr_write_iter (drivers/net/tun.c:2032)

The issue is completely latent until qdisc read transport header in
commit 7fb4c19670 ("net: pull headers in qdisc_pkt_len_segs_init()").
The crash requires four conditions to line up:

1. The incoming packet is encapsulated and carries GSO metadata. The outer
   transport header offset is stored in skb->transport_header while the
   packet is still in the outer tunnel context.
2. The tunnel receiver strips the outer headers. skb->data is advanced to
   the inner frame, but skb->transport_header is left pointing to the
   now-removed outer L4 header, so it becomes a negative offset relative to
   the new data.
3. The inner frame is not delivered to the local IP stack. Instead, it
   is forwarded at L2 by a bridge or HSR, so ip_rcv_core() never runs and
   the transport header is not reset to the inner L4 offset.
4. The forwarding path calls __dev_queue_xmit(), which enters
   qdisc_pkt_len_segs_init(). That function computes the GSO header length
   from skb_transport_offset(skb). Because the offset is negative, the
   unsigned cast overflows and pskb_may_pull(skb, hdr_len +
   sizeof(struct tcphdr)) reads past the end of the skb, triggering a
   KASAN fault or page fault.

The issue specifically requires GSO packets (shinfo->gso_size != 0), which
are processed/aggregated through gro_cells. Fix this by clearing
transport_header to the ~0U sentinel in gro_cell for all tunnnel driver.
GTP does not support GRO/GSO, drop the evil GSO packets in GTP directly.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Reported-by: syzbot+83181a31faf9455499c5@syzkaller.appspotmail.com
Closes: https://lore.kernel.org/all/69de2bee.a00a0220.475f0.0041.GAE@google.com/T/
Suggested-by: Eric Dumazet <edumazet@google.com>
Signed-off-by: Dong Chenchen <dongchenchen2@huawei.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260825123909.1463121-1-dongchenchen2@huawei.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-28 15:53:46 -07:00
Alice Mikityanska
0b13256ce3 net: Guard for gso_segs overflow in skb_segment
skb_segment calculates 32-bit partial_segs as len / gso_size, and then
assigns it to the 16-bit gso_segs field. The division might overflow in
some edge cases where the SKB is BIG TCP (65536 <= len <= 8*65535), and
gso_size < TCP_MIN_GSO_SIZE = 8. While normally this can't happen due to
TCP_MIN_GSO_SIZE, an AF_PACKET PACKET_VNET_HDR socket could generate
such a malformed packet until the previous patch.

Blocking malformed virtio_net packets was implemented in the previous
patch, but this patch clamps partial_segs in skb_segment itself for more
generic robustness. Should len / gso_size happen to be bigger than
65535 in partial GSO, skb_segment will now just produce more than two
output SKBs, all of which will be valid with gso_segs <= 65535.

In order to catch possible other cases of too many partial_segs, add a
DEBUG_NET_WARN_ON_ONCE when len / gso_size happens to be too big.

Signed-off-by: Alice Mikityanska <alice@isovalent.com>
Link: https://patch.msgid.link/20260822120117.1163423-3-alice.kernel@fastmail.im
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-08-27 15:47:18 +02:00
Tetsuo Handa
3220b62fbb net: fix a resource leak in copy_net_ns() error handling path
Currently, preinit_net() does two things:

  (1) call ns_common_init() which might fail
  (2) initialize resources which does not fail

However, preinit_net() is returning early when (1) fails, and copy_net_ns()
is jumping to the dec_ucounts: label. As a result, resources allocated by
net_alloc() are leaking. We need to call key_remove_domain() and
net_passive_dec() in order to release resources allocated by net_alloc().

We cannot simply jump to the put_userns: label when preinit_net() failed,
for (2) is not yet done. But we can reorder (1) and (2), for there is no
dependency between (1) and (2). Therefore, this patch decouples (1) from
preinit_net() and changes preinit_net() back to a void function, and calls
ns_common_init() after preinit_net() succeeded. Then, we can jump to
immediately after ns_common_free() of the put_userns: label.

Reported-by: sashiko (no mail address)
Closes: https://sashiko.dev/#/patchset/af7dabf3-d0d7-46dc-a878-e1715b3c9ac6%40I-love.SAKURA.ne.jp
Fixes: 08027f6b79 ("net: use ns_common_init()")
Signed-off-by: Tetsuo Handa <penguin-kernel@I-love.SAKURA.ne.jp>
Link: https://patch.msgid.link/c182cf90-1ed7-435b-88f7-9f00e88a0487@I-love.SAKURA.ne.jp
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-08-25 12:30:55 +02:00
Mina Almasry
97148bcb75 net: core: fix head-page leak in skb_zerocopy
When skb_orphan_frags() throws -ENOMEM, skb_copy_ubufs() may have
already reallocated and replaced 'from->head'. Accessing from->head to
drop the old refcount leaks the original head page, and erroneously
puts an unrelated new buffer. Use the local 'page' tracker variable
instead to drop the reference properly.

Fixes: 36d5fe6a00 ("core, nfqueue, openvswitch: Orphan frags in skb_zerocopy and handle errors")
Signed-off-by: Mina Almasry <almasrymina@google.com>
Link: https://patch.msgid.link/20260823183602.1051453-2-almasrymina@google.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-08-25 11:28:44 +02:00
Mina Almasry
00e11ee983 net: core: check skb_frags_readable before uncloning in skb_copy_ubufs
skb_copy_ubufs drops clones and modifies the SKB via pskb_expand_head()
before checking for !skb_frags_readable(skb). This alters the SKB
geometry prior to throwing an -EFAULT on an invalid SKB. Check
readability first.

Fixes: 65249feb6b ("net: add support for skbs with unreadable frags")
Signed-off-by: Mina Almasry <almasrymina@google.com>
Link: https://patch.msgid.link/20260823183602.1051453-1-almasrymina@google.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-08-25 11:28:40 +02:00
Norbert Szetei
f66bdb1cc0 net: skbuff: don't touch shared zerocopy state in skb_tx_error()
skb_tx_error() completes the zerocopy uarg and clears
SKBFL_ALL_ZEROCOPY, and skb_zcopy_downgrade_managed() clears
SKBFL_MANAGED_FRAG_REFS. Both live in skb_shinfo(), which every clone
shares, while the caller only owns the reference it is about to drop.
Through a clone it tells the producer its pages are free and drops
SKBFL_SHARED_FRAG for an skb that is still in flight.

Open vSwitch reaches this with a non-last OVS_ACTION_ATTR_RECIRC:
clone_execute() sends a skb_clone() into ovs_dp_process_packet() while
do_execute_actions() keeps forwarding the original, and skb_clone()
does not privatise the frags here -- skb_orphan_frags() returns early
on SKBFL_DONT_ORPHAN. A flow miss on the clone then strips the marker
from the packet still being forwarded, and a later local ESP delivery
decrypts in place over frags it does not own privately.

Skip it for a cloned skb. Nothing is lost: skb_release_data() clears
the zerocopy state once the last reference to the shared data goes.

Fixes: 25121173f7 ("skb: api to report errors for zero copy skbs")
Cc: stable@vger.kernel.org
Suggested-by: Ilya Maximets <i.maximets@ovn.org>
Signed-off-by: Norbert Szetei <norbert@doyensec.com>
Reviewed-by: Ilya Maximets <i.maximets@ovn.org>
Tested-by: Jongmin Jang <payload.jang@gmail.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/CFAB292A-674B-4C14-BB2C-BB8830AD5659@doyensec.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-08-25 09:36:47 +02:00
Norbert Szetei
8ece906150 net: skbuff: don't skb_tx_error() the source skb in skb_zerocopy()
skb_zerocopy() copies frags from @from into @to. On an
skb_orphan_frags() failure it calls skb_tx_error(@from), a destructive
operation on the source skb the copy helper does not own. That completes
@from's zerocopy uarg and clears SKBFL_ALL_ZEROCOPY, including the
SKBFL_SHARED_FRAG page-ownership marker.

Both callers already report the failure on their own drop path.
nfnetlink_queue does it at nla_put_failure, and Open vSwitch does it in
the flow-miss drop arm of ovs_dp_process_packet(), so nothing is lost by
dropping it here.

On Open vSwitch's OVS_ACTION_ATTR_USERSPACE path the skb is not freed on
this error: do_execute_actions() ignores output_userspace()'s return
value and, unless the upcall was the last action, keeps forwarding the
same skb through the flow's remaining actions. The uarg is completed
while that skb is still in flight, telling the producer its buffers are
free, and SKBFL_SHARED_FRAG is cleared on an skb the rest of the stack
still handles. That flag is what makes esp_input() call skb_cow_data()
instead of decrypting in place, so a later local ESP delivery can
decrypt over frags the skb does not own privately.

Leave error reporting to the callers.

Fixes: 36d5fe6a00 ("core, nfqueue, openvswitch: Orphan frags in skb_zerocopy and handle errors")
Cc: stable@vger.kernel.org
Suggested-by: Ilya Maximets <i.maximets@ovn.org>
Signed-off-by: Norbert Szetei <norbert@doyensec.com>
Reviewed-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/6E3A780D-FB87-421F-9964-B1D457D7D106@doyensec.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-08-25 09:36:47 +02:00
Rong Zhang
039f248a6c net: page_pool: Remove zone/policy GFP flags when allocating XArray entries
Net drivers request GFP flags according to both the current context and
the device constraints, but the XArray entry itself is by no mean used
by the device. Passing though device constraints to XArray allocation is
a bug and will be warned and fixed up by slab, e.g.:

    Unexpected gfp: 0x4 (GFP_DMA32). Fixing up to gfp: 0x82820 (GFP_ATOMIC|__GFP_NOWARN|__GFP_NOMEMALLOC). Fix your code!
    CPU: 2 UID: 0 PID: 1071629 Comm: kworker/u80:1 Not tainted 7.2.0-rc7+ #1 PREEMPT(lazy)
    Hardware name: LENOVO 21Q4/LNVNB161216, BIOS PXCN27WW 10/20/2025
    Workqueue: mt76 mt792x_pm_wake_work [mt792x_lib]
    Call Trace:
     <TASK>
     dump_stack_lvl+0x6e/0x90
     kmalloc_fix_flags+0x4d/0x6a
     refill_objects+0x10a/0x330
     __pcs_replace_empty_main+0x292/0x5c0
     kmem_cache_alloc_lru_noprof+0x4c2/0x680
     ? __xas_nomem+0x3a/0x120
     __xas_nomem+0x3a/0x120
     __xa_alloc+0xd4/0x190
     page_pool_dma_map+0xef/0x400
     __page_pool_alloc_netmems_slow+0xed/0x480
     ? lock_release+0x280/0x490
     page_pool_alloc_frag_netmem+0xe0/0x3a0
     page_pool_alloc_frag+0xe/0x20
     mt76_dma_rx_fill_buf+0x1f6/0x580 [mt76]
     mt76_dma_rx_reset+0x1cf/0x230 [mt76]
     mt792x_wpdma_reset+0x183/0x1b0 [mt792x_lib]
     mt792x_wpdma_reinit_cond+0x5e/0xa0 [mt792x_lib]
     mt792xe_mcu_drv_pmctrl+0x28/0x60 [mt792x_lib]
     mt792x_mcu_drv_pmctrl+0x3e/0x90 [mt792x_lib]
     mt792x_pm_wake_work+0x2d/0x1d0 [mt792x_lib]
     ? process_one_work+0x20e/0x600
     process_one_work+0x230/0x600
     ? process_one_work+0x256/0x600
     worker_thread+0x1ec/0x3c0
     ? rescuer_thread+0x610/0x610
     kthread+0xf2/0x130
     ? kthread_affine_node+0x140/0x140
     ret_from_fork+0x2a5/0x380
     ? kthread_affine_node+0x140/0x140
     ret_from_fork_asm+0x11/0x20
     </TASK>

Currently mt76 and stmmac may allocate page pool pages with GFP_DMA32.

Fix it by removing zone/policy GFP flags when allocating XArray entries.
This is inspired by commit 96d5780880 ("iommu/dma: Use the gfp
parameter in __iommu_dma_alloc_noncontiguous()").

Fixes: ee62ce7a1d ("page_pool: Track DMA-mapped pages and unmap them when destroying the pool")
Signed-off-by: Rong Zhang <i@rong.moe>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Link: https://patch.msgid.link/20260821-page-pool-xa-drop-dma32-v1-1-6eab295c3478@rong.moe
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:28:14 -07:00
Weiming Shi
71283aaa6c xdp: fix zero-copy frame layout
xdp_convert_zc_to_xdp_frame() clones an XSK packet into an order-0 page
and advertises PAGE_SIZE as its frame size.  It allows the copied frame
to occupy the page tail needed by skb_shared_info and records zero
headroom even when metadata separates the frame header from packet data.
An AF_XDP zero-copy packet redirected through cpumap can therefore make
the skb overlap skb_shared_info or place it beyond the allocated page.

Limit the copied layout to SKB_WITH_OVERHEAD(PAGE_SIZE) and include the
metadata length in frame headroom.  Redirect callers already handle a
NULL conversion result.

BUG: KASAN: slab-out-of-bounds in skb_gro_receive
Write of size 4 at addr ffff88800cf37004 by task cpumap/1/map:1/146
Call Trace:
 skb_gro_receive (net/core/gro.c:174)
 udp_gro_receive (net/ipv4/udp_offload.c:812)
 inet_gro_receive (net/ipv4/af_inet.c:1539)
 dev_gro_receive (net/core/gro.c:515)
 gro_receive_skb (net/core/gro.c:633)
 cpu_map_kthread_run (kernel/bpf/cpumap.c:395)
 kthread (kernel/kthread.c:436)
 ret_from_fork (arch/x86/kernel/process.c:164)
 ret_from_fork_asm (arch/x86/entry/entry_64.S:255)
Kernel panic - not syncing: KASAN: panic_on_warn set ...

Fixes: b0d1beeff2 ("xdp: implement convert_to_xdp_frame for MEM_TYPE_ZERO_COPY")
Cc: stable@vger.kernel.org
Reported-by: Xiang Mei <xmei5@asu.edu>
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Link: https://patch.msgid.link/20260818154516.793517-1-bestswngs@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:10:48 -07:00
Mina Almasry
68d8c65326 net: core: propagate unreadable flag in skb_zerocopy
skb_zerocopy() fails to propagate the unreadable flag when copying
unreadable fragments, causing target skbs to appear as readable memory.

This patch fixes the flag propagation. Additionally, it returns -EFAULT
if readable fragments are mixed with unreadable fragments during
extraction, and returns -EFAULT in openvswitch queue_userspace_packet().

Fixes: 65249feb6b ("net: add support for skbs with unreadable frags")
Cc: Stanislav Fomichev <sdf@fomichev.me>
Cc: Bobby Eshleman <bobbyeshleman@gmail.com>
Cc: Florian Westphal <fw@strlen.de>
Cc: Aaron Conole <aconole@redhat.com>
Cc: Eelco Chaudron <echaudro@redhat.com>
Cc: Willem de Bruijn <willemb@google.com>
Signed-off-by: Mina Almasry <almasrymina@google.com>
Reviewed-by: Pavel Begunkov <asml.silence@gmail.com>
Reviewed-by: Ilya Maximets <i.maximets@ovn.org>
Link: https://patch.msgid.link/20260814191336.187243-1-almasrymina@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:22:14 -07:00
Tetsuo Handa
f85dc137aa net: add missing ref_tracker_dir_exit() to net_passive_dec()
I found that trying to read /sys/kernel/debug/ref_tracker/* causes NULL
pointer dereference crash when alloc_netdev_mqs() via unshare() returned
NULL, for commit 9ba74e6c9e ("net: add networking namespace refcount
tracker") added ref_tracker_dir_exit(&net->refcnt_tracker) to only
__put_net() path whereas commit 65b584f536 ("ref_tracker: automatically
register a file in debugfs for a ref_tracker_dir") added
ref_tracker_dir_debugfs() to ref_tracker_dir_init() path.

Since preinit_net() calls ref_tracker_dir_init(&net->refcnt_tracker) and
ref_tracker_dir_init(&net->notrefcnt_tracker), we need to make sure that
both ref_tracker_dir_exit(&net->refcnt_tracker) and
ref_tracker_dir_exit(&net->notrefcnt_tracker) are called before
net_passive_dec() schedules for kmem_cache_free() via net_complete_free().

ref_tracker_dir_exit(&net->refcnt_tracker) is called via put_net() when
ns_ref_put() returned true. But put_net() is not called when copy_net_ns()
fails. Therefore, call ref_tracker_dir_exit() from net_passive_dec() if
put_net() is not yet called.

Link: https://sashiko.dev/#/patchset/b06ce35d-e7bc-47a5-8e0a-e82be7e4dd08%40I-love.SAKURA.ne.jp
Fixes: 9ba74e6c9e ("net: add networking namespace refcount tracker")
Reviewed-by: Eric Dumazet <edumazet@google.com>
Signed-off-by: Tetsuo Handa <penguin-kernel@I-love.SAKURA.ne.jp>
Link: https://patch.msgid.link/64254d80-9248-466c-8108-95f43bd71117@I-love.SAKURA.ne.jp
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 12:44:12 -07:00
Linus Torvalds
91ec203513 Networking changes for 7.3.
Core & protocols
 ----------------
 
  - A few steps lowering rtnl_lock dependence:
    - per-netns netdev unregistration for select SW drivers
      (e.g. veth, ipvlan, tunnels)
    - rtnl_lock-less FIB rule changes (RTM_NEWRULE and RTM_DELRULE)
    - prepare software drivers and TC qdiscs for rtnl_lock-less GET
 
  - Support BIG TCP (>64kB TSO) in UDP tunnels (vxlan, geneve).
 
  - Support buffers larger than PAGE_SIZE in devmem zero-copy API.
 
  - Improve MPTCP handling of extreme memory pressure handling,
    when out-of-order queue had to be pruned.
 
  - Report the per-group user count via RTM_GETMULTICAST.
 
  - Expose the route deletion reason in RTM_DELROUTE.
 
  - Add a SO_RIGHTS_NOTRUNC option to UNIX sockets to enable more useful
    handling of LSM denials when receiving SCM_RIGHTS messages: instead
    of truncating the message at the first blocked fd, keep every fd slot
    and store the LSM errno in the blocked slot.
 
  - IPv6 Segment Routing - support looking up the post-encap SID
    (address) in a different/specified routing table.
 
  - Support PRP RedBox (interlink) creation.
 
  - Support per-nexthop UDP dst port in VXLAN.
 
  - Continue converting getsockopt callbacks in a number of protocols
    to iov_iter.
 
 Ethernet
 --------
 
  - Marge initial CXL support for AMD/Solarflare NICs (shared branch
    with the CXL tree).
 
  - New drivers:
    - ADIN1140 10BASE-T1S MACPHY
    - Initial skeleton of Intel iXD and ZTE Dinghai drivers.
 
  - High-speed NICs:
    - AMD/Pensando:
      - support firmware flashing
    - Cisco (enic):
      - SR-IOV V2 admin channel and MBOX protocol
    - Huawei (hns3):
      - support for ethtool pfc_prevention_tout
    - nVidia/Mellanox:
      - support sharing bandwidth control across interfaces of
        the same device
    - Marvell (octeontx2-pf):
      - link RQ page pools to netdev for Netlink stats
    - Google vNIC:
      - XDP metadata support for DQ RDA
    - Microsoft vNIC:
      - support forcing full-page RX buffers
 
  - Other NICs:
    - Synopsys IP:
      - eic7700: support for eth1
    - Microchip (lan743x):
      - support for RMII interface
    - Wangxun:
      - support for ethtool -G and -C for VFs
      - add Tx timeout and PCIe error handling
    - Intel (igb/igc):
      - RSS key get/set support
      - support for forcing link speed without auto-negotiation
 
  - Switches:
    - NXP (dpaa2):
      - support bonding/LAG offload
    - Mediatek:
      - mt7530: EN7528 support
      - initial support for MT7628
    - Micrel (ksz8/9):
      - refactoring work to move towards library model
      - PTP support for KSZ8463
    - nVidia/Mellanox:
      - support rtnl-lock-less ethtool callbacks
    - Realtek:
      - rtl8366rb: use generic RTL83xx code
      - support SGMII and HSGMII for RTL8367S
 
  - PHYs:
    - Airoha:
      - EcoNet EN7528 PHY support
    - DAPU Telecom
      - DAPU Telecom DAP8211R(I) Gigabit PHY support
    - Realtek:
      - support RTL8261C_CG
      - support RTL8261D
 
 Wireless
 --------
 
  - nl80211: per-link statistics support for multi-link operation
 
  - mac80211: AQL/airtime-fairness support for multicast
 
  - Merge Peripheral Authentication Service (PAS) / TEE support
    for ath12k (shared branch with the firmware/qcom tree).
 
  - New drivers:
    - mm81x for Morse Micro Long-Range S1G devices
    - nxpwifi for NXP devices (mostly forked off from mwifiex)
 
  - Driver changes:
    - Broadcom (brcmfmac):
      - DPP support, some Cypress part update
    - MediaTek (mt76):
      - mt7928 support
      - mt7925 NAN support
      - mt7996 AP powersave improvements
    - Qualcomm (ath12k):
      - much kernel infrastructure integration work
      - AHB platform MultiPD support
    - Realtek (rt89):
      - LED support
      - RTL8922DE support
      - dual-BT coex for RTL8922D
    - Intel:
      - new FW version support
 
 Bluetooth
 ---------
 
  - HCI: add support for Shorter Connection Interval (SCI) feature.
 
  - af_bluetooth: add minimal context analysis annotations.
 
  - Driver changes:
    - Intel:
      - add Bluetooth SAR revision 2 support
      - add vendor_reset PCI sysfs for PLDR
    - Mediatek:
      - add USB IDs for MT7902 and MT7922 devices
    - Realtek:
      - add USB IDs for 8761CU and 8852BE devices
    - NXP:
      - add M.2 Bluetooth device support using pwrseq
 
 Misc
 ----
 
  - DPLL support for manual/numerical oscillator control (NCO)
    (implement in zl3073x).
 
  - MCTP support for MCTP over USB v1.1 (DMTF DSP0283).
 
  - Power-over-Ethernet: support Realtek PSE controllers.
 
  - Remove the IBM EHEA driver.
 
  - Remove tulip/xircom_cb driver.
 
 Signed-off-by: Jakub Kicinski <kuba@kernel.org>
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEE6jPA+I1ugmIBA4hXMUZtbf5SIrsFAmqEwJ4ACgkQMUZtbf5S
 Irsegw//fmHJae525nxg3DHoXhrUz8EDDOVoLH6oyWyLQnh5bmbReAY/+oWA4m54
 3KKKO0b2rtgRvmY/7rnjAt3bjecYgCSjvZT7I+NosB0QbbBYc14PtHfYig9HffYm
 uCXfNJOk+aJ2QK4ncEvU2SjgE89Ya7cC+yARFBAwYx4zi/Qx24RB+ziOyvkQ8ksX
 atvMOZrnhwqvYUFOwnOLNHTpvdxB/ZsNwWY6iXcx6EYp9xrtPusbh3FlushWkwxH
 8cI/dNla44TcIKXAzRn0znRdgiEVmCMyHvOv7LKaOfy8P3I+knmuIf/mScYQqOEF
 T143HdXhVSBZFRtLtFKXIja/KsvCjX9lCeMn/2ak0brQDUREcacXxYbuZKDsNAAK
 zXt/+5qAcm/mO8W1gKR9Ulfli5bhFN4HKXgXMLjo5ucPtzfPxFN7HGxTiC3Cxv1v
 lSXexKaj74pNBVFmADrb5jWbq7oG+GzIdjzx3ycvm2q39Fr4nJ2SzrSPPNwc/ItQ
 IHv3tGLQKXlr8dl0+p2mDkRInmHXrawVNsB1UgN8E/jtcwT2QMwyWOV6s5G3uEDl
 a+0U/XsrPvDYBTUCRs/KaOJQGB90QkzLe9DATt159mf+rPzAX2/oCDo8xIEe+kWV
 aivP+YutFfMH/CSC9PMuvdLE2KmoPY4mibAeE4/4AYLKtJnc/yU=
 =zDto
 -----END PGP SIGNATURE-----

Merge tag 'net-next-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net-next

Pull networking updates from Jakub Kicinski:
 "One of the 'small improvements all over the place' releases for us.

  It's hard to draw any direct comparisons because summer vacations
  disrupted our patch processing (and presumably - generation) quite a
  bit.

  Quick and dirty count suggests we (Paolo and I) merged a very similar
  number of net (632) and net-next (648) patches. This is not telling
  the full story either because 1/3 to 1/2 of the net-next patches also
  *seem* like AI-driven low priority fixes, cleanups and clarifications.

  We are completely overwhelmed, of course. The glimmer of hope is that
  we secured sufficient LLM budget and access (thank you Meta!) to run
  reviews with multiple frontier models on each patch. This eliminates
  some hallucinations. That said, in terms of review, the LLMs can only
  do so much.

  The sad truth is that our APIs (especially for rare events like PCIe
  errors, timeouts etc) have always been racy, and now LLMs don't let us
  ignore that. I expect our direction for the next release will be to
  tweak the reviews a little bit more, but start shifting focus to
  letting the LLMs take care of the busy work - managing patchwork,
  automating common process complaints, editing commit messages, and
  maybe applying patches which already got "reviewed-by" tags from
  people we trust...

  Core & protocols:

   - A few steps lowering rtnl_lock dependence:
      - per-netns netdev unregistration for select SW drivers (e.g.
        veth, ipvlan, tunnels)
      - rtnl_lock-less FIB rule changes (RTM_NEWRULE and RTM_DELRULE)
      - prepare software drivers and TC qdiscs for rtnl_lock-less GET

   - Support BIG TCP (>64kB TSO) in UDP tunnels (vxlan, geneve)

   - Support buffers larger than PAGE_SIZE in devmem zero-copy API

   - Improve MPTCP handling of extreme memory pressure handling, when
     out-of-order queue had to be pruned

   - Report the per-group user count via RTM_GETMULTICAST

   - Expose the route deletion reason in RTM_DELROUTE

   - Add a SO_RIGHTS_NOTRUNC option to UNIX sockets to enable more
     useful handling of LSM denials when receiving SCM_RIGHTS messages:
     instead of truncating the message at the first blocked fd, keep
     every fd slot and store the LSM errno in the blocked slot

   - IPv6 Segment Routing - support looking up the post-encap SID
     (address) in a different/specified routing table

   - Support PRP RedBox (interlink) creation

   - Support per-nexthop UDP dst port in VXLAN

   - Continue converting getsockopt callbacks in a number of protocols
     to iov_iter

  Ethernet:

   - Merge initial CXL support for AMD/Solarflare NICs (shared branch
     with the CXL tree)

   - New drivers:
      - ADIN1140 10BASE-T1S MACPHY
      - Initial skeleton of Intel iXD and ZTE Dinghai drivers

   - High-speed NICs:
      - AMD/Pensando:
         - support firmware flashing
      - Cisco (enic):
         - SR-IOV V2 admin channel and MBOX protocol
      - Huawei (hns3):
         - support for ethtool pfc_prevention_tout
      - nVidia/Mellanox:
         - support sharing bandwidth control across interfaces
           of the same device
      - Marvell (octeontx2-pf):
         - link RQ page pools to netdev for Netlink stats
      - Google vNIC:
         - XDP metadata support for DQ RDA
      - Microsoft vNIC:
         - support forcing full-page RX buffers

   - Other NICs:
      - Synopsys IP:
         - eic7700: support for eth1
      - Microchip (lan743x):
         - support for RMII interface
      - Wangxun:
         - support for ethtool -G and -C for VFs
         - add Tx timeout and PCIe error handling
      - Intel (igb/igc):
         - RSS key get/set support
         - support for forcing link speed without auto-negotiation

   - Switches:
      - NXP (dpaa2):
         - support bonding/LAG offload
      - Mediatek:
         - mt7530: EN7528 support
         - initial support for MT7628
      - Micrel (ksz8/9):
         - refactoring work to move towards library model
         - PTP support for KSZ8463
      - nVidia/Mellanox:
         - support rtnl-lock-less ethtool callbacks
      - Realtek:
         - rtl8366rb: use generic RTL83xx code
         - support SGMII and HSGMII for RTL8367S

   - PHYs:
      - Airoha:
         - EcoNet EN7528 PHY support
      - DAPU Telecom
         - DAPU Telecom DAP8211R(I) Gigabit PHY support
      - Realtek:
         - support RTL8261C_CG
         - support RTL8261D

  Wireless:

   - nl80211: per-link statistics support for multi-link operation

   - mac80211: AQL/airtime-fairness support for multicast

   - Merge Peripheral Authentication Service (PAS) / TEE support for
     ath12k (shared branch with the firmware/qcom tree)

   - New drivers:
      - mm81x for Morse Micro Long-Range S1G devices
      - nxpwifi for NXP devices (mostly forked off from mwifiex)

   - Driver changes:
      - Broadcom (brcmfmac):
         - DPP support, some Cypress part update
      - MediaTek (mt76):
         - mt7928 support
         - mt7925 NAN support
         - mt7996 AP powersave improvements
      - Qualcomm (ath12k):
         - much kernel infrastructure integration work
         - AHB platform MultiPD support
      - Realtek (rt89):
         - LED support
         - RTL8922DE support
         - dual-BT coex for RTL8922D
      - Intel:
         - new FW version support

  Bluetooth:

   - HCI: add support for Shorter Connection Interval (SCI) feature

   - af_bluetooth: add minimal context analysis annotations

   - Driver changes:
      - Intel:
         - add Bluetooth SAR revision 2 support
         - add vendor_reset PCI sysfs for PLDR
      - Mediatek:
         - add USB IDs for MT7902 and MT7922 devices
      - Realtek:
         - add USB IDs for 8761CU and 8852BE devices
      - NXP:
         - add M.2 Bluetooth device support using pwrseq

  Misc:

   - DPLL support for manual/numerical oscillator control (NCO)
     (implement in zl3073x)

   - MCTP support for MCTP over USB v1.1 (DMTF DSP0283)

   - Power-over-Ethernet: support Realtek PSE controllers

   - Remove the IBM EHEA driver

   - Remove tulip/xircom_cb driver"

* tag 'net-next-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net-next: (1433 commits)
  net/mlx5e: do not HW-GRO coalesce small frames
  net: openvswitch: fix nf_connlabels leak in ovs_ct_init
  net: add missing ref_tracker_dir_exit() to alloc_netdev_mqs()
  net: openvswitch: fix flow mask use-after-free on flow deletion
  sctp: stop processing a packet once its association is deleted
  dpll: zl3073x: add PTP clock support
  dpll: zl3073x: add channel ToD, phase step and TIE operations
  dpll: zl3073x: scale poll interval proportionally to timeout
  ptp: vmclock: prevent read-only mappings from becoming writable
  ipv4: reject undersized MTUs in ip_do_fragment()
  bonding: initialize err for empty target lists
  net: dsa: initial support for MT7628 embedded switch
  net: dsa: initial MT7628 tagging driver
  net: phy: mediatek: add phy driver for MT7628 built-in Fast Ethernet PHYs
  dt-bindings: net: dsa: add MT7628 ESW
  net: pse-pd: realtek-pse-mcu: add UART transport
  net: pse-pd: realtek-pse-mcu: add I2C transport
  net: pse-pd: add Realtek PSE MCU core
  dt-bindings: net: pse-pd: add bindings for Realtek PSE MCU
  vsock: use sock_error() to consume sk_err after a failed connect
  ...
2026-08-20 08:16:04 -07:00
Linus Torvalds
5a8cd539ac Major changes:
- Redesign the verifier error reporting: failures now carry source and
   instruction annotations along with the causal event history that led
   to them, making program rejections far easier to debug and repair
   (Kumar Kartikeya Dwivedi)
 
 - Add arena argument support to kfuncs and struct_ops through the new
   __arena and __arena__nullable suffixes (Tejun Heo, Puranjay Mohan,
   Kumar Kartikeya Dwivedi, Ihor Solodrai)
 
 - Signed BPF program loader rework to accommodate both BPF and security
   community needs where the kernel runs the signature verification at
   BPF_PROG_LOAD time before the LSM admission hook (Daniel Borkmann)
 
 - Add a set of ksock kfuncs which let BPF LSM and syscall programs
   create, connect and send on UDP sockets in order to emit telemetry
   data (Mahe Tardy)
 
 - Unify helper and kfunc call argument verification and classify kfunc
   arguments purely from BTF into a generated bpf_func_proto which is
   computed once at add-call time (Amery Hung)
 
 Other features and fixes:
 
 - Enable EXECMEM_ROX_CACHE for BPF allocations on x86 (Mike Rapoport)
 
 - Add bidirectional VLAN support to bpf_fib_lookup() through the new
   BPF_FIB_LOOKUP_VLAN and BPF_FIB_LOOKUP_VLAN_INPUT flags
   (Avinash Duduskar)
 
 - Infer zext_dst from static register liveness analysis to fix 32-bit
   zero-extension semantics, and remove the artificial limitations on
   pointer types eligible for spilling (Eduard Zingerman)
 
 - Inline the numeric open-coded iterator kfuncs so that bpf_for() loops
   no longer pay a kfunc call on every iteration (Puranjay Mohan)
 
 - Add an arena-based bitmap data structure to libarena along with
   serial and parallel selftests (Emil Tsalapatis)
 
 - Teach resolve_btfids to discover kfuncs from the kernel's BTF ID sets
   and to emit kfunc BTF decl tags, reducing the kernel build's
   dependency on pahole features (Ihor Solodrai)
 
 - Add BPF_F_ADJ_ROOM_DECAP_* flags to bpf_skb_adjust_room() so that
   tunnel decapsulation can update the GSO and encapsulation state of
   the skb (Nick Hudson)
 
 - Fix the ring buffer pending_pos walk and the available-data
   accounting on 32-bit position wrap (Israel Téllez García)
 
 - Add memory usage accounting for arena maps and fix an mmap_lock
   deadlock on arena lock failure (Jiayuan Chen)
 
 - Add tracing_multi link info support to the kernel UAPI and bpftool,
   and refactor the stack map code to run with preemption disabled
   (Jiri Olsa)
 
 - Support BPF_F_EGRESS in bpf_redirect_peer() to emit the skb in the
   egress direction of the target's peer device (Jordan Rife)
 
 - Add a KF_SPINLOCK_SAFE kfunc flag so that providers, in particular
   modules, can declare kfuncs safe to call under bpf_spin_lock instead
   of relying on the verifier's hard-coded allowlist (Kaitao Cheng)
 
 - Introduce global percpu data for BPF programs with libbpf probing
   and bpftool skeleton support, and stop exposing uninitialized kernel
   heap memory when copying per-CPU map values (Leon Hwang)
 
 - Add s390 JIT support for load-acquire and store-release instructions
   (Maxim Khmelevskii)
 
 - Fix a CFI mismatch in the task work callback and an arm64 KASAN
   false positive after bpf_throw() (Mykyta Yatsenko)
 
 - Reject writes through untrusted BTF pointers and bound the
   rdonly/rdwr_buf_size kfunc arguments (Nicholas Dudar)
 
 - Invalidate RCU pointers only after the final spin unlock and account
   for preempt and IRQ disabled regions as overlapping RCU protection
   (Ning Ding)
 
 - Support mixing bpf2bpf calls and tail calls on RV64, add signed
   operations and 32-bit atomics to the RV32 JIT, and add timed may_goto
   support (Pu Lehui, Kuan-Wei Chiu, Feng Jiang)
 
 - Fix a use-after-free on mm_struct in bpf_find_vma() for foreign tasks
   and an mmap_lock leak in the irq_work path (Sanghyun Park)
 
 - Populate mmap-able BPF array map memory lazily which makes mmap() O(1)
   instead of proportional to the map size (Song Liu)
 
 - Introduce a jit_required flag and reject programs with inlined
   helpers when no JIT is available, where the interpreter would
   otherwise jump into an invalid address (Tiezhu Yang)
 
 - Fix the x86 JIT per-CPU address resolution into an extended register
   where the REX prefix dropped the high destination register bit
   (Vineet Gupta)
 
 - Reject MEM_ALLOC BTF accesses past object bounds, arena frees below
   the arena base, and mixed arena and ordinary atomic paths
   (Yiyang Chen)
 
 - Fix the trampoline handling of 128-bit arguments and of return values
   larger than 8 bytes (Yonghong Song)
 
 - Ensure that any fault prone load is rewritten with exception table
   handling, and fix the arena load-acquire and atomic fetch handling
   in the x86, arm64, riscv and s390 JITs (Daniel Borkmann)
 
 - Many more fixes and cleanups across the verifier, arena, trampolines,
   sockmap, cgroup, ring buffer, x86/arm64/riscv/s390 JITs, libbpf,
   bpftool, resolve_btfids and selftests.
 
 Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
 -----BEGIN PGP SIGNATURE-----
 
 iIsEABYKADMWIQTFp0I1jqZrAX+hPRXbK58LschIgwUCaoNzBBUcZGFuaWVsQGlv
 Z2VhcmJveC5uZXQACgkQ2yufC7HISIOb3QEAy5cyrLXY+VWofhsC9wULkHyETOdj
 oTkdohQomZp4VhEA/1RZXdHVS1ANFgreWv0fMorUOHEKv2ZuNokfk3LWgW4L
 =VRyL
 -----END PGP SIGNATURE-----

Merge tag 'bpf-next-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf-next

Pull bpf updates from Daniel Borkmann:
 "Major changes:

   - Redesign the verifier error reporting: failures now carry source
     and instruction annotations along with the causal event history
     that led to them, making program rejections far easier to debug and
     repair (Kumar Kartikeya Dwivedi)

   - Add arena argument support to kfuncs and struct_ops through the new
     __arena and __arena__nullable suffixes (Tejun Heo, Puranjay Mohan,
     Kumar Kartikeya Dwivedi, Ihor Solodrai)

   - Signed BPF program loader rework to accommodate both BPF and
     security community needs where the kernel runs the signature
     verification at BPF_PROG_LOAD time before the LSM admission hook
     (Daniel Borkmann)

   - Add a set of ksock kfuncs which let BPF LSM and syscall programs
     create, connect and send on UDP sockets in order to emit telemetry
     data (Mahe Tardy)

   - Unify helper and kfunc call argument verification and classify
     kfunc arguments purely from BTF into a generated bpf_func_proto
     which is computed once at add-call time (Amery Hung)

  Other features and fixes:

   - Enable EXECMEM_ROX_CACHE for BPF allocations on x86 (Mike Rapoport)

   - Add bidirectional VLAN support to bpf_fib_lookup() through the new
     BPF_FIB_LOOKUP_VLAN and BPF_FIB_LOOKUP_VLAN_INPUT flags (Avinash
     Duduskar)

   - Infer zext_dst from static register liveness analysis to fix 32-bit
     zero-extension semantics, and remove the artificial limitations on
     pointer types eligible for spilling (Eduard Zingerman)

   - Inline the numeric open-coded iterator kfuncs so that bpf_for()
     loops no longer pay a kfunc call on every iteration (Puranjay
     Mohan)

   - Add an arena-based bitmap data structure to libarena along with
     serial and parallel selftests (Emil Tsalapatis)

   - Teach resolve_btfids to discover kfuncs from the kernel's BTF ID
     sets and to emit kfunc BTF decl tags, reducing the kernel build's
     dependency on pahole features (Ihor Solodrai)

   - Add BPF_F_ADJ_ROOM_DECAP_* flags to bpf_skb_adjust_room() so that
     tunnel decapsulation can update the GSO and encapsulation state of
     the skb (Nick Hudson)

   - Fix the ring buffer pending_pos walk and the available-data
     accounting on 32-bit position wrap (Israel Téllez García)

   - Add memory usage accounting for arena maps and fix an mmap_lock
     deadlock on arena lock failure (Jiayuan Chen)

   - Add tracing_multi link info support to the kernel UAPI and bpftool,
     and refactor the stack map code to run with preemption disabled
     (Jiri Olsa)

   - Support BPF_F_EGRESS in bpf_redirect_peer() to emit the skb in the
     egress direction of the target's peer device (Jordan Rife)

   - Add a KF_SPINLOCK_SAFE kfunc flag so that providers, in particular
     modules, can declare kfuncs safe to call under bpf_spin_lock
     instead of relying on the verifier's hard-coded allowlist (Kaitao
     Cheng)

   - Introduce global percpu data for BPF programs with libbpf probing
     and bpftool skeleton support, and stop exposing uninitialized
     kernel heap memory when copying per-CPU map values (Leon Hwang)

   - Add s390 JIT support for load-acquire and store-release
     instructions (Maxim Khmelevskii)

   - Fix a CFI mismatch in the task work callback and an arm64 KASAN
     false positive after bpf_throw() (Mykyta Yatsenko)

   - Reject writes through untrusted BTF pointers and bound the
     rdonly/rdwr_buf_size kfunc arguments (Nicholas Dudar)

   - Invalidate RCU pointers only after the final spin unlock and
     account for preempt and IRQ disabled regions as overlapping RCU
     protection (Ning Ding)

   - Support mixing bpf2bpf calls and tail calls on RV64, add signed
     operations and 32-bit atomics to the RV32 JIT, and add timed
     may_goto support (Pu Lehui, Kuan-Wei Chiu, Feng Jiang)

   - Fix a use-after-free on mm_struct in bpf_find_vma() for foreign
     tasks and an mmap_lock leak in the irq_work path (Sanghyun Park)

   - Populate mmap-able BPF array map memory lazily which makes mmap()
     O(1) instead of proportional to the map size (Song Liu)

   - Introduce a jit_required flag and reject programs with inlined
     helpers when no JIT is available, where the interpreter would
     otherwise jump into an invalid address (Tiezhu Yang)

   - Fix the x86 JIT per-CPU address resolution into an extended
     register where the REX prefix dropped the high destination register
     bit (Vineet Gupta)

   - Reject MEM_ALLOC BTF accesses past object bounds, arena frees below
     the arena base, and mixed arena and ordinary atomic paths (Yiyang
     Chen)

   - Fix the trampoline handling of 128-bit arguments and of return
     values larger than 8 bytes (Yonghong Song)

   - Ensure that any fault prone load is rewritten with exception table
     handling, and fix the arena load-acquire and atomic fetch handling
     in the x86, arm64, riscv and s390 JITs (Daniel Borkmann)

   - Many more fixes and cleanups across the verifier, arena,
     trampolines, sockmap, cgroup, ring buffer, x86/arm64/riscv/s390
     JITs, libbpf, bpftool, resolve_btfids and selftests"

* tag 'bpf-next-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf-next: (373 commits)
  selftests/bpf: Add tests for a store on a fault prone qdisc pointer
  selftests/bpf: Add tests for fault prone loads out of RCU pointers
  selftests/bpf: Add tests for pointer type merge at a shared load
  selftests/bpf: Remove duplicate copies of the arena spinlock qnodes
  selftests/bpf: Retry stat generation in cgroup_iter_memcg
  selftests/bpf: Test pseudo-function policy diagnostics
  bpf: Distinguish function references in policy diagnostics
  bpf: Preserve source attribution without source text
  selftests/bpf: Test kfunc argument diagnostics
  bpf: Correct kfunc argument diagnostics
  bpf: Use canonical stack argument names in diagnostics
  bpf: Preserve R0 lineage across helper calls
  selftests/bpf: Exercise negative optlen in cgroup getsockopt hook
  bpf: Reject negative optlen in cgroup getsockopt hook
  selftests/bpf: tc_tunnel - validate decap GSO and encapsulation state
  bpf: Clear decap state on skb_adjust_room shrink path
  bpf: Allow new DECAP flags and add guard rails
  bpf: Add BPF_F_ADJ_ROOM_DECAP_* flags for tunnel decapsulation
  bpf: Refactor masks for ADJ_ROOM flags and encap validation
  bpf: Name the enum for BPF_FUNC_skb_adjust_room flags
  ...
2026-08-20 07:36:20 -07:00
Jakub Kicinski
61eb236c41 Merge git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net
Merge in late fixes in preparation for the net-next PR.

Conflicts:

drivers/dpll/dpll_core.c
drivers/dpll/dpll_netlink.c
  33f016b23a ("dpll: fix NULL deref in dpll_device_ops() during teardown race")
  b1d0c41208 ("dpll: add STATE_CONNECTED_OVERRIDE pin capability")
https://lore.kernel.org/aoR9YYY2P5--3x0N@sirena.org.uk
https://lore.kernel.org/aoR9VmKllVGwmQn_@sirena.org.uk

No adjacent changes.

Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-18 10:42:41 -07:00
Tetsuo Handa
0b1c2af8a2 net: add missing ref_tracker_dir_exit() to alloc_netdev_mqs()
sashiko is reporting that trying to read /sys/kernel/debug/ref_tracker/*
causes use-afer-free crash when either alloc_percpu() or dev_addr_init()
in alloc_netdev_mqs() failed, for commit 4d92b95ff2 ("net: add net device
refcount tracker infrastructure") added ref_tracker_dir_exit() to only
free_netdev() path.

Closes: https://sashiko.dev/#/patchset/56c707e7-1fb0-43ec-b8fb-cf6f451e513e%40I-love.SAKURA.ne.jp
Fixes: 4d92b95ff2 ("net: add net device refcount tracker infrastructure")
Signed-off-by: Tetsuo Handa <penguin-kernel@I-love.SAKURA.ne.jp>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/b06ce35d-e7bc-47a5-8e0a-e82be7e4dd08@I-love.SAKURA.ne.jp
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-18 10:02:33 -07:00
Jori Koolstra
fd8756fa14 net: af_unix: useful handling of LSM denials on SCM_RIGHTS
Right now if some LSM such as Smack denies an AF_UNIX socket peer to
receive an SCM_RIGHTS fd, the SCM_RIGHTS fd array will be cut short at
that point, and MSG_CTRUNC is set on return of recvmsg(). This is
highly problematic behaviour, because it leaves the receiver
wondering what happened. As per man page MSG_CTRUNC is supposed to
indicate that the control buffer was sized too short, but suddenly
a permission error might result in the exact same flag being set.
Moreover, the receiver has no chance to determine how many fds got
originally sent and how many were suppressed.[1]

Add a SO_RIGHTS_NOTRUNC option to UNIX sockets to enable more useful
handling of LSM denials when receiving SCM_RIGHTS messages: instead of
truncating the message at the first blocked fd, keep every fd slot
and store the LSM errno in the blocked slot. The socket option is
inherited by the child accept() socket if set on the listen() socket.

[1]: https://github.com/uapi-group/kernel-features#useful-handling-of-lsm-denials-on-scm_rights

Reviewed-by: Christian Brauner (Amutable) <brauner@kernel.org>
Signed-off-by: Jori Koolstra <jkoolstra@xs4all.nl>
Link: https://patch.msgid.link/20260813162818.149248-4-jkoolstra@xs4all.nl
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-17 18:14:52 -07:00
Jori Koolstra
48b84acc5e net: scm: move scm_detach_fds() from common path to scm_recv_unix()
scm->fp can only be set when using UNIX sockets, therefore we should
move it out of the common path __scm_recv_common() into
scm_recv_unix().

Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Signed-off-by: Jori Koolstra <jkoolstra@xs4all.nl>
Link: https://patch.msgid.link/20260813162818.149248-3-jkoolstra@xs4all.nl
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-17 18:14:52 -07:00
Qi Zhang
f98ca137c2 net: pktgen: use a consistent flow count
pktgen_if_write() can update cflows while the packet generator thread is
inside mod_cur_headers(). The latter first tests cflows, but f_pick() then
reloads it when selecting a random flow.

This allows the following interleaving:

  CPU 0 (kpktgend)                 CPU 1 (proc write)
  if (pkt_dev->cflows) // 10
                                   pkt_dev->cflows = 0
  get_random_u32_below(pkt_dev->cflows)

get_random_u32_below(0) returns a full-width random value. Using that
value as an index into the fixed-size flows array causes an out-of-bounds
access. The kernel reported:

  BUG: unable to handle page fault for address: ffffc8fe2d2674bc
  #PF: supervisor read access in kernel mode
  Oops: Oops: 0000 [#1] SMP KASAN NOPTI
  CPU: 0 UID: 0 PID: 65 Comm: kpktgend_0
  RIP: 0010:mod_cur_headers+0x16f8/0x2840
  Call Trace:
   <TASK>
   pktgen_thread_worker+0x305a/0x6bc0
   kthread+0x2c6/0x3b0
   ret_from_fork+0x36e/0x5a0
   ret_from_fork_asm+0x1a/0x30
   </TASK>

Read cflows once at the start of mod_cur_headers(), pass the snapshot to
f_pick(), and use it for later flow-state decisions in the same packet.
Publish proc updates with WRITE_ONCE(). Flow selection then always uses a
nonzero count bounded by MAX_CFLOWS, while a concurrent update takes
effect on a later packet.

Cc: stable+noautosel@kernel.org # needs real net-admin (non-ns)
Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com>
Signed-off-by: Qi Zhang <marsy12010123@gmail.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-17 12:50:51 -07:00
Eric Dumazet
51b0aaafd9 net: add READ_ONCE()/WRITE_ONCE() annotations for dev->prio_tc_map
Concurrent fast-path readers access dev->prio_tc_map (e.g. via
skb_tx_hash(), netdev_get_prio_tc_map(), and qdiscs) while writers
update entries in dev->prio_tc_map or reset/clear the map via
netdev_reset_tc() and netdev_unbind_sb_channel().

Furthermore, memset() in netdev_reset_tc() and
netdev_unbind_sb_channel() provides no guarantee of performing
atomic word/byte stores.

Add READ_ONCE() and WRITE_ONCE() annotations to netdev_get_prio_tc_map()
and netdev_set_prio_tc_map(), replace memset() in dev.c with explicit
WRITE_ONCE() loops, and update direct array accesses in qdiscs to use
netdev_get_prio_tc_map().

Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260812085440.3917924-4-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-17 10:27:48 -07:00
Eric Dumazet
0c6c32a8c8 net: add READ_ONCE()/WRITE_ONCE() annotations for dev->num_tc
Several fast-path and control-path lockless readers access dev->num_tc
(e.g., skb_tx_hash(), netdev_txq_to_tc(), netdev_get_num_tc(), and
qdisc/driver lookups) while concurrent writers update dev->num_tc
during TC setup, device reset, or channel configuration.

Add READ_ONCE() and WRITE_ONCE() annotations to prevent compiler
reordering and load/store tearing when accessing dev->num_tc.

Update inline helpers in netdevice.h (netdev_get_num_tc(),
netdev_set_prio_tc_map(), and netdev_get_sb_channel()) as well as
writers and lockless readers in core networking code and drivers.

Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260812085440.3917924-3-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-17 10:27:48 -07:00
Eric Dumazet
21ef2d065a net: prevent torn reads in netdev_tc_txq
netdev_set_tc_queue() (and related helpers/drivers such as
netdev_bind_sb_channel_queue(), netdev_reset_tc(), and
netdev_unbind_sb_channel()) perform separate 16-bit writes to
dev->tc_to_txq[tc].count and dev->tc_to_txq[tc].offset.

Furthermore, memset() in netdev_reset_tc() and
netdev_unbind_sb_channel() provides no guarantee of performing
full 32-bit word stores.

Concurrent lockless readers (e.g. skb_tx_hash(), netdev_txq_to_tc(),
ixgbe_select_queue(), taprio, mqprio, FPE drivers) can observe torn
values where offset and count belong to inconsistent configurations.

Redefine struct netdev_tc_txq to embed count and offset inside a union
with a u32 combined field, allowing atomic manipulation via
READ_ONCE() and WRITE_ONCE().

Update all lockless readers and writers across the kernel to use
READ_ONCE() and WRITE_ONCE() on the combined field.

Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260812085440.3917924-2-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-17 10:27:48 -07:00
Jiayuan Chen
ad27ed7d23 bpf, xdp: move offload check into dev_xdp_install()
bpf_xdp_link_update() calls dev_xdp_install() directly and skips
dev_xdp_attach(), so the checks in dev_xdp_attach() do not run. A user can
make an XDP link with a normal program and then swap in an offloaded or
device-bound program with BPF_LINK_UPDATE, which puts it on the software
path.

dev_xdp_install() is the one place all three paths go through:
"ip link set xdp" and BPF_LINK_CREATE reach it via dev_xdp_attach(), and
BPF_LINK_UPDATE calls it directly. So move the program checks (offloaded,
bound to another device, device-bound in generic mode, native vs generic,
DEVMAP and CPUMAP) there, and keep only the netlink-flag check
(XDP_FLAGS_UPDATE_IF_NOEXIST) in dev_xdp_attach().

Fixes: 026a4c28e1 ("bpf, xdp: Implement LINK_UPDATE for BPF XDP link")
Signed-off-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-17 10:04:49 -07:00
Nick Hudson
ec20dee2f2 bpf: Clear decap state on skb_adjust_room shrink path
On shrink in bpf_skb_adjust_room(), apply decapsulation state updates
according to BPF_F_ADJ_ROOM_DECAP_* flags.

For GSO skbs, clear only the tunnel gso_type bits that correspond to
the requested decap layer:

- DECAP_L4_UDP: SKB_GSO_UDP_TUNNEL{,_CSUM}
- DECAP_L4_GRE: SKB_GSO_GRE{,_CSUM}
- DECAP_IPXIP4: SKB_GSO_IPXIP4
- DECAP_IPXIP6: SKB_GSO_IPXIP6

Then clear skb->encapsulation only if no tunnel GSO bits remain, keeping
encapsulation set for cases such as ESP-in-UDP where tunnel state remains.

For non-GSO skbs, there are no tunnel GSO bits to consult, so clear
skb->encapsulation directly when DECAP_L4_* or DECAP_IPXIP_* flags are set.

This keeps decap state handling consistent between GSO and non-GSO packets.

Co-developed-by: Max Tottenham <mtottenh@akamai.com>
Co-developed-by: Anna Glasgall <aglasgal@akamai.com>
Signed-off-by: Max Tottenham <mtottenh@akamai.com>
Signed-off-by: Anna Glasgall <aglasgal@akamai.com>
Signed-off-by: Nick Hudson <nhudson@akamai.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://lore.kernel.org/bpf/20260812083115.73100-6-nhudson@akamai.com
2026-08-17 11:30:13 +02:00
Nick Hudson
3a39c214fd bpf: Allow new DECAP flags and add guard rails
Add checks to require shrink-only decap, reject conflicting decap flag
combinations, and verify removed length is sufficient for claimed header
decapsulation.

Co-developed-by: Max Tottenham <mtottenh@akamai.com>
Co-developed-by: Anna Glasgall <aglasgal@akamai.com>
Signed-off-by: Max Tottenham <mtottenh@akamai.com>
Signed-off-by: Anna Glasgall <aglasgal@akamai.com>
Signed-off-by: Nick Hudson <nhudson@akamai.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://lore.kernel.org/bpf/20260812083115.73100-5-nhudson@akamai.com
2026-08-17 11:30:10 +02:00
Nick Hudson
7b2ea1151e bpf: Refactor masks for ADJ_ROOM flags and encap validation
Refactor the helper masks for bpf_skb_adjust_room() flags to simplify
validation logic and introduce:

- BPF_F_ADJ_ROOM_ENCAP_MASK
- BPF_F_ADJ_ROOM_DECAP_MASK

Refactor existing validation checks in bpf_skb_net_shrink() and
bpf_skb_adjust_room() to use the new masks (no behavior change).

This is in preparation for supporting the new decap flags.

Co-developed-by: Max Tottenham <mtottenh@akamai.com>
Co-developed-by: Anna Glasgall <aglasgal@akamai.com>
Signed-off-by: Max Tottenham <mtottenh@akamai.com>
Signed-off-by: Anna Glasgall <aglasgal@akamai.com>
Signed-off-by: Nick Hudson <nhudson@akamai.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://lore.kernel.org/bpf/20260812083115.73100-3-nhudson@akamai.com
2026-08-17 11:29:51 +02:00
Junseo Lim
84473a7e18 bpf: Disallow bpf_{g,s}etsockopt() in cgroup UNIX getname hooks
_bpf_setsockopt() and _bpf_getsockopt() call sock_owned_by_me() for
full sockets, so these helpers expect the socket lock to be held.

BPF_CGROUP_UNIX_GETPEERNAME and BPF_CGROUP_UNIX_GETSOCKNAME run BPF
programs without acquiring the socket lock. A program attached to
either hook can therefore trigger the sock_owned_by_me() warning by
calling bpf_setsockopt() or bpf_getsockopt().

Disallow bpf_setsockopt() and bpf_getsockopt() for CGROUP_UNIX_GETPEERNAME
and CGROUP_UNIX_GETSOCKNAME.

Fixes: 859051dd16 ("bpf: Implement cgroup sockaddr hooks for unix sockets")
Reported-by: Sechang Lim <rhkrqnwk98@gmail.com>
Signed-off-by: Junseo Lim <zirajs7@gmail.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Link: https://lore.kernel.org/bpf/20260812091654.244752-1-zirajs7@gmail.com
2026-08-17 11:10:09 +02:00
Junseo Lim
5fe7007aed lwt_bpf: Restore reserved headroom after xmit program
ip_finish_output2() expands an skb to LL_RESERVED_SPACE(dev) before LWT
xmit. An LWT_XMIT BPF program can then modify the skb head and still
return BPF_OK, so bpf_xmit() rechecks the remaining headroom before the
skb continues to neighbour output.

That recheck uses dst->dev->hard_header_len. This is not enough for the
neighbour cached-header path: neigh_hh_output() copies the cached hardware
header using the aligned hh_cache size, HH_DATA_MOD for short headers or
HH_DATA_ALIGN(hh_len) otherwise.

On Ethernet, hard_header_len is 14 but the cached copy needs 16 bytes. If
an LWT_XMIT BPF program calls bpf_skb_change_head(skb, 1, 0), the skb can
still have 15 bytes of headroom after the program. The existing check
accepts that, after which neigh_hh_output() hits its headroom warning and
drops the skb.

Use LL_RESERVED_SPACE(dst->dev) in the post-BPF headroom check to match
the reservation made before LWT xmit.

Fixes: 3a0af8fd61 ("bpf: BPF for lightweight tunnel infrastructure")
Reported-by: Sechang Lim <rhkrqnwk98@gmail.com>
Suggested-by: Daniel Borkmann <daniel@iogearbox.net>
Signed-off-by: Junseo Lim <zirajs7@gmail.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Link: https://lore.kernel.org/bpf/20260811044149.118235-1-zirajs7@gmail.com
2026-08-17 10:59:03 +02:00
Michal Luczaj
34e0eb763b bpf, sockmap: Use sock_hold() instead of refcount_inc_not_zero() in lookup
psock's hold on the looked up socket isn't dropped until sk_psock_drop() ->
queue_rcu_work() -> sk_psock_destroy() runs, which happens only after the
entry is unlinked and an RCU grace period elapses. Since the lookup runs
under RCU, a non-NULL result guarantees sk_refcnt >= 1:
refcount_inc_not_zero() can never fail here. Use sock_hold() instead.

Signed-off-by: Michal Luczaj <mhal@rbox.co>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
Reviewed-by: Jakub Sitnicki <jakub@cloudflare.com>
Link: https://lore.kernel.org/bpf/20260813-sockmap-lookup-get-ref-v1-2-31f5d55f44ac@rbox.co
2026-08-17 10:22:19 +02:00