Commit Graph

1483822 Commits

Author SHA1 Message Date
Linus Torvalds
e8dfd03a1c cgroup: Fixes for v7.3-rc4
- With local event accounting, a fork rejected by the pids controller
   updated pids.events without notifying its pollers.
 
 - A cgroup selftest failed to compile with fortification enabled because
   an O_TMPFILE open lacked its mode argument.
 -----BEGIN PGP SIGNATURE-----
 
 iIQEABYKACwWIQTfIjM1kS57o3GsC/uxYfJx3gVYGQUCarWSKg4cdGpAa2VybmVs
 Lm9yZwAKCRCxYfJx3gVYGYoiAQD6JUCqDjv2Hr1YeMFlsoYVQZ7tNmljVnPD2tZ6
 9WmD/AEAotdlmzM8egOuDqAi2s+UMJPm8vCZuxVzewiGqhYELgk=
 =0Ljp
 -----END PGP SIGNATURE-----

Merge tag 'cgroup-for-7.3-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup

Pull cgroup fixes from Tejun Heo:

 - With local event accounting, a fork rejected by the pids controller
   updated pids.events without notifying its pollers

 - A cgroup selftest failed to compile with fortification enabled
   because an O_TMPFILE open lacked its mode argument

* tag 'cgroup-for-7.3-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup:
  cgroup/pids: Restore pids.events notifications in local mode
  selftests: cgroup: give the O_TMPFILE open in get_temp_fd() a mode
2026-09-24 14:26:12 -07:00
Linus Torvalds
f2c53ea949 Including fixes from Bluetooth, NFC and Netfilter.
Every week in this release is record-setting for number of posted
 patches. It doesn't seem like we're creating any regressions with
 all these fixes, 3 Fixes tags here point to 7.2 commits but none
 are true regression fixes. We're trying to keep the count down,
 nonetheless.
 
 Previous releases - regressions:
 
  - net: don't require the hwtstamp NDOs when a PHY provides timestamping
 
  - ipv6: fix dst leak for uncached routes
 
  - vrf: stop corrupting skb->csum when capturing CHECKSUM_COMPLETE
    packets
 
 Previous releases - always broken:
 
  - packet: use ubuf_info completion for TX_RING packets
 
  - arp: terminate device name before lookup
 
  - ipv6: do not let ipv6_find_hdr() return an offset past the packet end
 
  - udp: remove a disconnected socket from the 4-tuple hash table
 
  - sctp: discard the rest of the packet on a stale-cookie error
 
  - eth: mlx5: Bridge, fix remaining switchdev ownership gaps on merged
    eswitch
 
 Signed-off-by: Jakub Kicinski <kuba@kernel.org>
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEE6jPA+I1ugmIBA4hXMUZtbf5SIrsFAmq1bCQACgkQMUZtbf5S
 IrsgCBAAiHOxV3pR0pNw+K/dO5maYhuUV4w+Xcd6CbjRnxzYtImVC1q2ieez5dEe
 JrQrVK76t2wuEgGSyiWcckIufPL8xYg5nYaOUbkYMC2QRtDaqhCo+u2jXu26/Lr8
 O3sI2cj86XDZN+xKTvmLld1zIRyD6MhYxWz+gIicArdBqTsr5JsWkwhJE7SH013a
 iVGu0sIxTgCNu/AusD6SdtwFte26dqvadelI8yuDhBR6rM9pi4EWXChleg7ohCsu
 DoZTlaEjEiFavat+6ni99oz8MAT0WVBX0IzTqZa6TCnbCw4qUEbqN8G+CNgC/VPe
 Sq5FRoo6vJjlFviwgWPbpje2FoOYtk6CX1eTavx8gZ4pPMTzGkT62crS+ETwLO5H
 mRbE+ZjG8UDTUKpIY5P/IQYdznHSc9Ny/pHNlZhFmDdntJOp1lnegx5CmZkyuZxC
 Qc0hoFxKdjUrpQ1n0ygT3/P6MT9vXMwbGP52VLIxaB1oaYCKFV71NOp0ET5fCnYb
 CBK8cdN0kcaSLrpi3MnPamyQoNZfH78BiuKsLDu/TWg3i2HGmFTlp0/LJ0nkQcCU
 4KlCmiJNOkHjbQKHNtYMUOZQcZu5witEv33kb7gr581b0RR+n0lZTaXcnRIyXYdb
 ORqIa1Khg5mW3DLaIc9IrzEGGAW5zXUHc5OX7t8J2rVRQWhBEh4=
 =YiiK
 -----END PGP SIGNATURE-----

Merge tag 'net-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net

Pull networking fixes from Jakub Kicinski:
 "Including fixes from Bluetooth, NFC and Netfilter.

  Every week in this release is record-setting for number of posted
  patches. It doesn't seem like we're creating any regressions with all
  these fixes, three 'Fixes' tags here point to 7.2 commits but none are
  true regression fixes. We're trying to keep the count down,
  nonetheless.

  Previous releases - regressions:

   - net: don't require the hwtstamp NDOs when a PHY provides
     timestamping

   - ipv6: fix dst leak for uncached routes

   - vrf: stop corrupting skb->csum when capturing CHECKSUM_COMPLETE
     packets

  Previous releases - always broken:

   - packet: use ubuf_info completion for TX_RING packets

   - arp: terminate device name before lookup

   - ipv6: do not let ipv6_find_hdr() return an offset past the packet
     end

   - udp: remove a disconnected socket from the 4-tuple hash table

   - sctp: discard the rest of the packet on a stale-cookie error

   - eth: mlx5: Bridge, fix remaining switchdev ownership gaps on merged
     eswitch"

[ And lots of other random network driver fixes ]

* tag 'net-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (189 commits)
  tcp: prevent collapsing skbs across boundary in rtx queue
  vlan: ensure sufficient headroom in vlan_dev_hard_header()
  net/sched: sch_teql: fix shadowed err in __teql_resolve()
  bridge: check llc_mac_hdr_init() return value in br_send_bpdu()
  llc: fix skb UAF and leaks on llc_mac_hdr_init() failure
  llc: reserve device headroom for allocated frames
  gve: DQO: reject TSO packets with an out of range MSS
  gve: fix TX drop when GSO MSS is too small for hw
  gve: DQO: fix header length used by gve_can_send_tso() for UDP GSO
  net: flush skb_defer_nodes in dev_cpu_dead()
  net: ethernet: stmmac: dwmac-rk: fix bulk clock leak when the PHY clock fails
  af_packet: fix integer overflow in prb_calc_retire_blk_tmo()
  tipc: Fix a data race on mon->peer_cnt in mon_timeout()
  net: phy: intel-xway: workaround 100BASE-TX Link-Up issue
  net/smc: fix UAF on lgr list traversal in smcr_port_err()
  net/rds: size a connection's path set by the transport it ends up with
  nfp: hold IPsec RX state under the XArray lock
  net: ena: fix MMIO read buffer leak on probe failure
  net: ena: fix PHC cleanup on probe failure
  net/sched: act_ct: fix helper UAF due to extensions realloc
  ...
2026-09-24 11:45:33 -07:00
Willem de Bruijn
fc6d80eb50 tcp: prevent collapsing skbs across boundary in rtx queue
tcp_write_collapse_fence() sets TCP_SKB_CB(skb)->eor = 1 on
tcp_write_queue_tail(sk) to prevent skbs queued after a switch to
device encryption from being collapsed into earlier skbs.

The fence is a no-op if all earlier data has already been transmitted
when the switch happens: sk->sk_write_queue is empty. The not yet
acknowledged earlier skbs wait in sk->tcp_rtx_queue with eor 0.

On a subsequent retransmit or SACK shift, tcp_retrans_try_collapse() or
tcp_shift_skb_data() can then merge an skb queued after the switch into
one queued before it.

Both users of the fence are affected:

- psp: devices only encrypt skbs with skb->decrypted set. The merged skb
  keeps decrypted = 0 from the earlier skb, so merged data sent after
  psp_sock_assoc_set_tx() is retransmitted in cleartext.

- tls device offload: the merged skb straddles the start marker set in
  tls_set_device_offload(). The software fallback (fill_sg_in() returns
  -EINVAL) and the mlx5, nfp and funeth drivers cannot handle such an
  skb and drop it. Every retransmit rebuilds the same skb, so the
  connection stalls.

Fix this in two places, for defense in depth:

1. Fall back to tcp_rtx_queue_tail(sk) in tcp_write_collapse_fence()
   when tcp_write_queue_tail(sk) is NULL.

2. Check !skb_cmp_decrypted(to, from) in tcp_skb_can_collapse(), as
   tcp_skb_can_collapse_rx() does on receive. skb_shift(), which both
   collapse paths call, already has a DEBUG_NET_WARN_ON_ONCE() for this
   condition.

Fixes: e8f6979981 ("net/tls: Add generic NIC offload infrastructure")
Cc: stable@vger.kernel.org
Signed-off-by: Willem de Bruijn <willemb@google.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Daniel Zahka <daniel.zahka@gmail.com>
Link: https://patch.msgid.link/20260924154427.953800-1-willemdebruijn.kernel@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 11:03:42 -07:00
Jakub Kicinski
1078a38344 Merge branch 'vlan-ensure-sufficient-headroom-in-vlan_dev_hard_header'
Eric Dumazet says:

====================
vlan: ensure sufficient headroom in vlan_dev_hard_header()

Callers that only reserve ETH_HLEN or less, or skbs allocated before
dynamic device/headroom changes (such as toggling VLAN_FLAG_REORDER_HDR
or bonding/team switching slaves), can reach vlan_dev_hard_header() with
insufficient headroom and trigger skb_under_panic().

When vlan_dev_hard_header() returns -ENOMEM upon skb_cow_head() failure,
a few callers of dev_hard_header() / llc_mac_hdr_init() had pre-existing
error-handling bugs:
====================

Link: https://patch.msgid.link/20260924082951.1599377-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 11:01:04 -07:00
Eric Dumazet
cd5dd68267 vlan: ensure sufficient headroom in vlan_dev_hard_header()
Callers that only reserve ETH_HLEN or less (such as llc_alloc_frame()),
or skbs allocated before dynamic device/headroom changes (e.g. toggling
VLAN_FLAG_REORDER_HDR or bonding/team switching slaves), can reach
vlan_dev_hard_header() with insufficient headroom and trigger
skb_under_panic().

Use skb_cow_head() in vlan_dev_hard_header() when VLAN_FLAG_REORDER_HDR
is not set to ensure sufficient headroom for the VLAN header(s) and the
underlying device hard header.

Use READ_ONCE() to read dev->hard_header_len and dev->needed_headroom as
they can be updated concurrently under RTNL (e.g. in
vlan_transfer_features()) while vlan_dev_hard_header() runs locklessly on
the transmit path. Also avoid LL_RESERVED_SPACE(dev) here so that the
extra HH_DATA_MOD alignment padding does not trigger unnecessary
pskb_expand_head() reallocations on inner stacked VLAN devices after the
outer VLAN header has been pushed.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Reported-by: Zixuan Chai <petalzu987@gmail.com>
Closes: https://lore.kernel.org/netdev/cover.1789987105.git.petalzu987@gmail.com/
Link: https://lore.kernel.org/netdev/179022851638.2160803.1808206741379444999@kernel.org/
Cc: Hangbin Liu <liuhangbin@gmail.com>
Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260924082951.1599377-5-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 11:00:59 -07:00
Eric Dumazet
907b978e82 net/sched: sch_teql: fix shadowed err in __teql_resolve()
__teql_resolve() declares an inner 'int err;' inside the
'if (neigh_event_send(n, skb_res) == 0)' block, shadowing the outer
'int err = 0;'. As a result, a negative return from dev_hard_header()
is written to the inner variable and __teql_resolve() still returns 0.

Remove the shadowed variable and set the outer err to -EINVAL when
dev_hard_header() returns a negative error.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Closes: https://lore.kernel.org/netdev/179022851638.2160803.1808206741379444999@kernel.org/
Cc: Jamal Hadi Salim <jhs@mojatatu.com>
Cc: Jiri Pirko <jiri@resnulli.us>
Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260924082951.1599377-4-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 11:00:59 -07:00
Eric Dumazet
ac704ff08e bridge: check llc_mac_hdr_init() return value in br_send_bpdu()
If llc_mac_hdr_init() fails (for instance if the port device type does
not support LLC or dev_hard_header() fails), br_send_bpdu() should drop
the skb instead of resetting the mac header to the LLC payload and
transmitting a malformed frame.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Closes: https://lore.kernel.org/netdev/179022851638.2160803.1808206741379444999@kernel.org/
Cc: Nikolay Aleksandrov <razor@blackwall.org>
Cc: Ido Schimmel <idosch@nvidia.com>
Cc: bridge@lists.linux.dev
Signed-off-by: Eric Dumazet <edumazet@google.com>
Acked-by: Nikolay Aleksandrov <razor@blackwall.org>
Link: https://patch.msgid.link/20260924082951.1599377-3-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 11:00:59 -07:00
Eric Dumazet
72f9dd522f llc: fix skb UAF and leaks on llc_mac_hdr_init() failure
In llc_conn_ac_resend_i_xxx_x_set_0_or_send_rr(), if llc_mac_hdr_init()
fails, kfree_skb(skb) is called instead of kfree_skb(nskb). This leaks
the newly allocated nskb, reads from the freed skb via LLC_I_GET_NR(pdu),
and double-frees skb when llc_conn_state_process() drops its reference.

In llc_sap_action_send_xid_r() and llc_sap_action_send_test_r(), nskb is
leaked if llc_mac_hdr_init() returns an error.

Free nskb in all three error paths.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Closes: https://lore.kernel.org/netdev/179022851638.2160803.1808206741379444999@kernel.org/
Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260924082951.1599377-2-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 11:00:59 -07:00
Zixuan Chai
72b5b9a28b llc: reserve device headroom for allocated frames
llc_alloc_frame() reserves link-layer headroom using the device type.
This is insufficient for stacked Ethernet devices such as VLAN devices,
where vlan_dev_hard_header() pushes a VLAN header before the lower
device's Ethernet header. An LLC response on such a device can
therefore underflow skb headroom in eth_header().

Use LL_RESERVED_SPACE() to account for the device's actual required
headroom while preserving the existing LLC device-type check.

Fixes: bf9ae5386b ("llc: use dev_hard_header")
Cc: stable@vger.kernel.org
Reported-by: VEGA <vega@nebusec.ai>
Signed-off-by: Zixuan Chai <petalzu987@gmail.com>
Signed-off-by: Ren Wei <weir@nebusec.ai>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260924012613.2533934-1-weir@nebusec.ai
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 10:58:16 -07:00
Jakub Kicinski
488089055b Merge branch 'gve-dqo-fix-handling-of-out-of-range-tso-mss'
Eric Dumazet says:

====================
gve: DQO: fix handling of out of range TSO MSS

The DQO TX path assumes that the MSS of a TSO packet is within the
range supported by the device, [88, 9728].

This holds for locally generated traffic, but not for packets coming
from a tap or from a packet socket: virtio_net_hdr_to_skb() takes
gso_size from user space and only enforces a minimum, layer 2
forwarding does not check the MTU of GSO packets, and
gso_features_check() bounds skb->len and gso_segs but never gso_size.

Patch 1, from Eddie Phillips, deals with the lower bound. It moves the
existing test out of gve_prep_tso() into gve_features_check_dqo(), so
that these packets are segmented in software instead of being dropped.

Patch 2 deals with the upper bound, which is currently not checked at
all. gve_tx_fill_tso_ctx_desc() stores gso_size into a 14 bits wide
field, so that an MSS of 16384 silently becomes zero. Falling back to
software segmentation is not an option here, because skb_segment()
splits at gso_size regardless of the MTU, and would only replace an
invalid TSO packet by non TSO packets larger than the 9728 bytes the
device supports. These packets are dropped instead.

As noted in patch 2, oversized non TSO packets can still reach the
device whenever the stack segments in software. This is not specific
to gve and is better fixed in the core, so a patch for
__is_skb_forwardable() will be sent separately for net-next.
====================

Link: https://patch.msgid.link/20260924004252.1196328-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 10:57:01 -07:00
Eric Dumazet
296c83b5cc gve: DQO: reject TSO packets with an out of range MSS
gve_prep_tso() notes that the device requires the MSS to be <= 9728,
but does not enforce it, assuming the 9K MTU enforced by the hypervisor
and the 64KB limit on TSO sizes are enough.

This does not hold for packets that were not generated locally.
A guest behind a tap, or any packet socket user, can provide an
arbitrary gso_size in virtio_net_hdr. Layer 2 forwarding does not check
the MTU for GSO packets (is_skb_forwardable()), and gso_features_check()
only bounds skb->len and gso_segs, never gso_size.

Such a packet reaches gve_tx_fill_tso_ctx_desc(), which puts gso_size
into the mss field of the TSO context descriptor. This field is 14 bits
wide, so a gso_size of 16384 is silently turned into an MSS of zero.

Drop these packets from gve_prep_tso(), and make sure that
gve_features_check_dqo() leaves their GSO bits alone: skb_segment()
splits at gso_size regardless of the MTU, so falling back to software
segmentation would give the device non TSO packets bigger than the
9728 bytes it supports.

Note that the device can still be given oversized non TSO packets when
the stack segments in software for other reasons, for instance after
TSO has been disabled with ethtool. This is a generic issue, because
the MTU check is skipped for GSO packets in the forwarding path, and
is addressed separately.

Fixes: a57e5de476 ("gve: DQO: Add TX path")
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Harshitha Ramamurthy <hramamurthy@google.com>
Link: https://patch.msgid.link/20260924004252.1196328-3-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 10:56:59 -07:00
Eddie Phillips
3b430ea623 gve: fix TX drop when GSO MSS is too small for hw
The device has a strict requirement that the minimum MSS
(gso_size) for TSO/GSO packets must be at least 88 bytes. If a packet
below this threshold is pushed to the hardware, it can cause
hardware to silently drop the packet, leading to increased latency
and retransmissions.

Currently, this is validated too late in the transmit pipeline
(gve_prep_tso), leading to silent drops.

Fix this by moving the validation into the .ndo_features_check
callback (gve_features_check_dqo). If we detect a GSO packet with
a gso_size smaller than GVE_TX_MIN_TSO_MSS_DQO, we clear the GSO
feature flags for this packet.

Fixes: a57e5de476 ("gve: DQO: Add TX path")
Signed-off-by: Eddie Phillips <eddiephillips@google.com>
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Harshitha Ramamurthy <hramamurthy@google.com>
Link: https://patch.msgid.link/20260924004252.1196328-2-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 10:56:59 -07:00
Linus Torvalds
415f204422 Landlock fix for v7.3-rc5
-----BEGIN PGP SIGNATURE-----
 
 iIYEABYKAC4WIQSVyBthFV4iTW/VU1/l49DojIL20gUCarVI0xAcbWljQGRpZ2lr
 b2QubmV0AAoJEOXj0OiMgvbSEVgA+gNbC9CVCbCo0oufZbVpQwlwtuuEKEVpvxZx
 q7oFYKCsAP9svCujGCXRHOmWhAAwe+wpXNb43l8coFn+pCfA8x3fCw==
 =NOli
 -----END PGP SIGNATURE-----

Merge tag 'landlock-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux

Pull Landlock fixes from Mickaël Salaün:
 "This mainly fixes the Landlock tracepoint support merged this cycle so
  that denial and rule events report the intended policy context,
  whether through tracefs or BTF-visible callbacks.

  The size of this all is mainly from propagating the corrected contract
  through event definitions and producers, adding new tests for the
  reported context, and updating the documentation.

  Also improve annotation and fix a GCC 16 build warning"

* tag 'landlock-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux:
  landlock: Widen ruleset versions to 64 bits
  landlock: Add counted_by in landlock_domain
  landlock: Fix tracepoint contract documentation
  selftests/landlock: Test network denial context
  selftests/landlock: Test filesystem denial blockers
  landlock: Report the effective signal number
  landlock: Report the actual ptrace tracer
  landlock: Fix network denial trace context
  landlock: Fix rule tracepoint context
  landlock: Fix filesystem denial blocker reporting
  landlock: Fix tracepoint fixed-width type names
  landlock: Work around gcc-16 -Wuninitialized warning
2026-09-24 10:53:59 -07:00
Eric Dumazet
83769c23fb gve: DQO: fix header length used by gve_can_send_tso() for UDP GSO
gve_can_send_tso() computes how many buffers each segment of a GSO
packet would span, and for this it needs the length of the headers
that the device replicates in front of every segment.

It unconditionally uses skb_tcp_all_headers(), which reads the doff
field of the TCP header. SKB_GSO_UDP_L4 packets have no TCP header:
tcp_hdrlen() then reads one byte of the UDP payload, and header_len
can be anything in [0, 60] instead of the transport offset plus the
eight bytes of the UDP header that gve_prep_tso() programs into the
TSO context descriptor.

A wrong header length shifts all the segment boundaries computed in
the loop, so the number of buffers per segment can be over or under
estimated. In the first case, GSO is needlessly disabled for this
packet by gve_features_check_dqo() and the stack has to segment it.
In the second case, the driver hands the device a packet whose
segments span more than GVE_TX_MAX_DATA_DESCS buffers.

Use the UDP header length for SKB_GSO_UDP_L4 packets, matching what
gve_prep_tso() does.

Fixes: 014c607f86 ("gve: add support for UDP GSO for DQO format")
Closes: https://lore.kernel.org/netdev/CANn89i+MS4L60sFQ49=-f-mibeveUfcrpVkD5X+Qy6SOnEpd6w@mail.gmail.com/
Signed-off-by: Eric Dumazet <edumazet@google.com>
Cc: Ankit Garg <nktgrg@google.com>
Cc: Harshitha Ramamurthy <hramamurthy@google.com>
Cc: Joshua Washington <joshwash@google.com>
Cc: Willem de Bruijn <willemb@google.com>
Reviewed-by: Ankit Garg <nktgrg@google.com>
Reviewed-by: Harshitha Ramamurthy <hramamurthy@google.com>
Link: https://patch.msgid.link/20260923145942.731365-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 10:51:29 -07:00
Eric Dumazet
06e3f54e8b net: flush skb_defer_nodes in dev_cpu_dead()
When a CPU goes offline, dev_cpu_dead() drains its softnet queues
(completion_queue, output_queue, poll_list, process_queue, and
input_pkt_queue), but leaves net_hotdata.skb_defer_nodes untouched.

If oldcpu goes offline while holding pending skbs in its
skb_defer_nodes lists (e.g. below the sysctl_skb_defer_max >> 1 IPI
threshold, or if the IPI races with CPU teardown), those skbs remain
stranded until oldcpu is brought back online. If any of these skbs
hold page_pool fragments, page_pool_destroy() will stall indefinitely
waiting for inflight pages to be returned when a netdev or driver is
torn down while oldcpu is offline.

Additionally, if smp_call_function_single_async() fails in
kick_defer_list_purge() because the target CPU went offline, reset
defer_ipi_scheduled to 0 so future IPI kicks are not blocked when the
CPU comes back online.

Also, if oldcpu was the last online CPU on its NUMA node, drain that
node's slot across all CPUs so no skbs deferred from that node remain
stranded on idle remote CPUs (or if the node itself is subsequently
offlined).

Finally, in skb_attempt_defer_free(), re-check cpu_online(cpu) and
whether the caller migrated CPUs after llist_add(), flushing the node
list if so, to close the preemption TOCTOU race against CPU/node
teardown.

Fixes: 68822bdf76 ("net: generalize skb freeing deferral to per-cpu lists")
Fixes: 5628f3fe3b ("net: add NUMA awareness to skb_attempt_defer_free()")
Closes: https://lore.kernel.org/netdev/20260916003430.3612956-1-kris.pan@intel.com/
Signed-off-by: Eric Dumazet <edumazet@google.com>
Cc: Kris Pan <kris.pan@intel.com>
Link: https://patch.msgid.link/20260923130318.607255-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 10:48:10 -07:00
Coia Prant
8db67bb6a1 net: ethernet: stmmac: dwmac-rk: fix bulk clock leak when the PHY clock fails
gmac_clk_enable() enables the bulk clocks first and then the optional
PHY clock. If clk_prepare_enable() on the PHY clock fails, the function
returns without rolling back the bulk clocks, and bsp_priv->clk_enabled
stays false, so the later gmac_clk_enable(bsp_priv, false) becomes a
no-op and the bulk clock references are leaked.

Add the missing clk_bulk_disable_unprepare() on that failure path.

Fixes: ea449f7fa0 ("net: ethernet: stmmac: dwmac-rk: rework optional clock handling")
Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Reviewed-by: Heiko Stuebner <heiko@sntech.de>
Acked-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Signed-off-by: Coia Prant <coiaprant@gmail.com>
Link: https://patch.msgid.link/20260923123713.3137146-1-coiaprant@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 10:47:24 -07:00
Dairui Zhang
56d82862a0 af_packet: fix integer overflow in prb_calc_retire_blk_tmo()
prb_calc_retire_blk_tmo() computes in 32-bit int arithmetic:

        mbits = (blk_size_in_bytes * 8) / (1024 * 1024);

If I'm reading the validation right, tp_block_size is user
controlled and packet_set_ring() only rejects values that are <= 0
as int or not page aligned, so a 256MiB block goes right through
(and alloc_one_pg_vec_page() even has a vzalloc fallback for it).
0x10000000 * 8 wraps to INT_MIN, and on a NIC reporting 1 Gbps
(div == 1) the function ends up returning -2047.

The condition is actually (8 * size) mod 2^32 >= 2^31 && div == 1,
so the trigger set is [256,512), [768,1024), [1280,1536) and
[1792,2048) MiB. Other sizes wrap to non-negative values and faster
links divide the unsigned value back below 2^31, which is why this
doesn't blow up for everyone.

What makes it fatal is what happens next in init_prb_bdqc():

        p1->interval_ktime = ms_to_ktime(prb_calc_retire_blk_tmo(...));
        hrtimer_start(&p1->retire_blk_timer, p1->interval_ktime,
                      HRTIMER_MODE_REL_SOFT);

A negative relative timeout expires immediately. The callback
unconditionally returns HRTIMER_RESTART, and hrtimer_forward() turns
the negative interval into hrtimer_resolution:

        if (interval < hrtimer_resolution)
                interval = hrtimer_resolution;

So the SOFT timer re-fires at the maximum rate forever, holding
sk_receive_queue.lock each pass. One CPU spins in softirq until the
socket is closed. Repeat with more rings and the machine is gone.

The overflow itself is ancient - it was introduced together with
TPACKET_V3 in f6fb8f100b ("af-packet: TPACKET_V3 flexible buffer
implementation."). Its effect prior to f7460d2989 ("net:
af_packet: Use hrtimer to do the retire operation", v6.18) was not
as clear-cut, though: the return value was stored into an unsigned
short retire_blk_tov, so a negative result was truncated, and a
0-jiffy delay loop could be programmed as well. Neither is nearly
as detrimental as the immediate maximum-rate spin the hrtimer
conversion turned it into.

(Unrelated to CVE-2019-20812 - that one was the ethtool failure path
returning 0, which now returns DEFAULT_PRB_RETIRE_TOV.)

Reproducer, needs CAP_NET_RAW (a --network host container has it by
default) and a 1 Gbps NIC (QEMU e1000 works):

        int fd = socket(AF_PACKET, SOCK_RAW, htons(ETH_P_ALL));
        bind(fd, ...);
        int v = TPACKET_V3;
        setsockopt(fd, SOL_PACKET, PACKET_VERSION, &v, sizeof(v));
        struct tpacket_req3 req = {
                .tp_block_size = 0x10000000,
                .tp_block_nr = 1,
                .tp_frame_size = 2048,
                .tp_frame_nr = 0x10000000 / 2048,
                .tp_retire_blk_tov = 0,
        };
        setsockopt(fd, SOL_PACKET, PACKET_RX_RING, &req, sizeof(req));

Compute in 64 bits instead. The operands are already bounded by the
existing validation, so nothing else changes. If you'd prefer a
different fix, just say so and I'll respin.

Fixes: f6fb8f100b ("af-packet: TPACKET_V3 flexible buffer implementation.")
Cc: stable@vger.kernel.org
Signed-off-by: Dairui Zhang <zhangdairui@gmail.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260923050101.1510064-1-zhangdairui@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 10:40:50 -07:00
Ginger Li
8e1937fed6 tipc: Fix a data race on mon->peer_cnt in mon_timeout()
mon_timeout() evaluates dom_size(mon->peer_cnt) before it takes mon->lock,
while mon->peer_cnt is updated under that lock by tipc_mon_add_peer() and
tipc_mon_remove_peer().  The value can therefore be stale, and the decision
whether the local domain has to be recomputed can be based on an outdated
member count.

Read mon->peer_cnt inside the write_lock_bh(&mon->lock) protected region.

Fixes: 35c55c9877 ("tipc: add neighbor monitoring framework")
Signed-off-by: Ginger Li <ginger.jzllee@gmail.com>
Reviewed-by: Tung Nguyen <tung.quang.nguyen@est.tech>
Link: https://patch.msgid.link/20260922080909.21123-1-ginger.jzllee@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 10:21:45 -07:00
Alexander Sverdlin
b94773dc4d net: phy: intel-xway: workaround 100BASE-TX Link-Up issue
MaxLinear GSW12x/GSW14x Ethernet Switch Errata Sheet states:
"An issue has been sporadically observed after device power-on on the first
link-up attempt in 100BASE-TX mode resulting in either the link-up taking a
long time, or failing to link-up altogether...

Workaround:
After power-on, enable Cable Diagnostic Mode for all ports and disable
it..."

Implement the proposed workaround unconditionally in the Intel XWAY driver
(MaxLinear GSW1xx switches incorporate Intel XWAY PHYs) because the
diagnostic bits have the same meaning even in older integral PHYs such as
GPY111/PEF7071/PHY11G. So it's not clear how to distinguish the affected
newer integrated PHYs, but the workaround should not hurt the older PHYs.

Cc: stable@vger.kernel.org
Fixes: 22335939ec ("net: dsa: add driver for MaxLinear GSW1xx switch family")
Signed-off-by: Alexander Sverdlin <alexander.sverdlin@siemens.com>
Reviewed-by: Andrew Lunn <andrew@lunn.ch>
Link: https://patch.msgid.link/20260922075251.23386-1-alexander.sverdlin@siemens.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 10:21:20 -07:00
Sidraya Jayagond
61cb282fe9 net/smc: fix UAF on lgr list traversal in smcr_port_err()
smcr_port_err() traverses smc_lgr_list.list without holding
smc_lgr_list.lock, allowing a concurrent smc_lgr_terminate_sched()
to free an lgr while it is still being dereferenced.

Hold smc_lgr_list.lock across the traversal. Update
smc_ib_gid_check() to call smcr_port_err() after releasing the lock.

Fixes: 541afa10c1 ("net/smc: add smcr_port_err() and smcr_link_down() processing")
Reviewed-by: Mahanta Jambigi <mjambigi@linux.ibm.com>
Signed-off-by: Sidraya Jayagond <sidraya@linux.ibm.com>
Reviewed-by: Dust Li <dust.li@linux.alibaba.com>
Link: https://patch.msgid.link/20260922073149.474762-1-sidraya@linux.ibm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 10:19:50 -07:00
Allison Henderson
ba3d1f480c net/rds: size a connection's path set by the transport it ends up with
__rds_conn_create() computes npaths from the caller's transport before
it decides whether a connection to one of the host's own addresses is
to be handled by the loopback transport instead.  That substitution is
what an RDS/TCP socket sending to a local address gets, and after it
the path init loop still runs for the TCP transport's RDS_MPATH_WORKERS
paths and allocates an ordered workqueue for each, while
rds_loop_conn_alloc() only ever provides transport data for path 0.

rds_conn_destroy() sizes its teardown from c_trans, by then the
loopback transport, so it visits path 0 only - and
rds_conn_path_destroy() would skip the other paths anyway, since it
returns before destroy_workqueue() for a path without transport data.
kfree(c_path) then drops the last pointers to seven workqueues.  That
repeats for every such connection, on every netns teardown or module
unload, and every distinct local destination address is a separate
connection.

Recompute npaths once the transport is final, so that creation and
destruction agree on the set of paths.  The c_path array stays sized
for the caller's transport; the unused entries are freed with it.

Fixes: 4716af3897 ("net/rds: Give each connection path its own workqueue")
Signed-off-by: Allison Henderson <achender@kernel.org>
Link: https://patch.msgid.link/20260921215027.174657-1-achender@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 10:03:23 -07:00
Sang-Hoon Choi
1a983a4e14 nfp: hold IPsec RX state under the XArray lock
nfp_net_ipsec_rx() drops the XArray lock before taking a reference to the
xfrm_state it found. The delete path can erase the entry and drop the last
state reference in that interval. RX can then try to increment a zero
refcount after the state has been queued for destruction.

The driver queues firmware invalidation asynchronously; the delete path
does not wait for the command to complete or drain pending RX processing.

The XFRM garbage collector waits for an RCU grace period before freeing
the state. That delays reclamation but does not make acquiring a reference
from zero valid.

Take the xfrm_state reference before releasing the XArray lock so
xa_erase() cannot run between lookup and reference acquisition.

Fixes: 57f273adbc ("nfp: add framework to support ipsec offloading")
Reported-by: Changyul Lee <lcy8047@gmail.com>
Signed-off-by: Sang-Hoon Choi <csh0052@gmail.com>
Link: https://patch.msgid.link/179001455912.44752.17153022439349797877.idr-bug-92@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 09:59:32 -07:00
Jakub Kicinski
5b49efa1e1 Merge branch 'net-ena-fix-resource-cleanup-on-probe-failure'
Guangshuo Li says:

====================
net: ena: fix resource cleanup on probe failure

This series fixes two resource leaks in the ena_probe() error path after
ena_device_init() has successfully initialized device resources.

Patch 1 adds the missing PHC cleanup.
Patch 2 adds the missing MMIO read request cleanup.
====================

Link: https://patch.msgid.link/20260921154202.471662-1-lgs201920130244@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 09:57:57 -07:00
Guangshuo Li
9476b44688 net: ena: fix MMIO read buffer leak on probe failure
ena_device_init() initializes the MMIO read mechanism with
ena_com_mmio_reg_read_request_init(), which allocates a coherent DMA
buffer for MMIO read responses.

The normal removal path releases this buffer through
ena_com_mmio_reg_read_request_destroy(). However, if ena_probe() fails
after ena_device_init() succeeds, the error path destroys the admin
resources and eventually frees ena_dev without destroying the MMIO read
request, leaving the coherent DMA buffer allocated.

Call ena_com_mmio_reg_read_request_destroy() in the probe error path
before releasing the remaining device resources.

This issue was found by manual code inspection.

Fixes: 1738cd3ed3 ("net: ena: Add a driver for Amazon Elastic Network Adapters (ENA)")
Cc: stable@vger.kernel.org
Signed-off-by: Guangshuo Li <lgs201920130244@gmail.com>
Link: https://patch.msgid.link/20260921154202.471662-3-lgs201920130244@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 09:57:55 -07:00
Guangshuo Li
0958ea4355 net: ena: fix PHC cleanup on probe failure
ena_probe() initializes the PHC as part of ena_device_init(), but the
probe failure path does not destroy it before freeing the PHC private
data.

The normal removal path calls ena_phc_destroy() through
ena_destroy_device() before ena_phc_free(). However, if probe fails
after ena_device_init() succeeds, the error path reaches ena_phc_free()
without unregistering the PTP clock or destroying the device PHC
resources.

Call ena_phc_destroy() in the probe error path before freeing the PHC
private data.

This issue was found by manual code inspection.

Cc: stable@vger.kernel.org tags and describe this as a consistency cleanup
Fixes: e0ea34158e ("net: ena: Add PHC support in the ENA driver")
Cc: stable@vger.kernel.org
Signed-off-by: Guangshuo Li <lgs201920130244@gmail.com>
Cc: stable
Link: https://patch.msgid.link/20260921154202.471662-2-lgs201920130244@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 09:57:55 -07:00
Jakub Kicinski
a08b39c5d5 Merge branch 'ovs-net-sched-fixes-for-uaf-after-conntrack-extension-realloc'
Ilya Maximets says:

====================
ovs, net/sched: fixes for UAF after conntrack extension realloc

One clean up change plus two fixes for the UAF on helper extension
realloc x 2.  First half for OVS and the second half for the similar
code in act_ct.  This should cover all the known cases of this problem
in these two modules.
====================

Link: https://patch.msgid.link/20260921145655.3167436-1-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 09:56:05 -07:00
Ilya Maximets
dad19b59da net/sched: act_ct: fix helper UAF due to extensions realloc
While calling the helpers, a raw pointer to the extensions area is
wired into expectations list:

  -> nf_ct_helper()
   -> helper->help()
    -> nf_ct_expect_related_report()
     -> nf_ct_expect_insert()
      -> hlist_add_head_rcu(&exp->lnode, &master_help->expectations)

In case the connection is not confirmed yet, more extensions can be
added afterwards with *_ext_add() calls reallocating the extension
space and leaving the now invalid pointer in the expectations list
that is later accessed while removing the expectation.

Make sure that helpers are called at the end after all the other
extensions are already added.

Note that the helper rejection now leaves the mark and labels set,
but that's not different from how the NAT was handled before or how
the mark and the labels were handled on confirmation failure.  And
there are no atomicity guarantees provided by the API anyway.

Fixes: a21b06e731 ("net: sched: add helper support in act_ct")
Cc: stable@vger.kernel.org
Reported-by: Axel Mierczuk <axel.mierczuk@1password.com>
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Xin Long <lucien.xin@gmail.com>
Reviewed-by: Jamal Hadi Salim <jhs@mojatatu.com>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Link: https://patch.msgid.link/20260921145655.3167436-7-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 09:56:02 -07:00
Ilya Maximets
00df72e39f net/sched: act_ct: remove 'add_helper' dead code
This variable can only become 'true' when the connection is not
confirmed, but it is only checked when it is confirmed.  So, it can be
treated as being always false and just removed.

Fixes: a21b06e731 ("net: sched: add helper support in act_ct")
Cc: stable@vger.kernel.org
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Reviewed-by: Xin Long <lucien.xin@gmail.com>
Reviewed-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/20260921145655.3167436-6-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 09:56:02 -07:00
Ilya Maximets
f85009dfcd net/sched: act_ct: avoid modifying shared unconfirmed ct entry
In a case where skb with an unconfirmed ct entry gets cloned, we may
end up processing both again but with different sets of extensions.

The series of events:

 1. The first clone wants to commit and runs the helpers wiring up
    the extension pointer into the expectation list.
 2. Then it looses the confirmation keeping the entry unconfirmed.
 3. Second clone now wants to commit labels or run NAT and adds the
    new extension for that breaking the pointer in the expectation
    list causing UAF on the destruction path later.

While this is possible to trigger, there should be no practical
network pipeline where we need to process both clones without
modifications in the same zone.  So, let's just reset the entry in
case for some reason we got an skb with a shared one.  This doesn't
affect any known use cases, but avoids any potential problems with
sharing and modification of the unconfirmed ct entry.

Unlike openvswitch module, act_ct allows for NAT without commit.
Changing that would be a uAPI break.  So, act_ct needs to reset on NAT
regardless of the commit flag to avoid reallocation of the extension
space.  This, however, doesn't really change the picture for sensible
networking cases as there should be no need to run the same packet
twice (before and after the clone) through conntrack without packet
header or zone changes and without commit.

The fixes tag points to the introduction of helpers, since that's the
main UAF trigger for the sharing.

Fixes: a21b06e731 ("net: sched: add helper support in act_ct")
Cc: stable@vger.kernel.org
Reported-by: Axel Mierczuk <axel.mierczuk@1password.com>
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Reviewed-by: Xin Long <lucien.xin@gmail.com>
Reviewed-by: Jamal Hadi Salim <jhs@mojatatu.com>
Link: https://patch.msgid.link/20260921145655.3167436-5-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 09:56:02 -07:00
Ilya Maximets
1a4151e6be net: openvswitch: conntrack: fix helper UAF due to extensions realloc
While calling the helpers, a raw pointer to the extensions area is
wired into expectations list:

  -> nf_ct_helper()
   -> helper->help()
    -> nf_ct_expect_related_report()
     -> nf_ct_expect_insert()
      -> hlist_add_head_rcu(&exp->lnode, &master_help->expectations)

In case the connection is not confirmed yet, more extensions can be
added afterwards with *_ext_add() calls reallocating the extension
space and leaving the now invalid pointer in the expectations list
that is later accessed while removing the expectation.

Make sure that helpers are called at the end after all the other
extensions are already added.

Note that the helper rejection now leaves the mark and labels set,
but that's not different from how the NAT was handled before or how
the mark and the labels were handled on confirmation failure.  And
there are no atomicity guarantees provided by the API anyway.

Fixes: cae3a26275 ("openvswitch: Allow attaching helpers to ct action")
Cc: stable@vger.kernel.org
Reported-by: Axel Mierczuk <axel.mierczuk@1password.com>
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Link: https://patch.msgid.link/20260921145655.3167436-4-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 09:56:01 -07:00
Ilya Maximets
5e6c14dd42 net: openvswitch: conntrack: remove 'add_helper' dead code
This variable can only become 'true' when the connection is not
confirmed, but it is only checked when it is confirmed.  So, it can be
treated as being always false and just removed.

Fixes: 3c1860543f ("openvswitch: add nf_ct_is_confirmed check before assigning the helper")
Cc: stable@vger.kernel.org
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Link: https://patch.msgid.link/20260921145655.3167436-3-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 09:56:01 -07:00
Ilya Maximets
26b2bd70d2 net: openvswitch: conntrack: avoid modifying shared unconfirmed ct entry
In a case where skb with an unconfirmed ct entry gets cloned, we may
end up committing both but with different sets of extensions.

The series of events:

 1. The first clone wants to commit and runs the helpers wiring up
    the extension pointer into the expectation list.
 2. Then it looses the confirmation keeping the entry unconfirmed.
 3. Second clone now wants to commit labels and adds the new extension
    for that breaking the pointer in the expectation list causing
    UAF on the destruction path later.

While this is possible to trigger, there should be no practical
network pipeline where committing both clones without modifications
into the same zone is needed.  So, let's just reset the entry in case
for some reason we got an skb with a shared one during commit.  This
doesn't affect any known use cases, but avoids any potential problems
with sharing and modification of the unconfirmed ct entry.

The fixes tag points to the introduction of helpers, since that's the
main UAF trigger for the sharing.

Fixes: cae3a26275 ("openvswitch: Allow attaching helpers to ct action")
Cc: stable@vger.kernel.org
Reported-by: Axel Mierczuk <axel.mierczuk@1password.com>
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Link: https://patch.msgid.link/20260921145655.3167436-2-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 09:56:01 -07:00
Jakub Kicinski
33a61f0923 Included fixes:
* add selftest coverage for peer VPN address validation
 * reject multicast, broadcast and loopback peer VPN addresses, which
   can never identify a peer
 * reject MP peers left with no usable VPN address, as they can never
   be selected for TX
 * reject duplicate peer VPN addresses, which made peer lookup return
   an arbitrary peer
 * fix stale entry left in the VPN address hashtable when an address
   is cleared
 * fix torn IPv6 address read on lockless TX when the unusable local
   source is cleared in place
 * fix torn IPv6 address read on lockless TX when a new local endpoint
   is learned in place
 * fix dst cache being populated with a route resolved from an already
   replaced bind
 * fix stale route being reused after the socket mark or UDP source
   port changed
 * fix bogus validation of an unspecified local source address, which
   must instead be left to route source autoselection
 * fix IPv6 link-local peer endpoints losing their scope id when
   configured via netlink, breaking route lookup
 -----BEGIN PGP SIGNATURE-----
 
 iJEEABYIADkWIQQr0db7q+Rc7Zog28Fc8QQzwdnOtwUCarECxxsUgAAAAAAEAA5t
 YW51MiwyLjUrMS4xMiwyLDIACgkQXPEEM8HZzrcvTAEA/r3oFnZ4EOtiOyA2OpZ/
 Iya+WAJ7LShj7tLymzujMe8BAIH6J2RyrWXnfDFvtSG8HGDEe4EBzVpP56tX0ZVn
 xc0N
 =fg0G
 -----END PGP SIGNATURE-----

Merge tag 'ovpn-net-20260921' of https://github.com/OpenVPN/ovpn-net-next

Antonio Quartulli says:

====================
Included fixes:
* add selftest coverage for peer VPN address validation
* reject multicast, broadcast and loopback peer VPN addresses, which
  can never identify a peer
* reject MP peers left with no usable VPN address, as they can never
  be selected for TX
* reject duplicate peer VPN addresses, which made peer lookup return
  an arbitrary peer
* fix stale entry left in the VPN address hashtable when an address
  is cleared
* fix torn IPv6 address read on lockless TX when the unusable local
  source is cleared in place
* fix torn IPv6 address read on lockless TX when a new local endpoint
  is learned in place
* fix dst cache being populated with a route resolved from an already
  replaced bind
* fix stale route being reused after the socket mark or UDP source
  port changed
* fix bogus validation of an unspecified local source address, which
  must instead be left to route source autoselection
* fix IPv6 link-local peer endpoints losing their scope id when
  configured via netlink, breaking route lookup

* tag 'ovpn-net-20260921' of https://github.com/OpenVPN/ovpn-net-next:
  selftests: ovpn: validate peer VPN addresses
  ovpn: reject invalid peer VPN addresses
  ovpn: reject multipeer peers without VPN addresses
  ovpn: reject duplicate peer VPN addresses
  ovpn: always unhash old VPN addresses before rehashing
  ovpn: replace bind when clearing stale local source
  ovpn: replace bind when learning local endpoint
  ovpn: validate peer state before caching UDP dst
  ovpn: track UDP socket route key for peer dst cache
  ovpn: skip UDP source validation for unspecified addresses
  ovpn: preserve IPv6 scope id for netlink peer endpoints
====================

Link: https://patch.msgid.link/20260921102215.3599702-1-antonio@openvpn.net
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 09:50:49 -07:00
Jiawen Wu
3173cba117 net: libwx: fix races in Tx timestamp handling
wx->ptp_tx_skb is shared between the Tx path, the PTP auxiliary
worker and the timestamp cleanup paths. The
WX_STATE_PTP_TX_IN_PROGRESS bit prevents multiple Tx paths from
submitting timestamp requests, but does not serialize the worker
against cleanup.

As a result, wx_ptp_clear_tx_timestamp() can free an skb after
wx_ptp_tx_hwtstamp_work() has obtained its pointer. The worker may
then pass the freed skb to skb_tstamp_tx() and release the same
reference again.

The cleanup path may also clear the in-progress bit while the worker
is still processing the old skb. This allows the Tx path to publish a
new skb which the worker can subsequently overwrite with NULL,
leaking its reference.

Add a dedicated spinlock to protect publication and consumption of
the Tx timestamp skb. Detach the skb and clear the in-progress bit
while holding the lock, then deliver the timestamp and release the skb
after dropping it. Use the same locked cleanup in the quiesce path,
but keep the detach there free of register accesses: quiesce runs
during PCIe error recovery, where MMIO is not reliable, and it
deliberately did not touch the device before. The lock is taken with
interrupts disabled, because netpoll can call ndo_start_xmit() with
hard interrupts already off.

When handling a Tx DMA mapping failure, keep the transmit path
reference until after comparing the skb under the lock. This prevents
skb address reuse from making the error path mistake a newer timestamp
request for the failed one.

Fixes: 06e75161b9 ("net: wangxun: Add support for PTP clock")
Reported-by: Sashiko <sashiko-bot@kernel.org>
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/6C7EC12D69217315%2B20260818074721.45536-1-jiawenwu%40trustnetic.com
Signed-off-by: Jiawen Wu <jiawenwu@trustnetic.com>
Link: https://patch.msgid.link/77431AF9A0E369F3+20260921071549.1141804-1-jiawenwu@trustnetic.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 09:40:29 -07:00
Aleksei Sviridkin
a940003f44 net: phylink: record the PHY only once bringup cannot fail
phylink_bringup_phy() stores the PHY in pl->phydev before its last
fallible step: on a MAC whose phylink ops implement LPI,
phy_eee_rx_clock_stop() can fail with a real MDIO error. The callers
unwind with phy_detach(), which knows nothing about pl->phydev, so a
pointer to a PHY that is no longer attached outlives the failed
connect.

What that costs depends on how the caller got here.
phylink_connect_phy() goes through phylink_attach_phy(), which refuses
to attach while pl->phydev is set, turning a transient MDIO error into
a permanent -EBUSY. The SFP path is worse than that: sfp_sm_probe_phy()
answers the failure with phy_device_remove() and phy_device_free(), and
it assigns sfp->mod_phy only past that error return, so nothing clears
pl->phydev and it is left pointing at a freed phy_device that
phylink_resolve() and the ethtool helpers go on reading.
phylink_fwnode_phy_connect() has no such check, so a later connect
overwrites the stale pointer and hides the problem. A disconnect does
not: phylink_disconnect_phy() hands that pointer to phy_disconnect(),
and the second phy_detach() on the same PHY drops references the first
one already released.

Found while making a DSA port survive a PHY whose driver arrives after
the switch probes: keeping the port across a failed connect and
retrying is what makes this window reachable.

Publish the pointer after the last call that can fail instead of
unwinding it afterwards. Nothing between the two points reads
pl->phydev, and the registration that follows cannot fail:
phy_request_interrupt() falls back to polling on its own. The PHY-side
state keeps the order it had, so no MDIO operation moves relative to
another.

Fixes: 03abf2a7c6 ("net: phylink: add EEE management")
Signed-off-by: Aleksei Sviridkin <f@lex.la>
Link: https://patch.msgid.link/20260920222044.1752860-1-f@lex.la
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 09:36:05 -07:00
Ming Wang
f75f21ef36 net: usb: cdc_mbim: add MeiG Smart SRM821 to ZLP whitelist
The MeiG Smart SRM821 5G module (0x2dee:0x4d53) crashes and drops off
the USB bus when it receives a Zero Length Packet (ZLP) after sending
or receiving an NTB of exactly 16384 bytes (tx_max).

According to the MBIM specification, devices do not require a ZLP
if the NTB size is exactly dwNtbOutMaxSize. However, the cdc_mbim
driver defaults to sending ZLPs for devices not explicitly whitelisted
to accommodate non-conformant hardware. This default behavior breaks
the strictly conformant MeiG SRM821 module.

Add this device to the ZLP conformance whitelist (cdc_mbim_info) so
the driver will pad the NTB to avoid sending ZLPs, preventing the
device firmware from crashing.

Cc: stable@vger.kernel.org
Signed-off-by: Ming Wang <wangming01@loongson.cn>
Link: https://patch.msgid.link/20260920074500.826121-1-wangming01@loongson.cn
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 09:27:47 -07:00
Guopeng Zhang
1765a153d9 cgroup/pids: Restore pids.events notifications in local mode
A fork rejected by the pids controller increments the counter reported by
pids.events. When local event accounting is selected, however, pids_event()
returns after notifying only events_local_file, leaving pids.events pollers
asleep.

On legacy hierarchies, pids.events.local does not exist. With
pids_localevents, pids.events reports the same local counter. In both
cases, pids.events changes without generating a notification.

This can be reproduced with a pids_localevents mount:

    mkdir /tmp/test
    mount -t cgroup2 -o pids_localevents none /tmp/test
    mkdir /tmp/test/t
    echo 1 > /tmp/test/t/pids.max
    cat /tmp/test/t/pids.events                 # max 0
    timeout 3 inotifywait -e modify /tmp/test/t/pids.events &
    sh -c 'echo $$ > /tmp/test/t/cgroup.procs; (true &)' 2>/dev/null
    wait
    cat /tmp/test/t/pids.events                 # max 1

Without this patch, inotifywait times out without reporting an event.
Notify pids.events before returning from the local event path.

Fixes: 3f26a885a0 ("cgroup/pids: Add pids.events.local")
Cc: stable@vger.kernel.org # v6.11+
Signed-off-by: Guopeng Zhang <zhangguopeng@kylinos.cn>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-09-24 06:22:11 -10:00
Yilin Zhang
fe99bbeee5 tcp: fix use-after-free of retransmit_skb_hint in tcp_send_synack()
When tcp_send_synack() replaces the cloned SYN skb at the head of the
retransmit queue with a copy, it frees the original with
tcp_rtx_queue_unlink_and_free() and only repairs tp->highest_sack.
tp->retransmit_skb_hint keeps pointing at the freed
skbuff_fclone_cache object.

The dangling hint is read in tcp_verify_retransmit_hint() and used as
the root of the rbtree walk in tcp_xmit_retransmit_queue().  An
unprivileged TFO client (sendmsg(MSG_FASTOPEN)) can arm the hint with
an attacker-supplied ICMP fragmentation-needed message, after which a
simultaneous open frees the armed SYN skb:

  BUG: KASAN: slab-use-after-free in tcp_mark_skb_lost (net/ipv4/tcp_input.c:1316)
  Read of size 4 at addr ffff88800604d928 by task swapper/1/0
  Call Trace:
   tcp_mark_skb_lost (net/ipv4/tcp_input.c:1316)
   tcp_simple_retransmit (net/ipv4/tcp_input.c:3158)
   tcp_v4_err (net/ipv4/tcp_ipv4.c:587)

Sync the hint to the copy.

Fixes: c31b70c996 ("tcp: Add logic to check for SYN w/ data in tcp_simple_retransmit")
Reported-by: Kimi Security Team <bug-report@moonshot.ai>
Tested-by: Weiming Shi <shiweiming@moonshot.ai>
Signed-off-by: Yilin Zhang <yilinzhang@moonshot.ai>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/8a9dff4063a2745653b7e88ceb745d75efa16e68.1790224474.git.yilinzhang@moonshot.ai
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 09:17:10 -07:00
Hui Peng
26cc0e69cc mctp: route: iterate socket tag list in mctp_lookup_prealloc_tag()
When a socket transmits a packet with MCTP_TAG_PREALLOC set,
mctp_lookup_prealloc_tag() iterates over the per-netns &mns->keys list
and matches netid, req_tag, peer_addr, and manual_alloc, without
checking whether tmp->sk == &msk->sk. This allows any MCTP socket in the
same network namespace to use and consume another socket's preallocated
tag.

Iterate the socket's own tag list (&msk->keys via sklist) instead of the
namespace-wide &mns->keys list in mctp_lookup_prealloc_tag(), ensuring
that only tags allocated by msk are matched.

Tested in QEMU against Linux 7.3.0-rc3 by allocating a manual tag
(0x18) on socket A via SIOCMCTPALLOCTAG for peer EID 9 and sending a
4-byte message with MCTP_TAG_PREALLOC from socket B in the same network
namespace. On the unfixed kernel, sendto(sock_b) using socket A's
preallocated tag succeeds (ret = 4); with this patch applied,
sendto(sock_b) fails with -ENOENT (errno = 2) while sendto(sock_a)
succeeds (ret = 4).

Fixes: 63ed1aab3d ("mctp: Add SIOCMCTP{ALLOC,DROP}TAG ioctls for tag control")
Suggested-by: Jeremy Kerr <jk@codeconstruct.com.au>
Cc: stable@vger.kernel.org
Signed-off-by: Hui Peng <benquike@gmail.com>
Link: https://patch.msgid.link/20260921051002.1656692-1-benquike@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 09:11:56 -07:00
Linus Torvalds
5fc5768c7c bpf-fixes
-----BEGIN PGP SIGNATURE-----
 
 iQJRBAABCgA7FiEE+soXsSLHKoYyzcli6rmadz2vbToFAmq1NvAdHGFsZXhlaS5z
 dGFyb3ZvaXRvdkBnbWFpbC5jb20ACgkQ6rmadz2vbTpmpA//fqEoC6Sq1zxo3ADH
 hV0Z9ewkNTjjH85QnispjcRkAhSHG3JscNXKRXm1NmNkwJsHJ4TDIKcjABYDjquD
 wqfvL9hLXPsvud0M/PR6/CZeBAWXpukkaxeYedY+83ttTHjzR0tDq0Ne9yvIV+Nc
 hS0qFIIHs8C6l2nyuNSxvgrv216orG8qd0Bi3tpDfsLqCsLLEmDyQ1H+ZzpJF2xW
 RV3oMcggCeo305m8+uiofQGf8hmmRrmA7SEfF+Qe08ab2GOn9glINfVTbTE3AdQW
 fkvAk8Zuio3hwwMBHDWWYXrKO0N3ykzcDk4V6JPUWiH1dOANf1tS3G1uJyxb4tsv
 dHVdA0xL3sg7YnuSywfb82vTXvQz5QEFeDYxxLB2fMJe8LcCfgF1lkIr1DbdrMjI
 hlozGCs3p/GVIVNhjGVPezWsUvzK4PIuKVM6U1qWojtAhXU/bJCvP0iBU83RlBL7
 S0GGSC/qvkuPed9hAp2pnrrRo/9GEp1PZq8AXsp3nti0OaRxgxpjjQQFzZEWlOrR
 bErtbKziHzl2xARvCRycNZ+QWT4ZdjmK++8pba7hOuFkRclOV4KRWk4OthBvcADu
 NfNsVXdXpOnesg7TNIUJ8heVAYPjKaBHJBuG2Prghrnc95uqEQwOfbYT15zAbe6I
 l3HqGerSa7WtbrnCwsf1DvyPPII=
 =ZxKg
 -----END PGP SIGNATURE-----

Merge tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf

Pull bpf fixes from Alexei Starovoitov:

 - Fix bpf_skb_change_tail() to drop the checksum offload instead of
   rejecting the trim of CHECKSUM_PARTIAL skbs (Daniel Borkmann)

 - Add KF_PERFMON kfunc flag and require CAP_PERFMON for kfuncs that
   read arbitrary memory and for untrusted read-only memory reads
   (Daniel Borkmann)

 - Clear scalar delta on narrowing stack spill (Daniel Borkmann)

 - Set up the frame pointer for the exception callback in arm64 JIT, and
   zero-fill other CPUs when BPF_F_CPU update creates a per-cpu hash
   element (Donggeun Yoo)

 - Various fixes (Emil Tsalapatis):
     - Fix bounds check underflow for skb-backed dynptrs
     - Fix rx_queue_mapping context access code generation in bpf_sock
     - Reject packet pointer arguments to subprogs that may mutate the
       packet
     - Reject ALU instructions that see arena and non-arena operands on
       different code paths

 - Fix copied_seq double-counting on sockmap self-redirect
   (Geliang Tang)

 - Fix divide-by-zero in btf_struct_walk() on a flexible array of
   zero-sized elements, fix out-of-bounds read of rtt_min in sock_ops
   (Jiayuan Chen)

 - Fix bpf_sock_destroy() out-of-bounds read of sk_protocol on TIME_WAIT
   and request socks, and sleeping under RCU when destroying a listener
   with pending children (Jiayuan Chen)

 - Fix JEQ/JNE with immediate operand in MIPS32 JIT and missing zero
   extension of BSWAP 16/32 in MIPS64 JIT (Johan Almbladh)

 - Avoid soft lockup in htab lookup[_and_delete] batch operations on
   large maps (Jose Fernandez)

 - Various fixes (Kumar Kartikeya Dwivedi):
     - Verify global subprogs in each sleepability context they are
       called from
     - Make post-verification instruction rewrites killable
     - Preserve packet pointer displacement in regsafe()
     - Apply CO-RE relocations before subprogram validation, restrict
       CO-RE poisoning to relocatable instructions, and reject truncated
       ldimm64 CO-RE relocations in libbpf
     - Assign lock identity to callback map values
     - Compare stack frames in regs_exact()
     - Bound ownership depth through local kptrs and graph roots

 - Fix u32 overflow in map batch operations when the map size exceeds
   4GB (Masoud Aghasi)

 - Fix UAF in bpf memalloc due to concurrent consumption of ttrace lists
   in alloc_bulk() (Pu Lehui)

 - Allow gotox as the terminal instruction of a program or a subprogram
   (Siddharth Chintamaneni)

 - Disallow bpf_skb_pull_data() for LWT_SEG6LOCAL, skip unsettled links
   in link iterator, and reject dev-bound-only programs on other devices
   (Weiming Shi)

 - Reject non-negative stack offsets in stack_slot_obj_get_spi()
   (Xu Yunxiang)

 - Check params size before reading reserved fields in
   bpf_crypto_ctx_create() (Yuqi Xu)

 - Reject max_entries > INT_MAX in sock_map_alloc() (Zhao Gongyi)

 - Use a 32-bit compare in xsk_map_gen_lookup() (Zhiling Zou)

 - Use kvfree() in xdp_test_run_teardown() (Zhixing Chen)

* tag 'bpf-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf: (58 commits)
  selftests/bpf: Test per-cpu initialization of a BPF_F_CPU created element
  bpf: Zero-fill other CPUs when BPF_F_CPU creates a per-cpu hash element
  bpf: Fix BSWAP 32 and 16 on MIPS64
  bpf: Fix immediate JMP JEQ/JNE on MIPS32
  bpf: Reject dev-bound-only programs on other devices
  bpf, sockmap: Reject max_entries > INT_MAX in sock_map_alloc
  selftests/bpf: Test for mixed arena/nonarena code paths
  bpf: Prevent variable arena/non-arena register contents
  selftests/bpf: Test rejection of pkt args to mutating subprogs
  bpf: Reject pkt arguments in mutating subprogs
  selftests/bpf: Add selftests for rx_queue_mapping context access
  bpf: Fix bpf_sock context code generation
  selftests/bpf: Test dynptr slices past end of skb
  bpf: Fix bounds check for skb-backed dynptrs
  selftests/bpf: Reject iterator destruction through fp+0
  bpf: Reject non-negative offsets in stack_slot_obj_get_spi()
  bpf: Check params size before reading reserved fields
  selftests/bpf: Check local object ownership depth
  bpf: Bound ownership depth through local kptrs and graph roots
  selftests/bpf: Cover frame changes in bounded loops
  ...
2026-09-24 08:25:26 -07:00
Jakub Kicinski
0b2bbfee2b Merge branch 'net-dsa-mt7530-fix-two-crashes-on-driver-unbind'
Aleksei Sviridkin says:

====================
net: dsa: mt7530: fix two crashes on driver unbind

Unbinding the MT7530 driver from an MT7531 dereferences NULL in
regulator_disable(). On a Netcraze NC-1012 (MT7981B + MT7531, 6.18.44):

  # echo mdio-bus:1f > /sys/bus/mdio_bus/drivers/mt7530-mdio/unbind

oopses there, and the build it was found on sets CONFIG_PANIC_ON_OOPS, so
the board goes down with it. Fix that and the same command gets as far as
mt7530_remove_common(), which disposes interrupt mappings the switch's own
regmap-irq chip still owns; the regmap-irq thread then faults in
handle_nested_irq() later in the same teardown. rmmod reaches both, since
mdio_module_driver() calls .remove on module exit.

Patch 1 is the regulator one. mt7530_probe() requests the core and io
supplies only for ID_MT7530 and mt7530_setup() enables them under the same
test, but mt7530_remove() disables them unconditionally, so on an MT7621 or
an MT7531 both pointers are still NULL from devm_kzalloc(). It reaches the
MDIO front end only.

Patch 2 is the interrupt one, and it reaches further.
mt7530_remove_common() disposes the per-PHY interrupt mappings by hand from
.remove, while the regmap-irq chip that owns the domain is devm-registered
and its parent interrupt is only freed once .remove has returned.
regmap_del_irq_chip() disposes the same mappings itself, in an order that
cannot race, so the driver's call adds nothing but a window. That helper is
called from both front ends, so the defect also covers the MMIO parts -
MT7988, EN7581, AN7583 and EN7528 - which have no regulators and never meet
the first defect at all.

The order is not arbitrary. On an MT7531 the regulator fault happens in the
first thing mt7530_remove() does with the switch, so execution never
reaches the interrupt defect. The second only became visible once the first
was fixed, which is also how both came to be found on one board.

Found and verified there. Without patch 1 the unbind panics in
regulator_disable(); with patch 1 alone the panic moves on to
handle_nested_irq(); with both, three unbind/bind cycles run, two back to
back and a third after a pause. In the two whose dmesg was captured, each
unbind removes the switch from the driver directory and takes lan1 to lan4
with it, each bind brings them back, and lan1 relinks at 1Gbps/full after
both binds, lan4 after the second. uptime rose from 58 to 202 seconds
across the three without resetting and pstore gained no new record. The
third cycle stayed unbound long enough to read the descriptors: no mt7530
line in /proc/interrupts and no irq/79, irq/80 or irq/81 directory, and the
next bind reuses those three numbers - regmap-irq freeing and disposing
what the driver no longer touches. The kernel under test was identified by
the sha256 of its ELF notes section, read from /sys/kernel/notes on the
running board and computed in advance from the flashed image.

What hardware could not answer here. There is no MT7530 or MT7621 part on
this bench, so the ID_MT7530 branch that patch 1 adds was checked by
reading the generated code rather than by running it, and no MMIO part was
available to exercise patch 2 on that front end either. One unrelated WARN
remains across the unbind, from sysfs_remove_link() under
dsa_user_destroy(); it is a separate DSA teardown-ordering defect and is
not addressed here.

v1: https://lore.kernel.org/20260914202421.2737079-1-f@lex.la
====================

Link: https://patch.msgid.link/20260918015020.2518315-1-f@lex.la
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 08:16:58 -07:00
Aleksei Sviridkin
0d80ba0a20 net: dsa: mt7530: leave the MDIO IRQ mappings to regmap-irq
mt7530_remove_common() disposes the per-PHY interrupt mappings from
.remove, but the regmap-irq chip that owns the domain is devm-registered,
so its parent interrupt is only freed once .remove has returned. The
switch's own regmap-irq thread can therefore still dispatch on a mapping
that is already gone: irq_find_mapping() returns 0, irq_to_desc() returns
NULL and handle_nested_irq() locks desc->lock without checking it. The
attached PHYs have not given those interrupts back yet either, which the
kernel warns about a moment before the fault.

regmap_del_irq_chip() disposes the same mappings itself, after freeing the
parent interrupt and before removing the domain, so there is nothing left
for the driver to do here. Until it runs the descriptors stay alive, and a
late dispatch on one of them is harmless: dsa_unregister_switch() has freed
the PHY handlers by then, so handle_nested_irq() finds no action and
returns.

Fixes: 254f6b272e ("dsa: mt7530: Utilize REGMAP_IRQ for interrupt handling")
Signed-off-by: Aleksei Sviridkin <f@lex.la>
Link: https://patch.msgid.link/20260918015020.2518315-3-f@lex.la
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 08:16:53 -07:00
Aleksei Sviridkin
c2cdef41e0 net: dsa: mt7530: fix NULL dereference on unbind of MT7531 and MT7621
The core and io supplies are only requested for ID_MT7530: both the
devm_regulator_get() in probe and the regulator_enable() in
mt7530_setup() are guarded by the switch id, but mt7530_remove()
disables them unconditionally. On an MT7621 or an MT7531 both pointers
are still NULL from devm_kzalloc(), so rmmod or a sysfs unbind calls
regulator_disable() on NULL.

Fixes: ddda1ac116 ("net: dsa: mt7530: support the 7530 switch on the Mediatek MT7621 SoC")
Signed-off-by: Aleksei Sviridkin <f@lex.la>
Link: https://patch.msgid.link/20260918015020.2518315-2-f@lex.la
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-09-24 08:16:53 -07:00
Alexei Starovoitov
1b1dca5682
Merge branch 'bpf-fix-per-cpu-initialization-of-a-bpf_f_cpu-created-hash-element'
Donggeun Yoo says:

====================
bpf: fix per-cpu initialization of a BPF_F_CPU created hash element

A BPF_F_CPU update that creates a [lru_]percpu_hash element writes the
named CPU's slot and leaves the others holding the recycled element's
values, so a lookup of the new key returns a deleted key's per-cpu
values.

Patch 1 zero-fills the other CPUs.  Patch 2 adds the selftest: the
existing cpu_flag subtests always prime a key with BPF_F_ALL_CPUS
first, so the create path is not covered today.

v1: https://lore.kernel.org/bpf/20260920093153.439743-1-donggeunyoo.kernel@gmail.com/
v2: https://lore.kernel.org/bpf/20260923000801.1764758-1-donggeunyoo.kernel@gmail.com/

Changes in v3:
 - patch 1: key init_cpu on BPF_F_CPU rather than on onallcpus (Leon
   Hwang)
 - patch 1: re-flow the paragraph naming pcpu_copy_value() (BPF CI AI
   review)
 - patch 2: pin the thread across the delete and the create, and name a
   CPU other than the pinned one, so the BPF_F_NO_PREALLOC arm does not
   depend on staying put (BPF CI AI review)

Changes in v2:
 - patch 1: cover the BPF_F_CPU entry condition in the block comment
   above pcpu_init_value() (BPF CI AI review, Alexei Starovoitov)
 - patch 2: skip the new subtests instead of failing them on a
   uniprocessor machine (Sashiko AI review)
====================

Link: https://patch.msgid.link/20260924102321.2120434-1-donggeunyoo.kernel@gmail.com
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
2026-09-24 14:25:01 +00:00
Donggeun Yoo
3422808f4e
selftests/bpf: Test per-cpu initialization of a BPF_F_CPU created element
The existing cpu_flag subtests always prime a key with BPF_F_ALL_CPUS
before any BPF_F_CPU write, so the create path is never covered.

Add a subtest that creates the element with BPF_F_CPU on a map with
max_entries 1, so the key can only reuse the element the previous key
released, and check that the CPUs the update did not name read back
zero.  Run it for PERCPU_HASH preallocated and BPF_F_NO_PREALLOC,
whose per-cpu areas come from different allocators, and for
LRU_PERCPU_HASH.

Under BPF_F_NO_PREALLOC the reuse is only guaranteed on the cpu that
ran the delete, so pin the thread across the pair, and name a cpu other
than that one in map_flags.

Signed-off-by: Donggeun Yoo <donggeunyoo.kernel@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260924102321.2120434-3-donggeunyoo.kernel@gmail.com
2026-09-24 14:24:52 +00:00
Donggeun Yoo
c3a66e5f5b
bpf: Zero-fill other CPUs when BPF_F_CPU creates a per-cpu hash element
pcpu_init_value() initializes the per-cpu area of a newly created
[lru_]percpu_hash element.  The area is recycled, so when the value
comes from a BPF program (onallcpus == false) it writes the running
CPU's slot and zeroes the rest.

bpf_percpu_hash_update() passes onallcpus == true, which delegates to
pcpu_copy_value().  pcpu_copy_value() writes only the CPU named in
map_flags when BPF_F_CPU is set, so on the create path the other slots
keep the recycled element's values:

  update(k1, 0xdeadc0de, BPF_F_ALL_CPUS)  every CPU holds 0xdeadc0de
  delete(k1)                              element back on the freelist
  update(k2, 0xc0ffee, BPF_F_CPU | 0)     creates, writes CPU 0 only
  lookup(k2)                              CPU 0 0xc0ffee, rest 0xdeadc0de

Zero-fill the other CPUs on that arm too.

Fixes: c6936161fd ("bpf: Add BPF_F_CPU and BPF_F_ALL_CPUS flags support for percpu_hash and lru_percpu_hash maps")
Signed-off-by: Donggeun Yoo <donggeunyoo.kernel@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260924102321.2120434-2-donggeunyoo.kernel@gmail.com
2026-09-24 14:24:51 +00:00
Christian Lamparter
7c9f391ec8 net: emac: move setting of netops to fix crash
fixes the following crash on driver initialization:

|BUG: Kernel NULL pointer dereference on read at 0x00000158
|Faulting instruction address: 0xc0566b40
|Oops: Kernel access of bad area, sig: 11 [#1]
|BE PAGE_SIZE=4K  PowerPC 44x Platform
|Modules linked in:
|CPU: 0 UID: 0 PID: 1 Comm: swapper/0 Tainted: GW 7.3.0-rc3+ #1
|Tainted: [W]=WARN
|Hardware name: MyBook Live APM821XX 0x12c41c83 PowerPC 44x Platform
|NIP:  c0566b40 LR: c05648f8 CTR: c04c1e9c
|REGS: c1053a20 TRAP: 0300   Tainted: GW  (7.3.0-rc3+)
|MSR:  0002b000 <CE,EE,FP,ME>  CR: 24008808  XER: 00000000
|DEAR: 00000158 ESR: 00000000
|GPR00: c05648f8 c1053b10 c1063600 c1030000 c5ab3000 00000000 [...]
|GPR08: 00000002 00000000 00000000 c1053b40 84002808 00000000 [...]
|GPR16: cfffd210 00000002 c0beafcc cfffc960 00000000 c1030644 [...]
|GPR24: c0beafbc c1037000 00000000 0000000a 00000000 c1030000 [...]
|NIP [c0566b40] phy_link_topo_add_phy+0x2c/0x1d0
|LR [c05648f8] phy_attach_direct+0x1a4/0x368
|Call Trace:
|[c1053b10] [c0811e04] klist_put+0x54/0xb4 (unreliable)
|[c1053b40] [c05648f8] phy_attach_direct+0x1a4/0x368
|[c1053b70] [c0564ae8] phy_connect_direct+0x2c/0x60
|[c1053b90] [c056c7a8] of_phy_connect+0x50/0x74
|[c1053bc0] [c0572a88] emac_probe+0xd50/0x119c
|[c1053c90] [c04cb770] platform_probe+0x74/0xa4
|[c1053cb0] [c04c8f08] really_probe+0x120/0x2b0
|[c1053cd0] [c04c9254] __driver_probe_device+0x1bc/0x1fc
|[c1053d00] [c04c9334] driver_probe_device+0x38/0xa8
|[c1053d30] [c04c9560] __driver_attach+0xf4/0x10c
|[c1053d50] [c04c6af8] bus_for_each_dev+0x68/0xd0
|[c1053d90] [c04c7c34] bus_add_driver+0xcc/0x1ec
|[c1053dc0] [c04c9f6c] driver_register+0xcc/0x110
|[c1053de0] [c0aa3f40] emac_init+0x1c4/0x200

This bug showed up starting with v7.3-rc1. At the NIP in
phy_link_topo_add_phy() is a netdev_need_ops_lock() check.
This was added by the following
commit ded86da4bb ("net: ethtool: relax ethnl_req_get_phydev() locking assertion")

The bug shows up because at the time of_phy_connect() was called, the
netdev_ops were *not yet* determined. My fix is to move the code that
sets netdev_ops+commac.ops+ethtool_ops further up as emac_init_config()
derives that by looking at the device-tree and sets the required
dev->phy_mode accordingly.

During review, the Sashiko bot's AI stated that the commac assignment
became a dead store. Great catch! To keep the original behavior as-is,
one mentioned option "set dev->commac.ops = &emac_commac_ops only in
the non-gige case?" sounded like a great plan. So the dev->commac.ops
assignment for the non-gige-case moves into the else block.

This patch was tested on a WD MyBook Live (RGMII). the device now works again.

Fixes: ded86da4bb ("net: ethtool: relax ethnl_req_get_phydev() locking assertion")
Signed-off-by: Christian Lamparter <chunkeey@gmail.com>
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Link: https://patch.msgid.link/49cd7343bc0e507c022071f3e2b5662b053dca73.1790007431.git.chunkeey@gmail.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-24 15:37:47 +02:00
Fang Xieyan
d6ec384c87 net/sched: act_ife: validate metadata length before decoding
skbmark_decode(), skbprio_decode() and skbtcindex_decode() read fixed-size
values from the TLV payload without validating its length.

A malformed IFE frame can declare a shorter payload, causing the decoders
to consume bytes beyond the declared metadata value:

[TLV type=IFE_META_SKBMARK len=4]
-> dlen == 0, but decode reads 4 bytes

The decoder may therefore set skb metadata from unintended input.

Validate the payload length before decoding and return -EINVAL for
invalid lengths. Read the values with get_unaligned_be32() and
get_unaligned_be16(), as TLV payloads are not guaranteed to be
aligned. Teach tcf_ife_decode() to log a decoder error separately
from an unknown metaid; both are counted as overlimits and decoding
continues with the remaining metadata.

The metadata length issue was found by an automated audit of the IFE
decode path at v6.18-rc7 and reproduced with a userspace sanitizer
model of the decode path. Compile-tested on x86_64 with defconfig and
NET_ACT_IFE=y: act_ife.o and the three act_meta_*.o build
warning-free.

Fixes: 084e2f6566 ("Support to encoding decoding skb mark on IFE action")
Fixes: 200e10f469 ("Support to encoding decoding skb prio on IFE action")
Fixes: 408fbc22ef ("net sched ife action: Introduce skb tcindex metadata encap decap")
Assisted-by: Hawkeye:GLM-5.3-flash
Assisted-by: Qoder:Qwen3.8-Max
Signed-off-by: Fang Xieyan <fangxy@xiaopeng.com>
Link: https://patch.msgid.link/20260921125441.81459-1-fangxy@xiaopeng.com
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-09-24 15:28:37 +02:00
Linus Torvalds
f49a343b30 Pin control fixes for the v7.3 series:
- Serialize the calls to pinctrl_generic_dt_nod_to_map()
 
 - Fix a typo in S4 group in the Meson driver.
 
 - Fix bank width and regmap usage in the MPFS-SSIO driver.
 
 - Data register output latch behavior and voltage encoding
   fixes in the Sunxi driver.
 
 - Free the IRQ domain on the error path in the single driver.
 
 - Fix some QUP1 SE2/SE3 groups and a missing OF module
   alias in the Qualcomm drivers.
 
 - Fix up the register banks for AON (I guess always-on)
   pins in the Tegra238 driver.
 
 Signed-off-by: Linus Walleij <linusw@kernel.org>
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEElDRnuGcz/wPCXQWMQRCzN7AZXXMFAmq1Bm4ACgkQQRCzN7AZ
 XXM36g//fk1LQHrRfCHR4VmPTP8SIielpN4J6RaTdVZYrOF5HB4FU1nkJvhgNH73
 YAiIeBR/XG7gevMnfuE3w9H6zSrw7bb4l1+piV+o4T5Zo1SNC2WAnHdIXJ/isrZq
 rHjwBcd5kvit+2VlaUxLZERc8V8L7+t6rD1+QDDrgqd2Ji3ekKB5bB4kdumuE5Yp
 c+oirXvcgC7Bi2Ql35HTeWFmie5IAbfAHFiXryUwJREZNbakS17sQRXQxesib0fi
 UPml71+vDVluKtluB065PwE8W4JpLi7+SkhMcB4xgT/ofMt46wzTtCU6VfBaTM/0
 dHLr6Kr9ajjmNzk5i3DZMFFY2C2ywqyGCCfgTO2m5wX9WaBNFoO8UsZ9Xb/Ij+RZ
 9T6Ub6QCJcvvZIScRXTAQEKiJ4I4cMOt1wL4rJfX8YQA0Aha9r7YXxnLxBM7AgvB
 8AzIJcgm5RSC2uJaRnNpSgOvBk6XYXGnC22MTbDR3CIdkYZcWy7M9XLYI21UtCOh
 ZmYvxiJGHoa2qe+AgHYgvO5lFIW9qTUOL11XXQJQwziTC7A/I/Qnf6Gwyq12KE1Q
 TuAGRqeWCo4UPEYhc5myKceF1J6ljJosWwsF3VH9Vj4N2/CHIcbikWE33g1qCIbf
 VCsPaPKegsLoK4nNau0z296nqNmNFjmuyTn8gcdEMlQYGP4WwhA=
 =PxJM
 -----END PGP SIGNATURE-----

Merge tag 'pinctrl-v7.3-2' of git://git.kernel.org/pub/scm/linux/kernel/git/linusw/linux-pinctrl

Pull pin control fixes from Linus Walleij:

 - Serialize the calls to pinctrl_generic_dt_nod_to_map()

 - Fix a typo in S4 group in the Meson driver

 - Fix bank width and regmap usage in the MPFS-SSIO driver

 - Data register output latch behavior and voltage encoding
   fixes in the Sunxi driver

 - Free the IRQ domain on the error path in the single driver

 - Fix some QUP1 SE2/SE3 groups and a missing OF module alias
   in the Qualcomm drivers

 - Fix up the register banks for AON (I guess always-on) pins
   in the Tegra238 driver

* tag 'pinctrl-v7.3-2' of git://git.kernel.org/pub/scm/linux/kernel/git/linusw/linux-pinctrl:
  pinctrl: tegra238: Fix register bank for AON pin groups
  pinctrl: qcom: ipq5210: Publish the OF module alias
  pinctrl: qcom: nord: Split QUP1 SE2/SE3 into lane-pair functions
  pinctrl: single: free the IRQ on domain creation failure
  pinctrl: sunxi: A523: fix voltage withstand encoding
  pinctrl: sunxi: keep a shadow copy of the data register output latches
  pinctrl: mpfs-mssio: use correct regmap function to set bank voltage
  pinctrl: mpfs-mssio: fix width of unused bank voltage setting
  pinctrl: generic: serialise pinctrl_generic_dt_node_to_map()
  pinctrl: meson: Fix typo in s4 group name
2026-09-24 06:19:09 -07:00
Linus Torvalds
e2936d228b memblock: add missing HugeTLB flag name
Rework of HugeTLB bootmem allocation forgot to add the name for a new
 RSV_HUGETLB flag for proper display in debugfs.
 
 Add the flag name now.
 -----BEGIN PGP SIGNATURE-----
 
 iQEzBAABCgAdFiEEeOVYVaWZL5900a/pOQOGJssO/ZEFAmq00GcACgkQOQOGJssO
 /ZFJxwf/TlhW+APk3p9UXHTMXTvsefWxrox7mlVEBAh+nqpiF43aDeZySV6IW4yr
 cHqBpI+yheyiuh1dOqfsPB/8b8jFV2GyBmknD+QbRhJi5Q8KXALVYgaKIi2WXaj5
 rHNuVn7aTwUbwEBaiwrFWfWce6GSLqFGoDYTb2iSg1Z/ZdOXnWeTil8iOONCAMSw
 3h8joiqPV7OlNDInxWYrlwDAvPo5Sr0VwBT7+GpuahjVB21l/XINPMwZUmLucF3t
 y3+WHYTAH8f98KF0pBzJeD7XJX/nupeyzKhvJHLbg6XxbvFh3pz3OTts4tQA3Khy
 REwmEB/v9Hbl1wsfA6XNROkgQN+l5g==
 =t0mt
 -----END PGP SIGNATURE-----

Merge tag 'fixes-2026-09-24' of git://git.kernel.org/pub/scm/linux/kernel/git/mm/memblock

Pull memblock fix from Mike Rapoport:
 "The rework of HugeTLB bootmem allocation forgot to add the name for a
  new RSV_HUGETLB flag for proper display in debugfs.

  Add the flag name now"

* tag 'fixes-2026-09-24' of git://git.kernel.org/pub/scm/linux/kernel/git/mm/memblock:
  mm: memblock: add missing HugeTLB flag name
2026-09-24 06:16:19 -07:00