The existing cpu_flag subtests always prime a key with BPF_F_ALL_CPUS
before any BPF_F_CPU write, so the create path is never covered.
Add a subtest that creates the element with BPF_F_CPU on a map with
max_entries 1, so the key can only reuse the element the previous key
released, and check that the CPUs the update did not name read back
zero. Run it for PERCPU_HASH preallocated and BPF_F_NO_PREALLOC,
whose per-cpu areas come from different allocators, and for
LRU_PERCPU_HASH.
Under BPF_F_NO_PREALLOC the reuse is only guaranteed on the cpu that
ran the delete, so pin the thread across the pair, and name a cpu other
than that one in map_flags.
Signed-off-by: Donggeun Yoo <donggeunyoo.kernel@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260924102321.2120434-3-donggeunyoo.kernel@gmail.com
pcpu_init_value() initializes the per-cpu area of a newly created
[lru_]percpu_hash element. The area is recycled, so when the value
comes from a BPF program (onallcpus == false) it writes the running
CPU's slot and zeroes the rest.
bpf_percpu_hash_update() passes onallcpus == true, which delegates to
pcpu_copy_value(). pcpu_copy_value() writes only the CPU named in
map_flags when BPF_F_CPU is set, so on the create path the other slots
keep the recycled element's values:
update(k1, 0xdeadc0de, BPF_F_ALL_CPUS) every CPU holds 0xdeadc0de
delete(k1) element back on the freelist
update(k2, 0xc0ffee, BPF_F_CPU | 0) creates, writes CPU 0 only
lookup(k2) CPU 0 0xc0ffee, rest 0xdeadc0de
Zero-fill the other CPUs on that arm too.
Fixes: c6936161fd ("bpf: Add BPF_F_CPU and BPF_F_ALL_CPUS flags support for percpu_hash and lru_percpu_hash maps")
Signed-off-by: Donggeun Yoo <donggeunyoo.kernel@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260924102321.2120434-2-donggeunyoo.kernel@gmail.com
The 16/32-bit byteswap implementations for MIPS64r1 and earlier do
not have an explicit zero extension afterwards. The input is first
sign-extended to 64 bits, and the byteswap sequence can then leave
the result sign-extended depending on the value of the low bits.
Add the missing zero-extension.
Found with test_bpf on MIPS64r1 emulated by QEMU.
Fixes: fbc802de6b ("mips, bpf: Add new eBPF JIT for 64-bit MIPS")
Signed-off-by: Johan Almbladh <johan.almbladh@anyfinetworks.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260923105158.3514342-2-johan.almbladh@anyfinetworks.com
An addu instruction was emitted instead of addiu, causing the immediate
value 1 to be interpreted as register $at. This made the comparison
result invalid when the immediate operand was negative. Note that $at
is mapped to BPF_REG_AX, which is used for constant blinding.
Fix the instruction to use the immediate form.
Found with test_bpf on MIPS32r1 emulated by QEMU.
Fixes: eb63cfcd2e ("mips, bpf: Add eBPF JIT for 32-bit MIPS")
Signed-off-by: Johan Almbladh <johan.almbladh@anyfinetworks.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260923105158.3514342-1-johan.almbladh@anyfinetworks.com
__bpf_offload_dev_match() falls back to comparing offdev pointers after an
exact netdev mismatch. Bound-only programs normally have NULL offdevs, so
unrelated netdevs compare equal. A bound-only program on an
offload-registered netdev can instead inherit a real offdev and match a
sibling port. With CAP_BPF and CAP_NET_ADMIN, a caller can use
bpf(BPF_LINK_CREATE) with a different target ifindex to run metadata kfuncs
specialized for the bound driver on the target driver's xdp_buff. Running a
veth-bound program on tun reads beyond tun's bare stack xdp_buff as a
veth_xdp_buff.
Oops: general protection fault, probably for non-canonical address
KASAN: null-ptr-deref in range [0x0000000000000010-0x0000000000000017]
RIP: 0010:veth_xdp_rx_timestamp (drivers/net/veth.c:1673)
Call Trace:
...
tun_build_skb (drivers/net/tun.c:1739)
tun_get_user (drivers/net/tun.c:1856)
tun_chr_write_iter (drivers/net/tun.c:2091)
vfs_write (fs/read_write.c:595 fs/read_write.c:687)
ksys_write (fs/read_write.c:739)
do_syscall_64 (arch/x86/entry/syscall_64.c:84)
entry_SYSCALL_64_after_hwframe (arch/x86/entry/entry_64.S:121)
Kernel panic - not syncing: Fatal exception in interrupt
Restrict non-offloaded programs to exact netdev matches and retain the
shared-offdev fallback only for genuinely offloaded multi-port programs.
Fixes: 2b3486bc2d ("bpf: Introduce device-bound XDP programs")
Reported-by: <co+ac0a8c41de69121d@bugs.sh>
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20260917161335.1020405-2-bestswngs@gmail.com/
Link: https://patch.msgid.link/20260920132303.4109240-3-bestswngs@gmail.com
sock_map_alloc() only rejects max_entries == 0 and otherwise allows any
u32 value. sock_map_free() then walks the sks[] array with a signed int
iterator:
int i;
for (i = 0; i < stab->map.max_entries; i++)
struct sock **psk = &stab->sks[i];
When a SOCKMAP is created with max_entries = 0xffffffff (UINT_MAX), the
allocation of 32 GiB can succeed on large-memory hosts. During free the
counter reaches 0x80000000, wraps to INT_MIN, is sign-extended by movslq
and turned into a ~16 GiB negative offset from stab->sks, pointing far
below the allocation.
The faulting access is an xchg() write in sock_map_free(). Without
KASAN, the same out-of-bounds write can fault on an unmapped vmalloc page
or corrupt an unrelated allocation if that vmalloc address is populated.
On a KASAN kernel with CONFIG_KASAN_VMALLOC=y, the shadow check for that
address hits an unmapped shadow page and oopses first:
BUG: unable to handle page fault for address: fffff521b59c5a00
RIP: 0010:kasan_check_range+0x107/0x190
Call Trace:
sock_map_free+0x93/0x190
map_create+0x68d/0xb30
__sys_bpf+0x21e/0x2e70
Vmcore confirmed stab->map.max_entries == 0xffffffff, stab->sks ==
0xffffc911ace2d000, and the faulting address sks + (s64)INT_MIN * 8
exactly at 0xffffc90dace2d000. The same buggy path is reached on the
normal close()/bpf_map_free_deferred() path whenever such a map is
destroyed.
sock_map_alloc() used to bound its allocation size through
bpf_map_charge_init(), but the bound was dropped when rlimit-based memory
accounting was removed. Reject max_entries > INT_MAX at creation time so
the signed iterator in sock_map_free() never sees a value that would
overflow.
Triggered by syzkaller and reproduced on both a 6.6-based KASAN kernel
and the upstream v7.3-rc2 kernel.
Fixes: 0d2c4f9640 ("bpf: Eliminate rlimit-based memory accounting for sockmap and sockhash maps")
Signed-off-by: Zhao Gongyi <zhaogongyi@BYTEDANCE.COM>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260917121016.48171-1-zhaogongyi@bytedance.com
Add a selftest to confirm the verifier rejects ALU operations
that return arena or non-arena results depending on code path.
Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260922172028.6269-9-emil@etsalapatis.com
The verifier marks ALU instructions that include at least
one arena operand with needs_zext: These instructions are
fixed up after verification to be ALU32 instructions to
ensure that the result is a valid offset into an arena.
However, different code paths may provide two non-arena
64-bit arguments to the same instruction. The result of
the operation in that code path is wrong, since it is
now unexpectedly truncated to 32 bits and zero-extended.
Add logic to the verifier to ensure every instruction either
always has at least one PTR_TO_ARENA argument, or never does.
Since needs_zext already tracks the first scenario, add a
prevent_zext field in bpf_insn_aux to track the latter.
Reject instructions that use arena arguments and have prevent_zext
set, or do not have arena arguments and have needs_zext set.
Fixes: 6082b6c328 ("bpf: Recognize addr_space_cast instruction in the verifier.")
Reported-by: Nicholas Carlini <nicholas@carlini.com>
Suggested-by: Nicholas Carlini <nicholas@carlini.com>
Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260922172028.6269-8-emil@etsalapatis.com
Add a selftests that ensures that PTR_TO_PACKET arguments can
only be passed to subprogs that will never adjust the underlying
packet memory, and are rejected otherwise.
Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260922172028.6269-7-emil@etsalapatis.com
The verifier tracks changes in how PTR_TO_PACKET registers'
bounds are modified across subprog boundaries. PTR_TO_PACKET
registers are actually passed as PTR_TO_MEM, which is assumed
valid for the entire call. This is not the case with packet memory,
where a pskb_* call may invalidate its memory region.
Reject BPF code that passes PTR_TO_PACKET pointers to subprogs that
may mutate a packet. We cannot pass the pointer as a true PTR_TO_PACKET
because we would also need to somehow pass the PTR_TO_PACKET_META
or PTR_TO_PACKET_END to the subprog. Since we cannot avoid representing
the pointer in the subprog as PTR_TO_MEM, only permit it if the
subprog is guaranteed not to mutate the packet.
Fixes: 80f281664f5a ("bpf: Support pointers in global func args")
Reported-by: Nicholas Carlini <nicholas@carlini.com>
Suggested-by: Nicholas Carlini <nicholas@carlini.com>
Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260922172028.6269-6-emil@etsalapatis.com
Add tests to ensure the verifier properly tracks the 0 bit state
and width of the rx_queue_mapping field read from struct sock.
Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260922172028.6269-5-emil@etsalapatis.com
Currently, the ctx access code reads the rx_queue_mapping
field with either a 4-byte or 2-byte load. The rest of the bits
in the register are marked known zero by the verifier. However,
the emitted ctx access code places in the register on certain
the special value (-1) using BPF_MOV_IMM64, which gets sign-extended
to turn on all the bits in the register. By shifting this value right,
the program ends up with a value at runtime above what the verifier
assumes is possible.
Fix this by ensuring the read value is as wide as the assumed size.
Use MOV32 instructions instead of MOV64 instructions to keep
the upper bits zero as assumed by the verifier. Also properly report
the size of the destination variable (the bpf_sock field, 4 bytes) instead
of the source (the socket field, 2 bytes).
Fixes: c3c16f2ea6 ("bpf: Add rx_queue_mapping to bpf_sock")
Reported-by: Nicholas Carlini <nicholas@carlini.com>
Suggested-by: Nicholas Carlini <nicholas@carlini.com>
Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Link: https://patch.msgid.link/20260922172028.6269-4-emil@etsalapatis.com
Add a selftest to ensure dynptr slices cannot include
past the end of the linear area of an skb.
Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260922172028.6269-3-emil@etsalapatis.com
The skb_pointer_if_linear() function checks whether a
memory region of length len starting at offset off into
the skb is in the linear area, and returns a pointer to
the region if so. The check currently subtracts between
skb_headlen and offset of the check, and since skb_headlen
is unsigned the subtraction can underflow. This causes the
bounds check to spuriously pass and generate an arbitrary
pointer of the form *(skb->data + off).
The only user of this helper is currently skb-backed BPF
dynptr code. Returning the wrong pointer leads to the
dynptr erroneously being backed with invalid memory.
Ensure the subtraction cannot underflow, and fail the check if
it would. Use u64 arithmetic to also prevent overflow when
calculating (skb_headlen(skb) - off) since off is unsigned.
Fixes: 6f5a630d7c ("bpf, net: Introduce skb_pointer_if_linear().")
Reported-by: Nicholas Carlini <nicholas@carlini.com>
Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Link: https://patch.msgid.link/20260922172028.6269-2-emil@etsalapatis.com
Xu Yunxiang says:
====================
bpf: Reject non-negative stack object offsets
Reject non-negative offsets before converting a stack object address to a
stack slot index. Add a regression test for iterator destruction through
fp+0.
Changes in v2:
- Split the kernel change and selftest as requested by Andrii.
- Rebase onto bpf/master at a11212910c.
- Keep the original code and test changes unchanged and retain Sun Jian's
Reviewed-by on both parts.
v1: https://lore.kernel.org/bpf/20260911084314.3481637-1-xyx2021@mail.ustc.edu.cn/
Split request: https://lore.kernel.org/bpf/CAEf4BzbYzt-Riekpc-A=MfQft9ZELwSDSE8kkQO1ZOyR3O_FAg@mail.gmail.com/
Validation on this exact candidate with a matching bpf_testmod:
- W=1 verifier, full kernel/modules, changed BPF objects and test_progs
builds passed.
- iters: 1/97 passed; 0 skipped.
- dynptr: 2/132 passed; 0 skipped.
- irq: 1/33 passed; 0 skipped.
- res_spin_lock: 3/14 passed; 1 skipped.
- file_reader: 1/8 passed; 0 skipped.
- kmem_cache_iter: 1/3 passed; 0 skipped.
- dmabuf_iter: 1/4 passed; 0 skipped.
No selected test failed. The VM ran with panic_on_warn and panic_on_oops;
no kernel WARN, Oops or panic was found.
res_spin_lock_stress skips because the VM has no hardware PMU.
Annotated verifier tests check load outcomes and diagnostics. The full
unfiltered suite, sanitizer configurations and architecture matrix were
not run.
Please queue this fix for stable after it reaches the BPF tree.
====================
Link: https://patch.msgid.link/20260920210423.345636-1-xyx2021@mail.ustc.edu.cn
Signed-off-by: Andrii Nakryiko <andrii@kernel.org>
Add a verifier regression test that initializes a numeric iterator at fp-8
and attempts to destroy it through fp+0. The verifier must reject the
non-negative offset instead of treating it as the initialized stack slot.
Check the offset diagnostic to ensure rejection happens at the stack
object address check. The numeric iterator destroy operation is a no-op;
this test checks verifier rejection and does not run the program.
Signed-off-by: Xu Yunxiang <xyx2021@mail.ustc.edu.cn>
Signed-off-by: Andrii Nakryiko <andrii@kernel.org>
Reviewed-by: Sun Jian <sun.jian.kdev@gmail.com>
Link: https://lore.kernel.org/bpf/20260920210423.345636-3-xyx2021@mail.ustc.edu.cn
bpf_get_spi() computes (-off - 1) / BPF_REG_SIZE using C division,
which truncates toward zero. For off == 0, this produces spi 0, the
same index used by the valid stack slot at fp-8.
stack_slot_obj_get_spi() currently checks alignment and the resulting
spi bounds, but does not reject the non-negative offset itself. It can
therefore validate a PTR_TO_STACK register holding fp+0 against an
iterator stored at fp-8 even though the runtime receives the actual fp+0
pointer. An effectful iterator kfunc can then interpret memory outside
the BPF stack as iterator state.
Reject non-negative offsets before converting the offset to an spi. All
valid stack objects begin at a negative offset from the frame pointer.
Fixes: 06accc8779 ("bpf: add support for open-coded iterator loops")
Signed-off-by: Xu Yunxiang <xyx2021@mail.ustc.edu.cn>
Signed-off-by: Andrii Nakryiko <andrii@kernel.org>
Reviewed-by: Sun Jian <sun.jian.kdev@gmail.com>
Link: https://lore.kernel.org/bpf/20260920210423.345636-2-xyx2021@mail.ustc.edu.cn
bpf_crypto_ctx_create() is a kfunc whose second argument is declared
with the __sz annotation, so the verifier only guarantees that
params__sz bytes of params are valid. The function nevertheless reads
params->reserved[0] and params->reserved[1] (offsets 14 and 15) before
comparing params__sz against the size of struct bpf_crypto_params, so a
BPF program can pass a shorter buffer and have the kernel read past the
region that was validated for it.
Move the size check in front of the reserved field reads.
Fixes: 3e1c6f3540 ("bpf: make common crypto API for TC/XDP programs")
Reported-by: Vega <vega@nebusec.ai>
Signed-off-by: Yuqi Xu <xuyuqiabc@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Reviewed-by: Ren Wei <weir@nebusec.ai>
Link: https://patch.msgid.link/4f3ab4b03e79017e215521743996555439bf0bb3.1789802413.git.xuyuqiabc@gmail.com
Kumar Kartikeya Dwivedi says:
====================
Fix acyclic ownership checks
Bound and ensure acyclic ownership graphs for native data structures to
fix a bug reported by Nicholas. See commit logs and tests for details.
The existing list/rbtree rule already rejects graph-only cycles and bounds
those chains conservatively. It misses ownership through local referenced
kptrs, which can produce unbounded synchronous field destruction. Validate
all local ownership edges together, with an explicit depth bound, and allow
longer acyclic graph-only layouts within that bound.
Changelog:
----------
v1 -> v2
v1: https://lore.kernel.org/bpf/20260905090750.4064411-1-memxor@gmail.com/
* Fold the graph-walk and local-kptr changes into one complete fix. (Alexei)
* Explain why the original rule catches graph-only cycles, its three-type
chain bound, and the missing local-kptr ownership edges. (Alexei)
* Distinguish synchronous recursive field destruction from the deferred
RCU freeing of object storage.
* Add depth-boundary tests with child-first BTF ordering and a shared
suffix reached with different remaining budgets.
* Cover list and rbtree chains at the original three-type bound and at the
new eight-type bound, including rejected over-limit cases.
====================
Link: https://patch.msgid.link/20260914132444.2564218-1-memxor@gmail.com
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Build raw program BTF records to exercise local object ownership without
constructing a runtime chain deep enough to threaten the kernel stack.
Cover referenced-kptr and percpu-kptr self-cycles, a two-type cycle, and a
cycle mixing a graph root with a referenced kptr. The existing graph-only
check accepts these local-kptr cycles and over-limit kptr chains; the fix
rejects them with -ELOOP.
Pin the eight-record depth boundary with a terminal plain object. Exercise
both parent-first and child-first BTF orders, and a shared suffix reached
first through a shorter path. These cases require cached suffix depths to
be checked against the remaining depth budget on each path.
Check list and rbtree chains of three, four, eight, and nine types. This
covers the old graph-only depth boundary and the new explicit bound. The
existing linked-list BTF tests still reject pure graph cycles and now accept
the longer acyclic layouts previously rejected by the conservative rule.
Although bpf_percpu_obj_new() currently rejects types with special fields,
require the percpu cycle to fail at BTF load so future support cannot bypass
the ownership bound. Keep a positive control for a non-owning kptr, which
remains outside the ownership graph.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260914132444.2564218-3-memxor@gmail.com
Program-allocated objects can own other local objects through referenced
kptrs. bpf_obj_free_fields() follows those pointers through
__bpf_obj_drop_impl() synchronously, before the object storage is freed
through RCU. A self-referential local kptr type therefore permits arbitrarily
deep object chains, and dropping the head can exhaust the kernel stack.
Long acyclic type chains have the same problem.
btf_check_and_fixup_fields() still assumes referenced kptrs only point to
kernel types and checks ownership through list and rbtree roots only. Its
existing rule is sufficient for graph-only cycles: the target of each graph
edge must contain a node, so every type in a cycle has both a root and a
node. The rule rejects such a type owning another root, breaking every
cycle. It also limits graph-only chains to three types, or two if the first
type contains a node, and conservatively rejects longer acyclic chains.
The missing local-kptr edges, rather than a missed graph-only cycle, are the
bug introduced by support for bpf_kptr_xchg() into local kptrs.
Replace that restriction with one bounded ownership walk covering graph
roots and local referenced kptrs. Run it after all BTF records have been
fixed up, reject cycles and paths deeper than eight record-bearing types,
and cache each type's suffix depth while checking it against the remaining
budget. This also permits the longer acyclic graph-only layouts rejected
by the old rule; update their existing BTF tests accordingly.
Keep the bound independent of MAX_CALL_FRAMES because recursive destruction
can run below a BPF call chain. A plain local pointee without special-field
metadata adds only a final non-recursing drop. Non-owning kptrs and
kernel-BTF kptrs do not recurse through local records and remain outside the
walk. Include local percpu-kptr edges too, although allocation of percpu
objects with special fields is currently forbidden, so that relaxing that
restriction cannot bypass the ownership bound.
btf_check_and_fixup_fields() continues to initialize graph_root.value_rec,
including for separately allocated map records. The ownership relationships
belong to immutable program BTF and only need validation at BTF load time.
Fixes: b0966c7245 ("bpf: Support bpf_kptr_xchg into local kptr")
Reported-by: Nicholas Carlini <npc@anthropic.com>
Suggested-by: Nicholas Carlini <npc@anthropic.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260914132444.2564218-2-memxor@gmail.com
Kumar Kartikeya Dwivedi says:
====================
Compare stack frames in exact register states
regs_exact() compares register values and their ID relationships, but it
does not compare frameno. regsafe() checks frameno for ordinary
PTR_TO_STACK comparisons, while its EXACT path returns through regs_exact()
before reaching that check. Infinite-loop detection can therefore mistake
pointers to the same offset in different stack frames for the same pointer
and reject a finite loop.
Move frameno into bpf_reg_state's type-specific metadata union so the
existing regs_exact() prefix comparison covers it. This avoids a separate
PTR_TO_STACK case and keeps the structure at 80 bytes. Adjust the
states_maybe_looping() comparison boundary for the new layout. Since
frameno now aliases other pointer metadata, bpf_func() returns NULL for
registers that are not stack pointers; the callers that look up the frame
before checking the register type dereference it only afterwards.
The selftest keeps a stack pointer live in a register across a loop
whose only change at the header is the pointer's frame number. On the
unfixed tree, the program is rejected with "infinite loop detected". With
the fix, it loads and returns the expected value.
Changelog:
----------
v3 -> v4
v3: https://lore.kernel.org/bpf/20260919004327.1403382-1-memxor@gmail.com
* Return NULL from bpf_func() for non-stack registers, since frameno now
aliases other pointer metadata and some callers look up the frame before
checking the register type. (Sashiko)
v2 -> v3
v2: https://lore.kernel.org/bpf/20260918011313.3053497-1-memxor@gmail.com
* Rebase on bpf/master.
* Drop the redundant spilled-pointer test, since existing tests already
cover the stacksafe() -> regsafe() path. (Eduard)
* Place asm labels on their own line in the selftest. (Eduard)
* Collect Acked-by and Tested-by tags.
v1 -> v2
v1: https://lore.kernel.org/bpf/20260914161340.3419141-1-memxor@gmail.com
* Rebase on bpf/master.
* Move frameno into the type-specific metadata union so regs_exact()'s
existing prefix comparison covers it without growing bpf_reg_state.
====================
Link: https://patch.msgid.link/20260919014213.1840880-1-memxor@gmail.com
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Add a finite loop whose progress is represented only by changing the
frame number of a stack pointer. The loop first reads zero from the
caller's stack, switches to the same offset in the callee's stack, and
exits after reading one on its next iteration.
Force frequent checkpoints so the test exercises infinite-loop detection,
and check that the program returns one when run.
Without the frameno comparison in regs_exact(), the program is rejected
with an "infinite loop detected" diagnostic instead of loading
successfully.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Tested-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260919014213.1840880-3-memxor@gmail.com
regs_exact() compares the register state up to id, followed by the ID
mappings, but does not compare frameno. The PTR_TO_STACK case in regsafe()
checks frameno separately, which is bypassed when exact comparison is
requested. Consequently, infinite-loop detection can treat pointers to
different stack frames as the same pointer and reject a finite loop.
For example, initialize fp-8 to zero in the caller and to one in the
callee, then pass the caller's fp-8 to the callee as r1:
loop:
r0 = *(u64 *)(r1 + 0);
if r0 != 0 goto done;
r1 = r10;
r1 += -8;
goto loop;
done:
exit;
The loop terminates after reading the callee's slot on its second
iteration. At the loop header, however, the only relevant difference is
r1's frameno, so exact comparison incorrectly reports an infinite loop.
The same problem occurs when the pointer is spilled to the stack.
Move frameno into the type-specific metadata union, ahead of id, so the
existing prefix comparison in regs_exact() covers it. Ordinary stack
pointers do not use another union member. Iterator and IRQ stack-slot
states use their dedicated union views and do not need a frame lookup.
This also keeps bpf_reg_state at 80 bytes.
Since frameno now shares storage with other pointer metadata, it is only
meaningful for PTR_TO_STACK registers. Return NULL from bpf_func() for
other register types. process_iter_arg(), get_constant_map_key() and
is_dynptr_reg_valid_init() look up the frame before checking the register
type and would otherwise index frame[] with a byte of the register's map
or BTF pointer. They dereference the frame only after their type check.
Move the states_maybe_looping() boundary from frameno to precise after the
field relocation. Its prefix comparison continues to cover the complete
value state and now includes frameno.
Continue to ignore precise. Precision marks control whether pruning may
ignore scalar ranges; they do not change the represented values, and exact
comparison already compares those ranges unconditionally. Marks can also
change through backtracking while an ancestor state is still being
explored.
Fixes: d5b892fd60 ("bpf: make infinite loop detection in is_state_visited() exact")
Reported-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260919014213.1840880-2-memxor@gmail.com
CO-RE relocation of an ldimm64 instruction operates on two instruction
slots. A malformed BPF ELF can end a function after the first slot and
attach a CO-RE relocation to it. libbpf allocates the instruction array
according to the function symbol size, so the shared relocation code would
then access beyond the allocation.
Reject a terminal ldimm64 in libbpf's relocation loop, where the program
length is available, before resolving or applying the relocation. Both
resolved and unresolved relocations validate the absent second slot, and
unresolved relocation poisoning would additionally write past the array.
The in-kernel caller is protected by the verifier's early instruction-stream
check before it applies CO-RE relocations.
Fixes: eacaaed784 ("libbpf: Implement enum value-based CO-RE relocations")
Reported-by: Sashiko <sashiko-bot@kernel.org>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Link: https://lore.kernel.org/20260914140852.03DA21F0089B@smtp.kernel.org
Link: https://patch.msgid.link/20260917233222.2542500-11-memxor@gmail.com
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
Add a verifier test which retains a map value from an outer callback and
then acquires a lock through an inner callback value before attempting to
release the outer callback value. Both values can denote different elements,
so the verifier must reject the mismatched unlock.
Also exercise callbacks reached through two inner-map lookups. The lookup
results share inner_map_meta but may refer to different one-element arrays,
so their callback values must retain distinct lock identities.
Extend the existing spin_lock failure table and reuse its array and
inner-map fixtures to keep these cases alongside the other lock identity
tests. Update the nested callback reference-leak expectation for the extra
callback value ID.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Link: https://patch.msgid.link/20260917233222.2542500-10-memxor@gmail.com
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
A nested bpf_for_each_map_elem() callback can unlock a different element
of the same map:
static long inner(void *map, int *key, struct value *v,
struct value **outer_value)
{
bpf_spin_lock(&v->lock);
bpf_spin_unlock(&(*outer_value)->lock);
return 0;
}
static long outer(void *map, int *key, struct value *v, void *ctx)
{
bpf_for_each_map_elem(map, inner, &v, 0);
return 0;
}
Both callback values currently have ID zero and the same map_ptr.
process_spin_lock() compares those two fields, so it accepts the unlock
even though the two callbacks can receive different map elements.
Assign a fresh ID to every callback map value in the for-each,
timer/workqueue, and task-work constructors. Copies of one callback
argument retain its ID, so locking and unlocking through that argument
continues to work. Distinct callbacks also get distinct IDs for
single-element arrays, including inner arrays sharing inner_map_meta.
Preserve map_uid for every inner-map lookup and compare it through
check_ids() during state pruning. This preserves relationships between
maps, keys, and values while allowing equivalent states with different
lookup IDs to match. It avoids field-specific rules for when an inner map
needs an identity.
Move map_uid out of the metadata union and next to the other IDs, so
register comparisons can use the existing memcmp() ranges and remap the
IDs separately. Clear it when resetting a register or converting a map
lookup result to a socket pointer. Shrink frameno to u8, which is enough
for MAX_CALL_FRAMES, to make room without growing bpf_reg_state.
Fixes: d0d78c1df9 ("bpf: Allow locking bpf_spin_lock global variables")
Reported-by: Nicholas Carlini <npc@anthropic.com>
Suggested-by: Nicholas Carlini <npc@anthropic.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260917233222.2542500-9-memxor@gmail.com
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
Add raw CO-RE relocations that fail to resolve their target enum value.
Place each supported and unsupported instruction form in dead code.
Unsupported targets must fail relocation with a diagnostic even when they
are unreachable. Supported ALU immediates, memory accesses, and ldimm64
instructions must still be poisoned and removed as dead code, allowing the
program to load. Check that both halves of ldimm64 are poisoned.
Load every instruction stream without relocations first to ensure that
rejection is caused by the relocation rather than the original program.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260917233222.2542500-8-memxor@gmail.com
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
CO-RE relocation records can name any instruction offset. When a
relocation cannot be resolved, bpf_core_patch_insn() currently poisons its
target before checking whether that instruction is a valid relocation
target. Malformed metadata can therefore replace jumps, calls, exits,
register-source arithmetic, or non-immediate loads instead of failing at
the relocation step.
Handle poisoning only after the instruction has passed the same class and
operand-form checks used for a resolved relocation. Route invalid forms
through the existing diagnostic and return a hard error. Keep poisoning
supported instructions, including both halves of a plain ldimm64, so an
unresolved relocation in dead code remains valid.
Extend bpf_core_poison_insn() to poison both halves of ldimm64, and return
its status directly from each validated instruction case. This avoids
routing the success path through a common label and leaves the helper free
to report errors.
The shared relocation code applies this restriction to both libbpf and
in-kernel CO-RE.
Fixes: d7a252708d ("libbpf: Improve handling of failed CO-RE relocations")
Reported-by: Nicholas Carlini <npc@anthropic.com>
Suggested-by: Nicholas Carlini <npc@anthropic.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260917233222.2542500-7-memxor@gmail.com
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
Add a raw program load with CO-RE relocation metadata but no func_info or
line_info. Place the relocation in dead code and require the poisoning log,
proving that the kernel processes standalone CO-RE metadata instead of
silently skipping it.
Also give a subprogram a relocatable immediate as its terminal instruction.
Require the relocation's poisoning log before check_subprogs() rejects the
resulting fall-through. With the old ordering, check_subprogs() rejects the
original terminal instruction before CO-RE can emit the substitution log, so
the test continues to distinguish the ordering after relocation target
validation is tightened.
Submit a trailing ldimm64 first slot with CO-RE metadata and require the early
structural diagnostic. This exercises the check that protects relocation
processing instead of the later regular instruction validation.
Load the standalone instruction stream without relocation metadata first to
ensure that CO-RE processing causes its poisoning diagnostic. Encode the fixed
BTF metadata directly with the selftest BTF helpers.
Suggested-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260917233222.2542500-6-memxor@gmail.com
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
check_subprogs() verifies that each subprogram ends in an exit or an
unconditional jump before in-kernel CO-RE relocations are applied. An
unresolved relocation can then replace that terminal instruction with an
invalid helper call. The resulting fall-through into another subprogram
breaks the CFG invariant used by postorder and stack liveness analysis,
which can write past their per-subprogram arrays.
Apply CO-RE relocations immediately after preparing the program BTF, before
subprogram discovery and validation. Keep func_info and line_info validation
after subprogram discovery because those records depend on the complete
subprogram layout.
Reject an ldimm64 first slot at the end of the instruction stream before
CO-RE can inspect its missing second slot. check_subprogs() previously
rejected this form before relocation processing because it is not a valid
subprogram terminator. Moving CO-RE ahead of check_subprogs() removes that
implicit protection, so perform an explicit check before applying
relocations.
Include core_relo_cnt when deciding whether to prepare program BTF. A load
that supplied only CO-RE relocation metadata previously skipped both BTF
setup and relocation processing.
Fixes: fbd94c7afc ("bpf: Pass a set of bpf_core_relo-s to prog_load command.")
Suggested-by: Andrii Nakryiko <andrii@kernel.org>
Suggested-by: Alexei Starovoitov <ast@kernel.org>
Suggested-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260917233222.2542500-5-memxor@gmail.com
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
Add two paths whose packet pointer ranges are individually compatible at
a join but whose members have different relative displacements. The first
path proves an eight-byte access through one member. On the second path,
the same guard only proves that the access starts before data_end.
An affected verifier prunes the second path and accepts the program. With
packet pointer class displacement preserved, it explores that path and
rejects the out-of-bounds access.
Read the unknown offset and branch selector directly from XDP context
fields, and force state checkpoints so the pruning attempt does not
depend on the verifier checkpoint heuristics.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Link: https://patch.msgid.link/20260917233222.2542500-4-memxor@gmail.com
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
regsafe() maps packet pointer IDs between states and checks that each
current register range is a subset of the corresponding explored
register range. It does not, however, preserve the displacement between
registers that share a packet pointer ID.
This is unsound because packet range is shared by ID. A bounds check on
one class member updates every member, and a later access can consume the
range through another member. Commit 022ac07508 ("bpf: use reg->var_off
instead of reg->off for pointers") folded the fixed pointer offset into
r64 and removed the old off equality check, so two individually narrower
registers can prune even when their displacement has changed. The
explored path can then license an out-of-bounds packet access on the
pruned path.
Require matching range bases for packet pointers with an ID. Together
with the existing ID mapping, this preserves the displacement between
members of each packet-pointer class without adding per-ID state.
Packet pointers without an ID remain unaffected.
Fixes: 022ac07508 ("bpf: use reg->var_off instead of reg->off for pointers")
Reported-by: Nicholas Carlini <npc@anthropic.com>
Suggested-by: Nicholas Carlini <npc@anthropic.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260917233222.2542500-3-memxor@gmail.com
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
After do_check() returns, the verifier runs several instruction rewrite
passes. Some of them patch or remove one instruction at a time. Each
operation moves the remaining instruction and auxiliary-data arrays and
adjusts all branch offsets, making the overall work quadratic in the
program length.
A privileged loader can submit 131072 unconditional jumps by zero followed
by a valid return. Verification finishes quickly, but bpf_opt_remove_nops()
then spends a long time removing each jump separately. Since this
post-verification work neither checks for signals nor reschedules, a pending
SIGKILL cannot terminate the task until the rewrite finishes.
Make bpf_patch_insn_data() and verifier_remove_insns() common cancellation
and rescheduling points. These helpers run from BPF_PROG_LOAD process
context, and bpf_patch_insn_data() can already sleep while reallocating
auxiliary data.
Report interrupted constant blinding as -EINTR and propagate it through
both JIT paths, including kernels that permit interpreter fallback.
Other blinding failures retain the existing fallback behavior.
This does not reduce the quadratic cost of the rewrite passes, but it makes
the work preemptible and allows a killed loader to be torn down promptly.
Fixes: 52875a04f4 ("bpf: verifier: remove dead code")
Reported-by: Nicholas Carlini <npc@anthropic.com>
Suggested-by: Nicholas Carlini <npc@anthropic.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260917233222.2542500-2-memxor@gmail.com
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
bpf_link_prime() inserts a link into link_idr before anon_inode_getfile()
succeeds and before bpf_link_settle() publishes the ID in link->id.
bpf_link_by_id() treats such an ID-zero link as unsettled, but the link
iterator takes a reference without this check.
If anon_inode_getfile() then fails, the creator removes the ID and frees
its still-private link directly. The iterator is left with a dangling
reference and its next bpf_link_put() accesses freed memory.
Treat ID-zero entries as transient in bpf_link_get_curr_or_next(), just as
bpf_link_by_id() does.
BUG: KASAN: slab-use-after-free in bpf_link_put
Write of size 8 by task exp/384
Call Trace:
bpf_link_put kernel/bpf/syscall.c:3372
bpf_link_seq_next kernel/bpf/link_iter.c:33
bpf_seq_read kernel/bpf/bpf_iter.c:158
vfs_read fs/read_write.c:572
ksys_read fs/read_write.c:716
do_syscall_64 arch/x86/entry/syscall_64.c:84
entry_SYSCALL_64_after_hwframe arch/x86/entry/entry_64.S:121
Kernel panic - not syncing: KASAN: panic_on_warn set ...
Fixes: 9f88361273 ("bpf: Add bpf_link iterator")
Reported-by: Xiang Mei <xmei5@asu.edu>
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Signed-off-by: Andrii Nakryiko <andrii@kernel.org>
Link: https://lore.kernel.org/bpf/20260914170206.170723-2-bestswngs@gmail.com
xsk_map_gen_lookup() loads a u32 key and compares it with max_entries
using BPF_JMP_IMM. BPF immediates are sign-extended to 64 bits, so a
max_entries value of 0x80000000 or higher becomes a threshold larger
than every zero-extended 32-bit key. An out-of-range index then skips
the bounds check and the generated lookup reads past xsk_map[].
Compare with BPF_JMP32_IMM so the check stays in 32-bit unsigned range.
Fixes: e65650f291 ("bpf: Implement map_gen_lookup() callback for XSKMAP")
Reported-by: Vega <vega@nebusec.ai>
Signed-off-by: Zhiling Zou <zhilinz@nebusec.ai>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
Link: https://patch.msgid.link/7d2cb8e8dfaa9eb8fdff85156987a60960787dc3.1789056660.git.zhilinz@nebusec.ai
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
Kumar Kartikeya Dwivedi says:
====================
Fix global subprog verification context
Global subprogs can be invoked in both sleepable and non-sleepable
contexts, but are verified based on context assumptions for the program
type as a whole and not the precise context in which they are invoked.
Nicholas identified that this can violate verifier safety assumptions.
Verify global subprogs once for each observed sleepability context and add
workqueue callback regression tests.
Changelog:
----------
v2 -> v3
v2: https://lore.kernel.org/bpf/20260905051224.2325381-1-memxor@gmail.com
* Clarify that globals are checked only in contexts from which they are
reached, with two passes only when both contexts occur. (Alexei)
* Implement in_sleepable_context() as !in_rcu_cs(). (Alexei, Eduard)
* Keep the original subprogram traversal order and repeat until all
called contexts have been verified. (Eduard)
* Explain why instruction accounting needs the saved total and why the
exception callback uses the main program's context. (Eduard)
* Rename the do_check_common() parameter to is_sleepable. (BPF CI)
* Preserve the harmless global calls with __weak __noinline and check
their instruction totals across both contexts. Drop the redundant
task-work global RCU test and retain the review acknowledgment. (Eduard)
v1 -> v2
v1: https://lore.kernel.org/bpf/20260905034018.2095649-1-memxor@gmail.com
* Preserve the harmless global subprog calls in the positive test with
barrier() so LLVM cannot eliminate the intended coverage.
* Return zero explicitly from workqueue callbacks after global subprog calls.
* Document cumulative instruction accounting across verification contexts.
* Fix BPF multi-line comment formatting.
====================
Link: https://patch.msgid.link/20260914131923.2544250-1-memxor@gmail.com
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
Exercise global subprogram verification from workqueue callbacks, which
can run in a sleepable context even when the containing program is not
sleepable. An unprotected callback must not let the global subprogram use
implicit RCU protection inherited from the program.
Add a negative case which loads an RCU-protected task kptr in a global
subprogram reached from a workqueue callback. It fails on an unfixed
kernel because the program is incorrectly accepted. Also cover a
workqueue callback protected by an explicit RCU read-side critical
section, where the same global subprogram remains valid.
Call the same harmless global subprogram directly from the main program
and from an unprotected callback. Mark it __weak __noinline so both calls
survive optimization, and check that its instruction statistics account
for both verification contexts. This also verifies that global calls
from callbacks are not rejected wholesale. Workqueue callbacks return
zero explicitly after the global call, as required by their contract.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260914131923.2544250-3-memxor@gmail.com
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
Global subprograms are verified independently with a fresh verifier root.
do_check_common() currently seeds that root's in_sleepable state from the
program, even though a global subprogram can also run from callbacks whose
execution context differs from the program's main entry point.
In particular, workqueue and task-work callbacks are sleepable even when
the containing program is not. A global subprogram of that program is
therefore verified as non-sleepable, making in_rcu_cs() true and allowing
loads of RCU-protected kptrs to produce trusted MEM_RCU pointers. The same
subprogram can then be called from a sleepable callback without a classic
RCU reader. It can retain such a pointer while the object is freed and use
it after free.
The verifier's execution-context predicates are complementary. A state is
sleepable only when in_sleepable is set and no RCU, preemption, IRQ, or lock
region is active. Each condition which prevents sleeping also provides RCU
protection, while in_rcu_cs() treats a non-sleepable state as implicitly
protected.
Use this relationship to represent a global subprogram caller with only the
result of in_sleepable_context(). A protected sleepable caller is normalized
to in_sleepable=false at the independent verification root. This both
prevents sleepable operations and makes in_rcu_cs() true without copying
caller-owned lock state.
Track only the contexts in which each global subprogram is actually
reached. Verify it once if all reachable calls use the same context, and
twice only if both sleepable and non-sleepable calls reach it. Calls found
while verifying globals or asynchronous callbacks mark further contexts
for checking. Repeat the existing subprogram walk until all called
contexts have been verified; unreachable global calls remain unchecked.
Accumulate instruction counts over those verification passes. Preserve
the total recorded before each pass, since path accounting has already
added this pass's synchronous instructions and its root total must also
include asynchronous subprograms.
This makes an unprotected callback verify the global subprogram as
sleepable, turning its RCU-protected kptr load into an untrusted pointer.
Protected callers and global subprograms which do not depend on implicit RCU
protection remain valid.
Fixes: 81f1d7a583 ("bpf: wq: add bpf_wq_set_callback_impl")
Fixes: 38aa7003e3 ("bpf: task work scheduling kfuncs")
Reported-by: Nicholas Carlini <npc@anthropic.com>
Suggested-by: Nicholas Carlini <npc@anthropic.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260914131923.2544250-2-memxor@gmail.com
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
Donggeun Yoo says:
====================
bpf, arm64: fix the exception callback's frame pointer
The arm64 JIT does not set BPF_REG_FP in the prologue of an exception
callback, so the callback runs with whatever x25 held when bpf_throw()
was called. A callback that materializes the register, for instance to
pass the address of a local variable to a helper, then works on the
frame of the subprogram that threw.
Patch 1 sets ctx->fp_used on that path, the same fix commit b114fcee76
("bpf, arm64: Fix fp initialization for exception boundary") made for
the exception boundary. Patch 2 adds a selftest that reaches the case.
Tested on aarch64 under QEMU with vmtest.sh, on the base below. Without
patch 1 the new test panics the kernel, because the address handed to
the helper lands on the helper's own saved return address:
pc : 0x1234
lr : 0x1234
Call trace:
0x1234 (P)
bpf_test_run+0x188/0x3e0
bpf_prog_test_run_skb+0x47c/0x998
__sys_bpf+0xbdc/0xdd8
Kernel panic - not syncing: Oops: Fatal exception in interrupt
0x1234 is the value the callback reads, so the helper wrote it over its
own return address. With patch 1 applied the whole group passes:
#117/11 exceptions/exception_throw_subprog_stack_cb:OK
#117 exceptions:OK
Summary: 1/118 PASSED, 0 SKIPPED, 0/0 FAILED
Not tested on other architectures.
v1: https://lore.kernel.org/bpf/20260904070210.4163193-1-donggeunyoo.kernel@gmail.com/
v2: https://lore.kernel.org/bpf/20260907054235.473103-1-donggeunyoo.kernel@gmail.com/
v2 -> v3:
- patch 2: use the standard multi-line comment style
- patch 1: no change, added the Acked-by
Nothing compiled changed, so the numbers above are still v2's.
v1 -> v2:
- rebase onto bpf/master; CI could not apply v1
- spell the three new declarations u64 rather than __u64, to match the
rest of progs/exceptions.c
- no change to patch 1
====================
Link: https://patch.msgid.link/20260907130624.611942-1-donggeunyoo.kernel@gmail.com
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
The existing exception tests do not reach a callback that materializes
BPF_REG_FP into a register. They either throw from the main program,
where BPF_REG_FP already holds the value the callback needs, or use a
callback whose only stack accesses are frame pointer relative, which the
arm64 JIT rewrites to be stack pointer relative.
Add a test that throws from a subprogram using its own BPF stack, with a
callback that hands the address of a local variable to
bpf_probe_read_kernel(). The helper and the callback have to name the
same slot for the value read back to be the one the helper stored.
Signed-off-by: Donggeun Yoo <donggeunyoo.kernel@gmail.com>
Link: https://lore.kernel.org/r/20260907130624.611942-3-donggeunyoo.kernel@gmail.com
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
A program acting as exception boundary saves all callee-saved registers,
so build_prologue() takes the exception_cb path and never calls
push_callee_regs(). That is the only place find_used_callee_regs() runs,
and with it the only place ctx->fp_used is set, so the callback prologue
does not emit the
mov x25, sp
that points BPF_REG_FP at the frame the callback runs on. x25 keeps
whatever it held when bpf_throw() was called. If the throw came from a
subprogram that uses its own BPF stack, that is the subprogram's frame
pointer, and since the subprogram never returns it never restores x25
either.
Stack accesses through BPF_REG_FP are rewritten to be stack pointer
relative, so those still land in the callback's own frame. Materializing
the register does not: a callback that passes the address of a local
variable to a helper hands over an address in the dead subprogram's
frame. That address is below the callback's stack pointer by then, and
the helper's own call chain covers it, so the helper can write over its
own return address. 0x1234 below is the value the helper was asked to
store:
pc : 0x1234
lr : 0x1234
Call trace:
0x1234 (P)
bpf_test_run+0x188/0x3e0
bpf_prog_test_run_skb+0x47c/0x998
__sys_bpf+0xbdc/0xdd8
Kernel panic - not syncing: Oops: Fatal exception in interrupt
Set ctx->fp_used on the exception callback path so that the existing code
further down sets x25 from the stack pointer. The epilogue restores it
from the main program's save area along with the other callee-saved
registers, as it already does. x86 sets the frame pointer for the
callback from the argument it is passed, and powerpc computes it from
the stack pointer.
Fixes: 5d4fa9ec56 ("bpf, arm64: Avoid blindly saving/restoring all callee-saved registers")
Acked-by: Xu Kuohai <xukuohai@huawei.com>
Signed-off-by: Donggeun Yoo <donggeunyoo.kernel@gmail.com>
Link: https://lore.kernel.org/r/20260907130624.611942-2-donggeunyoo.kernel@gmail.com
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Several map batch operation implementations such as
generic_map_lookup_batch() use calculations in the form of
"values + cp * map->value_size" to compute the desired userspace memory
address for reading or writing. This can overflow the u32 type
(the result of "cp * map->value_size") when the map size exceeds 4GB.
generic_map_lookup_batch() may corrupt values for some keys in
userspace memory, and in some cases it mismatches values for some keys
while still reporting success.
Other batch operations may fail to delete or update some keys,
or the syscall may return unexpected errors.
Add size_t casts to prevent the affected offset and size calculations
from overflowing.
Fixes: cb4d03ab49 ("bpf: Add generic support for lookup batch op")
Fixes: aa2e93b8e5 ("bpf: Add generic support for update and delete batch ops")
Fixes: 057996380a ("bpf: Add batch ops to all htab bpf map")
Signed-off-by: Masoud Aghasi <maghasi@disroot.org>
Link: https://lore.kernel.org/r/20260903082734.623904-1-maghasi@disroot.org
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
When a BPF stream_verdict program redirects an skb back to the same
socket (self-redirect with BPF_F_INGRESS), sk_psock_verdict_apply()
calls tcp_eat_skb() which advances tcp_sk->copied_seq. However, the
skb is then delivered to the socket's psock ingress queue and later
read by tcp_bpf_recvmsg_parser(), which also advances copied_seq via
the copied_from_self accounting path. This double-counting causes
copied_seq to advance by 2x the actual data length, triggering:
TCP recvmsg seq # bug 2: copied BF2E806, seq BF2E7FD, \
rcvnxt BF2E806, fl 0
WARNING: net/ipv4/tcp.c:2745 at tcp_recvmsg_locked+0x72b/0x2640
Call Trace:
tcp_recvmsg+0x10a/0x500
sock_recvmsg+0x168/0x1d0
__sys_recvfrom+0x19a/0x2a0
__x64_sys_recvfrom+0xe4/0x1f0
do_syscall_64+0xf7/0x530
entry_SYSCALL_64_after_hwframe+0x77/0x7f
cleanup rbuf bug: copied BF2E806 seq BF2E806 rcvnxt BF2E806
WARNING: net/ipv4/tcp.c:1609 at tcp_cleanup_rbuf+0xf2/0x1c0
Call Trace:
tcp_recvmsg_locked+0x8d1/0x2640
tcp_recvmsg+0x10a/0x500
sock_recvmsg+0x168/0x1d0
__sys_recvfrom+0x19a/0x2a0
__x64_sys_recvfrom+0xe4/0x1f0
do_syscall_64+0xf7/0x530
entry_SYSCALL_64_after_hwframe+0x77/0x7f
Fix this by converting self-redirect verdict to __SK_PASS at the
beginning of sk_psock_verdict_apply(). This bypasses the
__SK_REDIRECT case entirely (which calls sk_psock_eat_skb), letting
the __SK_PASS path queue the skb to the psock ingress queue. The
data is then read via tcp_bpf_recvmsg_parser(), which advances
copied_seq exactly once through copied_from_self. Cross-socket
redirects continue through __SK_REDIRECT with sk_psock_eat_skb()
unchanged.
Fixes: e5c6de5fa0 ("bpf, sockmap: Incorrectly handling copied_seq")
Suggested-by: Jakub Sitnicki <jakub@cloudflare.com>
Suggested-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Signed-off-by: Geliang Tang <tanggeliang@kylinos.cn>
Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Link: https://lore.kernel.org/r/1a8e797a1b26e2f695aaac22ac644c2862f63466.1788858299.git.tanggeliang@kylinos.cn
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Syzkaller repeatedly triggered UAF splats related to nodes in
waiting_for_gp_ttrace within the bpf memalloc:
BUG: KASAN: slab-use-after-free in llist_del_first+0x85/0x110 lib/llist.c:61
Read of size 8 at addr ffff8881572cd080 by task syz.4.470/5112
...
llist_del_first+0x85/0x110 lib/llist.c:61
alloc_bulk+0x193/0x460 kernel/bpf/memalloc.c:229
bpf_mem_refill+0x386/0x560 kernel/bpf/memalloc.c:436
Freed by task 14:
...
__free_rcu kernel/bpf/memalloc.c:281 [inline]
__free_rcu_tasks_trace+0x48/0xd0 kernel/bpf/memalloc.c:291
rcu_tasks_invoke_cbs+0x1ec/0x3e0 kernel/rcu/tasks.h:571
rcu_tasks_one_gp+0x13d/0x220 kernel/rcu/tasks.h:621
rcu_tasks_kthread+0xf3/0x120 kernel/rcu/tasks.h:651
The reason is that the UAF occurs after the RCU Tasks Trace GP expires:
when the __free_rcu() callback runs, there is no synchronization
protecting llist_del_all() against concurrent alloc_bulk() operating on
waiting_for_gp_ttrace, leading to the race condition below:
CPU0 CPU1
__free_rcu (RCU Tasks Trace callback)
alloc_bulk
llist_del_first(&c->waiting_for_gp_ttrace)
entry = smp_load_acquire(&head->first);
do {
if (entry == NULL)
return NULL;
free_all(llist_del_all(&c->waiting_for_gp_ttrace))
llist_for_each_safe(pos, t, llnode)
free_one(pos);
next = READ_ONCE(entry->next); <-- trigger UAF
} while (!try_cmpxchg(&head->first, &entry, next));
In addition, there is also a theoretical race condition on the
free_by_rcu_ttrace list. This race requires two preconditions: an
in-flight Tasks Trace GP keeping c->call_rcu_ttrace_in_progress == 1,
and concurrent cross-CPU frees repopulating c->free_by_rcu_ttrace with
new nodes. Under these conditions, the following scenario triggers UAF:
// CPU0
// irq work is still busy (on PREEMPT_RT)
alloc_bulk()
llist_del_first(&c->free_by_rcu_ttrace)
entry = smp_load_acquire(&head->first);
do {
if (entry == NULL)
return NULL;
// CPU1
bpf_mem_alloc_destroy()
WRITE_ONCE(c->draining, true)
// wait for CPU0
irq_work_sync()
// CPU2
do_call_rcu_ttrace(tgt(CPU0))
if (c->draining) {
llist_del_all(&c->free_by_rcu_ttrace)
free_all()
}
// CPU0 continue
next = READ_ONCE(entry->next); <-- trigger UAF
while (!try_cmpxchg(&head->first, &entry, next));
Fix this by introducing a raw spinlock to synchronize the concurrent
consumption on waiting_for_gp_ttrace and free_by_rcu_ttrace.
Fixes: 04fabf00b4 ("bpf: Allow reuse from waiting_for_gp_ttrace list.")
Suggested-by: Alexei Starovoitov <ast@kernel.org>
Suggested-by: Hou Tao <houtao1@huawei.com>
Signed-off-by: Pu Lehui <pulehui@huawei.com>
Acked-by: Hou Tao <houtao1@huawei.com>
Link: https://lore.kernel.org/r/20260905021139.4116529-1-pulehui@huaweicloud.com
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
xdp_test_run_setup() allocates xdp->frames and xdp->skbs with
kvmalloc_array(). The setup error path already releases both arrays with
kvfree(), while the normal teardown path still uses kfree().
Use kvfree() in xdp_test_run_teardown() as well, so the release helper
matches the allocator on both paths.
Signed-off-by: Zhixing Chen <running910@gmail.com>
Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
Link: https://lore.kernel.org/r/20260903104358.29228-1-running910@gmail.com
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
A sockops prog reading skops->rtt_min never checks the sk type: on the
tcp_conn_request() path sock_ops->sk is a request_sock (non-full), and the
ctx rewrite casts it to a tcp_sock (full) and reads rtt_min past the end of
the request_sock, returning dirty adjacent memory.
SEC("sockops")
int prog(struct bpf_sock_ops *skops)
{
switch (skops->op) {
case BPF_SOCK_OPS_RWND_INIT:
leak = skops->rtt_min; /* reads the request_sock OOB */
...
}
}
For instance one such read returned rtt_min=0xffff8881, the high half of a
leaked kernel pointer.
Guarding that cast is exactly what SOCK_OPS_GET_FIELD() does -- it checks
is_locked_tcp_sock and returns 0 when sock_ops->sk is not a locked full
socket. Every other tcp_sock field in sock_ops goes through it; rtt_min is
the only one open-coded, so it skips the check.
Read rtt_min through SOCK_OPS_GET_FIELD() too. rtt_min is a bit special:
it is a struct minmax and we only want the current min, so pass
rtt_min.s[0].v. That is equivalent to the old hand-computed offset
offsetof(struct tcp_sock, rtt_min) + sizeof_field(struct minmax_sample, t)
(s[0] sits at rtt_min + 0 and .v at + sizeof(.t), i.e. what minmax_get()
returns), so the loaded field is unchanged and only the full-sock guard is
added. The two BUILD_BUG_ON()s that protected the hand-computed offset
are no longer needed.
Before patch:
0: r1 = *(u64 *)(r1 +0) ; r1 = skops->sk
1: r1 = *(u32 *)(r1 +2324) ; ((tcp_sock *)sk)->rtt_min.s[0].v
After patch:
0: *(u64 *)(r1 +56) = r9
1: r9 = *(u8 *)(r1 +50) ; is_locked_tcp_sock
2: if r9 == 0 goto pc+4 ; not a locked full sock -> 0
3: r9 = *(u64 *)(r1 +56)
4: r1 = *(u64 *)(r1 +0) ; r1 = skops->sk
5: r1 = *(u32 *)(r1 +2324) ; rtt_min.s[0].v
6: goto pc+2
7: r9 = *(u64 *)(r1 +56)
8: r1 = 0
Fixes: 44f0e43037 ("bpf: Add support for reading sk_state and more")
Reported-by: VEGA <vega@nebusec.ai>
Signed-off-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
Link: https://lore.kernel.org/r/20260903100921.113374-1-jiayuan.chen@linux.dev
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
__htab_map_lookup_and_delete_batch() has no rescheduling point. The
batch count bounds how many entries are copied out, not how many
buckets are visited, so one BPF_MAP_LOOKUP_BATCH call can walk the
map end to end. The empty-bucket fast path is worse: it stays inside
a single rcu_read_lock() / bpf_disable_instrumentation() section for
any run of consecutive empty buckets.
That holds up on small maps, but it falls apart at scale. On a
144-CPU arm64 host running a CONFIG_PREEMPT_NONE kernel, periodic
BPF_MAP_LOOKUP_BATCH calls against an LRU hash map with 16,777,216
buckets held a CPU inside the batch op for 77+ seconds and triggered
the soft lockup watchdog.
Commit 75134f16e7 ("bpf: Add schedule points in batch ops") fixed this
same problem in the generic batch ops, but not in this htab-native path,
which every htab-based hash map variant uses for its lookup[_and_delete]
batch ops.
Complete that fix here. Leave the critical section after 64 consecutive
empty buckets, call cond_resched_tasks_rcu_qs(), and resume at the saved
bucket cursor. No locks are held at that point, and resuming from the
cursor is already the function's behavior for non-empty buckets. Add the
same call to the per-bucket loop after copy_to_user(), where every lock
has been dropped. cond_resched_rcu() is not enough here: sleeping with
bpf_prog_active elevated makes tracing programs on that CPU silently
skip their invocations.
Plain cond_resched() is not enough either. It is a no-op under PREEMPT
and PREEMPT_LAZY, the only models arm64 and x86 have offered since
commit 7dadeaa6e8 ("sched: Further restrict the preemption modes").
It is also never a Tasks RCU quiescent state, in any model: the
reschedule counts as a preemption. The walking task stays a holdout and
stalls every synchronize_rcu_tasks() caller, ftrace and BPF trampoline
teardown included, until the syscall returns [1].
cond_resched_tasks_rcu_qs() is the usual tool for that [2]. It reports
the quiescent state at each yield and still reschedules as
cond_resched() does on PREEMPT_NONE and PREEMPT_VOLUNTARY kernels.
Fixes: 057996380a ("bpf: Add batch ops to all htab bpf map")
Cc: "Paul E. McKenney" <paulmck@kernel.org>
Cc: Rik van Riel <riel@surriel.com>
Link: https://lore.kernel.org/bpf/20260715215314.44423f47@fangorn/ [1]
Link: https://lore.kernel.org/bpf/9d444098-7c03-4163-af12-bd0a79a51443@paulmck-laptop/ [2]
Assisted-by: LLM
Signed-off-by: Jose Fernandez (Anthropic) <jose.fernandez@linux.dev>
Signed-off-by: Josef Bacik <josef@toxicpanda.com>
Reviewed-by: Rik van Riel <riel@surriel.com>
Link: https://lore.kernel.org/r/20260909-b4-htab-batch-resched-v2-1-0cb529d8f95a@toxicpanda.com
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Add tests that place gotox at the end of the main program and a
subprogram, with each jump-table target preceding the gotox instruction.
This tests gotox as a valid non-fallthrough terminal instruction.
Signed-off-by: Siddharth Chintamaneni <sidchintamaneni@gmail.com>
Reviewed-by: Anton Protopopov <a.s.protopopov@gmail.com>
Link: https://lore.kernel.org/r/20260902171414.96165-2-sidchintamaneni@gmail.com
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
check_subprogs() treats gotox as a direct jump and validates its reserved
zero offset. When gotox is the final instruction, this produces a
synthetic successor one instruction past the end of the subprogram and
rejects an otherwise valid program.
Skip direct-offset validation for gotox and accept it as a
non-fallthrough terminal instruction. Its actual targets remain validated
from the instruction-array jump table during CFG construction.
Fixes: 493d9e0d60 ("bpf, x86: add support for indirect jumps")
Signed-off-by: Siddharth Chintamaneni <sidchintamaneni@gmail.com>
Reviewed-by: Anton Protopopov <a.s.protopopov@gmail.com>
Link: https://lore.kernel.org/r/20260902171414.96165-1-sidchintamaneni@gmail.com
Signed-off-by: Alexei Starovoitov <ast@kernel.org>