Merge branch 'bpf-introduce-global-percpu-data'

Leon Hwang says:

====================
bpf: Introduce global percpu data

This patch set introduces global percpu data, similar to commit
6316f78306 ("Merge branch 'support-global-data'"), to reduce restrictions
in C for BPF programs.

With this enhancement, it becomes possible to define and use global percpu
variables, like the DEFINE_PER_CPU() macro in the kernel
include/linux/percpu-defs.h.

The section name for global peurcpu data is ".percpu". Even though, a one-byte
percpu variable (e.g., char run SEC(".percpu") = 0;) can trigger a crash
with Clang 17 [1], users are expected to use such small variables as global
percpu data with newer Clang versions, which don't have the issue.

The idea stems from the bpfsnoop [2], which itself was inspired by
retsnoop [3]. During testing of bpfsnoop on the v6.6 kernel, two LBR
(Last Branch Record) entries were observed related to the
bpf_get_smp_processor_id() helper.

Since commit 1ae6921009 ("bpf: inline bpf_get_smp_processor_id() helper"),
the bpf_get_smp_processor_id() helper has been inlined on x86_64, reducing
the overhead and consequently minimizing these two LBR records.

However, the introduction of global percpu data offers a more robust
solution. By leveraging the percpu_array map and percpu instruction,
global percpu data can be implemented intrinsically.

This feature also facilitates sharing percpu information between tail
callers and callees or between freplace callers and callees through a
shared global percpu variable. Previously, this was achieved using a
1-entry percpu_array map, which this patch set aims to improve upon.

Links:
[1] https://lore.kernel.org/bpf/fd1b3f58-c27f-403d-ad99-644b7d06ecb3@linux.dev/
[2] https://github.com/bpfsnoop/bpfsnoop
[3] https://github.com/anakryiko/retsnoop

Changes:
v11 -> v12:
* Improve feature check in bpf_object__create_maps() in libbpf.
* Add percpu_array map support in bpf_map__set_value_size() in libbpf.
* Exercise bpf_map__set_value_size() in selftest.
* Drop dead warning in bpf_object__populate_internal_map() in libbpf.
  (Sashiko)
* v11: https://lore.kernel.org/bpf/20260806163125.11172-1-leon.hwang@linux.dev/

v10 -> v11:
* Drop env->prog->jit_requested check when inlining insns for global
  percpu data.
* Do not autocreate percpu_array map when kernel does not have global
  percpu data support in libbpf.
* Check map->btf_value_type_id in bpftool's is_skel_data().
* Exercise bpf_map__lookup_elem() in selftest.
* Collect Reviewed-by tags from Emil, thanks.
* Drop all duplicate blank lines in kernel/bpf/*.c. (Emil)
* Factor out check_map_mem_read() helper. (Emil)
* Check bpf_jit_supports_percpu_insn() first in
  percpu_array_map_direct_value_addr/meta(). (Emil)
* Add comment for 'map->libbpf_type == LIBBPF_MAP_PERCPU' in libbpf's
  map_is_mmapable(). (Emil)
* Init update_flags as a const var in libbpf's
  bpf_object__populate_internal_map(). (Emil)
* Keep is_mmapable_map() beyond is_skel_data() in bpftool. (Emil)
* Add 'run' and 'cpu_id' in selftest. (Emil)
* Drop subskel test. Verify the generated subskel manually. (Emil)
* Add comment to the raw insns in selftest. (Emil)
* v10: https://lore.kernel.org/bpf/20260715153254.92010-1-leon.hwang@linux.dev/

v9 -> v10:
* Rebase latest bpf-next tree to resolve code conflict in verifier in
  patch #1.
* v9: https://lore.kernel.org/bpf/20260713154024.30851-1-leon.hwang@linux.dev/

v8 -> v9:
* Use real name for percpu data maps in libbpf in patch #4.
* Add long map name test in patch #6.
* Move parse_cpu_mask_file() to test_percpu_data_on_cpus() in test in
  patch #6.
* Validate map type in get_map_ident() for percpu data maps in patch #5.
* Update code comment in verifier in patch #2. (per Andrii)
* Pass 'type' to internal_map_name in libbpf in patch #4. (per Andrii)
* Factor out the helper is_skel_data() in bpftool in patch #5.
  (per Quentin and Andrii)
* v8: https://lore.kernel.org/bpf/20260629152406.52582-1-leon.hwang@linux.dev/

v7 -> v8:
* Send patch #1 and #2 separately that fix interpreter fallback issues.
  (Andrii)
* Use 'array->elem_size' to avoid 'range' local variable in
  percpu_array_map_direct_value_meta(). (Andrii)
* Keep original map name for percpu data's map in libbpf. (Andrii)
* Factor out helper bpf_map_is_skel_data() in bpftool. (Andrii)
* Update commit message of direct access read-only percpu_array map.
  (Andrii)
* Add test to verify that it is disallowed to directly write data of
  read-only percpu_array map. (Andrii)
* Drop unused 'num_cpus' in test. (bot+bpf-ci)
* Factor out helper test_percpu_data_on_cpus() in test. (bot+bpf-ci)
* v7: https://lore.kernel.org/bpf/20260622143557.22955-1-leon.hwang@linux.dev/

v6 -> v7:
* Use tgt_endian() in bpf_gen__map_update_elem() in patch #6. (Sashiko)
* Use sizeof(args) in verifier_snprintf test in patch #10. (Sashiko)
* Drop xlated test of v6. (Alexei)
* v6: https://lore.kernel.org/bpf/20260615152646.27639-1-leon.hwang@linux.dev/

v5 -> v6:
* Prevent running user addr_space_cast and addr_percpu insns in
  interpreter. (Sashiko)
* Cast __percpu pointer to u64 with (__force unsigned long). (lkp)
* Exclude BPF_MAP_TYPE_PERCPU_ARRAY in check_mem_access() before calling
  bpf_map_direct_read(), and add a test to verify it.
  (Sashiko, bot+bpf-ci)
* Skip percpu data variables for subskeleton in bpftool. (Sashiko)
* Protect skel->percpu using mprotect(..., PROT_READ) in light skeleton.
  (Sashiko, bot+bpf-ci)
* Drop roundup() in tests. (Sashiko)
* Call test_global_percpu_data_verifier_log() without
  test__start_subtest(). (Sashiko)
* Cast insn->imm to __u64 with (__u32) in xlated test. (Sashiko)
* Check cnt using the new idx in xlated test. (Sashiko)
* v5: https://lore.kernel.org/bpf/20260608145113.65857-1-leon.hwang@linux.dev/

v4 -> v5:
* Add prog->jit_requested check to prevent running percpu data in
  interpreter in patch #1.
* Factor out verifier log tests using its own patch.
* Address comments from Alexei:
  * Move map_type check from check_mem_access() to bpf_map_direct_read()
    in patch #2.
  * Move BPF_MAP_TYPE_INSN_ARRAY map_type check from const_reg_xfer() to
    bpf_map_direct_read() in patch #2.
  * Add a test to verify that the off of xlated ldimm64 insn matches the
    off encoded in the ELF ldimm64 insn.
  * Drop patch #5 of v4.
* Address reviews from Sashiko:
  * Update commit message of patch #6 to indicate that maps.percpu->mmaped
    has been marked as read-only in libbpf.
  * Lookup elem on specified CPU using BPF_F_CPU in tests.
  * Drop unnecessary err == -EOPNOTSUPP in test.
  * Locate target field using its offset in the iter test.
* v4: https://lore.kernel.org/bpf/20260414132421.63409-1-leon.hwang@linux.dev/

v3 -> v4:
* Drop duplicate blank lines in verifier.
* Add percpu data feature probe in libbpf.
* Update percpu_array map using BPF_F_ALL_CPUS flag for lskel, if no cpu flag
  is set.
* Add two tests to verify verifier log.
* Add a test to verify mov64_percpu_reg instruction.
* Add a test to verify bpf_iter for percpu data map.
* Update percpu_array map using BPF_F_ALL_CPUS flag in libbpf
  (per Alexei and Andrii).
* Address comments from Andrii:
  * Use .percpu as section identifier.
  * Use bpf_jit_supports_percpu_insn() instead of CONFIG_SMP.
  * Drop bpf_map__is_internal_percpu() API.
  * Drop unnecessary __aligned(8) in libbpf, verified by selftest.
  * Make mmap data read-only after loading prog.
v3: https://lore.kernel.org/bpf/20250526162146.24429-1-leon.hwang@linux.dev/

v2 -> v3:
  * Use ".data..percpu" as PERCPU_DATA_SEC.
  * Address comment from Alexei:
    * Add u8, array of ints and struct { .. } vars to selftest.
v2: https://lore.kernel.org/bpf/20250213161931.46399-1-leon.hwang@linux.dev/

v1 -> v2:
  * Address comments from Andrii:
    * Use LIBBPF_MAP_PERCPU and SEC_PERCPU.
    * Reuse mmaped of libbpf's struct bpf_map for .percpu map data.
    * Set .percpu struct pointer to NULL after loading skeleton.
    * Make sure value size of .percpu map is __aligned(8).
    * Use raw_tp and opts.cpu to test global percpu variables on all CPUs.
  * Address comments from Alexei:
    * Test non-zero offset of global percpu variable.
    * Test case about BPF_PSEUDO_MAP_IDX_VALUE.
v1: https://lore.kernel.org/bpf/20250127162158.84906-1-leon.hwang@linux.dev/

rfc -> v1:
  * Address comments from Andrii:
    * Keep one image of global percpu variable for all CPUs.
    * Reject non-ARRAY map in bpf_map_direct_read(), check_reg_const_str(),
      and check_bpf_snprintf_call() in verifier.
    * Split out libbpf changes from kernel-side changes.
    * Use ".percpu" as PERCPU_DATA_SEC.
    * Use enum libbpf_map_type to distinguish BSS, DATA, RODATA and
      PERCPU_DATA.
    * Avoid using errno for checking err from libbpf_num_possible_cpus().
    * Use "map '%s': " prefix for error message.
rfc: https://lore.kernel.org/bpf/20250113152437.67196-1-leon.hwang@linux.dev/
====================

Link: https://patch.msgid.link/20260813152324.97937-1-leon.hwang@linux.dev
Signed-off-by: Andrii Nakryiko <andrii@kernel.org>
This commit is contained in:
Andrii Nakryiko 2026-08-13 10:27:41 -07:00
commit 806c1a1852
22 changed files with 725 additions and 95 deletions

View File

@ -259,6 +259,37 @@ static void *percpu_array_map_lookup_elem(struct bpf_map *map, void *key)
return this_cpu_ptr(array->pptrs[index & array->index_mask]);
}
static int percpu_array_map_direct_value_addr(const struct bpf_map *map, u64 *imm, u32 off)
{
struct bpf_array *array = container_of(map, struct bpf_array, map);
if (!bpf_jit_supports_percpu_insn())
return -EOPNOTSUPP;
if (map->max_entries != 1)
return -EOPNOTSUPP;
if (off >= map->value_size)
return -EINVAL;
*imm = (u64)(__force unsigned long) array->pptrs[0];
return 0;
}
static int percpu_array_map_direct_value_meta(const struct bpf_map *map, u64 imm, u32 *off)
{
struct bpf_array *array = container_of(map, struct bpf_array, map);
u64 base = (u64)(__force unsigned long) array->pptrs[0];
if (!bpf_jit_supports_percpu_insn())
return -EOPNOTSUPP;
if (map->max_entries != 1)
return -EOPNOTSUPP;
if (imm < base || imm >= base + array->elem_size)
return -ENOENT;
*off = imm - base;
return 0;
}
/* emit BPF instructions equivalent to C code of percpu_array_map_lookup_elem() */
static int percpu_array_map_gen_lookup(struct bpf_map *map, struct bpf_insn *insn_buf)
{
@ -551,9 +582,10 @@ static int array_map_check_btf(struct bpf_map *map,
const struct btf_type *key_type,
const struct btf_type *value_type)
{
/* One exception for keyless BTF: .bss/.data/.rodata map */
/* One exception for keyless BTF: .bss/.data/.rodata/.percpu map */
if (btf_type_is_void(key_type)) {
if (map->map_type != BPF_MAP_TYPE_ARRAY ||
if ((map->map_type != BPF_MAP_TYPE_ARRAY &&
map->map_type != BPF_MAP_TYPE_PERCPU_ARRAY) ||
map->max_entries != 1)
return -EINVAL;
@ -832,6 +864,8 @@ const struct bpf_map_ops percpu_array_map_ops = {
.map_get_next_key = bpf_array_get_next_key,
.map_lookup_elem = percpu_array_map_lookup_elem,
.map_gen_lookup = percpu_array_map_gen_lookup,
.map_direct_value_addr = percpu_array_map_direct_value_addr,
.map_direct_value_meta = percpu_array_map_direct_value_meta,
.map_update_elem = array_map_update_elem,
.map_delete_elem = array_map_delete_elem,
.map_lookup_percpu_elem = percpu_array_map_lookup_percpu_elem,

View File

@ -214,7 +214,6 @@ static inline bool bt_is_reg_set(struct backtrack_state *bt, u32 reg)
return bt->reg_masks[bt->frame] & (1 << reg);
}
/* format registers bitmask, e.g., "r0,r2,r4" for 0x15 mask */
static void fmt_reg_mask(char *buf, ssize_t buf_sz, u32 reg_mask)
{
@ -254,7 +253,6 @@ void bpf_fmt_stack_mask(char *buf, ssize_t buf_sz, u64 stack_mask)
}
}
/* For given verifier state backtrack_insn() is called from the last insn to
* the first insn. Its purpose is to compute a bitmask of registers and
* stack slots that needs precision in the parent verifier state.

View File

@ -2534,7 +2534,6 @@ static void btf_bitfield_show(void *data, u8 bits_offset,
btf_int128_print(show, print_num);
}
static void btf_int_bits_show(const struct btf *btf,
const struct btf_type *t,
void *data, u8 bits_offset,

View File

@ -47,7 +47,6 @@ enum {
BRANCH = 2,
};
static void mark_subprog_changes_pkt_data(struct bpf_verifier_env *env, int off)
{
struct bpf_subprog_info *subprog;

View File

@ -182,7 +182,6 @@ static void const_reg_xfer(struct bpf_verifier_env *env, struct const_arg_info *
u64 val = 0;
if (!bpf_map_is_rdonly(map) || !map->ops->map_direct_value_addr ||
map->map_type == BPF_MAP_TYPE_INSN_ARRAY ||
off < 0 || off + size > map->value_size ||
bpf_map_direct_read(map, off, size, &val, is_ldsx)) {
*dst = unknown;

View File

@ -1466,7 +1466,6 @@ int bpf_fixup_call_args(struct bpf_verifier_env *env)
return err;
}
/* The function requires that first instruction in 'patch' is insnsi[prog->len - 1] */
static int add_hidden_subprog(struct bpf_verifier_env *env, struct bpf_insn *patch, int len)
{
@ -1835,6 +1834,43 @@ int bpf_do_misc_fixups(struct bpf_verifier_env *env)
goto next_insn;
}
if (bpf_jit_supports_percpu_insn() &&
insn->code == (BPF_LD | BPF_IMM | BPF_DW) &&
(insn->src_reg == BPF_PSEUDO_MAP_VALUE ||
insn->src_reg == BPF_PSEUDO_MAP_IDX_VALUE)) {
struct bpf_map *map;
aux = &env->insn_aux_data[i + delta];
map = env->used_maps[aux->map_index];
if (map->map_type != BPF_MAP_TYPE_PERCPU_ARRAY)
goto next_insn;
prog->jit_required = true;
/*
* We are *skipping* first half of ld_imm64 insn
* with 'i++;', patching over second half of it
* with that same half + mov64_percpu_reg insn.
* All because bpf_patch_insn_data() can only
* replace one 8-byte insn, which does not work
* well for ld_imm64 insn.
*/
insn_buf[0] = insn[1];
insn_buf[1] = BPF_MOV64_PERCPU_REG(insn->dst_reg, insn->dst_reg);
cnt = 2;
i++;
new_prog = bpf_patch_insn_data(env, i + delta, insn_buf, cnt);
if (!new_prog)
return -ENOMEM;
delta += cnt - 1;
env->prog = prog = new_prog;
insn = new_prog->insnsi + i + delta;
goto next_insn;
}
if (insn->code != (BPF_JMP | BPF_CALL))
goto next_insn;
if (insn->src_reg == BPF_PSEUDO_CALL)

View File

@ -998,7 +998,6 @@ static void dec_elem_count(struct bpf_htab *htab)
atomic_dec(&htab->count);
}
static void free_htab_elem(struct bpf_htab *htab, struct htab_elem *l)
{
htab_put_fd_value(htab, l);
@ -2970,7 +2969,6 @@ static int rhtab_delete_elem(struct bpf_rhtab *rhtab, struct rhtab_elem *elem, v
return 0;
}
static long rhtab_map_delete_elem(struct bpf_map *map, void *key)
{
struct bpf_rhtab *rhtab = container_of(map, struct bpf_rhtab, map);

View File

@ -4871,7 +4871,6 @@ static const struct btf_kfunc_id_set generic_kfunc_set = {
.set = &generic_btf_ids,
};
BTF_ID_LIST(generic_dtor_ids)
BTF_ID(struct, task_struct)
BTF_ID(func, bpf_task_release_dtor)

View File

@ -269,7 +269,6 @@ bpf_insn_successors(struct bpf_verifier_env *env, u32 idx)
__diag_pop();
static inline bool update_insn(struct bpf_verifier_env *env,
struct func_instance *instance, u32 frame, u32 insn_idx)
{
@ -1862,7 +1861,6 @@ static int analyze_subprog(struct bpf_verifier_env *env,
if (need_resched())
cond_resched();
/*
* When an instance is reused (must_write_initialized == true),
* record into a fresh instance and merge afterward. This avoids

View File

@ -123,7 +123,6 @@ static long __queue_map_get(struct bpf_map *map, void *value, bool delete)
return err;
}
static long __stack_map_get(struct bpf_map *map, void *value, bool delete)
{
struct bpf_queue_stack *qs = bpf_queue_stack(map);

View File

@ -636,7 +636,6 @@ int bpf_map_alloc_pages(const struct bpf_map *map, int nid,
return ret;
}
static int btf_field_cmp(const void *a, const void *b)
{
const struct btf_field *f1 = a, *f2 = b;
@ -1830,7 +1829,6 @@ static int map_lookup_elem(union bpf_attr *attr)
return err;
}
#define BPF_MAP_UPDATE_ELEM_LAST_FIELD flags
static int map_update_elem(union bpf_attr *attr, bpfptr_t uattr)
@ -3497,7 +3495,6 @@ int bpf_link_prime(struct bpf_link *link, struct bpf_link_primer *primer)
if (fd < 0)
return fd;
id = bpf_link_alloc_id(link);
if (id < 0) {
put_unused_fd(fd);
@ -5505,7 +5502,6 @@ static int bpf_link_get_info_by_fd(struct file *file,
return 0;
}
static int token_get_info_by_fd(struct file *file,
struct bpf_token *token,
const union bpf_attr *attr,
@ -6507,7 +6503,6 @@ BPF_CALL_3(bpf_sys_bpf, int, cmd, union bpf_attr *, attr, u32, attr_size)
return __sys_bpf(cmd, KERNEL_BPFPTR(attr), attr_size, KERNEL_BPFPTR(NULL), 0);
}
/* To shut up -Wmissing-prototypes.
* This function is used by the kernel light skeleton
* to load bpf programs when modules are loaded or during kernel boot.

View File

@ -635,7 +635,6 @@ static void __mark_dynptr_reg(struct bpf_reg_state *reg,
enum bpf_dynptr_type type,
bool first_slot, int id, int parent_id);
static void mark_dynptr_stack_regs(struct bpf_verifier_env *env,
struct bpf_reg_state *sreg1,
struct bpf_reg_state *sreg2,
@ -1674,7 +1673,6 @@ static bool same_callsites(struct bpf_verifier_state *a, struct bpf_verifier_sta
return true;
}
void bpf_free_backedges(struct bpf_scc_visit *visit)
{
struct bpf_scc_backedge *backedge, *next;
@ -2291,7 +2289,6 @@ static struct bpf_verifier_state *push_async_cb(struct bpf_verifier_env *env,
return &elem->st;
}
static int cmp_subprogs(const void *a, const void *b)
{
return ((struct bpf_subprog_info *)a)->start -
@ -3969,7 +3966,6 @@ static int check_stack_read(struct bpf_verifier_env *env,
return err;
}
/* check_stack_write dispatches to check_stack_write_fixed_off or
* check_stack_write_var_off.
*
@ -4767,7 +4763,6 @@ static int check_sock_access(struct bpf_verifier_env *env, int insn_idx,
valid = false;
}
if (valid) {
env->insn_aux_data[insn_idx].ctx_field_size =
info.ctx_field_size;
@ -5587,6 +5582,8 @@ int bpf_map_direct_read(struct bpf_map *map, int off, int size, u64 *val,
u64 addr;
int err;
if (map->map_type == BPF_MAP_TYPE_INSN_ARRAY || map->map_type == BPF_MAP_TYPE_PERCPU_ARRAY)
return -EINVAL;
err = map->ops->map_direct_value_addr(map, &addr, off);
if (err)
return err;
@ -6083,6 +6080,51 @@ static void add_scalar_to_reg(struct bpf_reg_state *dst_reg, s64 val)
reg_bounds_sync(dst_reg);
}
static int check_map_mem_read(struct bpf_verifier_env *env, struct bpf_reg_state *reg, int off,
int bpf_size, int value_regno, bool is_ldsx)
{
struct bpf_reg_state *regs = cur_regs(env);
int size = bpf_size_to_bytes(bpf_size);
struct bpf_map *map = reg->map_ptr;
switch (map->map_type) {
case BPF_MAP_TYPE_INSN_ARRAY:
if (bpf_size != BPF_DW) {
verbose(env, "Invalid read of %d bytes from insn_array\n", size);
return -EACCES;
}
regs[value_regno] = *reg;
add_scalar_to_reg(&regs[value_regno], off);
regs[value_regno].type = PTR_TO_INSN;
return 0;
case BPF_MAP_TYPE_PERCPU_ARRAY:
goto reg_unknown;
default:
break;
}
/* If map is read-only, track its contents as scalars. */
if (tnum_is_const(reg->var_off) &&
bpf_map_is_rdonly(map) &&
map->ops->map_direct_value_addr) {
int map_off = off + reg->var_off.value;
u64 val = 0;
int err;
err = bpf_map_direct_read(map, map_off, size, &val, is_ldsx);
if (err)
return err;
regs[value_regno].type = SCALAR_VALUE;
__mark_reg_known(&regs[value_regno], val);
return 0;
}
reg_unknown:
mark_reg_unknown(env, regs, value_regno);
return 0;
}
/* check whether memory at (regno + off) is accessible for t = (read | write)
* if t==write, value_regno is a register which value is stored into memory
* if t==read, value_regno is a register which will receive the value from memory
@ -6137,38 +6179,7 @@ static int check_mem_access(struct bpf_verifier_env *env, int insn_idx, struct b
if (kptr_field) {
err = check_map_kptr_access(env, value_regno, insn_idx, kptr_field);
} else if (t == BPF_READ && value_regno >= 0) {
struct bpf_map *map = reg->map_ptr;
/*
* If map is read-only, track its contents as scalars,
* unless it is an insn array (see the special case below)
*/
if (tnum_is_const(reg->var_off) &&
bpf_map_is_rdonly(map) &&
map->ops->map_direct_value_addr &&
map->map_type != BPF_MAP_TYPE_INSN_ARRAY) {
int map_off = off + reg->var_off.value;
u64 val = 0;
err = bpf_map_direct_read(map, map_off, size,
&val, is_ldsx);
if (err)
return err;
regs[value_regno].type = SCALAR_VALUE;
__mark_reg_known(&regs[value_regno], val);
} else if (map->map_type == BPF_MAP_TYPE_INSN_ARRAY) {
if (bpf_size != BPF_DW) {
verbose(env, "Invalid read of %d bytes from insn_array\n",
size);
return -EACCES;
}
regs[value_regno] = *reg;
add_scalar_to_reg(&regs[value_regno], off);
regs[value_regno].type = PTR_TO_INSN;
} else {
mark_reg_unknown(env, regs, value_regno);
}
err = check_map_mem_read(env, reg, off, bpf_size, value_regno, is_ldsx);
}
} else if (base_type(reg->type) == PTR_TO_MEM) {
bool rdonly_mem = type_is_rdonly_mem(reg->type);
@ -6635,7 +6646,6 @@ static int check_stack_range_initialized(
if (err)
return err;
if (tnum_is_const(reg->var_off)) {
min_off = max_off = reg->var_off.value + off;
} else {
@ -7347,7 +7357,6 @@ static bool is_iter_new_kfunc(struct bpf_call_arg_meta *meta)
return meta->kfunc_flags & KF_ITER_NEW;
}
static bool is_iter_destroy_kfunc(struct bpf_call_arg_meta *meta)
{
return meta->kfunc_flags & KF_ITER_DESTROY;
@ -8125,6 +8134,12 @@ static int check_arg_const_str(struct bpf_verifier_env *env,
return -EACCES;
}
if (map->map_type == BPF_MAP_TYPE_PERCPU_ARRAY) {
verbose(env, "%s points to percpu_array map which cannot be used as const string\n",
reg_arg_name(env, argno));
return -EACCES;
}
if (!bpf_map_is_rdonly(map)) {
verbose(env, "%s does not point to a readonly map'\n", reg_arg_name(env, argno));
return -EACCES;
@ -11607,7 +11622,6 @@ static int process_irq_flag(struct bpf_verifier_env *env, struct bpf_reg_state *
return 0;
}
static int ref_set_non_owning(struct bpf_verifier_env *env, struct bpf_reg_state *reg)
{
struct btf_record *rec = reg_btf_record(reg);
@ -16412,7 +16426,6 @@ static int check_ld_abs(struct bpf_verifier_env *env, struct bpf_insn *insn)
return 0;
}
static bool return_retval_range(struct bpf_verifier_env *env, struct bpf_retval_range *range)
{
enum bpf_prog_type prog_type = resolve_prog_type(env->prog);
@ -18361,8 +18374,6 @@ static void release_insn_arrays(struct bpf_verifier_env *env)
bpf_insn_array_release(env->insn_array_maps[i]);
}
/* The verifier does more data flow analysis than llvm and will not
* explore branches that are dead at run time. Malicious programs can
* have dead code too. Therefore replace all dead at-run-time code
@ -18390,8 +18401,6 @@ static void sanitize_dead_code(struct bpf_verifier_env *env)
}
}
static void free_states(struct bpf_verifier_env *env)
{
struct bpf_verifier_state_list *sl;
@ -18678,7 +18687,6 @@ static int do_check_main(struct bpf_verifier_env *env)
return ret;
}
static void print_verification_stats(struct bpf_verifier_env *env)
{
/* Skip over hidden subprogs which are not verified. */

View File

@ -101,6 +101,12 @@ static bool get_map_ident(const struct bpf_map *map, char *buf, size_t buf_sz)
return true;
}
if (bpf_map__type(map) == BPF_MAP_TYPE_PERCPU_ARRAY) {
snprintf(buf, buf_sz, "%s", name + 1);
sanitize_identifier(buf);
return true;
}
for (i = 0, n = ARRAY_SIZE(sfxs); i < n; i++) {
const char *sfx = sfxs[i], *p;
@ -117,7 +123,7 @@ static bool get_map_ident(const struct bpf_map *map, char *buf, size_t buf_sz)
static bool get_datasec_ident(const char *sec_name, char *buf, size_t buf_sz)
{
static const char *pfxs[] = { ".data", ".rodata", ".bss", ".kconfig" };
static const char *pfxs[] = { ".data", ".rodata", ".bss", ".percpu", ".kconfig" };
int i, n;
/* recognize hard coded LLVM section name */
@ -254,7 +260,7 @@ static const struct btf_type *find_type_for_map(struct btf *btf, const char *map
return NULL;
}
static bool is_mmapable_map(const struct bpf_map *map, char *buf, size_t sz)
static bool is_skel_data(const struct bpf_map *map, char *buf, size_t sz)
{
size_t tmp_sz;
@ -263,13 +269,24 @@ static bool is_mmapable_map(const struct bpf_map *map, char *buf, size_t sz)
return true;
}
if (!bpf_map__is_internal(map) || !(bpf_map__map_flags(map) & BPF_F_MMAPABLE))
if (!bpf_map__is_internal(map))
return false;
if (!get_map_ident(map, buf, sz))
return false;
return true;
if (bpf_map__map_flags(map) & BPF_F_MMAPABLE)
return true;
if (bpf_map__type(map) == BPF_MAP_TYPE_PERCPU_ARRAY)
return bpf_map__btf_value_type_id(map) != 0;
return false;
}
static bool is_mmapable_map(const struct bpf_map *map, char *buf, size_t sz)
{
return is_skel_data(map, buf, sz) && bpf_map__type(map) != BPF_MAP_TYPE_PERCPU_ARRAY;
}
static int codegen_datasecs(struct bpf_object *obj, const char *obj_name)
@ -287,7 +304,7 @@ static int codegen_datasecs(struct bpf_object *obj, const char *obj_name)
bpf_object__for_each_map(map, obj) {
/* only generate definitions for memory-mapped internal maps */
if (!is_mmapable_map(map, map_ident, sizeof(map_ident)))
if (!is_skel_data(map, map_ident, sizeof(map_ident)))
continue;
sec = find_type_for_map(btf, map_ident);
@ -517,7 +534,7 @@ static void codegen_asserts(struct bpf_object *obj, const char *obj_name)
", obj_name);
bpf_object__for_each_map(map, obj) {
if (!is_mmapable_map(map, map_ident, sizeof(map_ident)))
if (!is_skel_data(map, map_ident, sizeof(map_ident)))
continue;
sec = find_type_for_map(btf, map_ident);
@ -668,8 +685,7 @@ static void codegen_destroy(struct bpf_object *obj, const char *obj_name)
bpf_object__for_each_map(map, obj) {
if (!get_map_ident(map, ident, sizeof(ident)))
continue;
if (bpf_map__is_internal(map) &&
(bpf_map__map_flags(map) & BPF_F_MMAPABLE))
if (is_skel_data(map, ident, sizeof(ident)))
printf("\tskel_free_map_data(skel->%1$s, skel->maps.%1$s.initial_value, %2$zu);\n",
ident, bpf_map_mmap_sz(map));
codegen("\
@ -741,7 +757,7 @@ static int gen_trace(struct bpf_object *obj, const char *obj_name, const char *h
const void *mmap_data = NULL;
size_t mmap_size = 0;
if (!is_mmapable_map(map, ident, sizeof(ident)))
if (!is_skel_data(map, ident, sizeof(ident)))
continue;
codegen("\
@ -849,9 +865,23 @@ static int gen_trace(struct bpf_object *obj, const char *obj_name, const char *h
bpf_object__for_each_map(map, obj) {
const char *mmap_flags;
if (!is_mmapable_map(map, ident, sizeof(ident)))
if (!is_skel_data(map, ident, sizeof(ident)))
continue;
if (bpf_map__type(map) == BPF_MAP_TYPE_PERCPU_ARRAY) {
codegen("\
\n\
err = skel_protect_map_data(skel->%1$s, &skel->maps.%1$s.initial_value, %2$zd);\n\
if (err) \n\
return err; \n\
#ifdef __KERNEL__ \n\
skel->%1$s = NULL; \n\
#endif \n\
",
ident, bpf_map_mmap_sz(map));
continue;
}
if (bpf_map__map_flags(map) & BPF_F_RDONLY_PROG)
mmap_flags = "PROT_READ";
else
@ -955,8 +985,7 @@ codegen_maps_skeleton(struct bpf_object *obj, size_t map_cnt, bool mmaped, bool
map->map = &obj->maps.%s; \n\
",
i, bpf_map__name(map), ident);
/* memory-mapped internal maps */
if (mmaped && is_mmapable_map(map, ident, sizeof(ident))) {
if (mmaped && is_skel_data(map, ident, sizeof(ident))) {
printf("\tmap->mmaped = (void **)&obj->%s;\n", ident);
}

View File

@ -65,7 +65,8 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
enum bpf_prog_type prog_type, const char *prog_name,
const char *license, struct bpf_insn *insns, size_t insn_cnt,
struct bpf_prog_load_opts *load_attr, int prog_idx);
void bpf_gen__map_update_elem(struct bpf_gen *gen, int map_idx, void *value, __u32 value_size);
void bpf_gen__map_update_elem(struct bpf_gen *gen, int map_idx, void *value, __u32 value_size,
__u64 flags);
void bpf_gen__map_freeze(struct bpf_gen *gen, int map_idx);
void bpf_gen__record_attach_target(struct bpf_gen *gen, const char *name, enum bpf_attach_type type);
void bpf_gen__record_extern(struct bpf_gen *gen, const char *name, bool is_weak,

View File

@ -620,6 +620,38 @@ static int probe_bpf_syscall_common_attrs(int token_fd)
return probe_sys_bpf_ext();
}
static int probe_kern_percpu_data(int token_fd)
{
struct bpf_insn insns[] = {
BPF_LD_MAP_VALUE(BPF_REG_1, 0, 0),
BPF_LDX_MEM(BPF_DW, BPF_REG_0, BPF_REG_1, 0),
BPF_EXIT_INSN(),
};
LIBBPF_OPTS(bpf_map_create_opts, map_opts,
.token_fd = token_fd,
.map_flags = token_fd ? BPF_F_TOKEN_FD : 0,
);
LIBBPF_OPTS(bpf_prog_load_opts, prog_opts,
.token_fd = token_fd,
.prog_flags = token_fd ? BPF_F_TOKEN_FD : 0,
);
int ret, map, insn_cnt = ARRAY_SIZE(insns);
map = bpf_map_create(BPF_MAP_TYPE_PERCPU_ARRAY, "libbpf_percpu", sizeof(int), 8, 1,
&map_opts);
if (map < 0) {
pr_warn("Error in %s(): %s. Couldn't create simple percpu_array map.\n",
__func__, errstr(map));
return map;
}
insns[0].imm = map;
ret = bpf_prog_load(BPF_PROG_TYPE_SOCKET_FILTER, NULL, "GPL", insns, insn_cnt, &prog_opts);
close(map);
return probe_fd(ret);
}
typedef int (*feature_probe_fn)(int /* token_fd */);
static struct kern_feature_cache feature_cache;
@ -707,6 +739,9 @@ static struct kern_feature_desc {
[FEAT_BPF_SYSCALL_COMMON_ATTRS] = {
"BPF syscall common attributes support", probe_bpf_syscall_common_attrs,
},
[FEAT_PERCPU_DATA] = {
"kernel supports percpu data", probe_kern_percpu_data,
},
};
bool feat_supported(struct kern_feature_cache *cache, enum kern_feature_id feat_id)

View File

@ -1128,7 +1128,7 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
}
void bpf_gen__map_update_elem(struct bpf_gen *gen, int map_idx, void *pvalue,
__u32 value_size)
__u32 value_size, __u64 flags)
{
int attr_size = offsetofend(union bpf_attr, flags);
int map_update_attr, value, key;
@ -1136,6 +1136,7 @@ void bpf_gen__map_update_elem(struct bpf_gen *gen, int map_idx, void *pvalue,
int zero = 0;
memset(&attr, 0, attr_size);
attr.flags = tgt_endian(flags);
value = add_data(gen, pvalue, value_size);
key = add_data(gen, &zero, sizeof(zero));

View File

@ -541,6 +541,7 @@ struct bpf_struct_ops {
};
#define DATA_SEC ".data"
#define PERCPU_SEC ".percpu"
#define BSS_SEC ".bss"
#define RODATA_SEC ".rodata"
#define KCONFIG_SEC ".kconfig"
@ -555,6 +556,7 @@ enum libbpf_map_type {
LIBBPF_MAP_BSS,
LIBBPF_MAP_RODATA,
LIBBPF_MAP_KCONFIG,
LIBBPF_MAP_PERCPU,
};
struct bpf_map_def {
@ -666,6 +668,7 @@ enum sec_type {
SEC_DATA,
SEC_RODATA,
SEC_ST_OPS,
SEC_PERCPU,
};
struct elf_sec_desc {
@ -1839,6 +1842,8 @@ static size_t bpf_map_mmap_sz(const struct bpf_map *map)
switch (map->def.type) {
case BPF_MAP_TYPE_ARRAY:
return array_map_mmap_sz(map->def.value_size, map->def.max_entries);
case BPF_MAP_TYPE_PERCPU_ARRAY:
return map->def.value_size;
case BPF_MAP_TYPE_ARENA:
return page_sz * map->def.max_entries;
default:
@ -1866,7 +1871,8 @@ static int bpf_map_mmap_resize(struct bpf_map *map, size_t old_sz, size_t new_sz
return 0;
}
static char *internal_map_name(struct bpf_object *obj, const char *real_name)
static char *internal_map_name(struct bpf_object *obj, const char *real_name,
enum libbpf_map_type type)
{
char map_name[BPF_OBJ_NAME_LEN], *p;
int pfx_len, sfx_len = max((size_t)7, strlen(real_name));
@ -1907,8 +1913,11 @@ static char *internal_map_name(struct bpf_object *obj, const char *real_name)
if (sfx_len >= BPF_OBJ_NAME_LEN)
sfx_len = BPF_OBJ_NAME_LEN - 1;
/* if there are two or more dots in map name, it's a custom dot map */
if (strchr(real_name + 1, '.') != NULL)
/*
* Don't prefix the bpf_object name if this is a custom dot map
* (containing two or more dots) or a percpu data map.
*/
if (strchr(real_name + 1, '.') != NULL || type == LIBBPF_MAP_PERCPU)
pfx_len = 0;
else
pfx_len = min((size_t)BPF_OBJ_NAME_LEN - sfx_len - 1, strlen(obj->name));
@ -1941,6 +1950,13 @@ static bool map_is_mmapable(struct bpf_object *obj, struct bpf_map *map)
if (!map->btf_value_type_id)
return false;
/*
* The internal PERCPU maps are not mmapble because the underlying
* percpu_array maps do not have mmap support.
*/
if (map->libbpf_type == LIBBPF_MAP_PERCPU)
return false;
t = btf__type_by_id(obj->btf, map->btf_value_type_id);
if (!btf_is_datasec(t))
return false;
@ -1962,6 +1978,7 @@ static int
bpf_object__init_internal_map(struct bpf_object *obj, enum libbpf_map_type type,
const char *real_name, int sec_idx, void *data, size_t data_sz)
{
bool is_percpu = type == LIBBPF_MAP_PERCPU;
struct bpf_map_def *def;
struct bpf_map *map;
size_t mmap_sz;
@ -1975,7 +1992,7 @@ bpf_object__init_internal_map(struct bpf_object *obj, enum libbpf_map_type type,
map->sec_idx = sec_idx;
map->sec_offset = 0;
map->real_name = strdup(real_name);
map->name = internal_map_name(obj, real_name);
map->name = internal_map_name(obj, real_name, type);
if (!map->real_name || !map->name) {
zfree(&map->real_name);
zfree(&map->name);
@ -1983,7 +2000,7 @@ bpf_object__init_internal_map(struct bpf_object *obj, enum libbpf_map_type type,
}
def = &map->def;
def->type = BPF_MAP_TYPE_ARRAY;
def->type = is_percpu ? BPF_MAP_TYPE_PERCPU_ARRAY : BPF_MAP_TYPE_ARRAY;
def->key_size = sizeof(int);
def->value_size = data_sz;
def->max_entries = 1;
@ -1996,8 +2013,9 @@ bpf_object__init_internal_map(struct bpf_object *obj, enum libbpf_map_type type,
if (map_is_mmapable(obj, map))
def->map_flags |= BPF_F_MMAPABLE;
pr_debug("map '%s' (global data): at sec_idx %d, offset %zu, flags %x.\n",
map->name, map->sec_idx, map->sec_offset, def->map_flags);
pr_debug("map '%s' (global %sdata): at sec_idx %d, offset %zu, flags %x.\n",
map->name, is_percpu ? "percpu " : "", map->sec_idx,
map->sec_offset, def->map_flags);
mmap_sz = bpf_map_mmap_sz(map);
map->mmaped = mmap(NULL, mmap_sz, PROT_READ | PROT_WRITE,
@ -2057,6 +2075,13 @@ static int bpf_object__init_global_data_maps(struct bpf_object *obj)
NULL,
sec_desc->data->d_size);
break;
case SEC_PERCPU:
sec_name = elf_sec_name(obj, elf_sec_by_idx(obj, sec_idx));
err = bpf_object__init_internal_map(obj, LIBBPF_MAP_PERCPU,
sec_name, sec_idx,
sec_desc->data->d_buf,
sec_desc->data->d_size);
break;
default:
/* skip */
break;
@ -4016,6 +4041,11 @@ static int bpf_object__elf_collect(struct bpf_object *obj)
sec_desc->sec_type = SEC_RODATA;
sec_desc->shdr = sh;
sec_desc->data = data;
} else if (strcmp(name, PERCPU_SEC) == 0 ||
str_has_pfx(name, PERCPU_SEC ".")) {
sec_desc->sec_type = SEC_PERCPU;
sec_desc->shdr = sh;
sec_desc->data = data;
} else if (strcmp(name, STRUCT_OPS_SEC) == 0 ||
strcmp(name, STRUCT_OPS_LINK_SEC) == 0 ||
strcmp(name, "?" STRUCT_OPS_SEC) == 0 ||
@ -4544,6 +4574,7 @@ static bool bpf_object__shndx_is_data(const struct bpf_object *obj,
case SEC_BSS:
case SEC_DATA:
case SEC_RODATA:
case SEC_PERCPU:
return true;
default:
return false;
@ -4569,6 +4600,8 @@ bpf_object__section_to_libbpf_map_type(const struct bpf_object *obj, int shndx)
return LIBBPF_MAP_DATA;
case SEC_RODATA:
return LIBBPF_MAP_RODATA;
case SEC_PERCPU:
return LIBBPF_MAP_PERCPU;
default:
return LIBBPF_MAP_UNSPEC;
}
@ -4944,7 +4977,7 @@ static int map_fill_btf_type_info(struct bpf_object *obj, struct bpf_map *map)
/*
* LLVM annotates global data differently in BTF, that is,
* only as '.data', '.bss' or '.rodata'.
* only as '.data', '.bss', '.percpu' or '.rodata'.
*/
if (!bpf_map__is_internal(map))
return -ENOENT;
@ -5293,18 +5326,20 @@ static int
bpf_object__populate_internal_map(struct bpf_object *obj, struct bpf_map *map)
{
enum libbpf_map_type map_type = map->libbpf_type;
bool is_percpu = map_type == LIBBPF_MAP_PERCPU;
const __u64 update_flags = is_percpu ? BPF_F_ALL_CPUS : 0;
int err, zero = 0;
size_t mmap_sz;
if (obj->gen_loader) {
bpf_gen__map_update_elem(obj->gen_loader, map - obj->maps,
map->mmaped, map->def.value_size);
map->mmaped, map->def.value_size, update_flags);
if (map_type == LIBBPF_MAP_RODATA || map_type == LIBBPF_MAP_KCONFIG)
bpf_gen__map_freeze(obj->gen_loader, map - obj->maps);
return 0;
}
err = bpf_map_update_elem(map->fd, &zero, map->mmaped, 0);
err = bpf_map_update_elem(map->fd, &zero, map->mmaped, update_flags);
if (err) {
err = -errno;
pr_warn("map '%s': failed to set initial contents: %s\n",
@ -5349,6 +5384,13 @@ bpf_object__populate_internal_map(struct bpf_object *obj, struct bpf_map *map)
return err;
}
map->mmaped = mmaped;
} else if (is_percpu) {
if (mprotect(map->mmaped, mmap_sz, PROT_READ)) {
err = -errno;
pr_warn("map '%s': failed to mprotect() contents: %s\n",
bpf_map__name(map), errstr(err));
return err;
}
} else if (map->mmaped) {
munmap(map->mmaped, mmap_sz);
map->mmaped = NULL;
@ -5624,9 +5666,16 @@ bpf_object__create_maps(struct bpf_object *obj)
* runtime due to bpf_program__set_autoload(prog, false),
* bpf_object loading will succeed just fine even on old
* kernels.
* Same skipping applies to percpu data.
*/
if (bpf_map__is_internal(map) && !kernel_supports(obj, FEAT_GLOBAL_DATA))
map->autocreate = false;
if (bpf_map__is_internal(map)) {
bool is_percpu = map->libbpf_type == LIBBPF_MAP_PERCPU;
enum kern_feature_id feat_id;
feat_id = is_percpu ? FEAT_PERCPU_DATA : FEAT_GLOBAL_DATA;
if (!kernel_supports(obj, feat_id))
map->autocreate = false;
}
if (!map->autocreate) {
pr_debug("map '%s': skipped auto-creating...\n", map->name);
@ -10807,11 +10856,16 @@ static bool map_uses_real_name(const struct bpf_map *map)
* such map's corresponding ELF section name as a map name.
* This check distinguishes .data/.rodata from .data.* and .rodata.*
* maps to know which name has to be returned to the user.
* Map name of the custom .percpu.* maps might be truncated to
* BPF_OBJ_NAME_LEN-1 chars in internal_map_name(). Hence, percpu data
* maps must use real name for their user-visible name.
*/
if (map->libbpf_type == LIBBPF_MAP_DATA && strcmp(map->real_name, DATA_SEC) != 0)
return true;
if (map->libbpf_type == LIBBPF_MAP_RODATA && strcmp(map->real_name, RODATA_SEC) != 0)
return true;
if (map->libbpf_type == LIBBPF_MAP_PERCPU)
return true;
return false;
}
@ -10976,7 +11030,8 @@ int bpf_map__set_value_size(struct bpf_map *map, __u32 size)
size_t mmap_old_sz, mmap_new_sz;
int err;
if (map->def.type != BPF_MAP_TYPE_ARRAY)
if (map->def.type != BPF_MAP_TYPE_ARRAY &&
map->def.type != BPF_MAP_TYPE_PERCPU_ARRAY)
return libbpf_err(-EOPNOTSUPP);
mmap_old_sz = bpf_map_mmap_sz(map);

View File

@ -401,6 +401,8 @@ enum kern_feature_id {
FEAT_BTF_LAYOUT,
/* Kernel supports BPF syscall common attributes */
FEAT_BPF_SYSCALL_COMMON_ATTRS,
/* Kernel supports percpu data */
FEAT_PERCPU_DATA,
__FEAT_CNT,
};

View File

@ -131,8 +131,10 @@ static inline void skel_free_map_data(void *p, __u64 addr, size_t sz)
{
if (addr != ~0ULL)
kvfree(p);
/* When addr == ~0ULL the 'p' points to
* ((struct bpf_array *)map)->value. See skel_finalize_map_data.
/*
* When addr == ~0ULL the init buffer has already been released.
* For skel_finalize_map_data(), 'p' points to
* ((struct bpf_array *)map)->value.
*/
}
@ -170,6 +172,15 @@ static inline void *skel_finalize_map_data(__u64 *init_val, size_t mmap_sz, int
return addr;
}
static inline int skel_protect_map_data(void *p, __u64 *init_val, size_t sz)
{
(void)sz;
kvfree(p);
*init_val = ~0ULL;
return 0;
}
#else
static inline void *skel_alloc(size_t size)
@ -208,6 +219,15 @@ static inline void *skel_finalize_map_data(__u64 *init_val, size_t mmap_sz, int
return NULL;
return addr;
}
static inline int skel_protect_map_data(void *p, __u64 *init_val, size_t sz)
{
(void)init_val;
if (mprotect(p, sz, PROT_READ))
return -errno;
return 0;
}
#endif
static inline int skel_closenz(int fd)

View File

@ -531,7 +531,7 @@ LSKELS_SIGNED := fentry_test.c fexit_test.c atomics.c
# Generate both light skeleton and libbpf skeleton for these
LSKELS_EXTRA := test_ksyms_module.c test_ksyms_weak.c kfunc_call_test.c \
kfunc_call_test_subprog.c
kfunc_call_test_subprog.c test_global_percpu_data.c
SKEL_BLACKLIST += $$(LSKELS) $$(LSKELS_SIGNED)
test_static_linked.skel.h-deps := test_static_linked1.bpf.o test_static_linked2.bpf.o

View File

@ -1,5 +1,8 @@
// SPDX-License-Identifier: GPL-2.0
#include <test_progs.h>
#include "bpf/libbpf_internal.h"
#include "test_global_percpu_data.skel.h"
#include "test_global_percpu_data.lskel.h"
void test_global_data_init(void)
{
@ -60,3 +63,336 @@ void test_global_data_init(void)
free(newval);
bpf_object__close(obj);
}
static void test_percpu_data_on_cpus(struct bpf_map *map, int map_fd, int prog_fd, int *runp)
{
struct test_global_percpu_data__percpu *data = NULL;
int i, err, key = 0, num_online, run = 0;
__u64 args[2] = {0x1234ULL, 0x5678ULL};
size_t data_sz;
bool *online;
LIBBPF_OPTS(bpf_test_run_opts, topts,
.ctx_in = args,
.ctx_size_in = sizeof(args),
.flags = BPF_F_TEST_RUN_ON_CPU,
);
err = parse_cpu_mask_file("/sys/devices/system/cpu/online", &online, &num_online);
if (!ASSERT_OK(err, "parse_cpu_mask_file"))
return;
data_sz = map ? bpf_map__value_size(map) : sizeof(*data);
data = calloc(1, data_sz);
if (!ASSERT_OK_PTR(data, "calloc percpu data"))
goto out;
/* run on every online-CPU */
for (i = 0; i < num_online; i++) {
__u64 flags;
if (!online[i])
continue;
topts.cpu = i;
topts.retval = -1;
err = bpf_prog_test_run_opts(prog_fd, &topts);
ASSERT_OK(err, "bpf_prog_test_run_opts");
ASSERT_EQ(topts.retval, 0, "bpf_prog_test_run_opts retval");
memset(data, 0, data_sz);
flags = ((__u64) i << 32) | BPF_F_CPU;
if (map)
err = bpf_map__lookup_elem(map, &key, sizeof(key), data, data_sz, flags);
else
err = bpf_map_lookup_elem_flags(map_fd, &key, data, flags);
if (!ASSERT_OK(err, "lookup_elem on cpu"))
break;
ASSERT_EQ(*runp, ++run, "run");
ASSERT_EQ(data->cpu_id[0], i, "cpu_id");
ASSERT_EQ(data->data, 1, "data");
ASSERT_TRUE(data->set, "set");
ASSERT_EQ(data->nums[6], 0xc0de, "nums[6]");
ASSERT_EQ(data->struct_data.i, 1, "struct_data.i");
ASSERT_TRUE(data->struct_data.set, "struct_data.set");
ASSERT_EQ(data->struct_data.nums[6], 0xc0de, "struct_data.nums[6]");
}
out:
free(data);
free(online);
}
static void test_global_percpu_data_init(void)
{
struct test_global_percpu_data__percpu init_value = {};
struct test_global_percpu_data__percpu *init_data;
const __u32 desired_sz = sysconf(_SC_PAGE_SIZE);
struct test_global_percpu_data *skel = NULL;
size_t init_data_sz;
struct bpf_map *map;
int prog_fd, err;
skel = test_global_percpu_data__open();
if (!ASSERT_OK_PTR(skel, "test_global_percpu_data__open"))
goto out;
if (!ASSERT_OK_PTR(skel->percpu, "skel->percpu"))
goto out;
if (!ASSERT_OK_PTR(skel->data_percpu, "skel->data_percpu"))
goto out;
if (!ASSERT_OK_PTR(skel->percpu_data, "skel->percpu_data"))
goto out;
if (!ASSERT_OK_PTR(skel->percpu_looooooooong, "skel->percpu_looooooooong"))
goto out;
ASSERT_STREQ(bpf_map__name(skel->maps.percpu_data), ".percpu.data",
".percpu.data map name");
ASSERT_STREQ(bpf_map__name(skel->maps.data_percpu), ".data.percpu",
".data.percpu map name");
ASSERT_STREQ(bpf_map__name(skel->maps.percpu_looooooooong), ".percpu.looooooooong",
"long map name");
ASSERT_STREQ(bpf_map__name(skel->maps.percpu), ".percpu", "map name");
ASSERT_EQ(skel->percpu->data, -1, "skel->percpu->data");
ASSERT_FALSE(skel->percpu->set, "skel->percpu->set");
ASSERT_EQ(skel->percpu->nums[6], 0, "skel->percpu->nums[6]");
ASSERT_EQ(skel->percpu->struct_data.i, -1, "struct_data.i");
ASSERT_FALSE(skel->percpu->struct_data.set, "struct_data.set");
ASSERT_EQ(skel->percpu->struct_data.nums[6], 0, "struct_data.nums[6]");
map = skel->maps.percpu;
if (!ASSERT_EQ(bpf_map__type(map), BPF_MAP_TYPE_PERCPU_ARRAY, "bpf_map__type"))
goto out;
init_value.data = 2;
init_value.nums[6] = -1;
init_value.struct_data.i = 2;
init_value.struct_data.nums[6] = -1;
err = bpf_map__set_initial_value(map, &init_value, sizeof(init_value));
if (!ASSERT_OK(err, "bpf_map__set_initial_value"))
goto out;
init_data = bpf_map__initial_value(map, &init_data_sz);
if (!ASSERT_OK_PTR(init_data, "bpf_map__initial_value"))
goto out;
ASSERT_EQ(init_data->data, init_value.data, "init_value data");
ASSERT_EQ(init_data->set, init_value.set, "init_value set");
ASSERT_EQ(init_data->struct_data.i, init_value.struct_data.i, "init_value struct_data.i");
ASSERT_EQ(init_data->struct_data.nums[6], init_value.struct_data.nums[6],
"init_value struct_data.nums[6]");
ASSERT_EQ(init_data_sz, sizeof(init_value), "init_value size");
ASSERT_EQ((void *) init_data, (void *) skel->percpu, "skel->percpu eq init_data");
ASSERT_EQ(skel->percpu->data, init_value.data, "skel->percpu->data");
ASSERT_EQ(skel->percpu->set, init_value.set, "skel->percpu->set");
ASSERT_EQ(skel->percpu->struct_data.i, init_value.struct_data.i,
"skel->percpu->struct_data.i");
ASSERT_EQ(skel->percpu->struct_data.nums[6], init_value.struct_data.nums[6],
"skel->percpu->struct_data.nums[6]");
ASSERT_GT(desired_sz, sizeof(init_value), "desired_sz");
err = bpf_map__set_value_size(map, desired_sz);
if (!ASSERT_OK(err, "bpf_map__set_value_size"))
goto out;
if (!ASSERT_EQ(bpf_map__value_size(map), desired_sz, "percpu value size"))
goto out;
if (!ASSERT_NEQ(bpf_map__btf_value_type_id(map), 0, "percpu BTF value type"))
goto out;
init_data = bpf_map__initial_value(map, &init_data_sz);
if (!ASSERT_OK_PTR(init_data, "resized bpf_map__initial_value"))
goto out;
if (!ASSERT_EQ(init_data_sz, desired_sz, "resized initial value size"))
goto out;
if (!ASSERT_EQ(init_data->data, init_value.data, "resized initial value data"))
goto out;
err = test_global_percpu_data__load(skel);
if (!ASSERT_OK(err, "test_global_percpu_data__load"))
goto out;
ASSERT_OK_PTR(skel->percpu, "skel->percpu");
prog_fd = bpf_program__fd(skel->progs.update_percpu_data);
test_percpu_data_on_cpus(map, bpf_map__fd(map), prog_fd, &skel->bss->run);
out:
test_global_percpu_data__destroy(skel);
}
static void test_global_percpu_data_lskel(void)
{
struct test_global_percpu_data_lskel *lskel = NULL;
int prog_fd, map_fd;
lskel = test_global_percpu_data_lskel__open_and_load();
if (!ASSERT_OK_PTR(lskel, "test_global_percpu_data_lskel__open_and_load"))
goto out;
map_fd = lskel->maps.percpu.map_fd;
prog_fd = lskel->progs.update_percpu_data.prog_fd;
test_percpu_data_on_cpus(NULL, map_fd, prog_fd, &lskel->bss->run);
out:
test_global_percpu_data_lskel__destroy(lskel);
}
static int create_rdonly_percpu_array(void)
{
LIBBPF_OPTS(bpf_map_create_opts, map_opts,
.map_flags = BPF_F_RDONLY_PROG,
);
int key = 0, map_fd, err;
__u64 value = 0;
map_fd = bpf_map_create(BPF_MAP_TYPE_PERCPU_ARRAY, "percpu_ro_map", sizeof(int),
sizeof(__u64), 1, &map_opts);
if (!ASSERT_GE(map_fd, 0, "bpf_map_create"))
return -1;
err = bpf_map_update_elem(map_fd, &key, &value, BPF_F_ALL_CPUS);
if (!ASSERT_OK(err, "bpf_map_update_elem"))
goto out;
err = bpf_map_freeze(map_fd);
if (!ASSERT_OK(err, "bpf_map_freeze"))
goto out;
return map_fd;
out:
close(map_fd);
return -1;
}
static void test_global_percpu_data_rdonly_direct_read(void)
{
/*
* Raw instructions with manually prepared rdonly percpu_array map
* for testing direct-read global percpu data, because libbpf
* doesn't have rdonly internal percpu_array map support for
* global percpu data.
*/
struct bpf_insn insns[] = {
BPF_LD_MAP_VALUE(BPF_REG_1, 0, 0),
BPF_LDX_MEM(BPF_DW, BPF_REG_0, BPF_REG_1, 0),
BPF_EXIT_INSN(),
};
int map_fd, prog_fd;
map_fd = create_rdonly_percpu_array();
if (map_fd < 0)
return;
insns[0].imm = map_fd;
prog_fd = bpf_prog_load(BPF_PROG_TYPE_SOCKET_FILTER, "percpu_ro_prog", "GPL", insns,
ARRAY_SIZE(insns), NULL);
if (ASSERT_GE(prog_fd, 0, "bpf_prog_load"))
close(prog_fd);
close(map_fd);
}
static void test_global_percpu_data_rdonly_direct_write(void)
{
LIBBPF_OPTS(bpf_prog_load_opts, prog_opts);
/* See the comment in test_global_percpu_data_rdonly_direct_read() */
struct bpf_insn insns[] = {
BPF_LD_MAP_VALUE(BPF_REG_1, 0, 0),
BPF_LDX_MEM(BPF_DW, BPF_REG_0, BPF_REG_1, 0),
BPF_ST_MEM(BPF_DW, BPF_REG_1, 0, 0),
BPF_EXIT_INSN(),
};
char log_buf[256] = {};
int map_fd, prog_fd;
prog_opts.log_buf = log_buf;
prog_opts.log_size = sizeof(log_buf);
prog_opts.log_level = 1;
map_fd = create_rdonly_percpu_array();
if (map_fd < 0)
return;
insns[0].imm = map_fd;
prog_fd = bpf_prog_load(BPF_PROG_TYPE_SOCKET_FILTER, "percpu_ro_prog", "GPL", insns,
ARRAY_SIZE(insns), &prog_opts);
if (!ASSERT_LT(prog_fd, 0, "bpf_prog_load"))
close(prog_fd);
else
ASSERT_HAS_SUBSTR(log_buf, "write into map forbidden", "verifier log");
close(map_fd);
}
static void test_global_percpu_data_verifier_log(void)
{
RUN_TESTS(test_global_percpu_data);
}
static void test_global_percpu_data_iter(void)
{
DECLARE_LIBBPF_OPTS(bpf_iter_attach_opts, opts);
struct test_global_percpu_data *skel;
union bpf_iter_link_info linfo = {};
struct bpf_link *link = NULL;
int fd, num_cpus, len, err;
char buf[16];
num_cpus = libbpf_num_possible_cpus();
if (!ASSERT_GT(num_cpus, 0, "libbpf_num_possible_cpus"))
return;
skel = test_global_percpu_data__open();
if (!ASSERT_OK_PTR(skel, "test_global_percpu_data__open"))
return;
skel->rodata->num_cpus = num_cpus;
skel->rodata->offsetof_num = offsetof(struct test_global_percpu_data__percpu, struct_data);
skel->rodata->offsetof_num += sizeof(skel->percpu->struct_data) - sizeof(int);
skel->rodata->elem_sz = roundup(sizeof(struct test_global_percpu_data__percpu), 8);
skel->percpu->struct_data.nums[6] = 0xc0de;
err = test_global_percpu_data__load(skel);
if (!ASSERT_OK(err, "test_global_percpu_data__load"))
goto out;
linfo.map.map_fd = bpf_map__fd(skel->maps.percpu);
opts.link_info = &linfo;
opts.link_info_len = sizeof(linfo);
link = bpf_program__attach_iter(skel->progs.dump_percpu_data, &opts);
if (!ASSERT_OK_PTR(link, "bpf_program__attach_iter"))
goto out;
fd = bpf_iter_create(bpf_link__fd(link));
if (!ASSERT_GE(fd, 0, "bpf_iter_create"))
goto out;
while ((len = read(fd, buf, sizeof(buf))) > 0)
do { } while (0);
ASSERT_EQ(len, 0, "read iter");
ASSERT_TRUE(skel->bss->run_iter, "run_iter");
ASSERT_EQ(skel->bss->percpu_data_sum, 0xc0de * num_cpus, "percpu_data_sum");
close(fd);
out:
bpf_link__destroy(link);
test_global_percpu_data__destroy(skel);
}
void test_global_percpu_data(void)
{
if (!feat_supported(NULL, FEAT_PERCPU_DATA)) {
test__skip();
return;
}
if (test__start_subtest("init"))
test_global_percpu_data_init();
if (test__start_subtest("lskel"))
test_global_percpu_data_lskel();
if (test__start_subtest("rdonly_direct_read"))
test_global_percpu_data_rdonly_direct_read();
if (test__start_subtest("rdonly_direct_write"))
test_global_percpu_data_rdonly_direct_write();
test_global_percpu_data_verifier_log();
if (test__start_subtest("iter"))
test_global_percpu_data_iter();
}

View File

@ -0,0 +1,89 @@
// SPDX-License-Identifier: GPL-2.0
#include <vmlinux.h>
#include <bpf/bpf_helpers.h>
#include "bpf_misc.h"
/* Used for testing map name. */
int loong SEC(".percpu.looooooooong");
int data3 SEC(".data.percpu");
int data2 SEC(".percpu.data");
int run;
/* cpu_id as array to verify map value resizing. */
int cpu_id[1] SEC(".percpu");
int data SEC(".percpu") = -1;
int nums[7] SEC(".percpu");
bool set SEC(".percpu") = false;
struct {
char set;
int i;
int nums[7];
} struct_data SEC(".percpu") = {
.set = 0,
.i = -1,
};
SEC("raw_tp/task_rename")
__auxiliary
int update_percpu_data(void *ctx)
{
struct_data.nums[6] = 0xc0de;
struct_data.set = 1;
struct_data.i = 1;
nums[6] = 0xc0de;
data = 1;
run++;
set = true;
cpu_id[0] = bpf_get_smp_processor_id();
return 0;
}
static const char fmt[] SEC(".percpu.fmt") = "data %d\n";
SEC("?kprobe")
__failure __msg("R{{[0-9]+}} points to percpu_array map which cannot be used as const string")
int verifier_strncmp(void *ctx)
{
return bpf_strncmp("test", 5, fmt);
}
SEC("?kprobe")
__failure __msg("R{{[0-9]+}} points to percpu_array map which cannot be used as const string")
int verifier_snprintf(void *ctx)
{
u64 args[] = { data };
char buf[128];
int len;
len = bpf_snprintf(buf, sizeof(buf), fmt, args, sizeof(args));
if (len > 0)
bpf_printk("snprintf: %s\n", buf);
return 0;
}
volatile const __u32 num_cpus = 0;
volatile const int offsetof_num;
volatile const int elem_sz;
__u32 percpu_data_sum = 0;
bool run_iter = false;
SEC("iter/bpf_map_elem")
__auxiliary
int dump_percpu_data(struct bpf_iter__bpf_map_elem *ctx)
{
void *pptr = ctx->value;
int i;
if (!pptr)
return 0;
run_iter = true;
for (i = 0; i < num_cpus; i++) {
percpu_data_sum += *(int *) (pptr + offsetof_num);
pptr += elem_sz;
}
return 0;
}
char _license[] SEC("license") = "GPL";