linux/tools/testing/selftests/bpf/benchs
Mykyta Yatsenko 84f7a49e76 selftests/bpf: Add resizable hashmap to benchmarks
Support resizable hashmap in BPF map benchmarks.

1. LOOKUP  (single producer, M events/sec)

  key | max | nr    |    htab |   rhtab | ratio | delta
  ----+-----+-------+---------+---------+-------+-------
    8 |  1K |   750 |   99.85 |   81.92 | 0.82x |  -18 %
    8 |  1K |    1K |  100.71 |   80.19 | 0.80x |  -20 %
    8 |  1M |  750K |   23.37 |   72.09 | 3.08x | +208 %
    8 |  1M |    1M |   13.39 |   53.72 | 4.01x | +301 %
   32 |  1K |   750 |   51.57 |   42.78 | 0.83x |  -17 %
   32 |  1K |    1K |   50.81 |   45.83 | 0.90x |  -10 %
   32 |  1M |  750K |   11.27 |   15.29 | 1.36x |  +36 %
   32 |  1M |    1M |    7.32 |    8.75 | 1.19x |  +19 %
  256 |  1K |   750 |    7.58 |    7.88 | 1.04x |   +4 %
  256 |  1K |    1K |    7.43 |    7.81 | 1.05x |   +5 %
  256 |  1M |  750K |    3.69 |    4.27 | 1.16x |  +16 %
  256 |  1M |    1M |    2.60 |    3.12 | 1.20x |  +20 %

Pattern:
  * Small map (1K): htab wins for 8 / 32 byte keys by 10-20%
  * Large map (1M): rhtab wins everywhere, up to 4x at high load
    factor with 8 byte keys.
  * Higher load factor amplifies rhtab's lead: rhtab grows the
    bucket array; htab stays at user-declared max.

2. FULL UPDATE  (M events/sec per producer)

  htab  per-producer:
    20.33   22.02   19.27   23.61   24.18   23.17   21.07
    mean  21.94   range  19.27 - 24.18

  rhtab per-producer:
   133.51  129.47   74.52  129.29  102.26  129.98  107.64
    mean 115.24   range  74.52 - 133.51

  speedup (mean): 5.25x   (+425 %)

In-place memcpy avoids the per-update alloc + RCU pointer swap
that htab pays.

3. MEMORY

  value_size |  htab ops/s | rhtab ops/s | htab mem | rhtab mem
  -----------+-------------+-------------+----------+----------
       32 B  |  122.87 k/s |  133.04 k/s | 2.47 MiB | 2.49 MiB
     4096 B  |   64.43 k/s |   65.38 k/s | 6.74 MiB | 6.44 MiB
  rhtab/htab :  +8 % ops, +0.8 % mem   (32 B)
                +1 % ops,  -4  % mem (4096 B)

Throughput effectively tied

SUMMARY

  * Small / well-fitting map: htab is faster (cache-friendly
    fixed bucket array), but only by ~10-20 %.
  * Large / high-load-factor map: rhtab is dramatically faster
    (1.2x to 4x) because rhashtable resizes to keep the load
    factor sane while htab stays stuck at user-declared max.
  * Update-heavy workloads: rhtab is ~5x faster per producer
    via in-place memcpy.
  * Memory benchmark: effectively on par.

Signed-off-by: Mykyta Yatsenko <yatsenko@meta.com>
Link: https://lore.kernel.org/r/20260605-rhash-v7-12-5b8e05f8630d@meta.com
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
2026-06-05 08:00:08 -07:00
..
bench_bloom_filter_map.c selftests/bpf: Set the default value of consumer_cnt as 0 2023-06-19 13:26:43 -07:00
bench_bpf_crypto.c selftests: bpf: crypto: add benchmark for crypto functions 2024-04-24 16:01:10 -07:00
bench_bpf_hashmap_full_update.c selftests/bpf: Add resizable hashmap to benchmarks 2026-06-05 08:00:08 -07:00
bench_bpf_hashmap_lookup.c selftests/bpf: Add resizable hashmap to benchmarks 2026-06-05 08:00:08 -07:00
bench_bpf_loop.c selftests/bpf: Set the default value of consumer_cnt as 0 2023-06-19 13:26:43 -07:00
bench_bpf_nop.c selftests/bpf: Add bpf-nop benchmark for timing overhead baseline 2026-05-11 15:25:24 -07:00
bench_bpf_timing.c selftests/bpf: Filter timing outliers with IQR in batch-timing library 2026-05-20 09:25:47 -07:00
bench_count.c selftests/bpf: Set the default value of consumer_cnt as 0 2023-06-19 13:26:43 -07:00
bench_htab_mem.c selftests/bpf: Add resizable hashmap to benchmarks 2026-06-05 08:00:08 -07:00
bench_local_storage_create.c selftests/bpf: Remove kmalloc tracing from local storage create bench 2026-04-10 21:22:31 -07:00
bench_local_storage_rcu_tasks_trace.c selftests/bpf: Set the default value of consumer_cnt as 0 2023-06-19 13:26:43 -07:00
bench_local_storage.c selftests/bpf: Set the default value of consumer_cnt as 0 2023-06-19 13:26:43 -07:00
bench_lpm_trie_map.c selftests/bpf: Add LPM trie microbenchmarks 2025-08-27 17:28:14 -07:00
bench_rename.c selftests/bpf: Set the default value of consumer_cnt as 0 2023-06-19 13:26:43 -07:00
bench_ringbufs.c selftests/bpf/benchs: Add overwrite mode benchmark for BPF ring buffer 2025-10-27 19:47:32 -07:00
bench_sockmap.c selftests/bpf: Fix incorrect array size calculation 2025-09-09 09:23:47 -07:00
bench_strncmp.c selftests/bpf: Set the default value of consumer_cnt as 0 2023-06-19 13:26:43 -07:00
bench_trigger.c selftests/bpf: Add usdt trigger bench 2026-03-03 08:39:22 -08:00
bench_xdp_lb.c selftests/bpf: Fix expired UDP LRU entries in XDP LB benchmark 2026-05-20 09:25:46 -07:00
run_bench_bloom_filter_map.sh bpf/benchs: Add benchmarks for comparing hashmap lookups w/ vs. w/out bloom filter 2021-10-28 13:22:49 -07:00
run_bench_bpf_hashmap_full_update.sh selftest/bpf/benchs: Fix a typo in bpf_hashmap_full_update 2023-02-15 16:29:31 -08:00
run_bench_bpf_loop.sh selftest/bpf/benchs: Add bpf_loop benchmark 2021-11-30 10:56:28 -08:00
run_bench_htab_mem.sh selftests/bpf: Add benchmark for bpf memory allocator 2023-07-05 18:36:19 -07:00
run_bench_local_storage_rcu_tasks_trace.sh selftest/bpf/benchs: Make quiet option common 2023-02-15 16:29:31 -08:00
run_bench_local_storage.sh selftests/bpf: Add benchmark for local_storage get 2022-06-22 19:14:33 -07:00
run_bench_rename.sh selftests/bpf: Clean up fmod_ret in bench_rename test script 2023-08-14 18:43:04 +02:00
run_bench_ringbufs.sh selftests/bpf: Add perfbuf multi-producer benchmark 2026-01-20 11:37:25 -08:00
run_bench_strncmp.sh selftests/bpf: Add benchmark for bpf_strncmp() helper 2021-12-11 17:40:23 -08:00
run_bench_trigger.sh selftests/bpf: add benchmark testing for kprobe-multi-all 2025-09-04 09:00:25 -07:00
run_bench_uprobes.sh selftests/bpf: Add usdt trigger bench 2026-03-03 08:39:22 -08:00
run_bench_xdp_lb.sh selftests/bpf: Add XDP load-balancer benchmark run script 2026-05-11 15:25:24 -07:00
run_common.sh selftests/bpf: Add benchmark for local_storage get 2022-06-22 19:14:33 -07:00