sock_map_alloc() only rejects max_entries == 0 and otherwise allows any
u32 value. sock_map_free() then walks the sks[] array with a signed int
iterator:
int i;
for (i = 0; i < stab->map.max_entries; i++)
struct sock **psk = &stab->sks[i];
When a SOCKMAP is created with max_entries = 0xffffffff (UINT_MAX), the
allocation of 32 GiB can succeed on large-memory hosts. During free the
counter reaches 0x80000000, wraps to INT_MIN, is sign-extended by movslq
and turned into a ~16 GiB negative offset from stab->sks, pointing far
below the allocation.
The faulting access is an xchg() write in sock_map_free(). Without
KASAN, the same out-of-bounds write can fault on an unmapped vmalloc page
or corrupt an unrelated allocation if that vmalloc address is populated.
On a KASAN kernel with CONFIG_KASAN_VMALLOC=y, the shadow check for that
address hits an unmapped shadow page and oopses first:
BUG: unable to handle page fault for address: fffff521b59c5a00
RIP: 0010:kasan_check_range+0x107/0x190
Call Trace:
sock_map_free+0x93/0x190
map_create+0x68d/0xb30
__sys_bpf+0x21e/0x2e70
Vmcore confirmed stab->map.max_entries == 0xffffffff, stab->sks ==
0xffffc911ace2d000, and the faulting address sks + (s64)INT_MIN * 8
exactly at 0xffffc90dace2d000. The same buggy path is reached on the
normal close()/bpf_map_free_deferred() path whenever such a map is
destroyed.
sock_map_alloc() used to bound its allocation size through
bpf_map_charge_init(), but the bound was dropped when rlimit-based memory
accounting was removed. Reject max_entries > INT_MAX at creation time so
the signed iterator in sock_map_free() never sees a value that would
overflow.
Triggered by syzkaller and reproduced on both a 6.6-based KASAN kernel
and the upstream v7.3-rc2 kernel.
Fixes: 0d2c4f9640 ("bpf: Eliminate rlimit-based memory accounting for sockmap and sockhash maps")
Signed-off-by: Zhao Gongyi <zhaogongyi@BYTEDANCE.COM>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260917121016.48171-1-zhaogongyi@bytedance.com