mirror of
https://github.com/torvalds/linux.git
synced 2026-10-05 11:24:03 +02:00
prb_calc_retire_blk_tmo() computes in 32-bit int arithmetic:
mbits = (blk_size_in_bytes * 8) / (1024 * 1024);
If I'm reading the validation right, tp_block_size is user
controlled and packet_set_ring() only rejects values that are <= 0
as int or not page aligned, so a 256MiB block goes right through
(and alloc_one_pg_vec_page() even has a vzalloc fallback for it).
0x10000000 * 8 wraps to INT_MIN, and on a NIC reporting 1 Gbps
(div == 1) the function ends up returning -2047.
The condition is actually (8 * size) mod 2^32 >= 2^31 && div == 1,
so the trigger set is [256,512), [768,1024), [1280,1536) and
[1792,2048) MiB. Other sizes wrap to non-negative values and faster
links divide the unsigned value back below 2^31, which is why this
doesn't blow up for everyone.
What makes it fatal is what happens next in init_prb_bdqc():
p1->interval_ktime = ms_to_ktime(prb_calc_retire_blk_tmo(...));
hrtimer_start(&p1->retire_blk_timer, p1->interval_ktime,
HRTIMER_MODE_REL_SOFT);
A negative relative timeout expires immediately. The callback
unconditionally returns HRTIMER_RESTART, and hrtimer_forward() turns
the negative interval into hrtimer_resolution:
if (interval < hrtimer_resolution)
interval = hrtimer_resolution;
So the SOFT timer re-fires at the maximum rate forever, holding
sk_receive_queue.lock each pass. One CPU spins in softirq until the
socket is closed. Repeat with more rings and the machine is gone.
The overflow itself is ancient - it was introduced together with
TPACKET_V3 in
|
||
|---|---|---|
| .. | ||
| af_packet.c | ||
| diag.c | ||
| internal.h | ||
| Kconfig | ||
| Makefile | ||