Files
linux/net/core
Florian Schauer dc0df5a0c6 page_pool: keep frag_offset aligned for odd-sized requests
page_pool_alloc_frag_netmem() rounds the requested fragment size with

	size = ALIGN(size, dma_get_cache_alignment());

dma_get_cache_alignment() returns 1 unless the architecture defines
ARCH_DMA_MINALIGN, which DMA-coherent architectures such as x86 do not.
There the ALIGN() is a no-op and pool->frag_offset advances by the raw,
unrounded size.

A single caller asking for an odd size then leaves frag_offset misaligned
for every fragment carved out of that page afterwards.  The pool is shared,
so the damage is not confined to the caller that caused it.

The per-cpu system_page_pool used by generic XDP hits this.
skb_pp_cow_data() allocates its fragments with the raw packet length:

	size = min_t(u32, len, PAGE_SIZE);
	truesize = size;
	page = page_pool_dev_alloc(pool, &page_off, &truesize);

leaving frag_offset odd for whatever is carved out of that page next.  Its
own head allocation is already aligned -- SKB_HEAD_ALIGN(size) plus the
XDP_PACKET_HEADROOM its callers pass -- so it is a later user of the shared
pool that pays: page_pool_dev_alloc_va() returns a misaligned buffer,
napi_build_skb() installs it as skb->head, and skb_shinfo(skb) ==
skb->head + skb->end is misaligned with it.

skb_shinfo()->dataref is a 4-byte atomic_t at offset 0x20, so the
atomic_inc() in __skb_clone() straddles a cache line.  On x86 with split
lock detection -- fatal for kernel split locks by default -- this panics
the machine:

  Oops: Split lock detected
  RIP: 0010:skb_clone+0x154/0x1e0
  Call Trace:
   <IRQ>
   raw_local_deliver+0x1ed/0x2c0
   ip_protocol_deliver_rcu+0x54/0x1c0
   ip_local_deliver_finish+0x85/0x100
   ip_local_deliver+0x67/0x100
   __netif_receive_skb_one_core+0x85/0xa0
   process_backlog+0x87/0x130

Reproduced by attaching any generic-mode XDP program to loopback and
opening a RAW IPPROTO_UDP socket, which makes raw_local_deliver() clone
every locally delivered UDP packet; ordinary DNS traffic then triggers it,
roughly once per 2500 clones.  Observed on 6.12.101 and 7.1.8.

Tracing page_pool_alloc_frag_netmem() over one such run shows the
amplification -- two odd-sized requests, nine misaligned offsets:

  requested size & 7:   0: 17035    5: 1    7: 1
  frag_offset & 7:      0: 17028    3: 1    4: 1    5: 1    6: 1    7: 5

and skb_pp_cow_data() returning heads that were aligned on entry:

  head 0xffff8f4c86aeac00 -> 0xffff8f4c53a9a9c4 (&7=4)
  head 0xffff8f4d6a8a42c0 -> 0xffff8f4c4f7b7a45 (&7=5)

Round the fragment size up to at least the alignment struct skb_shared_info
requires, so fragments are always suitably aligned for the objects callers
build on them.  Architectures needing a larger DMA alignment keep it.

This also makes the remainder computed in page_pool_alloc_netmem(),

	*size = max_size - *offset;

aligned, since max_size is a power of two -- which fixes the matching
misalignment of skb->end.

Verified with a controlled A/B under QEMU/KVM: same tree, same config,
same compiler, same rootfs and identical traffic, differing only by this
patch.  A SEC("xdp.frags") XDP_PASS program on lo plus UDP datagrams
larger than max_head_size drives skb_pp_cow_data()'s fragment loop, which
passes raw packet lengths to the pool.  Measured at the return of
skb_pp_cow_data():

                          unpatched   patched
  skb_pp_cow_data calls       40800     40800
  misaligned skb->head         1120         0
  dataref at line offset >60     80         0

The last row counts the accesses that actually fault:
skb_shinfo()->dataref sits at head+end+0x20 and is a 4-byte atomic, so
`lock incl` splits a 64-byte cache line only when that address lands at
offset 61..63.  All 80 occurrences were at offset 61; the panic reported
above was at offset 62.  Eliminating the misalignment removes every one
of them.

Same class of bug as commit 3bed3cc415 ("net: Do not allocate page
fragments that are not skb aligned"), which fixed the older
netdev_alloc_frag()/napi_alloc_frag() allocators.

Fixes: 53e0961da1 ("page_pool: add frag page recycling support in page pool")
Cc: stable@vger.kernel.org
Signed-off-by: Florian Schauer <florian@schauer.to>
Acked-by: Jesper Dangaard Brouer <hawk@kernel.org>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260828060822.2628276-1-florian@schauer.to
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-31 16:24:51 -07:00
..
2026-08-15 23:36:18 +02:00
2026-08-15 23:36:18 +02:00
2026-08-22 13:10:48 -07:00