x86: missing FRED #PF event data?
From: Sergey Senozhatsky
Date: Mon Aug 10 2026 - 02:53:04 EST
Greetings,
I'm currently looking at rather very strange crashes. At first they
looked like a bug in core networking, but then things started to look
quite interesting:
<1>[ 1998.756455][ T6928] BUG: kernel NULL pointer dereference, address: 0000000000000000
<1>[ 1998.756462][ T6928] #PF: supervisor read access in kernel mode
<1>[ 1998.756465][ T6928] #PF: error_code(0x0000) - not-present page
<6>[ 1998.756468][ T6928] PGD 0 P4D 0
<4>[ 1998.756471][ T6928] Oops: Oops: 0000 [#1] SMP NOPTI
<4>[ 1998.756475][ T6928] CPU: 2 UID: 1010221 PID: 6928 Comm: v_net:0 Tainted: G U W O 6.18.32 #1 PREEMPT
<4>[ 1998.756479][ T6928] Tainted: [U]=USER, [W]=WARN, [O]=OOT_MODULE
<4>[ 1998.756483][ T6928] RIP: 0010:csum_partial+0x8d/0x110
<4>[ 1998.756489][ T6928] Code: 10 48 13 57 18 48 13 57 20 48 83 d2 00 83 c1 d8 74 2d 48 83 c7 28 f6 c1 20 75 40 f6 c1 10 75 57 f6 c1 08 75 66 f6 c1 07 74 15 <48> 8b 07 f6 d9 c0 e1 03 48 d3 e0 48 d3 e8 48 01 c2 48 83 d2 00 48
<4>[ 1998.756491][ T6928] RSP: 0018:ffffb1eb88ecb608 EFLAGS: 00010202
<4>[ 1998.756493][ T6928] RAX: 540017b7b4eb9f91 RBX: 000000000000046c RCX: 000000000000000c
<4>[ 1998.756495][ T6928] RDX: b24ec1c172c93eee RSI: 000000000000046c RDI: ffff97cd9c6bbffc
<4>[ 1998.756497][ T6928] RBP: 0000000000000494 R08: 0000000000000000 R09: 0000000000000028
<4>[ 1998.756499][ T6928] R10: ffff97cda1a62a00 R11: 0000000000002140 R12: 0000000000000000
<4>[ 1998.756500][ T6928] R13: 0000000000000000 R14: ffff97ce2b800000 R15: 0000000000000000
<4>[ 1998.756502][ T6928] FS: 000075ea3ea45e78(0000) GS:ffff97d51b551000(0000) knlGS:0000000000000000
<4>[ 1998.756504][ T6928] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
<4>[ 1998.756505][ T6928] CR2: ffff97cd9c6bc000 CR3: 0000000164429002 CR4: 0000000100f72eb0
<4>[ 1998.756507][ T6928] PKRU: 55555554
<4>[ 1998.756508][ T6928] Call Trace:
<4>[ 1998.756510][ T6928] <TASK>
<4>[ 1998.756512][ T6928] skb_checksum+0x1bc/0x2f0
<4>[ 1998.756519][ T6928] skb_segment+0x729/0xde0
[..]
All the crashes are reported as NULL ptr derefs, however, I believe this
is not exactly the case. In all crashes CR2 is 0x1000 aligned (we always
crash accessing first byte of a page). It seems that csum_partial() calls
load_unaligned_zeropad() and we hit what load_unaligned_zeropad() comment
describes as very unlikely) case: "word being a page-crosser and the
next page not being mapped"). So instead of reading 4 remaining bytes
of the page and zeroes for trailing 4 bytes, we panic(). It appears that
FRED #PF is set to 0 while CR2 points to a correct page address. I added
a simple printk to exc_page_fault:
address = cpu_feature_enabled(X86_FEATURE_FRED) ? fred_event_data(regs) : read_cr2();
/* Fall back to CR2 if FRED event data was empty */
if (unlikely(!address)) {
address = read_cr2();
pr_err(":: fixed up address to %lx [[fred: %lx cr2: %lx]]\n", address, fred_event_data(regs), read_cr2());
}
and got the following while running my tests (and well, we don't crash
anymore):
[ 254.040223] :: fixed up address to ffff9c4d64af4000 [[fred: 0 cr2: ffff9c4d64af4000]]
...
[ 1821.904563] :: fixed up address to ffff9c4e9dd0a000 [[fred: 0 cr2: ffff9c4e9dd0a000]]
Does any of this make sense to you?