Re: hunting memory corruption bug in 6.18.x

From: Rik van Riel

Date: Sun Oct 04 2026 - 18:21:27 EST


On Fri, 2026-09-25 at 10:48 +0200, Nikola Ciprich wrote:
> Hi,
>
> I've been hunting a weird memory corruption bug for the last few
> weeks,
> without success so far, so I'd like to report it and kindly ask for
> help.
>
> We first hit it after a live VM migration between two KVM hosts:
> suddenly some dynamic libraries in the host appeared to be corrupted:
>
> Inconsistency detected by ld.so: ../sysdeps/x86_64/dl-machine.h: 548:
> elf_machine_rela_relative: Assertion `ELFW(R_TYPE) (reloc->r_info) ==
> R_X86_64_RELATIVE' failed!
>
> (Later we also hit this with libcrypto.so.3 etc.) The files on disk
> were OK; the problem seemed to exist only in RAM.
>
>

> I suspect two subsystems that have seen a lot of changes:
>
> - transparent hugepages
> - NUMA balancing
>
THP has a recent patch that may be relevant:

https://lore.kernel.org/all/20260903031608.1194238-1-vernon2gm@xxxxxxxxx/

> As a safety measure, we've disabled THP and NUMA balancing on all
> hosts.

Does the issue still happen with THP and NUMA balancing disabled?

>
> [1924553.414736] Oops: general protection fault, probably for non-
> canonical address 0xfffffff0c930038: 0000 [#1] SMP NOPTI

This is more than a little curious. That does not look
like a normal kernel address.

Documentation/arch/x86/x86_64/mm.rst confirms:

fffffc0000000000 | -4 TB | fffffdffffffffff | 2 TB | ...
unused hole
| | | |
vaddr_end for KASLR
fffffe0000000000 | -2 TB | fffffe7fffffffff | 0.5 TB |
cpu_entry_area mapping
fffffe8000000000 | -1.5 TB | fffffeffffffffff | 0.5 TB | ...
unused hole
ffffff0000000000 | -1 TB | ffffff7fffffffff | 0.5 TB | %esp
fixup stacks
ffffff8000000000 | -512 GB | ffffffeeffffffff | 444 GB | ...
unused hole
ffffffef00000000 | -68 GB | fffffffeffffffff | 64 GB | EFI
region mapping space
ffffffff00000000 | -4 GB | ffffffff7fffffff | 2 GB | ...
unused hole
ffffffff80000000 | -2 GB | ffffffff9fffffff | 512 MB | kernel
text mapping, mapped to physical address 0
ffffffff80000000 |-2048 MB | | |
ffffffffa0000000 |-1536 MB | fffffffffeffffff | 1520 MB | module
mapping space
ffffffffff000000 | -16 MB | | |
FIXADDR_START | ~-11 MB | ffffffffff5fffff | ~0.5 MB | kernel-
internal fixmap range, variable size and offset
ffffffffff600000 | -10 MB | ffffffffff600fff | 4 kB | legacy
vsyscall ABI
ffffffffffe00000 | -2 MB | ffffffffffffffff | 2 MB | ...
unused hole

That address is in the unused hole between EFI
region mapping space, and kernel text mapping.

That is not an address that should be overly
affected by TLB flushes, since the kernel never
maps anything there, and never has anything to
flush at that address, either.

> [1924553.434800] CPU: 23 UID: 189 PID: 7538 Comm: pacemaker-contr
> Kdump: loaded Tainted: G            E       6.18.44lb9.01 #1
> PREEMPT(voluntary)
> [1924553.456934] Tainted: [E]=UNSIGNED_MODULE
> [1924553.465551] Hardware name: ASUSTeK COMPUTER INC. RS720A-E12-
> RS12/K14PP-D24 Series, BIOS 2305 11/21/2025
> [1924553.484152] RIP: 0010:__d_lookup+0x4a/0xc0
> [1924553.492878] Code: ff 48 89 c5 c1 e8 07 48 8d 1c c2 e8 60 8f d1
> ff 48 8b 03 48 89 c3 48 83 e3 fe 48 83 f8 01 77 0a eb 2f 48 8b 1b 48
> 85 db 74 27 <39> 6b 18 75 f3 4c 8d 63 78 4c 89 e7 e8
> d5 e1 7c 00 4c 39 6b 10 74
> [1924553.525191] RSP: 0018:ff7532a13699fda0 EFLAGS: 00010212
> [1924553.534986] RAX: 0fffffff0c930020 RBX: 0fffffff0c930020 RCX:
> 0000000000000000

Which suggests it's probably following a bad
pointer in __d_lookup.

The dcache is allocated with kmalloc, and lives
in the kernel linear mapping. That is also not
a range where TLB flushes commonly happen, except
as a side effect of global flushes, or changes in
kernel mapping size (CPA flushes).

The filename lives in kernel stack memory, which
is often allocated with vmalloc. I'm not aware of
any recent vmalloc TLB issues, but maybe somebody
else knows something?

By the time memory is allocated to be used as a
kernel stack, the previous users of those pages
should be long gone. 

This does not feel like the INVLPGB bug, which
has a very short race window.

Lorenzo's CPA explanation seems like the most
likely right now.

If it still happens with that fix applied, we
need to do more digging.

--
All Rights Reversed.