Re: hunting memory corruption bug in 6.18.x

From: David Laight

Date: Thu Oct 01 2026 - 18:21:33 EST


On Thu, 1 Oct 2026 21:40:58 +0200
Nikola Ciprich <nikola.ciprich@xxxxxxxxxxx> wrote:

> Hello again,
>
> the good (?) news is, in the meantime we got another crash on different machine
> and I have a kdump including complete vmcore. This one was 6.18.53
>
> here are some details:
>
> Hardware: Supermicro AS-2024US-TRT / H12DSU-iN, BIOS 3.5 (2025-09-22). Dual-socket AMD EPYC 7343 16-Core, 32 CPUs, 1024 GB RAM.
>
> analyzed with crash + matching vmlinux debuginfo:
>
> Oops:
> general protection fault, probably for non-canonical address 0xfffffff0c930038
> RIP: __d_lookup+0x4a/0xc0
> Comm: systemd PID: 1274600 CPU: 23
> Call trace:
> __d_lookup
> lookup_fast
> walk_component
> link_path_walk
> path_openat
> do_filp_open
> do_sys_openat2
> __x64_sys_openat
> do_syscall_64
>
> Exception frame registers:
> RAX: 0fffffff0c930020 RBX: 0fffffff0c930020 RCX: 000000000000000b
> RDX: ffff985ffd223600 RSI: ffffb18962823d70 RDI: ffff989ddf09b5c0
> RBP: 000000005560450b R8: 000000007fffffff R9: fefefefefefefeff
> R10: 0000000000000000 R11: 93c3d2eb02a31dda R12: ffff989ddf09b5c0
> R13: ffff989ddf09b5c0 R14: ffffb18962823d70 R15: 0000000000000000
>
> Faulting instruction (cmp %ebp,0x18(%rbx)) dereferences RBX+0x18. RBX held
> the non-canonical value 0x0fffffff0c930020, giving fault address
> 0x0fffffff0c930038. __d_lookup was walking the d_hash (hlist_bl) bucket;
> RBX was the node pointer being dereferenced.

Surprisingly it looks like your compile matches the one I built from head.
The crash seems to be from the 'if (dentry->d_name.hash != hash) read.
Annoyingly the list is followed with 'mov (%rbx),%rbx' so you don't get the
address of the previous item.
However the same bad address is in %rax.
That would rather imply that it is the first time around the loop and
the 'bad address' came from the hash table itself.
(Unless the exception code manages to corrupt %rax.)

The list being corrupt would have to be memory reuse (for something else)
and the rcu protection not working.

I've just noticed that the RAX and RBX values (and the code RPC offset)
exactly match those in your original report from 25-sep.
That can't be a coincidence.
Has to be some kind of 'smoking gun'.
Possibly scanning the entire dump for 0x0c930020 might show it being
used somewhere?

David



>
> Observations from the vmcore:
>
> The faulting value 0x0fffffff0c930020 is non-canonical and is not a mapped
> kernel address:
> crash> kmem 0x0fffffff0c930020
> kmem: cannot determine page for fffffff0c930020
> fffffff0c930020: physical address not found in mem map
>
> The dentry being looked up (RDI/R12/R13 = 0xffff989ddf09b5c0) is intact and
> well-formed:
> name "app.slice", len 9, d_name.hash 0xCE973022 (consistent)
> d_op = kernfs_dops; valid d_parent, d_inode, d_sb
> d_hash.next = 0x0 (this node is the end of its bucket chain)
>
> The target dentry and its hash chain in the dump show no corruption; the
> chain terminates cleanly.
>
> No page migration, compaction, or THP activity was in progress on any CPU
> at panic. "bt -a" filtered for migrate*/compact*/khugepaged/kcompactd/
> kswapd/split_huge*/folio*/d_move/rename returned nothing.
>
> Automatic NUMA balancing was disabled at crash time (read from kernel memory):
> crash> p sysctl_numa_balancing_mode
> $ = 0
>
> Top-level (PMD) transparent hugepage policy was "never" at crash time:
> crash> p/x transparent_hugepage_flags
> $ = 0x1c0
> Bits set: 6 (DEFRAG_REQ_MADV), 7 (DEFRAG_KHUGEPAGED), 8 (USE_ZERO_PAGE).
> Bits 0 (TRANSPARENT_HUGEPAGE_FLAG) and 1 (REQ_MADV_FLAG) are clear, i.e.
> sysfs enabled = never. Per-order mTHP controls (huge_anon_orders_*) were not
> inspected for this dump, so mTHP state is not asserted here.
>
> No MCE/EDAC/hardware-error records are present in the kernel log for this host.
>
> I can provide the full vmcore and the matching vmlinux/debuginfo on request, and
> run further crash queries against it.
>
> not sure if this is of any help?
>
> with regards
>
> nik
>
>
>
> On Wed, Sep 30, 2026 at 08:40:23PM +0200, Nikola Ciprich wrote:
> > (CC Paolo Bonzini)
> >
> > Hello Lorenzo,
> >
> > >
> > > Yeah 6.18.15 is expected, I'd not say reverting is really worthwhile honestly,
> > > given what you've observed previously.
> > yes, I wasn't available at the time he was dealing with that..
> >
> >
> > >
> > > Maybe worth checking if commit 26505e1b5b54 ("KVM: SVM: make svm_flush_tlb_gva
> > > do a full asid flush if NPT enabled") helps in that case?
> > sure, I'll do that.. however, the patch doesn't apply cleanly on top of 6.18.54 at
> > all.. what do you guys recommend, is it OK to adjust the patch to this kernel
> > (I have to admit I'm able to do that, but without any deep knowledge of the subsystem)
> > or do you recommend to apply some of the previous patches?
> >
> > tried going through them, but its ~134 commits affecting svm.c between
> > v6.18 and 26505e1b5b54
> >
> > cheers
> >
> > nik
> >
> > >
> > > --
> > > Cheers, Lorenzo
> > >
> >
> > --
> > Ing. Nikola CIPRICH
> > technický ředitel
> >
> > +420 591 166 214
> > +420 777 093 799
> > nikola.ciprich@xxxxxxxxxxx
> >
> > www.linuxbox.cz
> >
>