Re: hunting memory corruption bug in 6.18.x

From: Nikola Ciprich

Date: Thu Oct 08 2026 - 16:50:00 EST


>
> What is that kernel?
>
> 6.18.55lb9.01
basically it is vanilla 6.18 + 6.18.55 + few mostly insignificant patches.
if you're willing to take a look, I tried to upload all important stuff here:

https://storage.linuxbox.cz/index.php/s/6rcp8oWCGzAqPEY

sorry for not making this available sooner, it completely slipped my mind

description of files:

linux-6.18.55-lb9.01.tar.xz - patched sources tarball
patches.tar.gz - tarball of patches applied to source, including description txt
(stable patch-6.18.55 is not included)
kernel-6.18.55-config - .config for this release
rpm/... - kernel, kernel-vmlinux etc binaries in RPM format, if this is OK for you

vmcore-dmesg.txt - kdump-dmesg, unfortunately the box got fenced before
being able to dump vmcore binary, I have to tweak this

lscpu.txt - lscpu output

if you prefer src.rpm, please let me know.

>
> The other machine has a 6.18.20lb9.03-something one.
>
> How can I look at the vmlinux you're running and the sources?
>
> rIP points to:
>
> [11402.958452] Code: 48 8d 04 c2 f6 07 02 0f 85 a0 00 00 00 48 8b 10 48 89 d0 48 83 e0 fe 48 83 fa 01 77 0d e9 80 00 00 00 48 8b 00 48 85 c0 74 78 <44> 8b 58 fc 48 39 78 10 75 ee 48 83 78 08
> All code
> ========
> 0: 48 8d 04 c2 lea (%rdx,%rax,8),%rax
> 4: f6 07 02 testb $0x2,(%rdi)
> 7: 0f 85 a0 00 00 00 jne 0xad
> d: 48 8b 10 mov (%rax),%rdx
> 10: 48 89 d0 mov %rdx,%rax
> 13: 48 83 e0 fe and $0xfffffffffffffffe,%rax
> 17: 48 83 fa 01 cmp $0x1,%rdx
> 1b: 77 0d ja 0x2a
> 1d: e9 80 00 00 00 jmp 0xa2
> 22: 48 8b 00 mov (%rax),%rax
> 25: 48 85 c0 test %rax,%rax
> 28: 74 78 je 0xa2
> 2a:* 44 8b 58 fc mov -0x4(%rax),%r11d <-- trapping instruction
> 2e: 48 39 78 10 cmp %rdi,0x10(%rax)
> 32: 75 ee jne 0x22
> 34: 48 rex.W
> 35: 83 .byte 0x83
> 36: 78 08 js 0x40
>
> I need to be able to pinpoint it back to the source.
>
> I asked the last time:
>
> "Just to rule out any other issues which got fixed in the meantime, can you try
> mainline Linux and see if you can reproduce your observation with it?
despite every effort, I wasn't able to reproduce this in lab and I couldn't easily
just upgrade customer production boxes to latest mainline...

however since we got the crash today on our own cluster node, I guess I can upgrade
at least some nodes of it to 7.2.x (possibly even 7.3-rc6 if necessary). but telling
whether it is ok after doing that is quite hard - we have tens of boxes running various
6.18 releases for months without issue. so I'm afraid the only info I'll be able
to get from this will be the problem is not fixed yet in case I get crash with latest kernel


>
> If so, you could share your crash core along with debug kernels yadda yadda so
> that I can poke at it.
>
> And before you do, make sure you have the latest BIOS and microcode installed on
> that machine.
OK

>
> Also, where can I find full dmesg and /proc/cpuinfo from those machines which
> trigger this?"
I'll collect this from those older crashes as well and upload.

>
> But still nothing.
>
> Imagine this issue has been fixed upstream but you don't have the fix in your
> kernels and we're basically chasing the same thing again...
I understand. That's why my initial question was whether this may be some known
problem for which the fixes just didn't get backported to LTS kernels yet.


>
> Sorry, but I have lost my debugging crystal ball which can help me guess what
> the machine does. :\
>
> > so we now know this didn't fixed it. however I didn't have tlbi=ipi set, so I'll
> > now try this.
>
> That won't help either but if you wanna try it.
>
> > any ideas on this new info?
>
> Yes, see above.
>
> Bottomline is: without sufficient debugging data and up-to-date hardware,
> there's not a lot I can do.
if the uploaded sources etc are of any help to you, I'll be grateful if you look
at them as well.. what else comes to my mind, if the full vmcore from one of previous
crashes including sources, debuginfo etc would be interesting for anyone, I can
either make it available for download as well (which I'll do anyways), or I can
prepare debugging VM where it all will be prepared including crash tool / gdb / whatever
can be of any use and allow access to it (supposing I'll get public part of SSH key),
just let me know if this makes sense..

thanks

BR

nik
>
> Thx.
>
> --
> Regards/Gruss,
> Boris.
>
> https://people.kernel.org/tglx/notes-about-netiquette
>