Re: hunting memory corruption bug in 6.18.x
From: Nikola Ciprich
Date: Mon Sep 28 2026 - 04:52:36 EST
Hi Luiz,
> > The problems started after we moved from 5.15.x to 6.18.x kernels.
>
> How long does it take to reproduce? Can you reliably distinguish good
> from bad?
Unfortunately it is painfully hard to reproduce and thus almost impossible
to bisect. A few times, after more than a week of successful tests, I
deployed a "fixed" kernel to production... and got another crash after
three weeks :(
If I were able to reproduce it more easily, bisecting would probably be
the first thing I'd try, but I still haven't found an easy way to trigger
it.
I tried heavily loaded guests running MSSQL being hammered by HammerDB
(one of the affected customers runs lots of Windows guests with MSSQL),
and others running kernel builds in a loop on top of a tmpfs ramdisk,
all of them being migrated back and forth.
>
> I know that Lorenzo jumped in and gave some good suggestions already,
> but in case you still find yourself without any further options you
> could consider if bisection is feasible: start with manual bisection
> to identify the first bad kernel between v5.15 and v6.18 and then the
> first bad -rc. You could go to git bisect from here, but it may take
> several weeks depending on how long it takes to reproduce.
>
> Another option is to try latest Linus tree to see if the issue is there.
> If it's not there then it might have been fixed, in this case you could
> bisect for the fix (if feasible, of course).
>
Yes, both approaches would be feasible if I were able to reproduce it
more easily :(
So for now I'm running another round of tests on 6.18.54, and I guess
I'll deploy it to the affected production clusters. That's still better
than waiting for a crash on the older release.
I'll report back once I have something new (from the lab or production).
cheers
nik
--
Ing. Nikola CIPRICH
technický ředitel
+420 591 166 214
+420 777 093 799
nikola.ciprich@xxxxxxxxxxx
www.linuxbox.cz