Re: hunting memory corruption bug in 6.18.x

From: Nikola Ciprich

Date: Mon Sep 28 2026 - 04:52:36 EST


Hi Luiz,

> > The problems started after we moved from 5.15.x to 6.18.x kernels.
>
> How long does it take to reproduce? Can you reliably distinguish good
> from bad?

Unfortunately it is painfully hard to reproduce and thus almost impossible
to bisect. A few times, after more than a week of successful tests, I
deployed a "fixed" kernel to production... and got another crash after
three weeks :(

If I were able to reproduce it more easily, bisecting would probably be
the first thing I'd try, but I still haven't found an easy way to trigger
it.

I tried heavily loaded guests running MSSQL being hammered by HammerDB
(one of the affected customers runs lots of Windows guests with MSSQL),
and others running kernel builds in a loop on top of a tmpfs ramdisk,
all of them being migrated back and forth.


>
> I know that Lorenzo jumped in and gave some good suggestions already,
> but in case you still find yourself without any further options you
> could consider if bisection is feasible: start with manual bisection
> to identify the first bad kernel between v5.15 and v6.18 and then the
> first bad -rc. You could go to git bisect from here, but it may take
> several weeks depending on how long it takes to reproduce.
>
> Another option is to try latest Linus tree to see if the issue is there.
> If it's not there then it might have been fixed, in this case you could
> bisect for the fix (if feasible, of course).
>

Yes, both approaches would be feasible if I were able to reproduce it
more easily :(

So for now I'm running another round of tests on 6.18.54, and I guess
I'll deploy it to the affected production clusters. That's still better
than waiting for a crash on the older release.

I'll report back once I have something new (from the lab or production).

cheers

nik



--
Ing. Nikola CIPRICH
technický ředitel

+420 591 166 214
+420 777 093 799
nikola.ciprich@xxxxxxxxxxx

www.linuxbox.cz