Re: hunting memory corruption bug in 6.18.x
From: Nikola Ciprich
Date: Mon Oct 05 2026 - 15:04:28 EST
> Two copies of the exact address you crashed on,
> in a vmalloc area, presumably the dentry_hashtable?
>
> Can you use vtop to check that the different occurrences
> of the 0fffffff0c930020 address in the same vmalloc
> area map to the pages you found in this crash dump?
Confirmed. Both occurrences in the vmalloc (dentry_hashtable) area map to
the same physical frame the physical search reported:
vtop 0xffffb18905f60440 -> 0x103360440
vtop 0xffffb18905f60450 -> 0x103360450 (frame 0x103360000)
(AI notes: )
One detail that may matter: this vmalloc mapping is backed by a 2MB PSE
hugepage (phys base 0x103200000); the crash frame 0x103360000 is a 4KB
slice at offset 0x160000 within it. So a single 2MB TLB entry covers the
whole bucket region.
The other frames the value appears in (0x114c32000, 0x1ae234000,
0x24f0e2000, 0x342c40000, ...) are scattered across physical memory and are
not slices of that 2MB hashtable page. So the value is not confined to
the dentry_hashtable — the hashtable frame is just the one that got walked
and crashed.
>
> >
> > crash> rd -p 0x103360000 512
> >
> > 103360440: 0fffffff0c930020 0000000000000000
> > ...............
> > 103360450: 0fffffff0c930020 fffff86eb0c30000
> > ...........n...
> >
> The same thing when reading the page directly.
>
> So the invalid value the CPU reads is being read from
> memory, and not hallucinated by the CPU.
>
> > crash> search -p 0x0fffffff0c930020
> > 103360440: fffffff0c930020
> > 103360450: fffffff0c930020
> > 114c32440: fffffff0c930020
> > 114c32450: fffffff0c930020
> > 1ae234400: fffffff0c930020
> > 1ae234420: fffffff0c930020
> > 1ae234430: fffffff0c930020
> > 1ae234450: fffffff0c930020
> > 24f0e2400: fffffff0c930020
> > 24f0e2420: fffffff0c930020
> > 24f0e2430: fffffff0c930020
> > 24f0e2450: fffffff0c930020
> > 342c40400: fffffff0c930020
> > 342c40420: fffffff0c930020
> > 342c40430: fffffff0c930020
> > 342c40450: fffffff0c930020
> > 77475b400: fffffff0c930020
> > 77475b420: fffffff0c930020
> > 77475b430: fffffff0c930020
> > 77475b450: fffffff0c930020
> > b2fabe400: fffffff0c930020
> > b2fabe420: fffffff0c930020
> > b2fabe430: fffffff0c930020
> > b2fabe450: fffffff0c930020
> > ea2d7c400: fffffff0c930020
> > ea2d7c420: fffffff0c930020
> > ea2d7c430: fffffff0c930020
> > ea2d7c450: fffffff0c930020
> > ef7a2d400: fffffff0c930020
> > ef7a2d420: fffffff0c930020
> > ef7a2d430: fffffff0c930020
> > ef7a2d450: fffffff0c930020
> > ef7a90400: fffffff0c930020
> > ef7a90420: fffffff0c930020
> > ef7a90430: fffffff0c930020
> > ef7a90450: fffffff0c930020
> > 10371fa400: fffffff0c930020
> > 10371fa420: fffffff0c930020
> > 10371fa430: fffffff0c930020
> > 10371fa450: fffffff0c930020
>
> These are not random offsets in the page,
> either.
>
> These all seem to be at offsets 0, 0x20,
> 0x30, 0x40, or 0x50 into the page.
>
> If you look at the other addresses, are
> there any that are not at one of these
> offsets?
I checked all 40 occurrences. Every one is at page offset 0x400, 0x420,
0x430, 0x440, or 0x450 — nothing elsewhere, and nothing at 0x410 (so the
hypothetical sixth slot at +0x10 is not populated in this dump).
Relative to a 0x400 base, the counts are:
+0x00 : 9
+0x10 : 0
+0x20 : 9
+0x30 : 9
+0x40 : 2
+0x50 : 11
There are two distinct per-frame patterns:
9 frames carry the value at {+0x00, +0x20, +0x30, +0x50}
2 frames (incl. the crash frame 0x103360000, and 0x114c32000) carry it
only at {+0x40, +0x50}
The value-holding frames are a consistent class: kmem reports them as
non-slab, no mapping, page flags 0x2ffff800000000 — the same as the crash
frame. For the crash frame I dumped the struct page earlier: pp_magic set,
pp_ref_count = 0 (i.e. stale page_pool metadata on a page now in use by the
hashtable). I have not dumped struct page for the other frames yet; happy to
if it helps.
Two observations for the code search, offered as observations only:
The offset signature (base 0x400, entries at +0x00/+0x20/+0x30/+0x40/
+0x50) looks like an array of 16- or 32-byte-stride objects starting
0x400 into a page, with ~5-6 slots — in case that narrows "writes into
one of 6 slots."
The value 0x0fffffff0c930020 may be a legitimate pointer with a sheared
top byte: 0xffffffff0c930020 -> 0x0fffffff0c930020 (ff -> 0f). If so it
would be a partial/torn write of a real address rather than a constructed
constant — might be worth having the code search consider both.
I can dump struct page for the other value-holding frames, or pull the
surrounding bytes of one of the 4-offset frames (they may show more context
than the sparse crash frame) if either is useful.
BR
nik
>
> If this is a case of "system writes
> fffffff0c930020 into one of 6 slots",
> there could be some at offset 0x10
> too.
>
> I'm having AI comb the kernel now for
> places where we could conceivably
> construct this value, and write it
> into one out of 6 slots.
>
> >
> --
> All Rights Reversed.
>
--
Ing. Nikola CIPRICH
technický ředitel
+420 591 166 214
+420 777 093 799
nikola.ciprich@xxxxxxxxxxx
www.linuxbox.cz