Re: [PATCH v5 9/9] mm/page_owner: use memcg_data snapshot instead of PageMemcgKmem() to avoid TOCTOU VM_BUG_ON

From: Vlastimil Babka (SUSE)

Date: Mon Jul 13 2026 - 13:38:21 EST


On 7/13/26 04:40, Ye Liu wrote:
>
>
> 在 2026/7/10 23:56, Zi Yan 写道:
>> On 10 Jul 2026, at 2:51, Ye Liu wrote:
>>
>>> 在 2026/7/2 10:02, Zi Yan 写道:
>>>> On 1 Jul 2026, at 2:10, Ye Liu wrote:
>>>>
>>>>> print_page_owner_memcg() takes a snapshot of page->memcg_data via
>>>>> READ_ONCE at the top of the function and guards against tail pages
>>>>> and NULL memcg_data. However, at the end it calls PageMemcgKmem(page)
>>>>> which internally calls folio_memcg_kmem() — and that function re-reads
>>>>> folio->memcg_data and page->compound_head locklessly, wrapping both
>>>>> in VM_BUG_ON assertions:
>>>>>
>>>>> VM_BUG_ON_PGFLAGS(PageTail(&folio->page), &folio->page);
>>>>> VM_BUG_ON_FOLIO(folio->memcg_data & MEMCG_DATA_OBJEXTS, folio);
>>>>>
>>>>> If the page is concurrently freed and reallocated as a THP tail page
>>>>> or a slab page between the initial guards and this final call, the
>>>>> VM_BUG_ON assertions can fire on debug builds (CONFIG_DEBUG_VM=y),
>>>>> causing a kernel panic.
>>>>>
>>>>> Fix by reusing the memcg_data snapshot already taken at function entry
>>>>> instead of calling PageMemcgKmem(), which is semantically equivalent:
>>>>> PageMemcgKmem()->folio_memcg_kmem()->folio->memcg_data & MEMCG_DATA_KMEM.
>>>>> This avoids both the TOCTOU window and the assertions entirely.
>>>>>
>>>>> Signed-off-by: Ye Liu <ye.liu@xxxxxxxxx>
>>>>> ---
>>>>> mm/page_owner.c | 2 +-
>>>>> 1 file changed, 1 insertion(+), 1 deletion(-)
>>>>>
>>>> LGTM.
>>>>
>>>> Reviewed-by: Zi Yan <ziy@xxxxxxxxxx>Hi,Zi,Vlastimil
>>>
>>> This patch still has warnings (see sashiko link[1]).
>>> I've made the following modifications, or do you have any better suggestions?
>>
>> Maybe use snapshot_page() from mm/debug.c, so that you do not need to replicate
>> the code from folio_memcg_check(). You probably still need the rcu_read_lock()
>> when reading memcg_data from the snapshot to prevent memcg going away.
>>
> But it does a full struct page + struct folio memcpy with retry loops,
> which is more than what we need here -- we only care about memcg_data.

Yeah seems too unnecessarily heavy.

> The READ_ONCE() snapshot is sufficient since the function already
> holds rcu_read_lock() to keep the memcg alive once resolved.
>
> The two-line inline of folio_memcg_check() logic is minimal, and we need
> the OBJEXTS check to be explicit anyway (to print "Slab cache page\n"
> -- page_memcg_check() just returns NULL silently for slab pages).

Yeah LGTM.