Re: [PATCH v5 9/9] mm/page_owner: use memcg_data snapshot instead of PageMemcgKmem() to avoid TOCTOU VM_BUG_ON

From: Zi Yan

Date: Mon Jul 13 2026 - 15:28:02 EST


On Mon Jul 13, 2026 at 1:34 PM EDT, Vlastimil Babka (SUSE) wrote:
> On 7/13/26 04:40, Ye Liu wrote:
>>
>>
>> 在 2026/7/10 23:56, Zi Yan 写道:
>>> On 10 Jul 2026, at 2:51, Ye Liu wrote:
>>>
>>>> 在 2026/7/2 10:02, Zi Yan 写道:
>>>>> On 1 Jul 2026, at 2:10, Ye Liu wrote:
>>>>>
>>>>>> print_page_owner_memcg() takes a snapshot of page->memcg_data via
>>>>>> READ_ONCE at the top of the function and guards against tail pages
>>>>>> and NULL memcg_data. However, at the end it calls PageMemcgKmem(page)
>>>>>> which internally calls folio_memcg_kmem() — and that function re-reads
>>>>>> folio->memcg_data and page->compound_head locklessly, wrapping both
>>>>>> in VM_BUG_ON assertions:
>>>>>>
>>>>>> VM_BUG_ON_PGFLAGS(PageTail(&folio->page), &folio->page);
>>>>>> VM_BUG_ON_FOLIO(folio->memcg_data & MEMCG_DATA_OBJEXTS, folio);
>>>>>>
>>>>>> If the page is concurrently freed and reallocated as a THP tail page
>>>>>> or a slab page between the initial guards and this final call, the
>>>>>> VM_BUG_ON assertions can fire on debug builds (CONFIG_DEBUG_VM=y),
>>>>>> causing a kernel panic.
>>>>>>
>>>>>> Fix by reusing the memcg_data snapshot already taken at function entry
>>>>>> instead of calling PageMemcgKmem(), which is semantically equivalent:
>>>>>> PageMemcgKmem()->folio_memcg_kmem()->folio->memcg_data & MEMCG_DATA_KMEM.
>>>>>> This avoids both the TOCTOU window and the assertions entirely.
>>>>>>
>>>>>> Signed-off-by: Ye Liu <ye.liu@xxxxxxxxx>
>>>>>> ---
>>>>>> mm/page_owner.c | 2 +-
>>>>>> 1 file changed, 1 insertion(+), 1 deletion(-)
>>>>>>
>>>>> LGTM.
>>>>>
>>>>> Reviewed-by: Zi Yan <ziy@xxxxxxxxxx>Hi,Zi,Vlastimil
>>>>
>>>> This patch still has warnings (see sashiko link[1]).
>>>> I've made the following modifications, or do you have any better suggestions?
>>>
>>> Maybe use snapshot_page() from mm/debug.c, so that you do not need to replicate
>>> the code from folio_memcg_check(). You probably still need the rcu_read_lock()
>>> when reading memcg_data from the snapshot to prevent memcg going away.
>>>
>> But it does a full struct page + struct folio memcpy with retry loops,
>> which is more than what we need here -- we only care about memcg_data.
>
> Yeah seems too unnecessarily heavy.
>
>> The READ_ONCE() snapshot is sufficient since the function already
>> holds rcu_read_lock() to keep the memcg alive once resolved.
>>
>> The two-line inline of folio_memcg_check() logic is minimal, and we need
>> the OBJEXTS check to be explicit anyway (to print "Slab cache page\n"
>> -- page_memcg_check() just returns NULL silently for slab pages).
>
> Yeah LGTM.

I agree that snapshot_page() is an overkill here and reusing the
memcg_data snapshot looks good to me.

--
Best Regards,
Yan, Zi