Re: [PATCH] drivers/char/mem: mmap readonly MAP_SHARED-/dev/zero correctly
From: David Hildenbrand (Arm)
Date: Fri Sep 25 2026 - 03:29:55 EST
On 9/25/26 04:58, Andrew Morton wrote:
> On Thu, 24 Sep 2026 17:37:29 +0200 "David Hildenbrand (Arm)" <david@xxxxxxxxxx> wrote:
>
>> On 9/24/26 16:48, Lorenzo Stoakes (ARM) wrote:
>>> Rather surprisingly, opening /dev/zero read-only then mmap()'ing it
>>> MAP_SHARED gets you true anonymous memory (albeit in a VMA with
>>> non-NULL vma->vm_file).
>>>
>>
>> ...
>>
>>> --- a/drivers/char/mem.c
>>> +++ b/drivers/char/mem.c
>>> @@ -503,7 +503,7 @@ static int mmap_zero_prepare(struct vm_area_desc *desc)
>>> #ifndef CONFIG_MMU
>>> return -ENOSYS;
>>> #endif
>>> - if (vma_desc_test(desc, VMA_SHARED_BIT))
>>> + if (vma_desc_test(desc, VMA_MAYSHARE_BIT))
>>> return shmem_zero_setup_desc(desc);
>>>
>>
>> So instead of shared zeropages we'd now get zero-filled shmem pages.
>
> "zeropage". Singular. Used to be!
Hey, leave that German native speaker alone! :P
Yes, you'd get the shared zeropage multiple times. (some architectures like
s390x do have multiple ones .... likely you could even get the huge zero folio here)
... unless the MM has the shared zeropage disabled, and fallback to anonymous
memory:
-> mm_forbids_zeropage()
... we end up using THPs and the huge zero folio is disallowed, so we fallback
to a anonymous THPs
-> transparent_hugepage_use_zero_page()
>
> The accounting differences, possible changes in reclaim, memcg
> charging, maybe swap behavior. Switching to a different fault handler.
> It's hard to foresee all the effects of this.
>
>> The alternative would be to just convert it to a proper read-only COW mapping in
>> mmap code:
>> * Not clearing VM_MAYWRITE, but keeping VM_WRITE clear
>> * Clearing VMA_SHARED and VMA_MAYSHARE
>>
>> Sure, someone could then mprotect(PROT_WRITE that thing) or
>> FOLL_FORCE|FOLL_WRITE to get anonymous memory. Just raising that as an alternative.
>
> I dunno, the whole thing feels imprudent. To alter such longstanding
> core(ish) behavior. And why? Because a shiny new assertion said "hey,
> that isn't quite right". Wouldn't it be better to squish the warning
> somehow and to set about this change in a very careful way?
We really shouldn't allow anonymous pages in non-cow mappings.
We can
a) Disallow allocating an anon_vma and fail gracefully. So only a shared
zeropage could ever get mapped there. Might break the s390x
mm_forbids_zeropage(). But given that's only used in hypervisors like QEMU,
unlikely.
b) Do what Lorenzo proposes. This will allocate real memory. Someone decided to
use MAP_SHARED, for unknown reasons, so I'd assume it's unlikely that
something breaks, but you have a point.
c) Convert them to proper COW mappings. After all, having the file read-only is
absolutely irrelevant, because we will never ever use that file. It's
anonymous memory.
--
Cheers,
David