Re: [PATCH v3 03/26] mm: introduce AS_NO_DIRECT_MAP
From: Yosry Ahmed
Date: Fri Aug 07 2026 - 14:12:51 EST
On Fri, Aug 7, 2026 at 7:26 AM Sean Christopherson <seanjc@xxxxxxxxxx> wrote:
>
> On Thu, Aug 06, 2026, Yosry Ahmed wrote:
> > On Thu, Aug 6, 2026 at 5:19 PM Sean Christopherson <seanjc@xxxxxxxxxx> wrote:
> > > > > 2. Always access guest memory through userspace mappings, i.e. through uaccess.
> > > > >
> > > > > #2 sounds nice, but the problem is that it effectively requires hand-coded assembly
> > > > > sequences for anything more complex than basic load/store operations. Which isn't
> > > > > a complete non-starter, but it's a pretty big blocker. E.g. see the mess that is
> > > > > record_steal_time(), and then imagine trying to convert something like
> > > > > nested_vmx_prepare_msr_bitmap() to use uaccess.
> > > > >
> > > > > So, unless someone comes up with a clever idea, KVM will need something GUP-like.
> > > > > Strictly speaking, it doesn't necessarily need to be exactly GUP, because KVM could
> > > > > poke into guest_memfd directly; KVM would "just" need to manually track its own
> > > > > mappings.
> > > >
> > > > Ideally we can have shared infrastructure for this (i.e. the mermap).
> > > >
> > > > > But on x86 at least, that's not really a viable option because it only
> > > > > works for map-rarely, read/write-many use cases. For one-off accesses, creating
> > > > > and destroying (very) shortlived mappings would be too costly, and so we'd want
> > > > > those to go through uaccess.
> > > >
> > > > Not necessarily (I hope). I think the mermap pre-allocates page tables
> > > > (or some of them) and defers some TLB flushes, it's aimed at
> > > > short-lived mappings (e.g. for read() syscalls to map a file page,
> > > > copy to buffer, then unmap).
> > >
> > > I highly recommend testing shadow paging if you have aspirations of replacing
> > > the get_user() in FNAME(walk_addr_generic) with an on-demand mapping. I would
> > > be (pleasantly) shocked if dynamic mappings can provide acceptable performance.
> >
> > Oh I was thinking of existing cases where KVM uses kernel mappings (e.g.
> > kvm_vcpu_map()), as these are the ones where KVM uses GUP now, and need to
> > work for AS_NO_DIRECT_MAP to be usable. I assume the get_user() calls are
> > already a problem for guest_memfd.
>
> They aren't. KVM doesn't yet support in-place conversion, so when guest_memfd is
> used for private memory, the backing for shared memory must come from something
> other than guest_memfd. If the guest does something to prompt a host/KVM access
> to memory that is private or doesn't have a valid backing, then it's either a
> guest bug or a host userspace VMM bug.
>
> When guest_memfd is being used for shared memory, i.e. was created with
> GUEST_MEMFD_FLAG_INIT_SHARED, then get_user() Just Works, because again it's
> userspace's responsibility to establish mappings for memory that KVM may need to
> access.
>
> All of that holds true for when in-place conversion comes along: if get_user()
> hits a fault, either the guest or userspace screwed up.
Yeah get_user() should be irrelevant here, it should still work
regardless of AS_NO_DIRECT_MAP. I think the discussion was prompted by
me saying that uaccess is not necessarily significantly faster than
the mermap (don't know, never measured it), but that's besides the
point.
The main difference is mappings that KVM access through kernel
mappings (e.g. kvm_vcpu_map()).
> > AS_NO_DIRECT_MAP will surely make it a bigger problem, but not a new one :P
>
> Well, if it disallows GUP, that will be a new problem.
Yeah I think we should check here and allow GUP on unmapped pages
(more below). One thing that confuses me is that the
GUEST_MEMFD_FLAG_NO_DIRECT_MAP series [1] seems to also have this
check that disallows GUP. So I am not sure if KVM needs GUP to work
for guest_memfd now (then how does [1] work?) or it will need it to
work in the future?
[1]https://lore.kernel.org/all/20260410151746.61150-7-kalyazin@xxxxxxxxxx/
>
> > > > > At that point, userspace is basically required to
> > > > > maintain mappings for all host-accessible guest memory, and if there are userspace
> > > > > mappings, then not using GUP doesn't make much sense.
> > > > >
> > > > > Note, I called out x86 because x86 has the most extensive emulator and shadow
> > > > > paging support, which is where the isolated, one-off accesses happen in spades.
> > > > > Other architectures might be able to squeak by without userspace mappings, at
> > > > > least for now.
> > > > >
> > > > > So, in all likelihood, KVM will want GUP.
> > > >
> > > > Yeah I am thinking that the check here to disallow GUP completely for
> > > > unmapped pages is aggressive. Maybe it works for now if KVM does not
> > > > currently have any use cases for accessing guest_memfd memory. But if it does
> > > > (or will very soon), we need to think more about it, otherwise
> > > > AS_NO_DIRECT_MAP is not really usable for guest_memfd. Since you said KVM
> > > > "will want" GUP, I assume it currently doesn't?
> > >
> > > Doesn't what? Have GUP? KVM heavily uses GUP, including for guest_memfd that
> > > can be mapped into userspace.
> >
> > Your wording made me think that KVM doesn't currently use GUP for
> > guest_memfd, but I was obviously wrong. So IIUC GUP needs to succeed for
> > guest_memfd pages with AS_NO_DIRECT_MAP.
>
> Yes, though as I said early, it doesn't *have* to be exactly GUP, just something
> GUP-like. E.g. it could be a new API, if that's easier/cleaner. What I don't
> think is a good idea though is handling this entirely in KVM/guest_memfd.
Just to clarify, you mean that GUP (or GUP-like) should work in terms
of pinning the page and handing KVM/guest_memfd the pfn/address, but
not actually making the page accessible or establishing mappings,
right?
Looking at [2], seems like the consensus was that AS_NO_DIRECT_MAP
means folios are not in the direct map, and callers are responsible
for establishing the mappings (e.g. using the mermap).
[2]https://lore.kernel.org/all/aeemS2wm38Cm4qAf@xxxxxxxxxx/
>
> > To actually access the memory, I assume the guest_memfd side will need to
> > handle this by either using ephemeral mappings (e.g. mermap), restoring and
> > zapping direct mappings, or using a userspace mapping. I suppose for the
> > purposes of AS_NO_DIRECT_MAP core support we just need GUP to succeed?
>
> And establish a (ephemeral?) kernel mapping, because general users of GUP will
> expect that they can access the physical memory through the direct map. That's
> why I didn't want to handle any of this in KVM[*], the rules and handling need
> to be kernel-wide.
See above, I am struggling to understand where you think establishing
mappings should lie. The current approach is that AS_NO_DIRECT_MAP
just means folios are not mapped, and users are responsible for
establishing the mappings. I assume you agree with this (since you
suggested this :P), but you don't want KVM to do this ad-hoc, but to
have a library for it.
This library should be the mermap. I imagine (for e.g.) kvm_vcpu_map()
using the mermap under the hood if it knows the mappings do not exist
and using the mermap virtual address instead of the direct map
address. This only works (with the current implementation) if we can
disable migration (or even better, preemption) while a mapping is
active, which I imagine would be tricky or just not possible.
The alternative could be destroying and recreating the mappings when
the vCPU moves between CPUs, which is.. interesting :)
I imagine we don't have to sort all of this out now. For the purposes
of AS_NO_DIRECT_MAP (and secretmem AFAICT), we just need to provide a
facility to allocate unmapped pages. None of this is user-facing at
this point.
>
> [*] https://lore.kernel.org/all/aeemS2wm38Cm4qAf@xxxxxxxxxx