[PATCH v11 00/46] guest_memfd: In-place conversion support

From: Ackerley Tng

Date: Wed Aug 26 2026 - 05:22:08 EST


Here's v11, thanks everyone for the reviewing and testing! We're at
~5 weeks to soft-close at 7.3-rc5 for the 7.4 merge.

v11 is based on kvm-x86/next, and a cherry-picked version of David's
refactoring patch [6], and is now dependent on another series [5],
which makes kvm_gmem_get_pfn() NOT return a refcounted page to KVM.

Here's everything stitched together for your convenience:

https://github.com/googleprodkernel/linux-cc/commits/guest_memfd-inplace-conversion-v11

v11 changes:

+ Added this to "KVM: guest_memfd: Introduce per-gmem attributes, use
to guard user mappings" as Yan requested In kvm_gmem_get_pfn(),
drop the folio refcount before releasing
filemap_invalidate_lock(). This ensures that a competing
conversion request from userspace (to be added in a later patch),
which also takes the filemap_invalidate_lock(), will never see an
elevated refcount due to kvm_gmem_get_pfn().
+ Added patch "KVM: Rename kvm_mem_is_private() to
kvm_is_private_gfn()"
+ I added it before "KVM: Provide generic interface for checking
memory private/shared status", so that this second patch can
clarify the VM version for kvm_is_private_gfn() to
kvm_vm_is_private_gfn()
+ In summary, we have
+ kvm_is_private_gfn(kvm, gfn), kvm_vm_is_private_gfn(kvm,
gfn), kvm_gmem_is_private_gfn(kvm, gfn)
+ kvm_gmem_is_private_mem(inode, index): called by
kvm_gmem_is_private_gfn(kvm, gfn)
+ Added documentation for module parameter
kvm.gmem_in_place_conversion in "KVM: Let userspace disable per-VM
mem attributes, enable per-gmem attributes" as discussed by Yan and
Sean in [1]. I was reminded after more discussions on v10 [2][3].
+ Added Sean's patch "KVM: guest_memfd: Invalidate both SHARED and
PRIVATE mappings", then a no-functional-change patch to allow filter
to be passed to kvm_gmem_invalidate_start(), then the actual patch
that introduces conversions.
+ Split out a patch that just updates private_mem_conversions_test to
support gmem_in_place_conversion, as Xiaoyao suggested
+ I realized that in the v10 patch "Update
private_mem_conversions_test to mmap() guest_memfd", the requirement
to have src_type be VM_MEM_SRC_SHMEM is an artificial
requirement. If private_mem_conversions_test and vm_mem_add()
figured out alignment correctly for guest_memfd in-place conversion,
that requirement could be removed. Added a new patch "KVM:
selftests: Set up page size and alignment independently for
guest_memfd".
+ Renamed v10 patch "Update private_mem_conversions_test to mmap()
guest_memfd" to "KVM: selftests: Test in-place conversions in
private_mem_conversions_test" to better reflect what changed.
+ Regarding conversions of the VMSA page, Sashiko pointed out
something in the prereq patch series "Stop returning struct page
from guest_memfd PFN lookup". That turned out to already be handled.
:) kvm_gmem_invalidate_start eventually calls
sev_gmem_invalidate_range(), which ensures that if a VMSA page is
being converted, the vCPU associated with the VMSA page will not be
allowed to enter the guest until kvm_gmem_invalidate_end(). This
allows kvm_gmem_make_shared() and eventually rmp_make_shared() to
never fail due to a VMSA page being in-use (since the vCPU using the
VMSA page was prevented from running). Conversions fit right into
the invalidate_start and invalidate_end model :).

Here's v11 with extra tests:

https://github.com/googleprodkernel/linux-cc/commits/guest_memfd-inplace-conversion-coco-selftests-v11

Tested with both CONFIG_KVM_VM_MEMORY_ATTRIBUTES enabled and disabled:

+ tools/testing/selftests/kvm/guest_memfd_test.c
+ tools/testing/selftests/kvm/pre_fault_memory_test.c
+ tools/testing/selftests/kvm/x86/guest_memfd_conversions_test.c
+ tools/testing/selftests/kvm/x86/private_mem_conversions_test.c
+ tools/testing/selftests/kvm/x86/private_mem_kvm_exits_test.c

Also tested ./gmem_convert_fault_race (this test is not for merging)

This test reproduces the race that Yan brought up in v10 of this
series (without TDX), where a vCPU faulting in a page using
kvm_gmem_get_pfn() would cause conversion failures.

The test has 4 vCPUs continuously try to fault in the pages while
conversion to private is attempted. Each attempt is a shared to
private conversion. In between attempts, the entire range is converted
back to shared in preparation for the next attempt.

Without patch series "Stop returning struct page from guest_memfd PFN
lookup", the race causes conversions to fail with -EAGAIN quite
quickly:

Random seed: 0x704e99ea
Reproduced -EAGAIN on attempt 1374 at error_offset 0x30000

Random seed: 0x3123e3de
Reproduced -EAGAIN on attempt 983 at error_offset 0x137000

Random seed: 0x2900e25d
Reproduced -EAGAIN on attempt 870 at error_offset 0x1e1000

Random seed: 0x721a2540
Reproduced -EAGAIN on attempt 1571 at error_offset 0x1ce000

Random seed: 0x46db97b4
Reproduced -EAGAIN on attempt 6601 at error_offset 0x1af000

Random seed: 0x7eca9fbd
Reproduced -EAGAIN on attempt 5894 at error_offset 0x19c000

Random seed: 0x41355a8a
Reproduced -EAGAIN on attempt 369 at error_offset 0x15a000

With patch series, within 10000 iterations there's no -EAGAIN:

Random seed: 0x51fb3075
__vm_create: mode='PA-bits:ANY, VA-bits:48 or 57, 4K pages' type='1', pages='674'
Guest physical address width detected: 46
Guest virtual address width detected: 48
==== Test Assertion Failure ====
gmem_convert_fault_race.c:170: iter < iterations
pid=243 tid=243 errno=0 - Success
1 0x0000000000243215: test_gmem_convert_fault_race at gmem_convert_fault_race.c:168
2 (inlined by) main at gmem_convert_fault_race.c:198
3 0x0000000000267022: __libc_start_call_main at dsa.c:?
4 0x000000000026919c: __libc_start_main at dsa.c:?
5 0x0000000000242ba0: _start at ??:?
Failed to reproduce -EAGAIN within 10000 attempts

[1] https://lore.kernel.org/all/aj7NwCRwWEfLK-gQ@xxxxxxxxxx/
[2] https://lore.kernel.org/all/ed483d25-d908-4248-b8c7-862dd35dc6d4@xxxxxxxxxx/
[3] https://lore.kernel.org/all/a410cca6-0776-44af-8353-9b7375e2fd4e@xxxxxxxxx/
[4] https://lore.kernel.org/all/CAEvNRgH5ZQPGYzg+YdtEktZ72DgK9PBbzcmC_QFN6M0_6wuo1w@xxxxxxxxxxxxxx/
[5] https://lore.kernel.org/all/20260820-gmem-no-return-page-v3-0-3bf8f80a7b4d@xxxxxxxxxx/
[6] https://lore.kernel.org/all/20260825015337.B58AD1F000E9@xxxxxxxxxxxxxxx/

Older series:

v10: https://lore.kernel.org/r/20260807-gmem-inplace-conversion-v10-0-2fc18ee6d3ba@xxxxxxxxxx
v9: https://lore.kernel.org/r/20260728-gmem-inplace-conversion-v9-0-35f9aec2aed2@xxxxxxxxxx
v8: https://lore.kernel.org/r/20260618-gmem-inplace-conversion-v8-0-9d2959357853@xxxxxxxxxx
v7: https://lore.kernel.org/r/20260522-gmem-inplace-conversion-v7-0-2f0fae496530@xxxxxxxxxx
v6: https://lore.kernel.org/r/20260507-gmem-inplace-conversion-v6-0-91ab5a8b19a4@xxxxxxxxxx
RFC v5: https://lore.kernel.org/r/20260428-gmem-inplace-conversion-v5-0-d8608ccfca22@xxxxxxxxxx
RFC v4: https://lore.kernel.org/all/20260326-gmem-inplace-conversion-v4-0-e202fe950ffd@xxxxxxxxxx/T/
RFC v3: https://lore.kernel.org/r/20260313-gmem-inplace-conversion-v3-0-5fc12a70ec89@xxxxxxxxxx/T/
RFC v2: https://lore.kernel.org/all/cover.1770071243.git.ackerleytng@xxxxxxxxxx/T/
RFC v1: https://lore.kernel.org/all/cover.1760731772.git.ackerleytng@xxxxxxxxxx/T/

Previous versions of this feature, part of other series, are available at:

+ https://lore.kernel.org/all/bd163de3118b626d1005aa88e71ef2fb72f0be0f.1726009989.git.ackerleytng@xxxxxxxxxx/
+ https://lore.kernel.org/all/20250117163001.2326672-6-tabba@xxxxxxxxxx/
+ https://lore.kernel.org/all/b784326e9ccae6a08388f1bf39db70a2204bdc51.1747264138.git.ackerleytng@xxxxxxxxxx/

Signed-off-by: Ackerley Tng <ackerleytng@xxxxxxxxxx>
---
Ackerley Tng (23):
KVM: Rename kvm_mem_is_private() to kvm_is_private_gfn()
KVM: guest_memfd: Pass mapping type filter to invalidation helper
KVM: guest_memfd: Add base support for KVM_SET_MEMORY_ATTRIBUTES2
KVM: guest_memfd: Ensure pages are not in use before conversion
KVM: guest_memfd: Call arch make_shared callback for to-shared conversion
KVM: guest_memfd: Return early if range already has requested attributes
KVM: guest_memfd: Handle lru_add fbatch refcounts during conversion safety check
KVM: guest_memfd: Zero page while getting pfn
KVM: TDX: Make source page optional for KVM_TDX_INIT_MEM_REGION
KVM: selftests: Test basic single-page conversion flow
KVM: selftests: Test conversion flow when INIT_SHARED
KVM: selftests: Test conversion precision in guest_memfd
KVM: selftests: Test conversion before allocation
KVM: selftests: Convert with allocated folios in different layouts
KVM: selftests: Test that truncation does not change shared/private status
KVM: selftests: Add helpers to pin pages with CONFIG_GUP_TEST
KVM: selftests: Test conversion with elevated page refcount
KVM: selftests: Reset shared memory after hole-punching
KVM: selftests: Provide function to look up guest_memfd details from gpa
KVM: selftests: Make TEST_EXPECT_SIGBUS thread-safe
KVM: selftests: Support guest_memfd attributes in private_mem_conversions_test
KVM: selftests: Set up page size and alignment independently for guest_memfd
KVM: selftests: Test in-place conversions in private_mem_conversions_test

David Hildenbrand (Arm) (1):
mm/gup: factor out LRU cache draining for folio into lru_cache_drain_for_folio()

Michael Roth (1):
KVM: SEV: Make 'uaddr' parameter optional for KVM_SEV_SNP_LAUNCH_UPDATE

Sean Christopherson (21):
KVM: guest_memfd: Optimize away conversion overheads via dead-code elimination
KVM: guest_memfd: Use kvm_mem_is_private() when populating guest_memfd memory
KVM: guest_memfd: Introduce per-gmem attributes, use to guard user mappings
KVM: Rename KVM_GENERIC_MEMORY_ATTRIBUTES to KVM_VM_MEMORY_ATTRIBUTES
KVM: Enumerate support for PRIVATE memory iff kvm_arch_has_private_mem is defined
KVM: Rename memory attribute APIs to prepare for in-place gmem conversion
KVM: Provide generic interface for checking memory private/shared status
KVM: guest_memfd: Stub in ability to enable in-place shared<=>private conversion
KVM: Consolidate private memory and guest_memfd ifdeffery in kvm_host.h
KVM: guest_memfd: Invalidate both SHARED and PRIVATE mappings for in-place conversions
KVM: Move KVM_VM_MEMORY_ATTRIBUTES config definition to x86
KVM: Let userspace disable per-VM mem attributes, enable per-gmem attributes
KVM: guest_memfd: Enable INIT_SHARED on guest_memfd for x86 Coco VMs
KVM: selftests: Create gmem fd before "regular" fd when adding memslot
KVM: selftests: Rename guest_memfd{,_offset} to gmem_{fd,offset}
KVM: selftests: Add support for mmap() on guest_memfd in core library
KVM: selftests: Add selftests global for guest memory attributes capability
KVM: selftests: Add helpers for calling ioctls on guest_memfd
KVM: selftests: Test that shared/private status is consistent across processes
KVM: selftests: Provide common function to set memory attributes
KVM: selftests: Update private memory exits test to work with per-gmem attributes

Documentation/admin-guide/kernel-parameters.txt | 25 +
Documentation/virt/kvm/api.rst | 85 +++-
.../virt/kvm/x86/amd-memory-encryption.rst | 14 +-
Documentation/virt/kvm/x86/intel-tdx.rst | 4 +
arch/x86/include/asm/kvm-x86-ops.h | 2 +-
arch/x86/include/asm/kvm_host.h | 9 +-
arch/x86/kvm/Kconfig | 15 +-
arch/x86/kvm/mmu/mmu.c | 28 +-
arch/x86/kvm/svm/sev.c | 13 +-
arch/x86/kvm/vmx/tdx.c | 8 +-
arch/x86/kvm/x86.c | 20 +-
include/linux/kvm_host.h | 83 ++--
include/linux/swap.h | 11 +-
include/trace/events/kvm.h | 6 +-
include/uapi/linux/kvm.h | 16 +
mm/gup.c | 20 +-
mm/swap.c | 48 ++
tools/testing/selftests/kvm/Makefile.kvm | 1 +
tools/testing/selftests/kvm/include/kvm_util.h | 139 +++++-
tools/testing/selftests/kvm/include/test_util.h | 32 +-
tools/testing/selftests/kvm/lib/kvm_util.c | 222 +++++----
tools/testing/selftests/kvm/lib/test_util.c | 7 -
.../kvm/x86/guest_memfd_conversions_test.c | 512 +++++++++++++++++++++
.../kvm/x86/private_mem_conversions_test.c | 82 +++-
.../selftests/kvm/x86/private_mem_kvm_exits_test.c | 36 +-
virt/kvm/Kconfig | 3 -
virt/kvm/guest_memfd.c | 467 +++++++++++++++++--
virt/kvm/kvm_main.c | 92 ++--
28 files changed, 1701 insertions(+), 299 deletions(-)
---
base-commit: 5619ae2be01e78f6a479706b659fae07823467c1
change-id: 20260225-gmem-inplace-conversion-bd0dbd39753a

Best regards,
--
Ackerley Tng <ackerleytng@xxxxxxxxxx>