Re: [RFC PATCH] mm/ksm: use checksum to speed up page comparison
From: David Hildenbrand (Arm)
Date: Wed Jul 29 2026 - 06:02:22 EST
On 7/16/26 14:20, Pedro Demarchi Gomes wrote:
> Use page checksums as the primary ordering key when traversing the stable
> and unstable trees and fall back to memcmp_pages() only when checksums
> match.
> Since struct ksm_stable_node does not have a checksum field, create one
> in a union with migration list fields, so when we encounter a migration
> page while scanning an address space we have to recalculate the page
> checksum. This avoids increasing the size of struct ksm_stable_node,
> which is maintained at 64 bytes, as show below.
>
> pedro@fedora:~/tmp/linux$ pahole -C ksm_stable_node ./vmlinux
> struct ksm_stable_node {
> union {
> struct {
> struct rb_node node __attribute__((__aligned__(8))); /* 0 24 */
> unsigned int checksum; /* 24 4 */
> } __attribute__((__aligned__(8))) __attribute__((__aligned__(8))); /* 0 32 */
> struct {
> struct list_head * head; /* 0 8 */
> struct {
> struct hlist_node hlist_dup; /* 8 16 */
> struct list_head list; /* 24 16 */
> }; /* 8 32 */
> }; /* 0 40 */
> } __attribute__((__aligned__(8))); /* 0 40 */
> struct hlist_head hlist; /* 40 8 */
> union {
> long unsigned int kpfn; /* 48 8 */
> long unsigned int chain_prune_time; /* 48 8 */
> }; /* 48 8 */
> int rmap_hlist_len; /* 56 4 */
> int nid; /* 60 4 */
>
> /* size: 64, cachelines: 1, members: 5 */
> /* forced alignments: 1 */
> } __attribute__((__aligned__(8)));
>
> To evaluate this change it was used two benchmarks, bench1.c and
> bench2.c. The first one allocates 8G of pages with different content,
> and the second one allocates 8G of pages where the first 4G are the same
> as the last 4G. The two benchmarks and the system ksm configuration are
> presented below.
>
> bench1.c:
>
> int main() {
> size_t size = 8ULL * 1024*1024*1024;
> unsigned long long int numpages = size/PAGESZ;
> char *pages = mmap(NULL, size, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
>
> // Generate #numpages pages with different contents
> for (unsigned long long i = 0; i < numpages; i++) {
> *((unsigned long long *) &pages[i*PAGESZ]) = i;
> }
>
> if (madvise(pages, size, MADV_MERGEABLE) != 0) {
> perror("madvise MADV_MERGEABLE failed");
> return 1;
> }
> printf("Wait...\n");
> getchar();
> return 0;
> }
>
> bench2.c:
>
> int main() {
> size_t size = 8ULL * 1024*1024*1024;
> unsigned long long int numpages = size/PAGESZ;
> char *pages = mmap(NULL, size, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
>
> // Generate #numpages pages with different contents
> for (unsigned long long i = 0; i < numpages/2; i++) {
> *((unsigned long long *) &pages[i*PAGESZ]) = i;
> *((unsigned long long *) &pages[(numpages-i-1)*PAGESZ]) = i;
> }
>
> if (madvise(pages, size, MADV_MERGEABLE) != 0) {
> perror("madvise MADV_MERGEABLE failed");
> return 1;
> }
> printf("Wait...\n");
> getchar();
> return 0;
> }
>
Hi!
> Configuration:
>
> echo never > /sys/kernel/mm/transparent_hugepage/enabled
> echo 1 > /sys/kernel/mm/ksm/sleep_millisecs
> echo 100000 > /sys/kernel/mm/ksm/pages_to_scan
Given that the default is 100, and sleep_millisecs is 200 ... and it is known
that frequent scanning is harmful for performance, what is the real world impact
in common setups?
IOW, do we even notice / care?
--
Cheers,
David