Re: [PATCH v2 0/6] alpha: hugetlb support using granularity hints
From: Magnus Lindholm
Date: Thu Oct 08 2026 - 15:08:55 EST
Hi Matt,
On Thu, Oct 8, 2026 at 5:24 AM Matt Turner <mattst88@xxxxxxxxx> wrote:
>
> Alpha has no leaf entry above the last page table level, so the usual
> PMD sized huge page is not available to it. What it does have is the
> granularity hint, bits <6:5> of the PTE, described in Table 22-3 of the
> Alpha Architecture Reference Manual: a hint of order N marks a PTE as
> one of 8^N physically contiguous, naturally aligned pages that the
> translation buffer is permitted to map with a single entry. With 8KB
> pages that gives 64KB, 512KB and 4MB blocks, and this series registers
> all three as hstates.
>
> This works like arm64's contiguous PTE support rather than RISC-V's
> NAPOT: every PTE of a block keeps the frame number of its own page, so a
> block is written with set_ptes() and the frame number advances across
> it. There is no PMD sized huge page, and hence no transparent huge
> pages, no PMD page table sharing and no gigantic pages.
>
> The architecture requires all PTEs of a block to agree in bits <15:0>
> (note 2 of Table 22-3), and both __ACCESS_BITS and __DIRTY_BITS reach
> into that range, so every update rewrites the whole block and
> huge_ptep_get() merges the young and dirty state back together.
> Rewriting a valid block in place would leave its PTEs disagreeing part
> way through, so changes to a valid block go through break before make,
> following the procedure in section 11.6.1. The invalidate there has to
> assume the hint is zero and so must cover every page of the block, which
> flush_tlb_range() on alpha already exceeds, since it rolls the address
> space number.
>
> On SMP, break before make is only as good as the flush: flush_tlb_mm()
> has to reach every CPU that may hold a translation. Magnus Lindholm's
> TLB shootdown series [1] fixes flushes that miss a CPU after fork, under
> lazy TLB, or on the calling CPU, and this series depends on it there.
>
> [1] https://lore.kernel.org/linux-alpha/20260923074903.862898-1-linmag7@xxxxxxxxx/
>
> The hint is advisory. An implementation that ignores it still translates
> correctly through the individual PTEs, so this is safe on every Alpha,
> and gup_fast needs no changes. The console block that would report which
> hint sizes the translation buffer implements was never filled in by any
> console through EV7, so which sizes an implementation honors can only be
> established by measuring, as the last section does for EV7.
>
> Patches 1 and 2 are independent fixes. Patch 1 adds a
> page_table_check_pte_clear() call missing from alpha's
> ptep_get_and_clear(). Patch 2 fixes a BUG() that testing this series
> turned up: do_page_fault() fell through to BUG() on VM_FAULT_HWPOISON,
> reachable through UFFDIO_POISON without any memory failure support.
> Patch 3 replaces the "xxx" comments on the PTE read and write enable
> bits with their descriptions from Table 22-3. Patches 4 and 5 are
> groundwork and patch 6 is the implementation.
>
> Testing
> =======
>
> Tested on megalith (EV7 Marvel, 8GB), with two CPUs and again with one,
> cross-built with alpha-unknown-linux-gnu-gcc, with CONFIG_DEBUG_VM=y,
> CONFIG_DEBUG_VM_PGTABLE=y, CONFIG_PAGE_TABLE_CHECK_ENFORCED=y and
> CONFIG_CGROUP_HUGETLB=y. All three hint sizes register as hstates.
>
> mm selftests, the same in both configurations: hugetlb 10/10,
> userfaultfd and cow 7 pass 1 skip (uffd-wp-mremap), 0 fail, including
> uffd-stress hugetlb and hugetlb-private at 128MB/32 threads.
> debug_vm_pgtable validates clean at boot. No page_table_check reports
> and no DEBUG_VM splats in any run.
>
> Ad hoc tests written for this series also pass at all three sizes: dense
> per-base-page write and read back across a block, which catches a wrong
> frame number inside a block that a strided pattern would alias over;
> mprotect down to PROT_READ and back, including a middle-block-only case,
> exercising break before make; fork COW; hugetlbfs shared mappings
> checked through pread and through a second independent mapping; hole
> punch of a middle block with both neighbors and the refaulted hole
> checked; ftruncate down and back up; and the 4MB -> 512KB -> 64KB demote
> chain with exact count checks. All of it again with four concurrent
> copies, and the whole suite ten times in a row.
>
> An earlier revision of the series was also tested on up1500 (EV68AL
> Nautilus, UP, 4GB) with CONFIG_DEBUG_VM=y and
> CONFIG_PAGE_TABLE_CHECK_ENFORCED=y. That machine has not been retested
> with this revision.
>
> Magnus Lindholm also tested the series on SMP, on a UP2000+ (2x EV68AL
> 833 MHz), including a multithreaded stress test. It found that
> huge_ptep_set_access_flags() broke and rewrote a block on every fault,
> so two threads on different CPUs could keep each other faulting: 17,316
> faults in 6 seconds after one permission change on a 4MB page. With the
> check now in patch 6 the same change causes one fault. No writes were
> lost with or without it.
>
> Magnus has since run v1 on the UP2000+ again, both as posted and on top
> of [1]: his test script 200 times, the hugetlb, userfaultfd and cow
> selftests (17 pass, 1 skip) and the stress test at all three sizes, with
> no lost writes and at most one fault after a permission change on either
> kernel. v2 changes only comments and changelogs; the code is unchanged.
>
> Hugepage migration is not enabled. It has no test coverage: the
> migration, rmap and ksm selftests do not cross-build for want of libnuma
> in the alpha sysroot, and neither machine is multi-node.
>
> Does the hint do anything?
> ==========================
>
> On EV7, yes. Since the hint is advisory there is no way to ask the
> hardware whether it implements one, and the console block that would
> report it was never filled in, so the only way to find out is to
> measure.
>
> Comparing the three hint sizes against each other rather than against
> normal pages avoids the hugetlb-versus-anonymous confound. A random
> pointer chase touches one cache line per 8KB base page, with the
> permutation seeded only from the page count, so every backing walks an
> identical sequence over an identical footprint and the only variable is
> how many DTB entries the working set needs. EV6 and EV7 have a 128 entry
> fully associative DTB, so coverage is 128 times the page size: 1MB with
> base pages, 8MB at 64KB, 64MB at 512KB, 512MB at 4MB.
>
> Nanoseconds per access on megalith with one CPU, 8,000,000 accesses per
> pass. Each figure is the best of ten runs, every run on freshly
> allocated memory: single runs differ by 40% or more from one allocation
> to the next, at every page size including base pages, so the best run is
> the one to compare.
>
> size base(8KB) 64KB 512KB 4096KB
> 1 MB 11.24 11.11 15.02 11.72
> 4 MB 155.47 104.54 105.03 104.17
> 16 MB 156.69 127.50 109.55 105.52
> 64 MB 154.61 146.79 110.16 105.76
> 256 MB 172.56 171.34 175.74 107.47
> 1024 MB 232.40 230.08 222.15 190.03
>
> Each size keeps its advantage until the working set passes its own
> coverage and then converges on the base page column: 64KB is useful to
> about 8MB, 512KB to about 64MB, 4MB to about 512MB.
>
> Magnus Lindholm reports that EV68AL honors all three hint sizes too. On
> the UP2000+, retired instructions per access, which include the PALcode
> DTB fill, drop to the no-miss level exactly where each size fits the 128
> entry DTB, whose entries can each map 1, 8, 64 or 512 pages (21264/EV67
> Hardware Reference Manual, section 2.1.6.4).
>
> Signed-off-by: Matt Turner <mattst88@xxxxxxxxx>
> ---
> Changes in v2:
> - Collect Reviewed-by and Tested-by tags from Magnus Lindholm.
> - Patch 1: say which BUG_ON() the missing call leads to, add Fixes.
> - Patch 2: Cc stable, note that any local user can reach the BUG() and
> what a backport has to drop.
> - Patch 3: describe the enable bits as the Alpha Linux PTE (Table 22-3)
> defines them rather than the OpenVMS one, move the 21264 Executive
> mode detail to the changelog, and retitle to match (Magnus).
> - Patch 4: cite Table 22-3 rather than its Tru64 twin, Table 17-3.
> - Patch 5: name the commits that made the alignment the architecture's
> job (Magnus).
> - Patch 6: cite note 2 of Table 22-3 as the reason for break before
> make and section 11.6.1 as the procedure followed, in the changelog
> and in set_huge_pte_at() (Magnus).
> - Cover letter: state the SMP dependency on the TLB shootdown series,
> add the UP2000+ results for v1 and the EV68AL hint measurement.
> - Link to v1: https://lore.kernel.org/r/20261006-alpha-hugepages-v1-0-a673a18aaa70@xxxxxxxxx
>
> ---
> Matt Turner (6):
> alpha: add missing page_table_check_pte_clear() to ptep_get_and_clear()
> alpha: handle VM_FAULT_HWPOISON in do_page_fault()
> alpha: describe the PTE read and write enable bits
> alpha: define granularity hint PTE bits
> alpha: align hugetlb mappings in arch_get_unmapped_area()
> alpha: implement hugetlb support
>
> arch/alpha/Kconfig | 1 +
> arch/alpha/include/asm/hugetlb.h | 43 ++++++
> arch/alpha/include/asm/page.h | 13 ++
> arch/alpha/include/asm/pgtable.h | 61 +++++++-
> arch/alpha/kernel/osf_sys.c | 14 +-
> arch/alpha/mm/Makefile | 2 +
> arch/alpha/mm/fault.c | 19 +++
> arch/alpha/mm/hugetlbpage.c | 308 +++++++++++++++++++++++++++++++++++++++
> 8 files changed, 450 insertions(+), 11 deletions(-)
> ---
> base-commit: e946efcc89066c5d80acbae42d015a4da33a11de
> change-id: 20260912-alpha-hugepages-0c6cb6c0995d
>
> Best regards,
> --
> Matt Turner <mattst88@xxxxxxxxx>
>
Applied to my for-next
Thanks
Magnus