Re: [PATCH v7 2/4] s390/mm: Batch PTE updates in lazy MMU mode
From: Alexander Gordeev
Date: Thu Oct 01 2026 - 04:01:15 EST
On Mon, Aug 24, 2026 at 12:40:48PM +0200, Heiko Carstens wrote:
> > +static __always_inline bool is_lazy_mmu_active(void)
> > +{
> > + if (__is_defined(__DECOMPRESSOR))
> > + return false;
> > + if (!get_lowcore()->lazy_mmu_count)
> > + return false;
>
> I guess there is opportunity to generate better code here using an
> alternative and using a flag output constraint too.
It is only CY and LT instructions that make sense for me.
Here is what I came up with:
static __always_inline bool is_lazy_mmu_active(void)
{
unsigned long lc_lazy_mmu_count;
int cc;
if (__is_defined(__DECOMPRESSOR))
return false;
lc_lazy_mmu_count = offsetof(struct lowcore, lazy_mmu_count);
asm_inline(
ALTERNATIVE(" cy %[zero],%[offzero](%%r0)\n",
" cy %[zero],%[offalt](%%r0)\n",
ALT_FEATURE(MFEATURE_LOWCORE))
CC_IPM(cc)
: CC_OUT(cc, cc)
: [offzero] "i" (lc_lazy_mmu_count),
[offalt] "i" (lc_lazy_mmu_count + LOWCORE_ALT_ADDRESS),
[zero] "d" (0), "m" (((struct lowcore *)0)->lazy_mmu_count)
: CC_CLOBBER);
return CC_TRANSFORM(cc);
}
As result instead of gcc code:
lghi %r8,0
lt %r1,1032(%r8)
je ...
We get more or less the same:
lhi %r1,0
cy %r1,1032
jne ...
The only benefit is [zero] "d" (0) could sometimes avoid LHI if a
zeroed register is already around. But also the KASAN coverage is
not generated.
Does it make sense to go ahead whith the alternative?