Re: [PATCH v7 2/4] s390/mm: Batch PTE updates in lazy MMU mode

From: Alexander Gordeev

Date: Thu Oct 01 2026 - 04:01:15 EST


On Mon, Aug 24, 2026 at 12:40:48PM +0200, Heiko Carstens wrote:
> > +static __always_inline bool is_lazy_mmu_active(void)
> > +{
> > + if (__is_defined(__DECOMPRESSOR))
> > + return false;
> > + if (!get_lowcore()->lazy_mmu_count)
> > + return false;
>
> I guess there is opportunity to generate better code here using an
> alternative and using a flag output constraint too.

It is only CY and LT instructions that make sense for me.
Here is what I came up with:

static __always_inline bool is_lazy_mmu_active(void)
{
unsigned long lc_lazy_mmu_count;
int cc;

if (__is_defined(__DECOMPRESSOR))
return false;
lc_lazy_mmu_count = offsetof(struct lowcore, lazy_mmu_count);
asm_inline(
ALTERNATIVE(" cy %[zero],%[offzero](%%r0)\n",
" cy %[zero],%[offalt](%%r0)\n",
ALT_FEATURE(MFEATURE_LOWCORE))
CC_IPM(cc)
: CC_OUT(cc, cc)
: [offzero] "i" (lc_lazy_mmu_count),
[offalt] "i" (lc_lazy_mmu_count + LOWCORE_ALT_ADDRESS),
[zero] "d" (0), "m" (((struct lowcore *)0)->lazy_mmu_count)
: CC_CLOBBER);
return CC_TRANSFORM(cc);
}

As result instead of gcc code:

lghi %r8,0
lt %r1,1032(%r8)
je ...

We get more or less the same:

lhi %r1,0
cy %r1,1032
jne ...

The only benefit is [zero] "d" (0) could sometimes avoid LHI if a
zeroed register is already around. But also the KASAN coverage is
not generated.

Does it make sense to go ahead whith the alternative?