Re: [BUG] Splat during modest hazptrtorture run
From: KunWu Chan
Date: Thu Oct 08 2026 - 00:16:08 EST
On Thu, Oct 8, 2026 at 11:53 AM Paul E. McKenney <paulmck@xxxxxxxxxx> wrote:
>
> On Thu, Oct 08, 2026 at 09:39:50AM +0800, KunWu Chan wrote:
> > On Wed, Oct 7, 2026 at 11:59 PM Paul E. McKenney <paulmck@xxxxxxxxxx> wrote:
> > >
> > > On Wed, Oct 07, 2026 at 03:09:27AM -0400, Mathieu Desnoyers wrote:
> > > > On 2026-10-06 19:40, Paul E. McKenney wrote:
> > > > > Hello!
> > > > >
> > > > > This is all new code, so the bug could be anywhere. So you all need
> > > > > to know. ;-)
>
> TL;DR: Adding the first patch in Mathieu's four-patch series cures this
> for me.
>
> > > > > I got this from the NOPREEMPT variant of hazptr during a nominal
> > > > > "--duration 60" run of torture.sh on x86 with the guest OS split across
> > > > > two NUMA nodes:
> > > > >
> > > > > [ 137.765104] WARNING: kernel/rcu/hazptrtorture.c:658 at hazptr_torture_stats_print+0x25b/0x4f0, CPU#1: hazptr_torture_/136
> > > > > [ 137.768647] Modules linked in:
> > > > > [ 137.768891] CPU: 1 UID: 0 PID: 136 Comm: hazptr_torture_ Not tainted 7.3.0-rc5-00661-ga1f1e4890d1d-dirty #193 PREEMPTLAZY
> > > > > [ 137.769683] Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-6.el9 11/05/2023
> > > > > [ 137.770319] RIP: 0010:hazptr_torture_stats_print+0x25b/0x4f0
> > > > > [ 137.770735] Code: 48 83 c4 20 e8 86 44 fe ff 41 83 fd 01 7e 23 48 c7 c6 6e 85 17 b0 48 c7 c7 db 27 17 b0 e8 6d 44 fe ff f0 ff 05 76 6e 0a 02 90 <0f> 0b 90 4c 8b 64 24 70 48 c7 c7 73 85 17 b0 e8 51 44 fe ff 48 8b
> > > > > [ 137.772061] RSP: 0018:ffffa738004ffdf0 EFLAGS: 00010202
> > > > > [ 137.772435] RAX: 0000000000000004 RBX: ffffa738004ffe08 RCX: 0000000000000027
> > > > > [ 137.772950] RDX: 0000000000000000 RSI: 0000000000000000 RDI: 0000000000000001
> > > > > [ 137.773455] RBP: ffffa738004ffe60 R08: 4000000000001236 R09: ffffffffb01727df
> > > > > [ 137.773980] R10: 0000000020212121 R11: 0000000020212121 R12: 0000000000000000
> > > > > [ 137.774479] R13: 0000000000000002 R14: 000000000002ffd6 R15: 0000000000034f03
> > > > > [ 137.775000] FS: 0000000000000000(0000) GS:ffff999c2e502000(0000) knlGS:0000000000000000
> > > > > [ 137.775578] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> > > > > [ 137.776001] CR2: 0000000000000000 CR3: 0000000008830000 CR4: 00000000000006f0
> > > > > [ 137.776512] Call Trace:
> > > > > [ 137.776700] <TASK>
> > > > > [ 137.776863] ? __pfx_hazptr_torture_stats+0x10/0x10
> > > > > [ 137.777207] hazptr_torture_stats+0x25/0x70
> > > > > [ 137.777508] kthread+0xd9/0x110
> > > > > [ 137.777751] ? __pfx_kthread+0x10/0x10
> > > > > [ 137.778017] ret_from_fork+0x1bd/0x220
> > > > > [ 137.778285] ? __pfx_kthread+0x10/0x10
> > > > > [ 137.778560] ret_from_fork_asm+0x1a/0x30
> > > > > [ 137.778849] </TASK>
> > > > > [ 137.779007] ---[ end trace 0000000000000000 ]---
> > > > >
> > > > > I have not yet seen this during a PREEMPT run. Line 658 of
> > > > > hazptrtorture.c is the WARN_ON_ONCE() below:
> > > > >
> > > > > if (i > 1) {
> > > > > pr_cont("%s", "!!! ");
> > > > > atomic_inc(&n_hazptr_torture_error);
> > > > > WARN_ON_ONCE(i > 1); // Too-short grace period
> > > > > }
> > > > >
> > > > > Any thoughts on what might be causing this?
> > > >
> > > > Hi Paul!
> > > >
> > > > Can you share which tree and branch (commit) this is running ?
> > > >
> > > > Does it include my fix from this series ? That would be patch 1 of:
> > > >
> > > > https://lore.kernel.org/lkml/20260927155134.4740-1-mathieu.desnoyers@xxxxxxxxxxxx/
> > >
> > > Ah, no, I somehow got the impression that this series was going to be
> > > updated, so did not apply it. Apologies!
> > >
> > > I will pull these in.
> > >
> > > > Does it run the additional series from Kunwu Chan found at
> > > > https://lore.kernel.org/lkml/20261002170847.3653663-1-kunwu.chan@xxxxxxxxx/ ?
> > >
> >
> > Hi Paul,
> >
> > I have posted v3 of the series since then. The link in Mathieu's
> > question points to v2.
> >
> > Please use the latest v3 series here:
> > https://lore.kernel.org/all/20261005171529.1378809-1-kunwu.chan@xxxxxxxxx/
>
> I do it the old way with out b4, so I would have bot there. I will
> hold off a bit for merge-window preparation, though.
>
> > I wasn't able to reproduce the warning with my current setup(without
> > my series).
> > Could you share the .config and the exact command line you used to run the
> > test? That may help me reproduce it.
>
> I used this command line, which creates the .config:
>
> tools/testing/selftests/rcutorture/bin/kvm.sh --torture hazptr --allcpus --configs "2*CFLIST" --duration 90m --trust-make
>
> This was on an ARM system, in case that matters.
>
> But pulling in Mathieu's first patch took care of this issue for me.
>
Hi Paul,
Good to hear Mathieu's first patch fixed it. That is consistent with
the race we were looking at.
The race is between the wildcard flip and a context-switch promotion:
hazptr_chain_backup_slot() selected the overflow list based on the
wildcard phase, so a backup slot promoted during the second scan could
end up in a list that was no longer covered by that scan.
Thanks for sharing the command line and test setup. The 90-minute run
also explains why it was difficult for me to reproduce with shorter
runs.
Thanks,
KunWu
> Thanx, Paul
>
> > Thanks,
> > Kunwu
> >
> > > No, but if you are good with it I will pull it in.
> > >
> > > May I add your Acked-by or Reviewed-by?
> > >
> > > Thanx, Paul
> > >
> > > > Thanks,
> > > >
> > > > Mathieu
> > > >
> > > > >
> > > > > For that matter, can anyone else reproduce this?
> > > > >
> > > > > Thanx, Paul
> > > >
> > > >
> > > > --
> > > > Mathieu Desnoyers
> > > > EfficiOS Inc.
> > > > https://www.efficios.com