* Re: [BUG] Splat during modest hazptrtorture run
2026-10-07 15:59 ` Paul E. McKenney
@ 2026-10-07 16:03 ` Mathieu Desnoyers
2026-10-07 16:24 ` Paul E. McKenney
2026-10-07 16:08 ` Bradley Morgan
2026-10-08 1:39 ` KunWu Chan
2 siblings, 1 reply; 9+ messages in thread
From: Mathieu Desnoyers @ 2026-10-07 16:03 UTC (permalink / raw)
To: paulmck
Cc: Boqun Feng, Kunwu Chan, Wang Lian, Zqiang, Bradley Morgan, rcu,
lkmm, linux-kernel, Joel Fernandes
On 2026-10-07 11:59, Paul E. McKenney wrote:
> On Wed, Oct 07, 2026 at 03:09:27AM -0400, Mathieu Desnoyers wrote:
>> On 2026-10-06 19:40, Paul E. McKenney wrote:
>>> Hello!
>>>
>>> This is all new code, so the bug could be anywhere. So you all need
>>> to know. ;-)
>>>
>>> I got this from the NOPREEMPT variant of hazptr during a nominal
>>> "--duration 60" run of torture.sh on x86 with the guest OS split across
>>> two NUMA nodes:
>>>
>>> [ 137.765104] WARNING: kernel/rcu/hazptrtorture.c:658 at hazptr_torture_stats_print+0x25b/0x4f0, CPU#1: hazptr_torture_/136
>>> [ 137.768647] Modules linked in:
>>> [ 137.768891] CPU: 1 UID: 0 PID: 136 Comm: hazptr_torture_ Not tainted 7.3.0-rc5-00661-ga1f1e4890d1d-dirty #193 PREEMPTLAZY
>>> [ 137.769683] Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-6.el9 11/05/2023
>>> [ 137.770319] RIP: 0010:hazptr_torture_stats_print+0x25b/0x4f0
>>> [ 137.770735] Code: 48 83 c4 20 e8 86 44 fe ff 41 83 fd 01 7e 23 48 c7 c6 6e 85 17 b0 48 c7 c7 db 27 17 b0 e8 6d 44 fe ff f0 ff 05 76 6e 0a 02 90 <0f> 0b 90 4c 8b 64 24 70 48 c7 c7 73 85 17 b0 e8 51 44 fe ff 48 8b
>>> [ 137.772061] RSP: 0018:ffffa738004ffdf0 EFLAGS: 00010202
>>> [ 137.772435] RAX: 0000000000000004 RBX: ffffa738004ffe08 RCX: 0000000000000027
>>> [ 137.772950] RDX: 0000000000000000 RSI: 0000000000000000 RDI: 0000000000000001
>>> [ 137.773455] RBP: ffffa738004ffe60 R08: 4000000000001236 R09: ffffffffb01727df
>>> [ 137.773980] R10: 0000000020212121 R11: 0000000020212121 R12: 0000000000000000
>>> [ 137.774479] R13: 0000000000000002 R14: 000000000002ffd6 R15: 0000000000034f03
>>> [ 137.775000] FS: 0000000000000000(0000) GS:ffff999c2e502000(0000) knlGS:0000000000000000
>>> [ 137.775578] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
>>> [ 137.776001] CR2: 0000000000000000 CR3: 0000000008830000 CR4: 00000000000006f0
>>> [ 137.776512] Call Trace:
>>> [ 137.776700] <TASK>
>>> [ 137.776863] ? __pfx_hazptr_torture_stats+0x10/0x10
>>> [ 137.777207] hazptr_torture_stats+0x25/0x70
>>> [ 137.777508] kthread+0xd9/0x110
>>> [ 137.777751] ? __pfx_kthread+0x10/0x10
>>> [ 137.778017] ret_from_fork+0x1bd/0x220
>>> [ 137.778285] ? __pfx_kthread+0x10/0x10
>>> [ 137.778560] ret_from_fork_asm+0x1a/0x30
>>> [ 137.778849] </TASK>
>>> [ 137.779007] ---[ end trace 0000000000000000 ]---
>>>
>>> I have not yet seen this during a PREEMPT run. Line 658 of
>>> hazptrtorture.c is the WARN_ON_ONCE() below:
>>>
>>> if (i > 1) {
>>> pr_cont("%s", "!!! ");
>>> atomic_inc(&n_hazptr_torture_error);
>>> WARN_ON_ONCE(i > 1); // Too-short grace period
>>> }
>>>
>>> Any thoughts on what might be causing this?
>>
>> Hi Paul!
>>
>> Can you share which tree and branch (commit) this is running ?
>>
>> Does it include my fix from this series ? That would be patch 1 of:
>>
>> https://lore.kernel.org/lkml/20260927155134.4740-1-mathieu.desnoyers@efficios.com/
>
> Ah, no, I somehow got the impression that this series was going to be
> updated, so did not apply it. Apologies!
>
> I will pull these in.
About the patches introducing ptr_eq(), I plan to handle Linus' feedback
perhaps in an additional patch and have the "generic" ptr_eq()
implemented under asm-generic, and a x86-specific implementation of it
in assembly. Are you OK with this being a follow up or would you prefer
I fold this in the original patch ?
>
>> Does it run the additional series from Kunwu Chan found at
>> https://lore.kernel.org/lkml/20261002170847.3653663-1-kunwu.chan@gmail.com/ ?
>
> No, but if you are good with it I will pull it in.
>
> May I add your Acked-by or Reviewed-by?
With LPC going on, I did not have time to review it, sorry. Not yet
please.
Thanks,
Mathieu
>
> Thanx, Paul
>
>> Thanks,
>>
>> Mathieu
>>
>>>
>>> For that matter, can anyone else reproduce this?
>>>
>>> Thanx, Paul
>>
>>
>> --
>> Mathieu Desnoyers
>> EfficiOS Inc.
>> https://www.efficios.com
--
Mathieu Desnoyers
EfficiOS Inc.
https://www.efficios.com
^ permalink raw reply [flat|nested] 9+ messages in thread* Re: [BUG] Splat during modest hazptrtorture run
2026-10-07 16:03 ` Mathieu Desnoyers
@ 2026-10-07 16:24 ` Paul E. McKenney
0 siblings, 0 replies; 9+ messages in thread
From: Paul E. McKenney @ 2026-10-07 16:24 UTC (permalink / raw)
To: Mathieu Desnoyers
Cc: Boqun Feng, Kunwu Chan, Wang Lian, Zqiang, Bradley Morgan, rcu,
lkmm, linux-kernel, Joel Fernandes
On Wed, Oct 07, 2026 at 12:03:37PM -0400, Mathieu Desnoyers wrote:
> On 2026-10-07 11:59, Paul E. McKenney wrote:
> > On Wed, Oct 07, 2026 at 03:09:27AM -0400, Mathieu Desnoyers wrote:
> > > On 2026-10-06 19:40, Paul E. McKenney wrote:
> > > > Hello!
> > > >
> > > > This is all new code, so the bug could be anywhere. So you all need
> > > > to know. ;-)
> > > >
> > > > I got this from the NOPREEMPT variant of hazptr during a nominal
> > > > "--duration 60" run of torture.sh on x86 with the guest OS split across
> > > > two NUMA nodes:
> > > >
> > > > [ 137.765104] WARNING: kernel/rcu/hazptrtorture.c:658 at hazptr_torture_stats_print+0x25b/0x4f0, CPU#1: hazptr_torture_/136
> > > > [ 137.768647] Modules linked in:
> > > > [ 137.768891] CPU: 1 UID: 0 PID: 136 Comm: hazptr_torture_ Not tainted 7.3.0-rc5-00661-ga1f1e4890d1d-dirty #193 PREEMPTLAZY
> > > > [ 137.769683] Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-6.el9 11/05/2023
> > > > [ 137.770319] RIP: 0010:hazptr_torture_stats_print+0x25b/0x4f0
> > > > [ 137.770735] Code: 48 83 c4 20 e8 86 44 fe ff 41 83 fd 01 7e 23 48 c7 c6 6e 85 17 b0 48 c7 c7 db 27 17 b0 e8 6d 44 fe ff f0 ff 05 76 6e 0a 02 90 <0f> 0b 90 4c 8b 64 24 70 48 c7 c7 73 85 17 b0 e8 51 44 fe ff 48 8b
> > > > [ 137.772061] RSP: 0018:ffffa738004ffdf0 EFLAGS: 00010202
> > > > [ 137.772435] RAX: 0000000000000004 RBX: ffffa738004ffe08 RCX: 0000000000000027
> > > > [ 137.772950] RDX: 0000000000000000 RSI: 0000000000000000 RDI: 0000000000000001
> > > > [ 137.773455] RBP: ffffa738004ffe60 R08: 4000000000001236 R09: ffffffffb01727df
> > > > [ 137.773980] R10: 0000000020212121 R11: 0000000020212121 R12: 0000000000000000
> > > > [ 137.774479] R13: 0000000000000002 R14: 000000000002ffd6 R15: 0000000000034f03
> > > > [ 137.775000] FS: 0000000000000000(0000) GS:ffff999c2e502000(0000) knlGS:0000000000000000
> > > > [ 137.775578] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> > > > [ 137.776001] CR2: 0000000000000000 CR3: 0000000008830000 CR4: 00000000000006f0
> > > > [ 137.776512] Call Trace:
> > > > [ 137.776700] <TASK>
> > > > [ 137.776863] ? __pfx_hazptr_torture_stats+0x10/0x10
> > > > [ 137.777207] hazptr_torture_stats+0x25/0x70
> > > > [ 137.777508] kthread+0xd9/0x110
> > > > [ 137.777751] ? __pfx_kthread+0x10/0x10
> > > > [ 137.778017] ret_from_fork+0x1bd/0x220
> > > > [ 137.778285] ? __pfx_kthread+0x10/0x10
> > > > [ 137.778560] ret_from_fork_asm+0x1a/0x30
> > > > [ 137.778849] </TASK>
> > > > [ 137.779007] ---[ end trace 0000000000000000 ]---
> > > >
> > > > I have not yet seen this during a PREEMPT run. Line 658 of
> > > > hazptrtorture.c is the WARN_ON_ONCE() below:
> > > >
> > > > if (i > 1) {
> > > > pr_cont("%s", "!!! ");
> > > > atomic_inc(&n_hazptr_torture_error);
> > > > WARN_ON_ONCE(i > 1); // Too-short grace period
> > > > }
> > > >
> > > > Any thoughts on what might be causing this?
> > >
> > > Hi Paul!
> > >
> > > Can you share which tree and branch (commit) this is running ?
> > >
> > > Does it include my fix from this series ? That would be patch 1 of:
> > >
> > > https://lore.kernel.org/lkml/20260927155134.4740-1-mathieu.desnoyers@efficios.com/
> >
> > Ah, no, I somehow got the impression that this series was going to be
> > updated, so did not apply it. Apologies!
> >
> > I will pull these in.
>
> About the patches introducing ptr_eq(), I plan to handle Linus' feedback
> perhaps in an additional patch and have the "generic" ptr_eq()
> implemented under asm-generic, and a x86-specific implementation of it
> in assembly. Are you OK with this being a follow up or would you prefer
> I fold this in the original patch ?
So I did get the right impression, then! Don't worry, it won't happen
again. ;-)
Ah, and my memory is also colored by the fact that we really need to
dispense with the wildcards altogether to avoid inflicting one of the
weaknesses of RCU (susceptibility to delay) onto hazard pointers.
Please fold the fix into the original. I will pull in your patch 1/1
in as an EXP patch for testing purposes in the meantime, just in case
further events affect 1/1 as well as 1/2.
> > > Does it run the additional series from Kunwu Chan found at
> > > https://lore.kernel.org/lkml/20261002170847.3653663-1-kunwu.chan@gmail.com/ ?
> >
> > No, but if you are good with it I will pull it in.
> >
> > May I add your Acked-by or Reviewed-by?
>
> With LPC going on, I did not have time to review it, sorry. Not yet
> please.
No problem! Given that hazard pointers isn't likely to make the v7.4
merge window, we do have time.
Thanx, Paul
> Thanks,
>
> Mathieu
>
> >
> > Thanx, Paul
> >
> > > Thanks,
> > >
> > > Mathieu
> > >
> > > >
> > > > For that matter, can anyone else reproduce this?
> > > >
> > > > Thanx, Paul
> > >
> > >
> > > --
> > > Mathieu Desnoyers
> > > EfficiOS Inc.
> > > https://www.efficios.com
>
>
> --
> Mathieu Desnoyers
> EfficiOS Inc.
> https://www.efficios.com
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [BUG] Splat during modest hazptrtorture run
2026-10-07 15:59 ` Paul E. McKenney
2026-10-07 16:03 ` Mathieu Desnoyers
@ 2026-10-07 16:08 ` Bradley Morgan
2026-10-08 1:39 ` KunWu Chan
2 siblings, 0 replies; 9+ messages in thread
From: Bradley Morgan @ 2026-10-07 16:08 UTC (permalink / raw)
To: paulmck, Paul E. McKenney, Mathieu Desnoyers
Cc: Boqun Feng, Kunwu Chan, Wang Lian, Zqiang, rcu, lkmm,
linux-kernel, Joel Fernandes
On 7 October 2026 16:59:36 BST, "Paul E. McKenney" <paulmck@kernel.org>
wrote:
>On Wed, Oct 07, 2026 at 03:09:27AM -0400, Mathieu Desnoyers wrote:
>> On 2026-10-06 19:40, Paul E. McKenney wrote:
>> > Hello!
>> >
>> > This is all new code, so the bug could be anywhere. So you all need
>> > to know. ;-)
>> >
>> > I got this from the NOPREEMPT variant of hazptr during a nominal
>> > "--duration 60" run of torture.sh on x86 with the guest OS split
>across
>> > two NUMA nodes:
>> >
>> > [ 137.765104] WARNING: kernel/rcu/hazptrtorture.c:658 at
>hazptr_torture_stats_print+0x25b/0x4f0, CPU#1: hazptr_torture_/136
>> > [ 137.768647] Modules linked in:
>> > [ 137.768891] CPU: 1 UID: 0 PID: 136 Comm: hazptr_torture_ Not
>tainted 7.3.0-rc5-00661-ga1f1e4890d1d-dirty #193 PREEMPTLAZY
>> > [ 137.769683] Hardware name: QEMU Standard PC (Q35 + ICH9, 2009),
>BIOS 1.16.3-6.el9 11/05/2023
>> > [ 137.770319] RIP: 0010:hazptr_torture_stats_print+0x25b/0x4f0
>> > [ 137.770735] Code: 48 83 c4 20 e8 86 44 fe ff 41 83 fd 01 7e 23 48
>c7 c6 6e 85 17 b0 48 c7 c7 db 27 17 b0 e8 6d 44 fe ff f0 ff 05 76 6e 0a 02
>90 <0f> 0b 90 4c 8b 64 24 70 48 c7 c7 73 85 17 b0 e8 51 44 fe ff 48 8b
>> > [ 137.772061] RSP: 0018:ffffa738004ffdf0 EFLAGS: 00010202
>> > [ 137.772435] RAX: 0000000000000004 RBX: ffffa738004ffe08 RCX:
>0000000000000027
>> > [ 137.772950] RDX: 0000000000000000 RSI: 0000000000000000 RDI:
>0000000000000001
>> > [ 137.773455] RBP: ffffa738004ffe60 R08: 4000000000001236 R09:
>ffffffffb01727df
>> > [ 137.773980] R10: 0000000020212121 R11: 0000000020212121 R12:
>0000000000000000
>> > [ 137.774479] R13: 0000000000000002 R14: 000000000002ffd6 R15:
>0000000000034f03
>> > [ 137.775000] FS: 0000000000000000(0000) GS:ffff999c2e502000(0000)
>knlGS:0000000000000000
>> > [ 137.775578] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
>> > [ 137.776001] CR2: 0000000000000000 CR3: 0000000008830000 CR4:
>00000000000006f0
>> > [ 137.776512] Call Trace:
>> > [ 137.776700] <TASK>
>> > [ 137.776863] ? __pfx_hazptr_torture_stats+0x10/0x10
>> > [ 137.777207] hazptr_torture_stats+0x25/0x70
>> > [ 137.777508] kthread+0xd9/0x110
>> > [ 137.777751] ? __pfx_kthread+0x10/0x10
>> > [ 137.778017] ret_from_fork+0x1bd/0x220
>> > [ 137.778285] ? __pfx_kthread+0x10/0x10
>> > [ 137.778560] ret_from_fork_asm+0x1a/0x30
>> > [ 137.778849] </TASK>
>> > [ 137.779007] ---[ end trace 0000000000000000 ]---
>> >
>> > I have not yet seen this during a PREEMPT run. Line 658 of
>> > hazptrtorture.c is the WARN_ON_ONCE() below:
>> >
>> > if (i > 1) {
>> > pr_cont("%s", "!!! ");
>> > atomic_inc(&n_hazptr_torture_error);
>> > WARN_ON_ONCE(i > 1); // Too-short grace period
>> > }
>> >
>> > Any thoughts on what might be causing this?
>>
>> Hi Paul!
>>
>> Can you share which tree and branch (commit) this is running ?
>>
>> Does it include my fix from this series ? That would be patch 1 of:
>>
>>
>https://lore.kernel.org/lkml/20260927155134.4740-1-mathieu.desnoyers@efficios.com/
>
>Ah, no, I somehow got the impression that this series was going to be
>updated, so did not apply it. Apologies!
>
>I will pull these in.
>
>> Does it run the additional series from Kunwu Chan found at
>>
>https://lore.kernel.org/lkml/20261002170847.3653663-1-kunwu.chan@gmail.com/
>?
>
>No, but if you are good with it I will pull it in.
>
>May I add your Acked-by or Reviewed-by?
Sorry to impede but I'd be fine with it going in, I'll add my A-B for this:
Acked-by: Bradley Morgan <brads@mainlining.org>
>
> Thanx, Paul
>
>> Thanks,
>>
>> Mathieu
>>
>> >
>> > For that matter, can anyone else reproduce this?
>> >
>> > Thanx, Paul
>>
>>
>> --
>> Mathieu Desnoyers
>> EfficiOS Inc.
>> https://www.efficios.com
--- Thanks!
"I'm not a very positive person" - Linus torvalds
^ permalink raw reply [flat|nested] 9+ messages in thread* Re: [BUG] Splat during modest hazptrtorture run
2026-10-07 15:59 ` Paul E. McKenney
2026-10-07 16:03 ` Mathieu Desnoyers
2026-10-07 16:08 ` Bradley Morgan
@ 2026-10-08 1:39 ` KunWu Chan
2026-10-08 3:53 ` Paul E. McKenney
2 siblings, 1 reply; 9+ messages in thread
From: KunWu Chan @ 2026-10-08 1:39 UTC (permalink / raw)
To: paulmck
Cc: Mathieu Desnoyers, Boqun Feng, Wang Lian, Zqiang, Bradley Morgan,
rcu, lkmm, linux-kernel, Joel Fernandes
On Wed, Oct 7, 2026 at 11:59 PM Paul E. McKenney <paulmck@kernel.org> wrote:
>
> On Wed, Oct 07, 2026 at 03:09:27AM -0400, Mathieu Desnoyers wrote:
> > On 2026-10-06 19:40, Paul E. McKenney wrote:
> > > Hello!
> > >
> > > This is all new code, so the bug could be anywhere. So you all need
> > > to know. ;-)
> > >
> > > I got this from the NOPREEMPT variant of hazptr during a nominal
> > > "--duration 60" run of torture.sh on x86 with the guest OS split across
> > > two NUMA nodes:
> > >
> > > [ 137.765104] WARNING: kernel/rcu/hazptrtorture.c:658 at hazptr_torture_stats_print+0x25b/0x4f0, CPU#1: hazptr_torture_/136
> > > [ 137.768647] Modules linked in:
> > > [ 137.768891] CPU: 1 UID: 0 PID: 136 Comm: hazptr_torture_ Not tainted 7.3.0-rc5-00661-ga1f1e4890d1d-dirty #193 PREEMPTLAZY
> > > [ 137.769683] Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-6.el9 11/05/2023
> > > [ 137.770319] RIP: 0010:hazptr_torture_stats_print+0x25b/0x4f0
> > > [ 137.770735] Code: 48 83 c4 20 e8 86 44 fe ff 41 83 fd 01 7e 23 48 c7 c6 6e 85 17 b0 48 c7 c7 db 27 17 b0 e8 6d 44 fe ff f0 ff 05 76 6e 0a 02 90 <0f> 0b 90 4c 8b 64 24 70 48 c7 c7 73 85 17 b0 e8 51 44 fe ff 48 8b
> > > [ 137.772061] RSP: 0018:ffffa738004ffdf0 EFLAGS: 00010202
> > > [ 137.772435] RAX: 0000000000000004 RBX: ffffa738004ffe08 RCX: 0000000000000027
> > > [ 137.772950] RDX: 0000000000000000 RSI: 0000000000000000 RDI: 0000000000000001
> > > [ 137.773455] RBP: ffffa738004ffe60 R08: 4000000000001236 R09: ffffffffb01727df
> > > [ 137.773980] R10: 0000000020212121 R11: 0000000020212121 R12: 0000000000000000
> > > [ 137.774479] R13: 0000000000000002 R14: 000000000002ffd6 R15: 0000000000034f03
> > > [ 137.775000] FS: 0000000000000000(0000) GS:ffff999c2e502000(0000) knlGS:0000000000000000
> > > [ 137.775578] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> > > [ 137.776001] CR2: 0000000000000000 CR3: 0000000008830000 CR4: 00000000000006f0
> > > [ 137.776512] Call Trace:
> > > [ 137.776700] <TASK>
> > > [ 137.776863] ? __pfx_hazptr_torture_stats+0x10/0x10
> > > [ 137.777207] hazptr_torture_stats+0x25/0x70
> > > [ 137.777508] kthread+0xd9/0x110
> > > [ 137.777751] ? __pfx_kthread+0x10/0x10
> > > [ 137.778017] ret_from_fork+0x1bd/0x220
> > > [ 137.778285] ? __pfx_kthread+0x10/0x10
> > > [ 137.778560] ret_from_fork_asm+0x1a/0x30
> > > [ 137.778849] </TASK>
> > > [ 137.779007] ---[ end trace 0000000000000000 ]---
> > >
> > > I have not yet seen this during a PREEMPT run. Line 658 of
> > > hazptrtorture.c is the WARN_ON_ONCE() below:
> > >
> > > if (i > 1) {
> > > pr_cont("%s", "!!! ");
> > > atomic_inc(&n_hazptr_torture_error);
> > > WARN_ON_ONCE(i > 1); // Too-short grace period
> > > }
> > >
> > > Any thoughts on what might be causing this?
> >
> > Hi Paul!
> >
> > Can you share which tree and branch (commit) this is running ?
> >
> > Does it include my fix from this series ? That would be patch 1 of:
> >
> > https://lore.kernel.org/lkml/20260927155134.4740-1-mathieu.desnoyers@efficios.com/
>
> Ah, no, I somehow got the impression that this series was going to be
> updated, so did not apply it. Apologies!
>
> I will pull these in.
>
> > Does it run the additional series from Kunwu Chan found at
> > https://lore.kernel.org/lkml/20261002170847.3653663-1-kunwu.chan@gmail.com/ ?
>
Hi Paul,
I have posted v3 of the series since then. The link in Mathieu's
question points to v2.
Please use the latest v3 series here:
https://lore.kernel.org/all/20261005171529.1378809-1-kunwu.chan@gmail.com/
I wasn't able to reproduce the warning with my current setup(without
my series).
Could you share the .config and the exact command line you used to run the
test? That may help me reproduce it.
Thanks,
Kunwu
> No, but if you are good with it I will pull it in.
>
> May I add your Acked-by or Reviewed-by?
>
> Thanx, Paul
>
> > Thanks,
> >
> > Mathieu
> >
> > >
> > > For that matter, can anyone else reproduce this?
> > >
> > > Thanx, Paul
> >
> >
> > --
> > Mathieu Desnoyers
> > EfficiOS Inc.
> > https://www.efficios.com
^ permalink raw reply [flat|nested] 9+ messages in thread* Re: [BUG] Splat during modest hazptrtorture run
2026-10-08 1:39 ` KunWu Chan
@ 2026-10-08 3:53 ` Paul E. McKenney
2026-10-08 4:15 ` KunWu Chan
0 siblings, 1 reply; 9+ messages in thread
From: Paul E. McKenney @ 2026-10-08 3:53 UTC (permalink / raw)
To: KunWu Chan
Cc: Mathieu Desnoyers, Boqun Feng, Wang Lian, Zqiang, Bradley Morgan,
rcu, lkmm, linux-kernel, Joel Fernandes
On Thu, Oct 08, 2026 at 09:39:50AM +0800, KunWu Chan wrote:
> On Wed, Oct 7, 2026 at 11:59 PM Paul E. McKenney <paulmck@kernel.org> wrote:
> >
> > On Wed, Oct 07, 2026 at 03:09:27AM -0400, Mathieu Desnoyers wrote:
> > > On 2026-10-06 19:40, Paul E. McKenney wrote:
> > > > Hello!
> > > >
> > > > This is all new code, so the bug could be anywhere. So you all need
> > > > to know. ;-)
TL;DR: Adding the first patch in Mathieu's four-patch series cures this
for me.
> > > > I got this from the NOPREEMPT variant of hazptr during a nominal
> > > > "--duration 60" run of torture.sh on x86 with the guest OS split across
> > > > two NUMA nodes:
> > > >
> > > > [ 137.765104] WARNING: kernel/rcu/hazptrtorture.c:658 at hazptr_torture_stats_print+0x25b/0x4f0, CPU#1: hazptr_torture_/136
> > > > [ 137.768647] Modules linked in:
> > > > [ 137.768891] CPU: 1 UID: 0 PID: 136 Comm: hazptr_torture_ Not tainted 7.3.0-rc5-00661-ga1f1e4890d1d-dirty #193 PREEMPTLAZY
> > > > [ 137.769683] Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-6.el9 11/05/2023
> > > > [ 137.770319] RIP: 0010:hazptr_torture_stats_print+0x25b/0x4f0
> > > > [ 137.770735] Code: 48 83 c4 20 e8 86 44 fe ff 41 83 fd 01 7e 23 48 c7 c6 6e 85 17 b0 48 c7 c7 db 27 17 b0 e8 6d 44 fe ff f0 ff 05 76 6e 0a 02 90 <0f> 0b 90 4c 8b 64 24 70 48 c7 c7 73 85 17 b0 e8 51 44 fe ff 48 8b
> > > > [ 137.772061] RSP: 0018:ffffa738004ffdf0 EFLAGS: 00010202
> > > > [ 137.772435] RAX: 0000000000000004 RBX: ffffa738004ffe08 RCX: 0000000000000027
> > > > [ 137.772950] RDX: 0000000000000000 RSI: 0000000000000000 RDI: 0000000000000001
> > > > [ 137.773455] RBP: ffffa738004ffe60 R08: 4000000000001236 R09: ffffffffb01727df
> > > > [ 137.773980] R10: 0000000020212121 R11: 0000000020212121 R12: 0000000000000000
> > > > [ 137.774479] R13: 0000000000000002 R14: 000000000002ffd6 R15: 0000000000034f03
> > > > [ 137.775000] FS: 0000000000000000(0000) GS:ffff999c2e502000(0000) knlGS:0000000000000000
> > > > [ 137.775578] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> > > > [ 137.776001] CR2: 0000000000000000 CR3: 0000000008830000 CR4: 00000000000006f0
> > > > [ 137.776512] Call Trace:
> > > > [ 137.776700] <TASK>
> > > > [ 137.776863] ? __pfx_hazptr_torture_stats+0x10/0x10
> > > > [ 137.777207] hazptr_torture_stats+0x25/0x70
> > > > [ 137.777508] kthread+0xd9/0x110
> > > > [ 137.777751] ? __pfx_kthread+0x10/0x10
> > > > [ 137.778017] ret_from_fork+0x1bd/0x220
> > > > [ 137.778285] ? __pfx_kthread+0x10/0x10
> > > > [ 137.778560] ret_from_fork_asm+0x1a/0x30
> > > > [ 137.778849] </TASK>
> > > > [ 137.779007] ---[ end trace 0000000000000000 ]---
> > > >
> > > > I have not yet seen this during a PREEMPT run. Line 658 of
> > > > hazptrtorture.c is the WARN_ON_ONCE() below:
> > > >
> > > > if (i > 1) {
> > > > pr_cont("%s", "!!! ");
> > > > atomic_inc(&n_hazptr_torture_error);
> > > > WARN_ON_ONCE(i > 1); // Too-short grace period
> > > > }
> > > >
> > > > Any thoughts on what might be causing this?
> > >
> > > Hi Paul!
> > >
> > > Can you share which tree and branch (commit) this is running ?
> > >
> > > Does it include my fix from this series ? That would be patch 1 of:
> > >
> > > https://lore.kernel.org/lkml/20260927155134.4740-1-mathieu.desnoyers@efficios.com/
> >
> > Ah, no, I somehow got the impression that this series was going to be
> > updated, so did not apply it. Apologies!
> >
> > I will pull these in.
> >
> > > Does it run the additional series from Kunwu Chan found at
> > > https://lore.kernel.org/lkml/20261002170847.3653663-1-kunwu.chan@gmail.com/ ?
> >
>
> Hi Paul,
>
> I have posted v3 of the series since then. The link in Mathieu's
> question points to v2.
>
> Please use the latest v3 series here:
> https://lore.kernel.org/all/20261005171529.1378809-1-kunwu.chan@gmail.com/
I do it the old way with out b4, so I would have bot there. I will
hold off a bit for merge-window preparation, though.
> I wasn't able to reproduce the warning with my current setup(without
> my series).
> Could you share the .config and the exact command line you used to run the
> test? That may help me reproduce it.
I used this command line, which creates the .config:
tools/testing/selftests/rcutorture/bin/kvm.sh --torture hazptr --allcpus --configs "2*CFLIST" --duration 90m --trust-make
This was on an ARM system, in case that matters.
But pulling in Mathieu's first patch took care of this issue for me.
Thanx, Paul
> Thanks,
> Kunwu
>
> > No, but if you are good with it I will pull it in.
> >
> > May I add your Acked-by or Reviewed-by?
> >
> > Thanx, Paul
> >
> > > Thanks,
> > >
> > > Mathieu
> > >
> > > >
> > > > For that matter, can anyone else reproduce this?
> > > >
> > > > Thanx, Paul
> > >
> > >
> > > --
> > > Mathieu Desnoyers
> > > EfficiOS Inc.
> > > https://www.efficios.com
^ permalink raw reply [flat|nested] 9+ messages in thread* Re: [BUG] Splat during modest hazptrtorture run
2026-10-08 3:53 ` Paul E. McKenney
@ 2026-10-08 4:15 ` KunWu Chan
0 siblings, 0 replies; 9+ messages in thread
From: KunWu Chan @ 2026-10-08 4:15 UTC (permalink / raw)
To: paulmck
Cc: Mathieu Desnoyers, Boqun Feng, Wang Lian, Zqiang, Bradley Morgan,
rcu, lkmm, linux-kernel, Joel Fernandes
On Thu, Oct 8, 2026 at 11:53 AM Paul E. McKenney <paulmck@kernel.org> wrote:
>
> On Thu, Oct 08, 2026 at 09:39:50AM +0800, KunWu Chan wrote:
> > On Wed, Oct 7, 2026 at 11:59 PM Paul E. McKenney <paulmck@kernel.org> wrote:
> > >
> > > On Wed, Oct 07, 2026 at 03:09:27AM -0400, Mathieu Desnoyers wrote:
> > > > On 2026-10-06 19:40, Paul E. McKenney wrote:
> > > > > Hello!
> > > > >
> > > > > This is all new code, so the bug could be anywhere. So you all need
> > > > > to know. ;-)
>
> TL;DR: Adding the first patch in Mathieu's four-patch series cures this
> for me.
>
> > > > > I got this from the NOPREEMPT variant of hazptr during a nominal
> > > > > "--duration 60" run of torture.sh on x86 with the guest OS split across
> > > > > two NUMA nodes:
> > > > >
> > > > > [ 137.765104] WARNING: kernel/rcu/hazptrtorture.c:658 at hazptr_torture_stats_print+0x25b/0x4f0, CPU#1: hazptr_torture_/136
> > > > > [ 137.768647] Modules linked in:
> > > > > [ 137.768891] CPU: 1 UID: 0 PID: 136 Comm: hazptr_torture_ Not tainted 7.3.0-rc5-00661-ga1f1e4890d1d-dirty #193 PREEMPTLAZY
> > > > > [ 137.769683] Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-6.el9 11/05/2023
> > > > > [ 137.770319] RIP: 0010:hazptr_torture_stats_print+0x25b/0x4f0
> > > > > [ 137.770735] Code: 48 83 c4 20 e8 86 44 fe ff 41 83 fd 01 7e 23 48 c7 c6 6e 85 17 b0 48 c7 c7 db 27 17 b0 e8 6d 44 fe ff f0 ff 05 76 6e 0a 02 90 <0f> 0b 90 4c 8b 64 24 70 48 c7 c7 73 85 17 b0 e8 51 44 fe ff 48 8b
> > > > > [ 137.772061] RSP: 0018:ffffa738004ffdf0 EFLAGS: 00010202
> > > > > [ 137.772435] RAX: 0000000000000004 RBX: ffffa738004ffe08 RCX: 0000000000000027
> > > > > [ 137.772950] RDX: 0000000000000000 RSI: 0000000000000000 RDI: 0000000000000001
> > > > > [ 137.773455] RBP: ffffa738004ffe60 R08: 4000000000001236 R09: ffffffffb01727df
> > > > > [ 137.773980] R10: 0000000020212121 R11: 0000000020212121 R12: 0000000000000000
> > > > > [ 137.774479] R13: 0000000000000002 R14: 000000000002ffd6 R15: 0000000000034f03
> > > > > [ 137.775000] FS: 0000000000000000(0000) GS:ffff999c2e502000(0000) knlGS:0000000000000000
> > > > > [ 137.775578] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> > > > > [ 137.776001] CR2: 0000000000000000 CR3: 0000000008830000 CR4: 00000000000006f0
> > > > > [ 137.776512] Call Trace:
> > > > > [ 137.776700] <TASK>
> > > > > [ 137.776863] ? __pfx_hazptr_torture_stats+0x10/0x10
> > > > > [ 137.777207] hazptr_torture_stats+0x25/0x70
> > > > > [ 137.777508] kthread+0xd9/0x110
> > > > > [ 137.777751] ? __pfx_kthread+0x10/0x10
> > > > > [ 137.778017] ret_from_fork+0x1bd/0x220
> > > > > [ 137.778285] ? __pfx_kthread+0x10/0x10
> > > > > [ 137.778560] ret_from_fork_asm+0x1a/0x30
> > > > > [ 137.778849] </TASK>
> > > > > [ 137.779007] ---[ end trace 0000000000000000 ]---
> > > > >
> > > > > I have not yet seen this during a PREEMPT run. Line 658 of
> > > > > hazptrtorture.c is the WARN_ON_ONCE() below:
> > > > >
> > > > > if (i > 1) {
> > > > > pr_cont("%s", "!!! ");
> > > > > atomic_inc(&n_hazptr_torture_error);
> > > > > WARN_ON_ONCE(i > 1); // Too-short grace period
> > > > > }
> > > > >
> > > > > Any thoughts on what might be causing this?
> > > >
> > > > Hi Paul!
> > > >
> > > > Can you share which tree and branch (commit) this is running ?
> > > >
> > > > Does it include my fix from this series ? That would be patch 1 of:
> > > >
> > > > https://lore.kernel.org/lkml/20260927155134.4740-1-mathieu.desnoyers@efficios.com/
> > >
> > > Ah, no, I somehow got the impression that this series was going to be
> > > updated, so did not apply it. Apologies!
> > >
> > > I will pull these in.
> > >
> > > > Does it run the additional series from Kunwu Chan found at
> > > > https://lore.kernel.org/lkml/20261002170847.3653663-1-kunwu.chan@gmail.com/ ?
> > >
> >
> > Hi Paul,
> >
> > I have posted v3 of the series since then. The link in Mathieu's
> > question points to v2.
> >
> > Please use the latest v3 series here:
> > https://lore.kernel.org/all/20261005171529.1378809-1-kunwu.chan@gmail.com/
>
> I do it the old way with out b4, so I would have bot there. I will
> hold off a bit for merge-window preparation, though.
>
> > I wasn't able to reproduce the warning with my current setup(without
> > my series).
> > Could you share the .config and the exact command line you used to run the
> > test? That may help me reproduce it.
>
> I used this command line, which creates the .config:
>
> tools/testing/selftests/rcutorture/bin/kvm.sh --torture hazptr --allcpus --configs "2*CFLIST" --duration 90m --trust-make
>
> This was on an ARM system, in case that matters.
>
> But pulling in Mathieu's first patch took care of this issue for me.
>
Hi Paul,
Good to hear Mathieu's first patch fixed it. That is consistent with
the race we were looking at.
The race is between the wildcard flip and a context-switch promotion:
hazptr_chain_backup_slot() selected the overflow list based on the
wildcard phase, so a backup slot promoted during the second scan could
end up in a list that was no longer covered by that scan.
Thanks for sharing the command line and test setup. The 90-minute run
also explains why it was difficult for me to reproduce with shorter
runs.
Thanks,
KunWu
> Thanx, Paul
>
> > Thanks,
> > Kunwu
> >
> > > No, but if you are good with it I will pull it in.
> > >
> > > May I add your Acked-by or Reviewed-by?
> > >
> > > Thanx, Paul
> > >
> > > > Thanks,
> > > >
> > > > Mathieu
> > > >
> > > > >
> > > > > For that matter, can anyone else reproduce this?
> > > > >
> > > > > Thanx, Paul
> > > >
> > > >
> > > > --
> > > > Mathieu Desnoyers
> > > > EfficiOS Inc.
> > > > https://www.efficios.com
^ permalink raw reply [flat|nested] 9+ messages in thread