mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
To: Nikola Ciprich <nikola.ciprich@linuxbox.cz>
Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org,
	 akpm@linux-foundation.org, david@kernel.org,
	Mike Rapoport <rppt@kernel.org>,
	 Dave Hansen <dave.hansen@linux.intel.com>,
	Pedro Falcato <pfalcato@suse.de>,
	 Kiryl Shutsemau <kas@kernel.org>,
	luizcap@redhat.com, pbonzini@redhat.com,
	 Borislav Petkov <bp@alien8.de>,
	Tal Zussman <tz2294@columbia.edu>,
	 Rik van Riel <riel@surriel.com>,
	Matt Fleming <matt@readmodwrite.com>
Subject: Re: hunting memory corruption bug in 6.18.x
Date: Fri, 2 Oct 2026 20:27:22 +0100	[thread overview]
Message-ID: <asAE3kLFDUx5v0Nw@gremlin> (raw)
In-Reply-To: <ar/F6+OUzj8qQ8YC@pcnci.linuxbox.cz>

On Fri, Oct 02, 2026 at 04:55:39PM +0200, Nikola Ciprich wrote:

> > The fact you've seen this on a Milan machine (EPYC 7343) is new information
> > so that could be useful for the report, so if you can reproduce _without
> > the fix_ first that'd be very useful to know!
>
> well, not so good news here.. I tried the reproducer on multiple
> lab machines and also on one drained production box on which we've
> experienced one crash and wasn't able to get a single hit so far.
> Here's the list of machines:
>
> hostname   	 CPU	   	RAM    kernel		microcode
> pocstdv1b	EPYC 9124	384G	6.18.31lb9.01	0x0a101158
> labtest		EPYC 9274F	64G	6.18.20lb9.01	0x0a101158
> lbxovav6d	EPYC 9124	1.5TB	6.18.31lb9.01	0x0a101158
> prfjazv1g	EPYC 7343	1TB	6.18.53lb9.02	0x0a0011de
> nrbphav4a	EPYC 7343	1TB	6.18.15lb9.01	0x0a0011de
>
> I'll leave it running for few hours, and report back. If I'm able to reproduce,
> I'll try recommended workarounds and report as well.

Thanks for trying that!

Yeah, it might have been tuned to the zen arch rather than milan so that could
be making it less effective unfortunately.

> (and lots of more addresses, truncated)
>
> crash> vtop 0xffff986002783a50
> VIRTUAL           PHYSICAL
> ffff986002783a50  582783a50
>
> PGD DIRECTORY: ffffffffa0836000
> PAGE DIRECTORY: 143a7801067
>    PUD: 143a7801c00 => 80000005800001e3
>   PAGE: 580000000  (1GB)
>
>       PTE         PHYSICAL   FLAGS
> 80000005800001e3  580000000  (PRESENT|RW|ACCESSED|DIRTY|PSE|GLOBAL|NX)
>
>       PAGE         PHYSICAL      MAPPING       INDEX CNT FLAGS
> fffff2715609e0c0  582783000 ffff996e5a8932d9 7f6a1589e  1 2ffff800020938 uptodate,dirty,lru,active,owner_2,swapbacked
>
> crash> search -p -m 0xfff0000000000fff 582783a50
> 3c600964f0: 8000000582783067
> 11e238a54f0: 582783e67
>
> should you need anything else, please let me know

Thanks! The LLM has informed me that it got that wrong but it was interesting
data (...!) which suggests guest memory page tables might have somehow ended up
there.

Could you try:

    crash> p d_hash_shift
    crash> eval (0x5560450b >> D) * 8 + 0xffffb18900a00000
                                          D = the d_hash_shift value
    crash> eval (B & 0xfffffffffffff000)
                                          B = the hex result above
    crash> vtop B
    crash> rd -64 P 512
                                          P = the hex result of the second eval
    crash> search -p -m 0xfff0000000000fff PHYS
                                          PHYS = the PHYSICAL column from vtop

  And for the pages the physical search found:

    crash> kmem 0x1ae234000
    crash> kmem 0x24f0e2000
    crash> kmem 0x103360000
    crash> rd -p 0x1ae234000 512

Thanks!

--
Cheers, Lorenzo

  reply	other threads:[~2026-10-02 19:27 UTC|newest]

Thread overview: 16+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-25  8:48 Nikola Ciprich
2026-09-25 10:05 ` Lorenzo Stoakes (ARM)
2026-09-25 12:13 ` Lorenzo Stoakes (ARM)
2026-09-26  5:58   ` Nikola Ciprich
2026-09-26  9:32     ` Lorenzo Stoakes (ARM)
2026-09-28  9:05       ` Nikola Ciprich
2026-09-29 19:11         ` Nikola Ciprich
2026-09-30  9:07           ` Lorenzo Stoakes (ARM)
2026-09-30 18:40             ` Nikola Ciprich
2026-10-01 19:40               ` Nikola Ciprich
2026-10-01 22:21                 ` David Laight
2026-10-02  9:50                 ` Lorenzo Stoakes (ARM)
2026-10-02 14:55                   ` Nikola Ciprich
2026-10-02 19:27                     ` Lorenzo Stoakes (ARM) [this message]
2026-09-26 16:02 ` Luiz Capitulino
2026-09-28  8:47   ` Nikola Ciprich

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=asAE3kLFDUx5v0Nw@gremlin \
    --to=ljs@kernel.org \
    --cc=akpm@linux-foundation.org \
    --cc=bp@alien8.de \
    --cc=dave.hansen@linux.intel.com \
    --cc=david@kernel.org \
    --cc=kas@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=luizcap@redhat.com \
    --cc=matt@readmodwrite.com \
    --cc=nikola.ciprich@linuxbox.cz \
    --cc=pbonzini@redhat.com \
    --cc=pfalcato@suse.de \
    --cc=riel@surriel.com \
    --cc=rppt@kernel.org \
    --cc=tz2294@columbia.edu \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®