From: "Huang, Kai" <kai.huang@intel.com>
To: "tglx@linutronix.de" <tglx@linutronix.de>,
"peterz@infradead.org" <peterz@infradead.org>,
"pbonzini@redhat.com" <pbonzini@redhat.com>,
"Hansen, Dave" <dave.hansen@intel.com>,
"mingo@redhat.com" <mingo@redhat.com>,
"bp@alien8.de" <bp@alien8.de>
Cc: "Edgecombe, Rick P" <rick.p.edgecombe@intel.com>,
"seanjc@google.com" <seanjc@google.com>,
"x86@kernel.org" <x86@kernel.org>,
"sagis@google.com" <sagis@google.com>,
"hpa@zytor.com" <hpa@zytor.com>,
"Chatre, Reinette" <reinette.chatre@intel.com>,
"kirill.shutemov@linux.intel.com"
<kirill.shutemov@linux.intel.com>,
"Williams, Dan J" <dan.j.williams@intel.com>,
"linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>,
"thomas.lendacky@amd.com" <thomas.lendacky@amd.com>,
"Yamahata, Isaku" <isaku.yamahata@intel.com>,
"ashish.kalra@amd.com" <ashish.kalra@amd.com>,
"nik.borisov@suse.com" <nik.borisov@suse.com>
Subject: Re: [RFC PATCH v2 6/5] KVM: TDX: Explicitly do WBINVD upon reboot notifier
Date: Mon, 16 Jun 2025 10:39:24 +0000 [thread overview]
Message-ID: <8327b8ebab52a7b38431169193e7281e92e9bf01.camel@intel.com> (raw)
In-Reply-To: <3a7c0856-6e7b-4d3d-b966-6f17f1aca42e@redhat.com>
On Fri, 2025-06-13 at 13:36 +0200, Paolo Bonzini wrote:
> On 5/10/25 13:25, Kai Huang wrote:
> > On TDX platforms, during kexec, the kernel needs to make sure there's no
> > dirty cachelines of TDX private memory before booting to the new kernel
> > to avoid silent memory corruption to the new kernel.
> >
> > During kexec, the kexec-ing CPU firstly invokes native_stop_other_cpus()
> > to stop all remote CPUs before booting to the new kernel. The remote
> > CPUs will then execute stop_this_cpu() to stop themselves.
> >
> > The kernel has a percpu boolean to indicate whether the cache of a CPU
> > may be in incoherent state. In stop_this_cpu(), the kernel does WBINVD
> > if that percpu boolean is true.
> >
> > TDX turns on that percpu boolean on a CPU when the kernel does SEAMCALL.
> > This makes sure the cahces will be flushed during kexec.
> >
> > However, the native_stop_other_cpus() and stop_this_cpu() have a "race"
> > which is extremely rare to happen but if did could cause system to hang.
>
> s/if did//
>
> > Specifically, the native_stop_other_cpus() firstly sends normal reboot
> > IPI to remote CPUs and wait one second for them to stop. If that times
> > out, native_stop_other_cpus() then sends NMIs to remote CPUs to stop
> > them.
> >
> > The aforementioned race happens when NMIs are sent. Doing WBINVD in
> > stop_this_cpu() makes each CPU take longer time to stop and increases
> > the chance of the race to happen.
> >
> > Register reboot notifier in KVM to explcitly flush caches upon reboot
> > for TDX. This brings doing WBINVD at earlier stage and aovids the
> > WBINVD in stop_this_cpu(), eliminating the possibility of increasing the
> > chance of the aforementioned race.
>
> "This moves the WBINVD to an earlier stage than stop_this_cpus(),
> avoiding a possibly lengthy operation at a time where it could cause
> this race."
>
> Acked-by: Paolo Bonzini <pbonzini@redhat.com>
>
> Waiting for v3. :)
Thanks Paolo, and will do!
next prev parent reply other threads:[~2025-06-16 10:39 UTC|newest]
Thread overview: 15+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-05-10 11:20 [PATCH v2 0/5] TDX host: kexec/kdump support Kai Huang
2025-05-10 11:20 ` [PATCH v2 1/5] x86/sme: Use percpu boolean to control wbinvd during kexec Kai Huang
2025-05-10 11:20 ` [PATCH v2 2/5] x86/virt/tdx: Mark memory cache state incoherent when making SEAMCALL Kai Huang
2025-05-14 10:10 ` [PATCH v2.1 " Kai Huang
2025-05-14 15:52 ` Dave Hansen
2025-05-14 22:13 ` Huang, Kai
2025-05-10 11:20 ` [PATCH v2 3/5] x86/kexec: Disable kexec/kdump on platforms with TDX partial write erratum Kai Huang
2025-05-10 11:20 ` [PATCH v2 4/5] x86/virt/tdx: Remove the !KEXEC_CORE dependency Kai Huang
2025-05-10 11:20 ` [PATCH v2 5/5] x86/virt/tdx: Update the kexec section in the TDX documentation Kai Huang
2025-05-10 11:25 ` [RFC PATCH v2 6/5] KVM: TDX: Explicitly do WBINVD upon reboot notifier Kai Huang
2025-06-13 11:36 ` Paolo Bonzini
2025-06-16 10:39 ` Huang, Kai [this message]
2025-05-14 12:37 ` [PATCH v2 0/5] TDX host: kexec/kdump support Tom Lendacky
2025-05-14 22:09 ` Huang, Kai
2025-05-29 11:06 ` Huang, Kai
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=8327b8ebab52a7b38431169193e7281e92e9bf01.camel@intel.com \
--to=kai.huang@intel.com \
--cc=ashish.kalra@amd.com \
--cc=bp@alien8.de \
--cc=dan.j.williams@intel.com \
--cc=dave.hansen@intel.com \
--cc=hpa@zytor.com \
--cc=isaku.yamahata@intel.com \
--cc=kirill.shutemov@linux.intel.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mingo@redhat.com \
--cc=nik.borisov@suse.com \
--cc=pbonzini@redhat.com \
--cc=peterz@infradead.org \
--cc=reinette.chatre@intel.com \
--cc=rick.p.edgecombe@intel.com \
--cc=sagis@google.com \
--cc=seanjc@google.com \
--cc=tglx@linutronix.de \
--cc=thomas.lendacky@amd.com \
--cc=x86@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®