From: Kai Huang <kai.huang@intel.com>
To: dave.hansen@intel.com, bp@alien8.de, tglx@linutronix.de,
peterz@infradead.org, mingo@redhat.com, hpa@zytor.com,
thomas.lendacky@amd.com
Cc: x86@kernel.org, kas@kernel.org, rick.p.edgecombe@intel.com,
dwmw@amazon.co.uk, linux-kernel@vger.kernel.org,
pbonzini@redhat.com, seanjc@google.com, kvm@vger.kernel.org,
reinette.chatre@intel.com, isaku.yamahata@intel.com,
dan.j.williams@intel.com, ashish.kalra@amd.com,
nik.borisov@suse.com, chao.gao@intel.com, sagis@google.com,
Farrah Chen <farrah.chen@intel.com>,
Binbin Wu <binbin.wu@linux.intel.com>
Subject: [PATCH v5 7/7] KVM: TDX: Explicitly do WBINVD when no more TDX SEAMCALLs
Date: Tue, 29 Jul 2025 00:28:41 +1200 [thread overview]
Message-ID: <c29f7a3348a95f687c83ac965ebc92ff5f253e87.1753679792.git.kai.huang@intel.com> (raw)
In-Reply-To: <cover.1753679792.git.kai.huang@intel.com>
On TDX platforms, during kexec, the kernel needs to make sure there are
no dirty cachelines of TDX private memory before booting to the new
kernel to avoid silent memory corruption to the new kernel.
During kexec, the kexec-ing CPU firstly invokes native_stop_other_cpus()
to stop all remote CPUs before booting to the new kernel. The remote
CPUs will then execute stop_this_cpu() to stop themselves.
The kernel has a percpu boolean to indicate whether the cache of a CPU
may be in incoherent state. In stop_this_cpu(), the kernel does WBINVD
if that percpu boolean is true.
TDX turns on that percpu boolean on a CPU when the kernel does SEAMCALL.
This makes sure the caches will be flushed during kexec.
However, the native_stop_other_cpus() and stop_this_cpu() have a "race"
which is extremely rare to happen but could cause the system to hang.
Specifically, the native_stop_other_cpus() firstly sends normal reboot
IPI to remote CPUs and waits one second for them to stop. If that times
out, native_stop_other_cpus() then sends NMIs to remote CPUs to stop
them.
The aforementioned race happens when NMIs are sent. Doing WBINVD in
stop_this_cpu() makes each CPU take longer time to stop and increases
the chance of the race happening.
Explicitly flush cache in tdx_disable_virtualization_cpu() after which
no more TDX activity can happen on this cpu. This moves the WBINVD to
an earlier stage than stop_this_cpus(), avoiding a possibly lengthy
operation at a time where it could cause this race.
Signed-off-by: Kai Huang <kai.huang@intel.com>
Acked-by: Paolo Bonzini <pbonzini@redhat.com>
Tested-by: Farrah Chen <farrah.chen@intel.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
---
v4 -> v5:
- No change
v3 -> v4:
- Change doing wbinvd() from rebooting notifier to
tdx_disable_virtualization_cpu() to cover the case where more
SEAMCALL can be made after cache flush, i.e., doing kexec when
there's TD alive. - Chao.
- Add check to skip wbinvd if the boolean is false. -- Chao
- Fix typo in the comment -- Binbin.
---
arch/x86/include/asm/tdx.h | 2 ++
arch/x86/kvm/vmx/tdx.c | 12 ++++++++++++
arch/x86/virt/vmx/tdx/tdx.c | 12 ++++++++++++
3 files changed, 26 insertions(+)
diff --git a/arch/x86/include/asm/tdx.h b/arch/x86/include/asm/tdx.h
index 488274959cd5..b7c978281934 100644
--- a/arch/x86/include/asm/tdx.h
+++ b/arch/x86/include/asm/tdx.h
@@ -217,6 +217,7 @@ u64 tdh_mem_page_remove(struct tdx_td *td, u64 gpa, u64 level, u64 *ext_err1, u6
u64 tdh_phymem_cache_wb(bool resume);
u64 tdh_phymem_page_wbinvd_tdr(struct tdx_td *td);
u64 tdh_phymem_page_wbinvd_hkid(u64 hkid, struct page *page);
+void tdx_cpu_flush_cache(void);
#else
static inline void tdx_init(void) { }
static inline int tdx_cpu_enable(void) { return -ENODEV; }
@@ -224,6 +225,7 @@ static inline int tdx_enable(void) { return -ENODEV; }
static inline u32 tdx_get_nr_guest_keyids(void) { return 0; }
static inline const char *tdx_dump_mce_info(struct mce *m) { return NULL; }
static inline const struct tdx_sys_info *tdx_get_sysinfo(void) { return NULL; }
+static inline void tdx_cpu_flush_cache(void) { }
#endif /* CONFIG_INTEL_TDX_HOST */
#endif /* !__ASSEMBLER__ */
diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c
index ec79aacc446f..93477233baae 100644
--- a/arch/x86/kvm/vmx/tdx.c
+++ b/arch/x86/kvm/vmx/tdx.c
@@ -442,6 +442,18 @@ void tdx_disable_virtualization_cpu(void)
tdx_flush_vp(&arg);
}
local_irq_restore(flags);
+
+ /*
+ * No more TDX activity on this CPU from here. Flush cache to
+ * avoid having to do WBINVD in stop_this_cpu() during kexec.
+ *
+ * Kexec calls native_stop_other_cpus() to stop remote CPUs
+ * before booting to new kernel, but that code has a "race"
+ * when the normal REBOOT IPI times out and NMIs are sent to
+ * remote CPUs to stop them. Doing WBINVD in stop_this_cpu()
+ * could potentially increase the possibility of the "race".
+ */
+ tdx_cpu_flush_cache();
}
#define TDX_SEAMCALL_RETRIES 10000
diff --git a/arch/x86/virt/vmx/tdx/tdx.c b/arch/x86/virt/vmx/tdx/tdx.c
index d6ee4e5a75d2..c098a6e0382b 100644
--- a/arch/x86/virt/vmx/tdx/tdx.c
+++ b/arch/x86/virt/vmx/tdx/tdx.c
@@ -1870,3 +1870,15 @@ u64 tdh_phymem_page_wbinvd_hkid(u64 hkid, struct page *page)
return seamcall(TDH_PHYMEM_PAGE_WBINVD, &args);
}
EXPORT_SYMBOL_GPL(tdh_phymem_page_wbinvd_hkid);
+
+void tdx_cpu_flush_cache(void)
+{
+ lockdep_assert_preemption_disabled();
+
+ if (!this_cpu_read(cache_state_incoherent))
+ return;
+
+ wbinvd();
+ this_cpu_write(cache_state_incoherent, false);
+}
+EXPORT_SYMBOL_GPL(tdx_cpu_flush_cache);
--
2.50.1
next prev parent reply other threads:[~2025-07-28 12:29 UTC|newest]
Thread overview: 20+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-07-28 12:28 [PATCH v5 0/7] TDX host: kexec/kdump support Kai Huang
2025-07-28 12:28 ` [PATCH v5 1/7] x86/kexec: Consolidate relocate_kernel() function parameters Kai Huang
2025-08-06 6:53 ` Huang, Kai
2025-08-06 13:00 ` Tom Lendacky
2025-08-06 22:29 ` Huang, Kai
2025-07-28 12:28 ` [PATCH v5 2/7] x86/sme: Use percpu boolean to control WBINVD during kexec Kai Huang
2025-07-28 12:28 ` [PATCH v5 3/7] x86/virt/tdx: Mark memory cache state incoherent when making SEAMCALL Kai Huang
2025-08-01 8:23 ` Chao Gao
2025-08-04 12:47 ` Huang, Kai
2025-08-12 0:51 ` Edgecombe, Rick P
2025-08-12 1:32 ` Huang, Kai
2025-08-12 1:34 ` Edgecombe, Rick P
2025-08-12 2:03 ` Huang, Kai
2025-08-14 0:09 ` Huang, Kai
2025-07-28 12:28 ` [PATCH v5 4/7] x86/kexec: Disable kexec/kdump on platforms with TDX partial write erratum Kai Huang
2025-07-28 12:28 ` [PATCH v5 5/7] x86/virt/tdx: Remove the !KEXEC_CORE dependency Kai Huang
2025-07-28 12:28 ` [PATCH v5 6/7] x86/virt/tdx: Update the kexec section in the TDX documentation Kai Huang
2025-07-28 12:28 ` Kai Huang [this message]
2025-08-01 8:30 ` [PATCH v5 7/7] KVM: TDX: Explicitly do WBINVD when no more TDX SEAMCALLs Chao Gao
2025-08-04 12:48 ` Huang, Kai
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=c29f7a3348a95f687c83ac965ebc92ff5f253e87.1753679792.git.kai.huang@intel.com \
--to=kai.huang@intel.com \
--cc=ashish.kalra@amd.com \
--cc=binbin.wu@linux.intel.com \
--cc=bp@alien8.de \
--cc=chao.gao@intel.com \
--cc=dan.j.williams@intel.com \
--cc=dave.hansen@intel.com \
--cc=dwmw@amazon.co.uk \
--cc=farrah.chen@intel.com \
--cc=hpa@zytor.com \
--cc=isaku.yamahata@intel.com \
--cc=kas@kernel.org \
--cc=kvm@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=mingo@redhat.com \
--cc=nik.borisov@suse.com \
--cc=pbonzini@redhat.com \
--cc=peterz@infradead.org \
--cc=reinette.chatre@intel.com \
--cc=rick.p.edgecombe@intel.com \
--cc=sagis@google.com \
--cc=seanjc@google.com \
--cc=tglx@linutronix.de \
--cc=thomas.lendacky@amd.com \
--cc=x86@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®