From: Andrei Vagin <avagin@google.com>
To: Borislav Petkov <bp@alien8.de>,
"Chang S. Bae" <chang.seok.bae@intel.com>
Cc: linux-kernel@vger.kernel.org, criu@lists.linux.dev,
Thomas Gleixner <tglx@kernel.org>,
Ingo Molnar <mingo@redhat.com>,
Dave Hansen <dave.hansen@linux.intel.com>,
x86@kernel.org, Andrei Vagin <avagin@google.com>,
Alexander Mikhalitsyn <alexander@mihalicyn.com>,
"H. Peter Anvin" <hpa@zytor.com>
Subject: [PATCH 1/2] x86/fpu: Allow restoring signal frames with larger xstate_size
Date: Tue, 29 Sep 2026 23:26:20 +0000 [thread overview]
Message-ID: <20260929232621.3745312-2-avagin@google.com> (raw)
In-Reply-To: <20260929232621.3745312-1-avagin@google.com>
check_xstate_in_sigframe() currently enforces that fx_sw->xstate_size in
the signal frame must not exceed the current task's fpstate->user_size.
This prevents checkpoint/restore tools (such as CRIU) from migrating a
process that was saved on a CPU with a larger set of enabled xstate
features to a CPU with fewer features, even if the process did not
actively use any of the unsupported features.
Commit fd14edd82077 ("x86/fpu: Pre-fault only required size of xstate
buffer") introduced validation in check_xstate_in_sigframe() to
calculate the required buffer size from the intersection of
fx_sw->xfeatures and fpstate->user_xfeatures, and to shrink
fx_sw->xstate_size to that required size.
Remove the fx_sw->xstate_size > fpstate->user_size check so that signal
frames with a larger xstate_size can be restored when all active
features are supported. Because fx_sw->xstate_size is no longer bounded
by fpstate->user_size before reading the trailing FP_XSTATE_MAGIC2
marker, check for unsigned overflow when comparing fx_sw->xstate_size
against fx_sw->extended_size and use get_user() instead of __get_user()
when reading FP_XSTATE_MAGIC2.
When allowing a larger xstate_size, the kernel must still reject the
signal frame (resulting in SIGSEGV) if the XSAVE header (xstate_bv)
marks any unsupported feature as active. On the direct restore path,
XRSTOR raises #GP if any bit in xstate_bv is 1 while the corresponding
bit in XCR0 is 0, but this does not cover dynamically enabled features
(such as AMX XFEATURE_XTILE_DATA) that are enabled in XCR0 while
disabled for the task via MSR_IA32_XFD and fpstate->user_xfeatures.
Because __restore_fpregs_from_user() masks xrestore_mask (EDX:EAX) with
fpstate->user_xfeatures before executing XRSTOR (as kernel-mode XRSTOR
must not trigger an XFD #NM trap), XRSTOR would ignore XFD-disabled
components and silently drop their active state.
Therefore, explicitly validate in check_xstate_in_sigframe() that
xbuf->header.xfeatures (xstate_bv) only contains features present in the
intersection of fx_sw->xfeatures and fpstate->user_xfeatures. A
concurrent user-space modification of xbuf->header.xfeatures between
this check and XRSTOR cannot compromise the kernel: the restore mask
(EDX:EAX) is derived from the kernel copy of fx_sw and masked with
fpstate->user_xfeatures, so XRSTOR will either fault with #GP (if a bit
outside XCR0 is set) or ignore any bit not enabled in the restore mask
without touching its memory area or raising a kernel-mode #NM.
Update Documentation/arch/x86/xstate.rst to reflect that signal frames
with a larger xstate_size can be restored if all active features are
supported.
Signed-off-by: Andrei Vagin <avagin@google.com>
---
Documentation/arch/x86/xstate.rst | 10 +++++++---
arch/x86/kernel/fpu/signal.c | 31 ++++++++++++++++++++++++++-----
2 files changed, 33 insertions(+), 8 deletions(-)
diff --git a/Documentation/arch/x86/xstate.rst b/Documentation/arch/x86/xstate.rst
index e2944f744255..f34bd06706d9 100644
--- a/Documentation/arch/x86/xstate.rst
+++ b/Documentation/arch/x86/xstate.rst
@@ -179,8 +179,9 @@ Signal Frame Layout and Portability
The signal frame is designed to be self-describing and portable. This is
especially important for checkpoint/restore tools like CRIU, which may restore
a process on a different host than where it was checkpointed. A signal frame
-created on a machine with fewer CPU features can be successfully restored on a
-machine with more CPU features, but not vice-versa.
+can be successfully restored across machines with different sets of enabled
+CPU features (whether the frame's ``xstate_size`` is smaller or larger than the
+destination host's default), provided all active features are supported.
Signal Frame Software Reserved Bytes
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
@@ -222,7 +223,10 @@ Portability Constraints
Signal frame portability is constrained by the architectural XSAVE layout.
Restoration is supported only if the destination host supports all features
-present in the frame and uses matching component offsets and sizes for them.
+active in the XSAVE header (``XSTATE_BV``) and uses matching component offsets
+and sizes for them. (The frame's ``xstate_size`` and ``_fpx_sw_bytes.xfeatures``
+may exceed the destination host's default if the source host had additional
+features enabled that were not actively used by the task.)
While layout compatibility is generally maintained across CPUs from the same
vendor, differences can occur across vendors or if the XSAVE space of a
deprecated feature (e.g. MPX) is repurposed for a newer feature (e.g. APX).
diff --git a/arch/x86/kernel/fpu/signal.c b/arch/x86/kernel/fpu/signal.c
index 1f721ac84283..ce8939a9ce8f 100644
--- a/arch/x86/kernel/fpu/signal.c
+++ b/arch/x86/kernel/fpu/signal.c
@@ -36,11 +36,18 @@ static inline bool check_xstate_in_sigframe(struct fxregs_state __user *buf_fx,
if (__copy_from_user(fx_sw, &buf_fx->sw_reserved[0], sizeof(*fx_sw)))
return false;
- /* Check for the first magic field and other error scenarios. */
+ /*
+ * Check for the first magic field and other error scenarios.
+ *
+ * fx_sw->xstate_size can exceed fpstate->user_size if the frame was
+ * saved on a CPU with a larger set of enabled xstate features.
+ * Reject the buffer below if the XSAVE header contains any active
+ * features that are not in both fx_sw->xfeatures and user_xfeatures.
+ */
if (fx_sw->magic1 != FP_XSTATE_MAGIC1 ||
fx_sw->xstate_size < min_xstate_size ||
- fx_sw->xstate_size > fpstate->user_size ||
- fx_sw->extended_size < fx_sw->xstate_size + FP_XSTATE_MAGIC2_SIZE)
+ fx_sw->xstate_size > fx_sw->extended_size ||
+ fx_sw->extended_size - fx_sw->xstate_size < FP_XSTATE_MAGIC2_SIZE)
goto err_setfx;
/*
@@ -49,19 +56,33 @@ static inline bool check_xstate_in_sigframe(struct fxregs_state __user *buf_fx,
* fpstate layout with out copying the extended state information
* in the memory layout.
*/
- if (__get_user(magic2, (__u32 __user *)(buf + fx_sw->xstate_size)))
+ if (get_user(magic2, (__u32 __user *)(buf + fx_sw->xstate_size)))
return false;
if (unlikely(magic2 != FP_XSTATE_MAGIC2))
goto err_setfx;
if (fx_sw->xstate_size != fpstate->user_size ||
fx_sw->xfeatures != fpstate->user_xfeatures) {
+ struct xregs_state __user *xbuf = buf;
+ u64 xstate_bv, xfeatures;
unsigned int xsize;
- u64 xfeatures;
+
+ if (__get_user(xstate_bv, &xbuf->header.xfeatures))
+ return false;
/* Calculate size of enabled features only. */
xfeatures = fx_sw->xfeatures & fpstate->user_xfeatures;
+ /*
+ * Reject XFD-disabled features present in XCR0 that XRSTOR
+ * would otherwise ignore when masked out of EDX:EAX, as well
+ * as any other active features not in xfeatures. Concurrent
+ * user-space changes after this check cannot affect the kernel
+ * because XRSTOR is masked with xfeatures.
+ */
+ if (xstate_bv & ~xfeatures)
+ return false;
+
xsize = xstate_calculate_size(xfeatures, false);
if (fx_sw->xstate_size < xsize)
return false;
--
2.56.0.rc1.315.gc6ed9934b7-goog
next prev parent reply other threads:[~2026-09-29 23:26 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-29 23:26 [PATCH 0/2] " Andrei Vagin
2026-09-29 23:26 ` Andrei Vagin [this message]
2026-09-29 23:26 ` [PATCH 2/2] selftests/x86: Check restoring FPU state " Andrei Vagin
2026-09-29 23:57 ` [PATCH 0/2] x86/fpu: Allow restoring signal frames " Borislav Petkov
2026-09-30 0:02 ` Andrei Vagin
2026-09-30 2:51 ` Borislav Petkov
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260929232621.3745312-2-avagin@google.com \
--to=avagin@google.com \
--cc=alexander@mihalicyn.com \
--cc=bp@alien8.de \
--cc=chang.seok.bae@intel.com \
--cc=criu@lists.linux.dev \
--cc=dave.hansen@linux.intel.com \
--cc=hpa@zytor.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mingo@redhat.com \
--cc=tglx@kernel.org \
--cc=x86@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®