mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH] arm64: scrub vector and pointer auth state on task death
@ 2026-09-30  6:36 Bradley Morgan
  2026-09-30  8:02 ` Will Deacon
  0 siblings, 1 reply; 4+ messages in thread
From: Bradley Morgan @ 2026-09-30  6:36 UTC (permalink / raw)
  To: Catalin Marinas, Will Deacon
  Cc: Mark Rutland, Mark Brown, Vladimir Murzin, Ard Biesheuvel,
	Kees Cook, linux-arm-kernel, linux-kernel, Bradley Morgan

When a task dies its SVE, SME and FPSIMD register state, and its
pointer auth keys, stay in freed slab pages until the allocator hands
them out again. init_on_free=1 users get this wiped for the whole heap
but the option is expensive, so it is off in most builds.

The vector registers are not idle memory: userspace crypto runs
AES-GCM and Argon2 through SVE, so key material ends up in
sve_state, sme_state and thread_struct. On task exit and on the
exec and vector length reallocation paths those buffers are
dropped with a plain kfree(), and on task death the register file
copy embedded in thread_struct is not wiped at all. Anything that
reads the freed slab back (slab bugs, cold boot, a leaked page)
gets the dead task's keys.

Scrub it. kfree_sensitive() already exists for this exact job,
so the freeing sites just switch to it. The register file copy
and the pointer auth keys embedded in thread_struct cannot go
through kfree_sensitive(), so arch_release_task_struct() zeroes
them directly with memzero_explicit().

Scrubbed state:
  - sve_state, sme_state allocations, freed on task death, exec
    flush and vector length change
  - the uw.fpsimd_state register file copy embedded in
    thread_struct
  - pointer auth user and kernel keys embedded in thread_struct

This is a cold path. The cost is one memset of at most a few KB
plus the register file per dying task.

Signed-off-by: Bradley Morgan <brads@mainlining.org>
---
 arch/arm64/kernel/fpsimd.c  | 12 ++++++------
 arch/arm64/kernel/process.c | 16 +++++++++++++++-
 2 files changed, 21 insertions(+), 7 deletions(-)

diff --git a/arch/arm64/kernel/fpsimd.c b/arch/arm64/kernel/fpsimd.c
index e7f1682a3059..c149c12ba9f9 100644
--- a/arch/arm64/kernel/fpsimd.c
+++ b/arch/arm64/kernel/fpsimd.c
@@ -737,7 +737,7 @@ void cpu_enable_fpmr(const struct arm64_cpu_capabilities *__always_unused p)
 #ifdef CONFIG_ARM64_SVE
 static void sve_free(struct task_struct *task)
 {
-	kfree(task->thread.sve_state);
+	kfree_sensitive(task->thread.sve_state);
 	task->thread.sve_state = NULL;
 }
 
@@ -846,13 +846,13 @@ static int change_live_vector_length(struct task_struct *task,
 	 */
 	fpsimd_sync_from_effective_state(task);
 	task_set_vl(task, type, vl);
-	kfree(task->thread.sve_state);
+	kfree_sensitive(task->thread.sve_state);
 	task->thread.sve_state = sve_state;
 	fpsimd_sync_to_effective_state_zeropad(task);
 
 	if (type == ARM64_VEC_SME) {
 		task->thread.svcr &= ~SVCR_ZA_MASK;
-		kfree(task->thread.sme_state);
+		kfree_sensitive(task->thread.sme_state);
 		task->thread.sme_state = sme_state;
 	}
 
@@ -1195,7 +1195,7 @@ void sme_alloc(struct task_struct *task, bool flush)
 
 static void sme_free(struct task_struct *task)
 {
-	kfree(task->thread.sme_state);
+	kfree_sensitive(task->thread.sme_state);
 	task->thread.sme_state = NULL;
 }
 
@@ -1687,8 +1687,8 @@ void fpsimd_flush_thread(void)
 	current->thread.fp_type = FP_STATE_FPSIMD;
 
 	put_cpu_fpsimd_context();
-	kfree(sve_state);
-	kfree(sme_state);
+	kfree_sensitive(sve_state);
+	kfree_sensitive(sme_state);
 }
 
 /*
diff --git a/arch/arm64/kernel/process.c b/arch/arm64/kernel/process.c
index 581f80e9b9b7..c1d5d7c0fd3f 100644
--- a/arch/arm64/kernel/process.c
+++ b/arch/arm64/kernel/process.c
@@ -344,6 +344,20 @@ void flush_thread(void)
 void arch_release_task_struct(struct task_struct *tsk)
 {
 	fpsimd_release_task(tsk);
+
+	/*
+	 * Wipe the register file copy and the pointer auth keys, they
+	 * can hold key material the dead task used.
+	 */
+	memzero_explicit(&tsk->thread.uw.fpsimd_state,
+			 sizeof(tsk->thread.uw.fpsimd_state));
+#ifdef CONFIG_ARM64_PTR_AUTH
+	memzero_explicit(&tsk->thread.keys_user, sizeof(tsk->thread.keys_user));
+#ifdef CONFIG_ARM64_PTR_AUTH_KERNEL
+	memzero_explicit(&tsk->thread.keys_kernel,
+			 sizeof(tsk->thread.keys_kernel));
+#endif
+#endif
 }
 
 int arch_dup_task_struct(struct task_struct *dst, struct task_struct *src)
@@ -397,7 +411,7 @@ static int copy_thread_za(struct task_struct *dst, struct task_struct *src)
 					sme_state_size(src),
 					GFP_KERNEL);
 	if (!dst->thread.sme_state) {
-		kfree(dst->thread.sve_state);
+		kfree_sensitive(dst->thread.sve_state);
 		dst->thread.sve_state = NULL;
 		return -ENOMEM;
 	}
-- 
2.53.0


^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH] arm64: scrub vector and pointer auth state on task death
  2026-09-30  6:36 [PATCH] arm64: scrub vector and pointer auth state on task death Bradley Morgan
@ 2026-09-30  8:02 ` Will Deacon
  2026-09-30 15:43   ` Bradley Morgan
  0 siblings, 1 reply; 4+ messages in thread
From: Will Deacon @ 2026-09-30  8:02 UTC (permalink / raw)
  To: Bradley Morgan
  Cc: Catalin Marinas, Mark Rutland, Mark Brown, Vladimir Murzin,
	Ard Biesheuvel, Kees Cook, linux-arm-kernel, linux-kernel

On Wed, Sep 30, 2026 at 06:36:47AM +0000, Bradley Morgan wrote:
> When a task dies its SVE, SME and FPSIMD register state, and its
> pointer auth keys, stay in freed slab pages until the allocator hands
> them out again. init_on_free=1 users get this wiped for the whole heap
> but the option is expensive, so it is off in most builds.
> 
> The vector registers are not idle memory: userspace crypto runs
> AES-GCM and Argon2 through SVE, so key material ends up in
> sve_state, sme_state and thread_struct. On task exit and on the
> exec and vector length reallocation paths those buffers are
> dropped with a plain kfree(), and on task death the register file
> copy embedded in thread_struct is not wiped at all. Anything that
> reads the freed slab back (slab bugs, cold boot, a leaked page)
> gets the dead task's keys.
> 
> Scrub it. kfree_sensitive() already exists for this exact job,
> so the freeing sites just switch to it. The register file copy
> and the pointer auth keys embedded in thread_struct cannot go
> through kfree_sensitive(), so arch_release_task_struct() zeroes
> them directly with memzero_explicit().

I don't really find this very compelling, tbh. Presumably this data can
end up all over the place: on the stack, in a vCPU structure, in a GPR
so it really just feels like doing something for the sake of feeling like
we're making the kernel more secure rather than actually adding any
tangible benefits. There are also lots of things you're not covering,
so it's not clear why this is either necessary or sufficient.

What prompted you to do this?

Will

P.S. This ended up in my spam for some reason.

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH] arm64: scrub vector and pointer auth state on task death
  2026-09-30  8:02 ` Will Deacon
@ 2026-09-30 15:43   ` Bradley Morgan
  2026-09-30 15:57     ` Bradley Morgan
  0 siblings, 1 reply; 4+ messages in thread
From: Bradley Morgan @ 2026-09-30 15:43 UTC (permalink / raw)
  To: Will Deacon
  Cc: Catalin Marinas, Mark Rutland, Mark Brown, Vladimir Murzin,
	Ard Biesheuvel, Kees Cook, linux-arm-kernel, linux-kernel

On 30 September 2026 09:02:04 BST, Will Deacon <will@kernel.org> wrote:
>On Wed, Sep 30, 2026 at 06:36:47AM +0000, Bradley Morgan wrote:
>> When a task dies its SVE, SME and FPSIMD register state, and its
>> pointer auth keys, stay in freed slab pages until the allocator hands
>> them out again. init_on_free=1 users get this wiped for the whole heap
>> but the option is expensive, so it is off in most builds.
>> 
>> The vector registers are not idle memory: userspace crypto runs
>> AES-GCM and Argon2 through SVE, so key material ends up in
>> sve_state, sme_state and thread_struct. On task exit and on the
>> exec and vector length reallocation paths those buffers are
>> dropped with a plain kfree(), and on task death the register file
>> copy embedded in thread_struct is not wiped at all. Anything that
>> reads the freed slab back (slab bugs, cold boot, a leaked page)
>> gets the dead task's keys.
>> 
>> Scrub it. kfree_sensitive() already exists for this exact job,
>> so the freeing sites just switch to it. The register file copy
>> and the pointer auth keys embedded in thread_struct cannot go
>> through kfree_sensitive(), so arch_release_task_struct() zeroes
>> them directly with memzero_explicit().
>
>I don't really find this very compelling, tbh. Presumably this data can
>end up all over the place: on the stack, in a vCPU structure, in a GPR
>so it really just feels like doing something for the sake of feeling like
>we're making the kernel more secure rather than actually adding any
>tangible benefits. There are also lots of things you're not covering,
>so it's not clear why this is either necessary or sufficient.
>
>What prompted you to do this?
>

Hi Will,

I may be completely wrong here, but the "it ends up all over the place
anyway" argument applies to the keys subsystem too, and we still scrub
there. big_key, dh and friends use kfree_sensitive() on free, with the
same acknowledgment that the key material passed through other memory
on the way. iirc nobody argued those were pointless because the key
also touched a stack or a GPR.

The buffers here are not spill corners either:

- sve_state is the task's whole vector register file, up to ~8KB, and
  it holds whatever userspace crypto left there. OpenSSL has SVE2 AES
  GCM assembly these days, so keys do end up in exactly this buffer
- uw.fpsimd_state is 528 bytes of register file embedded in every
  task_struct, one slab bug in that file away from being read

And init_on_free, KSTACK_ERASE and kfree_sensitive itself are all the
same idea, memory that held crypto state should not outlive its task.
This is just the scoped version, a few KB of memset on task death and
nothing anywhere else. Still true on today's next btw, sve_free() and
friends are plain kfree() there.


Honestly what prompted me to do this is, an audit of where a dead task's
secrets
stay.
sve_state and
sme_state free with plain kfree() on a path every exiting task takes,
that was the whole trigger.

But if "not necessary and not sufficient" is the bar, it bounces the
keys precedent too, and I don't think you want to argue that. If the
bar for arm64 is different from the one keys lives with, that's fine,
I'll drop it, I'd just like to know where it is.


>Will
>
>P.S. This ended up in my spam for some reason.

Maybe this ended up on spam too?

https://lore.kernel.org/all/20260919121044.13883-1-brads@mainlining.org/


--- Thanks!
"I'm not a very positive person" - Linus torvalds

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH] arm64: scrub vector and pointer auth state on task death
  2026-09-30 15:43   ` Bradley Morgan
@ 2026-09-30 15:57     ` Bradley Morgan
  0 siblings, 0 replies; 4+ messages in thread
From: Bradley Morgan @ 2026-09-30 15:57 UTC (permalink / raw)
  To: Will Deacon
  Cc: Catalin Marinas, Mark Rutland, Mark Brown, Vladimir Murzin,
	Ard Biesheuvel, Kees Cook, linux-arm-kernel, linux-kernel

On 30 September 2026 16:43:09 BST, Bradley Morgan <brads@mainlining.org>
wrote:
>On 30 September 2026 09:02:04 BST, Will Deacon <will@kernel.org> wrote:
>>On Wed, Sep 30, 2026 at 06:36:47AM +0000, Bradley Morgan wrote:
>>> When a task dies its SVE, SME and FPSIMD register state, and its
>>> pointer auth keys, stay in freed slab pages until the allocator hands
>>> them out again. init_on_free=1 users get this wiped for the whole heap
>>> but the option is expensive, so it is off in most builds.
>>> 
>>> The vector registers are not idle memory: userspace crypto runs
>>> AES-GCM and Argon2 through SVE, so key material ends up in
>>> sve_state, sme_state and thread_struct. On task exit and on the
>>> exec and vector length reallocation paths those buffers are
>>> dropped with a plain kfree(), and on task death the register file
>>> copy embedded in thread_struct is not wiped at all. Anything that
>>> reads the freed slab back (slab bugs, cold boot, a leaked page)
>>> gets the dead task's keys.
>>> 
>>> Scrub it. kfree_sensitive() already exists for this exact job,
>>> so the freeing sites just switch to it. The register file copy
>>> and the pointer auth keys embedded in thread_struct cannot go
>>> through kfree_sensitive(), so arch_release_task_struct() zeroes
>>> them directly with memzero_explicit().
>>
>>I don't really find this very compelling, tbh. Presumably this data can
>>end up all over the place: on the stack, in a vCPU structure, in a GPR
>>so it really just feels like doing something for the sake of feeling like
>>we're making the kernel more secure rather than actually adding any
>>tangible benefits. There are also lots of things you're not covering,
>>so it's not clear why this is either necessary or sufficient.
>>
>>What prompted you to do this?
>>
>
>Hi Will,
>
>I may be completely wrong here, but the "it ends up all over the place
>anyway" argument applies to the keys subsystem too, and we still scrub
>there. big_key, dh and friends use kfree_sensitive() on free, with the
>same acknowledgment that the key material passed through other memory
>on the way. iirc nobody argued those were pointless because the key
>also touched a stack or a GPR.
>
>The buffers here are not spill corners either:
>
>- sve_state is the task's whole vector register file, up to ~8KB, and
>  it holds whatever userspace crypto left there. OpenSSL has SVE2 AES
>  GCM assembly these days, so keys do end up in exactly this buffer
>- uw.fpsimd_state is 528 bytes of register file embedded in every
>  task_struct, one slab bug in that file away from being read
>
>And init_on_free, KSTACK_ERASE and kfree_sensitive itself are all the
>same idea, memory that held crypto state should not outlive its task.
>This is just the scoped version, a few KB of memset on task death and
>nothing anywhere else. Still true on today's next btw, sve_free() and
>friends are plain kfree() there.
>
>
>Honestly what prompted me to do this is, an audit of where a dead task's
>secrets
>stay.
>sve_state and
>sme_state free with plain kfree() on a path every exiting task takes,
>that was the whole trigger.
>
>But if "not necessary and not sufficient" is the bar, it bounces the
>keys precedent too, and I don't think you want to argue that. If the
>bar for arm64 is different from the one keys lives with, that's fine,
>I'll drop it, I'd just like to know where it is.
>
>
>>Will
>>
>>P.S. This ended up in my spam for some reason.
>
>Maybe this ended up on spam too?
>
>https://lore.kernel.org/all/20260919121044.13883-1-brads@mainlining.org/
>
>
>--- Thanks!
>"I'm not a very positive person" - Linus torvalds
Oh and if it helps, my mentality for these, is for like someone really
paranoid. Imagine that, if you don't like this, what would you suggest?
--- Thanks!
"I'm not a very positive person" - Linus torvalds

^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-09-30 15:58 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-30  6:36 [PATCH] arm64: scrub vector and pointer auth state on task death Bradley Morgan
2026-09-30  8:02 ` Will Deacon
2026-09-30 15:43   ` Bradley Morgan
2026-09-30 15:57     ` Bradley Morgan

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®