From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from foss.arm.com (foss.arm.com [217.140.110.172]) by smtp.subspace.kernel.org (Postfix) with ESMTP id B32522C21EB for ; Mon, 24 Nov 2025 17:50:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=217.140.110.172 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1764006621; cv=none; b=jgtwYBO8AdingpVC4nAuMFDlywvGG1d0YdbVEkxaAPoNpsxqdl4jnWSvbR7TmNDlaSdd8nTwjMwI4LvqnqnMQ3/MPZixn43TvkMo4nyD+BtO6BQ0QBM2o23rXKkZWXLl7qdqlKPkmG7PWvUMUCGqIQM8/1HV31HznydJhMRpjAI= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1764006621; c=relaxed/simple; bh=6rrEWUlXiKE/RH6IaxTHlsYJ7qZXeSt/8zAMwWSL5yc=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=b9f2kUdkq2poLxyaEiCET0qSqWM/b9F0iBxumxLSNB29ukT56s/+DBMGls3LCjztiNBZMD8F5Iy3oHvXHyvO96fi6KYLozw1+GUzfKV2dlLW743ZhcSShzjNTT0A3g2D8/meGSLXXElMPwmKd1kDLY4F11VFSWeJtjnT2ypYNYs= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=arm.com; spf=pass smtp.mailfrom=arm.com; arc=none smtp.client-ip=217.140.110.172 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=arm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=arm.com Received: from usa-sjc-imap-foss1.foss.arm.com (unknown [10.121.207.14]) by usa-sjc-mx-foss1.foss.arm.com (Postfix) with ESMTP id 4374D1477; Mon, 24 Nov 2025 09:50:10 -0800 (PST) Received: from [10.57.88.238] (unknown [10.57.88.238]) by usa-sjc-imap-foss1.foss.arm.com (Postfix) with ESMTPSA id 4D8623F73B; Mon, 24 Nov 2025 09:50:16 -0800 (PST) Message-ID: Date: Mon, 24 Nov 2025 17:50:14 +0000 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [DISCUSSION] kstack offset randomization: bugs and performance Content-Language: en-GB To: Kees Cook , Will Deacon Cc: Arnd Bergmann , Ard Biesheuvel , Jeremy Linton , Catalin Marinas , Mark Rutland , "linux-arm-kernel@lists.infradead.org" , Linux Kernel Mailing List References: <66c4e2a0-c7fb-46c2-acce-8a040a71cd8e@arm.com> From: Ryan Roberts In-Reply-To: Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 7bit On 24/11/2025 17:11, Kees Cook wrote: > > > On November 24, 2025 6:36:25 AM PST, Will Deacon wrote: >> On Mon, Nov 17, 2025 at 11:31:22AM +0000, Ryan Roberts wrote: >>> On 17/11/2025 11:30, Ryan Roberts wrote: >>>> Could this give us a middle ground between strong-crng and >>>> weak-timestamp-counter? Perhaps the main issue is that we need to store the >>>> secret key for a long period? >>>> >>>> >>>> Anyway, I plan to work up a series with the bugfixes and performance >>>> improvements. I'll add the siphash approach as an experimental addition and get >>>> some more detailed numbers for all the options. But wanted to raise it all here >>>> first to get any early feedback. >> >> FWIW, I share Mark's concerns about using a counter for this. Given that >> the feature currently appears to be both slow _and_ broken I'd vote for >> either removing it or switching over to per-thread offsets as a first >> step. > > That it has potential weaknesses doesn't mean it should be entirely removed. > >> We already have a per-task stack canary with >> CONFIG_STACKPROTECTOR_PER_TASK so I don't understand the reluctance to >> do something similar here. > > That's not a reasonable comparison: the stack canary cannot change arbitrarily for a task or it would immediately crash on a function return. :) > >> Speeding up the crypto feels like something that could happen separately. > > Sure. But let's see what Ryan's patches look like. The suggested changes sound good to me. Just to say I haven't forgotten about this; I ended up having to switch to something more urgent. Hoping to get back to it later this week. I don't think this is an urgent issue, so hopefully folks are ok waiting. I propose to post whatever I end up with then we can all disscuss from there. But the rough shape so far: Fixes: - Remove choose_random_kstack_offset() - arch passes random into add_random_kstack_offset() (fixes migration bypass) - Move add_random_kstack_offset() to el0_svc()/el0_svc_compat() (before enabling interrupts) to fix non-preemption requirement (arm64) Perf Improvements: - Based on Jeremy's prng, but buffer the 32 bits and use 6 bits per syscall (so cost of prng generation is amortized over 5 syscalls) - Reseed prng using get_random_u64() every 64K prng invocations (so cost of get_random_u64() is amortized over 64K*5 syscalls) - So while get_random_u64() still has a latency spike, it's so infrequent that it doesn't show up in p99.9 for my benchmarks. - If we want to change it to per-task, I think it's all amenable. - I'll leave the timer off limits for arm64. Although I'm seeing some inconsistencies in the performance measurements, so need to get that understood properly first. Thanks, Ryan > > -Kees > >