From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from foss.arm.com (foss.arm.com [217.140.110.172]) by smtp.subspace.kernel.org (Postfix) with ESMTP id C080428E5 for ; Tue, 18 Nov 2025 17:21:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=217.140.110.172 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1763486482; cv=none; b=em1sPq21Lc2eTMUjKoiUfFZjbU0z7VzDZgkFK+1RyWbtrm/rqfU90kpn+3Iwkeq8hsVYoYkuvW99mtQSQh+q2c0s4kej2s2KURrxYCGtZ+EugvY3b7viRsq6VOoU2QyXzTRyR/K1MHVL1xSNLkr7xWVt3hGz3Bx0q+WaCj1nyIU= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1763486482; c=relaxed/simple; bh=QqvqYK94BlMUEcPM6rHDCu+Tk7Dh/Ic+ebceM3I7CuY=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=DMS7iZwEK14thXXJJyoPScB3fbNwc/pMZknmY+Aeh1eiPvc2MKL5HXvicsmpMFXV2TJ7uxdnYdjwXpLUbam3K4MBApT338/rAL7i2rGTz/JufUxNhiV6W7gc2AnFxABz0O8rGNZwoayR2K1glwByo+wjE/hEF3RYGIadCzLCznM= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=arm.com; spf=pass smtp.mailfrom=arm.com; arc=none smtp.client-ip=217.140.110.172 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=arm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=arm.com Received: from usa-sjc-imap-foss1.foss.arm.com (unknown [10.121.207.14]) by usa-sjc-mx-foss1.foss.arm.com (Postfix) with ESMTP id 81EB0169C; Tue, 18 Nov 2025 09:21:12 -0800 (PST) Received: from [10.1.25.191] (XHFQ2J9959.cambridge.arm.com [10.1.25.191]) by usa-sjc-imap-foss1.foss.arm.com (Postfix) with ESMTPSA id A26CA3F66E; Tue, 18 Nov 2025 09:21:18 -0800 (PST) Message-ID: Date: Tue, 18 Nov 2025 17:21:17 +0000 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [DISCUSSION] kstack offset randomization: bugs and performance Content-Language: en-GB To: "Jason A. Donenfeld" , Arnd Bergmann Cc: Kees Cook , Ard Biesheuvel , Jeremy Linton , Will Deacon , Catalin Marinas , Mark Rutland , "linux-arm-kernel@lists.infradead.org" , Linux Kernel Mailing List References: <66c4e2a0-c7fb-46c2-acce-8a040a71cd8e@arm.com> From: Ryan Roberts In-Reply-To: Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 7bit On 18/11/2025 17:15, Jason A. Donenfeld wrote: > On Mon, Nov 17, 2025 at 05:47:05PM +0100, Arnd Bergmann wrote: >> On Mon, Nov 17, 2025, at 12:31, Ryan Roberts wrote: >>> On 17/11/2025 11:30, Ryan Roberts wrote: >>>> Hi All, >>>> >>>> Over the last few years we had a few complaints that syscall performance on >>>> arm64 is slower than x86. Most recently, it was observed that a certain Java >>>> benchmark that does a lot of fstat and lseek is spending ~10% of it's time in >>>> get_random_u16(). Cue a bit of digging, which led me to [1] and also to some new >>>> ideas about how performance could be improved. >> >> >>>> I believe this helps the mean latency significantly without sacrificing any >>>> strength. But it doesn't reduce the tail latency because we still have to call >>>> into the crng eventually. >>>> >>>> So here's another idea: Could we use siphash to generate some random bits? We >>>> would generate the secret key at boot using the crng. Then generate a 64 bit >>>> siphash of (cntvct_el0 ^ tweak) (where tweak increments every time we generate a >>>> new hash). As long as the key remains secret, the hash is unpredictable. >>>> (perhaps we don't even need the timer value). For every hash we get 64 bits, so >>>> that would last for 10 syscalls at 6 bits per call. So we would still have to >>>> call siphash every 10 syscalls, so there would still be a tail, but from my >>>> experiements, it's much less than the crng: >> >> IIRC, Jason argued against creating another type of prng inside of the >> kernel for a special purpose. > > Yes indeed... I'm really not a fan of adding bespoke crypto willynilly > like that. Let's make get_random_u*() faster. If you're finding that the > issue with it is the locking, and that you're calling this from irq > context anyway, then your proposal (if I read this discussion correctly) > to add a raw_get_random_u*() seems like it could be sensible. Those > functions are generated via macro anyway, so it wouldn't be too much to > add the raw overloads. Feel free to send a patch to my random.git tree > if you'd like to give that a try. Thanks Jason; that's exactly what I did, and it helps. But I think ultimately the get_random_uXX() slow path is too slow; that's the part that causes the tail latency problem. I doubt there are options for speeding that up? Anyway, I'm currently prototyping a few options and getting clear performance numbers. I'll be back in a couple of days and we can continue the discussion in light of the data. Thanks, Ryan > > Jason