From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-dy1-f181.google.com (mail-dy1-f181.google.com [74.125.82.181]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2E61778F3A for ; Fri, 22 May 2026 03:31:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.82.181 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1779420666; cv=none; b=PYBfYQvkjVVEbZKhWvssRX6r6K1tz6C+g3ical/aeu6x8p49Brzi52Nv/4klYsTmrvgfOJ6R1GX0QRHl2Ki4mvwK3gXj1bf9WU/OpDn3efO9bMNOy+Jyrigpny8GE9nlkDwXWMlu0MuFqqMbEvNNmK0NGnZeQO/+dbJXbifilDQ= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1779420666; c=relaxed/simple; bh=Ij/zJqsaaW0WYrYIVXaBGkdZCJREWqw1zOVpDLlYswc=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=qcIKvwv6hbogGPEMBKERF/qQMvW7nTRr6XEum9P7FAtIl6b3Pa+/6QfQvbtuCc7zvsqIZjyoXIWMpq4rU0SXimF6izxqHWUk0RxGrtMbUZtaZGabiEhtEZDu2Cz18ty7MuwyrMhMvFpET8k2iBFNVNLHhXDzEoAMscSXYmOfMiA= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=lixom.net; spf=none smtp.mailfrom=lixom.net; dkim=pass (2048-bit key) header.d=lixom-net.20251104.gappssmtp.com header.i=@lixom-net.20251104.gappssmtp.com header.b=rfLrNLv+; arc=none smtp.client-ip=74.125.82.181 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=lixom.net Authentication-Results: smtp.subspace.kernel.org; spf=none smtp.mailfrom=lixom.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=lixom-net.20251104.gappssmtp.com header.i=@lixom-net.20251104.gappssmtp.com header.b="rfLrNLv+" Received: by mail-dy1-f181.google.com with SMTP id 5a478bee46e88-304106b1204so3422602eec.0 for ; Thu, 21 May 2026 20:31:04 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=lixom-net.20251104.gappssmtp.com; s=20251104; t=1779420664; x=1780025464; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=gtaHIZdTtoWRwcagu2t4Vb/Fs0NvogRICoTuD3wnfwg=; b=rfLrNLv+UG/ZZ52ap9gN1aiJyIYxRlqBPgHOH1/l3hez44NZDljdPkBdAitKloaiEL a2YEQzWoNG+iShYKHJ4pAwrRjUAPz/uToRY/0qZ56RCwtsusBwqfAX3FmTK+3jMm+q9i 10Jh0YrfVGO4xzZw/0GFQmSv8VkJfWdhO5EmjkQgLMe4yDlQNxm9OuzWEdcvV2/Whkkc wC+FVZxEMY7JUe/YJnMKzCumpQ3QDLsgpg+hvWUht05CQ96WjTnINZhAteITTvHGw9P2 GgL5choYsUdpQrlCTy9oXIoiH8ABNxQDMsuIGs8jugc8GBv/bUQB8/nK10DvGepRIw+B Xv1Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1779420664; x=1780025464; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-gg:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=gtaHIZdTtoWRwcagu2t4Vb/Fs0NvogRICoTuD3wnfwg=; b=S4GyinSCzUMO1f0Njza7gRzWwfBQK2tifGw8DLtnFyyOY4bYBi5or4qdw3FvNqY/D8 jkPzXJhGkF6+bsD6JjHTnI7ErMaapxUFKiBVwGPR3h6YETn9NczA3kTHKfmrh9o78H2U aO4EFl4dhODP1Ai/jlsv8+W34oqrSXv10+olgCW58WkW2LyuMpEsMjSmV54dSC/4yBxX tLfH495/8s+u5nNQ1KI0pdv33qkKlgwT/0Q9mAmt6PBJKBeyKA8fG3nK5p3mmSMQieG8 IQf6MpH0nUAGhJMlqAWMo0c4GKgX6d2a+kdiE6gUtCglRup45SZg/r2d861rIpqWxlYT AUEQ== X-Forwarded-Encrypted: i=1; AFNElJ8ehfC/7A30Szv68r/IfMIIrrwsDx8Lw8TBlkPMyuCi76cdetcBGRAMqlc92E0E6vyflSPKbUqs8YdiP8Q=@vger.kernel.org X-Gm-Message-State: AOJu0Yz5MGF0woZYARIbV4ynU+9m9DXnhrMx3TzwqmvKaPciQi/9ZCLR oeXNhH/+1mHKEaC8ZumXKG64MeHuQZILLFS7qDFyyW/jpAwhBY5XqaTqGoQDcpA10CI= X-Gm-Gg: Acq92OHCVMtagzptSufBWkBSwG7knpjtoBrfiaLgMxYB50WIuI8Lf27cMU02Hc3nIVG oOsXoKPS9aJvUrNDo0JCKz5po4MQ9HcqOxvVSmNJMnFsIJEpq7D15eA2ujq5BBi4cuX8tHNNiMi /5jgPSx8zcNcXnlXUsA3HNcsYKIwtKoSaZ1yUIGOxNxPb5jOx75MqQqIBDjMQypaCKSS8LYeo0P /HKEZ+YqixaqxFM6xMmKt0bHZQerBHwDis5UiMSgQzUzBL1CkGDh2UHM+THvHW9mf6yrzWWOVmN XTC7PmNwglAIGjKJm7x+WV94fWn/YFzXoaVPC/U0nvZfvZiwmkzJ0d0G+hM4Qti9b539S4spMz0 yHuJ/vyHOxQ/cvet7w/2oIqOA6ACm4DJvHuAZKe8xlOzwY3SJscU84Gfjy/oMAco807dYJXYQ0o Qlfb6AZL3ReWTL X-Received: by 2002:a05:7301:fa8b:b0:2c4:4363:3742 with SMTP id 5a478bee46e88-30448ffc881mr969872eec.9.1779420663873; Thu, 21 May 2026 20:31:03 -0700 (PDT) Received: from localhost ([99.152.116.91]) by smtp.gmail.com with UTF8SMTPSA id 5a478bee46e88-30451ef4af0sm30644eec.3.2026.05.21.20.31.02 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 21 May 2026 20:31:02 -0700 (PDT) Date: Thu, 21 May 2026 20:30:56 -0700 From: Olof Johansson To: Andy Chiu Cc: linux-riscv@lists.infradead.org, linux-kernel@vger.kernel.org, dfustini@oss.tenstorrent.com, bjorn@kernel.org, greentime.hu@sifive.com, vincent.chen@sifive.com Subject: Re: Re: [PATCH v3 0/4] riscv: optimize Vector context restore on syscall Message-ID: References: <20260521162521.188629-1-tchiu@tenstorrent.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: On Thu, May 21, 2026 at 04:09:40PM -0500, Andy Chiu wrote: > On Thu, May 21, 2026 at 12:15:07PM -0700, Olof Johansson wrote: > > Hi Andy, > > > > On Thu, May 21, 2026 at 11:25:16AM -0500, Andy Chiu wrote: > > > This patch series optimizes riscv vector state handling across syscall > > > boundaries and context switches. The kernel now keeps track of the > > > INITIAL state in sstatus.vs to optimize unnecessary context management > > > operations. > > > > > > This version merges daichengrong's RFC patch [1] for the state tracking > > > code as it looks cleaner than my v2/v1. > > > > > > [1]: https://lore.kernel.org/linux-riscv/7ba2f4b7-8475-4ec3-ab31-58b332bda47e@iscas.ac.cn/#r > > > Link to v2: https://lore.kernel.org/linux-riscv/20260402043414.2421916-1-andybnac@gmail.com/ > > > > A patchset like this would be really helped by some kind of numbers in the > > cover letter to indicate how much performance moved, given a claim of > > optimization. > > > > Thanks for pointing it out, I totally agree with you. I had included a > test result on sifive's hardware in v2[1]. But that was on FPGA, I will > test it on a real silicon as soon as I have an access. Sorry for the > confusing claim here. > > My test was running on a vector enabled version of lat_ctx. I modified > the main function to make sure the process touches vector, then run with > 2 threads. Since lat_ctx uses syscall interface to notify another > process, the kernel will trash their vector registers instead of wasting > cycles on saving/restoring them. > > > > > Just for kicks I tried a simple microbenchmark for syscalls from > > a vector-enabled process: > > > > #define _GNU_SOURCE > > #include > > #include > > #include > > #include > > #include > > #include > > > > static inline uint64_t ns_now(void) { > > struct timespec t; > > clock_gettime(CLOCK_MONOTONIC, &t); > > return t.tv_sec * 1000000000ull + t.tv_nsec; > > } > > > > int main(int argc, char **argv) { > > int iters = argc > 1 ? atoi(argv[1]) : 10000000; > > int use_v = argc > 2 ? atoi(argv[2]) : 1; > > > > if (use_v) { > > asm volatile( > > ".option push\n\t.option arch, +v\n\t" > > "vsetivli x0, 1, e32, m1, ta, ma\n\t" > > "vmv.v.i v0, 1\n\t" > > ".option pop\n\t" ::: "memory"); > > } > > > > for (int i = 0; i < 10000; i++) syscall(SYS_getppid); // warmup > > > > uint64_t t0 = ns_now(); > > for (int i = 0; i < iters; i++) syscall(SYS_getppid); > > uint64_t t1 = ns_now(); > > > > printf("V=%d %.1f ns/call (%lu ns / %d iters)\n", > > use_v, (double)(t1 - t0) / iters, t1 - t0, iters); > > return 0; > > } > > > > > > I compiled with gcc -O3, default GCC 14.2 on Debian 13. Host is x280 > > (Blackhole). Base kernel sources is 7.1.0-rc4-next-20260520 defconfig. Ran > > with taskset to pin to one of the CPUs. > > > > The testcase doesn't use vector inbetween each syscall, but will obviously > > have initiated the state (if started with '1' as second argument). > > > > Without this patchset: > > V=1 242.9 ns/call (12144527848 ns / 50000000 iters) > > > > With this patchset: > > V=1 264.5 ns/call (13226852900 ns / 50000000 iters) > > This 9% regression is suprising to me as this patch set (without the fix > below) should be equivalent in performance on this code path (and better > if getpid takes one process switch). > > Before the patch, nulling v resgisters happens at the entry point. After > this patchset we mark sstatus.vs to INITIAL at the entry and nulling > happens right before getting back to the user space. > > > > > Interestingly enough, with V=0 test it sped up slightly (194.3 -> 189.5 ns). > > > > This result is expected as with V=0 the kernel doesn't have to maintain > vstate at all. But we are also looking into ways to improve mode switch > latencies. > > > I repeated the runs a few times, with similar results so I don't think it's > > explainable as noise. > > > > Thanks for carrying out the experiment, it's very sound! I actually > missed one thing on this patch for it to be optimized on this specific > case: > > diff --git a/arch/riscv/include/asm/vector.h b/arch/riscv/include/asm/vector.h > index 8c1e64e0dd0b..5d1282870a20 100644 > --- a/arch/riscv/include/asm/vector.h > +++ b/arch/riscv/include/asm/vector.h > @@ -40,6 +40,15 @@ > _res; \ > }) > > +#define __riscv_v_vstate_check_gt(_val, TYPE) ({ \ > + bool _res; \ > + if (has_xtheadvector()) \ > + _res = ((_val) & SR_VS_THEAD) > SR_VS_##TYPE##_THEAD; \ > + else \ > + _res = ((_val) & SR_VS) > SR_VS_##TYPE; \ > + _res; \ > +}) > + > extern unsigned long riscv_v_vsize; > int riscv_v_setup_vsize(void); > bool insn_is_vector(u32 insn_buf); > @@ -323,7 +332,7 @@ static inline void riscv_v_vstate_set_restore(struct task_struct *task, > > static inline void riscv_v_vstate_discard(struct pt_regs *regs) > { > - if (riscv_v_vstate_query(regs)) { > + if (__riscv_v_vstate_check_gt(regs->status, INITIAL)) { > riscv_v_vstate_set_restore(current, regs); > riscv_v_vstate_init(regs); > } > > In this way if the user never touch vregs again then we will not null > out the context at syscall exit. Significantly better indeed, now ~191ns without vector, ~192ns with -- so a proper optimization. > Again, I only have it functionally tested at the moment. I appreciate it > if you could get the number with the above diff. Meanwhile, I am going > souce and run on an actual hardware and hopefully find the reason for > the above regression, before rolling out v4. Sounds good. -Olof