From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wr1-f43.google.com (mail-wr1-f43.google.com [209.85.221.43]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B41D63C1FEB for ; Thu, 15 Jan 2026 18:46:27 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.221.43 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1768502793; cv=none; b=WGI+vfmvbkin1gdqlQTH8zKKkSPkdDfx1byjkQeLLFHIeJhjF0T9Ex313o75Ke+1tz4gL1jRhKqOED8G25d2f3OkCQi2mtFAHm5Ak7klM6A1/Rt1Kn1azMxdqJeqLsv25NnopGLwSlqIp79ntT5iH4fkv0AklZzAb2x3OEbGPLU= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1768502793; c=relaxed/simple; bh=ys9RofZHGUwyCtoznH+eR3kK37AdNd0tsqflsTwFHKU=; h=Date:From:To:Cc:Subject:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=gBw8vr8ICMu/TGeh10oVKB0/GjKqMaVhtByW6FLzSj7c9Rg1hF1L3cV3o5LWK2u45TLsewjR6nGBkhCawVk2WS9WtxCni2zIbgHwDnRRtloujkWbq0FdEJil3M/iecJf2iviSBINRvX7YynP/PKR/TuhQoq6OVz43+mecgF88/8= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=eNf6Un2b; arc=none smtp.client-ip=209.85.221.43 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="eNf6Un2b" Received: by mail-wr1-f43.google.com with SMTP id ffacd0b85a97d-432d2c7dd52so1116637f8f.2 for ; Thu, 15 Jan 2026 10:46:26 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1768502781; x=1769107581; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:subject:cc:to:from:date:from:to:cc:subject:date :message-id:reply-to; bh=y01hcEYTZNFeQclPQMnDmXUrt0mXOtnGHL4WtYQ+0+U=; b=eNf6Un2bepfK7NLm8QX3mNGXqiBn+dBPNOBl83YY4N8EBkCFUrJkGOFtV4IJNzZLhA pKDBC5nN+AVEoOrfbb4a9+cRWesAkA7VnJfZRJnTwbBIK5ny7NE7X9Z6XnqB4CviKhf7 Q3juuLqIUkJfVeCACYSPJkXW9iWPhoOretOL7Ze1vKKfw80C3i9rk+3sB5QxeYTRYQWc CKYGhtmRl/KuqwnSl9g1ccxuOrjfF9fOriQLG9HzxPZykqPz8pqJCfP364ztnhQzac+Z sjuyNJ/hi2IQeKB7guRbmLBFejebQDpkmty9XZCAr6s3oP8lCMQIaPcMLsxRYxYk9wOa nQIw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1768502781; x=1769107581; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:subject:cc:to:from:date:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to; bh=y01hcEYTZNFeQclPQMnDmXUrt0mXOtnGHL4WtYQ+0+U=; b=I2kFV1+IbvuaPpha+CUWTtEmv6SoZauUkkpQ7ZP6+svYQIhlVPH7Es9sU+sbGn3aZX T0uxPOgRgzTCzBL0RZth2XIc7RAulugvVtN4To356h+Wi3rKWXBsiC9LVtfOG8GO1lN0 iPk7EOxM/rWZf1XC2eiJqNQcjrdBRp6BRnC9La3cH4tzjwI1MpfFbsu4RYLh0H82FpHA HhOfUENhjIwbTQK6EPBaWwKbCOATWO44HiKuR5m2rYm1Hkt3t7CpRZUHq2cEcRyky4pT 5H2JqfEcTEoxPzNJpr2xbhzBYYAcT6ntAt08cl97AT9nZ7q73vZBkp1xjTu5yJTJr3+5 0RfA== X-Forwarded-Encrypted: i=1; AJvYcCWSGplePQcHIZogcUNlauHCp5kFyIrafvijJO3NqAaTyMQJYNqXeR2y92droiAbfnluhIgAyxb+hApSuGg=@vger.kernel.org X-Gm-Message-State: AOJu0YzHU2OkJzU4kdD/7tYoZhaP9bCyHsTSPKwFSZzcbB9Ws+jII/WB aSYA09w6coNzsDHtDiWSc868TpVmBOtbjoayDkYzkK8THeCzeoTK8CeU X-Gm-Gg: AY/fxX4gJ1cIsDu1ju42+71dGV+IMLAIl0BBNraZAe8OXr9UhYQPQQ5/WxTX4LlUoUN 1G/L+b0kY4Rcey4ADSiyZ501a2ibp/IpNNpFZABxe7CutietlQLsTnWwLmunq/BCKem59bFS2z8 W5yKixP20ZRU3r/zBcfwtnpdtKeDLo8iEaNRrS7dThaCPQeFWKSRnAylWyK3xIRWrw5rknfqyZz k5goQ10dgAZql1ASvsJWE0WXMd3hGrog9XRgvC+90eH/U4Pxo9E6jQmCYbTjiBr5xSF30wOGJd9 e9QcnZu9pOlWKBAbPu621PoeCVkTmzXaOMMPbAGgD7QsJ+9Y8DvMTULZLGjqd7NhmrYvBMwfpXK 3jwwbw+2qXPIQZH50jJLB6tU+7raR6QsCeEVyEUyhuaZy6T+Qeo6pCQ6gQ0dxQwTyeGBazNIh78 5gF562kiLY9TfRzn7tbqAe2MId7dwEsJitHRWVAvYoCPY+d/wbQR3p X-Received: by 2002:a05:6000:200f:b0:432:c0b8:ee58 with SMTP id ffacd0b85a97d-435699173f5mr400692f8f.0.1768502781108; Thu, 15 Jan 2026 10:46:21 -0800 (PST) Received: from pumpkin (82-69-66-36.dsl.in-addr.zen.co.uk. [82.69.66.36]) by smtp.gmail.com with ESMTPSA id ffacd0b85a97d-4356997e6cdsm488913f8f.31.2026.01.15.10.46.20 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 15 Jan 2026 10:46:20 -0800 (PST) Date: Thu, 15 Jan 2026 18:46:19 +0000 From: David Laight To: Paul Walmsley Cc: Feng Jiang , palmer@dabbelt.com, aou@eecs.berkeley.edu, alex@ghiti.fr, samuel.holland@sifive.com, charlie@rivosinc.com, conor.dooley@microchip.com, linux-riscv@lists.infradead.org, linux-kernel@vger.kernel.org Subject: Re: [PATCH] riscv: lib: optimize strlen loop efficiency Message-ID: <20260115184619.574f1b36@pumpkin> In-Reply-To: <20260115111947.54929ed0@pumpkin> References: <20251218032614.57356-1-jiangfeng@kylinos.cn> <20260115111947.54929ed0@pumpkin> X-Mailer: Claws Mail 4.1.1 (GTK 3.24.38; arm-unknown-linux-gnueabihf) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit On Thu, 15 Jan 2026 11:19:47 +0000 David Laight wrote: > For 64bit you can do a lot better (in C) by loading 64bit words and doing > the correct 'shift and mask' sequence to detect a zero byte. > It usually isn't worth in for 32bit. > > Does need to handle a mis-aligned base - eg by masking the bits off > the base pointer and or'ing in non-zero values to the value read from > the base pointer. > > David The version below seems to work https://www.godbolt.org/z/sME3Ts6vW It actually looks ok for x86-32, the loop is 8 instructions plus the branch but the 'register dependency chain' is only 4 instructions. So maybe better than byte compares for moderate to long strings. (Especially if the cpu starts speculatively executing the next loop iteration.) The OPTIMIZER_HIDE_VAR() helps a lot on (eg) MIPS-64 and a bit elsewhere since most 64bit cpu can't load 64bit immediates. I can't get gcc and clang to reliably have a loop with a conditional jump at the bottom, especially with an unconditional jump into the loop (to remove the '| mask' from the loop body). Also KASAN (or one of its friends) wont like the code reading entire words that hold the string. And it does need ffs/clz instructions - or a different loop bottom. (For BE one with clzl() returning 0 will work.) While I suspect the per-byte cost is 'two bytes/clock' on x86-64 the fixed cost may move the break-even point above the length of the average strlen() in the kernel. Of course, x86 probably falls back to 'rep scasb' at (maybe) (40 + 2n) clocks for 'n' bytes. A carefully written slightly unrolled asm loop might manage one byte per clock! I could spend weeks benchmarking different versions. David #define OPTIMIZER_HIDE_VAR(var) \ __asm__ ("" : "=r" (var) : "0" (var)) /* Set BE to test big-endian on little-endian. * For real BE either do a byteswapping read or use the BE code. */ #ifdef BE #define SWP(x) __builtin_bswap64(x) #define SHIFT << #else #define SWP(x) (x) #define SHIFT >> #endif unsigned long my_strlen(const char *s) { unsigned int off = (unsigned long)s % sizeof (long); const unsigned long *p = (void *)(s - off); unsigned long val; unsigned long mask; unsigned long ones = 0x01010101; /* Force the compiler to generate the related constants sanely. */ OPTIMIZER_HIDE_VAR(ones); ones |= ones << 16 << 16; mask = ((~0ul SHIFT 8) SHIFT 8 * (sizeof (long) - 1 - off)); do { val = SWP(*p++) | mask; mask = (val - ones) & ~val & ones << 7; } while (!mask); #ifdef BE off = __builtin_clzl(mask); /* Correct for "...\x01" */ val <<= off; for (off /= 8; val > (~0ul >> 8); off++) val <<= 8; #else off = (__builtin_ffsl(mask) - 1)/8; #endif return (const char *)(p - 1) + off - s; }