From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wr1-f44.google.com (mail-wr1-f44.google.com [209.85.221.44]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EB8A332F770 for ; Thu, 30 Jul 2026 08:15:02 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.221.44 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785399306; cv=none; b=I1ZpX1drH84ZP8kREiibBbMe0lgIcWXFn07GkKz4DCZRfrIhP9EKxcCkcHpJFNGMMmXLf9b7NhLyDMhn9mSGgqPk6cNz2KvL5yYXgGTf2fPBoXFHfkIroam8CsmWgzcsRHCZNwlSyP9E29IVWTsAXd42e2bhuY04UOgH81YeADA= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785399306; c=relaxed/simple; bh=zIgGXFhPhcmCwRhLNlAk+r8l30LriLi3g1dxUs263rc=; h=Date:From:To:Cc:Subject:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=JkTfB1zj6SaOlQ6euWVquuSByagRx5v5KmVGI21pjix8u42H4BvHBH8XuKE0TVFmVwgOM8CESFP8jHgubh41QR07x0pMl28BV+rfCNKRz3qNSP06d4R+l+JonUJcT19E0dOx/hQm3Sgo4I+ms3wkroBam9RCs+oWt+rnTxHDZAI= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=q8f/YB4Q; arc=none smtp.client-ip=209.85.221.44 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="q8f/YB4Q" Received: by mail-wr1-f44.google.com with SMTP id ffacd0b85a97d-47c6e9a694bso1188988f8f.1 for ; Thu, 30 Jul 2026 01:15:02 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1785399301; x=1786004101; darn=vger.kernel.org; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:subject:cc:to:from:date:from:to:cc:subject :date:message-id:reply-to:content-type; bh=gU8t2nLiO6KyR9mHd5oQe5CKsSlNQJu7tk5IuJVGRSw=; b=q8f/YB4Q2BvmLLnS3Mo3pHYQEvh9tBDhwsK0NjUV7UF+2N/jm19XI8Z8ux+RHZcLAI 8Z9AtOgv0EevMC7q+gL37iW236gJkxN0FKcKV1177/e6GsYjI0cnigY7OhoWb3cjMWZO b8rioRaWei4mf8lFVNN4PuzNC/aNFyx6wSlCtX6CiJj6FZnw8OxZZYLs9TzFMcMzBa6x 0YTjyKo7GQMtEjAqUpXVubnpsJpPIHq7L+L8CZdfWFNSmlOwLHP4VeDeXE1JSB6yYIBY qlZNf2G2rvEyzr4r5vWqMRhTazEvUllBRLMUbKsBmQN//o+c09vmx6waEwmWXeKgYrf7 yMPQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785399301; x=1786004101; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:subject:cc:to:from:date:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=gU8t2nLiO6KyR9mHd5oQe5CKsSlNQJu7tk5IuJVGRSw=; b=azV6NU97i44Y/XsTFdGwLTgpBehyDxHLhIg8QC3wIChqxGV0/K1rB0OKFPvOZ9ERFj ZVrnTdmNRrgtt+UxrOC/g8MeU8l+0N+dJuucGtaeHhQuIZ1MvuH6msWQIS0KF3SqnlCP wDHqjbRmrbYO6rFtZYWxZQZPgnjIKKW5kpAfhqMFSRQi3iljHrvjMLNXi81P8kh8lOFp NasIahozaqir1dhaBSD2kXGii4GY3rdlo0GwJhPgS36m7re0+M2qxpB7SJdDZNYE0kDx tjcgQsZGZKIz4xGM4Qj/0xLHd528mv5kGw1dMjGGYxgFaqnJ/D/aFWJRAWtQ9k6M9qTH b/Uw== X-Forwarded-Encrypted: i=1; AHgh+RovmCYjJUAlce6FTklCj7EN9Dsdzn5YZyT8lPgFDYngcS/KMMW7ddnRKNP4nbD4r4GGkAhQuT/+pxuWcjs=@vger.kernel.org X-Gm-Message-State: AOJu0YwvOnvZU82b69QftiV1T0xNX0seUT2bQxiZ9Y51YUh5xfEiFLSL 8bCB+X/b9N7F6mWSEjicZrW2m7XxKjewU+YwT8gVLMa0o9EMXjmbmPWE X-Gm-Gg: AR+sD11GBiRmHB3eWSNDXuaN++4vk+eD4G5Nd/2lfdn6/s5V28gvverqRBBGOgL2eGW 2LJHvlkUXoKeQRvTyPNHRsOqWJkdJCWMKTogSU+Wt77PqftqHu88sQhlo9f/qef+1ZIuGIWXcUH Se/LCAtTXM7Sx42BbeLlVnKde/HlvNOGxzXvBwEavKfVpALYokiH+1VCMXhiIY2j3y5RZYhDNYB uywz3I1/9/1rcvz/vlHrOKYaQKEU4zAzKESkgqfBz051dQpFRY/KIficR0NG7daY2eIeQM0c/Vj /Wn8DofGDMrpXfia0eu89U0PefKx3ozFXcs7KgffBs77dKs/6wSwDJNdZq5LFL2KLZcsJsSwqxp gCHufqHHH0syagoDgjZ186+QVsrPMqzVjrvdZgZOZ2KLY4fyQkZ/DIzu3C3nzsNRHz6g8vleATF S4ixzJNaDIT40rjwdRwW1/ZpWfqrBLlnWiziLwLGkQkhFbgfXk+7WefrBFT8pdF28J3neJD4QYE Y8FhO0mOltiVj45Rau3P80TFw== X-Received: by 2002:a05:600c:154b:b0:495:78ea:2687 with SMTP id 5b1f17b1804b1-49800eae778mr19575005e9.29.1785399300682; Thu, 30 Jul 2026 01:15:00 -0700 (PDT) Received: from pumpkin (82-69-66-36.dsl.in-addr.zen.co.uk. [82.69.66.36]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49800f0ebfasm37683065e9.2.2026.07.30.01.14.59 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 30 Jul 2026 01:14:59 -0700 (PDT) Date: Thu, 30 Jul 2026 09:14:55 +0100 From: David Laight To: Borislav Petkov Cc: Li Zhe , akpm@linux-foundation.org, apopple@nvidia.com, arnd@arndb.de, balbirs@nvidia.com, dave.hansen@linux.intel.com, david@kernel.org, kees@kernel.org, mingo@redhat.com, muchun.song@linux.dev, rppt@kernel.org, tglx@kernel.org, linux-arch@vger.kernel.org, linux-hardening@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, x86@kernel.org Subject: Re: [PATCH v8 7/9] x86/string: extend memcpy_flushcache() fixed-size fastpaths Message-ID: <20260730091455.1242d01a@pumpkin> In-Reply-To: <20260729234842.GFamqRWva8h7X59ccN@fat_crate.local> References: <20260727123429.5673-1-lizhe.67@bytedance.com> <20260727123429.5673-8-lizhe.67@bytedance.com> <20260729234842.GFamqRWva8h7X59ccN@fat_crate.local> X-Mailer: Claws Mail 4.1.1 (GTK 3.24.38; arm-unknown-linux-gnueabihf) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit On Wed, 29 Jul 2026 16:48:42 -0700 Borislav Petkov wrote: ... > at least the code is making a lot more sense now. > > The fact that you had to axe off so much cruft off of it tells me that you > haven't really measured it right. Especially since if you do actually measure the clock counts (non-trivial) you'll find that loops are often completely free. The out-of-order execution unit will (effectively) execute the loop control instructions to generate a list of instructions that get executed at a later time. So provided the loop control doesn't use more clocks than the loop body (and there are spare ALU units - usually true) loops really make little difference. This also means that unrolling loops often doesn't make things faster. You do need to minimise the loop control instructions (and gcc doesn't like the best loop that uses negative offsets from the end), and intel cpu can't execute single clock loops (amd ones can). Inlining also increases the code size, the I-cache reads are actually likely to be significant. You need to time a single 'cold-cache' call not just loops for long enough that the result is also skewed by timer ticks (etc). David > > So why do I really want your patch? > > Thx. >