From: David Laight <david.laight.linux@gmail.com>
To: Demian Shulhan <demyansh@gmail.com>
Cc: Catalin Marinas <catalin.marinas@arm.com>,
Will Deacon <will@kernel.org>,
Mark Rutland <mark.rutland@arm.com>,
Eric Biggers <ebiggers@kernel.org>,
Andrew Morton <akpm@linux-foundation.org>,
Marco Elver <elver@google.com>, Ard Biesheuvel <ardb@kernel.org>,
Robin Murphy <robin.murphy@arm.com>,
David Gow <davidgow@google.com>,
Brendan Higgins <brendan.higgins@linux.dev>,
Nathan Chancellor <nathan@kernel.org>,
linux-arm-kernel@lists.infradead.org,
linux-kernel@vger.kernel.org, kunit-dev@googlegroups.com,
netdev@vger.kernel.org, llvm@lists.linux.dev
Subject: Re: [PATCH 0/2] arm64: csum: Add fused copy and Internet checksum
Date: Sun, 27 Sep 2026 18:44:14 +0100 [thread overview]
Message-ID: <20260927184414.6c0c8867@pumpkin> (raw)
In-Reply-To: <20260927131838.6774-1-demyansh@gmail.com>
On Sun, 27 Sep 2026 15:17:56 +0200
Demian Shulhan <demyansh@gmail.com> wrote:
> arm64 currently uses the generic csum_partial_copy_nocheck(), which
> performs memcpy() followed by a second pass for csum_partial(). This
> double pass exerts unnecessary pressure on the L1 cache.
Which workload actually needs this?
Most modern ethernet MAC support checksum setting on transmit and
checking on receive.
So the software checksum shouldn't be needed very often.
IIRC there is also code to defer UDP checksum validation until the
copy_to_user().
I'd bet (a few pints of beer) that the complication this adds isn't
actually worth while.
Even Linus can't remember why it was done, my guess is it improved the
performance of the userspace NFS (over UDP) daemon that would be
doing 8k UDP send/receive (fragmented by IP).
There is certainly still code to checksum data during copy_from_user()
in send().
Last time I looked I couldn't see why send on TCP sockets didn't go through it.
On x86 (in particular) copies can be done far faster than ones that
include a checksum.
David
>
> Replace it with a single-pass implementation. The new implementation
> provides a general-purpose register path for short buffers and atomic
> contexts, and a kernel-mode NEON path for lengths >= 1024 bytes.
>
> Measured in-kernel on an Ampere Altra (Neoverse-N1):
> - Scalar path: 1.2x-1.6x faster for lengths < 1024 bytes.
> - NEON path: 1.2x faster at 1024 bytes, scaling up to 1.6x-1.8x at
> 4096 bytes.
> On Apple M-series cores, gains are 1.3-1.7x below 1024 bytes and
> 1.6-2.4x above. No length or alignment regresses on either
> microarchitecture.
>
> Patch 1 implements the fused routines and the dispatcher.
> Patch 2 adds KUnit test coverage for the new API and internal paths.
>
> Tested: in-kernel benchmark module on Neoverse-N1 with both
> implementations cross-checked (0 mismatches); KUnit suite under QEMU
> (with/without KASAN, with PREEMPT_RT), exhaustive and random userspace
> testing of both routines against a naive reference with PROT_NONE guard
> pages, gcc 13 and clang 18 W=1 builds, checkpatch --strict.
>
> Demian Shulhan (2):
> arm64: csum: Add fused copy and Internet checksum
> lib/tests: checksum: Add KUnit tests for csum_partial_copy_nocheck()
>
> arch/arm64/include/asm/checksum.h | 3 +
> arch/arm64/lib/Makefile | 7 +-
> arch/arm64/lib/csum-copy-neon.c | 168 +++++++++++
> arch/arm64/lib/csum-copy.c | 108 +++++++
> arch/arm64/lib/csum-copy.h | 84 ++++++
> arch/arm64/lib/csum.c | 54 ++++
> lib/Kconfig.debug | 10 +
> lib/tests/checksum_kunit.c | 463 ++++++++++++++++++++++++++++++
> 8 files changed, 896 insertions(+), 1 deletion(-)
> create mode 100644 arch/arm64/lib/csum-copy-neon.c
> create mode 100644 arch/arm64/lib/csum-copy.c
> create mode 100644 arch/arm64/lib/csum-copy.h
>
>
> base-commit: 93f51579e7df248780214094418f205253383cc5
prev parent reply other threads:[~2026-09-27 17:44 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-27 13:17 Demian Shulhan
2026-09-27 13:17 ` [PATCH 1/2] " Demian Shulhan
2026-09-27 13:17 ` [PATCH 2/2] lib/tests: checksum: Add KUnit tests for csum_partial_copy_nocheck() Demian Shulhan
2026-09-27 17:44 ` David Laight [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260927184414.6c0c8867@pumpkin \
--to=david.laight.linux@gmail.com \
--cc=akpm@linux-foundation.org \
--cc=ardb@kernel.org \
--cc=brendan.higgins@linux.dev \
--cc=catalin.marinas@arm.com \
--cc=davidgow@google.com \
--cc=demyansh@gmail.com \
--cc=ebiggers@kernel.org \
--cc=elver@google.com \
--cc=kunit-dev@googlegroups.com \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=llvm@lists.linux.dev \
--cc=mark.rutland@arm.com \
--cc=nathan@kernel.org \
--cc=netdev@vger.kernel.org \
--cc=robin.murphy@arm.com \
--cc=will@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®