mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Leonid Ravich <lravich@amazon.com>
To: Christoph Hellwig <hch@infradead.org>
Cc: Herbert Xu <herbert@gondor.apana.org.au>,
	<linux-crypto@vger.kernel.org>, <dm-devel@lists.linux.dev>,
	<linux-kernel@vger.kernel.org>, <davem@davemloft.net>,
	<ebiggers@kernel.org>, <agk@redhat.com>, <snitzer@kernel.org>,
	<mpatocka@redhat.com>, <bmarzins@redhat.com>
Subject: Re: [PATCH v6 0/6] crypto: skcipher - multi-data-unit request splitting
Date: Sun, 27 Sep 2026 07:14:14 +0000	[thread overview]
Message-ID: <20260927071422.12143-1-lravich@amazon.com> (raw)
In-Reply-To: <arY6lzQtvVzUGngU@infradead.org>

On Fri, Sep 25, 2026 at 02:10:47AM -0700, Christoph Hellwig wrote:
> On Thu, Sep 24, 2026 at 07:58:40AM +0000, Leonid Ravich wrote:
> > * No throughput *win* is claimed for software AES.  The win is for
> >   accelerators that amortise setup across units; the software split
> >   exists so the interface works on every existing skcipher today and
> >   goes quiet as algorithms gain CRYPTO_ALG_REQ_SEG.
>
> So what is the use case?  So far all crypto driver we've seen have
> shown to be slower than cpu.  Which one are you using that isn't?

The engine we use is out of tree, but it is not a special case.  It
is an SoC-integrated, DMA-driven, asynchronous xts(aes) engine of the
same class as caam, ccree, qce, hisilicon sec2, inside-secure and the
Marvell CPT drivers already in the tree.  Like those, it pays a fixed
cost per request (descriptor setup, doorbell, completion) that
dm-crypt's one-request-per-sector pattern multiplies by the number of
sectors.  So the problem is common to the whole class, not specific to
our hardware.

You are right that such engines are generally not faster than the CPU
in raw throughput, and I am not claiming ours is.  What they give is
offload: when dm-crypt batches a bio into a single request, the engine
does the work that the CPU would otherwise do.  In our measurements
that cuts CPU utilisation by roughly 20-40% for the same I/O.  On
systems with few or small cores, those cycles are worth more than
peak throughput.  Without batching the per-sector request overhead
eats that gain, which is why the series targets the request pattern
rather than the cipher.

> And if you have a genuinely useful one, should we have a proper
> interface to it that doesn't pay the scatterlist overhead to start
> with?

For a DMA engine the scatterlist is not overhead: it is what the
hardware consumes.  The per-sector cost it pays is the per-request
cost, and unit_size is what removes it.

For the software path, v6 follows what Herbert suggested in the v4
thread [1]: the per-unit loop moves from the caller into the Crypto
API, and it can then move further down into each algorithm, so the
per-unit calls become direct calls and only the single call into the
API stays indirect.  That can improve on the status quo rather than
just match it.  v6 is the first step (the mid-layer split);
per-algorithm native splitting via CRYPTO_ALG_REQ_SEG is the next.

The interface is also deliberately the one being introduced for acomp
in the batching series [2]: the same unit_size field and setter
semantics, and the same CRYPTO_ALG_REQ_SEG bit (patch 2 here is Herbert's
patch from that series).  skcipher and acomp would share one model for
multi-unit requests, with a native path for hardware and a mid-layer
fallback for everything else, rather than skcipher growing a separate
interface.

If you had a different shape in mind for the hardware path, I would
rather build that than guess.

[1] https://lore.kernel.org/linux-block/ajDNT5jVGgRtiNH6@gondor.apana.org.au/
[2] https://lore.kernel.org/all/20260125033537.334628-15-kanchana.p.sridhar@intel.com/

Thanks,
Leonid

  reply	other threads:[~2026-09-27  7:14 UTC|newest]

Thread overview: 10+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-24  7:58 Leonid Ravich
2026-09-24  7:58 ` [PATCH v6 1/6] crypto: skcipher - add per-request unit_size Leonid Ravich
2026-09-24  7:58 ` [PATCH v6 2/6] crypto: acomp - Add bit to indicate segmentation support Leonid Ravich
2026-09-24  7:58 ` [PATCH v6 3/6] crypto: skcipher - add crypto_skcipher_req_seg() helper Leonid Ravich
2026-09-24  7:58 ` [PATCH v6 4/6] crypto: skcipher - split multi-unit requests in the API layer Leonid Ravich
2026-09-24  7:58 ` [PATCH v6 5/6] crypto: testmgr - test multi-unit dispatch Leonid Ravich
2026-09-24  7:58 ` [PATCH v6 6/6] dm crypt: batch a bio segment's sectors via multi-unit requests Leonid Ravich
2026-09-25  9:10 ` [PATCH v6 0/6] crypto: skcipher - multi-data-unit request splitting Christoph Hellwig
2026-09-27  7:14   ` Leonid Ravich [this message]
2026-09-28  5:33     ` Christoph Hellwig

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260927071422.12143-1-lravich@amazon.com \
    --to=lravich@amazon.com \
    --cc=agk@redhat.com \
    --cc=bmarzins@redhat.com \
    --cc=davem@davemloft.net \
    --cc=dm-devel@lists.linux.dev \
    --cc=ebiggers@kernel.org \
    --cc=hch@infradead.org \
    --cc=herbert@gondor.apana.org.au \
    --cc=linux-crypto@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mpatocka@redhat.com \
    --cc=snitzer@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®