On Fri, Sep 25, 2026 at 02:10:47AM -0700, Christoph Hellwig wrote:
> On Thu, Sep 24, 2026 at 07:58:40AM +0000, Leonid Ravich wrote:
> > * No throughput *win* is claimed for software AES.  The win is for
> >   accelerators that amortise setup across units; the software split
> >   exists so the interface works on every existing skcipher today and
> >   goes quiet as algorithms gain CRYPTO_ALG_REQ_SEG.
>
> So what is the use case?  So far all crypto driver we've seen have
> shown to be slower than cpu.  Which one are you using that isn't?

The engine we use is out of tree, but it is not a special case.  It
is an SoC-integrated, DMA-driven, asynchronous xts(aes) engine of the
same class as caam, ccree, qce, hisilicon sec2, inside-secure and the
Marvell CPT drivers already in the tree.  Like those, it pays a fixed
cost per request (descriptor setup, doorbell, completion) that
dm-crypt's one-request-per-sector pattern multiplies by the number of
sectors.  So the problem is common to the whole class, not specific to
our hardware.

You are right that such engines are generally not faster than the CPU
in raw throughput, and I am not claiming ours is.  What they give is
offload: when dm-crypt batches a bio into a single request, the engine
does the work that the CPU would otherwise do.  In our measurements
that cuts CPU utilisation by roughly 20-40% for the same I/O.  On
systems with few or small cores, those cycles are worth more than
peak throughput.  Without batching the per-sector request overhead
eats that gain, which is why the series targets the request pattern
rather than the cipher.

> And if you have a genuinely useful one, should we have a proper
> interface to it that doesn't pay the scatterlist overhead to start
> with?

For a DMA engine the scatterlist is not overhead: it is what the
hardware consumes.  The per-sector cost it pays is the per-request
cost, and unit_size is what removes it.

For the software path, v6 follows what Herbert suggested in the v4
thread [1]: the per-unit loop moves from the caller into the Crypto
API, and it can then move further down into each algorithm, so the
per-unit calls become direct calls and only the single call into the
API stays indirect.  That can improve on the status quo rather than
just match it.  v6 is the first step (the mid-layer split);
per-algorithm native splitting via CRYPTO_ALG_REQ_SEG is the next.

The interface is also deliberately the one being introduced for acomp
in the batching series [2]: the same unit_size field and setter
semantics, and the same CRYPTO_ALG_REQ_SEG bit (patch 2 here is Herbert's
patch from that series).  skcipher and acomp would share one model for
multi-unit requests, with a native path for hardware and a mid-layer
fallback for everything else, rather than skcipher growing a separate
interface.

If you had a different shape in mind for the hardware path, I would
rather build that than guess.

[1] https://lore.kernel.org/linux-block/[email protected]/
[2] 
https://lore.kernel.org/all/[email protected]/

Thanks,
Leonid

Reply via email to