On Fri, Sep 25, 2026 at 02:10:47AM -0700, Christoph Hellwig wrote: > On Thu, Sep 24, 2026 at 07:58:40AM +0000, Leonid Ravich wrote: > > * No throughput *win* is claimed for software AES. The win is for > > accelerators that amortise setup across units; the software split > > exists so the interface works on every existing skcipher today and > > goes quiet as algorithms gain CRYPTO_ALG_REQ_SEG. > > So what is the use case? So far all crypto driver we've seen have > shown to be slower than cpu. Which one are you using that isn't?
The engine we use is out of tree, but it is not a special case. It is an SoC-integrated, DMA-driven, asynchronous xts(aes) engine of the same class as caam, ccree, qce, hisilicon sec2, inside-secure and the Marvell CPT drivers already in the tree. Like those, it pays a fixed cost per request (descriptor setup, doorbell, completion) that dm-crypt's one-request-per-sector pattern multiplies by the number of sectors. So the problem is common to the whole class, not specific to our hardware. You are right that such engines are generally not faster than the CPU in raw throughput, and I am not claiming ours is. What they give is offload: when dm-crypt batches a bio into a single request, the engine does the work that the CPU would otherwise do. In our measurements that cuts CPU utilisation by roughly 20-40% for the same I/O. On systems with few or small cores, those cycles are worth more than peak throughput. Without batching the per-sector request overhead eats that gain, which is why the series targets the request pattern rather than the cipher. > And if you have a genuinely useful one, should we have a proper > interface to it that doesn't pay the scatterlist overhead to start > with? For a DMA engine the scatterlist is not overhead: it is what the hardware consumes. The per-sector cost it pays is the per-request cost, and unit_size is what removes it. For the software path, v6 follows what Herbert suggested in the v4 thread [1]: the per-unit loop moves from the caller into the Crypto API, and it can then move further down into each algorithm, so the per-unit calls become direct calls and only the single call into the API stays indirect. That can improve on the status quo rather than just match it. v6 is the first step (the mid-layer split); per-algorithm native splitting via CRYPTO_ALG_REQ_SEG is the next. The interface is also deliberately the one being introduced for acomp in the batching series [2]: the same unit_size field and setter semantics, and the same CRYPTO_ALG_REQ_SEG bit (patch 2 here is Herbert's patch from that series). skcipher and acomp would share one model for multi-unit requests, with a native path for hardware and a mid-layer fallback for everything else, rather than skcipher growing a separate interface. If you had a different shape in mind for the hardware path, I would rather build that than guess. [1] https://lore.kernel.org/linux-block/[email protected]/ [2] https://lore.kernel.org/all/[email protected]/ Thanks, Leonid

