On Thu, Sep 24, 2026 at 07:58:40AM +0000, Leonid Ravich wrote:
>
> 1. "Could you send me the patch so I can take a look?" (per-unit cost)
> ---------------------------------------------------------------------
> 
> Patch 4 is that patch -- the API-layer split, self-contained in
> crypto/skcipher.c.  The cost I measured, and where I believe it goes:
> 
> An in-kernel microbench of skcipher_crypt_unit() against the legacy
> per-unit loop (identical inner AES, VAES-AVX512, r7i.metal) shows a
> fixed ~48-52 ns per data unit, CV <1%, and it is the *same* ~50 ns
> for a 512 B unit and a 4096 B unit -- so it is per-call setup, not a
> crypto effect.  Against ~70 ns of VAES work for a 512 B sector that
> is ~+70% on the crypto call itself; end-to-end in dm-crypt it is
> ~2% of an ~18 us I/O and does not show up in fio at all (see the
> Performance section below).

I think the issue is that we're generating the IV twice.  Once
in the Crypto API and once again in the DM layer.  Not only is
this slow, but it is actually wrong for decryption.  You're ignoring
the IVs on the disk.

We really should only do it once, either in the DM layer or in
the Crypto API.  It should be transmitted via memory to the other
side.

My suggestion is to allocate memory for the IVs.  Of course
memory allocation can fail, but we have an easy fallback, which
is to use the existing single-unit path.

IOW if you succeed in allocating memory for storing the IVs,
then invoke the multi-unit code path, otherwise fall back to
the single-unit code path which iterates over the sectors one-
by-one.

To pass the IVs to the Crypto API (or back), just use the existing
IV pointer and extend it by the number of units.

Thanks,
-- 
Email: Herbert Xu <[email protected]>
Home Page: http://gondor.apana.org.au/~herbert/
PGP Key: http://gondor.apana.org.au/~herbert/pubkey.txt

Reply via email to