Zoltán Borók-Nagy created IMPALA-15424:
------------------------------------------

             Summary: sha1/sha2/md5/mask_hash hit global OpenSSL 3 locks and 
build a stringstream on every row
                 Key: IMPALA-15424
                 URL: https://issues.apache.org/jira/browse/IMPALA-15424
             Project: IMPALA
          Issue Type: Improvement
          Components: Backend
            Reporter: Zoltán Borók-Nagy


Every row of sha1(), sha2(), md5() and mask_hash() does three expensive things
(be/src/exprs/utility-functions-ir.cc:266-339, mask-functions-ir.cc:959-968):

# IsFIPSMode() (util/openssl-util.h:81). On OpenSSL 3 this calls
{{EVP_default_properties_is_fips_enabled(nullptr)}}, which takes a lock on the 
global
property store.
# A digest call with a legacy EVP_MD: {{EVP_Digest(..., EVP_sha256(), ...)}} or 
the
SHA1()/SHA256()/SHA512() one-shots. On OpenSSL 3 each call does an implicit 
provider
fetch, which takes a global lock and bumps a shared refcount, and it allocates 
and frees
a digest context.
# Hex encoding through MathFunctions::HexString(), which builds a 
std::stringstream, then
StringFunctions::Lower() makes another copy.

Items 1 and 2 serialize all threads in the process. Items 1–3 make the 
single-thread
cost high. mask_hash() is what Ranger MASK_HASH policies evaluate on every row 
of a
masked column, so this also affects masked tables.

h3. Measurements

Standalone replica, OpenSSL 3.0.2 (Ubuntu 22.04), 32-byte input, sha2(x, 256):

||Threads||Current||Fixed||Speedup||
|1|1267 ns/row|96 ns/row|12.8x|
|16|about 2.5M rows/s in total|about 101M rows/s in total|40x|
|32|about 1.65M rows/s in total|about 135M rows/s in total|81x|

With the current code, total throughput stays at about 1.5–2.5M rows/s from 8 
threads
up. At 16 threads, mask_hash and sha1 are 41–46x faster and md5 is 11x faster.
The output was byte-identical over 1M rows × 3 input lengths plus known-answer 
tests.

Both parts of the fix are needed:
* Digest fix alone: 1.3x on 1 thread, 4–7x at 16–32 threads.
* Hex fix alone: 3–5x on 1 thread, but about 1x from 8 threads up.

h3. Proposed fix

* Cache the FIPS flag once at startup.
* Fetch each EVP_MD once per algorithm (EVP_MD_fetch) and reuse a thread_local
EVP_MD_CTX. Keep the OpenSSL < 3 path.
* Hex-encode directly into the result StringVal with a lower-case lookup table. 
This
removes the stringstream and the extra Lower() copy, and also speeds up hex().
* Handle digest failures. Today the EVP_Digest() return value is ignored and an
unfilled buffer is hex-encoded.

Risks: the thread_local context is freed at thread exit, which can run after
OPENSSL_cleanup. The digest fetch must happen after the FIPS configuration is 
loaded.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to