[ 
https://issues.apache.org/jira/browse/IMPALA-15425?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Zoltán Borók-Nagy updated IMPALA-15425:
---------------------------------------
    Labels: performance ramp-up  (was: performance)

> lower()/upper()/initcap() call libc tolower/toupper for every byte
> ------------------------------------------------------------------
>
>                 Key: IMPALA-15425
>                 URL: https://issues.apache.org/jira/browse/IMPALA-15425
>             Project: IMPALA
>          Issue Type: Improvement
>          Components: Backend
>            Reporter: Zoltán Borók-Nagy
>            Priority: Major
>              Labels: performance, ramp-up
>
> With the default UTF8_MODE=false, StringFunctions::LowerAscii(), UpperAscii() 
> and
> InitCapAscii() (be/src/exprs/string-functions-ir.cc:304-358), and the 
> lower-casing loop
> in udf-builtins-ir.cc:52, convert one byte at a time:
> {code:cpp}
> for (int i = 0; i < str.len; ++i) {
>   result.ptr[i] = ::tolower(str.ptr[i]);
> }
> {code}
> ::tolower/::toupper is an out-of-line libc call through the PLT with a locale 
> table
> lookup. It is not inlined, even in the codegen'd IR. A plain branchless loop 
> would not
> help the codegen path either: the IR module is compiled with -Os, and LLVM 
> does not
> vectorize the loop there (upper() at 100 bytes measured 0.95x).
> lower() and upper() are among the most common string functions: 
> case-insensitive
> filters, join keys and GROUP BY keys.
> h3. Proposed fix
> Add a small shared header, e.g. util/ascii-util.h, with ASCII case conversion 
> that
> processes 8 bytes per step with 64-bit word arithmetic and handles the tail 
> with scalar
> code. Use it in lower/upper/initcap and in udf-builtins. {{#pragma clang loop 
> vectorize}}
> is not an option because it breaks the -Werror IR build.
> h3. Measurements
> Standalone replica; gcc-compiled and IR (-Os, JIT approximated with opt/llc) 
> paths
> measured separately:
> ||Input length||lower/upper speedup||
> |8 bytes|2.4–2.5x|
> |24 bytes|6.5–7x|
> |100 bytes|11–14x|
> initcap is 2.2–2.9x faster. The output was byte-identical under the C and 
> en_US.UTF-8
> locales.
> h3. Behaviour note
> The output differs from today's only if impalad runs under a single-byte 
> locale such as
> ISO-8859-1, where libc tolower() also maps bytes ≥ 0x80. Under C and UTF-8 
> locales, libc
> leaves those bytes unchanged, which is what the new code does too.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to