Max Gekk created SPARK-59783:
--------------------------------

             Summary: Group the calls to split functions so the method holding 
them stays JIT-compilable
                 Key: SPARK-59783
                 URL: https://issues.apache.org/jira/browse/SPARK-59783
             Project: Spark
          Issue Type: Improvement
          Components: SQL
    Affects Versions: 5.0.0
            Reporter: Max Gekk


CodegenContext.splitExpressions splits the generated code of expressions into 
functions, but leaves one call per split function in the calling method. When 
the code splits into hundreds of functions, as a CASE WHEN of a thousand 
branches does, those calls alone take the calling method past the 8000 bytes 
HotSpot compiles (-XX:+DontCompileHugeMethods), so the method every row goes 
through runs interpreted.

Example: a CASE WHEN of 1000 branches, `SELECT CASE WHEN v = 1 THEN v * 1 ... 
WHEN v = 1000 THEN v * 1000 ELSE 0 END FROM (SELECT id % 1000 AS v FROM 
range(2000000))`, with whole-stage codegen off. The generated 
`SpecificUnsafeProjection$NestedClass_0.CaseWhen_0` that holds the calls is 
8060 bytes, Spark logs it as too long to be JIT compiled, and it is never 
compiled. A run takes about 11.1 s on a laptop.

Proposal: group the calls into functions of their own wherever their fold is 
longer than spark.sql.codegen.methodSplitThreshold, level by level, folding 
each group with the caller's foldFunctions and wrapping it with its 
makeSplitFunction, as generateInnerClassesFunctionCalls already does for the 
functions of one inner class. With it, the same query takes about 1.7 s a run.




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to