Max Gekk created SPARK-60049:
--------------------------------

             Summary: Keep each method's bytecode size in the codegen compile 
cache to avoid the trial's second compile
                 Key: SPARK-60049
                 URL: https://issues.apache.org/jira/browse/SPARK-60049
             Project: Spark
          Issue Type: Improvement
          Components: SQL
    Affects Versions: 5.0.0
            Reporter: Max Gekk


When whole-stage codegen decides whether to split a stage (SPARK-33301, 
https://github.com/apache/spark/pull/59225), it compiles the stage's code as a 
trial and reads the bytecode size of each method from the compile. A hit in the 
compile cache produces no method sizes, because they are computed when a class 
is compiled, so WholeStageCodegenExec.trialCompile drops the class and compiles 
it again. A hit happens where the class is cached while the decision for its 
code is not: the decision was evicted, or the same body was compiled earlier 
without being decided (with the split off, or under another limit), or another 
thread loaded the class in between.

Proposed: keep each method's size with ByteCodeStats in the compile cache, so 
that a hit gives the sizes and needs no second compile.

Care is needed: a trial holds back what its compile would report (the codegen 
metric updates, the "Code generated in" log line and the reports of methods 
past the JIT limit) until the caller keeps the code. A hit makes none of these, 
so they would have to be rebuilt from the cached stats for the code that is 
kept, or the kept code's second compile would report them as today.




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to