FrankChen021 opened a new pull request, #20362:
URL: https://github.com/apache/druid/pull/20362

   Related to #20326.
   
   ### Description
   
   Large `IN` lists can be translated into native filters multiple times during 
SQL planning. Each supported literal previously created a temporary 
`DruidLiteral` wrapper during every translation pass.
   
   This change directly extracts non-null Calcite integer and string literals 
into their native values. Other literal types, casts, and nulls continue 
through the existing general conversion path.
   
   The existing `InPlanningBenchmark` now has separate plan-only long and 
string variants. They prebuild the SQL string and stop after `planner.plan()`, 
excluding query-string construction, execution, and EXPLAIN serialization.
   
   ### Benchmark
   
   The comparison uses Calcite 1.42.0 with 
[apache/calcite#5263](https://github.com/apache/calcite/pull/5263) applied 
locally on both sides. The Calcite patch removes the dominant quadratic 
string-ARRAY validation cost, allowing this benchmark to isolate the 
incremental allocation reduction from the Druid change. This PR does not change 
Druid's Calcite dependency.
   
   JMH 1.37, JDK 25, one fork, one warmup iteration, and three measurement 
iterations:
   
   | Literal type | Count | Allocation before | Allocation after | Change | 
Time before | Time after |
   |---|---:|---:|---:|---:|---:|---:|
   | Long | 100,000 | 355,871,750 B/op | 341,459,277 B/op | -4.05% | 215.806 
ms/op | 200.867 ms/op |
   | Long | 1,000,000 | 3,535,105,248 B/op | 3,391,244,437 B/op | -4.07% | 
2,750.551 ms/op | 2,845.355 ms/op |
   | String | 100,000 | 1,532,068,185 B/op | 1,519,832,159 B/op | -0.80% | 
719.258 ms/op | 729.453 ms/op |
   | String | 1,000,000 | 15,664,264,261 B/op | 15,434,139,336 B/op | -1.47% | 
7,693.313 ms/op | 7,659.831 ms/op |
   
   The allocation reduction is stable at both sizes. Planning-time confidence 
intervals overlap, so this PR does not claim a demonstrated latency improvement.
   
   Benchmark command:
   
   ```bash
   java -jar benchmarks/target/benchmarks.jar \
     
'org.apache.druid.benchmark.query.InPlanningBenchmark.query(Long|String)InSqlPlanOnly'
 \
     -p inClauseLiteralsCount=100000,1000000 \
     -p inSubQueryThreshold=2147483647 \
     -wi 1 -i 3 -w 1s -r 1s -f 1 -prof gc \
     -jvmArgsAppend --add-exports=java.base/sun.nio.ch=ALL-UNNAMED
   ```
   
   #### Release note
   
   SQL planning allocates fewer temporary objects when translating large 
integer or string `IN` lists to native filters.
   
   <hr>
   
   ##### Key changed/added classes in this PR
   
   * `ScalarInArrayOperatorConversion`
   * `InPlanningBenchmark`
   
   <hr>
   
   This PR has:
   
   - [x] been self-reviewed.
   - [x] added comments explaining the intent of the optimization.
   - [x] added benchmarks covering integer and string literals.
   - [x] a release note entry in the PR description.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to