FrankChen021 opened a new pull request, #20362: URL: https://github.com/apache/druid/pull/20362
Related to #20326. ### Description Large `IN` lists can be translated into native filters multiple times during SQL planning. Each supported literal previously created a temporary `DruidLiteral` wrapper during every translation pass. This change directly extracts non-null Calcite integer and string literals into their native values. Other literal types, casts, and nulls continue through the existing general conversion path. The existing `InPlanningBenchmark` now has separate plan-only long and string variants. They prebuild the SQL string and stop after `planner.plan()`, excluding query-string construction, execution, and EXPLAIN serialization. ### Benchmark The comparison uses Calcite 1.42.0 with [apache/calcite#5263](https://github.com/apache/calcite/pull/5263) applied locally on both sides. The Calcite patch removes the dominant quadratic string-ARRAY validation cost, allowing this benchmark to isolate the incremental allocation reduction from the Druid change. This PR does not change Druid's Calcite dependency. JMH 1.37, JDK 25, one fork, one warmup iteration, and three measurement iterations: | Literal type | Count | Allocation before | Allocation after | Change | Time before | Time after | |---|---:|---:|---:|---:|---:|---:| | Long | 100,000 | 355,871,750 B/op | 341,459,277 B/op | -4.05% | 215.806 ms/op | 200.867 ms/op | | Long | 1,000,000 | 3,535,105,248 B/op | 3,391,244,437 B/op | -4.07% | 2,750.551 ms/op | 2,845.355 ms/op | | String | 100,000 | 1,532,068,185 B/op | 1,519,832,159 B/op | -0.80% | 719.258 ms/op | 729.453 ms/op | | String | 1,000,000 | 15,664,264,261 B/op | 15,434,139,336 B/op | -1.47% | 7,693.313 ms/op | 7,659.831 ms/op | The allocation reduction is stable at both sizes. Planning-time confidence intervals overlap, so this PR does not claim a demonstrated latency improvement. Benchmark command: ```bash java -jar benchmarks/target/benchmarks.jar \ 'org.apache.druid.benchmark.query.InPlanningBenchmark.query(Long|String)InSqlPlanOnly' \ -p inClauseLiteralsCount=100000,1000000 \ -p inSubQueryThreshold=2147483647 \ -wi 1 -i 3 -w 1s -r 1s -f 1 -prof gc \ -jvmArgsAppend --add-exports=java.base/sun.nio.ch=ALL-UNNAMED ``` #### Release note SQL planning allocates fewer temporary objects when translating large integer or string `IN` lists to native filters. <hr> ##### Key changed/added classes in this PR * `ScalarInArrayOperatorConversion` * `InPlanningBenchmark` <hr> This PR has: - [x] been self-reviewed. - [x] added comments explaining the intent of the optimization. - [x] added benchmarks covering integer and string literals. - [x] a release note entry in the PR description. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
