cloud-fan commented on code in PR #57742: URL: https://github.com/apache/spark/pull/57742#discussion_r3730386201
########## sql/core/src/test/scala/org/apache/spark/sql/execution/benchmark/AdaptivePartialAggregationBenchmark.scala: ########## @@ -0,0 +1,137 @@ +/* + * Licensed to the Apache Software Foundation (ASF) under one or more + * contributor license agreements. See the NOTICE file distributed with + * this work for additional information regarding copyright ownership. + * The ASF licenses this file to You under the Apache License, Version 2.0 + * (the "License"); you may not use this file except in compliance with + * the License. You may obtain a copy of the License at + * + * http://www.apache.org/licenses/LICENSE-2.0 + * + * Unless required by applicable law or agreed to in writing, software + * distributed under the License is distributed on an "AS IS" BASIS, + * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. + * See the License for the specific language governing permissions and + * limitations under the License. + */ + +package org.apache.spark.sql.execution.benchmark + +import org.apache.spark.benchmark.Benchmark +import org.apache.spark.sql.DataFrame +import org.apache.spark.sql.internal.SQLConf + +/** + * Benchmark comparing runtime adaptive partial aggregation (see + * [[SQLConf.ADAPTIVE_PARTIAL_AGGREGATION_ENABLED]]) against the static pre-shuffle partial + * aggregation. When the partial aggregation is not reducing rows, the operator streams the + * remaining rows through as single-row partial buffers instead of maintaining (and possibly + * spilling) a large aggregation map. + * + * Each scenario runs the query across the full matrix of whole-stage codegen on/off and the + * feature disabled (`adaptive = F`, the pre-change baseline) vs enabled (`adaptive = T`), over a + * {high, low}-cardinality x {no-spill, on-spill} grid: + * - high-cardinality, no spill: the periodic check bypasses, which should win. + * - low-cardinality, no spill: nothing bypasses, which must not regress. + * - high-cardinality, forced regular-map spill: the spill check bypasses instead of spilling, + * which should win. + * - low-cardinality, forced regular-map spill: the ratio is too low for the spill check to Review Comment: Please flip the ratio direction here and in the scenario comment at line 119. With only 1,000 keys over a large input, rows per key is high, which is why the spill check keeps aggregating; calling it low or tiny contradicts the compaction-ratio definition used by the implementation. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
