Akash3121 commented on code in PR #9390: URL: https://github.com/apache/paimon/pull/9390#discussion_r4099196361
########## paimon-spark/paimon-spark-ut/src/test/scala/org/apache/paimon/spark/sql/MergeRowsCodegenTestBase.scala: ########## @@ -0,0 +1,153 @@ +/* + * Licensed to the Apache Software Foundation (ASF) under one + * or more contributor license agreements. See the NOTICE file + * distributed with this work for additional information + * regarding copyright ownership. The ASF licenses this file + * to you under the Apache License, Version 2.0 (the + * "License"); you may not use this file except in compliance + * with the License. You may obtain a copy of the License at + * + * http://www.apache.org/licenses/LICENSE-2.0 + * + * Unless required by applicable law or agreed to in writing, software + * distributed under the License is distributed on an "AS IS" BASIS, + * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. + * See the License for the specific language governing permissions and + * limitations under the License. + */ + +package org.apache.paimon.spark.sql + +import org.apache.paimon.spark.{PaimonSparkTestBase, SparkConnectorOptions} + +import org.apache.spark.sql.{PaimonUtils, Row} +import org.apache.spark.sql.catalyst.expressions.{AttributeReference, Expression, LessThan, Literal} +import org.apache.spark.sql.catalyst.expressions.Literal.TrueLiteral +import org.apache.spark.sql.catalyst.plans.logical.MergeRows +import org.apache.spark.sql.catalyst.plans.logical.MergeRows.{Discard, Instruction, Split} +import org.apache.spark.sql.execution.{QueryExecution, SparkPlan, WholeStageCodegenExec} +import org.apache.spark.sql.execution.datasources.v2.MergeRowsExec +import org.apache.spark.sql.paimon.Utils +import org.apache.spark.sql.types.IntegerType +import org.apache.spark.sql.util.QueryExecutionListener + +import java.util.concurrent.atomic.AtomicBoolean + +abstract class MergeRowsCodegenTestBase extends PaimonSparkTestBase { + + import testImplicits._ + + private val paimonCodegenKey = + s"spark.paimon.${SparkConnectorOptions.MERGE_CODEGEN_ENABLED.key()}" + + protected def keepInstruction(condition: Expression, output: Seq[Expression]): Instruction + + test("merge row codegen requires Spark and Paimon flags") { Review Comment: Since `write.merge.codegen.enabled` defaults to `false`, the existing MERGE tests continue to exercise only the interpreted iterator. This new test validates one constructed Keep/Discard/Split combination, but it does not establish semantic parity for the broader MERGE surface - for example duplicate-match cardinality violations, nullable conditions and assignments, multiple ordered clauses, and `NOT MATCHED BY SOURCE` behavior. Upstream Spark addressed the same concern by running MERGE behavior through both codegen and interpreted configurations. Could we similarly reuse or parameterize the existing Paimon MERGE tests with this option enabled, rather than relying on this single synthetic case? This would protect the new execution path against producing different rows or exceptions from the established implementation. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
