unikdahal opened a new pull request, #6696: URL: https://github.com/apache/datafusion-comet/pull/6696
## Which issue does this PR close? Closes #6612. ## Rationale for this change Spark 4.2 rewrites a MERGE that has only `NOT MATCHED` clauses as `InsertOnlyMergeExec`. With several clauses its query contains a `MergeRowsExec`. Since Spark 4.1, Comet keeps `MergeRowsExec` on the JVM when a stock V2 writer needs it to build the `MergeSummary`. Here the outer `InsertOnlyMergeExec` owns the summary independently of that child's metrics, so the child can run natively while the write stays on Spark's V2 writer. ## What changes are included in this PR? - Serve the insert-only `MergeRowsExec` child through the shared `CometMergeRows` serde, schema validation and `createExec`. - A Spark-version shim lets it run without a native merge summary only for the exact Spark 4.2 insert-only shape: no matched or target-only instructions, only `Keep` insert actions, and constant source/target presence expressions. Spark 4.1 always rejects that exception. - Any other `MergeRowsExec` on Spark 4.1+ stays on the JVM when a stock V2 writer requires it for `MergeSummary`; when the plan has native merge-summary support, semantic metrics are still required. - The insert-only shape keeps its no-cardinality-check semantics, and the schema, missing-child and unsupported-expression fallbacks keep their specific reasons. - Update the compatibility, operators, Iceberg writes and plan docs, and run the new suite in the Linux and macOS PR builds. ## How are these changes tested? - Spark vs Comet parity for multiple `NOT MATCHED` clauses, with AQE on and off. - Duplicate unmatched source keys and `MergeSummary` counters. - A single-clause insert-only rewrite, which has no `MergeRowsExec` child. - First-match short-circuiting, so later predicates do not evaluate rows already handled. - Scalar-subquery assignments remain discoverable by Comet. - A general Spark 4.2 MERGE still keeps `MergeRowsExec` on the JVM. - Near-miss shapes keep Spark execution: matched or target-only instructions, non-insert `Keep` actions, `Discard`/`Split`, wrong presence expressions, and zero or one instruction. - Malformed output schemas and missing children keep their specific fallback reasons; unsupported predicates and assignments fall back with row and summary parity. - Shared Spark 4.1/4.2 checks cover the version gate, the semantic-metrics flag, the insert action context and the no-cardinality-check fields. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
