andygrove opened a new pull request, #6664: URL: https://github.com/apache/datafusion-comet/pull/6664
## Which issue does this PR close? Part of #5644. This is step 1 of its rollout, the split-operator default. Step 2, the native-writer default, stays open there. ## Rationale for this change #5644 makes the Iceberg write path the default in two steps: the split-operator plan first, then the native writer one release later. The split plan changes only the plan shape. iceberg-java still writes and commits the data files, so running it by default for a release gives the plan shape real use before the native writer, which builds on it, is turned on. With 1.2.0 targeted for late October or early November (#6550), it needs to land now to get nightly runs first. What blocked it is fixed: #6142 (the planner strategy ignored `spark.comet.enabled`) and #6586 (Spark 4.2 catalog transactions, fixed by #6587). The Iceberg Spark test jobs have run with the split plan on since it was added in #4658. ## What changes are included in this PR? - `spark.comet.write.iceberg.splitOperator.enabled` defaults to `true` on every Spark version. It moves from the testing config category to query execution, since it is now the switch that restores Spark's own operator, and its description says so. - Docs: the Iceberg writes user guide (intro, configuration example, and the fallback list, which now also names `spark.comet.enabled=false`), `iceberg.md`, the operators table, a 1.2.0 entry in the upgrade guide, the Iceberg writes contributor guide, and the Iceberg write review skill. - A test in `CometIcebergWriteActionSuite` that unsets the flag and checks that an Iceberg append plans `IcebergCommit` over `IcebergWrite`. ## How are these changes tested? - The new test failed before the default changed, with no `IcebergCommitExec` in the plan. - The existing test that `spark.comet.enabled=false` keeps Spark's own write plan still passes. - Every Iceberg suite that does not need MinIO, plus `CometConfSuite` and the Iceberg SQL file tests, passes locally on Spark 4.1 (379 tests, one canceled by an existing SPARK-55626 assumption) and on Spark 3.5 (365 tests, 11 canceled by Spark 4.0+, Spark 4.1+ and Iceberg version assumptions). Most of these suites create their tables with `INSERT`, which now goes through the split plan. - The Iceberg Spark test diffs already set the flag, so those jobs run the same plan as before. `run-iceberg-tests` and `run-all-spark-profiles` are applied for the other Spark versions. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
