andygrove opened a new pull request, #2566: URL: https://github.com/apache/datafusion-ballista/pull/2566
# Which issue does this PR close? N/A. This is a scope decision, so I'm happy to open a tracking issue if people would like to discuss it there first. # Rationale for this change Spark compatibility is a very large undertaking. Matching Spark's function semantics, edge cases and behavior across versions is an ongoing commitment that never really finishes, and it would pull a lot of maintainer time away from Ballista's core job as a distributed execution engine for DataFusion. I think the project is better served by not taking that burden on. Users who need Spark-compatible execution already have [Apache DataFusion Comet](https://github.com/apache/datafusion-comet), which accelerates Spark itself and keeps full Spark semantics. The optional `spark-compat` feature is also thin. It only registers the functions from `datafusion-spark`, and it is off by default. # What changes are included in this PR? - Remove the `spark-compat` feature from `ballista-core`, `ballista-scheduler` and `ballista-executor`, and drop the `datafusion-spark` dependency (workspace `Cargo.toml` and `Cargo.lock`) - Simplify `BallistaFunctionRegistry::default()` and `ballista_{scalar,aggregate,window}_functions()` to register DataFusion functions only, and remove the feature-gated tests - Remove the `remote-spark-functions` example and its `Cargo.toml` entry - Remove the Spark-Compatible Functions user guide page, its index entry, and the "Installing with Optional Features" section in the `cargo install` guide, which only covered this feature - Remove `spark-compat` from the feature tables and build examples in the READMEs - Remove the `spark_support` field from the scheduler state REST response, the history server's scheduler state response, and the TUI scheduler info popup, since it only reported this feature Not changed: comparisons with Spark in the introduction, FAQ and benchmark docs, and code comments noting that a Ballista setting mirrors a Spark one. These describe the execution model or the performance comparison, not Spark compatibility. `python/Cargo.lock` is also untouched because the Python package pins published `54.0.0` crates. # Are there any user-facing changes? Yes, this is a breaking change (`api-change` label): - Building with `--features spark-compat` will fail because the feature no longer exists - Spark-only functions such as `sha1` and `expm1` are no longer available - The `spark_support` field is gone from the scheduler state REST response. Clients that require that field, such as an older `ballista-cli` TUI against a newer scheduler, will need to be updated -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
