mbutrovich commented on code in PR #201:
URL: https://github.com/apache/datafusion-site/pull/201#discussion_r3738395693
##########
content/blog/2026-08-07-datafusion-comet-1.0.0.md:
##########
@@ -0,0 +1,224 @@
+---
+layout: post
+title: Apache DataFusion Comet 1.0.0 Release
+date: 2026-08-07
+author: pmc
+categories: [subprojects]
+---
+
+<!--
+{% comment %}
+Licensed to the Apache Software Foundation (ASF) under one or more
+contributor license agreements. See the NOTICE file distributed with
+this work for additional information regarding copyright ownership.
+The ASF licenses this file to you under the Apache License, Version 2.0
+(the "License"); you may not use this file except in compliance with
+the License. You may obtain a copy of the License at
+
+http://www.apache.org/licenses/LICENSE-2.0
+
+Unless required by applicable law or agreed to in writing, software
+distributed under the License is distributed on an "AS IS" BASIS,
+WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+See the License for the specific language governing permissions and
+limitations under the License.
+{% endcomment %}
+-->
+
+[TOC]
+
+The Apache DataFusion PMC is pleased to announce version 1.0.0 of the
[Comet](https://datafusion.apache.org/comet/) subproject.
+
+Comet is an accelerator for Apache Spark that translates Spark physical plans
to DataFusion physical plans for
+improved performance and efficiency without requiring any code changes.
+
+This release covers roughly six weeks of development since 0.17.0 and consists
of 244 commits from 23
+contributors. See the [change log] for the full list of changes.
+
+[change log]:
https://github.com/apache/datafusion-comet/blob/main/docs/source/changelog/1.0.0.md
+
+## The Road to 1.0
+
+Comet was [donated] to the Apache DataFusion project in March 2024 and cut its
first release, 0.1.0, five
+months later with support for 13 operators and 106 expressions. Since then,
the project has shipped 20
+releases and drawn contributions from more than 120 developers, and the
codebase now recognizes over 400
+Spark expressions. Operator coverage has grown alongside it: 1.0 accelerates
each of Spark's four join
+operators, window functions, generators (`explode`, `explode_outer`,
`posexplode`, and `posexplode_outer`
+over arrays), sampling, in-memory table scans, and a fully native shuffle.
+
+[donated]: https://datafusion.apache.org/blog/2024/03/06/comet-donation/
+
+The 1.0 release marks the point at which Comet begins following [semantic
versioning]. Users upgrading
+within the 1.x line can expect backward-compatible changes only; features
slated for removal will be
+deprecated in a minor release before being dropped in the next major version.
This is why the deprecations
+of JDK 11 and Spark 3.4 announced below are scheduled for 1.1 rather than
landing in 1.0 itself.
+
+[semantic versioning]:
https://datafusion.apache.org/comet/about/versioning_policy.html
+
+### Support for Spark 4.0+ with ANSI mode
+
+Comet 1.0.0 supports Spark versions 3.4 through 4.1, with experimental support
for 4.2. Comet fully supports Spark's ANSI mode, which is enabled by default
starting with Spark 4.0.
+
+### Correctness Testing
+
+It is important that queries accelerated by Comet produce the same results as
Spark. Correctness checking has always been a large effort in Comet
development, but the approach has evolved over time.
+
+- **Upstream Spark tests**: Comet runs Spark's own test suite with Comet
enabled, providing more than 24,000 unit tests effectively for free. These
tests run in Comet's CI for all supported Spark versions.
+- **Scala tests**: end-to-end queries that run with Comet enabled versus
disabled, checking that results match.
+- **Fuzz testing**: many of the Scala tests generate randomized data to catch
regressions around edge cases such as nulls, NaN, Infinity, and timezone issues.
+- **Comet SQL tests**: a sqllogictest-inspired approach that makes end-to-end
tests easier to write.
+- **Generative AI audits**: agentic skills sweep every expression, comparing
Comet's implementation to Spark's source and ensuring tests cover important
edge cases.
+
+### Performance
+
+The early Comet releases provided a very modest speedup and the published
benchmark results were based on running TPC workloads at small scale factors on
a single node. There are now independent benchmark results published by AWS
Labs that show significant speedups for [TPC-DS @ 3TB running in
EKS](https://awslabs.github.io/data-on-eks/docs/benchmarks/spark-datafusion-comet-benchmark).
+
+### Codegen Dispatch
+
+Comet 0.17.0 introduced a new approach to filling gaps in expression coverage.
In earlier releases,
+whenever Comet's planner encountered an expression that lacked a native Rust
implementation, it fell back
+to executing an entire subtree of the plan in Spark. That required converting
Arrow columns back to Spark
+rows before the expression ran and back to Arrow after, and the cost was often
enough to erase the speedup
+Comet had bought elsewhere in the plan.
Review Comment:
These ideas seem in conflict with each other. It talks about the an entire
subtree executing in Spark, but then it mentioned converting back and forth to
Arrow. Is it fair to say that subsequent stage operators ran in Spark, and then
the next stage would pay the cost of converting back, or is that still not
quite it? Maybe we just generalize it a bit and just say that fallbacks
incurred expensive Arrow transitions and not worry about stages and subtrees.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]