andygrove opened a new pull request, #6309:
URL: https://github.com/apache/datafusion-comet/pull/6309

   ## Which issue does this PR close?
   
   No issue. This is a documentation refresh of the contributor roadmap, 
following the last pass in #5064.
   
   ## Rationale for this change
   
   Several roadmap sections still describe work as unstarted or blocked when it 
has since landed, and a few references point at issues that have closed or code 
that has moved. Contributors use the roadmap to find where work is coordinated, 
so stale entries send them to the wrong place.
   
   ## What changes are included in this PR?
   
   Only `docs/source/contributor-guide/roadmap.md` changes. Each section was 
checked against main and the linked issues and PRs:
   
   - **Iceberg Table Writes**: the section said no design was committed. The 
split writer/committer plan (#4658) and the iceberg-rust writer (#5361) have 
landed behind two experimental, off-by-default settings. The section now 
describes them, notes that merge-on-read writes aren't intercepted yet (#6240), 
and points at the production-quality epic (#5649) and the default flip (#5644). 
The closed draft #4487 is dropped.
   - **Iceberg Table Format V3 Support**: native deletion vector reads landed 
in #5853, which replaced the closed draft #4887, and the upstream iceberg-rust 
epic has closed. The section now lists what still falls back: row lineage 
metadata columns, column default values, and the 
`variant`/`geometry`/`geography`/`unknown` types. The HDFS note referred to a 
scheme match in `iceberg_scan.rs` that has since moved to `iceberg_common.rs`. 
It now describes the supported storage backends without naming a file.
   - **Native Coverage for Codegen-Dispatched Expressions**: the section said 
aggregate functions go through the codegen-dispatch bridge. The dispatcher only 
handles scalar expressions (`CometBatchKernelCodegen` rejects 
`AggregateFunction`), so incompatible aggregates fall back to Spark. The link 
now points at the expression reference, whose Implementation column lists the 
dispatched expressions; the compatibility guide it pointed at doesn't have that 
list.
   - **Java/Scala UDF Support**: the section said Python UDFs always fall back. 
`mapInArrow` and `mapInPandas` have had experimental support since #4234.
   - **TPC-H and TPC-DS Performance**: the TPC-DS epic #858 is closed, so the 
section now points at #2551.
   - **Upstream Work in DataFusion**: nearly every `datafusion-spark` function 
now has a Comet serde (#4150); the remaining exception, `json_tuple`, is 
tracked in #3160.
   - **Spillable Hash Join**: adds links to #2545 and to the upstream 
DataFusion design for spilling `HashJoinExec` (apache/datafusion#24768).
   - **Delta Lake Support**: describes the delta-spark contrib scan in review 
(#5365) and the proposal to converge it with the `delta-kernel-rs` path 
(#5411). The dormant plain-table draft #4669 is dropped because #5365 
supersedes its approach; I can add it back if we'd rather keep it listed.
   
   The window, lambda, native Parquet write, and memory management sections are 
still accurate and are unchanged.
   
   ## How are these changes tested?
   
   This is a documentation-only change. `prettier --check` passes on the file, 
and every reference-style link it uses is defined, with no unused definitions.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to