andygrove opened a new issue, #6665:
URL: https://github.com/apache/datafusion-comet/issues/6665

   Triage pass over the open `requires-triage` queue, per the project [Bug 
Triage 
Guide](https://github.com/apache/datafusion-comet/blob/main/docs/source/contributor-guide/bug_triage.md).
   
   - Date: 2026-10-05
   - Total issues processed: 87 (82 triaged, 5 skipped, 0 failed)
   - Type counts: 31 bugs, 51 enhancements
   - Priority counts applied: `priority:critical` 10, `priority:high` 0, 
`priority:medium` 17, `priority:low` 4
   - Guide: 
[docs/source/contributor-guide/bug_triage.md](https://github.com/apache/datafusion-comet/blob/main/docs/source/contributor-guide/bug_triage.md)
   
   Labels have already been applied. A reviewer should spot-check the calls 
below and close this issue when satisfied. Corrections should be made directly 
on the affected issue.
   
   Notes on this pass:
   
   - **Reviewer corrections since the last two passes.** No priority was 
changed on any issue from the 2026-09-21 or 2026-09-28 passes. After the 
2026-09-28 pass, kazuyukitanimura added `correctness` to five issues that pass 
had labeled critical (#6086, #6131, #6136, #6264 and #6278), and `regression` 
was added to #6254. Neither label is in the guide, so this pass adds neither. 
The critical bugs below that don't carry `correctness` are #6470, #6476, #6477, 
#6519, #6549, #6570 and #6573.
   - **No type reclassifications.** Every `bug` or `enhancement` label an 
author had already applied matched the issue's content. The 18 issues that 
arrived without a type label were classified from their bodies.
   - **No author priority was changed.** Author priorities were confirmed on 
#6333, #6335, #6385, #6414, #6454, #6525, #6562 and #6613.
   - **Opt-in paths.** #6573 needs `spark.comet.convert.parquet.enabled`, which 
is off by default, and is labeled critical, as #4793 was under the same 
setting. The native Iceberg writer metadata bugs #6562 and #6659 stay at medium 
with #6146, since query results are unaffected. #6482 and #6532 are seen only 
through native `to_json`, an `Incompatible` opt-in, so they follow #5131 at 
medium. See the escalations.
   - The guide lists `spark 4` as an area indicator, but the repository only 
has `spark 4.0` / `spark 4.1` / `spark 4.2` / `spark 3.x`, so nothing was 
applied for it to the Spark 4-only issues #6470, #6549, #6612 and #6633.
   - `regression` is not in the guide, so this pass didn't add it. It fits 
#6482 and #6645, which come from changes on `main` after 1.1.0 (#6350 and 
#6547).
   - Three sub-issues of #6615 describe suspected product bugs that have no 
issue of their own, and this pass didn't file them. #6622 cites a 
`CometNativeCastSuite` TODO: native `try_cast` of a float or double equal to 
2^63 to `BIGINT` returns NULL where Spark returns `Long.MaxValue`. #6637 says 
EXPLAIN on Spark 4.x may misreport dispatched `json_array_length` and 
`to_json`. #6634 says `bit_get`'s out-of-range error text matches neither Spark 
3.x nor 4.x. The first would be silent wrong results if confirmed.
   - #6661 was opened while this pass was running, and is included.
   
   ## Bugs
   
   ### priority:critical
   
   - [EPIC] Timezone handling bugs 
([#6335](https://github.com/apache/datafusion-comet/issues/6335))
     - Area labels: `area:scan`, `area:expressions`; also carries `EPIC`, 
`correctness`
     - Rationale: An EPIC for timezone bugs, several of which returned silent 
wrong results. Confirms the author's label; see the escalations, since its only 
critical child is closed.
   - [EPIC] Match Spark's -0.0 and NaN semantics systematically instead of per 
expression ([#6385](https://github.com/apache/datafusion-comet/issues/6385))
     - Area labels: `area:expressions`; also carries `EPIC`, `correctness`
     - Rationale: The reproductions return different answers from Spark on 
`main` with default configs (`max`/`min`, `hash`, comparisons in aggregates and 
joins, `greatest`/`least`, array functions), and children such as #6157 and 
#6519 are still open. Confirms the author's label.
   - array_contains, arrays_overlap, array_distinct and array_union ignore 
string collation 
([#6470](https://github.com/apache/datafusion-comet/issues/6470))
     - Area labels: `area:expressions`
     - Rationale: On Spark 4.x with default configs, the four native kernels 
compare collated strings by raw bytes, so `array_contains` returns `false` 
where Spark returns `true` and the set functions keep values that Spark merges. 
Step 1, the same class as #6158.
   - Native sort orders null elements of array and struct keys by the key's 
null order, unlike Spark 
([#6476](https://github.com/apache/datafusion-comet/issues/6476))
     - Area labels: none
     - Rationale: Under `ASC NULLS LAST` or `DESC NULLS FIRST`, a key holding a 
null element sorts at the opposite end from Spark, which also changes which 
rows a TopK keeps and how window order keys rank. That is silent wrong results 
at step 1, like the collated sort key in #6158.
   - RANGE window frames over an array or struct key with a null element span 
the whole partition 
([#6477](https://github.com/apache/datafusion-comet/issues/6477))
     - Area labels: none
     - Rationale: `SUM(id) OVER (ORDER BY array(i))` returns the whole 
partition's total on the null-element row and every row after it, where Spark 
returns running sums. The plan is native and raises no error, so this is step 1.
   - Native corr, covariance, variance and stddev return wrong values for a 
constant fractional column merged from several partitions 
([#6481](https://github.com/apache/datafusion-comet/issues/6481))
     - Area labels: `area:aggregation`; also carries `correctness`
     - Rationale: On default configs `corr` returns 0.878 where Spark returns 
NULL, or raises `DIVIDE_BY_ZERO` under ANSI, and the variance and covariance 
functions return tiny non-zero values instead of 0.0. That is silent wrong 
results at step 1, including the "Comet returns a value where Spark raises" 
shape. It has been present since 1.0.0.
   - percentile_approx returns different percentiles from Spark when its input 
holds a NaN with the sign bit set 
([#6519](https://github.com/apache/datafusion-comet/issues/6519))
     - Area labels: `area:aggregation`
     - Rationale: `CometApproxPercentile` reports `Compatible` for floats and 
returns different percentiles at every percentage once a partition merges more 
than one head buffer. Arithmetic produces sign-bit NaNs on x86-64, so this is 
step 1, the same class as the other #6385 children.
   - Native map construction doesn't match Spark 4.0+ float key normalization 
(-0.0 keys, missing DUPLICATED_MAP_KEY) 
([#6549](https://github.com/apache/datafusion-comet/issues/6549))
     - Area labels: `area:expressions`
     - Rationale: On Spark 4.0+ with default configs, `map_from_entries` keeps 
a `-0.0` key where Spark returns `0.0`, and both `map_from_entries` and 
`map_from_arrays` return a map where Spark raises `DUPLICATED_MAP_KEY`. Both 
serdes report `Compatible`, so this is step 1, and it is the "Comet returns a 
value where Spark raises" shape that reviewers raised to critical in #5801 and 
#5936.
   - Codegen dispatcher initializes kernels with an incorrect partition index 
under UNION ALL and coalesce 
([#6570](https://github.com/apache/datafusion-comet/issues/6570))
     - Area labels: `area:expressions`
     - Rationale: The dispatcher initializes each kernel with 
`TaskContext.partitionId()`, so a dispatched `spark_partition_id()`, 
`monotonically_increasing_id()`, `rand` or `uuid` in a later `UNION ALL` 
branch, or below `CometCoalesceExec`, sees the task's partition index instead 
of the index of the partition being computed, which Spark uses. The dispatcher 
is on by default (`spark.comet.exec.scalaUDF.codegen.enabled=true`), so this is 
silent wrong results at step 1 (see escalations).
   - input_file_name() returns empty values above a converted Spark Parquet 
scan ([#6573](https://github.com/apache/datafusion-comet/issues/6573))
     - Area labels: `area:scan`
     - Rationale: With `spark.comet.convert.parquet.enabled=true`, 
`input_file_name()` and the two block columns return `""`, `-1` and `-1` for 
every row and the query succeeds, which is step 1. The conversion is off by 
default; #4793, a silent wrong count under the same setting, was labeled 
critical (see escalations).
   
   ### priority:medium
   
   - Native aggregate fails after spilling when groups are few and large 
(collect_list, collect_set) 
([#6363](https://github.com/apache/datafusion-comet/issues/6363))
     - Area labels: `area:aggregation`
     - Rationale: The task fails with `Additional allocation failed` where 
Spark completes, and more off-heap memory or turning off the native aggregate 
avoids it. A visible failure with workarounds is step 3, as with #6254 (see 
escalations).
   - [EPIC] Regressions in 1.1.0 since 1.0.0 
([#6402](https://github.com/apache/datafusion-comet/issues/6402))
     - Area labels: none; also carries `EPIC`
     - Rationale: An EPIC of confirmed regressions, which are bugs under the 
guide's type table. Every entry except #6254 is fixed on `branch-1.1`, and 
#6254 is a visible failure under memory pressure with workarounds, so step 3.
   - Comet shuffle does not report Spark 4.1's order-independent shuffle 
checksum ([#6414](https://github.com/apache/datafusion-comet/issues/6414))
     - Area labels: `area:shuffle`; also carries `spark 4.1`, `spark 4.2`
     - Rationale: With the opt-in `orderIndependentChecksum` configs, every 
Comet map output reports a checksum of 0, so Spark's protection against mixed 
retry output is silently off. The configs are off by default and nothing is 
wrong unless an indeterminate stage is retried, so step 3. Confirms the 
author's label.
   - Investigate TPC-DS q54 and q68 slowdowns in the 1.1.0 benchmark on Spark 
4.2 ([#6418](https://github.com/apache/datafusion-comet/issues/6418))
     - Area labels: none; also carries `performance`, `spark 4.2`
     - Rationale: Comet's q68 went from 1.98x faster than Spark (1.0.0 on Spark 
3.5.8) to 0.60x (1.1.0 on Spark 4.2.0). That is significant performance 
degradation at step 3, like the Spark 4.2 fallbacks #4949 and #5834. The cause 
isn't pinned on Comet yet, and q54 is a long-standing gap.
   - AQE does not coalesce shuffle partitions under a CometUnion when a branch 
is not a shuffle 
([#6454](https://github.com/apache/datafusion-comet/issues/6454))
     - Area labels: `area:shuffle`; also carries `performance`
     - Rationale: A small union runs hundreds of near-empty tasks, because 
`CoalesceShufflePartitions` matches Spark's `UnionExec` and not 
`CometUnionExec`. Results are correct, so this is performance degradation at 
step 3. Confirms the author's label.
   - Native CASE WHEN names its struct result's fields after the ELSE branch 
instead of the first THEN branch 
([#6482](https://github.com/apache/datafusion-comet/issues/6482))
     - Area labels: `area:expressions`; also carries `correctness`
     - Rationale: Values are right, and only native code that reads field names 
sees the ELSE branch's names. The reproduction needs native `to_json` under 
`StructsToJson.allowIncompatible=true`; default `to_json` dispatches and is 
correct. So it follows the #5131 rule at step 3 (see escalations).
   - Reduce retained buffer allocation when map lookups select short nested 
values ([#6525](https://github.com/apache/datafusion-comet/issues/6525))
     - Area labels: `area:expressions`
     - Rationale: The map counterpart of #6225: values are correct, but the 
output retains about 2 MiB of child buffer per 8,192-row batch, which native 
shuffle charges against its reservation. Step 3, matching #6225, and confirms 
the author's label.
   - AQE skew join split makes the join fall back to Spark with no fallback 
reason ([#6530](https://github.com/apache/datafusion-comet/issues/6530))
     - Area labels: `area:shuffle`; also carries `performance`
     - Rationale: With AQE's skew join handling, which is on by default, a 
split join and everything above it in the stage run in Spark with no fallback 
reason. That cost 38% to 87% more task time in the aggregation cases of the 
reporter's benchmark. Results are correct, so this is significant performance 
degradation at step 3, like the plan-shape fallbacks #4949 and #5834.
   - Native CASE WHEN and COALESCE reconcile struct fields by name instead of 
position ([#6532](https://github.com/apache/datafusion-comet/issues/6532))
     - Area labels: `area:expressions`; also carries `correctness`
     - Rationale: When branch structs have case-distinct field names in swapped 
positions, the native common type can widen the wrong field. The issue observes 
it only through native `to_json`, which needs 
`StructsToJson.allowIncompatible=true`, so it follows the #5131 rule for 
`Incompatible` opt-ins (see escalations).
   - S3 credential SPI: follow-ups from #6478 (hostless Iceberg metadata 
locations, docs, tests) 
([#6536](https://github.com/apache/datafusion-comet/issues/6536))
     - Area labels: `area:scan`; also carries `documentation`, `test`, 
`area:Iceberg`
     - Rationale: Item 1 is a defect: a native Iceberg scan of a table with a 
hostless alias metadata location never calls the configured credential provider 
and signs with the default chain, so reads fail with 403 or use broader 
credentials than the provider vends. It needs an opted-in alias scheme and was 
found by reading the code, so step 3 (see escalations). The other items are 
docs and tests.
   - Native Iceberg writer counts NaNs under NULL structs and NULL list or map 
entries ([#6562](https://github.com/apache/datafusion-comet/issues/6562))
     - Area labels: `area:writer`; also carries `area:Iceberg`
     - Rationale: `nan_value_counts` is wrong for struct fields on every 
profile, and for list elements before Iceberg 1.10, while the data and query 
results match. Wrong metadata from the off-by-default native writer is step 3, 
as with #6146. Confirms the author's label.
   - Native scans resolve the AWS SDK ProfileCredentialsProvider names 
differently from Hadoop 
([#6575](https://github.com/apache/datafusion-comet/issues/6575))
     - Area labels: `area:scan`
     - Rationale: For the AWS SDK profile provider names on Spark 4.x, a few 
profile shapes make native reads sign with different credentials, or call a 
different STS endpoint, than Hadoop. The reporter notes that these are uncommon 
profile shapes and that a plain profile works, so step 3 (see escalations).
   - Native Azure credential lookup still differs from Hadoop ABFS in a few 
configurations ([#6605](https://github.com/apache/datafusion-comet/issues/6605))
     - Area labels: `area:scan`
     - Rationale: In the configurations listed (container-scoped OAuth keys on 
Hadoop 3.4.2+, Fabric hosts, stale account-scoped key forms, credential 
provider stores), the native scan either fails or authenticates with a 
different credential than Hadoop's ABFS driver would. Each needs an uncommon 
setup, so step 3 (see escalations).
   - Array lookup functions evaluate their second argument on rows where the 
array is NULL, raising ANSI errors that Spark skips 
([#6613](https://github.com/apache/datafusion-comet/issues/6613))
     - Area labels: `area:expressions`
     - Rationale: Under ANSI, Comet fails a query that Spark completes. The 
failure is visible, and ANSI off or `try_cast` avoids it, so step 3. This is 
the opposite direction from the "Comet returns a value where Spark raises" 
corrections. Confirms the author's label.
   - Spark 3.4 SQL test SPARK-34637 fails since #6547 because AQE re-plans flip 
the DPP join's build side 
([#6645](https://github.com/apache/datafusion-comet/issues/6645))
     - Area labels: `area:shuffle`, `spark sql tests`
     - Rationale: Since #6547, a re-plan on Spark 3.4 flips the join's build 
side, so the DPP subquery no longer reuses the join's broadcast. The test is 
right and the plan changed, so this is a product regression on default configs 
rather than a test-only failure. It has a workaround 
(`spark.comet.shuffle.directRead.enabled=false`), so step 3 (see escalations).
   - Native Iceberg writer's value and null counts for a float under a NULL 
struct differ from iceberg-java 1.9+ 
([#6659](https://github.com/apache/datafusion-comet/issues/6659))
     - Area labels: `area:writer`; also carries `area:Iceberg`
     - Rationale: The data file metrics differ from iceberg-java's, but the 
issue shows that the native counts never prune more files than iceberg-java's, 
so query results are unaffected. Wrong metadata from the off-by-default native 
writer is step 3, as with #6146.
   - ANSI integral SUM overflow reports "integer overflow" without Spark's 
try_add suggestion 
([#6661](https://github.com/apache/datafusion-comet/issues/6661))
     - Area labels: `area:aggregation`
     - Rationale: Comet throws where Spark throws, but with `integer overflow` 
and no `try_add` suggestion instead of Spark's `long overflow` parameters. Only 
the error parameters differ, which is the tier of #6217 and its predecessor 
#5071, so step 3.
   
   ### priority:low
   
   - days transform is evaluated in the session timezone, while hours and 
Iceberg use UTC 
([#6333](https://github.com/apache/datafusion-comet/issues/6333))
     - Area labels: `area:expressions`
     - Rationale: Spark never evaluates these partition transforms, so there is 
no Spark answer to match, and the author rated it low. Confirms the author's 
label (see escalations).
   - Broadcast/hash join fallback reasons are lost from the AQE-final plan 
([#6442](https://github.com/apache/datafusion-comet/issues/6442))
     - Area labels: none
     - Rationale: Only the `EXPLAIN` annotation is lost under AQE. The join 
still falls back, and the reason still reaches the fallback log. Diagnostics 
only, step 4.
   - test: `query tolerance=` in Comet SQL tests passes when either side is NaN 
([#6616](https://github.com/apache/datafusion-comet/issues/6616))
     - Area labels: none; also carries `test`
     - Rationale: A test harness defect: NaN results under `tolerance=` are 
never compared, so those fixtures can't catch a wrong NaN. Test-only, step 4, 
as with #6203. Any real mismatch that the fix exposes should be filed 
separately.
   - Flaky test: CometIcebergWriteActionSuite "a failed write job deletes the 
data files of tasks that completed" 
([#6643](https://github.com/apache/datafusion-comet/issues/6643))
     - Area labels: `area:writer`; also carries `test`, `area:Iceberg`
     - Rationale: An intermittent test failure, step 4. The likely cause is in 
the test: its UDF blocks a native runtime thread while it waits for the other 
tasks.
   
   ## Enhancements
   
   - [EPIC] DataFusion 56 upgrade: regressions found by tracking DataFusion 
main ([#6410](https://github.com/apache/datafusion-comet/issues/6410))
     - Area labels: none; also carries `EPIC`
     - Rationale: Tracks a dependency upgrade. The regression it lists exists 
only on the unmerged draft PR #6404.
   - Separate JNI-free core logic and error classification from native JNI 
entry points ([#6434](https://github.com/apache/datafusion-comet/issues/6434))
     - Area labels: `area:ffi`
     - Rationale: A refactor with no user-visible change.
   - to_csv: ignoreLeadingWhiteSpace / ignoreTrailingWhiteSpace trim Unicode 
whitespace instead of univocity's chars <= ' ' 
([#6446](https://github.com/apache/datafusion-comet/issues/6446))
     - Area labels: `area:expressions`
     - Rationale: `to_csv` is `Incompatible`, so its divergences are filed as 
enhancements. Confirms the author's label.
   - Trying to reduce CI time 
([#6465](https://github.com/apache/datafusion-comet/issues/6465))
     - Area labels: `area:ci`
     - Rationale: A CI speedup.
   - Investigate nested TPC-H q21 slowdown: Comet 37% slower than Spark at 
SF1000 ([#6467](https://github.com/apache/datafusion-comet/issues/6467))
     - Area labels: none; also carries `performance`
     - Rationale: A performance gap against Spark, measured on a downstream 
build and not reproduced on `main`, with no regression identified. If the 
rising iteration times turn out to be a leak, that part is a bug.
   - Native existence join enumerates every duplicate build match (N*M 
candidates for M markers) 
([#6484](https://github.com/apache/datafusion-comet/issues/6484))
     - Area labels: none
     - Rationale: Results are correct. This is a performance limit of 
DataFusion's mark join on skewed build keys, deferred by design in #4587.
   - Assess native (Rust) code generation for fused expression evaluation 
([#6485](https://github.com/apache/datafusion-comet/issues/6485))
     - Area labels: `area:expressions`; also carries `performance`
     - Rationale: Performance research.
   - Experimental local execution mode that runs a whole query as one 
DataFusion graph 
([#6499](https://github.com/apache/datafusion-comet/issues/6499))
     - Area labels: none
     - Rationale: A new opt-in execution mode.
   - Native Parquet scan assigns row groups to splits differently from Spark 
([#6512](https://github.com/apache/datafusion-comet/issues/6512))
     - Area labels: `area:scan`
     - Rationale: Matching Spark's row group assignment, through an upstream 
change. The wrong `_metadata` values it caused already fall back (#6505, 
#6510). Confirms the author's label.
   - Support native existence sort-merge joins 
([#6514](https://github.com/apache/datafusion-comet/issues/6514))
     - Area labels: none
     - Rationale: New operator support. The fallback is intentional today.
   - Run array_contains on float elements natively with Spark's equality 
instead of through the codegen dispatcher 
([#6520](https://github.com/apache/datafusion-comet/issues/6520))
     - Area labels: `area:expressions`
     - Rationale: Moves an `Incompatible` path, which dispatches today, to a 
native kernel.
   - Clarify how to run native Comet -> Celeborn shuffle end-to-end 
([#6523](https://github.com/apache/datafusion-comet/issues/6523))
     - Area labels: `area:shuffle`
     - Rationale: A request to document what a Celeborn client must provide for 
the native path. The fallback with released clients is intended and documented.
   - perf: Comet hash joins use more task time than Spark (skewed SHJ with 
nested payload, BHJ) 
([#6528](https://github.com/apache/datafusion-comet/issues/6528))
     - Area labels: none; also carries `performance`
     - Rationale: A performance gap against Spark with correct results and no 
regression identified, so an optimization request. Its skew-split part is the 
bug #6530.
   - Comet 1.2.0 Release 
([#6550](https://github.com/apache/datafusion-comet/issues/6550))
     - Area labels: none
     - Rationale: Release planning, labeled like the 1.1.0 tracker #5327. 
Confirms the author's label.
   - Support OneRowRelation as a native Comet source 
([#6555](https://github.com/apache/datafusion-comet/issues/6555))
     - Area labels: none
     - Rationale: A new native source for `OneRowRelation`.
   - [EPIC] Faster Spark-to-Arrow conversion in CometSparkToColumnarExec 
([#6565](https://github.com/apache/datafusion-comet/issues/6565))
     - Area labels: `area:scan`; also carries `performance`, `EPIC`
     - Rationale: A performance EPIC for `CometSparkToColumnarExec`.
   - My priorities for 1.2.0 
([#6580](https://github.com/apache/datafusion-comet/issues/6580))
     - Area labels: none
     - Rationale: A personal planning tracker. Confirms the author's label, as 
with release trackers such as #5327.
   - Remove the spill replay workaround from the memory pools once DataFusion 
includes apache/datafusion#25383 
([#6583](https://github.com/apache/datafusion-comet/issues/6583))
     - Area labels: `area:aggregation`; also carries `area:memory`
     - Rationale: Cleanup once Comet is on a DataFusion release with 
apache/datafusion#25383.
   - Compare floats in Spark's order without normalizing copies of the operands 
([#6590](https://github.com/apache/datafusion-comet/issues/6590))
     - Area labels: `area:expressions`; also carries `performance`
     - Rationale: A performance optimization of the comparisons that #6447 made 
correct.
   - JVM columnar shuffle spends most of its string and map conversion time on 
per-value overhead 
([#6596](https://github.com/apache/datafusion-comet/issues/6596))
     - Area labels: `area:shuffle`; also carries `performance`
     - Rationale: A performance optimization of the JVM columnar shuffle's 
row-to-Arrow converter.
   - JVM columnar shuffle's dictionary trial costs about 10 ns per row on 
high-cardinality strings 
([#6597](https://github.com/apache/datafusion-comet/issues/6597))
     - Area labels: `area:shuffle`; also carries `performance`
     - Rationale: A performance optimization. It also asks for a correction to 
a config's doc.
   - Use native shuffle for row input by converting it with 
CometSparkToColumnarExec 
([#6598](https://github.com/apache/datafusion-comet/issues/6598))
     - Area labels: `area:shuffle`; also carries `performance`
     - Rationale: A performance optimization behind a new config that is off by 
default.
   - Use RoaringTreemap for MergeRows cardinality tracking 
([#6608](https://github.com/apache/datafusion-comet/issues/6608))
     - Area labels: `area:writer`
     - Rationale: A memory optimization of native MergeRows' cardinality state.
   - A Spark operator reading a Comet shuffle through AQE's coalesced read gets 
Spark's ColumnarToRow instead of Comet's 
([#6610](https://github.com/apache/datafusion-comet/issues/6610))
     - Area labels: `area:shuffle`
     - Rationale: Using Comet's transition over a coalesced shuffle read. The 
issue finds no measurable cost today. Confirms the author's label.
   - Optimize MergeRows predicate evaluation and row selection 
([#6611](https://github.com/apache/datafusion-comet/issues/6611))
     - Area labels: `area:writer`
     - Rationale: A performance optimization of native MergeRows.
   - Support Spark 4.2 insert-only MERGE execution 
([#6612](https://github.com/apache/datafusion-comet/issues/6612))
     - Area labels: `area:writer`
     - Rationale: New native support for Spark 4.2's insert-only MERGE.
   - Share one Spark-equality match kernel between array_contains, 
array_position and array_remove 
([#6614](https://github.com/apache/datafusion-comet/issues/6614))
     - Area labels: `area:expressions`
     - Rationale: A refactor to one shared kernel. The ANSI bug it would also 
fix is tracked separately as #6613.
   - [EPIC] Move expression test coverage from Scala suites to Comet SQL tests 
([#6615](https://github.com/apache/datafusion-comet/issues/6615))
     - Area labels: `area:expressions`; also carries `test`, `EPIC`
     - Rationale: A test migration EPIC with no product change.
   - test: add a Comet SQL test error mode that checks native execution and 
error-class parity 
([#6617](https://github.com/apache/datafusion-comet/issues/6617))
     - Area labels: none; also carries `test`
     - Rationale: A new SQL test framework query mode.
   - test: let Comet SQL tests keep ConstantFolding, exclude optimizer rules 
and change configs mid-file 
([#6618](https://github.com/apache/datafusion-comet/issues/6618))
     - Area labels: none; also carries `test`
     - Rationale: New SQL test framework directives. The `Config` parser fix it 
includes affects only fixture values that contain `=`.
   - docs: document the ways a Comet SQL test can silently miss Comet 
([#6619](https://github.com/apache/datafusion-comet/issues/6619))
     - Area labels: none; also carries `documentation`, `test`
     - Rationale: A contributor guide addition. Documentation is an enhancement 
under the guide's type table.
   - test: move math and arithmetic coverage from Scala suites to Comet SQL 
tests ([#6620](https://github.com/apache/datafusion-comet/issues/6620))
     - Area labels: `area:expressions`; also carries `test`
     - Rationale: Test migration under #6615.
   - test: move decimal arithmetic coverage from Scala suites to Comet SQL 
tests ([#6621](https://github.com/apache/datafusion-comet/issues/6621))
     - Area labels: `area:expressions`; also carries `test`
     - Rationale: Test migration under #6615.
   - test: move cast coverage for boolean, integral, floating-point and decimal 
sources to Comet SQL tests 
([#6622](https://github.com/apache/datafusion-comet/issues/6622))
     - Area labels: `area:expressions`; also carries `test`
     - Rationale: Test migration under #6615. About 20 of the Scala tests it 
replaces compare Comet with Comet, which the move fixes.
   - test: move cast coverage for string sources to Comet SQL tests 
([#6623](https://github.com/apache/datafusion-comet/issues/6623))
     - Area labels: `area:expressions`; also carries `test`
     - Rationale: Test migration under #6615.
   - test: move cast coverage for date, timestamp, binary and complex-type 
sources to Comet SQL tests 
([#6624](https://github.com/apache/datafusion-comet/issues/6624))
     - Area labels: `area:expressions`; also carries `test`
     - Rationale: Test migration under #6615.
   - test: move datetime expression coverage from Scala suites to Comet SQL 
tests ([#6625](https://github.com/apache/datafusion-comet/issues/6625))
     - Area labels: `area:expressions`; also carries `test`
     - Rationale: Test migration under #6615.
   - test: move string expression coverage from Scala suites to Comet SQL tests 
([#6626](https://github.com/apache/datafusion-comet/issues/6626))
     - Area labels: `area:expressions`; also carries `test`
     - Rationale: Test migration under #6615.
   - test: move regex coverage to Comet SQL tests and fix fixtures whose 
patterns lose their backslashes 
([#6627](https://github.com/apache/datafusion-comet/issues/6627))
     - Area labels: `area:expressions`; also carries `test`
     - Rationale: Test migration under #6615. Seven regex fixtures never run 
the patterns they name, but that is a test gap that the move fixes, not a 
product defect.
   - test: move array expression coverage from Scala suites to Comet SQL tests 
([#6628](https://github.com/apache/datafusion-comet/issues/6628))
     - Area labels: `area:expressions`; also carries `test`
     - Rationale: Test migration under #6615.
   - test: move map and struct expression coverage from Scala suites to Comet 
SQL tests ([#6629](https://github.com/apache/datafusion-comet/issues/6629))
     - Area labels: `area:expressions`; also carries `test`
     - Rationale: Test migration under #6615.
   - test: move explode/posexplode coverage from CometGenerateExecSuite to 
Comet SQL tests 
([#6630](https://github.com/apache/datafusion-comet/issues/6630))
     - Area labels: `area:expressions`; also carries `test`
     - Rationale: Test migration under #6615.
   - test: move aggregate function coverage from CometAggregateSuite to Comet 
SQL tests ([#6631](https://github.com/apache/datafusion-comet/issues/6631))
     - Area labels: `area:aggregation`; also carries `test`
     - Rationale: Test migration under #6615.
   - test: move window function coverage from CometWindowExecSuite to Comet SQL 
tests ([#6632](https://github.com/apache/datafusion-comet/issues/6632))
     - Area labels: `area:expressions`; also carries `test`
     - Rationale: Test migration under #6615.
   - test: move the Spark 4 collation suites to Comet SQL tests 
([#6633](https://github.com/apache/datafusion-comet/issues/6633))
     - Area labels: `area:expressions`; also carries `test`
     - Rationale: Test migration under #6615, for the Spark 4 collation suites.
   - test: move hash and bitwise expression coverage from Scala suites to Comet 
SQL tests ([#6634](https://github.com/apache/datafusion-comet/issues/6634))
     - Area labels: `area:expressions`; also carries `test`
     - Rationale: Test migration under #6615.
   - test: move conditional, JSON, CSV, Variant and other expression coverage 
to Comet SQL tests 
([#6635](https://github.com/apache/datafusion-comet/issues/6635))
     - Area labels: `area:expressions`; also carries `test`
     - Rationale: Test migration under #6615.
   - test: move CometFloatSemanticsSuite's matching cases to Comet SQL tests 
([#6636](https://github.com/apache/datafusion-comet/issues/6636))
     - Area labels: `area:expressions`; also carries `test`
     - Rationale: Test migration under #6615.
   - test: move expression-level codegen dispatcher tests to Comet SQL tests 
([#6637](https://github.com/apache/datafusion-comet/issues/6637))
     - Area labels: `area:expressions`; also carries `test`
     - Rationale: Test migration under #6615.
   - [EPIC] Adaptive runtime pruning for native Iceberg scans 
([#6641](https://github.com/apache/datafusion-comet/issues/6641))
     - Area labels: `area:scan`
     - Rationale: New runtime pruning for native Iceberg scans.
   - perf: explore reducing copies in native Celeborn shuffle writes 
([#6654](https://github.com/apache/datafusion-comet/issues/6654))
     - Area labels: `area:shuffle`
     - Rationale: A performance exploration of the native Celeborn write path. 
Nothing is broken.
   
   ## Escalations to consider
   
   - Codegen dispatcher initializes kernels with an incorrect partition index 
under UNION ALL and coalesce 
([#6570](https://github.com/apache/datafusion-comet/issues/6570))
     - Labeled critical from the code: `CometScalaUDFCodegen` initializes 
kernels with `TaskContext.partitionId()` on `main`, and the dispatcher is on by 
default. The report comes from the #5875 review and has no standalone 
reproducer yet, so the reviewer may want one before relying on the label.
   - input_file_name() returns empty values above a converted Spark Parquet 
scan ([#6573](https://github.com/apache/datafusion-comet/issues/6573))
     - Labeled critical, but it needs 
`spark.comet.convert.parquet.enabled=true`, which is off by default. #4793 had 
the same precondition and was labeled critical, as #5719 was for another 
off-by-default feature. The reviewer may prefer `priority:high` under the "core 
path over experimental" principle.
   - [EPIC] Timezone handling bugs 
([#6335](https://github.com/apache/datafusion-comet/issues/6335))
     - The author's `priority:critical` is confirmed, but its only critical 
child, #6328, is closed. The open children are #5633 (`priority:high`) and 
#6333 (`priority:low`), so `priority:high` fits if an EPIC's priority should 
follow its open children.
   - Native aggregate fails after spilling when groups are few and large 
(collect_list, collect_set) 
([#6363](https://github.com/apache/datafusion-comet/issues/6363))
     - Matches the trigger "A `priority:medium` bug is reported by multiple 
users or affects a common workload → consider escalating to `priority:high`". 
Any `collect_list` or `collect_set` over a low-cardinality key that spills can 
hit it. The last pass made the same note for #6254.
   - Native CASE WHEN and COALESCE reconcile struct fields by name instead of 
position ([#6532](https://github.com/apache/datafusion-comet/issues/6532))
     - Held at medium because the issue observes the widened field only through 
native `to_json`, which needs `allowIncompatible`. The wrong field type could 
also reach native consumers that run by default, such as `hash` or a cast to 
string. If one of them returns a different answer from Spark, step 1 applies.
   - Native CASE WHEN names its struct result's fields after the ELSE branch 
instead of the first THEN branch 
([#6482](https://github.com/apache/datafusion-comet/issues/6482))
     - Held at medium on the same basis as #6532: only native `to_json` under 
`allowIncompatible` exposes the ELSE branch's field names. If a default-config 
consumer that reads field names turns up, step 1 applies.
   - Native Azure credential lookup still differs from Hadoop ABFS in a few 
configurations ([#6605](https://github.com/apache/datafusion-comet/issues/6605))
     - Held at medium. It is the same class as #5542, which the 2026-08-31 pass 
labeled `priority:high`: the native Azure store authenticates with a credential 
other than the one Hadoop's ABFS driver would use. #5542 hit the common AKS 
workload identity setup, while each case here needs an uncommon configuration. 
If one of them turns out to be common, `priority:high` fits.
   - Native scans resolve the AWS SDK ProfileCredentialsProvider names 
differently from Hadoop 
([#6575](https://github.com/apache/datafusion-comet/issues/6575))
     - Held at medium on the same basis as #6605, for S3 profile providers. The 
reporter's proposed fix routes these names through 
`HadoopS3ACredentialProviderAdapter`, which adds a classpath requirement to 
configurations that work today.
   - S3 credential SPI: follow-ups from #6478 (hostless Iceberg metadata 
locations, docs, tests) 
([#6536](https://github.com/apache/datafusion-comet/issues/6536))
     - Held at medium on the same basis as #6605. Here the configured 
location-scoped provider is bypassed rather than mismatched, so a read can 
succeed with broader credentials than the provider would vend. That is closer 
to the security case in the guide's `priority:critical` row, but it needs an 
opted-in alias scheme with hostless locations, and it hasn't been reproduced.
   - Spark 3.4 SQL test SPARK-34637 fails since #6547 because AQE re-plans flip 
the DPP join's build side 
([#6645](https://github.com/apache/datafusion-comet/issues/6645))
     - Labeled medium rather than low, because the failing Spark test reflects 
a plan change on Spark 3.4 with default configs. It doesn't block the merge 
queue, since the Spark 3.4 SQL tests run only with `run-spark-3.4-tests`. If 
the reviewer reads it as a test failure, `priority:low` fits.
   - days transform is evaluated in the session timezone, while hours and 
Iceberg use UTC 
([#6333](https://github.com/apache/datafusion-comet/issues/6333))
     - Confirms the author's `priority:low`. Comet returns a value where Spark 
raises `PARTITION_TRANSFORM_EXPRESSION_NOT_IN_PARTITIONED_BY`, the shape that 
reviewers raised to critical in #5801 and #5936. It is held at low because 
Spark's error rejects calling a partition transform as an expression rather 
than reporting something about the data, and it only matters when something 
evaluates `days()` directly.
   
   ## Skipped (needs more info)
   
   #6576 has no reproduction on a supported path. The others are tracking or 
record issues rather than bug reports or feature requests. `requires-triage` 
was left in place on all five, so they reappear in the next pass until they are 
closed or classified.
   
   - Arrow struct writer should respect projected schema width 
([#6576](https://github.com/apache/datafusion-comet/issues/6576))
     - There is no reproduction, and the issue and its PR #6578 describe the 
problem only on the in-progress native Iceberg merge-on-read path (#6240). If a 
supported path can reach it, the walk past the Arrow children fails the task, 
which is step 2. If none can, it is hardening for #6240 and may be an 
enhancement.
   - Audit the PRs in 1.1.0 for regressions since 1.0.0 
([#6399](https://github.com/apache/datafusion-comet/issues/6399))
     - An audit task rather than a bug report or a feature request. Every phase 
is checked off and the results are in #6402, so it can be closed.
   - [EPIC] Bug fixes to consider backporting to branch-1.0 
([#6201](https://github.com/apache/datafusion-comet/issues/6201))
     - A release-management tracker, skipped on the same grounds as in the 
2026-09-28 pass.
   - Bug triage results: 2026-09-28 
([#6321](https://github.com/apache/datafusion-comet/issues/6321))
     - The summary of the 2026-09-28 pass. Since then, its issues have had no 
priority changes, only the `correctness` and `regression` additions described 
above, so it can be closed once reviewed.
   - Bug triage results: 2026-08-24 
([#5454](https://github.com/apache/datafusion-comet/issues/5454))
     - The summary of the 2026-08-24 pass. The 2026-09-28 pass found nothing 
outstanding in it, so it can be closed.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to