felipepessoto opened a new issue, #13138:
URL: https://github.com/apache/gluten/issues/13138
### Backend
VL (Velox)
### Bug description
**Expected behavior:** A valid Delta write partitioned by a Spark `DOUBLE`
column should succeed. If Velox cannot write that partition type, Gluten should
validate the partition columns and fall back to Spark's Delta writer before
starting the write.
**Actual behavior:** With native Delta writes enabled, the write reaches
Velox's `PartitionIdGenerator` and fails with `Unsupported partition type:
DOUBLE.` Spark wraps the exception as `[TASK_WRITE_FAILED]`.
The configuration migration in #13131 exposed this limitation in
`DeltaUpdateCatalogSuite`, which uses `DeltaHiveTest` rather than the
previously patched `DeltaSQLCommandTest`. The migration now enables the native
Delta writer for that suite too. These failures are separate from the
task-resource initialization issue and the adaptive-prefetch SIGFPE.
The two affected tests in
`org.apache.spark.sql.delta.DeltaUpdateCatalogSuite` are:
- `creating and replacing a table puts the schema and table properties in
the metastore`
- `partitioned table + add column`
Both derive `part` as `id / 2`, producing a `DOUBLE`, and then create a
Delta table partitioned by `part`. The initial write fails before the catalog
assertions.
Minimal equivalent write shape from the failing tests, using a Spark session
configured for Delta and Gluten:
```scala
import org.apache.spark.sql.functions.col
val df = spark.range(10)
.withColumn("part", col("id") / 2)
.withColumn("id2", col("id"))
df.writeTo("delta_double_partition_repro")
.partitionedBy(col("part"))
.using("delta")
.create()
```
Upstream test sources at Delta `v4.2.0`:
- [Create/replace
test](https://github.com/delta-io/delta/blob/v4.2.0/spark/src/test/scala/org/apache/spark/sql/delta/DeltaUpdateCatalogSuiteBase.scala#L180-L207)
- [Add-column
test](https://github.com/delta-io/delta/blob/v4.2.0/spark/src/test/scala/org/apache/spark/sql/delta/DeltaUpdateCatalogSuite.scala#L259-L279)
**Suggested fix:** Check native partition-type support before entering the
native Delta write path, and fall back for unsupported partition columns. Do
not reject `DOUBLE` as an ordinary data column or disable native writing
globally. Add regression coverage for unsupported partition fallback and
continued native writing for supported partitions and ordinary `DOUBLE` data
columns. Cover the relevant Delta versions through shared validation where
practical.
The runtime fix is deferred. #13131 will temporarily record exactly these
two cases in its known-failures baseline; the tests will continue to run.
Remove those entries when this issue is fixed.
This issue was written with assistance from GitHub Copilot 1.0.87-0 (GPT-6
Astra).
### Gluten version
main branch / `1.8.0-SNAPSHOT`.
Observed in #13131 at migration commit
`9388f6307c6e618ee0fab25375e1d21b006a91c4`, CI merge commit
`a28c4edc8b4ed35c7b02dfab3fba74a2b88c765b`. No upstream Velox PR override was
applied in this run.
### Spark version
Spark-4.1.x: Delta `v4.2.0` tests selected Spark `4.1.0`; the CI bundle's
build metadata reports Spark `4.1.1` and Scala `2.13.17`.
### Spark configurations
Relevant effective defaults:
```properties
spark.plugins=org.apache.gluten.GlutenPlugin
spark.shuffle.manager=org.apache.spark.shuffle.sort.ColumnarShuffleManager
spark.memory.offHeap.enabled=true
spark.memory.offHeap.size=2g
spark.default.parallelism=1
spark.sql.shuffle.partitions=5
spark.sql.ansi.enabled=false
spark.gluten.sql.ansiFallback.enabled=false
spark.gluten.sql.columnar.backend.velox.delta.enableNativeWrite=true
spark.databricks.delta.snapshotPartitions=2
```
The Delta fixture additionally configures
`io.delta.sql.DeltaSparkSessionExtension` and
`org.apache.spark.sql.delta.catalog.DeltaCatalog`.
### System information
GitHub Actions Delta Spark UT run `36165436135`, shard 3. JDK `17.0.20`;
Linux runner, `apache/gluten:centos-9-jdk17` test container. Native build used
CentOS 8 and GCC `13.3.1`. IBM Velox pin: `dft-2026_09_25`
(`340ab366bdf5a7c25c913c84730227675aad72c5`).
`dev/info.sh` output from the CI runner is not available; the details above
come from the workflow and build metadata.
### Relevant logs
[Failing CI job and test
step](https://github.com/apache/gluten/actions/runs/36165436135/job/108178306323#step:10:1),
with failures at `2026-09-25T18:52:38Z` and `2026-09-25T18:53:04Z`. Both XML
reports and the job log contain the same underlying error:
```text
org.apache.spark.SparkException: [TASK_WRITE_FAILED] Task failed while
writing rows ...
Cause: org.apache.gluten.exception.GlutenException: Exception: VeloxUserError
Error Source: USER
Error Code: INVALID_ARGUMENT
Reason: Unsupported partition type: DOUBLE.
Retriable: False
Expression: hashers_.back()->typeSupportsValueIds()
Function: PartitionIdGenerator
File:
/work/ep/build-velox/build/velox_ep/velox/connectors/hive/PartitionIdGenerator.cpp
Line: 34
```
The shard gate reported exactly these two new regressions; neither was in
the baseline when this run executed.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]