felipepessoto opened a new issue, #13138:
URL: https://github.com/apache/gluten/issues/13138

   ### Backend
   
   VL (Velox)
   
   ### Bug description
   
   **Expected behavior:** A valid Delta write partitioned by a Spark `DOUBLE` 
column should succeed. If Velox cannot write that partition type, Gluten should 
validate the partition columns and fall back to Spark's Delta writer before 
starting the write.
   
   **Actual behavior:** With native Delta writes enabled, the write reaches 
Velox's `PartitionIdGenerator` and fails with `Unsupported partition type: 
DOUBLE.` Spark wraps the exception as `[TASK_WRITE_FAILED]`.
   
   The configuration migration in #13131 exposed this limitation in 
`DeltaUpdateCatalogSuite`, which uses `DeltaHiveTest` rather than the 
previously patched `DeltaSQLCommandTest`. The migration now enables the native 
Delta writer for that suite too. These failures are separate from the 
task-resource initialization issue and the adaptive-prefetch SIGFPE.
   
   The two affected tests in 
`org.apache.spark.sql.delta.DeltaUpdateCatalogSuite` are:
   
   - `creating and replacing a table puts the schema and table properties in 
the metastore`
   - `partitioned table + add column`
   
   Both derive `part` as `id / 2`, producing a `DOUBLE`, and then create a 
Delta table partitioned by `part`. The initial write fails before the catalog 
assertions.
   
   Minimal equivalent write shape from the failing tests, using a Spark session 
configured for Delta and Gluten:
   
   ```scala
   import org.apache.spark.sql.functions.col
   
   val df = spark.range(10)
     .withColumn("part", col("id") / 2)
     .withColumn("id2", col("id"))
   
   df.writeTo("delta_double_partition_repro")
     .partitionedBy(col("part"))
     .using("delta")
     .create()
   ```
   
   Upstream test sources at Delta `v4.2.0`:
   
   - [Create/replace 
test](https://github.com/delta-io/delta/blob/v4.2.0/spark/src/test/scala/org/apache/spark/sql/delta/DeltaUpdateCatalogSuiteBase.scala#L180-L207)
   - [Add-column 
test](https://github.com/delta-io/delta/blob/v4.2.0/spark/src/test/scala/org/apache/spark/sql/delta/DeltaUpdateCatalogSuite.scala#L259-L279)
   
   **Suggested fix:** Check native partition-type support before entering the 
native Delta write path, and fall back for unsupported partition columns. Do 
not reject `DOUBLE` as an ordinary data column or disable native writing 
globally. Add regression coverage for unsupported partition fallback and 
continued native writing for supported partitions and ordinary `DOUBLE` data 
columns. Cover the relevant Delta versions through shared validation where 
practical.
   
   The runtime fix is deferred. #13131 will temporarily record exactly these 
two cases in its known-failures baseline; the tests will continue to run. 
Remove those entries when this issue is fixed.
   
   This issue was written with assistance from GitHub Copilot 1.0.87-0 (GPT-6 
Astra).
   
   ### Gluten version
   
   main branch / `1.8.0-SNAPSHOT`.
   
   Observed in #13131 at migration commit 
`9388f6307c6e618ee0fab25375e1d21b006a91c4`, CI merge commit 
`a28c4edc8b4ed35c7b02dfab3fba74a2b88c765b`. No upstream Velox PR override was 
applied in this run.
   
   ### Spark version
   
   Spark-4.1.x: Delta `v4.2.0` tests selected Spark `4.1.0`; the CI bundle's 
build metadata reports Spark `4.1.1` and Scala `2.13.17`.
   
   ### Spark configurations
   
   Relevant effective defaults:
   
   ```properties
   spark.plugins=org.apache.gluten.GlutenPlugin
   spark.shuffle.manager=org.apache.spark.shuffle.sort.ColumnarShuffleManager
   spark.memory.offHeap.enabled=true
   spark.memory.offHeap.size=2g
   spark.default.parallelism=1
   spark.sql.shuffle.partitions=5
   spark.sql.ansi.enabled=false
   spark.gluten.sql.ansiFallback.enabled=false
   spark.gluten.sql.columnar.backend.velox.delta.enableNativeWrite=true
   spark.databricks.delta.snapshotPartitions=2
   ```
   
   The Delta fixture additionally configures 
`io.delta.sql.DeltaSparkSessionExtension` and 
`org.apache.spark.sql.delta.catalog.DeltaCatalog`.
   
   ### System information
   
   GitHub Actions Delta Spark UT run `36165436135`, shard 3. JDK `17.0.20`; 
Linux runner, `apache/gluten:centos-9-jdk17` test container. Native build used 
CentOS 8 and GCC `13.3.1`. IBM Velox pin: `dft-2026_09_25` 
(`340ab366bdf5a7c25c913c84730227675aad72c5`).
   
   `dev/info.sh` output from the CI runner is not available; the details above 
come from the workflow and build metadata.
   
   ### Relevant logs
   
   [Failing CI job and test 
step](https://github.com/apache/gluten/actions/runs/36165436135/job/108178306323#step:10:1),
 with failures at `2026-09-25T18:52:38Z` and `2026-09-25T18:53:04Z`. Both XML 
reports and the job log contain the same underlying error:
   
   ```text
   org.apache.spark.SparkException: [TASK_WRITE_FAILED] Task failed while 
writing rows ...
   Cause: org.apache.gluten.exception.GlutenException: Exception: VeloxUserError
   Error Source: USER
   Error Code: INVALID_ARGUMENT
   Reason: Unsupported partition type: DOUBLE.
   Retriable: False
   Expression: hashers_.back()->typeSupportsValueIds()
   Function: PartitionIdGenerator
   File: 
/work/ep/build-velox/build/velox_ep/velox/connectors/hive/PartitionIdGenerator.cpp
   Line: 34
   ```
   
   The shard gate reported exactly these two new regressions; neither was in 
the baseline when this run executed.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to