felipepessoto opened a new pull request, #13123:
URL: https://github.com/apache/gluten/pull/13123
<!--
Thank you for submitting a pull request! Here are some tips:
1. For first-time contributors, please read our contributing guide:
https://github.com/apache/gluten/blob/main/CONTRIBUTING.md
2. If necessary, create a GitHub issue for discussion beforehand to avoid
duplicate work.
3. If the PR is specific to a single backend, include [VL] or [CH] in the PR
title to indicate the
Velox or ClickHouse backend, respectively.
4. If the PR is not ready for review, please mark it as a draft.
-->
## What changes are proposed in this pull request?
<!--
Provide a clear and concise description of the changes introduced in this PR.
Ensure the PR description aligns with the code changes, especially after
updates.
If applicable, include "Fixes #<GitHub_Issue_ID>" to automatically close the
corresponding issue
when the PR is merged.
-->
Partially addresses #12743. This is a **standalone draft targeting main**,
exploring a native alternative/follow-up to the JVM correctness fallback in
#13042. It does not resolve the full issue, is not stacked on #13042, and does
not modify that PR.
**Native C++ compilation and the new end-to-end DV tests remain unverified
locally.** This draft needs Linux native/Delta CI before it can be considered
ready; it makes no performance claim.
For `spark.databricks.delta.deletionVectors.useMetadataRowIndex=false`:
- Introduce distinct Substrait column kinds for generated Delta row indexes
and deleted-row flags. Recognize exact Delta names without requiring
generated-metadata markers, checking both output and required schema.
- Generate absolute file positions through Velox's row-index reader. A
private BIGINT channel supplies positions for bitmap lookup; the public
deleted-status output remains TINYINT. No extra native plan operator is
inserted, preserving the existing metrics mapping.
- **Mark, do not prematurely remove, rows** when a deleted-status output is
requested. `IF_CONTAINED` marks bitmap members; `IF_NOT_CONTAINED` marks
non-members. An absent DV yields `0`, while a present empty inverse bitmap
yields `1` for every row.
- Preserve consumer predicates and projections, and keep generated-field
predicates out of physical Parquet pushdown. Enforce generated-key dynamic
filters after materialization, including Velox's join-replacement optimization,
empty batches and preloaded-source handoff. Restrict legacy predicate/column
stripping to Delta's injected unary shape so an ordinary scan in another join
branch cannot strip generated outputs.
- Remove Spark 4.0/4.1's generic classification of the deleted-status name
as a row-index column. Delta-specific classification stays in the Delta
transformer.
The existing metadata-row-index=true native masking path, executor-deferred
checksum-validating DV payload reads, memoization, metrics, and authoritative
PreparedDeltaFileIndex AddFiles are retained. The specialized Tahoe metadata
route still supplies inverse filter semantics. This builds on the handoff
merged in #12836; the closed, unmerged native range-read proposal #12867 is not
included.
JVM fallback remains for Spark 3.4 DV scans, CDF scans touching DVs, the DML
configuration escape hatch, DV-bearing scans without a deleted-status output,
and unsupported generated-field/schema shapes (including required-only fields
absent from output, ambiguous/malformed fields, bucketed scans and
mapped/partition-name collisions). DV-free row-index-only scans are supported.
Adds shared Delta 3.3/4.0 integration cases with native-plan assertions,
native connector/conversion regressions, a gluten-ut contract/shim suite, and
related Delta documentation. Coverage includes the nine DV-free field/reader
combinations, full value/index association, combined fields, raw marked rows,
inverse/mixed-file semantics, row-group pruning, split offsets, small batches,
name mapping, mixed-scan projections, repeated DML and fallback. Native
unique-key inner/semi-join cases cover generated keys and split preloading;
inner cases assert `replacedWithDynamicFilterRows > 0` so correctness cannot be
supplied only by a remaining hash lookup.
**The known-failures baseline is deliberately unchanged until CI provides
evidence.** Rebuild and deploy matching native libraries and JVM artifacts for
the new column-kind protocol.
## How was this patch tested?
<!--
Describe how the changes were tested, if applicable.
Include new tests to validate the functionality, if necessary.
For UI-related changes, attach screenshots to demonstrate the updates.
-->
Local environment: Windows, Git Bash, Temurin JDK 17; WSL Ubuntu 24.04 was
also used for available checks.
**Passed:**
- Maven-wrapper Spark 3.5 / Scala 2.12 / Delta 3.3 production and
test-source compilation. The final compilation was `bash .\build\mvn -B -q -pl
gluten-ut/test -am test-compile
-Pdelta,spark-ut,backends-velox,spark-3.5,scala-2.12,fast-build -DskipTests`. A
matching reactor `install -DskipTests` also completed during validation.
- `org.apache.gluten.sql.shims.DeltaMetadataColumnSuite`: **2 tests
passed**, using ScalaTest with a Windows-safe classpath JAR generated from this
worktree and Maven-resolved dependencies. The repository build-info script
supplied the resource whose Maven generation is Unix-only.
- Constructor-level registration check: **all 27 Delta 3.3 Handoff tests
registered**, including the mapping, DML, fallback and mixed-scan cases. This
is a registration check, not a runtime DV test pass.
- `dev/format-scala-code.sh` and scoped Spotless check across all changed
Scala files/profiles.
- `dev/format-cpp-code.sh` in an isolated Linux validation copy, plus
explicit clang-format **15.0.7** formatting and `--dry-run --Werror` checks for
every touched C++ file, including `.cpp` files not selected by the script.
- Actual CI license-header fixer/checker,
`.github/workflows/util/license-header.py`, on every changed file; `git diff
--check`. The documented `dev/check.py header main --fix` entry point was
attempted but refers to a missing `dev/license-header.py`, so the existing CI
implementation was used instead.
**Blocked / not claimed as passing:**
- Native build attempt, after normalizing shell line endings only in the
isolated validation copy: `dev/builddeps-veloxbe.sh build_gluten_cpp
--enable_s3=OFF --enable_gcs=OFF --enable_hdfs=OFF --enable_abfs=OFF
--build_tests=ON` reached CMake and stopped at `cmake: command not found`. The
WSL native compiler/Ninja/dependency toolchain is not provisioned. **No C++
compile or native test pass is claimed.**
- Delta Handoff suite runtime startup aborts at missing
`linux/amd64/gluten.dll`; **zero integration tests executed**.
- Spark 4.0 / Scala 2.13 / JDK17 wrapper attempt with
`-Dmaven.compiler.release=17` reaches the backend module and fails on six
existing `src-delta40` symlink placeholders parsed as Scala source. Full Delta
4.0 compilation/runtime is not claimed. A narrowed standalone test-source
attempt also lacked the vendored Delta test utilities normally compiled by that
module.
- An initial reactor `test` attempt additionally encountered the existing
Windows symlink-loop assertion in `JniLibLoaderTest`; the focused ScalaTest
execution above avoids running unrelated reactor tests.
Linux follow-up targets are `velox_delta_read_test` and
`velox_plan_conversion_test`, plus `DeltaDeletionVectorHandoffSuite` and the
existing Delta file-format/DV/DML/CDF regressions under both Spark 3.5/Delta
3.3 and Spark 4.0/Delta 4.0. The native join-replacement and preloading
assertions are included but have not been executed locally.
## Was this patch authored or co-authored using generative AI tooling?
<!--
If generative AI tooling has been used in the process of authoring this
patch, please include the
phrase: 'Generated-by: ' followed by the name of the tool and its version.
If no, write 'No'.
Please refer to the [ASF Generative Tooling
Guidance](https://www.apache.org/legal/generative-tooling.html) for details.
-->
Generated-by: GitHub Copilot CLI 1.0.87-0 (GPT-6 Astra, gpt-6-astra)
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]