felipepessoto opened a new pull request, #13136:
URL: https://github.com/apache/gluten/pull/13136
<!--
Thank you for submitting a pull request! Here are some tips:
1. For first-time contributors, please read our contributing guide:
https://github.com/apache/gluten/blob/main/CONTRIBUTING.md
2. If necessary, create a GitHub issue for discussion beforehand to avoid
duplicate work.
3. If the PR is specific to a single backend, include [VL] or [CH] in the PR
title to indicate the
Velox or ClickHouse backend, respectively.
4. If the PR is not ready for review, please mark it as a draft.
-->
## What changes are proposed in this pull request?
<!--
Provide a clear and concise description of the changes introduced in this PR.
Ensure the PR description aligns with the code changes, especially after
updates.
If applicable, include "Fixes #<GitHub_Issue_ID>" to automatically close the
corresponding issue
when the PR is merged.
-->
Separate the physical reader schema from the logical scan-output schema in
the Velox plan converter. A metadata-only query such as `SELECT DISTINCT
_metadata.file_name` over a Parquet table containing a user `file_name BIGINT`
must not request that physical field as `VARCHAR` before the connector installs
file metadata constants.
- Without an explicit split table schema, build
`HiveTableHandle::dataColumns()` from **regular scan slots by role and index**.
Partition, synthesized, and generated row-index outputs remain available to the
scan but do not become requested physical fields. Metadata-only projections
retain a non-null empty physical schema.
- With an explicit table schema, retain genuine physical names, types, and
ordinals, applying only the existing partition exclusion. Do not remove a real
field just because a synthetic output has the same name.
- Preserve logical output slots, assignment roles, filter binding, split
metadata, and the existing Spark fallback for simultaneous real and same-named
metadata outputs. No production Scala/Substrait protocol changes or permissive
cast changes.
- Add programmatic native converter/Parquet regressions and a bounded Spark
4.1 regression matrix for all six file constants, incompatible same-name
physical types, aliases/whole structs, filters, partitions, case normalization,
and explicit positional schemas. The native fixture now forwards parsed split
metadata and writes actual Parquet fixtures.
This is independent of the generated Delta DV metadata work in #13123; it
does not change that feature's protocol or implement the separate Parquet
date-widening fix.
## How was this patch tested?
<!--
Describe how the changes were tested, if applicable.
Include new tests to validate the functionality, if necessary.
For UI-related changes, attach screenshots to demonstrate the updates.
-->
**Draft: compilation and runtime validation are pending.**
Completed:
- Repository Scala formatter (`dev/format-scala-code.sh`) through the Maven
wrapper in the Gluten JDK 17 container.
- clang-format **15.0.7** on the changed C++ files, including a final
dry-run check.
- Repository CI license-header checker on all changed files, and `git diff
--check`. The documented `dev/check.py header` wrapper refers to a missing
`dev/license-header.py`, so the existing
`.github/workflows/util/license-header.py` checker was used directly.
Added but **not yet compiled/executed**:
- Six native converter/Parquet tests: regular-slot selection, empty/untagged
schemas, explicit physical schema preservation, per-file metadata and row
counts (including an empty file), metadata/regular filters, genuine
regular-type rejection, and positional mapping.
- Twelve Spark 4.1 metadata regression cases, checking actual metadata
values and native scans for supported projections, while explicitly retaining
same-name combined-output fallback.
An exact-source native build was attempted in an isolated
`apache/gluten:centos-9-jdk17` container. CMake stopped before compilation
because the IBM-pinned Velox dependency source/build tree is not prepared. No
sibling PR native binaries were substituted. The native and JVM regression
suites, baseline/fixed red-green reproduction, and the original twelve Delta
aliasing failures have not been run locally. Native PR CI and focused runtime
regressions remain acceptance gates before marking this ready.
## Was this patch authored or co-authored using generative AI tooling?
<!--
If generative AI tooling has been used in the process of authoring this
patch, please include the
phrase: 'Generated-by: ' followed by the name of the tool and its version.
If no, write 'No'.
Please refer to the [ASF Generative Tooling
Guidance](https://www.apache.org/legal/generative-tooling.html) for details.
-->
Generated-by: GitHub Copilot 1.0.87-0 (GPT-6 Astra)
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]