malinjawi opened a new pull request, #13160: URL: https://github.com/apache/gluten/pull/13160
### What changes were proposed in this pull request? Fix native partitioned Delta writes that exceed `maxRecordsPerFile` or roll files too early. Account rows per written chunk, fill the current file before opening another, and slice partition stripes that exceed its remaining capacity. Send each chunk to the statistics trackers for its destination file and release slices and remaining stripe batches on failure. Apply the fix to the Delta 3.3 and Delta 4.0 source paths. Add 13 shared backend regression cases and a Spark 3.5 `gluten-ut` case for null partition keys with file limits of 1 and 2. Declare its Delta runtime dependency in the test profile so standalone module runs exercise native writing. Include Delta suites in the x86 Spark 3.4–4.1 and enhanced Spark 4.0 CI selections. Related issue: #10215. Revives #12016; partition-column preservation remains separate in #12069. ### Why are the changes needed? The writer currently adds the original batch's row count to the last partition's file after writing every stripe. For example, three rows for partition A and one for B count as four rows against B's file, causing a later batch for B to roll the file early. A stripe larger than the file limit is also written without being split. ### Does this PR introduce any user-facing change? Native partitioned Delta writes honor `maxRecordsPerFile`, with file row counts and statistics updated for each written chunk. No public API or configuration change. ### How was this patch tested? The 13-case `DeltaNativeWriteLayoutSuite` covers partial files, partition boundaries across batches, oversized stripes, exact limits, empty batches, unlimited files, and file-open/write failures. Its integration matrix toggles statistics and optimized writes, asserts native execution, and checks physical Parquet rows, per-file statistics, duplicate-preserving readback, and history row counts using Spark as the read oracle. Completed on macOS arm64 with JDK 17: - Clean native build against pinned Velox `5cf370d2637d64feac516b49608daa9f95b0e120` passed. - Spark 4.1.1 / Delta 4.1.0 / Scala 2.13.17: the Delta package run passed 301 ScalaTest tests and 15 JUnit tests. The final shared layout suite also passed 13/13 after correcting its handling of URI-encoded filenames. - Regression check: the old writer implementation produced six expected failures; restoring the fix passed all 13 layout cases. - Spark 3.5.5 / Delta 3.3.2 / Scala 2.12.18: clean reactor build passed. All 277 existing Delta cases passed in the package run; after correcting the test path helper, the final layout suite passed 13/13 in a focused rerun. The separate `gluten-ut` null-partition test passed 1/1. Both package runs used the normal CI tag exclusions and reported 262 ignored cases. The Spark 3.5 package initially failed the eight layout integration cases because the new helper treated URI-encoded filenames as literal paths; the focused rerun validates that correction. No existing Delta suite failed in the fresh test warehouse. Scala/POM formatting, license-header checks, workflow selection and shell syntax checks, and `git diff --check` passed. ### Was this patch authored or co-authored using generative AI tooling? Generated-by: OpenAI Codex (GPT-6) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
