rdtr opened a new pull request, #13184:
URL: https://github.com/apache/gluten/pull/13184

   ## What changes are proposed in this pull request?
   
   The ANSI CI (`velox_backend_ansi.yml`) runs the backends-velox tests with 
ANSI mode enabled and ANSI fallback disabled. 17 tests fail there because they 
use inputs that ANSI mode rejects, such as malformed strings cast to numbers. 
Vanilla Spark itself throws for these queries, so the failures come from the 
tests, not from Gluten.
   
   This PR adds `runQueryAndCompareOrBothFail` to `GlutenQueryComparisonTest`. 
It behaves like `runQueryAndCompare`. In ANSI mode, it also accepts a query 
that fails on vanilla Spark, as long as Gluten fails too. The affected tests 
use it, so in ANSI mode they now check that Gluten rejects the same inputs as 
Spark.
   
   Two tests whose subject is unrelated to ANSI mode run with ANSI mode 
disabled instead:
   
   - `FallbackSuite` ("fallback when nested loop join has unsupported 
expression") checks the fallback reason.
   - `MiscOperatorSuite` ("Support multi-children count with row construct"): 
Spark before 4.3 (SPARK-58213) throws DIVIDE_BY_ZERO for `corr` on a column 
with zero variance in ANSI mode. Velox returns NULL, as Spark 4.3 does.
   
   With this PR, 4 tests still fail in the ANSI CI. They are real ANSI gaps, 
where vanilla Spark fails but Gluten returns a result, and each has a fix in 
review:
   
   - `cast`: Velox ignores ANSI mode for casts between complex types (arrays, 
maps and structs). Fixed by 
https://github.com/facebookincubator/velox/pull/19311. With that change applied 
locally, the `cast` test passes.
   - `make_date`: https://github.com/facebookincubator/velox/pull/18764
   - `decimal arithmetic` and `decimal arithmetic respects allowPrecisionLoss 
captured at view analysis time`: https://github.com/apache/gluten/pull/13170 
for add, subtract and multiply, and 
https://github.com/facebookincubator/velox/pull/19290 for divide. Gluten will 
map divide to `checked_divide` once that lands.
   
   ## How was this patch tested?
   
   Ran the 6 affected suites on Spark 4.1:
   
   - ANSI mode disabled (`SPARK_ANSI_SQL_MODE=false`, as in the regular CI): 
290 passed.
   - ANSI mode enabled with `spark.gluten.sql.ansiFallback.enabled=false`, as 
in the ANSI CI: 17 failures before this PR, 4 after, and 3 after also applying 
the Velox fix for complex type casts. The remaining failures are listed above.
   
   ## Was this patch authored or co-authored using generative AI tooling?
   
   Generated-by: Claude Code (Claude Opus 5.5)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to