goutamadwant opened a new issue, #12317: URL: https://github.com/apache/seatunnel/issues/12317
### Search before asking - [x] I searched existing issues and pull requests and found no matching fix for exact numeric predicates in `ZetaSQLFilter`. ### What happened SQLTransform converts numeric operands to `double` before evaluating comparisons and numeric `IN` membership. Distinct BIGINT or high-precision DECIMAL values can consequently compare equal. A job can complete successfully while filtering or classifying the wrong records. For example, BIGINT `9007199254740993` incorrectly matches `WHERE a = 9007199254740992`. The two values become the same double. This also affects ordering, `IN`/`NOT IN` and the comparison path used by CASE expressions. ### SeaTunnel Version Reproduced against unreleased `3.0.0-SNAPSHOT`, commit `0d9f9e2303da0b40b20e55fd341d564757e86dc0`. Current `dev` at `f4a9665e8457239dd30c04a8738b506fa3fce180` leaves the affected implementation unchanged. Released-version coverage has not been established. ### SeaTunnel Config and reproduction The reproduction invokes `SQLTransform.transformRow` with a typed SeaTunnel row, rather than submitting a standalone engine job: ```text Input schema: a BIGINT, b BIGINT Input row: a = 9007199254740993, b = 9007199254740992 ``` ```sql SELECT a, b FROM dual WHERE a = 9007199254740992 ``` Expected: no output row. Actual: the input row is retained. Further observed cases: - `a > b` for the same input should retain the row but filters it out. - `a IN (0, 9007199254740992)` should not match but does. - DECIMAL `123456789012345678.99` and `123456789012345678.98` lose their ordering after conversion to double. - Searched and simple CASE expressions can choose the wrong result for adjacent large integers. ### Running Command The accompanying regression class is [SQLNumericComparisonTest.java](https://github.com/goutamadwant/seatunnel/blob/ca5abddfbab30a38853127f754d61dbac133a292/seatunnel-transforms-v2/src/test/java/org/apache/seatunnel/transform/sql/SQLNumericComparisonTest.java). To verify the before behavior, apply only that test file to the reproduction revision, leaving production code unchanged, and run: ```sh ./mvnw -B -pl seatunnel-transforms-v2 -am package \ -Dtest=SQLNumericComparisonTest \ -Dsurefire.failIfNoSpecifiedTests=false -Dskip.spotless=true -Dskip.ui=true ``` Repeat with Java 8 and Java 11 selected through `JAVA_HOME`. Run the same command with the production fix applied to verify the after behavior. ### Error Exception The production defect is a wrong result, not an exception. Regression assertions fail because rows are retained or discarded incorrectly. Against unchanged production code, the initial Java 8 run had 99 failing cases out of 317. Java 11 had 100 failures out of 318; the extra test covers CASE expressions and was added after the first Java 8 baseline. With the fix, all 318 cases pass on both JDKs. ### Before and after - Before: exact numeric operands are rounded through double for equality, inequality, ordering and numeric membership checks. - After: Byte, Short, Integer and Long operands use `Long.compare`; comparisons involving BigDecimal use the existing `NumericFunction.toBigDecimal` conversion and `BigDecimal.compareTo`. - FLOAT/DOUBLE coercion, NaN, infinity, signed zero, non-numeric comparisons and existing NULL behavior remain unchanged. Arithmetic and separate NULL-semantics work are outside this fix. ### Advantages - Prevents silently selecting or classifying incorrect records containing large identifiers or precise decimal values. - Fixes the shared predicate path, including CASE and membership evaluation, instead of adding connector-specific handling. - Integral comparisons introduce no allocation; the fix adds no dependency, option or public API. ### Breaking changes and migration This is an observable correctness change: affected predicates may now select different rows. Review such predicates and reconcile previously written data where necessary. There are no configuration, schema, state-format or wire-format changes. English and Chinese documentation and upgrade notes accompany the fix. ### Validation After rebasing onto `f4a9665e8457`, all 486 affected SQL tests and formatting checks passed on Java 8 and Java 11. Production and test changes are unchanged from the broader validation below. The new regression cases and full transform suite passed on Java 8 and Java 11 before rebasing the patch. The SQL-branch transform suite contains 1,475 cases. Java 11 root `./mvnw -q -DskipTests verify` also passed, including distribution packaging, on the original base. An initial broader Java 8 run encountered an unrelated Python executable-path/allowlist mismatch. Putting the real interpreter directory first on PATH resolved it without changing the allowlist. The retry reran all transform tests; only an already-passing dependency date/time performance method was omitted. Java 11 had no test exclusions. No Docker-backed E2E or deployed workload result is claimed. ### Zeta or Flink or Spark Version Direct SQLTransform tests using the default ZETA SQL engine. The INTERNAL setting resolves to the same implementation. No external Flink/Spark runtime was used for this reproduction. ### Java or Scala Version Java 8 and Java 11. No Scala changes. ### Are you willing to submit PR? - [x] Yes, I am willing to submit a PR. ### Code of Conduct - [x] I agree to follow the project's Code of Conduct. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
