JingsongLi commented on PR #10175:
URL: https://github.com/apache/paimon/pull/10175#issuecomment-5831535460
Reviewed 1e800e9. The ARRAY predicates have real end-to-end functional
value: the new read-path tests return the exact matching row IDs, and the
row-level implementations agree with Java for null and empty-literal cases. One
correctness gap needs attention:
**[P2] Reject non-array fields instead of silently applying Python
containment.** The new builder methods only check the field name, and
`ArrayContains.test_by_value` / `ArraysOverlap.test_by_value` then use `literal
in val`. For a STRING column `s = "cat"`, `array_contains("s", "a")` and
`arrays_overlap("s", ["a"])` both return true (I reproduced this locally),
whereas Java `ArrayContains.elementType` rejects a non-ARRAY field. A wrong
field name or schema change can therefore silently select incorrect rows.
Retain field types in `PredicateBuilder` and validate these three methods
against `ArrayType`, with a negative test.
Validation: `array_predicate_test` 7 passed; adjacent projection/predicate
tests 15 passed; `git diff --check` passed; CI run 36100713243 is green.
Because the new stats testers always return true and these methods are
Arrow-unsafe, this is exact row-level filtering but does not prune files; the
PR description should not imply reduced table I/O. Please remove the `#
placeholder-e2e` comment too.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]