JingsongLi commented on PR #10175:
URL: https://github.com/apache/paimon/pull/10175#issuecomment-5831535460

   Reviewed 1e800e9. The ARRAY predicates have real end-to-end functional 
value: the new read-path tests return the exact matching row IDs, and the 
row-level implementations agree with Java for null and empty-literal cases. One 
correctness gap needs attention:
   
   **[P2] Reject non-array fields instead of silently applying Python 
containment.** The new builder methods only check the field name, and 
`ArrayContains.test_by_value` / `ArraysOverlap.test_by_value` then use `literal 
in val`. For a STRING column `s = "cat"`, `array_contains("s", "a")` and 
`arrays_overlap("s", ["a"])` both return true (I reproduced this locally), 
whereas Java `ArrayContains.elementType` rejects a non-ARRAY field. A wrong 
field name or schema change can therefore silently select incorrect rows. 
Retain field types in `PredicateBuilder` and validate these three methods 
against `ArrayType`, with a negative test.
   
   Validation: `array_predicate_test` 7 passed; adjacent projection/predicate 
tests 15 passed; `git diff --check` passed; CI run 36100713243 is green. 
Because the new stats testers always return true and these methods are 
Arrow-unsafe, this is exact row-level filtering but does not prune files; the 
PR description should not imply reduced table I/O. Please remove the `# 
placeholder-e2e` comment too.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to