jackylee-ch opened a new pull request, #10175:
URL: https://github.com/apache/paimon/pull/10175
## Purpose
Java's `PredicateBuilder` has `arrayContains` / `arraysOverlap` /
`arrayContainsAll` (Spark pushes them down), but PyPaimon's
`PredicateBuilder` had no array predicates. AI datasets commonly filter an
`ARRAY<STRING>` label column ("samples tagged `cat`"); without these,
PyPaimon could only read the whole table and filter client-side.
## Change
- Add `array_contains` / `arrays_overlap` / `array_contains_all` to
`PredicateBuilder` and the matching `Predicate` testers, mirroring Java
`ArrayContains` / `ArraysOverlap` / `ArrayContainsAll` row semantics
(null array or null element → no match; `arrays_overlap` needs one
shared non-null element; `array_contains_all` is vacuously true for an
empty literal list on a non-null array).
- Arrays have no dataset-expression form, so the three methods are marked
arrow-unsafe and run on Paimon's exact row-level filter path (like
`startsWith`/`contains`/`like`) instead of a no-op truthy Arrow
expression. Stats testers return `True` (array element bounds aren't
tracked, so no file is pruned).
## Tests
- Row-level tester parity with Java (match/no-match, null array, null
element, empty literals) and that the three methods are not
arrow-pushable.
- End-to-end: filter an `ARRAY<STRING>` column and assert the exact ids
returned (a dropped/ignored filter would return every row).
Written with Claude Code; verification is mine.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]