jackylee-ch opened a new pull request, #10175:
URL: https://github.com/apache/paimon/pull/10175

   ## Purpose
   
   Java's `PredicateBuilder` has `arrayContains` / `arraysOverlap` /
   `arrayContainsAll` (Spark pushes them down), but PyPaimon's
   `PredicateBuilder` had no array predicates. AI datasets commonly filter an
   `ARRAY<STRING>` label column ("samples tagged `cat`"); without these,
   PyPaimon could only read the whole table and filter client-side.
   
   ## Change
   
   - Add `array_contains` / `arrays_overlap` / `array_contains_all` to
     `PredicateBuilder` and the matching `Predicate` testers, mirroring Java
     `ArrayContains` / `ArraysOverlap` / `ArrayContainsAll` row semantics
     (null array or null element → no match; `arrays_overlap` needs one
     shared non-null element; `array_contains_all` is vacuously true for an
     empty literal list on a non-null array).
   - Arrays have no dataset-expression form, so the three methods are marked
     arrow-unsafe and run on Paimon's exact row-level filter path (like
     `startsWith`/`contains`/`like`) instead of a no-op truthy Arrow
     expression. Stats testers return `True` (array element bounds aren't
     tracked, so no file is pruned).
   
   ## Tests
   
   - Row-level tester parity with Java (match/no-match, null array, null
     element, empty literals) and that the three methods are not
     arrow-pushable.
   - End-to-end: filter an `ARRAY<STRING>` column and assert the exact ids
     returned (a dropped/ignored filter would return every row).
   
   Written with Claude Code; verification is mine.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to