JingsongLi opened a new issue, #1052:
URL: https://github.com/apache/paimon-rust/issues/1052

   ### Problem
   
   Primary-key BLOB reads resolve merged Arrow batches before applying a read 
limit. A selected row with a valid descriptor followed by an unselected row 
with a missing URI makes `with_limit(1)` fail opening the second payload. 
Payload predicates also resolve output-only BLOB columns on rejected rows and 
fetch BLOBs in Boolean branches that Java short-circuits. `IS NULL` / `IS NOT 
NULL` must inspect reference validity without accessing payloads.
   
   The raw primary-key streaming path used by row-kind reads exposes serialized 
descriptors instead of resolved payloads. Mixed materialized and streaming 
split groups need one shared quota.
   
   ### Expected behavior
   
   Follow Java's exact filtering before LIMIT, with lazy descriptor resolution: 
merge where required, evaluate predicates, stop after enough matching rows, and 
resolve only output payloads that will be returned. Keep raw change-event row 
kinds on streaming splits. Unselected or unused payload URIs must remain 
unopened, while an invalid selected payload must still fail.
   
   ### Reproduction
   
   1. Create a Parquet primary-key table with `payload BLOB` and 
`blob-descriptor-field=payload`.
   2. Write two keys: a valid descriptor first, a descriptor targeting a 
missing URI second.
   3. Read `id,payload` with LIMIT 1: current core attempts the missing URI.
   4. Project only `id` with `id = 1 OR payload = <bytes>` or a BLOB null 
check: current core unnecessarily accesses the payload.
   5. Consume a delta split with row kinds: current core returns descriptor 
bytes.
   
   ARRAY/MAP BLOB outputs additionally need to drop children hidden by a sliced 
or NULL parent without violating non-nullable element/value schemas.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to