jackylee-ch opened a new pull request, #10128:
URL: https://github.com/apache/paimon/pull/10128

   ### Purpose
   
   `FormatAvroReader.read_arrow_batch` reads the fastavro generator 
`batch_size` (default 1024) rows at a time and applies the pushed-down 
predicate in Python. When a whole batch matched nothing it returned `None` — 
but a `RecordBatchReader` returns `None` only at end of input, and 
`ConcatBatchReader` treats `None` as "reader exhausted" and advances to the 
next file.
   
   So a filtered read of an Avro file with a full non-matching batch silently 
dropped every remaining row: an append-only `file.format = avro` table of 3000 
rows read with `user_id >= 2000` returns **0 rows** instead of 1000 (the 
leading 1024-row block fails the predicate → `None` → EOF).
   
   The reader now loops to the next batch on an empty filtered batch, returning 
`None` only when the generator is exhausted. A loop (not recursion) avoids a 
`RecursionError` on long filtered runs; the no-predicate path is unchanged. 
`format_row_reader` already handles this the same way.
   
   ### Tests
   
   
`reader_append_only_test.test_avro_ao_reader_filter_keeps_rows_past_first_batch`:
 a 3000-row Avro table read with `user_id >= 2000` asserts the 1000 matching 
rows. Fails on master (`[]`); passes here.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to