jackylee-ch opened a new pull request, #10128: URL: https://github.com/apache/paimon/pull/10128
### Purpose `FormatAvroReader.read_arrow_batch` reads the fastavro generator `batch_size` (default 1024) rows at a time and applies the pushed-down predicate in Python. When a whole batch matched nothing it returned `None` — but a `RecordBatchReader` returns `None` only at end of input, and `ConcatBatchReader` treats `None` as "reader exhausted" and advances to the next file. So a filtered read of an Avro file with a full non-matching batch silently dropped every remaining row: an append-only `file.format = avro` table of 3000 rows read with `user_id >= 2000` returns **0 rows** instead of 1000 (the leading 1024-row block fails the predicate → `None` → EOF). The reader now loops to the next batch on an empty filtered batch, returning `None` only when the generator is exhausted. A loop (not recursion) avoids a `RecursionError` on long filtered runs; the no-predicate path is unchanged. `format_row_reader` already handles this the same way. ### Tests `reader_append_only_test.test_avro_ao_reader_filter_keeps_rows_past_first_batch`: a 3000-row Avro table read with `user_id >= 2000` asserts the 1000 matching rows. Fails on master (`[]`); passes here. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
