jackylee-ch opened a new pull request, #10113: URL: https://github.com/apache/paimon/pull/10113
### Purpose `ShardBatchReader.read_arrow_batch` skips each out-of-range batch by recursively calling itself, so recursion depth grows one per skipped batch. A `with_shard`/`with_slice` read whose range sits deep in a data file then overflows the stack with `RecursionError`: the pyarrow reader yields 1024-row batches by default, so a slice starting ~1M rows in skips >1000 batches (Python's recursion limit). A slice covering only the head of a large file hits it too — the batches after `end_pos` are drained the same recursive way. The fix iterates over skipped batches with a `while` loop, matching `ConcatBatchReader` and `ApplyDeletionVectorReader`. Behavior is otherwise unchanged. ### Tests `shard_batch_reader_test`: 2000 single-row batches, slice `(1999, 2000)` — fails on master with `RecursionError`, passes here. Two more cases cover range filtering (`[2,5)`) and the straddle branches (`[2,9)` over 4-row batches). Written with Claude Code; verification is mine. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
