jackylee-ch opened a new pull request, #10113:
URL: https://github.com/apache/paimon/pull/10113

   ### Purpose
   
   `ShardBatchReader.read_arrow_batch` skips each out-of-range batch by 
recursively calling itself, so recursion depth grows one per skipped batch. A 
`with_shard`/`with_slice` read whose range sits deep in a data file then 
overflows the stack with `RecursionError`: the pyarrow reader yields 1024-row 
batches by default, so a slice starting ~1M rows in skips >1000 batches 
(Python's recursion limit). A slice covering only the head of a large file hits 
it too — the batches after `end_pos` are drained the same recursive way.
   
   The fix iterates over skipped batches with a `while` loop, matching 
`ConcatBatchReader` and `ApplyDeletionVectorReader`. Behavior is otherwise 
unchanged.
   
   ### Tests
   
   `shard_batch_reader_test`: 2000 single-row batches, slice `(1999, 2000)` — 
fails on master with `RecursionError`, passes here. Two more cases cover range 
filtering (`[2,5)`) and the straddle branches (`[2,9)` over 4-row batches).
   
   Written with Claude Code; verification is mine.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to