leaves12138 commented on code in PR #10003:
URL: https://github.com/apache/paimon/pull/10003#discussion_r4056522916


##########
paimon-python/pypaimon/read/table_read.py:
##########
@@ -219,7 +219,7 @@ def to_arrow_batch_reader(
         table reads do not guarantee row order. Python fallback reads remain
         serial.
         """
-        effective = self._resolve_parallelism(parallelism, len(splits))
+        effective = self._effective_parallelism(parallelism, len(splits))

Review Comment:
   [P2] Preserve explicit iterator cleanup when LIMIT reduces concurrency to one
   
   With two or more native splits, `parallelism=2`, and an unfiltered 
`limit=1`, the new effective parallelism becomes 1. Consequently, the check 
below skips `_ClosableArrowBatchReader` and returns the raw 
`RecordBatchReader.from_batches(...)` reader. On PyArrow 19.0.1, calling 
`read_next_batch()` once and then `close()` does not close the suspended 
`_convert_native_batches` generator, so its `finally` never calls the 
underlying native reader's `close()`. This retains the native stream until 
later exhaustion/garbage collection instead of releasing it at explicit close.
   
   I reproduced this through `to_arrow_batch_reader` with close-tracking native 
readers: the base revision closes both readers after that sequence, whereas 
this head leaves its sole native reader open. `read_all()` hides the regression 
by exhausting the iterator.
   
   Could we preserve the explicit iterator-close chain independently of the 
effective worker count (including the newly capped single-reader path), and add 
a regression covering `limit=1`, one `read_next_batch()`, then public `close()`?



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to