leaves12138 commented on code in PR #10003:
URL: https://github.com/apache/paimon/pull/10003#discussion_r4056522916
##########
paimon-python/pypaimon/read/table_read.py:
##########
@@ -219,7 +219,7 @@ def to_arrow_batch_reader(
table reads do not guarantee row order. Python fallback reads remain
serial.
"""
- effective = self._resolve_parallelism(parallelism, len(splits))
+ effective = self._effective_parallelism(parallelism, len(splits))
Review Comment:
[P2] Preserve explicit iterator cleanup when LIMIT reduces concurrency to one
With two or more native splits, `parallelism=2`, and an unfiltered
`limit=1`, the new effective parallelism becomes 1. Consequently, the check
below skips `_ClosableArrowBatchReader` and returns the raw
`RecordBatchReader.from_batches(...)` reader. On PyArrow 19.0.1, calling
`read_next_batch()` once and then `close()` does not close the suspended
`_convert_native_batches` generator, so its `finally` never calls the
underlying native reader's `close()`. This retains the native stream until
later exhaustion/garbage collection instead of releasing it at explicit close.
I reproduced this through `to_arrow_batch_reader` with close-tracking native
readers: the base revision closes both readers after that sequence, whereas
this head leaves its sole native reader open. `read_all()` hides the regression
by exhausting the iterator.
Could we preserve the explicit iterator-close chain independently of the
effective worker count (including the newly capped single-reader path), and add
a regression covering `limit=1`, one `read_next_batch()`, then public `close()`?
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]