maximilliangrand opened a new pull request, #51543: URL: https://github.com/apache/arrow/pull/51543
### Rationale for this change `fragment_readahead=0` is documented to disable fragment readahead, but `Scanner.to_table()` hangs even on an empty dataset. The unordered scan creates a merged generator with zero subscriptions, so no fragment is pulled and the first consumer future never completes. Fixes #48111. ### What changes are included in this PR? Keep at least one active fragment in the unordered merge. All positive readahead limits and the sequenced scan path retain their existing behavior. Add regressions for empty and multiple-fragment datasets, sequenced and unordered C++ scans, and threaded and unthreaded Python Parquet scans. ### Are these changes tested? - The new C++ regression fails against the unmodified source-built dataset library: its scan future does not finish. The fixed build passes the complete dataset scanner suite (137 tests; one existing disabled test). - Source-built PyArrow: `pytest python/pyarrow/tests/test_dataset.py -k "scanner or to_table or fragment"` passes 63 tests, with 2 skips for S3 and Substrait support not enabled in the local build. - The public API reproducer completes for zero readahead with threads on/off and empty datasets, plus `head`, `to_batches`, and positive-readahead controls. Tested on macOS arm64 with Python 3.13. - C++ formatting checked with clang-format 18.1.8; Python formatting checked with the repository's autopep8 configuration. ### Are there any user-facing changes? Table scans finish when fragment readahead is disabled. No public API changes. ### Was AI used for this PR? Codex investigated the issue, wrote the implementation and regression tests, and ran local validation. The changes also received independent AI review. **PR code and description written by:** - [ ] Human - [x] AI **Reviewed before submission by:** - [ ] Human - [x] AI - [ ] Not reviewed -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
