JingsongLi opened a new pull request, #9996: URL: https://github.com/apache/paimon/pull/9996
## Purpose Make Python's streaming Arrow reader use the same effective split parallelism as `to_arrow`, while keeping memory and BLOB I/O concurrency bounded. Depends on apache/paimon-rust#884 for native `TIMESTAMP(0)` / `TIMESTAMP_LTZ(0)` Arrow-second support. `paimon-rust` has not released this native-read API yet, so this intentionally does not retain a compatibility fallback for older builds. ## Changes - Resolve `to_arrow_batch_reader` parallelism from the explicit argument, `read.parallelism`, or the automatic CPU/split limit. - Partition splits across multiple Rust readers and stream batches in completion order. Table scans do not guarantee row order. - Keep at most one pending batch per Rust reader, propagate streaming failures, close readers on cancellation/setup failure, and enforce the global row limit. - Share the existing total BLOB concurrency cap across the Rust readers. - Remove the Python fallback for precision-zero timestamps now that paimon-rust emits Arrow seconds. - Add native Data Evolution coverage for nested `ARRAY<BLOB>` and `MAP<STRING, BLOB>` values and assert that the Python split reader is not used. - Keep Vortex outside the native-read formats in this change. ## Verification - `native_read_test.py` + `native_plan_integration_test.py`: 53 passed, 11 subtests passed. - Native/streaming compatibility suite: 98 passed, 1 skipped, 19 subtests passed. - Full native plan/read suite: 334 passed, 124 subtests passed. - BLOB suites: 238 passed, 2 skipped, 58 subtests passed. - `flake8` passed for all changed files. - Artificial 8-split / 40 ms-per-split benchmark: 0.351 s serial vs 0.089 s with parallelism 4 (3.93x faster), with identical rows after order-insensitive comparison. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
