JingsongLi opened a new pull request, #9996:
URL: https://github.com/apache/paimon/pull/9996

   ## Purpose
   
   Make Python's streaming Arrow reader use the same effective split 
parallelism as `to_arrow`, while keeping memory and BLOB I/O concurrency 
bounded.
   
   Depends on apache/paimon-rust#884 for native `TIMESTAMP(0)` / 
`TIMESTAMP_LTZ(0)` Arrow-second support. `paimon-rust` has not released this 
native-read API yet, so this intentionally does not retain a compatibility 
fallback for older builds.
   
   ## Changes
   
   - Resolve `to_arrow_batch_reader` parallelism from the explicit argument, 
`read.parallelism`, or the automatic CPU/split limit.
   - Partition splits across multiple Rust readers and stream batches in 
completion order. Table scans do not guarantee row order.
   - Keep at most one pending batch per Rust reader, propagate streaming 
failures, close readers on cancellation/setup failure, and enforce the global 
row limit.
   - Share the existing total BLOB concurrency cap across the Rust readers.
   - Remove the Python fallback for precision-zero timestamps now that 
paimon-rust emits Arrow seconds.
   - Add native Data Evolution coverage for nested `ARRAY<BLOB>` and 
`MAP<STRING, BLOB>` values and assert that the Python split reader is not used.
   - Keep Vortex outside the native-read formats in this change.
   
   ## Verification
   
   - `native_read_test.py` + `native_plan_integration_test.py`: 53 passed, 11 
subtests passed.
   - Native/streaming compatibility suite: 98 passed, 1 skipped, 19 subtests 
passed.
   - Full native plan/read suite: 334 passed, 124 subtests passed.
   - BLOB suites: 238 passed, 2 skipped, 58 subtests passed.
   - `flake8` passed for all changed files.
   - Artificial 8-split / 40 ms-per-split benchmark: 0.351 s serial vs 0.089 s 
with parallelism 4 (3.93x faster), with identical rows after order-insensitive 
comparison.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to