JingsongLi opened a new pull request, #10003:
URL: https://github.com/apache/paimon/pull/10003

   ## Summary
   
   - close every native reader before waiting for parallel workers, so early 
batch-reader close interrupts workers blocked in their first read
   - cap unfiltered LIMIT fan-out to avoid eagerly constructing readers that 
cannot contribute result rows, while retaining full parallelism for filtered 
limits
   - balance contiguous native split groups by physical file bytes instead of 
split count, preserving input order while reducing tail skew
   - update native-read documentation and add deterministic concurrency, LIMIT, 
and grouping regressions
   
   ## Dependency
   
   Depends on apache/paimon-rust#885 for the cancellable and idempotent 
RecordBatchReader.close API. The Rust Python binding has not been released, so 
this intentionally targets current paimon-rust main rather than adding 
compatibility logic for older builds.
   
   ## Performance
   
   A local native-read benchmark used four valid Paimon splits of about 39.5 
MiB, 39.5 MiB, 1.58 MiB, and 1.58 MiB with parallelism 2.
   
   - count grouping: 79.0 MiB vs 3.2 MiB; median 10.83 ms
   - byte grouping: 39.5 MiB vs 42.7 MiB; median 7.55 ms
   
   The byte-balanced path was about 30 percent faster in this cached local 
decode benchmark.
   
   ## Tests
   
   - native read, native integration, Data Evolution, formats, row rolling, 
reader parallelism, and deferred BLOB suites: 177 passed, 15 subtests passed
   - paimon-rust binding test_read.py against the rebuilt extension: 82 passed
   - focused Rust pending-read cancellation test and clippy with warnings denied
   - flake8 on the changed Python files


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to