JingsongLi opened a new pull request, #880:
URL: https://github.com/apache/paimon-rust/pull/880

   ### Purpose
   
   Reduce scalar inline BLOB read amplification. The reader previously issued 
one FileRead request per selected BLOB entry, so a 16,000-row scan performed 
roughly 16,000 small range reads even when entries were adjacent.
   
   ### Changes
   
   - Coalesce nearby scalar BLOB entry ranges with the same defaults as the 
Python reader: 1 MiB maximum gap and 8 MiB maximum span.
   - Fetch merged spans with the configured blob parallelism, then restore the 
requested row order.
   - Preserve per-entry magic, embedded-length, and CRC32 validation after 
slicing each merged response.
   - Strengthen the data-evolution fallback test to assert exact payload ranges 
instead of request counts.
   
   ### Performance
   
   Release Python binding, local filesystem, 16,000 rows, 4 KiB payloads, 8 
splits, 5 measured iterations:
   
   - native parallelism=1, blob_parallelism=1: 152.1 ms -> 17.5 ms
   - native parallelism=4, blob_parallelism=1: 136.8 ms -> 5.8 ms
   - native parallelism=4, blob_parallelism=4: 134.6 ms -> 6.0 ms
   
   For 512 rows with 128 KiB payloads in one split, increasing blob_parallelism 
from 1 to 4 reduced the native median from 16.5 ms to 13.8 ms, confirming that 
separate 8 MiB spans still read concurrently.
   
   All benchmark results matched Python row counts, schemas, and checksums.
   
   ### Verification
   
   - cargo fmt --all -- --check
   - cargo clippy --locked --all-targets --workspace --features fulltext,vortex 
-- -D warnings
   - cargo test --locked -p paimon --all-targets --features fulltext,vortex
     - 3009 unit tests passed, 2 ignored
     - all integration and example targets passed
   - complete BLOB module tests: 38 passed


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to