JingsongLi opened a new pull request, #880:
URL: https://github.com/apache/paimon-rust/pull/880
### Purpose
Reduce scalar inline BLOB read amplification. The reader previously issued
one FileRead request per selected BLOB entry, so a 16,000-row scan performed
roughly 16,000 small range reads even when entries were adjacent.
### Changes
- Coalesce nearby scalar BLOB entry ranges with the same defaults as the
Python reader: 1 MiB maximum gap and 8 MiB maximum span.
- Fetch merged spans with the configured blob parallelism, then restore the
requested row order.
- Preserve per-entry magic, embedded-length, and CRC32 validation after
slicing each merged response.
- Strengthen the data-evolution fallback test to assert exact payload ranges
instead of request counts.
### Performance
Release Python binding, local filesystem, 16,000 rows, 4 KiB payloads, 8
splits, 5 measured iterations:
- native parallelism=1, blob_parallelism=1: 152.1 ms -> 17.5 ms
- native parallelism=4, blob_parallelism=1: 136.8 ms -> 5.8 ms
- native parallelism=4, blob_parallelism=4: 134.6 ms -> 6.0 ms
For 512 rows with 128 KiB payloads in one split, increasing blob_parallelism
from 1 to 4 reduced the native median from 16.5 ms to 13.8 ms, confirming that
separate 8 MiB spans still read concurrently.
All benchmark results matched Python row counts, schemas, and checksums.
### Verification
- cargo fmt --all -- --check
- cargo clippy --locked --all-targets --workspace --features fulltext,vortex
-- -D warnings
- cargo test --locked -p paimon --all-targets --features fulltext,vortex
- 3009 unit tests passed, 2 ignored
- all integration and example targets passed
- complete BLOB module tests: 38 passed
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]