XiaoHongbo-Hope opened a new pull request, #994:
URL: https://github.com/apache/paimon-rust/pull/994
### Purpose
Linked issue: close #993
Small Parquet objects are read twice today: metadata loading prefetches the
complete object, then the Arrow stream requests its data range from storage
again. This is especially expensive on object stores where request rate,
rather
than transferred bytes, is the limiting resource.
In an anonymized one-minute production sample, 48,872 distinct Parquet
objects
produced 387,280 client/file pairs and 801,034 GETs. Every object was at most
512 KiB, and 77.7% were at most 16 KiB. At the peak second, eliminating the
redundant request would have removed approximately 37k Parquet GET/s (about
49%
of Parquet GETs and 8.5% of all bucket requests in that second), assuming one
metadata/data pair per client and file.
### Brief change log
- Wrap active Parquet readers for files no larger than the existing 512 KiB
metadata prefetch hint.
- Retain a successful whole-file prefetch only for that file reader's
lifetime.
- Serve later byte ranges as zero-copy `Bytes` slices.
- Preserve the underlying reader's cache identity and metadata-cache context.
- Leave larger files on the existing range-read path.
The retained allocation is bounded to 512 KiB per active small-file reader.
It
is neither a process-wide body cache nor an additional copy of the response.
### Tests
- Added an end-to-end Parquet test that verifies metadata plus data decoding
performs one underlying read for a small file.
- Added a threshold test that verifies files larger than the prefetch hint
keep
the existing read path.
- Kept row-group concurrency tests on files above the threshold so they still
exercise actual underlying reads.
- `cargo test -p paimon --lib` (3441 passed, 6 ignored)
- `cargo clippy -p paimon --lib --tests -- -D warnings`
- `cargo fmt --all -- --check`
### API and Format
No public API or storage-format change.
### Documentation
No user-facing documentation change is required.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]