XiaoHongbo-Hope opened a new pull request, #994:
URL: https://github.com/apache/paimon-rust/pull/994

   ### Purpose
   
   Linked issue: close #993
   
   Small Parquet objects are read twice today: metadata loading prefetches the
   complete object, then the Arrow stream requests its data range from storage
   again. This is especially expensive on object stores where request rate, 
rather
   than transferred bytes, is the limiting resource.
   
   In an anonymized one-minute production sample, 48,872 distinct Parquet 
objects
   produced 387,280 client/file pairs and 801,034 GETs. Every object was at most
   512 KiB, and 77.7% were at most 16 KiB. At the peak second, eliminating the
   redundant request would have removed approximately 37k Parquet GET/s (about 
49%
   of Parquet GETs and 8.5% of all bucket requests in that second), assuming one
   metadata/data pair per client and file.
   
   ### Brief change log
   
   - Wrap active Parquet readers for files no larger than the existing 512 KiB
     metadata prefetch hint.
   - Retain a successful whole-file prefetch only for that file reader's 
lifetime.
   - Serve later byte ranges as zero-copy `Bytes` slices.
   - Preserve the underlying reader's cache identity and metadata-cache context.
   - Leave larger files on the existing range-read path.
   
   The retained allocation is bounded to 512 KiB per active small-file reader. 
It
   is neither a process-wide body cache nor an additional copy of the response.
   
   ### Tests
   
   - Added an end-to-end Parquet test that verifies metadata plus data decoding
     performs one underlying read for a small file.
   - Added a threshold test that verifies files larger than the prefetch hint 
keep
     the existing read path.
   - Kept row-group concurrency tests on files above the threshold so they still
     exercise actual underlying reads.
   - `cargo test -p paimon --lib` (3441 passed, 6 ignored)
   - `cargo clippy -p paimon --lib --tests -- -D warnings`
   - `cargo fmt --all -- --check`
   
   ### API and Format
   
   No public API or storage-format change.
   
   ### Documentation
   
   No user-facing documentation change is required.
   
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to