XiaoHongbo-Hope opened a new issue, #993:
URL: https://github.com/apache/paimon-rust/issues/993

   ### Search before asking
   
   I searched the existing issues and did not find a report for this behavior.
   
   ### Description
   
   The native Parquet reader configures a 512 KiB metadata prefetch hint. When a
   Parquet object is no larger than that hint, metadata loading reads the 
complete
   object. The subsequent Arrow data read nevertheless requests the selected 
byte
   range from storage again.
   
   For object storage this turns one logical read of a small immutable Parquet 
file
   into two GET requests and retransmits part of the file.
   
   An anonymized one-minute production sample contained:
   
   - 48,872 distinct Parquet objects;
   - 387,280 distinct client/file pairs;
   - 801,034 Parquet GET requests; and
   - no Parquet object larger than 512 KiB (77.7% were at most 16 KiB).
   
   ### Expected behavior
   
   When metadata prefetch returns the complete object, retain that `Bytes`
   allocation for the lifetime of the active file reader and serve subsequent 
data
   ranges by zero-copy slicing. Large files should keep the existing range-read
   path, and the optimization should not introduce a process-wide body cache.
   
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to