XiaoHongbo-Hope opened a new pull request, #972:
URL: https://github.com/apache/paimon-rust/pull/972

   ### Purpose
   Closes #971. Cache BLOB metadata independently of value bodies, so repeated 
descriptor scans avoid metadata I/O. Related: apache/paimon#10225.
   
   [RocksDB 
BlobDB](https://github.com/facebook/rocksdb/wiki/BlobDB#column-family-options) 
similarly caches SST blocks containing blob references separately from blob 
values. Here, the cached bytes are footer/index and ARRAY/MAP metadata 
(including keys) used to build descriptors.
   
   ### Brief change log
   Add `blob-meta` to the default whitelist; local caching remains disabled by 
default. Exact ranges reuse cache invalidation and concurrent load coalescing. 
Account for small-entry overhead, cap built-in caches at 65,536 entries, and 
bypass oversized entries. Body reads stay direct unless `data` is enabled.
   
   ### Tests
   185 IO/cache, BLOB reader and data-file reader tests passed with 
`storage-memory,storage-fs`. Tests cover descriptor/body ranges, repeated 
scans, disk reopen, invalidation, capacity and 16 concurrent cold readers 
sharing one read. Formatting, diff checks and scoped Clippy passed (allowing 
the existing unused-import/variable warnings in this feature subset).
   
   ### API and Format
   Add a default `FileRead::read_blob_metadata()` method. Table formats are 
unchanged. The disposable local-cache format/directory moves to v3 to 
distinguish exact ranges from aligned blocks; v2 entries are not reused.
   
   ### Documentation
   Document `blob-meta` and its capacity accounting in the getting-started 
guide.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to