XiaoHongbo-Hope opened a new pull request, #929: URL: https://github.com/apache/paimon-rust/pull/929
## What - Reuse complete parsed `BlobFileIndex` values across independent readers in one process. - Scope cache keys by `FileIO` context, storage identity, path, and file size. - Bound the LRU cache by an estimated 64 MiB of parsed-index memory. - Coalesce concurrent first loads and allow retries after failed loads. - Bypass caching when a reader has no reliable file identity. ## Trade-offs The cache is enabled by default and retains up to 64 MiB of parsed index metadata per process. It caches no BLOB payloads and does not change selection or result semantics. Sharing is process-local only; it does not span Ray workers or other processes. Paimon data files are immutable, and file size is included in the cache identity. With a tracking `FileRead`, opening the same file through 10 independent readers changes repeated index loading from 10 footer + 10 index reads to 1 + 1. These are logical `FileRead` calls, not a claim about production OSS request reduction. ## Tests - `cargo test -p paimon --lib blob` (150 passed) - `cargo test -p paimon --lib blob_index_cache` (8 passed) - `cargo test -p paimon --lib test_file_read_cache_key_is_scoped_to_file_io_context` (1 passed) - `cargo clippy -p paimon --lib -- -D warnings` - `cargo fmt --all --check` - `git diff --check` -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
