wangyong9999 opened a new pull request, #314:
URL: https://github.com/apache/paimon-cpp/pull/314

   ### Purpose
   
   Related to #273; this narrower exact-range cache does not close that issue. 
Add default-off `parquet.read.enable-data-cache`, reusing the caller-provided 
bounded Cache across reader lifetimes. Data uses CacheKind::DEFAULT; metadata 
retains its budget. URI/offset/length identify immutable bytes. Missing URI 
bypasses caching, failed and short reads are not published, and async cached 
reads use Arrow's IO executor. No global cache or whole-file prefetch is added; 
partial overlaps are not reused.
   
   ### Tests
   
   Three new tests cover reader destruction, eviction/reload, both pre-buffer 
modes, async reuse, missing URI and invalid options. Those plus two existing 
page-index tests passed in a focused C++17 executable built from changed 
headers/test source against an existing dependency bundle; warm readers report 
zero storage-read bytes. Full upstream pre-commit and diff-check pass. This is 
not a full CMake build; existing tests/older GoogleTest emit signed-comparison 
warnings. Full warning-free build and end-to-end capacity gains remain 
unverified.
   
   ### API and Format
   
   No public include headers, storage format or protocol change. One 
default-off Parquet option using the existing CacheKind::DEFAULT budget.
   
   ### Documentation
   
   Documented configuration, immutable URI contract, exact-range limitations, 
capacity/eviction ownership, allocator lifetime and storage-read-byte 
accounting in parquet_metadata_cache.rst.
   
   ### Generative AI tooling
   
   Generated-by: OpenAI Codex (GPT-5)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to