wangyong9999 opened a new pull request, #314: URL: https://github.com/apache/paimon-cpp/pull/314
### Purpose Related to #273; this narrower exact-range cache does not close that issue. Add default-off `parquet.read.enable-data-cache`, reusing the caller-provided bounded Cache across reader lifetimes. Data uses CacheKind::DEFAULT; metadata retains its budget. URI/offset/length identify immutable bytes. Missing URI bypasses caching, failed and short reads are not published, and async cached reads use Arrow's IO executor. No global cache or whole-file prefetch is added; partial overlaps are not reused. ### Tests Three new tests cover reader destruction, eviction/reload, both pre-buffer modes, async reuse, missing URI and invalid options. Those plus two existing page-index tests passed in a focused C++17 executable built from changed headers/test source against an existing dependency bundle; warm readers report zero storage-read bytes. Full upstream pre-commit and diff-check pass. This is not a full CMake build; existing tests/older GoogleTest emit signed-comparison warnings. Full warning-free build and end-to-end capacity gains remain unverified. ### API and Format No public include headers, storage format or protocol change. One default-off Parquet option using the existing CacheKind::DEFAULT budget. ### Documentation Documented configuration, immutable URI contract, exact-range limitations, capacity/eviction ownership, allocator lifetime and storage-read-byte accounting in parquet_metadata_cache.rst. ### Generative AI tooling Generated-by: OpenAI Codex (GPT-5) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
