gripleaf opened a new pull request, #412:
URL: https://github.com/apache/paimon-cpp/pull/412

   ### Purpose
   
   Linked issue: none (performance improvement).
   
   Repeated Parquet pre-buffer reads of the same immutable file range currently 
fetch the bytes from storage again. Add an **opt-in** exact-range cache for 
asynchronous reads. A cache hit returns an owned Arrow buffer; a miss preserves 
the filesystem's asynchronous read and attempts admission only after successful 
completion.
   
   The option defaults to disabled. Entries share the caller's bounded cache, 
with a configurable per-range admission limit (4 MiB by default). Cache 
eviction cannot invalidate a returned buffer, and cached bytes retain their 
allocator. Missing cache capability, ineligible ranges and admission failures 
do not fail ordinary reads. Failed reads are not cached. No query results, 
decoded columns, mutable reader state or snapshot discovery are cached.
   
   Synthetic tests make the avoided work observable: a repeated exact-range 
read completes from the cache without a second storage call, while a different 
file or byte range still issues I/O. The reader exposes hit/miss/bypass and 
byte counters so users can evaluate benefits on their own cold/warm workloads. 
This draft does not claim a universal latency improvement.
   
   ### Tests
   
   - CMake Debug build with shared libraries and repository-bundled 
dependencies, `-Wall -Werror`.
   - `cmake --build build --target paimon-parquet-format-test 
paimon-common-test -j 24`
   - `./build/debug/paimon-parquet-format-test`: 235 tests passed.
   - `./build/debug/paimon-common-test 
--gtest_filter='*Cache*:*ArrowInputStream*:*Metrics*'`: 114 passed, 2 existing 
conditional skips.
   - Full `./build/debug/paimon-common-test`: 1,694 passed, the same 2 existing 
conditional skips.
   - New synthetic tests cover non-blocking misses, storage failures, rejected 
admissions, disabled/unsupported/oversized bypasses, eviction and allocator 
lifetime, independent readers, file/range key isolation and concurrent warm 
readers.
   - Real Parquet reader regression covers cold/warm metric propagation and 
snapshots after close; invalid admission-size configuration is rejected.
   - Full-repository pre-commit hooks passed using an explicit list of all 
tracked files and new files (the host Git lacks `--deduplicate`, required by 
pre-commit's `--all-files` implementation); final added tests also passed 
separately.
   - `git diff --check` passed.
   
   ### API and Format
   
   Adds `Cache::GetIfPresent()` with a default `NotImplemented` implementation, 
preserving source compatibility for custom cache subclasses. It must not invoke 
a loader or wait for storage I/O. `LruCache` implements it; unsupported custom 
caches bypass data caching.
   
   **C++ ABI impact:** the added virtual member requires rebuilding 
applications and custom cache implementations against the updated 
headers/library. No storage format or protocol changes.
   
   New read options: `parquet.read.enable-data-cache` (default false) and 
`parquet.read.data-cache.max-range-bytes` (default 4194304, positive). New 
cumulative reader metrics cover hits, misses, bypasses, hit bytes, attempted 
admission bytes and admission failures.
   
   ### Documentation
   
   Added the Parquet data cache guide and user-guide navigation. It documents 
immutable URI requirements, total-budget versus retained-buffer memory, 
configuration, custom cache integration, ABI impact and counter semantics. 
Attempted admission bytes are not resident cache size.
   
   ### Generative AI tooling
   
   Generated-by: OpenAI Codex (GPT-6)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to