gripleaf opened a new pull request, #412: URL: https://github.com/apache/paimon-cpp/pull/412
### Purpose Linked issue: none (performance improvement). Repeated Parquet pre-buffer reads of the same immutable file range currently fetch the bytes from storage again. Add an **opt-in** exact-range cache for asynchronous reads. A cache hit returns an owned Arrow buffer; a miss preserves the filesystem's asynchronous read and attempts admission only after successful completion. The option defaults to disabled. Entries share the caller's bounded cache, with a configurable per-range admission limit (4 MiB by default). Cache eviction cannot invalidate a returned buffer, and cached bytes retain their allocator. Missing cache capability, ineligible ranges and admission failures do not fail ordinary reads. Failed reads are not cached. No query results, decoded columns, mutable reader state or snapshot discovery are cached. Synthetic tests make the avoided work observable: a repeated exact-range read completes from the cache without a second storage call, while a different file or byte range still issues I/O. The reader exposes hit/miss/bypass and byte counters so users can evaluate benefits on their own cold/warm workloads. This draft does not claim a universal latency improvement. ### Tests - CMake Debug build with shared libraries and repository-bundled dependencies, `-Wall -Werror`. - `cmake --build build --target paimon-parquet-format-test paimon-common-test -j 24` - `./build/debug/paimon-parquet-format-test`: 235 tests passed. - `./build/debug/paimon-common-test --gtest_filter='*Cache*:*ArrowInputStream*:*Metrics*'`: 114 passed, 2 existing conditional skips. - Full `./build/debug/paimon-common-test`: 1,694 passed, the same 2 existing conditional skips. - New synthetic tests cover non-blocking misses, storage failures, rejected admissions, disabled/unsupported/oversized bypasses, eviction and allocator lifetime, independent readers, file/range key isolation and concurrent warm readers. - Real Parquet reader regression covers cold/warm metric propagation and snapshots after close; invalid admission-size configuration is rejected. - Full-repository pre-commit hooks passed using an explicit list of all tracked files and new files (the host Git lacks `--deduplicate`, required by pre-commit's `--all-files` implementation); final added tests also passed separately. - `git diff --check` passed. ### API and Format Adds `Cache::GetIfPresent()` with a default `NotImplemented` implementation, preserving source compatibility for custom cache subclasses. It must not invoke a loader or wait for storage I/O. `LruCache` implements it; unsupported custom caches bypass data caching. **C++ ABI impact:** the added virtual member requires rebuilding applications and custom cache implementations against the updated headers/library. No storage format or protocol changes. New read options: `parquet.read.enable-data-cache` (default false) and `parquet.read.data-cache.max-range-bytes` (default 4194304, positive). New cumulative reader metrics cover hits, misses, bypasses, hit bytes, attempted admission bytes and admission failures. ### Documentation Added the Parquet data cache guide and user-guide navigation. It documents immutable URI requirements, total-budget versus retained-buffer memory, configuration, custom cache integration, ABI impact and counter semantics. Attempted admission bytes are not resident cache size. ### Generative AI tooling Generated-by: OpenAI Codex (GPT-6) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
