gripleaf opened a new pull request, #411:
URL: https://github.com/apache/paimon-cpp/pull/411

   ### Purpose
   
   Linked issue: none (performance improvement).
   
   Data-evolution scans with row ID ranges currently deserialize manifest 
entries and their file metadata before discarding non-overlapping files. This 
change probes the aligned Arrow row-count/first-row-ID columns first and 
materializes only candidates. It also reuses aligned immutable manifest 
batches, as Arrow IPC bytes, in the caller's existing bounded manifest cache. 
These two changes form one manifest-read optimization; no snapshot selection or 
query result is cached.
   
   Unknown or invalid ranges are retained conservatively, serialization 
versions are checked before pruning, Add/Delete entries use the same selection, 
and ordinary entry filters still run. Cache values retain their allocator, 
readers own their decoding state, and optional cache-loading failures fall back 
to normal reads. Existing lazy-decode and data-evolution gates are preserved.
   
   The regression tests make the work reduction observable without a private 
dataset: a disjoint entry is pruned before materialization, repeated reads 
reuse the aligned cache without another load, and eviction forces a reload. The 
new counters and histogram allow callers to evaluate latency and cache 
effectiveness on their own workloads. This draft does not claim a universal 
speedup.
   
   ### Tests
   
   - CMake Debug build with tests, shared libraries, Avro and ORC enabled, 
using the repository's bundled dependencies and `-Wall -Werror`.
   - `cmake --build build --target paimon-core-test -j 24`
   - `./build/debug/paimon-core-test 
--gtest_filter='*Manifest*:*FileStoreScan*:*RowRange*:*DataEvolution*'`: 153 
tests passed.
   - `./build/debug/paimon-core-test`: all 2,273 core tests passed.
   - New `RowRangeManifestFileTest` coverage: boundary/disjoint ranges, 
Add/Delete merge equivalence, unknown/overflow ranges, schema evolution, 
unsupported versions, missing files/filter failures, bounded cache 
fallback/eviction, concurrent readers, allocator lifetime, counters and 
histogram snapshots.
   - Full-repository pre-commit hooks passed by explicitly enumerating `git 
ls-files` plus the new test. The host Git is too old for the `--deduplicate` 
command used by `pre-commit run --all-files`; the equivalent explicit file-list 
invocation was used.
   - `git diff --check` passed.
   
   ### API and Format
   
   No storage format or protocol changes. Additive `ScanMetrics` constants 
expose cumulative row-range scan/pruning/materialization work, cache 
hit/miss/fallback counts and a per-file duration histogram. No new pure virtual 
method or user configuration is introduced.
   
   ### Documentation
   
   Updated the metrics guide with counter semantics, cache ownership/scope and 
per-file wall-time units. Parallel file durations overlap and must not be 
summed into request latency.
   
   ### Generative AI tooling
   
   Generated-by: OpenAI Codex (GPT-6)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to