gripleaf opened a new pull request, #323:
URL: https://github.com/apache/paimon-cpp/pull/323
### Purpose
We have a Paimon table with 90 partitions and approximately 16K buckets
per partition. At this scale, manifest processing becomes a bottleneck during
scan planning: even when a query targets a specific bucket, reading the
relevant manifests can materialize many entries belonging to other buckets.
This change reduces unnecessary materialization for bucket-specific scans:
- Probe version and bucket columns, validate every entry's version, and
materialize only matching entries.
- Perform the extra pass only when manifest bytes are retained in memory,
avoiding repeated remote reads.
- Use a single pass when manifest metadata identifies only the requested
bucket.
- Support precise Avro bitmap selection while preserving physical row IDs.
- Reuse block positions across schema resets, with a limit of 65,536
blocks and sequential fallback beyond that limit.
No performance benchmark was run for this patch. High selection ratios in
mixed-bucket manifests may incur additional probe and decompression costs.
### Tests
- Debug build with `-Wall -Werror`: passed.
- Full Avro suite: 74 tests passed.
- Full core suite: 2,065 tests passed.
- Focused manifest, scan, and metadata-reader tests: 47 tests passed.
- Changed-file pre-commit checks and `git diff --check`: passed.
- Independent local subagent review: no blocking findings.
Regression coverage includes sparse and empty selections, out-of-range row
IDs, compressed blocks, schema resets, empty files, block-index capacity
fallback, cold and warm caches, cache rejection, invalid versions, null bucket
fields, and Java 0.9/1.1 manifest compatibility.
The full integration suite was not run.
### API and Format
No changes to public headers, storage formats, or protocols.
Avro implements the existing precise bitmap-selection contract. Legacy
manifest schema evolution remains supported.
### Documentation
No new configuration options. Code comments document the in-memory-only
preparation and bounded block-index reuse.
### Generative AI tooling
Generated-by: OpenAI Codex (GPT-6)
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]