lxy-9602 opened a new pull request, #324:
URL: https://github.com/apache/paimon-cpp/pull/324
<!-- PR titles must follow Conventional Commits: <type>(<optional-scope>):
<description> -->
### Purpose
<!-- Linking this pull request to the issue -->
Linked issue: #158
This change improves real-time primary-key reads, commit conflict handling,
and observability.
<!-- What is the purpose of the change -->
#### Main changes
- Retry real-time snapshot conflicts using the configured retry limit,
timeout, and backoff.
- Rebase file changes and merge partition-bucket offsets with the latest
snapshot on every retry.
- Preserve idempotency by validating both the commit identifier and
requested offset ranges.
- Clean up unreferenced offset files after known snapshot conflicts.
- Add real-time metrics for building, sealed, and total memory usage and
physical row counts.
- Collect batch statistics for primary-key real-time stores.
- Push primary-key predicates into level-0 files and in-memory batches while
leaving non-key predicates out of these readers.
#### Failure handling and limitations
A failed real-time commit is terminal for the writer state that produced it.
The caller must discard the `RealtimeContext` and `FileStoreWrite`, recover
durable offsets from the latest snapshot, and replay input from those offsets.
This change does not retry conflicts reported after an external REST catalog
submission. Concurrent rollback or partition deletion from another process must
still be fenced by the upstream coordinator.
### API and Format
<!-- List UT and IT cases to verify this change -->
Adds `RealtimeContext::GetMetrics()`.
Adds `RealtimeStoreDataUsage` and `RealtimeStore::GetDataUsage()`.
Custom RealtimeStore implementations must implement the new data-usage API.
No persisted table, snapshot, manifest, deletion-vector, or offset-file
format changes.
<!-- Does this change affect API in include dir or storage format or
protocol -->
### Tests
Added or updated coverage for:
- real-time commit retry and offset rebasing;
- conflicts between real-time append and compaction commits;
- idempotent and inconsistent offset handling;
- building, sealed, committed, and aggregated metrics;
- primary-key predicate pruning for level-0 and in-memory data;
- Arrow statistics lifetime after the context is released.
### Documentation
<!-- Does this change introduce a new feature -->
### Generative AI tooling
Generated-by: OpenAI Codex (GPT-5)
<!--
If generative AI tooling has been used in the process of authoring this
patch, please include the
phrase: 'Generated-by: ' followed by the name of the tool and its version.
If no, write 'No'.
Please refer to the [ASF Generative Tooling
Guidance](https://www.apache.org/legal/generative-tooling.html) for details.
-->
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]