gripleaf commented on code in PR #323:
URL: https://github.com/apache/paimon-cpp/pull/323#discussion_r4002313748
##########
src/paimon/core/operation/file_store_scan.cpp:
##########
@@ -464,6 +482,24 @@ Status FileStoreScan::ReadAndMergeBucketFileEntries(
return MergeLiveEntries(unmerged_entries, merged_entries);
}
+Result<bool> FileStoreScan::CheckHistoricalBucketCompatibility(
+ const std::vector<ManifestFileMeta>& manifest_metas) const {
+ int64_t max_schema_id = -1;
+ for (const auto& meta : manifest_metas) {
+ max_schema_id = std::max(max_schema_id, meta.SchemaId());
+ }
+ for (int64_t schema_id = 0; schema_id <= max_schema_id; ++schema_id) {
+ PAIMON_ASSIGN_OR_RAISE(bool compatible,
HasCompatibleBucketKeys(schema_id));
+ if (!compatible) {
+ return false;
+ }
+ if (schema_id == max_schema_id) {
+ break;
+ }
+ }
+ return true;
+}
Review Comment:
Thanks for pointing this out. I agree that repeatedly loading historical
schemas can be costly. My plan is to reduce this overhead through #329, which
allows scans of the same table and branch to share a `SchemaManager` through
`TableScanResources`. After warm-up, subsequent scans can reuse cached schema
objects, avoiding repeated schema-file reads and JSON parsing.
I would retain the compatibility check to preserve correctness for the
schema histories currently supported by the C++ implementation. Unchanged
bucket-key names do not necessarily imply unchanged hashing: the regression
test covers an `INT`-to-`BIGINT` change with the same bucket count. Pruning
solely by the current bucket could discard matching historical files.
Sharing the manager amortizes schema-loading costs across scans, but it does
not eliminate the initial loading cost or the traversal of historical versions,
including unused ones. I propose keeping this correctness safeguard here and
using #329 to reduce the repeated I/O overhead.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]