gripleaf commented on code in PR #323:
URL: https://github.com/apache/paimon-cpp/pull/323#discussion_r4002313748


##########
src/paimon/core/operation/file_store_scan.cpp:
##########
@@ -464,6 +482,24 @@ Status FileStoreScan::ReadAndMergeBucketFileEntries(
     return MergeLiveEntries(unmerged_entries, merged_entries);
 }
 
+Result<bool> FileStoreScan::CheckHistoricalBucketCompatibility(
+    const std::vector<ManifestFileMeta>& manifest_metas) const {
+    int64_t max_schema_id = -1;
+    for (const auto& meta : manifest_metas) {
+        max_schema_id = std::max(max_schema_id, meta.SchemaId());
+    }
+    for (int64_t schema_id = 0; schema_id <= max_schema_id; ++schema_id) {
+        PAIMON_ASSIGN_OR_RAISE(bool compatible, 
HasCompatibleBucketKeys(schema_id));
+        if (!compatible) {
+            return false;
+        }
+        if (schema_id == max_schema_id) {
+            break;
+        }
+    }
+    return true;
+}

Review Comment:
   Thanks for pointing this out. I agree that repeatedly loading historical 
schemas can be costly. My plan is to reduce this overhead through #329, which 
allows scans of the same table and branch to share a `SchemaManager` through 
`TableScanResources`. After warm-up, subsequent scans can reuse cached schema 
objects, avoiding repeated schema-file reads and JSON parsing.
   
   I would retain the compatibility check to preserve correctness for the 
schema histories currently supported by the C++ implementation. Unchanged 
bucket-key names do not necessarily imply unchanged hashing: the regression 
test covers an `INT`-to-`BIGINT` change with the same bucket count. Pruning 
solely by the current bucket could discard matching historical files.
   
   Sharing the manager amortizes schema-loading costs across scans, but it does 
not eliminate the initial loading cost or the traversal of historical versions, 
including unused ones. I propose keeping this correctness safeguard here and 
using #329 to reduce the repeated I/O overhead.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to