gripleaf commented on code in PR #323:
URL: https://github.com/apache/paimon-cpp/pull/323#discussion_r4002313748


##########
src/paimon/core/operation/file_store_scan.cpp:
##########
@@ -464,6 +482,24 @@ Status FileStoreScan::ReadAndMergeBucketFileEntries(
     return MergeLiveEntries(unmerged_entries, merged_entries);
 }
 
+Result<bool> FileStoreScan::CheckHistoricalBucketCompatibility(
+    const std::vector<ManifestFileMeta>& manifest_metas) const {
+    int64_t max_schema_id = -1;
+    for (const auto& meta : manifest_metas) {
+        max_schema_id = std::max(max_schema_id, meta.SchemaId());
+    }
+    for (int64_t schema_id = 0; schema_id <= max_schema_id; ++schema_id) {
+        PAIMON_ASSIGN_OR_RAISE(bool compatible, 
HasCompatibleBucketKeys(schema_id));
+        if (!compatible) {
+            return false;
+        }
+        if (schema_id == max_schema_id) {
+            break;
+        }
+    }
+    return true;
+}

Review Comment:
   Thanks for pointing this out. I agree that repeatedly loading historical 
schemas can be costly. My plan is to reduce this overhead through #329, which 
allows scans of the same table and branch to share a `SchemaManager` through 
`TableScanResources`. After warm-up, subsequent scans can reuse cached schema 
objects, avoiding repeated schema-file reads and JSON parsing.
   
   Sharing the manager amortizes schema-loading costs across scans, but it does 
not eliminate the initial loading cost or the traversal of historical versions, 
including unused ones. I propose keeping this correctness safeguard here and 
using #329 to reduce the repeated I/O overhead.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to