gripleaf commented on code in PR #323:
URL: https://github.com/apache/paimon-cpp/pull/323#discussion_r4011857679


##########
src/paimon/core/operation/file_store_scan.cpp:
##########
@@ -464,6 +482,24 @@ Status FileStoreScan::ReadAndMergeBucketFileEntries(
     return MergeLiveEntries(unmerged_entries, merged_entries);
 }
 
+Result<bool> FileStoreScan::CheckHistoricalBucketCompatibility(
+    const std::vector<ManifestFileMeta>& manifest_metas) const {
+    int64_t max_schema_id = -1;
+    for (const auto& meta : manifest_metas) {
+        max_schema_id = std::max(max_schema_id, meta.SchemaId());
+    }
+    for (int64_t schema_id = 0; schema_id <= max_schema_id; ++schema_id) {
+        PAIMON_ASSIGN_OR_RAISE(bool compatible, 
HasCompatibleBucketKeys(schema_id));
+        if (!compatible) {
+            return false;
+        }
+        if (schema_id == max_schema_id) {
+            break;
+        }
+    }
+    return true;
+}

Review Comment:
   Thanks for the clarification. I confirmed that Java’s bucket-pruning path 
does not perform an equivalent historical schema compatibility check. I also 
discussed this with @wangyong9999, and we agreed that HasCompatibleBucketKeys() 
was overly defensive and could be removed.
   
   The latest commit removes both HasCompatibleBucketKeys() and the full 
historical schema traversal introduced by this PR, along with the compatibility 
cache and schema-ID-based fallback. Bucket pruning now follows Java’s approach 
and no longer reads historical schemas for bucket-key compatibility checks. The 
related tests have been updated accordingly.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to