zhang-arvin commented on PR #9497: URL: https://github.com/apache/paimon/pull/9497#issuecomment-6060185024
Rebased onto latest `master` (conflicts in `IcebergManifestFile` / `AvroFileFormat` resolved) and addressed the remaining positive-ID [P1] on head `c416e65fd`: - One consistent positive-ID mapping. `IcebergSchema#create` now shifts every Paimon ID - top-level and nested ROW/ARRAY/MAP/MULTISET alike - by one, so the manifest header, the table metadata schemas, the partition source IDs and the manifest metrics maps (`null_value_counts` / `lower_bounds` / `upper_bounds`) all share the same ID space. The previous top-level-only `withPositiveFieldIds` is removed, because deriving everything from `IcebergSchema#create` makes the mapping recursive by construction. - Added `IcebergSchemaTest`, asserting that a schema with nested types emits unique, positive IDs with the nested struct ids following their parent (no duplicate ids). - The `partition-spec` array (earlier round) and the per-`Content` writer factories are unchanged. - The `IcebergManifestFile` conflict is resolved by keeping master's `FileFormat#manifestFormat` for manifest blocks and building one writer factory per `Content`. Note: `IcebergCompatibilityTest` needs Hadoop/JDK 8-11, so the end-to-end `ManifestFiles.read` interop is validated by CI rather than locally. This comment was generated by an AI agent (Hermes Agent). -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
