SteNicholas opened a new pull request, #380: URL: https://github.com/apache/paimon-cpp/pull/380
### Purpose Linked issue: close #347 Java's `DataTypeJsonParser` numbers the ROW fields of a parsed type when none of them carries an id, so a configured `variant.shreddingSchema` whose fields all omit their ids is accepted by Java Paimon but was rejected here while deserializing the first `DataField` (`key 'id' must exist`). An embedding engine therefore had to route an otherwise supported Variant write back to the Java SDK only because the configured schema omitted optional field ids. This threads a field id assigner through the data type JSON parsing, mirroring Java's all-or-none rule: - ids are generated in traversal order starting at 0 when no ROW field carries one; - explicit ids are preserved at every level; - a type that mixes both forms is rejected with `Partial field id is not allowed.` One deliberate deviation from Java: Java rejects the mix only when a generated id comes first, so an explicit id followed by an omitted one silently produces duplicate ids there, while both orders are rejected here. Deserializing a stored table schema still requires every field to carry its own id, as before. ### Tests - `DataTypeJsonParserTest.ParseTypeRowTypeGeneratesMissingFieldIds`: ids generated in traversal order across nested ROW, ARRAY and MAP descendants. - `DataTypeJsonParserTest.ParseTypeRowTypeSuccess`: non-sequential explicit ids preserved at every level. - `DataTypeJsonParserTest.ParseTypeComplexTypeFailure`: partial ids rejected in either order, alongside the malformed complex type cases. - `DataFieldTest.FromJson` and `DataFieldTest.FromJsonFailed`: a stored schema still requires an id on every field, names and descriptions keep embedded NULs, and each malformed field key reports its own diagnostic. - `VariantShreddingWritePlanFactoryTest.ConfiguredSchema`: `variant.shreddingSchema` and its `parquet.variant.shreddingSchema` fallback plan the same physical schema, field ids included, whether the configured schema carries ids or omits them. - `VariantShreddingWritePlanFactoryTest.ConfiguredSchemaRejectsPartialFieldIds`. - `VariantParquetTest.ShreddedWriteAndReadRoundTrip`: the new `configured-without-ids` mode asserts the generated ids reach the Parquet physical schema and that the shredded values round trip. ### API and Format `include/paimon/defs.h` only gains documentation for `variant.shreddingSchema`; there is no public symbol, storage format or protocol change. A configured schema that already carries field ids produces exactly the same physical schema as before. ### Documentation `docs/source/user_guide/data_types.rst` documents that the configured shredding schema's fields either all carry an `id` or all omit one, in which case the ids are assigned in traversal order starting at 0. ### Generative AI tooling Generated-by: Claude Opus 5 🤖 Generated with [Claude Code](https://claude.com/claude-code) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
