SteNicholas opened a new pull request, #380:
URL: https://github.com/apache/paimon-cpp/pull/380

   ### Purpose
   
   Linked issue: close #347
   
   Java's `DataTypeJsonParser` numbers the ROW fields of a parsed type when 
none of them carries an id, so a configured `variant.shreddingSchema` whose 
fields all omit their ids is accepted by Java Paimon but was rejected here 
while deserializing the first `DataField` (`key 'id' must exist`). An embedding 
engine therefore had to route an otherwise supported Variant write back to the 
Java SDK only because the configured schema omitted optional field ids.
   
   This threads a field id assigner through the data type JSON parsing, 
mirroring Java's all-or-none rule:
   
   - ids are generated in traversal order starting at 0 when no ROW field 
carries one;
   - explicit ids are preserved at every level;
   - a type that mixes both forms is rejected with `Partial field id is not 
allowed.`
   
   One deliberate deviation from Java: Java rejects the mix only when a 
generated id comes first, so an explicit id followed by an omitted one silently 
produces duplicate ids there, while both orders are rejected here. 
Deserializing a stored table schema still requires every field to carry its own 
id, as before.
   
   ### Tests
   
   - `DataTypeJsonParserTest.ParseTypeRowTypeGeneratesMissingFieldIds`: ids 
generated in traversal order across nested ROW, ARRAY and MAP descendants.
   - `DataTypeJsonParserTest.ParseTypeRowTypeSuccess`: non-sequential explicit 
ids preserved at every level.
   - `DataTypeJsonParserTest.ParseTypeComplexTypeFailure`: partial ids rejected 
in either order, alongside the malformed complex type cases.
   - `DataFieldTest.FromJson` and `DataFieldTest.FromJsonFailed`: a stored 
schema still requires an id on every field, names and descriptions keep 
embedded NULs, and each malformed field key reports its own diagnostic.
   - `VariantShreddingWritePlanFactoryTest.ConfiguredSchema`: 
`variant.shreddingSchema` and its `parquet.variant.shreddingSchema` fallback 
plan the same physical schema, field ids included, whether the configured 
schema carries ids or omits them.
   - 
`VariantShreddingWritePlanFactoryTest.ConfiguredSchemaRejectsPartialFieldIds`.
   - `VariantParquetTest.ShreddedWriteAndReadRoundTrip`: the new 
`configured-without-ids` mode asserts the generated ids reach the Parquet 
physical schema and that the shredded values round trip.
   
   ### API and Format
   
   `include/paimon/defs.h` only gains documentation for 
`variant.shreddingSchema`; there is no public symbol, storage format or 
protocol change. A configured schema that already carries field ids produces 
exactly the same physical schema as before.
   
   ### Documentation
   
   `docs/source/user_guide/data_types.rst` documents that the configured 
shredding schema's fields either all carry an `id` or all omit one, in which 
case the ids are assigned in traversal order starting at 0.
   
   ### Generative AI tooling
   
   Generated-by: Claude Opus 5
   
   🤖 Generated with [Claude Code](https://claude.com/claude-code)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to