Looks good to me. Thanks for the updates!

On Tue, Sep 29, 2026, at 13:31, Dewey Dunnington wrote:
> Hi Kent,
>
> Thank you for putting this revision together. I left a few minor comments
> in the document but in general this looks great and I am looking forward to
> implementing it and/or using it!
>
> Cheers,
>
> -dewey
>
> On Thu, Sep 17, 2026 at 10:06 PM Kent Wu <[email protected]> wrote:
>
>> Apologies for going quiet on this, and thank you Dewey, David, Rusty,
>> and Matt for the feedback.
>>
>> I incorporated some of the feedback into the doc [1], which I summarized
>> under draft-02.
>>
>> David:
>> > But I think "simple" string types is reasonable for unparameterized
>> > types.
>>
>> Agreed. I updated the doc with a proposed mapping of types represented
>> by simple string form vs object form. The mapping is unambiguous, so
>> each type has only one allowed representation (a type cannot have both a
>> valid simple string type and an object type).
>>
>> Matt:
>> > I wouldn't object to making metadata an object
>>
>> I agree, and incorporated it into the proposal. Surveying the ecosystem,
>> I found that many implementations expose metadata as a map, which cannot
>> represent duplicate keys. The object form aligns well with that.
>>
>> David:
>> > raises too many questions about what happens to the canonical
>> > representation/equality/serialization/round-tripping
>>
>> Most of these questions already exist and aren't specified by Arrow.
>> Equality and round-tripping for metadata aren't defined anywhere today,
>> and the escaped-JSON-in-JSON case is currently handled by
>> domain-specific convention at the application layer.
>>
>> After considering this and Dewey's feedback, I don't think this version
>> needs to define equality at all, either over documents or over schemas,
>> and I'm in agreement that the format should have an answer for
>> JSON-valued metadata.
>>
>> A specification of the mapping and a conformance suite are sufficient.
>> Losslessness is a property of a reader and writer pair, certified by the
>> suite. The suite has provenance the schemas themselves don't, since it
>> knows which metadata values a fixture wrote as a non-string, so it can
>> compare those as JSON and everything else byte-for-byte. Bindings don't
>> need to make that comparison - only the suite does, once, to certify the
>> pair round-trips a schema.
>>
>> To be clear about scope: this specifies the mapping between a document
>> and an Arrow schema, not the behavior of Arrow implementations. "reader"
>> and "writer" below mean the code performing that conversion in either
>> direction. What an implementation does with the metadata afterwards is
>> out of scope.
>>
>> This is the language I've proposed in the doc:
>>
>>     A metadata value MAY be any JSON value. A reader MUST map a string
>>     value to its contents. A reader MUST map a value of any other type
>>     (object, array, number, boolean, or null) to its serialized string
>>     form, for which the RFC 8785 serialization SHOULD be used. A writer
>>     MAY emit a string or non-string value.
>>
>> Nothing here requires an implementation to detect which metadata values
>> are JSON, or to change its extension APIs. A writer that only emits
>> strings is fully conformant.
>>
>> I've left RFC 8785 [2] as a recommendation rather than a requirement.
>> The conformance requirement is that both forms map to the same Arrow
>> metadata, which the suite checks structurally.
>>
>> I'm not opposed to making it a requirement. I just don't have a good
>> sense of how much of a burden an RFC 8785 implementation is across
>> Arrow target languages.
>>
>> [1]:
>> https://docs.google.com/document/d/1ho0FKy9ge0tUSRzebq1AFi28KR1H_utecQjHx4CNCEs/edit?usp=sharing
>> [2]: https://www.rfc-editor.org/info/rfc8785/
>>

Reply via email to