[
https://issues.apache.org/jira/browse/FLINK-40262?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated FLINK-40262:
-----------------------------------
Labels: pull-request-available (was: )
> Support matching Avro record fields by name in the Avro RowData converters
> --------------------------------------------------------------------------
>
> Key: FLINK-40262
> URL: https://issues.apache.org/jira/browse/FLINK-40262
> Project: Flink
> Issue Type: Improvement
> Components: Formats (JSON, Avro, Parquet, ORC, SequenceFile)
> Reporter: Pritam Kumar
> Priority: Major
> Labels: pull-request-available
>
> h2. Problem
> {{RowDataToAvroConverters}} and {{AvroToRowDataConverters}} pair a
> {{RowType}} field with an Avro record field by *ordinal position*:
> {code:java}
> for (int i = 0; i < length; ++i) {
> final Schema.Field schemaField = fields.get(i);
> record.put(i, fieldConverters[i].convert(
> schemaField.schema(), fieldGetters[i].getFieldOrNull(row)));
> }
> {code}
> * a schema registry subject;
> * the {{avro-confluent.schema}} option;
> * the {{AvroRowDataSerializationSchema(rowType, nestedSchema, converter)}}
> and {{AvroRowDataDeserializationSchema(nestedSchema, converter, typeInfo)}}
> constructors, which downstream connectors use directly.
> In those cases column {{i}} and Avro field {{i}} need not be the same
> field. Writing puts each value into whichever field happens to sit at that
> position, and reading does the mirror image. Where the types at a position
> differ the job fails
> with a {{ClassCastException}}; where they happen to agree, the data is
> silently corrupted.
> h3. Example
> Table {{a STRING, b STRING}} against a registry schema that declares {{b}}
> before {{a}}. Both are strings, so nothing fails: {{a}}'s value is written
> into {{b}} and vice versa.
> h2. Proposal
> Add an explicit strategy and thread it through both converters:
> {code:java}
> public enum FieldMatching { INDEX, NAME }
> {code}
> * {{INDEX}} — today's behaviour, and the default. Unchanged.
> * {{NAME}} — pair fields by name. Names are compared exactly first, then
> against Avro field {{aliases}}, and finally ignoring case ({{Locale.ROOT}}).
> Each stage runs to completion before the next, so a fuzzy match can never
> claim an Avro field
> that another column matches exactly.
> Every existing overload is kept and delegates with {{INDEX}}, so the change
> is binary compatible (japicmp clean).
> h3. What NAME refuses
> Name matching should not paper over a schema that does not line up.
> Resolution fails, naming the offending field, when:
> * a column matches more than one Avro field;
> * two columns match the same Avro field;
> * a column has no Avro counterpart and we are writing — the column would be
> dropped;
> * a {{NOT NULL}} column has no Avro counterpart and we are reading — it
> could only be read as {{NULL}};
> * an Avro field that no column maps to is neither nullable nor declares a
> default — the record cannot be encoded at all, and Avro surfaces that as a
> bare {{NullPointerException}}.
> Tolerated:
> * an Avro field that no column maps to but that declares a default is
> written with that default;
> * an Avro field that no column reads is ignored, which is ordinary
> projection.
> h3. Cost
> Resolution runs once per (row type, schema) pair and is memoized, so
> {{NAME}} costs one array lookup per field per record: no hashing, no
> lowercasing, no allocation on the hot path. The memo is {{transient}} and
> rebuilt after the converter is
> shipped to a task.
> h2. Scope
> This issue covers the converters only, which makes the capability reachable
> from Java but not from SQL.
> Exposing the strategy as an {{avro-confluent.field-matching}} option is
> deliberately left to a follow-up so that the two can be reviewed
> independently. That part also has to relax the order sensitive schema check in
> {{RegistryAvroFormatFactory#getAvroSchema}}, which compares the converted
> {{DataType}} with {{equals()}} and therefore rejects a reordered schema
> during planning with _"Schema provided for 'avro-confluent' format does not
> match the table
> schema"_ — so today such a schema never reaches the converters at all.
> Happy to fold that into this issue instead if reviewers would rather have it
> in one go.
> The plain {{avro}} format needs no option: it always derives its schema
> from the table, so field order agrees by construction.
> h2. Compatibility
> {{INDEX}} is the default everywhere, no existing signature or default value
> changes, and japicmp reports no incompatibility.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)