[ 
https://issues.apache.org/jira/browse/FLINK-40262?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

ASF GitHub Bot updated FLINK-40262:
-----------------------------------
    Labels: pull-request-available  (was: )

> Support matching Avro record fields by name in the Avro RowData converters
> --------------------------------------------------------------------------
>
>                 Key: FLINK-40262
>                 URL: https://issues.apache.org/jira/browse/FLINK-40262
>             Project: Flink
>          Issue Type: Improvement
>          Components: Formats (JSON, Avro, Parquet, ORC, SequenceFile)
>            Reporter: Pritam Kumar
>            Priority: Major
>              Labels: pull-request-available
>
> h2. Problem
>   {{RowDataToAvroConverters}} and {{AvroToRowDataConverters}} pair a 
> {{RowType}} field with an Avro record field by *ordinal position*:
>   {code:java}
>   for (int i = 0; i < length; ++i) {
>       final Schema.Field schemaField = fields.get(i);
>       record.put(i, fieldConverters[i].convert(
>               schemaField.schema(), fieldGetters[i].getFieldOrNull(row)));
>   }
>   {code}
>   * a schema registry subject;
>   * the {{avro-confluent.schema}} option;
>   * the {{AvroRowDataSerializationSchema(rowType, nestedSchema, converter)}} 
> and {{AvroRowDataDeserializationSchema(nestedSchema, converter, typeInfo)}} 
> constructors, which downstream connectors use directly.
>   In those cases column {{i}} and Avro field {{i}} need not be the same 
> field. Writing puts each value into whichever field happens to sit at that 
> position, and reading does the mirror image. Where the types at a position 
> differ the job fails
>   with a {{ClassCastException}}; where they happen to agree, the data is 
> silently corrupted.
>   h3. Example
>   Table {{a STRING, b STRING}} against a registry schema that declares {{b}} 
> before {{a}}. Both are strings, so nothing fails: {{a}}'s value is written 
> into {{b}} and vice versa.
>   h2. Proposal
>   Add an explicit strategy and thread it through both converters:
>   {code:java}
>   public enum FieldMatching { INDEX, NAME }
>   {code}
>   * {{INDEX}} — today's behaviour, and the default. Unchanged.
>   * {{NAME}} — pair fields by name. Names are compared exactly first, then 
> against Avro field {{aliases}}, and finally ignoring case ({{Locale.ROOT}}). 
> Each stage runs to completion before the next, so a fuzzy match can never 
> claim an Avro field
>   that another column matches exactly.
>   Every existing overload is kept and delegates with {{INDEX}}, so the change 
> is binary compatible (japicmp clean).
>   h3. What NAME refuses
>   Name matching should not paper over a schema that does not line up. 
> Resolution fails, naming the offending field, when:
>   * a column matches more than one Avro field;
>   * two columns match the same Avro field;
>   * a column has no Avro counterpart and we are writing — the column would be 
> dropped;
>   * a {{NOT NULL}} column has no Avro counterpart and we are reading — it 
> could only be read as {{NULL}};
>   * an Avro field that no column maps to is neither nullable nor declares a 
> default — the record cannot be encoded at all, and Avro surfaces that as a 
> bare {{NullPointerException}}.
>   Tolerated:
>   * an Avro field that no column maps to but that declares a default is 
> written with that default;
>   * an Avro field that no column reads is ignored, which is ordinary 
> projection.
>   h3. Cost
>   Resolution runs once per (row type, schema) pair and is memoized, so 
> {{NAME}} costs one array lookup per field per record: no hashing, no 
> lowercasing, no allocation on the hot path. The memo is {{transient}} and 
> rebuilt after the converter is
>   shipped to a task.
>   h2. Scope
>   This issue covers the converters only, which makes the capability reachable 
> from Java but not from SQL.
>   Exposing the strategy as an {{avro-confluent.field-matching}} option is 
> deliberately left to a follow-up so that the two can be reviewed 
> independently. That part also has to relax the order sensitive schema check in
>   {{RegistryAvroFormatFactory#getAvroSchema}}, which compares the converted 
> {{DataType}} with {{equals()}} and therefore rejects a reordered schema 
> during planning with _"Schema provided for 'avro-confluent' format does not 
> match the table
>   schema"_ — so today such a schema never reaches the converters at all. 
> Happy to fold that into this issue instead if reviewers would rather have it 
> in one go.
>   The plain {{avro}} format needs no option: it always derives its schema 
> from the table, so field order agrees by construction.
>   h2. Compatibility
>   {{INDEX}} is the default everywhere, no existing signature or default value 
> changes, and japicmp reports no incompatibility.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to