goutamadwant opened a new issue, #12008:
URL: https://github.com/apache/seatunnel/issues/12008

   ### Search before asking
   
   - [x] I searched the issues and open and closed pull requests and found no 
report or implementation covering this behavior.
   
   ### What happened
   
   Avro serialization fails when a `SeaTunnelRowType` contains a mixed-case 
field name.
   
   `SeaTunnelRowTypeToAvroSchemaConverter` preserves the original field name 
when it creates the Avro schema, but `RowToAvroConverter` lowercases the field 
name before passing it to `GenericRecordBuilder.set(...)`.
   
   For example, a row type containing `CustomerID` produces an Avro schema 
field named `CustomerID`, while serialization attempts to set `customerid`. 
Avro field lookup is case-sensitive, so the builder cannot find the field and 
serialization fails before a record is written.
   
   The same mismatch occurs for fields inside a nested `ROW`.
   
   Expected behavior:
   
   - The writer uses the same field names as the generated Avro schema.
   - Mixed-case fields serialize at both the top level and inside nested rows.
   - Existing lowercase fields continue to work unchanged.
   
   Actual behavior:
   
   - Top-level and nested mixed-case fields fail in 
`GenericRecordBuilder.set(...)`.
   - The failure is in the shared Avro serialization path used by connectors 
including Kafka and Pulsar.
   
   ### SeaTunnel Version
   
   Current `dev` branch at commit `96048a7b81c92d743acc23b5bed652bdfef9d6e5`.
   
   ### SeaTunnel Config
   
   Not configuration-dependent. The issue can be reproduced directly through 
the Avro conversion API with a `SeaTunnelRowType` containing mixed-case field 
names.
   
   ### Running Command
   
   ```shell
   ./mvnw -pl seatunnel-formats/seatunnel-format-avro \
     -Dskip.spotless=true \
     -Dcheckstyle.skip=true \
     -Dlicense.skip=true \
     -Dtest=AvroConverterTest test
   ```
   
   ### Error Exception
   
   With focused top-level and nested regression cases added to 
`AvroConverterTest`, both new cases fail:
   
   ```log
   Tests run: 3, Failures: 0, Errors: 2, Skipped: 0
   
   java.lang.NullPointerException:
   Cannot invoke "org.apache.avro.Schema$Field.pos()" because "field" is null
   ```
   
   The missing field is the lowercased lookup name, which is not present in the 
case-preserving schema.
   
   ### Minimal reproduction
   
   Top-level field:
   
   ```java
   SeaTunnelRowType rowType =
           new SeaTunnelRowType(
                   new String[] {"CustomerID"},
                   new SeaTunnelDataType<?>[] {BasicType.INT_TYPE});
   SeaTunnelRow row = new SeaTunnelRow(1);
   row.setField(0, 42);
   
   new RowToAvroConverter(rowType).convertRowToGenericRecord(row);
   ```
   
   Nested field:
   
   ```java
   SeaTunnelRowType nestedType =
           new SeaTunnelRowType(
                   new String[] {"InnerID"},
                   new SeaTunnelDataType<?>[] {BasicType.INT_TYPE});
   SeaTunnelRowType rowType =
           new SeaTunnelRowType(
                   new String[] {"payload"},
                   new SeaTunnelDataType<?>[] {nestedType});
   
   SeaTunnelRow nestedRow = new SeaTunnelRow(1);
   nestedRow.setField(0, 42);
   SeaTunnelRow row = new SeaTunnelRow(1);
   row.setField(0, nestedRow);
   
   new RowToAvroConverter(rowType).convertRowToGenericRecord(row);
   ```
   
   ### Root cause
   
   - The schema converter passes each original field name to `new 
Schema.Field(...)`: 
[SeaTunnelRowTypeToAvroSchemaConverter.java](https://github.com/apache/seatunnel/blob/96048a7b81c92d743acc23b5bed652bdfef9d6e5/seatunnel-formats/seatunnel-format-avro/src/main/java/org/apache/seatunnel/format/avro/SeaTunnelRowTypeToAvroSchemaConverter.java#L37-L53).
   - The top-level writer calls `builder.set(fieldName.toLowerCase(), ...)`: 
[RowToAvroConverter.java](https://github.com/apache/seatunnel/blob/96048a7b81c92d743acc23b5bed652bdfef9d6e5/seatunnel-formats/seatunnel-format-avro/src/main/java/org/apache/seatunnel/format/avro/RowToAvroConverter.java#L76-L83).
   - The nested `ROW` writer also calls 
`recordBuilder.set(fieldNames[i].toLowerCase(), ...)`: 
[RowToAvroConverter.java](https://github.com/apache/seatunnel/blob/96048a7b81c92d743acc23b5bed652bdfef9d6e5/seatunnel-formats/seatunnel-format-avro/src/main/java/org/apache/seatunnel/format/avro/RowToAvroConverter.java#L139-L153).
   
   The schema and writer therefore disagree whenever a field name contains an 
uppercase character.
   
   ### Impact
   
   The converter is shared by connector serialization paths. Kafka constructs 
`AvroSerializationSchema` in its default row serializer, and Pulsar does the 
same when Avro format is selected:
   
   - [Kafka 
serializer](https://github.com/apache/seatunnel/blob/96048a7b81c92d743acc23b5bed652bdfef9d6e5/seatunnel-connectors-v2/connector-kafka/src/main/java/org/apache/seatunnel/connectors/seatunnel/kafka/serialize/DefaultSeaTunnelRowSerializer.java#L430-L432)
   - [Pulsar sink 
writer](https://github.com/apache/seatunnel/blob/96048a7b81c92d743acc23b5bed652bdfef9d6e5/seatunnel-connectors-v2/connector-pulsar/src/main/java/org/apache/seatunnel/connectors/seatunnel/pulsar/sink/PulsarSinkWriter.java#L276-L287)
   
   Any upstream schema that legitimately preserves names such as `CustomerID`, 
`eventTime`, or nested mixed-case fields can reach this failure.
   
   ### Proposed fix and regression coverage
   
   Use the exact field name in both `GenericRecordBuilder.set(...)` calls 
instead of lowercasing it. Add regression coverage for:
   
   1. A top-level mixed-case field.
   2. A mixed-case field inside a nested `ROW`.
   3. Conversion back to `SeaTunnelRow` to confirm a complete round trip.
   
   As a validation of this direction, removing the two lowercase conversions 
makes both regressions pass. The complete `seatunnel-format-avro` module then 
passes with 6 tests, 0 failures, and 0 errors.
   
   This is backward-compatible for existing lowercase field names. It does not 
change the generated schema, configuration, dependencies, serialized format, or 
public API; it only makes the writer use the schema's actual field names.
   
   ### Zeta or Flink or Spark Version
   
   Not engine-specific.
   
   ### Java or Scala Version
   
   Java 21 was used for reproduction.
   
   ### Screenshots
   
   Not applicable.
   
   ### Are you willing to submit PR?
   
   - [x] Yes I am willing to submit a PR!
   
   ### Code of Conduct
   
   - [x] I agree to follow this project's [Code of 
Conduct](https://www.apache.org/foundation/policies/conduct).
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to