sarkaramrit1993 opened a new pull request, #29450:
URL: https://github.com/apache/flink/pull/29450

   ## What is the purpose of the change
   
   `AvroSerializationSchema` resolves the schema of a specific record with 
`SpecificData.get().getSchema(clazz)`, which reads the static `SCHEMA$` field. 
Classes generated by avrohugger don't have that field (the schema lives on the 
Scala companion object), so serialization fails with `AvroRuntimeException: Not 
a Specific class`.
   
   `AvroDeserializationSchema` already extracts the schema through 
`AvroFactory` since FLINK-18478, but on Avro 1.11 it still fails the same way: 
`SpecificDatumReader.setSchema()` looks the schema up from the class again when 
no expected schema is set.
   
   ## Brief change log
   
   - `AvroSerializationSchema` gets the `SpecificData` and schema through 
`AvroFactory.getSpecificDataForClass` and 
`AvroFactory.extractAvroSpecificSchema`, the same way the deserializer and 
`AvroSerializer` do.
   - `AvroDeserializationSchema` passes the reader schema to 
`SpecificDatumReader` up front.
   - `AvroFactory.getSpecificDataForClass` takes a `Class<?>`. It was declared 
as `Class<T extends SpecificData>` but is always called with record classes, so 
every caller needed an unchecked cast. A separate hotfix commit removes those 
casts.
   
   For regular avro-generated classes nothing changes: the writer picks up the 
same `MODEL$` and `SCHEMA$` as before, and the existing byte-level tests still 
pass.
   
   ## Verifying this change
   
   This change added tests and can be verified as follows:
   
   - Added `NoSchemaFieldRecord`, a test record without static `SCHEMA$` and 
`MODEL$` fields.
   - Added `testSpecificRecordWithoutStaticSchemaField` to 
`AvroSerializationSchemaTest` and `AvroDeserializationSchemaTest`, for binary 
and JSON encoding. Both fail on master with `Not a Specific class` and pass 
with this change.
   - `flink-avro` and `flink-avro-confluent-registry` tests pass.
   
   ## Does this pull request potentially affect one of the following parts:
   
     - Dependencies (does it add or upgrade a dependency): no
     - The public API, i.e., is any changed class annotated with 
`@Public(Evolving)`: no
     - The serializers: yes (Avro (de)serialization schemas, wire format 
unchanged for generated classes)
     - The runtime per-record code paths (performance sensitive): no
     - Anything that affects deployment or recovery: JobManager (and its 
components), Checkpointing, Kubernetes/Yarn, ZooKeeper: no
     - The S3 file system connector: no
   
   ## Documentation
   
     - Does this pull request introduce a new feature? no
     - If yes, how is the feature documented? not applicable
   
   `AvroWriters`, `AvroParquetWriters`, `AvroSchemaConverter`, 
`AvroOutputFormat` and `AvroInputFormat` resolve the schema the same way. I 
kept them out to keep this PR focused and can follow up on them separately.
   
   ---
   
   ##### Was generative AI tooling used to co-author this PR?
   
   - [ ] Yes (please specify the tool below)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to