tonymtu opened a new pull request, #10303: URL: https://github.com/apache/paimon/pull/10303
### Purpose Format Tables currently use the table-level file format for every partition. A table configured as Parquet therefore cannot correctly read a catalog-managed partition containing ORC files. This PR supports `file.format` in partition options. The effective format is resolved from the explicit partition option, then the table option, then the existing Parquet default. Supported formats match the table-level formats: ORC, Parquet, CSV, TEXT, JSON and MOSAIC. The resolved format travels with each split and is used for file splitting, reader selection and partition statistics. Spark ANALYZE carries the partition metadata to the statistics collector. Splits serialized without a format continue to inherit the table format. Writers continue to use the table format. Appends reject conflicting partition formats before publishing files. Overwrites update `file.format` only when an existing override differs from the write format; partitions without an override continue to inherit the table configuration. This reuses the existing `partitionOptions` request field. Catalog implementations must apply an explicitly supplied `file.format` to existing partitions while preserving other options; omitting it preserves the stored override. Follow-up to #8750 and #9540. Related to #9710 for partition metadata updates during overwrite. ### Tests - Focused core tests cover format inheritance and validation, mixed-format reads, split serialization and planning, statistics, append rejection, and static, dynamic and empty overwrites. The core build passed with normal Maven checks enabled. - `PaimonFormatTableTest` on Spark 4: 23 tests passed, including mixed-format reads, ANALYZE and writes that preserve inherited formats. - `FormatTableITCase#testPartitionFileFormats` and `FlinkFormatTableDataStreamSinkTest` on Flink 1.20: 7 tests passed. - `FormatTableMosaicReadTest`: 5 tests passed with the native library available. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
