zjw1111 opened a new issue, #321: URL: https://github.com/apache/paimon-cpp/issues/321
### Problem `ParquetFileBatchReader` currently relies on `ParquetTimestampConverter`, which mixes file-schema normalization, compatibility checking, and array conversion. The reader also recursively inspects the schema for every batch. Rebuilding nested arrays from inferred child types can lose requested field metadata, nullability, and MAP attributes. The compatibility fixtures also group `ARRAY<BLOB>` and `MAP<K, BLOB>` together even though Paimon C++ rejects the former and supports the latter only when it uses the standard Paimon BLOB storage. In particular, the Rust fixture stores raw BLOB bytes inline as Parquet `binary`; Java cannot interpret those bytes as a Paimon `BlobDescriptor`, so C++ should not add a special conversion for that representation. ### Proposed changes - Replace `ParquetTimestampConverter` with a reusable read-type adaptation plan built once per read schema. - Distinguish representation-compatible fields from timestamp casts and timezone views. - Preserve target field metadata, nullability, and nested container attributes while rebuilding STRUCT, LIST, and MAP arrays. - Keep the existing millisecond-to-second timestamp conversion and timezone normalization behavior. - Split the compatibility fixtures into explicit ARRAY BLOB and MAP BLOB tables. - Verify Python and Java standard `MAP<K, BLOB>` tables can be read, and explicitly reject the Rust raw-binary representation. - Update BLOB placement errors to mention the supported direct value of a top-level MAP field. ### Acceptance criteria - The type adaptation plan is reused across batches. - Existing timestamp compatibility tests continue to pass. - Nested target field attributes are preserved. - Python and Java MAP BLOB fixtures return the expected values. - Rust inline raw-binary MAP BLOB and `ARRAY<BLOB>` fixtures return explicit errors. - No public API or storage format is changed. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
