JingsongLi commented on PR #10336: URL: https://github.com/apache/paimon/pull/10336#issuecomment-5992021991
Please complete the internal refactor around `read_type` in this PR. The reader projection should have one canonical representation, following [Java's `ReadBuilderImpl`](https://github.com/apache/paimon/blob/414520cb55cf0936cd0e6c8d33826ffe16e8609a/paimon-core/src/main/java/org/apache/paimon/table/source/ReadBuilderImpl.java#L126), where `withProjection` resolves the requested fields and delegates to `withReadType`. The Python implementation still persists `_variant_fields` alongside the source projection and passes it through `TableRead`, the native adapter, stream reads, split providers, and Ray. It then separately reconstructs the native read type and Arrow schema from that dictionary. This leaves several layers responsible for keeping the same extraction request consistent. Please make `with_projection` compile the reader request once and store the resolved `read_type` in `ReadBuilder`: - Resolve ordinary and nested column selections into the reader's field types. - Encode Variant extraction paths, target types, and per-extraction error policies using the existing Variant `DataField` metadata in the read type. - Pass that same read type through batch, stream, and Ray reads to native `with_read_type`. Derive the reader's Arrow schema and native-only capability checks from it. - Remove the persistent `variant_fields` state and its cross-layer parameters. Any source-column paths needed by a reader adapter should be derived from the read type and table schema. Keep the final output projection as a separate layer: aliases, flattening, duplicate outputs, and output order are SQL projection semantics that the standard Paimon reader `read_type` does not fully represent. That layer can map the reader's result to the requested flat columns without carrying a second Variant extraction specification. This distinction also applies to operations such as MAP-key lookup when they are not representable in the reader type. Please cover the refactor with tests verifying the same resolved read type across batch/stream/Ray, the existing named-output semantics, and replacement of a projection on the same builder without stale extraction state. The goal is to finish the read-type-based implementation throughout the read path. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
