JingsongLi commented on PR #10336:
URL: https://github.com/apache/paimon/pull/10336#issuecomment-5992021991

   Please complete the internal refactor around `read_type` in this PR. The 
reader projection should have one canonical representation, following [Java's 
`ReadBuilderImpl`](https://github.com/apache/paimon/blob/414520cb55cf0936cd0e6c8d33826ffe16e8609a/paimon-core/src/main/java/org/apache/paimon/table/source/ReadBuilderImpl.java#L126),
 where `withProjection` resolves the requested fields and delegates to 
`withReadType`.
   
   The Python implementation still persists `_variant_fields` alongside the 
source projection and passes it through `TableRead`, the native adapter, stream 
reads, split providers, and Ray. It then separately reconstructs the native 
read type and Arrow schema from that dictionary. This leaves several layers 
responsible for keeping the same extraction request consistent.
   
   Please make `with_projection` compile the reader request once and store the 
resolved `read_type` in `ReadBuilder`:
   
   - Resolve ordinary and nested column selections into the reader's field 
types.
   - Encode Variant extraction paths, target types, and per-extraction error 
policies using the existing Variant `DataField` metadata in the read type.
   - Pass that same read type through batch, stream, and Ray reads to native 
`with_read_type`. Derive the reader's Arrow schema and native-only capability 
checks from it.
   - Remove the persistent `variant_fields` state and its cross-layer 
parameters. Any source-column paths needed by a reader adapter should be 
derived from the read type and table schema.
   
   Keep the final output projection as a separate layer: aliases, flattening, 
duplicate outputs, and output order are SQL projection semantics that the 
standard Paimon reader `read_type` does not fully represent. That layer can map 
the reader's result to the requested flat columns without carrying a second 
Variant extraction specification. This distinction also applies to operations 
such as MAP-key lookup when they are not representable in the reader type.
   
   Please cover the refactor with tests verifying the same resolved read type 
across batch/stream/Ray, the existing named-output semantics, and replacement 
of a projection on the same builder without stale extraction state. The goal is 
to finish the read-type-based implementation throughout the read path.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to