XiaoHongbo-Hope opened a new pull request, #10398:
URL: https://github.com/apache/paimon/pull/10398

   ### Purpose
   
   Wide named Variant projections repeatedly construct the output schema and 
extract Struct children for every Arrow batch. This adds Python CPU overhead 
even when native decoding is already cheap.
   
   ### Changes
   
   - Apply named projection once after `to_arrow()` assembles the physical 
columns, without mutating the reader.
   - Reuse the expression output schema and flatten each Struct source once for 
streaming batches.
   - Preserve aliases, ordering, parent/child nulls, slices and row kinds. No 
API changes.
   
   ### Validation
   
   - 195 tests and 22 subtests passed across native reads, Variant and nested 
projection tests.
   - Added table/streaming equivalence, empty/sliced/null inputs, repeated 
reads, schema reuse and value-buffer sharing tests.
   - Full read-only workload A/B is pending; no end-to-end speedup claimed yet.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to