XiaoHongbo-Hope opened a new pull request, #10398: URL: https://github.com/apache/paimon/pull/10398
### Purpose Wide named Variant projections repeatedly construct the output schema and extract Struct children for every Arrow batch. This adds Python CPU overhead even when native decoding is already cheap. ### Changes - Apply named projection once after `to_arrow()` assembles the physical columns, without mutating the reader. - Reuse the expression output schema and flatten each Struct source once for streaming batches. - Preserve aliases, ordering, parent/child nulls, slices and row kinds. No API changes. ### Validation - 195 tests and 22 subtests passed across native reads, Variant and nested projection tests. - Added table/streaming equivalence, empty/sliced/null inputs, repeated reads, schema reuse and value-buffer sharing tests. - Full read-only workload A/B is pending; no end-to-end speedup claimed yet. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
