juntaozhang opened a new issue, #10302: URL: https://github.com/apache/paimon/issues/10302
### Search before asking - [x] I searched in the [issues](https://github.com/apache/paimon/issues) and found nothing similar. ### Motivation Paimon can shred a Variant column into many Parquet sub-columns: ``` v.metadata v.value v.typed_value.a v.typed_value.b v.typed_value.c ... ``` A query may only need $.a. Today Python reads the whole v column and extracts the path in memory. With this feature, the reader skips unused sub-columns at IO time. Before vs After (pseudocode) ```python # User query read_fields = [RowType("v", ["a" : INT with metadata "$.a"])] # BEFORE: reads everything columns = ["v"] # all sub-columns batch = parquet.read(columns) a = variant_get(batch.v, "$.a") # extract after full read # AFTER: reads only what is needed columns = ["v.metadata", "v.typed_value.a"] batch = parquet.read(columns) a = assemble_projection(batch) # reader returns {"a": ...} directly ``` ### Solution Make FormatPyArrowReader compute the minimal column set from a Variant projection RowType, so queries that touch only a few Variant fields do not pull the entire shredded structure from disk. 1. Refactor Variant path segments to ObjectExtraction / ArrayExtraction dataclasses. 2. Add variant_metadata parser and variant_shredding_pruner library. 3. Integrate pruning into FormatPyArrowReader with projection assembly. ### Anything else? _No response_ ### Are you willing to submit a PR? - [x] I'm willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
