juntaozhang opened a new issue, #10302:
URL: https://github.com/apache/paimon/issues/10302

   ### Search before asking
   
   - [x] I searched in the [issues](https://github.com/apache/paimon/issues) 
and found nothing similar.
   
   
   ### Motivation
   
      Paimon can shred a Variant column into many Parquet sub-columns:          
                                                                                
   
                                                                                
                                                                                
   
      ```                                                                       
                                                                                
   
        v.metadata                                                              
                                                                                
   
        v.value                                                                 
                                                                                
   
        v.typed_value.a                                                         
                                                                                
   
        v.typed_value.b                                                         
                                                                                
   
        v.typed_value.c                                                         
                                                                                
   
        ...                                                                     
                                                                                
   
      ```                                                                       
                                                                                
   
                                                                                
                                                                                
   
      A query may only need $.a. Today Python reads the whole v column and 
extracts the path in memory.                                                    
        
                                                                                
                                                                                
   
      With this feature, the reader skips unused sub-columns at IO time.        
                                                                                
   
                                                                                
                                                                                
   
      Before vs After (pseudocode)                                              
                                                                                
   
                                                                                
                                                                                
   
      ```python                                                                 
                                                                                
   
        # User query                                                            
                                                                                
   
        read_fields = [RowType("v", ["a" : INT with metadata "$.a"])]           
                                                                                
   
                                                                                
                                                                                
   
        # BEFORE: reads everything                                              
                                                                                
   
        columns = ["v"]                    # all sub-columns                    
                                                                                
   
        batch = parquet.read(columns)                                           
                                                                                
   
        a = variant_get(batch.v, "$.a")    # extract after full read            
                                                                                
   
                                                                                
                                                                                
   
        # AFTER: reads only what is needed                                      
                                                                                
   
        columns = ["v.metadata", "v.typed_value.a"]                             
                                                                                
   
        batch = parquet.read(columns)                                           
                                                                                
   
        a = assemble_projection(batch)     # reader returns {"a": ...} directly 
                                                                                
   
      ```  
   
   ### Solution
   
   Make FormatPyArrowReader compute the minimal column set from a Variant 
projection RowType, so queries that touch only a few Variant fields do not pull 
the  entire shredded structure from disk.  
   
      1. Refactor Variant path segments to ObjectExtraction / ArrayExtraction 
dataclasses.                                                                    
     
      2. Add variant_metadata parser and variant_shredding_pruner library.      
                                                                                
   
      3. Integrate pruning into FormatPyArrowReader with projection assembly.   
   
   ### Anything else?
   
   _No response_
   
   ### Are you willing to submit a PR?
   
   - [x] I'm willing to submit a PR!


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to