marin-ma opened a new pull request, #13109:
URL: https://github.com/apache/gluten/pull/13109

   Velox's Hive connector can declare required subfields on a column handle, 
and its Parquet map reader then materialises only the entries whose key 
matches, so that reading m['k'] no longer decodes every entry of every row's 
map. This change lets Gluten derive those declarations for map columns read by 
a native Parquet scan.
   
   Planner rule ScanMapKeyPruning (Velox backend, post-transform, off by 
default, spark.gluten.sql.columnar.backend.velox.scanMapKeyPruningEnabled): for 
each map-typed output of a native Parquet scan it walks up through Filter, 
Project and Generate nodes only. Every reference to the map on the way must be 
a constant-key path (GetMapValue, ElementAt on a map, GetStructField steps) or 
a null check, and some Project or Generate must drop the attribute, and every 
alias still holding map data, before the chain ends. Since an operator can only 
reference what its child outputs, nothing above that point can read the map and 
needs no analysis. Any other shape leaves the map whole: the map still live at 
an exchange, join, aggregate, union, write or the fragment output, a whole-map 
use such as size(m) or explode(m), a non-constant key, a key type without a 
Velox subscript form (date, decimal), or a string key whose bytes are not valid 
UTF-8. Null checks need no map entries and are declared only
  when no value path covers them; a declaration is emitted only when it lets 
the reader skip something. Spark's NestedColumnAliasing alias of a struct field 
holding the map (c.m AS _extract_m) is followed with a path prefix.
   
   Transport and native side: paths travel as a structured protobuf 
(RequiredSubfieldsExtension: column, then field / string key / long key 
elements) in the ReadRel advanced extension and are rebuilt as common::Subfield 
on the HiveColumnHandle, with case folded like the schema's names; declared 
columns that match no scan column are logged.
   
   Spark's GetMapValue is emitted as get_map_value, a Gluten-registered 
map-only subscript that reports canPushdown(), so a remaining filter such as 
m['k'].x = v extracts m["k"] instead of the whole map and does not defeat the 
declaration. Delta scans participate when the table uses no column mapping; 
BatchScan and FileSourceScan transformers carry the declaration and include it 
in their equality.
   
   Tests: ScanMapKeyPruningSuite runs each case with adaptive execution off and 
on and asserts the exact declaration of every native scan, including a 
direct-declaration test that the reader returns only the declared key; Delta 
and Hive UDF suites cover column mapping and partial generates.
   
   
   ## Was this patch authored or co-authored using generative AI tooling?
   
   Claude Fable 5.1
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to