yihua opened a new issue, #20101:
URL: https://github.com/apache/hudi/issues/20101

   For every parquet base file, 
`HoodieParquetFileFormatHelper.buildImplicitSchemaChangeInfo` converts the 
whole footer schema to a Spark schema and compares it with the requested 
schema, matching nested fields by name at every level, to find implicit type 
changes. The conversion covers every column of the file even when the query 
reads a few. The files of a table share a few schemas, so scans over many files 
repeat the same work.
   
   Proposal: cache the result per executor, keyed by everything the computation 
reads (the file schema, the requested schema and the relevant conf values).
   
   part of #20064
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to