yihua opened a new issue, #20100: URL: https://github.com/apache/hudi/issues/20100
For every parquet base file, the Spark readers copy the Hadoop conf three times (in `SparkParquetReaderBase.read`, in `ParquetSchemaEvolutionUtils`, and in `TaskAttemptContextImpl`, because the copy is not a `JobConf`) and render the requested schema to JSON three times, although the keys they set depend only on the requested schema. Every task also converts the table schema to parquet with a new default-loaded `Configuration` for logical timestamp repair, even for tables without a timestamp-millis field, where repair is off. With a large Hadoop conf and a wide schema this is a noticeable share of the CPU of small or selective file reads. Proposal: prepare the requested-schema read conf once per scan, copy it per file only when the file needs keys of its own, and convert the table schema only when timestamp repair is on. part of #20064 -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
