yihua opened a new issue, #20094: URL: https://github.com/apache/hudi/issues/20094
Several table services and write utilities ship heavy objects with every task or redo table-wide work per task. Clean, marker reconciliation, clustering and compaction planning and metadata table file group initialization capture the object running the service (with the table, the write config and the meta client, 80 to 150 KB per task). Marker reconciliation always runs `hoodie.finalize.write.parallelism` tasks, clustering planning reads the pending table services of the whole table for every partition, listing-based rollback reloads the timeline per partition, and `WriteMarkersFactory` copies the Hadoop configuration per written file. Separately, the bootstrap source listing lists directories on executors with a default Hadoop configuration, ignoring the job's settings, and fails for a file system only the job configures (`No FileSystem for scheme`). Proposal: broadcast the state these tasks read instead of capturing the service object, size the reconciliation tasks to the work, read the pending table services once per plan, and list the bootstrap source with the source storage's configuration. part of #20064 -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
