yihua opened a new issue, #20094:
URL: https://github.com/apache/hudi/issues/20094

   Several table services and write utilities ship heavy objects with every 
task or redo table-wide work per task. Clean, marker reconciliation, clustering 
and compaction planning and metadata table file group initialization capture 
the object running the service (with the table, the write config and the meta 
client, 80 to 150 KB per task). Marker reconciliation always runs 
`hoodie.finalize.write.parallelism` tasks, clustering planning reads the 
pending table services of the whole table for every partition, listing-based 
rollback reloads the timeline per partition, and `WriteMarkersFactory` copies 
the Hadoop configuration per written file.
   
   Separately, the bootstrap source listing lists directories on executors with 
a default Hadoop configuration, ignoring the job's settings, and fails for a 
file system only the job configures (`No FileSystem for scheme`).
   
   Proposal: broadcast the state these tasks read instead of capturing the 
service object, size the reconciliation tasks to the work, read the pending 
table services once per plan, and list the bootstrap source with the source 
storage's configuration.
   
   part of #20064
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to