SEZ9 commented on issue #11663: URL: https://github.com/apache/seatunnel/issues/11663#issuecomment-5215603453
This is a substantial design effort — could you draft a design doc (SEIP-style) covering architecture (where the autoscaler loop runs, e.g., master node), metric collection, decision policy, and configuration surface? That would help move the discussion forward. Please clarify the initial target environment: will the first iteration only emit scaling recommendations (metrics/REST) for external actuators like Kubernetes HPA/operator, or also actively provision/terminate workers? Actuation differs significantly between K8s and bare-metal/standalone deployments. Scale-in safety needs careful design: how will the autoscaler drain a worker without disrupting running jobs? Please describe interaction with checkpointing/task restoration, and whether slot decommissioning (blocking new allocations on a draining worker) will be a prerequisite. <!-- streview-comment:65 --> -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
