GitHub user luganofer edited a comment on the discussion: Additional Zone vs
Region - US East and West
Hi @daviftorres
came across your post regarding the CloudStack management server architecture
and had a couple of specific questions about the failover behavior you
described:
You mentioned:
"1st primarily responds to UI/API. If unhealthy, it fails over to 3rd, 4th,
and 2nd."
Could you elaborate on how this failover is actually implemented?
Are you using a load balancer (e.g., HAProxy, NGINX) with active health
checks and a defined priority/weight order?
Or is this handled at the application level (e.g., CloudStack's own
internal failover logic, or a custom script)?
How do you detect "unhealthy" – is it based on TCP port checks, HTTP status
codes, or something more application-aware (e.g., API responsiveness, DB
connectivity)?
Similarly, for the agents:
"2nd first priority for Agents to connect. Followed by 4th, 3rd, and 1st."
How do you enforce this priority ordering for agent connections?
Is this configured via the agent.properties file (e.g., host list with
order), or do you use DNS with a prioritized record set, or something like a
service discovery tool?
Do agents automatically retry the next in the list if the primary is
unreachable, and if so, how is that handled in practice (timeouts, fallback
logic)?
I'm particularly interested in the operational details – e.g., what stack/tools
you're using, how you handle stateful connections, and whether you've
encountered any pitfalls with this priority-based failover approach.
GitHub link:
https://github.com/apache/cloudstack/discussions/12115#discussioncomment-17976707
----
This is an automatically sent email for [email protected].
To unsubscribe, please send an email to: [email protected]