nabassa1323 opened a new issue, #14172:
URL: https://github.com/apache/cloudstack/issues/14172

   ### problem
   
   With Database High Availability enabled (db.ha.enabled=true), after the 
MySQL source becomes unreachable the Management Server UI becomes extremely 
slow / barely usable.
   
   Failover to the configured replica seems to happen via the MySQL connector, 
but each (or almost each) request appears to retry the dead source first, 
causing severe UI latency.
   
   Workaround that restores normal performance:
   - disable db.ha.enabled
   - manually promote the replica (STOP REPLICA / RESET REPLICA ALL / 
read_only=OFF)
   - point db.cloud.host / db.usage.host to the promoted DB in db.properties
   - restart cloudstack-management
   
   This matches the install guide "Failover" procedure better than the 
client-side db.ha mechanism.
   
   Note: docs still say "Tested with MySQL 5.1 and 5.5" and suggest two-way 
replication.
   Ref: 
https://docs.cloudstack.apache.org/en/4.22.1.0/adminguide/reliability.html#configuring-database-high-availability
   
   ### versions
   
   CloudStack: 4.22.1.0 (Ubuntu packages, download.cloudstack.org noble)
   OS: Ubuntu 24.04
   MySQL: 8.0.46 (async replication source→replica, GTID)
   Hypervisor: VMware
   Topology: 2 Management Servers (multi-site), reverse proxy in front of UI
   
   ### The steps to reproduce the bug
   
   1. Deploy 2 MS pointing to MySQL source; configure two-way replica
   2. Set in db.properties:
      db.ha.enabled=true
      db.cloud.replicas=<replica-ip>
      db.usage.replicas=<replica-ip>
   3. Restart cloudstack-management on both nodes
   4. Stop / isolate the MySQL source (simulate site/DB failure)
   5. Use UI/API through the remaining Management Server
   
   ### What to do about it?
   
   Expected: MS fail over to replica promptly; UI remains usable.
   
   Actual: UI becomes very slow after source outage (connector retries against 
unreachable source; defaults 
secondsBeforeRetrySource/queriesBeforeRetrySource/initialTimeout = 
3600/5000/3600).
   
   Suggestions:
   - fix connector failover so replica is used without per-request source 
retries hanging the UI
   - and/or document clearly that db.ha.enabled is not recommended for 
production on MySQL 8.x and that manual promotion (install guide Failover) is 
the supported path


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to