Despite tests and maintenance work on the host, cloudvirt1071 just crashed again. This time, the VMs that rebooted as a result were:

orchestrator-1.mobileappsperformance
osmit-estratti-trixie.osmit
pki-db-1.pki
pki-root-1.pki
syslog-client06-new.auditlogging
tools-opensearch-1.tools
tools-k8s-haproxy-7.tools

All affected VMs are now back up and moved to different hardware.

-Andrew


On 7/6/26 11:37 PM, Andrew Bogott wrote:
One of the cloud-vps compute servers, cloudvirt1071, crashed a few minutes ago. I have now restarted it and moved all VMs to other hosts, but a small number of VMs suffered 15 or so minutes of downtime followed by a reboot. Those VMs are:

kube-2js6t-default-worker-5ff26-rg5n9-nxtmp.zuul
kube-2js6t-default-worker-5ff26-rg5n9-r722f.zuul
k3s.catalyst-dev
mediawiki2latex.collection-alt-renderer
toolsbeta-test-k8s-worker-nfs-8.toolsbeta
tools-k8s-worker-nfs-44.tools
tools-k8s-worker-nfs-32.tools

No data was lost, and no action should be needed; this is just informational for anyone who noticed the downtime. Toolforge tools should have been unaffected, although some pods were likely rescheduled after a few seconds of hesitation.

Any followup on this incident will be tracked on https://phabricator.wikimedia.org/T431374

-Andrew



_______________________________________________
Cloud-announce mailing list -- [email protected]
List information: 
https://lists.wikimedia.org/postorius/lists/cloud-announce.lists.wikimedia.org/

Reply via email to