VishnuVuggepalli opened a new issue, #14185:
URL: https://github.com/apache/cloudstack/issues/14185
##### ISSUE TYPE
* Bug Report
##### COMPONENT NAME
~~~
VM Autoscale Feature
~~~
##### CLOUDSTACK VERSION
~~~
4.22.0.0 (observed)
Also present in 4.22.1.0, 4.23.0.0 and main (verified by source inspection)
~~~
##### CONFIGURATION
Advanced zone, KVM, Ceph/RBD-only primary storage. One AutoScale VM group
(`min_members=1`, `max_members=2`, `interval=30`) on an isolated network
with a
/24 guest CIDR.
##### OS / ENVIRONMENT
Linux (Debian 12), KVM hosts.
##### SUMMARY
The infinite-autoscaling guard added in #11244 (fixing #9318) only counts
instances in `State.Error`. When scale-up VMs fail to **start** — as opposed
to
failing to be **created** — they land in `State.Stopped`, not `State.Error`.
The
guard therefore never trips, and the group scales up on every interval
indefinitely.
In our incident this produced **2,296 VMs from a group whose `max_members`
is 2**,
over roughly 21 hours, until the guest subnet was exhausted.
Two independent counters are involved and both exclude `Stopped`:
1. `AutoScaleVmGroupVmMapDaoImpl.getErroredInstanceCount()` — the #11244
guard —
counts `State.Error` only:
```java
public int getErroredInstanceCount(long vmGroupId) {
SearchCriteria<Integer> sc = CountBy.create();
sc.setParameters("vmGroupId", vmGroupId);
sc.setJoinParameters("vmSearch", "states", State.Error); // Stopped
not counted
...
}
```
2. `AutoScaleVmGroupVmMapDaoImpl.countAvailableVmsByGroup()` — used by every
scaling decision in `AutoScaleManagerImpl` (`checkConditionUp`,
`checkConditionDown`, `checkAutoScaleVmGroup`, and the group-state
handlers) —
counts only `Starting`, `Running`, `Stopping`, `Migrating`:
```java
sc.setJoinParameters("vmSearch", "states",
State.Starting, State.Running, State.Stopping, State.Migrating); //
Stopped not counted
```
So with N leaked `Stopped` members, `currentVM == 0`:
* `checkAutoScaleVmGroup`: `if (currentVM < minMembers)` -> `0 < 1` -> scale
up, every interval
* `checkAutoScaleVmGroup`: `if (currentVM > maxMembers)` -> `0 > 2` ->
scale-down never fires
* `checkConditionDown`: `if (currentVM - 1 < minVm)` -> `-1 < 1` ->
scale-down additionally blocked
* `checkConditionUp`: errored-instance guard -> `0 > 10` false -> guard
never trips
A third defect prevents the failed VM from being cleaned up, which is what
allows
the leak to accumulate in the first place. `doScaleUp` persists the group
map row
*before* attempting the start, and its cleanup is guarded on
`ServerApiException`:
```java
// AutoScaleManagerImpl.doScaleUp
autoScaleVmGroupVmMapDao.persist(groupVmMapVO); // persisted BEFORE the
start attempt
try {
startNewVM(vm.getId());
...
} catch (ServerApiException e) {
...
destroyVm(vm.getId()); // never reached, see below
break;
}
```
`startNewVM` does convert `InsufficientCapacityException` into
`ServerApiException`,
but it never sees that exception, because `VirtualMachineManagerImpl.start()`
has already wrapped it into an unchecked `CloudRuntimeException`:
```java
try {
advanceStart(vmUuid, params, planToDeploy, planner);
} catch (ConcurrentOperationException | InsufficientCapacityException e) {
throw new CloudRuntimeException(String.format("Unable to start a VM [%s]
due to [%s].", vmUuid, e.getMessage()), e);
}
```
`CloudRuntimeException` matches none of `startNewVM`'s typed catches and is
not a
`ServerApiException`, so it propagates past `doScaleUp`'s handler to
`AutoScaleManagerImpl$MonitorTask`, and `destroyVm()` is never called. The
observed log line is exactly this:
```
WARN [c.c.n.a.A.MonitorTask] Caught the following exception on monitoring
AutoScale Vm Group
com.cloud.utils.exception.CloudRuntimeException: Unable to start a VM
[...]
```
Note that PR #9574 ("Prevent infinite retries of autoscaling"), which
proposed a
one-line change to `AutoScaleVmGroupVmMapDaoImpl`, was closed unmerged; the
merged
#11244 took the threshold approach instead, which is what leaves this variant
uncovered.
##### STEPS TO REPRODUCE
1. Create an AutoScale VM group (`min_members=1`, `max_members=2`, short
interval).
2. Let it stabilise at 1 running VM.
3. Break VM **start** (not creation) in a way that returns an
`InsufficientCapacityException` or `ResourceUnavailableException` from
`advanceStart`.
The trigger we actually hit was the group network's Virtual Router
becoming
unreachable, so `VirtualRouterElement.applyDhcpEntries` failed with:
`ResourceUnavailableException: Resource [DataCenter:1] is unreachable:
Unable to apply dhcp entry on router`.
That was an observed failure rather than a deliberate test, so I have not
confirmed that stopping the VR is a minimal reproducer — any start-path
failure
that surfaces as `CloudRuntimeException` out of
`VirtualMachineManagerImpl.start()`
should exhibit the same leak.
4. Observe: one new VM per interval, each landing in `Stopped`, each
retaining its
`autoscale_vmgroup_vm_map` row, indefinitely.
##### EXPECTED RESULTS
Scale-up stops after a bounded number of consecutive failed starts, and/or
failed
instances are cleaned up, and/or `Stopped` members count toward
`max_members`.
##### ACTUAL RESULTS
Unbounded VM creation. In our case, one VM per 30s for ~21 hours:
* 2,043 VMs in `Error` (created after the guest subnet was exhausted — these
fail
at IP allocation, before a NIC is assigned)
* 253 VMs in `Stopped`, each holding a NIC and therefore a guest IP
* The /24 guest network reached 254/254 NICs and all subsequent VM
deployments —
including unrelated, non-autoscale ones — failed with
`InsufficientVirtualNetworkCapacityException: Unable to acquire Guest IP
address`
`autoscale.errored.instance.threshold` was at its default of 10 throughout,
and
`getErroredInstanceCount()` returned 0 the entire time, because none of the
leaked
instances were in `Error` — they were in `Stopped`.
##### SUGGESTED FIX
Any one of these would break the loop; the first two seem most direct:
1. Include `State.Stopped` in `getErroredInstanceCount()` (or add a separate
"failed instance" counter covering both `Error` and `Stopped`).
2. Broaden `doScaleUp`'s catch from `ServerApiException` to also handle
`CloudRuntimeException`, so `destroyVm()` runs and the member is not
leaked.
3. Count `Stopped` members in `countAvailableVmsByGroup()` so that leaked
members
contribute to `max_members` and become eligible for scale-down.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]