shashank created CAMEL-25007:
--------------------------------
Summary: Supervising route controller: a cancelled restart task
completes a second time and removes the route's new restart task, which then
restarts the route unsupervised
Key: CAMEL-25007
URL: https://issues.apache.org/jira/browse/CAMEL-25007
Project: Camel
Issue Type: Bug
Components: camel-core
Reporter: shashank
{{BackOffTimerTask.complete()}} calls every registered consumer each time it is
called ({{BackOffTimerTask.java:218-226}}). A task that is cancelled while its
attempt runs is completed twice: once by {{cancel()}} ({{:144}}) and once more
by {{run()}} when the attempt returns ({{:183}} if it failed, {{:197}} if it
succeeded).
The consumer that {{DefaultSupervisingRouteController.RouteManager.start}}
registers ends with an unconditional {{routes.remove(r)}}
({{DefaultSupervisingRouteController.java:744}}). If the route got a new
restart task in between (a user {{startRoute}} that failed calls
{{routeManager.start}}, {{:487}}), the second completion of the old task
removes the new task from {{routes}}. The new task keeps running, but:
* {{getRestartingRoutes()}} and {{getRestartingRouteState(id)}} no longer show
it, and {{hasUnhealthyRoutes()}} reports healthy while the route is down and
being restarted;
* {{stopRoute(id)}} cannot cancel it ({{release}} finds nothing to cancel), so
it starts the route after the user stopped it.
*Reproduction*: r1 fails to start; its first attempt is held at
{{RouteRestartingEvent}}; the user retries {{startRoute("r1")}} while the
broker is still down (fails, new task registered); the old attempt is released
and fails too; later the user stops r1 and the broker comes back:
{noformat}
[orphan] user startRoute(r1) failed as expected: FailedToStartRouteException
[orphan] after user startRoute: restarting=1 state=BackOffTimerTask[...
status=Active, currentAttempts=1 ...]
[orphan] after the old attempt finished: restarting=0 hasUnhealthyRoutes=false
state=null
[orphan] user stopRoute(r1): status=Stopped
[orphan] 1.5s later: status=Started restartAttempts=2 consumerStarts=1
{noformat}
TLA+: {{user_ops2_orphan}} violates {{NoOrphanTask}} (every task that can still
run is the route's current task): TRun(1) -> TCheck(1) -> TStart(1, FALSE) ->
UAcquire/URelease/UAct("startFail") -> TPost(1) -> TCbA(1) -> TCbB(1).
{{user_ops2_cbonce}} shows the callback running twice for one task
({{CallbackOnce}}).
*Proposed fix:*
* in the callback, remove only the task's own entry: {{routes.remove(r, task)}}
(the lambda has the task through {{backOffTask}}), and do the exhausted
bookkeeping only for that task;
* make {{BackOffTimerTask.complete()}} run the consumers once (e.g. an
{{AtomicBoolean completed}}), so a cancelled task is not completed again by
{{run()}}.
The model with these changes ({{fix_user_ops2}}, {{fix_user_ops3}}) keeps
{{NoOrphanTask}} and {{NoRestartAfterStop}}.
_Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)