shashank created CAMEL-25007:
--------------------------------

             Summary: Supervising route controller: a cancelled restart task 
completes a second time and removes the route's new restart task, which then 
restarts the route unsupervised
                 Key: CAMEL-25007
                 URL: https://issues.apache.org/jira/browse/CAMEL-25007
             Project: Camel
          Issue Type: Bug
          Components: camel-core
            Reporter: shashank


{{BackOffTimerTask.complete()}} calls every registered consumer each time it is 
called ({{BackOffTimerTask.java:218-226}}). A task that is cancelled while its 
attempt runs is completed twice: once by {{cancel()}} ({{:144}}) and once more 
by {{run()}} when the attempt returns ({{:183}} if it failed, {{:197}} if it 
succeeded).

The consumer that {{DefaultSupervisingRouteController.RouteManager.start}} 
registers ends with an unconditional {{routes.remove(r)}} 
({{DefaultSupervisingRouteController.java:744}}). If the route got a new 
restart task in between (a user {{startRoute}} that failed calls 
{{routeManager.start}}, {{:487}}), the second completion of the old task 
removes the new task from {{routes}}. The new task keeps running, but:

* {{getRestartingRoutes()}} and {{getRestartingRouteState(id)}} no longer show 
it, and {{hasUnhealthyRoutes()}} reports healthy while the route is down and 
being restarted;
* {{stopRoute(id)}} cannot cancel it ({{release}} finds nothing to cancel), so 
it starts the route after the user stopped it.

*Reproduction*: r1 fails to start; its first attempt is held at 
{{RouteRestartingEvent}}; the user retries {{startRoute("r1")}} while the 
broker is still down (fails, new task registered); the old attempt is released 
and fails too; later the user stops r1 and the broker comes back:
{noformat}
[orphan] user startRoute(r1) failed as expected: FailedToStartRouteException
[orphan] after user startRoute: restarting=1 state=BackOffTimerTask[... 
status=Active, currentAttempts=1 ...]
[orphan] after the old attempt finished: restarting=0 hasUnhealthyRoutes=false 
state=null
[orphan] user stopRoute(r1): status=Stopped
[orphan] 1.5s later: status=Started restartAttempts=2 consumerStarts=1
{noformat}

TLA+: {{user_ops2_orphan}} violates {{NoOrphanTask}} (every task that can still 
run is the route's current task): TRun(1) -> TCheck(1) -> TStart(1, FALSE) -> 
UAcquire/URelease/UAct("startFail") -> TPost(1) -> TCbA(1) -> TCbB(1). 
{{user_ops2_cbonce}} shows the callback running twice for one task 
({{CallbackOnce}}).

*Proposed fix:*
* in the callback, remove only the task's own entry: {{routes.remove(r, task)}} 
(the lambda has the task through {{backOffTask}}), and do the exhausted 
bookkeeping only for that task;
* make {{BackOffTimerTask.complete()}} run the consumers once (e.g. an 
{{AtomicBoolean completed}}), so a cancelled task is not completed again by 
{{run()}}.

The model with these changes ({{fix_user_ops2}}, {{fix_user_ops3}}) keeps 
{{NoOrphanTask}} and {{NoRestartAfterStop}}.

_Filed with Claude Code on behalf of allthingssecurity._




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to