[
https://issues.apache.org/jira/browse/CAMEL-25007?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18118914#comment-18118914
]
shashank commented on CAMEL-25007:
----------------------------------
PR: https://github.com/apache/camel/pull/26867
_Claude Code on behalf of allthingssecurity_
> Supervising route controller: a cancelled restart task completes a second
> time and removes the route's new restart task, which then restarts the route
> unsupervised
> -------------------------------------------------------------------------------------------------------------------------------------------------------------------
>
> Key: CAMEL-25007
> URL: https://issues.apache.org/jira/browse/CAMEL-25007
> Project: Camel
> Issue Type: Bug
> Components: camel-core
> Reporter: shashank
> Priority: Minor
>
> {{BackOffTimerTask.complete()}} calls every registered consumer each time it
> is called ({{BackOffTimerTask.java:218-226}}). A task that is cancelled while
> its attempt runs is completed twice: once by {{cancel()}} ({{:144}}) and once
> more by {{run()}} when the attempt returns ({{:183}} if it failed, {{:197}}
> if it succeeded).
> The consumer that {{DefaultSupervisingRouteController.RouteManager.start}}
> registers ends with an unconditional {{routes.remove(r)}}
> ({{DefaultSupervisingRouteController.java:744}}). If the route got a new
> restart task in between (a user {{startRoute}} that failed calls
> {{routeManager.start}}, {{:487}}), the second completion of the old task
> removes the new task from {{routes}}. The new task keeps running, but:
> * {{getRestartingRoutes()}} and {{getRestartingRouteState(id)}} no longer
> show it, and {{hasUnhealthyRoutes()}} reports healthy while the route is down
> and being restarted;
> * {{stopRoute(id)}} cannot cancel it ({{release}} finds nothing to cancel),
> so it starts the route after the user stopped it.
> *Reproduction*: r1 fails to start; its first attempt is held at
> {{RouteRestartingEvent}}; the user retries {{startRoute("r1")}} while the
> broker is still down (fails, new task registered); the old attempt is
> released and fails too; later the user stops r1 and the broker comes back:
> {noformat}
> [orphan] user startRoute(r1) failed as expected: FailedToStartRouteException
> [orphan] after user startRoute: restarting=1 state=BackOffTimerTask[...
> status=Active, currentAttempts=1 ...]
> [orphan] after the old attempt finished: restarting=0
> hasUnhealthyRoutes=false state=null
> [orphan] user stopRoute(r1): status=Stopped
> [orphan] 1.5s later: status=Started restartAttempts=2 consumerStarts=1
> {noformat}
> TLA+: {{user_ops2_orphan}} violates {{NoOrphanTask}} (every task that can
> still run is the route's current task): TRun(1) -> TCheck(1) -> TStart(1,
> FALSE) -> UAcquire/URelease/UAct("startFail") -> TPost(1) -> TCbA(1) ->
> TCbB(1). {{user_ops2_cbonce}} shows the callback running twice for one task
> ({{CallbackOnce}}).
> *Proposed fix:*
> * in the callback, remove only the task's own entry: {{routes.remove(r,
> task)}} (the lambda has the task through {{backOffTask}}), and do the
> exhausted bookkeeping only for that task;
> * make {{BackOffTimerTask.complete()}} run the consumers once (e.g. an
> {{AtomicBoolean completed}}), so a cancelled task is not completed again by
> {{run()}}.
> The model with these changes ({{fix_user_ops2}}, {{fix_user_ops3}}) keeps
> {{NoOrphanTask}} and {{NoRestartAfterStop}}.
> _Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)