[ 
https://issues.apache.org/jira/browse/CAMEL-25008?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18118915#comment-18118915
 ] 

shashank commented on CAMEL-25008:
----------------------------------

PR: https://github.com/apache/camel/pull/26867

_Claude Code on behalf of allthingssecurity_

> Supervising route controller deadlocks when a route is stopped or started 
> while its restart attempt completes
> -------------------------------------------------------------------------------------------------------------
>
>                 Key: CAMEL-25008
>                 URL: https://issues.apache.org/jira/browse/CAMEL-25008
>             Project: Camel
>          Issue Type: Bug
>          Components: camel-core
>            Reporter: shashank
>            Priority: Minor
>
> Two locks are taken in opposite orders:
> * Restart thread: {{BackOffTimerTask.run()}} -> {{complete()}} takes the task 
> lock ({{BackOffTimerTask.java:220}}) and, while holding it, calls the 
> consumer registered by {{RouteManager.start}}, which takes the controller 
> lock ({{DefaultSupervisingRouteController.java:711}}).
> * User thread: {{stopRoute}}/{{startRoute}}/{{suspendRoute}}/{{resumeRoute}} 
> -> {{doStopRoute}}/{{doStartRoute}} take the controller lock ({{:445}}, 
> {{:467}}) and call {{routeManager.release(route)}} -> {{task.cancel()}} -> 
> {{complete()}}, which takes the task lock ({{BackOffTimerTask.java:144}}, 
> {{:220}}).
> If a user operation on the route takes the controller lock between the 
> restart thread's {{complete()}} taking the task lock and its consumer taking 
> the controller lock, both threads wait forever. {{complete()}} is reached 
> when a restart attempt succeeds ({{:197}}), when the attempts are exhausted 
> or the task was cancelled ({{:183}}). The window is small, but it is hit 
> every time the restart thread has to wait for the controller lock at 
> {{:711}}, which happens whenever any route operation holds the lock (the lock 
> is shared by all routes).
> After the deadlock the route controller is unusable: every later 
> {{startRoute}}/{{stopRoute}}/{{suspendRoute}}/{{resumeRoute}} blocks on the 
> controller lock, and the restart thread of the supervisor (single thread by 
> default) is stuck, so no route is restarted any more.
> *Reproduction*: route r1 fails to start (back-off 400 ms), route r2 has a 
> consumer that takes 600 ms to stop. At +150 ms the broker comes back and an 
> operator thread calls {{stopRoute("r2")}} and then {{stopRoute("r1")}}. r1's 
> attempt fires during the first stop, starts r1 when the lock is free, and 
> completes its task while {{stopRoute("r1")}} holds the controller lock:
> {noformat}
> [deadlock] operator thread still blocked after 5s; JVM-detected deadlocked 
> threads: 2
> [deadlock]   'Camel (camel-1) thread #1 - SupervisingRouteController' waits 
> for ReentrantLock$NonfairSync@5ba3f27a held by 'operator'
> [deadlock]       at 
> DefaultSupervisingRouteController$RouteManager.lambda$start$2(DefaultSupervisingRouteController.java:711)
> [deadlock]       at BackOffTimerTask.complete(BackOffTimerTask.java:222)
> [deadlock]       at BackOffTimerTask.run(BackOffTimerTask.java:197)
> [deadlock]   'operator' waits for ReentrantLock$NonfairSync@7bd4937b held by 
> 'Camel (camel-1) thread #1 - SupervisingRouteController'
> [deadlock]       at BackOffTimerTask.complete(BackOffTimerTask.java:220)
> [deadlock]       at BackOffTimerTask.cancel(BackOffTimerTask.java:144)
> [deadlock]       at 
> DefaultSupervisingRouteController$RouteManager.release(DefaultSupervisingRouteController.java:756)
> [deadlock]       at 
> DefaultSupervisingRouteController.doStopRoute(DefaultSupervisingRouteController.java:451)
> [deadlock]       at 
> DefaultSupervisingRouteController.stopRoute(DefaultSupervisingRouteController.java:318)
> {noformat}
> 3 of 3 runs deadlocked ({{ThreadMXBean.findDeadlockedThreads}} reports both 
> threads).
> TLA+: {{user_stop_deadlock}} violates {{NoDeadlock}} in 7 states (TRun -> 
> TCheck -> TStart(ok) -> TPost -> UAcquire("stop") -> TCbA), and 
> {{user_stop_liveness}} violates {{UserTerminates}}.
> *Proposed fix:* do not call foreign code while holding the task lock. In 
> {{BackOffTimerTask.complete()}}, copy the consumers under the lock and invoke 
> them after unlocking (the lock only protects the list):
> {code:java}
> void complete(Throwable throwable) {
>     this.cause = throwable;
>     List<BiConsumer<BackOffTimer.Task, Throwable>> copy;
>     lock.lock();
>     try {
>         copy = new ArrayList<>(consumers);
>     } finally {
>         lock.unlock();
>     }
>     copy.forEach(c -> c.accept(this, throwable));
> }
> {code}
> With this change the model ({{fix_user_ops2}}, {{fix_user_ops3}}, 
> {{fix_ctx_stop}}) has no deadlock and {{UserTerminates}} holds.
> With this change the restart thread no longer blocks a user operation, so the 
> interleaving of CAMEL-25007 (a completion of an old task removes the route's 
> new restart task) becomes reachable where it deadlocked before. The two fixes 
> should go together.
> _Filed with Claude Code on behalf of allthingssecurity._



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to