[
https://issues.apache.org/jira/browse/CAMEL-25008?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18118915#comment-18118915
]
shashank commented on CAMEL-25008:
----------------------------------
PR: https://github.com/apache/camel/pull/26867
_Claude Code on behalf of allthingssecurity_
> Supervising route controller deadlocks when a route is stopped or started
> while its restart attempt completes
> -------------------------------------------------------------------------------------------------------------
>
> Key: CAMEL-25008
> URL: https://issues.apache.org/jira/browse/CAMEL-25008
> Project: Camel
> Issue Type: Bug
> Components: camel-core
> Reporter: shashank
> Priority: Minor
>
> Two locks are taken in opposite orders:
> * Restart thread: {{BackOffTimerTask.run()}} -> {{complete()}} takes the task
> lock ({{BackOffTimerTask.java:220}}) and, while holding it, calls the
> consumer registered by {{RouteManager.start}}, which takes the controller
> lock ({{DefaultSupervisingRouteController.java:711}}).
> * User thread: {{stopRoute}}/{{startRoute}}/{{suspendRoute}}/{{resumeRoute}}
> -> {{doStopRoute}}/{{doStartRoute}} take the controller lock ({{:445}},
> {{:467}}) and call {{routeManager.release(route)}} -> {{task.cancel()}} ->
> {{complete()}}, which takes the task lock ({{BackOffTimerTask.java:144}},
> {{:220}}).
> If a user operation on the route takes the controller lock between the
> restart thread's {{complete()}} taking the task lock and its consumer taking
> the controller lock, both threads wait forever. {{complete()}} is reached
> when a restart attempt succeeds ({{:197}}), when the attempts are exhausted
> or the task was cancelled ({{:183}}). The window is small, but it is hit
> every time the restart thread has to wait for the controller lock at
> {{:711}}, which happens whenever any route operation holds the lock (the lock
> is shared by all routes).
> After the deadlock the route controller is unusable: every later
> {{startRoute}}/{{stopRoute}}/{{suspendRoute}}/{{resumeRoute}} blocks on the
> controller lock, and the restart thread of the supervisor (single thread by
> default) is stuck, so no route is restarted any more.
> *Reproduction*: route r1 fails to start (back-off 400 ms), route r2 has a
> consumer that takes 600 ms to stop. At +150 ms the broker comes back and an
> operator thread calls {{stopRoute("r2")}} and then {{stopRoute("r1")}}. r1's
> attempt fires during the first stop, starts r1 when the lock is free, and
> completes its task while {{stopRoute("r1")}} holds the controller lock:
> {noformat}
> [deadlock] operator thread still blocked after 5s; JVM-detected deadlocked
> threads: 2
> [deadlock] 'Camel (camel-1) thread #1 - SupervisingRouteController' waits
> for ReentrantLock$NonfairSync@5ba3f27a held by 'operator'
> [deadlock] at
> DefaultSupervisingRouteController$RouteManager.lambda$start$2(DefaultSupervisingRouteController.java:711)
> [deadlock] at BackOffTimerTask.complete(BackOffTimerTask.java:222)
> [deadlock] at BackOffTimerTask.run(BackOffTimerTask.java:197)
> [deadlock] 'operator' waits for ReentrantLock$NonfairSync@7bd4937b held by
> 'Camel (camel-1) thread #1 - SupervisingRouteController'
> [deadlock] at BackOffTimerTask.complete(BackOffTimerTask.java:220)
> [deadlock] at BackOffTimerTask.cancel(BackOffTimerTask.java:144)
> [deadlock] at
> DefaultSupervisingRouteController$RouteManager.release(DefaultSupervisingRouteController.java:756)
> [deadlock] at
> DefaultSupervisingRouteController.doStopRoute(DefaultSupervisingRouteController.java:451)
> [deadlock] at
> DefaultSupervisingRouteController.stopRoute(DefaultSupervisingRouteController.java:318)
> {noformat}
> 3 of 3 runs deadlocked ({{ThreadMXBean.findDeadlockedThreads}} reports both
> threads).
> TLA+: {{user_stop_deadlock}} violates {{NoDeadlock}} in 7 states (TRun ->
> TCheck -> TStart(ok) -> TPost -> UAcquire("stop") -> TCbA), and
> {{user_stop_liveness}} violates {{UserTerminates}}.
> *Proposed fix:* do not call foreign code while holding the task lock. In
> {{BackOffTimerTask.complete()}}, copy the consumers under the lock and invoke
> them after unlocking (the lock only protects the list):
> {code:java}
> void complete(Throwable throwable) {
> this.cause = throwable;
> List<BiConsumer<BackOffTimer.Task, Throwable>> copy;
> lock.lock();
> try {
> copy = new ArrayList<>(consumers);
> } finally {
> lock.unlock();
> }
> copy.forEach(c -> c.accept(this, throwable));
> }
> {code}
> With this change the model ({{fix_user_ops2}}, {{fix_user_ops3}},
> {{fix_ctx_stop}}) has no deadlock and {{UserTerminates}} holds.
> With this change the restart thread no longer blocks a user operation, so the
> interleaving of CAMEL-25007 (a completion of an old task removes the route's
> new restart task) becomes reachable where it deadlocked before. The two fixes
> should go together.
> _Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)