shashank created CAMEL-25008:
--------------------------------
Summary: Supervising route controller deadlocks when a route is
stopped or started while its restart attempt completes
Key: CAMEL-25008
URL: https://issues.apache.org/jira/browse/CAMEL-25008
Project: Camel
Issue Type: Bug
Components: camel-core
Reporter: shashank
Two locks are taken in opposite orders:
* Restart thread: {{BackOffTimerTask.run()}} -> {{complete()}} takes the task
lock ({{BackOffTimerTask.java:220}}) and, while holding it, calls the consumer
registered by {{RouteManager.start}}, which takes the controller lock
({{DefaultSupervisingRouteController.java:711}}).
* User thread: {{stopRoute}}/{{startRoute}}/{{suspendRoute}}/{{resumeRoute}} ->
{{doStopRoute}}/{{doStartRoute}} take the controller lock ({{:445}}, {{:467}})
and call {{routeManager.release(route)}} -> {{task.cancel()}} ->
{{complete()}}, which takes the task lock ({{BackOffTimerTask.java:144}},
{{:220}}).
If a user operation on the route takes the controller lock between the restart
thread's {{complete()}} taking the task lock and its consumer taking the
controller lock, both threads wait forever. {{complete()}} is reached when a
restart attempt succeeds ({{:197}}), when the attempts are exhausted or the
task was cancelled ({{:183}}). The window is small, but it is hit every time
the restart thread has to wait for the controller lock at {{:711}}, which
happens whenever any route operation holds the lock (the lock is shared by all
routes).
After the deadlock the route controller is unusable: every later
{{startRoute}}/{{stopRoute}}/{{suspendRoute}}/{{resumeRoute}} blocks on the
controller lock, and the restart thread of the supervisor (single thread by
default) is stuck, so no route is restarted any more.
*Reproduction*: route r1 fails to start (back-off 400 ms), route r2 has a
consumer that takes 600 ms to stop. At +150 ms the broker comes back and an
operator thread calls {{stopRoute("r2")}} and then {{stopRoute("r1")}}. r1's
attempt fires during the first stop, starts r1 when the lock is free, and
completes its task while {{stopRoute("r1")}} holds the controller lock:
{noformat}
[deadlock] operator thread still blocked after 5s; JVM-detected deadlocked
threads: 2
[deadlock] 'Camel (camel-1) thread #1 - SupervisingRouteController' waits for
ReentrantLock$NonfairSync@5ba3f27a held by 'operator'
[deadlock] at
DefaultSupervisingRouteController$RouteManager.lambda$start$2(DefaultSupervisingRouteController.java:711)
[deadlock] at BackOffTimerTask.complete(BackOffTimerTask.java:222)
[deadlock] at BackOffTimerTask.run(BackOffTimerTask.java:197)
[deadlock] 'operator' waits for ReentrantLock$NonfairSync@7bd4937b held by
'Camel (camel-1) thread #1 - SupervisingRouteController'
[deadlock] at BackOffTimerTask.complete(BackOffTimerTask.java:220)
[deadlock] at BackOffTimerTask.cancel(BackOffTimerTask.java:144)
[deadlock] at
DefaultSupervisingRouteController$RouteManager.release(DefaultSupervisingRouteController.java:756)
[deadlock] at
DefaultSupervisingRouteController.doStopRoute(DefaultSupervisingRouteController.java:451)
[deadlock] at
DefaultSupervisingRouteController.stopRoute(DefaultSupervisingRouteController.java:318)
{noformat}
3 of 3 runs deadlocked ({{ThreadMXBean.findDeadlockedThreads}} reports both
threads).
TLA+: {{user_stop_deadlock}} violates {{NoDeadlock}} in 7 states (TRun ->
TCheck -> TStart(ok) -> TPost -> UAcquire("stop") -> TCbA), and
{{user_stop_liveness}} violates {{UserTerminates}}.
*Proposed fix:* do not call foreign code while holding the task lock. In
{{BackOffTimerTask.complete()}}, copy the consumers under the lock and invoke
them after unlocking (the lock only protects the list):
{code:java}
void complete(Throwable throwable) {
this.cause = throwable;
List<BiConsumer<BackOffTimer.Task, Throwable>> copy;
lock.lock();
try {
copy = new ArrayList<>(consumers);
} finally {
lock.unlock();
}
copy.forEach(c -> c.accept(this, throwable));
}
{code}
With this change the model ({{fix_user_ops2}}, {{fix_user_ops3}},
{{fix_ctx_stop}}) has no deadlock and {{UserTerminates}} holds.
With this change the restart thread no longer blocks a user operation, so the
interleaving of CAMEL-25007 (a completion of an old task removes the route's
new restart task) becomes reachable where it deadlocked before. The two fixes
should go together.
_Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)