[ 
https://issues.apache.org/jira/browse/CAMEL-25286?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Claus Ibsen updated CAMEL-25286:
--------------------------------
    Fix Version/s: 4.23.0

> camel-consul - ConsulClusterView stops taking part in the leader election 
> after a failed query or when Consul invalidates its session, and can keep the 
> leadership while another node takes it
> ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
>
>                 Key: CAMEL-25286
>                 URL: https://issues.apache.org/jira/browse/CAMEL-25286
>             Project: Camel
>          Issue Type: Bug
>          Components: camel-consul
>            Reporter: shashank
>            Priority: Major
>             Fix For: 4.23.0
>
>
> {{ConsulClusterView}} (used by {{ConsulClusterService}}, {{camel-master}} and 
> {{ClusteredRoutePolicy}}) holds the leadership with a lock on 
> {{<rootPath>/<namespace>}} owned by a Consul session. A {{Watcher}} watches 
> that key with a chain of blocking queries; each query also renews the 
> session. Three paths end the election for the node until the cluster service 
> is restarted:
> h3. 1. A failed query: the watch ends and the node can stay leader while 
> another node takes the lock
> {code:java}
> public void onFailure(Throwable throwable) {
>     LOGGER.debug("{}", throwable.getMessage(), throwable);
>     if (sessionId.get() != null) {
>         keyValueClient.releaseLock(configuration.getRootPath(), 
> sessionId.get());   // wrong key, synchronous
>     }
>     localMember.setMaster(false);
>     watch();
> }
> {code}
> * It releases the lock of the root path ({{/camel}}), not of the namespace 
> path that the session holds (the same mistake was fixed in {{getLeader()}} in 
> 2021, PR #4937).
> * {{releaseLock}} is a synchronous HTTP call. When the query failed because 
> Consul cannot be reached (the agent restarts, a network error), that call 
> throws a {{ConsulException}}, so {{setMaster(false)}} and {{watch()}} never 
> run: no more queries and no more session renewals. A leader keeps its 
> leadership locally (its clustered routes keep running); its session expires 
> after the TTL, Consul releases the lock, another node acquires it, and both 
> nodes run the clustered routes. A node that was not leader never takes part 
> again.
> * The same happens when handling an answer fails ({{acquireLock}} and the 
> leadership event call Consul synchronously from {{onComplete}}), and 
> {{setMaster(false)}} itself calls Consul ({{getLeader()}}) to fire the 
> leadership event, so the event was lost while Consul was down.
> h3. 2. An invalidated session is never replaced
> When Consul invalidates the session (its TTL expired during a network 
> partition, the local agent restarted or left so its {{serfHealth}} check 
> failed, Consul lost its data), every renewal fails with {{404 Session id 
> '...' not found}} (the stack trace of CAMEL-15423), and {{acquireLock}} 
> returns {{false}} for ever because {{getSessionInfo}} finds no session. The 
> node can never be leader again; if all nodes lost their sessions (for example 
> a rolling restart of the Consul agents) no node is leader any more and the 
> clustered routes stay stopped everywhere.
> h3. 3. A key that does not exist is ignored
> {{onComplete}} only acts when the key exists ({{if (value.isPresent())}}). 
> After Consul lost its data the key is gone and no node acquires the lock 
> again; a node that was leader also keeps its leadership locally.
> h3. Reproduction (unit test, no Consul)
> {{ConsulClusterViewRecoveryTest}} runs the real {{ConsulClusterService}} and 
> view against mocked {{SessionClient}} / {{KeyValueClient}} that simulate 
> Consul (sessions, lock holder, key, reachability; a renewal of an unknown 
> session throws {{ConsulException}} with code 404 as the client does). On main 
> all 4 tests fail:
> * Consul cannot be reached, the query fails: the node stays leader 
> ({{expected: <false> but was: <true>}}) and never queries again.
> * A failed query: the lock released is not the one of the namespace path 
> (Mockito: {{Argument(s) are different! Wanted: releaseLock("/camel/my-ns", 
> "session-1")}}, actual {{"/camel"}}).
> * Consul invalidated the session: no new session is created ({{expected: <2> 
> but was: <1>}}) and the node never takes the free lock again.
> * Consul lost its data: the node stays leader without a lock ({{expected: 
> <false> but was: <true>}}) and never creates the key again.
> h3. Proposed fix
> * {{onFailure}}: give the leadership up first, release the lock of the 
> namespace path (ignoring a failure), and query again after 
> {{sessionRefreshInterval}} (at least one second) on a single-thread scheduler 
> of the view, so that an agent that cannot be reached is not queried in a loop.
> * Renewal: a session that Consul does not know any more (404, or no session 
> returned) is replaced by a new one (under the session lock, only while the 
> view runs), after giving the leadership up; other renewal failures are logged 
> at debug and retried with the next query.
> * {{onComplete}}: a missing key is handled like a free key (try to acquire 
> it); an exception while handling the answer gives the leadership up instead 
> of ending the watch. The index is stored before handling the answer.
> * The leadership event is fired with no leader when the current leader cannot 
> be read.
> A failed query while Consul can be reached now really releases the lock (no 
> {{lock-delay}} on an explicit release), so the leadership can move to another 
> node; the node had already given it up locally on main. While the Consul 
> cluster has no leader the release fails too and the node takes its own lock 
> back with the next answer. File lock and ZooKeeper cluster views also give 
> the leadership up on the first failure.
> No new option. With the fix the 4 new tests pass and the module's unit tests 
> pass (8); the cluster ITs need Docker (not run).
> Affected: 4.14.x, 4.18.x and main (same code).
> Duplicate check (2026-10-03): JIRA text "ConsulClusterView" / 
> "ConsulClusterService" / "consul" + "leader" / "consul" + "session" (14 
> issues): only CAMEL-15423 (2020, closed Incomplete, same symptom, no fix), 
> CAMEL-22523 (test automation), CAMEL-25275 (consumers, open, ours). GitHub 
> pull requests "ConsulClusterView": only #25842 (ZooKeeper cluster shutdown) 
> and #4937 (getLeader path, 2021).
> _Filed with Claude Code on behalf of allthingssecurity._



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to