[ 
https://issues.apache.org/jira/browse/CAMEL-25052?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Claus Ibsen updated CAMEL-25052:
--------------------------------
    Fix Version/s: 4.23.0

> camel-file - FileLockClusterView: stopping the view no longer ends a 
> leadership check that is already running (regression from CAMEL-22784), so a 
> stopped view can take the cluster lock, and a quick restart doubles the checks
> --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
>
>                 Key: CAMEL-25052
>                 URL: https://issues.apache.org/jira/browse/CAMEL-25052
>             Project: Camel
>          Issue Type: Bug
>          Components: camel-file
>            Reporter: shashank
>            Priority: Minor
>             Fix For: 4.23.0
>
>
> Since CAMEL-22784 ("Use predictable scheduling for FileLockClusterService 
> lock acquisition", PR #20686, commit 7ee5afcd2f), {{FileLockClusterView}} 
> runs its leadership check as a chain of one-shot tasks: {{scheduleTryLock}} 
> ({{FileLockClusterView.java:296-324}}) calls 
> {{executor.schedule(this::tryLock, ...)}}, and {{tryLock}} schedules the next 
> run in its {{finally}} block ({{:280}}). The {{ScheduledFuture}} is no longer 
> stored in the {{task}} field ({{:60}}). {{task}} therefore stays null, and 
> {{closeInternal()}} ({{:156-163}}), which {{doStop}} ({{:133-154}}) calls, 
> cancels nothing. Before that commit, {{task.cancel(true)}} interrupted a 
> running check.
> {{tryLock}} checks {{isStarting() || isStarted()}} only once, when it begins 
> ({{:196}}). This has two consequences.
> # *A stop during a running check.* A follower that sees a stale or absent 
> leader opens the lock and data files and calls {{FileChannel.tryLock}} 
> ({{:250-263}}). If the view is stopped meanwhile, the check carries on after 
> {{doStop}} has closed the files. It takes the file lock, sets the member to 
> LEADER, fires the leadership event and writes its heartbeat once, on a view 
> that is Stopped. The next run of the chain sees the view stopped and ends. 
> The stopped node reports {{isLeader(namespace)=true}}, and the OS lock stays 
> held until the view is started again ({{doStart}} closes leftover files at 
> {{:101-104}}) or the JVM exits. Meanwhile no other node can become the 
> leader, so the clustered routes run nowhere.
> # *A stop and start within one {{acquireLockInterval}}.* The chain of the 
> previous start is still scheduled, sees the view started again and keeps 
> running next to the chain of the new start. Every quick restart adds one more 
> chain.
> The first case needs the view to be stopped while its node is inside the 
> acquisition I/O (it has just seen the old leader go away), and the JVM to 
> keep running afterwards. The view is stopped when its last user releases it: 
> a camel-master route stops, the last route with a {{ClusteredRoutePolicy}} is 
> removed, or the CamelContext stops in a JVM that keeps running (an 
> application server, a Spring context refresh). JMX {{stopView}} also stops 
> it. All of these can coincide with a leadership handover. The window is the 
> acquisition I/O: small on a local disk, larger on the network storage this 
> service targets ({{clusterDataTaskTimeout}} defaults to 10 s, with 5 attempts 
> per task).
> h3. Reproduction
> Each node is a CamelContext with its own {{FileLockClusterService}} on a 
> shared root ({{acquireLockInterval}} 500 ms) and a {{master:}} route.
> * A is the leader. B follows. A's context stops, and B's next check is held 
> (test hook in {{createRandomAccessFile}}) just before it opens the lock file. 
> B's master route is stopped, which stops B's view, then the hold is released. 
> Result: B's view is Stopped but {{isLeader(ns)=true}}. A probe from a 
> separate JVM finds the lock held, and a node C in a separate JVM does not 
> become the leader within 6 s.
> * The unit test in the PR does the same in one JVM, with B's view released 
> directly: B reports leadership and the lock file stays locked by B's channel.
> * Without a hook, with 200 ms file opens (as on network storage) and B's 
> master route stopped at a random time during the takeover: 3 of 15 rounds 
> ended with the stopped view holding the lock. With local file opens: 0 of 15.
> * One node, three quick {{stopRoute}}/{{startRoute}} of its master route: the 
> leadership check runs 10 times per 5 s before and 40 times per 5 s after.
> A TLA+ model of two or three nodes (the check, view stop and start, the OS 
> lock, and the camel-master listener) finds the same traces: 
> {{NoStoppedHolder}} and {{StoppedNotLeader}} are violated by the trace 
> TickAcquire(a), ConsumerStop(a), Acquire(a); {{EventuallyLeader}} is 
> violated; {{OneChain}} is violated by a quick stop and start. The fixed model 
> holds all properties.
> h3. Proposed fix
> * Give every start a generation number, incremented on every start and stop 
> under a small state lock. A check ends without rescheduling itself once its 
> generation is over, which ends the chain on stop and keeps one chain per 
> start. This replaces the dead {{task}} field.
> * Open the files and take the OS lock into local variables, without holding 
> any lock. Then re-check the generation under the state lock and only then 
> publish the lock and files to the view and set LEADER. If the view was 
> stopped meanwhile, release the lock and close the files instead.
> * {{doStop}} increments the generation and takes over the lock and files 
> under the same state lock, then does the file I/O (truncate, release, close) 
> outside it.
> * The leadership event is fired outside the state lock, as before. The state 
> lock is never held during file I/O or listener callbacks, so a stop does not 
> wait for a slow (NFS) check.
> Affected: 4.17.0 and later, and 4.14.5 and later (the backport PR #20780 
> brought the change to 4.14.x). Checked at the release tags: the {{task =}} 
> assignment is present in 4.16.0 and 4.14.4, and absent in 4.14.5, 4.17.0, 
> 4.18.0, 4.18.4, 4.22.0 and main.
> Duplicate check (2026-09-27): JIRA text "FileLockClusterView" (6 issues), 
> "FileLockClusterService" (4) and "file lock cluster" (15) return only the 
> CAMEL-22430 and CAMEL-22784 work and CAMEL-22541 (a flaky test). GitHub PRs 
> for these names are the CAMEL-22784 series (#20433, #20452, #20526, #20578, 
> #20590, #20686) and #20780. Nothing covers the stop.
> _Filed with Claude Code on behalf of allthingssecurity._



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to