[
https://issues.apache.org/jira/browse/CAMEL-25209?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18121459#comment-18121459
]
Claus Ibsen commented on CAMEL-25209:
-------------------------------------
Merged in https://github.com/apache/camel/pull/27158 (commit 97e434aa597d),
thanks!
_Claude Code on behalf of davsclaus_
> camel-kubernetes - a failed Lease renewal stops the
> KubernetesLeadershipController for good: the leader gives up its routes and
> no other pod takes over
> -------------------------------------------------------------------------------------------------------------------------------------------------------
>
> Key: CAMEL-25209
> URL: https://issues.apache.org/jira/browse/CAMEL-25209
> Project: Camel
> Issue Type: Bug
> Components: camel-kubernetes
> Reporter: shashank
> Assignee: shashank
> Priority: Major
> Fix For: 4.23.0
>
>
> {{KubernetesLeadershipController}} runs the leader election as a chain of
> tasks on a single-thread scheduled executor: each {{refreshStatus()}} run
> schedules the next one ({{rescheduleAfterDelay()}} or
> {{serializedExecutor.execute(this::refreshStatus)}}). Every Kubernetes call
> in it catches its exceptions and reschedules ({{lookupNewLeaderInfo}},
> {{tryAcquireLeadership}}, {{yieldLeadership}}), except the Lease renewal in
> the LEADER state:
> {code:java}
> HasMetadata newLease =
> this.leaseManager.refreshLeaseRenewTime(kubernetesClient,
> this.latestLeaseResource,
> this.lockConfiguration.getRenewDeadlineSeconds());
> updateLatestLeaderInfo(newLease, this.latestMembers);
> rescheduleAfterDelay();
> {code}
> With the default resource type {{Lease}},
> {{NativeLeaseResourceManager.refreshLeaseRenewTime}} does a PUT with
> {{lockResourceVersion}} once per renew deadline. When that PUT fails, the
> {{KubernetesClientException}} leaves {{refreshStatus}}, the executor keeps it
> in the (unread) future, and the next run is never scheduled. Nothing is
> logged. The Kubernetes client retries a 429, a 5xx or an I/O error itself (by
> default up to 10 times with an exponential backoff, about 20 s), so the PUT
> fails when the API server cannot be reached or answers with errors for longer
> than that (an API server restart, a control plane upgrade, a network
> problem), or with an error that is not retried (a 409 conflict, 401/403, 404
> when the Lease was deleted).
> Then:
> * on the leader pod, the {{TimedLeaderNotifier}} is no longer refreshed, so
> after the renew deadline it reports "no leader": the pod stops its master
> routes / clustered routes;
> * the other pods keep reading a Lease held by a pod that is running and
> ready, which {{LeaderInfo.hasValidLeader()}} accepts (it does not look at
> {{renewTime}}), so none of them tries to acquire it.
> No pod of the group leads until the former leader pod is restarted (or
> becomes not ready). One failed renewal is enough; a failed lookup of the
> Lease during the same outage is harmless, as it is caught and retried.
> The {{ConfigMap}} resource type is not affected ({{refreshLeaseRenewTime}}
> does nothing there). The unguarded call came with the Lease support in
> CAMEL-15881 (2020); the other calls were guarded from the start.
> h3. Reproduction
> {{KubernetesClusterServiceTest}} with the existing {{LockTestServer}} (mock
> API server), Lease type, two pods. When a leader is elected, the leader's
> server refuses PUT requests on the Lease (500) until one renewal has been
> refused, then accepts them again (the test client has the client retries
> disabled, as in the other tests of the module). Without the fix the leader's
> recorder reports {{null}} after the renew deadline and stays so ({{expected:
> <mypod1> but was: <null>}}, three runs), and the test log has no warning
> about the failed renewal.
> h3. Proposed fix
> {{refreshStatus()}} wraps the state machine in a {{try/catch}}: on an
> exception it logs a WARN (stack trace at DEBUG) and schedules the next run
> after the retry period, like the other failure paths; after {{stop()}} it
> only logs at DEBUG. In the LEADER state the next run reads the Lease again
> and renews it.
> Test:
> {{KubernetesClusterServiceTest.testLeaderKeepsLeadershipAfterFailedLeaseRenewal}},
> with a new option of {{LockTestServer}} to refuse only the update (PUT)
> requests. It fails without the fix.
> Affected: all versions with the Lease resource type (the same unguarded call
> at camel-3.20.0, 4.18.0 and main).
> Duplicate check (2026-09-30): JIRA text "KubernetesLeadershipController",
> "refreshLeaseRenewTime", "kubernetes leadership renew", and component
> camel-kubernetes with leader/leadership/cluster/lease in the summary: only
> CAMEL-21720, CAMEL-15881, CAMEL-19825, CAMEL-11837 (different problems).
> GitHub pull requests "KubernetesLeadershipController", "kubernetes
> leadership": none. No open pull request touches these files.
> _Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)