shashank created CAMEL-25209:
--------------------------------

             Summary: camel-kubernetes - a failed Lease renewal stops the 
KubernetesLeadershipController for good: the leader gives up its routes and no 
other pod takes over
                 Key: CAMEL-25209
                 URL: https://issues.apache.org/jira/browse/CAMEL-25209
             Project: Camel
          Issue Type: Bug
          Components: camel-kubernetes
            Reporter: shashank


{{KubernetesLeadershipController}} runs the leader election as a chain of tasks 
on a single-thread scheduled executor: each {{refreshStatus()}} run schedules 
the next one ({{rescheduleAfterDelay()}} or 
{{serializedExecutor.execute(this::refreshStatus)}}). Every Kubernetes call in 
it catches its exceptions and reschedules ({{lookupNewLeaderInfo}}, 
{{tryAcquireLeadership}}, {{yieldLeadership}}), except the Lease renewal in the 
LEADER state:

{code:java}
HasMetadata newLease = 
this.leaseManager.refreshLeaseRenewTime(kubernetesClient, 
this.latestLeaseResource,
        this.lockConfiguration.getRenewDeadlineSeconds());
updateLatestLeaderInfo(newLease, this.latestMembers);
rescheduleAfterDelay();
{code}

With the default resource type {{Lease}}, 
{{NativeLeaseResourceManager.refreshLeaseRenewTime}} does a PUT with 
{{lockResourceVersion}} once per renew deadline. When that PUT fails, the 
{{KubernetesClientException}} leaves {{refreshStatus}}, the executor keeps it 
in the (unread) future, and the next run is never scheduled. Nothing is logged. 
The Kubernetes client retries a 429, a 5xx or an I/O error itself (by default 
up to 10 times with an exponential backoff, about 20 s), so the PUT fails when 
the API server cannot be reached or answers with errors for longer than that 
(an API server restart, a control plane upgrade, a network problem), or with an 
error that is not retried (a 409 conflict, 401/403, 404 when the Lease was 
deleted).

Then:
* on the leader pod, the {{TimedLeaderNotifier}} is no longer refreshed, so 
after the renew deadline it reports "no leader": the pod stops its master 
routes / clustered routes;
* the other pods keep reading a Lease held by a pod that is running and ready, 
which {{LeaderInfo.hasValidLeader()}} accepts (it does not look at 
{{renewTime}}), so none of them tries to acquire it.

No pod of the group leads until the former leader pod is restarted (or becomes 
not ready). One failed renewal is enough; a failed lookup of the Lease during 
the same outage is harmless, as it is caught and retried.

The {{ConfigMap}} resource type is not affected ({{refreshLeaseRenewTime}} does 
nothing there). The unguarded call came with the Lease support in CAMEL-15881 
(2020); the other calls were guarded from the start.

h3. Reproduction

{{KubernetesClusterServiceTest}} with the existing {{LockTestServer}} (mock API 
server), Lease type, two pods. When a leader is elected, the leader's server 
refuses PUT requests on the Lease (500) until one renewal has been refused, 
then accepts them again (the test client has the client retries disabled, as in 
the other tests of the module). Without the fix the leader's recorder reports 
{{null}} after the renew deadline and stays so ({{expected: <mypod1> but was: 
<null>}}, three runs), and the test log has no warning about the failed renewal.

h3. Proposed fix

{{refreshStatus()}} wraps the state machine in a {{try/catch}}: on an exception 
it logs a WARN (stack trace at DEBUG) and schedules the next run after the 
retry period, like the other failure paths; after {{stop()}} it only logs at 
DEBUG. In the LEADER state the next run reads the Lease again and renews it.

Test: 
{{KubernetesClusterServiceTest.testLeaderKeepsLeadershipAfterFailedLeaseRenewal}},
 with a new option of {{LockTestServer}} to refuse only the update (PUT) 
requests. It fails without the fix.

Affected: all versions with the Lease resource type (the same unguarded call at 
camel-3.20.0, 4.18.0 and main).

Duplicate check (2026-09-30): JIRA text "KubernetesLeadershipController", 
"refreshLeaseRenewTime", "kubernetes leadership renew", and component 
camel-kubernetes with leader/leadership/cluster/lease in the summary: only 
CAMEL-21720, CAMEL-15881, CAMEL-19825, CAMEL-11837 (different problems). GitHub 
pull requests "KubernetesLeadershipController", "kubernetes leadership": none. 
No open pull request touches these files.

_Filed with Claude Code on behalf of allthingssecurity._




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to