[
https://issues.apache.org/jira/browse/CAMEL-25273?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Claus Ibsen updated CAMEL-25273:
--------------------------------
Fix Version/s: 4.23.0
> camel-kubernetes - the watch consumers stop receiving events for good when
> the Kubernetes client closes their watch with an error (410 Gone), while the
> route stays started
> ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------
>
> Key: CAMEL-25273
> URL: https://issues.apache.org/jira/browse/CAMEL-25273
> Project: Camel
> Issue Type: Bug
> Components: camel-kubernetes
> Reporter: shashank
> Priority: Major
> Fix For: 4.23.0
>
>
> The watch consumers of camel-kubernetes (kubernetes-pods, -services,
> -deployments, -config-maps, -custom-resources, -events, -hpa, -namespaces,
> -nodes, -replication-controllers and openshift-deploymentconfigs) create one
> fabric8 {{Watch}} when they start, and their
> {{Watcher.onClose(WatcherException)}} only logs:
> {code:java}
> @Override
> public void onClose(WatcherException cause) {
> if (cause != null) {
> LOG.error(cause.getMessage(), cause);
> }
> }
> {code}
> The fabric8 client reconnects a watch by itself after transient errors, but
> closes it for good, calling {{onClose}} with an exception, when it gives up:
> * the API server answers *410 Gone* ("too old resource version"):
> {{AbstractWatchManager.onStatus}} closes the watch with the comment "The
> resource version no longer exists - this has to be handled by the caller".
> This happens to long-running watches when the client reconnects with a
> resource version that etcd has compacted (or after an API server restart);
> * the reconnect limit ({{watchReconnectLimit}}) is reached, "Exhausted
> reconnects".
> After that the consumer receives no event anymore: one ERROR log line, the
> route stays started, health checks are green.
> h3. Reproduction
> {{KubernetesPodsConsumerWatchClosedTest}} (fabric8 mock server, as the
> producer tests use): the first watch request of {{kubernetes-pods}} gets an
> {{ERROR}} event with a 410 {{Status}}, a second watch request would get an
> {{ADDED}} pod. On main the consumer logs "too old resource version" and never
> watches again: {{mock://result Received message count. Expected: <1> but was:
> <0>}}.
> h3. Proposed fix
> {{onClose}} with an exception runs the watch task of the consumer again
> ({{KubernetesHelper.watchAgain}}: only when the consumer is not stopping and
> its executor is running). A new watch without resource version starts with
> the current state (the API server sends {{ADDED}} for the existing resources,
> as when the route starts), then the changes. A close without exception (the
> consumer closing its own watch on stop) is unchanged. The same two lines in
> the 11 consumers. With the fix the new test passes and the camel-kubernetes
> unit tests pass (156 tests; the ITs need a cluster).
> Affected: 4.14.x, 4.18.x and main (same code).
> Not in scope: the consumers watch without first listing, so changes between
> the closed watch and the new one are only seen as the state of the new watch;
> a full fix would use informers. Also, when the new watch cannot be started
> (for example the API server is still unreachable after the reconnect limit
> was reached), the task fails as it does when the route starts.
> Duplicate check (2026-10-02): JIRA text "kubernetes" with
> "watch"/"onClose"/"reconnect" (16 issues): only CAMEL-11020 (closing the
> watchers on stop), CAMEL-13978 and features. GitHub pull requests "kubernetes
> watch", "kubernetes consumer watch": none.
> _Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)