[ 
https://issues.apache.org/jira/browse/CAMEL-25273?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Claus Ibsen updated CAMEL-25273:
--------------------------------
    Fix Version/s: 4.23.0

> camel-kubernetes - the watch consumers stop receiving events for good when 
> the Kubernetes client closes their watch with an error (410 Gone), while the 
> route stays started
> ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------
>
>                 Key: CAMEL-25273
>                 URL: https://issues.apache.org/jira/browse/CAMEL-25273
>             Project: Camel
>          Issue Type: Bug
>          Components: camel-kubernetes
>            Reporter: shashank
>            Priority: Major
>             Fix For: 4.23.0
>
>
> The watch consumers of camel-kubernetes (kubernetes-pods, -services, 
> -deployments, -config-maps, -custom-resources, -events, -hpa, -namespaces, 
> -nodes, -replication-controllers and openshift-deploymentconfigs) create one 
> fabric8 {{Watch}} when they start, and their 
> {{Watcher.onClose(WatcherException)}} only logs:
> {code:java}
> @Override
> public void onClose(WatcherException cause) {
>     if (cause != null) {
>         LOG.error(cause.getMessage(), cause);
>     }
> }
> {code}
> The fabric8 client reconnects a watch by itself after transient errors, but 
> closes it for good, calling {{onClose}} with an exception, when it gives up:
> * the API server answers *410 Gone* ("too old resource version"): 
> {{AbstractWatchManager.onStatus}} closes the watch with the comment "The 
> resource version no longer exists - this has to be handled by the caller". 
> This happens to long-running watches when the client reconnects with a 
> resource version that etcd has compacted (or after an API server restart);
> * the reconnect limit ({{watchReconnectLimit}}) is reached, "Exhausted 
> reconnects".
> After that the consumer receives no event anymore: one ERROR log line, the 
> route stays started, health checks are green.
> h3. Reproduction
> {{KubernetesPodsConsumerWatchClosedTest}} (fabric8 mock server, as the 
> producer tests use): the first watch request of {{kubernetes-pods}} gets an 
> {{ERROR}} event with a 410 {{Status}}, a second watch request would get an 
> {{ADDED}} pod. On main the consumer logs "too old resource version" and never 
> watches again: {{mock://result Received message count. Expected: <1> but was: 
> <0>}}.
> h3. Proposed fix
> {{onClose}} with an exception runs the watch task of the consumer again 
> ({{KubernetesHelper.watchAgain}}: only when the consumer is not stopping and 
> its executor is running). A new watch without resource version starts with 
> the current state (the API server sends {{ADDED}} for the existing resources, 
> as when the route starts), then the changes. A close without exception (the 
> consumer closing its own watch on stop) is unchanged. The same two lines in 
> the 11 consumers. With the fix the new test passes and the camel-kubernetes 
> unit tests pass (156 tests; the ITs need a cluster).
> Affected: 4.14.x, 4.18.x and main (same code).
> Not in scope: the consumers watch without first listing, so changes between 
> the closed watch and the new one are only seen as the state of the new watch; 
> a full fix would use informers. Also, when the new watch cannot be started 
> (for example the API server is still unreachable after the reconnect limit 
> was reached), the task fails as it does when the route starts.
> Duplicate check (2026-10-02): JIRA text "kubernetes" with 
> "watch"/"onClose"/"reconnect" (16 issues): only CAMEL-11020 (closing the 
> watchers on stop), CAMEL-13978 and features. GitHub pull requests "kubernetes 
> watch", "kubernetes consumer watch": none.
> _Filed with Claude Code on behalf of allthingssecurity._



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to