gnodet-bot commented on code in PR #27302:
URL: https://github.com/apache/camel/pull/27302#discussion_r4170824117


##########
components/camel-kubernetes/src/main/java/org/apache/camel/component/kubernetes/KubernetesHelper.java:
##########
@@ -90,6 +93,27 @@ public static void close(Runnable runnable, Supplier<Watch> 
watchGetter) {
         }
     }
 
+    /**
+     * Watches again when the Kubernetes client closed the watch of a consumer 
with an error. The client reconnects a
+     * watch by itself after transient errors, and only closes it with an 
exception when it gives up: when the API
+     * server answers 410 Gone because the resource version of the watch is 
too old (which happens to long-running
+     * watches), or when the reconnect limit is reached. The consumer would 
then not receive any event anymore.
+     *
+     * @param consumer the consumer of the watch
+     * @param executor the executor of the consumer
+     * @param task     the task that creates the watch of the consumer
+     */
+    public static void watchAgain(ServiceSupport consumer, ExecutorService 
executor, Runnable task) {
+        if (consumer.isRunAllowed() && executor != null && 
!executor.isShutdown()) {
+            LOG.info("Watching again for {} after its watch was closed", 
consumer);
+            try {
+                executor.submit(task);
+            } catch (RejectedExecutionException e) {
+                LOG.debug("Cannot watch again for {} as it is stopping", 
consumer, e);
+            }
+        }
+    }

Review Comment:
   ⚠️ **No backoff / retry cap:** If the API server keeps returning 410 Gone 
immediately (or any terminal error that triggers `onClose` with an exception), 
this creates a tight retry loop: `submit → run → watch → onClose → submit → run 
→ ...` with zero delay between iterations. In production under API server 
pressure, this could hammer the server and spike CPU.
   
   Consider adding at minimum a short delay before re-watching (e.g. 
`ScheduledExecutorService.schedule` with a 1-5 second delay), or a bounded 
retry count with exponential backoff. The fabric8 client already does its own 
internal retries with backoff before giving up and calling `onClose` — so 
`watchAgain` is the outer retry layer and should have its own protection.
   
   That said, this is an improvement over the status quo (consumer going 
permanently dead), and backoff could be a follow-up enhancement if the 
maintainers agree.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to