Ashfaqbs opened a new issue, #1105:
URL: https://github.com/apache/flink-agents/issues/1105

   ### Search before asking
   
   - [x] I searched in the 
[issues](https://github.com/apache/flink-agents/issues) and found nothing 
similar.
   
   ### Description
   
   `ChatModelInvoker.chatWithRetries()`'s retry loop 
(`plan/src/main/java/org/apache/flink/agents/plan/actions/ChatModelInvoker.java`,
 around line 155: `for (int attempt = 0; attempt < numRetries + 1; attempt++)`) 
uses a plain local loop variable for the attempt count. It is not stored in 
Flink's checkpointed state.
   
   That means `numRetries` only bounds retries within one uninterrupted 
execution of `chatWithRetries`. If the job restarts mid-retry-sequence — a 
checkpoint recovery triggered by something unrelated, a rolling deploy, a 
TaskManager failure — the method call restarts from the top on resume, and the 
attempt counter resets to zero. For a chat call using 
`ErrorHandlingStrategy.RETRY` against a *permanent* failure (not a transient 
one the retry was meant to ride out), the configured retry ceiling isn't 
actually a ceiling across the job's lifetime — only within a single 
non-interrupted attempt sequence. A permanently-failing call could in principle 
keep retrying indefinitely across enough restarts.
   
   **Question for maintainers:** is this intentional? Giving a fresh retry 
budget after a restart is a defensible design choice — the restart itself may 
indicate the environment changed, so discarding prior attempt history isn't 
obviously wrong. I haven't found anything in the code, docs, or existing issues 
stating which behavior is intended, and wanted to ask before assuming either 
way.
   
   If it turns out to be an unintended gap, the fix would presumably mean 
persisting the attempt count in the same durable/keyed state mechanism 
`ActionState` already uses elsewhere, rather than a plain loop variable — happy 
to help build that if it's the direction the team wants.
   
   ### Version and environment
   
   Current `main` (`0.4-SNAPSHOT`). Applies to the Java runtime whenever 
durable execution is enabled with `ErrorHandlingStrategy.RETRY`, independent of 
the specific durable store backend.
   
   ### Are you willing to submit a PR?
   
   - [ ] I'm willing to submit a PR!
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to