pltbkd commented on PR #1138:
URL: https://github.com/apache/flink-agents/pull/1138#issuecomment-6040946018

   I have a proposal for the recovery consistency model.
   
   Given that:
   
   1. Internal sub-agent recovery can provide the same recovery strength as 
durable execution under deterministic replay.
   2. `ActionState` serves an action-level role analogous to reconciliation: 
one record is initialized before execution, updated as durable calls finish, 
and marked completed with the action's memory updates and emitted events when 
the action finishes.
   3. The deterministic attributes of `InternalSubagentCallEvent`, including 
the sub-agent identity, form the event-identity component of the `ActionState` 
key. Together with the same business key, sequence number, and action identity, 
replay can therefore locate the original `ActionState` entry after failover.
   
   We do not need to restore child `ActionTask`s. We can retain only unfinished 
root callers and let them issue the calls again. The regenerated child tasks 
then reconcile against `ActionState`: an incomplete action reruns and reuses 
its persisted durable-call results, while a completed action skips its body and 
replays the framework-visible events needed to reconstruct the remaining action 
graph. Calls to a deeper-level sub-agent initiated by that completed action can 
be skipped because their results have already been consumed by the action. 
Without `ActionState`, the unfinished subtree is recomputed.
   
   This makes the recovery strategy clear and simple. I have sent the code to 
@yunfengzhou-hub  for review. Do you think this strategy is feasible?
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to