yunfengzhou-hub commented on PR #1138:
URL: https://github.com/apache/flink-agents/pull/1138#issuecomment-6057176401

   > I have a proposal for the recovery consistency model.
   > 
   > Given that:
   > 
   > 1. Internal sub-agent recovery can provide the same recovery strength as 
durable execution under deterministic replay.
   > 2. `ActionState` serves an action-level role analogous to reconciliation: 
one record is initialized before execution, updated as durable calls finish, 
and marked completed with the action's memory updates and emitted events when 
the action finishes.
   > 3. The deterministic attributes of `InternalSubagentCallEvent`, including 
the sub-agent identity, form the event-identity component of the `ActionState` 
key. Together with the same business key, sequence number, and action identity, 
replay can therefore locate the original `ActionState` entry after failover.
   > 
   > We do not need to restore child `ActionTask`s. We can retain only 
unfinished root callers and let them issue the calls again. The regenerated 
child tasks then reconcile against `ActionState`: an incomplete action reruns 
and reuses its persisted durable-call results, while a completed action skips 
its body and replays the framework-visible events needed to reconstruct the 
remaining action graph. Calls to a deeper-level sub-agent initiated by that 
completed action can be skipped because their results have already been 
consumed by the action. Without `ActionState`, the unfinished subtree is 
recomputed.
   > 
   > This makes the recovery strategy clear and simple. I have sent the code to 
@yunfengzhou-hub for review. Do you think this strategy is feasible?
   
   @pltbkd  Feasible — adopted and implemented(co-authored to you, now pushed 
to the PR). It follows your model faithfully: drop child tasks on resume, 
replay from the root caller, and reconcile completed actions against their 
persisted `ActionState`. It also closes @purushah 's concern #3 (partial output 
after failover) as a side effect — a yield now holds same-call and ordinary 
events until the action finishes, so `ActionState` captures them and replay 
reconstructs the full output instead of losing mid-yield events.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to