surafel58 commented on PR #11827: URL: https://github.com/apache/seatunnel/pull/11827#issuecomment-5366357668
Thanks @SEZ9. I have taken the documentation half of both findings on this PR, and consolidated the two behavioral changes into the follow-up #11911. Docs (commit 2ce10e84, EN and ZH): added a "Checkpoint Flush and Failure Handling" section covering both points: - Replay safety depends on the receiver accepting an exact duplicate (same labels, timestamp, value); a receiver that rejects a same-timestamp-different-value or out-of-order sample (Prometheus TSDB, and Cortex/Mimir/Thanos, return 400) can fail the replayed flush. - Transient failures fail the checkpoint; note to raise the engine's tolerableCheckpointFailureNumber on Spark/Flink, with the bounded retry tracked in #11911. Behavioral fixes -> #11911: both remaining items are flush() error-handling changes, so I grouped them in the follow-up rather than expanding this PR: (a) bounded retry with backoff for retryable failures (5xx/429/transport), and (b) treat duplicate/out-of-order 400 rejections as delivered so a replay after restore does not loop the checkpoint. I have widened #11911 to cover both and will implement it right after this PR merges (both change flush(), so doing it here would conflict). F3 is resolved, and F1/F2 now have the documentation you named as the minimum, with the behavioral work scoped and tracked in #11911. Would you be comfortable merging this PR with the behavioral changes handled in that follow-up? Happy to bring any of it into this PR instead if you would prefer. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
