SEZ9 commented on issue #12585:
URL: https://github.com/apache/seatunnel/issues/12585#issuecomment-5965495753

   Thanks @goutamadwant for reproducing this on both PostgreSQL 18 
(`idle_replication_slot_timeout`) and 17 (`max_slot_wal_keep_size`) — that 
confirms the two failure modes described in the report: the restore either 
hangs in Debezium's slot lookup or fails after several minutes with a generic 
replication error, neither of which tells the user the slot was invalidated.
   
   Good to see the fix PR you linked. A few things I'd like to see covered 
there so we can close this issue against it:
   
   1. The pre-flight check should handle both invalidation causes 
(`wal_status='lost'` from WAL removal and from idle timeout) and fail fast with 
a dedicated error code (the proposed `POSTGRES-04`), with a message that names 
the slot and the reason so users know to drop/recreate it rather than retry.
   2. Please make sure the check runs before Debezium's slot lookup on job 
restore, so the ~30 min hang path is avoided entirely and Zeta doesn't burn 
retries on an unrecoverable slot.
   3. Tests for both the PG 17 and PG 18 scenarios you reproduced (or at least 
a unit test for the slot-status query plus a note on which versions were 
verified manually).
   4. If you still have them handy, could you paste the exact Debezium error 
text you saw in each case into this issue? That helps anyone searching for the 
symptom land here.
   
   I'll follow up on the PR once those are in.
   
   <!-- streview-comment:1489 -->


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to