SEZ9 commented on issue #12585: URL: https://github.com/apache/seatunnel/issues/12585#issuecomment-5965495753
Thanks @goutamadwant for reproducing this on both PostgreSQL 18 (`idle_replication_slot_timeout`) and 17 (`max_slot_wal_keep_size`) — that confirms the two failure modes described in the report: the restore either hangs in Debezium's slot lookup or fails after several minutes with a generic replication error, neither of which tells the user the slot was invalidated. Good to see the fix PR you linked. A few things I'd like to see covered there so we can close this issue against it: 1. The pre-flight check should handle both invalidation causes (`wal_status='lost'` from WAL removal and from idle timeout) and fail fast with a dedicated error code (the proposed `POSTGRES-04`), with a message that names the slot and the reason so users know to drop/recreate it rather than retry. 2. Please make sure the check runs before Debezium's slot lookup on job restore, so the ~30 min hang path is avoided entirely and Zeta doesn't burn retries on an unrecoverable slot. 3. Tests for both the PG 17 and PG 18 scenarios you reproduced (or at least a unit test for the slot-status query plus a note on which versions were verified manually). 4. If you still have them handy, could you paste the exact Debezium error text you saw in each case into this issue? That helps anyone searching for the symptom land here. I'll follow up on the PR once those are in. <!-- streview-comment:1489 --> -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
