goutamadwant opened a new issue, #12585: URL: https://github.com/apache/seatunnel/issues/12585
### Search before asking - [X] I had searched in the [issues](https://github.com/apache/seatunnel/issues?q=is%3Aissue+label%3A%22bug%22) and found no similar issues. ### What happened PostgreSQL 13+ can invalidate a logical replication slot (`max_slot_wal_keep_size`, and `idle_replication_slot_timeout` on PostgreSQL 18). `pg_replication_slots` then shows `wal_status = 'lost'` and, on 17+, an `invalidation_reason`. Postgres-CDC does not check this and hands the slot to Debezium: - `wal_removed`: `restart_lsn` is NULL, so Debezium's slot-state loop retries 900 × 2 s (`Cannot obtain valid replication slot ... concurrent tx probably blocks taking snapshot`). The job stays RUNNING for about 30 minutes per attempt. - `idle_timeout`: the replication stream fails after 6 × 10 s retries with `can no longer access replication slot`, then the engine restores the pipeline and repeats this 3 times. Neither message says the slot is invalidated or what to do about it. **Steps (PostgreSQL 18)** 1. `ALTER SYSTEM SET idle_replication_slot_timeout = '1min'; SELECT pg_reload_conf();` 2. Start the job below, let it stream, then stop it with a savepoint. 3. Wait until `SELECT wal_status, invalidation_reason FROM pg_replication_slots` shows `lost` / `idle_timeout` (after a checkpoint). 4. Restore the job from the savepoint. On PostgreSQL 17, use `max_slot_wal_keep_size = '1MB'` and generate WAL while the job is stopped (`wal_removed`). **Before / After** (dev, Zeta, default `job.retry.times = 3`) | Case | dev | with fix | |---|---|---| | PG 18.6 idle_timeout, restore | FAILED after ~4 min 20 s, Debezium error only | FAILED in 13 s with `POSTGRES-04 (reason: idle_timeout)` | | PG 17.9 wal_removed, restore | still RUNNING after 152 s at attempt 70/900 (~30 min per attempt) | FAILED in 13 s with `POSTGRES-04 (reason: wal_removed)` | | PG 17.9 slot invalidated while running | hung 4+ min after the pipeline restore | FAILED in 15 s | | Healthy slot restore, fresh start, snapshot-only | works | unchanged | ### SeaTunnel Version dev (4c4fd615d), 3.0.0 ### SeaTunnel Config ```conf env { parallelism = 1, job.mode = "STREAMING", checkpoint.interval = 5000 } source { Postgres-CDC { url = "jdbc:postgresql://localhost:5432/shop", username = "postgres", password = "***", database-names = ["shop"], schema-names = ["public"], table-names = ["shop.public.orders"], slot.name = "st_idle" } } sink { Jdbc { url = "jdbc:postgresql://localhost:5432/shop", driver = "org.postgresql.Driver", username = "postgres", password = "***", generate_sink_sql = true, database = "shop", table = "public.orders_sink", primary_keys = ["id"] } } ``` ### Running Command ```shell bin/seatunnel.sh --config pg_cdc.conf -s <jobId> bin/seatunnel.sh --config pg_cdc.conf -r <jobId> ``` ### Error Exception ```log # idle_timeout org.postgresql.util.PSQLException: ERROR: can no longer access replication slot "st_idle" Detail: This replication slot has been invalidated due to "idle_timeout". # wal_removed (job stays RUNNING) Cannot obtain valid replication slot 'st_wal' for plugin 'pgoutput' and database 'shop' [during attempt 70 out of 900, concurrent tx probably blocks taking snapshot. ``` ### Zeta or Flink or Spark Version Zeta ### Java or Scala Version Java 8, Java 11 ### Screenshots _No response_ ### Are you willing to submit PR? - [X] Yes I am willing to submit a PR! ### Code of Conduct - [X] I agree to follow this project's [Code of Conduct](https://www.apache.org/foundation/policies/conduct) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
