goutamadwant opened a new issue, #12585:
URL: https://github.com/apache/seatunnel/issues/12585

   ### Search before asking
   
   - [X] I had searched in the 
[issues](https://github.com/apache/seatunnel/issues?q=is%3Aissue+label%3A%22bug%22)
 and found no similar issues.
   
   ### What happened
   
   PostgreSQL 13+ can invalidate a logical replication slot 
(`max_slot_wal_keep_size`, and `idle_replication_slot_timeout` on PostgreSQL 
18). `pg_replication_slots` then shows `wal_status = 'lost'` and, on 17+, an 
`invalidation_reason`. Postgres-CDC does not check this and hands the slot to 
Debezium:
   
   - `wal_removed`: `restart_lsn` is NULL, so Debezium's slot-state loop 
retries 900 × 2 s (`Cannot obtain valid replication slot ... concurrent tx 
probably blocks taking snapshot`). The job stays RUNNING for about 30 minutes 
per attempt.
   - `idle_timeout`: the replication stream fails after 6 × 10 s retries with 
`can no longer access replication slot`, then the engine restores the pipeline 
and repeats this 3 times.
   
   Neither message says the slot is invalidated or what to do about it.
   
   **Steps (PostgreSQL 18)**
   
   1. `ALTER SYSTEM SET idle_replication_slot_timeout = '1min'; SELECT 
pg_reload_conf();`
   2. Start the job below, let it stream, then stop it with a savepoint.
   3. Wait until `SELECT wal_status, invalidation_reason FROM 
pg_replication_slots` shows `lost` / `idle_timeout` (after a checkpoint).
   4. Restore the job from the savepoint.
   
   On PostgreSQL 17, use `max_slot_wal_keep_size = '1MB'` and generate WAL 
while the job is stopped (`wal_removed`).
   
   **Before / After** (dev, Zeta, default `job.retry.times = 3`)
   
   | Case | dev | with fix |
   |---|---|---|
   | PG 18.6 idle_timeout, restore | FAILED after ~4 min 20 s, Debezium error 
only | FAILED in 13 s with `POSTGRES-04 (reason: idle_timeout)` |
   | PG 17.9 wal_removed, restore | still RUNNING after 152 s at attempt 70/900 
(~30 min per attempt) | FAILED in 13 s with `POSTGRES-04 (reason: wal_removed)` 
|
   | PG 17.9 slot invalidated while running | hung 4+ min after the pipeline 
restore | FAILED in 15 s |
   | Healthy slot restore, fresh start, snapshot-only | works | unchanged |
   
   ### SeaTunnel Version
   
   dev (4c4fd615d), 3.0.0
   
   ### SeaTunnel Config
   
   ```conf
   env { parallelism = 1, job.mode = "STREAMING", checkpoint.interval = 5000 }
   source { Postgres-CDC { url = "jdbc:postgresql://localhost:5432/shop", 
username = "postgres", password = "***",
     database-names = ["shop"], schema-names = ["public"], table-names = 
["shop.public.orders"], slot.name = "st_idle" } }
   sink { Jdbc { url = "jdbc:postgresql://localhost:5432/shop", driver = 
"org.postgresql.Driver", username = "postgres",
     password = "***", generate_sink_sql = true, database = "shop", table = 
"public.orders_sink", primary_keys = ["id"] } }
   ```
   
   ### Running Command
   
   ```shell
   bin/seatunnel.sh --config pg_cdc.conf -s <jobId>
   bin/seatunnel.sh --config pg_cdc.conf -r <jobId>
   ```
   
   ### Error Exception
   
   ```log
   # idle_timeout
   org.postgresql.util.PSQLException: ERROR: can no longer access replication 
slot "st_idle"
     Detail: This replication slot has been invalidated due to "idle_timeout".
   
   # wal_removed (job stays RUNNING)
   Cannot obtain valid replication slot 'st_wal' for plugin 'pgoutput' and 
database 'shop' [during attempt 70 out of 900, concurrent tx probably blocks 
taking snapshot.
   ```
   
   ### Zeta or Flink or Spark Version
   
   Zeta
   
   ### Java or Scala Version
   
   Java 8, Java 11
   
   ### Screenshots
   
   _No response_
   
   ### Are you willing to submit PR?
   
   - [X] Yes I am willing to submit a PR!
   
   ### Code of Conduct
   
   - [X] I agree to follow this project's [Code of 
Conduct](https://www.apache.org/foundation/policies/conduct)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to