davidzollo opened a new issue, #12127:
URL: https://github.com/apache/seatunnel/issues/12127

   ### Search before asking
   - [x] I had searched in the issues and found no similar issues.
   
   ### What happened
   `CouchbaseIT#testFakeSourceToCouchbaseSink` fails intermittently in CI with 
the **sink** unable to bootstrap its Couchbase client within 30 seconds, while 
the test-side client in the same JVM has been connected to the same container 
for minutes:
   
   ```
   com.couchbase.client.core.error.UnambiguousTimeoutException: WaitUntilReady 
timed out in stage WAIT_FOR_CONFIG (spent PT30.00S in that stage) 
{"bucket":"test_bucket","checkedServices":[],"desiredState":"ONLINE", ...}
   ```
   
   Observed (all on `dev`-based branches that do not touch connector-couchbase):
   - PR #11077, fork run `hesam-oxe/seatunnel` 33971836407, 
`all-connectors-it-2 (8)`: invocations [2] Flink 1.13.6, [4] Flink 1.18.0 and 
[6] the second Zeta leg failed; [1] Zeta, [3] Flink 1.15.3, [5] Flink 1.20.1, 
[7] Spark passed — same container, same bucket, alternating.
   - PR #11757, fork run `waterWang/seatunnel` 33945211319, 
`all-connectors-it-2 (8)`: same `WAIT_FOR_CONFIG` timeout.
   
   So it is not engine-specific: a *fresh* client from inside the engine 
container intermittently does not receive the bucket configuration from 
`e2e_couchbase:8091` within 30 s, even though the container is up and 
previously served jobs.
   
   ### Root cause / gap (`dev`, `CouchbaseWriter.java` ~112–118)
   ```java
   Cluster connectedCluster = Cluster.connect(connectionString, username, 
password);
   
connectedCluster.bucket(options.getBucket()).waitUntilReady(Duration.ofSeconds(30));
   ```
   The readiness budget is hard-coded to 30 seconds, not configurable, and the 
initial connect/bootstrap is not retried — while the same writer already 
exposes `retry.max` / `retry.interval` for writes. A production cluster that is 
busy or rebalancing can legitimately take longer than 30 s to serve a bucket 
config to a new client, and the job then fails outright at writer construction.
   
   The e2e side already hardens container start (#11845: startup attempts + 
3-minute HTTP wait) and wraps DDL in Awaitility retries; there is no 
non-weakening test-side change left for a timeout that lives inside the writer.
   
   ### Suggested fix (connector-couchbase, `src/main`)
   - Make the `waitUntilReady` budget configurable (e.g. `connect.timeout` / 
`ready.timeout`, default ≥ 60 s), and/or
   - Apply the existing `retry.max` / `retry.interval` to the initial 
`Cluster.connect` + `waitUntilReady` bootstrap.
   
   ### SeaTunnel Version
   dev (reproduced 2026-09-05 on Zeta, Flink 1.13.6 and Flink 1.18.0 e2e legs).
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to