DanielLeens commented on issue #12127:
URL: https://github.com/apache/seatunnel/issues/12127#issuecomment-5554212730
Two more sightings today, with server-side detail that narrows the failure
mode:
- PR #11727, fork run `abdessalems/seatunnel` 33978895542, job
`all-connectors-it-2 (11)` (job 101340798239): `testFakeSourceToCouchbaseSink`
failed on the Zeta leg; the Flink 1.18 and 1.20 legs hit the same window but
recovered through the job's restart.
- PR #11077, fork run `hesam-oxe/seatunnel` 33971836407, job
`all-connectors-it-2 (8)` (job 101321961110): failed on the Flink 1.15, Flink
1.20 and Spark 2.4 legs.
In every case the writer's `waitUntilReady` times out in stage
`WAIT_FOR_CONFIG` after 30 s, and the KV endpoint keeps failing SASL for the
whole window:
```
[com.couchbase.io][SaslAuthenticationFailedEvent][95ms] Authentication
Failure - Potential causes: invalid credentials or if LDAP is enabled ensure
PLAIN SASL mechanism is exclusively used ...
{"bucket":"test_bucket","remote":"e2e_couchbase:11210","status":"UNKNOWN","type":"KV"}
[com.couchbase.endpoint][EndpointConnectionFailedEvent][203ms] Connect
attempt 1 failed because of AuthenticationFailureException ...
UnambiguousTimeoutException: WaitUntilReady timed out in stage
WAIT_FOR_CONFIG (spent PT30.001S in that stage) {"bucket":"test_bucket", ...
"services":{"mgmt":[{"state":"connected",
"remote":"e2e_couchbase:8091"}],"kv":[{"lastConnectAttemptFailure":"Authentication
Failure ..."}]}}
```
The management port accepts the credentials (`mgmt: connected`), only the KV
`SELECT_BUCKET` step is refused, and the same credentials work from the test
JVM (`CouchbaseIT#startUp` already passed `waitUntilReady(2 min)` and ran DDL
before the job started). The KV service therefore intermittently reports the
bucket as not selectable for well over 30 s after the IT has verified
readiness, and the hard-coded `Duration.ofSeconds(30)` in `CouchbaseWriter`
with no retry turns that into a job failure. The IT cannot pre-empt this from
the test side: readiness had been confirmed 17-60 s before each failure (the
Flink 1.18 leg failed 17 s after `Couchbase cluster ready`).
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]