zhangshenghang opened a new pull request, #12380: URL: https://github.com/apache/seatunnel/pull/12380
### Purpose of this PR Stabilize two E2E timing-sensitive failure signatures that keep burning CI runs on shared runners. Test-only change; no production code is touched. **1. `OpengaussCDCIT.testAddFieldWithRestore` (also `testMultiTableWithRestore`)** Fails with an opaque `ConditionTimeout` at the post-restore data assertion, e.g. from fork CI (reproduced across 5+ runs with the same signature): ``` [ERROR] OpengaussCDCIT.testAddFieldWithRestore:476 ยป ConditionTimeout Assertion condit... ``` The test allows 60s for the restored job to reconnect the WAL replication connection and replicate the inserted row. On loaded shared runners, checkpoint restore + replication-slot reconnect alone can consume most of that budget. This is exactly why PostgresCDCIT already gives the same stage a 180s budget (`RESTORE_ASSERT_TIMEOUT_MILLIS`, comment: *"Restoring the checkpoint and reconnecting the existing replication slot can take longer on shared GitHub runners than the initial CDC startup"*), but the Opengauss sibling suite was left at 60s. Changes in `OpengaussCDCIT`: - Post-restore data assertions use the same 180s `RESTORE_ASSERT_TIMEOUT_MILLIS` budget as `PostgresCDCIT`. - The asynchronously submitted job futures are now captured, and every data await polls `assertJobHasNoAsyncFailure(job)` (same helper/semantics as `PostgresCDCIT`) so a failed or stopped job is reported as the root cause instead of a generic data-comparison timeout. **2. `NebulaGraphIT.startUp`** New failure signature seen on fork CI (`all-connectors-it-1`): ``` [ERROR] NebulaGraphIT.startUp:109 expected: <true> but was: <false> ``` Line 109 is `assertTrue(adminPool.init(...))`. `graphd` answers its HTTP `/status` startup probe (port 19669) before the graph service on port 9669 accepts connections, so the one-shot `NebulaPool.init` + `getSession` can transiently fail immediately after startup. The suite now retries pool initialization and session creation (fresh pool per attempt) for up to 2 minutes before proceeding. ### Does this PR introduce any user-facing interface change? No. ### Related issues/PRs - Same hardening rationale as already-merged PostgresCDCIT restore budget (180s) and `assertJobHasNoAsyncFailure` pattern. - Complements the in-flight engine-side flaky fixes (#12311, #12316, #12313); this PR only makes the affected test suites tolerant and diagnosable. ### Should this PR be highlighted in the ` incompatible-changes.md`? No. ### Verification - `./mvnw spotless:apply` and `./mvnw test-compile` pass for both modules (`connector-cdc-opengauss-e2e`, `connector-nebulagraph-e2e`). - Runtime validation lands with this PR's CI (`all-connectors-it-2` runs `OpengaussCDCIT` on ubuntu runners; nebula suite runs in its part). -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
