krylosov-aa opened a new pull request, #2054: URL: https://github.com/apache/cloudberry/pull/2054
### What does this PR do? Problem: during teardown, the test restarts content 1 with `pg_ctl restart -w` and then immediately resets the `checkpoint` fault on all primaries. `pg_ctl -w` may return as soon as crash recovery starts, even if the segment is not ready to accept connections yet. `gp_inject_fault` connects directly to each segment and does not retry, so the reset can sometimes fail with an error like `connection to ... failed`. Regular queries do not have this issue because the dispatcher retries gang creation. Solution: reset the `checkpoint` fault before the second restart. The second restart is only used for teardown and does not need the fault. Also, content 1 loses its fault state after the restart anyway. ### Type of Change - [x] Bug fix (non-breaking change) ### Test Plan Local checks on Linux ARM64 ### Checklist - [x] Followed [contribution guide](https://cloudberry.apache.org/contribute/code) - [ ] Requested review from [cloudberry committers](https://github.com/orgs/apache/teams/cloudberry-committers) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
