joseluisll opened a new pull request, #8748: URL: https://github.com/apache/hbase/pull/8748
https://issues.apache.org/jira/browse/HBASE-30468 `TestProcDispatcher.testRetryLimitOnConnClosedErrors` waits until a `ServerCrashProcedure` shows up in `getMasterProcedureExecutor().getProcedures()`. An SCP does not wait for a client ack, so once it finishes, the next `CompletedProcedureCleaner` run (every 30s) evicts it. When that run falls between the SCPs finishing and the test's next poll, the test never sees an SCP and times out at line 141. This is why the test is still flaky on master after HBASE-30265. With this change, the test reads the master's SCP submitted counter (`getServerCrashProcMetrics().getSubmittedCounter()`) before injecting errors and waits for it to increase. The counter is not affected by eviction. The other wait conditions and the hbck check are unchanged. ### Testing - With `hbase.procedure.cleaner.interval=1000` set in the test, the old version fails every time and the new one passes. - Looping both versions locally with default settings: the old one failed 6 of 39 runs, the new one 0 of 39. - `TestProcDispatcher` passes on this branch. 🤖 Generated with [Claude Code](https://claude.com/claude-code) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
