joseluisll opened a new pull request, #8792: URL: https://github.com/apache/hadoop/pull/8792
### Description of PR https://issues.apache.org/jira/browse/HDFS-17987 `TestDataNodeLifeline` is flaky because of two separate problems in the test. Both appeared together in the CI run for https://github.com/apache/hadoop/pull/8704. **1. Mockito stubbing race.** `setup()` gave the NameNode and lifeline spies to the running `BPServiceActor` before the tests stubbed them, so `doAnswer(...).when(namenode).sendHeartbeat(...)` ran while the heartbeat thread was calling the same spy every second. Mockito stubbing is not safe against a concurrent call on the same mock. A heartbeat landing mid-stubbing kills the actor, either with an `AssertionError` from `InvocationContainerImpl.setMethodForStubbing` or by getting `null` back and failing `assert resp != null` in `BPServiceActor#offerService`. The block pool is then removed, and `testSendLifelineIfHeartbeatBlocked` fails when it reconfigures the data dirs: java.lang.IndexOutOfBoundsException: Index 0 out of bounds for length 0 at org.apache.hadoop.hdfs.server.datanode.DataStorage.prepareVolume(DataStorage.java:334) `setup()` now only creates the spies. A new `installSpies()` hands them to the actor, and the two tests that stub them (`testSendLifelineIfHeartbeatBlocked`, `testNoLifelineSentIfHeartbeatsOnTime`) call it once their stubs are in place. **2. Leaked fault injector.** `testHeartbeatAndLifelineOnError` sets the static `BlockManagerFaultInjector.instance` to an injector that throws from every `BlockManager#updateHeartbeat` and never restores it. When surefire reruns a failed test of this class in the same JVM, the new cluster's DataNode never gets a heartbeat through, and `setup()` waits in `MiniDFSCluster#waitActive` until the 600 s timeout. The injector is now restored in `finally`. ### How was this patch tested? All runs on JDK 17, `mvn -pl hadoop-hdfs-project/hadoop-hdfs test -Dtest=TestDataNodeLifeline`, one fresh fork per run, no reruns: | | Runs | Failed | |---|---|---| | trunk | 10 | 4 (all from the stubbing race) | | this patch | 10 | 0 | Rerun hang: with `testSendLifelineIfHeartbeatBlocked` temporarily forced to fail and `-Dsurefire.rerunFailingTestsCount=2` (as in CI), trunk reproduced the CI result exactly. Both reruns hit `setup() timed out after 600 seconds`, and the NameNode logged the injector's `UnknownError` on every heartbeat. With this patch both reruns get through `setup()` and finish in about 1.3 s. Checkstyle: no new violations. The two existing `RedundantModifier` warnings in the file are on untouched lines. ### For code changes: - [x] Does the title of this PR start with the corresponding JIRA issue id (e.g. 'HADOOP-17799. Your PR title ...')? - [ ] Object storage: Have the integration tests been executed and the endpoint declared according to the connector-specific documentation? - [ ] If adding new dependencies to the code, are these dependencies licensed in a way that is compatible for inclusion under [ASF 2.0](http://www.apache.org/legal/resolved.html#category-a)? - [ ] If applicable, have you updated the `LICENSE`, `LICENSE-binary`, `NOTICE-binary` files? ### AI Tooling Contains content generated by Claude Code. - [x] The PR includes the phrase "Contains content generated by <tool>" where <tool> is the name of the AI tool used. - [x] My use of AI contributions follows the ASF legal policy https://www.apache.org/legal/generative-tooling.html 🤖 Generated with [Claude Code](https://claude.com/claude-code) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
