[ 
https://issues.apache.org/jira/browse/HADOOP-19979?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18118825#comment-18118825
 ] 

ASF GitHub Bot commented on HADOOP-19979:
-----------------------------------------

joseluisll commented on code in PR #8717:
URL: https://github.com/apache/hadoop/pull/8717#discussion_r4093301870


##########
hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-timelineservice/src/main/java/org/apache/hadoop/yarn/server/timelineservice/reader/TimelineReaderServer.java:
##########
@@ -147,10 +147,16 @@ private void join() {
 
   @Override
   protected void serviceStop() throws Exception {
-    if (readerWebServer != null) {
-      readerWebServer.stop();
+    try {
+      if (readerWebServer != null) {
+        readerWebServer.stop();
+      }
+    } finally {
+      // super.serviceStop() is what stops the reader and, with it, the
+      // storage monitor's polling executor.  A web server that fails to
+      // stop must not leave those behind.
+      super.serviceStop();

Review Comment:
   Added, in 
`TestTimelineReaderServer.testChildServicesStoppedWhenWebAppStopFails`.
   
   It needed a seam, since `readerWebServer` is private and built inside 
`startTimelineReaderWebApp()`. `serviceStop()` now stops it through a 
package-private `@VisibleForTesting stopTimelineReaderWebApp()`, which the test 
overrides in an anonymous subclass to call `super` and then throw. The test 
asserts `stop()` surfaces the `ServiceStateException` and that every child 
service is still `STOPPED`, `FileSystemTimelineReaderImpl` standing in for the 
HBase reader that owns the storage monitor.
   
   Confirmed it is a regression test and not just a passing one: reverting 
`serviceStop()` to the sequential form gives
   
   ```
   AssertionFailedError: 
org.apache.hadoop.yarn.server.timelineservice.storage.FileSystemTimelineReaderImpl
     was left running by the failed stop ==> expected: <STOPPED> but was: 
<STARTED>
   ```
   
   and it passes with the `finally` restored. Full class: 4/4.





> Fix four tests that assert on work owned by another thread
> ----------------------------------------------------------
>
>                 Key: HADOOP-19979
>                 URL: https://issues.apache.org/jira/browse/HADOOP-19979
>             Project: Hadoop Common
>          Issue Type: Test
>          Components: common, test
>            Reporter: Jose Luis López
>            Priority: Critical
>              Labels: pull-request-available
>
> Four tests assert on, or tear down around, work owned by another thread 
> without
> waiting for it or stopping it. All four fail intermittently, and none of the
> failures say anything about the code under test.
>  * {{TestSSLHttpServerMTLS.testUntrustedClientIsRejected}} (common) expects an
> SSLHandshakeException, but the server's close races the client's last 
> handshake
> flight; when the close wins the client gets a SocketException instead. 7 of 25
> runs fail. Assert that the request is refused rather than which exception
> carries it.
>  * {{TestLogAggregationService.testLocalFileDeletionAfterUpload}} 
> (nodemanager)
> waits for each log file to go, then asserts on the parent directory with no
> wait; DeletionService removes files before the directories holding them. Hit 4
> of the 60 most recent PRs, including unrelated ones. Wait for the directory 
> too.
>  * {{TestStandbyCheckpoints.testLastCheckpointTime}} (hdfs) waits for the 
> active
> to hold the new image, then reads a standby's checkpoint time, which that wait
> does not cover: any standby may be the uploader, and it stamps
> lastCheckpointTime only after doCheckpoint() returns, so the interval can 
> read 0
> against an expected 3000. Wait for that value to move.
>  * {{TestTimelineReaderHBaseDown}} (timelineservice-hbase-tests) starts a
> TimelineReaderServer in all five tests and never stops one, leaking the
> TimelineStorageMonitor it schedules: non-daemon threads polling a minicluster
> the test has torn down. The module builds with forkCount 0, so these 
> accumulate
> across its eleven test classes. Stop the server in a finally, as every other
> test in the module already does.
> Test-only change. 



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to