[
https://issues.apache.org/jira/browse/HADOOP-19979?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18118824#comment-18118824
]
ASF GitHub Bot commented on HADOOP-19979:
-----------------------------------------
joseluisll commented on code in PR #8717:
URL: https://github.com/apache/hadoop/pull/8717#discussion_r4093299735
##########
hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/http/TestSSLHttpServerMTLS.java:
##########
@@ -145,6 +147,21 @@ public void testUntrustedClientIsRejected() throws
Exception {
HttpsURLConnection conn = (HttpsURLConnection) url.openConnection();
// presents untrustedCert; server cert is trusted via no-op TrustManager
KeyStoreTestUtil.setAllowAllSSL(conn, untrustedCert, untrustedKeyPair);
- assertThrows(SSLHandshakeException.class, () -> conn.getInputStream());
+ // The server rejects the certificate as soon as it arrives and drops the
+ // connection, which races the client's own last handshake flight, and how
+ // the refusal surfaces depends on who wins and on the protocol in play.
+ // Under TLSv1.2 it is an SSLHandshakeException; when the close wins the
+ // client fails writing that flight and gets a SocketException instead;
+ // under TLSv1.3 the handshake completes client-side before the server has
+ // verified the cert, so the failure lands on the request write as a bare
+ // IOException with no cause. What the server guarantees is that the
+ // request is refused, not which of those the client gets to see. Assert
+ // the refusal and exclude only ConnectException, which would mean we
+ // never reached the server at all.
+ IOException e =
+ assertThrows(IOException.class, () -> conn.getInputStream());
Review Comment:
Adopted, and it is a real hole rather than a theoretical one: with
`getInputStream()` the assertion is satisfied by any `IOException`, and an HTTP
error status raises one, so a 403 or a 500 reached over a handshake the server
should have refused would have passed. `getResponseCode()` returns such a
status instead of throwing, and only throws when no status line was ever read.
Now `assertThrows(IOException.class, () -> conn.getResponseCode())`, comment
updated to say why. Ran it 30x under `hadoop.ssl.enabled.protocols=TLSv1.2` and
30x under `TLSv1.3` on JDK 21: 30/30 both times. (The plumbing was worth
checking - a bogus protocol value fails `testTrustedClientCanConnect`,
confirming the setting reaches the server.)
> Fix four tests that assert on work owned by another thread
> ----------------------------------------------------------
>
> Key: HADOOP-19979
> URL: https://issues.apache.org/jira/browse/HADOOP-19979
> Project: Hadoop Common
> Issue Type: Test
> Components: common, test
> Reporter: Jose Luis López
> Priority: Critical
> Labels: pull-request-available
>
> Four tests assert on, or tear down around, work owned by another thread
> without
> waiting for it or stopping it. All four fail intermittently, and none of the
> failures say anything about the code under test.
> * {{TestSSLHttpServerMTLS.testUntrustedClientIsRejected}} (common) expects an
> SSLHandshakeException, but the server's close races the client's last
> handshake
> flight; when the close wins the client gets a SocketException instead. 7 of 25
> runs fail. Assert that the request is refused rather than which exception
> carries it.
> * {{TestLogAggregationService.testLocalFileDeletionAfterUpload}}
> (nodemanager)
> waits for each log file to go, then asserts on the parent directory with no
> wait; DeletionService removes files before the directories holding them. Hit 4
> of the 60 most recent PRs, including unrelated ones. Wait for the directory
> too.
> * {{TestStandbyCheckpoints.testLastCheckpointTime}} (hdfs) waits for the
> active
> to hold the new image, then reads a standby's checkpoint time, which that wait
> does not cover: any standby may be the uploader, and it stamps
> lastCheckpointTime only after doCheckpoint() returns, so the interval can
> read 0
> against an expected 3000. Wait for that value to move.
> * {{TestTimelineReaderHBaseDown}} (timelineservice-hbase-tests) starts a
> TimelineReaderServer in all five tests and never stops one, leaking the
> TimelineStorageMonitor it schedules: non-daemon threads polling a minicluster
> the test has torn down. The module builds with forkCount 0, so these
> accumulate
> across its eleven test classes. Stop the server in a finally, as every other
> test in the module already does.
> Test-only change.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]