SteNicholas commented on code in PR #3801:
URL: https://github.com/apache/celeborn/pull/3801#discussion_r3793872934
##########
cpp/celeborn/network/MessageDispatcher.cpp:
##########
@@ -176,12 +218,88 @@ folly::Future<std::unique_ptr<Message>>
MessageDispatcher::operator()(
return p.getFuture();
});
- this->pipeline_->write(std::move(toSendMsg));
+ // Observe the write future, like Java's TransportClient does with
+ // StdChannelListener. wangle's AsyncSocketHandler::write returns an
+ // already-failed future when the socket is no longer good, and otherwise
+ // fails it from AsyncTransport::WriteCallback::writeErr. Neither necessarily
+ // flips closed_ before the check below, so dropping the future would leave
+ // the promise pending until the request timeout.
+ this->pipeline_->write(std::move(toSendMsg))
+ .thenError([this, requestId](const folly::exception_wrapper& e) {
Review Comment:
@yugan95, guard the dispatcher lifetime in this asynchronous continuation.
The write future may complete after `MessageDispatcher` has been destroyed, but
the callback captures a raw `this`. The default `TransportClient` destruction
order destroys `dispatcher_` before `client_` and its pipeline, so a delayed
write failure during pipeline teardown can call `failPendingRequest` through a
dangling pointer. Please keep lifetime-safe shared state, cancel/drain these
continuations before destroying the dispatcher, or otherwise guarantee the
pipeline is quiescent first. A test whose mock write returns a deferred promise
and fails it during teardown would exercise this path; the fetch continuation
has the same issue.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]