comphead commented on code in PR #6805:
URL: https://github.com/apache/datafusion-comet/pull/6805#discussion_r4238684488
##########
spark/src/main/java/org/apache/comet/CometShuffleBlockIterator.java:
##########
@@ -77,17 +81,15 @@ public int hasNext() throws IOException {
}
// Read 16-byte header: clear() resets position=0, limit=capacity,
- // preparing the buffer for channel.read() to fill it
+ // preparing the buffer for readFully() to fill it
headerBuf.clear();
- while (headerBuf.hasRemaining()) {
- int bytesRead = channel.read(headerBuf);
- if (bytesRead < 0) {
- if (headerBuf.position() == 0) {
- close();
- return -1;
- }
- throw new EOFException("Data corrupt: unexpected EOF while reading
batch header");
+ readFully(inputStream, headerBuf, readBufferSize);
Review Comment:
Fixed in `5c025a8acb` as you suggested. I measured it locally with one 1.4
GiB shuffle block read by a native project and cancelled 1 s into the read. The
direct-read task now stops 1 ms after the cancel, with or without an interrupt.
Before the fix it kept reading for about 3.4 s either way, and on `main` it did
so for kills without an interrupt. The table against `main` is in the
description.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]