wombatu-kun opened a new pull request, #9920:
URL: https://github.com/apache/paimon/pull/9920

   ### Purpose
   
   `readVectored` already reads a single range inline, but only under 
`sequentialReadFallback`. That flag guards `fallbackToReadSequence`, which 
calls `SeekableInputStream.seek`, so it really answers "may I move the stream 
position?". A caller that shares one stream across threads has to turn the flag 
off, and thereby also loses the inline path: a lone 4 KiB read is submitted to 
`IO_THREAD_POOL` through a `BlockingExecutor` and joined back. This reads it on 
the calling thread through the same `readSingleRange` the pool worker would 
have run. A single range never coalesced with anything and never reached 
`splitBatches`, so the I/O is identical and only the hand-off is gone.
   
   **When the new branch is taken.** Exactly one range, the calling thread is 
not interrupted, and the existing fallback did not already claim the call - 
that is, either `sequentialReadFallback` is off, or the readable is not a 
`SeekableInputStream` and so could never have used the sequential fallback 
anyway. Parquet and ORC keep the flag on and hand in a `SeekableInputStream`, 
so they are untouched.
   
   **How often.** The only production caller that turns the flag off is the 
vector index reader, 
`NativeVectorGlobalIndexReader.SeekableStreamVectorIndexInput`, and 
single-range callbacks dominate its traffic: paimon-vindex reads the index 
header at open, streams resident sections chunk by chunk during 
`optimizeForSearch()`, preloads the DiskANN adjacency, and reads every probed 
IVF cluster payload with a one-range `pread`. An IVF search with `nprobe = 32` 
pays 32 hand-offs per shard per query on top of the reads themselves.
   
   An interrupted caller keeps the executor path on purpose: `FileChannel` is 
an `AbstractInterruptibleChannel`, so an inline read on an interrupted thread 
closes the channel and breaks the stream for every other reader sharing it, 
while `BlockingExecutor.submit` fails fast on `semaphore.acquire()` and leaves 
it intact.
   
   One 4 KiB range through `LocalFileIO` over a 128 MiB page-cached file at 
random offsets, with the vector index reader's options, 20k warm-up plus 100k 
measured calls per JVM, median of 3 runs before and 8 after (JDK 8):
   
   | per call | before | after |
   |---|---|---|
   | mean | 4.09 us | 1.56 us |
   | p50 | 2.02 us | 1.39 us |
   | p90 | 10.67 us | 1.98 us |
   | p99 | 17.41 us | 3.30 us |
   
   What is saved is the hand-off, so it is the same number of microseconds on a 
slower device but a smaller share of the total.
   
   ### Tests
   
   `VectoredReadUtilsTest.testSingleRangeIsReadOnCallingThread` records the 
thread that served `pread` and asserts it is the calling thread; it fails 
without the new branch. `testInterruptedCallerDoesNotReadInline` asserts an 
interrupted caller performs no inline read; it fails without the interrupt 
guard. Existing tests are unchanged.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to