SteNicholas opened a new pull request, #420:
URL: https://github.com/apache/paimon-cpp/pull/420

   <!-- PR titles must follow Conventional Commits: <type>(<optional-scope>): 
<description> -->
   
   ### Purpose
   
   <!-- Linking this pull request to the issue -->
   Linked issue: close #419
   
   <!-- What is the purpose of the change -->
   `ConcatBatchReader` never called `Warmup()`, so the read paths that 
concatenate data files without
   merging issued the first read of every file only after the previous file had 
reached EOF. #286
   overlapped that first read on the merge-on-read path through 
`ConcatKeyValueRecordReader` and
   `LoserTree`; this PR does the same in `ConcatBatchReader`, which 
append-table reads, append
   compaction (`AppendOnlyFileStoreWrite::CreateFilesReader` through 
`RawFileSplitRead`) and
   raw-convertible primary-key splits read through.
   
   - Before reading a batch, `NextBatchWithBitmap()` warms up the current child 
and the
     `kWarmupLookahead = 1` child after it, as 
`ConcatKeyValueRecordReader::NextBatch()` does.
     `NextBatch()` goes through `NextBatchWithBitmap()` and gets the same 
behavior.
   - `Warmup()` is declared on `FileBatchReader` while the children are held as 
`BatchReader`, so each
     child's `FileBatchReader*` is resolved once at construction and cleared 
when the child is
     released. Children that are not `FileBatchReader`s, such as per-split 
readers, are skipped, and
     the lookahead window still counts them.
   - The public `BatchReader` is unchanged.
   
   Limits, also documented in the user guide and in #419:
   
   - Only files that are actually prefetched are warmed. The adaptive strategy 
reads a file without
     prefetching when its first read range holds more batches than one prefetch 
queue can buffer.
     Append compaction uses a queue of three batches, so it only warms files 
whose first read range
     spans no more than three full batches.
   - `WarmupLevel::RAW` only overlaps I/O with the read-ahead cache enabled and 
a file system whose
     `ReadAsync` is asynchronous. The local file system reads synchronously, so 
there `RAW` warmup adds
     latency to the current batch; `WarmupLevel::DECODED` or 
`WarmupLevel::NONE` avoid that.
   - Warmup does not cross split boundaries.
   
   ### Tests
   
   <!-- List UT and IT cases to verify this change -->
   UT, `src/paimon/common/reader/concat_batch_reader_test.cpp` (`common_test`):
   
   - `TestWarmupLooksOneReaderAhead`: nothing is warmed before the first read, 
the next child is
     warmed before the current one starts reading, and the lookahead does not 
move on while the read
     is still inside the same file.
   - `TestWarmupSkipsReaderThatIsNotFileBatchReader`: a child that is not a 
`FileBatchReader` is never
     warmed, the lookahead window still counts it, and the returned rows are 
unchanged.
   
   The change has not been compiled or tested locally and relies on CI for both.
   
   ### API and Format
   
   <!-- Does this change affect API in include dir or storage format or 
protocol -->
   No. `ConcatBatchReader` is internal, and no public API, storage format or 
protocol changes.
   
   ### Documentation
   
   <!-- Does this change introduce a new feature -->
   Yes. `docs/source/user_guide/prefetch.rst` gains a "Warmup" section that 
describes the warmup
   levels, where warmup applies, the adaptive-strategy and 
synchronous-file-system limits, and the
   memory cost.
   
   ### Generative AI tooling
   
   Generated-by: Claude Code 2.1.289 (Claude Opus 5.5)
   
   🤖 Generated with [Claude Code](https://claude.com/claude-code)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to