gripleaf opened a new pull request, #413: URL: https://github.com/apache/paimon-cpp/pull/413
### Purpose Linked issue: none (performance improvement). Sparse Parquet row-group reads currently decode selected top-level fields serially even when Arrow threaded reading is enabled. Decode independent projected fields on the existing Arrow CPU pool and assemble results in projection order. Page-index readers lazily initialize shared buffers, so offset indexes are resolved serially before dispatch and only immutable indexes are shared with field tasks. Every submitted task is drained before returning an error or releasing captured state. Encrypted files, single-field projections, empty ranges and callers already executing on that CPU pool use the serial path; the last case avoids nested submission/wait deadlocks on a bounded pool. This is intended to reduce the serial decode critical path for multi-field batch reads. Task scheduling can cost more than it saves for small requests, and throughput improvements do not imply universal P99 improvements. This draft deliberately makes no universal latency claim. New metrics distinguish the serial/parallel paths and measure per-group wall time so users can assess representative workloads. ### Tests - CMake Debug build with shared libraries and repository-bundled dependencies, `-Wall -Werror`. - `cmake --build build --target paimon-parquet-format-test -j 24` - `./build/debug/paimon-parquet-format-test`: all 230 tests passed. - Added serial/parallel equivalence tests for reversed projections, disjoint sparse selections, dictionary data, partial final pages, nested structs/lists/maps, nulls, partial nested projections and multiple row groups. - Added nested Arrow CPU-pool fallback and metric propagation/snapshot-after-close coverage. - Full-repository pre-commit checks passed using an explicit list of tracked files because the host Git lacks `--deduplicate`. - `git diff --check` passed. ### API and Format No public API, storage format or protocol change. Uses the existing `parquet.read.executor.thread-count` / Arrow threaded-read configuration; does not create another executor or resize it during decoding. Adds reader counters for parallel/serial filtered row groups and projected fields, plus per-group index-preparation and decode wall-time histograms. Submitted task execution overlaps; durations are not sums of worker CPU times. ### Documentation Updated the metrics guide with metric names/units, scheduling and fallback conditions, error accounting and the small-request scheduling tradeoff. Metrics remain readable after closing the reader. ### Generative AI tooling Generated-by: OpenAI Codex (GPT-6) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
