wangyong9999 commented on PR #314: URL: https://github.com/apache/paimon-cpp/pull/314#issuecomment-5700853683
Thanks @lxy-9602. Added `parquet.read.enable-offset-index-cache`, disabled by default, in 8f56331b. With it off, parsed OffsetIndex objects are not retained. The per-leaf direct-plan decision reuse remains independent of the switch. Filesystem caching cannot remove this particular CPU cost: Arrow 17 invokes OffsetIndex::Make on each lookup after obtaining the index bytes. With warm local files, the new reproducible benchmark measures 2.445 → 0.892 ms for a single-row read over 6,250 pages (same 162,389 bytes read); a 98-page case improves only 3.5%, and full scan is unchanged. Separate perf samples show stacks containing OffsetIndex::Make dropping from 79.9% to 45.0%. The first parse still happens. This makes the feature useful for selective reads over many pages, not a general improvement for wide scans. I have not benchmarked thousand-column projections, and those should remain off by default because retained memory scales with accessed columns and pages. The PR includes the measurement setup, controls, reproduction command and memory estimate. Parquet 227/227, read integration 302/302 and read-with-index 68/68 pass. A focused ASan benchmark exposed an exit-order problem in the existing static fixture cleanup; fixtures now release before filesystem singletons, and all six cases pass the probe. CI is pending for the new revision. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
