wangyong9999 commented on PR #314:
URL: https://github.com/apache/paimon-cpp/pull/314#issuecomment-5700853683

   Thanks @lxy-9602. Added `parquet.read.enable-offset-index-cache`, disabled 
by default, in 8f56331b. With it off, parsed OffsetIndex objects are not 
retained. The per-leaf direct-plan decision reuse remains independent of the 
switch.
   
   Filesystem caching cannot remove this particular CPU cost: Arrow 17 invokes 
OffsetIndex::Make on each lookup after obtaining the index bytes. With warm 
local files, the new reproducible benchmark measures 2.445 → 0.892 ms for a 
single-row read over 6,250 pages (same 162,389 bytes read); a 98-page case 
improves only 3.5%, and full scan is unchanged. Separate perf samples show 
stacks containing OffsetIndex::Make dropping from 79.9% to 45.0%. The first 
parse still happens.
   
   This makes the feature useful for selective reads over many pages, not a 
general improvement for wide scans. I have not benchmarked thousand-column 
projections, and those should remain off by default because retained memory 
scales with accessed columns and pages. The PR includes the measurement setup, 
controls, reproduction command and memory estimate.
   
   Parquet 227/227, read integration 302/302 and read-with-index 68/68 pass. A 
focused ASan benchmark exposed an exit-order problem in the existing static 
fixture cleanup; fixtures now release before filesystem singletons, and all six 
cases pass the probe. CI is pending for the new revision.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to