lilei1128 commented on PR #9987: URL: https://github.com/apache/paimon/pull/9987#issuecomment-5759482462
> Requirement fit: SUPPORTED (triage: GO) > > Implementation: FINDINGS > > [P2] Preserve the cheap file-index rejection before loading the deletion vector > > `DataEvolutionSplitRead#createReader` now calls `readDeletionVector` at lines 253-254 before `skipByFileIndex`. `DeletionVector.Factory#create` opens and reads the DV sidecar, so every merged row-id group with a DV now pays that storage read even when the ordinary file index would have rejected the group without opening any data/DV file. This regresses the most selective scans: a table with many merged groups can add one remote sidecar read per group that used to be skipped cheaply. > > Please keep a first file-index-only rejection before loading the DV (return immediately if it is already `SKIP`), and only load/intersect the DV when the raw index still has candidates. The new deleted-candidate behavior and its test can remain as the second-stage evaluation. Fixed by adding a two-stage check for merged groups: run the regular file-index check first, then load and apply the deletion vector only when candidates remain. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
