liangjie3138 opened a new pull request, #10146:
URL: https://github.com/apache/paimon/pull/10146

   ### Purpose
   close https://github.com/apache/paimon/issues/10145
   
   ## Solution
   
     - Plan index queries from metadata without reading index results on the 
JobManager. Attach the plan (including the predicate, row ranges and 
GlobalIndexIOMetas) to data splits and evaluate it in Flink taskmanagers; BTree 
reads are constrained to each split’s row ranges.
     - Preserve index coverage and unindexed-row handling. For mixed indexes, 
an unsupported `AND` branch does not prevent an independent supported branch 
from running distributed; an `OR` query that cannot be safely pruned uses a 
full data scan.
     - Keep the existing path when no supported index is available. The feature 
is disabled by default via `scan.index-distributed-query.enabled`.
   
     We also considered splitting work by index group (the index files sharing 
a row range): query each group in the source stage, shuffle the resulting row 
IDs according to data-file row ranges, then construct index split in downstream 
tasks. This would avoid
     repeated index queries, but requires an additional shuffle and a more 
complex execution pipeline.
   
     In our measurements with an index at the 100-billion-row scale and a query 
recalling one billion rows, the two approaches performed similarly. Repeated 
index queries were acceptable, so we chose the simpler split-local approach.
   
   ## Tests
   
     - **Planning and results:** Verify that planning does not open index 
files, and compare distributed reads with regular reads across BTree/Bitmap 
indexes, predicate types, and `fast`/`full`/`detail` search modes.
     - **Large results and split boundaries:** Use a synthetic result 
containing billions of matches to verify that results are clipped to a split 
before row IDs are materialized or offset. Verify that BTree readers receive 
split-local row ranges.
     - **Coverage and correctness:** Test unindexed tails, null and negative 
predicates, metadata-pruned empty splits, deletion vectors, and predicates 
whose literal types must survive split serialization.
     - **Mixed-index fallback:** Verify that an unsupported `AND` branch does 
not disable an independent distributed index branch; an unsupported `OR` branch 
falls back to a full data scan; and queries with only unsupported indexes 
retain the existing fallback.
     - **Recovery and Flink integration:** Verify split serialization, pinned 
index files and tagged snapshots, reader position recovery, partition 
filtering, and dedicated split generation.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to