liangjie3138 opened a new issue, #10145:
URL: https://github.com/apache/paimon/issues/10145

   ### Search before asking
   
   - [x] I searched in the [issues](https://github.com/apache/paimon/issues) 
and found nothing similar.
   
   
   ### Motivation
   
    The goal is to enable distributed global-index querying for Flink query 
jobs. Currently, global-index predicates are evaluated on the JobManager during 
scan planning. In our use cases, this causes problems:
   
     1. JobManager OOM: In a real workload, a query needs to retrieve about one 
billion matching rows from a table at the trillion-row scale. Processing the 
index result on the JobManager makes it prone to out-of-memory failures.
     2. Interference between jobs: We use a Flink session cluster for ad-hoc 
queries. Its JobManager is long-lived and shared by multiple jobs. Evaluating a 
BTree index there can affect other queries and cluster stability.
   
   ### Solution
   
     Distribute global-index queries across the tasks that read data. During 
planning, identify the relevant index files and assign the necessary query 
information to each data split. Each task evaluates the relevant indexes before 
reading its data, while preserving rows not covered by an index. Unsupported 
cases retain the existing query path.
   
   ### Anything else?
   
   _No response_
   
   ### Are you willing to submit a PR?
   
   - [x] I'm willing to submit a PR!


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to