liangjie3138 opened a new issue, #10145: URL: https://github.com/apache/paimon/issues/10145
### Search before asking - [x] I searched in the [issues](https://github.com/apache/paimon/issues) and found nothing similar. ### Motivation The goal is to enable distributed global-index querying for Flink query jobs. Currently, global-index predicates are evaluated on the JobManager during scan planning. In our use cases, this causes problems: 1. JobManager OOM: In a real workload, a query needs to retrieve about one billion matching rows from a table at the trillion-row scale. Processing the index result on the JobManager makes it prone to out-of-memory failures. 2. Interference between jobs: We use a Flink session cluster for ad-hoc queries. Its JobManager is long-lived and shared by multiple jobs. Evaluating a BTree index there can affect other queries and cluster stability. ### Solution Distribute global-index queries across the tasks that read data. During planning, identify the relevant index files and assign the necessary query information to each data split. Each task evaluates the relevant indexes before reading its data, while preserving rows not covered by an index. Unsupported cases retain the existing query path. ### Anything else? _No response_ ### Are you willing to submit a PR? - [x] I'm willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
