liangjie3138 commented on PR #10146: URL: https://github.com/apache/paimon/pull/10146#issuecomment-5811975573
> Let me take a look again. Could you share some performance test data? With the latest version, our single-node setup is already capable of supporting tables exceeding 10 PB in size. This performance issue is not determined by table size alone; it depends more heavily on the number of matching row IDs. In my tests, querying the index on a single node caused no problems when the match count was in the tens of millions. When it reached several hundred million, the single node showed significant performance problems: the JobManager ran out of memory and stalled. My test dataset is a 100-billion-row, 400-TB table. The query applies a simple filter on one indexed column in `fast` mode. I built a BTree index on `self_merge_content_length` covering the entire table; the index itself is 369 GB. <img width="2630" height="586" alt="image" src="https://github.com/user-attachments/assets/c5e4dbf1-e900-4119-86ae-f36d1dd2a509" /> The SQL is: ```sql SELECT * FROM html_feature_source /*+ OPTIONS('scan.index-distributed-query.enabled' = 'true', 'btree-index.fallback-scan-max-size'='1000tb')*/ WHERE `self_merge_content_length` > 200000; ``` This SQL is expected to match around 450 million rows. At job startup, the JobManager's memory is exhausted. The job then remains in the pending state for hours and cannot complete. The JobManager has 8 CPU cores and 16 GB of memory. <img width="2596" height="914" alt="image" src="https://github.com/user-attachments/assets/253d4bf7-c81c-4cdd-9af1-7b6b516bc961" /> With the approach in this branch, the same SQL and resources - an 8-core, 16-GB JobManager and a source parallelism of 1,024 - complete the job in approximately 40 minutes. <img width="2560" height="1476" alt="image" src="https://github.com/user-attachments/assets/d445c0a5-461a-4215-8128-80bd9e0fe00d" /> -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
