liangjie3138 commented on PR #10146: URL: https://github.com/apache/paimon/pull/10146#issuecomment-5806956595
> This is a significant change that introduces a considerable amount of coupling; I need substantial input regarding the business logic, so please describe your business scenario in detail. We have a use case involving the ingestion of a web-page feature table into a data lake. Because we need to add columns frequently and filter datasets by feature columns, we chose a data evolution table and use BTree indexes to accelerate queries. The web-page feature table is expected to reach one trillion rows and PB-scale storage. During a PoC using data at the hundred-billion-row scale, I found that when the number of rows matched by the BTree index reaches the billion-row scale (filtering a dataset by feature fields to obtain a subset often selects a large amount of data), JobManager memory usage becomes excessively high, preventing the query from completing successfully. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
