liangjie3138 commented on PR #10146: URL: https://github.com/apache/paimon/pull/10146#issuecomment-5808833488
> The Flink framework is quite unique, and the current approach is rather hacky; have you considered using Spark SQL? @JingsongLi We plan to use Spark for adding columns via MERGE INTO because it offers better performance and more flexible syntax. However, on the query side, Spark has a similar issue : when querying a BTree index, the index lookup runs on the driver, which is a single point of execution. At our company, we run Spark as a service, where one driver executes multiple SQL queries concurrently. Paimon’s index lookup can consume so much memory that it crashes the driver, causing all jobs running on that driver to fail. And for ecosystem considering, we hope to add support for distributed BTree index queries in both Flink and Spark. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
