liangjie3138 commented on PR #10146:
URL: https://github.com/apache/paimon/pull/10146#issuecomment-5808833488

   > The Flink framework is quite unique, and the current approach is rather 
hacky; have you considered using Spark SQL?
   
   @JingsongLi  We plan to use Spark for adding columns via MERGE INTO because 
it offers better performance and more flexible syntax. However, on the query 
side, Spark has a similar issue : when querying a BTree index, the index lookup 
runs on the driver, which is a single point of execution. At our company, we 
run Spark as a service, where one driver executes multiple SQL queries 
concurrently. Paimon’s index lookup can consume so much memory that it crashes 
the driver, causing all jobs running on that driver to fail. And for ecosystem 
considering, we hope to add support for distributed BTree index queries in both 
Flink and Spark.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to