andygrove commented on issue #6124: URL: https://github.com/apache/datafusion-comet/issues/6124#issuecomment-5861059160
The worker-starvation explanation in the last paragraph holds up in a standalone test. Only Tokio workers drive the I/O and timer driver, so an I/O future polled from a Spark task thread's `block_on` gets no wake-ups while every worker is busy. With tokio 1.53 and one worker stuck in a 2 s poll, a 10 ms sleep in another thread's `block_on` took 1.9 s, and so did a socket read whose peer wrote after 300 ms. With two workers both busy it was the same. Two things make that more likely than the executor-cores sizing suggests. On a standalone cluster without `spark.executor.cores` the runtime has a single worker (#6292). And several JVM calls block a worker from inside a poll (#6293), including the S3 credential provider, which is called on every request. Running the failing query with `COMET_WORKER_THREADS` set to the executor's core count would show whether this is what you're hitting. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
