lifumao created SPARK-58591:
-------------------------------
Summary: Spark Kubernetes Operator: OutOfMemoryError Caused by
Thread Leak
Key: SPARK-58591
URL: https://issues.apache.org/jira/browse/SPARK-58591
Project: Spark
Issue Type: Bug
Components: Bug
Affects Versions: 0.9.0
Reporter: lifumao
h1. Description
The `KubernetesClientFactory` in spark-kubernetes-operator uses
`VertxHttpClientFactory` with a no-arg constructor and overrides
`newBuilder(Config)`. This causes fabric8 internally to create a **new Vert.x
instance (with 3 threads: eventloop, acceptor, blocked-thread-checker) for each
HTTP connection**, leading to unbounded thread growth and eventual
`OutOfMemoryError: unable to create native thread`.
h1. Environment
- Spark Kubernetes Operator: 0.9.0
- Java: 21.0.10+1-LTS
h1. Steps to Reproduce
1. Deploy spark-kubernetes-operator watching all namespaces
2. Submit Spark applications continuously (triggering reconcile and K8s API
calls)
3. Monitor thread count: `ls /proc/1/task | wc -l`
h1. Observed Behavior
Thread count grows linearly with API calls (~40 threads/minute under moderate
load):
```
After 0 min: threads=90
After 10 min: threads=400
After 30 min: threads=1200
After 2 hr: threads=4000+
After 8 hr: threads=9000+ → OOM
```
Thread dump shows thousands of duplicate Vert.x instances:
```
3031 × "vert.x-eventloop-thread-0"
3026 × "vert.x-acceptor-thread-0"
3031 × "vertx-blocked-thread-checker"
```
Final error:
```
java.lang.OutOfMemoryError: unable to create native thread: possibly out of
memory or process/resource limits reached
[warning][os,thread] Failed to start thread - pthread_create failed (EAGAIN)
```
h1. Root Cause
In `KubernetesClientFactory.java`:
```java
new VertxHttpClientFactory() { // no-arg constructor → sharedVertx = null
@Override
publicHttpClient.BuildernewBuilder(Configconfig) {
VertxHttpClientBuilderbuilder=super.newBuilder(); // creates new Vertx instance
each time
HttpClientUtils.applyCommonConfiguration(config, builder, this);
...
}
}
```
Although fabric8
[#6726](https://github.com/fabric8io/kubernetes-client/pull/6726) (merged in
7.0.1) fixed the default no-arg constructor to reuse a singleton Vertx via
`VertxHolder.INSTANCE`, the spark-operator's override of `newBuilder(Config)`
triggers internal fabric8 code paths (derived clients for watch/informer
reconnections) that bypass this fix, still creating new Vertx instances.
Passing a shared Vertx instance via `new VertxHttpClientFactory(sharedVertx)`
reduces but does **not eliminate** the leak, because fabric8's internal
derived-client creation paths still allocate new Vert.x instances independently.
h1. Fix
We replace `VertxHttpClientFactory` with `JdkHttpClientFactory`. After fix,
thread count remains stable at 81-123 under continuous workload over 3 days
with zero leakage.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]