[ 
https://issues.apache.org/jira/browse/SPARK-58591?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

lifumao updated SPARK-58591:
----------------------------
    Attachment: spark-kubernetes-operator.log

>  Spark Kubernetes Operator: OutOfMemoryError Caused by Thread Leak
> ------------------------------------------------------------------
>
>                 Key: SPARK-58591
>                 URL: https://issues.apache.org/jira/browse/SPARK-58591
>             Project: Spark
>          Issue Type: Bug
>          Components: Bug
>    Affects Versions: 0.9.0
>            Reporter: lifumao
>            Priority: Major
>         Attachments: spark-kubernetes-operator.log
>
>
> h1. Description
>  
> The `KubernetesClientFactory` in spark-kubernetes-operator uses 
> `VertxHttpClientFactory` with a no-arg constructor and overrides 
> `newBuilder(Config)`. This causes fabric8 internally to create a *{*}new 
> Vert.x instance (with 3 threads: eventloop, acceptor, blocked-thread-checker) 
> for each HTTP connection{*}*, leading to unbounded thread growth and eventual 
> `OutOfMemoryError: unable to create native thread`.
> h1. Environment
>  
> Spark Kubernetes Operator: 0.9.0
> Java: 21.0.10+1-LTS
>  
> h1. Steps to Reproduce
>  
> 1. Deploy spark-kubernetes-operator watching all namespaces
> 2. Submit Spark applications continuously (triggering reconcile and K8s API 
> calls)
> 3. Monitor thread count: `ls /proc/1/task | wc -l`
>  
> h1. Observed Behavior
>  
> Thread count grows linearly with API calls (~40 threads/minute under moderate 
> load):
>  
> ```
> After 0 min: threads=90
> After 10 min: threads=400
> After 30 min: threads=1200
> After 2 hr: threads=4000+
> After 8 hr: threads=9000+ → OOM
> ```
>  
> Thread dump shows thousands of duplicate Vert.x instances:
>  
> ```
> 3031 × "vert.x-eventloop-thread-0"
> 3026 × "vert.x-acceptor-thread-0"
> 3031 × "vertx-blocked-thread-checker"
> ```
>  
> Final error:
>  
> ```
> java.lang.OutOfMemoryError: unable to create native thread: possibly out of 
> memory or process/resource limits reached
> [warning][os,thread] Failed to start thread - pthread_create failed (EAGAIN)
> ```
> h1. Root Cause
>  
> In `KubernetesClientFactory.java`:
>  
> ```java
> new VertxHttpClientFactory() { // no-arg constructor → sharedVertx = null
> @Override
> publicHttpClient.BuildernewBuilder(Configconfig)
> { VertxHttpClientBuilderbuilder=super.newBuilder(); // creates new Vertx 
> instance each time HttpClientUtils.applyCommonConfiguration(config, builder, 
> this); ... }
> }
> ```
>  
> Although fabric8 
> [#6726]([https://github.com/fabric8io/kubernetes-client/pull/6726]) (merged 
> in 7.0.1) fixed the default no-arg constructor to reuse a singleton Vertx via 
> `VertxHolder.INSTANCE`, the spark-operator's override of `newBuilder(Config)` 
> triggers internal fabric8 code paths (derived clients for watch/informer 
> reconnections) that bypass this fix, still creating new Vertx instances.
>  
> Passing a shared Vertx instance via `new VertxHttpClientFactory(sharedVertx)` 
> reduces but does *{*}not eliminate{*}* the leak, because fabric8's internal 
> derived-client creation paths still allocate new Vert.x instances 
> independently.
>  
> h1. Fix
>  
> We replace `VertxHttpClientFactory` with `JdkHttpClientFactory`. After fix, 
> thread count remains stable at 81-123 under continuous workload over 3 days 
> with zero leakage.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to