[
https://issues.apache.org/jira/browse/SPARK-58967?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Shazeb Khan updated SPARK-58967:
--------------------------------
Description:
We observed unnecessary DNS lookups when running multiple concurrent Spark
pipelines using {{spark-submit}} in Kubernetes cluster mode.
The issue occurs when executors attempt to connect to the Spark driver. The
driver hostname generated by {{DriverServiceFeatureStep}} contains the
namespace and {{svc}} components:
{{sample-app-driver-svc.default.svc}}
Because this hostname contains fewer dots than the configured {{ndots}} value
({{{}ndots:5{}}}), the pod's DNS resolver first attempts to resolve the
hostname using the configured search domains.
This results in two failed DNS lookups ({{{}NXDOMAIN{}}}) before the driver
service is successfully resolved.
h2. Environment
* *Apache Spark:* 3.5.2 and 4.1.1
* *Deployment Mode:* Cluster mode on Kubernetes
* *CoreDNS:* Standard Kubernetes configuration
* *Pod DNS resolver:* {{ndots:5}}
* *DNS search domains:*
{{search default.svc.cluster.local svc.cluster.local cluster.local}}
h2. Current Behaviour
{{DriverServiceFeatureStep.scala}} currently constructs {{spark.driver.host}}
as:
{{val driverHostname =
s"$resolvedServiceName.${kubernetesConf.namespace}.svc"}}
For example:
{{sample-app-driver-svc.default.svc}}
The DNS resolver then performs the following lookups:
{{Attempt 1:
sample-app-driver-svc.default.svc.default.svc.cluster.local. → NXDOMAIN
Attempt 2:
sample-app-driver-svc.default.svc.svc.cluster.local. → NXDOMAIN
Attempt 3:
sample-app-driver-svc.default.svc.cluster.local. → NOERROR}}
As a result, two unnecessary DNS lookups are performed before the driver
service is resolved.
h2. Impact
We are running multiple Spark pipelines concurrently in our production
Kubernetes environment.
Each executor startup can generate these unnecessary failed DNS queries while
attempting to connect to the driver. With a large number of concurrent
pipelines and executors, this results in additional DNS traffic and can
contribute to delays during executor-to-driver connection establishment.
It also generates additional {{NXDOMAIN}} traffic to CoreDNS, which is a shared
cluster-wide component.
h2. Observed DNS Queries
With the current hostname construction, CoreDNS receives:
{{[INFO] "A IN sample-app-driver-svc.default.svc.default.svc.cluster.local."
NXDOMAIN
[INFO] "AAAA IN sample-app-driver-svc.default.svc.default.svc.cluster.local."
NXDOMAIN
[INFO] "A IN sample-app-driver-svc.default.svc.svc.cluster.local." NXDOMAIN
[INFO] "AAAA IN sample-app-driver-svc.default.svc.svc.cluster.local." NXDOMAIN
[INFO] "A IN sample-app-driver-svc.default.svc.cluster.local." NOERROR
[INFO] "AAAA IN sample-app-driver-svc.default.svc.cluster.local." NOERROR}}
h2. Proposed Solution
One possible solution we have tested is to use only the Kubernetes Service name
as {{{}spark.driver.host{}}}:
{{sample-app-driver-svc}}
The Kubernetes DNS search path would then resolve it directly as:
{{sample-app-driver-svc.default.svc.cluster.local.}}
This avoids the intermediate {{NXDOMAIN}} lookups and allows the service to be
resolved on the first DNS attempt.
The corresponding code change would be:
*Current:*
{{val driverHostname =
s"$resolvedServiceName.${kubernetesConf.namespace}.svc"}}
*Proposed:*
{{val driverHostname = s"$resolvedServiceName"}}
h2. Questions
Before proposing a code change to {{{}DriverServiceFeatureStep{}}}, we would
like to understand whether there is an existing Spark configuration or external
property that can address this behaviour.
Specifically:
1. Is there any Spark configuration/property that can be used to avoid these
unnecessary DNS lookups without modifying {{{}DriverServiceFeatureStep{}}}?
2. Is there a specific reason why {{DriverServiceFeatureStep}} constructs
{{spark.driver.host}} using {{.namespace.svc}} instead of relying on the
standard Kubernetes DNS search path?
3. Would using only the Kubernetes Service name as {{spark.driver.host}} be a
safe and supported approach for Kubernetes cluster mode?
4. If there is no existing configuration option, would changing the default
hostname construction to use the Service name only be considered an appropriate
fix?
was:
We observed unnecessary DNS lookups when running multiple concurrent Spark
pipelines using {{spark-submit}} in Kubernetes cluster mode.
The issue occurs when executors attempt to connect to the Spark driver. The
driver hostname generated by {{DriverServiceFeatureStep}} contains the
namespace and {{svc}} components:
{{sample-app-driver-svc.default.svc}}
Because this hostname contains fewer dots than the configured {{ndots}} value
({{{}ndots:5{}}}), the pod's DNS resolver first attempts to resolve the
hostname using the configured search domains.
This results in two failed DNS lookups ({{{}NXDOMAIN{}}}) before the driver
service is successfully resolved.
h2. Environment
* *Apache Spark:* 3.5.2 and 4.1.1
* *Deployment Mode:* Cluster mode on Kubernetes
* *CoreDNS:* Standard Kubernetes configuration
* *Pod DNS resolver:* {{ndots:5}}
* *DNS search domains:*
{{search default.svc.cluster.local svc.cluster.local cluster.local}}
h2. Current Behaviour
{{DriverServiceFeatureStep.scala}} currently constructs {{spark.driver.host}}
as:
{{val driverHostname =
s"$resolvedServiceName.${kubernetesConf.namespace}.svc"}}
For example:
{{sample-app-driver-svc.default.svc}}
The DNS resolver then performs the following lookups:
{{Attempt 1:
sample-app-driver-svc.default.svc.default.svc.cluster.local. → NXDOMAIN
Attempt 2:
sample-app-driver-svc.default.svc.svc.cluster.local. → NXDOMAIN
Attempt 3:
sample-app-driver-svc.default.svc.cluster.local. → NOERROR}}
As a result, two unnecessary DNS lookups are performed before the driver
service is resolved.
h2. Impact
We are running multiple Spark pipelines concurrently in our production
Kubernetes environment.
Each executor startup can generate these unnecessary failed DNS queries while
attempting to connect to the driver. With a large number of concurrent
pipelines and executors, this results in additional DNS traffic and can
contribute to delays during executor-to-driver connection establishment.
It also generates additional {{NXDOMAIN}} traffic to CoreDNS, which is a shared
cluster-wide component.
h2. Observed DNS Queries
With the current hostname construction, CoreDNS receives:
{{[INFO] "A IN sample-app-driver-svc.default.svc.default.svc.cluster.local."
NXDOMAIN
[INFO] "AAAA IN sample-app-driver-svc.default.svc.default.svc.cluster.local."
NXDOMAIN
[INFO] "A IN sample-app-driver-svc.default.svc.svc.cluster.local." NXDOMAIN
[INFO] "AAAA IN sample-app-driver-svc.default.svc.svc.cluster.local." NXDOMAIN
[INFO] "A IN sample-app-driver-svc.default.svc.cluster.local." NOERROR
[INFO] "AAAA IN sample-app-driver-svc.default.svc.cluster.local." NOERROR}}
h2. Proposed Solution
One possible solution we have tested is to use only the Kubernetes Service name
as {{{}spark.driver.host{}}}:
{{sample-app-driver-svc}}
The Kubernetes DNS search path would then resolve it directly as:
{{sample-app-driver-svc.default.svc.cluster.local.}}
This avoids the intermediate {{NXDOMAIN}} lookups and allows the service to be
resolved on the first DNS attempt.
The corresponding code change would be:
*Current:*
{{val driverHostname =
s"$resolvedServiceName.${kubernetesConf.namespace}.svc"}}
*Proposed:*
{{val driverHostname = s"$resolvedServiceName"}}
h2. Questions
Before proposing a code change to {{{}DriverServiceFeatureStep{}}}, we would
like to understand whether there is an existing Spark configuration or external
property that can address this behaviour.
Specifically:
# Is there any Spark configuration/property that can be used to avoid these
unnecessary DNS lookups without modifying {{{}DriverServiceFeatureStep{}}}?
# Is there a specific reason why {{DriverServiceFeatureStep}} constructs
{{spark.driver.host}} using {{.namespace.svc}} instead of relying on the
standard Kubernetes DNS search path?
# Would using only the Kubernetes Service name as {{spark.driver.host}} be a
safe and supported approach for Kubernetes cluster mode?
# If there is no existing configuration option, would changing the default
hostname construction to use the Service name only be considered an appropriate
fix?
> [apache-spark] [K8S] Multiple Failed DNS Lookups When Resolving the Driver
> Hostname in Spark Applications on Kubernetes
> -----------------------------------------------------------------------------------------------------------------------
>
> Key: SPARK-58967
> URL: https://issues.apache.org/jira/browse/SPARK-58967
> Project: Spark
> Issue Type: Bug
> Components: Kubernetes
> Affects Versions: 3.5.2, 4.1.1
> Reporter: Shazeb Khan
> Priority: Major
> Labels: CoreDNS, kubernetes, spark
>
> We observed unnecessary DNS lookups when running multiple concurrent Spark
> pipelines using {{spark-submit}} in Kubernetes cluster mode.
> The issue occurs when executors attempt to connect to the Spark driver. The
> driver hostname generated by {{DriverServiceFeatureStep}} contains the
> namespace and {{svc}} components:
> {{sample-app-driver-svc.default.svc}}
> Because this hostname contains fewer dots than the configured {{ndots}} value
> ({{{}ndots:5{}}}), the pod's DNS resolver first attempts to resolve the
> hostname using the configured search domains.
> This results in two failed DNS lookups ({{{}NXDOMAIN{}}}) before the driver
> service is successfully resolved.
> h2. Environment
> * *Apache Spark:* 3.5.2 and 4.1.1
> * *Deployment Mode:* Cluster mode on Kubernetes
> * *CoreDNS:* Standard Kubernetes configuration
> * *Pod DNS resolver:* {{ndots:5}}
> * *DNS search domains:*
> {{search default.svc.cluster.local svc.cluster.local cluster.local}}
> h2. Current Behaviour
> {{DriverServiceFeatureStep.scala}} currently constructs {{spark.driver.host}}
> as:
> {{val driverHostname =
> s"$resolvedServiceName.${kubernetesConf.namespace}.svc"}}
> For example:
> {{sample-app-driver-svc.default.svc}}
> The DNS resolver then performs the following lookups:
> {{Attempt 1:
> sample-app-driver-svc.default.svc.default.svc.cluster.local. → NXDOMAIN
> Attempt 2:
> sample-app-driver-svc.default.svc.svc.cluster.local. → NXDOMAIN
> Attempt 3:
> sample-app-driver-svc.default.svc.cluster.local. → NOERROR}}
> As a result, two unnecessary DNS lookups are performed before the driver
> service is resolved.
> h2. Impact
> We are running multiple Spark pipelines concurrently in our production
> Kubernetes environment.
> Each executor startup can generate these unnecessary failed DNS queries while
> attempting to connect to the driver. With a large number of concurrent
> pipelines and executors, this results in additional DNS traffic and can
> contribute to delays during executor-to-driver connection establishment.
> It also generates additional {{NXDOMAIN}} traffic to CoreDNS, which is a
> shared cluster-wide component.
> h2. Observed DNS Queries
> With the current hostname construction, CoreDNS receives:
>
> {{[INFO] "A IN sample-app-driver-svc.default.svc.default.svc.cluster.local."
> NXDOMAIN
> [INFO] "AAAA IN sample-app-driver-svc.default.svc.default.svc.cluster.local."
> NXDOMAIN
> [INFO] "A IN sample-app-driver-svc.default.svc.svc.cluster.local." NXDOMAIN
> [INFO] "AAAA IN sample-app-driver-svc.default.svc.svc.cluster.local." NXDOMAIN
> [INFO] "A IN sample-app-driver-svc.default.svc.cluster.local." NOERROR
> [INFO] "AAAA IN sample-app-driver-svc.default.svc.cluster.local." NOERROR}}
> h2. Proposed Solution
> One possible solution we have tested is to use only the Kubernetes Service
> name as {{{}spark.driver.host{}}}:
> {{sample-app-driver-svc}}
> The Kubernetes DNS search path would then resolve it directly as:
> {{sample-app-driver-svc.default.svc.cluster.local.}}
> This avoids the intermediate {{NXDOMAIN}} lookups and allows the service to
> be resolved on the first DNS attempt.
> The corresponding code change would be:
> *Current:*
> {{val driverHostname =
> s"$resolvedServiceName.${kubernetesConf.namespace}.svc"}}
> *Proposed:*
> {{val driverHostname = s"$resolvedServiceName"}}
> h2. Questions
> Before proposing a code change to {{{}DriverServiceFeatureStep{}}}, we would
> like to understand whether there is an existing Spark configuration or
> external property that can address this behaviour.
> Specifically:
> 1. Is there any Spark configuration/property that can be used to avoid these
> unnecessary DNS lookups without modifying {{{}DriverServiceFeatureStep{}}}?
> 2. Is there a specific reason why {{DriverServiceFeatureStep}} constructs
> {{spark.driver.host}} using {{.namespace.svc}} instead of relying on the
> standard Kubernetes DNS search path?
> 3. Would using only the Kubernetes Service name as {{spark.driver.host}} be a
> safe and supported approach for Kubernetes cluster mode?
> 4. If there is no existing configuration option, would changing the default
> hostname construction to use the Service name only be considered an
> appropriate fix?
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]