amaechler commented on code in PR #20404:
URL: https://github.com/apache/druid/pull/20404#discussion_r4084077787
##########
docs/operations/dynamic-config-provider.md:
##########
@@ -76,4 +76,96 @@ When you define the consumer properties in the supervisor
spec, use the dynamic
},
...
```
-When connecting to Kafka, Druid replaces the environment variables with their
corresponding values.
\ No newline at end of file
+When connecting to Kafka, Druid replaces the environment variables with their
corresponding values.
+
+## Kubernetes node label dynamic config provider
+
+The
[`druid-kubernetes-extensions`](../development/extensions-core/kubernetes.md)
extension provides a dynamic config provider
(`K8sNodeLabelDynamicConfigProvider`) that reads configuration values from the
labels of the Kubernetes node that the Druid process runs on. Use it for values
that depend on Kubernetes runtime information, for example where a process was
scheduled, which cannot be written into a spec ahead of time.
+
+The Kubernetes node label dynamic config provider uses the following syntax:
+
+```json
+druid.dynamic.config.provider={"type": "k8sNodeLabel","labels":{"property1":
"example.com/label-one","property2": "example.com/label-two"}}
+```
+
+|Field|Type|Description|Required|
+|-----|----|-----------|--------|
+|`type`|String|dynamic config provider type|Yes: `k8sNodeLabel`|
+|`labels`|Map|configuration keys to resolve, each mapped to the name of the
node label that holds its value|Yes|
+|`nodeNameVariable`|String|environment variable that holds the name of the
node the process runs on|No (default: `HOST_NODE_NAME`)|
+
+You can use it anywhere Druid accepts a dynamic config provider, for example
Kafka [`consumerProperties`](../ingestion/kafka-ingestion.md), Iceberg
[`catalogProperties`](../development/extensions-contrib/iceberg.md), and Schema
Registry [`config` and `headers`](../ingestion/data-formats.md).
+
+### Prerequisites
+
+Include `druid-kubernetes-extensions` in the [extensions load
list](../configuration/extensions.md#loading-extensions) of every service that
resolves the spec. For a supervisor spec, that is the Overlord and the Peon
services.
+
+Each pod must know the node it runs on. Expose the node name from the downward
API as the environment variable named by `nodeNameVariable`, or set
`nodeNameVariable` to a variable your pod template already provides:
+
+```yaml
+env:
+ - name: HOST_NODE_NAME
+ valueFrom:
+ fieldRef:
+ fieldPath: spec.nodeName
+```
+
+Nodes are cluster-scoped, so the service account the Druid pods run as needs a
ClusterRole granting `get` on `nodes`:
+
+```yaml
+kind: ClusterRole
+apiVersion: rbac.authorization.k8s.io/v1
+metadata:
+ name: druid-node-reader
+rules:
+- apiGroups:
+ - ""
+ resources:
+ - nodes
+ verbs:
+ - get
+---
+kind: ClusterRoleBinding
+apiVersion: rbac.authorization.k8s.io/v1
+metadata:
+ name: druid-node-reader
+subjects:
+- kind: ServiceAccount
+ name: default # the service account your Druid pods use
+ namespace: druid # their namespace
+roleRef:
+ kind: ClusterRole
+ name: druid-node-reader
+ apiGroup: rbac.authorization.k8s.io
+```
+
+### Behavior
+
+- If the node name variable is unset, the API server cannot be reached or
refuses the request, or the node does not carry the label, Druid logs a warning
and omits the key. The consumer then applies its own default. A label with an
empty value counts as missing. As with other dynamic config providers, a
resolved value takes precedence over the same key specified directly.
+- Druid reads the node once per process and keeps its labels until the process
restarts. A failed lookup is not kept, so it is retried the next time the
configuration is resolved.
+
+### Examples
+
+Kafka consumers fetch from the nearest replica when `client.rack` matches a
broker's `broker.rack`
([KIP-392](https://cwiki.apache.org/confluence/display/KAFKA/KIP-392%3A+Allow+consumers+to+fetch+from+closest+replica)).
On EKS, `topology.k8s.aws/zone-id` holds the zone ID that MSK uses for
`broker.rack`; on GKE, `topology.kubernetes.io/zone` holds the zone name that
Google Cloud Managed Service for Apache Kafka uses:
+
+```json
+"consumerProperties": {
+ "bootstrap.servers": "localhost:9092",
+ "druid.dynamic.config.provider": {
+ "type": "k8sNodeLabel",
+ "labels": { "client.rack": "topology.k8s.aws/zone-id" }
+ }
+}
+```
+
+An Iceberg REST catalog can read from S3 in the task's own region by taking
`client.region` from the node's region label:
+
+```json
+"catalogProperties": {
+ "uri": "https://iceberg-rest.example.com",
Review Comment:
Good find, updated.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]