FrankChen021 commented on code in PR #20404:
URL: https://github.com/apache/druid/pull/20404#discussion_r4082377991


##########
docs/operations/dynamic-config-provider.md:
##########
@@ -76,4 +76,96 @@ When you define the consumer properties in the supervisor 
spec, use the dynamic
       },
 ...
 ```
-When connecting to Kafka, Druid replaces the environment variables with their 
corresponding values.
\ No newline at end of file
+When connecting to Kafka, Druid replaces the environment variables with their 
corresponding values.
+
+## Kubernetes node label dynamic config provider
+
+The 
[`druid-kubernetes-extensions`](../development/extensions-core/kubernetes.md) 
extension provides a dynamic config provider 
(`K8sNodeLabelDynamicConfigProvider`) that reads configuration values from the 
labels of the Kubernetes node that the Druid process runs on. Use it for values 
that depend on Kubernetes runtime information, for example where a process was 
scheduled, which cannot be written into a spec ahead of time.
+
+The Kubernetes node label dynamic config provider uses the following syntax:
+
+```json
+druid.dynamic.config.provider={"type": "k8sNodeLabel","labels":{"property1": 
"example.com/label-one","property2": "example.com/label-two"}}
+```
+
+|Field|Type|Description|Required|
+|-----|----|-----------|--------|
+|`type`|String|dynamic config provider type|Yes: `k8sNodeLabel`|
+|`labels`|Map|configuration keys to resolve, each mapped to the name of the 
node label that holds its value|Yes|
+|`nodeNameVariable`|String|environment variable that holds the name of the 
node the process runs on|No (default: `HOST_NODE_NAME`)|
+
+You can use it anywhere Druid accepts a dynamic config provider, for example 
Kafka [`consumerProperties`](../ingestion/kafka-ingestion.md), Iceberg 
[`catalogProperties`](../development/extensions-contrib/iceberg.md), and Schema 
Registry [`config` and `headers`](../ingestion/data-formats.md).
+
+### Prerequisites
+
+Include `druid-kubernetes-extensions` in the [extensions load 
list](../configuration/extensions.md#loading-extensions) of every service that 
resolves the spec. For a supervisor spec, that is the Overlord and the Peon 
services.
+
+Each pod must know the node it runs on. Expose the node name from the downward 
API as the environment variable named by `nodeNameVariable`, or set 
`nodeNameVariable` to a variable your pod template already provides:
+
+```yaml
+env:
+  - name: HOST_NODE_NAME
+    valueFrom:
+      fieldRef:
+        fieldPath: spec.nodeName
+```
+
+Nodes are cluster-scoped, so the service account the Druid pods run as needs a 
ClusterRole granting `get` on `nodes`:
+
+```yaml
+kind: ClusterRole
+apiVersion: rbac.authorization.k8s.io/v1
+metadata:
+  name: druid-node-reader
+rules:
+- apiGroups:
+  - ""
+  resources:
+  - nodes
+  verbs:
+  - get
+---
+kind: ClusterRoleBinding
+apiVersion: rbac.authorization.k8s.io/v1
+metadata:
+  name: druid-node-reader
+subjects:
+- kind: ServiceAccount
+  name: default        # the service account your Druid pods use
+  namespace: druid     # their namespace
+roleRef:
+  kind: ClusterRole
+  name: druid-node-reader
+  apiGroup: rbac.authorization.k8s.io
+```
+
+### Behavior
+
+- If the node name variable is unset, the API server cannot be reached or 
refuses the request, or the node does not carry the label, Druid logs a warning 
and omits the key. The consumer then applies its own default. A label with an 
empty value counts as missing. As with other dynamic config providers, a 
resolved value takes precedence over the same key specified directly.
+- Druid reads the node once per process and keeps its labels until the process 
restarts. A failed lookup is not kept, so it is retried the next time the 
configuration is resolved.
+
+### Examples
+
+Kafka consumers fetch from the nearest replica when `client.rack` matches a 
broker's `broker.rack` 
([KIP-392](https://cwiki.apache.org/confluence/display/KAFKA/KIP-392%3A+Allow+consumers+to+fetch+from+closest+replica)).
 On EKS, `topology.k8s.aws/zone-id` holds the zone ID that MSK uses for 
`broker.rack`; on GKE, `topology.kubernetes.io/zone` holds the zone name that 
Google Cloud Managed Service for Apache Kafka uses:
+
+```json
+"consumerProperties": {
+  "bootstrap.servers": "localhost:9092",
+  "druid.dynamic.config.provider": {
+    "type": "k8sNodeLabel",
+    "labels": { "client.rack": "topology.k8s.aws/zone-id" }
+  }
+}
+```
+
+An Iceberg REST catalog can read from S3 in the task's own region by taking 
`client.region` from the node's region label:
+
+```json
+"catalogProperties": {
+  "uri": "https://iceberg-rest.example.com";,

Review Comment:
   [P2] Use the required REST catalog URI field
   
   **Finding:** This example places the REST endpoint in 
`catalogProperties.uri`, but `RestIcebergCatalog` requires a non-null top-level 
`catalogUri` (and the `rest` catalog type) before it resolves the provider; its 
setup then overwrites any `catalogProperties.uri`. Following this new example 
therefore fails with `catalogUri cannot be null` instead of configuring the 
REST catalog.
   
   **Suggestion:** Show the example in its `icebergCatalog` context with the 
required `type` and `catalogUri` fields, leaving only the dynamic 
`client.region` mapping inside `catalogProperties`.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to