sanjayk9r opened a new issue, #20246:
URL: https://github.com/apache/druid/issues/20246

   ### Description
   
   Add a configurable sampling mechanism to the Kafka emitter community 
extension.
   The mechanism should allow operators to selectively drop a percentage of
   ServiceMetricEvents before they are produced to Kafka.
   
   Suggested configuration properties for KafkaEmitterConfig:
   
   | Property | Type | Default | Description |
   | --- | --- | --- | --- |
   | `druid.emitter.kafka.samplingRate` | int (0-100) | 100 | Percentage of 
matching metric events to keep. `100` keeps all matching events; `0` drops all 
matching events. |
   | `druid.emitter.kafka.sampledMetrics` | Set<String> | `[]` | Metric names 
subject to sampling, such as `query/time`. If empty, sampling applies to all 
metrics. |
   | `druid.emitter.kafka.sampledNodeTypes` | Set<String> | `[]` | Node types 
subject to sampling, such as `service/broker`. If empty, sampling applies to 
all services. |
   
   Expected filtering behavior:
   
   1. Only `ServiceMetricEvent` instances are eligible for sampling.
   2. If `sampledMetrics` is non-empty, sampling applies only when the metric
      name is present in that set.
   3. If `sampledNodeTypes` is non-empty, sampling applies only when the event's
      node type is present in that set.
   4. For matching events, use pseudo-random selection to retain approximately
      `samplingRate` percent of events.
   5. Events that do not match the configured metric or node-type filters must 
be
      retained.
   6. Alert, segment metadata, and request events, when configured, must remain
      exempt from sampling and continue to be emitted at 100%.
   7. Events should be dropped before they reach the Kafka producer buffer so
      that the feature reduces Kafka produce volume, network traffic, and
      infrastructure cost.
   
   ### Motivation
   
   High-traffic Druid deployments can generate millions of metric events. High 
cardinality metrics such as `query/time` can create substantial Kafka 
throughput even when the events are primarily used for trend analysis or 
anomaly detection.
   
   For example, retaining 10% of per-query latency metrics can provide useful 
statistical distributions such as p50 and p99 while substantially reducing 
Kafka ingestion, storage, and network overhead. Consumer-side filtering does 
not provide the same benefit because the full event volume has already been
   produced and stored in Kafka.
   
   This feature would let operators choose the trade-off between metric 
granularity and infrastructure cost while preserving complete delivery of 
alerts and other critical event types.
   
   ### Rationale
   
   Sampling at the emitter is preferable to consumer-side filtering because it 
avoids the cost of serializing, producing, transmitting, and storing events 
that will eventually be discarded. The proposed matching operation is an O(1) 
set lookup followed by random selection and should have negligible overhead 
compared with event serialization.
   
   The default sampling rate of 100 preserves existing behavior for users who 
do not configure the feature.
   
   ### Operational impact
   
   This change is opt-in in effect because the default is `100`, meaning no 
events are sampled. Operators using sampling should account for its effect on 
absolute metric counts. Consumers and BI tools may need to scale sampled 
metrics when estimating total event volume.
   
   Alerts remain exempt from sampling to avoid reducing monitoring and alerting
   reliability.
   
   ### Configuration example
   
   ```properties
   # Basic Kafka emitter configuration
   druid.emitter=kafka
   druid.emitter.kafka.bootstrap.servers=kafka-broker:9092
   druid.emitter.kafka.event.types=["metrics","alerts"]
   
   # Keep 10% of selected high-volume metrics
   druid.emitter.kafka.samplingRate=10
   
druid.emitter.kafka.sampledMetrics=["query/time","query/bytes","sqlQuery/time"]
   ```
   
   ### Test plan
   
   Add coverage for:
   
   - Default configuration retaining all events.
   - Sampling rates of 0 and 100.
   - Sampling of matching and non-matching metric names.
   - Sampling of matching and non-matching node types.
   - Empty filter sets applying sampling to all eligible metrics.
   - Non-`ServiceMetricEvent` events always being retained.
   - Approximate retention behavior over a sufficiently large event sample.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to