GGraziadei commented on PR #9154: URL: https://github.com/apache/storm/pull/9154#issuecomment-6022905907
Hi @dpol1, thanks for this contribution. The idea is solid and there is a real need for it; Latency has a stragic value for Storm, so I think we should aim to do better on that side. The PR measures throughput and CPU per ack, but not latency: could you add complete latency (p50/p99) to the benchmark, with tracing off, at ratio 0, 1% and 100%? On the export side, could you publish the `otel.bsp.max.export.batch.size`, `otel.bsp.max.queue.size` and `otel.bsp.schedule.delay` values you used, and the ones you tried when tuning? Freshness is not a strict requirement for telemetry, so larger batches, a larger queue and a longer schedule delay should be acceptable; it would be useful to see how far they go and where they stop helping. A further idea: collect the telemetry on a dedicated stream into a collector bolt, flushed at a fixed interval with a tick tuple, instead of exporting from every worker. It changes where the cost lands, so it is worth keeping in mind here (but better producing benchmarks also for this approach). -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
