GGraziadei commented on PR #9154:
URL: https://github.com/apache/storm/pull/9154#issuecomment-6022905907

   Hi @dpol1, thanks for this contribution. The idea is solid and there is a 
real need for it;
   
   Latency has a stragic value for Storm, so I think we should aim to do better 
on that side. The PR measures throughput and CPU per ack, but not latency: 
could you add complete latency (p50/p99) to the benchmark, with tracing off, at 
ratio 0, 1% and 100%?
   
   On the export side, could you publish the `otel.bsp.max.export.batch.size`, 
`otel.bsp.max.queue.size` and `otel.bsp.schedule.delay` values you used, and 
the ones you tried when tuning? Freshness is not a strict requirement for 
telemetry, so larger batches, a larger queue and a longer schedule delay should 
be acceptable; it would be useful to see how far they go and where they stop 
helping.
   
   A further idea: collect the telemetry on a dedicated stream into a collector 
bolt, flushed at a fixed interval with a tick tuple, instead of exporting from 
every worker. It changes where the cost lands, so it is worth keeping in mind 
here (but better producing benchmarks also for this approach). 


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to