nzw921rx commented on PR #12248:
URL: https://github.com/apache/seatunnel/pull/12248#issuecomment-5632743268
Hello, I have some thoughts that I would like to discuss with you regarding
the original design idea of the report:
A. I first asked myself several questions:
1. How can we use the most concise daily report to intuitively identify
problems at a glance?
My answer to myself is that the fewer metrics we focus on, and the more
those metrics can truly express meaningful information and provide long-term
value, the more appropriate they are to be reflected in the report. The
detailed process should be handled by downloading the detailed artifacts for
investigation once a problem has been identified.
2. How much runner fluctuation do we consider acceptable?
Based on my own long-term observation, fluctuations within 5% are very
common.
3. Are different JDKs comparable?
They are not. The servers they run on almost always have different
CPUs/models/core counts/memory for each run, so the scores between different
JDKs are naturally not comparable, and Error/CV fluctuations are also related
to machine load.
B. Have we currently found that there are somewhat too many scenarios in the
long-term `benchmarks_core` report?
Yes, they can be simplified. I think the long-term value lies in stable
benchmarks, for example:
CheckpointingTime.checkpointSingleInput
CheckpointingTime.checkpointSingleInput
Pipeline.sourceSink
Pipeline.sourceTransformSink
They have been very stable recently, with both Error/CV around 1%. I think
these are suitable to remain in the long-term report, and they help us identify
regressions.
Summary:
I think the intention of this PR is very clear. It wants people who see CV
fluctuations to further diagnose the fluctuation based on the details, but
there is a maintainability issue here, as well as the question of what kind of
problem it ultimately provides value in identifying. In fact, it is trying to
identify the case where only one fork is high while the others are very stable,
but here I want to give an example: in my observation, the case where one fork
is high has almost never occurred.
1.Pipeline.sourceTransformSink
This JMH benchmark has been very stable during my observations over the past
few weeks and has not shown this kind of problem, so this further reduces the
intention of the new report to reduce investigation costs.
2.IntermediateQueue.disruptorRecordHandoff
This JMH benchmark has been consistently unstable during my observations
over the past few weeks, which is intended to demonstrate that the current CV
metric already fully expresses this situation.
I look forward to your reply. I would like to seriously discuss this with
you, and I also look forward to your suggestions.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]