[
https://issues.apache.org/jira/browse/SPARK-59949?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated SPARK-59949:
-----------------------------------
Labels: pull-request-available (was: )
> Record the `operator.sdk` controller execution histograms with sub-second
> precision
> -----------------------------------------------------------------------------------
>
> Key: SPARK-59949
> URL: https://issues.apache.org/jira/browse/SPARK-59949
> Project: Spark
> Issue Type: Sub-task
> Components: Kubernetes
> Affects Versions: kubernetes-operator-1.0.0
> Reporter: Peter Toth
> Priority: Major
> Labels: pull-request-available
>
> - {{OperatorJosdkMetrics.timeControllerExecution}} updates its histograms
> with {{toSeconds(startTime)}}, which is
> {{TimeUnit.MILLISECONDS.toSeconds(...)}}. So the elapsed time is cut down to
> whole seconds, and a reconcile under one second records 0.
> - Most reconciles take well under a second. So the quantiles of these
> histograms are mostly 0, and their Prometheus {{_sum}} (SPARK-59935)
> undercounts.
> - Recording nanoseconds in a histogram whose name contains {{nanos}} would
> fix it, since {{PrometheusPullModelHandler}} then exports seconds. The name
> change is user-facing, so it needs a migration guide entry.
> - This is there since SPARK-48984 (0.1.0).
> Found while reviewing
> https://github.com/apache/spark-kubernetes-operator/pull/922.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]