[
https://issues.apache.org/jira/browse/FLINK-40372?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18103972#comment-18103972
]
liting liu commented on FLINK-40372:
------------------------------------
I'm glad to fix this
> [Kubernetes Operator] JOSDK per-resource metrics are not unregistered after
> resource cleanup
> --------------------------------------------------------------------------------------------
>
> Key: FLINK-40372
> URL: https://issues.apache.org/jira/browse/FLINK-40372
> Project: Flink
> Issue Type: Bug
> Components: Kubernetes Operator
> Affects Versions: 1.13
> Reporter: liting liu
> Priority: Critical
>
> ## Description
> `OperatorJosdkMetrics` creates Flink metric groups and counters whose scopes
> contain
> the namespace and name of every reconciled Kubernetes custom resource.
> After a resource is successfully deleted, `cleanupDone()` only increments a
> cleanup
> counter. It does not close or remove the metric group, and it does not remove
> the
> associated entries from the internal maps.
> As resource names change over time, especially for automatically generated
> `FlinkStateSnapshot` resources such as periodic savepoints, the number of
> registered
> metrics can grow monotonically.
> This increases metric cardinality and retains metric groups, counters and
> their scope
> strings for resources that no longer need to be monitored. When the SLF4J
> reporter is
> enabled, every reporting cycle also iterates over and formats the accumulated
> metrics.
> ## Affected versions
> Observed with an Apache Flink Kubernetes Operator 1.13 based build.
> The same lifecycle issue is still present on Apache `main` at:
> - commit: `fdfcc695daba16e4f4a8c9a6c08b7f470787de00`
> - date: 2026-08-11
> JOSDK version in the 1.13 based build: 5.1.2.
> ## Relevant configuration
> The SLF4J reporter was configured with a five-minute interval:
> kubernetes.operator.metrics.reporter.slf4j.factory.class:
> org.apache.flink.metrics.slf4j.Slf4jReporterFactory
> kubernetes.operator.metrics.reporter.slf4j.interval: 5 MINUTE
> The Operator container had a 1 GiB memory limit.
> ## Source analysis
> `OperatorJosdkMetrics` keeps the following maps for the lifetime of the
> Operator:
> ```java
> private final Map<ResourceID, KubernetesResourceNamespaceMetricGroup>
> resourceNsMetricGroups = new ConcurrentHashMap<>();
> private final Map<ResourceID, KubernetesResourceMetricGroup>
> resourceMetricGroups = new ConcurrentHashMap<>();
> private final Map<List<String>, Histogram> histograms = new
> ConcurrentHashMap<>();
> private final Map<List<String>, Counter> counters = new ConcurrentHashMap<>();
--
This message was sent by Atlassian Jira
(v8.20.10#820010)