[ 
https://issues.apache.org/jira/browse/IGNITE-29047?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18121810#comment-18121810
 ] 

Ignite TC Bot commented on IGNITE-29047:
----------------------------------------

{panel:title=Branch: [pull/13574/head] Base: [master] : No blockers 
found!|borderStyle=dashed|borderColor=#ccc|titleBGColor=#D6F7C1}{panel}
{panel:title=Branch: [pull/13574/head] Base: [master] : New Tests 
(2)|borderStyle=dashed|borderColor=#ccc|titleBGColor=#D6F7C1}
{color:#00008b}Basic 4{color} [[tests 
2|https://ci2.ignite.apache.org/viewLog.html?buildId=9378510]]
* {color:#013220}IgniteBasicTestSuite2: 
FailureProcessorMetricsTest.testPerTypeIgnoredFailuresCountMetrics - 
PASSED{color}
* {color:#013220}IgniteBasicTestSuite2: 
FailureProcessorMetricsTest.testNoMetricsRegisteredWhenNothingIgnored - 
PASSED{color}

{panel}
[TeamCity *--> Run :: All* 
Results|https://ci2.ignite.apache.org/viewLog.html?buildId=9378652&buildTypeId=IgniteTests24Java8_RunAll]
{color:#ffffff}tcbot-analysis-comment chainBuildId=9378652 
rerunBuildIds=none{color}

> Add ignored failure count metric to FailureProcessor
> ----------------------------------------------------
>
>                 Key: IGNITE-29047
>                 URL: https://issues.apache.org/jira/browse/IGNITE-29047
>             Project: Ignite
>          Issue Type: Task
>            Reporter: Dmitry Werner
>            Assignee: Dmitry Werner
>            Priority: Major
>              Labels: from-log-to-metric
>          Time Spent: 0.5h
>  Remaining Estimate: 0h
>
> *Problem*
> When a FailureHandler is configured to ignore certain failure types (via 
> AbstractFailureHandler#setIgnoredFailureTypes), the corresponding failures 
> are silently suppressed — only a warning is logged. There is no way to 
> observe, through the metrics system, how many failures of each type have been 
> suppressed on a node. Operators monitoring a cluster have no visibility into 
> the rate/volume of ignored failures, which makes it hard to detect recurring 
> critical conditions (e.g. a repeatedly blocked system worker) that are being 
> deliberately tolerated.
> *Proposed Solution*
> Expose one long counter per ignored failure type, registered under a 
> dedicated metrics registry failure.ignored. Each counter is incremented every 
> time FailureProcessor suppresses a failure of the matching type. Counters are 
> created only for the failure types the configured handler actually ignores, 
> so no metrics are registered when nothing is ignored.
> *Proposed change*
> - Register a new metrics registry "failure.ignored" in FailureProcessor.
> - Add a long counter per ignored failure type, incremented every time a 
> failure is suppressed (ignored) by the configured failure handler — i.e. 
> exactly where IGNORED_FAILURE_LOG_MSG is printed.
> - Keep the existing log message unchanged.
> - Document the new metric in docs/_docs/monitoring-metrics/new-metrics.adoc.
> - Add FailureProcessorMetricsTest covering per-type counting and the 
> no-metrics-when-nothing-ignored case.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to