[ 
https://issues.apache.org/jira/browse/IGNITE-29047?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Dmitry Werner updated IGNITE-29047:
-----------------------------------
    Description: 
*Problem*
When a FailureHandler is configured to ignore certain failure types (via 
AbstractFailureHandler#setIgnoredFailureTypes), the corresponding failures are 
silently suppressed — only a warning is logged. There is no way to observe, 
through the metrics system, how many failures of each type have been suppressed 
on a node. Operators monitoring a cluster have no visibility into the 
rate/volume of ignored failures, which makes it hard to detect recurring 
critical conditions (e.g. a repeatedly blocked system worker) that are being 
deliberately tolerated.

*Proposed Solution*
Expose one long counter per ignored failure type, registered under a dedicated 
metrics registry failure.ignored. Each counter is incremented every time 
FailureProcessor suppresses a failure of the matching type. Counters are 
created only for the failure types the configured handler actually ignores, so 
no metrics are registered when nothing is ignored.

*Proposed change*
- Register a new metrics registry "failure.ignored" in FailureProcessor.
- Add a long counter per ignored failure type, incremented every time a failure 
is suppressed (ignored) by the configured failure handler — i.e. exactly where 
IGNORED_FAILURE_LOG_MSG is printed.
- Keep the existing log message unchanged.
- Document the new metric in docs/_docs/monitoring-metrics/new-metrics.adoc.
- Add FailureProcessorMetricsTest covering per-type counting and the 
no-metrics-when-nothing-ignored case.

  was:
Currently, the event "Possible failure suppressed accordingly to a configured 
handler" is only reported via the log by FailureProcessor. There is no way to 
track how often such failures occur without parsing (grepping) the log.

We need to expose this as a metric so the count of suppressed (ignored) 
failures can be monitored through the metrics subsystem (e.g. JMX, system 
views, exporters) instead of relying on log parsing.

*Proposed change*
- Register a new metric group failure (register name: failure) in 
FailureProcessor.
- Add a counter metric IgnoredFailuresCount that is incremented every time a 
failure is suppressed (ignored) by the configured failure handler — i.e. 
exactly where IGNORED_FAILURE_LOG_MSG is printed.
- Keep the existing log message unchanged.
- Document the new metric in docs/_docs/monitoring-metrics/new-metrics.adoc.

*Acceptance criteria*
- The metric is exposed and readable through the standard metrics subsystem.
- The metric increases by 1 for each suppressed failure and is not affected by 
failures processed normally (not ignored).
- A unit test covers the above behavior.


> Add ignored failure count metric to FailureProcessor
> ----------------------------------------------------
>
>                 Key: IGNITE-29047
>                 URL: https://issues.apache.org/jira/browse/IGNITE-29047
>             Project: Ignite
>          Issue Type: Task
>            Reporter: Dmitry Werner
>            Assignee: Dmitry Werner
>            Priority: Major
>              Labels: from-log-to-metric
>          Time Spent: 0.5h
>  Remaining Estimate: 0h
>
> *Problem*
> When a FailureHandler is configured to ignore certain failure types (via 
> AbstractFailureHandler#setIgnoredFailureTypes), the corresponding failures 
> are silently suppressed — only a warning is logged. There is no way to 
> observe, through the metrics system, how many failures of each type have been 
> suppressed on a node. Operators monitoring a cluster have no visibility into 
> the rate/volume of ignored failures, which makes it hard to detect recurring 
> critical conditions (e.g. a repeatedly blocked system worker) that are being 
> deliberately tolerated.
> *Proposed Solution*
> Expose one long counter per ignored failure type, registered under a 
> dedicated metrics registry failure.ignored. Each counter is incremented every 
> time FailureProcessor suppresses a failure of the matching type. Counters are 
> created only for the failure types the configured handler actually ignores, 
> so no metrics are registered when nothing is ignored.
> *Proposed change*
> - Register a new metrics registry "failure.ignored" in FailureProcessor.
> - Add a long counter per ignored failure type, incremented every time a 
> failure is suppressed (ignored) by the configured failure handler — i.e. 
> exactly where IGNORED_FAILURE_LOG_MSG is printed.
> - Keep the existing log message unchanged.
> - Document the new metric in docs/_docs/monitoring-metrics/new-metrics.adoc.
> - Add FailureProcessorMetricsTest covering per-type counting and the 
> no-metrics-when-nothing-ignored case.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to