mirror of
https://github.com/prometheus-operator/runbooks.git
synced 2026-08-27 12:37:20 +00:00
Update AlertmanagerFailedToSendAlerts.md
This commit is contained in:
@@ -11,12 +11,25 @@ At least one instance is unable to routed alert to the corresponding integration
|
||||
|
||||
## Impact
|
||||
|
||||
No impact since another instance will be able to send the notification.
|
||||
No impact since another instance should be able to send the notification, unless `AlertmanagerClusterFailedToSendAlerts` is also triggerd for the same integration.
|
||||
|
||||
## Diagnosis
|
||||
|
||||
Verify that alerts send by each instance have equivalent alert distribution per integration.
|
||||
Verify the amount of failed notification per alert-manager-[instance] for a specific integration.
|
||||
|
||||
You can look metrics exposed in prometheus console using promQL. For exemple the following query will display the number of failed notifications per instance for pager duty integration. We have 3 instances involved in the example bellow.
|
||||
|
||||
```
|
||||
rate(alertmanager_notifications_total{integration="pagerduty"}[5m])
|
||||
```
|
||||
|
||||

|
||||
|
||||
|
||||
## Mitigation
|
||||
|
||||
Depending on the integration, correct the integration with the faulty instance (network, authorization token, firewall...)
|
||||
Depending on the integration, you can have a look to alert-manager logs and act (network, authorization token, firewall...)
|
||||
|
||||
```
|
||||
kubectl -n monitoring logs -l 'alertmanager=main' -c alertmanager
|
||||
```
|
||||
|
||||
Reference in New Issue
Block a user