mirror of
https://github.com/prometheus-operator/runbooks.git
synced 2026-08-27 12:37:20 +00:00
Merge pull request #7 from swissquote/alertmanager
This commit is contained in:
@@ -0,0 +1,35 @@
|
||||
---
|
||||
title: Alertmanager Cluster Crashlooping
|
||||
weight: 20
|
||||
---
|
||||
|
||||
# AlertmanagerClusterCrashlooping
|
||||
|
||||
## Meaning
|
||||
|
||||
Half or more of the Alertmanager instances within the same cluster are crashlooping.
|
||||
|
||||
## Impact
|
||||
|
||||
Alerts could be notified multiple time unless pods are crashing to fast and no alerts can be sent.
|
||||
|
||||
## Diagnosis
|
||||
|
||||
```bash
|
||||
kubectl get pod -l app=alertmanager
|
||||
|
||||
NAMESPACE NAME READY STATUS RESTARTS AGE
|
||||
default alertmanager-main-0 1/2 CrashLoopBackOff 37107 2d
|
||||
default alertmanager-main-1 2/2 Running 0 43d
|
||||
default alertmanager-main-2 2/2 Running 0 43d
|
||||
```
|
||||
|
||||
Find the root cause by looking to events for a given pod/deployement
|
||||
|
||||
```
|
||||
kubectl get events --field-selector involvedObject.name=alertmanager-main-0
|
||||
```
|
||||
|
||||
## Mitigation
|
||||
|
||||
Make sure pods have enough resources (CPU, MEM) to work correctly.
|
||||
@@ -0,0 +1,28 @@
|
||||
---
|
||||
title: Alertmanager Cluster Down
|
||||
weight: 20
|
||||
---
|
||||
|
||||
# AlertmanagerClusterDown
|
||||
|
||||
## Meaning
|
||||
|
||||
Half or more of the Alertmanager instances within the same cluster are down.
|
||||
|
||||
## Impact
|
||||
|
||||
You have an unstable cluster, if everything goes wrong you will lose the whole cluster.
|
||||
|
||||
## Diagnosis
|
||||
|
||||
Verify why pods are not running.
|
||||
You can get a big picture with `events`.
|
||||
|
||||
```bash
|
||||
$ kubectl get events --field-selector involvedObject.kind=Pod | grep alertmanager
|
||||
```
|
||||
|
||||
## Mitigation
|
||||
|
||||
There are no cheap options to mitigate this risk.
|
||||
Verifying any new changes in preprod before production environment should improve stability.
|
||||
@@ -0,0 +1,22 @@
|
||||
---
|
||||
title: Alertmanager Cluster Failed To Send Alerts
|
||||
weight: 20
|
||||
---
|
||||
|
||||
# AlertmanagerClusterFailedToSendAlerts
|
||||
|
||||
## Meaning
|
||||
|
||||
All instances failed to send notification to an integration.
|
||||
|
||||
## Impact
|
||||
|
||||
You will not receive a notification when an alert is raised.
|
||||
|
||||
## Diagnosis
|
||||
|
||||
No alerts are received at the integration level from the cluster.
|
||||
|
||||
## Mitigation
|
||||
|
||||
Depending on the integration, correct the integration with the faulty instance (network, authorization token, firewall...)
|
||||
@@ -0,0 +1,24 @@
|
||||
---
|
||||
title: Alertmanager ConfigInconsistent
|
||||
weight: 20
|
||||
---
|
||||
|
||||
# AlertmanagerConfigInconsistent
|
||||
|
||||
## Meaning
|
||||
|
||||
The configuration between instances inside a cluster is inconsistent.
|
||||
|
||||
## Impact
|
||||
|
||||
Configuration inconsistency can be multiple and impact is hard to predict.
|
||||
Nevertheless, in most cases the alert might be lost or routed to the incorrect integration.
|
||||
|
||||
## Diagnosis
|
||||
|
||||
Run a `diff` tool between all `alertmanager.yml` that are deployed to find what is wrong.
|
||||
You could run a job within your CI to avoid this issue in the future.
|
||||
|
||||
## Mitigation
|
||||
|
||||
Delete the incorrect secret and deploy the correct one.
|
||||
@@ -1,3 +1,8 @@
|
||||
---
|
||||
title: AlertmanagerFailedReload
|
||||
weight: 20
|
||||
---
|
||||
|
||||
# AlertmanagerFailedReload
|
||||
|
||||
## Meaning
|
||||
@@ -8,16 +13,23 @@ configuration for a certain period.
|
||||
|
||||
## Impact
|
||||
|
||||
Alerts for cluster components may not be delivered as expected.
|
||||
The impact depends on the type of the error you will find in the logs.
|
||||
Most of the time, previous configuration is still working, thanks to multiple instances, so avoid deleting existing pods.
|
||||
|
||||
## Diagnosis
|
||||
|
||||
Check the logs for the `alertmanager-main` pods in the `monitoring` namespace:
|
||||
Verify if there is an error in `config-reloader` container logs.
|
||||
Here an example with network issues.
|
||||
|
||||
```console
|
||||
$ kubectl -n monitoring logs -l 'alertmanager=main'
|
||||
```bash
|
||||
$ kubectl logs sts/alertmanager-main -c config-reloader
|
||||
|
||||
level=error ts=2021-09-24T11:24:52.69629226Z caller=runutil.go:101 msg="function failed. Retrying in next tick" err="trigger reload: reload request failed: Post \"http://localhost:9093/alertmanager/-/reload\": dial tcp [::1]:9093: connect: connection refused"
|
||||
```
|
||||
|
||||
You can also verify directly `alertmanager.yaml` file (default: `/etc/alertmanager/config/alertmanager.yaml`).
|
||||
|
||||
## Mitigation
|
||||
|
||||
The resolution depends on the particular issue reported in the logs.
|
||||
Running [amtool check-config alertmanager.yaml](https://github.com/prometheus/alertmanager#amtool) on your configuration file will help you detect problem related to syntax.
|
||||
You could also rollback `alertmanager.yaml` to the previous version in order to get back to a stable version.
|
||||
|
||||
@@ -0,0 +1,35 @@
|
||||
---
|
||||
title: Alertmanager Failed To Send Alerts
|
||||
weight: 20
|
||||
---
|
||||
|
||||
# AlertmanagerFailedToSendAlerts
|
||||
|
||||
## Meaning
|
||||
|
||||
At least one instance is unable to routed alert to the corresponding integration.
|
||||
|
||||
## Impact
|
||||
|
||||
No impact since another instance should be able to send the notification, unless `AlertmanagerClusterFailedToSendAlerts` is also triggerd for the same integration.
|
||||
|
||||
## Diagnosis
|
||||
|
||||
Verify the amount of failed notification per alert-manager-[instance] for a specific integration.
|
||||
|
||||
You can look metrics exposed in prometheus console using promQL. For exemple the following query will display the number of failed notifications per instance for pager duty integration. We have 3 instances involved in the example bellow.
|
||||
|
||||
```
|
||||
rate(alertmanager_notifications_total{integration="pagerduty"}[5m])
|
||||
```
|
||||
|
||||

|
||||
|
||||
|
||||
## Mitigation
|
||||
|
||||
Depending on the integration, you can have a look to alert-manager logs and act (network, authorization token, firewall...)
|
||||
|
||||
```
|
||||
kubectl -n monitoring logs -l 'alertmanager=main' -c alertmanager
|
||||
```
|
||||
Reference in New Issue
Block a user