Skip to main content

Alerting — The Hospital Alarm System

Collecting metrics is like monitoring heart rate on a screen. Alerting is when the system actually sounds the alarm — sending a notification to the nurse (or the engineer) when something goes wrong.

The Alert Pipeline​

Alert Rules — Defining Abnormal Vitals​

Alert rules are PromQL expressions that Prometheus evaluates every evaluation_interval. When the expression returns any result, the alert fires.

prometheus.yml — include alert rules
rule_files:
- "alerts/*.yml"
alerts/app_alerts.yml
groups:
- name: campus-library
rules:
- alert: HighErrorRate
expr: |
sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))
> 0.05
for: 2m
labels:
severity: critical
annotations:
summary: "High error rate on Campus Library"
description: "Error rate is {{ $value | humanizePercentage }} (threshold: 5%)"

- alert: HighMemoryUsage
expr: |
(1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100
> 90
for: 5m
labels:
severity: warning
annotations:
summary: "High memory usage"
description: "Memory usage is {{ $value | humanizePercentage }}"

- alert: ServiceDown
expr: up == 0
for: 1m
labels:
severity: critical
annotations:
summary: "Service is down"
description: "{{ $labels.instance }} has been down for more than 1 minute"

Each field explained​

FieldPurpose
alertName of the alert
exprPromQL expression — when true, the alert fires
forHow long the condition must be true before firing (avoids flapping)
labelsExtra labels attached to the alert
annotationsHuman-readable message (supports {{ $value }} and {{ $labels }})
Remember

for prevents false alarms. A spike for 5 seconds isn't a problem. A spike for 2 minutes is.

Alertmanager — Routing the Alarm​

Alertmanager receives alerts from Prometheus and routes them to the right people.

alertmanager.yml
global:
resolve_timeout: 5m

route:
receiver: 'default'
group_by: ['alertname', 'severity']
group_wait: 30s # Wait 30s before sending (group alerts)
group_interval: 5m # Wait 5m between grouped notifications
repeat_interval: 4h # Repeat every 4h if not resolved

routes:
- match:
severity: critical
receiver: 'pagerduty'

- match:
severity: warning
receiver: 'slack'

receivers:
- name: 'default'
email_configs:
- to: 'team@campuslibrary.dev'
from: 'prometheus@campuslibrary.dev'
smarthost: 'smtp.gmail.com:587'

- name: 'slack'
slack_configs:
- api_url: 'https://hooks.slack.com/services/xxx/yyy/zzz'
channel: '#alerts'
title: '{{ .GroupLabels.alertname }}'
text: '{{ range .Alerts }}{{ .Annotations.description }}{{ end }}'

- name: 'pagerduty'
pagerduty_configs:
- service_key: '<your-key>'

Routing explained​

FieldPurpose
group_byGroup alerts by name and severity (avoid spam)
group_waitWait before first notification (collect related alerts)
group_intervalWait between grouped notifications
repeat_intervalRepeat if the alert hasn't been resolved
routesMatch alerts to receivers by labels
Remember

Alertmanager groups related alerts. If 5 instances go down simultaneously, you get one notification, not five.

Inhibition — Suppressing Noise​

Suppress less critical alerts when a critical one is firing:

alertmanager.yml
inhibit_rules:
- source_match:
severity: 'critical'
target_match:
severity: 'warning'
equal: ['alertname', 'instance']

If a critical "ServiceDown" alert fires, warning alerts for the same instance are suppressed.

Remember

Inhibition prevents alert storms. If the service is down, you don't need to be told about high memory usage on that same server.

Silence — Muting During Maintenance​

Silence alerts during planned maintenance:

Create a silence via amtool
amtool silence add alertname=HighMemoryUsage instance=server1 duration=2h reason="Scheduled maintenance"
List active silences
amtool silence query
Expire a silence
amtool silence expire <silence-id>
Common mistake

Forgetting to remove silences after maintenance. Old silences can suppress real alerts. Always check active silences.

Testing Alert Rules​

Test a rule expression
curl -G http://localhost:9090/api/v1/query \
--data-urlencode 'query=sum(rate(http_requests_total{status=~"5.."}[5m])) / sum(rate(http_requests_total[5m])) > 0.05'
Check which alerts are firing
curl http://localhost:9090/api/v1/alerts
Remember

Test alert rules before deploying. A wrong expression can either miss real problems or spam you with false alarms.