Alerting — The Hospital Alarm System
Collecting metrics is like monitoring heart rate on a screen. Alerting is when the system actually sounds the alarm — sending a notification to the nurse (or the engineer) when something goes wrong.
The Alert Pipeline
Alert Rules — Defining Abnormal Vitals
Alert rules are PromQL expressions that Prometheus evaluates every evaluation_interval. When the expression returns any result, the alert fires.
rule_files:
- "alerts/*.yml"
groups:
- name: campus-library
rules:
- alert: HighErrorRate
expr: |
sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))
> 0.05
for: 2m
labels:
severity: critical
annotations:
summary: "High error rate on Campus Library"
description: "Error rate is {{ $value | humanizePercentage }} (threshold: 5%)"
- alert: HighMemoryUsage
expr: |
(1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100
> 90
for: 5m
labels:
severity: warning
annotations:
summary: "High memory usage"
description: "Memory usage is {{ $value | humanizePercentage }}"
- alert: ServiceDown
expr: up == 0
for: 1m
labels:
severity: critical
annotations:
summary: "Service is down"
description: "{{ $labels.instance }} has been down for more than 1 minute"
Each field explained
| Field | Purpose |
|---|---|
alert | Name of the alert |
expr | PromQL expression — when true, the alert fires |
for | How long the condition must be true before firing (avoids flapping) |
labels | Extra labels attached to the alert |
annotations | Human-readable message (supports {{ $value }} and {{ $labels }}) |
for prevents false alarms. A spike for 5 seconds isn't a problem. A spike for 2 minutes is.
Alertmanager — Routing the Alarm
Alertmanager receives alerts from Prometheus and routes them to the right people.
global:
resolve_timeout: 5m
route:
receiver: 'default'
group_by: ['alertname', 'severity']
group_wait: 30s # Wait 30s before sending (group alerts)
group_interval: 5m # Wait 5m between grouped notifications
repeat_interval: 4h # Repeat every 4h if not resolved
routes:
- match:
severity: critical
receiver: 'pagerduty'
- match:
severity: warning
receiver: 'slack'
receivers:
- name: 'default'
email_configs:
- to: 'team@campuslibrary.dev'
from: 'prometheus@campuslibrary.dev'
smarthost: 'smtp.gmail.com:587'
- name: 'slack'
slack_configs:
- api_url: 'https://hooks.slack.com/services/xxx/yyy/zzz'
channel: '#alerts'
title: '{{ .GroupLabels.alertname }}'
text: '{{ range .Alerts }}{{ .Annotations.description }}{{ end }}'
- name: 'pagerduty'
pagerduty_configs:
- service_key: '<your-key>'
Routing explained
| Field | Purpose |
|---|---|
group_by | Group alerts by name and severity (avoid spam) |
group_wait | Wait before first notification (collect related alerts) |
group_interval | Wait between grouped notifications |
repeat_interval | Repeat if the alert hasn't been resolved |
routes | Match alerts to receivers by labels |
Alertmanager groups related alerts. If 5 instances go down simultaneously, you get one notification, not five.
Inhibition — Suppressing Noise
Suppress less critical alerts when a critical one is firing:
inhibit_rules:
- source_match:
severity: 'critical'
target_match:
severity: 'warning'
equal: ['alertname', 'instance']
If a critical "ServiceDown" alert fires, warning alerts for the same instance are suppressed.
Inhibition prevents alert storms. If the service is down, you don't need to be told about high memory usage on that same server.
Silence — Muting During Maintenance
Silence alerts during planned maintenance:
amtool silence add alertname=HighMemoryUsage instance=server1 duration=2h reason="Scheduled maintenance"
amtool silence query
amtool silence expire <silence-id>
Forgetting to remove silences after maintenance. Old silences can suppress real alerts. Always check active silences.
Testing Alert Rules
curl -G http://localhost:9090/api/v1/query \
--data-urlencode 'query=sum(rate(http_requests_total{status=~"5.."}[5m])) / sum(rate(http_requests_total[5m])) > 0.05'
curl http://localhost:9090/api/v1/alerts
Test alert rules before deploying. A wrong expression can either miss real problems or spam you with false alarms.