Skip to main content

Alerting in Grafana — Warning Lights on the Dashboard

Your dashboard shows the gauges. Alerting is when a gauge turns red and a warning light flashes — the system notifies the driver (engineer) before the engine fails.

Grafana Alerting vs Prometheus Alerting​

FeatureGrafana AlertingPrometheus Alerting
Where configuredGrafana UI or provisioning filesprometheus.yml + Alertmanager
Query languagePromQL, SQL, etc.PromQL only
Notification routingBuilt into GrafanaSeparate Alertmanager
Multi-sourceWorks with any data sourcePrometheus only
Best forDashboard-centric alertsApplication-level monitoring
Remember

Use Prometheus alerting for application health (error rates, latency). Use Grafana alerting for dashboard-centric alerts (when a specific panel goes red).

Creating an Alert from a Panel​

  1. Edit a panel → Alert tab → Create alert rule
  2. Set the condition (when the query exceeds a threshold)
  3. Configure evaluation interval and for duration
  4. Add notification channels
Alert rule — High CPU
alert:
name: High CPU Usage
message: "CPU usage is above 80% on {{ $labels.instance }}"
frequency: 1m
conditions:
- evaluator:
type: gt
params: [80]
query:
params: ['A', '5m', 'now']
reducer:
type: avg
params: []
type: query
Remember

Grafana alerts are evaluated by Grafana, not Prometheus. This means they work even if Prometheus is down (for cached data).

Notification Channels — Where Alerts Go​

ChannelType
EmailTraditional email notifications
SlackSlack channel messages
PagerDutyIncident management
OpsGenieOn-call management
WebhookCustom HTTP endpoint
Microsoft TeamsTeams channel messages

Configure Slack notifications​

  1. Alerting → Contact points → Add contact point
  2. Select Slack
  3. Enter webhook URL and channel name

Configure email notifications​

  1. Alerting → Contact points → Add contact point
  2. Select Email
  3. Enter recipient email addresses
  4. Configure SMTP in grafana.ini:
grafana.ini
[smtp]
enabled = true
host = smtp.gmail.com:587
user = your-email@gmail.com
password = your-app-password
from_address = grafana@campuslibrary.dev
Remember

Set up at least two notification channels. Email for non-urgent alerts, Slack/PagerDuty for critical.

Alert Rules — Provisioning​

Alert rules can be defined as code and provisioned automatically:

provisioning/alerting/rules.yml
apiVersion: 1

groups:
- orgId: 1
name: Campus Library
folder: Monitoring
interval: 1m
rules:
- uid: high-cpu
title: High CPU Usage
condition: C
data:
- refId: A
datasourceUid: prometheus
model:
expr: 100 - (avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
instant: false
for: 5m
labels:
severity: warning
annotations:
summary: "High CPU usage"
description: "CPU is {{ $values.A }}%"
Remember

Provisioned alerts are version-controlled and reviewed like any other code. This is the recommended approach for production.

Silence Rules — Muting During Maintenance​

provisioning/alerting/silences.yml
apiVersion: 1

silences:
- orgId: 1
name: Scheduled maintenance
matchers:
- name: alertname
value: HighCPU
operator: =
startsAt: '2024-10-15T02:00:00Z'
endsAt: '2024-10-15T04:00:00Z'
createdBy: riya
comment: "Scheduled maintenance window"
Common mistake

Forgetting to remove silences after maintenance. Old silences can suppress real alerts. Always check active silences.

On-Call Rotation — Who Gets Paged​

Grafana has built-in on-call rotation management:

  1. Alerting → OnCall → Schedules
  2. Define rotation schedules (who is on-call when)
  3. Link notification policies to schedules
Remember

On-call rotation ensures alerts reach the right person at the right time. Critical alerts during business hours → team lead. After hours → on-call engineer.

Notification Policies — Routing Rules​

Notification policies route alerts to the right contact point based on labels:

provisioning/alerting/policies.yml
apiVersion: 1

policies:
- orgId: 1
receiver: default-slack
group_by: ['alertname', 'severity']
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
routes:
- receiver: pagerduty-critical
matchers:
- name: severity
value: critical
- receiver: slack-warning
matchers:
- name: severity
value: warning
Remember

Route critical alerts to PagerDuty (immediate page). Route warnings to Slack (non-urgent notification).