Skip to main content

Capstone — Full Monitoring for Campus Library

Time for the real thing. This challenge chains everything you learned into one complete monitoring setup. Plan for about one hour.

Goal

By the end, you should have Prometheus scraping a Node Exporter, with alert rules for CPU/memory/disk, PromQL dashboards, and Alertmanager sending notifications.

Step 1 · The Stack​

Create a docker-compose.yml:

docker-compose.yml
services:
prometheus:
image: prom/prometheus:latest
ports:
- "9090:9090"
volumes:
- ./prometheus.yml:/etc/prometheus/prometheus.yml
- ./alerts.yml:/etc/prometheus/alerts.yml
- prometheus-data:/prometheus
command:
- '--config.file=/etc/prometheus/prometheus.yml'
- '--storage.tsdb.retention.time=30d'

node-exporter:
image: prom/node-exporter:latest
ports:
- "9100:9100"
volumes:
- /proc:/host/proc:ro
- /sys:/host/sys:ro
- /:/rootfs:ro
command:
- '--path.procfs=/host/proc'
- '--path.sysfs=/host/sys'
- '--path.rootfs=/rootfs'

alertmanager:
image: prom/alertmanager:latest
ports:
- "9093:9093"
volumes:
- ./alertmanager.yml:/etc/alertmanager/alertmanager.yml

volumes:
prometheus-data:

Step 2 · Configuration​

prometheus.yml
global:
scrape_interval: 15s
evaluation_interval: 15s

rule_files:
- "alerts.yml"

alerting:
alertmanagers:
- static_configs:
- targets: ['alertmanager:9093']

scrape_configs:
- job_name: 'prometheus'
static_configs:
- targets: ['localhost:9090']

- job_name: 'node'
static_configs:
- targets: ['node-exporter:9100']

Step 3 · Alert Rules​

alerts.yml
groups:
- name: campus-alerts
rules:
- alert: HighCPU
expr: 100 - (avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 80
for: 2m
labels:
severity: warning
annotations:
summary: "High CPU on {{ $labels.instance }}"
description: "CPU usage is {{ $value | humanizePercentage }}"

- alert: HighMemory
expr: (1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100 > 90
for: 5m
labels:
severity: warning
annotations:
summary: "High memory on {{ $labels.instance }}"
description: "Memory usage is {{ $value | humanizePercentage }}"

- alert: DiskSpaceLow
expr: (1 - node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) * 100 > 85
for: 5m
labels:
severity: critical
annotations:
summary: "Disk space low on {{ $labels.instance }}"
description: "Disk usage is {{ $value | humanizePercentage }}"

- alert: TargetDown
expr: up == 0
for: 1m
labels:
severity: critical
annotations:
summary: "{{ $labels.job }} is down"
description: "{{ $labels.instance }} has been down for more than 1 minute"

Step 4 · Alertmanager​

alertmanager.yml
global:
resolve_timeout: 5m

route:
receiver: 'default'
group_by: ['alertname', 'severity']
group_wait: 30s
group_interval: 5m
repeat_interval: 4h

receivers:
- name: 'default'
slack_configs:
- api_url: 'https://hooks.slack.com/services/YOUR/WEBHOOK/URL'
channel: '#alerts'
title: '{{ .GroupLabels.alertname }}'
text: '{{ range .Alerts }}{{ .Annotations.description }}{{ end }}'

Step 5 · Start and Verify​

Start everything
docker compose up -d
Verify targets are up
curl -s http://localhost:9090/api/v1/targets | python3 -c "import sys,json; data=json.load(sys.stdin); [print(f'{t[\"labels\"][\"job\"]}: {t[\"health\"]}') for t in data['data']['activeTargets']]"
prometheus: up
node: up
Test a PromQL query
curl -s 'http://localhost:9090/api/v1/query?query=up'

Step 6 · Explore the UI​

  1. http://localhost:9090 — Prometheus UI

    • Try the queries from Exercise 4
    • Check the Alerts tab
    • Check the Targets tab
  2. http://localhost:9093 — Alertmanager UI

    • See the routing tree
    • Check active alerts and silences

Step 7 · Clean Up​

Stop everything
docker compose down -v
rm -f prometheus.yml alerts.yml alertmanager.yml docker-compose.yml
Remember

This capstone chained: scrape config → node_exporter → alert rules → PromQL → Alertmanager. This is exactly how production monitoring works.