Prometheus Hands-on Exercises — Set Up the Monitoring Room
Time to stop reading and start doing. These exercises use Docker to run Prometheus locally.
Run each command, check the output, and verify in the Prometheus UI at http://localhost:9090.
Exercise 0 · Setup — Open the Monitoring Room
docker run -d --name prometheus -p 9090:9090 \
-v $(pwd)/prometheus.yml:/etc/prometheus/prometheus.yml \
prom/prometheus:latest
First, create a minimal config:
cat > prometheus.yml << 'EOF'
global:
scrape_interval: 15s
scrape_configs:
- job_name: 'prometheus'
static_configs:
- targets: ['localhost:9090']
EOF
docker rm -f prometheus
docker run -d --name prometheus -p 9090:9090 \
-v $(pwd)/prometheus.yml:/etc/prometheus/prometheus.yml \
prom/prometheus:latest
Open http://localhost:9090/targets — you should see Prometheus scraping itself.
Prometheus scrapes itself by default. This is a good sanity check — if Prometheus can't scrape itself, something is fundamentally wrong.
Exercise 1 · First Query — Check the Heartbeat
Open http://localhost:9090 and type in the query bar:
count(up)
up
Both should return 1 — one target is up.
curl -s 'http://localhost:9090/api/v1/query?query=up' | python3 -m json.tool
up == 1 means the target is healthy. up == 0 means it's down.
Exercise 2 · Node Exporter — Add a Server
docker run -d --name node-exporter -p 9100:9100 \
--net=host \
prom/node-exporter:latest
Update prometheus.yml:
cat > prometheus.yml << 'EOF'
global:
scrape_interval: 15s
scrape_configs:
- job_name: 'prometheus'
static_configs:
- targets: ['localhost:9090']
- job_name: 'node'
static_configs:
- targets: ['localhost:9100']
EOF
docker exec prometheus kill -HUP 1
Check http://localhost:9090/targets — both targets should be up.
100 - (avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
node_cpu_seconds_total is a counter. rate() converts it to a rate. Subtracting from 100 gives you CPU usage percentage.
Exercise 3 · Alert Rules — Sound the Alarm
cat > alerts.yml << 'EOF'
groups:
- name: node-alerts
rules:
- alert: HighCPU
expr: 100 - (avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 80
for: 2m
labels:
severity: warning
annotations:
summary: "High CPU usage"
description: "CPU usage is {{ $value | humanizePercentage }}"
- alert: TargetDown
expr: up == 0
for: 1m
labels:
severity: critical
annotations:
summary: "Target is down"
description: "{{ $labels.job }} is not responding"
EOF
Update prometheus.yml to include the rules:
cat > prometheus.yml << 'EOF'
global:
scrape_interval: 15s
rule_files:
- "alerts.yml"
scrape_configs:
- job_name: 'prometheus'
static_configs:
- targets: ['localhost:9090']
- job_name: 'node'
static_configs:
- targets: ['localhost:9100']
EOF
docker rm -f prometheus
docker run -d --name prometheus -p 9090:9090 \
-v $(pwd)/prometheus.yml:/etc/prometheus/prometheus.yml \
-v $(pwd)/alerts.yml:/etc/prometheus/alerts.yml \
prom/prometheus:latest
Check http://localhost:9090/alerts — the rules should be listed (firing or inactive).
Rules are evaluated every evaluation_interval. If the expression is true for longer than for, the alert fires.
Exercise 4 · PromQL Practice — The Diagnostic Toolkit
Try these queries at http://localhost:9090:
node_memory_MemAvailable_bytes / 1024 / 1024 / 1024
node_filesystem_avail_bytes{mountpoint="/"} / 1024 / 1024 / 1024
rate(node_network_receive_bytes_total{device="lo"}[5m])
node_open_filefds_allocated
PromQL is like SQL for time series. Use rate() for counters, direct access for gauges, and histogram_quantile() for histograms.
Exercise 5 · Visualization — Built-in Graphs
Prometheus has a built-in graph UI. Select a query and click Graph:
100 - (avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
(1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100
Adjust the time range to "1h" to see recent trends.
Prometheus graphs are useful for quick debugging. For dashboards, use Grafana.
Exercise 6 · Cleanup
docker rm -f prometheus node-exporter
rm -f prometheus.yml alerts.yml
Remove containers after exercises. In production, use Docker Compose or Kubernetes for Prometheus.