Troubleshooting — The Monitoring Problem Desk
Everyone eventually sees a confusing Prometheus error. This is the Problem Desk: what the message means, why it happened, and how to get out.
1 · Target shows "DOWN" in the Targets page
http://localhost:9090/targets shows a target with health "DOWN."
Cause: Prometheus can't reach the target's /metrics endpoint. The service is down, the port is wrong, or a firewall is blocking the connection.
Fix:
curl http://localhost:9100/metrics
ss -tlnp | grep 9100
cat prometheus.yml | grep targets
The target URL must be reachable from the Prometheus server. If Prometheus runs in Docker, localhost inside the container is not your host machine. Use the Docker network IP or host.docker.internal.
2 · No data in PromQL queries
A query returns no results even though the target is up.
Cause: The metric name is wrong, the labels don't match, or the metric isn't exposed by the target.
Fix:
curl http://localhost:9100/metrics | grep "node_cpu"
curl -s 'http://localhost:9090/api/v1/label/__name__/values' | python3 -m json.tool
Metric names are case-sensitive. http_requests_total ≠ HttpRequestsTotal. Check the exact name the exporter exposes.
3 · Prometheus is using too much memory
Prometheus consumes more and more memory over time.
Cause: Too many time series (high cardinality), scrape interval too short, or retention too long.
Fix:
curl -s 'http://localhost:9090/api/v1/status/tsdb' | python3 -m json.tool
global:
scrape_interval: 30s # was 15s
retention.time: 15d # was 30d
scrape_configs:
- job_name: 'app'
scrape_interval: 30s
metric_relabel_configs:
- source_labels: [__name__]
regex: 'unwanted_metric_.*'
action: drop
High cardinality (millions of unique label combinations) is the #1 cause of Prometheus memory issues. Avoid unbounded label values like user IDs or request URLs.
4 · "context deadline exceeded" in targets
Targets show "context deadline exceeded" errors.
Cause: The target is too slow to respond. Prometheus times out before getting the metrics.
Fix:
scrape_configs:
- job_name: 'slow-app'
scrape_timeout: 30s # default is 10s
static_configs:
- targets: ['slow-app:8080']
time curl http://slow-app:8080/metrics
If a target takes more than 10 seconds to respond, the application is usually overloaded. Increasing the timeout is a band-aid — fix the root cause.
5 · Alert is firing but shouldn't be
An alert fires even though everything seems fine.
Cause: The PromQL expression is too sensitive, the for duration is too short, or there's a brief spike.
Fix:
curl -s 'http://localhost:9090/api/v1/query?query=YOUR_EXPRESSION'
- alert: HighCPU
expr: ... > 80
for: 5m # was 2m — requires 5 minutes of sustained high CPU
- alert: HighCPU
expr: ... > 90 # was 80
for: 5m
False alarms erode trust in the monitoring system. Tune thresholds and for durations until alerts are meaningful.
6 · Alertmanager not sending notifications
Prometheus shows alerts firing, but Alertmanager isn't sending emails/Slack.
Cause: Alertmanager isn't receiving alerts from Prometheus, or the receiver configuration is wrong.
Fix:
curl http://localhost:9093/api/v2/alerts
curl http://localhost:9090/api/v1/alertmanagers
alerting:
alertmanagers:
- static_configs:
- targets: ['alertmanager:9093'] # must match the container name
The Alertmanager URL in prometheus.yml must be reachable from the Prometheus container. Use Docker service names, not localhost.