Skip to main content

Troubleshooting — The Monitoring Problem Desk

Everyone eventually sees a confusing Prometheus error. This is the Problem Desk: what the message means, why it happened, and how to get out.

1 · Target shows "DOWN" in the Targets page​

Problem

http://localhost:9090/targets shows a target with health "DOWN."

Cause: Prometheus can't reach the target's /metrics endpoint. The service is down, the port is wrong, or a firewall is blocking the connection.

Fix:

Test the target manually
curl http://localhost:9100/metrics
Check if the port is listening
ss -tlnp | grep 9100
Check the target URL in prometheus.yml
cat prometheus.yml | grep targets
Remember

The target URL must be reachable from the Prometheus server. If Prometheus runs in Docker, localhost inside the container is not your host machine. Use the Docker network IP or host.docker.internal.

2 · No data in PromQL queries​

Problem

A query returns no results even though the target is up.

Cause: The metric name is wrong, the labels don't match, or the metric isn't exposed by the target.

Fix:

Check what metrics are available
curl http://localhost:9100/metrics | grep "node_cpu"
Check the exact metric name in Prometheus
curl -s 'http://localhost:9090/api/v1/label/__name__/values' | python3 -m json.tool
Remember

Metric names are case-sensitive. http_requests_total ≠ HttpRequestsTotal. Check the exact name the exporter exposes.

3 · Prometheus is using too much memory​

Problem

Prometheus consumes more and more memory over time.

Cause: Too many time series (high cardinality), scrape interval too short, or retention too long.

Fix:

Check the number of time series
curl -s 'http://localhost:9090/api/v1/status/tsdb' | python3 -m json.tool
Reduce cardinality or retention
global:
scrape_interval: 30s # was 15s
retention.time: 15d # was 30d
Or limit series per job
scrape_configs:
- job_name: 'app'
scrape_interval: 30s
metric_relabel_configs:
- source_labels: [__name__]
regex: 'unwanted_metric_.*'
action: drop
Remember

High cardinality (millions of unique label combinations) is the #1 cause of Prometheus memory issues. Avoid unbounded label values like user IDs or request URLs.

4 · "context deadline exceeded" in targets​

Problem

Targets show "context deadline exceeded" errors.

Cause: The target is too slow to respond. Prometheus times out before getting the metrics.

Fix:

Increase the scrape timeout
scrape_configs:
- job_name: 'slow-app'
scrape_timeout: 30s # default is 10s
static_configs:
- targets: ['slow-app:8080']
Check if the target is actually slow
time curl http://slow-app:8080/metrics
Remember

If a target takes more than 10 seconds to respond, the application is usually overloaded. Increasing the timeout is a band-aid — fix the root cause.

5 · Alert is firing but shouldn't be​

Problem

An alert fires even though everything seems fine.

Cause: The PromQL expression is too sensitive, the for duration is too short, or there's a brief spike.

Fix:

Test the expression manually
curl -s 'http://localhost:9090/api/v1/query?query=YOUR_EXPRESSION'
Increase the for duration
- alert: HighCPU
expr: ... > 80
for: 5m # was 2m — requires 5 minutes of sustained high CPU
Or raise the threshold
- alert: HighCPU
expr: ... > 90 # was 80
for: 5m
Remember

False alarms erode trust in the monitoring system. Tune thresholds and for durations until alerts are meaningful.

6 · Alertmanager not sending notifications​

Problem

Prometheus shows alerts firing, but Alertmanager isn't sending emails/Slack.

Cause: Alertmanager isn't receiving alerts from Prometheus, or the receiver configuration is wrong.

Fix:

Check Alertmanager is running
curl http://localhost:9093/api/v2/alerts
Check Prometheus is connected to Alertmanager
curl http://localhost:9090/api/v1/alertmanagers
Verify the alertmanager URL in prometheus.yml
alerting:
alertmanagers:
- static_configs:
- targets: ['alertmanager:9093'] # must match the container name
Remember

The Alertmanager URL in prometheus.yml must be reachable from the Prometheus container. Use Docker service names, not localhost.