Skip to main content

Troubleshooting — The Dashboard Problem Desk

Everyone eventually sees a confusing Grafana error. This is the Problem Desk: what the message means, why it happened, and how to get out.

1 · "No data" in panels​

Problem

Panels show "No data" even though Prometheus has data.

Cause: The data source is wrong, the query returns no results, or the time range doesn't match.

Fix:

Test the query directly in Prometheus
curl -s 'http://localhost:9090/api/v1/query?query=up'
Check the data source URL
# In Grafana: Settings → Data Sources → Prometheus
# URL must be reachable from the Grafana container
Check the time range
# Make sure the dashboard time range covers data that exists
# Click the time range picker → select "Last 1 hour"
Remember

"No data" usually means the query returns nothing. Test the query in Prometheus first. If Prometheus has data, the data source URL is wrong.

2 · "Datasource refused connection"​

Problem

Grafana can't connect to the data source.

Cause: The data source URL is wrong, the service isn't running, or the network is blocked.

Fix:

Check if Prometheus is running
curl http://localhost:9090/api/v1/status/config
Check the URL in Grafana data source config
# Inside Docker: use service names (http://prometheus:9090)
# Outside Docker: use localhost (http://localhost:9090)
Check they're on the same Docker network
docker network ls
docker network inspect <network-name>
Remember

Inside Docker, localhost refers to the container itself, not the host. Use Docker service names for inter-container communication.

3 · Panel shows wrong values​

Problem

The panel shows values that don't seem right.

Cause: The query aggregation is wrong, the time range is misleading, or units are incorrect.

Fix:

Check the query is aggregating correctly
# Bad: showing per-series values
http_requests_total

# Good: aggregated across all series
sum(rate(http_requests_total[5m]))
Check the time range
# A narrow time range (5 minutes) can show misleading spikes
# A wide range (7 days) can smooth out important details
Remember

Always aggregate with sum by (label) or avg by (label). Raw metric names show one line per label combination — usually too many.

4 · Dashboard is slow to load​

Problem

The dashboard takes a long time to load or refresh.

Cause: Too many panels, too many queries per panel, or the time range is too wide.

Fix:

Reduce the number of panels
# Keep dashboards focused — one dashboard per service or signal type
Increase the refresh interval
# Default: 30s. Change to 60s for less load.
# Dashboard Settings → Refresh rate
Narrow the time range
# Don't show 30 days by default — use 1 hour
# Users can always zoom out
Remember

Each panel runs a query on every refresh. 20 panels × 30s refresh = 40 queries per minute. Reduce panels or increase refresh interval.

5 · Alerts not firing​

Problem

Alert rules are defined but never fire.

Cause: The condition is never true, the for duration is too long, or the data source is not returning data.

Fix:

Check the alert rule state
# Alerting → Alert rules → check if the rule is "Normal" or "No data"
Test the condition manually
# In the panel editor, run the query and check if it returns a result
Check evaluation history
# Alerting → Alert rules → click the rule → see evaluation timeline
Remember

"No data" means the alert can't evaluate. Make sure the data source is returning data before debugging the alert rule.

6 · Variables not populating​

Problem

The variable dropdown is empty.

Cause: The variable query returns no results, or the data source is wrong.

Fix:

Test the query in Prometheus
curl -s 'http://localhost:9090/api/v1/label/job/values'
Check the variable query in Grafana
# Dashboard Settings → Variables → click the variable → test the query
Remember

Variable queries must return label values. label_values(metric, label) is the most common pattern.