Dashboard Building — Assembling the Instrument Cluster
The mechanic has the dashboard frame. Now it's time to install the gauges. Building a Grafana dashboard is about choosing the right panel type, writing the right query, and arranging everything so the driver (engineer) sees the right information at a glance.
Creating Your First Dashboard
- Click + → New Dashboard
- Click Add visualization
- Select Prometheus as the data source
- Enter a PromQL query
- Choose a visualization type
- Click Apply
Start with the four golden signals: Latency, Traffic, Errors, Saturation. One panel for each gives you a complete overview.
Panel Types — Choosing the Right Gauge
Time Series — The Speedometer
Shows values changing over time. This is the most common panel type.
sum(rate(http_requests_total[5m])) by (method)
Best for: Anything that changes over time — CPU, memory, request rate, latency.
Stat — The Odometer
Shows a single current value with optional color thresholds.
sum(http_requests_total{status=~"5.."})
Best for: Key metrics you want to see at a glance — total errors, uptime percentage, active connections.
Gauge — The Fuel Gauge
Shows a value within a min/max range with color bands.
100 - (avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
Settings:
- Min: 0, Max: 100
- Thresholds: green < 70, yellow 70-90, red > 90
Best for: Percentages — CPU usage, memory usage, disk usage.
Table — The Diagnostic Report
Shows structured data in rows and columns.
topk(10, sum(rate(http_requests_total[5m])) by (path))
Best for: Comparing many items — top endpoints, slowest queries, largest pods.
Bar Gauge — The Comparison Chart
Compares values across categories.
sum(rate(http_requests_total[5m])) by (service)
Best for: Comparing values across labels — requests per service, memory per pod, CPU per node.
Match the panel type to the question. "How is it changing?" → Time series. "What is it now?" → Stat. "How do they compare?" → Bar gauge.
Query Tips — Writing Better PromQL
Use legends for readability
sum(rate(http_requests_total[5m])) by (method)
# Legend: {{method}} → shows "GET", "POST", etc.
Multiple queries in one panel
# Query A: Request rate
sum(rate(http_requests_total[5m]))
# Query B: Error rate
sum(rate(http_requests_total{status=~"5.."}[5m]))
Math operations between queries
# Query A: total requests
# Query B: error requests
# Then use: B / A * 100
Legends turn cryptic metric names into human-readable labels. Always add a legend to your queries.
Thresholds — Color-Coding Health
Thresholds change panel colors based on values:
thresholds:
steps:
- color: green
value: null # 0 to first threshold
- color: yellow
value: 70 # 70% to 90%
- color: red
value: 90 # 90% and above
| Color | Meaning |
|---|---|
| Green | Healthy |
| Yellow | Warning — investigate soon |
| Red | Critical — act now |
Thresholds are the visual equivalent of alert rules. They make problems visible at a glance.
Dashboard Organization — Best Practices
1. Top row: Overview panels
Put the most important metrics at the top. Use Stat panels for:
- Total requests
- Error rate
- Average latency
- Uptime
2. Group by signal type
- Row 1: Overview (stat panels)
- Row 2: Traffic (request rate graphs)
- Row 3: Latency (response time graphs)
- Row 4: Errors (error rate graphs)
- Row 5: Saturation (CPU, memory, disk)
3. Use variables for multi-service dashboards
templating:
list:
- name: service
type: query
query: label_values(http_requests_total, service)
- name: instance
type: query
query: label_values(http_requests_total{service="$service"}, instance)
A well-organized dashboard answers questions in order: "Is everything healthy?" → "Which service has a problem?" → "What's the root cause?"
Export and Import — Sharing Dashboards
# In Grafana UI: Dashboard → Settings → JSON Model → Copy
# In Grafana UI: + → Import → Paste JSON or enter dashboard ID
Grafana has a community dashboard library at grafana.com/grafana/dashboards. Popular IDs:
- Node Exporter Full: 1860
- Docker: 893
- Kubernetes: 315
Don't build from scratch if a community dashboard exists. Import it and customize.