PromQL — The Monitoring Language
PromQL (Prometheus Query Language) is how you ask questions about your metrics. Think of it as the nurse's diagnostic toolkit — you use it to check vitals, spot trends, and trigger alarms.
Basic Queries
Instant vector — the current reading
node_cpu_seconds_total
This returns the latest value for every time series matching that name.
Range vector — the last hour of readings
http_requests_total[5m]
Returns all data points from the last 5 minutes.
Label filtering — narrow the search
http_requests_total{method="GET"}
http_requests_total{method="GET", status="200"}
http_requests_total{path=~"/api/.*"}
= means exact match. =~ means regex match. != means not equal.
Core Functions
rate() — Speed of the Counter
Converts a counter (total requests) into a rate (requests per second):
rate(http_requests_total[5m])
rate(http_requests_total{status=~"5.."}[5m])
Always use rate() with counters. Without it, you see the total count, not the current speed.
increase() — Growth Over a Period
increase(http_requests_total[1h])
histogram_quantile() — Percentiles
histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m]))
histogram_quantile(0.50, rate(http_request_duration_seconds_bucket[5m]))
sum() — Aggregate Across Labels
sum(rate(http_requests_total[5m]))
sum by (path) (rate(http_requests_total[5m]))
avg(), min(), max()
avg(node_cpu_seconds_total)
max(node_memory_MemAvailable_bytes)
Common Patterns
Error rate
sum(rate(http_requests_total{status=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))
* 100
Availability (uptime)
100 - (
sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))
* 100
)
Saturation (resource usage)
(1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100
(1 - node_filesystem_avail_bytes / node_filesystem_size_bytes) * 100
The four golden signals (Google SRE): Latency (how slow), Traffic (how much), Errors (how broken), Saturation (how full). PromQL queries for all four.
Aggregation Operators
| Operator | Meaning | Example |
|---|---|---|
sum | Add all values | Total requests across all pods |
avg | Average of all values | Average CPU per node |
min | Smallest value | Minimum available memory |
max | Largest value | Maximum request latency |
count | Number of series | How many targets are up |
topk | Top K values | Top 5 busiest servers |
bottomk | Bottom K values | Bottom 3 servers by memory |
topk(5, rate(node_cpu_seconds_total[5m]))
count(up == 1)
sum by (label) groups by a label. without (label) removes a label from grouping.