Core Concepts — The Hospital Monitoring Blueprint
Before the head nurse can set up the monitoring room, everyone needs to learn the vocabulary. Prometheus has a small set of core concepts — once you understand these, every configuration makes sense.
The Big Picture
Metrics — The Patient's Vitals
Metrics are numeric measurements collected over time. Think of them as heart rate, blood pressure, and temperature — but for servers and applications.
There are four metric types:
| Type | What it measures | Example |
|---|---|---|
| Counter | A value that only goes up | Total HTTP requests, total errors |
| Gauge | A value that goes up and down | CPU usage, memory usage, queue size |
| Histogram | Distribution of values (buckets) | Request duration, response sizes |
| Summary | Similar to histogram (pre-calculated) | Request latency percentiles |
Counter — Total Heartbeats
http_requests_total
The value only increases. To get "requests per second," use rate():
rate(http_requests_total[5m])
Gauge — Current Temperature
node_memory_MemAvailable_bytes
The value can go up or down. No special function needed.
Histogram — Blood Pressure Distribution
http_request_duration_seconds_bucket
Histograms group values into buckets (e.g., "100 requests took 0.1-0.5s"). Use histogram_quantile() to get percentiles:
histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m]))
Counters go up → use rate(). Gauges fluctuate → use directly. Histograms → use histogram_quantile() for percentiles.
Targets — The Patients
A target is a system that Prometheus monitors. Each target exposes metrics over HTTP. Prometheus periodically scrapes (fetches) these metrics.
scrape_configs:
- job_name: 'campus-library'
static_configs:
- targets: ['localhost:3000']
Prometheus will HTTP GET http://localhost:3000/metrics every scrape interval.
Every application must expose a /metrics endpoint. This is the "heartbeat" that Prometheus reads.
Exporters — The Medical Instruments
An exporter converts metrics from a system that doesn't natively expose Prometheus metrics into a format Prometheus can scrape.
| Exporter | Monitors |
|---|---|
node_exporter | Linux servers (CPU, memory, disk, network) |
mysqld_exporter | MySQL databases |
redis_exporter | Redis |
blackbox_exporter | HTTP endpoints, TCP ports, DNS, ICMP |
kubernetes-pod | Kubernetes pods (auto-discovered) |
scrape_configs:
- job_name: 'node'
static_configs:
- targets: ['server1:9100', 'server2:9100']
If a system doesn't expose Prometheus metrics, find or build an exporter. Exporters are the "translation layer" between your system and Prometheus.
Labels — The Patient's Name Tag
Labels are key-value pairs attached to metrics. They let you filter and group metrics.
http_requests_total{method="GET", path="/api/books"}
sum by (status) (rate(http_requests_total[5m]))
Labels are how you ask questions: "How many GET requests returned 500 errors in the last 5 minutes?"
Scrape Interval — The Check-up Frequency
The scrape interval defines how often Prometheus fetches metrics from targets. The default is 15 seconds.
global:
scrape_interval: 15s # Check every 15 seconds
evaluation_interval: 15s # Evaluate alert rules every 15 seconds
Shorter intervals = more data, more storage, more load. 15s is fine for most use cases. Use 5s only for critical metrics.
Retention — How Long to Keep Records
global:
retention.time: 30d # Keep data for 30 days
retention.size: 10GB # Or until 10GB, whichever comes first
Prometheus is not designed for long-term storage. For years of data, use Thanos, Cortex, or VictoriaMetrics.