Skip to main content

Core Concepts — The Hospital Monitoring Blueprint

Before the head nurse can set up the monitoring room, everyone needs to learn the vocabulary. Prometheus has a small set of core concepts — once you understand these, every configuration makes sense.

The Big Picture​

Metrics — The Patient's Vitals​

Metrics are numeric measurements collected over time. Think of them as heart rate, blood pressure, and temperature — but for servers and applications.

There are four metric types:

TypeWhat it measuresExample
CounterA value that only goes upTotal HTTP requests, total errors
GaugeA value that goes up and downCPU usage, memory usage, queue size
HistogramDistribution of values (buckets)Request duration, response sizes
SummarySimilar to histogram (pre-calculated)Request latency percentiles

Counter — Total Heartbeats​

Counter — total HTTP requests
http_requests_total

The value only increases. To get "requests per second," use rate():

Requests per second
rate(http_requests_total[5m])

Gauge — Current Temperature​

Gauge — current memory usage
node_memory_MemAvailable_bytes

The value can go up or down. No special function needed.

Histogram — Blood Pressure Distribution​

Histogram — request duration buckets
http_request_duration_seconds_bucket

Histograms group values into buckets (e.g., "100 requests took 0.1-0.5s"). Use histogram_quantile() to get percentiles:

95th percentile latency
histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m]))
Remember

Counters go up → use rate(). Gauges fluctuate → use directly. Histograms → use histogram_quantile() for percentiles.

Targets — The Patients​

A target is a system that Prometheus monitors. Each target exposes metrics over HTTP. Prometheus periodically scrapes (fetches) these metrics.

prometheus.yml — define targets
scrape_configs:
- job_name: 'campus-library'
static_configs:
- targets: ['localhost:3000']

Prometheus will HTTP GET http://localhost:3000/metrics every scrape interval.

Remember

Every application must expose a /metrics endpoint. This is the "heartbeat" that Prometheus reads.

Exporters — The Medical Instruments​

An exporter converts metrics from a system that doesn't natively expose Prometheus metrics into a format Prometheus can scrape.

ExporterMonitors
node_exporterLinux servers (CPU, memory, disk, network)
mysqld_exporterMySQL databases
redis_exporterRedis
blackbox_exporterHTTP endpoints, TCP ports, DNS, ICMP
kubernetes-podKubernetes pods (auto-discovered)
Scrape node_exporter
scrape_configs:
- job_name: 'node'
static_configs:
- targets: ['server1:9100', 'server2:9100']
Remember

If a system doesn't expose Prometheus metrics, find or build an exporter. Exporters are the "translation layer" between your system and Prometheus.

Labels — The Patient's Name Tag​

Labels are key-value pairs attached to metrics. They let you filter and group metrics.

Labels — filter by method and path
http_requests_total{method="GET", path="/api/books"}
Labels — group by status code
sum by (status) (rate(http_requests_total[5m]))
Remember

Labels are how you ask questions: "How many GET requests returned 500 errors in the last 5 minutes?"

Scrape Interval — The Check-up Frequency​

The scrape interval defines how often Prometheus fetches metrics from targets. The default is 15 seconds.

prometheus.yml
global:
scrape_interval: 15s # Check every 15 seconds
evaluation_interval: 15s # Evaluate alert rules every 15 seconds
Remember

Shorter intervals = more data, more storage, more load. 15s is fine for most use cases. Use 5s only for critical metrics.

Retention — How Long to Keep Records​

prometheus.yml
global:
retention.time: 30d # Keep data for 30 days
retention.size: 10GB # Or until 10GB, whichever comes first
Remember

Prometheus is not designed for long-term storage. For years of data, use Thanos, Cortex, or VictoriaMetrics.