Glossary — Speak the Language
The five friends may use hospital words, but in job interviews and documentation people use these. This page has everything you need to talk about Prometheus with confidence: a quick-reference table, a set of interview-ready answers, and a Classic Interview Q&A section.
Quick Reference
| Term | Hospital meaning | Plain meaning |
|---|---|---|
| Prometheus | The monitoring room | An open-source systems monitoring and alerting toolkit that collects and queries metrics |
| Metric | A patient's vital sign | A numeric measurement collected over time (CPU usage, request count, etc.) |
| Time Series | A patient's health record over time | A sequence of data points for a specific metric with a unique set of labels |
| Label | The patient's name tag | Key-value pairs attached to metrics for filtering and grouping |
| Scrape | Taking a patient's readings | Prometheus fetching metrics from a target's HTTP endpoint |
| Scrape Target | A patient being monitored | An endpoint that exposes metrics for Prometheus to collect |
| Scrape Interval | Check-up frequency | How often Prometheus fetches metrics (default: 15s) |
| Target | A system being monitored | A service or endpoint configured for Prometheus to scrape |
| Exporter | A medical instrument | A translator that converts a system's native metrics into Prometheus format |
| PromQL | The diagnostic language | Prometheus Query Language — used to query and analyze metrics |
| Counter | Total heartbeats (only goes up) | A metric that only increases (total requests, total errors) |
| Gauge | Current temperature (goes up and down) | A metric that can increase or decrease (CPU usage, queue size) |
| Histogram | Blood pressure distribution | A metric that measures distribution of values in buckets |
| Summary | Pre-calculated statistics | Similar to histogram but calculates quantiles on the client side |
| Alert Rule | Diagnostic threshold | A PromQL expression that triggers an alert when the condition is true |
| Alert | An alarm sounding | A notification triggered when an alert rule fires |
| Alertmanager | The alarm routing system | Receives alerts from Prometheus, deduplicates, groups, and routes them |
| Silence | Muting an alarm during maintenance | Temporarily suppressing alerts (e.g., during planned downtime) |
| Inhibition | Suppressing lower alarms | Preventing less critical alerts when a critical one is firing |
| Receiver | Who gets the alarm | The destination for alerts (email, Slack, PagerDuty, etc.) |
| Evaluation Interval | How often the nurse checks thresholds | How often Prometheus evaluates alert rules (default: 15s) |
| Retaining | How long to keep records | How long Prometheus stores data (default: 15 days) |
| TSDB | The patient record database | Time Series Database — Prometheus's storage engine for metrics |
| rate() | Heart rate calculator | Converts a counter to a per-second rate over a time window |
| increase() | Growth over a period | The total increase of a counter over a time range |
| histogram_quantile() | Percentile calculator | Calculates a quantile from histogram buckets |
| Recording Rule | Pre-computed vital sign | A rule that pre-computes expensive queries and saves them as new metrics |
| Federation | Sharing records between hospitals | Scraping metrics from one Prometheus server into another |
| Remote Write | Sending records to a central hospital | Sending metrics to a remote storage system (Thanos, Cortex) |
| Service Discovery | Automatic patient registration | Automatically finding targets to scrape (Kubernetes, Consul, DNS) |
Interview-Ready Answers
Core Concepts
Prometheus Prometheus is an open-source monitoring and alerting system designed for reliability. It scrapes metrics from HTTP endpoints, stores them in a time series database, and evaluates alert rules using PromQL. It follows a pull-based model — Prometheus pulls metrics from targets, rather than targets pushing to Prometheus. The gotcha: Prometheus is not designed for long-term storage. Use Thanos, Cortex, or VictoriaMetrics for years of data.
Metric Types
There are four types: Counter (monotonically increasing, like total requests), Gauge (can go up or down, like CPU usage), Histogram (distribution of values in buckets), and Summary (pre-calculated quantiles). The gotcha: always use rate() with counters. Using a counter directly gives you the total count, not the current speed.
PromQL
PromQL is the query language for Prometheus. It lets you filter, aggregate, and compute metrics. Common patterns include rate() for counters, histogram_quantile() for percentiles, and sum by (label) for aggregation. The gotcha: PromQL operates on time ranges. A query without a range selector ([5m]) returns the latest instant value.
Scrape Model (Pull-based) Prometheus pulls metrics from targets over HTTP. This means Prometheus must be able to reach the target's network. If the target is behind a firewall, metrics won't be collected. The gotcha: pull-based means Prometheus is not in the critical path. If Prometheus goes down, your application keeps running — you just lose monitoring data.
Alerting
Alert Rules
Alert rules are PromQL expressions evaluated at a regular interval. When the expression returns a result, the alert fires. The for field requires the condition to be true for a sustained period before firing, preventing false alarms from brief spikes. The gotcha: always test alert rules before deploying. A wrong expression can miss real problems or create alert fatigue.
Alertmanager Alertmanager receives alerts from Prometheus and handles deduplication, grouping, routing, and notification. It groups related alerts (so you get one notification instead of five), routes by severity (critical → PagerDuty, warning → Slack), and supports silences for maintenance windows. The gotcha: Alertmanager is a separate component. If it's not running, Prometheus fires alerts but nobody hears them.
Prometheus vs Alertmanager Prometheus evaluates rules and detects problems. Alertmanager handles notifications. They communicate over HTTP. Prometheus sends alerts to Alertmanager, which routes them to the right people. The gotcha: both must be running for alerting to work. A common mistake is configuring alerts in Prometheus but forgetting to set up Alertmanager.
Classic Interview Q&A
Q1: Explain the difference between pull-based and push-based monitoring.
Answer: In pull-based monitoring (Prometheus), the monitoring system connects to the target and fetches metrics over HTTP. In push-based monitoring (StatsD, Datadog agent), the application sends metrics to the monitoring server. Pull-based advantages: no client configuration needed, Prometheus knows exactly what's being monitored, and the monitoring system controls the collection rate. Push-based advantages: works through firewalls, doesn't require the target to be reachable. The gotcha: Prometheus supports push gateway for short-lived jobs that can't be scraped.
Q2: What is the difference between a Counter, Gauge, Histogram, and Summary?
Answer: A Counter only increases (total requests, total errors). A Gauge can go up and down (CPU usage, queue size). A Histogram measures the distribution of values in configurable buckets (request duration). A Summary is similar to Histogram but calculates quantiles on the client side. Use Counters with rate(), Gauges directly, and Histograms with histogram_quantile(). The gotcha: never use rate() on a Gauge — it's meaningless.
Q3: How does Prometheus handle high availability?
Answer: Prometheus alone is not highly available — it's a single process. For HA, run two identical Prometheus servers scraping the same targets. Use load balancing in front, or use Thanos/Cortex for a multi-replica setup. For alerting, multiple Alertmanagers can be run in a cluster for HA. The gotcha: without Thanos/Cortex, each Prometheus instance has its own data. Queries against one instance won't see data from the other.
Q4: What is the difference between rate() and increase()?
Answer: rate() calculates the per-second average rate of increase of a counter over a time window. It's like measuring heart rate (beats per second). increase() calculates the total increase of a counter over a time window. It's like counting total heartbeats in the last hour. Use rate() for dashboards (smooth, per-second values). Use increase() for alerts (total count over a period).
Q5: How do you reduce alert fatigue?
Answer: Increase the for duration so brief spikes don't fire. Tune thresholds to meaningful levels. Use inhibition to suppress lower-severity alerts when critical ones fire. Group related alerts so you get one notification instead of five. Use silence for planned maintenance. The gotcha: alert fatigue is a real problem. If engineers ignore alerts because there are too many false alarms, they'll miss the real ones.
Q6: What is service discovery and why does Prometheus need it?
Answer: Service discovery automatically finds targets to scrape. Instead of hardcoding every target's IP address, Prometheus queries a registry (Kubernetes API, Consul, DNS) to discover targets dynamically. This is critical in modern environments where containers are created and destroyed constantly. The gotcha: service discovery requires the target to be registered. If a target isn't in the registry, Prometheus won't scrape it.
Q7: Explain the four golden signals of monitoring.
Answer: Google's SRE book defines four golden signals: Latency (how long requests take), Traffic (how many requests), Errors (how many fail), and Saturation (how full the system is). These four metrics give a complete picture of system health. Monitor latency percentiles (not just averages), total traffic, error rates, and resource utilization (CPU, memory, disk). The gotcha: averages hide problems. Always use percentiles for latency and percentiles for error rates.