Skip to main content

Glossary — Speak the Language

The five friends may use hospital words, but in job interviews and documentation people use these. This page has everything you need to talk about Prometheus with confidence: a quick-reference table, a set of interview-ready answers, and a Classic Interview Q&A section.

Quick Reference​

TermHospital meaningPlain meaning
PrometheusThe monitoring roomAn open-source systems monitoring and alerting toolkit that collects and queries metrics
MetricA patient's vital signA numeric measurement collected over time (CPU usage, request count, etc.)
Time SeriesA patient's health record over timeA sequence of data points for a specific metric with a unique set of labels
LabelThe patient's name tagKey-value pairs attached to metrics for filtering and grouping
ScrapeTaking a patient's readingsPrometheus fetching metrics from a target's HTTP endpoint
Scrape TargetA patient being monitoredAn endpoint that exposes metrics for Prometheus to collect
Scrape IntervalCheck-up frequencyHow often Prometheus fetches metrics (default: 15s)
TargetA system being monitoredA service or endpoint configured for Prometheus to scrape
ExporterA medical instrumentA translator that converts a system's native metrics into Prometheus format
PromQLThe diagnostic languagePrometheus Query Language — used to query and analyze metrics
CounterTotal heartbeats (only goes up)A metric that only increases (total requests, total errors)
GaugeCurrent temperature (goes up and down)A metric that can increase or decrease (CPU usage, queue size)
HistogramBlood pressure distributionA metric that measures distribution of values in buckets
SummaryPre-calculated statisticsSimilar to histogram but calculates quantiles on the client side
Alert RuleDiagnostic thresholdA PromQL expression that triggers an alert when the condition is true
AlertAn alarm soundingA notification triggered when an alert rule fires
AlertmanagerThe alarm routing systemReceives alerts from Prometheus, deduplicates, groups, and routes them
SilenceMuting an alarm during maintenanceTemporarily suppressing alerts (e.g., during planned downtime)
InhibitionSuppressing lower alarmsPreventing less critical alerts when a critical one is firing
ReceiverWho gets the alarmThe destination for alerts (email, Slack, PagerDuty, etc.)
Evaluation IntervalHow often the nurse checks thresholdsHow often Prometheus evaluates alert rules (default: 15s)
RetainingHow long to keep recordsHow long Prometheus stores data (default: 15 days)
TSDBThe patient record databaseTime Series Database — Prometheus's storage engine for metrics
rate()Heart rate calculatorConverts a counter to a per-second rate over a time window
increase()Growth over a periodThe total increase of a counter over a time range
histogram_quantile()Percentile calculatorCalculates a quantile from histogram buckets
Recording RulePre-computed vital signA rule that pre-computes expensive queries and saves them as new metrics
FederationSharing records between hospitalsScraping metrics from one Prometheus server into another
Remote WriteSending records to a central hospitalSending metrics to a remote storage system (Thanos, Cortex)
Service DiscoveryAutomatic patient registrationAutomatically finding targets to scrape (Kubernetes, Consul, DNS)

Interview-Ready Answers​

Core Concepts​

Prometheus Prometheus is an open-source monitoring and alerting system designed for reliability. It scrapes metrics from HTTP endpoints, stores them in a time series database, and evaluates alert rules using PromQL. It follows a pull-based model — Prometheus pulls metrics from targets, rather than targets pushing to Prometheus. The gotcha: Prometheus is not designed for long-term storage. Use Thanos, Cortex, or VictoriaMetrics for years of data.

Metric Types There are four types: Counter (monotonically increasing, like total requests), Gauge (can go up or down, like CPU usage), Histogram (distribution of values in buckets), and Summary (pre-calculated quantiles). The gotcha: always use rate() with counters. Using a counter directly gives you the total count, not the current speed.

PromQL PromQL is the query language for Prometheus. It lets you filter, aggregate, and compute metrics. Common patterns include rate() for counters, histogram_quantile() for percentiles, and sum by (label) for aggregation. The gotcha: PromQL operates on time ranges. A query without a range selector ([5m]) returns the latest instant value.

Scrape Model (Pull-based) Prometheus pulls metrics from targets over HTTP. This means Prometheus must be able to reach the target's network. If the target is behind a firewall, metrics won't be collected. The gotcha: pull-based means Prometheus is not in the critical path. If Prometheus goes down, your application keeps running — you just lose monitoring data.

Alerting​

Alert Rules Alert rules are PromQL expressions evaluated at a regular interval. When the expression returns a result, the alert fires. The for field requires the condition to be true for a sustained period before firing, preventing false alarms from brief spikes. The gotcha: always test alert rules before deploying. A wrong expression can miss real problems or create alert fatigue.

Alertmanager Alertmanager receives alerts from Prometheus and handles deduplication, grouping, routing, and notification. It groups related alerts (so you get one notification instead of five), routes by severity (critical → PagerDuty, warning → Slack), and supports silences for maintenance windows. The gotcha: Alertmanager is a separate component. If it's not running, Prometheus fires alerts but nobody hears them.

Prometheus vs Alertmanager Prometheus evaluates rules and detects problems. Alertmanager handles notifications. They communicate over HTTP. Prometheus sends alerts to Alertmanager, which routes them to the right people. The gotcha: both must be running for alerting to work. A common mistake is configuring alerts in Prometheus but forgetting to set up Alertmanager.

Classic Interview Q&A​

Q1: Explain the difference between pull-based and push-based monitoring.​

Answer: In pull-based monitoring (Prometheus), the monitoring system connects to the target and fetches metrics over HTTP. In push-based monitoring (StatsD, Datadog agent), the application sends metrics to the monitoring server. Pull-based advantages: no client configuration needed, Prometheus knows exactly what's being monitored, and the monitoring system controls the collection rate. Push-based advantages: works through firewalls, doesn't require the target to be reachable. The gotcha: Prometheus supports push gateway for short-lived jobs that can't be scraped.

Q2: What is the difference between a Counter, Gauge, Histogram, and Summary?​

Answer: A Counter only increases (total requests, total errors). A Gauge can go up and down (CPU usage, queue size). A Histogram measures the distribution of values in configurable buckets (request duration). A Summary is similar to Histogram but calculates quantiles on the client side. Use Counters with rate(), Gauges directly, and Histograms with histogram_quantile(). The gotcha: never use rate() on a Gauge — it's meaningless.

Q3: How does Prometheus handle high availability?​

Answer: Prometheus alone is not highly available — it's a single process. For HA, run two identical Prometheus servers scraping the same targets. Use load balancing in front, or use Thanos/Cortex for a multi-replica setup. For alerting, multiple Alertmanagers can be run in a cluster for HA. The gotcha: without Thanos/Cortex, each Prometheus instance has its own data. Queries against one instance won't see data from the other.

Q4: What is the difference between rate() and increase()?​

Answer: rate() calculates the per-second average rate of increase of a counter over a time window. It's like measuring heart rate (beats per second). increase() calculates the total increase of a counter over a time window. It's like counting total heartbeats in the last hour. Use rate() for dashboards (smooth, per-second values). Use increase() for alerts (total count over a period).

Q5: How do you reduce alert fatigue?​

Answer: Increase the for duration so brief spikes don't fire. Tune thresholds to meaningful levels. Use inhibition to suppress lower-severity alerts when critical ones fire. Group related alerts so you get one notification instead of five. Use silence for planned maintenance. The gotcha: alert fatigue is a real problem. If engineers ignore alerts because there are too many false alarms, they'll miss the real ones.

Q6: What is service discovery and why does Prometheus need it?​

Answer: Service discovery automatically finds targets to scrape. Instead of hardcoding every target's IP address, Prometheus queries a registry (Kubernetes API, Consul, DNS) to discover targets dynamically. This is critical in modern environments where containers are created and destroyed constantly. The gotcha: service discovery requires the target to be registered. If a target isn't in the registry, Prometheus won't scrape it.

Q7: Explain the four golden signals of monitoring.​

Answer: Google's SRE book defines four golden signals: Latency (how long requests take), Traffic (how many requests), Errors (how many fail), and Saturation (how full the system is). These four metrics give a complete picture of system health. Monitor latency percentiles (not just averages), total traffic, error rates, and resource utilization (CPU, memory, disk). The gotcha: averages hide problems. Always use percentiles for latency and percentiles for error rates.