Data Model
Prometheus' core components—scrape targets, time series data model, and storage architecture—form the foundation of its observability capabilities. These elements work together to enable efficient, scalable metric collection and querying, even as systems grow in complexity and scale. Understanding how they interoperate is critical for designing robust monitoring pipelines.
Scrape Targets and Metric Collection¶
Scrape targets are the sources from which Prometheus pulls metrics. These can be HTTP endpoints, exporters, or other metric sources. Prometheus uses scrape configurations to define how and where to collect metrics, including intervals, job names, and authentication details.
Example: Static Scrape Configuration¶
Prometheus supports service discovery (e.g., Kubernetes, Consul) to dynamically discover targets, reducing manual configuration. Exporters like node_exporter or blackbox_exporter expose metrics in a format compatible with Prometheus' scraping mechanism.
Time Series Data Model¶
Prometheus uses a time series data model to store metrics. Each time series is uniquely identified by:
- Metric name (e.g., http_requests_total)
- Labels (key-value pairs, e.g., {job="node_exporter", instance="localhost"})
- Timestamp (seconds since epoch)
This model allows for flexible filtering and aggregation. Labels act as metadata, enabling users to group metrics by dimensions like environment, service, or region.
Example: Querying Time Series¶
This query filters metrics for HTTP 200 responses from the node_exporter job, leveraging labels for precision.
Storage Architecture¶
Prometheus stores metrics in a disk-based storage engine optimized for time series data. Key aspects include:
- Chunk storage: Metrics are stored in chunks (fixed-size blocks) by timestamp, enabling efficient compaction and retrieval.
- Retention: Data is retained indefinitely by default, but retention periods can be configured (e.g., via storage.tsdb.retention.time).
- Compaction: Old chunks are merged into larger chunks to reduce disk usage and improve query performance.
The storage engine is designed for write-heavy workloads, with asynchronous compaction and indexing to balance performance and disk space.
Enabling Scalable Metric Collection¶
The combination of these components ensures scalability: 1. Efficient data model: The time series model avoids schema changes, allowing metrics to grow without performance degradation. 2. Storage optimizations: Chunking and compaction manage disk usage, while retention policies balance data longevity with resource constraints. 3. Scrape target flexibility: Dynamic discovery and exporters enable monitoring of diverse systems without overloading the Prometheus server.
For large-scale deployments, Prometheus can be horizontally scaled using remote storage (e.g., Thanos, Cortex) to offload data and improve query performance.
Key takeaways¶
- Scrape targets define how metrics are collected, with static and dynamic discovery options.
- The time series data model enables flexible querying via metric names and labels.
- Storage architecture uses chunking and compaction to handle large datasets efficiently.
- Together, these components support scalable, reliable metric collection for complex systems.