Storage Retention
Prometheus metrics retention is a critical aspect of observability, balancing data longevity with storage efficiency. Properly configuring retention policies and storage limits ensures you retain sufficient historical data to meet SLOs while avoiding unnecessary costs or performance degradation. This section covers how to configure Prometheus' retention policies and storage settings for long-term metric storage.
Retention Policies Configuration¶
Prometheus uses retention policies to determine how long metrics are stored. These policies are defined in the storage section of the Prometheus configuration file (prometheus.yml). Each policy specifies a retention duration and a maximum storage size (if applicable).
Default Retention¶
By default, Prometheus retains metrics for 6 months when using local storage. However, this can be adjusted via the retention_time parameter. For example:
Multiple Retention Policies¶
You can define multiple policies to manage data lifecycle. The longest retention time is used as the default, but policies can be prioritized by order:
storage:
retention_time: 6m # Default retention
policies:
- name: short-term
retention_time: 1h
storage_limit: 10GB
- name: long-term
retention_time: 1y
storage_limit: 100GB
Storage Configuration¶
Local vs. Remote Storage¶
- Local storage: Metrics are stored on the host's filesystem. Use this for short-term retention or testing.
- Remote storage: For long-term retention, configure Prometheus to write metrics to remote systems like Thanos, Cortex, or VictoriaMetrics. This offloads storage management and enables scalable retention.
Example: Remote Write Configuration¶
remote_write:
- url: http://thanos-store:10901/api/v1/write
queue_config:
max_samples_per_send: 10000
Storage Limits and Data Dropping¶
Prometheus enforces storage limits via the storage_limit parameter. When storage reaches this threshold, it begins dropping the oldest data first. For example:
- Local storage: If
storage_limitis set, Prometheus will delete metrics once the limit is reached. - Remote storage: The remote system (e.g., Thanos) manages storage limits and retention policies independently.
Best Practices¶
- Align with SLOs: Set retention times based on your SLO requirements. For example, a 99.9% SLO might require 1 year of historical data.
- Use Remote Storage for Long-Term: For multi-year retention, always use remote storage to avoid local disk exhaustion.
- Monitor Storage Usage: Track storage metrics (e.g.,
prometheus_remote_storage_used_bytes) to proactively adjust limits. - Tier Data Strategically: Use short-term local storage for real-time analysis and remote storage for archival.
Diagram: Retention Policy Workflow¶
[Metrics Collected]
↓
[Local/Remote Storage]
↓
[Retention Policy Enforcement]
↓
[Data Retention (e.g., 1 year)]
↓
[Queryable Historical Data]
Key takeaways¶
- Configure
retention_timeinprometheus.ymlto define how long metrics are stored. - Use remote storage (e.g., Thanos) for long-term retention and scalability.
- Set
storage_limitto prevent local disk overflow and enable controlled data dropping. - Align retention policies with SLO requirements and monitor storage usage regularly.
- Prioritize remote storage for multi-year archival to avoid local storage constraints.