深色模式
采集性能优化
指标越来越多,Prometheus 越跑越慢。本文给出降负载、降基数的几招实用手段。
适用环境
- Prometheus 已运行且资源吃紧(CPU/内存/磁盘高)
- 能修改 prometheus.yml 与规则
bash
# 看 TSDB 头信息,了解序列数与 chunk 情况
curl -s http://localhost:9090/api/v1/status/tsdb | head -c 5001
2
2
操作步骤
1. 降低抓取频率
非关键任务拉长 scrape_interval:
yaml
- job_name: node
scrape_interval: 30s # 从 15s 提到 30s,负载减半1
2
2
2. 用 metric_relabel 削减标签
yaml
metric_relabel_configs:
- regex: '(container_label_.*|pod_template_hash)'
action: labeldrop1
2
3
2
3
3. 降采样:只保留聚合结果
把高频原始指标通过 recording rules 聚合成低基数指标,查询只查聚合值。
4. 分片:按 target 拆多个 Prometheus
用 hashmod 把大量 target 均分到多个实例:
yaml
relabel_configs:
- source_labels: [__address__]
modulus: 2
target_label: __tmp_shard
action: hashmod
- source_labels: [__tmp_shard]
regex: 0
action: keep1
2
3
4
5
6
7
8
2
3
4
5
6
7
8
验证
优化前后对比:
promql
# 序列总数
prometheus_tsdb_head_series
# 抓取耗时
scrape_duration_seconds1
2
3
4
2
3
4
常见坑
降基数别误删关键标签
labeldrop 删错标签会导致无法区分实例,先在小范围验证。
分片 hashmod 需一致
所有分片用相同 modulus 与 source_labels,否则目标会被重复抓或漏抓。