深色模式
监控自身成本
观测系统本身也是一套生产系统,而且成本随业务量线性甚至超线性增长。常见情况是:监控费用占到总账单的 5%~15%。
适用环境
- Prometheus / Loki / Elasticsearch 或商业可观测平台
- 指标、日志、链路三类数据至少有一类
- 有采集端配置权限
bash
# 先看监控数据占了多少空间
du -sh /var/lib/prometheus /var/lib/loki 2>/dev/null1
2
2
操作步骤
1. 定位成本构成
text
指标成本 = 时间序列数(基数)× 采样间隔 × 保留期 × 单价
日志成本 = 写入量(GB) × 单价 + 存储量 × 单价 + 检索次数 × 单价
链路成本 = Span 数 × 单价 + 存储量1
2
3
2
3
三者里基数和日志量是最关键的两个杠杆。
2. 治理指标基数(最重要)
promql
# 找出指标数量最多的 job / target
topk(10, count by (__name__)({__name__=~".+"}))1
2
2
promql
# 找出基数最高的指标(时间序列数)
topk(20, count by (__name__)({__name__=~".+"}))1
2
2
bash
# 查询 Prometheus 的 TSDB 状态
curl -s http://localhost:9090/api/v1/status/tsdb | jq '.data | {numSeries, headSeries}'1
2
2
高基数典型来源:
text
1. 把 URL path、user_id、trace_id 写进 label
2. histogram 分桶过多
3. 每个 Pod/容器实例产生大量短生命周期序列
4. 未使用的 exporter 指标全量采集1
2
3
4
2
3
4
用 relabel 丢弃:
yaml
scrape_configs:
- job_name: app
metrics_path: /metrics
relabel_configs:
# 丢弃高基数 label
- regex: 'url|user_id|trace_id|request_id'
action: labeldrop
metric_relabel_configs:
# 丢弃不需要的指标
- source_labels: [__name__]
regex: 'go_gc_.*|process_open_fds'
action: drop1
2
3
4
5
6
7
8
9
10
11
12
2
3
4
5
6
7
8
9
10
11
12
bash
promtool check config /etc/prometheus/prometheus.yml1
3. 调整采集间隔与保留期
yaml
global:
scrape_interval: 30s # 非核心指标从 15s 放宽到 30s,数据量减半
scrape_timeout: 10s
# 通过启动参数控制本地保留期1
2
3
4
5
2
3
4
5
bash
# 启动参数:保留 15 天(默认通常 15d)
prometheus --storage.tsdb.retention.time=15d --storage.tsdb.retention.size=50GB1
2
2
text
保留期决策:
实时排障用:7~15 天(放本地/热存储)
趋势与容量分析用:3~13 个月(用 recording rule 降精度后远程写)1
2
3
2
3
用 recording rule 降低长期数据的精度:
yaml
groups:
- name: long_term
interval: 5m
rules:
- record: job:cpu_usage:avg5m
expr: avg by (job) (rate(process_cpu_seconds_total[5m]))1
2
3
4
5
6
2
3
4
5
6
4. 日志分级与采样
yaml
# Fluent Bit:只采集 warn 以上 + 采样 info
processors:
- name: sampling
sampling:
rate: 10 # 每 10 条保留 1 条1
2
3
4
5
2
3
4
5
bash
# 应用层直接降噪:把高频 debug/info 关掉
# 统计各级别日志量占比
awk '{print $3}' app.log | sort | uniq -c | sort -rn1
2
3
2
3
text
降噪优先级:
1. 健康检查日志(/healthz 每 5 秒一条,量大且无用)
2. 重复的成功日志("request success" 类)
3. 全量 SQL 日志
4. debug 级别1
2
3
4
5
2
3
4
5
nginx
# 健康检查不记日志
location = /healthz {
access_log off;
return 200 "ok";
}1
2
3
4
5
2
3
4
5
5. 链路追踪采样
yaml
# OpenTelemetry Collector:尾部采样,保留错误与慢请求
processors:
tail_sampling:
decision_wait: 10s
policies:
- name: keep-errors
type: status_code
status_code: {status_codes: [ERROR]}
- name: keep-slow
type: latency
latency: {threshold_ms: 1000}
- name: baseline
type: probabilistic
probabilistic: {sampling_percentage: 1}1
2
3
4
5
6
7
8
9
10
11
12
13
14
2
3
4
5
6
7
8
9
10
11
12
13
14
text
基线采样 1% + 错误全采 + 慢请求全采
既保留排障能力,又把 Span 量降到 1% 量级1
2
2
6. 分层存储与估算收益
text
监控数据分层:< 7 天热存储(快查),7~90 天低频(Loki/ES 配 ILM 自动降冷),
> 90 天归档或聚合后删除。
优化前:序列数 500 万,15s 采样,保留 15 天
优化后:labeldrop 减 30% 序列 → 350 万
采样间隔 15s → 30s,数据量再减 50%
保留期分层,热数据只留 7 天
综合存储与写入成本可下降 50%~70%(以实际账单为准)1
2
3
4
5
6
7
8
2
3
4
5
6
7
8
验证
bash
# 1) 序列数是否下降
curl -s http://localhost:9090/api/v1/status/tsdb | jq '.data.headSeries'
# 2) 关键告警与看板是否仍然正常(不能因为 labeldrop 导致看板空白)
# 3) 日志检索是否满足排障需要(抽样后仍能定位问题)1
2
3
4
2
3
4
优化后必须抽查:历史对比类看板、容量趋势查询是否受影响。
常见坑
为省成本关掉关键告警
监控成本优化绝不能触碰:SLO 相关指标、错误率、饱和度告警。只能优化"从来没人看"的数据。
高基数 label 事后才发现
一个 user_id label 就能让序列数从几千涨到千万。上线新指标前必须先评估基数,并在 CI 里加基数上限检查。
采样后无法排障
均匀采样 1% 会让低频错误恰好被丢掉。必须采用"错误与慢请求全采 + 正常请求采样"的尾部采样策略。
删了历史数据才发现要做同比分析
容量预测需要 90 天以上数据。缩短保留期前,先用 recording rule 把关键聚合指标降精度长期保存。