深色模式
高基数问题治理
摘要:基数(cardinality)是可观测性成本与稳定性的头号杀手——一个
user_id标签就能让时间序列从几千涨到几千万。本文教你定位高基数标签、用 Collector 在入口处拦截,并通过聚合降维把规模拉回可控范围。
适用环境
bash
# Prometheus 当前时间序列总数
curl -s http://127.0.0.1:9090/api/v1/status/tsdb | python3 -c \
"import sys,json;print('numSeries:',json.load(sys.stdin)['data']['headStats']['numSeries'])"
# 按指标统计时间序列数(找出最占地方的指标)
curl -s 'http://127.0.0.1:9090/api/v1/query' --data-urlencode \
'query=topk(20, count by (__name__)({__name__=~".+"}))'1
2
3
4
5
6
7
2
3
4
5
6
7
操作步骤
1. 找出高基数元凶
promql
# 各指标的时间序列数量排行
topk(20, count by (__name__)({__name__=~".+"}))
# 某个指标的标签基数(看哪个标签取值最多)
count(count by (user_id) (http_server_request_duration_seconds_count))1
2
3
4
5
2
3
4
5
bash
# 用 promtool 分析 TSDB 块,输出基数最高的标签对
promtool tsdb analyze /data/prometheus | head -301
2
2
2. 识别典型的高基数字段
bash
# 这些字段一旦进入指标标签就是灾难:
# user_id / order_id / request_id / trace_id / session_id / ip / url_full / timestamp
# 应留在:日志与链路(适合存明细),而不是指标(只存聚合值)1
2
3
2
3
3. 在 Collector 入口处拦截(推荐,最省事)
yaml
processors:
# 3.1 直接删除高危标签
attributes/drop:
actions:
- key: user.id
action: delete
- key: net.peer.ip
action: delete
# 3.2 把高基数值替换为低基数分组(如 /user/123 -> /user/{id})
attributes/normalize:
actions:
- key: http.route
pattern: '^/user/[0-9]+'
replacement: '/user/{id}'
action: regex_replace # 需使用 transform 处理器实现
# 3.3 聚合掉不需要细粒度的标签
metrics/aggregate:
transforms:
- include: http_server_request_duration_seconds
action: update
operations:
- action: aggregate_labels
label_set: [service_name, http_route, http_method]
aggregation_type: sum1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
4. 用 Prometheus 侧的 relabel 兜底
yaml
scrape_configs:
- job_name: app
metric_relabel_configs:
- source_labels: [__name__]
regex: 'debug_.*'
action: drop # 丢弃调试指标
- regex: 'user_id|request_id|trace_id'
action: labeldrop # 删除指定标签1
2
3
4
5
6
7
8
2
3
4
5
6
7
8
5. 对链路做同样的治理
yaml
# span 属性同样受语义约定约束,避免把完整 SQL、请求体写进属性
processors:
attributes/trace:
actions:
- key: db.statement
action: delete # 或截断后再保留
- key: http.request.body
action: delete1
2
3
4
5
6
7
8
2
3
4
5
6
7
8
6. 设置基数预算与监控
promql
# Prometheus 自身指标:当前序列数与增长速率
prometheus_tsdb_head_series
# 抓取样本速率(也能反映基数压力)
rate(prometheus_tsdb_head_samples_appended_total[5m])1
2
3
4
2
3
4
yaml
groups:
- name: cardinality
rules:
- alert: TSDBSeriesTooHigh
expr: prometheus_tsdb_head_series > 5000000
for: 10m
labels: {severity: warning}
annotations: {summary: "时间序列数超过预算,需排查高基数标签"}
- alert: SeriesGrowthFast
expr: deriv(prometheus_tsdb_head_series[1h]) > 100000
for: 30m
labels: {severity: warning}1
2
3
4
5
6
7
8
9
10
11
12
2
3
4
5
6
7
8
9
10
11
12
验证
bash
# 1) 治理前后对比序列总数
curl -s http://127.0.0.1:9090/api/v1/status/tsdb | python3 -c \
"import sys,json;print(json.load(sys.stdin)['data']['headStats']['numSeries'])"
# 2) 高危标签是否已消失
curl -s http://127.0.0.1:9090/api/v1/label/user_id/values | head -c 100
# 3) Collector 配置合法
otelcol-contrib validate --config /etc/otelcol/config.yaml
# 4) 治理后看板仍能正常聚合(人工核对关键面板不为空)1
2
3
4
5
6
7
8
9
10
11
2
3
4
5
6
7
8
9
10
11
常见坑
WARNING
只删标签不降维会让指标失去分析能力。正确做法是"转成低基数分组":用户 ID 换成哈希分桶、URL 换成路由模板、IP 换成网段或地域。
WARNING
在 Prometheus 侧做 labeldrop 只能治标——数据已经传到了抓取端,网络与解析开销已经发生。优先在 Collector 或应用埋点处解决。
DANGER
为了降基数把 instance、pod 这类定位标签也聚合掉,会导致故障时无法定位到具体实例。降基数前请明确哪些标签是排查必需的,保留最小可定位维度。