深色模式
SLI 采集实现
摘要:SLI 再好定义,采不到也是空谈。本文讲三类采集方式:应用埋点(counter)、访问日志解析、客户端上报,给出 Prometheus 侧落库与聚合示例。
适用环境
- 应用可埋点(如 Prometheus client)或能产生结构化日志/访问日志
- 已部署 Prometheus
操作步骤
1. 应用埋点(最准)
python
# Python 示例,用 prometheus_client
from prometheus_client import Counter
req = Counter('http_requests_total', '请求数', ['code'])
# 请求结束时打标
req.labels(code='200').inc()
req.labels(code='500').inc()1
2
3
4
5
6
2
3
4
5
6
2. 访问日志解析(无埋点时)
yaml
# promtail + loki,或从 nginx log 提取
# 用 Grafana Loki 统计 5xx 占比作为 SLI 来源
sum(rate({job="nginx"} | json | status>=500 [5m]))1
2
3
2
3
3. 客户端/合成上报
javascript
// 前端上报真实用户体验
navigator.sendBeacon('/metrics', JSON.stringify({slate:'ok', dur: ttfb}));1
2
2
4. 在 Prometheus 聚合成 SLI
promql
# 可用性 SLI
sum(rate(http_requests_total{code!~"5.."}[5m])) / sum(rate(http_requests_total[5m]))1
2
2
DANGER
埋点的“成功”判定要和 SLI 定义一致:业务失败但 HTTP 200 不能算成功,否则 SLI 虚高。
验证
bash
# 制造若干 5xx,确认 counter 增长且 SLI 表达式下降
curl -s prom:9090/api/v1/query?query=http_requests_total | jq '.data.result'1
2
2
常见坑
WARNING
- Counter 重启归零会导致 SLI 瞬间跳变,用
rate/irate而非直接比值缓解。 - 只采服务端忽略客户端体验(如渲染慢),SLI 会高估真实质量。