深色模式
合成监控告警
摘要:白盒指标(CPU、QPS)正常不代表用户能用。合成监控(Synthetic Monitoring)从外部定时访问关键链路(登录、下单、支付),失败即告警,补齐“用户视角”的盲区。
适用环境
- 已部署 Prometheus + Blackbox Exporter(或类似拨测工具)
- 有若干个关键 URL/接口需要主动探测
操作步骤
1. 部署 Blackbox Exporter
bash
docker run -d -p 9115:9115 prom/blackbox-exporter1
2. 在 Prometheus 配探测任务
yaml
scrape_configs:
- job_name: blackbox
metrics_path: /probe
params:
module: [http_2xx]
static_configs:
- targets:
- https://www.example.com/health
relabel_configs:
- source_labels: [__address__]
target_label: __param_target
- source_labels: [__param_target]
target_label: instance
- target_label: __address__
replacement: localhost:91151
2
3
4
5
6
7
8
9
10
11
12
13
14
15
2
3
4
5
6
7
8
9
10
11
12
13
14
15
3. 告警规则:探测失败即响
yaml
groups:
- name: synthetic
rules:
- alert: ProbeFailed
expr: probe_success == 0
for: 1m
labels:
severity: critical
annotations:
summary: "拨测失败:{{ $labels.instance }}"1
2
3
4
5
6
7
8
9
10
2
3
4
5
6
7
8
9
10
4. 接 Alertmanager 通知
yaml
# 路由到 oncall/IM(见前文 Webhook/IM 配置)
route:
matchers: ['alertname="ProbeFailed"']
receiver: 'page'1
2
3
4
2
3
4
DANGER
拨测机要与用户同网络位置(同 IDC/同公网出口),否则“探测机通、用户不通”或反之,失去意义。
验证
bash
# 手动停掉目标服务,1 分钟后 probe_success 应为 0 并触发告警
curl -s 'localhost:9115/probe?module=http_2xx&target=https://www.example.com/health' \
| grep probe_success1
2
3
2
3
常见坑
WARNING
- 探测频率太低(如 5 分钟一次)会漏掉短时故障;关键链路建议 30s-1min。
- 只探
/health不够,要探“真实业务链路”(带登录态的下单),否则健康页通但实际不可用。