深色模式
Prometheus 高可用
单个 Prometheus 挂了监控就盲区。双副本 + Alertmanager 集群,让监控自身也高可用。
适用环境
- 两台及以上服务器
- 已能独立部署单节点 Prometheus
bash
# 确认两台机器时间同步(HA 依赖时间一致)
timedatectl status | grep "System clock"1
2
2
操作步骤
1. 部署两份独立 Prometheus
在两台机器各装一份 Prometheus,配置相同 prometheus.yml。它们各自抓取相同目标,数据相互独立。
bash
# 节点 A 与节点 B 都运行,监听 9090
systemctl status prometheus --no-pager1
2
2
2. 部署 Alertmanager 集群
yaml
# alertmanager.yml 开启集群通信
cluster:
listen_address: '0.0.0.0:9094'
peers:
- 'am-node-a:9094'
- 'am-node-b:9094'1
2
3
4
5
6
2
3
4
5
6
3. Prometheus 指向多个 Alertmanager
yaml
alerting:
alertmanagers:
- static_configs:
- targets: ['am-node-a:9093', 'am-node-b:9093']1
2
3
4
2
3
4
验证
bash
# 查看 Alertmanager 集群成员
curl -s http://localhost:9093/api/v2/status | head -c 3001
2
2
触发一条测试告警,确认只收到一次通知(Gossip 去重生效)。
常见坑
双副本不等于去重查询
两份 Prometheus 数据独立,前端查询(Grafana)需做去重或只查一份,否则图表翻倍。
Alertmanager 必须奇数节点
Gossip 集群建议 3 个节点,2 节点在脑裂时无法选举,反而更易出问题。