深色模式
云监控使用
摘要:云监控提供的是「平台侧视角」的指标——CPU、网络、磁盘、托管服务的健康状态,这些是主机内 agent 拿不到的。本文讲清该看哪些指标、告警怎么分级才不吵,以及告警如何落到真正有人响应的通道。
适用环境
bash
# 一台云主机,用于验证 agent 与自定义指标上报
systemctl status cloudwatch-agent 2>/dev/null || echo "agent 未安装"
# 本地
which curl jq1
2
3
4
2
3
4
操作步骤
一、分清两类指标
| 类型 | 来源 | 例子 | 特点 |
|---|---|---|---|
| 平台指标 | 云厂商采集,无需装 agent | 实例 CPU(部分厂商需 agent)、网络出入流量、LB 5xx、RDS 连接数 | 覆盖广,粒度通常 60 秒 |
| 主机内指标 | 需装 agent | 内存、磁盘使用率、进程数、TCP 连接数 | 更细,可自定义 |
注意
很多厂商的「内存使用率」和「磁盘使用率」默认没有,必须在主机内安装监控 agent 才能上报。不装 agent 会导致「磁盘写满了但云监控一切正常」。
二、必配的四条基础告警
给每台生产机器配这四条,能覆盖绝大多数突发故障:
bash
# 1) CPU 持续高位
aws cloudwatch put-metric-alarm --alarm-name web-cpu-high \
--namespace AWS/EC2 --metric-name CPUUtilization --statistic Average \
--period 300 --evaluation-periods 2 --threshold 85 \
--comparison-operator GreaterThanThreshold \
--dimensions Name=InstanceId,Value=i-0abc --treat-missing-data notBreaching
# 2) 磁盘使用率(需 agent 上报的 mem_used_percent / disk_used_percent 命名空间)
aws cloudwatch put-metric-alarm --alarm-name web-disk-high \
--namespace CWAgent --metric-name disk_used_percent --statistic Average \
--period 300 --evaluation-periods 1 --threshold 85 \
--comparison-operator GreaterThanThreshold \
--dimensions Name=InstanceId,Value=i-0abc Name=path,Value=/ Name=fstype,Value=xfs
# 3) 网络出流量突增(可能被打流量或被当肉鸡)
aws cloudwatch put-metric-alarm --alarm-name web-netout-spike \
--namespace AWS/EC2 --metric-name NetworkOut --statistic Sum \
--period 300 --evaluation-periods 1 --threshold 100000000 \
--comparison-operator GreaterThanThreshold \
--dimensions Name=InstanceId,Value=i-0abc
# 4) 实例状态检查失败(底层宿主机故障)
aws cloudwatch put-metric-alarm --alarm-name web-status-failed \
--namespace AWS/EC2 --metric-name StatusCheckFailed --statistic Maximum \
--period 60 --evaluation-periods 2 --threshold 1 \
--comparison-operator GreaterThanOrEqualToThreshold \
--dimensions Name=InstanceId,Value=i-0abc1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
三、业务侧告警:看负载均衡而非机器
bash
# LB 5xx 比例告警(用 ALB 的 HTTPCode_Target_5XX_Count)
aws cloudwatch put-metric-alarm --alarm-name alb-5xx \
--namespace AWS/ApplicationELB --metric-name HTTPCode_Target_5XX_Count \
--statistic Sum --period 60 --evaluation-periods 3 --threshold 10 \
--comparison-operator GreaterThanThreshold \
--dimensions Name=LoadBalancer,Value=app/web-alb/abc123
# 目标响应时间(TargetResponseTime)超过 1 秒
aws cloudwatch put-metric-alarm --alarm-name alb-slow \
--namespace AWS/ApplicationELB --metric-name TargetResponseTime \
--statistic p95 --period 300 --evaluation-periods 2 --threshold 1.0 \
--comparison-operator GreaterThanThreshold \
--dimensions Name=LoadBalancer,Value=app/web-alb/abc1231
2
3
4
5
6
7
8
9
10
11
12
13
2
3
4
5
6
7
8
9
10
11
12
13
四、告警分级与通知通道
text
P0(电话/短信,7x24) :全站不可用、数据丢失风险、安全事件
P1(IM 群 @ 值班) :核心功能降级、5xx 升高、主库延迟
P2(工单/邮件,工作时间):磁盘 80%、证书 30 天内过期、成本异常
P3(周报汇总) :资源闲置、标签缺失1
2
3
4
2
3
4
危险
不要把所有告警都设为 P0。告警疲劳是监控失效的根本原因——一旦值班人习惯了忽略告警,真正严重的告警也会被一起忽略。
五、验证告警真的会响
配置完必须验证一次,否则等于没配:
bash
# 1) 人为触发:临时把阈值改到必然触发的水平,或压测制造真实负载
stress --cpu 8 --timeout 600
# 2) 查看告警状态
aws cloudwatch describe-alarms --alarm-names web-cpu-high \
--query 'MetricAlarms[].[AlarmName,StateValue,StateReason]' --output table
# 3) 手动测试通知链路(把指标直接推到阈值以上)
aws cloudwatch put-metric-data --namespace Test --metric-name Probe --value 1001
2
3
4
5
6
7
8
9
2
3
4
5
6
7
8
9
验证
- [ ] 每台生产机都有 CPU / 磁盘 / 状态检查三条基础告警
- [ ] 业务侧有 5xx 与响应时间的告警
- [ ] 告警分级文档存在,且 P0 有 7x24 通知通道
- [ ] 至少做过一次真实的告警触发演练,并记录响应时间
常见坑
- 不装 agent:以为云监控已经覆盖内存和磁盘,结果磁盘满到服务崩溃都没告警。
- 只用平均值:CPU 平均值 50% 掩盖了某台机器 100% 的事实,关键指标应同时看最大值。
treat-missing-data默认当成正常:agent 挂了导致数据缺失,告警反而不触发,建议设为breaching或notBreaching按场景明确指定。- 告警没有 Runbook:收到告警不知道该干什么,值班人只能重启了事。每条 P0/P1 应附一句「先看什么、怎么处理」。
- 只在月底看一次:监控面板应该是日常巡检入口,而不是出事才打开。