深色模式
弹性伸缩 ASG
摘要:弹性伸缩让机器数量跟着负载走,忙时自动扩容、闲时自动缩容。本文给出以自定义镜像为基础的伸缩组搭建方法、三种常用伸缩策略(指标/定时/预测)的参数设置,以及如何避免「抖动扩容」这类典型问题。
适用环境
bash
# 一台已完成基线配置的云服务器,用于制作启动模板镜像
which stress || sudo apt install -y stress
# 观察工具
which watch htop || sudo apt install -y htop1
2
3
4
2
3
4
操作步骤
一、前置条件:机器必须能「无人值守启动」
弹性伸缩创建出的机器不能靠人去配环境,必须做到开机即用。检查清单:
bash
# 1) 服务开机自启
systemctl is-enabled myapp
# 2) 开机自动挂载数据盘(本次伸缩一般为无状态机,尽量不依赖数据盘)
grep -c "^UUID\|^/dev" /etc/fstab
# 3) 监控 agent、日志 agent 开机自启
systemctl is-enabled node_exporter filebeat 2>/dev/null
# 4) 启动脚本(user-data)能拉起应用
cat /var/log/cloud-init-output.log1
2
3
4
5
6
7
8
2
3
4
5
6
7
8
把验证通过的机器打成镜像,作为伸缩组的启动模板源。
二、创建启动模板与伸缩组
bash
# 启动模板:指定镜像、实例类型、user-data、实例角色
aws ec2 create-launch-template --launch-template-name web-lt \
--launch-template-data '{
"ImageId":"ami-0abc123",
"InstanceType":"c6.large",
"SecurityGroupIds":["sg-app"],
"UserData":"'"$(base64 -w0 user-data.sh)"'",
"IamInstanceProfile":{"Name":"app-oss-role"}
}'
# 伸缩组:跨两个子网,最小 2 最大 10
aws autoscaling create-auto-scaling-group --auto-scaling-group-name web-asg \
--launch-template LaunchTemplateName=web-lt,Version='$Latest' \
--min-size 2 --max-size 10 --desired-capacity 2 \
--vpc-zone-identifier "subnet-app-a,subnet-app-b" \
--target-group-arns arn:aws:...web-tg \
--health-check-type ELB --health-check-grace-period 3001
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
--health-check-type ELB 让伸缩组以负载均衡的健康检查为准,能自动替换掉应用起不来的机器。
三、三种伸缩策略
bash
# 1) 目标跟踪:把 CPU 平均维持在 60%
aws autoscaling put-scaling-policy --auto-scaling-group-name web-asg \
--policy-name cpu60 --policy-type TargetTrackingScaling \
--target-tracking-configuration '{
"TargetValue":60.0,
"PredefinedMetricSpecification":{"PredefinedMetricType":"ASGAverageCPUUtilization"},
"ScaleInCooldown":300,"ScaleOutCooldown":60
}'
# 2) 定时伸缩:每天 09:00 扩到 8 台
aws autoscaling put-scheduled-update-group-action --auto-scaling-group-name web-asg \
--scheduled-action-name morning-up --recurrence "0 1 * * *" --desired-capacity 8
aws autoscaling put-scheduled-update-group-action --auto-scaling-group-name web-asg \
--scheduled-action-name night-down --recurrence "0 15 * * *" --desired-capacity 2
# 3) 步长伸缩:按告警幅度分档扩
aws cloudwatch put-metric-alarm --alarm-name high-cpu \
--metric-name CPUUtilization --namespace AWS/EC2 --statistic Average \
--period 60 --threshold 80 --comparison-operator GreaterThanThreshold \
--evaluation-periods 2 --dimensions Name=AutoScalingGroupName,Value=web-asg \
--alarm-actions arn:aws:autoscaling:...:scalingPolicy:.../policy/step-out1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
注意
--recurrence 用的是 UTC 时间。上述 0 1 * * * 对应北京时间 09:00,别写错成 0 9 * * *。
四、冷却时间:防止抖动扩容
扩容冷却(ScaleOutCooldown)建议 60–120 秒,让新机器有时间真正承载流量;缩容冷却(ScaleInCooldown)建议 300 秒以上,避免刚缩完又扩,来回震荡。
五、观察与验证
bash
# 制造负载触发扩容
stress --cpu 4 --timeout 300
# 观察实例数量变化
watch -n 10 "aws autoscaling describe-auto-scaling-groups --auto-scaling-group-name web-asg \
--query 'AutoScalingGroups[].[MinSize,MaxSize,DesiredCapacity]' --output text"
# 查看伸缩活动历史(最重要,能看到为什么扩/为什么失败)
aws autoscaling describe-scaling-activities --auto-scaling-group-name web-asg --max-records 10
# 确认新机器已注册进负载均衡且健康
aws elbv2 describe-target-health --target-group-arn arn:aws:...web-tg \
--query 'TargetHealthDescriptions[].[Target.Id,TargetHealth.State]' --output table1
2
3
4
5
6
7
8
9
10
11
12
13
2
3
4
5
6
7
8
9
10
11
12
13
危险
缩容会直接释放实例。如果机器上保存了本地会话或临时状态,缩容瞬间用户会被踢下线。伸缩组内的机器必须是无状态的,会话应放到 Redis 等外部存储。
验证
- [ ] 人为打高 CPU 后,实例数在 5 分钟内增加,且新实例健康检查为 healthy
- [ ] 停止负载后,实例数在冷却期后回落到最小值
- [ ]
describe-scaling-activities中无失败记录(如配额不足、镜像不可用) - [ ] 实例分布在至少两个可用区
常见坑
- 没做镜像直接伸缩:新机器起来没有应用,负载均衡全部 unhealthy,反而引发故障。
- 忽略了启动耗时:健康检查宽限期(grace period)太短,机器还在拉镜像就被判死,陷入无限替换。
- 只扩不缩:忘了配缩容策略或冷却时间过长,机器越攒越多,账单失控。
- 伸缩组配额:实例数上限受账号配额限制,扩容失败时先看
describe-scaling-activities里的配额错误。 - 有状态服务进伸缩组:本地 session、本地缓存导致请求落在不同机器上结果不一致。