深色模式
CPU 内存故障注入
摘要:资源打满是最常见的真实故障形态。本文用 Chaos Mesh 的 StressChaos 与命令行 stress-ng 注入 CPU/内存压力,观察 HPA 扩容、限流与降级是否符合预期。
适用环境
bash
# 命令行方式(非 K8s 环境亦可用)
command -v stress-ng || echo "需安装:apt-get install -y stress-ng"
nproc; free -m | head -2
# K8s 方式
kubectl -n chaos-mesh get pod | head -31
2
3
4
5
2
3
4
5
操作步骤
第 1 步:记录基线
bash
# 注入前先记录,否则无对比
uptime
kubectl -n test top pod -l app=demo 2>/dev/null || top -bn1 | head -12
curl -sS --data-urlencode 'query=avg(rate(container_cpu_usage_seconds_total{pod=~"demo.*"}[2m]))' \
http://127.0.0.1:9090/api/v1/query | jq -r '.data.result[].value[1]'1
2
3
4
5
2
3
4
5
第 2 步:命令行注入(物理机/容器均可)
bash
# 4 个 CPU 满载,持续 60 秒
stress-ng --cpu 4 --cpu-load 80 --timeout 60s --metrics-brief
# 内存占用 512MB,持续 60 秒
stress-ng --vm 1 --vm-bytes 512M --vm-hang 0 --timeout 60s --metrics-brief1
2
3
4
5
2
3
4
5
第 3 步:K8s 中用 StressChaos
yaml
apiVersion: chaos-mesh.org/v1alpha1
kind: StressChaos
metadata:
name: stress-cpu-demo
namespace: test
spec:
mode: one
selector:
namespaces: [test]
labelSelectors: { app: demo }
stressors:
cpu:
workers: 2
load: 80
memory:
workers: 1
size: '256MB'
duration: '120s'1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
bash
kubectl -n test apply -f stress-cpu-demo.yaml
kubectl -n test get stresschaos1
2
2
第 4 步:观察系统反应
bash
# HPA 是否扩容
kubectl -n test get hpa demo -w
# 是否触发 OOM 或重启
kubectl -n test get events --sort-by=.lastTimestamp | tail -10
# 业务指标是否守住稳态
watch -n 5 'kubectl -n test top pod -l app=demo'1
2
3
4
5
6
7
8
2
3
4
5
6
7
8
第 5 步:停止并核对恢复
bash
kubectl -n test delete -f stress-cpu-demo.yaml
kubectl -n test get pod -l app=demo
kubectl -n test top pod -l app=demo1
2
3
2
3
验证
bash
# 1. 压力确实打上去了(CPU 使用率显著高于基线)
kubectl -n test top pod -l app=demo --no-headers | awk '{print $2}'
# 2. HPA 有反应或触发了限流(二者至少其一)
kubectl -n test get hpa demo -o jsonpath='{.status.currentReplicas}{"\n"}'
# 3. 实验结束后资源回落
kubectl -n test top pod -l app=demo --no-headers | awk '{print $3}'1
2
3
4
5
6
7
8
2
3
4
5
6
7
8
常见坑
内存注入超过容器 limit 直接 OOMKill
这是预期内的"实验结果"而非失败,但要先确认业务是否需要能扛住 OOM。想测"内存紧张但不死"就注入到 limit 的 70% 左右。
在跳板机/生产机上跑 stress-ng
stress-ng --cpu 0 会打满所有核,导致同机其他服务受损。务必指定 worker 数量与 timeout。
对无 HPA 且副本为 1 的服务持续高压
会直接造成服务不可用。资源类实验务必先确认有扩容或限流兜底,并设短 duration。