深色模式
依赖失效演练
摘要:现代故障大多来自依赖而非自身。本文按"强依赖/弱依赖"梳理调用链,用 HTTP 故障注入、断网与返回错误码三种方式模拟下游失效,验证熔断与降级。
适用环境
bash
# 先梳理依赖清单
kubectl -n test get svc
kubectl -n test get cm app-config -o yaml | grep -iE 'url|host|addr' | head
command -v iptables >/dev/null && echo "可用 iptables 断链"1
2
3
4
2
3
4
操作步骤
第 1 步:给依赖分类
bash
cat > deps.md <<'EOF'
| 依赖 | 类型 | 失效后应有行为 |
| --- | --- | --- |
| 数据库 | 强 | 快速失败 + 友好报错 |
| Redis 缓存 | 弱 | 降级直连 DB |
| 推荐服务 | 弱 | 隐藏该模块 |
| 短信网关 | 弱 | 异步重试 |
EOF1
2
3
4
5
6
7
8
2
3
4
5
6
7
8
第 2 步:让下游直接返回错误(最简单可控)
yaml
# 用服务网格注入 HTTP 500(以 Istio 为例)
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata:
name: rec-fault
spec:
hosts: [rec]
http:
- fault:
abort:
percentage: { value: 100 }
httpStatus: 500
route:
- destination: { host: rec }1
2
3
4
5
6
7
8
9
10
11
12
13
14
2
3
4
5
6
7
8
9
10
11
12
13
14
bash
kubectl apply -f rec-fault.yaml
curl -sS -o /dev/null -w 'code=%{http_code}\n' http://<入口>/home1
2
2
第 3 步:模拟依赖完全不可达
yaml
# Chaos Mesh:注入对特定目标的丢包
apiVersion: chaos-mesh.org/v1alpha1
kind: NetworkChaos
metadata:
name: redis-unreachable
namespace: test
spec:
action: loss
mode: all
selector:
namespaces: [test]
labelSelectors: { app: demo }
loss:
loss: '100'
correlation: '100'
direction: to
target:
mode: all
selector:
namespaces: [test]
labelSelectors: { app: redis }
duration: '60s'1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
bash
kubectl -n test apply -f redis-unreachable.yaml
kubectl -n test logs deploy/demo --tail=30 | grep -iE 'redis|connect|refused'1
2
2
第 4 步:观察是否按预期降级
bash
# 期望:弱依赖失效时主流程仍可用(非 5xx)
curl -sS -o /dev/null -w 'code=%{http_code} time=%{time_total}s\n' http://<入口>/order
# 期望:强依赖失效时快速失败,而不是超时挂住
time curl -sS -o /dev/null http://<入口>/order1
2
3
4
5
2
3
4
5
第 5 步:清理并确认恢复
bash
kubectl delete -f rec-fault.yaml
kubectl -n test delete -f redis-unreachable.yaml
curl -sS -o /dev/null -w 'code=%{http_code}\n' http://<入口>/home1
2
3
2
3
验证
bash
# 1. 弱依赖失效时主流程 HTTP 200 或明确降级码
curl -sS -o /dev/null -w '%{http_code}\n' http://<入口>/home
# 2. 强依赖失效时响应在超时时间内返回
curl -sS -m 3 -o /dev/null -w '%{time_total}\n' http://<入口>/order
# 3. 依赖恢复后功能完整
kubectl -n test get networkchaos1
2
3
4
5
6
7
8
2
3
4
5
6
7
8
常见坑
只测"依赖慢"不测"依赖挂"
慢会触发超时,挂会触发连接拒绝,二者代码路径不同。两种都要测。
降级开关从未被验证
很多降级开关写了但从没打开过,真正需要用时不生效。演练必须实际触发一次开关。
对数据库做写入故障注入
可能造成脏数据或数据不一致。数据层实验优先在只读场景或从库进行,并提前备份。